Ian D. Reid 0001

dblp:r/IanDReid1 · also Ian David Reid, Ian Reid 0001 · DBLP profile ↗
← Back
280ranked-venue papers
18as first author
42since 2021 · last 2025
0000-0001-7790-6423ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 263 · 18 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 161 · 9 first-author · 14 since 2021Systems, architecture and hardware · 51 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2025 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
abstract
Current 3D Large Multimodal Models (3D LMMs) have shown tremendous potential in 3D-vision-based dialogue and reasoning. However, how to further enhance 3D LMMs to achieve fine-grained scene understanding and facilitate flexible human-agent interaction remains a challenging problem. In this work, we introduce 3D-LLaVA, a simple yet highly powerful 3D LMM designed to act as an intelligent assistant in comprehending, reasoning, and interacting with the 3D world. Unlike existing top-performing methods that rely on complicated pipelines—such as offline multi-view feature extraction or additional task-specific heads—3D-LLaVA adopts a minimalist design with integrated architecture and only takes point clouds as input. At the core of 3D-LLaVA is a new Omni Superpoint Transformer (OST), which integrates three functionalities: (1) a visual feature selector that converts and selects visual tokens, (2) a visual prompt encoder that embeds interactive visual prompts into the visual token space, and (3) a referring mask decoder that produces 3D masks based on text description. This versatile OST is empowered by the hybrid pretraining to obtain perception priors and leveraged as the visual connector that bridges the 3D data to the LLM. After performing unified instruction tuning, our 3D-LLaVA reports impressive results on various benchmarks. The code and model will be released at https://github.com/djiajunustc/3D-LLaVA.
Jiajun Deng, Tianyu He, Tianyu Wang 0035, Feras Dayoub, Ian D. Reid 0001
CVPR6
2025 Social-MAE: Social Masked Autoencoder for Multi-Person Motion Representation Learning
Mahsa Ehsanpour, Ian D. Reid 0001, Seyed Hamid Rezatofighi
ICRA2
2025 Hier-SLAM: Scaling-Up Semantics in SLAM with a Hierarchically Categorical Gaussian Splatting
abstract
We propose Hier-SLAM, a semantic 3D Gaussian Splatting SLAM method featuring a novel hierarchical categorical representation, which enables accurate global 3D semantic mapping, scaling-up capability, and explicit semantic label prediction in the 3D world. The parameter usage in semantic SLAM systems increases significantly with the growing complexity of the environment, making it particularly challenging and costly for scene understanding. To address this problem, we introduce a novel hierarchical representation that encodes semantic information in a compact form into 3D Gaussian Splatting, leveraging the capabilities of large language models (LLMs). We further introduce a novel semantic loss designed to optimize hierarchical semantic information through both inter-level and cross-level optimization. Furthermore, we enhance the whole SLAM system, resulting in improved tracking and mapping performance. Our Hier-SLAM outperforms existing dense SLAM methods in both mapping and tracking accuracy, while achieving a 2x operation speed-up. Additionally, it achieves on-par semantic rendering performance compared to existing methods while significantly reducing storage and training time requirements. Rendering FPS impressively reaches 2,000 with semantic information and 3,000 without it. Most notably, it showcases the capability of handling the complex real-world scene with more than 500 semantic classes, highlighting its valuable scaling-up capability. The open-source code is available at https://github.com/LeeBY68/Hier-SLAM.
Boying Li, Zhixi Cai, Yuan-Fang Li, Ian D. Reid 0001, Seyed Hamid Rezatofighi
ICRA4
2025 TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
abstract
Visual navigation in robotics traditionally relies on globally-consistent 3D maps or learned controllers, which can be computationally expensive and difficult to generalize across diverse environments. In this work, we present a novel RGB-only, object-level topometric navigation pipeline that enables zero-shot, long-horizon robot navigation without requiring 3D maps or pre-trained controllers. Our approach integrates global topological path planning with local metric trajectory control, allowing the robot to navigate towards object-level sub-goals while avoiding obstacles. We address key limitations of previous methods by continuously predicting local trajectory using monocular depth and traversability estimation, and in-corporating an auto-switching mechanism that falls back to a baseline controller when necessary. The system operates using foundational models, ensuring open-set applicability without the need for domain-specific fine-tuning. We demonstrate the effectiveness of our method in both simulated environments and real-world tests, highlighting its robustness and deployability. Our approach outperforms existing state-of-the-art methods, offering a more adaptable and effective solution for visual navigation in open-set environments. The source code is made publicly available: https://github.com/podgorki/TANGO.
Stefan Podgorski, Sourav Garg, Mehdi Hosseinzadeh 0003, Lachlan Mares, Feras Dayoub, Ian D. Reid 0001
ICRA6
2025 GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning
abstract
Grasp detection is a fundamental robotic task critical to the success of many industrial applications. However, current language-driven models for this task often struggle with cluttered images, lengthy textual descriptions, or slow inference speed. We introduce GraspMamba, a new language-driven grasp detection method that employs hierarchical feature fusion with Mamba vision to tackle these challenges. By leveraging rich visual features of the Mamba-based backbone alongside textual information, our approach effectively enhances the fusion of multimodal features. GraspMamba represents the first Mamba-based grasp detection model to extract vision and language features at multiple scales, delivering robust performance and rapid inference time. Intensive experiments show that GraspMamba outperforms recent methods by a clear margin. We validate our approach through real-world robotic experiments, highlighting its fast inference speed.
An Vuong, Anh Nguyen 0003, Ian D. Reid 0001, Minh Nhat Vu
IROS4
2025 Action Tokenizer Matters in In-Context Imitation Learning
abstract
In-context imitation learning (ICIL) is a new paradigm that enables robots to generalize from demonstrations to unseen tasks without retraining. A well-structured action representation is the key to capturing demonstration information effectively, yet action tokenizer (the process of discretizing and encoding actions) remains largely unexplored in ICIL. In this work, we first systematically evaluate existing action tokenizer methods in ICIL and reveal a critical limitation: while they effectively encode action trajectories, they fail to preserve temporal smoothness, which is crucial for stable robotic execution. To address this, we propose LipVQ-VAE, a variational autoencoder that enforces the Lipschitz condition in the latent action space via weight normalization. By propagating smoothness constraints from raw action inputs to a quantized latent codebook, LipVQ-VAE generates smoother actions. When integrating into ICIL, LipVQ-VAE improves performance by more than 5.3% in high-fidelity simulators, with real-world experiments confirming its ability to produce smoother, more reliable trajectories. Code and checkpoints are available at https://action-tokenizer-matters.github.io/.
An Dinh Vuong, Minh Nhat Vu, Dong An 0002, Ian D. Reid 0001
IROS4
2025 FlashMo: Geometric Interpolants and Frequency-Aware Sparsity for Scalable Efficient Motion Generation
abstract
Diffusion models have recently advanced 3D human motion generation by producing smoother and more realistic sequences from natural language. However, existing approaches face two major challenges: high computational cost during training and inference, and limited scalability due to reliance on U-Net inductive bias. To address these challenges, we propose **FlashMo**, a frequency-aware sparse motion diffusion model that prunes low-frequency tokens to enhance efficiency without custom kernel design. We further introduce *MotionSiT*, a scalable diffusion transformer based on a joint-temporal factorized interpolant with Lie group geodesics over $\mathrm{SO}(3)$ manifolds, enabling principled generation of joint rotations. Extensive experiments on the large-scale MotionHub V2 dataset and standard benchmarks including HumanML3D and KIT-ML demonstrate that our method significantly outperforms previous approaches in motion quality, efficiency, and scalability. Compared to the state-of-the-art 1-step distillation baseline, FlashMo reduces **12.9%** inference time and FID by **34.1%**. Project website: https://steve-zeyu-zhang.github.io/FlashMo.
Zeyu Zhang 0006, Danning Li, Dong Gong, Ian D. Reid 0001, Richard I. Hartley
NeurIPS5
2025 Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
abstract
We address the task of Vision-Language Navigation in Continuous Environments (VLN-CE) under the zero-shot setting. Zero-shot VLN-CE is particularly challenging due to the absence of expert demonstrations for training and minimal environment structural prior to guide navigation. To confront these challenges, we propose a Constraint-Aware Navigator (CA-Nav), which reframes zero-shot VLN-CE as a sequential, constraint-aware sub-instruction completion process. CA-Nav continuously translates sub-instructions into navigation plans using two core modules: the Constraint-Aware Sub-instruction Manager (CSM) and the Constraint-Aware Value Mapper (CVM). CSM defines the completion criteria for decomposed sub-instructions as constraints and tracks navigation progress by switching sub-instructions in a constraint-aware manner. CVM, guided by CSM's constraints, generates a value map on the fly and refines it using superpixel clustering to improve navigation stability. CA-Nav achieves the state-of-the-art performance on two VLN-CE benchmarks, surpassing the previous best method by 12% and 13% in Success Rate on the validation unseen splits of R2R-CE and RxR-CE, respectively. Moreover, CA-Nav demonstrates its effectiveness in real-world robot deployments across various indoor scenes and instructions.
Dong An 0002, Yan Huang 0008, Rongtao Xu, Yifei Su, Yonggen Ling, Ian D. Reid 0001, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 JRDB-PanoTrack: An Open-World Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments
abstract
Autonomous robot systems have attracted increasing research attention in recent years, where environment understanding is a crucial step for robot navigation, human-robot interaction, and decision. Real-world robot systems usually collect visual data from multiple sensors and are required to recognize numerous objects and their movements in complex human-crowded settings. Traditional benchmarks, with their reliance on single sensors and limited object classes and scenarios, fail to provide the comprehensive environmental understanding robots need for accurate navigation, interaction, and decision-making. As an extension of JRDB dataset, we unveil JRDB-PanoTrack, a novel open-world panoptic segmentation and tracking benchmark, towards more comprehensive environmental perception. JRDB-PanoTrack includes (1) various data involving indoor and outdoor crowded scenes, as well as comprehensive 2D and 3D synchronized data modalities; (2) high-quality 2D spatial panoptic segmentation and temporal tracking annotations, with additional 3D label projections for further spatial understanding; (3) diverse object classes for closed- and open-world recognition benchmarks, with OSPA-based metrics for evaluation. Extensive evaluation of leading methods shows significant challenges posed by our dataset.
Duy-Tho Le, Chenhui Gou, Stavya Datta, Hengcan Shi, Ian D. Reid 0001, Jianfei Cai 0001, Seyed Hamid Rezatofighi
CVPR5
2024 ItTakesTwo: Leveraging Peer Representations for Semi-supervised LiDAR Semantic Segmentation
Yuyuan Liu, Yuanhong Chen, Hu Wang 0005, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001
ECCV (1)5
2024 GaussCtrl: Multi-view Consistent Text-Driven 3D Gaussian Splatting Editing
Jiawang Bian, Xinghui Li, Guangrun Wang, Ian D. Reid 0001, Philip Torr 0001, Victor Adrian Prisacariu
ECCV (14)5
2024 Motion Mamba: Efficient and Long Sequence Motion Generation
Zeyu Zhang 0006, Akide Liu, Ian D. Reid 0001, Richard I. Hartley, Bohan Zhuang, Hao Tang 0005
ECCV (1)3
2024 RoboHop: Segment-based Topological Map Representation for Open-World Visual Navigation
abstract
Mapping is crucial for spatial reasoning, planning and robot navigation. Existing approaches range from metric, which require precise geometry-based optimization, to purely topological, where image-as-node based graphs lack explicit object-level reasoning and interconnectivity. In this paper, we propose a novel topological representation of an environment based on , which are semantically meaningful and open-vocabulary queryable, conferring several advantages over previous works based on pixel-level features. Unlike 3D scene graphs, we create a purely topological graph with segments as nodes, where edges are formed by a) associating segment-level descriptors between pairs of consecutive images and b) connecting neighboring segments within an image using their pixel centroids. This unveils a continuous sense of a place, defined by inter-image persistence of segments along with their intra-image neighbours. It further enables us to represent and update segment-level descriptors through neighborhood aggregation using graph convolution layers, which improves robot localization based on segment-level retrieval. Using real-world data, we show how our proposed map representation can be used to i) generate navigation plans in the form of hops over segments and ii) search for target objects using natural language queries describing spatial relations of objects. Furthermore, we quantitatively analyze data association at the segment level, which underpins inter-image connectivity during mapping and segment-level localization when revisiting the same place. Finally, we show preliminary trials on segment-level ‘hopping’ based zero-shot real-world navigation. Project page with supplementary details: oravus.github.io/RoboHop/.
Sourav Garg, Krishan Rana, Mehdi Hosseinzadeh 0003, Lachlan Mares, Niko Sünderhauf, Feras Dayoub, Ian D. Reid 0001
ICRA7
2024 BEVPose: Unveiling Scene Semantics through Pose-Guided Multi-Modal BEV Alignment
abstract
In the field of autonomous driving and mobile robotics, there has been a significant shift in the methods used to create Bird’s Eye View (BEV) representations. This shift is characterised by using transformers and learning to fuse measurements from disparate vision sensors, mainly lidar and cameras, into a 2D planar ground-based representation. However, these learning-based methods for creating such maps often rely heavily on extensive annotated data, presenting notable challenges, particularly in diverse or non-urban environments where large-scale datasets are scarce. In this work, we present BEVPose, a framework that integrates BEV representations from camera and lidar data, using sensor pose as a guiding supervisory signal. This method notably reduces the dependence on costly annotated data. By leveraging pose information, we align and fuse multi-modal sensory inputs, facilitating the learning of latent BEV embeddings that capture both geometric and semantic aspects of the environment. Our pretraining approach demonstrates promising performance in BEV map segmentation tasks, outperforming fully-supervised state-of the-art methods, while necessitating only a minimal amount of annotated data. This development not only confronts the challenge of data efficiency in BEV representation learning but also broadens the potential for such techniques in a variety of domains, including off-road and indoor environments.
Mehdi Hosseinzadeh 0003, Ian D. Reid 0001
IROS2
2024 Assessing domain gap for continual domain adaptation in object detection
abstract
To ensure reliable object detection in autonomous systems, the detector must be able to adapt to changes in appearance caused by environmental factors such as time of day, weather, and seasons. Continually adapting the detector to incorporate these changes is a promising solution, but it can be computationally costly. Our proposed approach is to selectively adapt the detector only when necessary, using new data that does not have the same distribution as the current training data. To this end, we investigate three popular metrics for domain gap evaluation and find that there is a correlation between the domain gap and detection accuracy. Therefore, we apply the domain gap as a criterion to decide when to adapt the detector. Our experiments show that our approach has the potential to improve the efficiency of the detector’s operation in real-world scenarios, where environmental conditions change in a cyclical manner, without sacrificing the overall performance of the detector. Our code is publicly available https://github.com/dadung/DGE-CDA.
Anh-Dzung Doan, Nguyen Bach Long, Ian D. Reid 0001, Markus Wagner 0007, Tat-Jun Chin
Comput. Vis. Image Underst.4
2024 SC-DepthV3: Robust Self-Supervised Monocular Depth Estimation for Dynamic Scenes
abstract
Self-supervised monocular depth estimation has shown impressive results in static scenes. It relies on the multi-view consistency assumption for training networks, however, that is violated in dynamic object regions and occlusions. Consequently, existing methods show poor accuracy in dynamic scenes, and the estimated depth map is blurred at object boundaries because they are usually occluded in other training views. In this paper, we propose SC-DepthV3 for addressing the challenges. Specifically, we introduce an external pretrained monocular depth estimation model for generating single-image depth prior, namely pseudo-depth, based on which we propose novel losses to boost self-supervised training. As a result, our model can predict sharp and accurate depth maps, even when training from monocular videos of highly dynamic scenes. We demonstrate the significantly superior performance of our method over previous methods on six challenging datasets, and we provide detailed ablation studies for the proposed terms.
Libo Sun 0002, Jiawang Bian, Huangying Zhan, Wei Yin 0006, Ian D. Reid 0001, Chunhua Shen
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Sensor Allocation and Online-Learning-Based Path Planning for Maritime Situational Awareness Enhancement: A Multi-Agent Approach
abstract
Countries with access to large bodies of water often aim to protect their maritime transport by employing maritime surveillance systems. However, the number of available sensors (e.g., cameras) is typically small compared to the to-be-monitored targets, and their Field of View (FOV) and range are often limited. This makes improving the situational awareness of maritime transports challenging. To this end, we propose a method that not only distributes multiple sensors but also plans paths for them to observe multiple targets, while minimizing the time needed to achieve situational awareness. In particular, we provide a formulation of this sensor allocation and path planning problem which considers the partial awareness of the targets’ state, as well as the unawareness of the targets’ trajectories. To solve the problem we present two algorithms: 1) a greedy algorithm for assigning sensors to targets, and 2) a distributed multi-agent path planning algorithm based on regret-matching learning. Because a quick convergence is a requirement for algorithms developed for high mobility environments, we employ a forgetting factor to quickly converge to correlated equilibrium solutions. Experimental results show that our combined approach achieves situational awareness more quickly than related work.
Nguyen Bach Long, Anh-Dzung Doan, Tat-Jun Chin, Christophe Guettier, Estelle Parra, Ian D. Reid 0001, Markus Wagner 0007
IEEE Trans. Intell. Transp. Syst.7
2023 Residual Pattern Learning for Pixel-wise Out-of-Distribution Detection in Semantic Segmentation
abstract
Semantic segmentation models classify pixels into a set of known ("in-distribution") visual classes. When deployed in an open world, the reliability of these models depends on their ability to not only classify in-distribution pixels but also to detect out-of-distribution (OoD) pixels. Historically, the poor OoD detection performance of these models has motivated the design of methods based on model re-training using synthetic training images that include OoD visual objects. Although successful, these re-trained methods have two issues: 1) their in-distribution segmentation accuracy may drop during re-training, and 2) their OoD detection accuracy does not generalise well to new contexts outside the training set (e.g., from city to country context). In this paper, we mitigate these issues with: (i) a new residual pattern learning (RPL) module that assists the segmentation model to detect OoD pixels with minimal deterioration to inlier segmentation accuracy; and (ii) a novel context-robust contrastive learning (CoroCL) that enforces RPL to robustly detect OoD pixels in various contexts. Our approach improves by around 10% FPR and 7% AuPRC previous state-of-the-art in Fishyscapes, Segment-Me-If-You-Can, and RoadAnomaly datasets.
Yuyuan Liu, Choubo Ding, Yu Tian 0001, Guansong Pang, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001
ICCV6
2023 How Trustworthy are Performance Evaluations for Basic Vision Tasks?
abstract
This article examines performance evaluation criteria for basic vision tasks involving sets of objects namely, object detection, instance-level segmentation and multi-object tracking. The rankings of algorithms by a criterion can fluctuate with different choices of parameters, e.g. Intersection over Union (IoU) threshold, making their evaluations unreliable. More importantly, there is no means to verify whether we can trust the evaluations of a criterion. This work suggests a notion of trustworthiness for performance criteria, which requires (i) robustness to parameters for reliability, (ii) contextual meaningfulness in sanity tests, and (iii) consistency with mathematical requirements such as the metric properties. We observe that these requirements were overlooked by many widely-used criteria, and explore alternative criteria using metrics for sets of shapes. We also assess all these criteria based on the suggested requirements for trustworthiness.
Tran Thien Dat Nguyen, Seyed Hamid Rezatofighi, Ba-Ngu Vo, Ba-Tuong Vo, Silvio Savarese, Ian D. Reid 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Looking Beyond Two Frames: End-to-End Multi-Object Tracking Using Spatial and Temporal Transformers
abstract
Tracking a time-varying indefinite number of objects in a video sequence over time remains a challenge despite recent advances in the field. Most existing approaches are not able to properly handle multi-object tracking challenges such as occlusion, in part because they ignore long-term temporal information. To address these shortcomings, we present MO3TR: a truly end-to-end Transformer-based online multi-object tracking (MOT) framework that learns to handle occlusions, track initiation and termination without the need for an explicit data association module or any heuristics. MO3TR encodes object interactions into long-term temporal embeddings using a combination of spatial and temporal Transformers, and recursively uses the information jointly with the input data to estimate the states of all tracked objects over time. The spatial attention mechanism enables our framework to learn implicit representations between all the objects and the objects to the measurements, while the temporal attention mechanism focuses on specific parts of past information, allowing our approach to resolve occlusions over multiple frames. Our experiments demonstrate the potential of this new approach, achieving results on par with or better than the current state-of-the-art on multiple MOT metrics for several popular multi-object tracking benchmarks.
Tianyu Zhu 0001, Markus Hiller, Mahsa Ehsanpour, Rongkai Ma, Tom Drummond, Ian D. Reid 0001, Seyed Hamid Rezatofighi
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 LongReMix: Robust learning with high confidence samples in a noisy label environment
Filipe R. Cordeiro, Ragav Sachdeva, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001
Pattern Recognit.4
2023 ScanMix: Learning from Severe Label Noise via Semantic Clustering and Semi-Supervised Learning
abstract
We propose a new training algorithm, ScanMix, that explores semantic clustering and semi-supervised learning (SSL) to allow superior robustness to severe label noise and competitive robustness to non-severe label noise problems, in comparison to the state of the art (SOTA) methods. ScanMix is based on the expectation maximisation framework, where the E-step estimates the latent variable to cluster the training images based on their appearance and classification results, and the M-step optimises the SSL classification and learns effective feature representations via semantic clustering. We present a theoretical result that shows the correctness and convergence of ScanMix, and an empirical result that shows that ScanMix has SOTA results on CIFAR-10/-100 (with symmetric, asymmetric and semantic label noise), Red Mini-ImageNet (from the Controlled Noisy Web Labels), Clothing1M and WebVision. In all benchmarks with severe label noise, our results are competitive to the current SOTA.
Ragav Sachdeva, Filipe R. Cordeiro, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001
Pattern Recognit.4
2022 JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity Detection
abstract
The availability of large-scale video action understanding datasets has facilitated advances in the interpretation of visual scenes containing people. However, learning to recognise human actions and their social interactions in an unconstrained real-world environment comprising numerous people, with potentially highly unbalanced and longtailed distributed action labels from a stream of sensory data captured from a mobile robot platform remains a significant challenge, not least owing to the lack of a reflective large-scale dataset. In this paper, we introduce JRDB-Act, as an extension of the existing JRDB, which is captured by a social mobile manipulator and reflects a real distribution of human daily-life actions in a university campus environment. JRDB-Act has been densely annotated with atomic actions, comprises over 2.8M action labels, constituting a large-scale spatio-temporal action detection dataset. Each human bounding box is labeled with one pose-based action label and multiple (optional) interaction-based action labels. Moreover JRDB-Act provides social group annotation, conducive to the task of grouping individuals based on their interactions in the scene to infer their social activities (common activities in each social group). Each annotated label in JRDB-Act is tagged with the annotators' confidence level which contributes to the development of reliable evaluation strategies. In order to demonstrate how one can effectively utilise such annotations, we develop an end-to-end trainable pipeline to learn and infer these tasks, i.e. individual action and social group detection. The data and the evaluation code will be publicly available at https://jrdb.erc.monash.edu/
Mahsa Ehsanpour, Fatemehsadat Saleh, Silvio Savarese, Ian D. Reid 0001, Seyed Hamid Rezatofighi
CVPR4
2022 You Only Cut Once: Boosting Data Augmentation with a Single Cut
abstract
We present You Only Cut Once (YOCO) for performing data augmentations. YOCO cuts one image into two pieces and performs data augmentations individually within each piece. Applying YOCO improves the diversity of the augmentation per sample and encourages neural networks to recognize objects from partial information. YOCO enjoys the properties of parameter-free, easy usage, and boosting almost all augmentations for free. Thorough experiments are conducted to evaluate its effectiveness. We first demonstrate that YOCO can be seamlessly applied to varying data augmentations, neural network architectures, and brings performance gains on CIFAR and ImageNet classification tasks, sometimes surpassing conventional image-level augmentation by large margins. Moreover, we show YOCO benefits contrastive pre-training toward a more powerful representation that can be better transferred to multiple downstream tasks. Finally, we study a number of variants of YOCO and empirically analyze the performance for respective settings.
Junlin Han, Pengfei Fang, Weihao Li 0005, Mohammad Ali Armin, Ian D. Reid 0001, Lars Petersson, Hongdong Li
ICML6
2022 Asynchronous Optimisation for Event-based Visual Odometry
abstract
Event cameras open up new possibilities for robotic perception due to their low latency and high dynamic range. On the other hand, developing effective event-based vision algorithms that fully exploit the beneficial properties of event cameras remains work in progress. In this paper, we focus on event-based visual odometry (VO). While existing event-driven VO pipelines have adopted continuous-time representations to asynchronously process event data, they either assume a known map, restrict the camera to planar trajectories, or integrate other sensors into the system. Towards map-free event-only monocular VO in SE(3), we propose an asynchronous structure-from-motion optimisation back-end. Our formulation is underpinned by a principled joint optimisation problem involving non-parametric Gaussian Process motion modelling and incremental maximum a posteriori inference. A high-performance incremental computation engine is employed to reason about the camera trajectory with every incoming event. We demonstrate the robustness of our asynchronous back-end in comparison to frame-based methods which depend on accurate temporal accumulation of measurements.
Daqi Liu, Álvaro Parra Bustos, Yasir Latif, Bo Chen 0009, Tat-Jun Chin, Ian D. Reid 0001
ICRA6
2022 Autonomy and Perception for Space Mining
abstract
Future Moon bases will likely be constructed using resources mined from the surface of the Moon. The difficulty of maintaining a human workforce on the Moon and communications lag with Earth means that mining will need to be conducted using collaborative robots with a high degree of autonomy. In this paper, we describe our solution for Phase 2 of the NASA Space Robotics Challenge, which provided a simulated lunar environment in which teams were tasked to develop software systems to achieve autonomous collaborative robots for mining on the Moon. Our 3rd place and innovation award winning solution shows how machine learning-enabled vision could alleviate major challenges posed by the lunar environment towards autonomous space mining, chiefly the lack of satellite positioning systems, hazardous terrain, and delicate robot interactions. A robust multi-robot coordinator was also developed to achieve long-term operation and effective collaboration between robots11A recording of our robots in action is available at [1]..
Ragav Sachdeva, Ravi Hammond, James Bockman, Alec Arthur, Brandon Smart, Dustin Craggs, Anh-Dzung Doan, T. Rowntree, Elijah Schutz, Adrian Orenstein, Andy Yu, Tat-Jun Chin, Ian D. Reid 0001
ICRA13
2022 Dual-Attention-Guided Network for Ghost-Free High Dynamic Range Imaging
Qingsen Yan, Dong Gong, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Ian D. Reid 0001, Yanning Zhang 0001
Int. J. Comput. Vis.6
2022 Structured Binary Neural Networks for Image Recognition
Bohan Zhuang, Chunhua Shen, Mingkui Tan, Peng Chen 0037, Lingqiao Liu, Ian D. Reid 0001
Int. J. Comput. Vis.6
2022 Auto-Rectify Network for Unsupervised Indoor Depth Estimation
abstract
Single-View depth estimation using the CNNs trained from unlabelled videos has shown significant promise. However, excellent results have mostly been obtained in street-scene driving scenarios, and such methods often fail in other settings, particularly indoor videos taken by handheld devices. In this work, we establish that the complex ego-motions exhibited in handheld settings are a critical obstacle for learning depth. Our fundamental analysis suggests that the rotation behaves as noise during training, as opposed to the translation (baseline) which provides supervision signals. To address the challenge, we propose a data pre-processing method that rectifies training images by removing their relative rotations for effective learning. The significantly improved performance validates our motivation. Towards end-to-end learning without requiring pre-processing, we propose an Auto-Rectify Network with novel loss functions, which can automatically learn to rectify images during training. Consequently, our results outperform the previous unsupervised SOTA method by a large margin on the challenging NYUv2 dataset. We also demonstrate the generalization of our trained model in ScanNet and Make3D, and the universality of our proposed learning method on 7-Scenes and KITTI datasets.
Jiawang Bian, Huangying Zhan, Naiyan Wang, Tat-Jun Chin, Chunhua Shen, Ian D. Reid 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Deep Object Tracking With Shrinkage Loss
abstract
In this paper, we address the issue of data imbalance in learning deep models for visual object tracking. Although it is well known that data distribution plays a crucial role in learning and inference models, considerably less attention has been paid to data imbalance in visual tracking. For the deep regression trackers that directly learn a dense mapping from input images of target objects to soft response maps, we identify their performance is limited by the extremely imbalanced pixel-to-pixel differences when computing regression loss. This prevents existing end-to-end learnable deep regression trackers from performing as well as discriminative correlation filters (DCFs) trackers. For the deep classification trackers that draw positive and negative samples to learn discriminative classifiers, there exists heavy class imbalance due to a limited number of positive samples when compared to the number of negative samples. To balance training data, we propose a novel shrinkage loss to penalize the importance of easy training data mostly coming from the background, which facilitates both deep regression and classification trackers to better distinguish target objects from the background. We extensively validate the proposed shrinkage loss function on six benchmark datasets, including the OTB-2013, OTB-2015, UAV-123, VOT-2016, VOT-2018 and LaSOT. Equipped with our shrinkage loss, the proposed one-stage deep regression tracker achieves favorable results against state-of-the-art methods, especially in comparison with DCFs trackers. Meanwhile, our shrinkage loss generalizes well to deep classification trackers. When replacing the original binary cross entropy loss with our shrinkage loss, three representative baseline trackers achieve large performance gains, even setting new state-of-the-art results.
Xiankai Lu, Chao Ma 0004, Jianbing Shen, Xiaokang Yang 0001, Ian D. Reid 0001, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Learn to Predict Sets Using Feed-Forward Neural Networks
abstract
This paper addresses the task of set prediction using deep feed-forward neural networks. A set is a collection of elements which is invariant under permutation and the size of a set is not fixed in advance. Many real-world problems, such as image tagging and object detection, have outputs that are naturally expressed as sets of entities. This creates a challenge for traditional deep neural networks which naturally deal with structured outputs such as vectors, matrices or tensors. We present a novel approach for learning to predict sets with unknown permutation and cardinality using deep neural networks. In our formulation we define a likelihood for a set distribution represented by a) two discrete distributions defining the set cardinally and permutation variables, and b) a joint distribution over set elements with a fixed cardinality. Depending on the problem under consideration, we define different training models for set prediction using deep neural networks. We demonstrate the validity of our set formulations on relevant vision problems such as: 1) multi-label image classification where we outperform the other competing methods on the PASCAL VOC and MS COCO datasets, 2) object detection, for which our formulation outperforms popular state-of-the-art detectors, and 3) a complex CAPTCHA test, where we observe that, surprisingly, our set-based network acquired the ability of mimicking arithmetics without any rules being coded.
Seyed Hamid Rezatofighi, Tianyu Zhu 0001, Roman Kaskman, Farbod T. Motlagh, Qinfeng Shi, Anton Milan, Daniel Cremers, Laura Leal-Taixé, Ian D. Reid 0001
IEEE Trans. Pattern Anal. Mach. Intell.9
2022 Effective Training of Convolutional Neural Networks With Low-Bitwidth Weights and Activations
abstract
This paper tackles the problem of training a deep convolutional neural network of both low-bitwidth weights and activations. Optimizing a low-precision network is very challenging due to the non-differentiability of the quantizer, which may result in substantial accuracy loss. To address this, we propose three practical approaches, including (i) progressive quantization; (ii) stochastic precision; and (iii) joint knowledge distillation to improve the network training. First, for progressive quantization, we propose two schemes to progressively find good local minima. Specifically, we propose to first optimize a network with quantized weights and subsequently quantize activations. This is in contrast to the traditional methods which optimize them simultaneously. Furthermore, we propose a second progressive quantization scheme which gradually decreases the bitwidth from high-precision to low-precision during training. Second, to alleviate the excessive training burden due to the multi-round training stages, we further propose a one-stage stochastic precision strategy to randomly sample and quantize sub-networks while keeping other parts in full-precision. Finally, we adopt a novel learning scheme to jointly train a full-precision model alongside the low-precision one. By doing so, the full-precision model provides hints to guide the low-precision model training and significantly improves the performance of the low-precision network. Extensive experiments on various datasets (e.g., CIFAR-100, ImageNet) show the effectiveness of the proposed methods.
Bohan Zhuang, Mingkui Tan, Jing Liu 0048, Lingqiao Liu, Ian D. Reid 0001, Chunhua Shen
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 NVSS: High-quality Novel View Selfie Synthesis
abstract
We present a novel method to synthesize novel view selfies from a mobile phone captured video. This is challenging due to the inconsistent geometry that is caused by the person’s unavoidable movement. Recent methods reconstruct the whole deformable scene implicitly with a deformation field. We argue that they are inefficient and hard to fit diverse real-world videos. In contrast, we use an explicit reconstruction for generalization and efficiency, where we separately track, reconstruct, and synthesize the foreground and background to overcome the geometry inconsistency. Several novel and effective modules are proposed for better performance and visual results. We demonstrate the advantage of the proposed method against the existing alternatives in a collection of our captured selfie videos with the support of quantitative and qualitative results.
Jiawang Bian, Huangying Zhan, Ian D. Reid 0001
3DV3
2021 PropMix: Hard Sample Filtering and Proportional MixUp for Learning with Noisy Labels
Filipe R. Cordeiro, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001
BMVC3
2021 Weakly Supervised Training of Monocular 3D Object Detectors Using Wide Baseline Multi-view Traffic Camera Data
Matthew Howe, Ian D. Reid 0001, Jamie Mackenzie
BMVC2
2021 Rotation Coordinate Descent for Fast Globally Optimal Rotation Averaging
abstract
Under mild conditions on the noise level of the measurements, rotation averaging satisfies strong duality, which enables global solutions to be obtained via semidefinite programming (SDP) relaxation. However, generic solvers for SDP are rather slow in practice, even on rotation averaging instances of moderate size, thus developing specialised algorithms is vital. In this paper, we present a fast algorithm that achieves global optimality called rotation coordinate descent (RCD). Unlike block coordinate descent (BCD) which solves SDP by updating the semidefinite matrix in a row-by-row fashion, RCD directly maintains and updates all valid rotations throughout the iterations. This obviates the need to store a large dense semidefinite matrix. We mathematically prove the convergence of our algorithm and empirically show its superior efficiency over state-of-the-art global methods on a variety of problem configurations. Maintaining valid rotations also facilitates incorporating local optimisation routines for further speed-ups. Moreover, our algorithm is simple to implement1.
Álvaro Parra Bustos, Shin-Fang Ch'ng, Tat-Jun Chin, Anders P. Eriksson, Ian D. Reid 0001
CVPR5
2021 TRiPOD: Human Trajectory and Pose Dynamics Forecasting in the Wild
abstract
Joint forecasting of human trajectory and pose dynamics is a fundamental building block of various applications ranging from robotics and autonomous driving to surveillance systems. Predicting body dynamics requires capturing subtle information embedded in the humans’ interactions with each other and with the objects present in the scene. In this paper, we propose a novel TRajectory and POse Dynamics (nicknamed TRiPOD) method based on graph attentional networks to model the human-human and human-object interactions both in the input space and the output space (decoded future output). The model is supplemented by a message passing interface over the graphs to fuse these different levels of interactions efficiently. Furthermore, to incorporate a real-world challenge, we propound to learn an indicator representing whether an estimated body joint is visible/invisible at each frame, e.g. due to occlusion or being outside the sensor field of view. Finally, we introduce a new benchmark for this joint task based on two challenging datasets (PoseTrack and 3DPW) and propose evaluation metrics to measure the effectiveness of predictions in the global space, even when there are invisible cases of joints. Our evaluation shows that TRiPOD outperforms all prior work and state-of-the-art specifically designed for each of the trajectory and pose forecasting tasks.
Vida Adeli, Mahsa Ehsanpour, Ian D. Reid 0001, Juan Carlos Niebles, Silvio Savarese, Ehsan Adeli-Mosabbeb, Seyed Hamid Rezatofighi
ICCV3
2021 ODAM: Object Detection, Association, and Mapping using Posed RGB Video
abstract
Localizing objects and estimating their extent in 3D is an important step towards high-level 3D scene understanding, which has many applications in Augmented Reality and Robotics. We present ODAM, a system for 3D Object Detection, Association, and Mapping using posed RGB videos. The proposed system relies on a deep learning front-end to detect 3D objects from a given RGB frame and associate them to a global object-based map using a graph neural network (GNN). Based on these frame-to-model associations, our back-end optimizes object bounding volumes, represented as super-quadrics, under multi-view geometry constraints and the object scale prior. We validate the proposed system on ScanNet where we show a significant improvement over existing RGB-only methods.
Kejie Li, Daniel DeTone, Steven Chen, Minh Vo, Ian D. Reid 0001, Seyed Hamid Rezatofighi, Chris Sweeney, Julian Straub, Richard A. Newcombe
ICCV5
2021 EvidentialMix: Learning with Combined Open-set and Closed-set Noisy Labels
abstract
The efficacy of deep learning depends on large-scale data sets that have been carefully curated with reliable data acquisition and annotation processes. However, acquiring such large-scale data sets with precise annotations is very expensive and time-consuming, and the cheap alternatives often yield data sets that have noisy labels. The field has addressed this problem by focusing on training models under two types of label noise: 1) closed-set noise, where some training samples are incorrectly annotated to a training label other than their known true class; and 2) open-set noise, where the training set includes samples that possess a true class that is (strictly) not contained in the set of known training labels. In this work, we study a new variant of the noisy label problem that combines the open-set and closed-set noisy labels, and introduce a benchmark evaluation to assess the performance of training algorithms under this setup. We argue that such problem is more general and better reflects the noisy label scenarios in practice. Furthermore, we propose a novel algorithm, called EvidentialMix, that addresses this problem and compare its performance with the state-of-the-art methods for both closed-set and open-set noise on the proposed benchmark. Our results show that our method produces superior classification results and better feature representations than previous state-of-the-art methods. The code is available at https:/github.com/ragavsachdeva/EvidentialMix.
Ragav Sachdeva, Filipe R. Cordeiro, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001
WACV4
2021 Unsupervised Scale-Consistent Depth Learning from Video
Jiawang Bian, Huangying Zhan, Naiyan Wang, Le Zhang 0001, Chunhua Shen, Ming-Ming Cheng, Ian D. Reid 0001
Int. J. Comput. Vis.8
2021 MOTChallenge: A Benchmark for Single-Camera Multiple Target Tracking
abstract
Abstract Standardized benchmarks have been crucial in pushing the performance of computer vision algorithms, especially since the advent of deep learning. Although leaderboards should not be over-claimed, they often provide the most objective measure of performance and are therefore important guides for research. We present MOTChallenge , a benchmark for single-camera Multiple Object Tracking (MOT) launched in late 2014, to collect existing and new data and create a framework for the standardized evaluation of multiple object tracking methods. The benchmark is focused on multiple people tracking, since pedestrians are by far the most studied object in the tracking community, with applications ranging from robot navigation to self-driving cars. This paper collects the first three releases of the benchmark: (i) MOT15 , along with numerous state-of-the-art results that were submitted in the last years, (ii) MOT16 , which contains new challenging videos, and (iii) MOT17 , that extends MOT16 sequences with more precise labels and evaluates tracking performance on three different object detectors. The second and third release not only offers a significant increase in the number of labeled boxes, but also provide labels for multiple object classes beside pedestrians, as well as the level of visibility for every single object of interest. We finally provide a categorization of state-of-the-art trackers and a broad error analysis. This will help newcomers understand the related work and research trends in the MOT community, and hopefully shed some light into potential future research directions.
Patrick Dendorfer, Aljosa Osep, Anton Milan, Konrad Schindler, Daniel Cremers, Ian D. Reid 0001, Stefan Roth 0001, Laura Leal-Taixé
Int. J. Comput. Vis.6
2021 Visual localization under appearance change: filtering approaches
Anh-Dzung Doan, Yasir Latif, Tat-Jun Chin, Yu Liu 0029, Shin-Fang Ch'ng, Thanh-Toan Do, Ian D. Reid 0001
Neural Comput. Appl.7
2020 A Generalized Framework for Edge-Preserving and Structure-Preserving Image Smoothing
abstract
Image smoothing is a fundamental procedure in applications of both computer vision and graphics. The required smoothing properties can be different or even contradictive among different tasks. Nevertheless, the inherent smoothing nature of one smoothing operator is usually fixed and thus cannot meet the various requirements of different applications. In this paper, a non-convex non-smooth optimization framework is proposed to achieve diverse smoothing natures where even contradictive smoothing behaviors can be achieved. To this end, we first introduce the truncated Huber penalty function which has seldom been used in image smoothing. A robust framework is then proposed. When combined with the strong flexibility of the truncated Huber penalty function, our framework is capable of a range of applications and can outperform the state-of-the-art approaches in several tasks. In addition, an efficient numerical solution is provided and its convergence is theoretically guaranteed even the optimization framework is non-convex and non-smooth. The effectiveness and superior performance of our approach are validated through comprehensive experimental results in a range of applications.
Wei Liu 0044, Yinjie Lei, Xiaolin Huang, Jie Yang 0002, Ian D. Reid 0001
AAAI6
2020 Augmentation Network for Generalised Zero-Shot Learning
Rafael Felix, Michele Sasdelli, Ian D. Reid 0001, Gustavo Carneiro 0001
ACCV (4)3
2020 Reconstruct Locally, Localize Globally: A Model Free Method for Object Pose Estimation
abstract
Six degree-of-freedom pose estimation of a known object in a single image is a long-standing computer vision objective. It is classically posed as a correspondence problem between a known geometric model, such as a CAD model, and image locations. If a CAD model is not available, it is possible to use multi-view visual reconstruction methods to create a geometric model, and use this in the same manner. Instead, we propose a learning-based method whose input is a collection of images of a target object, and whose output is the pose of the object in a novel view. At inference time, our method maps from the RoI features of the input image to a dense collection of object-centric 3D coordinates, one per pixel. This dense 2D-3D mapping is then used to determine 6dof pose using standard PnP plus RANSAC. The model that maps 2D to object 3D coordinates is established at training time by automatically discovering and matching image landmarks that are consistent across multiple views. We show that this method eliminates the requirement for a 3D CAD model (needed by classical geometry-based methods and state-of-the-art learning-based methods alike) but still achieves performance on a par with the prior art.
Ian D. Reid 0001
CVPR2
2020 FroDO: From Detections to 3D Objects
abstract
Object-oriented maps are important for scene understanding since they jointly capture geometry and semantics, allow individual instantiation and meaningful reasoning about objects. We introduce FroDO, a method for accurate 3D reconstruction of object instances from RGB video that infers their location, pose and shape in a coarse to fine manner. Key to FroDO is to embed object shapes in a novel learnt shape space that allows seamless switching between sparse point cloud and dense DeepSDF decoding. Given an input sequence of localized RGB frames, FroDO first aggregates 2D detections to instantiate a 3D bounding box per object. A shape code is regressed using an encoder network before optimizing shape and pose further under the learnt shape priors using sparse or dense shape representations. The optimization uses multi-view geometric, photometric and silhouette losses. We evaluate on real-world datasets, including Pix3D, Redwood-OS, and ScanNet, for single-view, multi-view, and multi-object reconstruction.
Martin Rünz, Kejie Li, Meng Tang 0001, Lingni Ma, Chen Kong, Tanner Schmidt, Ian D. Reid 0001, Lourdes Agapito, Julian Straub, Steven Lovegrove, Richard A. Newcombe
CVPR7
2020 Training Quantized Neural Networks With a Full-Precision Auxiliary Module
abstract
In this paper, we seek to tackle a challenge in training low-precision networks: the notorious difficulty in propagating gradient through a low-precision network due to the non-differentiable quantization function. We propose a solution by training the low-precision network with a full-precision auxiliary module. Specifically, during training, we construct a mix-precision network by augmenting the original low-precision network with the full precision auxiliary module. Then the augmented mix-precision network and the low-precision network are jointly optimized. This strategy creates additional full-precision routes to update the parameters of the low-precision model, thus making the gradient back-propagates more easily. At the inference time, we discard the auxiliary module without introducing any computational complexity to the low-precision network. We evaluate the proposed method on image classification and object detection over various quantization approaches and show consistent performance increase. In particular, we achieve near lossless performance to the full-precision model by using a 4-bit detector, which is of great practical value.
Bohan Zhuang, Lingqiao Liu, Mingkui Tan, Chunhua Shen, Ian D. Reid 0001
CVPR5
2020 Joint Learning of Social Groups, Individuals Action and Sub-group Activities in Videos
Mahsa Ehsanpour, Alireza Abedin Varamin, Fatemehsadat Saleh, Qinfeng Shi, Ian D. Reid 0001, Seyed Hamid Rezatofighi
ECCV (9)5
2020 NeuRoRA: Neural Robust Rotation Averaging
Pulak Purkait, Tat-Jun Chin, Ian D. Reid 0001
ECCV (24)3
2020 SG-VAE: Scene Grammar Variational Autoencoder to Generate New Indoor Scenes
Pulak Purkait, Christopher Zach, Ian D. Reid 0001
ECCV (24)3
2020 SPRINT: Subgraph Place Recognition for INtelligent Transportation
abstract
Visual place recognition is an important problem in mobile robotics which aims to localize a robot using image information alone. Recent methods have shown promising results for place recognition under varying environmental conditions by exploiting the sequential nature of the image acquisition process. We show that by using k nearest neighbours based image retrieval as the backend, and exploiting the structure of the image acquisition process which introduces temporal relations between images in the database, the location of possible matches can be restricted to a subset of all the images seen so far. In effect, the original problem space can thus be restricted to a significantly smaller subspace, reducing the inference time significantly. This is particularly important for scalable place recognition over databases containing millions of images. We present large scale experiments using publicly sourced data that show the computational performance of the proposed method under varying environmental conditions.
Yasir Latif, Anh-Dzung Doan, Tat-Jun Chin, Ian D. Reid 0001
ICRA4
2020 Visual Odometry Revisited: What Should Be Learnt?
abstract
In this work we present a monocular visual odometry (VO) algorithm which leverages geometry-based methods and deep learning. Most existing VO/SLAM systems with superior performance are based on geometry and have to be carefully designed for different application scenarios. Moreover, most monocular systems suffer from scale-drift issue. Some recent deep learning works learn VO in an end-to-end manner but the performance of these deep systems is still not comparable to geometry-based methods. In this work, we revisit the basics of VO and explore the right way for integrating deep learning with epipolar geometry and Perspective-n-Point (PnP) method. Specifically, we train two convolutional neural networks (CNNs) for estimating single-view depths and two-view optical flows as intermediate outputs. With the deep predictions, we design a simple but robust frame-to-frame VO algorithm (DF-VO) which outperforms pure deep learning-based and geometry-based methods. More importantly, our system does not suffer from the scale-drift issue being aided by a scale consistent single-view depth CNN. Extensive experiments on KITTI dataset shows the robustness of our system and a detailed ablation study shows the effect of different factors in our system. Code is available at here: DF-VO.
Huangying Zhan, Chamara Saroj Weerasekera, Jiawang Bian, Ian D. Reid 0001
ICRA4
2020 Meta Learning with Differentiable Closed-form Solver for Fast Video Object Segmentation
abstract
Video object segmentation plays a vital role to many robotic tasks, beyond the satisfied accuracy, quickly adapt to the new scenario with very limited annotations and conduct a quick inference are also important. In this paper, we are specifically concerned with the task of fast segmenting all pixels of a target object in all frames, given the annotation mask in the first frame. Even when such annotation is available, this remains a challenging problem because of the changing appearance and shape of the object over time. In this paper, we tackle this task by formulating it as a meta-learning problem, where the base learner grasping the semantic scene understanding for a general type of objects, and the meta learner quickly adapting the appearance of the target object with a few examples. Our proposed meta-learning method uses a closed form optimizer, the so-called "ridge regression", which has been shown to be conducive for fast and better training convergence. Moreover, we propose a mechanism, named "block splitting", to further speed up the training process as well as to reduce the number of learning parameters. In comparison with the state-of-the art methods, our proposed framework achieves significant boost up in processing speed, while having highly comparable performance compared to the best performing methods on the widely used datasets. Video demo can be found here1.
Yu Liu 0029, Lingqiao Liu, Haokui Zhang, Seyed Hamid Rezatofighi, Qingsen Yan, Ian D. Reid 0001
IROS6
2020 Architecture Search of Dynamic Cells for Semantic Video Segmentation
abstract
In semantic video segmentation the goal is to acquire consistent dense semantic labelling across image frames. To this end, recent approaches have been reliant on manually arranged operations applied on top of static semantic segmentation networks - with the most prominent building block being the optical flow able to provide information about scene dynamics. Related to that is the line of research concerned with speeding up static networks by approximating expensive parts of them with cheaper alternatives, while propagating information from previous frames. In this work we attempt to come up with generalisation of those methods, and instead of manually designing contextual blocks that connect per-frame outputs, we propose a neural architecture search solution, where the choice of operations together with their sequential arrangement are being predicted by a separate neural network. We showcase that such generalisation leads to stable and accurate results across common benchmarks, such as CityScapes and CamVid datasets. Importantly, the proposed methodology takes only 2 GPU-days, finds high-performing cells and does not rely on the expensive optical flow computation.
Vladimir Nekrasov, Hao Chen 0041, Chunhua Shen, Ian D. Reid 0001
WACV4
2020 Template-Based Automatic Search of Compact Semantic Segmentation Architectures
abstract
Automatic search of neural architectures for various vision and natural language tasks is becoming a prominent tool as it allows to discover high-performing structures on any dataset of interest. Nevertheless, on more difficult domains, such as dense per-pixel classification, current automatic approaches are limited in their scope - due to their strong reliance on existing image classifiers they tend to search only for a handful of additional layers with discovered architectures still containing a large number of parameters. In contrast, in this work we propose a novel solution able to find light-weight and accurate segmentation architectures starting from only few blocks of a pre-trained classification network. To this end, we progressively build up a methodology that relies on templates of sets of operations, predicts which template and how many times should be applied at each step, while also generating the connectivity structure and downsampling factors. All these decisions are being made by a recurrent neural network that is rewarded based on the score of the emitted architecture on the holdout set and trained using reinforcement learning. One discovered architecture achieves 63.2% mean IoU on CamVid and 67.8% on CityScapes having only 270K parameters.
Vladimir Nekrasov, Chunhua Shen, Ian D. Reid 0001
WACV3
2020 GMS: Grid-Based Motion Statistics for Fast, Ultra-robust Feature Correspondence
abstract
Abstract Feature matching aims at generating correspondences across images, which is widely used in many computer vision tasks. Although considerable progress has been made on feature descriptors and fast matching for initial correspondence hypotheses, selecting good ones from them is still challenging and critical to the overall performance. More importantly, existing methods often take a long computational time, limiting their use in real-time applications. This paper attempts to separate true correspondences from false ones at high speed. We term the proposed method (GMS) grid-based motion Statistics, which incorporates the smoothness constraint into a statistic framework for separation and uses a grid-based implementation for fast calculation. GMS is robust to various challenging image changes, involving in viewpoint, scale, and rotation. It is also fast, e.g., take only 1 or 2 ms in a single CPU thread, even when 50K correspondences are processed. This has important implications for real-time applications. What’s more, we show that incorporating GMS into the classic feature matching and epipolar geometry estimation pipeline can significantly boost the overall performance. Finally, we integrate GMS into the well-known ORB-SLAM system for monocular initialization, resulting in a significant improvement.
Jiawang Bian, Wen-Yan Lin, Yun Liu 0011, Le Zhang 0001, Sai-Kit Yeung, Ming-Ming Cheng, Ian D. Reid 0001
Int. J. Comput. Vis.7
2020 Approximate Fisher Information Matrix to Characterize the Training of Deep Neural Networks
abstract
In this paper, we introduce a novel methodology for characterizing the performance of deep learning networks (ResNets and DenseNet) with respect to training convergence and generalization as a function of mini-batch size and learning rate for image classification. This methodology is based on novel measurements derived from the eigenvalues of the approximate Fisher information matrix, which can be efficiently computed even for high capacity deep models. Our proposed measurements can help practitioners to monitor and control the training process (by actively tuning the mini-batch size and learning rate) to allow for good training convergence and generalization. Furthermore, the proposed measurements also allow us to show that it is possible to optimize the training process with a new dynamic sampling training approach that continuously and automatically change the mini-batch size and learning rate during the training process. Finally, we show that the proposed dynamic sampling training approach has a faster training time and a competitive classification accuracy compared to the current state of the art.
Zhibin Liao, Tom Drummond, Ian D. Reid 0001, Gustavo Carneiro 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 RefineNet: Multi-Path Refinement Networks for Dense Prediction
abstract
Recently, very deep convolutional neural networks (CNNs) have shown outstanding performance in object recognition and have also been the first choice for dense prediction problems such as semantic segmentation and depth estimation. However, repeated subsampling operations like pooling or convolution striding in deep CNNs lead to a significant decrease in the initial image resolution. Here, we present RefineNet, a generic multi-path refinement network that explicitly exploits all the information available along the down-sampling process to enable high-resolution prediction using long-range residual connections. In this way, the deeper layers that capture high-level semantic features can be directly refined using fine-grained features from earlier convolutions. The individual components of RefineNet employ residual connections following the identity mapping mindset, which allows for effective end-to-end training. Further, we introduce chained residual pooling, which captures rich background context in an efficient manner. We carry out comprehensive experiments on semantic segmentation which is a dense classification problem and achieve good performance on seven public datasets. We further apply our method for depth estimation and demonstrate the effectiveness of our method on dense regression problems.
Guosheng Lin, Fayao Liu, Anton Milan, Chunhua Shen, Ian D. Reid 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2020 Model-Free Tracker for Multiple Objects Using Joint Appearance and Motion Inference
abstract
Model-free tracking is a widely-accepted approach to track an arbitrary object in a video using a single frame annotation with no further prior knowledge about the object of interest. Extending this problem to track multiple objects is really challenging because: a) the tracker is not aware of the objects' type while trying to distinguish them from background (detection task), and b) The tracker needs to distinguish one object from other potentially similar objects (data association task) to generate stable trajectories. In order to track multiple arbitrary objects, most existing model-free tracking approaches rely on tracking each target individually by updating their appearance model independently. Therefore, in this scenario they often fail to perform well due to confusion between the appearance of similar objects, their sudden appearance changes and occlusion. To tackle this problem, we propose to use both appearance and motion models, and to learn them jointly using graphical models and deep neural networks features. We introduce an indicator variable to predict sudden appearance change and/or occlusion. When these happen, our model does not update the appearance model thus avoiding using the background and/or incorrect object to update the appearance of the object of interest mistakenly, and relies on our motion model to track. Moreover, we consider the correlation among all targets, and seek the joint optimal locations for all targets simultaneously as a graphical model inference problem. We learn the joint parameters for both appearance model and motion model in an online fashion under the framework of LaRank. Experiment results show that our method achieved superior performance compared to the competitive methods.
Chongyu Liu, Rui Yao 0006, Seyed Hamid Rezatofighi, Ian D. Reid 0001, Qinfeng Shi
IEEE Trans. Image Process.4
2020 Real-time Image Smoothing via Iterative Least Squares
abstract
Edge-preserving image smoothing is a fundamental procedure for many computer vision and graphic applications. There is a tradeoff between the smoothing quality and the processing speed: the high smoothing quality usually requires a high computational cost, which leads to the low processing speed. In this article, we propose a new global optimization based method, named iterative least squares (ILS), for efficient edge-preserving image smoothing. Our approach can produce high-quality results but at a much lower computational cost. Comprehensive experiments demonstrate that the proposed method can produce results with little visible artifacts. Moreover, the computation of ILS can be highly parallel, which can be easily accelerated through either multi-thread computing or the GPU hardware. With the acceleration of a GTX 1080 GPU, it is able to process images of 1080p resolution (1920 × 1080) at the rate of 20fps for color images and 47fps for gray images. In addition, the ILS is flexible and can be modified to handle more applications that require different smoothing properties. Experimental results of several applications show the effectiveness and efficiency of the proposed method. The code is available at https://github.com/wliusjtu/Real-time-Image-Smoothing-via-Iterative-Least-Squares.
Wei Liu 0044, Xiaolin Huang, Jie Yang 0002, Chunhua Shen, Ian D. Reid 0001
ACM Trans. Graph.6
2019 An Evaluation of Feature Matchers for Fundamental Matrix Estimation
Jiawang Bian, Yu-Huan Wu, Ji Zhao 0001, Yun Liu 0011, Le Zhang 0001, Ming-Ming Cheng, Ian D. Reid 0001
BMVC7
2019 Single-view Object Shape Reconstruction Using Deep Shape Prior and Silhouette
Kejie Li, Ravi Garg, Ian D. Reid 0001
BMVC4
2019 A Theoretically Sound Upper Bound on the Triplet Loss for Improving the Efficiency of Deep Distance Metric Learning
abstract
We propose a method that substantially improves the efficiency of deep distance metric learning based on the optimization of the triplet loss function. One epoch of such training process based on a na¨ıve optimization of the triplet loss function has a run-time complexity O(N^3), where N is the number of training samples. Such optimization scales poorly, and the most common approach proposed to address this high complexity issue is based on sub-sampling the set of triplets needed for the training process. Another approach explored in the field relies on an ad-hoc linearization (in terms of N) of the triplet loss that introduces class centroids, which must be optimized using the whole training set for each mini-batch – this means that a na¨ıve implementation of this approach has run-time complexity O(N^2). This complexity issue is usually mitigated with poor, but computationally cheap, approximate centroid optimization methods. In this paper, we first propose a solid theory on the linearization of the triplet loss with the use of class centroids, where the main conclusion is that our new linear loss represents a tight upper-bound to the triplet loss. Furthermore, based on the theory above, we propose a training algorithm that no longer requires the centroid optimization step, which means that our approach is the first in the field with a guaranteed linear run-time complexity. We show that the training of deep distance metric learning methods using the proposed upper-bound is substantially faster than triplet-based methods, while producing competitive retrieval accuracy results on benchmark datasets (CUB-200-2011 and CAR196).
Thanh-Toan Do, Toan Tran 0002, Ian D. Reid 0001, Tuan Hoang, Gustavo Carneiro 0001
CVPR3
2019 RGBD Based Dimensional Decomposition Residual Network for 3D Semantic Scene Completion
abstract
RGB images differentiate from depth as they carry more details about the color and texture information, which can be utilized as a vital complement to depth for boosting the performance of 3D semantic scene completion (SSC). SSC is composed of 3D shape completion (SC) and semantic scene labeling while most of the existing approaches use depth as the sole input which causes the performance bottleneck. Moreover, the state-of-the-art methods employ 3D CNNs which have cumbersome networks and tremendous parameters. We introduce a light-weight Dimensional Decomposition Residual network (DDR) for 3D dense prediction tasks. The novel factorized convolution layer is effective for reducing the network parameters, and the proposed multi-scale fusion mechanism for depth and color image can improve the completion and segmentation accuracy simultaneously. Our method demonstrates excellent performance on two public datasets. Compared with the latest method SSCNet, we achieve 5.9% gains in SC-IoU and 5.7% gains in SSC-IOU, albeit with only 21% network parameters and 16.6% FLOPs employed compared with that of SSCNet.
Jie Li 0040, Yu Liu 0029, Dong Gong, Qinfeng Shi, Xia Yuan, Chunxia Zhao, Ian D. Reid 0001
CVPR7
2019 Fast Neural Architecture Search of Compact Semantic Segmentation Models via Auxiliary Cells
abstract
Automated design of neural network architectures tailored for a specific task is an extremely promising, albeit inherently difficult, avenue to explore. While most results in this domain have been achieved on image classification and language modelling problems, here we concentrate on dense per-pixel tasks, in particular, semantic image segmentation using fully convolutional networks. In contrast to the aforementioned areas, the design choices of a fully convolutional network require several changes, ranging from the sort of operations that need to be used - e.g., dilated convolutions - to a solving of a more difficult optimisation problem. In this work, we are particularly interested in searching for high-performance compact segmentation architectures, able to run in real-time using limited resources. To achieve that, we intentionally over-parameterise the architecture during the training time via a set of auxiliary cells that provide an intermediate supervisory signal and can be omitted during the evaluation phase. The design of the auxiliary cell is emitted by a controller, a neural network with the fixed structure trained using reinforcement learning. More crucially, we demonstrate how to efficiently search for these architectures within limited time and computational budgets. In particular, we rely on a progressive strategy that terminates non-promising architectures from being further trained, and on Polyak averaging coupled with knowledge distillation to speed-up the convergence. Quantitatively, in 8 GPU-days our approach discovers a set of architectures performing on-par with state-of-the-art among compact models on the semantic segmentation, pose estimation and depth prediction tasks. Code will be made available here: https://github.com/drsleep/nas-segm-pytorch.
Vladimir Nekrasov, Hao Chen 0041, Chunhua Shen, Ian D. Reid 0001
CVPR4
2019 Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression
abstract
Intersection over Union (IoU) is the most popular evaluation metric used in the object detection benchmarks. However, there is a gap between optimizing the commonly used distance losses for regressing the parameters of a bounding box and maximizing this metric value. The optimal objective for a metric is the metric itself. In the case of axis-aligned 2D bounding boxes, it can be shown that IoU can be directly used as a regression loss. However, IoU has a plateau making it infeasible to optimize in the case of non-overlapping bounding boxes. In this paper, we address the this weakness by introducing a generalized version of IoU as both a new loss and a new metric. By incorporating this generalized IoU (GIoU) as a loss into the state-of-the art object detection frameworks, we show a consistent improvement on their performance using both the standard, IoU based, and new, GIoU based, performance measures on popular object detection benchmarks such as PASCAL VOC and MS COCO.
Seyed Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian D. Reid 0001, Silvio Savarese
CVPR5
2019 TopNet: Structural Point Cloud Decoder
abstract
3D point cloud generation is of great use for 3D scene modeling and understanding. Real-world 3D object point clouds can be properly described by a collection of low-level and high-level structures such as surfaces, geometric primitives, semantic parts,etc. In fact, there exist many different representations of a 3D object point cloud as a set of point groups. Existing frameworks for point cloud genera-ion either do not consider structure in their proposed solutions, or assume and enforce a specific structure/topology,e.g. a collection of manifolds or surfaces, for the generated point cloud of a 3D object. In this work, we pro-pose a novel decoder that generates a structured point cloud without assuming any specific structure or topology on the underlying point set. Our decoder is softly constrained to generate a point cloud following a hierarchical rooted tree structure. We show that given enough capacity and allowing for redundancies, the proposed decoder is very flexible and able to learn any arbitrary grouping of points including any topology on the point set. We evaluate our decoder on the task of point cloud generation for 3D point cloud shape completion. Combined with encoders from existing frameworks, we show that our proposed decoder significantly outperforms state-of-the-art 3D point cloud completion methods on the Shapenet dataset.
Lyne P. Tchapmi, Vineet Kosaraju, Seyed Hamid Rezatofighi, Ian D. Reid 0001, Silvio Savarese
CVPR4
2019 Attention-Guided Network for Ghost-Free High Dynamic Range Imaging
abstract
Ghosting artifacts caused by moving objects or misalignments is a key challenge in high dynamic range (HDR) imaging for dynamic scenes. Previous methods first register the input low dynamic range (LDR) images using optical flow before merging them, which are error-prone and cause ghosts in results. A very recent work tries to bypass optical flows via a deep network with skip-connections, however, which still suffers from ghosting artifacts for severe movement. To avoid the ghosting from the source, we propose a novel attention-guided end-to-end deep neural network (AHDRNet) to produce high-quality ghost-free HDR images. Unlike previous methods directly stacking the LDR images or features for merging, we use attention modules to guide the merging according to the reference image. The attention modules automatically suppress undesired components caused by misalignments and saturation and enhance desirable fine details in the non-reference images. In addition to the attention model, we use dilated residual dense block (DRDB) to make full use of the hierarchical features and increase the receptive field for hallucinating the missing details. The proposed AHDRNet is a non-flow-based method, which can also avoid the artifacts generated by optical-flow estimation error. Experiments on different datasets show that the proposed AHDRNet can achieve state-of-the-art quantitative and qualitative results.
Qingsen Yan, Dong Gong, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Ian D. Reid 0001, Yanning Zhang 0001
CVPR6
2019 Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
abstract
In this paper, we propose to train convolutional neural networks (CNNs) with both binarized weights and activations, leading to quantized models specifically for mobile devices with limited power capacity and computation resources. By assuming the same architecture to full-precision networks, previous works on quantizing CNNs seek to preserve the floating-point information using a set of discrete values, which we call value approximation. However, we take a novel ``structure approximation'' view for quantization--- it is very likely that a different architecture may be better for best performance. In particular, we propose a ``network decomposition'' strategy, named Group-Net, in which we divide the network into groups. In this way, each full-precision group can be effectively reconstructed by aggregating a set of homogeneous binary branches. In addition, we learn effect connections among groups to improve the representational capability. Moreover, the proposed Group-Net shows strong generalization to other tasks. For instance, we extend Group-Net for highly accurate semantic segmentation by embedding rich context into the binary structure. Experiments on both classification and semantic segmentation tasks demonstrate the superior performance of the proposed methods over various popular architectures. In particular, we outperform the previous best binary neural networks in terms of accuracy and huge computation saving.
Bohan Zhuang, Chunhua Shen, Mingkui Tan, Lingqiao Liu, Ian D. Reid 0001
CVPR5
2019 Scalable Place Recognition Under Appearance Change for Autonomous Driving
abstract
A major challenge in place recognition for autonomous driving is to be robust against appearance changes due to short-term (e.g., weather, lighting) and long-term (seasons, vegetation growth, etc.) environmental variations. A promising solution is to continuously accumulate images to maintain an adequate sample of the conditions and incorporate new changes into the place recognition decision. However, this demands a place recognition technique that is scalable on an ever growing dataset. To this end, we propose a novel place recognition technique that can be efficiently retrained and compressed, such that the recognition of new queries can exploit all available data (including recent changes) without suffering from visible growth in computational cost. Underpinning our method is a novel temporal image matching technique based on Hidden Markov Models. Our experiments show that, compared to state-of-the-art techniques, our method has much greater potential for large-scale place recognition for autonomous driving.
Anh-Dzung Doan, Yasir Latif, Tat-Jun Chin, Yu Liu 0029, Thanh-Toan Do, Ian D. Reid 0001
ICCV6
2019 Bayesian Generative Active Deep Learning
abstract
Deep learning models have demonstrated outstanding performance in several problems, but their training process tends to require immense amounts of computational and human resources for training and labeling, constraining the types of problems that can be tackled. Therefore, the design of effective training methods that require small labeled training sets is an important research direction that will allow a more effective use of resources. Among current approaches designed to address this issue, two are particularly interesting: data augmentation and active learning. Data augmentation achieves this goal by artificially generating new training points, while active learning relies on the selection of the “most informative” subset of unlabeled training samples to be labelled by an oracle. Although successful in practice, data augmentation can waste computational resources because it indiscriminately generates samples that are not guaranteed to be informative, and active learning selects a small subset of informative samples (from a large un-annotated set) that may be insufficient for the training process. In this paper, we propose a Bayesian generative active deep learning approach that combines active learning with data augmentation – we provide theoretical and empirical evidence (MNIST, CIFAR-$\{10,100\}$, and SVHN) that our approach has more efficient training and better classification results than data augmentation and active learning.
Toan Tran 0002, Thanh-Toan Do, Ian D. Reid 0001, Gustavo Carneiro 0001
ICML3
2019 Visual SLAM: Why Bundle Adjust?
abstract
Bundle adjustment plays a vital role in feature-based monocular SLAM. In many modern SLAM pipelines, bundle adjustment is performed to estimate the 6DOF camera trajectory and 3D map (3D point cloud) from the input feature tracks. However, two fundamental weaknesses plague SLAM systems based on bundle adjustment. First, the need to carefully initialise bundle adjustment means that all variables, in particular the map, must be estimated as accurately as possible and maintained over time, which makes the overall algorithm cumbersome. Second, since estimating the 3D structure (which requires sufficient baseline) is inherent in bundle adjustment, the SLAM algorithm will encounter difficulties during periods of slow motion or pure rotational motion. We propose a different SLAM optimisation core: instead of bundle adjustment, we conduct rotation averaging to incrementally optimise only camera orientations. Given the orientations, we estimate the camera positions and 3D points via a quasi-convex formulation that can be solved efficiently and globally optimally. Our approach not only obviates the need to estimate and maintain the positions and 3D map at keyframe rate (which enables simpler SLAM systems), it is also more capable of handling slow motions or pure rotational motions.
Álvaro Parra Bustos, Tat-Jun Chin, Anders P. Eriksson, Ian D. Reid 0001
ICRA4
2019 Real-Time Monocular Object-Model Aware Sparse SLAM
abstract
Simultaneous Localization And Mapping (SLAM) is a fundamental problem in mobile robotics. While sparse point-based SLAM methods provide accurate camera localization, the generated maps lack semantic information. On the other hand, state of the art object detection methods provide rich information about entities present in the scene from a single image. This work incorporates a real-time deep-learned object detector to the monocular SLAM framework for representing generic objects as quadrics that permit detections to be seamlessly integrated while allowing the real-time performance. Finer reconstruction of an object, learned by a CNN network, is also incorporated and provides a shape prior for the quadric leading further refinement. To capture the structure of the scene, additional planar landmarks are detected by a CNN-based plane detector and modelled as independent landmarks in the map. Extensive experiments support our proposed inclusion of semantic objects and planar structures directly in the bundle-adjustment of SLAM - Semantic SLAM- that enriches the reconstructed map semantically, while significantly improving the camera localization.
Mehdi Hosseinzadeh 0003, Kejie Li, Yasir Latif, Ian D. Reid 0001
ICRA4
2019 Real-Time Joint Semantic Segmentation and Depth Estimation Using Asymmetric Annotations
abstract
Deployment of deep learning models in robotics as sensory information extractors can be a daunting task to handle, even using generic GPU cards. Here, we address three of its most prominent hurdles, namely, i) the adaptation of a single model to perform multiple tasks at once (in this work, we consider depth estimation and semantic segmentation crucial for acquiring geometric and semantic understanding of the scene), while ii) doing it in real-time, and iii) using asymmetric datasets with uneven numbers of annotations per each modality. To overcome the first two issues, we adapt a recently proposed real-time semantic segmentation network, making changes to further reduce the number of floating point operations. To approach the third issue, we embrace a simple solution based on hard knowledge distillation under the assumption of having access to a powerful `teacher' network. We showcase how our system can be easily extended to handle more tasks, and more datasets, all at once, performing depth estimation and segmentation both indoors and outdoors with a single model. Quantitatively, we achieve results equivalent to (or better than) current state-of-the-art approaches with one forward pass costing just 13ms and 6.5 GFLOPs on 640×480 inputs. This efficiency allows us to directly incorporate the raw predictions of our network into the SemanticFusion framework [1] for dense 3D semantic reconstruction of the scene.
Vladimir Nekrasov, Thanuja Dharmasiri, Andrew Spek, Tom Drummond, Chunhua Shen, Ian D. Reid 0001
ICRA6
2019 Self-supervised Learning for Single View Depth and Surface Normal Estimation
abstract
In this work we present a self-supervised learning framework to simultaneously train two Convolutional Neural Networks (CNNs) to predict depth and surface normals from a single image. In contrast to most existing frameworks which represent outdoor scenes as fronto-parallel planes at piece-wise smooth depth, we propose to predict depth with surface orientation while assuming that natural scenes have piece-wise smooth normals. We show that a simple depth-normal consistency as a soft-constraint on the predictions is sufficient and effective for training both these networks simultaneously. The trained normal network provides state-of-the-art predictions while the depth network, relying on much realistic smooth normal assumption, outperforms the traditional self-supervised depth prediction network by a large margin on the KITTI benchmark.
Huangying Zhan, Chamara Saroj Weerasekera, Ravi Garg, Ian D. Reid 0001
ICRA4
2019 Seeing Behind Things: Extending Semantic Segmentation to Occluded Regions
abstract
Semantic segmentation and instance level segmentation made substantial progress in recent years due to the emergence of deep neural networks (DNNs). A number of deep architectures with Convolution Neural Networks (CNNs) were proposed that surpass the traditional machine learning approaches for segmentation by a large margin. These architectures predict the directly observable semantic category of each pixel by usually optimizing a cross-entropy loss. In this work we push the limit of semantic segmentation towards predicting semantic labels of directly visible as well as occluded objects or objects parts, where the network's input is a single depth image. We group the semantic categories into one background and multiple foreground object groups, and we propose a modification of the standard cross-entropy loss to cope with the settings. In our experiments we demonstrate that a CNN trained by minimizing the proposed loss is able to predict semantic categories for visible and occluded object parts without requiring to increase the network size (compared to a standard segmentation task). The results are validated on a newly generated dataset (augmented from SUNCG) dataset.
Pulak Purkait, Christopher Zach, Ian D. Reid 0001
IROS3
2019 Unsupervised Scale-consistent Depth and Ego-motion Learning from Monocular Video
abstract
Recent work has shown that CNN-based depth and ego-motion estimators can be learned using unlabelled monocular videos. However, the performance is limited by unidentified moving objects that violate the underlying static scene assumption in geometric image reconstruction. More significantly, due to lack of proper constraints, networks output scale-inconsistent results over different samples, i.e., the ego-motion network cannot provide full camera trajectories over a long video sequence because of the per-frame scale ambiguity. This paper tackles these challenges by proposing a geometry consistency loss for scale-consistent predictions and an induced self-discovered mask for handling moving objects and occlusions. Since we do not leverage multi-task learning like recent works, our framework is much simpler and more efficient. Comprehensive evaluation results demonstrate that our depth estimator achieves the state-of-the-art performance on the KITTI dataset. Moreover, we show that our ego-motion network is able to predict a globally scale-consistent camera trajectory for long video sequences, and the resulting visual odometry accuracy is competitive with the recent model that is trained using stereo videos. To the best of our knowledge, this is the first work to show that deep networks trained using unlabelled monocular videos can predict globally scale-consistent camera trajectories over a long video sequence.
Jiawang Bian, Naiyan Wang, Huangying Zhan, Chunhua Shen, Ming-Ming Cheng, Ian D. Reid 0001
NeurIPS7
2019 Social-BiGAT: Multimodal Trajectory Forecasting using Bicycle-GAN and Graph Attention Networks
abstract
Predicting the future trajectories of multiple interacting pedestrians in a scene has become an increasingly important problem for many different applications ranging from control of autonomous vehicles and social robots to security and surveillance. This problem is compounded by the presence of social interactions between humans and their physical interactions with the scene. While the existing literature has explored some of these cues, they mainly ignored the multimodal nature of each human's future trajectory which is noticeably influenced by the intricate social interactions. In this paper, we present Social-BiGAT, a graph-based generative adversarial network that generates realistic, multimodal trajectory predictions for multiple pedestrians in a scene. Our method is based on a graph attention network (GAT) that learns feature representations that encode the social interactions between humans in the scene, and a recurrent encoder-decoder architecture that is trained adversarially to predict, based on the features, the humans' paths. We explicitly account for the multimodal nature of the prediction problem by forming a reversible transformation between each scene and its latent noise vector, as in Bicycle-GAN. We show that our framework achieves state-of-the-art performance comparing it to several baselines on existing trajectory forecasting benchmarks.
Vineet Kosaraju, Amir Sadeghian, Roberto Martin Martin, Ian D. Reid 0001, Seyed Hamid Rezatofighi, Silvio Savarese
NeurIPS4
2019 Binary Constrained Deep Hashing Network for Image Retrieval Without Manual Annotation
abstract
Learning compact binary codes for image retrieval task using deep neural networks has attracted increasing attention recently. However, training deep hashing networks for the task is challenging due to the binary constraints on the hash codes, the similarity preserving property, and the requirement for a vast amount of labelled images. To the best of our knowledge, none of the existing methods has tackled all of these challenges completely in a unified framework. In this work, we propose a novel end-to-end deep learning approach for the task, in which the network is trained to produce binary codes directly from image pixels without the need o f manual annotation. In particular, to deal with the non-smoothness of binary constraints, we propose a novel pairwise constrained loss function, which simultaneously encodes the distances between pairs of hash codes, and the binary quantization error. In order to train the network with the proposed loss function, we propose an efficient parameter learning algorithm. In addition, to provide similar / dissimilar training images to train the network, we exploit 3D models reconstructed from unlabelled images for automatic generation of enormous training image pairs. The extensive experiments on image retrieval benchmark datasets demonstrate the improvements of the proposed method over the state-of-the-art compact representation methods on the image retrieval problem.
Thanh-Toan Do, Tuan Hoang, Dang-Khoa Le Tan, Trung Pham, Huu Le, Ngai-Man Cheung, Ian D. Reid 0001
WACV7
2019 Multi-Scale Dense Networks for Deep High Dynamic Range Imaging
abstract
Generating a high dynamic range (HDR) image from a set of sequential exposures is a challenging task for dynamic scenes. The most common approaches are aligning the input images to a reference image before merging them into an HDR image, but artifacts often appear in cases of large scene motion. The state-of-the-art method using deep learning can solve this problem effectively. In this paper, we propose a novel deep convolutional neural network to generate HDR, which attempts to produce more vivid images. The key idea of our method is using the coarse-to-fine scheme to gradually reconstruct the HDR image with the multi-scale architecture and residual network. By learning the relative changes of inputs and ground truth, our method can produce not only artificial free image but also restore missing information. Furthermore, we compare to existing methods for HDR reconstruction, and show high-quality results from a set of low dynamic range (LDR) images. We evaluate the results in qualitative and quantitative experiments, our method consistently produces excellent results than existing state-of-the-art approaches in challenging scenes.
Qingsen Yan, Dong Gong, Qinfeng Shi, Jinqiu Sun, Ian D. Reid 0001, Yanning Zhang 0001
WACV6
2019 Pre and post-hoc diagnosis and interpretation of malignancy from breast DCE-MRI
Gabriel Maicas, Andrew P. Bradley, Jacinto C. Nascimento, Ian D. Reid 0001, Gustavo Carneiro 0001
Medical Image Anal.4
2019 Multi-Task Structure-Aware Context Modeling for Robust Keypoint-Based Object Tracking
abstract
In the fields of computer vision and graphics, keypoint-based object tracking is a fundamental and challenging problem, which is typically formulated in a spatio-temporal context modeling framework. However, many existing keypoint trackers are incapable of effectively modeling and balancing the following three aspects in a simultaneous manner: temporal model coherence across frames, spatial model consistency within frames, and discriminative feature construction. To address this problem, we propose a robust keypoint tracker based on spatio-temporal multi-task structured output optimization driven by discriminative metric learning. Consequently, temporal model coherence is characterized by multi-task structured keypoint model learning over several adjacent frames; spatial model consistency is modeled by solving a geometric verification based structured learning problem; discriminative feature construction is enabled by metric learning to ensure the intra-class compactness and inter-class separability. To achieve the goal of effective object tracking, we jointly optimize the above three modules in a spatio-temporal multi-task learning scheme. Furthermore, we incorporate this joint learning scheme into both single-object and multi-object tracking scenarios, resulting in robust tracking results. Experiments over several challenging datasets have justified the effectiveness of our single-object and multi-object trackers against the state-of-the-art.
Xi Li 0001, Wei Ji 0008, Yiming Wu 0005, Fei Wu 0001, Ming-Hsuan Yang 0001, Dacheng Tao, Ian D. Reid 0001
IEEE Trans. Pattern Anal. Mach. Intell.8
2018 Joint Learning of Set Cardinality and State Distribution
abstract
We present a novel approach for learning to predict sets using deep learning. In recent years, deep neural networks have shown remarkable results in computer vision, natural language processing and other related problems. Despite their success,traditional architectures suffer from a serious limitation in that they are built to deal with structured input and output data,i.e. vectors or matrices. Many real-world problems, however, are naturally described as sets, rather than vectors. Existing techniques that allow for sequential data, such as recurrent neural networks, typically heavily depend on the input and output order and do not guarantee a valid solution. Here, we derive in a principled way, a mathematical formulation for set prediction where the output is permutation invariant. In particular, our approach jointly learns both the cardinality and the state distribution of the target set. We demonstrate the validity of our method on the task of multi-label image classification and achieve a new state of the art on the PASCAL VOC and MS COCO datasets.
Seyed Hamid Rezatofighi, Anton Milan, Qinfeng Shi, Anthony R. Dick, Ian D. Reid 0001
AAAI5
2018 HCVRD: A Benchmark for Large-Scale Human-Centered Visual Relationship Detection
abstract
Visual relationship detection aims to capture interactions between pairs of objects in images. Relationships between objects and humans represent a particularly important subset of this problem, with implications for challenges such as understanding human behavior, and identifying affordances, amongst others. In addressing this problem we first construct a large-scale human-centric visual relationship detection dataset (HCVRD), which provides many more types of relationship annotations (nearly 10K categories) than the previous released datasets. This large label space better reflects the reality of human-object interactions, but gives rise to a long-tail distribution problem, which in turn demands a zero-shot approach to labels appearing only in the test set. This is the first time this issue has been addressed. We propose a webly-supervised approach to these problems and demonstrate that the proposed model provides a strong baseline on our HCVRD dataset.
Bohan Zhuang, Qi Wu 0001, Chunhua Shen, Ian D. Reid 0001, Anton van den Hengel
AAAI4
2018 Structure Aware SLAM Using Quadrics and Planes
Mehdi Hosseinzadeh 0003, Yasir Latif, Trung Pham, Niko Sünderhauf, Ian D. Reid 0001
ACCV (3)5
2018 Learning Deeply Supervised Good Features to Match for Dense Monocular Reconstruction
Chamara Saroj Weerasekera, Ravi Garg, Yasir Latif, Ian D. Reid 0001
ACCV (5)4
2018 Scalable Deep k-Subspace Clustering
Tong Zhang 0023, Pan Ji, Mehrtash Harandi, Richard I. Hartley, Ian D. Reid 0001
ACCV (5)5
2018 A Hybrid Probabilistic Model for Camera Relocalization
Chunhua Shen, Ian D. Reid 0001
BMVC3
2018 LieNet: Real-time Monocular Object Instance 6D Pose Estimation
Thanh-Toan Do, Trung Pham, Ian D. Reid 0001
BMVC4
2018 Light-Weight RefineNet for Real-Time Semantic Segmentation
Vladimir Nekrasov, Chunhua Shen, Ian D. Reid 0001
BMVC3
2018 Visual Question Answering With Memory-Augmented Networks
abstract
In this paper, we exploit memory-augmented neural networks to predict accurate answers to visual questions, even when those answers rarely occur in the training set. The memory network incorporates both internal and external memory blocks and selectively pays attention to each training exemplar. We show that memory-augmented neural networks are able to maintain a relatively long-term memory of scarce training exemplars, which is important for visual question answering due to the heavy-tailed distribution of answers in a general VQA setting. Experimental results in two large-scale benchmark datasets show the favorable performance of the proposed algorithm with the comparison to state of the art.
Chao Ma 0004, Chunhua Shen, Anthony R. Dick, Qi Wu 0001, Peng Wang 0023, Anton van den Hengel, Ian D. Reid 0001
CVPR7
2018 Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments
abstract
A robot that can carry out a natural-language instruction has been a dream since before the Jetsons cartoon series imagined a life of leisure mediated by a fleet of attentive robot helpers. It is a dream that remains stubbornly distant. However, recent advances in vision and language methods have made incredible progress in closely related areas. This is significant because a robot interpreting a natural-language navigation instruction on the basis of what it sees is carrying out a vision and language process that is similar to Visual Question Answering. Both tasks can be interpreted as visually grounded sequence-to-sequence translation problems, and many of the same methods are applicable. To enable and encourage the application of vision and language methods to the problem of interpreting visually-grounded navigation instructions, we present the Matter-port3D Simulator - a large-scale reinforcement learning environment based on real imagery [11]. Using this simulator, which can in future support a range of embodied vision and language tasks, we provide the first benchmark dataset for visually-grounded natural language navigation in real buildings - the Room-to-Room (R2R) dataset1.
Peter Anderson 0001, Qi Wu 0001, Damien Teney, Jake Bruce, Mark Johnson 0001, Niko Sünderhauf, Ian D. Reid 0001, Stephen Gould, Anton van den Hengel
CVPR7
2018 Bootstrapping the Performance of Webly Supervised Semantic Segmentation
abstract
Fully supervised methods for semantic segmentation require pixel-level class masks to train, the creation of which is expensive in terms of manual labour and time. In this work, we focus on weak supervision, developing a method for training a high-quality pixel-level classifier for semantic segmentation, using only image-level class labels as the provided ground-truth. Our method is formulated as a two-stage approach in which we first aim to create accurate pixel-level masks for the training images via a bootstrapping process, and then use these now-accurately segmented images as a proxy ground-truth in a more standard supervised setting. The key driver for our work is that in the target dataset we typically have reliable ground-truth image-level labels, while data crawled from the web may have unreliable labels, but can be filtered to comprise only easy images to segment, therefore having reliable boundaries. These two forms of information are complementary and we use this observation to build a novel bi-directional transfer learning framework. This framework transfers knowledge between two domains, target domain and web domain, bootstrapping the performance of weakly supervised semantic segmentation. Conducting experiments on the popular benchmark dataset PASCAL VOC 2012 based on both a VGG16 network and on ResNet50, we reach state-of-the-art performance with scores of 60.2% IoU and 63.9% IoU respectively1.
Guosheng Lin, Chunhua Shen, Ian D. Reid 0001
CVPR4
2018 Are You Talking to Me? Reasoned Visual Dialog Generation Through Adversarial Learning
abstract
The visual dialog task requires an agent to engage in a conversation about an image with a human. It represents an extension of the visual question answering task in that the agent needs to answer a question about an image, but it needs to do so in light of the previous dialog that has taken place. The key challenge in visual dialog is thus maintaining a consistent, and natural dialog while continuing to answer questions correctly. We present a novel approach that combines Reinforcement Learning and Generative Adversarial Networks (GANS) to generate more human-like responses to questions. The GAN helps overcome the relative paucity of training data, and the tendency of the typical MLE-based approach to generate overly terse answers. Critically, the GAN is tightly integrated into the attention mechanism that generates human-interpretable reasons for each answer. This means that the discriminative model of the GAN has the task of assessing whether a candidate answer is generated by a human or not, given the provided reason. This is significant because it drives the generative model to produce high quality answers that are well supported by the associated reasoning. The method also generates the state-of-the-art results on the primary benchmark.
Qi Wu 0001, Peng Wang 0015, Chunhua Shen, Ian D. Reid 0001, Anton van den Hengel
CVPR4
2018 Unsupervised Learning of Monocular Depth Estimation and Visual Odometry With Deep Feature Reconstruction
abstract
Despite learning based methods showing promising results in single view depth estimation and visual odometry, most existing approaches treat the tasks in a supervised manner. Recent approaches to single view depth estimation explore the possibility of learning without full supervision via minimizing photometric error. In this paper, we explore the use of stereo sequences for learning depth and visual odometry. The use of stereo sequences enables the use of both spatial (between left-right pairs) and temporal (forward backward) photometric warp error, and constrains the scene depth and camera motion to be in a common, real-world scale. At test time our framework is able to estimate single view depth and two-view odometry from a monocular sequence. We also show how we can improve on a standard photometric warp loss by considering a warp of deep features. We show through extensive experiments that: (i) jointly training for single view depth and visual odometry improves depth prediction because of the additional constraint imposed on depths and achieves competitive results for visual odometry; (ii) deep feature-based warping loss improves upon simple photometric warp loss for both single view depth estimation and visual odometry. Our method outperforms existing learning based methods on the KITTI driving dataset in both tasks. The source code is available at https://github.com/Huangying-Zhan/Depth-VO-Feat.
Huangying Zhan, Ravi Garg, Chamara Saroj Weerasekera, Kejie Li, Ian D. Reid 0001
CVPR6
2018 Towards Effective Low-Bitwidth Convolutional Neural Networks
abstract
This paper tackles the problem of training a deep convolutional neural network with both low-precision weights and low-bitwidth activations. Optimizing a low-precision network is very challenging since the training process can easily get trapped in a poor local minima, which results in substantial accuracy loss. To mitigate this problem, we propose three simple-yet-effective approaches to improve the network training. First, we propose to use a two-stage optimization strategy to progressively find good local minima. Specifically, we propose to first optimize a net with quantized weights and then quantized activations. This is in contrast to the traditional methods which optimize them simultaneously. Second, following a similar spirit of the first method, we propose another progressive optimization approach which progressively decreases the bit-width from high-precision to low-precision during the course of training. Third, we adopt a novel learning scheme to jointly train a full-precision model alongside the low-precision one. By doing so, the full-precision model provides hints to guide the low-precision model training. Extensive experiments on various datasets (i.e., CIFAR-100 and ImageNet) show the effectiveness of the proposed methods. To highlight, using our methods to train a 4-bit precision network leads to no performance decrease in comparison with its full-precision counterpart with standard network architectures (i.e., AlexNet and ResNet-50).
Bohan Zhuang, Chunhua Shen, Mingkui Tan, Lingqiao Liu, Ian D. Reid 0001
CVPR5
2018 Parallel Attention: A Unified Framework for Visual Object Discovery Through Dialogs and Queries
abstract
Recognising objects according to a pre-defined fixed set of class labels has been well studied in the Computer Vision. There are a great many practical applications where the subjects that may be of interest are not known beforehand, or so easily delineated, however. In many of these cases natural language dialog is a natural way to specify the subject of interest, and the task achieving this capability (a.k.a, Referring Expression Comprehension) has recently attracted attention. To this end we propose a unified framework, the ParalleL AttentioN (PLAN) network, to discover the object in an image that is being referred to in variable length natural expression descriptions, from short phrases query to long multi-round dialogs. The PLAN network has two attention mechanisms that relate parts of the expressions to both the global visual content and also directly to object candidates. Furthermore, the attention mechanisms are recurrent, making the referring process visualizable and explainable. The attended information from these dual sources are combined to reason about the referred object. These two attention mechanisms can be trained in parallel and we find the combined system outperforms the state-of-art on several benchmarked datasets with different length language input, such as RefCOCO, RefCOCO+ and GuessWhat?!.
Bohan Zhuang, Qi Wu 0001, Chunhua Shen, Ian D. Reid 0001, Anton van den Hengel
CVPR4
2018 Multi-modal Cycle-Consistent Generalized Zero-Shot Learning
Rafael Felix, Ian D. Reid 0001, Gustavo Carneiro 0001
ECCV (6)3
2018 Efficient Dense Point Cloud Object Reconstruction Using Deformation Vector Fields
Kejie Li, Trung Pham, Huangying Zhan, Ian D. Reid 0001
ECCV (12)4
2018 Deep Regression Tracking with Shrinkage Loss
Xiankai Lu, Chao Ma 0004, Bingbing Ni, Xiaokang Yang 0001, Ian D. Reid 0001, Ming-Hsuan Yang 0001
ECCV (14)5
2018 Bayesian Semantic Instance Segmentation in Open Set World
Trung Pham, Thanh-Toan Do, Gustavo Carneiro 0001, Ian D. Reid 0001
ECCV (10)5
2018 AffordanceNet: An End-to-End Deep Learning Approach for Object Affordance Detection
abstract
We propose AffordanceNet, a new deep learning approach to simultaneously detect multiple objects and their affordances from RGB images. Our AffordanceNet has two branches: an object detection branch to localize and classify the object, and an affordance detection branch to assign each pixel in the object to its most probable affordance label. The proposed framework employs three key components for effectively handling the multiclass problem in the affordance mask: a sequence of deconvolutional layers, a robust resizing strategy, and a multi-task loss function. The experimental results on the public datasets show that our AffordanceNet outperforms recent state-of-the-art methods by a fair margin, while its end-to-end architecture allows the inference at the speed of 150ms per image. This makes our AffordanceNet well suitable for real-time robotic applications. Furthermore, we demonstrate the effectiveness of AffordanceNet in different testing environments and in real robotic applications. The source code is available at https://github.com/nqanh/affordance-net.
Thanh-Toan Do, Anh Nguyen 0003, Ian D. Reid 0001
ICRA3
2018 Addressing Challenging Place Recognition Tasks Using Generative Adversarial Networks
abstract
Place recognition is an essential component of Simultaneous Localization And Mapping (SLAM). Under severe appearance change, reliable place recognition is a difficult perception task since the same place is perceptually very different in the morning, at night, or over different seasons. This work addresses place recognition as a domain translation task. Using a pair of coupled Generative Adversarial Networks (GANs), we show that it is possible to generate the appearance of one domain (such as summer) from another (such as winter) without requiring image-to-image correspondences across the domains. Mapping between domains is learned from sets of images in each domain without knowing the instance-to-instance correspondence by enforcing a cyclic consistency constraint. In the process, meaningful feature spaces are learned for each domain, the distances in which can be used for the task of place recognition. Experiments show that learned features correspond to visual similarity and can be effectively used for place recognition across seasons.
Yasir Latif, Ravi Garg, Michael Milford, Ian D. Reid 0001
ICRA4
2018 Semantic Segmentation from Limited Training Data
abstract
We present our approach for robotic perception in cluttered scenes that led to winning the recent Amazon Robotics Challenge (ARC) 2017. Next to small objects with shiny and transparent surfaces, the biggest challenge of the 2017 competition was the introduction of unseen categories. In contrast to traditional approaches which require large collections of annotated data and many hours of training, the task here was to obtain a robust perception pipeline with only few minutes of data acquisition and training time. To that end, we present two strategies that we explored. One is a deep metric learning approach that works in three separate steps: semantic-agnostic boundary detection, patch classification and pixel-wise voting. The other is a fully-supervised semantic segmentation approach with efficient dataset collection. We conduct an extensive analysis of the two methods on our ARC 2017 dataset. Interestingly, only few examples of each class are sufficient to fine-tune even very deep convolutional neural networks for this specific task.
Anton Milan, Trung Pham, Kumar Vijay, Douglas Morrison, Adam W. Tow, Lingqiao Liu, Jordan Erskine, Riccardo Grinover, Alec Gurman, Thomas Hunn, Norton Kelly-Boxall, Darryl Qijun Lee, Matthew McTaggart, Gerald Rallos, Andrew Razjigaev, Thomas James Rowntree, Rohan Smith, Sean Wade-McCue, Zheyu Zhuang, Chris Lehnert, Guosheng Lin, Ian D. Reid 0001, Peter I. Corke, Jürgen Leitner
ICRA23
2018 Cartman: The Low-Cost Cartesian Manipulator that Won the Amazon Robotics Challenge
abstract
The Amazon Robotics Challenge enlisted sixteen teams to each design a pick-and-place robot for autonomous warehousing, addressing development in robotic vision and manipulation. This paper presents the design of our custom-built, cost-effective, Cartesian robot system Cartman, which won first place in the competition finals by stowing 14 (out of 16) and picking all 9 items in 27 minutes, scoring a total of 272 points. We highlight our experience-centred design methodology and key aspects of our system that contributed to our competitiveness. We believe these aspects are crucial to building robust and effective robotic systems.
Douglas Morrison, Adam W. Tow, M. McTaggart, Norton Kelly-Boxall, Sean Wade-McCue, Jordan Erskine, R. Grinover, A. Gurman, T. Hunn, Anton Milan, Trung Pham, G. Rallos, A. Razjigaev, T. Rowntree, K. Vijay, Zheyu Zhuang, Chris Lehnert, Ian D. Reid 0001, Peter I. Corke, Jürgen Leitner
ICRA20
2018 SceneCut: Joint Geometric and Object Segmentation for Indoor Scenes
abstract
This paper presents SceneCut, a novel approach to jointly discover previously unseen objects and non-object surfaces using a single RGB-D image. SceneCut's joint reasoning over scene semantics and geometry allows a robot to detect and segment object instances in complex scenes where modern deep learning-based methods either fail to separate object instances, or fail to detect objects that were not seen during training. SceneCut automatically decomposes a scene into meaningful regions which either represent objects or scene surfaces. The decomposition is qualified by an unified energy function over objectness and geometric fitting. We show how this energy function can be optimized efficiently by utilizing hierarchical segmentation trees. Moreover, we leverage a pre-trained convolutional oriented boundary network to predict accurate boundaries from images, which are used to construct high-quality region hierarchies. We evaluate SceneCut on several different indoor environments, and the results show that SceneCut significantly outperforms all the existing methods.
Trung T. Pham, Thanh-Toan Do, Niko Sünderhauf, Ian D. Reid 0001
ICRA4
2018 Just-in-Time Reconstruction: Inpainting Sparse Maps Using Single View Depth Predictors as Priors
abstract
We present “just-in-time reconstruction” as realtime image-guided inpainting of a map with arbitrary scale and sparsity to generate a fully dense depth map for the image. In particular, our goal is to inpaint a sparse map - obtained from either a monocular visual SLAM system or a sparse sensor - using a single-view depth prediction network as a virtual depth sensor. We adopt a fairly standard approach to data fusion, to produce a fused depth map by performing inference over a novel fully-connected Conditional Random Field (CRF) which is parameterized by the input depth maps and their pixel-wise confidence weights. Crucially, we obtain the confidence weights that parameterize the CRF model in a data-dependent manner via Convolutional Neural Networks (CNNs) which are trained to model the conditional depth error distributions given each source of input depth map and the associated RGB image. Our CRF model penalises absolute depth error in its nodes and pairwise scale-invariant depth error in its edges, and the confidence-based fusion minimizes the impact of outlier input depth values on the fused result. We demonstrate the flexibility of our method by real-time inpainting of ORB-SLAM, Kinect, and LIDAR depth maps acquired both indoors and outdoors at arbitrary scale and varied amount of irregular sparsity.
Chamara Saroj Weerasekera, Thanuja Dharmasiri, Ravi Garg, Tom Drummond, Ian D. Reid 0001
ICRA5
2018 Training Medical Image Analysis Systems like Radiologists
Gabriel Maicas, Andrew P. Bradley, Jacinto C. Nascimento, Ian D. Reid 0001, Gustavo Carneiro 0001
MICCAI (1)4
2018 Dense 3D Face Correspondence
abstract
We present an algorithm that automatically establishes dense correspondences between a large number of 3D faces. Starting from automatically detected sparse correspondences on the outer boundary of 3D faces, the algorithm triangulates existing correspondences and expands them iteratively by matching points of distinctive surface curvature along the triangle edges. After exhausting keypoint matches, further correspondences are established by generating evenly distributed points within triangles by evolving level set geodesic curves from the centroids of large triangles. A deformable model (K3DM) is constructed from the dense corresponded faces and an algorithm is proposed for morphing the K3DM to fit unseen faces. This algorithm iterates between rigid alignment of an unseen face followed by regularized morphing of the deformable model. We have extensively evaluated the proposed algorithms on synthetic data and real 3D faces from the FRGCv2, Bosphorus, BU3DFE and UND Ear databases using quantitative and qualitative benchmarks. Our algorithm achieved dense correspondences with a mean localisation error of 1.28 mm on synthetic faces and detected 14 anthropometric landmarks on unseen real faces from the FRGCv2 database with 3 mm precision. Furthermore, our deformable model fitting algorithm achieved 98.5 percent face recognition accuracy on the FRGCv2 and 98.6 percent on Bosphorus database. Our dense model is also able to generalize to unseen datasets.
Syed Zulqarnain Gilani, Ajmal Mian, Faisal Shafait, Ian D. Reid 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Exploring Context with Deep Structured Models for Semantic Segmentation
abstract
We propose an approach for exploiting contextual information in semantic image segmentation, and particularly investigate the use of patch-patch context and patch-background context in deep CNNs. We formulate deep structured models by combining CNNs and Conditional Random Fields (CRFs) for learning the patch-patch context between image regions. Specifically, we formulate CNN-based pairwise potential functions to capture semantic correlations between neighboring patches. Efficient piecewise training of the proposed deep structured model is then applied in order to avoid repeated expensive CRF inference during the course of back propagation. For capturing the patch-background context, we show that a network design with traditional multi-scale image inputs and sliding pyramid pooling is very effective for improving performance. We perform comprehensive evaluation of the proposed method. We achieve new state-of-the-art performance on a number of challenging semantic segmentation datasets.
Guosheng Lin, Chunhua Shen, Anton van den Hengel, Ian D. Reid 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2017 Online Multi-Target Tracking Using Recurrent Neural Networks
abstract
We present a novel approach to online multi-target tracking based on recurrent neural networks (RNNs). Tracking multiple objects in real-world scenes involves many challenges, including a) an a-priori unknown and time-varying number of targets, b) a continuous state estimation of all present targets, and c) a discrete combinatorial problem of data association. Most previous methods involve complex models that require tedious tuning of parameters. Here, we propose for the first time, an end-to-end learning approach for online multi-target tracking. Existing deep learning methods are not designed for the above challenges and cannot be trivially applied to the task. Our solution addresses all of the above points in a principled way. Experiments on both synthetic and real data show promising results obtained at ~300 Hz on a standard CPU, and pave the way towards future research in this direction.
Anton Milan, Seyed Hamid Rezatofighi, Anthony R. Dick, Ian D. Reid 0001, Konrad Schindler
AAAI4
2017 Data-Driven Approximations to NP-Hard Problems
abstract
There exist a number of problem classes for which obtaining the exact solution becomes exponentially expensive with increasing problem size. The quadratic assignment problem (QAP) or the travelling salesman problem (TSP) are just two examples of such NP-hard problems. In practice, approximate algorithms are employed to obtain a suboptimal solution, where one must face a trade-off between computational complexity and solution quality. In this paper, we propose to learn to solve these problem from approximate examples, using recurrent neural networks (RNNs). Surprisingly, such architectures are capable of producing highly accurate solutions at minimal computational cost. Moreover, we introduce a simple, yet effective technique for improving the initial (weak) training set by incorporating the objective cost into the training procedure. We demonstrate the functionality of our approach on three exemplar applications: marginal distributions of a joint matching space, feature point matching and the travelling salesman problem. We show encouraging results on synthetic and real data in all three cases.
Anton Milan, Seyed Hamid Rezatofighi, Ravi Garg, Anthony R. Dick, Ian D. Reid 0001
AAAI5
2017 Weakly Supervised Semantic Segmentation Based on Co-segmentation
Guosheng Lin, Lingqiao Liu, Chunhua Shen, Ian D. Reid 0001
BMVC5
2017 From Motion Blur to Motion Flow: A Deep Learning Solution for Removing Heterogeneous Motion Blur
abstract
Removing pixel-wise heterogeneous motion blur is challenging due to the ill-posed nature of the problem. The predominant solution is to estimate the blur kernel by adding a prior, but extensive literature on the subject indicates the difficulty in identifying a prior which is suitably informative, and general. Rather than imposing a prior based on theory, we propose instead to learn one from the data. Learning a prior over the latent image would require modeling all possible image content. The critical observation underpinning our approach, however, is that learning the motion flow instead allows the model to focus on the cause of the blur, irrespective of the image content. This is a much easier learning task, but it also avoids the iterative process through which latent image priors are typically applied. Our approach directly estimates the motion flow from the blurred image through a fully-convolutional deep neural network (FCN) and recovers the unblurred image from the estimated motion flow. Our FCN is the first universal end-to-end mapping from the blurred image to the dense motion flow. To train the FCN, we simulate motion flows to generate synthetic blurred-image-motion-flow pairs thus avoiding the need for human labeling. Extensive experiments on challenging realistic blurred images demonstrate that the proposed method outperforms the state-of-the-art.
Dong Gong, Jie Yang 0002, Lingqiao Liu, Yanning Zhang 0001, Ian D. Reid 0001, Chunhua Shen, Anton van den Hengel, Qinfeng Shi
CVPR5
2017 RefineNet: Multi-path Refinement Networks for High-Resolution Semantic Segmentation
abstract
Recently, very deep convolutional neural networks (CNNs) have shown outstanding performance in object recognition and have also been the first choice for dense classification problems such as semantic segmentation. However, repeated subsampling operations like pooling or convolution striding in deep CNNs lead to a significant decrease in the initial image resolution. Here, we present RefineNet, a generic multi-path refinement network that explicitly exploits all the information available along the down-sampling process to enable high-resolution prediction using long-range residual connections. In this way, the deeper layers that capture high-level semantic features can be directly refined using fine-grained features from earlier convolutions. The individual components of RefineNet employ residual connections following the identity mapping mindset, which allows for effective end-to-end training. Further, we introduce chained residual pooling, which captures rich background context in an efficient manner. We carry out comprehensive experiments and set new state-of-the-art results on seven public datasets. In particular, we achieve an intersection-over-union score of 83.4 on the challenging PASCAL VOC 2012 dataset, which is the best reported result to date.
Guosheng Lin, Anton Milan, Chunhua Shen, Ian D. Reid 0001
CVPR4
2017 Attend in Groups: A Weakly-Supervised Deep Learning Framework for Learning from Web Data
abstract
Large-scale datasets have driven the rapid development of deep neural networks for visual recognition. However, annotating a massive dataset is expensive and time-consuming. Web images and their labels are, in comparison, much easier to obtain, but direct training on such automatically harvested images can lead to unsatisfactory performance, because the noisy labels of Web images adversely affect the learned recognition models. To address this drawback we propose an end-to-end weakly-supervised deep learning framework which is robust to the label noise in Web images. The proposed framework relies on two unified strategies - random grouping and attention - to effectively reduce the negative impact of noisy web image annotations. Specifically, random grouping stacks multiple images into a single training instance and thus increases the labeling accuracy at the instance level. Attention, on the other hand, suppresses the noisy signals from both incorrectly labeled images and less discriminative image regions. By conducting intensive experiments on two challenging datasets, including a newly collected fine-grained dataset with Web images of different car models,1, the superior performance of the proposed methods over competitive baselines is clearly demonstrated.
Bohan Zhuang, Lingqiao Liu, Yao Li 0003, Chunhua Shen, Ian D. Reid 0001
CVPR5
2017 Smart Mining for Deep Metric Learning
abstract
To solve deep metric learning problems and producing feature embeddings, current methodologies will commonly use a triplet model to minimise the relative distance between samples from the same class and maximise the relative distance between samples from different classes. Though successful, the training convergence of this triplet model can be compromised by the fact that the vast majority of the training samples will produce gradients with magnitudes that are close to zero. This issue has motivated the development of methods that explore the global structure of the embedding and other methods that explore hard negative/positive mining. The effectiveness of such mining methods is often associated with intractable computational requirements. In this paper, we propose a novel deep metric learning method that combines the triplet model and the global structure of the embedding space. We rely on a smart mining procedure that produces effective training samples for a low computational cost. In addition, we propose an adaptive controller that automatically adjusts the smart mining hyper-parameters and speeds up the convergence of the training process. We show empirically that our proposed method allows for fast and more accurate training of triplet ConvNets than other competing mining methods. Additionally, we show that our method achieves new state-of-the-art embedding results for CUB-200-2011 and Cars196 datasets.
Ben Harwood, Gustavo Carneiro 0001, Ian D. Reid 0001, Tom Drummond
ICCV4
2017 "Maximizing Rigidity" Revisited: A Convex Programming Approach for Generic 3D Shape Reconstruction from Multiple Perspective Views
abstract
Rigid structure-from-motion (RSfM) and non-rigid structure-from-motion (NRSfM) have long been treated in the literature as separate (different) problems. Inspired by a previous work which solved directly for 3D scene structure by factoring the relative camera poses out, we revisit the principle of “maximizing rigidity” in structure-from-motion literature, and develop a unified theory which is applicable to both rigid and non-rigid structure reconstruction in a rigidity-agnostic way. We formulate these problems as a convex semi-definite program, imposing constraints that seek to apply the principle of minimizing non-rigidity. Our results demonstrate the efficacy of the approach, with stateof- the-art accuracy on various 3D reconstruction problems.
Pan Ji, Hongdong Li, Yuchao Dai, Ian D. Reid 0001
ICCV4
2017 DeepSetNet: Predicting Sets with Deep Neural Networks
abstract
This paper addresses the task of set prediction using deep learning. This is important because the output of many computer vision tasks, including image tagging and object detection, are naturally expressed as sets of entities rather than vectors. As opposed to a vector, the size of a set is not fixed in advance, and it is invariant to the ordering of entities within it. We define a likelihood for a set distribution and learn its parameters using a deep neural network. We also derive a loss for predicting a discrete distribution corresponding to set cardinality. Set prediction is demonstrated on the problem of multi-class image classification. Moreover, we show that the proposed cardinality loss can also trivially be applied to the tasks of object counting and pedestrian detection. Our approach outperforms existing methods in all three cases on standard datasets.
Seyed Hamid Rezatofighi, Anton Milan, Ehsan Abbasnejad, Anthony R. Dick, Ian D. Reid 0001
ICCV6
2017 Towards Context-Aware Interaction Recognition for Visual Relationship Detection
abstract
Recognizing how objects interact with each other is a crucial task in visual recognition. If we define the context of the interaction to be the objects involved, then most current methods can be categorized as either: (i) training a single classifier on the combination of the interaction and its context; or (ii) aiming to recognize the interaction independently of its explicit context. Both methods suffer limitations: the former scales poorly with the number of combinations and fails to generalize to unseen combinations, while the latter often leads to poor interaction recognition performance due to the difficulty of designing a contextindependent interaction classifier.,,To mitigate those drawbacks, this paper proposes an alternative, context-aware interaction recognition framework. The key to our method is to explicitly construct an interaction classifier which combines the context, and the interaction. The context is encoded via word2vec into a semantic space, and is used to derive a classification result for the interaction. The proposed method still builds one classifier for one interaction (as per type (ii) above), but the classifier built is adaptive to context via weights which are context dependent. The benefit of using the semantic space is that it naturally leads to zero-shot generalizations in which semantically similar contexts (subject-object pairs) can be recognized as suitable contexts for an interaction, even if they were not observed in the training set. Our method also scales with the number of interaction-context pairs since our model parameters do not increase with the number of interactions. Thus our method avoids the limitation of both approaches. We demonstrate experimentally that the proposed framework leads to improved performance for all investigated interaction representations and datasets.
Bohan Zhuang, Lingqiao Liu, Chunhua Shen, Ian D. Reid 0001
ICCV4
2017 Deep learning features at scale for visual place recognition
abstract
The success of deep learning techniques in the computer vision domain has triggered a range of initial investigations into their utility for visual place recognition, all using generic features from networks that were trained for other types of recognition tasks. In this paper, we train, at large scale, two CNN architectures for the specific place recognition task and employ a multi-scale feature encoding method to generate condition- and viewpoint-invariant features. To enable this training to occur, we have developed a massive Specific PlacEs Dataset (SPED) with hundreds of examples of place appearance change at thousands of different places, as opposed to the semantic place type datasets currently available. This new dataset enables us to set up a training regime that interprets place recognition as a classification problem. We comprehensively evaluate our trained networks on several challenging benchmark place recognition datasets and demonstrate that they achieve an average 10% increase in performance over other place recognition algorithms and pre-trained CNNs. By analyzing the network responses and their differences from pre-trained networks, we provide insights into what a network learns when training for place recognition, and what these results signify for future research in this area.
Zetao Chen, Adam Jacobson, Niko Sünderhauf, Ben Upcroft, Lingqiao Liu, Chunhua Shen, Ian D. Reid 0001, Michael Milford
ICRA7
2017 A branch-and-bound algorithm for checkerboard extraction in camera-laser calibration
abstract
We address the problem of camera-to-laserscanner calibration using a checkerboard and multiple imagelaser scan pairs. Distinguishing which laser points measure the checkerboard and which lie on the background is essential to any such system. We formulate the checkerboard extraction as a combinatorial optimization problem with a clear cut objective function. We propose a branch-and-bound technique that deterministically and globally optimizes the objective. Unlike what is available in the literature, the proposed method is not heuristic and does not require assumptions such as constraints on the background or relying on discontinuity of the range measurements to partition the data into line segments. The proposed approach is generic and can be applied to both 3D or 2D laser scanners as well as the cases where multiple checkerboards are present. We demonstrate the effectiveness of the proposed approach by providing numerical simulations as well as experimental results.
Alireza Khosravian, Tat-Jun Chin, Ian D. Reid 0001
ICRA3
2017 A discrete-time attitude observer on SO(3) for vision and GPS fusion
abstract
This paper proposes a discrete-time geometric attitude observer for fusing monocular vision with GPS velocity measurements. The observer takes the relative transformations obtained from processing monocular images with any visual odometry algorithm and fuses them with GPS velocity measurements. The objectives of this sensor fusion are twofold; first to mitigate the inherent drift of the attitude estimates of the visual odometry, and second, to estimate the orientation directly with respect to the North-East-Down frame. A key contribution of the paper is to present a rigorous stability analysis showing that the attitude estimates of the observer converge exponentially to the true attitude and to provide a lower bound for the convergence rate of the observer. Through experimental studies, we demonstrate that the observer effectively compensates for the inherent drift of the pure monocular vision based attitude estimation and is able to recover the North-East-Down orientation even if it is initialized with a very large attitude error.
Alireza Khosravian, Tat-Jun Chin, Ian D. Reid 0001, Robert E. Mahony
ICRA3
2017 RRD-SLAM: Radial-distorted rolling-shutter direct SLAM
abstract
In this paper, we present a monocular direct semi-dense SLAM (Simultaneous Localization And Mapping) method that can handle both radial distortion and rolling-shutter distortion. Such distortions are common in, but not restricted to, situations when an inexpensive wide-angle lens and a CMOS sensor are used, and leads to significant inaccuracy in the map and trajectory estimates if not modeled correctly. The apparent naive solution of simply undistorting the images using pre-calibrated parameters does not apply to this case since rows in the undistorted image are no longer captured at the same time. To address this we develop an algorithm that incorporates radial distortion into an existing state-of-the-art direct semi-dense SLAM system that takes rolling-shutters into account. We propose a method for finding the generalized epipolar curve for each rolling-shutter radially distorted image. Our experiments demonstrate the efficacy of our approach and compare it favorably with the state-of-the-art in direct semi-dense rolling-shutter SLAM.
Jae-Hak Kim, Yasir Latif, Ian D. Reid 0001
ICRA3
2017 Dense monocular reconstruction using surface normals
abstract
This paper presents an efficient framework for dense 3D scene reconstruction using input from a moving monocular camera. Visual SLAM (Simultaneous Localisation and Mapping) approaches based solely on geometric methods have proven to be quite capable of accurately tracking the pose of a moving camera and simultaneously building a map of the environment in real-time. However, most of them suffer from the 3D map being too sparse for practical use. The missing points in the generated map correspond mainly to areas lacking texture in the input images, and dense mapping systems often rely on hand-crafted priors like piecewise-planarity or piecewise-smooth depth. These priors do not always provide the required level of scene understanding to accurately fill the map. On the other hand, Convolutional Neural Networks (CNNs) have had great success in extracting high-level information from images and regressing pixel-wise surface normals, semantics, and even depth. In this work we leverage this high-level scene context learned by a deep CNN in the form of a surface normal prior. We show, in particular, that using the surface normal prior leads to better reconstructions than the weaker smoothness prior.
Chamara Saroj Weerasekera, Yasir Latif, Ravi Garg, Ian D. Reid 0001
ICRA4
2017 Learning Multi-level Region Consistency with Dense Multi-label Networks for Semantic Segmentation
abstract
Semantic image segmentation is a fundamental task in image understanding. Per-pixel semantic labelling of an image benefits greatly from the ability to consider region consistency both locally and globally. However, many Fully Convolutional Network based methods do not impose such consistency, which may give rise to noisy and implausible predictions. We address this issue by proposing a dense multi-label network module that is able to encourage the region consistency at different levels. This simple but effective module can be easily integrated into any semantic segmentation systems. With comprehensive experiments, we show that the dense multi-label can successfully remove the implausible labels and clear the confusion so as to boost the performance of semantic segmentation systems.
Guosheng Lin, Chunhua Shen, Ian D. Reid 0001
IJCAI4
2017 Deep learning for 2D scan matching and loop closure
abstract
Although 2D LiDAR based Simultaneous Localization and Mapping (SLAM) is a relatively mature topic nowadays, the loop closure problem remains challenging due to the lack of distinctive features in 2D LiDAR range scans. Existing research can be roughly divided into correlation based approaches e.g. scan-to-submap matching and feature based methods e.g. bag-of-words (BoW). In this paper, we solve loop closure detection and relative pose transformation using 2D LiDAR within an end-to-end Deep Learning framework. The algorithm is verified with simulation data and on an Unmanned Aerial Vehicle (UAV) flying in indoor environment. The loop detection ConvNet alone achieves an accuracy of 98.2% in loop closure detection. With a verification step using the scan matching ConvNet, the false positive rate drops to around 0.001%. The proposed approach processes 6000 pairs of raw LiDAR scans per second on a Nvidia GTX1080 GPU.
Huangying Zhan, Ben M. Chen, Ian D. Reid 0001, Gim Hee Lee
IROS4
2017 Meaningful maps with object-oriented semantic mapping
abstract
For intelligent robots to interact in meaningful ways with their environment, they must understand both the geometric and semantic properties of the scene surrounding them. The majority of research to date has addressed these mapping challenges separately, focusing on either geometric or semantic mapping. In this paper we address the problem of building environmental maps that include both semantically meaningful, object-level entities and point- or mesh-based geometrical representations. We simultaneously build geometric point cloud models of previously unseen instances of known object classes and create a map that contains these object models as central entities. Our system leverages sparse, feature-based RGB-D SLAM, image-based deep-learning object detection and 3D unsupervised segmentation.
Niko Sünderhauf, Trung T. Pham, Yasir Latif, Michael Milford, Ian D. Reid 0001
IROS5
2017 Deep Reinforcement Learning for Active Breast Lesion Detection from DCE-MRI
Gabriel Maicas, Gustavo Carneiro 0001, Andrew P. Bradley, Jacinto C. Nascimento, Ian D. Reid 0001
MICCAI (3)5
2017 Deep Subspace Clustering Networks
abstract
We present a novel deep neural network architecture for unsupervised subspace clustering. This architecture is built upon deep auto-encoders, which non-linearly map the input data into a latent space. Our key idea is to introduce a novel self-expressive layer between the encoder and the decoder to mimic the "self-expressiveness" property that has proven effective in traditional subspace clustering. Being differentiable, our new self-expressive layer provides a simple but effective way to learn pairwise affinities between all data points through a standard back-propagation procedure. Being nonlinear, our neural-network based method is able to cluster data points having complex (often nonlinear) structures. We further propose pre-training and fine-tuning strategies that let us effectively learn the parameters of our subspace clustering networks. Our experiments show that the proposed method significantly outperforms the state-of-the-art unsupervised subspace clustering methods.
Pan Ji, Tong Zhang 0023, Hongdong Li, Mathieu Salzmann, Ian D. Reid 0001
NIPS5
2017 A Bayesian Data Augmentation Approach for Learning Deep Models
abstract
Data augmentation is an essential part of the training process applied to deep learning models. The motivation is that a robust training process for deep learning models depends on large annotated datasets, which are expensive to be acquired, stored and processed. Therefore a reasonable alternative is to be able to automatically generate new annotated training samples using a process known as data augmentation. The dominant data augmentation approach in the field assumes that new training samples can be obtained via random geometric or appearance transformations applied to annotated training samples, but this is a strong assumption because it is unclear if this is a reliable generative model for producing new training samples. In this paper, we provide a novel Bayesian formulation to data augmentation, where new annotated training points are treated as missing variables and generated based on the distribution learned from the training set. For learning, we introduce a theoretically sound algorithm --- generalised Monte Carlo expectation maximisation, and demonstrate one possible implementation via an extension of the Generative Adversarial Network (GAN). Classification results on MNIST, CIFAR-10 and CIFAR-100 show the better performance of our proposed method compared to the current dominant data augmentation approach mentioned above --- the results also show that our approach produces better classification results than similar GAN models.
Toan Tran 0002, Trung Pham, Gustavo Carneiro 0001, Lyle John Palmer, Ian D. Reid 0001
NIPS5
2017 Real-Time Tracking of Single and Multiple Objects from Depth-Colour Imagery Using 3D Signed Distance Functions
abstract
We describe a novel probabilistic framework for real-time tracking of multiple objects from combined depth-colour imagery. Object shape is represented implicitly using 3D signed distance functions. Probabilistic generative models based on these functions are developed to account for the observed RGB-D imagery, and tracking is posed as a maximum a posteriori problem. We present first a method suited to tracking a single rigid 3D object, and then generalise this to multiple objects by combining distance functions into a shape union in the frame of the camera. This second model accounts for similarity and proximity between objects, and leads to robust real-time tracking without recourse to bolt-on or ad-hoc collision detection.
Carl Yuheng Ren, Victor Adrian Prisacariu, Olaf Kähler, Ian D. Reid 0001, David William Murray 0001
Int. J. Comput. Vis.4
2016 Learning Local Image Descriptors with Deep Siamese and Triplet Convolutional Networks by Minimizing Global Loss Functions
abstract
Recent innovations in training deep convolutional neural network (ConvNet) models have motivated the design of new methods to automatically learn local image descriptors. The latest deep ConvNets proposed for this task consist of a siamese network that is trained by penalising misclassification of pairs of local image patches. Current results from machine learning show that replacing this siamese by a triplet network can improve the classification accuracy in several problems, but this has yet to be demonstrated for local image descriptor learning. Moreover, current siamese and triplet networks have been trained with stochastic gradient descent that computes the gradient from individual pairs or triplets of local image patches, which can make them prone to overfitting. In this paper, we first propose the use of triplet networks for the problem of local image descriptor learning. Furthermore, we also propose the use of a global loss that minimises the overall classification error in the training set, which can improve the generalisation capability of the model. Using the UBC benchmark dataset for comparing local image descriptors, we show that the triplet network produces a more accurate embedding than the siamese network in terms of the UBC dataset errors. Moreover, we also demonstrate that a combination of the triplet and global losses produces the best embedding in the field, using this triplet network. Finally, we also show that the use of the central-surround siamese network trained with the global loss produces the best result of the field on the UBC dataset.
Gustavo Carneiro 0001, Ian D. Reid 0001
CVPR3
2016 Efficient Piecewise Training of Deep Structured Models for Semantic Segmentation
abstract
Recent advances in semantic image segmentation have mostly been achieved by training deep convolutional neural networks (CNNs). We show how to improve semantic segmentation through the use of contextual information, specifically, we explore 'patch-patch' context between image regions, and 'patch-background' context. For learning from the patch-patch context, we formulate Conditional Random Fields (CRFs) with CNN-based pairwise potential functions to capture semantic correlations between neighboring patches. Efficient piecewise training of the proposed deep structured model is then applied to avoid repeated expensive CRF inference for back propagation. For capturing the patch-background context, we show that a network design with traditional multi-scale image input and sliding pyramid pooling is effective for improving performance. Our experimental results set new state-of-the-art performance on a number of popular semantic segmentation datasets, including NYUDv2, PASCAL VOC 2012, PASCAL-Context, and SIFT-flow. In particular, we achieve an intersection-overunion score of 78:0 on the challenging PASCAL VOC 2012 dataset.
Guosheng Lin, Chunhua Shen, Anton van den Hengel, Ian D. Reid 0001
CVPR4
2016 Efficient Point Process Inference for Large-Scale Object Detection
abstract
We tackle the problem of large-scale object detection in images, where the number of objects can be arbitrarily large, and can exhibit significant overlap/occlusion. A successful approach to modelling the large-scale nature of this problem has been via point process density functions which jointly encode object qualities and spatial interactions. But the corresponding optimisation problem is typically difficult or intractable, and many of the best current methods rely on Monte Carlo Markov Chain (MCMC) simulation, which converges slowly in a large solution space. We propose an efficient point process inference for largescale object detection using discrete energy minimization. In particular, we approximate the solution space by a finite set of object proposals and cast the point process density function to a corresponding energy function of binary variables whose values indicate which object proposals are accepted. We resort to the local submodular approximation (LSA) based trust-region optimisation to find the optimal solution. Furthermore we analyse the error of LSA approximation, and show how to adjust the point process energy to dramatically speed up the convergence without harming the optimality. We demonstrate the superior efficiency and accuracy of our method using a variety of large-scale object detection applications such as crowd human detection, birds, cells counting/localization.
Trung T. Pham, Seyed Hamid Rezatofighi, Ian D. Reid 0001, Tat-Jun Chin
CVPR3
2016 Joint Probabilistic Matching Using m-Best Solutions
abstract
Matching between two sets of objects is typically approached by finding the object pairs that collectively maximize the joint matching score. In this paper, we argue that this single solution does not necessarily lead to the optimal matching accuracy and that general one-to-one assignment problems can be improved by considering multiple hypotheses before computing the final similarity measure. To that end, we propose to utilize the marginal distributionsfor each entity. Previously, this idea has been neglected mainly because exact marginalization is intractable due to a combinatorial number of all possible matching permutations. Here, we propose a generic approach to efficiently approximate the marginal distributions by exploiting the m-best solutions of the original problem. This approach not only improves the matching solution, but also provides more accurate ranking of the results, because of the extra information included in the marginal distribution. We validate our claim on two distinct objectives: (i) person re-identification and temporal matching modeled as an integer linear program, and (ii) feature point matching using a quadratic cost function. Our experiments confirm that marginalization indeed leads to superior performance compared to the single (nearly) optimal solution, yielding state-of-the-art results in both applications on standard benchmarks.
Seyed Hamid Rezatofighi, Anton Milan, Zhen Zhang 0008, Qinfeng Shi, Anthony R. Dick, Ian D. Reid 0001
CVPR6
2016 Fast Training of Triplet-Based Deep Binary Embedding Networks
abstract
In this paper, we aim to learn a mapping (or embedding) from images to a compact binary space in which Hamming distances correspond to a ranking measure for the image retrieval task. We make use of a triplet loss because this has been shown to be most effective for ranking problems. However, training in previous works can be prohibitively expensive due to the fact that optimization is directly performed on the triplet space, where the number of possible triplets for training is cubic in the number of training examples. To address this issue, we propose to formulate high-order binary codes learning as a multi-label classification problem by explicitly separating learning into two interleaved stages. To solve the first stage, we design a large-scale high-order binary codes inference algorithm to reduce the high-order objective to a standard binary quadratic problem such that graph cuts can be used to efficiently infer the binary codes which serve as the labels of each training datum. In the second stage we propose to map the original image to compact binary codes via carefully designed deep convolutional neural networks (CNNs) and the hashing function fitting can be solved by training binary CNN classifiers. An incremental/interleaved optimization strategy is proffered to ensure that these two steps are interactive with each other during training for better accuracy. We conduct experiments on several benchmark datasets, which demonstrate both improved training time (by as much as two orders of magnitude) as well as producing state-of-the-art hashing for various retrieval tasks.
Bohan Zhuang, Guosheng Lin, Chunhua Shen, Ian D. Reid 0001
CVPR4
2016 Unsupervised CNN for Single View Depth Estimation: Geometry to the Rescue
Ravi Garg, Gustavo Carneiro 0001, Ian D. Reid 0001
ECCV (8)4
2016 Direct semi-dense SLAM for rolling shutter cameras
abstract
In this paper, we present a monocular Direct and Semi-dense SLAM (Simultaneous Localization And Mapping) system for rolling shutter cameras. In a rolling shutter camera, the pose is different for each row of each image, and this yields poor pose estimates and poor structure estimates when using a state-of-the-art semi-dense direct method designed for global shutter cameras. To address this issue in tracking, we model the smooth and continuous camera trajectory using a B-spline curve of degree k??1 for poses in the Lie algebra, se(3).We solve for the camera poses at each row-time by a direct optimisation of photometric error as a function of the control points of the spline. Likewise for mapping, we develop generalised epipolar geometry for the rolling shutter case and solve for point depths using photometric error. Although each of these issues has been previously tackled, to the best of our knowledge ours is the first full solution to monocular, direct (feature-less) SLAM. We benchmark our method for pose accuracy and map accuracy against the state-of-the-art semi-dense SLAM system, LSD-SLAM, demonstrating the improved efficacy of our approach when using rolling shutter cameras via synthetic sequences with known ground-truth and real sequences.
Jae-Hak Kim, Cesar Dario Cadena Lerma, Ian D. Reid 0001
ICRA3
2016 Measuring the performance of single image depth estimation methods
abstract
We consider the question of benchmarking the performance of methods used for estimating the depth of a scene from a single image. We describe various measures that have been used in the past, discuss their limitations and demonstrate that each is deficient in one or more ways. We propose a new measure of performance for depth estimation that overcomes these deficiencies, and has a number of desirable properties. We show that in various cases of interest the new measure enables visualisation of the performance of a method that is otherwise obfuscated by existing metrics. Our proposed method is capable of illuminating the relative performance of different algorithms on different kinds of data, such as the difference in efficacy of a method when estimating the depth of the ground plane versus estimating the depth of other generic scene structure. We showcase the method by comparing a number of existing single-view methods against each other and against more traditional depth estimation methods such as binocular stereo.
Cesar Dario Cadena Lerma, Yasir Latif, Ian D. Reid 0001
IROS3
2016 Geometrically consistent plane extraction for dense indoor 3D maps segmentation
abstract
Modern SLAM systems with a depth sensor are able to reliably reconstruct dense 3D geometric maps of indoor scenes. Representing these maps in terms of meaningful entities is a step towards building semantic maps for autonomous robots. One approach is to segment the 3D maps into semantic objects using Conditional Random Fields (CRF), which requires large 3D ground truth datasets to train the classification model. Additionally, the CRF inference is often computationally expensive. In this paper, we present an unsupervised geometric-based approach for the segmentation of 3D point clouds into objects and meaningful scene structures. We approximate an input point cloud by an adjacency graph over surface patches, whose edges are then classified as being either on or off. We devise an effective classifier which utilises both global planar surfaces and local surface convexities for edge classification. More importantly, we propose a novel global plane extraction algorithm for robustly discovering the underlying planes in the scene. Our algorithm is able to enforce the extracted planes to be mutually orthogonal or parallel which conforms usually with human-made indoor environments. We reconstruct 654 3D indoor scenes from NYUv2 sequences to validate the efficiency and effectiveness of our segmentation method.
Trung T. Pham, Markus Eich, Ian D. Reid 0001, Gordon F. Wyeth
IROS3
2016 12th Asian conference on computer vision
Ian D. Reid 0001
Comput. Vis. Image Underst.1
2016 Online unsupervised feature learning for visual tracking
Fayao Liu, Chunhua Shen, Ian D. Reid 0001, Anton van den Hengel
Image Vis. Comput.3
2016 Learning Depth from Single Monocular Images Using Deep Convolutional Neural Fields
abstract
In this article, we tackle the problem of depth estimation from single monocular images. Compared with depth estimation using multiple images such as stereo depth perception, depth from monocular images is much more challenging. Prior work typically focuses on exploiting geometric priors or additional sources of information, most using hand-crafted features. Recently, there is mounting evidence that features from deep convolutional neural networks (CNN) set new records for various vision applications. On the other hand, considering the continuous characteristic of the depth values, depth estimation can be naturally formulated as a continuous conditional random field (CRF) learning problem. Therefore, here we present a deep convolutional neural field model for estimating depths from single monocular images, aiming to jointly explore the capacity of deep CNN and continuous CRF. In particular, we propose a deep structured learning scheme which learns the unary and pairwise potentials of continuous CRF in a unified deep CNN framework. We then further propose an equally effective model based on fully convolutional networks and a novel superpixel pooling method, which is about 10 times faster, to speedup the patch-wise convolutions in the deep model. With this more efficient model, we are able to design deeper networks to pursue better performance. Our proposed method can be used for depth estimation of general scenes with no geometric priors nor any extra information injected. In our case, the integral of the partition function can be calculated in a closed form such that we can exactly solve the log-likelihood maximization. Moreover, solving the inference problem for predicting depths of a test image is highly efficient as closed-form solutions exist. Experiments on both indoor and outdoor scene datasets demonstrate that the proposed method outperforms state-of-the-art depth estimation approaches.
Fayao Liu, Chunhua Shen, Guosheng Lin, Ian D. Reid 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Past, Present, and Future of Simultaneous Localization and Mapping: Toward the Robust-Perception Age
abstract
Simultaneous localization and mapping (SLAM) consists in the concurrent construction of a model of the environment (the map), and the estimation of the state of the robot moving within it. The SLAM community has made astonishing progress over the last 30 years, enabling large-scale real-world applications and witnessing a steady transition of this technology to industry. We survey the current state of SLAM and consider future directions. We start by presenting what is now the de-facto standard formulation for SLAM. We then review related work, covering a broad set of topics including robustness and scalability in long-term mapping, metric and semantic representations for mapping, theoretical performance guarantees, active SLAM and exploration, and other new frontiers. This paper simultaneously serves as a position paper and tutorial to those who are users of SLAM. By looking at the published research with a critical eye, we delineate open challenges and new research issues, that still deserve careful scientific investigation. The paper also contains the authors' take on two questions that often animate discussions during robotics conferences: Do robots need SLAM? and Is SLAM solved?
Cesar Dario Cadena Lerma, Luca Carlone, Henry Carrillo, Yasir Latif, Davide Scaramuzza 0001, José Neira, Ian D. Reid 0001, John J. Leonard
IEEE Trans. Robotics7
2015 The k-support norm and convex envelopes of cardinality and rank
abstract
Sparsity, or cardinality, as a tool for feature selection is extremely common in a vast number of current computer vision applications. The k-support norm is a recently proposed norm with the proven property of providing the tightest convex bound on cardinality over the Euclidean norm unit ball. In this paper we present a re-derivation of this norm, with the hope of shedding further light on this particular surrogate function. In addition, we also present a connection between the rank operator, the nuclear norm and the k-support norm. Finally, based on the results established in this re-derivation, we propose a novel algorithm with significantly improved computational efficiency, empirically validated on a number of different problems, using both synthetic and real world data.
Anders P. Eriksson, Trung-Thanh Pham, Tat-Jun Chin, Ian D. Reid 0001
CVPR4
2015 Joint tracking and segmentation of multiple targets
abstract
Tracking-by-detection has proven to be the most successful strategy to address the task of tracking multiple targets in unconstrained scenarios [e.g. 40, 53, 55]. Traditionally, a set of sparse detections, generated in a preprocessing step, serves as input to a high-level tracker whose goal is to correctly associate these “dots” over time. An obvious short-coming of this approach is that most information available in image sequences is simply ignored by thresholding weak detection responses and applying non-maximum suppression. We propose a multi-target tracker that exploits low level image information and associates every (super)-pixel to a specific target or classifies it as background. As a result, we obtain a video segmentation in addition to the classical bounding-box representation in unconstrained, real-world videos. Our method shows encouraging results on many standard benchmark sequences and significantly outperforms state-of-the-art tracking-by-detection approaches in crowded scenes with long-term partial occlusions.
Anton Milan, Laura Leal-Taixé, Konrad Schindler, Ian D. Reid 0001
CVPR4
2015 Hierarchical Higher-Order Regression Forest Fields: An Application to 3D Indoor Scene Labelling
abstract
This paper addresses the problem of semantic segmentation of 3D indoor scenes reconstructed from RGB-D images. Traditionally label prediction for 3D points is tackled by employing graphical models that capture scene features and complex relations between different class labels. However, the existing work is restricted to pairwise conditional random fields, which are insufficient when encoding rich scene context. In this work we propose models with higher-order potentials to describe complex relational information from the 3D scenes. Specifically, we relax the labelling problem to a regression, and generalize the higher-order associative P n Potts model to a new family of arbitrary higher-order models based on regression forests. We show that these models, like the robust P n models, can still be decomposed into the sum of pairwise terms by introducing auxiliary variables. Moreover, our proposed higher-order models also permit extension to hierarchical random fields, which allows for the integration of scene context and features computed at different scales. Our potential functions are constructed based on regression forests encoding Gaussian densities that admit efficient inference. The parameters of our model are learned from training data using a structured learning approach. Results on two datasets show clear improvements over current state-of-the-art methods.
Trung-Thanh Pham, Ian D. Reid 0001, Yasir Latif, Stephen Gould
ICCV2
2015 Joint Probabilistic Data Association Revisited
abstract
In this paper, we revisit the joint probabilistic data association (JPDA) technique and propose a novel solution based on recent developments in finding the m-best solutions to an integer linear program. The key advantage of this approach is that it makes JPDA computationally tractable in applications with high target and/or clutter density, such as spot tracking in fluorescence microscopy sequences and pedestrian tracking in surveillance footage. We also show that our JPDA algorithm embedded in a simple tracking framework is surprisingly competitive with state-of-the-art global tracking methods in these two applications, while needing considerably less processing time.
Seyed Hamid Rezatofighi, Anton Milan, Zhen Zhang 0008, Qinfeng Shi, Anthony R. Dick, Ian D. Reid 0001
ICCV6
2015 A fast, modular scene understanding system using context-aware object detection
abstract
We propose a semantic scene understanding system that is suitable for real robotic operations. The system solves different tasks (semantic segmentation and object detections) in an opportunistic and distributed fashion but still allows communication between modules to improve their respective performances. We propose the use of the semantic space to improve specific out-of-the-box object detectors and an update model to take the evidence from different detection into account in the semantic segmentation process. Our proposal is evaluated with the KITTI dataset, on the object detection benchmark and on five different sequences manually annotated for the semantic segmentation task, demonstrating the efficacy of our approach.
Cesar Dario Cadena Lerma, Anthony R. Dick, Ian D. Reid 0001
ICRA3
2015 Deeply Learning the Messages in Message Passing Inference
abstract
Deep structured output learning shows great promise in tasks like semantic image segmentation. We proffer a new, efficient deep structured model learning scheme, in which we show how deep Convolutional Neural Networks (CNNs) can be used to directly estimate the messages in message passing inference for structured prediction with Conditional Random Fields CRFs). With such CNN message estimators, we obviate the need to learn or evaluate potential functions for message calculation. This confers significant efficiency for learning, since otherwise when performing structured learning for a CRF with CNN potentials it is necessary to undertake expensive inference for every stochastic gradient iteration. The network output dimension of message estimators is the same as the number of classes, rather than exponentially growing in the order of the potentials. Hence it is more scalable for cases that a large number of classes are involved. We apply our method to semantic image segmentation and achieve impressive performance, which demonstrates the effectiveness and usefulness of our CNN message learning method.
Guosheng Lin, Chunhua Shen, Ian D. Reid 0001, Anton van den Hengel
NIPS3
2015 Real-Time 3D Tracking and Reconstruction on Mobile Phones
abstract
We present a novel framework for jointly tracking a camera in 3D and reconstructing the 3D model of an observed object. Due to the region based approach, our formulation can handle untextured objects, partial occlusions, motion blur, dynamic backgrounds and imperfect lighting. Our formulation also allows for a very efficient implementation which achieves real-time performance on a mobile phone, by running the pose estimation and the shape optimisation in parallel. We use a level set based pose estimation but completely avoid the, typically required, explicit computation of a global distance. This leads to tracking rates of more than 100 Hz on a desktop PC and 30 Hz on a mobile phone. Further, we incorporate additional orientation information from the phone's inertial sensor which helps us resolve the tracking ambiguities inherent to region based formulations. The reconstruction step first probabilistically integrates 2D image statistics from selected keyframes into a 3D volume, and then imposes coherency and compactness using a total variational regularisation term. The global optimum of the overall energy function is found using a continuous max-flow algorithm and we show that, similar to tracking, the integration of per voxel posteriors instead of likelihoods improves the precision and accuracy of the reconstruction.
Victor Adrian Prisacariu, Olaf Kähler, David William Murray 0001, Ian D. Reid 0001
IEEE Trans. Vis. Comput. Graph.4
2014 3D Tracking of Multiple Objects with Identical Appearance Using RGB-D Input
abstract
Most current approaches for 3D object tracking rely on distinctive object appearances. While several such trackers can be instantiated to track multiple objects independently, this not only neglects that objects should not occupy the same space in 3D, but also fails when objects have highly similar or identical appearances. In this paper we develop a probabilistic graphical model that accounts for similarity and proximity and leads to robust real-time tracking of multiple objects from RGB-D data, without recourse to bolton collision detection.
Carl Yuheng Ren, Victor Adrian Prisacariu, Olaf Kähler, Ian D. Reid 0001, David William Murray 0001
3DV4
2014 Towards semantic visual SLAM
abstract
Summary form only given. Visual Simultaneous Localisation and Mapping is the process whereby a camera builds a map of a previously unseen environment, and localises itself with respect to that environment, often in real-time. Although there has been remarkable progress, and it is now possible, for example, to build dense maps in real-time using high-end commodity hardware, most SLAM research has remained rooted in geometry. Geometric representations are limited though, in that they rarely encode higher-level information to describe the scene, are wasteful of storage, and brittle to changes. I am therefore interested extending SLAM beyond geometry to more semantically meaningful representations in which a scene can be segmented into components, and compactly described and represented. In work towards that end, in this talk I will describe work i my group from the last few years that progresses to this end:. First, I will describe a system that combines detection, segmentation and tracking of instances of a known 3D object class with a system for real-time dense visual mapping. We learn a low dimensional space of shapes to encode prior shape knowledge, and then perform simultaneous segmentation and tracking of an instance of a 3D shape class by optimising for shape and pose. This tracking methodology is then incorporated into a system for real-time dense visual mapping of a scene demonstrating that prior knowledge of objects can be incorporated into a SLAM map to improve map fidelity. Second I will present work that shows how structured learning methods can be used within SLAM scene understanding methods. I will discuss a method for reconstructing building interiors using a combination of point-based parallel tracking and mapping, with single-view reconstruction techniques. Here, by leveraging prior knowledge of shape (namely that the building interior conforms to a "Manhattan" model), we develop an efficient global optimisation method for inferring key semantic properties of the scene, namely its boundaries: the floor, ceiling and walls. This method makes use of single-view pixel-level texture cues, as well as opportunistic use of 3D information such as photo-consistency, and sparse 3d map data. Finally I will discuss progress in using structured learning methods for interpreting RGB-D data using Decision Tree Fields.
Ian D. Reid 0001
ICARCV1
2014 Hybrid Inference Optimization for robust pose graph estimation
abstract
In this paper we introduce a new optimization algorithm for networks of switched nonlinear objectives and apply this to the important problem of pose graph estimation for robot localization and mapping. The key insight is to replace the linear solver typically used in Gauss-Newton style methods with hybrid inference over switched discrete/continuous linear Gaussian networks. Since exact inference in these networks is known to be NP-hard, we also propose an approximate inference algorithm for the linearized hybrid networks based on message passing. We apply the new algorithm to the problem of robust pose graph estimation in the presence of incorrect loop closures and compare against three recently published approaches to the same problem. Evaluation is performed on ten sequences from two different datasets and shows that our approach performs substantially better than the state of the art.
Aleksandr V. Segal, Ian D. Reid 0001
IROS2
2014 Regressing Local to Global Shape Properties for Online Segmentation and Tracking
Carl Yuheng Ren, Victor Adrian Prisacariu, Ian D. Reid 0001
Int. J. Comput. Vis.3
2013 Dense Reconstruction Using 3D Object Shape Priors
abstract
We propose a formulation of monocular SLAM which combines live dense reconstruction with shape priors-based 3D tracking and reconstruction. Current live dense SLAM approaches are limited to the reconstruction of visible surfaces. Moreover, most of them are based on the minimisation of a photo-consistency error, which usually makes them sensitive to specularities. In the 3D pose recovery literature, problems caused by imperfect and ambiguous image information have been dealt with by using prior shape knowledge. At the same time, the success of depth sensors has shown that combining joint image and depth information drastically increases the robustness of the classical monocular 3D tracking and 3D reconstruction approaches. In this work we link dense SLAM to 3D object pose and shape recovery. More specifically, we automatically augment our SLAM system with object specific identity, together with 6D pose and additional shape degrees of freedom for the object(s) of known class in the scene, combining image data and depth information for the pose and shape recovery. This leads to a system that allows for full scaled 3D reconstruction with the known object(s) segmented from the scene. The segmentation enhances the clarity, accuracy and completeness of the maps built by the dense SLAM system, while the dense 3D data aids the segmentation process, yielding faster and more reliable convergence than when using 2D image data alone.
Amaury Dame, Victor Adrian Prisacariu, Carl Yuheng Ren, Ian D. Reid 0001
CVPR4
2013 Efficient 3D Scene Labeling Using Fields of Trees
abstract
We address the problem of 3D scene labeling in a structured learning framework. Unlike previous work which uses structured Support Vector Machines, we employ the recently described Decision Tree Field and Regression Tree Field frameworks, which learn the unary and binary terms of a Conditional Random Field from training data. We show this has significant advantages in terms of inference speed, while maintaining similar accuracy. We also demonstrate empirically the importance for overall labeling accuracy of features that make use of prior knowledge about the coarse scene layout such as the location of the ground plane. We show how this coarse layout can be estimated by our framework automatically, and that this information can be used to bootstrap improved accuracy in the detailed labeling.
Olaf Kähler, Ian D. Reid 0001
ICCV2
2013 STAR3D: Simultaneous Tracking and Reconstruction of 3D Objects Using RGB-D Data
abstract
We introduce a probabilistic framework for simultaneous tracking and reconstruction of 3D rigid objects using an RGB-D camera. The tracking problem is handled using a bag-of-pixels representation and a back-projection scheme. Surface and background appearance models are learned online, leading to robust tracking in the presence of heavy occlusion and outliers. In both our tracking and reconstruction modules, the 3D object is implicitly embedded using a 3D level-set function. The framework is initialized with a simple shape primitive model (e.g. a sphere or a cube), and the real 3D object shape is tracked and reconstructed online. Unlike existing depth-based 3D reconstruction works, which either rely on calibrated/fixed camera set up or use the observed world map to track the depth camera, our framework can simultaneously track and reconstruct small moving objects. We use both qualitative and quantitative results to demonstrate the superior performance of both tracking and reconstruction of our method.
Carl Yuheng Ren, Victor Adrian Prisacariu, David William Murray 0001, Ian D. Reid 0001
ICCV4
2013 Latent Data Association: Bayesian Model Selection for Multi-target Tracking
abstract
We propose a novel parametrization of the data association problem for multi-target tracking. In our formulation, the number of targets is implicitly inferred together with the data association, effectively solving data association and model selection as a single inference problem. The novel formulation allows us to interpret data association and tracking as a single Switching Linear Dynamical System (SLDS). We compute an approximate posterior solution to this problem using a dynamic programming/message passing technique. This inference-based approach allows us to incorporate richer probabilistic models into the tracking system. In particular, we incorporate inference over inliers/outliers and track termination times into the system. We evaluate our approach on publicly available datasets and demonstrate results competitive with, and in some cases exceeding the state of the art.
Aleksandr V. Segal, Ian D. Reid 0001
ICCV2
2013 Simultaneous 3D tracking and reconstruction on a mobile phone
abstract
A novel framework for joint monocular 3D tracking and reconstruction is described that can handle untextured objects, occlusions, motion blur, changing background and imperfect lighting, and that can run at frame rate on a mobile phone. The method runs in parallel (i) level set based pose estimation and (ii) continuous max flow based shape optimisation. By avoiding a global computation of distance transforms typically used in level set methods, tracking rates here exceed 100Hz and 20Hz on a desktop and mobile phone, respectively, without needing a GPU. Tracking ambiguities are reduced by augmenting orientation information from the phone's inertial sensor. Reconstruction involves probabilistic integration of the 2D image statistics from keyframes into a 3D volume. Per-voxel posteriors are used instead of the standard likelihoods, giving increased accuracy and robustness. Shape coherency and compactness is then imposed using a total variational approach solved using globally optimal continuous max flow.
Victor Adrian Prisacariu, Olaf Kähler, David William Murray 0001, Ian D. Reid 0001
ISMAR4
2012 Simultaneous Monocular 2D Segmentation, 3D Pose Recovery and 3D Reconstruction
Victor Adrian Prisacariu, Aleksandr V. Segal, Ian D. Reid 0001
ACCV (1)3
2012 On the comparison of uncertainty criteria for active SLAM
abstract
In this paper, we consider the computation of the D-optimality criterion as a metric for the uncertainty of a SLAM system. Properties regarding the use of this uncertainty criterion in the active SLAM context are highlighted, and comparisons against the A-optimality criterion and entropy are presented. This paper shows that contrary to what has been previously reported, the D-optimality criterion is indeed capable of giving fruitful information as a metric for the uncertainty of a robot performing SLAM. Finally, through various experiments with simulated and real robots, we support our claims and show that the use of D-opt has desirable effects in various SLAM related tasks such as active mapping and exploration.
Henry Carrillo, Ian D. Reid 0001, José A. Castellanos 0001
ICRA2
2012 Cognitive active vision for human identification
abstract
We describe an integrated, real-time multi-camera surveillance system that is able to find and track individuals, acquire and archive facial image sequences, and perform face recognition. The system is based around an inference engine that can extract high-level information from an observed scene, and generate appropriate commands for a set of pan-tilt-zoom (PTZ) cameras. The incorporation of a reliable facial recognition into the high-level feedback is a main novelty of our work, showing how high-level understanding of a scene can be used to deploy PTZ sensing resources effectively. The system comprises a distributed camera system using SQL tables as virtual communication channels, Situation Graph Trees for knowledge representation, inference and high-level camera control, and a variety of visual processing algorithms including an on-line acquisition of facial images, and on-line recognition of faces by comparing image sets using subspace distance. We provide an extensive evaluation of this method using our system for both acquisition of training data, and later recognition. A set of experiments in a surveillance scenario show the effectiveness of our approach and its potential for real applications of cognitive vision.
Yuzuko Utsumi, Eric Sommerlade, Nicola Bellotto, Ian D. Reid 0001
ICRA4
2012 Cognitive visual tracking and camera control
Nicola Bellotto, Ben Benfold, Hanno Harland, Hans-Hellmut Nagel, Nicola Pirlo, Ian D. Reid 0001, Eric Sommerlade
Comput. Vis. Image Underst.6
2012 PWP3D: Real-Time Segmentation and Tracking of 3D Objects
Victor Adrian Prisacariu, Ian D. Reid 0001
Int. J. Comput. Vis.2
2012 3D hand tracking for human computer interaction
Victor Adrian Prisacariu, Ian D. Reid 0001
Image Vis. Comput.2
2012 Structured Learning of Human Interactions in TV Shows
abstract
The objective of this work is recognition and spatiotemporal localization of two-person interactions in video. Our approach is person-centric. As a first stage we track all upper bodies and heads in a video using a tracking-by-detection approach that combines detections with KLT tracking and clique partitioning, together with occlusion detection, to yield robust person tracks. We develop local descriptors of activity based on the head orientation (estimated using a set of pose-specific classifiers) and the local spatiotemporal region around them, together with global descriptors that encode the relative positions of people as a function of interaction type. Learning and inference on the model uses a structured output SVM which combines the local and global descriptors in a principled manner. Inference using the model yields information about which pairs of people are interacting, their interaction class, and their head orientation (which is also treated as a variable, enabling mistakes in the classifier to be corrected using global context). We show that inference can be carried out with polynomial complexity in the number of people, and describe an efficient algorithm for this. The method is evaluated on a new dataset comprising 300 video clips acquired from 23 different TV shows and on the benchmark UT--Interaction dataset.
Alonso Patron-Perez, Marcin Marszalek, Ian D. Reid 0001, Andrew Zisserman
IEEE Trans. Pattern Anal. Mach. Intell.3
2011 Regressing Local to Global Shape Properties for Online Segmentation and Tracking
abstract
We propose a regression based learning framework that learns a set of shapes online, which can then be used to recover occluded object shapes. We represent shapes using their 2D discrete cosine transforms, and the key insight we propose is to regress low frequency harmonics, which represent the global properties of the shape, from high frequency harmonics, that encode the details of the object's shape. We learn the regression model using Locally Weighted Projection Regression (LWPR) which expedites online, incremental learning. After sufficient observation of a set of unoccluded shapes, the learned model can detect occlusion and recover the full shapes from the occluded ones. We demonstrate the ideas using a level-set based tracking system that provides shape and pose, however, the framework could be embedded in any segmentation-based tracking system. Our experiments demonstrate the efficacy of the method on a variety of objects using both real data and artificial data.
Carl Yuheng Ren, Victor Adrian Prisacariu, Ian D. Reid 0001
BMVC3
2011 Stable multi-target tracking in real-time surveillance video
abstract
The majority of existing pedestrian trackers concentrate on maintaining the identities of targets, however systems for remote biometric analysis or activity recognition in surveillance video often require stable bounding-boxes around pedestrians rather than approximate locations. We present a multi-target tracking system that is designed specifically for the provision of stable and accurate head location estimates. By performing data association over a sliding window of frames, we are able to correct many data association errors and fill in gaps where observations are missed. The approach is multi-threaded and combines asynchronous HOG detections with simultaneous KLT tracking and Markov-Chain Monte-Carlo Data Association (MCM-CDA) to provide guaranteed real-time tracking in high definition video. Where previous approaches have used ad-hoc models for data association, we use a more principled approach based on a Minimal Description Length (MDL) objective which accurately models the affinity between observations. We demonstrate by qualitative and quantitative evaluation that the system is capable of providing precise location estimates for large crowds of pedestrians in real-time. To facilitate future performance comparisons, we make a new dataset with hand annotated ground truth head locations publicly available.
Ben Benfold, Ian D. Reid 0001
CVPR2
2011 Nonlinear shape manifolds as shape priors in level set segmentation and tracking
abstract
We propose a novel nonlinear, probabilistic and variational method for adding shape information to level set-based segmentation and tracking. Unlike previous work, we represent shapes with elliptic Fourier descriptors and learn their lower dimensional latent space using Gaussian Process Latent Variable Models. Segmentation is done by a nonlinear minimisation of an image-driven energy function in the learned latent space. We combine it with a 2D pose recovery stage, yielding a single, one shot, optimisation of both shape and pose. We demonstrate the performance of our method, both qualitatively and quantitatively, with multiple images, video sequences and latent spaces, capturing both shape kinematics and object class variance.
Victor Adrian Prisacariu, Ian D. Reid 0001
CVPR2
2011 Robust 3D hand tracking for human computer interaction
abstract
We propose a system for human computer interaction via 3D hand movements, based on a combination of visual tracking and a cheap, off-the-shelf, accelerometer. We use a 3D model and region based tracker, resulting in robustness to variations in illumination, motion blur and occlusions. At the same time the accelerometer allows us to deal with the multimodality in the silhouette to pose function. We synchronise the accelerometer and tracker online, by casting the calibration problem as a maximum covariance problem, which we then solve probabilistically. We show the effectiveness of our solution with multiple real-world tests and demonstration scenarios.
Victor Adrian Prisacariu, Ian D. Reid 0001
FG2
2011 Unsupervised learning of a scene-specific coarse gaze estimator
abstract
We present a method to estimate the coarse gaze directions of people from surveillance data. Unlike previous work we aim to do this without recourse to a large hand-labelled corpus of training data. In contrast we propose a method for learning a classifier without any hand labelled data using only the output from an automatic tracking system. A Conditional Random Field is used to model the interactions between the head motion, walking direction, and appearance to recover the gaze directions and simultaneously train randomised decision tree classifiers. Experiments demonstrate performance exceeding that of conventionally trained classifiers on two large surveillance datasets.
Ben Benfold, Ian D. Reid 0001
ICCV2
2011 Manhattan scene understanding using monocular, stereo, and 3D features
abstract
This paper addresses scene understanding in the context of a moving camera, integrating semantic reasoning ideas from monocular vision with 3D information available through structure-from-motion. We combine geometric and photometric cues in a Bayesian framework, building on recent successes leveraging the indoor Manhattan assumption in monocular vision. We focus on indoor environments and show how to extract key boundaries while ignoring clutter and decorations. To achieve this we present a graphical model that relates photometric cues learned from labeled data, stereo photo-consistency across multiple views, and depth cues derived from structure-from-motion point clouds. We show how to solve MAP inference using dynamic programming, allowing exact, global inference in ~100 ms (in addition to feature computation of under one second) without using specialized hardware. Experiments show our system out-performing the state-of-the-art.
Alex Flint, David William Murray 0001, Ian D. Reid 0001
ICCV3
2011 Shared shape spaces
abstract
We propose a method for simultaneous shape-constrained segmentation and parameter recovery. The parameters can describe anything from 3D shape to 3D pose and we place no restriction on the topology of the shapes, i.e. they can have holes or be made of multiple parts. We use Shared Gaussian Process Latent Variable Models to learn multimodal shape-parameter spaces. These allow non-linear embeddings of the high-dimensional shape and parameter spaces in low dimensional spaces in a fully probabilistic manner. We propose a method for exploring the multimodality in the joint space in an efficient manner, by learning a mapping from the latent space to a space that encodes the similarity between shapes. We further extend the SGP-LVM to a model that makes use of a hierarchy of embeddings and show that this yields faster convergence and greater accuracy over the standard non-hierarchical embedding. Shapes are represented implicitly using level sets, and inference is made tractable by compressing the level set embedding functions with discrete cosine transforms. We show state of the art results in various fields, ranging from pose recovery to gaze tracking and to monocular 3D reconstruction.
Victor Adrian Prisacariu, Ian D. Reid 0001
ICCV2
2011 Hidden view synthesis using real-time visual SLAM for simplifying video surveillance analysis
abstract
Understanding and analysing video data from static or mobile surveillance cameras often requires knowledge of the scene and the camera placement. In this article, we provide a way to simplify the user's task of understanding the scene by rendering the camera view as if observed from the user's perspective by estimating his position using a real-time visual SLAM system. Augmenting the view is referred to as hidden view synthesis. Compared to previous work, the current approach improves by simplifying the setup and requiring minimal user input. This is achieved by building a map of the environment using a visual SLAM system and then registering the surveillance camera in this map. By exploiting the map, a different moving camera can render hidden views in real-time at 30Hz. We discuss some of the challenges remaining for full automation. Results are shown in an indoor environment for surveillance applications and outdoors with application to improved safety in transport.
Christopher Mei, Eric Sommerlade, Gabe Sibley, Paul Newman 0001, Ian D. Reid 0001
ICRA5
2011 Gaze directed camera control for face image acquisition
abstract
Face recognition in surveillance situations usually requires high resolution face images to be captured from remote active cameras. Since the recognition accuracy is typically a function of the face direction with frontal faces more likely to lead to reliable recognition we propose a system which optimises the capturing of such images by using coarse gaze estimates from a static camera. By considering the potential information gain from observing each target, our system automatically sets the pan, tilt and zoom values (i.e. the field of view) of multiple cameras observing different tracked targets in order to maximise the likelihood of correct identification. The expected gain in information is influenced by the controllable field of view, and by the false positive and negative rates of the identification process, which are in turn a function of the gaze angle. We validate the approach using a combination of simulated situations and real tracking output to demonstrate superior performance over alternative approaches, notably using no gaze information, or using gaze inferred from direction of travel (i.e. assuming each person is always looking directly ahead).We also show results from a live implementation with a static camera and two pan-tilt-zoom devices, involving real-time tracking, processing and control.
Eric Sommerlade, Ben Benfold, Ian D. Reid 0001
ICRA3
2011 RSLAM: A System for Large-Scale Mapping in Constant-Time Using Stereo
Christopher Mei, Gabe Sibley, Mark Joseph Cummins, Paul Newman 0001, Ian D. Reid 0001
Int. J. Comput. Vis.5
2011 Automatic Relocalization and Loop Closing for Real-Time Monocular SLAM
abstract
Monocular SLAM has the potential to turn inexpensive cameras into powerful pose sensors for applications such as robotics and augmented reality. We present a relocalization module for such systems which solves some of the problems encountered by previous monocular SLAM systems--tracking failure, map merging, and loop closure detection. This module extends recent advances in keypoint recognition to determine the camera pose relative to the landmarks within a single frame time of 33 ms. We first show how this module can be used to improve the robustness of these systems. Blur, sudden motion, and occlusion can all cause tracking to fail, leading to a corrupted map. Using the relocalization module, the system can automatically detect and recover from tracking failure while preserving map integrity. Extensive tests show that the system can then reliably generate maps for long sequences even in the presence of frequent tracking failure. We then show that the relocalization module can be used to recognize overlap in maps, i.e., when the camera has returned to a previously mapped area. Having established an overlap, we determine the relative pose of the maps using trajectory alignment so that independent maps can be merged and loop closure events can be recognized. The system combining all of these abilities is able to map larger environments and for significantly longer periods than previous systems.
Brian Patrick Williams, Georg Klein, Ian D. Reid 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2010 High Five: Recognising human interactions in TV shows
abstract
In this paper we address the problem of recognising interactions between two people in realistic scenarios for video retrieval purposes. We develop a per-person descriptor that uses attention (head orientation) and the local spatial and temporal context in a neighbourhood of each detected person. Using head orientation mitigates camera view ambiguities, while the local context, comprised of histograms of gradients and motion, aims to capture cues such as hand and arm movement. We also employ structured learning to capture spatial relationships between interacting individuals. We train an initial set of one-vs-the-rest linear SVM classifiers, one for each interaction, using this descriptor. Noting that people generally face each other while interacting, we learn a structured SVM that combines head orientation and the relative location of people in a frame to improve upon the initial classification obtained with our descriptor. To test the efficacy of our method, we have created a new dataset of realistic human interactions comprised of clips extracted from TV shows, which represents a very difficult challenge. Our experiments show that using structured learning improves the retrieval results compared to using the interaction classifiers independently.
Alonso Patron-Perez, Marcin Marszalek, Andrew Zisserman, Ian D. Reid 0001
BMVC4
2010 Mean-Shift Visual Tracking with NP-Windows Density Estimates
abstract
The mean-shift algorithm is a robust and easy method of finding local extrema in the density distribution of a data set. It has been used successfully for visual tracking in which the target is modelled using a colour histogram, and the image window with best matching histogram is sought. However a histogram is potentially a poor estimate of the underlying colour distribution: it is not invariant to the image scale, the number of histogram bins or the number of samples, and this can have an adverse affect on the speed and accuracy of convergence of the mean-shift algorithm. We apply a general non-parametric PDF estimation method [7] to replace the histogram in mean-shift tracking to improve its accuracy. This algorithm uses an interpolation scheme which fits piecewise functions to the signal samples, calculating the PDF by accumulating the contribution of each piecewise function. Its accuracy is dependent only on the accuracy of the piecewise representation, not the number of samples or the number of bins. Experiments are conducted to demonstrate that we improve mean-shift visual tracking both in accuracy and speed.
Ian D. Reid 0001
BMVC1
2010 Real-time tracking of multiple occluding objects using level sets
abstract
We derive a probabilistic framework for robust, realtime, visual tracking of multiple previously unseen objects from a moving camera. This framework models the discrete depth ordering of the objects being tracked in the scene. The method uses the observed image data to compute a posterior over the objects' poses, shapes and relative depths. The poses are group transformations, the shapes are implicit contours represented using level-sets and the relative depths give the discrete depth ordering of the objects. All nuisance variables are marginalised out at the pixel-level resulting in a pixel-wise posterior, as opposed to a pixel-wise likelihood, and we show using quantitative results that this provides increased resilience to noise. We also demonstrate how motion models can be incorporated within the same probabilistic framework and show how this enables the system to track complete occlusions. The effectiveness of our method is demonstrated on a variety of challenging video sequences.
Charles Bibby, Ian D. Reid 0001
CVPR2
2010 Growing semantically meaningful models for visual SLAM
abstract
Though modern Visual Simultaneous Localisation and Mapping (vSLAM) systems are capable of localising robustly and efficiently even in the case of a monocular camera, the maps produced are typically sparse point-clouds that are difficult to interpret and of little use for higher-level reasoning tasks such as scene understanding or human- machine interaction. In this paper we begin to address this deficiency, presenting progress on expanding the competency of visual SLAM systems to build richer maps. Specifically, we concentrate on modelling indoor scenes using semantically meaningful surfaces and accompanying labels, such as “floor”, “wall”, and “ceiling” - an important step towards a representation that can support higher-level reasoning and planning. We leverage the Manhattan world assumption and show how to extract vanishing directions jointly across a video stream. We then propose a guided line detector that utilises known vanishing points to extract extremely subtle axis- aligned edges. We utilise recent advances in single view structure recovery to building geometric scene models and demonstrate our system operating on-line.
Alex Flint, Christopher Mei, Ian D. Reid 0001, David William Murray 0001
CVPR3
2010 A Dynamic Programming Approach to Reconstructing Building Interiors
Alex Flint, Christopher Mei, David William Murray 0001, Ian D. Reid 0001
ECCV (5)4
2010 Integrating Object Detection with 3D Tracking Towards a Better Driver Assistance System
abstract
Driver assistance helps save lives. Accurate 3D pose is required to establish if a traffic sign is relevant to the driver. We propose a real-time system that integrates single view detection with region-based 3D tracking of road signs. The optimal set of candidate detections is found, followed by AdaBoost cascades and SVMs. The 2D detections are then employed in simultaneous 2D segmentation and 3D pose tracking, using the known 3D model of the recognised traffic sign. We demonstrate the abilities of our system by tracking multiple road signs in real world scenarios.
Victor Adrian Prisacariu, Radu Timofte, Karel Zimmermann, Ian D. Reid 0001, Luc Van Gool
ICPR4
2010 A hybrid SLAM representation for dynamic marine environments
abstract
We present a hybrid SLAM system for marine environments that combines cubic splines to represent the trajectories of dynamic objects, point features to represent stationary objects and an occupancy grid to represent land masses. This hybrid representation enables SLAM to be applied in environments with moving objects, where solutions using point features alone are computationally prohibitive or where dense objects e.g. landmasses can not be represented correctly using point features. Estimation is achieved using a sliding window framework with reversible data-association and reversible model-selection. Our main contributions are: (i) a hybrid representation of the environment; (ii) occupancy grid fusion is continually refined for the duration of the sliding window; (iii) the trajectories of dynamic objects are represented using cubic splines and (iv) radar scans are re-rendered at a sub-scan resolution to compensate for the egomotion during the scan acquisition period. We show that the continual refinement of the occupancy grid greatly improves the quality of the resultant map, leading to a better estimate of the egomotion and therefore better estimates of the trajectories of dynamic objects. We also demonstrate that the use of cubic splines to represent trajectories has two major advantages: (i) the state space is compressed i.e. many vehicle poses can be represented using a single spline section and (ii) the trajectory becomes continuous and so fusing information from asynchronous sensors running at multiple frequencies becomes trivial. The efficacy of our system is demonstrated using real marine radar data, showing that it can successfully estimate the positions/velocities of objects and landmasses observed during a typical voyage on a small boat.
Charles Bibby, Ian D. Reid 0001
ICRA2
2010 Planes, trains and automobiles - autonomy for the modern robot
abstract
We are concerned with enabling truly large scale autonomous navigation in typical human environments. To this end we describe the acquisition and modeling of large urban spaces from data that reflects human sensory input. Over 181GB of image and inertial data are captured using head-mounted stereo cameras. This data is processed into a relative map covering 121 km of Southern England. We point out the numerous challenges we encounter, and highlight in particular the problem of undetected ego-motion, which occurs when the robot finds itself on-or-within a moving frame of reference. In contrast to global-frame representations, we find that the continuous relative representation naturally accommodates moving-reference-frames - without having to identify them first, and without inconsistency. Within a moving-reference-frame, and without drift-less global exteroceptive sensing, motion with respect to the global-frame is effectively unobservable. This underlying truth drives us towards relative topometric solutions like relative bundle adjustment (RBA), which has no problem representing distance and metric Euclidean structure, yet does not suffer inconsistency introduced by the attempt to solve in the global-frame.
Gabe Sibley, Christopher Mei, Ian D. Reid 0001, Paul Newman 0001
ICRA3
2010 Probabilistic surveillance with multiple active cameras
abstract
In this work we present a consistent probabilistic approach to control multiple, but diverse pan-tilt-zoom cameras concertedly observing a scene. There are disparate goals to this control: the cameras are not only to react to objects moving about, arbitrating conflicting interests of target resolution and trajectory accuracy, they are also to anticipate the appearance of new targets. We base our control function on maximisation of expected mutual information gain, which to our knowledge is novel to the field of computer vision in the context of multiple pan-tilt-zoom camera control. This information theoretic measure yields a utility for each goal and parameter setting, making the use of physical or computational resources comparable. Weighting this utility allows to prioritise certain objectives or targets in the control. The resulting behaviours in typical situations for multi-camera systems, such as camera hand-off, acquisition of close-ups and scene exploration, are emergent but intuitive. We quantitatively show that without the need for hand crafted rules they address the given objectives.
Eric Sommerlade, Ian D. Reid 0001
ICRA2
2010 On combining visual SLAM and visual odometry
abstract
Sequential monocular SLAM systems perform drift free tracking of the pose of a camera relative to a jointly estimated map of landmarks. To allow real-time operation in moderately sized environments, the map is kept quite spare with usually only tens of landmarks visible in each frame. In contrast, visual odometry techniques track hundreds of visual features per frame. This leads to a very accurate estimate of the relative camera motion, but without a persistent map, the estimate tends to drift over time. We demonstrate a new monocular SLAM system which combines the benefits of these two techniques. In addition to maintaining a sparse map of landmarks in the world, our system finds as many inter-frame point matches as possible. These point matches provide additional constraints on the inter-frame motion of the camera leading to a more accurate pose estimate, and, since they are not maintained as full map landmarks, they do not cause a large increase in the computational cost. Our results in both a simulated environment and in real video demonstrate the improvement in estimation accuracy gained by the inclusion of visual odometry style observations. The constraints available from pairwise point matches are most naturally cast in the context of a camera-centric rather than world-centric frame. To that end we recast the usual world-centric EKF implementation of visual SLAM in a robo-centric frame. We show that this robo-centric visual SLAM, as expected, leads to the estimated uncertainty more closely matching the ideal uncertainty; i.e., that robo-centric visual SLAM yields a more consistent estimate than the traditional world-centric EKF algorithm.
Brian Patrick Williams, Ian D. Reid 0001
ICRA2
2010 Multiview segmentation and tracking of dynamic occluding layers
Ian D. Reid 0001, Keith Richard Connor
Image Vis. Comput.1
2009 Guiding Visual Surveillance by Tracking Human Attention
abstract
We describe a novel method for directing the attention of an automated surveillance system. Our starting premise is that the attention of people in a scene can be used as an indicator of interesting areas and events. To determine people’s attention from passive visual observations we develop a system for automatic tracking and detection of individual heads to infer their gaze direction. The former is achieved by combining a histogram of oriented gradient (HOG) based head detector with frame-to-frame tracking using multiple point features to provide stable head images. The latter is achieved using a head pose classification method which uses randomised ferns with decision branches based on both HOG and colour based features to determine a coarse gaze direction for each person in the scene. By building both static and temporally varying maps of areas where people look we are able to identify interesting regions.
Ben Benfold, Ian D. Reid 0001
BMVC2
2009 A Constant-Time Efficient Stereo SLAM System
abstract
Continuous, real-time mapping of an environment using a camera requires a constant-time estimation engine. This rules out optimal global solving such as bundle adjustment. In this article, we investigate the precision that can be achieved with only local estimation of motion and structure provided by a stereo pair. We introduce a simple but novel representation of the environment in terms of a sequence of relative locations. We demonstrate precise local mapping and easy navigation using the relative map, and importantly show that this can be done without requiring a global minimisation after loop closure. We discuss some of the issues that arise from using a relative representation, and evaluate our system on long sequences processed at a constant 30-45 Hz, obtaining precisions down to a few metres over distances of a few kilometres.
Christopher Mei, Gabe Sibley, Mark Joseph Cummins, Paul Newman 0001, Ian D. Reid 0001
BMVC5
2009 PWP3D: Real-time Segmentation and Tracking of 3D Objects
abstract
We formulate a probabilistic framework for simultaneous 2D segmentation and 2D– 3D pose tracking, using a known 3D model (of arbitrary shape) of the segmented object. Our technique is region-based; at each frame we maximise the discrimination between statistical foreground and background models, by adjusting the pose parameters itera-tively. Unlike all previous work in 3D tracking, we use posterior membership proba-bilities for foreground and background pixels, rather than pixel likelihoods, and during periods of stable tracking we allow adaptation of the statistical foreground and back-ground models. We support our ideas with a real-time implementation, and use this to generate experimental results on both real and artificial video sequences, with a number of 3D models, to showcase the qualities of our tracker, and to demonstrate the benefit of using pixel-wise posteriors rather than likelihoods. 1
Victor Adrian Prisacariu, Ian D. Reid 0001
BMVC2
2009 Video synchronization from human motion using rank constraints
Philip A. Tresadern, Ian D. Reid 0001
Comput. Vis. Image Underst.2
2009 Global Stereo Reconstruction under Second-Order Smoothness Priors
abstract
Second-order priors on the smoothness of 3D surfaces are a better model of typical scenes than first-order priors. However, stereo reconstruction using global inference algorithms, such as graph cuts, has not been able to incorporate second-order priors because the triple cliques needed to express them yield intractable (nonsubmodular) optimization problems. This paper shows that inference with triple cliques can be effectively performed. Our optimization strategy is a development of recent extensions to alpha -- expansion, based on the "QPBO" algorithm. The strategy is to repeatedly merge proposal depth maps using a novel extension of QPBO. Proposal depth maps can come from any source, for example, frontoparallel planes as in alpha-expansion, or indeed any existing stereo algorithm, with arbitrary parameter settings.
Oliver J. Woodford, Philip Torr 0001, Ian D. Reid 0001, Andrew W. Fitzgibbon
IEEE Trans. Pattern Anal. Mach. Intell.3
2008 Colour Invariant Head Pose Classification in Low Resolution Video
abstract
This paper presents an algorithm for the classification of head pose in low resolution video. Invariance to skin, hair and background colours is achieved by classifying using an ensemble of randomised ferns which have been trained on labelled images. The ferns are used to simultaneously classify the head pose and to identify the most likely hypothesis for the mapping between colours and labels. Results from video sequences demonstrate that an improved posterior estimation using learnt colour distributions reduces classification error and provides accurate pose information in images where the head occupies as little as 10 pixels square. 1
Ben Benfold, Ian D. Reid 0001
BMVC2
2008 Probabilistic Parameter Selection for Learning Scene Structure from Video
abstract
We present an online learning approach for robustly combining unreliable observations from a pedestrian detector to estimate the rough 3D scene geometry from video sequences of a static camera. Our approach is based on an entropy modelling framework, which allows to simultaneously adapt the detector parameters, such that the expected information gain about the scene structure is maximised. As a result, our approach automatically restricts the detector scale range for each image region as the estimation results become more confident, thus improving detector run-time and limiting false positives. 1
Michael D. Breitenstein, Eric Sommerlade, Bastian Leibe, Luc Van Gool, Ian D. Reid 0001
BMVC5
2008 From Visual Query to Visual Portrayal
abstract
In this paper we show how online images can be automatically exploited for scene visualization and reconstruction starting from a mere visual query provided by the user. A visual query is used to retrieve images of a landmark place using a visual search engine. These images are used to reconstruct robust 3–D features and camera poses in projective space. Novel views are then rendered corresponding to a virtual camera flying smoothly through the projective space by triangulation of the projected points in the output view. We introduce a method to fuse the rendered novel views from all input images at each virtual view point by computing their intrinsic image and illuminations. This approach allows us to remove the occlusions and maintain consistent and controlled illumination throughout the rendered sequence. We demonstrate the performance of our prototype system on two landmark structures. 1
Ali Shahrokni, Christopher Mei, Philip Torr 0001, Ian D. Reid 0001
BMVC4
2008 Modeling and generating complex motion blur for real-time tracking
abstract
This article addresses the problem of real-time visual tracking in presence of complex motion blur. Previous authors have observed that efficient tracking can be obtained by matching blurred images instead of applying the computationally expensive task of deblurring (H. Jin et al., 2005). The study was however limited to translational blur. In this work, we analyse the problem of tracking in presence of spatially variant motion blur generated by a planar template. We detail how to model the blur formation and parallelise the blur generation, enabling a real-time GPU implementation. Through the estimation of the camera exposure time, we discuss how tracking initialisation can be improved. Our algorithm is tested on challenging real data with complex motion blur where simple models fail. The benefit of blur estimation is shown for structure and motion.
Christopher Mei, Ian D. Reid 0001
CVPR2
2008 Information-theoretic active scene exploration
abstract
Studies support the need for high resolution imagery to identify persons in surveillance videos. However, the use of telephoto lenses sacrifices a wider field of view and thereby increases the uncertainty of other, possibly more interesting events in the scene. Using zoom lenses offers the possibility of enjoying the benefits of both wide field of view and high resolution, but not simultaneously. We approach this problem of balancing these finite imaging resources - or of exploration vs exploitation - using an information-theoretic approach. We argue that the camera parameters - pan, tilt and zoom - should be set to maximise information gain, or equivalently minimising conditional entropy of the scene model, comprised of multiple targets and a yet unobserved one. The information content of the former is supplied directly by the uncertainties computed using a Kalman filter tracker, while the latter is modelled using a rdquobackgroundrdquo Poisson process whose parameters are learned from extended scene observations; together these yield an entropy for the scene. We support our argument with quantitative and qualitative analyses in simulated and real-world environments, demonstrating that this approach yields sensible exploration behaviours in which the camera alternates between obtaining close-up views of the targets while paying attention to the background, especially to areas of known high activity.
Eric Sommerlade, Ian D. Reid 0001
CVPR2
2008 Global stereo reconstruction under second order smoothness priors
abstract
Second-order priors on the smoothness of 3D surfaces are a better model of typical scenes than first-order priors. However, stereo reconstruction using global inference algorithms, such as graph-cuts, has not been able to incorporate second-order priors because the triple cliques needed to express them yield intractable (non-submodular) optimization problems. This paper shows that inference with triple cliques can be effectively optimized. Our optimization strategy is a development of recent extensions to a-expansion, based on the "QPBO" algorithm [5, 14, 26]. The strategy is to repeatedly merge proposal depth maps using a novel extension of QPBO. Proposal depth maps can come from any source, for example fronto-parallel planes as in a-expansion, or indeed any existing stereo algorithm, with arbitrary parameter settings. Experimental results demonstrate the usefulness of the second-order prior and the efficacy of our optimization framework. An implementation of our stereo framework is available online [34].
Oliver J. Woodford, Philip Torr 0001, Ian D. Reid 0001, Andrew W. Fitzgibbon
CVPR3
2008 Robust Real-Time Visual Tracking Using Pixel-Wise Posteriors
Charles Bibby, Ian D. Reid 0001
ECCV (2)2
2008 Target tracking using mean-shift and affine structure
abstract
In this paper, we present a new approach for tracking targets with their size and shape time-varying, based on a combination of mean-shift and affine structure. Although the well-known mean-shift colour-based tracking algorithm is an effective tracking tool, difficulties arise when it is applied to track a size-changing visual target due to the fixed kernel-bandwidth. To improve this, the present study employs a corner detector on the object candidate from mean-shift and reconstructs the target position and relative scale between frames using the affine structure available from two or three views. In comparison experiments against previous algorithms, the present model shows better tracking-consistency and good efficiency. Our algorithm is also demonstrated in a real-time implement controlling the pan-tilt-zoom parameters of an active camera. The results indicate the modelpsilas tracking capability in the presence of scale change and partial occlusions.
Ian D. Reid 0001
ICPR3
2008 Influence of zoom selection on a Kalman filter
abstract
The use of a single camera with a zoom lens for tracking involves a continuous arbitration of accuracy vs. reliability. We address this problem with an information-theoretic approach, where we extend zoom selection based on conditional entropy by incorporating the fixation errors into the observation likelihood. We present a thorough analysis of previous approaches, revealing zoom and speed limits, especially how the ratio of process to measurement noise effectively limits the maximally usable zoom for any system tracking with a Kalman filter. This work finally presents means to circumvent aforementioned limitations.
Eric Sommerlade, Ian D. Reid 0001
IROS2
2008 An image-to-map loop closing method for monocular SLAM
abstract
In this paper we present a loop closure method for a handheld single-camera SLAM system based on our previous work on relocalization. By finding correspondences between the current image and the map, our system is able to reliably detect loop closures. We compare our algorithm to existing techniques for loop closure in single-camera SLAM based on both image-to-image and map-to-map correspondences and discuss both the reliability and suitability of each algorithm in the context of monocular SLAM.
Brian Patrick Williams, Mark Joseph Cummins, José Neira, Paul Newman 0001, Ian D. Reid 0001, Juan D. Tardós
IROS5
2008 Camera calibration from human motion
Philip A. Tresadern, Ian D. Reid 0001
Image Vis. Comput.2
2007 Temporal Priors for Novel Video Synthesis
Ali Shahrokni, Oliver J. Woodford, Ian D. Reid 0001
ACCV (2)3
2007 A Probabilistic Framework for Recognizing Similar Actions using Spatio-Temporal Features
abstract
One of the challenges found in recent methods for action recognition has been to classify ambiguous actions successfully. In the case of methods that use spatio-temporal features this phenomenon is observed when two actions generate similar feature types. Ideally, a probabilistic classification method would be based on a model of the full joint distribution of features, but this is computationally intractable. In this paper we propose using an approximation of the full joint via first order dependencies between feature types using so-called Chow-Liu trees. We obtain promising results and achieve an improvement in the classification accuracy over naive Bayes and other simple classifiers. Our implementation of the method makes use of a binary descriptor for a video analogous to one previously used in location recognition for mobile robots. Because of the simplicity of the algorithm, once the offline learning phase is over, real-time action recognition is possible and we present an adaptation of this method that works in real-time. 1
Alonso Patron-Perez, Ian D. Reid 0001
BMVC2
2007 An Evaluation of Shape Descriptors for Image Retrieval in Human Pose Estimation
abstract
This paper presents an empirical comparison of several shape representations in order to search a database of training examples (silhouettes) for the task of human pose estimation. In particular, we compare the Discrete Cosine Transform (DCT), Lipschitz embeddings and the Histogram of Shape Contexts that has previously demonstrated some success in this task. Our results suggest that a simple linear transformation of the image (such as the DCT) is as effective as the more complex, non-linear methods. 1
Philip A. Tresadern, Ian D. Reid 0001
BMVC2
2007 On New View Synthesis Using Multiview Stereo
abstract
We show that application of modern multiview stereo techniques to the newview synthesis (NVS) problem introduces a number of non-trivial complexities. By simultaneously solving for the colour and depth of the new-view pixels we can eliminate the visual artefacts that conventional NVS-via-stereo suffers. The global occlusion reasoning which has led to considerable improvements in recent stereo algorithms can easily be included in the new algorithm, using a recently improved graph-cut-based optimizer for general multi-label conditional random fields (CRFs). However, the CRF priors that are important to success in stereo cannot be easily applied if the reconstruction is to be computed in the reference frame of the novel view. We address this problem by extending recent work on the fast optimization of texture priors in NVS to model the image edge structure, yielding a synthesis of the two approaches which yields good results on difficult image sequences. 1
Oliver J. Woodford, Ian D. Reid 0001, Philip Torr 0001, Andrew W. Fitzgibbon
BMVC2
2007 Efficient new-view synthesis using pairwise dictionary priors
abstract
New-view synthesis (NVS) using texture priors (as opposed to surface-smoothness priors) can yield high quality results, but the standard formulation is in terms of large-clique Markov random fields (MRFs). Only local optimization methods such as iterated conditional modes, which are prone to fall into local minima close to the initial estimate, are practical for solving these problems. In this paper we replace the large-clique energies with pairwise potentials, by restricting the patch dictionary for each clique to image regions suitable for that clique. This enables for the first time the use of a global optimization method, such as tree-reweighted message passing, to solve the NVS problem with image-based priors. We employ a robust, truncated quadratic kernel to reject outliers caused by occlusions, specularities and moving objects, within our global optimization. Because the MRF optimization is thus fast, computing the unary potentials becomes the new performance bottleneck. An additional contribution of this paper is a novel, fast method for enumerating color modes of the per-pixel unary potentials, despite the non-convex nature of our robust kernel. We compare the results of our technique with other rendering methods, and discuss the relative merits and flaws of regularizing color, and of local versus global dictionaries.
Oliver J. Woodford, Ian D. Reid 0001, Andrew W. Fitzgibbon
CVPR2
2007 Real-Time SLAM Relocalisation
abstract
Monocular SLAM has the potential to turn inexpensive cameras into powerful pose sensors for applications such as robotics and augmented reality. However, current implementations lack the robustness required to be useful outside laboratory conditions: blur, sudden motion and occlusion all cause tracking to fail and corrupt the map. Here we present a system which automatically detects and recovers from tracking failure while preserving map integrity. By extending recent advances in keypoint recognition the system can quickly resume tracking - i.e. within a single frame time of 33 ms - using any of the features previously stored in the map. Extensive tests show that the system can reliably generate maps for long sequences even in the presence of frequent tracking failure.
Brian Patrick Williams, Georg Klein, Ian D. Reid 0001
ICCV3
2007 Automatic Relocalisation for a Single-Camera Simultaneous Localisation and Mapping System
abstract
We describe a fast method to relocalise a monocular visual SLAM (simultaneous localisation and mapping) system after tracking failure. The monocular SLAM system stores the 3D locations of visual landmarks, together with a local image patch. When the system becomes lost, candidate matches are obtained using correlation, then the pose of the camera is solved via an efficient implementation of RANSAC using a three-point-pose algorithm. We demonstrate the usefulness of this method within visual SLAM: (i) we show tracking can reliably resume after tracking failure due to occlusions, motion blur or unmodelled rapid motions; (ii) we show how the method can be used as an adjunct for a proposal distribution in a particle filter framework; (iii) during successful tracking we use idle cycles to test if the current map overlaps with a previously-built map, and we provide a solution to aligning the two maps by splicing the camera trajectories in a consistent and optimal way.
Brian Patrick Williams, Ian D. Reid 0001
ICRA3
2007 MonoSLAM: Real-Time Single Camera SLAM
abstract
We present a real-time algorithm which can recover the 3D trajectory of a monocular camera, moving rapidly through a previously unknown scene. Our system, which we dub MonoSLAM, is the first successful application of the SLAM methodology from mobile robotics to the "pure vision" domain of a single uncontrolled camera, achieving real time but drift-free performance inaccessible to Structure from Motion approaches. The core of the approach is the online creation of a sparse but persistent map of natural landmarks within a probabilistic framework. Our key novel contributions include an active approach to mapping and measurement, the use of a general motion model for smooth camera movement, and solutions for monocular feature initialization and feature orientation estimation. Together, these add up to an extremely efficient and robust algorithm which runs at 30 Hz with standard PC and camera hardware. This work extends the range of robotic systems in which SLAM can be usefully applied, but also opens up new areas. We present applications of MonoSLAM to real-time 3D localization and mapping for a high-performance full-size humanoid robot and live augmented reality with a hand-held camera.
Andrew J. Davison, Ian D. Reid 0001, Nicholas Molton, Olivier Stasse
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 Real-Time Monocular SLAM with Straight Lines
abstract
The use of line features in real-time visual tracking applications is commonplace when a prior map is available, but building the map while tracking in real-time is much more difficult. We describe how straight lines can be added to a monocular Extended Kalman Filter Simultaneous Mapping and Localisation (EKF SLAM) system in a manner that is both fast and which integrates easily with point features. To achieve real-time operation, we present a fast straight-line detector that hypothesises and tests straight lines connecting detected seed points. We demonstrate that the resulting system provides good camera localisation and mapping in real-time on a standard workstation, using either line features alone, or lines and points combined.
Ian D. Reid 0001, Andrew J. Davison
BMVC2
2006 Fields of Experts for Image-based Rendering
abstract
Image priors for novel view synthesis have traditionally been non-parametric models based on large libraries of image patch exemplars, producing highquality results but making inference very slow. Recently a parametric framework, called Fields of Experts, has been proposed for image restoration that promises to speed up inference dramatically. In this paper we apply Fields of Experts for the first time to the problem of novel view synthesis, posed as a Markov random field labelling problem with very large cliques. Additionally, we introduce to computer vision for the first time a new optimization algorithm from statistical physics which reaches better minima than the ICM and simulated annealing algorithms to which such large-clique problems have previously been restricted. 1
Oliver J. Woodford, Ian D. Reid 0001, Philip Torr 0001, Andrew W. Fitzgibbon
BMVC2
2006 Estimating Gaze Direction from Low-Resolution Faces in Video
Neil Robertson 0002, Ian D. Reid 0001
ECCV (2)2
2006 A general method for human activity recognition in video
Neil Robertson 0002, Ian D. Reid 0001
Comput. Vis. Image Underst.2
2006 Automated Alignment of Robotic Pan-Tilt Camera Units Using Vision
Joss Knight, Ian D. Reid 0001
Int. J. Comput. Vis.2
2005 Multiview Segmentation and Tracking of Dynamic Occluding Layers
abstract
We present an algorithm for the layered segmentation of video data in multiple views. The approach is based on computing the parameters of a layered representation of the scene in which each layer is modelled by its motion, appearance and occupancy, where occupancy describes, probabilistically, the layer's spatial extent and not simply its segmentation in a particular view. The problem is formulated as the MAP estimation of all layer parameters conditioned on those at the previous time step; i.e., a sequential estimation problem that is equivalent to tracking multiple objects in a given number views. Expectation-Maximisation is used to establish layer posterior probabilities for both occupancy and visibility, which are represented distinctly. Evidence from areas in each view which are described poorly under the model is used to propose new layers automatically. Since these potential new layers often occur at the fringes of images, the algorithm is able to segment and track these in a single view until such time as a suitable candidate match is discovered in the other views. The algorithm is shown to be very effective at segmenting and tracking non-rigid objects and can cope with extreme occlusion. We demonstrate an application of this representation to dynamic novel view synthesis.
Ian D. Reid 0001, Keith Richard Connor
BMVC1
2005 Articulated Structure from Motion by Factorization
abstract
Multibody affine structure from motion (SFM) methods commonly assume independent motion between objects such that the 'measurement matrix' has rank 4k. When multiple views are available, each object is then independently calibrated to a metric co-ordinate frame. However, articulated motion results in a further decrease in rank - a fact that we exploit to detect articulated objects and determine their degrees of freedom using simple linear methods. Furthermore, these objects cannot be recovered and calibrated independently since this violates articulation constraints. We show that articulation constraints can be imposed during factorization and self-calibration to recover consistent 3D structure and motion, from which link lengths and joint angles can be computed. The stability of the method is evaluated using synthetic data for comparison with ground truth and results are also presented for real image sequences.
Philip A. Tresadern, Ian D. Reid 0001
CVPR (2)2
2005 Behaviour Understanding in Video: A Combined Method
abstract
In this paper we develop a system for human behaviour recognition in video sequences. Human behaviour is modelled as a stochastic sequence of actions. Actions are described by a feature vector comprising both trajectory information (position and velocity), and a set of local motion descriptors. Action recognition is achieved via probabilistic search of image feature databases representing previously seen actions. A HMM which encodes the rules of the scene is used to smooth sequences of actions. High-level behaviour recognition is achieved by computing the likelihood that a set of predefined hidden Markov models explains the current action sequence. Thus, human actions and behaviour are represented using a hierarchy of abstraction: from simple actions, to actions with spatio-temporal context, to action sequences and finally general behaviours. While the upper levels all use (parametric) Bayes networks and belief propagation, the lowest level uses nonparametric sampling from a previously learned database of actions. The combined method represents a general framework for human behaviour modelling. In this paper we demonstrate the results chiefly on broadcast tennis sequences for automated video annotation.
Neil Robertson 0002, Ian D. Reid 0001
ICCV2
2005 Visual Tracking at Sea
abstract
We describe progress in the design, build and test of a mechatronic system capable of visually tracking objects at sea from a moving platform, such as a boat. The mechanism comprises three controllable degrees of freedom that can carry a small camera as payload. The camera position is stabilised using inertial sensing, and the stabilised images are processed using a published colour-based tracking algorithm to achieve visual pursuit of the target [1], [3]. We describe novel improvements to the tracker that enhance its performance in the case of tracking a target at sea. The system is demonstrated on real footage and its performance assessed.
Charles Bibby, Ian D. Reid 0001
ICRA2
2005 Articulated Body Motion Capture by Stochastic Search
Jonathan Deutscher, Ian D. Reid 0001
Int. J. Comput. Vis.2
2005 Image interpolation for virtual sports scenarios
Tomás Rodríguez, Ian D. Reid 0001, Radu Horaud, Navneet Dalal, Marcelo Götz
Mach. Vis. Appl.2
2004 Dynamic Classifier for Non-rigid Human motion analysis
abstract
Automatic analysis (parsing) of non-rigid human motion in a cluttered outdoor enviroment is a useful but challenging task. In a single view point, the lack of depth order relations causes a major ambiguity of the object identities. Coupled with the non-rigidity of articulation, 3D human motion tracking/pose estimation in one view is a formidable problem. In this paper, we present a novel solution that directly address this depth ambiguity, in which we extend a discriminative analysis (Support Vector Machine (SVM)) to non-rigid human motion classification with a temporal generative motion model (Hidden Markov Model (HMM)). This method can discriminate dynamic depth ordering as well as 3D articulated motion automatically from 2D images. Experiments with this method have demonstrated promising results. 1
Huang Fei, Ian D. Reid 0001
BMVC2
2004 Locally Planar Patch Features for Real-Time Structure from Motion
abstract
The performance of sequential structure from motion systems, where scene mapping is sparse to permit real-time operation, depends greatly on the ability to repeatedly measure the same visual features from a wide range of viewpoints. While previous systems have tracked features as 2D templates in image space, we show that long-term tracking is improved by treating salient feature patches as observations of locally planar regions on 3D world surfaces. Within a SLAM framework for motion and structure estimation, a gradient-based image alignment method is used to deduce estimates feature surface normal estimates, enabling pre-warping of templates for matching. As an added benefit these normals provide a richer description of the scene. 1
Nicholas Molton, Andrew J. Davison, Ian D. Reid 0001
BMVC3
2004 Uncalibrated and Unsynchronized Human Motion Capture: A Stereo Factorization Approach
Philip A. Tresadern, Ian D. Reid 0001
CVPR (1)2
2004 Joint Bayes Filter: A Hybrid Tracker for Non-rigid Hand Motion Recognition
Huang Fei, Ian D. Reid 0001
ECCV (3)2
2003 A Multiple View Layered Representation for Dynamic Novel View Synthesis
abstract
We propose a multiple view layered representation for tracking and segmentation of multiple objects in a scene. Existing layered approaches are dominated by the single view case and generally exploit only motion cues. We extend this to integrate static, dynamic and structural cues over a pair of views. The goal is to update coherent correspondence information sequentially, producing a multi-object tracker as a natural byproduct. We formulate a MAP solution for estimating layer parameters which are consistent across views, with the EM algorithm used to determine both the hidden segmentation labelling and motion parameters. A persistent representation of occupancy is maintained in spite of occlusion without enforcing a particular parametric shape model. An immediate application is dynamic novel view synthesis, for which our layered approach offers a direct and convenient representation. 1
Keith Richard Connor, Ian D. Reid 0001
BMVC2
2003 Synchronizing Image Sequences of Non-Rigid Objects
abstract
For stereopsis, images of a given scene must be captured at the same instant to ensure temporal consistency. For sequences of images (i.e. video streams) this requires the potentially costly and technically complex process of synchronizing cameras. We present a simple but effective method for automatically recovering the sub-frame temporal offset between image sequences taken using unsynchronized cameras. Having recovered the offset, we obtain the affine structure of a non-rigid motion. The technique is demonstrated for the application of human motion capture. 1
Philip A. Tresadern, Ian D. Reid 0001
BMVC2
2003 Linear Auto-Calibration for Ground Plane Motion
abstract
Planar scenes would appear to be ideally suited for self-calibration because, by eliminating the problems of occlusion and parallax, high accuracy two-view relationships can be calculated without restricting motion to pure rotation. Unfortunately, the only monocular solutions so far devised involve costly nonlinear minimizations, which must be initialized with educated guesses for the calibration parameters. So far, this problem has been circumvented by using stereo or a known calibration object. In this work we show that when there is some control over the motion of the camera, a fast linear solution is available without these restrictions. For a camera undergoing a motion about a plane-normal rotation axis (typified for instance by a motion in the plane of the scene), the complex eigenvectors of a plane-induced homography are coincident with the circular points of the motion. Three such homographies provide sufficient information to solve for the image of the absolute conic (IAC), and therefore the calibration parameters. The required situation arises most commonly when the camera is viewing the ground plane, and either moving along it, or rotating about some vertical axis. We demonstrate a number of useful applications, and show the algorithm to be simple, fast, and accurate.
Joss Knight, Andrew Zisserman, Ian D. Reid 0001
CVPR (1)3
2002 Novel View Specification and Synthesis
abstract
Given a set of real images, Novel View Synthesis (NVS) aims to produce views of a scene that would correspond to that of a virtual camera. There exist many approaches to solving this problem. We consider physically valid NVS methods, in particular those based on epipolar and trifocal transfer. We review a number of methods and place them into a common framework of view specification and mapping method. We present a new method for di-rectly specifying the novel camera motion for epipolar transfer. We also de-velop a backward mapping scheme for trifocal transfer which overcomes the problems associated with standard forward mapping methods. 1
Keith Richard Connor, Ian D. Reid 0001
BMVC2
2002 Self-Calibration of Rotating and Zooming Cameras
Lourdes Agapito, Eric Hayman, Ian D. Reid 0001
Int. J. Comput. Vis.3
2001 Automatic Partitioning of High Dimensional Search Spaces Associated with Articulated Body Motion Capture
abstract
Particle filters have proven to be an effective tool for visual tracking in non-Gaussian, cluttered environments. Conventional particle filters, however, do not scale to the problem of human motion capture (HMC) because of the large number of degrees of freedom involved. Annealed Particle Filtering (APF), introduced by J. Deutscher et al. (2000), tackled this by layering the search space and was shown to be a very effective tool for HMC. We improve upon and extend the APF in two ways. First we develop a hierarchical search strategy which automatically partitions the search space without any explicit representation of the partitions. Then we introduce a crossover operator (similar to that found in genetic algorithms) which improves the ability of the tracker to search different partitions in parallel. We present results for a simple example to demonstrate the new algorithm's implementation and then apply it to the considerably more complex problem of human motion capture with 34 degrees of freedom.
Jonathan Deutscher, Andrew J. Davison, Ian D. Reid 0001
CVPR (2)3
2001 Towards constant time SLAM using postponement
abstract
Many recent approaches to simultaneous localisation and mapping (SLAM) use an extended Kalman filter (EKF) to update and maintain a map of vehicle location. and multiple feature positions as a sensor moves through a scene. Although it is a highly powerful and well-used tool, it suffers from a well-known complexity problem. In this paper we outline the postponement technique which allows for much greater flexibility on when to use the available processing time, while not affecting the optimality of the filter. It works by updating a constant-sized data set based on current measurements, which can be used to affect the updates on all unobserved parts of the map at a later stage. By expanding the set of updated features when each new feature is observed we show that the full map update can be postponed indefinitely. We also demonstrate how postponement can be used to improve the performance of sub-optimal algorithms by applying it to a simple constant time method.
Joss Knight, Andrew J. Davison, Ian D. Reid 0001
IROS3
2001 Self-Calibration of Rotating and Zooming Cameras
Lourdes Agapito, Eric Hayman, Ian D. Reid 0001
Int. J. Comput. Vis.3
2001 Providing synthetic views for teleoperation using visual pose tracking in multiple cameras
abstract
This paper describes a visual tool for teleoperative experimentation involving remote manipulation and contact tasks. Using modest hardware, it recovers in real time the pose of moving polyhedral objects, and presents a synthetic view of the scene to the operator of a teleoperated robot using any chosen viewpoint and viewing direction. To recover pose, the method of line tracking first introduced by Harris (1992) is extended to multiple calibrated cameras, and its dynamic performance improved using robust methods and iterative filtering. Experiments are reported which determine the static and dynamic performance of the vision system, and its use in teleoperation is illustrated in two experiments, a peg-in-hole manipulation task and an impact control task.
Richard L. Thompson, Ian D. Reid 0001, L. A. Muñoz, David William Murray 0001
IEEE Trans. Syst. Man Cybern. Part A2
2000 Articulated Body Motion Capture by Annealed Particle Filtering
abstract
The main challenge in articulated body motion tracking is the large number of degrees of freedom (around 30) to be recovered. Search algorithms, either deterministic or stochastic, that search such a space without constraint, fall foul of exponential computational complexity. One approach is to introduce constraints: either labelling using markers or colour coding, prior assumptions about motion trajectories or view restrictions. Another is to relax constraints arising from articulation, and track limbs as if their motions were independent. In contrast, we aim for general tracking without special preparation of objects or restrictive assumptions. The principal contribution of the paper is the development of a modified particle filter for search in high dimensional configuration spaces. It uses a continuation principle based on annealing to introduce the influence of narrow peaks in the fitness function, gradually. The new algorithm, termed annealed particle filtering, is shown to be capable of recovering full articulated body motion efficiently.
Jonathan Deutscher, Andrew Blake 0001, Ian D. Reid 0001
CVPR3
2000 The Role of Self-Calibration in Euclidean Reconstruction from Two Rotating and Zooming Cameras
Eric Hayman, Lourdes Agapito, Ian D. Reid 0001, David William Murray 0001
ECCV (2)3
2000 Binocular Self-Alignment and Calibration from Planar Scenes
Joss Knight, Ian D. Reid 0001
ECCV (2)2
2000 Motion Estimation Using the Differential Epipolar Equation
abstract
We consider the motion estimation problem in the case of very closely spaced views. We revisit the differential epipolar equation providing an interpretation of it. On the basis of this interpretation we introduce a cost function to estimate the parameters of the differential epipolar equation, which enables us to compute the camera extrinsics and some intrinsics. In the synthetic tests performed we compare this continuous method with traditional discrete motion estimation and, contrary to previous findings by Vieville et al. (1996), show that the continuous method did not perceive any computational advantage.
Luis Baumela, Lourdes Agapito, Ian D. Reid 0001, Pablo Bustos
ICPR3
2000 Self-Calibration of a Stereo Rig in a Planar Scene by Data Combination
abstract
We present a very simple and effective method for eliminating the degeneracy inherent in a planar scene, and demonstrate its performance in a useful application - binocular self-calibration. The projective geometry of planar scenes suffers from a two-fold ambiguity: projective structure cannot be recovered from a pair of images of the scene, and transformations in projective 3-space are underconstrained for features lying on a plane. We overcome these problems by combining the feature data from different pairs of images where we know the relationship between the images is identical. In essence this generates images of a non-degenerate scene which can then be processed using standard algorithms. The only constraint is that the binocular head must be able to make repeated identical rotations about its axes. The benefits of this technique are demonstrated by its use in solving the problem of self-calibration. The use of data combination enables standard general scene stereo calibration algorithms to be used.
Joss Knight, Ian D. Reid 0001
ICPR2
2000 Active Visual Alignment of a Mobile Stereo Camera Platform
abstract
We present a complete system for automatic alignment and calibration of a stereo pan-tilt camera platform on a mobile robot. The system uses visual data from one or two controlled rotations of the head, and a single forward motion of the robot. We show how the images alone provide head alignment information, camera calibration, and head geometry. We also discuss automatic zeroing of steering angle for a single steering wheel AGV. Results are provided from tests on the working system.
Joss Knight, Ian D. Reid 0001
ICRA2
2000 Single View Metrology
Antonio Criminisi, Ian D. Reid 0001, Andrew Zisserman
Int. J. Comput. Vis.2
1999 Single View Metrology
abstract
We describe how 3D affine measurements may be computed from a single perspective view of a scene given only minimal geometric information determined from the image. This minimal information is typically the vanishing line of a reference plane and a vanishing point for a direction not parallel to the plane. It is shown that affine scene structure may then be determined from the image, without knowledge of the camera's internal calibration (e.g. focal length), nor of the explicit relation between camera and world (pose). In particular we show how to: compute the distance between planes parallel to the reference plane (up to a common scale factor); compute area and length ratios on any plane parallel to the reference plane; determine the camera's (viewer's) location. Simple geometric derivations are given for these results. We also develop an algebraic representation which unifies the three types of measurement and, amongst other advantages, permits a first order error propagation analysis to be performed, associating an uncertainty with each measurement. We demonstrate the technique for a variety of applications, including height measurements in forensic images and 3D graphical modelling from single images.
Antonio Criminisi, Ian D. Reid 0001, Andrew Zisserman
ICCV2
1999 Camera Calibration and the Search for Infinity
abstract
This paper considers the problem of self-calibration of a camera from an image sequence in the case where the camera's internal parameters (most notably focal length) may change. The problem of camera self-calibration from a sequence of images has proven to be a difficult one in practice, due to the need ultimately to resort to non-linear methods, which have often proven to be unreliable. In a stratified approach to self-calibration, a projective reconstruction is obtained first and this is successively refined first to an affine and then to a Euclidean (or metric) reconstruction. It has been observed that the difficult step is to obtain the affine reconstruction, or equivalently to locate the plane at infinity in the projective coordinate frame. The problem is inherently non-linear and requires iterative methods that risk not finding the optimal solution. The present paper overcomes this difficulty by imposing chirality constraints to limit the search for the plane at infinity to a 3-dimensional cubic region of parameter space. It is then possible to carry out a dense search over this cube in reasonable time. For each hypothesised placement of the plane at infinity, the calibration problem is reduced to one of calibration of a nontranslating camera, for which fast non-iterative algorithms exist. A cost function based on the result of the trial calibration is used to determine the best placement of the plane at infinity. Because of the simplicity of each trial, speeds of over 10,000 trials per second are achieved on a 256 MHz processor. It is shown that this dense search allows one to avoid areas of local minima effectively and find global minima of the cost function.
Richard I. Hartley, Lourdes Agapito, Ian D. Reid 0001, Eric Hayman
ICCV3
1999 A plane measuring device
Antonio Criminisi, Ian D. Reid 0001, Andrew Zisserman
Image Vis. Comput.2
1998 Self-Calibration of a Rotating Camera with Varying Intrinsic Parameters
abstract
We present a method for self-calibration of a camera which is free to rotate and change its intrinsic parameters, but which cannot translate. The method is based on the so-called infinite homography constraint which leads to a non-linear minimisation routine to find the unknown camera intrinsics over an extended sequence of images. We give experimental results using real image sequences for which ground truth data was available. 1
Lourdes Agapito, Eric Hayman, Ian D. Reid 0001
BMVC3
1998 Real-time Visual Recovery of Pose using Line Tracking in Multiple Cameras
abstract
This paper describes a system for the recovery in real-time of the pose of moving polyhedral objects using modest hardware. A method of line track-ing first introduced by Harris is extended to multiple calibrated cameras, and afforced by robust methods and filtering. The system, which uses three cam-eras a low-end commercial framegrabber and runs on single PC, has been devised to provide simple visual feedback for the tele-operator of a force re-flecting robot manipulator. Experimental results are given which demonstrate the accuracy of the vision system. 1
David William Murray 0001, Ian D. Reid 0001, Richard L. Thompson
BMVC2
1998 3D Trajectories from a Single Viewpoint using Shadows
abstract
We consider the problem of obtaining the 3D trajectory of a ball from a sequence of images taken with a camera which is possibly rotating and zooming (but not translating). Techniques are developed to compute the component of image motion of the ball due to camera rotation and zoom, using optic flow. The 3D location of the ball in each frame of the sequence is then determined using a novel geometric construction which makes use of shadows on the known ground plane in order to compute the vertical projection of the ball onto the ground, and the height of the ball above the ground. 1 Introduction In the absence of any other constraints, the image projections of world points in a single view of a scene are insufficient to compute a 3D reconstruction of the scene. The most obvious way to obtain 3D structure is therefore to consider multiple views separated spatially (and possibly temporally). An alternative to using multiple viewpoints is to enforce physical and geometric constrai...
Ian D. Reid 0001, A. North
BMVC1
1998 Duality, Rigidity and Planar Parallax
Antonio Criminisi, Ian D. Reid 0001, Andrew Zisserman
ECCV (2)2
1998 Transfer of Fixation Using Affine Structure: Extending the Analysis to Stereo
Stuart M. Fairley, Ian D. Reid 0001, David William Murray 0001
Int. J. Comput. Vis.2
1997 A Plane Measuring Device
Antonio Criminisi, Ian D. Reid 0001, Andrew Zisserman
BMVC2
1997 The Active Recovery of 3D Motion Trajectories and Their Use in Prediction
abstract
This paper describes the theory and real-time implementation using an active camera platform of a method of planar trajectory recovery, and of the use of those trajectories to facilitate prediction over delays in the visual feedback loop. Image-based position and velocity demands for tracking are generated by detecting and segmenting optical flow within a central region of the image, and a projective construct is used to map the camera platform's joint angles into a Euclidean coordinate system within a plane, typically the ground plane, in the scene. A set of extended Kalman filters with different dynamics is implemented to analyze the trajectories, and these compete to provide the best description of the motion within an interacting multiple model. Prediction from the optimum motion model is used within the visual feedback loop to overcome visual latency. It is demonstrated that prediction from the 3D planar description gives better tracking performance than prediction based on a filtered description of observer-based 2D motion trajectories.
Kevin J. Bradshaw, Ian D. Reid 0001, David William Murray 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
1996 Zooming while Tracking Using Affine Transfer
abstract
Zoom interacts strongly with both vision and control processes in an active visual system, causing problems for many commonly used tracking methods. This paper demonstrates the use of afne transfer to track while zooming, using clusters of corner features. Afne transfer not only is fundamentally invariant to zoom but also provides a natural mechanism to allow features to appear and disappear while tracking, events which will occur as detail sharpens and dissolves during zooming. The paper demonstrates ofine 3D afne transfer during zoom for objects undergoing substantial rotation, and describes real-time experiments using 2D afne transfer while zooming and tracking using an active camera platform.
Eric Hayman, Ian D. Reid 0001, David William Murray 0001
BMVC2
1996 Steering and Navigation Behaviours Using Fixation
abstract
Steering a motor vehicle around a winding but otherwise uncluttered road has been observed by Land and Lee (1994) to involve repeated periods of visual fixation upon the tangent point of the inside of each bend. We demonstrate a similar use of `active' fixation in the autonomous navigation of a robot vehicle around an obstacle, and show how the control law devised for steering in the robotic example is applicable to the observed human performance data. We discuss the merits of fixation for mobile robot localization.
David William Murray 0001, Ian D. Reid 0001, Andrew J. Davison
BMVC2
1996 Goal-directed Video Metrology
Ian D. Reid 0001, Andrew Zisserman
ECCV (2)1
1996 Active tracking of foveated feature clusters using affine structure
Ian D. Reid 0001, David William Murray 0001
Int. J. Comput. Vis.1
1996 Projective calibration of a laser-stripe range finder
Ian D. Reid 0001
Image Vis. Comput.1
1996 Self-alignment of a binocular robot
Ian D. Reid 0001, Paul A. Beardsley
Image Vis. Comput.1
1995 The Active Camera as a Projective Pointing Device
abstract
This paper demonstrates an approach which exploits an active camera as a projective pointing mechanism. The optical centre of a static camera is notionally substituted by the centre of rotation of the active camera, as is the image plane by a frontal plane, a plane perpendicular to the optical axis of the active camera in its resting direction. Algorithms devised for 3D motion and 3D structure recovery using a single passive camera become immediately applicable to the active camera without need for reformulation. Furthermore, because the active camera can access a panoramic field of view, instabilities which may arise when the field of view is small, or because the shared field of view between successive after movement is small, are lessened. Two quite different applications of the idea are presented. In the first, the homography between a planar surface in the scene and the frontal plane is recovered and used to recover scene trajectories. In the second, the essential matrix between points in two frontal-plane views is recovered and used to determine the motion of a mobile vehicle.
Andrew J. Davison, Ian D. Reid 0001, David William Murray 0001
BMVC2
1995 Self-alignment of a Binocular Robot
Ian D. Reid 0001, Paul A. Beardsley
BMVC1
1995 Active Visual Navigation Using Non-Metric Structure
abstract
Demonstrates a method of using nonmetric visual information derived from an uncalibrated active vision system to navigate an autonomous vehicle through free-space regions detected in a cluttered environment. The structure of 3-space is recovered modulo an affine transformation using an uncalibrated active stereo head carried by the vehicle. The plane at infinity, necessary for recovering affine structure from projective structure, is found in a novel manner by making controlled rotations of the head. The structure is composed of 3D points obtained by detecting and matching image corners through the stereo image sequence. Considerable care has been taken to ensure that the processing is reliable, robust and automatic. Driveable regions are determined from the projection of the affine structure onto a plane parallel to the ground determined using projective constructs. Two methods of negotiating the regions are explored. The first introduces metric information to allow control of a Euclidean vehicle. The second uses visual servoing of the active head to navigate in the affinely described free-space regions.>
Paul A. Beardsley, Ian D. Reid 0001, Andrew Zisserman, David William Murray 0001
ICCV2
1995 Transfer of Fixation for an Active Stereo Platform via Affine Structure Recovery
abstract
This paper describes an algorithm for stereo tracking using 3D affine transfer of a body-centred fixation point. Transfer is based on corners detected in the image and matched over time and in stereo. The paper presents a method of basing the transfer on all the available data, providing immunity to noise and poor conditioning. The paper also shows an implementation at video rates on a four axis active camera platform. Graceful degradation in the presence of insufficient data and fixed latency tracking in parallel with the structure calculation provide robust performance. Recovered trajectories are shown in an approximately Euclidean frame while structure transfer is demonstrated by the evolution of the target's convex hull.>
Stuart M. Fairley, Ian D. Reid 0001, David William Murray 0001
ICCV2
1995 Recognition of Object Classes from Range Data
Ian D. Reid 0001, J. Michael Brady
Artif. Intell.1
1995 Driving saccade to pursuit using image motion
David William Murray 0001, Kevin J. Bradshaw, Philip F. McLauchlan, Ian D. Reid 0001, Paul M. Sharkey
Int. J. Comput. Vis.4
1994 Stereo Fixation using Affne Transfer
abstract
We describe an algorithm which uses affine transfer of the fixation point in a stereo pair to obtain stereo fixation of a moving object. The method runs in real-time using corners tracked temporally and in stereo. The fixation point is transferred to new left and right views using affine structure from four views old left and right and new left and right. We also present the singular value decomposition as a means by which all possible feature points contribute to the transfer, providing immunity to poor choice of basis features. Early results are presented for the method, showing its speed and reliability.
Stuart M. Fairley, Ian D. Reid 0001, David William Murray 0001
BMVC2
1994 Recursive Affine Structure and Motion from Image Sequences
Philip F. McLauchlan, Ian D. Reid 0001, David William Murray 0001
ECCV (1)2
1994 Towards Active Exploration of Static and Dynamic Scene Geometry
abstract
In this paper we describe the use of a steerable camera and active vision system in the real-time recovery of trajectories of objects moving on a planar surface in the scene. The system first self-calibrates the transformation from camera platform joint angles to scene plane coordinates by active observation of static scene geometry which has been stored as a model. The transformation is encoded as a homography. When distracted by motion, the camera saccades to the object and subsequently tracks it, and the 2D Cartesian components of the motion trajectory are recovered within the scene plane.>
Ian D. Reid 0001, David William Murray 0001, Kevin J. Bradshaw
ICRA1
1994 Saccade and pursuit on an active head/eye platform
Kevin J. Bradshaw, Philip F. McLauchlan, Ian D. Reid 0001, David William Murray 0001
Image Vis. Comput.3
1994 Recognition of parameterized objects from 3D data: a parallel implementation
Frédéric Chenavier, Ian D. Reid 0001, J. Michael Brady
Image Vis. Comput.2
1993 Saccade and Pursuit on an Active Head/Eye Platform
abstract
We describe the implementation of, and results from, a realtime active surveillance vision system which detects moving objects in an everyday environment, directs the gaze of a headkye platform towards the objects and subsequently pursues them smoothly. Target detection and pursuit arc performed purely on the basis of image motion, and can continue over extended periods. Two independent parallel processes derive (i) coarse resolution motion over the entire image to direct saccadic shifts in attention over a wide field of view, and (ii) fine resolution motion in a small central region of the image used to perform smooth-pursuit. A gaze controller which selects results from the two visual processes and controls the movement of the head platform is implemented as a finite state machine.
Kevin J. Bradshaw, Philip F. McLauchlan, Ian D. Reid 0001, David William Murray 0001
BMVC3
1993 Reactions to peripheral image motion using a head/eye platform
abstract
The authors demonstrate four real-time reactive responses to movement in everyday scenes using an active head/eye platform. They first describe the design and realization of a high-bandwidth four-degree-of-freedom head/eye platform and visual feedback loop for the exploration of motion processing within active vision. The vision system divides processing into two scales and two broad functions. At a coarse, quasi-peripheral scale, detection and segmentation of new motion occurs across the whole image, and at fine scale, tracking of already detected motion takes place within a foveal region. Several simple coarse scale motion sensors which run concurrently at 25 Hz with latencies around 100 ms are detailed. The use of these sensors are discussed to drive the following real-time responses: (1) head/eye saccades to moving regions of interest; (2) a panic response to looming motion; (3) an opto-kinetic response to continuous motion across the image and (4) smooth pursuit of a moving target using motion alone.>
David William Murray 0001, Philip F. McLauchlan, Ian D. Reid 0001, Paul M. Sharkey
ICCV3
1993 Recognition of object classes from range data
abstract
The authors present techniques for recognizing instances of 3-D object classes from sets of 3-D feature observations. Recognition of a class instance is structured as a search of an interpretation tree in which geometric constraints on pairs of sensed features not only prune the tree, but are used to determine upper and lower bounds on the model parameter values of the instance. A real-valued constraint propagation network unifies the representations of the model parameters, model constraints and feature constraints, and provides a simple and effective mechanism for accessing and updating parameter values. Recognition of objects with multiple internal degrees of freedom, including non-uniform scaling and stretching, articulations, and subpart repetitions, is demonstrated for two different types of real range data: 3-D edge fragments from a stereo vision system, and position/surface normal data derived from planar patches extracted from a range image.>
Ian D. Reid 0001, J. Michael Brady
ICCV1
1993 Tracking foveated corner clusters using affine structure
abstract
The authors describe a novel method of obtaining a fixation point on a moving object for a real-time gaze control system. The method makes use of a real-time implementation of a corner detector and tracker and reconstructs the image position of the desired fixation point from a cluster of corners detected on the object using the affine structure available from two or three views. The method is fast, reliable, viewpoint invariant, and insensitive to occlusion and/or individual corner dropout or reappearance. Results are presented for the method used with a high performance head/eye platform. The results are compared with two naive fixation methods.>
Ian D. Reid 0001, David William Murray 0001
ICCV1
1993 From saccades to smooth pursuit: real-time gaze control using motion feedback
abstract
The authors present an active vision system which performs a surveillance task in everyday dynamic scenes. The system is based around simple, rapid motion processors and a control strategy which uses both position and velocity information. The surveillance task is defined in terms of two separate behavioral subsystems, saccade and smooth pursuit, which are demonstrated individually on the system. It is shown how these and other elementary responses to 2D motion can be built up into behavior sequences, and how judicious close cooperation between vision and control results in smooth transitions between the behaviors. These ideas are demonstrated by an implementation of a saccade to smooth pursuit surveillance system on a high-performance robotic hand/eye platform.
Ian D. Reid 0001, Kevin J. Bradshaw, Philip F. McLauchlan, Paul M. Sharkey, David William Murray 0001
IROS1
1992 Coarse Image Motion for Saccade Control
Philip F. McLauchlan, Ian D. Reid 0001, David William Murray 0001
BMVC2
1992 Model-based recognition and range imaging for a guided vehicle
Ian D. Reid 0001, J. Michael Brady
Image Vis. Comput.1
1991 Recognizing Parameterized Objects Using 3D Edges
Ian D. Reid 0001, J. Michael Brady
BMVC1