EDBT 2026 Demo / reviewers in the wild / expert
Tom Drummond
dblp:50/1633
· DBLP profile ↗
133ranked-venue papers
9as first author
28since 2021 · last 2026
0000-0001-8204-5904ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 102 · 8 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 84 · 5 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 19 · 1 first-author · 4 since 2021Systems, architecture and hardware · 17Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Admitting Ignorance Helps the Video Question Answering Models to AnswerabstractSignificant progress has been made in the field of video question answering (VideoQA) thanks to deep learning and large-scale pretraining. Despite the presence of sophisticated model structures and powerful video-text foundation models, most existing methods focus solely on maximizing the correlation between answers and video-question pairs during training. We argue that these models often establish shortcuts, resulting in spurious correlations between questions and answers, especially when the alignment between video and text data is suboptimal. To address these spurious correlations, we propose a novel training framework in which the model is compelled to acknowledge its ignorance when presented with an intervened question, rather than making guesses solely based on superficial question-answer correlations. We introduce methodologies for intervening in questions, utilizing techniques such as displacement and perturbation, and design frameworks for the model to admit its lack of knowledge in both multi-choice VideoQA and open-ended settings. In practice, we integrate a state-of-the-art model into our framework to validate its effectiveness. The results clearly demonstrate that our framework can significantly enhance the performance of VideoQA models with minimal structural modifications. Haopeng Li 0001, Tom Drummond, Mingming Gong, Mohammed Bennamoun, Qiuhong Ke |
IEEE Trans. Multim. | 2 |
| 2025 | TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion ModelabstractWe introduce TCAM-Diff, a novel 3D medical image generation model that reduces the memory requirements to encode and generate high-resolution 3D data. This model utilizes a decoder-only autoencoder method to learn triplane representation from dense volume and leverages generalization operations to prevent overfitting. Subsequently, it uses a triplane-aware cross-attention diffusion model to learn and integrate these features effectively. Furthermore, the features generated by the diffusion model can be rapidly transformed into 3D volumes using a pre-trained decoder module. Our experiments on three different scales of medical datasets, BrainTumour 128x128x128, Pancreas 256x256x256, and Colon 512x512x512, demonstrated outstanding results. We utilized MSE and SSIM to evaluate reconstruction quality and leveraged the Wasserstein Generative Adversarial Network (W-GAN) critic to assess generative quality. Comparisons to existing approaches show that our method gives better reconstruction and generation results than other encoder-decoder methods with similar-sized latent spaces. Zhenkai Zhang 0001, Krista A. Ehinger, Tom Drummond |
AAAI | 3 |
| 2025 | Sound Judgment: Properties of Consequential Sounds Affecting Human-Perception of RobotsabstractPositive human-perception of robots is critical to achieving sustained use of robots in shared environments. One key factor affecting human-perception of robots are their sounds, especially the consequential sounds which robots (as machines) must produce as they operate. This paper explores qualitative responses from 182 participants to gain insight into human-perception of robot consequential sounds. Participants viewed videos of different robots performing their typical movements, and responded to an online survey regarding their perceptions of robots and the sounds they produce. Topic analysis was used to identify common properties of robot consequential sounds that participants expressed liking, disliking, wanting or wanting to avoid being produced by robots. Alongside expected reports of disliking high pitched and loud sounds, many participants preferred informative and audible sounds (over no sound) to provide predictability of purpose and trajectory of the robot. Rhythmic sounds were preferred over acute or continuous sounds, and many participants wanted more natural sounds (such as wind or cat purrs) in-place of machine-like noise. The results presented in this paper support future research on methods to improve consequential sounds produced by robots by highlighting features of sounds that cause negative perceptions, and providing insights into sound profile changes for improvement of human-perception of robots, thus enhancing human robot interaction. Aimee Allen, Tom Drummond, Dana Kulic |
HRI | 2 |
| 2025 | Few-Shot Multilingual Open-Domain QA from Five ExamplesabstractAbstract Recent approaches to multilingual open- domain question answering (MLODQA) have achieved promising results given abundant language-specific training data. However, the considerable annotation cost limits the application of these methods for underrepresented languages. We introduce a few-shot learning approach to synthesize large-scale multilingual data from large language models (LLMs). Our method begins with large-scale self-supervised pre-training using WikiData, followed by training on high-quality synthetic multilingual data generated by prompting LLMs with few-shot supervision. The final model, FsModQA, significantly outperforms existing few-shot and supervised baselines in MLODQA and cross-lingual and monolingual retrieval. We further show our method can be extended for effective zero-shot adaptation to new languages through a cross-lingual prompting strategy with only English-supervised data, making it a general and applicable solution for MLODQA tasks without costly large-scale annotation. Fan Jiang 0014, Tom Drummond, Trevor Cohn |
Trans. Assoc. Comput. Linguistics | 2 |
| 2024 | Pre-training Cross-lingual Open Domain Question Answering with Large-scale Synthetic SupervisionabstractCross-lingual open domain question answering (CLQA) is a complex problem, comprising cross-lingual retrieval from a multilingual knowledge base, followed by answer generation in the query language.Both steps are usually tackled by separate models, requiring substantial annotated datasets, and typically auxiliary resources, like machine translation systems to bridge between languages.In this paper, we show that CLQA can be addressed using a single encoder-decoder model.To effectively train this model, we propose a selfsupervised method based on exploiting the cross-lingual link structure within Wikipedia.We demonstrate how linked Wikipedia pages can be used to synthesise supervisory signals for cross-lingual retrieval, through a form of cloze query, and generate more natural questions to supervise answer generation.Together, we show our approach, CLASS, outperforms comparable methods on both supervised and zero-shot language adaptation settings, including those using machine translation.𝓒 𝐸𝑛 𝓒 𝐸𝑛 𝓒 𝐸𝑛 𝓒 𝑀𝑢𝑙𝑡𝑖 𝓒 𝑀𝑢𝑙𝑡𝑖 Parallel Sentence Mining Once in I n d i a , hippies went to many different destinations, on the beaches of Goa and Kovalam in Trivandrum (Kerala), or crossed the border into Nepal to spend months in Kathmandu.インドでは、ヒッピーは多くの異な る目的地へいったが、トリヴァンド ラム(ケーララ州)のゴアとコバラ ムのビーチに大量に集まったり、 国境を越えたネパールのカトマン ズで数ヶ月過ごしたりした。 q En : Once in [Mask], hippies went to many different destinations... Fan Jiang 0014, Tom Drummond, Trevor Cohn |
EMNLP | 2 |
| 2024 | Perceiving Longer Sequences With Bi-Directional Cross-Attention TransformersabstractWe present a novel bi-directional Transformer architecture (BiXT) which scales linearly with input size in terms of computational cost and memory consumption, but does not suffer the drop in performance or limitation to only one input modality seen with other efficient Transformer-based approaches. BiXT is inspired by the Perceiver architectures but replaces iterative attention with an efficient bi-directional cross-attention module in which input tokens and latent variables attend to each other simultaneously, leveraging a naturally emerging attention-symmetry between the two. This approach unlocks a key bottleneck experienced by Perceiver-like architectures and enables the processing and interpretation of both semantics ('what') and location ('where') to develop alongside each other over multiple layers -- allowing its direct application to dense and instance-based tasks alike. By combining efficiency with the generality and performance of a full Transformer architecture, BiXT can process longer sequences like point clouds, text or images at higher feature resolutions and achieves competitive performance across a range of tasks like point cloud part segmentation, semantic image segmentation, image classification, hierarchical sequence modeling and document retrieval. Our experiments demonstrate that BiXT models outperform larger competitors by leveraging longer sequences more efficiently on vision tasks like classification and segmentation, and perform on par with full Transformer variants on sequence modeling and document retrieval -- but require 28\% fewer FLOPs and are up to $8.4\times$ faster. Markus Hiller, Krista A. Ehinger, Tom Drummond |
NeurIPS | 3 |
| 2023 | Improving Denoising Diffusion Models via Simultaneous Estimation of Image and Noise
Zhenkai Zhang 0001, Krista A. Ehinger, Tom Drummond |
ACML | 3 |
| 2023 | Knowledge Combination to Learn Rotated Detection without Rotated AnnotationabstractRotated bounding boxes drastically reduce output ambiguity of elongated objects, making it superior to axis-aligned bounding boxes. Despite the effectiveness, rotated detectors are not widely employed. Annotating rotated bounding boxes is such a laborious process that they are not provided in many detection datasets where axis-aligned annotations are used instead. In this paper, we propose a framework that allows the model to predict precise rotated boxes only requiring cheaper axis-aligned annotation of the target dataset1. To achieve this, we leverage the fact that neural networks are capable of learning richer representation of the target domain than what is utilized by the task. The under-utilized representation can be exploited to address a more detailed task. Our framework combines task knowledge of an out-of-domain source dataset with stronger annotation and domain knowledge of the target dataset with weaker annotation. A novel assignment process and projection loss are used to enable the cotraining on the source and target datasets. As a result, the model is able to solve the more detailed task in the target domain, without additional computation overhead during inference. We extensively evaluate the method on various target datasets including fresh-produce dataset, HRSC2016 and SSDD. Results show that the proposed method consistently performs on par with the fully supervised approach. Tianyu Zhu 0001, Bryce Ferenczi, Pulak Purkait, Tom Drummond, Seyed Hamid Rezatofighi, Anton van den Hengel |
CVPR | 4 |
| 2023 | Don't Mess with Mister-in-Between: Improved Negative Search for Knowledge Graph CompletionabstractThe best methods for knowledge graph completion use a 'dual-encoding' framework, a form of neural model with a bottleneck that facilitates fast approximate search over a vast collection of candidates.These approaches are trained using contrastive learning to differentiate between known positive examples and sampled negative instances.The mechanism for sampling negatives to date has been very simple, driven by pragmatic engineering considerations (e.g., using mismatched instances from the same batch).We propose several novel means of finding more informative negatives, based on searching for candidates with high lexical overlaps, from the dual-encoder model and according to knowledge graph structures.Experimental results on four benchmarks show that our best single model improves consistently over previous methods and obtains new state-of-the-art performance, including the challenging large-scale Wikidata5M dataset.Combing different strategies through model ensembling results in a further performance boost. Fan Jiang 0014, Tom Drummond, Trevor Cohn |
EACL | 2 |
| 2023 | Progressive Video Summarization via Multimodal Self-supervised LearningabstractModern video summarization methods are based on deep neural networks that require a large amount of annotated data for training. However, existing datasets for video summarization are small-scale, easily leading to over-fitting of the deep models. Considering that the annotation of large-scale datasets is time-consuming, we propose a multimodal self-supervised learning framework to obtain semantic representations of videos, which benefits the video summarization task. Specifically, the self-supervised learning is conducted by exploring the semantic consistency between the videos and text in both coarse-grained and fine-grained fashions, as well as recovering masked frames in the videos. The multimodal framework is trained on a newly-collected dataset that consists of video-text pairs. Additionally, we introduce a progressive video summarization method, where the important content in a video is pinpointed progressively to generate better summaries. Extensive experiments have proved the effectiveness and superiority of our method in rank correlation coefficients and F-score1. Haopeng Li 0001, Qiuhong Ke, Mingming Gong, Tom Drummond |
WACV | 4 |
| 2023 | Looking Beyond Two Frames: End-to-End Multi-Object Tracking Using Spatial and Temporal TransformersabstractTracking a time-varying indefinite number of objects in a video sequence over time remains a challenge despite recent advances in the field. Most existing approaches are not able to properly handle multi-object tracking challenges such as occlusion, in part because they ignore long-term temporal information. To address these shortcomings, we present MO3TR: a truly end-to-end Transformer-based online multi-object tracking (MOT) framework that learns to handle occlusions, track initiation and termination without the need for an explicit data association module or any heuristics. MO3TR encodes object interactions into long-term temporal embeddings using a combination of spatial and temporal Transformers, and recursively uses the information jointly with the input data to estimate the states of all tracked objects over time. The spatial attention mechanism enables our framework to learn implicit representations between all the objects and the objects to the measurements, while the temporal attention mechanism focuses on specific parts of past information, allowing our approach to resolve occlusions over multiple frames. Our experiments demonstrate the potential of this new approach, achieving results on par with or better than the current state-of-the-art on multiple MOT metrics for several popular multi-object tracking benchmarks. Tianyu Zhu 0001, Markus Hiller, Mahsa Ehsanpour, Rongkai Ma, Tom Drummond, Ian D. Reid 0001, Seyed Hamid Rezatofighi |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Adaptive Poincaré Point to Set Distance for Few-Shot ClassificationabstractLearning and generalizing from limited examples, i.e., few-shot learning, is of core importance to many real-world vision applications. A principal way of achieving few-shot learning is to realize an embedding where samples from different classes are distinctive. Recent studies suggest that embedding via hyperbolic geometry enjoys low distortion for hierarchical and structured data, making it suitable for few-shot learning. In this paper, we propose to learn a context-aware hyperbolic metric to characterize the distance between a point and a set associated with a learned set to set distance. To this end, we formulate the metric as a weighted sum on the tangent bundle of the hyperbolic space and develop a mechanism to obtain the weights adaptively, based on the constellation of the points. This not only makes the metric local but also dependent on the task in hand, meaning that the metric will adapt depending on the samples that it compares. We empirically show that such metric yields robustness in the presence of outliers and achieves a tangible improvement over baseline models. This includes the state-of-the-art results on five popular few-shot classification benchmarks, namely mini-ImageNet, tiered-ImageNet, Caltech-UCSD Birds-200-2011(CUB), CIFAR-FS, and FC100. Rongkai Ma, Pengfei Fang, Tom Drummond, Mehrtash Harandi |
AAAI | 3 |
| 2022 | A Differentiable Distance Approximation for Fairer Image Classification
Nicholas Rosa, Tom Drummond, Mehrtash Harandi |
ACCV (6) | 2 |
| 2022 | Flynet: Max it, Excite it, Quantize it
Luis Guerra, Tom Drummond |
BMVC | 2 |
| 2022 | Implicit Motion Handling for Video Camouflaged Object DetectionabstractWe propose a new video camouflaged object detection (VCOD) framework that can exploit both short-term dynamics and long-term temporal consistency to detect camouflaged objects from video frames. An essential property of camouflaged objects is that they usually exhibit patterns similar to the background and thus make them hard to identify from still images. Therefore, effectively handling temporal dynamics in videos becomes the key for the VCOD task as the camouflaged objects will be noticeable when they move. However, current VCOD methods often leverage homography or optical flows to represent motions, where the detection error may accumulate from both the motion estimation error and the segmentation error. On the other hand, our method unifies motion estimation and object segmentation within a single optimization framework. Specifically, we build a dense correlation volume to implicitly capture motions between neighbouring frames and utilize the final segmentation supervision to optimize the implicit motion estimation and segmentation jointly. Furthermore, to enforce temporal consistency within a video sequence, we jointly utilize a spatio-temporal transformer to refine the short-term predictions. Extensive experiments on VCOD benchmarks demonstrate the architectural effectiveness of our approach. We also provide a large-scale VCOD dataset named MoCA-Mask with pixel-level handcrafted ground-truth masks and construct a comprehensive VCOD bench-mark with previous methods to facilitate research in this direction. Dataset Link: https://xueliancheng.github.io/SLT-Net-project. Xuelian Cheng, Huan Xiong, Deng-Ping Fan, Yiran Zhong, Mehrtash Harandi, Tom Drummond, ZongYuan Ge |
CVPR | 6 |
| 2022 | Learning Instance and Task-Aware Dynamic Kernels for Few-Shot Learning
Rongkai Ma, Pengfei Fang, Gil Avraham, Tianyu Zhu 0001, Tom Drummond, Mehrtash Harandi |
ECCV (20) | 6 |
| 2022 | Training 1-Bit Networks on a Sphere: A Geometric Approach
Luis Guerra, Thalaiyasingam Ajanthan, Gil Avraham, Yan Zou, Tom Drummond |
ICANN (3) | 5 |
| 2022 | Deep Laparoscopic Stereo Matching with Transformers
Xuelian Cheng, Yiran Zhong, Mehrtash Harandi, Tom Drummond, Zhiyong Wang 0001, ZongYuan Ge |
MICCAI (8) | 4 |
| 2022 | On Enforcing Better Conditioned Meta-Learning for Rapid Few-Shot AdaptationabstractInspired by the concept of preconditioning, we propose a novel method to increase adaptation speed for gradient-based meta-learning methods without incurring extra parameters. We demonstrate that recasting the optimisation problem to a non-linear least-squares formulation provides a principled way to actively enforce a well-conditioned parameter space for meta-learning models based on the concepts of the condition number and local curvature. Our comprehensive evaluations show that the proposed method significantly outperforms its unconstrained counterpart especially during initial adaptation steps, while achieving comparable or better overall results on several few-shot classification tasks – creating the possibility of dynamically choosing the number of adaptation steps at inference time. Markus Hiller, Mehrtash Harandi, Tom Drummond |
NeurIPS | 3 |
| 2022 | Rethinking Generalization in Few-Shot ClassificationabstractSingle image-level annotations only correctly describe an often small subset of an image’s content, particularly when complex real-world scenes are depicted. While this might be acceptable in many classification scenarios, it poses a significant challenge for applications where the set of classes differs significantly between training and test time. In this paper, we take a closer look at the implications in the context of few-shot learning. Splitting the input samples into patches and encoding these via the help of Vision Transformers allows us to establish semantic correspondences between local regions across images and independent of their respective class. The most informative patch embeddings for the task at hand are then determined as a function of the support set via online optimization at inference time, additionally providing visual interpretability of ‘what matters most’ in the image. We build on recent advances in unsupervised training of networks via masked image modelling to overcome the lack of fine-grained labels and learn the more general statistical structure of the data while avoiding negative image-level annotation influence, aka supervision collapse. Experimental results show the competitiveness of our approach, achieving new state-of-the-art results on four popular few-shot classification benchmarks for 5-shot and 1-shot scenarios. Markus Hiller, Rongkai Ma, Mehrtash Harandi, Tom Drummond |
NeurIPS | 4 |
| 2022 | Visualizing Robot Intent for Object Handovers with Augmented RealityabstractHumans are highly skilled in communicating their intent for when and where a handover would occur. However, even the state-of-the-art robotic implementations for handovers typically lack of such communication skills. This study investigates visualization of the robot’s internal state and intent for Human-to-Robot Handovers using Augmented Reality. Specifically, we explore the use of visualized 3D models of the object and the robotic gripper to communicate the robot’s estimation of where the object is and the pose in which the robot intends to grasp the object. We tested this design via a user study with 16 participants, in which each participant handed over a cube-shaped object to the robot 12 times. Results show communicating robot intent via augmented reality substantially improves the perceived experience of the users for handovers. Results also indicate that the effectiveness of augmented reality is even more pronounced for the perceived safety and fluency of the interaction when the robot makes errors in localizing the object. Rhys Newbury, Akansel Cosgun, Tysha Crowley-Davis, Wesley P. Chan, Tom Drummond, Elizabeth A. Croft |
RO-MAN | 5 |
| 2021 | Learning Online for Unified Segmentation and Tracking ModelsabstractTracking requires building a discriminative model for the target in the inference stage. An effective way to achieve this is online learning, which can comfortably outperform models that are only trained offline. Recent research shows that visual tracking benefits significantly from the unification of visual tracking and segmentation due to its pixel-level discrimination. However, it imposes a great challenge to perform online learning for such a unified model. A segmentation model cannot easily learn from prior information given in the visual tracking scenario. In this paper, we propose TrackMLP: a novel meta-learning method optimized to learn from only partial information to resolve the imposed challenge. Our model is capable of extensively exploiting limited prior information hence possesses much stronger target-background discriminability than other online learning methods. Empirically, we show that our model achieves state-of-the-art performance and tangible improvement over competing models. Our model achieves improved average overlaps of 66.0%,67.1%, and 68.5% in VOT2019, VOT2018, and VOT2016 datasets, which are 6.4%, 7.3%, and 6.4% higher than our baseline. Code will be made publicly available. Tianyu Zhu 0001, Mehrtash Harandi, Rongkai Ma, Tom Drummond |
IJCNN | 4 |
| 2021 | Relational Subsets Knowledge Distillation for Long-Tailed Retinal Diseases Recognition
Lie Ju, Xin Wang 0094, Lin Wang 0027, Tongliang Liu, Tom Drummond, Dwarikanath Mahapatra, ZongYuan Ge |
MICCAI (8) | 6 |
| 2021 | Seeing Thru Walls: Visualizing Mobile Robots in Augmented RealityabstractWe present an approach for visualizing mobile robots through an Augmented Reality headset when there is no line-of-sight visibility between the robot and the human. Three elements are visualized in Augmented Reality: 1) Robot’s 3D model to indicate its position, 2) An arrow emanating from the robot to indicate its planned movement direction, and 3) A 2D grid to represent the ground plane. We conduct a user study with 18 participants, in which each participant are asked to retrieve objects, one at a time, from stations at the two sides of a T-junction at the end of a hallway where a mobile robot is roaming. The results show that visualizations improved the perceived safety and efficiency of the task and led to participants being more comfortable with the robot within their personal spaces. Furthermore, visualizing the motion intent in addition to the robot model was found to be more effective than visualizing the robot model alone. The proposed system can improve the safety of automated warehouses by increasing the visibility and predictability of robots. Morris Gu, Akansel Cosgun, Wesley P. Chan, Tom Drummond, Elizabeth A. Croft |
RO-MAN | 4 |
| 2021 | Demonstrating Cloth Folding to Robots: Design and Evaluation of a 2D and a 3D User InterfaceabstractAn appropriate user interface to collect human demonstration data for deformable object manipulation has been mostly overlooked in the literature. We present an inter-action design for demonstrating cloth folding to robots. Users choose pick and place points on the cloth and can preview a visualization of a simulated cloth before real-robot execution. Two interfaces are proposed: A 2D display-and-mouse interface where points are placed by clicking on an image of the cloth, and a 3D Augmented Reality interface where the chosen points are placed by hand gestures. We conduct a user study with 18 participants, in which each user completed two sequential folds to achieve a cloth goal shape. Results show that while both interfaces were acceptable, the 3D interface was more suitable for understanding the task, and the 2D interface was suitable for repetition. Results also found that fold previews improve three key metrics: task efficiency, the ability to predict the final shape of the cloth, and overall user satisfaction. Benjamin Waymouth, Akansel Cosgun, Rhys Newbury, Tin Tran, Wesley P. Chan, Tom Drummond, Elizabeth A. Croft |
RO-MAN | 6 |
| 2021 | Driving among Flatmobiles: Bird-Eye-View occupancy grids from a monocular camera for holistic trajectory planningabstractCamera-based end-to-end driving neural networks bring the promise of a low-cost system that maps camera images to driving control commands. These networks are appealing because they replace laborious hand engineered building blocks but their black-box nature makes them difficult to delve in case of failure. Recent works have shown the importance of using an explicit intermediate representation that has the benefits of increasing both the interpretability and the accuracy of networks' decisions. Nonetheless, these camera-based networks reason in camera view where scale is not homogeneous and hence not directly suitable for motion forecasting. In this paper, we introduce a novel monocular camera-only holistic end-to-end trajectory planning network with a Bird-Eye-View (BEV) intermediate representation that comes in the form of binary Occupancy Grid Maps (OGMs). To ease the prediction of OGMs in BEV from camera images, we introduce a novel scheme where the OGMs are first predicted as semantic masks in camera view and then warped in BEV using the homography between the two planes. The key element allowing this transformation to be applied to 3D objects such as vehicles, consists in predicting solely their footprint in camera-view, hence respecting the flat world hypothesis implied by the homography. Abdelhak Loukkal, Yves Grandvalet, Tom Drummond, You Li 0005 |
WACV | 3 |
| 2021 | Improved Training of Generative Adversarial Networks Using Decision ForestsabstractWhilst Generative Adversarial Networks (GANs) have gained a reputation as powerful generative models, they are notoriously difficult to train and suffer from instability in optimisation. Recent methods for tackling this drawback have typically approached it by inducing better behaviour on the discriminator component of the GAN; these include loss function modification, gradient regularisation and weight normalisation to create a discriminator that is well-behaved from a Lipschitz perspective. In this paper, we propose a novel and orthogonal contribution which modifies the architecture of a GAN. Our method embeds the powerful discriminating capabilities inherent in decision forests within the discriminator of a GAN. Empirically, we test the effectiveness of our approach on the CIFAR-10, Oxford Flowers and CUB Birds datasets. We show that our technique is easy to incorporate into existing GAN baselines and offers improvements on Fréchet-Inception Distance (FID) scores by as high as 56.1% over several GAN baselines. Gil Avraham, Tom Drummond |
WACV | 3 |
| 2021 | Leveraging Regular Fundus Images for Training UWF Fundus Diagnosis Models via Adversarial Learning and Pseudo-LabelingabstractRecently, ultra-widefield (UWF) 200° fundus imaging by Optos cameras has gradually been introduced because of its broader insights for detecting more information on the fundus than regular 30° - 60° fundus cameras. Compared with UWF fundus images, regular fundus images contain a large amount of high-quality and well-annotated data. Due to the domain gap, models trained by regular fundus images to recognize UWF fundus images perform poorly. Hence, given that annotating medical data is labor intensive and time consuming, in this paper, we explore how to leverage regular fundus images to improve the limited UWF fundus data and annotations for more efficient training. We propose the use of a modified cycle generative adversarial network (CycleGAN) model to bridge the gap between regular and UWF fundus and generate additional UWF fundus images for training. A consistency regularization term is proposed in the loss of the GAN to improve and regulate the quality of the generated data. Our method does not require that images from the two domains be paired or even that the semantic labels be the same, which provides great convenience for data collection. Furthermore, we show that our method is robust to noise and errors introduced by the generated unlabeled data with the pseudo-labeling technique. We evaluated the effectiveness of our methods on several common fundus diseases and tasks, such as diabetic retinopathy (DR) classification, lesion detection and tessellated fundus segmentation. The experimental results demonstrate that our proposed method simultaneously achieves superior generalizability of the learned representations and performance improvements in multiple tasks. Lie Ju, Xin Wang 0094, C. Paul Bonnington, Tom Drummond, ZongYuan Ge |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Localising In Complex Scenes Using Balanced Adversarial AdaptationabstractDomain adaptation and generative modelling have collectively mitigated the expensive nature of data collection and labelling by leveraging the rich abundance of accurate, labelled data in simulation environments. In this work, we study the performance gap that exists between representations optimised for localisation on simulation environments and the application of such representations in a real-world setting. Our method exploits the shared geometric similarities between simulation and real-world environments whilst maintaining invariance towards visual discrepancies. This is achieved by optimising a representation extractor to project both simulated and real representations into a shared representation space. Our method uses a symmetrical adversarial approach which encourages the representation extractor to conceal the domain that features are extracted from and simultaneously preserves robust attributes between source and target domains that are beneficial for localisation. We evaluate our method by adapting representations optimised for indoor Habitat simulated environments (Matterport3D and Replica) to a real-world indoor environment (Active Vision Dataset), showing that it compares favourably against fully-supervised approaches. Gil Avraham, Tom Drummond |
3DV | 3 |
| 2020 | OpenGAN: Open Set Generative Adversarial Networks
Luke Ditria, Benjamin J. Meyer 0001, Tom Drummond |
ACCV (4) | 3 |
| 2020 | Residual Likelihood Forests
Tom Drummond |
BMVC | 2 |
| 2020 | Reducing the Sim-to-Real Gap for Event Cameras
Timo Stoffregen, Cedric Scheerlinck, Davide Scaramuzza 0001, Tom Drummond, Nick Barnes, Lindsay Kleeman, Robert E. Mahony |
ECCV (27) | 4 |
| 2020 | Supportive Actions for Manipulation in Human-Robot Coworker TeamsabstractThe increasing presence of robots alongside humans, such as in human-robot teams in manufacturing, gives rise to research questions about the kind of behaviors people prefer in their robot counterparts. We term actions that support interaction by reducing future interference with others as supportive robot actions and investigate their utility in a co-located manipulation scenario. We compare two robot modes in a shared table pick-and-place task: (1) Task-oriented: the robot only takes actions to further its task objective and (2) Supportive: the robot sometimes prefers supportive actions to task-oriented ones when they reduce future goal-conflicts. Our experiments in simulation, using a simplified human model, reveal that supportive actions reduce the interference between agents, especially in more difficult tasks, but also cause the robot to take longer to complete the task. We implemented these modes on a physical robot in a user study where a human and a robot perform object placement on a shared table. Our results show that a supportive robot was perceived more favorably as a coworker and also reduced interference with the human in one of two scenarios. However, it also took longer to complete the task highlighting an interesting trade-off between task-efficiency and human-preference that needs to be considered before designing robot behavior for close-proximity manipulation scenarios. Shray Bansal, Rhys Newbury, Wesley P. Chan, Akansel Cosgun, Aimee Allen, Dana Kulic, Tom Drummond, Charles L. Isbell Jr. |
IROS | 7 |
| 2020 | Learning to Take Good Pictures of People with a Robot PhotographerabstractWe present a robotic system capable of navigating autonomously by following a line and taking good quality pictures of people. When a group of people is detected, the robot rotates towards them and then back to line while continuously taking pictures from different angles. Each picture is processed in the cloud where its quality is estimated in a two-stage algorithm. First, features such as the face orientation and likelihood of facial emotions are input to a fully connected neural network to assign a quality score to each face. Second, a representation is extracted by abstracting faces from the image and it is input to a Convolutional Neural Network (CNN) to classify the quality of the overall picture. We collected a dataset in which a picture was labeled as good quality if subjects are well-positioned in the image and oriented towards the camera with a pleasant expression. Our approach detected the quality of pictures with 78.4% accuracy in this dataset and received a better mean user rating (3.71/5) than a heuristic method that uses photographic composition procedures in a study where 97 human judges rated each picture. Statistical analysis against the state-of-the-art verified the quality of the resulting pictures. Rhys Newbury, Akansel Cosgun, Tom Drummond |
IROS | 4 |
| 2020 | Hierarchical Neural Architecture Search for Deep Stereo MatchingabstractTo reduce the human efforts in neural network design, Neural Architecture Search (NAS) has been applied with remarkable success to various high-level vision tasks such as classification and semantic segmentation. The underlying idea for the NAS algorithm is straightforward, namely, to allow the network the ability to choose among a set of operations (\eg convolution with different filter sizes), one is able to find an optimal architecture that is better adapted to the problem at hand. However, so far the success of NAS has not been enjoyed by low-level geometric vision tasks such as stereo matching. This is partly due to the fact that state-of-the-art deep stereo matching networks, designed by humans, are already sheer in size. Directly applying the NAS to such massive structures is computationally prohibitive based on the currently available mainstream computing resources. In this paper, we propose the first \emph{end-to-end} hierarchical NAS framework for deep stereo matching by incorporating task-specific human knowledge into the neural architecture search framework. Specifically, following the gold standard pipeline for deep stereo matching (\ie, feature extraction -- feature volume construction and dense matching), we optimize the architectures of the entire pipeline jointly. Extensive experiments show that our searched network outperforms all state-of-the-art deep stereo matching architectures and is ranked at the top 1 accuracy on KITTI stereo 2012, 2015, and Middlebury benchmarks, as well as the top 1 on SceneFlow dataset with a substantial improvement on the size of the network and the speed of inference. Code available at https://github.com/XuelianCheng/LEAStereo. Xuelian Cheng, Yiran Zhong, Mehrtash Harandi, Yuchao Dai, Xiaojun Chang, Hongdong Li, Tom Drummond, ZongYuan Ge |
NeurIPS | 7 |
| 2020 | Approximate Fisher Information Matrix to Characterize the Training of Deep Neural NetworksabstractIn this paper, we introduce a novel methodology for characterizing the performance of deep learning networks (ResNets and DenseNet) with respect to training convergence and generalization as a function of mini-batch size and learning rate for image classification. This methodology is based on novel measurements derived from the eigenvalues of the approximate Fisher information matrix, which can be efficiently computed even for high capacity deep models. Our proposed measurements can help practitioners to monitor and control the training process (by actively tuning the mini-batch size and learning rate) to allow for good training convergence and generalization. Furthermore, the proposed measurements also allow us to show that it is possible to optimize the training process with a new dynamic sampling training approach that continuously and automatically change the mini-batch size and learning rate during the training process. Finally, we show that the proposed dynamic sampling training approach has a faster training time and a competitive classification accuracy compared to the current state of the art. Zhibin Liao, Tom Drummond, Ian D. Reid 0001, Gustavo Carneiro 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Parallel Optimal Transport GANabstractAlthough Generative Adversarial Networks (GANs) are known for their sharp realism in image generation, they often fail to estimate areas of the data density. This leads to low modal diversity and at times distorted generated samples. These problems essentially arise from poor estimation of the distance metric responsible for training these networks. To address these issues, we introduce an additional regularisation term which performs optimal transport in parallel within a low dimensional representation space. We demonstrate that operating in a low dimension representation of the data distribution benefits from convergence rate gains in estimating the Wasserstein distance, resulting in more stable GAN training. We empirically show that our regulariser achieves a stabilising effect which leads to higher quality of generated samples and increased mode coverage of the given data distribution. Our method achieves significant improvements on the CIFAR-10, Oxford Flowers and CUB Birds datasets over several GAN baselines both qualitatively and quantitatively. Gil Avraham, Tom Drummond |
CVPR | 3 |
| 2019 | EMPNet: Neural Localisation and Mapping Using Embedded Memory PointsabstractContinuously estimating an agent's state space and a representation of its surroundings has proven vital towards full autonomy. A shared common ground among systems which successfully achieve this feat is the integration of previously encountered observations into the current state being estimated. This necessitates the use of a memory module for incorporating previously visited states whilst simultaneously offering an internal representation of the observed environment. In this work we develop a memory module which contains rigidly aligned point-embeddings that represent a coherent scene structure acquired from an RGB-D sequence of observations. The point-embeddings are extracted using modern convolutional neural network architectures, and alignment is performed by computing a dense correspondence matrix between a new observation and the current embeddings residing in the memory module. The whole framework is end-to-end trainable, resulting in a recurrent joint optimisation of the point-embeddings contained in the memory. This process amplifies the shared information across states, providing increased robustness and accuracy. We show significant improvement of our method across a set of experiments performed on the synthetic VIZDoom environment and a real world Active Vision Dataset. Gil Avraham, Thanuja Dharmasiri, Tom Drummond |
ICCV | 4 |
| 2019 | Event-Based Motion Segmentation by Motion CompensationabstractIn contrast to traditional cameras, whose pixels have a common exposure time, event-based cameras are novel bio-inspired sensors whose pixels work independently and asynchronously output intensity changes (called "events"), with microsecond resolution. Since events are caused by the apparent motion of objects, event-based cameras sample visual information based on the scene dynamics and are, therefore, a more natural fit than traditional cameras to acquire motion, especially at high speeds, where traditional cameras suffer from motion blur. However, distinguishing between events caused by different moving objects and by the camera's ego-motion is a challenging task. We present the first per-event segmentation method for splitting a scene into independently moving objects. Our method jointly estimates the event-object associations (i.e., segmentation) and the motion parameters of the objects (or the background) by maximization of an objective function, which builds upon recent results on event-based motion-compensation. We provide a thorough evaluation of our method on a public dataset, outperforming the state-of-the-art by as much as 10%. We also show the first quantitative evaluation of a segmentation algorithm for event cameras, yielding around 90% accuracy at 4 pixels relative displacement. Timo Stoffregen, Guillermo Gallego 0002, Tom Drummond, Lindsay Kleeman, Davide Scaramuzza 0001 |
ICCV | 3 |
| 2019 | Learning Factorized Representations for Open-Set Domain Adaptation
Mahsa Baktash, Masoud Faraki, Tom Drummond, Mathieu Salzmann |
ICLR (Poster) | 3 |
| 2019 | Look No Deeper: Recognizing Places from Opposing Viewpoints under Varying Scene Appearance using Single-View Depth EstimationabstractVisual place recognition (VPR) - the act of recognizing a familiar visual place - becomes difficult when there is extreme environmental appearance change or viewpoint change. Particularly challenging is the scenario where both phenomena occur simultaneously, such as when returning for the first time along a road at night that was previously traversed during the day in the opposite direction. While such problems can be solved with panoramic sensors, humans solve this problem regularly with limited field-of-view vision and without needing to constantly turn around. In this paper, we present a new depth- and temporal-aware visual place recognition system that solves the opposing viewpoint, extreme appearance-change visual place recognition problem. Our system performs sequence-to-single frame matching by extracting depth-filtered keypoints using a state-of-the-art depth estimation pipeline, constructing a keypoint sequence over multiple frames from the reference dataset, and comparing these keypoints to the keypoints extracted from a single query image. We evaluate the system on a challenging benchmark dataset and show that it consistently outperforms state-of-the-art techniques. We also develop a range of diagnostic simulation experiments that characterize the contribution of depth-filtered keypoint sequences with respect to key domain parameters including the degree of appearance change and camera motion. Sourav Garg, V. Madhu Babu, Thanuja Dharmasiri, Stephen Hausler, Niko Sünderhauf, Swagat Kumar, Tom Drummond, Michael Milford |
ICRA | 7 |
| 2019 | The Importance of Metric Learning for Robotic Vision: Open Set Recognition and Active LearningabstractState-of-the-art deep neural network recognition systems are designed for a static and closed world. It is usually assumed that the distribution at test time will be the same as the distribution during training. As a result, classifiers are forced to categorise observations into one out of a set of predefined semantic classes. Robotic problems are dynamic and open world; a robot will likely observe objects that are from outside of the training set distribution. Classifier outputs in robotic applications can lead to real-world robotic action and as such, a practical recognition system should not silently fail by confidently misclassifying novel observations. We show how a deep metric learning classification system can be applied to such open set recognition problems, allowing the classifier to label novel observations as unknown. Further to detecting novel examples, we propose an open set active learning approach that allows a robot to efficiently query a user about unknown observations. Our approach enables a robot to improve its understanding of the true distribution of data in the environment, from a small number of label queries. Experimental results show that our approach significantly outperforms comparable methods in both the open set recognition and active learning problems. Benjamin J. Meyer 0001, Tom Drummond |
ICRA | 2 |
| 2019 | Real-Time Joint Semantic Segmentation and Depth Estimation Using Asymmetric AnnotationsabstractDeployment of deep learning models in robotics as sensory information extractors can be a daunting task to handle, even using generic GPU cards. Here, we address three of its most prominent hurdles, namely, i) the adaptation of a single model to perform multiple tasks at once (in this work, we consider depth estimation and semantic segmentation crucial for acquiring geometric and semantic understanding of the scene), while ii) doing it in real-time, and iii) using asymmetric datasets with uneven numbers of annotations per each modality. To overcome the first two issues, we adapt a recently proposed real-time semantic segmentation network, making changes to further reduce the number of floating point operations. To approach the third issue, we embrace a simple solution based on hard knowledge distillation under the assumption of having access to a powerful `teacher' network. We showcase how our system can be easily extended to handle more tasks, and more datasets, all at once, performing depth estimation and segmentation both indoors and outdoors with a single model. Quantitatively, we achieve results equivalent to (or better than) current state-of-the-art approaches with one forward pass costing just 13ms and 6.5 GFLOPs on 640×480 inputs. This efficiency allows us to directly incorporate the raw predictions of our network into the SemanticFusion framework [1] for dense 3D semantic reconstruction of the scene. Vladimir Nekrasov, Thanuja Dharmasiri, Andrew Spek, Tom Drummond, Chunhua Shen, Ian D. Reid 0001 |
ICRA | 4 |
| 2019 | Adversarial Pulmonary Pathology Translation for Pairwise Chest X-Ray Data Augmentation
Yunyan Xing, ZongYuan Ge, Dwarikanath Mahapatra, Jarrel Seah, Meng Law, Tom Drummond |
MICCAI (6) | 7 |
| 2018 | ENG: End-to-End Neural Geometry for Robust Depth and Pose Estimation Using CNNs
Thanuja Dharmasiri, Andrew Spek, Tom Drummond |
ACCV (1) | 3 |
| 2018 | Traversing Latent Space Using Decision Ferns
Gil Avraham, Tom Drummond |
ACCV (1) | 3 |
| 2018 | Efficient Subpixel Refinement With Symbolic Linear PredictorsabstractWe present an efficient subpixel refinement method using a learning-based approach called Linear Predictors. Two key ideas are shown in this paper. Firstly, we present a novel technique, called Symbolic Linear Predictors, which makes the learning step efficient for subpixel refinement. This makes our approach feasible for online applications without compromising accuracy, while taking advantage of the run-time efficiency of learning based approaches. Secondly, we show how Linear Predictors can be used to predict the expected alignment error, allowing us to use only the best keypoints in resource constrained applications. We show the efficiency and accuracy of our method through extensive experiments. Vincent Lui, Jonathon Geeves, Winston Yii, Tom Drummond |
CVPR | 4 |
| 2018 | Deep Metric Learning and Image Classification with Nearest Neighbour Gaussian KernelsabstractWe present a Gaussian kernel loss function and training algorithm for convolutional neural networks that can be directly applied to both distance metric learning and image classification problems. Our method treats all training features from a deep neural network as Gaussian kernel centres and computes loss by summing the influence of a feature's nearby centres in the feature embedding space. Our approach is made scalable by treating it as an approximate nearest neighbour search problem. We show how to make end-to-end learning feasible, resulting in a well formed embedding space, in which semantically related instances are likely to be located near one another, regardless of whether or not the network was trained on those classes. Our approach outperforms state-of-the-art deep metric learning approaches on embedding learning challenges, as well as conventional softmax classification on several datasets. Benjamin J. Meyer 0001, Ben Harwood, Tom Drummond |
ICIP | 3 |
| 2018 | Just-in-Time Reconstruction: Inpainting Sparse Maps Using Single View Depth Predictors as PriorsabstractWe present “just-in-time reconstruction” as realtime image-guided inpainting of a map with arbitrary scale and sparsity to generate a fully dense depth map for the image. In particular, our goal is to inpaint a sparse map - obtained from either a monocular visual SLAM system or a sparse sensor - using a single-view depth prediction network as a virtual depth sensor. We adopt a fairly standard approach to data fusion, to produce a fused depth map by performing inference over a novel fully-connected Conditional Random Field (CRF) which is parameterized by the input depth maps and their pixel-wise confidence weights. Crucially, we obtain the confidence weights that parameterize the CRF model in a data-dependent manner via Convolutional Neural Networks (CNNs) which are trained to model the conditional depth error distributions given each source of input depth map and the associated RGB image. Our CRF model penalises absolute depth error in its nodes and pairwise scale-invariant depth error in its edges, and the confidence-based fusion minimizes the impact of outlier input depth values on the fused result. We demonstrate the flexibility of our method by real-time inpainting of ORB-SLAM, Kinect, and LIDAR depth maps acquired both indoors and outdoors at arbitrary scale and varied amount of irregular sparsity. Chamara Saroj Weerasekera, Thanuja Dharmasiri, Ravi Garg, Tom Drummond, Ian D. Reid 0001 |
ICRA | 4 |
| 2018 | CReaM: Condensed Real-time Models for Depth Prediction using Convolutional Neural NetworksabstractSince the resurgence of CNNs the robotic vision community has developed a range of algorithms that perform classification, semantic segmentation and structure prediction (depths, normals, surface curvature) using neural networks. While some of these models achieve state-of-the art results and super human level performance, deploying these models in a time critical robotic environment remains an ongoing challenge. Real-time frameworks are of paramount importance to build a robotic society where humans and robots integrate seamlessly. To this end, we present a novel real-time structure prediction framework that predicts depth at 30 frames per second on an NVIDIA-TX2. At the time of writing, this is the first piece of work to showcase such a capability on a mobile platform. We also demonstrate with extensive experiments that neural networks with very large model capacities can be leveraged in order to train accurate condensed model architectures in a “from teacher to student” style knowledge transfer. Andrew Spek, Thanuja Dharmasiri, Tom Drummond |
IROS | 3 |
| 2017 | A Compact Parametric Solution to Depth Sensor Calibration
Andrew Spek, Tom Drummond |
BMVC | 2 |
| 2017 | FPGA acceleration of multilevel ORB feature extraction for computer visionabstractIn this paper, we present the first multilevel implementation of the Harris-Stephens corner detector and the ORB feature extractor running on FPGA hardware, for computer vision and robotics applications. ORB is a fundamental component of many robotics applications, and requires significant computation. The design has been validated both in behavioural simulation and in implementation on an Arria V FPGA connected to a desktop PC via PCI-Express. A Linux kernel-mode driver and userspace library allow integration of the acceleration hardware into C++ programs. The device has significantly higher throughput than a CPU implementation (150 MPixel/s vs 27 MPixel/s) and a GPU implementation (40 MPixel/s), with much lower power draw (5.3 W vs 145 W). This throughput is equivalent to 72 fps at 1920 × 1080 or 488 fps at 640 × 480. Josh Weberruss, Lindsay Kleeman, David Boland, Tom Drummond |
FPL | 4 |
| 2017 | Smart Mining for Deep Metric LearningabstractTo solve deep metric learning problems and producing feature embeddings, current methodologies will commonly use a triplet model to minimise the relative distance between samples from the same class and maximise the relative distance between samples from different classes. Though successful, the training convergence of this triplet model can be compromised by the fact that the vast majority of the training samples will produce gradients with magnitudes that are close to zero. This issue has motivated the development of methods that explore the global structure of the embedding and other methods that explore hard negative/positive mining. The effectiveness of such mining methods is often associated with intractable computational requirements. In this paper, we propose a novel deep metric learning method that combines the triplet model and the global structure of the embedding space. We rely on a smart mining procedure that produces effective training samples for a low computational cost. In addition, we propose an adaptive controller that automatically adjusts the smart mining hyper-parameters and speeds up the convergence of the training process. We show empirically that our proposed method allows for fast and more accurate training of triplet ConvNets than other competing mining methods. Additionally, we show that our method achieves new state-of-the-art embedding results for CUB-200-2011 and Cars196 datasets. Ben Harwood, Gustavo Carneiro 0001, Ian D. Reid 0001, Tom Drummond |
ICCV | 5 |
| 2017 | Improved semantic segmentation for robotic applications with hierarchical conditional random fieldsabstractConventional approaches to semantic segmentation are inappropriate for robotic applications, as they focus on pixel-level performance and give little significance to spurious object detections. This paper presents a region-based conditional random field model for semantic segmentation that focuses on object-level performance, recognising that in a robotics context, false object detections can have costly consequences. We show how optimising at the semantic region-level results in significantly fewer false positive object detections than conventional approaches. We further show how both object and pixel-level performance can be improved over conventional methods by combining region random fields with dense pixel random fields in a hierarchical manner. An object-aware performance metric is introduced that heavily penalises false positive and false negative object detections, as appropriate for robotic applications. Our approach is evaluated on the challenging NYU v2 and Pascal VOC datasets, outperforming comparable conventional methods in terms of object and pixel-level performance. Benjamin J. Meyer 0001, Tom Drummond |
ICRA | 2 |
| 2017 | Joint pose and principal curvature refinement using quadricsabstractIn this paper we present a novel joint approach for optimising surface curvature and pose alignment. We present two implementations of this joint optimisation strategy, including a fast implementation that uses two frames and an offline multi-frame approach. We demonstrate an order of magnitude improvement in simulation over state of the art dense relative point-to-plane Iterative Closest Point (ICP) pose alignment using our dense joint frame-to-frame approach and show comparable pose drift to dense point-to-plane ICP bundle adjustment using low-cost depth sensors. Additionally our improved joint quadric based approach can be used to more accurately estimate surface curvature on noisy point clouds than previous approaches. Andrew Spek, Tom Drummond |
ICRA | 2 |
| 2017 | Joint prediction of depths, normals and surface curvature from RGB images using CNNsabstractUnderstanding the 3D structure of a scene is of vital importance, when it comes to developing fully autonomous robots. To this end, we present a novel deep learning based framework that estimates depth, surface normals and surface curvature by only using a single RGB image. To the best of our knowledge this is the first work to estimate surface curvature from colour using a machine learning approach. Additionally, we demonstrate that by tuning the network to infer well designed features, such as surface curvature, we can achieve improved performance at estimating depth and normals. This indicates that network guidance is still a useful aspect of designing and training a neural network. We run extensive experiments where the network is trained to infer different tasks while the model capacity is kept constant resulting in different feature maps based on the tasks at hand. We outperform the previous state-of-the-art benchmarks which jointly estimate depths and surface normals while predicting surface curvature in parallel. Thanuja Dharmasiri, Andrew Spek, Tom Drummond |
IROS | 3 |
| 2017 | Solving Robust Regularization Problems Using Iteratively Re-weighted Least SquaresabstractMany computer vision problems are formulated as an objective function consisting of a sum of functions. In the case of ill-constrained problems, regularization terms are included in the objective function to reduce the ambiguity and noise in the solution. The most commonly used regularization terms are the L2norm and the L1norm. Since the last two decades, the class of regularized problems, especially the L1-regularized problems, has received much attention but still many regularized problems are either difficult to solve, or require complex optimization techniques. We propose a method based on an Iteratively Re-weighted Least Squares approach to minimize an objective function comprising a mixture of m-estimator regularization terms. In addition to the proof of convergence of the algorithm to the desired minimum, we show the applicability of the proposed algorithm by solving the problems of edge-preserved image denoising and image super-resolution. In both the cases, our experimental results show that the proposed algorithm gives superior results to the state-of-art regularization methods. Khurrum Aftab Kiani, Tom Drummond |
WACV | 2 |
| 2016 | FANNG: Fast Approximate Nearest Neighbour GraphsabstractWe present a new method for approximate nearest neighbour search on large datasets of high dimensional feature vectors, such as SIFT or GIST descriptors. Our approach constructs a directed graph that can be efficiently explored for nearest neighbour queries. Each vertex in this graph represents a feature vector from the dataset being searched. The directed edges are computed by exploiting the fact that, for these datasets, the intrinsic dimensionality of the local manifold-like structure formed by the elements of the dataset is significantly lower than the embedding space. We also provide an efficient search algorithm that uses this graph to rapidly find the nearest neighbour to a query with high probability. We show how the method can be adapted to give a strong guarantee of 100% recall where the query is within a threshold distance of its nearest neighbour. We demonstrate that our method is significantly more efficient than existing state of the art methods. In particular, our GPU implementation can deliver 90% recall for queries on a data set of 1 million SIFT descriptors at a rate of over 1.2 million queries per second on a Titan X. Finally we also demonstrate how our method scales to datasets of 5M and 20M entries. Ben Harwood, Tom Drummond |
CVPR | 2 |
| 2016 | A distributed robotic vision serviceabstractRobotic vision is limited by line of sight and on-board camera capabilities. Robots can acquire video or images from remote cameras, but processing additional data has a computational burden. This paper applies the Distributed Robotic Vision Service, DRVS, to robot path planning using data outside line-of-sight of the robot. DRVS implements a distributed visual object detection service to distributes the computation to remote camera nodes with processing capabilities. Robots request task-specific object detection from DRVS by specifying a geographic region of interest and object type. The remote camera nodes perform the visual processing and send the high-level object information to the robot. Additionally, DRVS relieves robots of sensor discovery by dynamically distributing object detection requests to remote camera nodes. Tested over two different indoor path planning tasks DRVS showed dramatic reduction in mobile robot compute load and wireless network utilization. William Chamberlain, Jürgen Leitner, Tom Drummond, Peter I. Corke |
ICRA | 3 |
| 2016 | MO-SLAM: Multi object SLAM with run-time object discovery through duplicatesabstractIn this paper, we present MO-SLAM, a novel visual SLAM system that is capable of detecting duplicate objects in the scene during run-time without requiring an offline training stage to pre-populate a database of objects. Instead, we propose a novel method to detect landmarks that belong to duplicate objects. Further, we show how landmarks belonging to duplicate objects can be converted to first-order entities which generate additional constraints for optimizing the map. We evaluate the performance of MO-SLAM with extensive experiments on both synthetic and real data, where the experimental results verify the capabilities of MO-SLAM in detecting duplicate objects and using these constraints to improve the accuracy of the map. Thanuja Dharmasiri, Vincent Lui, Tom Drummond |
IROS | 3 |
| 2016 | Fast Depth Video Compression for Mobile RGB-D SensorsabstractWe propose a new method, called 3-D image warping-based depth video compression (IW-DVC), for fast and efficient compression of depth images captured by mobile RGB-D sensors. The emergence of low-cost RGB-D sensors has created opportunities to find new solutions for a number of computer vision and networked robotics problems, such as 3-D map building, immersive telepresence, or remote sensing. However, efficient transmission and storage of depth data still presents a challenging task to the research community in these applications. Image/video compression has been comprehensively studied and several methods have already been developed. However, these methods result in unacceptably suboptimal outcomes when applied to the depth images. We have designed the IW-DVC method to exploit the special properties of the depth data to achieve a high compression ratio while preserving the quality of the captured depth images. Our solution combines the egomotion estimation and 3-D image warping techniques and includes a lossless coding scheme that is capable of adapting to depth data with a high dynamic range. IW-DVC operates at a high speed, suitable for real-time applications, and is able to attain an enhanced motion compensation accuracy compared with the conventional approaches. In addition, it removes the existing redundant information between the depth frames to further increase compression efficiency. Our experiments show that IW-DVC attains a very high performance yielding significant compression ratios without sacrificing image quality. Y. Ahmet Sekercioglu, Tom Drummond, Enrico Natalizio, Isabelle Fantoni, Vincent Frémont |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Fast Inverse Compositional Image Alignment with Missing Data and Re-weightingabstractThis paper proposes a novel method of performing inverse compositional image alignment which elegantly deals with missing data and re-weighting, and does not require the Jacobians and Hessian to be re-computed at every iteration. We show how missing data and re-weighting can be handled through preconditioning. We propose a few preconditioning techniques and analyse how each technique models the effects of missing data and re-weighting for inverse composition. We show through extensive experiments on different applications that our method improves the convergence rate of the conventional re-weighted inverse compositional method while remaining robust to outliers. We also show that the the update parameters are usually underestimated and how this can be used to further speed up convergence of image alignment methods. Vincent Lui, Dinesh Gamage, Tom Drummond |
BMVC | 3 |
| 2015 | Self-calibration in visual sensor networks equipped with RGB-D camerasabstractWe consider the self-calibration problem (estimation of location and orientation of multiple camera sensors), in visual sensor networks equipped with RGB-D cameras. We propose two algorithms based on feature matching and relative pose estimation. First one uses Floyd-Warshall algorithm, and can accurately estimate the camera locations. On the other hand, the second algorithm is more scalable than the first one, and can be used in large networks if high accuracy is not an issue. Numerical results demonstrate that both algorithms, depending on the network topology, can be used for self-calibration in RGB-D equipped visual sensor networks. Y. Ahmet Sekercioglu, Tom Drummond |
ICASSP | 3 |
| 2015 | Image based optimisation without global consistency for constant time monocular visual SLAMabstractThis paper presents a monocular visual SLAM system that does not require a globally consistent 3D model. Instead of generating a globally consistent 3D model and localising the camera from the 3D model, the system merely optimises relative pose parameters for pairs of keyframes that overlap on the scene, providing accurate local information at the expense of global consistency. During run-time, the camera is localised using only 2D measurements from nearby keyframes instead of using correspondences between 2D measurements and 3D features of a 3D model. Extensive experiments using both synthetic and real data sets were performed to evaluate the system's performance. Results show that our system is accurate and runs in real time at an average of 25 frames per second on a standard computer. Finally, we also show how useful applications can be easily developed on top of a framework without global consistency. Vincent Lui, Tom Drummond |
ICRA | 2 |
| 2015 | Monocular image space tracking on a computationally limited MAVabstractWe propose a method of monocular camera-inertial based navigation for computationally limited micro air vehicles (MAVs). Our approach is derived from the recent development of parallel tracking and mapping algorithms, but unlike previous results, we show how the tracking and mapping processes operate using different representations. The separation of representations allows us not only to move the computational load of full map inference to a ground station, but to further reduce the computational cost of on-board tracking for pose estimation. Our primary contribution is to show how the cost of tracking the vehicle pose on-board can be substantially reduced by estimating the camera motion directly in the image frame, rather than in the world co-ordinate frame. We demonstrate our method on an Ascending Technologies Pelican quad-rotor, and show that we can track the vehicle pose with reduced on-board computation but without compromised navigation accuracy. Kyel Ok, Dinesh Gamage, Tom Drummond, Frank Dellaert, Nicholas Roy |
ICRA | 3 |
| 2015 | Reduced dimensionality extended Kalman Filter for SLAM in a relative formulationabstractModern approaches to monocular SLAM retain all observations and repeatedly perform bundle adjustment in order to overcome the inconsistency problem that arises in approaches that marginalise out camera positions. Bundle adjustment is inherently an expensive operation and so sparse matrix techniques and double window optimisation on sparsely sampled key-frames are employed to minimize the computational cost. Dinesh Gamage, Tom Drummond |
IROS | 2 |
| 2014 | Vision-based robot-assisted biological cell micromanipulationabstractThis paper presents a modular design for a robot-assisted biological cell microinjection system. The proposed design is composed of injection, vision, force measurement and control units that provides sufficient flexibility to observe and control cell microinjection by monitoring and regulating position and force simultaneously. Methodologies have been presented for automation of the laborious tasks associated with microinjection to improve the repeatability and reliability of the process. The system is capable of automatic positioning and focusing of the microcapillary tip as well as automatic realization of the cell piercing during the microinjection process with vision-based approaches. The proposed methods were tested for 100 zebrafish embryos micromanipulation experiments at Blastula stage. 97% success rate was achieved which shows high capability of the proposed design and methods. Fatemeh Karimirad, Sunita Chauhan, Bijan Shirinzadeh, Tom Drummond, Saeid Nahavandi |
RO-MAN | 4 |
| 2013 | Reduced Dimensionality Extended Kalman Filter for SLAMabstractComputational complexity of the Kalman filter grows at least quadratically with the number of dimensions in the filter. This is a particular problem for applications like monocular simultaneous localization and mapping (SLAM) where it is not possible to run a single filter on a large map with many thousands of landmarks. The filtering approach for SLAM, maintains only the current camera pose with all landmarks of interest as the state [2]. This paper presents a method for reducing the computational complexity of the Kalman filters by reducing the dimensionality as information is acquired. The method reduces the dimensionality of the extended Kalman filter (EKF) for SLAM by identifying dominant modes of the filter. This can be used in general to reduce the dimensionality of the EKF irrespectively of its application, without being limited to SLAM. For a filter with zero process noise, the mean of this distribution can be represented as a point in a nD space with its uncertainty as a hyper ellipse. Further information will reduce this uncertainty along some directions. After some time uncertainty along many directions becomes comparatively small, making further information redundant. So the covariance matrix can be decompose into certain and uncertain dimensions. Dinesh Gamage, Tom Drummond |
BMVC | 2 |
| 2013 | An Iterative 5-pt Algorithm for Fast and Robust Essential Matrix EstimationabstractThe essential matrix, first introduced by Longuet-Higgins [5], is a 3× 3 matrix encoding the relative pose information between two views. Conventional approaches for relative pose estimation is to solve a system of linear equations. The 5-pt algorithm [6] is the current state-of-the-art algorithm in relative pose estimation. It is a minimal-set direct solver which solves the essential matrix as a system of polynomial equations. We show in this paper an iterative method which provides robust and real-time essential matrix estimation, capable of 30Hz, frame-rate performance. The benefit of an iterative approach lies in its simplicity and speed. The use of high degree polynomials may lead to ill-conditioning [1] and are often difficult to solve, leading to alternative methods which sacrifice speed for simplicity [1, 3, 4]. Although convergence is not guaranteed, when used within RANSAC, more hypotheses can be evaluated in the same block of time, yielding improved performance. While iterative solvers which minimizes the algebraic epipolar reprojection error have previously been proposed [2, 7], our parametrization is based on a novel geometric error which incorporates the half-plane constraint [8] and thus enforces orientation consistency between points. Figure 1 illustrates the concept of our iterative 5-pt algorithm. A coordinate frame is chosen such that the z-axis ez joins the two camera centres. In this frame, vectors vi and vi and ez are coplanar. The goal is to find a rotation for each of the two cameras that maps from their internal coordinate frame to that of Figure 1. We parametrize the image for each camera as a unit 2-sphere, mapping image points in normalized camera coordinates [x,y,1]T to unit vectors by dividing by √ x2 + y2 +1. At each iteration, the normalized point correspondences ui ↔ ui are left multiplied with the rotations R and R′, giving the rotated point correspondences v = Ru, v′ = R′u′. This rotates the two unit spheres, changing the direction of the epipoles. By rotating the unit spheres such that the z-axis ez is aligned with the epipoles, i.e. ez = Re = R′e′, vi↔ vi become coplanar with the epipoles e,e′. The epipoles can then be computed as e = RT ez, e′ = R′T ez. (1) Vincent Lui, Tom Drummond |
BMVC | 2 |
| 2013 | A Unified Rolling Shutter and Motion Blur Model for 3D Visual RegistrationabstractMotion blur and rolling shutter deformations both inhibit visual motion registration, whether it be due to a moving sensor or a moving target. Whilst both deformations exist simultaneously, no models have been proposed to handle them together. Furthermore, neither deformation has been considered previously in the context of monocular full-image 6 degrees of freedom registration or RGB-D structure and motion. As will be shown, rolling shutter deformation is observed when a camera moves faster than a single pixel in parallax between subsequent scan-lines. Blur is a function of the pixel exposure time and the motion vector. In this paper a complete dense 3D registration model will be derived to account for both motion blur and rolling shutter deformations simultaneously. Various approaches will be compared with respect to ground truth and live real-time performance will be demonstrated for complex scenarios where both blur and shutter deformations are dominant. Maxime Meilland, Tom Drummond, Andrew I. Comport |
ICCV | 2 |
| 2013 | Tutorial chairsabstractIt is our great pleasure to present the ISMAR Tutorials. We proudly host two tutorials that provide sharing of knowledge from seasoned researchers. Our tutorials cover open standards and development using HTML and instantAR. Through these exciting tutorials we hope to expand the minds of ISMAR 2013 attendees and help to foster the next generation of Mixed and Augmented Reality researchers, practitioners, and artists. Tom Drummond, Matt Adcock |
ISMAR | 1 |
| 2013 | Algorithmic methodologies for FPGA-based vision
Yoong Kang Lim, Lindsay Kleeman, Tom Drummond |
Mach. Vis. Appl. | 3 |
| 2012 | Corner Matching Refinement for Monocular Pose EstimationabstractMany tasks in computer vision rely on accurate detection and matching of visual landmarks (e.g. image corners) between two images. In particular, for the calculation of epipolar geometry from a minimal set of five correspondences the spatial accuracy of matched landmarks is critical because the result is very sensitive to errors. The most common way of improving the accuracy is to calculate a sub-pixel location independently for each landmark in the hope that this reduces the re-projection error of the point in space to which they refer. This paper presents a method for refining the coordinates of correspondences directly. Thus given some coordinates in the first image, our goal is to maximise the accuracy of the estimate of the coordinates in second image corresponding to the same real world point without being too concerned about which real world point is being matched. We show how this can be achieved as a frequency domain optimisation between two image patches to refine the correspondence by estimating affine parameters. We select the correct frequency range for optimisation by identifying a direct relationship between the Gabor phase based approach and the frequency response of a patch. Further, we show how parametric estimation can be made accurate by operating in the frequency domain. Finally, we present experiments which demonstrate the accuracy of this approach, its robustness to changes in scale and orientation and its superior performance by comparison to other sub-pixel methods. Dinesh Gamage, Tom Drummond |
BMVC | 2 |
| 2012 | Robust egomotion estimation using ICP in inverse depth coordinatesabstractThis paper presents a 6 degrees of freedom egomotion estimation method using Iterative Closest Point (ICP) for low cost and low accuracy range cameras such as the Microsoft Kinect. Instead of Euclidean coordinates, the method uses inverse depth coordinates which better conforms to the error characteristics of raw sensor data. Novel inverse depth formulations of point-to-point and point-to-plane error metrics are derived as part of our implementation. The implemented system runs in real time at an average of 28 frames per second (fps) on a standard computer. Extensive experiments were performed to evaluate different combinations of error metrics and parameters. Results show that our system is accurate and robust across a variety of motion trajectories. The point-to-plane error metric was found to be the best at coping with large inter-frame motion while remaining accurate and maintaining real time performance. Wen Lik Dennis Lui, Titus Jia Jie Tang, Tom Drummond, Wai Ho Li |
ICRA | 3 |
| 2012 | Distributed visual processing for augmented realityabstractRecent advances have made augmented reality on smartphones possible but these applications are still constrained by the limited computational power available. This paper presents a system which combines smartphones with networked infrastructure and fixed sensors and shows how these elements can be combined to deliver real-time augmented reality. A key feature of this framework is the asymmetric nature of the distributed computing environment. Smartphones have high bandwidth video cameras but limited computational ability. Our system connects multiple smartphones through relatively low bandwidth network links to a server with large computational resources connected to fixed sensors that observe the environment. By contrast to other systems that use preprocessed static models or markers, our system has the ability to rapidly build dynamic models of the environment on the fly at frame rate. We achieve this by processing data from a Microsoft Kinect, to build a trackable point cloud model of each frame. The smartphones process their video camera data on-board to extract their own set of compact and efficient feature descriptors which are sent via WiFi to a server. The server runs computationally intensive algorithms including feature matching, pose estimation and occlusion testing for each smartphone. Our system demonstrates real-time performance for two smartphones. Winston Yii, Wai Ho Li, Tom Drummond |
ISMAR | 3 |
| 2012 | Rapidly constructed appearance models for tracking in augmented reality applications
Jeremiah Neubert, John Pretlove, Tom Drummond |
Mach. Vis. Appl. | 3 |
| 2011 | Transformative reality: Augmented reality for visual prosthesesabstractVisual prostheses such as retinal implants provide bionic vision that is limited in spatial and intensity resolution. This limitation is a fundamental challenge of bionic vision as it severely truncates salient visual information. We propose to address this challenge by performing real time transformations of visual and non-visual sensor data into symbolic representations that are then rendered as low resolution vision; a concept we call Transformative Reality. For example, a depth camera allows the detection of empty ground in cluttered environments that is then visually rendered as bionic vision to enable indoor navigation. Such symbolic representations are similar to virtual content overlays used in Augmented Reality but are registered to the 3D world via the user's sense of touch. Preliminary user trials, where a head mounted display artificially constrains vision to a 25×25 grid of binary dots, suggest that Transformative Reality provides practical and significant improvements over traditional bionic vision in tasks such as indoor navigation, object localisation and people detection. Wen Lik Dennis Lui, Damien Browne, Lindsay Kleeman, Tom Drummond, Wai Ho Li |
ISMAR | 4 |
| 2011 | Rapid scene reconstruction on mobile phones from panoramic imagesabstractRapid 3D reconstruction of environments has become an active research topic due to the importance of 3D models in a huge number of applications, be it in Augmented Reality (AR), architecture or other commercial areas. In this paper we present a novel system that allows for the generation of a coarse 3D model of the environment within several seconds on mobile smartphones. By using a very fast and flexible algorithm a set of panoramic images is captured to form the basis of wide field-of-view images required for reliable and robust reconstruction. A cheap on-line space carving approach based on Delaunay triangulation is employed to obtain dense, polygonal, textured representations. The use of an intuitive method to capture these images, as well as the efficiency of the reconstruction approach allows for an application on recent mobile phone hardware, giving visually pleasing results almost instantly. Qi Pan, Clemens Arth, Edward Rosten, Gerhard Reitmayr, Tom Drummond |
ISMAR | 5 |
| 2011 | Binary Histogrammed Intensity Patches for Efficient and Robust Matching
Simon Taylor, Tom Drummond |
Int. J. Comput. Vis. | 2 |
| 2010 | Deterministic Sample Consensus with Multiple Match HypothesesabstractRANSAC (Random Sample Consensus) is a popular and effective technique for estimating model parameters in the presence of outliers. Efficient algorithms are necessary for both frame-rate vision tasks and offline tasks with difficult data. We present a deterministic scheme for selecting samples to generate hypotheses, applied to data from feature matching. This method combines matching scores, ambiguity and past performance of hypotheses generated by the matches to estimate the probability that a match is correct. At every stage the best matches are chosen to generate a hypothesis. This method will therefore only spend time on bad matches when the best ones have proven themselves to be unsuitable. The result is a system that is able to operate very efficiently on ambiguous data and is suitable for implementation on devices with limited computing resources. 1 Paul McIlroy, Edward Rosten, Simon Taylor, Tom Drummond |
BMVC | 4 |
| 2010 | Faster and Better: A Machine Learning Approach to Corner DetectionabstractThe repeatability and efficiency of a corner detector determines how likely it is to be useful in a real-world application. The repeatability is important because the same scene viewed from different positions should yield features which correspond to the same real-world 3D locations. The efficiency is important because this determines whether the detector combined with further processing can operate at frame rate. Three advances are described in this paper. First, we present a new heuristic for feature detection and, using machine learning, we derive a feature detector from this which can fully process live PAL video using less than 5 percent of the available processing time. By comparison, most other detectors cannot even operate at frame rate (Harris detector 115 percent, SIFT 195 percent). Second, we generalize the detector, allowing it to be optimized for repeatability, with little loss of efficiency. Third, we carry out a rigorous comparison of corner detectors based on the above repeatability criterion applied to 3D scenes. We show that, despite being principally constructed for speed, on these stringent tests, our heuristic detector significantly outperforms existing feature detectors. Finally, the comparison demonstrates that using machine learning produces significant improvements in repeatability, yielding a detector that is both very fast and of very high quality. Edward Rosten, Reid B. Porter, Tom Drummond |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2010 | Real-Time Detection and Tracking for Augmented Reality on Mobile PhonesabstractIn this paper, we present three techniques for 6DOF natural feature tracking in real time on mobile phones. We achieve interactive frame rates of up to 30 Hz for natural feature tracking from textured planar targets on current generation phones. We use an approach based on heavily modified state-of-the-art feature descriptors, namely SIFT and Ferns plus a template-matching-based tracker. While SIFT is known to be a strong, but computationally expensive feature descriptor, Ferns classification is fast, but requires large amounts of memory. This renders both original designs unsuitable for mobile phones. We give detailed descriptions on how we modified both approaches to make them suitable for mobile phones. The template-based tracker further increases the performance and robustness of the SIFT- and Ferns-based approaches. We present evaluations on robustness and performance and discuss their appropriateness for Augmented Reality applications. Daniel Wagner 0003, Gerhard Reitmayr, Alessandro Mulloni, Tom Drummond, Dieter Schmalstieg |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2009 | Reconstruction from Uncalibrated Affine SilhouettesabstractA method for recovering structure and motion from a set of scaled orthographic silhouette views with unconstrained camera motion is presented. The outer epipolar tangencies between six or more views are used simultaneously to recover the relative pose of the cameras and hence the visual hull of the object. The camera representation proposed permits a closed form solution in the optimization of the camera parameters. The resulting system is applied to the problem of reconstructing aircraft in flight. Paul McIlroy, Tom Drummond |
BMVC | 2 |
| 2009 | ProFORMA: Probabilistic Feature-based On-line Rapid Model AcquisitionabstractOff-line model reconstruction relies on an image collection phase and a slow reconstruction phase, requiring a long time to verify a model obtained from an image sequence is acceptable. We propose a new model acquisition system, called ProFORMA, which generates a 3D model on-line as the input sequence is being collected. As the user rotates the object in front of a stationary camera, a partial model is reconstructed and displayed to the user to assist view planning. The model is also used by the system to robustly track the pose of the object. Models are rapidly produced through a Delaunay tetrahe-dralisation of points obtained from on-line structure from motion estimation, followed by a probabilistic tetrahedron carving step to obtain a textured surface mesh of the object. © 2009. The copyright of this document resides with its authors. Qi Pan, Gerhard Reitmayr, Tom Drummond |
BMVC | 3 |
| 2009 | Multiple Target Localisation at over 100 FPSabstractThis paper presents a method for fast feature-based matching which enables 7 independent targets to be localised in a video sequence with an average total processing time of 7.46ms per frame. We extend recent work [14] on fast matching using Histogrammed Intensity Patches (HIPs) by adding a rotation invariant framework and a treebased lookup scheme. Compared to state-of-the-art fast localisation schemes [15] we achieve better matching robustness in under a quarter of the computation time and requiring 5-10 times less memory. Simon Taylor, Tom Drummond |
BMVC | 2 |
| 2009 | Interactive model reconstruction with user guidanceabstractProFORMA, an on-line reconstruction system for textured objects rotated by a user's hand, can be coupled with augmented reality (AR) to allow users to rapidly generate textured 3D models. We demonstrate how the use of an overlaid mesh model and 3D arrow can be used to assist the user in view planning, guiding the user to collect new keyframes from desirable views. The method described is particularly suited for use with AR headsets, providing guidance with minimal user input and allowing in situ modelling using the head-mounted camera (ProFORMA does not require a completely stationary camera). Qi Pan, Gerhard Reitmayr, Tom Drummond |
ISMAR | 3 |
| 2009 | Edge landmarks in monocular SLAM
Ethan Eade, Tom Drummond |
Image Vis. Comput. | 2 |
| 2009 | Classification-Based Probabilistic Modeling of Texture Transition for Fast Line Search Tracking and DelineationabstractWe introduce a classification-based approach to finding occluding texture boundaries. The classifier is composed of a set of weak learners which operate on image intensity discriminative features which are defined on small patches and fast to compute. A database which is designed to simulate digitized occluding contours of textured objects in natural images is used to train the weak learners. The trained classifier score is then used to obtain a probabilistic model for the presence of texture transitions which can readily be used for line search texture boundary detection in the direction normal to an initial boundary estimate. This method is fast and therefore suitable for real-time and interactive applications. It works as a robust estimator which requires a ribbon like search region and can handle complex texture structures without requiring a large number of observations. We demonstrate results both in the context of interactive 2-D delineation and fast 3-D tracking and compare its performance with other existing methods for line search boundary detection. Ali Shahrokni, Tom Drummond, François Fleuret, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Guest Editors' Introduction: Special Section on The International Symposium on Mixed and Augmented Reality (ISMAR)abstractThe two papers in this special section are extended versions of papers originally presented at the International Symposium on Mixed and Augmented Reality (ISMAR) 2007. These two papers won awards at the symposium. Mark A. Livingston, Reinhold Behringer, Hirokazu Kato 0001, Tom Drummond |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2008 | Unified Loop Closing and Recovery for Real Time Monocular SLAMabstractWe present a unified method for recovering from tracking failure and closing loops in real time monocular simultaneous localisation and mapping. Within a graph-based map representation, we show that recovery and loop closing both reduce to the creation of a graph edge. We describe and implement a bag-of-words appearance model for ranking potential loop closures, and a robust method for using both structure and image appearance to confirm likely matches. The resulting system closes loops and recovers from failures while mapping thousands of landmarks, all in real time. 1 Ethan Eade, Tom Drummond |
BMVC | 2 |
| 2008 | Student-tMixture Filter for Robust, Real-Time Visual Tracking
James Loxam, Tom Drummond |
ECCV (3) | 2 |
| 2008 | Pose tracking from natural features on mobile phonesabstractIn this paper we present two techniques for natural feature tracking in real-time on mobile phones. We achieve interactive frame rates of up to 20 Hz for natural feature tracking from textured planar targets on current-generation phones. We use an approach based on heavily modified state-of-the-art feature descriptors, namely SIFT and Ferns. While SIFT is known to be a strong, but computationally expensive feature descriptor, Ferns classification is fast, but requires large amounts of memory. This renders both original designs unsuitable for mobile phones. We give detailed descriptions on how we modified both approaches to make them suitable for mobile phones. We present evaluations on robustness and performance on various devices and finally discuss their appropriateness for augmented reality applications. Daniel Wagner 0003, Gerhard Reitmayr, Alessandro Mulloni, Tom Drummond, Dieter Schmalstieg |
ISMAR | 4 |
| 2008 | Multi-modal tracking using texture changes
Christopher Kemp, Tom Drummond |
Image Vis. Comput. | 2 |
| 2007 | Monocular SLAM as a Graph of Coalesced ObservationsabstractWe present a monocular SLAM system that avoids inconsistency by coalescing observations into independent local coordinate frames, building a graph of the local frames, and optimizing the resulting graph. We choose coordinates that minimize the nonlinearity of the updates in the nodes, and suggest a heuristic measure of such nonlinearity, using it to guide our traversal of the graph. The system operates in real-time on sequences with several hundreds of landmarks while performing global graph optimization, yielding accurate and nearly consistent estimation relative to offline bundle adjustment, and considerably better consistency than EKF SLAM and FastSLAM. Ethan Eade, Tom Drummond |
ICCV | 2 |
| 2007 | Semi-Autonomous Generation of Appearance-based Edge Models from Image SequencesabstractMany of the robust visual tracking techniques utilized by augmented reality applications rely on 3D models and information extracted from images. Models enhanced with image information make it possible to initialize tracking and detect poor registration. Unfortunately, generating 3D CAD models and registering them to image information can be a time consuming operation. Regularly the process requires multiple trips between the site being modeled and the workstation used to create the model. The system presented in this work eliminates the need for a separately generated 3D model by utilizing modern structure-from-motion techniques to extract the model and associated image information directly from an image sequence. The technique can be implemented on any handheld device instrumented with a camera and network connection. The process of creating the model requires minimal user interaction in the form of a few cues to identify planar regions on the object of interest. In addition the system selects a set of keyframes for each region to capture viewpoint based appearance changes. This work also presents a robust tracking framework to take advantage of these new edge models. Performance of both the modeling technique and the tracking system are verified on several different objects. Jeremiah Neubert, John Pretlove, Tom Drummond |
ISMAR | 3 |
| 2007 | Initialisation for Visual Tracking in Urban EnvironmentsabstractOutdoor augmented reality systems often rely on GPS to cover large environments. Visual tracking approaches can provide more accurate location estimates but typically require a manual initialisation procedure. This paper describes the combination of both techniques to create an accurate localisation system that does not require any additional input for (re-)initialisation. The 2D GPS position together with average user height is used as an initial estimate for the visual tracking. The large gap in available GPS accuracy versus required accuracy for initialisation is overcome through a search procedure that tries to minimise search time by improving the likelihood of finding the correct estimate early. Re-initialisation of the visual tracking system after catastrophic failures is further improved by modelling the GPS error with a Gaussian process to provide a better estimate of the current location, thereby decreasing search time. Gerhard Reitmayr, Tom Drummond |
ISMAR | 2 |
| 2007 | Semi-automatic Annotations in Unknown EnvironmentsabstractUnknown environments pose a particular challenge for augmented reality applications because the 3D models required for tracking, rendering and interaction are not available ahead of time. Consequently, authoring of AR content must take place on-line. This work describes a set of techniques to simplify the online authoring of annotations in unknown environments using a simultaneous localisation and mapping (SLAM) system. The point-based SLAM system is extended to specifically track and estimate high-level features indicated by the user. The automatic estimation of these complex landmarks by the system relieves the user from the burden of manually specifying the full 3D pose of annotations while improving accuracy. These properties are especially interesting for remote collaboration applications where either user interfaces on handhelds or camera control by the remote expert are limited. Gerhard Reitmayr, Ethan Eade, Tom Drummond |
ISMAR | 3 |
| 2006 | Edge Landmarks in Monocular SLAMabstractWhile many visual simultaneous localization and mapping (SLAM) systems use point features as landmarks, few take advantage of the edge information in images. Those SLAM systems that do observe edge features do not consider edges with all degrees of freedom. Edges are difficult to use in vision SLAM because of selection, observation, initialization and data association challenges. A map that includes edge features, however, contains higher-order geometric information useful both during and after SLAM. We define a well-localized edge landmark and present an efficient algorithm for selecting such landmarks. Further, we describe how to initialize new landmarks, observe mapped landmarks in subsequent images, and address the data association challenges of edges. Our methods, implemented in a particle-filter SLAM system, operate at frame rate on live video sequences. Ethan Eade, Tom Drummond |
BMVC | 2 |
| 2006 | Vision-Based Augmented Reality Visual Guidance with Keyframes
Timothy S. Y. Gan, Tom Drummond |
Computer Graphics International | 2 |
| 2006 | Scalable Monocular SLAMabstractLocalization and mapping in unknown environments becomes more difficult as the complexity of the environment increases. With conventional techniques, the cost of maintaining estimates rises rapidly with the number of landmarks mapped. We present a monocular SLAM system that employs a particle filter and top-down search to allow realtime performance while mapping large numbers of landmarks. To our knowledge, we are the first to apply this FastSLAM-type particle filter to single-camera SLAM. We also introduce a novel partial initialization procedure that efficiently determines the depth of new landmarks. Moreover, we use information available in observations of new landmarks to improve camera pose estimates. Results show the system operating in real-time on a standard workstation while mapping hundreds of landmarks. Ethan Eade, Tom Drummond |
CVPR (1) | 2 |
| 2006 | Machine Learning for High-Speed Corner Detection
Edward Rosten, Tom Drummond |
ECCV (1) | 2 |
| 2006 | Using backlight intensity for device identificationabstractThis paper presents a method for identifying handheld devices (e.g. smart phones and pocket PCs) to facilitate the use of these devices as tangible interfaces for desktop augmented reality systems. The proposed system leverages the ability of these handheld devices to programmatically control their backlight intensity to display a binary code. The codes produced are non-intrusive, require no specialized hardware, and can be generated with most handheld devices. This technique is shown to accurately and robustly identify up to 16 different devices in under 500 msec and is easily expandable to 256 or more devices. Jeremiah Neubert, Tom Drummond |
ISMAR | 2 |
| 2006 | Going out: robust model-based tracking for outdoor augmented realityabstractThis paper presents a model-based hybrid tracking system for outdoor augmented reality in urban environments enabling accurate, realtime overlays for a handheld device. The system combines several well-known approaches to provide a robust experience that surpasses each of the individual components alone: an edge-based tracker for accurate localisation, gyroscope measurements to deal with fast motions, measurements of gravity and magnetic field to avoid drift, and a back store of reference frames with online frame selection to re-initialize automatically after dynamic occlusions or failures. A novel edge-based tracker dispenses with the conventional edge model, and uses instead a coarse, but textured, 3D model. This yields several advantages: scale-based detail culling is automatic, appearance-based edge signatures can be used to improve matching and the models needed are more commonly available. The accuracy and robustness of the resulting system is demonstrated with comparisons to map-based ground truth data. Gerhard Reitmayr, Tom Drummond |
ISMAR | 2 |
| 2005 | A Single-frame Visual GyroscopeabstractRapid camera rotations (e.g. camera shake) are a significant problem when real-time computer vision algorithms are applied to video from a handheld or head-mounted camera. Such camera motions cause image features to move large distances in the image and cause significant motion blur. Here we propose a very fast method of estimating the camera rotation from a single frame which does not require any detection, matching or extraction of feature points and can be used as a motion estimator to reduce the search range for feature matching algorithms that may be subsequently applied to the image. This method exploits the motion blur in the frame, using features which remain sharp to rapidly compute the axis of rotation of the camera, and using blurred features to estimate the magnitude of the camera’s rotation. 1 Georg S. W. Klein, Tom Drummond |
BMVC | 2 |
| 2005 | Dynamic Measurement Clustering to Aid Real Time TrackingabstractWe present a technique/or clustering measurements such that high-dimensional parameter estimation problems can be simplified. The key idea is to find rows of the measurement Jacobian whose rank is significantly less than its width. Such a set of rows gives a cluster of measurements which is affected only by a subset of the parameter space. This cluster can be used independently from other measurements to isolate parameter decisions. Unlike static partitioning techniques, the method presented dynamically generates clusters at each step of the estimation. This achieves substantial computational reductions, even for problems which cannot be partitioned in the traditional sense. The technique is applied to the task of tracking camera motions in real-time and video sequences are used to compare the resulting system to previous methods. Christopher Kemp, Tom Drummond |
ICCV | 2 |
| 2005 | Fusing Points and Lines for High Performance TrackingabstractThis paper addresses the problem of real-time 3D model-based tracking by combining point-based and edge-based tracking systems. We present a careful analysis of the properties of these two sensor systems and show that this leads to some non -trivial design choices that collectively yield extremely high performance. In particular, we present a method for integrating the two systems and robustly combining the pose estimates they produce. Further we show how on-line learning can be used to improve the performance of feature tracking. Finally, to aid real-time performance, we introduce the FAST feature detector which can perform full-frame feature detection at 400Hz. The combination of these techniques results in a system which is capable of tracking average prediction errors of 200 pixels. This level of robustness allows us to track very rapid motions, such as 50deg camera shake at 6Hz Edward Rosten, Tom Drummond |
ICCV | 2 |
| 2005 | Fast Texture-Based Tracking and Delineation Using Texture EntropyabstractWe propose a fast texture-segmentation approach to the problem of 2D and 3D model-based contour tracking, which is suitable for real-time or interactive applications. Our approach relies on detecting texture boundaries in the direction normal to the contour boundaries and on using a hidden Markov model to link these boundary points in the other direction. The probabilities that appear in this computation closely relate to texture entropy and Kullback-Leibler divergence, a property we use to compute and update dynamic texture models. We demonstrate results both in the context of interactive 2D delineation and fast 3D tracking Ali Shahrokni, Tom Drummond, Pascal Fua |
ICCV | 2 |
| 2005 | Localisation and Interaction for Augmented MapsabstractPaper-based cartographic maps provide highly detailed information visualisation with unrivalled fidelity and information density. Moreover, the physical properties of paper afford simple interactions for browsing a map or focusing on individual details, managing concurrent access for multiple users and general malleability. However, printed maps are static displays and while computer-based map displays can support dynamic information, they lack the nice properties of real maps identified above. We address these shortcomings by presenting a system to augment printed maps with digital graphical information and user interface components. These augmentations complement the properties of the printed information in that they are dynamic, permit layer selection and provide complex computer mediated interactions with geographically embedded information and user interface controls. Two methods are presented which exploit the benefits of using tangible artifacts for such interactions. Gerhard Reitmayr, Ethan Eade, Tom Drummond |
ISMAR | 3 |
| 2004 | Multi-Modal Tracking using Texture ChangesabstractWe present a method for efficiently generating a representation of a multi-modal posterior probability distribution. The technique combines ideas from RANSAC and particle filtering such that the 3D visual tracking problem can be partitioned into two levels, while maintaining multiple hypotheses throughout. A simple texture change-point detector finds multiple hypotheses for the position of image edgels. From these, multiple locations for each scene edge are generated. Finally, we determine the best pose of the whole structure. While the multi-modal representation is strongly related to particle filtering techniques, this approach is driven by data from the image. Hence the resulting system is able to perform robust visual tracking of all six degrees of freedom in real time. Real video sequences are used to compare the complete tracking system to previous systems. Christopher Kemp, Tom Drummond |
BMVC | 2 |
| 2004 | Markov-based Silhouette Extraction for Three--Dimensional Body Tracking in Presence of Cluttered BackgroundabstractWe propose a novel method to detect human body contours in presence of clutter and complex texture. Contours are extracted using a novel Markovbased approach which learns a texture along a given scanline in order to detect texture crossings. In contrast to conventional silhouette detection algorithms based on gradient, our texture boundary detection method allows extraction of silhouettes of textured and non-textured objects under difficult conditions such as having a cluttered/moving background. We demonstrate on demanding examples of monocular body tracking that our proposed method yields better results than gradient-based techniques. 1. Ali Shahrokni, Vincent Lepetit, Tom Drummond, Pascal Fua |
BMVC | 3 |
| 2004 | Texture Boundary Detection for Real-Time Tracking
Ali Shahrokni, Tom Drummond, Pascal Fua |
ECCV (2) | 2 |
| 2004 | Sensor Fusion and Occlusion Refinement for Tablet-Based ARabstractThis paper presents a set of technologies which enable robust, accurate, high resolution augmentation of live video, delivered via a tablet PC to which a video camera has been attached. By combining several technologies, this is achieved without the use of contrived markers in the environment: An outside-in tracker observes the tablet to generate robust, low-accuracy pose estimates. An inside-out tracker running on the tablet observes the video feed from the tablet-mounted camera and provides high-accuracy pose estimates by tracking natural features in the environment. Information from both of these trackers is combined in an extended Kalman filter. Finally, to maximise the quality of the augmented imagery, boundaries where the real world occludes the virtual imagery are identified and another tracker is used to refine the boundaries between real and virtual imagery so that their synthesis is as convincing as possible. Georg S. W. Klein, Tom Drummond |
ISMAR | 2 |
| 2004 | Tightly integrated sensor fusion for robust visual tracking
Georg S. W. Klein, Tom Drummond |
Image Vis. Comput. | 2 |
| 2004 | Layered Motion Segmentation and Depth Ordering by Tracking EdgesabstractThis paper presents a new Bayesian framework for motion segmentation--dividing a frame from an image sequence into layers representing different moving objects--by tracking edges between frames. Edges are found using the Canny edge detector, and the Expectation-Maximization algorithm is then used to fit motion models to these edges and also to calculate the probabilities of the edges obeying each motion model. The edges are also used to segment the image into regions of similar color. The most likely labeling for these regions is then calculated by using the edge probabilities, in association with a Markov Random Field-style prior. The identification of the relative depth ordering of the different motion layers is also determined, as an integral part of the process. An efficient implementation of this framework is presented for segmenting two motions (foreground and background) using two frames. It is then demonstrated how, by tracking the edges into further frames, the probabilities may be accumulated to provide an even more accurate and robust estimate, and segment an entire sequence. Further extensions are then presented to address the segmentation of more than two motions. Here, a hierarchical method of initializing the Expectation-Maximization algorithm is described, and it is demonstrated that the Minimum Description Length principle may be used to automatically select the best number of motion layers. The results from over 30 sequences (demonstrating both two and three motions) are presented and discussed. Tom Drummond, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | Rapid rendering of apparent contours of implicit surfaces for real-time trackingabstractThis paper addresses the problem of real-time visual tracking of structures with curved surfaces by localising the apparent contour in each frame of an image sequence. A scheme is presented for rapidly rendering the apparent contour from a predicted pose. Errors between this contour and the observed contour are then used to update the pose estimate for tracking. The rendering algorithm makes use of two contributions. Firstly, a differential equation is derived which traces out the contour generators on an iso-surface of a scalar field. Secondly a set of rules for determining the visibility of each part of the apparent contour is presented. These techniques are used to render structures of moderate complexity in under 30ms which permits real-time tracking at video frame rate. Edward Rosten, Tom Drummond |
BMVC | 2 |
| 2003 | Computing MAP trajectories by representing, propagating and combining PDFs over groupsabstractThis paper addresses the problem of computing the trajectory of a camera from sparse positional measurements that have been obtained from visual localisation, and dense differential measurements from odometry or inertial sensors. A fast method is presented for fusing these two sources of information to obtain the maximum a posteriori estimate of the trajectory. A formalism is introduced for representing probability density functions over Euclidean transformations, and it is shown how these density functions can be propagated along the data sequence and how multiple estimates of a transformation can be combined. A three-pass algorithm is described which makes use of these results to yield the trajectory of the camera. Simulation results are presented which are validated against a physical analogue of the vision problem, and results are then shown from sequences of approximately 1,800 frames captured from a video camera mounted on a go-kart. Several of these frames are processed using computer vision to obtain estimates of the position of the go-kart. The algorithm fuses these estimates with odometry from the entire sequence in 150 ms to obtain the trajectory of the kart. Tom Drummond, Kimon Roussopoulos |
ICCV | 2 |
| 2003 | Robust Visual Tracking for Non-Instrumented Augmented RealityabstractThis paper presents a robust and flexible framework for augmented reality which does not require instrumenting either the environment or the workpiece. A model-based visual tracking system is combined with rate gyroscopes to produce a system which can track the rapid camera rotations generated by a head-mounted camera, even if images are substantially degraded by motion blur. This tracking yields estimates of head position at video field rate (50Hz) which are used to align computer-generated graphics on an optical see-through display. Nonlinear optimisation is used for the calibration of display parameters which include a model of optical distortion. Rendered visuals are pre-distorted to correct the optical distortion of the display. Georg S. W. Klein, Tom Drummond |
ISMAR | 2 |
| 2002 | Tightly Integrated Sensor Fusion for Robust Visual TrackingabstractThis paper presents novel methods for increasing the robustness of visual tracking systems by incorporating information from inertial sensors. We show that more can be achieved than simply combining the sensor data within a statistical filter. In particular we show how, in addition to using inertial data to provide predictions for the visual sensor, this data can also be used to provide an estimate of motion blur for each feature and this can be used to dynamically tune the parameters of each feature detector in the visual sensor. This allows the system to obtain useful information from the visual sensor even in the presence of substantial motion blur. Finally, the visual sensor can be used to calibrate the parameters of the inertial sensor to eliminate drift. Georg S. W. Klein, Tom Drummond |
BMVC | 2 |
| 2002 | Real-time tracking of complex structures with on-line camera calibration
Tom Drummond, Roberto Cipolla |
Image Vis. Comput. | 1 |
| 2002 | Real-Time Visual Tracking of Complex StructuresabstractPresents a framework for three-dimensional model-based tracking. Graphical rendering technology is combined with constrained active contour tracking to create a robust wire-frame tracking system. It operates in real time at video frame rate (25 Hz) on standard hardware. It is based on an internal CAD model of the object to be tracked which is rendered using a binary space partition tree to perform hidden line removal. A Lie group formalism is used to cast the motion computation problem into simple geometric terms so that tracking becomes a simple optimization problem solved by means of iterative reweighted least squares. A visual servoing system constructed using this framework is presented together with results showing the accuracy of the tracker. The paper then describes how this tracking system has been extended to provide a general framework for tracking in complex configurations. The adjoint representation of the group is used to transform measurements into common coordinate frames. The constraints are then imposed by means of Lagrange multipliers. Results from a number of experiments performed using this framework are presented and discussed. Tom Drummond, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | Using occlusions to aid position estimation for visual motion captureabstractA new method for estimating the pose of a person in a marker-based visual motion capture system is presented. In such systems, certain markers are not detected at each time frame due to occlusion by the subject. It is proposed to estimate the person's pose subject to the constraint that non-detected markers remain occluded and that detected markers remain visible. In this paper, the complete posterior PDF for the person's pose is developed using the constraints mentioned above, as is a least squares solution for determining its maximum. It is shown that the new technique is able to correctly estimate a variety of poses in which many marker occlusions occur and in which typical motion capture systems fail. Maurice Ringer, Tom Drummond, Joan Lasenby |
CVPR (2) | 2 |
| 2001 | A Probabilistic Framework for Space Carving
Adrian Broadhurst, Tom Drummond, Roberto Cipolla |
ICCV | 2 |
| 2001 | Real-Time Tracking of Highly Articulated Structures in the Presence of Noisy MeasurementsabstractThis paper presents a novel approach for model-based real-time tracking of highly articulated structures such as humans. This approach is based on an algorithm which efficiently propagates statistics of probability distributions through a kinematic chain to obtain maximum a posteriori estimates of the motion of the entire structure. This algorithm yields the least squares solution in linear time (in the number of components of the model) and can also be applied to non-Gaussian statistics using a simple but powerful trick. The resulting implementation runs in real-time on standard hardware without any pre-processing of the video data and can thus operate on live video. Results from experiments performed using this system are presented and discussed. Tom Drummond, Roberto Cipolla |
ICCV | 1 |
| 2000 | 3D Model Acquisition by Tracking 2D WireframesabstractThis paper presents a semi-automatic wireframe acquisition system. The sys-tem uses real-time (25Hz) tracking of a user specified 2D wireframe and in-termittent camera pose parameters to accumulate 3D position information. The 2D tracking framework enables the application of model-based con-straints in an intuitive way which interacts naturally with the Kalman filter formulation used. In particular, this is used to introduce feedback from the current 3D shape estimate to improve the robustness of the 2D tracking. The scheme allows wireframe models of simple edge based objects to be built in around 5 minutes. 1 Matthew A. Brown, Tom Drummond, Roberto Cipolla |
BMVC | 2 |
| 2000 | Segmentation of Multiple Motions by Edge Tracking between Two FramesabstractThis paper presents a method for segmenting multiple motions using edges. Recent work in this field has been constrained to the case of two motions, and this paper demonstrates that the approach can be extended to more than two motions. The image is first segmented into regions, and then the framework determines the motions present and labels the edges in the image. Initial-isation is particularly difficult, and a novel scheme is proposed which re-cursively splits motions to provide the Expectation-Maximisation algorithm with a reasonable guess, and a Minimum Description Length approach is used to determine the best number of models to use. The edge labels are then used to determine the the region labelling. A global optimisation is intro-duced to refine the motions and provide the most likely region labelling. 1 Tom Drummond, Roberto Cipolla |
BMVC | 2 |
| 2000 | Real-Time Tracking of Multiple Articulated Structures in Multiple Views
Tom Drummond, Roberto Cipolla |
ECCV (2) | 1 |
| 2000 | Motion Segmentation by Tracking Edge Information over Multiple Frames
Tom Drummond, Roberto Cipolla |
ECCV (2) | 2 |
| 2000 | Learning Task-Specific Object Recognition and Scene Understanding
Tom Drummond, Terry Caelli |
Comput. Vis. Image Underst. | 1 |
| 2000 | Application of Lie Algebras to Visual Servoing
Tom Drummond, Roberto Cipolla |
Int. J. Comput. Vis. | 1 |
| 1999 | Camera Calibration from Vanishing Points in Image of Architectural Scenes
Roberto Cipolla, Tom Drummond, Duncan P. Robertson |
BMVC | 2 |
| 1999 | Real-time Tracking of Complex Structures with On-line Camera CalibrationabstractThis paper presents a novel three-dimensional model-based tracking system which has been incorporated into a visual servoing system. The tracking system combines modern graphical rendering technology with constrained active contour tracking techniques to create wireframe -snakes. It operates in real time at video frame rate (25 Hz) and is based on an internal CAD model of the object to be tracked. This model is rendered using a binary space partition tree to perform hidden line removal and the visible features are identified on-line at each frame and are tracked in the video feed. The tracking system has been extended to incorporate real-time on-line calibration and tracking of internal camera parameters. Results from on-line calibration and visual servoing experiments are presented. 1 Introduction The tracking of rigid three-dimensional objects is useful for numerous applications, including motion analysis, surveillance and robotic control tasks. This paper tackles two p... Tom Drummond, Roberto Cipolla |
BMVC | 1 |
| 1999 | Edge Tracking for Motion Segmentation and Depth OrderingabstractThis paper presents a new theoretical framework for motion segmentation based on the motion of tracked region edges. By considering the visible edges of an object, constraints may be placed on the motion labelling of edges. This enables the most likely region labelling and layer ordering to be established, thus producing a segmentation. An implementation is outlined and demonstrated on test sequences containing two motions. The image is divided into regions using a colour edgebased segmentation scheme and the normal motion of these edges is tracked. The EM algorithm is used to partition the edges and fit the best two motions accordingly. Hypothesising each motion in turn to be the foreground motion, the labelling constraints can be applied and the frame segmented. The hypothesis which best fits the observed edge motions indicates the layer ordering and leads to a very accurate segmentation. 1 Introduction The segmentation of a video sequence into moving objects is a first... Tom Drummond, Roberto Cipolla |
BMVC | 2 |
| 1999 | Visual Tracking and Control using Lie AlgebrasabstractA novel approach to visual servoing is presented, which takes advantage of the structure of the Lie algebra of affine transformations. The aim of this project is to use feedback from a visual sensor to guide a robot arm to a target position. The sensor is placed in the end effector of the robot, the 'camera-in-hand' approach, and thus provides direct feedback of the robot motion relative to the target scene via observed transformations of the scene. These scene transformations are obtained by measuring the affine deformations of a target planar contour, captured by use of an active contour, or snake. Deformations of the snake are constrained using the Lie groups of affine and projective transformations. Properties of the Lie algebra of affine transformations are exploited to integrate observed deformations to the target contour which can be compensated with appropriate robot motion using a non-linear control structure. These techniques have been implemented using a video camera to control a 5 DoF robot arm. Experiments with this implementation are presented, together with a discussion of the results. Tom Drummond, Roberto Cipolla |
CVPR | 1 |