VLDB 2026 Research / reviewers in the wild / expert
Roberto Cipolla
dblp:c/RobertoCipolla
· DBLP profile ↗
302ranked-venue papers
12as first author
31since 2021 · last 2025
0000-0002-8999-2151ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 280 · 12 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 225 · 7 first-author · 28 since 2021Systems, architecture and hardware · 4Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FOCUS - Multi-View Foot Reconstruction from Synthetically Trained Dense CorrespondencesabstractSurface reconstruction from multiple, calibrated images is a challenging task - often requiring a large number of collected images with significant overlap. We look at the specific case of human foot reconstruction. As with previous successful foot reconstruction work, we seek to extract rich per-pixel geometry cues from multi-view RGB images, and fuse these into a final 3D object. Our method, FOCUS, tackles this problem with 3 main contributions: (i) SynFoot2, an extension of an existing synthetic foot dataset to include a new data type: dense correspondence with the parameterized foot model FIND; (ii) an uncertainty-aware dense correspondence predictor trained on our synthetic dataset; (iii) two methods for reconstructing a 3D surface from dense correspondence predictions: one inspired by Structure-from-Motion, and one optimization-based using the FIND model. We show that our reconstruction achieves state-of-the-art reconstruction quality in a few-view setting, performing comparably to state-of-the-art when many views are available, runs substantially faster, and can run without a GPU. We release our synthetic dataset to the research community. Code is available at: https://github.com/OllieBoyne/FOCUS Oliver Boyne, Roberto Cipolla |
3DV | 2 |
| 2025 | NPL-MVPS: Neural Point-Light Multi-View Photometric StereoabstractIn this work we present a novel multi-view photometric stereo (MVPS) method. Like many works in 3D reconstruction we are leveraging neural shape representations and learnt renderers. However, our work differs from the state-of-the-art multi-view PS methods such as PS-NeRF [47] or Supernormal [4] in that we explicitly leverage per-pixel intensity renderings rather than relying mainly on estimated normals. We model point light attenuation and explicitly raytrace cast shadows in order to best approximate the incoming radiance for each point. The estimated incoming radiance is used as input to a fully neural material renderer that uses minimal prior assumptions and it is jointly optimised with the surface. Estimated normals and segmentation maps are also incorporated in order to maximise the surface accuracy. Our method is among the first (along with Supernormal [4]) to outperform the classical MVPS approach proposed by the DiLiGenT-MV benchmark and achieves average 0.2mm Chamfer distance for objects imaged at approx 1.5m distance away with approximate 400 × 400 resolution. Moreover, our method shows high robustness to the sparse MVPS setup (6 views, 6 lights) greatly outperforming the SOTA competitor (0.38mm vs 0.61mm), illustrating the importance of neural rendering in multi-view photometric stereo. Fotios Logothetis, Ignas Budvytis, Roberto Cipolla |
WACV | 3 |
| 2024 | ReCoRe: Regularized Contrastive Representation Learning of World ModelabstractWhile recent model-free Reinforcement Learning (RL) methods have demonstrated human-level effectiveness in gaming environments, their success in everyday tasks like visual navigation has been limited, particularly under significant appearance variations. This limitation arises from (i) poor sample efficiency and (ii) over-fitting to training scenarios. To address these challenges, we present a world model that learns invariant features using (i) contrastive unsupervised learning and (ii) an intervention-invariant regularizer. Learning an explicit representation of the world dynamics i.e. a world model, improves sample efficiency while contrastive learning implicitly enforces learning of invariant features, which improves generalization. However, the naïve integration of contrastive loss to world models is not good enough, as world-model-based RL methods independently optimize representation learning and agent policy. To overcome this issue, we propose an intervention-invariant regularizer in the form of an auxiliary task such as depth prediction, image denoising, image segmentation, etc., that explicitly enforces invariance to style interventions. Our method outperforms current state-of-the-art model-based and model-free RL methods and significantly improves on out-of-distribution point navigation tasks evaluated on the iGibson benchmark. With only visual observations, we further demonstrate that our approach outperforms recent language-guided foundation models for point navigation, which is essential for deployment on robots with limited computation capabilities. Finally, we demonstrate that our proposed model excels at the sim-to-real transfer of its perception module on the Gibson benchmark. Rudra P. K. Poudel, Harit Pandya, Stephan Liwicki, Roberto Cipolla |
CVPR | 4 |
| 2024 | FOUND: Foot Optimization with Uncertain Normals for Surface Deformation Using Synthetic DataabstractSurface reconstruction from multi-view images is a challenging task, with solutions often requiring a large number of sampled images with high overlap. We seek to develop a method for few-view reconstruction, for the case of the human foot. To solve this task, we must extract rich geometric cues from RGB images, before carefully fusing them into a final 3D object. Our FOUND approach tackles this, with 4 main contributions: (i) SynFoot, a synthetic dataset of 50,000 photorealistic foot images, paired with ground truth surface normals and keypoints; (ii) an uncertainty-aware surface normal predictor trained on our synthetic dataset; (iii) an optimization scheme for fitting a generative foot model to a series of images; and (iv) a benchmark dataset of calibrated images and high resolution ground truth geometry. We show that our normal predictor outperforms all off-the-shelf equivalents significantly on real images, and our optimization scheme outperforms state-of-the-art photogrammetry pipelines, especially for a few-view setting. We release our synthetic dataset and baseline 3D scans to the research community. Oliver Boyne, Gwangbin Bae, James Charles, Roberto Cipolla |
WACV | 4 |
| 2024 | A Neural Height-Map Approach for the Binocular Photometric Stereo ProblemabstractIn this work we propose a novel, highly practical, binocular photometric stereo (PS) framework, which has same acquisition speed as single view PS, however significantly improves the quality of the estimated geometry.As in recent neural multi-view shape estimation frameworks such as NeRF [29], SIREN [35] and inverse graphics approaches to multi-view photometric stereo (e.g. PS-NeRF [38]) we formulate shape estimation task as learning of a differentiable surface and texture representation by minimising surface normal discrepancy for normals estimated from multiple varying light images for two views as well as discrepancy between rendered surface intensity and observed images. Our method differs from typical multi-view shape estimation approaches in two key ways. First, our surface is represented not as a volume but as a neural heightmap where heights of points on a surface are computed by a deep neural network. Second, instead of predicting an average intensity as PS-NeRF or introducing lambertian material assumptions as Guo et al. [7], we use a learnt BRDF and perform near-field per point intensity rendering.Our method achieves the state-of-the-art performance on the DiLiGenT-MV dataset adapted to binocular stereo setup as well as a new binocular photometric stereo dataset - LUCES-ST. Fotios Logothetis, Ignas Budvytis, Roberto Cipolla |
WACV | 3 |
| 2023 | Sparse Multi-Object Render-and-Compare
Florian Langer, Ignas Budvytis, Roberto Cipolla |
BMVC | 3 |
| 2023 | HuManiFlow: Ancestor-Conditioned Normalising Flows on SO(3) Manifolds for Human Pose and Shape Distribution EstimationabstractMonocular 3D human pose and shape estimation is an illposed problem since multiple 3D solutions can explain a 2D image of a subject. Recent approaches predict a probability distribution over plausible 3D pose and shape parameters conditioned on the image. We show that these approaches exhibit a trade-off between three key properties: (i) accuracy - the likelihood of the ground-truth 3D solution under the predicted distribution, (ii) sample-input consistency - the extent to which 3D samples from the predicted distribution match the visible 2D image evidence, and (iii) sample diversity - the range of plausible 3D solutions modelled by the predicted distribution. Our method, HuManiFlow, predicts simultaneously accurate, consistent and diverse distributions. We use the human kinematic tree to factorise full body pose into ancestor-conditioned per-body-part pose distributions in an autoregressive manner. Per-body-part distributions are implemented using normalising flows that respect the manifold structure of SO(3), the Lie group of per-body-part poses. We show that ill-posed, but ubiquitous, 3D point estimate losses reduce sample diversity, and employ only probabilistic training losses. HuManiFlow outperforms state-of-the-art probabilistic approaches on the 3DPW and SSP-3D datasets. Akash Sengupta, Ignas Budvytis, Roberto Cipolla |
CVPR | 3 |
| 2023 | SFD2: Semantic-Guided Feature Detection and DescriptionabstractVisual localization is a fundamental task for various applications including autonomous driving and robotics. Prior methods focus on extracting large amounts of often redundant locally reliable features, resulting in limited efficiency and accuracy, especially in large-scale environments under challenging conditions. Instead, we propose to extract globally reliable features by implicitly embedding high-level semantics into both the detection and description processes. Specifically, our semantic-aware detector is able to detect keypoints from reliable regions (e.g. building, traffic lane) and suppress unreliable areas (e.g. sky, car) implicitly instead of relying on explicit semantic labels. This boosts the accuracy of keypoint matching by reducing the number of features sensitive to appearance changes and avoiding the need of additional segmentation networks at test time. Moreover, our descriptors are augmented with semantics and have stronger discriminative ability, providing more inliers at test time. Particularly, experiments on long-term large-scale visual localization Aachen Day Night and RobotCar-Seasons datasets demonstrate that our model outperforms previous local features and gives competitive accuracy to advanced matchers but is about 2 and 3 times faster when using 2k and 4k keypoints, respectively. Code is available at https://github.com/feixue94/sfd2. Ignas Budvytis, Roberto Cipolla |
CVPR | 3 |
| 2023 | IMP: Iterative Matching and Pose Estimation with Adaptive PoolingabstractPrevious methods solve feature matching and pose estimation using a two-stage process by first finding matches and then estimating the pose. As they ignore the geometric relationships between the two tasks, they focus on either improving the quality of matches or filtering potential outliers, leading to limited efficiency or accuracy. In contrast, we propose an iterative matching and pose estimation framework (IMP) leveraging the geometric connections between the two tasks: a few good matches are enough for a roughly accurate pose estimation; a roughly accurate pose can be used to guide the matching by providing geometric constraints. To this end, we implement a geometry-aware recurrent attention-based module which jointly outputs sparse matches and camera poses. Specifically, for each iteration, we first implicitly embed geometric information into the module via a pose-consistency loss, allowing it to predict geometry-aware matches progressively. Second, we introduce an efficient IMP, called EIMP, to dynamically discard keypoints without potential matches, avoiding redundant updating and significantly reducing the quadratic time complexity of attention computation in transformers. Experiments on YFCC100m, Scannet, and Aachen Day-Night datasets demonstrate that the proposed method outperforms previous approaches in terms of accuracy and efficiency. Code is available at https://github.com/feixue94/imp-release Ignas Budvytis, Roberto Cipolla |
CVPR | 3 |
| 2023 | DigiFace-1M: 1 Million Digital Face Images for Face RecognitionabstractState-of-the-art face recognition models show impressive accuracy, achieving over 99.8% on Labeled Faces in the Wild (LFW) dataset. Such models are trained on large-scale datasets that contain millions of real human face images collected from the internet. Web-crawled face images are severely biased (in terms of race, lighting, makeup, etc) and often contain label noise. More importantly, the face images are collected without explicit consent, raising ethical concerns. To avoid such problems, we introduce a large-scale synthetic dataset for face recognition, obtained by rendering digital faces using a computer graphics pipeline1. We first demonstrate that aggressive data augmentation can significantly reduce the synthetic-to-real domain gap. Having full control over the rendering pipeline, we also study how each attribute (e.g., variation in facial pose, accessories and textures) affects the accuracy. Compared to Syn-Face, a recent method trained on GAN-generated synthetic faces, we reduce the error rate on LFW by 52.5% (accuracy from 91.93% to 96.17%). By fine-tuning the network on a smaller number of real face images that could reason-ably be obtained with consent, we achieve accuracy that is comparable to the methods trained on millions of real face images. Gwangbin Bae, Martin de La Gorce, Tadas Baltrusaitis, Charlie Hewitt, Dong Chen 0003, Julien P. C. Valentin, Roberto Cipolla, Jingjing Shen |
WACV | 7 |
| 2023 | A CNN Based Approach for the Point-Light Photometric Stereo Problem
Fotios Logothetis, Roberto Mecca, Ignas Budvytis, Roberto Cipolla |
Int. J. Comput. Vis. | 4 |
| 2023 | HexNet: An Orientation-Aware Deep Learning Framework for Omni-Directional InputabstractWhile omni-directional sensors provide holistic representations typical deep learning frameworks reduce the benefits by introducing distortions and discontinuities as spherical data is supplied as planar input. On the other hand, recent spherical convolutional neural networks (CNNs) often require significant memory and parameters, thus enabling execution only at very low resolutions and shallow architectures. We propose HexNet, an orientation-aware deep learning framework for spherical signals, that allows for fast computation as we exploit standard planar network operations on an efficiently arranged projection of the sphere. Furthermore, we introduce a graph-based version for partial spheres, allowing us to compete at high-resolution with planar CNNs using residual network architectures. Our kernels operate on the tangent of the sphere and thus standard feature weights, pretrained on perspective data, can be transferred, enabling spherical pretraining on ImageNet. As our design is free of distortions and discontinuity, our orientation-aware CNN becomes a new state of the art for semantic segmentation on the recent 2D3DS dataset, and the omni-directional version of SYNTHIA introduced in this work. Moreover, we experimentally show the benefit of our spherical representation over standard images on the Cityscapes dataset by reducing distortion effects of planar CNNs. We implement object detection for the spherical domain. Rotation invariant classification and segmentation tasks are additionally presented for comparison to prior art. Chao Zhang 0023, Stephan Liwicki, Sen He 0001, William A. P. Smith, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Beyond the CLS Token: Image Reranking using Pretrained Vision Transformers
Chao Zhang 0023, Stephan Liwicki, Roberto Cipolla |
BMVC | 3 |
| 2022 | IronDepth: Iterative Refinement of Single-View Depth using Surface Normal and its Uncertainty
Gwangbin Bae, Ignas Budvytis, Roberto Cipolla |
BMVC | 3 |
| 2022 | FIND: An Unsupervised Implicit 3D Model of Articulated Human Feet
Oliver Boyne, James Charles, Roberto Cipolla |
BMVC | 3 |
| 2022 | Style2NeRF: An Unsupervised One-Shot NeRF for Semantic 3D Reconstruction
James Charles, Wim Abbeloos, Daniel Olmeda Reino, Roberto Cipolla |
BMVC | 4 |
| 2022 | SPARC: Sparse Render-and-Compare for CAD model alignment in a single RGB Image
Florian Langer, Gwangbin Bae, Ignas Budvytis, Roberto Cipolla |
BMVC | 4 |
| 2022 | Multi-View Depth Estimation by Fusing Single-View Depth Probability with Multi-View GeometryabstractMulti-view depth estimation methods typically require the computation of a multi-view cost-volume, which leads to huge memory consumption and slow inference. Furthermore, multi-view matching can fail for texture-less surfaces, reflective surfaces and moving objects. For such failure modes, single-view depth estimation methods are often more reliable. To this end, we propose MaGNet, a novel framework for fusing single-view depth probability with multi-view geometry, to improve the accuracy, robustness and efficiency of multi-view depth estimation. For each frame, MaGNet estimates a single-view depth probability distribution, parameterized as a pixel-wise Gaussian. The distribution estimated for the reference frame is then used to sample per-pixel depth candidates. Such probabilistic sampling enables the network to achieve higher accuracy while evaluating fewer depth candidates. We also propose depth consistency weighting for the multi-view matching score, to ensure that the multi-view depth is consistent with the single-view predictions. The proposed method achieves state-of-the-art performance on ScanNet [8], 7- Scenes [38] and KITTI [15]. Qualitative evaluation demonstrates that our method is more robust against challenging artifacts such as texture-less/reflective surfaces and moving objects. Our code and model weights are available at https://github.com/baegwangbin/MaGNet. Gwangbin Bae, Ignas Budvytis, Roberto Cipolla |
CVPR | 3 |
| 2022 | Efficient Large-scale Localization by Global Instance RecognitionabstractHierarchical frameworks consisting of both coarse and fine localization are often used as the standard pipeline for large-scale visual localization. Despite their promising performance in simple environments, they still suffer from low efficiency and accuracy in large-scale scenes, especially under challenging conditions. In this paper, we propose an efficient and accurate large-scale localization framework based on the recognition of buildings, which are not only discriminative for coarse localization but also robust for fine localization. Specifically, we assign each building instance a global ID and perform pixel-wise recognition of these global instances in the localization process. For coarse localization, we employ an efficient reference search strategy to find candidates progressively from the local map observing recognized instances instead of the whole database. For fine localization, predicted labels are further used for instance-wise feature detection and matching, allowing our model to focus on fewer but more robust keypoints for establishing correspondences. The experiments in long-term large-scale localization datasets including Aachen and RobotCar-Seasons demonstrate that our method outperforms previous approaches consistently in terms of both efficiency and accuracy. Ignas Budvytis, Daniel Olmeda Reino, Roberto Cipolla |
CVPR | 4 |
| 2022 | Model-Based Imitation Learning for Urban DrivingabstractAn accurate model of the environment and the dynamic agents acting in it offers great potential for improving motion planning. We present MILE: a Model-based Imitation LEarning approach to jointly learn a model of the world and a policy for autonomous driving. Our method leverages 3D geometry as an inductive bias and learns a highly compact latent space directly from high-resolution videos of expert demonstrations. Our model is trained on an offline corpus of urban driving data, without any online interaction with the environment. MILE improves upon prior state-of-the-art by 31% in driving score on the CARLA simulator when deployed in a completely new town and new weather conditions. Our model can predict diverse and plausible states and actions, that can be interpretably decoded to bird's-eye view semantic segmentation. Further, we demonstrate that it can execute complex driving manoeuvres from plans entirely predicted in imagination. Our approach is the first camera-only method that models static scene, dynamic scene, and ego-behaviour in an urban driving environment. The code and model weights are available at https://github.com/wayveai/mile. Anthony Hu 0001, Gianluca Corrado, Nicolas Griffiths, Zachary Murez, Corina Gurau, Hudson Yeo, Alex Kendall, Roberto Cipolla, Jamie Shotton |
NeurIPS | 8 |
| 2022 | Discrete neural representations for explainable anomaly detectionabstractThe aim of this work is to detect and automatically generate high-level explanations of anomalous events in video. Understanding the cause of an anomalous event is crucial as the required response is dependant on its nature and severity. Recent works typically use object or action classifier to detect and provide labels for anomalous events. However, this constrains detection systems to a finite set of known classes and prevents generalisation to unknown objects or behaviours. Here we show how to robustly detect anomalies without the use of object or action classifiers yet still recover the high level reason behind the event. We make the following contributions: (1) a method using saliency maps to decouple the explanation of anomalous events from object and action classifiers, (2) show how to improve the quality of saliency maps using a novel neural architecture for learning discrete representations of video by predicting future frames and (3) beat the state-of-the-art anomaly explanation methods by 60% on a subset of the public benchmark X-MAN dataset [25]. Stanislaw Szymanowicz, James Charles, Roberto Cipolla |
WACV | 3 |
| 2021 | Lifted Semantic Graph Embedding for Omnidirectional Place RecognitionabstractTypical place recognition is dependent on the visual appearance and camera position of query images, without explicit use of domain knowledge and geometric relationships between key features in the scene. We exploit semantic grouping of pixels, and camera-pose robust scene graphs to perform structure-based visual localization for place recognition. In particular, we first formulate place recognition as an image retrieval task. Then, we lift the omnidirectional input images into 3D space, and compute a rotation and translation invariant semantic graph embedding to encode query and reference images. Finally, place information is obtained through graph similarity matching. Our graph representation is a simple addition to standard image embeddings with minimal overhead, but contains awareness of objects and their geometric relationships. In our experiments, we show improvement over typical place recognition, especially in environments with repetitions and dynamic appearance changes. Chao Zhang 0023, Ignas Budvytis, Stephan Liwicki, Roberto Cipolla |
3DV | 4 |
| 2021 | Leveraging Geometry for Shape Estimation from a Single RGB Image
Florian Langer, Ignas Budvytis, Roberto Cipolla |
BMVC | 3 |
| 2021 | LUCES: A Dataset for Near-Field Point Light Source Photometric Stereo
Roberto Mecca, Fotios Logothetis, Ignas Budvytis, Roberto Cipolla |
BMVC | 4 |
| 2021 | Probabilistic Estimation of 3D Human Shape and Pose with a Semantic Local Parametric Model
Akash Sengupta, Ignas Budvytis, Roberto Cipolla |
BMVC | 3 |
| 2021 | Probabilistic 3D Human Shape and Pose Estimation From Multiple Unconstrained Images in the WildabstractThis paper addresses the problem of 3D human body shape and pose estimation from RGB images. Recent progress in this field has focused on single images, video or multi-view images as inputs. In contrast, we propose a new task: shape and pose estimation from a group of multiple images of a human subject, without constraints on subject pose, camera viewpoint or background conditions between images in the group. Our solution to this task predicts distributions over SMPL body shape and pose parameters conditioned on the input images in the group. We probabilistically combine predicted body shape distributions from each image to obtain a final multi-image shape prediction. We show that the additional body shape information present in multi-image input groups improves 3D human shape estimation metrics compared to single-image inputs on the SSP-3D dataset and a private dataset of tape-measured humans. In addition, predicting distributions over 3D bodies allows us to quantify pose prediction uncertainty, which is useful when faced with challenging input images with significant occlusion. Our method demonstrates meaningful pose uncertainty on the 3DPW dataset and is competitive with the state-of-the-art in terms of pose estimation metrics. Akash Sengupta, Ignas Budvytis, Roberto Cipolla |
CVPR | 3 |
| 2021 | Estimating and Exploiting the Aleatoric Uncertainty in Surface Normal EstimationabstractSurface normal estimation from a single image is an important task in 3D scene understanding. In this paper, we address two limitations shared by the existing methods: the inability to estimate the aleatoric uncertainty and lack of detail in the prediction. The proposed network estimates the per-pixel surface normal probability distribution. We introduce a new parameterization for the distribution, such that its negative log-likelihood is the angular loss with learned attenuation. The expected value of the angular error is then used as a measure of the aleatoric uncertainty. We also present a novel decoder framework where pixel-wise multi-layer perceptrons are trained on a subset of pixels sampled based on the estimated uncertainty. The proposed uncertainty-guided sampling prevents the bias in training towards large planar surfaces and improves the quality of prediction, especially near object boundaries and on small structures. Experimental results show that the proposed method outperforms the state-of-the-art in ScanNet [4] and NYUv2 [33], and that the estimated uncertainty correlates well with the prediction error. Code is available at https://github.com/baegwangbin/surface_normal_uncertainty. Gwangbin Bae, Ignas Budvytis, Roberto Cipolla |
ICCV | 3 |
| 2021 | FIERY: Future Instance Prediction in Bird's-Eye View from Surround Monocular CamerasabstractDriving requires interacting with road agents and predicting their future behaviour in order to navigate safely. We present FIERY: a probabilistic future prediction model in bird’s-eye view from monocular cameras. Our model predicts future instance segmentation and motion of dynamic agents that can be transformed into non-parametric future trajectories. Our approach combines the perception, sensor fusion and prediction components of a traditional autonomous driving stack by estimating bird’s-eye-view prediction directly from surround RGB monocular camera inputs. FIERY learns to model the inherent stochastic nature of the future solely from camera driving data in an end-to-end manner, without relying on HD maps, and predicts multimodal future trajectories. We show that our model outperforms previous prediction baselines on the NuScenes and Lyft datasets. The code and trained models are available at https://github.com/wayveai/fiery. Anthony Hu 0001, Zak Murez, Nikhil Mohan, Sofía Dudas, Jeffrey Hawke, Vijay Badrinarayanan, Roberto Cipolla, Alex Kendall |
ICCV | 7 |
| 2021 | PX-NET: Simple and Efficient Pixel-Wise Training of Photometric Stereo NetworksabstractRetrieving accurate 3D reconstructions of objects from the way they reflect light is a very challenging task in computer vision. Despite more than four decades since the definition of the Photometric Stereo problem, most of the literature has had limited success when global illumination effects such as cast shadows, self-reflections and ambient light come into play, especially for specular surfaces. Recent approaches have leveraged the capabilities of deep learning in conjunction with computer graphics in order to cope with the need of a vast number of training data to invert the image irradiance equation and retrieve the geometry of the object. However, rendering global illumination effects is a slow process which can limit the amount of training data that can be generated.In this work we propose a novel pixel-wise training procedure for normal prediction by replacing the training data (observation maps) of globally rendered images with independent per-pixel generated data. We show that global physical effects can be approximated on the observation map domain and this simplifies and speeds up the data creation procedure. Our network, PX-NET, achieves state-of-the-art performance compared to other pixelwise methods on synthetic datasets, as well as the DiLiGenT real dataset on both dense and sparse light settings. Fotios Logothetis, Ignas Budvytis, Roberto Mecca, Roberto Cipolla |
ICCV | 4 |
| 2021 | Hierarchical Kinematic Probability Distributions for 3D Human Shape and Pose Estimation from Images in the WildabstractThis paper addresses the problem of 3D human body shape and pose estimation from an RGB image. This is often an ill-posed problem, since multiple plausible 3D bodies may match the visual evidence present in the input - particularly when the subject is occluded. Thus, it is desirable to estimate a distribution over 3D body shape and pose conditioned on the input image instead of a single 3D re-construction. We train a deep neural network to estimate a hierarchical matrix-Fisher distribution over relative 3D joint rotation matrices (i.e. body pose), which exploits the human body’s kinematic tree structure, as well as a Gaussian distribution over SMPL body shape parameters. To further ensure that the predicted shape and pose distributions match the visual evidence in the input image, we implement a differentiable rejection sampler to impose a reprojection loss between ground-truth 2D joint coordinates and samples from the predicted distributions, projected onto the image plane. We show that our method is competitive with the state-of-the-art in terms of 3D shape and pose metrics on the SSP-3D and 3DPW datasets, while also yielding a structured probability distribution over 3D body shape and pose, with which we can meaningfully quantify prediction uncertainty and sample multiple plausible 3D reconstructions to explain a given input image. Akash Sengupta, Ignas Budvytis, Roberto Cipolla |
ICCV | 3 |
| 2021 | Scaling digital screen reading with one-shot learning and re-identificationabstractUsing only a mobile phone app, our objective is to cheaply retro-fit digital meters (e.g blood pressure, blood glucose or industrial gauges) with `smart' data transfer capabilities. Using the mobile phone camera we build an app to securely and accurately transcribe information from digital meter screens. Only a single labelled training image of a target meter is required to build a custom screen reading module. Here we show how this can scale to potentially hundreds of different meters by learning to recognising the meter type so that the reading module can be automatically selected. This makes the system very easy for a user who would need to scan multiple different meter types. To this end, we build a CNN based system which runs in real-time on mobile device with very high read accuracy and meter recognition. Our contributions include (i) a method of one-shot training by synthesis through domain shift reduction, (ii) a deep embedding network for scale, translation and rotation invariant re-identification of digital meters, (iii) a highly accurate and efficient mobile phone app for recognising and parsing digital meter screens and (iv) release of a new digital meter re-identification dataset. James Charles, Stefano Bucciarelli, Roberto Cipolla |
WACV | 3 |
| 2020 | FootNet: An Efficient Convolutional Network for Multiview 3D Foot Reconstruction
Felix Kok, James Charles, Roberto Cipolla |
ACCV (6) | 3 |
| 2020 | Efficient Large-Scale Semantic Visual Localization in 2D Maps
Tomás Vojír, Ignas Budvytis, Roberto Cipolla |
ACCV (3) | 3 |
| 2020 | Rotation Equivariant Orientation Estimation for Omnidirectional Localization
Chao Zhang 0023, Ignas Budvytis, Stephan Liwicki, Roberto Cipolla |
ACCV (4) | 4 |
| 2020 | Real-time screen reading: reducing domain shift for one-shot learning
James Charles, Stefano Bucciarelli, Roberto Cipolla |
BMVC | 3 |
| 2020 | A CNN Based Approach for the Near-Field Photometric Stereo Problem
Fotios Logothetis, Ignas Budvytis, Roberto Mecca, Roberto Cipolla |
BMVC | 4 |
| 2020 | Synthetic Training for Accurate 3D Human Pose and Shape Estimation in the Wild
Akash Sengupta, Roberto Cipolla, Ignas Budvytis |
BMVC | 2 |
| 2020 | Predicting Semantic Map Representations From Images Using Pyramid Occupancy NetworksabstractAutonomous vehicles commonly rely on highly detailed birds-eye-view maps of their environment, which capture both static elements of the scene such as road layout as well as dynamic elements such as other cars and pedestrians. Generating these map representations on the fly is a complex multi-stage process which incorporates many important vision-based elements, including ground plane estimation, road segmentation and 3D object detection. In this work we present a simple, unified approach for estimating these map representations directly from monocular images using a single end-to-end deep learning architecture. For the maps themselves we adopt a semantic Bayesian occupancy grid framework, allowing us to trivially accumulate information over multiple cameras and timesteps. We demonstrate the effectiveness of our approach by evaluating against several challenging baselines on the NuScenes and Argoverse datasets, and show that we are able to achieve a relative improvement of 9.1% and 22.3% respectively compared to the best-performing existing method. Thomas Roddick, Roberto Cipolla |
CVPR | 2 |
| 2020 | Who Left the Dogs Out? 3D Animal Reconstruction with Expectation Maximization in the Loop
Benjamin Biggs, Oliver Boyne, James Charles, Andrew W. Fitzgibbon, Roberto Cipolla |
ECCV (11) | 5 |
| 2020 | Deep Multi-view Stereo for Dense 3D Reconstruction from Monocular Endoscopic Video
Gwangbin Bae, Ignas Budvytis, Chung-Kwong Yeung, Roberto Cipolla |
MICCAI (3) | 4 |
| 2019 | Large scale joint semantic re-localisation and scene understanding via globally unique instance coordinate regression
Ignas Budvytis, Marvin Teichmann, Tomás Vojír, Roberto Cipolla |
BMVC | 4 |
| 2019 | Fast-SCNN: Fast Semantic Segmentation Network
Rudra P. K. Poudel, Stephan Liwicki, Roberto Cipolla |
BMVC | 3 |
| 2019 | Orthographic Feature Transform for Monocular 3D Object Detection
Thomas Roddick, Alex Kendall, Roberto Cipolla |
BMVC | 3 |
| 2019 | Convolutional CRFs for Semantic Segmentation
Marvin Teichmann, Roberto Cipolla |
BMVC | 2 |
| 2019 | A Differential Volumetric Approach to Multi-View Photometric StereoabstractHighly accurate 3D volumetric reconstruction is still an open research topic where the main difficulty is usually related to merging some rough estimations with high frequency details. One of the most promising methods is the fusion between multi-view stereo and photometric stereo images. Beside the intrinsic difficulties that multi-view stereo and photometric stereo in order to work reliably, supplementary problems arise when considered together. In this work, we present a volumetric approach to the multi-view photometric stereo problem. The key point of our method is the signed distance field parameterisation and its relation to the surface normal. This is exploited in order to obtain a linear partial differential equation which is solved in a variational framework, that combines multiple images from multiple points of view in a single system. In addition, the volumetric approach is naturally implemented on an octree, which allows for fast ray-tracing that reliably alleviates occlusions and cast shadows. Our approach is evaluated on synthetic and real data-sets and achieves state-of-the-art results. Fotios Logothetis, Roberto Mecca, Roberto Cipolla |
ICCV | 3 |
| 2019 | Orientation-Aware Semantic Segmentation on Icosahedron SpheresabstractWe address semantic segmentation on omnidirectional images, to leverage a holistic understanding of the surrounding scene for applications like autonomous driving systems. For the spherical domain, several methods recently adopt an icosahedron mesh, but systems are typically rotation invariant or require significant memory and parameters, thus enabling execution only at very low resolutions. In our work, we propose an orientation-aware CNN framework for the icosahedron mesh. Our representation allows for fast network operations, as our design simplifies to standard network operations of classical CNNs, but under consideration of north-aligned kernel convolutions for features on the sphere. We implement our representation and demonstrate its memory efficiency up-to a level-8 resolution mesh (equivalent to 640 x 1024 equirectangular images). Finally, since our kernels operate on the tangent of the sphere, standard feature weights, pretrained on perspective data, can be directly transferred with only small need for weight refinement. In our evaluation our orientation-aware CNN becomes a new state of the art for the recent 2D3DS dataset, and our Omni-SYNTHIA version of SYNTHIA. Rotation invariant classification and segmentation tasks are additionally presented for comparison to prior art. Chao Zhang 0023, Stephan Liwicki, William A. P. Smith, Roberto Cipolla |
ICCV | 4 |
| 2019 | A Differential Approach to Shape from Polarisation: A Level-Set CharacterisationabstractDespite the longtime research aimed at retrieving geometrical information of an object from polarimetric imaging, physical limitations in the polarisation phenomena constrain current approaches to provide ambiguous depth estimation. As an additional constraint, polarimetric imaging formulation differs when light is reflected off the object specularly or diffusively. This introduces another source of ambiguity that current formulations cannot overcome. With the aim of deriving a formulation capable of dealing with as many heterogeneous effects as possible, we propose a differential formulation of the Shape from Polarisation problem that depends only on polarimetric images. This allows the direct geometrical characterisation of the level-set of the object keeping consistent mathematical formulation for diffuse and specular reflection. We show via synthetic and real-world experiments that diffuse and specular reflection can be easily distinguished in order to extract meaningful geometrical features from just polarimetric imaging. The inherent ambiguity of the Shape from Polarization problem becomes evident through the impossibility of reconstructing the whole surface with this differential approach. To overcome this limitation, we consider shading information elegantly embedding this new formulation into a two-light calibrated photometric stereo approach.. Fotios Logothetis, Roberto Mecca, Fiorella Sgallari, Roberto Cipolla |
Int. J. Comput. Vis. | 4 |
| 2018 | Creatures Great and SMAL: Recovering the Shape and Motion of Animals from Video
Benjamin Biggs, Thomas Roddick, Andrew W. Fitzgibbon, Roberto Cipolla |
ACCV (5) | 4 |
| 2018 | Semantic Localisation via Globally Unique Instance Segmentation
Ignas Budvytis, Patrick Sauer, Roberto Cipolla |
BMVC | 3 |
| 2018 | Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and SemanticsabstractNumerous deep learning applications benefit from multitask learning with multiple regression and classification objectives. In this paper we make the observation that the performance of such systems is strongly dependent on the relative weighting between each task's loss. Tuning these weights by hand is a difficult and expensive process, making multi-task learning prohibitive in practice. We propose a principled approach to multi-task deep learning which weighs multiple loss functions by considering the homoscedastic uncertainty of each task. This allows us to simultaneously learn various quantities with different units or scales in both classification and regression settings. We demonstrate our model learning per-pixel depth regression, semantic and instance segmentation from a monocular input image. Perhaps surprisingly, we show our model can learn multi-task weightings and outperform separate models trained individually on each task. Alex Kendall, Yarin Gal, Roberto Cipolla |
CVPR | 3 |
| 2018 | Adaptation of an Expressive Single Speaker Deep Neural Network Speech Synthesis SystemabstractOne of the advantages of statistical parametric speech synthesis is the ability to alter some of the characteristics of the speech e.g. change the speaker, expression etc. In this paper we present a technique to adapt an expressive single speaker deep neural network (DNN) speech synthesis model to a new speaker, allowing for both neutral and expressive speech in the new speaker's voice. Experiments show that the proposed adaptation technique achieves higher MOS scores on both neutral and expressive speech, and higher speaker similarity and slightly lower expression similarity scores on the expressive speech when compared with another DNN speaker adaptation technique. Jonathan Parker, Yannis Stylianou, Roberto Cipolla |
ICASSP | 3 |
| 2018 | MultiNet: Real-time Joint Semantic Reasoning for Autonomous DrivingabstractWhile most approaches to semantic reasoning have focused on improving performance, in this paper we argue that computational times are very important in order to enable real time applications such as autonomous driving. Towards this goal, we present an approach to joint classification, detection and semantic segmentation using a unified architecture where the encoder is shared amongst the three tasks. Our approach is very simple, can be trained end-to-end and performs extremely well in the challenging KITTI dataset. Our approach is also very efficient, allowing us to perform inference at more then 23 frames per second. Training scripts and trained weights to reproduce our results can be found here: https://github.com/MarvinTeichmann/MultiNet Marvin Teichmann, Michael Weber 0009, Johann Marius Zöllner, Roberto Cipolla, Raquel Urtasun |
Intelligent Vehicles Symposium | 4 |
| 2017 | Real-time Factored ConvNets: Extracting the X Factor in Human Parsing
James Charles, Ignas Budvytis, Roberto Cipolla |
BMVC | 3 |
| 2017 | Bayesian SegNet: Model Uncertainty in Deep Convolutional Encoder-Decoder Architectures for Scene Understanding
Alex Kendall, Vijay Badrinarayanan, Roberto Cipolla |
BMVC | 3 |
| 2017 | A Differential Approach to Shape from Polarization
Roberto Mecca, Fotios Logothetis, Roberto Cipolla |
BMVC | 3 |
| 2017 | Indirect deep structured learning for 3D human body shape and pose prediction
Vince Tan, Ignas Budvytis, Roberto Cipolla |
BMVC | 3 |
| 2017 | Deep Roots: Improving CNN Efficiency with Hierarchical Filter GroupsabstractWe propose a new method for creating computationally efficient and compact convolutional neural networks (CNNs) using a novel sparse connection structure that resembles a tree root. This allows a significant reduction in computational cost and number of parameters compared to state-of-the-art deep CNNs, without compromising accuracy, by exploiting the sparsity of inter-layer filter dependencies. We validate our approach by using it to train more efficient variants of state-of-the-art CNN architectures, evaluated on the CIFAR10 and ILSVRC datasets. Our results show similar or higher accuracy than the baseline architectures with much less computation, as measured by CPU and GPU timings. For example, for ResNet 50, our model has 40% fewer parameters, 45% fewer floating point operations, and is 31% (12%) faster on a CPU (GPU). For the deeper ResNet 200 our model has 48% fewer parameters and 27% fewer floating point operations, while maintaining state-of-the-art accuracy. For GoogLeNet, our model has 7% fewer parameters and is 21% (16%) faster on a CPU (GPU). Yani Ioannou, Duncan P. Robertson, Roberto Cipolla, Antonio Criminisi |
CVPR | 3 |
| 2017 | Geometric Loss Functions for Camera Pose Regression with Deep LearningabstractDeep learning has shown to be effective for robust and real-time monocular image relocalisation. In particular, PoseNet [22] is a deep convolutional neural network which learns to regress the 6-DOF camera pose from a single image. It learns to localize using high level features and is robust to difficult lighting, motion blur and unknown camera intrinsics, where point based SIFT registration fails. However, it was trained using a naive loss function, with hyper-parameters which require expensive tuning. In this paper, we give the problem a more fundamental theoretical treatment. We explore a number of novel loss functions for learning camera pose which are based on geometry and scene reprojection error. Additionally we show how to automatically learn an optimal weighting to simultaneously regress position and orientation. By leveraging geometry, we demonstrate that our technique significantly improves PoseNets performance across datasets ranging from indoor rooms to a small city. Alex Kendall, Roberto Cipolla |
CVPR | 2 |
| 2017 | Semi-Calibrated Near Field Photometric Stereoabstract3D Reconstruction from shading information through Photometric Stereo is considered a very challenging problem in Computer Vision. Although this technique can potentially provide highly detailed shape recovery, its accuracy is critically dependent on a numerous set of factors among them the reliability of the light sources in emitting a constant amount of light. In this work, we propose a novel variational approach to solve the so called semi-calibrated near field Photometric Stereo problem, where the positions but not the brightness of the light sources are known. Additionally, we take into account realistic modeling features such as perspective viewing geometry and heterogeneous scene composition, containing both diffuse and specular objects. Furthermore, we also relax the point light source assumption that usually constraints the near field formulation by explicitly calculating the light attenuation maps. Synthetic experiments are performed for quantitative evaluation for a wide range of cases whilst real experiments provide comparisons, qualitatively outperforming the state of the art. Fotios Logothetis, Roberto Mecca, Roberto Cipolla |
CVPR | 3 |
| 2017 | Expressive visual text to speech and expression adaptation using deep neural networksabstractIn this paper, we present an expressive visual text to speech system (VTTS) based on a deep neural network (DNN). Given an input text sentence and a set of expression tags, the VTTS is able to produce not only the audio speech, but also the accompanying facial movements. The expressions can either be one of the expressions in the training corpus or a blend of expressions from the training corpus. Furthermore, we present a method of adapting a previously trained DNN to include a new expression using a small amount of training data. Experiments show that the proposed DNN-based VTTS is preferred by 57.9% over the baseline hidden Markov model based VTTS which uses cluster adaptive training. Jonathan Parker, Ranniery Maia, Yannis Stylianou, Roberto Cipolla |
ICASSP | 4 |
| 2017 | Concrete Problems for Autonomous Vehicle Safety: Advantages of Bayesian Deep LearningabstractAutonomous vehicle (AV) software is typically composed of a pipeline of individual components, linking sensor inputs to motor outputs. Erroneous component outputs propagate downstream, hence safe AV software must consider the ultimate effect of each component’s errors. Further, improving safety alone is not sufficient. Passengers must also feel safe to trust and use AV systems. To address such concerns, we investigate three under-explored themes for AV research: safety, interpretability, and compliance. Safety can be improved by quantifying the uncertainties of component outputs and propagating them forward through the pipeline. Interpretability is concerned with explaining what the AV observes and why it makes the decisions it does, building reassurance with the passenger. Compliance refers to maintaining some control for the passenger. We discuss open challenges for research within these themes. We highlight the need for concrete evaluation metrics, propose example problems, and highlight possible solutions. Rowan McAllister, Yarin Gal, Alex Kendall, Mark van der Wilk, Amar Shah 0001, Roberto Cipolla, Adrian Weller |
IJCAI | 6 |
| 2017 | SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image SegmentationabstractWe present a novel and practical deep fully convolutional neural network architecture for semantic pixel-wise segmentation termed SegNet. This core trainable segmentation engine consists of an encoder network, a corresponding decoder network followed by a pixel-wise classification layer. The architecture of the encoder network is topologically identical to the 13 convolutional layers in the VGG16 network [1] . The role of the decoder network is to map the low resolution encoder feature maps to full input resolution feature maps for pixel-wise classification. The novelty of SegNet lies is in the manner in which the decoder upsamples its lower resolution input feature map(s). Specifically, the decoder uses pooling indices computed in the max-pooling step of the corresponding encoder to perform non-linear upsampling. This eliminates the need for learning to upsample. The upsampled maps are sparse and are then convolved with trainable filters to produce dense feature maps. We compare our proposed architecture with the widely adopted FCN [2] and also with the well known DeepLab-LargeFOV [3] , DeconvNet [4] architectures. This comparison reveals the memory versus accuracy trade-off involved in achieving good segmentation performance. SegNet was primarily motivated by scene understanding applications. Hence, it is designed to be efficient both in terms of memory and computational time during inference. It is also significantly smaller in the number of trainable parameters than other competing architectures and can be trained end-to-end using stochastic gradient descent. We also performed a controlled benchmark of SegNet and other architectures on both road scenes and SUN RGB-D indoor scene segmentation tasks. These quantitative assessments show that SegNet provides good performance with competitive inference time and most efficient inference memory-wise as compared to other architectures. We also provide a Caffe implementation of SegNet and a web demo at http://mi.eng.cam.ac.uk/projects/segnet. Vijay Badrinarayanan, Alex Kendall, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | Near-Field Photometric Stereo in Ambient Light
Fotios Logothetis, Roberto Mecca, Yvain Quéau, Roberto Cipolla |
BMVC | 4 |
| 2016 | Understanding RealWorld Indoor Scenes with Synthetic DataabstractScene understanding is a prerequisite to many high level tasks for any automated intelligent machine operating in real world environments. Recent attempts with supervised learning have shown promise in this direction but also highlighted the need for enormous quantity of supervised data- performance increases in proportion to the amount of data used. However, this quickly becomes prohibitive when considering the manual labour needed to collect such data. In this work, we focus our attention on depth based semantic per-pixel labelling as a scene understanding problem and show the potential of computer graphics to generate virtually unlimited labelled data from synthetic 3D scenes. By carefully synthesizing training data with appropriate noise models we show comparable performance to state-of-the-art RGBD systems on NYUv2 dataset despite using only depth data as input and set a benchmark on depth-based segmentation on SUN RGB-D dataset. Ankur Handa, Viorica Patraucean, Vijay Badrinarayanan, Simon Stent, Roberto Cipolla |
CVPR | 5 |
| 2016 | Refining Architectures of Deep Convolutional Neural NetworksabstractDeep Convolutional Neural Networks (CNNs) have recently evinced immense success for various image recognition tasks [11, 27]. However, a question of paramount importance is somewhat unanswered in deep learning research - is the selected CNN optimal for the dataset in terms of accuracy and model size? In this paper, we intend to answer this question and introduce a novel strategy that alters the architecture of a given CNN for a specified dataset, to potentially enhance the original accuracy while possibly reducing the model size. We use two operations for architecture refinement, viz. stretching and symmetrical splitting. Stretching increases the number of hidden units (nodes) in a given CNN layer, while a symmetrical split of say K between two layers separates the input and output channels into K equal groups, and connects only the corresponding input-output channel groups. Our procedure starts with a pre-trained CNN for a given dataset, and optimally decides the stretch and split factors across the network to refine the architecture. We empirically demonstrate the necessity of the two operations. We evaluate our approach on two natural scenes attributes datasets, SUN Attributes [16] and CAMIT-NSAD [20], with architectures of GoogleNet and VGG-11, that are quite contrasting in their construction. We justify our choice of datasets, and show that they are interestingly distinct from each other, and together pose a challenge to our architectural refinement algorithm. Our results substantiate the usefulness of the proposed method. Sukrit Shankar, Duncan P. Robertson, Yani Ioannou, Antonio Criminisi, Roberto Cipolla |
CVPR | 5 |
| 2016 | Projective Bundle Adjustment from Arbitrary Initialization Using the Variable Projection Method
Je Hyeong Hong, Christopher Zach, Andrew W. Fitzgibbon, Roberto Cipolla |
ECCV (1) | 4 |
| 2016 | SceneNet: An annotated model generator for indoor scene understandingabstractWe introduce SceneNet, a framework for generating high-quality annotated 3D scenes to aid indoor scene understanding. SceneNet leverages manually-annotated datasets of real world scenes such as NYUv2 to learn statistics about object co-occurrences and their spatial relationships. Using a hierarchical simulated annealing optimisation, these statistics are exploited to generate a potentially unlimited number of new annotated scenes, by sampling objects from various existing databases of 3D objects such as ModelNet, and textures such as OpenSurfaces and ArchiveTextures. Depending on the task, SceneNet can be used directly in the form of annotated 3D models for supervised training and 3D reconstruction benchmarking, or in the form of rendered annotated sequences of RGB-D frames or videos. Ankur Handa, Viorica Patraucean, Simon Stent, Roberto Cipolla |
ICRA | 4 |
| 2016 | Modelling uncertainty in deep learning for camera relocalizationabstractWe present a robust and real-time monocular six degree of freedom visual relocalization system. We use a Bayesian convolutional neural network to regress the 6-DOF camera pose from a single RGB image. It is trained in an end-to-end manner with no need of additional engineering or graph optimisation. The algorithm can operate indoors and outdoors in real time, taking under 6ms to compute. It obtains approximately 2m and 6° accuracy for very large scale outdoor scenes and 0.5m and 10° accuracy indoors. Using a Bayesian convolutional neural network implementation we obtain an estimate of the model's relocalization uncertainty and improve state of the art localization accuracy on a large scale outdoor dataset. We leverage the uncertainty measure to estimate metric relocalization error and to detect the presence or absence of the scene in the input image. We show that the model's uncertainty is caused by images being dissimilar to the training dataset in either pose or appearance. Alex Kendall, Roberto Cipolla |
ICRA | 2 |
| 2016 | Precise deterministic change detection for smooth surfacesabstractWe introduce a precise deterministic approach for pixel-wise change detection in images taken of a scene of interest over time. Our motivation is for applications such as artefact condition monitoring and structural inspection, where a common problem is the need to efficiently and accurately identify subtle signs of damage and deterioration. The approach we describe is designed to compensate for the three most common sources of nuisance variation encountered when tackling the problem of change detection, namely: viewpoint variation due to camera motion between images, photometric variation due to lighting differences, and changes in image resolution/focal settings. To tackle viewpoint variation, particularly in areas of low texture, we propose the use of the generalised PatchMatch (PM) correspondence algorithm to compute a dense flow field. The flow field is regularized using a Thin Plate Spline (TPS) model which assumes a smooth underlying geometry and allows registration to be interpolated precisely through areas of low texture or uncertain flow. To compensate for low-frequency lighting variation, we fit a second TPS model to the photometric differences between registered images. Finally, to account for changes in focal settings, we estimate and apply a blurring kernel via optimisation over image differences. We provide a thorough evaluation of the performance of our method on an illustrative toy dataset and on two recent, real-world inspection datasets. Our approach performs favourably versus state-of-the-art baselines in both cases, while remaining relatively transparent to understand and simple to compute. Simon Stent, Riccardo Gherardi, Björn Stenger, Roberto Cipolla |
WACV | 4 |
| 2016 | Expressive visual text-to-speech as an assistive technology for individuals with autism spectrum conditionsabstractAdults with Autism Spectrum Conditions (ASC) experience marked difficulties in recognising the emotions of others and responding appropriately. The clinical characteristics of ASC mean that face to face or group interventions may not be appropriate for this clinical group. This article explores the potential of a new interactive technology, converting text to emotionally expressive speech, to improve emotion processing ability and attention to faces in adults with ASC. We demonstrate a method for generating a near-videorealistic avatar (XpressiveTalk), which can produce a video of a face uttering inputted text, in a large variety of emotional tones. We then demonstrate that general population adults can correctly recognize the emotions portrayed by XpressiveTalk. Adults with ASC are significantly less accurate than controls, but still above chance levels for inferring emotions from XpressiveTalk. Both groups are significantly more accurate when inferring sad emotions from XpressiveTalk compared to the original actress, and rate these expressions as significantly more preferred and realistic. The potential applications for XpressiveTalk as an assistive technology for adults with ASC is discussed. Sarah A. Cassidy, Björn Stenger, L. Van Dongen, Kayoko Yanagisawa, Vincent Wan, Simon Baron-Cohen, Roberto Cipolla |
Comput. Vis. Image Underst. | 8 |
| 2016 | Visual change detection on tunnel linings
Simon Stent, Riccardo Gherardi, Björn Stenger, Kenichi Soga, Roberto Cipolla |
Mach. Vis. Appl. | 5 |
| 2016 | A Single-Lobe Photometric Stereo Approach for Heterogeneous MaterialabstractShape from shading with multiple light sources is an active research area, and a diverse range of approaches have been proposed in recent decades. However, devising a robust reconstruction technique still remains a challenging goal, as the image acquisition process is highly nonlinear. Recent Photometric Stereo variants rely on simplifying assumptions in order to make the problem solvable: light propagation is still commonly assumed to be uniform, and the Bidirectional Reflectance Distribution Function is assumed to be diffuse, with limited interest for specular materials. In this work, we introduce a well-posed formulation based on partial differential equations (PDEs) for a unified reflectance function that can model both diffuse and specular reflections. We base our derivation on ratio of images, which makes the model independent from photometric invariants and yields a well-posed differential problem based on a system of quasi-linear PDEs with discontinuous coefficients. In addition, we directly solve a differential problem for the unknown depth, thus avoiding the intermediate step of approximating the normal field. A variational approach is presented ensuring robustness to noise and outliers (such as black shadows), and this is confirmed with a wide range of experiments on both synthetic and real data, where we compare favorably to the state of the art. Roberto Mecca, Yvain Quéau, Fotios Logothetis, Roberto Cipolla |
SIAM J. Imaging Sci. | 4 |
| 2015 | Detecting Change for Multi-View, Long-Term Surface InspectionabstractWe describe a system for the detection of changes in multiple views of a tunnel surface. From data gathered by a robotic inspection rig, we use a structure-from-motion pipeline to build panoramas of the surface and register images from different time instances. Reliably detecting changes such as hairline cracks, water ingress and other surface damage between the registered images is a challenging problem: achieving the best possible performance for a given set of data requires sub-pixel precision and careful modelling of the noise sources. The task is further complicated by factors such as unavoidable registration error and changes in image sensors, capture settings and lighting. Our contribution is a novel approach to change detection using a two-channel convolutional neural network. The network accepts pairs of approximately registered image patches taken at different times and classifies them to detect anomalous changes. To train the network, we take advantage of synthetically generated training examples and the homogeneity of the tunnel surfaces to eliminate most of the manual labelling effort. We evaluate our method on field data gathered from a live tunnel over several months, demonstrating it to outperform existing approaches from recent literature and industrial practice. Simon Stent, Riccardo Gherardi, Björn Stenger, Roberto Cipolla |
BMVC | 4 |
| 2015 | DEEP-CARVING: Discovering visual attributes by carving deep neural netsabstractMost of the approaches for discovering visual attributes in images demand significant supervision, which is cumbersome to obtain. In this paper, we aim to discover visual attributes in a weakly supervised setting that is commonly encountered with contemporary image search engines. For instance, given a noun (say forest) and its associated attributes (say dense, sunlit, autumn), search engines can now generate many valid images for any attribute-noun pair (dense forests, autumn forests, etc). However, images for an attributenoun pair do not contain any information about other attributes (like which forests in the autumn are dense too). Thus, a weakly supervised scenario occurs: each of the M attributes corresponds to a class such that a training image in class m ∈ {1, . . . , M} contains a single label that indicates the presence of the mth attribute only. The task is to discover all the attributes present in a test image. Deep Convolutional Neural Networks (CNNs) [20] have enjoyed remarkable success in vision applications recently. However, in a weakly supervised scenario, widely used CNN training procedures do not learn a robust modelfor predicting multiple attribute labels simultaneously. The primary reason is that the attributes highly co-occur within the training data, and unlike objects, do not generally exist as well-defined spatial boundaries within the image. To ameliorate this limitation, we propose Deep-Carving, a novel training procedure with CNNs, that helps the net efficiently carve itselffor the task of multiple attribute prediction. During training, the responses of the feature maps are exploited in an ingenious way to provide the net with multiple pseudo-labels (for training images) for subsequent iterations. The process is repeated periodically after a fixed number of iterations, and enables the net carve itself iteratively for efficiently disentangling features. Additionally, we contribute a noun-adjective pairing inspired Natural Scenes Attributes Dataset to the research community, CAMITNSAD, containing a number of co-occurring attributes within a noun category. We describe, in detail, salient aspects of this dataset. Our experiments on CAMITNSAD and the SUN Attributes Dataset [29], with weak supervision, clearly demonstrate that the Deep-Carved CNNs consistently achieve considerable improvement in the precision of attribute prediction over popular baseline methods. Sukrit Shankar, Vikas Garg 0001, Roberto Cipolla |
CVPR | 3 |
| 2015 | PoseNet: A Convolutional Network for Real-Time 6-DOF Camera RelocalizationabstractWe present a robust and real-time monocular six degree of freedom relocalization system. Our system trains a convolutional neural network to regress the 6-DOF camera pose from a single RGB image in an end-to-end manner with no need of additional engineering or graph optimisation. The algorithm can operate indoors and outdoors in real time, taking 5ms per frame to compute. It obtains approximately 2m and 3 degrees accuracy for large scale outdoor scenes and 0.5m and 5 degrees accuracy indoors. This is achieved using an efficient 23 layer deep convnet, demonstrating that convnets can be used to solve complicated out of image plane regression problems. This was made possible by leveraging transfer learning from large scale classification data. We show that the PoseNet localizes from high level features and is robust to difficult lighting, motion blur and different camera intrinsics where point based SIFT registration fails. Furthermore we show how the pose feature that is produced generalizes to other scenes allowing us to regress pose with only a few dozen training examples. Alex Kendall, Matthew Koichi Grimes, Roberto Cipolla |
ICCV | 3 |
| 2015 | Distances and Means of Direct Similarities
Minh-Tri Pham, Oliver J. Woodford, Frank Perbet, Atsuto Maki, Riccardo Gherardi, Björn Stenger, Roberto Cipolla |
Int. J. Comput. Vis. | 7 |
| 2014 | Bi-label Propagation for Generic Multiple Object TrackingabstractIn this paper, we propose a label propagation framework to handle the multiple object tracking (MOT) problem for a generic object type (cf. pedestrian tracking). Given a target object by an initial bounding box, all objects of the same type are localized together with their identities. We treat this as a problem of propagating bi-labels, i.e. a binary class label for detection and individual object labels for tracking. To propagate the class label, we adopt clustered Multiple Task Learning (cMTL) while enforcing spatio-temporal consistency and show that this improves the performance when given limited training data. To track objects, we propagate labels from trajectories to detections based on affinity using appearance, motion, and context. Experiments on public and challenging new sequences show that the proposed method improves over the current state of the art on this task. Wenhan Luo, Tae-Kyun Kim 0001, Björn Stenger, Roberto Cipolla |
CVPR | 5 |
| 2014 | Robust Instance Recognition in Presence of Occlusion and Clutter
Ujwal Bonde, Vijay Badrinarayanan, Roberto Cipolla |
ECCV (2) | 3 |
| 2014 | Part Bricolage: Flow-Assisted Part-Based Graphs for Detecting Activities in Videos
Sukrit Shankar, Vijay Badrinarayanan, Roberto Cipolla |
ECCV (6) | 3 |
| 2014 | Using Bounded Diameter Minimum Spanning Trees to Build Dense Active Appearance Models
Björn Stenger, Roberto Cipolla |
Int. J. Comput. Vis. | 3 |
| 2014 | Mixture of Trees Probabilistic Graphical Model for Video Segmentation
Vijay Badrinarayanan, Ignas Budvytis, Roberto Cipolla |
Int. J. Comput. Vis. | 3 |
| 2014 | Special Issue on Large-Scale Computer Vision: Geometry, Inference, and Learning
Roberto Cipolla, Carlo Colombo, Alberto Del Bimbo |
Int. J. Comput. Vis. | 1 |
| 2013 | Expressive Visual Text-to-Speech Using Active Appearance ModelsabstractThis paper presents a complete system for expressive visual text-to-speech (VTTS), which is capable of producing expressive output, in the form of a 'talking head', given an input text and a set of continuous expression weights. The face is modeled using an active appearance model (AAM), and several extensions are proposed which make it more applicable to the task of VTTS. The model allows for normalization with respect to both pose and blink state which significantly reduces artifacts in the resulting synthesized sequences. We demonstrate quantitative improvements in terms of reconstruction error over a million frames, as well as in large-scale user studies, comparing the output of different systems. Björn Stenger, Vincent Wan, Roberto Cipolla |
CVPR | 4 |
| 2013 | Unconstrained Monocular 3D Human Pose Estimation by Action Detection and Cross-Modality Regression ForestabstractThis work addresses the challenging problem of unconstrained 3D human pose estimation (HPE) from a novel perspective. Existing approaches struggle to operate in realistic applications, mainly due to their scene-dependent priors, such as background segmentation and multi-camera network, which restrict their use in unconstrained environments. We therfore present a framework which applies action detection and 2D pose estimation techniques to infer 3D poses in an unconstrained video. Action detection offers spatiotemporal priors to 3D human pose estimation by both recognising and localising actions in space-time. Instead of holistic features, e.g. silhouettes, we leverage the flexibility of deformable part model to detect 2D body parts as a feature to estimate 3D poses. A new unconstrained pose dataset has been collected to justify the feasibility of our method, which demonstrated promising results, significantly outperforming the relevant state-of-the-arts. Tsz-Ho Yu, Tae-Kyun Kim 0001, Roberto Cipolla |
CVPR | 3 |
| 2013 | Semantic Transform: Weakly Supervised Semantic Inference for Relating Visual AttributesabstractRelative (comparative) attributes are promising for thematic ranking of visual entities, which also aids in recognition tasks. However, attribute rank learning often requires a substantial amount of relational supervision, which is highly tedious, and apparently impractical for real-world applications. In this paper, we introduce the Semantic Transform, which under minimal supervision, adaptively finds a semantic feature space along with a class ordering that is related in the best possible way. Such a semantic space is found for every attribute category. To relate the classes under weak supervision, the class ordering needs to be refined according to a cost function in an iterative procedure. This problem is ideally NP-hard, and we thus propose a constrained search tree formulation for the same. Driven by the adaptive semantic feature space representation, our model achieves the best results to date for all of the tasks of relative, absolute and zero-shot classification on two popular datasets. Sukrit Shankar, Joan Lasenby, Roberto Cipolla |
ICCV | 3 |
| 2013 | Photo-realistic expressive text to talking head synthesis
Vincent Wan, Art Blokland, Norbert Braunschweiler, Langzhou Chen, BalaKrishna Kolluru, Javier Latorre, Ranniery Maia, Björn Stenger, Kayoko Yanagisawa, Yannis Stylianou, Masami Akamine, Mark J. F. Gales, Roberto Cipolla |
INTERSPEECH | 14 |
| 2013 | A Performance Evaluation of Volumetric 3D Interest Point Detectors
Tsz-Ho Yu, Oliver J. Woodford, Roberto Cipolla |
Int. J. Comput. Vis. | 3 |
| 2013 | Semi-Supervised Video Segmentation Using Tree Structured Graphical ModelsabstractWe present a novel patch-based probabilistic graphical model for semi-supervised video segmentation. At the heart of our model is a temporal tree structure that links patches in adjacent frames through the video sequence. This permits exact inference of pixel labels without resorting to traditional short time window-based video processing or instantaneous decision making. The input to our algorithm is labeled key frame(s) of a video sequence and the output is pixel-wise labels along with their confidences. We propose an efficient inference scheme that performs exact inference over the temporal tree, and optionally a per frame label smoothing step using loopy BP, to estimate pixel-wise labels and their posteriors. These posteriors are used to learn pixel unaries by training a Random Decision Forest in a semi-supervised manner. These unaries are used in a second iteration of label inference to improve the segmentation quality. We demonstrate the efficacy of our proposed algorithm using several qualitative and quantitative tests on both foreground/background and multiclass video segmentation problems using publicly available and our own datasets. Vijay Badrinarayanan, Ignas Budvytis, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2013 | Achieving robust face recognition from video by combining a weak photometric model and a learnt generic face invariant
Ognjen Arandjelovic, Roberto Cipolla |
Pattern Recognit. | 2 |
| 2013 | Detecting bipedal motion from correlated probabilistic trajectories
Atsuto Maki, Frank Perbet, Björn Stenger, Roberto Cipolla |
Pattern Recognit. Lett. | 4 |
| 2012 | Video Segmentation with Superpixels
Fabio Galasso, Roberto Cipolla, Bernt Schiele |
ACCV (1) | 2 |
| 2012 | Dense Active Appearance Models Using a Bounded Diameter Minimum Spanning TreeabstractWe present a method for producing dense Active Appearance Models (AAMs), suitable for video-realistic synthesis. To this end we estimate a joint alignment of all training images using a set of pairwise registrations and ensure that these pairwise registrations are only calculated between similar images. This is achieved by defining a graph on the image set whose edge weights correspond to registration errors and computing a bounded diameter minimum spanning tree (BDMST). Dense optical flow is used to compute pairwise registration and we introduce a flow refinement method to align small scale texture. Once registration between training images has been established we propose a method to add vertices to the AAM in a way that minimises error between the observed flow fields and a flow field interpolated between the AAM mesh points. We demonstrate a significant improvement in model compactness using the proposed method and show it dealing with cases that are problematic for current state-of-the-art approaches. Björn Stenger, Roberto Cipolla |
BMVC | 3 |
| 2012 | MoT - Mixture of Trees Probabilistic Graphical Model for Video SegmentationabstractWe present a novel mixture of trees (MoT) graphical model for video segmentation.Each component in this mixture represents a tree structured temporal linkage between super-pixels from the first to the last frame of a video sequence.Our time-series model explicitly captures the uncertainty in temporal linkage between adjacent frames which improves segmentation accuracy.We provide a variational inference scheme for this model to estimate super-pixel labels and their confidences in nearly realtime.The efficacy of our approach is demonstrated via quantitative comparisons on the challenging SegTrack joint segmentation and tracking dataset [23]. Ignas Budvytis, Vijay Badrinarayanan, Roberto Cipolla |
BMVC | 3 |
| 2012 | A unifying resolution-independent formulation for early visionabstractWe present a model for early vision tasks such as denoising, super-resolution, deblurring, and demosaicing. The model provides a resolution-independent representation of discrete images which admits a truly rotationally invariant prior. The model generalizes several existing approaches: variational methods, finite element methods, and discrete random fields. The primary contribution is a novel energy functional which has not previously been written down, which combines the discrete measurements from pixels with a continuous-domain world viewed through continous-domain point-spread functions. The value of the functional is that simple priors (such as total variation and generalizations) on the continous-domain world become realistic priors on the sampled images. We show that despite its apparent complexity, optimization of this model depends on just a few computational primitives, which although tedious to derive, can now be reused in many domains. We define a set of optimization algorithms which greatly overcome the apparent complexity of this model, and make possible its practical application. New experimental results include infinite-resolution upsampling, and a method for obtaining “subpixel superpixels”. Fabio Viola, Andrew W. Fitzgibbon, Roberto Cipolla |
CVPR | 3 |
| 2012 | 3D shape and its applicationsabstractSummary form only given. The talk will begin with an overview of the state-of-the-art in the 3R's of computer vision: registration, reconstruction and recognition and will include demonstrations of research which has been recently commercialised at Cambrdige (Zappar, Metail, Toshiba gesture interfaces and Microsoft Kinect). This will be followed by a review of more advanced techniques in multi-view stereo and photometric stereo for recovering accurate and complete 3D models from uncalibrated images. I will then look in more detail at techniques for recovering the shape of deforming objects such as the human body and face and the challenges of large scale reconstruction of outdoor scenes and the application to ageing infrastructure. Roberto Cipolla |
ICARCV | 1 |
| 2012 | Making a Shallow Network Deep: Conversion of a Boosting Classifier into a Decision Tree by Boolean Optimisation
Tae-Kyun Kim 0001, Ignas Budvytis, Roberto Cipolla |
Int. J. Comput. Vis. | 3 |
| 2011 | A Practical System for Modelling Body Shapes from Single View MeasurementsabstractThis paper describes an interactive system for quickly modelling 3D body shapes from a single image. It provides the user with a convenient way to obtain their 3D body shapes so as to try on virtual garments online. For the ease of use, we first introduce a novel interface for users to conveniently extract anthropometric measurements from a single photo, while using readily available scene cues for automatic image rectification. Then, we propose a unified probabilistic framework using Gaussian processes, which predict the body parameters from input measurements while correcting the aspect ratio ambiguity resulting from photo rectification. Extensive experiments and user studies have supported the efficacy of our system. This system is now being exploited commercially online1. © 2011. The copyright of this document resides with its authors. Yu Chen 0009, Duncan P. Robertson, Roberto Cipolla |
BMVC | 3 |
| 2011 | High-level scene structure using visibility and occlusionabstractWe demonstrate a new method for extracting high-level scene information from the type of data available from simultaneous localisation and mapping systems. We model the scene with a collection of primitives (such as bounded planes), and make explicit use of both visible and occluded points in order to refine the model. Since our formulation allows for different kinds of primitives and an arbitrary number of each, we use Bayesian model evidence to compare very different models on an even footing. Additionally, by making use of Bayesian techniques we can also avoid explicitly finding the optimal assignment of map landmarks to primitives. The results show that explicit reasoning about occlusion improves model accuracy and yields models which are suitable for aiding data association. © 2011. The copyright of this document resides with its authors. Paul McIlroy, Roberto Cipolla, Edward Rosten |
BMVC | 2 |
| 2011 | Semi-supervised video segmentation using tree structured graphical modelsabstractWe present a novel, implementation friendly and occlusion aware semi-supervised video segmentation algorithm using tree structured graphical models, which delivers pixel labels along with their uncertainty estimates. Our motivation to employ supervision is to tackle a task-specific segmentation problem where the semantic objects are pre-defined by the user. The video model we propose for this problem is based on a tree structured approximation of a patch based undirected mixture model, which includes a novel time-series and a soft label Random Forest classifier participating in a feedback mechanism. We demonstrate the efficacy of our model in cutting out foreground objects and multi-class segmentation problems in lengthy and complex road scene sequences. Our results have wide applicability, including harvesting labelled video data for training discriminative models, shape/pose/articulation learning and large scale statistical analysis to develop priors for video segmentation. Ignas Budvytis, Vijay Badrinarayanan, Roberto Cipolla |
CVPR | 3 |
| 2011 | Color photometric stereo for multicolored surfacesabstractWe present a multispectral photometric stereo method for capturing geometry of deforming surfaces. A novel photometric calibration technique allows calibration of scenes containing multiple piecewise constant chromaticities. This method estimates per-pixel photometric properties, then uses a RANSAC-based approach to estimate the dominant chromaticities in the scene. A likelihood term is developed linking surface normal, image intensity and photometric properties, which allows estimating the number of chromaticities present in a scene to be framed as a model estimation problem. The Bayesian Information Criterion is applied to automatically estimate the number of chromaticities present during calibration. A two-camera stereo system provides low resolution geometry, allowing the likelihood term to be used in segmenting new images into regions of constant chromaticity. This segmentation is carried out in a Markov Random Field framework and allows the correct photometric properties to be used at each pixel to estimate a dense normal map. Results are shown on several challenging real-world sequences, demonstrating state-of-the-art results using only two cameras and three light sources. Quantitative evaluation is provided against synthetic ground truth data. Björn Stenger, Roberto Cipolla |
ICCV | 3 |
| 2011 | Silhouette-based object phenotype recognition using 3D shape priorsabstractThis paper tackles the novel challenging problem of 3D object phenotype recognition from a single 2D silhouette. To bridge the large pose (articulation or deformation) and camera viewpoint changes between the gallery images and query image, we propose a novel probabilistic inference algorithm based on 3D shape priors. Our approach combines both generative and discriminative learning. We use latent probabilistic generative models to capture 3D shape and pose variations from a set of 3D mesh models. Based on these 3D shape priors, we generate a large number of projections for different phenotype classes, poses, and camera viewpoints, and implement Random Forests to efficiently solve the shape and pose inference problems. By model selection in terms of the silhouette coherency between the query and the projections of 3D shapes synthesized using the galleries, we achieve the phenotype recognition result as well as a fast approximate 3D reconstruction of the query. To verify the efficacy of the proposed approach, we present new datasets which contain over 500 images of various human and shark phenotypes and motions. The experimental results clearly show the benefits of using the 3D priors in the proposed method over previous 2D-based methods. Yu Chen 0009, Tae-Kyun Kim 0001, Roberto Cipolla |
ICCV | 3 |
| 2011 | Spatio-temporal clustering of probabilistic region trajectoriesabstractWe propose a novel model for the spatio-temporal clustering of trajectories based on motion, which applies to challenging street-view video sequences of pedestrians captured by a mobile camera. A key contribution of our work is the introduction of novel probabilistic region trajectories, motivated by the non-repeatability of segmentation of frames in a video sequence. Hierarchical image segments are obtained by using a state-of-the-art hierarchical segmentation algorithm, and connected from adjacent frames in a directed acyclic graph. The region trajectories and measures of confidence are extracted from this graph using a dynamic programming-based optimisation. Our second main contribution is a Bayesian framework with a twofold goal: to learn the optimal, in a maximum likelihood sense, Random Forests classifier of motion patterns based on video features, and construct a unique graph from region trajectories of different frames, lengths and hierarchical levels. Finally, we demonstrate the use of Isomap for effective spatio-temporal clustering of the region trajectories of pedestrians. We support our claims with experimental results on new and existing challenging video sequences. Fabio Galasso, Masahiro Iwasaki, Kunio Nobori, Roberto Cipolla |
ICCV | 4 |
| 2011 | A new distance for scale-invariant 3D shape recognition and registrationabstractThis paper presents a method for vote-based 3D shape recognition and registration, in particular using mean shift on 3D pose votes in the space of direct similarity transforms for the first time. We introduce a new distance between poses in this space-the SRT distance. It is left-invariant, unlike Euclidean distance, and has a unique, closed-form mean, in contrast to Riemannian distance, so is fast to compute. We demonstrate improved performance over the state of the art in both recognition and registration on a real and challenging dataset, by comparing our distance with others in a mean shift framework, as well as with the commonly used Hough voting approach. Minh-Tri Pham, Oliver J. Woodford, Frank Perbet, Atsuto Maki, Björn Stenger, Roberto Cipolla |
ICCV | 6 |
| 2011 | Co-occurrence flow for pedestrian detectionabstractThe last few years have seen considerable progress in pedestrian detection. Recent work has established a combination of oriented gradients and optic flow as effective features although the detection rates are still unsatisfactory for practical use. This paper introduces a new type of motion feature, the co-occurrence flow (CoF). The advance is to capture relative movements of different parts of the entire body, unlike existing motion features which extract internal motion in a local fashion. Through evaluations on the TUD-Brussels pedestrian dataset, we show that our motion feature based on co-occurrence flow contributes to boost the performance of existing methods. Atsuto Maki, Akihito Seki, Tomoki Watanabe, Roberto Cipolla |
ICIP | 4 |
| 2011 | Single and sparse view 3D reconstruction by learning shape priors
Yu Chen 0009, Roberto Cipolla |
Comput. Vis. Image Underst. | 2 |
| 2011 | Incremental Linear Discriminant Analysis Using Sufficient Spanning Sets and Its Applications
Tae-Kyun Kim 0001, Björn Stenger, Josef Kittler, Roberto Cipolla |
Int. J. Comput. Vis. | 4 |
| 2011 | Video Normals from Colored LightsabstractWe present an algorithm and the associated single-view capture methodology to acquire the detailed 3D shape, bends, and wrinkles of deforming surfaces. Moving 3D data has been difficult to obtain by methods that rely on known surface features, structured light, or silhouettes. Multispectral photometric stereo is an attractive alternative because it can recover a dense normal field from an untextured surface. We show how to capture such data, which in turn allows us to demonstrate the strengths and limitations of our simple frame-to-frame registration over time. Experiments were performed on monocular video sequences of untextured cloth and faces with and without white makeup. Subjects were filmed under spatially separated red, green, and blue lights. Our first finding is that the color photometric stereo setup is able to produce smoothly varying per-frame reconstructions with high detail. Second, when these 3D reconstructions are augmented with 2D tracking results, one can register both the surfaces and relax the homogenous-color restriction of the single-hue subject. Quantitative and qualitative experiments explore both the practicality and limitations of this simple multispectral capture system. Gabriel J. Brostow, Carlos Hernández 0002, George Vogiatzis, Björn Stenger, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2011 | Overcoming Shadows in 3-Source Photometric StereoabstractLight occlusions are one of the most significant difficulties of photometric stereo methods. When three or more images are available without occlusion, the local surface orientation is overdetermined so that shape can be computed and the shadowed pixels can be discarded. In this paper, we look at the challenging case when only two images are available without occlusion, leading to a one degree of freedom ambiguity per pixel in the local orientation. We show that, in the presence of noise, integrability alone cannot resolve this ambiguity and reconstruct the geometry in the shadowed regions. As the problem is ill-posed in the presence of noise, we describe two regularization schemes that improve the numerical performance of the algorithm while preserving the data. Finally, the paper describes how this theory applies in the framework of color photometric stereo where one is restricted to only three images and light occlusions are common. Experiments on synthetic and real image sequences are presented. Carlos Hernández 0002, George Vogiatzis, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2010 | Label propagation in complex video sequences using semi-supervised learningabstractWe propose a novel directed graphical model for label propagation in lengthy and complex video sequences. Given hand-labelled start and end frames of a video sequence, a variational EM based inference strategy propagates either one of several class labels or assigns an unknown class (void) label to each pixel in the video. These labels are used to train a multi-class classifier. The pixel labels estimated by this classifier are injected back into the Bayesian network for another iteration of label inference. The novel aspect of this iterative scheme, as compared to a recent approach [1], is its ability to handle occlusions. This is attributed to a hybrid of generative propagation and discriminative classification in a pseudo time-symmetric video model. The end result is a conservative labelling of the video; large parts of the static scene are labelled into known classes, and a void label is assigned to moving objects and remaining parts of the static scene. These labels can be used as ground truth data to learn the static parts of a scene from videos of it or more generally for semantic video segmentation. We demonstrate the efficacy of the proposed approach using extensive qualitative and quantitative tests over six challenging sequences. We bring out the advantages and drawbacks of our approach, both to encourage its repeatability and motivate future research directions. Ignas Budvytis, Vijay Badrinarayanan, Roberto Cipolla |
BMVC | 3 |
| 2010 | Making a Shallow Network Deep: Growing a Tree from Decision Regions of a Boosting ClassifierabstractThis paper presents a novel way to speed up the classification time of a boosting classifier. We make the shallow (flat) network deep (hierarchical) by growing a tree from the decision regions of a given boosting classifier. This provides many short paths for speeding up and preserves the reasonably smooth decision regions of the boosting classifier for good generalisation. We express the conversion as a Boolean optimisation problem, which has been previously studied for circuit design but limited to a small number of binary variables. In this work, a novel optimisation method is proposed for several tens of variables, i.e. weak-learners of a boosting classifier. The method is then used in a two stage cascade allowing the speed-up of a boosting classifier with any larger number of weak-learners. Experiments on the synthetic and face image data sets show that the obtained tree significantly speeds up both a standard boosting classifier and Fast-exit, a prior-art for fast boosting classification, at the same accuracy. The proposed method as a general meta-algorithm is also shown useful for a boosting cascade, since it speeds up individual stage classifiers by different gains. The proposed method is further demonstrated for rapid object tracking and segmentation problems. Tae-Kyun Kim 0001, Ignas Budvytis, Roberto Cipolla |
BMVC | 3 |
| 2010 | Real-time Action Recognition by Spatiotemporal Semantic and Structural ForestsabstractThis paper presents a novel real-time action recogniser by utilising both local appearance and structural information. Our method is able to recognise actions continuously in real-time while achieving comparably high accuracy over state-of-the-arts. Run-time speed is of vital importance in real-world action recognition systems, but existing methods seldom take computational complexity into full consideration. A class label is assigned after an entire query video is analysed, or a large lookahead is required to recognise an action. In addition, the “bag of words”(BOW) has proven effective for action recognition [5]. However, the standard BOW model ignores the spatiotemporal relationships among feature descriptors, which are useful for describing actions. Addressing these challenges, we present a novel approach for action recognition. The major contributions include the followings: Efficient Spatiotemporal Codebook Learning: We extend the use of semantic texton forests [6] (STFs) from 2D image segmentation to spatiotemporal analysis. As well as being much faster than a traditional flat codebook such as k-means clustering, STFs achieve high accuracy comparable to that of existing approaches. STFs are ensembles of random decision trees that textonise input video patches into semantic textons. Since only a small number of simple features are used to traverse the trees, STFs are extremely fast to evaluate. They also serve a powerful discriminative codebook by multiple decision trees. Figure 1 illustrates how visual codewords are generated using STFs in the proposed method. Combined Structural and Appearance Information: We propose a richer description of features, hence actions can be classified in very short video sequences. Based on [3], we introduce the pyramidal spatiotemporal relationship match (PSRM) to encapsulate both local appearance and structural information efficiently. Subsequences are sampled from an input video in short intervals (e.g. ≤ 10 frames). After spatiotemporal interest points are localised, the trained STFs assign visual codewords to the features. A set of pairwise spatiotemporal associations are designed to capture the structural relationships among features (i.e. pairwise distances along space-time axes). All possible pairs in the bag of features are analysed by the association rules and stored in the 3-D histogram. PSRM leverages the properties of semantic trees and pyramidal match kernels. Multiple pyramidal histograms are then combined to classify a query video. Figure 2 illustrates how the relationship histograms are constructed and matched using PSRM. For each tree in STFs, the threedimensional histogram is constructed according to their spatiotemporal structures (see figure 2 (left)). Its hierarchical structure offers a time efficient way to perform the pyramid match kernel [1] for codeword matching (figure 2 (right)). Enhanced Efficiency and Combined Classification: Several techniques are employed to improve the recognition speed and accuracy. A novel spatiotemporal interest point detector, called V-FAST, is designed based on the FAST 2D corners [2]. The recognition accuracy is enhanced by adaptively combining PSRM and the bag of semantic texton (BOST) method [6]: the k-means forest classifier is learned using PSRM as a matching kernel. The task of action recognition is performed separately Spatiotemporal Relationship Match of visual codewords from Semantic Texton Forest Pyramid Match Kernel is utilised to match the histograms Feature Extraction Feature Matching Tsz-Ho Yu, Tae-Kyun Kim 0001, Roberto Cipolla |
BMVC | 3 |
| 2010 | Label propagation in video sequencesabstractThis paper proposes a probabilistic graphical model for the problem of propagating labels in video sequences, also termed the label propagation problem. Given a limited amount of hand labelled pixels, typically the start and end frames of a chunk of video, an EM based algorithm propagates labels through the rest of the frames of the video sequence. As a result, the user obtains pixelwise labelled video sequences along with the class probabilities at each pixel. Our novel algorithm provides an essential tool to reduce tedious hand labelling of video sequences, thus producing copious amounts of useable ground truth data. A novel application of this algorithm is in semi-supervised learning of discriminative classifiers for video segmentation and scene parsing. The label propagation scheme can be based on pixel-wise correspondences obtained from motion estimation, image patch based similarities as seen in epitomic models or even the more recent, semantically consistent hierarchical regions. We compare the abilities of each of these variants, both via quantitative and qualitative studies against ground truth data. We then report studies on a state of the art Random forest classifier based video segmentation scheme, trained using fully ground truth data and with data obtained from label propagation. The results of this study strongly support and encourage the use of the proposed label propagation algorithm. Vijay Badrinarayanan, Fabio Galasso, Roberto Cipolla |
CVPR | 3 |
| 2010 | Inferring 3D Shapes and Deformations from Single Views
Yu Chen 0009, Tae-Kyun Kim 0001, Roberto Cipolla |
ECCV (3) | 3 |
| 2010 | Automatic 3D object segmentation in multiple views using volumetric graph-cuts
Neill D. F. Campbell, George Vogiatzis, Carlos Hernández 0002, Roberto Cipolla |
Image Vis. Comput. | 4 |
| 2010 | Thermal and reflectance based personal identification methodology under variable illumination
Ognjen Arandjelovic, Riad I. Hammoud, Roberto Cipolla |
Pattern Recognit. | 3 |
| 2010 | On-line Learning of Mutually Orthogonal Subspaces for Face Recognition by Image SetsabstractWe address the problem of face recognition by matching image sets. Each set of face images is represented by a subspace (or linear manifold) and recognition is carried out by subspace-to-subspace matching. In this paper, 1) a new discriminative method that maximises orthogonality between subspaces is proposed. The method improves the discrimination power of the subspace angle based face recognition method by maximizing the angles between different classes. 2) We propose a method for on-line updating the discriminative subspaces as a mechanism for continuously improving recognition accuracy. 3) A further enhancement called locally orthogonal subspace method is presented to maximise the orthogonality between competing classes. Experiments using 700 face image sets have shown that the proposed method outperforms relevant prior art and effectively boosts its accuracy by online learning. It is shown that the method for online learning delivers the same solution as the batch computation at far lower computational cost and the locally orthogonal method exhibits improved accuracy. We also demonstrate the merit of the proposed face recognition method on portal scenarios of multiple biometric grand challenge. Tae-Kyun Kim 0001, Josef Kittler, Roberto Cipolla |
IEEE Trans. Image Process. | 3 |
| 2009 | Obtaining the Shape of a Moving Object with a Specular SurfaceabstractThis paper addresses the basic problem of recovering the 3D surface of an object that is observed in motion by a single camera and under a static but unknown lighting condition. We propose a method to establish pixelwise correspondence between input images by way of depth search by investigating optimal subsets of intensities rather than employing all the relevant pixel values. The thrust of our algorithm is that it is capable of dealing with specularities which appear on the top of shading variance that is caused due to object motion. This is in terms of both stages of finding sparse point correspondence and dense depth search. We also propose that a linearised image basis can be directly computed by the procudure of finding the correspondence. We illustrate the performance of the theoretical propositions using images of real objects. © 2009. The copyright of this document resides with its authors. Atsuto Maki, Roberto Cipolla |
BMVC | 2 |
| 2009 | Learning to track with multiple observersabstractWe propose a novel approach to designing algorithms for object tracking based on fusing multiple observation models. As the space of possible observation models is too large for exhaustive on-line search, this work aims to select models that are suitable for a particular tracking task at hand. During an off-line training stage observation models from various off-the-shelf trackers are evaluated. From this data different methods of fusing the observers on-line are investigated, including parallel and cascaded evaluation. Experiments on test sequences show that this evaluation is useful for automatically designing and assessing algorithms for a particular tracking task. Results are shown for face tracking with a handheld camera and hand tracking for gesture interaction. We show that for these cases combining a small number of observers in a sequential cascade results in efficient algorithms that are both robust and precise. Björn Stenger, Thomas Woodley, Roberto Cipolla |
CVPR | 3 |
| 2009 | Image mosaicing via quadric surface estimation with priors for tunnel inspectionabstractIn this paper, a system which constructs a mosaic image of the tunnel surface with little distortion is presented. The tunnel surface is typically composed of a roughly cylindrical surface and protuberant regions containing objects such as pipes, pans and tunnel ridges. Since the true surface is neither planar nor quadric, existing mosaicing methods, which assume either homography or quadratic motion models, suffer from distortion. The proposed system obtains a sparse 3D model of the tunnel by multi-view reconstruction. Then, the Support Vector Machine (SVM) classifier is applied in order to separate image features lying on the cylindrical surface from those of the non-surface. The reconstructed 3D points are reprojected into images to retrieve the priors given by the SVM classifier for accurate cylindrical surface estimation. The final mosaic image is obtained by flattening the estimated textured surface onto a plane. The results suggest that the mosaic quality depends critically on the surface estimation accuracy and the proposed system is able to produce the mosaic image that preserves all physical sense, e.g. line parallelism and straightness, which is important for tunnel inspection. Krisada Chaiyasarn, Tae-Kyun Kim 0001, Fabio Viola, Roberto Cipolla, Kenichi Soga |
ICIP | 4 |
| 2009 | A pose-wise linear illumination manifold model for face recognition using video
Ognjen Arandjelovic, Roberto Cipolla |
Comput. Vis. Image Underst. | 2 |
| 2009 | A methodology for rapid illumination-invariant face recognition using image processing filters
Ognjen Arandjelovic, Roberto Cipolla |
Comput. Vis. Image Underst. | 2 |
| 2009 | Canonical Correlation Analysis of Video Volume Tensors for Action Categorization and DetectionabstractThis paper addresses a spatiotemporal pattern recognition problem. The main purpose of this study is to find a right representation and matching of action video volumes for categorization. A novel method is proposed to measure video-to-video volume similarity by extending Canonical Correlation Analysis (CCA), a principled tool to inspect linear relations between two sets of vectors, to that of two multiway data arrays (or tensors). The proposed method analyzes video volumes as inputs avoiding the difficult problem of explicit motion estimation required in traditional methods and provides a way of spatiotemporal pattern matching that is robust to intraclass variations of actions. The proposed matching is demonstrated for action classification by a simple Nearest Neighbor classifier. We, moreover, propose an automatic action detection method, which performs 3D window search over an input video with action exemplars. The search is speeded up by dynamic learning of subspaces in the proposed CCA. Experiments on a public action data set (KTH) and a self-recorded hand gesture data showed that the proposed method is significantly better than various state-of-the-art methods with respect to accuracy. Our method has low time complexity and does not require any major tuning parameters. Tae-Kyun Kim 0001, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Semantic object classes in video: A high-definition ground truth database
Gabriel J. Brostow, Julien Fauqueur, Roberto Cipolla |
Pattern Recognit. Lett. | 3 |
| 2008 | Efficiently Combining Contour and Texture Cues for Object RecognitionabstractThis paper proposes an efficient fusion of contour and texture cues for im-age categorization and object detection. Our work confirms and strengthens recent results that combining complementary feature types improves perfor-mance. We obtain a similar improvement in accuracy and additionally an improvement in efficiency. We use a boosting algorithm to learn models that use contour and texture features. Our main contributions are (i) the use of dense generic texture features to complement contour fragments, and (ii) a simple feature selection mechanism that includes the computational costs of features in order to learn a run-time efficient model. Our evaluation on 17 challenging and varied object classes confirms that the synergy of the two feature types performs significantly better than either alone, and that computational efficiency is substantially improved using our feature selection mechanism. An investigation of the boosted features shows a fascinating emergent property: the absence of certain textures often con-tributes towards object detection. Comparison with recent work shows that performance is state of the art. 1 Jamie Shotton, Andrew Blake 0001, Roberto Cipolla |
BMVC | 3 |
| 2008 | AIDIA - Adaptive Interface for Display InterActionabstractThis paper presents a vision-based system for interaction with a display via hand pointing. An attention mechanism based on face and hand detection allows users in the camera’s field of view to take control of the interface. Face recognition is used for identification and customisation. The system allows the user to control the screen pointer by tracking their fist. On-screen items can be selected using one of four activation mechanisms. Current sample applications include browsing image and video collections as well as viewing a gallery of 3D objects. In experiments we demonstrate the performance of the vision components in challenging conditions and compare it to that of other systems. 1 Björn Stenger, Thomas Woodley, Tae-Kyun Kim 0001, Carlos Hernández 0002, Roberto Cipolla |
BMVC | 5 |
| 2008 | Semantic texton forests for image categorization and segmentationabstractWe propose semantic texton forests, efficient and powerful new low-level features. These are ensembles of decision trees that act directly on image pixels, and therefore do not need the expensive computation of filter-bank responses or local descriptors. They are extremely fast to both train and test, especially compared with k-means clustering and nearest-neighbor assignment of feature descriptors. The nodes in the trees provide (i) an implicit hierarchical clustering into semantic textons, and (ii) an explicit local classification estimate. Our second contribution, the bag of semantic textons, combines a histogram of semantic textons over an image region with a region prior category distribution. The bag of semantic textons is computed over the whole image for categorization, and over local rectangular regions for segmentation. Including both histogram and region prior allows our segmentation algorithm to exploit both textural and semantic context. Our third contribution is an image-level prior for segmentation that emphasizes those categories that the automatic categorization believes to be present. We evaluate on two datasets including the very challenging VOC 2007 segmentation dataset. Our results significantly advance the state-of-the-art in segmentation accuracy, and furthermore, our use of efficient decision forests gives at least a five-fold increase in execution speed. Jamie Shotton, Matthew Johnson 0003, Roberto Cipolla |
CVPR | 3 |
| 2008 | Principled fusion of high-level model and low-level cues for motion segmentationabstractHigh-level generative models provide elegant descriptions of videos and are commonly used as the inference framework in many unsupervised motion segmentation schemes. However, approximate inference in these models often require ad-hoc initialization to avoid local minima issues. Low-level cues, obtained independently from the high-level model, can constrain the search space and reduce the chance of inference algorithms falling into a local minima. This paper introduces a novel principled fusion framework where, local hierarchical superpixels segmentation of images are used to capture local motion. The low-level cues such as local motion, on their own, not adequate to obtain full motion segmentation as occlusion needs to be handled globally. We fuse the low-level motion cues with the high-level model in a principled manner to surmount the shortcomings of using only the high-level model or low-level cues to perform motion segmentation. The fused model contains both continuous and discrete variables which forms a number of Markov Random fields. Variational approximation or belief propagation algorithms cannot be applied due to the complex interactions between the variables. Hence, approximate inference is performed using expectation propagation (EP) algorithm. The scheme is demonstrated by performing motion segmentation in two video sequences. Arasanathan Thayananthan, Masahiro Iwasaki, Roberto Cipolla |
CVPR | 3 |
| 2008 | Segmentation and Recognition Using Structure from Motion Point Clouds
Gabriel J. Brostow, Jamie Shotton, Julien Fauqueur, Roberto Cipolla |
ECCV (1) | 4 |
| 2008 | Using Multiple Hypotheses to Improve Depth-Maps for Multi-View Stereo
Neill D. F. Campbell, George Vogiatzis, Carlos Hernández 0002, Roberto Cipolla |
ECCV (1) | 4 |
| 2008 | Shadows in Three-Source Photometric Stereo
Carlos Hernández 0002, George Vogiatzis, Roberto Cipolla |
ECCV (1) | 3 |
| 2008 | Colour invariants for machine face recognitionabstractIllumination invariance remains the most researched, yet the most challenging aspect of automatic face recognition. In this paper we investigate the discriminative power of colour-based invariants in the presence of large illumination changes between training and test data, when appearance changes due to cast shadows and non-Lambertian effects are significant. Specifically, there are three main contributions: (i) we employ a more sophisticated photometric model of the camera and show how its parameters can be estimated, (ii) we derive several novel colour-based face invariants, and (iii) on a large database of video sequences we examine and evaluate the largest number of colour-based representations in the literature. Our results suggest that colour invariants do have a substantial discriminative power which may increase the robustness and accuracy of recognition from low resolution images. Ognjen Arandjelovic, Roberto Cipolla |
FG | 2 |
| 2008 | MCBoost: Multiple Classifier Boosting for Perceptual Co-clustering of Images and Visual FeaturesabstractWe present a new co-clustering problem of images and visual features. The problem involves a set of non-object images in addition to a set of object images and features to be co-clustered. Co-clustering is performed in a way of maximising discrimination of object images from non-object images, thus emphasizing discriminative features. This provides a way of obtaining perceptual joint-clusters of object images and features. We tackle the problem by simultaneously boosting multiple strong classifiers which compete for images by their expertise. Each boosting classifier is an aggregation of weak-learners, i.e. simple visual features. The obtained classifiers are useful for multi-category and multi-view object detection tasks. Experiments on a set of pedestrian images and a face data set demonstrate that the method yields intuitive image clusters with associated features and is much superior to conventional boosting classifiers in object detection tasks. Tae-Kyun Kim 0001, Roberto Cipolla |
NIPS | 2 |
| 2008 | Reconstructing relief surfaces
George Vogiatzis, Philip Torr 0001, Steven M. Seitz, Roberto Cipolla |
Image Vis. Comput. | 4 |
| 2008 | Multiview Photometric StereoabstractThis paper addresses the problem of obtaining complete, detailed reconstructions of textureless shiny objects. We present an algorithm which uses silhouettes of the object, as well as images obtained under changing illumination conditions. In contrast with previous photometric stereo techniques, ours is not limited to a single viewpoint but produces accurate reconstructions in full 3D. A number of images of the object are obtained from multiple viewpoints, under varying lighting conditions. Starting from the silhouettes, the algorithm recovers camera motion and constructs the object's visual hull. This is then used to recover the illumination and initialise a multi-view photometric stereo scheme to obtain a closed surface reconstruction. There are two main contributions in this paper: Firstly we describe a robust technique to estimate light directions and intensities and secondly, we introduce a novel formulation of photometric stereo which combines multiple viewpoints and hence allows closed surface reconstructions. The algorithm has been implemented as a practical model acquisition system. Here, a quantitative evaluation of the algorithm on synthetic data is presented together with complete reconstructions of challenging real objects. Finally, we show experimentally how even in the case of highly textured objects, this technique can greatly improve on correspondence-based multi-view stereo results. Carlos Hernández Esteban, George Vogiatzis, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2008 | Multiscale Categorical Object Recognition Using Contour FragmentsabstractPsychophysical studies [9], [17] show that we can recognize objects using fragments of outline contour alone. This paper proposes a new automatic visual recognition system based only on local contour features, capable of localizing objects in space and scale. The system first builds a class-specific codebook of local fragments of contour using a novel formulation of chamfer matching. These local fragments allow recognition that is robust to within-class variation, pose changes, and articulation. Boosting combines these fragments into a cascaded sliding-window classifier, and mean shift is used to select strong responses as a final set of detections. We show how learning can be performed iteratively on both training and test sets to boot-strap an improved classifier. We compare with other methods based on contour and local descriptors in our detailed evaluation over 17 challenging categories, and obtain highly competitive results. The results confirm that contour is indeed a powerful cue for multi-scale and multi-class visual object recognition. Jamie Shotton, Andrew Blake 0001, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2008 | Pose estimation and tracking using multivariate regression
Arasanathan Thayananthan, Ramanan Navaratnam, Björn Stenger, Philip Torr 0001, Roberto Cipolla |
Pattern Recognit. Lett. | 5 |
| 2007 | Gesture Recognition Under Small Sample Size
Tae-Kyun Kim 0001, Roberto Cipolla |
ACCV (1) | 2 |
| 2007 | Automatic 3D Object Segmentation in Multiple Views using Volumetric Graph-CutsabstractWe propose an algorithm for automatically obtaining a segmentation of a rigid object in a sequence of images that are calibrated for camera pose and intrinsic parameters. Until recently, the best segmentation results have been obtained by interactive methods that require manual labelling of image regions. Our method requires no user input but instead relies on the camera fixating on the object of interest during the sequence. We begin by learning a model of the object is colour, from the image pixels around the fixation points. We then extract image edges and combine these with the object colour information in a volumetric binary MRF model. The globally optimal segmentation of 3D space is obtained by a graph-cut optimisation. From this segmentation an improved colour model is extracted and the whole process is iterated until convergence. Our first finding is that the fixation constraint, which requires that the object of interest is more or less central in the image, is enough to determine what to segment and initialise an automatic segmentation process. Second, we find that by performing a single segmentation in 3D, we implicitly exploit a 3D rigidity constraint, expressed as silhouette coherency, which significantly improves silhouette quality over independent 2D segmentations. We demonstrate the validity of our approach by providing segmentation results on real sequences. Neill D. F. Campbell, George Vogiatzis, Carlos Hernández Esteban, Roberto Cipolla |
BMVC | 4 |
| 2007 | Tracking Using Online Feature Selection and a Local Generative ModelabstractThis paper proposes an algorithm for online feature selection which improves robustness to occlusions by referring to a localized generative appearance model. Discriminative classifiers based on feature extraction have classically either prepared a fixed prior model by training offline, or continually adapted their classification parameters to any apparent appearance changes. By combining the attractive qualities of each approach, our framework can cope with appearance changes of a target object and will maintain proximity to a static appearance model. Our main contribution is the use of a generative model to guide the online feature selection to regions of an image which maintain a valid appearance. The generative model exhibits the properties of non-negativity, localization and orthogonality. We demonstrate the system in a tracking framework to show improved tracking performance through occlusions. 1 Thomas Woodley, Björn Stenger, Roberto Cipolla |
BMVC | 3 |
| 2007 | Probabilistic visibility for multi-view stereoabstractWe present a new formulation to multi-view stereo that treats the problem as probabilistic 3D segmentation. Previous work has used the stereo photo-consistency criterion as a detector of the boundary between the 3D scene and the surrounding empty space. Here we show how the same criterion can also provide a foreground/background model that can predict if a 3D location is inside or outside the scene. This model replaces the commonly used naive foreground model based on ballooning which is known to perform poorly in concavities. We demonstrate how the probabilistic visibility is linked to previous work on depth-map fusion and we present a multi-resolution graph-cut implementation using the new ballooning term that is very efficient both in terms of computation time and memory requirements. Carlos Hernández 0002, George Vogiatzis, Roberto Cipolla |
CVPR | 3 |
| 2007 | Tensor Canonical Correlation Analysis for Action ClassificationabstractWe introduce a new framework, namely tensor canonical correlation analysis (TCCA) which is an extension of classical canonical correlation analysis (CCA) to multidimensional data arrays (or tensors) and apply this for action/gesture classification in videos. By tensor CCA, joint space-time linear relationships of two video volumes are inspected to yield flexible and descriptive similarity features of the two videos. The TCCA features are combined with a discriminative feature selection scheme and a nearest neighbor classifier for action classification. In addition, we propose a time-efficient action detection method based on dynamic learning of subspaces for tensor CCA for the case that actions are not aligned in the space-time domain. The proposed method delivered significantly better accuracy and comparable detection speed over state-of-the-art methods on the KTH action data set as well as self-recorded hand gesture data sets. Tae-Kyun Kim 0001, Shu-Fai Wong, Roberto Cipolla |
CVPR | 3 |
| 2007 | Incremental Linear Discriminant Analysis Using Sufficient Spanning Set ApproximationsabstractThis paper presents a new incremental learning solution for linear discriminant analysis (LDA). We apply the concept of the sufficient spanning set approximation in each update step, i.e. for the between-class scatter matrix, the projected data matrix as well as the total scatter matrix. The algorithm yields a more general and efficient solution to incremental LDA than previous methods. It also significantly reduces the computational complexity while providing a solution which closely agrees with the batch LDA result. The proposed algorithm has a time complexity of O(Nd2) and requires O(Nd) space, where d is the reduced subspace dimension and N the data dimension. We show two applications of incremental LDA: First, the method is applied to semi-supervised learning by integrating it into an EM framework. Secondly, we apply it to the task of merging large databases which were collected during MPEG standardization for face image retrieval. Tae-Kyun Kim 0001, Shu-Fai Wong, Björn Stenger, Josef Kittler, Roberto Cipolla |
CVPR | 5 |
| 2007 | Learning Motion Categories using both Semantic and Structural InformationabstractCurrent approaches to motion category recognition typically focus on either full spatiotemporal volume analysis (holistic approach) or analysis of the content of spatiotemporal interest points (part-based approach). Holistic approaches tend to be more sensitive to noise e.g. geometric variations, while part-based approaches usually ignore structural dependencies between parts. This paper presents a novel generative model, which extends probabilistic latent semantic analysis (pLSA), to capture both semantic (content of parts) and structural (connection between parts) information for motion category recognition. The structural information learnt can also be used to infer the location of motion for the purpose of motion detection. We test our algorithm on challenging datasets involving human actions, facial expressions and hand gestures and show its performance is better than existing unsupervised methods in both tasks of motion localisation and recognition. Shu-Fai Wong, Tae-Kyun Kim 0001, Roberto Cipolla |
CVPR | 3 |
| 2007 | Assisted Video Object Labeling By Joint Tracking of Regions and KeypointsabstractManual labeling of objects in videos is a tedious task. We present an approach which automatically propagates the labels from a single frame to the next ones. We tackle the challenging problem of tracking segmented regions by combining keypoint tracking with an advanced multiple region matching strategy, based on inclusion similarity and connected regions. We ran experiments on a 101 frame driving video sequence for which we produced the corresponding hand- labeled groundtruth. We make this valuable dataset available for the research community. We show our technique can accommodate variations in segmentation (and correct them), even in presence of multiple independent motions and partial occlusion. Results show that most of the labeled pixels can be correctly propagated even after a hundred frames. The performance of this automatic propagation mechanism over many frames can greatly reduce the user effort in the task of video object labeling. Julien Fauqueur, Gabriel J. Brostow, Roberto Cipolla |
ICCV | 3 |
| 2007 | Non-rigid Photometric Stereo with Colored LightsabstractWe present an algorithm and the associated capture methodology to acquire and track the detailed 3D shape, bends, and wrinkles of deforming surfaces. Moving 3D data has been difficult to obtain by methods that rely on known surface features, structured light, or silhouettes. Multispec- tral photometric stereo is an attractive alternative because it can recover a dense normal field from an un-textured surface. We show how to capture such data and register it over time to generate a single deforming surface. Experiments were performed on video sequences of un- textured cloth, filmed under spatially separated red, green, and blue light sources. Our first finding is that using zero- depth-silhouettes as the initial boundary condition already produces rather smoothly varying per-frame reconstructions with high detail. Second, when these 3D reconstructions are augmented with 2D optical flow, one can register the first frame's reconstruction to every subsequent frame. Carlos Hernández 0002, George Vogiatzis, Gabriel J. Brostow, Björn Stenger, Roberto Cipolla |
ICCV | 5 |
| 2007 | The Joint Manifold Model for Semi-supervised Multi-valued RegressionabstractMany computer vision tasks may be expressed as the problem of learning a mapping between image space and a parameter space. For example, in human body pose estimation, recent research has directly modelled the mapping from image features (z) to joint angles (θ). Fitting such models requires training data in the form of labelled (zθ) pairs, from which are learned the conditional densities p(zθ). Inference is then simple: given test image featuresz, the conditional (zθ) is immediately computed. However large amounts of training data are required to fit the models, particularly in the case where the spaces are high dimensional. We show how the use of unlabelled data—samples from the marginal distributions p(z) and p(θ)—may be used to improve fitting. This is valuable because it is often significantly easier to obtain unlabelled than labelled samples. We use a Gaussian process latent variable model to learn the mapping from a shared latent low-dimensional manifold to the feature and parameter spaces. This extends existing approaches to (a) use unlabelled data, and (b) represent one-to-many mappings. Experiments on synthetic and real problems demonstrate how the use of unlabelled data improves over existing techniques. In our comparisons, we include existing approaches that are explicitly semi-supervised as well as those which implicitly make use of unlabelled examples. Ramanan Navaratnam, Andrew W. Fitzgibbon, Roberto Cipolla |
ICCV | 3 |
| 2007 | Extracting Spatiotemporal Interest Points using Global InformationabstractLocal spatiotemporal features or interest points provide compact but descriptive representations for efficient video analysis and motion recognition. Current local feature extraction approaches involve either local filtering or entropy computation which ignore global information (e.g. large blobs of moving pixels) in video inputs. This paper presents a novel extraction method which utilises global information from each video input so that moving parts such as a moving hand can be identified and are used to select relevant interest points for a condensed representation. The proposed method involves obtaining a small set of subspace images, which can synthesise frames in the video input from their corresponding coefficient vectors, and then detecting interest points from the subspaces and the coefficient vectors. Experimental results indicate that the proposed method can yield a sparser set of interest points for motion recognition than existing methods. Shu-Fai Wong, Roberto Cipolla |
ICCV | 2 |
| 2007 | Estimating 3D hand pose using hierarchical multi-label classification
Björn Stenger, Arasanathan Thayananthan, Philip Torr 0001, Roberto Cipolla |
Image Vis. Comput. | 4 |
| 2007 | Silhouette Coherence for Camera Calibration under Circular MotionabstractWe present a new approach to camera calibration as a part of a complete and practical system to recover digital copies of sculpture from uncalibrated image sequences taken under turntable motion. In this paper, we introduce the concept of the silhouette coherence of a set of silhouettes generated by a 3D object. We show how the maximization of the silhouette coherence can be exploited to recover the camera poses and focal length. Silhouette coherence can be considered as a generalization of the well-known epipolar tangency constraint for calculating motion from silhouettes or outlines alone. Further, silhouette coherence exploits all the geometric information encoded in the silhouette (not just at epipolar tangency points) and can be used in many practical situations where point correspondences or outer epipolar tangents are unavailable. We present an algorithm for exploiting silhouette coherence to efficiently and reliably estimate camera motion. We use this algorithm to reconstruct very high quality 3D models from uncalibrated circular motion sequences, even when epipolar tangency points are not available or the silhouettes are truncated. The algorithm has been integrated into a practical system and has been tested on more than 50 uncalibrated sequences to produce high quality photo-realistic models. Three illustrative examples are included in this paper. The algorithm is also evaluated quantitatively by comparing it to a state-of-the-art system that exploits only epipolar tangents. Carlos Hernández 0002, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2007 | Discriminative Learning and Recognition of Image Set Classes Using Canonical CorrelationsabstractWe address the problem of comparing sets of images for object recognition, where the sets may represent variations in an object's appearance due to changing camera pose and lighting conditions. Canonical Correlations (also known as principal or canonical angles), which can be thought of as the angles between two d-dimensional subspaces, have recently attracted attention for image set matching. Canonical correlations offer many benefits in accuracy, efficiency, and robustness compared to the two main classical methods: parametric distribution-based and nonparametric sample-based matching of sets. Here, this is first demonstrated experimentally for reasonably sized data sets using existing methods exploiting canonical correlations. Motivated by their proven effectiveness, a novel discriminative learning method over sets is proposed for set classification. Specifically, inspired by classical Linear Discriminant Analysis (LDA), we develop a linear discriminant function that maximizes the canonical correlations of within-class sets and minimizes the canonical correlations of between-class sets. Image sets transformed by the discriminant function are then compared by the canonical correlations. Classical orthogonal subspace method (OSM) is also investigated for the similar purpose and compared with the proposed method. The proposed method is evaluated on various object recognition problems using face image sets with arbitrary motion captured under different illuminations and image sets of 500 general objects taken at different views. The method is also applied to object category recognition using ETH-80 database. The proposed method is shown to outperform the state-of-the-art methods in terms of accuracy and efficiency. Tae-Kyun Kim 0001, Josef Kittler, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2007 | Multiview Stereo via Volumetric Graph-Cuts and Occlusion Robust Photo-ConsistencyabstractThis paper presents a volumetric formulation for the multi-view stereo problem which is amenable to a computationally tractable global optimisation using Graph-cuts. Our approach is to seek the optimal partitioning of 3D space into two regions labelled as "object" and "empty" under a cost functional consisting of the following two terms: (1) A term that forces the boundary between the two regions to pass through photo-consistent locations and (2) a ballooning term that inflates the "object" region. To take account of the effect of occlusion on the first term we use an occlusion robust photo-consistency metric based on Normalised Cross Correlation, which does not assume any geometric knowledge about the reconstructed object. The globally optimal 3D partitioning can be obtained as the minimum cut solution of a weighted graph. George Vogiatzis, Carlos Hernández Esteban, Philip Torr 0001, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2007 | Boosted manifold principal angles for image set-based recognition
Tae-Kyun Kim 0001, Ognjen Arandjelovic, Roberto Cipolla |
Pattern Recognit. | 3 |
| 2006 | On Person Authentication by Fusing Visual and Thermal Face BiometricsabstractRecognition algorithms that use data obtained by imaging faces in the thermal spectrum are promising in achieving invariance to extreme illumination changes that are often present in practice. In this paper we analyze the performance of a recently proposed face recognition algorithm that combines visual and thermal modalities by decision level fusion. We examine (i) the effects of the proposed data preprocessing in each domain, (ii) the contribution to improved recognition of different types of features, (iii) the importance of prescription glasses detection, in the context of both 1-to-N and 1-to-1 matching (recognition vs. verification performance). Finally, we discuss the significance of our results and, in particular, identify a number of limitations of the current state-of-the-art and propose promising directions for future research. Ognjen Arandjelovic, Riad I. Hammoud, Roberto Cipolla |
AVSS | 3 |
| 2006 | Incremental Learning of Locally Orthogonal Subspaces for Set-based Object RecognitionabstractOrthogonal subspaces are effective models to represent object image sets (generally any high-dimensional vector sets). Canonical correlation analysis of the orthogonal subspaces provides a good solution to discriminate objects with sets of images. In such a recognition task involving image sets, an efficient learning over a large volume of image sets, which may be increasing over time, is important. In this paper, an incremental learning method of orthogonal subspaces is proposed by updating the principal components of the class correlation and total correlation matrices separately, yielding the same solution as the batch computation with far lower computational cost. A novel concept of local orthogonality is further proposed to cope with non-linear manifolds of data vectors and find a more optimal solution of orthogonal subspaces for a certain neighbouring object image sets. In the experiments using 700 face image sets, the locally orthogonal subspaces outperformed the orthogonal subspaces as well as relevant state-of-the-art methods in accuracy. Note that the locally orthogonal subspaces are also amenable to incremental updating due to their linear property. 1 Tae-Kyun Kim 0001, Josef Kittler, Roberto Cipolla |
BMVC | 3 |
| 2006 | Semi-supervised Learning of Joint Density Models for Human Pose EstimationabstractLearning regression models (for example for body pose estimation, or BPE) currently requires large numbers of training examples—pairs of the form (image, pose parameters). These examples are difficult to obtain for many problems, demanding considerable effort in manual labelling. However it is easy to obtain unlabelled examples—in BPE, simply by collecting many images, and by sampling many poses using motion capture. We show how the use of unlabelled examples can improve the performance of such estimators, making better use of the difficult-to-obtain training examples. Because the distribution of parameters conditioned on a given image is often multimodal, conventional regression models must be extended to allow for multiple modes. Such extensions have to date had a pre-set number of modes, independent of the contents of the input image, and amount to fitting several regressors simultaneously. Our framework models instead the joint distribution of images and poses, so the conditional estimates are inherently multimodal, and the number of modes is a function of the joint-space complexity, rather than of the maximum number of output modes. We demonstrate the improvements obtainable by using unlabelled samples on synthetic examples and on a real pose estimation problem, and demonstrate in both cases the additional accuracy provided by the use of unlabelled data. 1 Ramanan Navaratnam, Andrew W. Fitzgibbon, Roberto Cipolla |
BMVC | 3 |
| 2006 | Automatic Cast Listing in Feature-Length Films with Anisotropic Manifold SpaceabstractOur goal is to automatically determine the cast of a feature-length film. This is challenging because the cast size is not known, with appearance changes of faces caused by extrinsic imaging factors (illumination, pose, expression) often greater than due to differing identities. The main contribution of this paper is an algorithm for clustering over face appearance manifolds. Specifically: (i) we develop a novel algorithm for exploiting coherence of dissimilarities between manifolds, (ii) we show how to estimate the optimal dataset-specific discriminant manifold starting from a generic one, and (iii) we describe a fully automatic, practical system based on the proposed algorithm. The performance of the system is evaluated on well-known featurelength films and situation comedies on which it is shown to produce good results. Ognjen Arandjelovic, Roberto Cipolla |
CVPR (2) | 2 |
| 2006 | Unsupervised Bayesian Detection of Independent Motion in CrowdsabstractWhile crowds of various subjects may offer applicationspecific cues to detect individuals, we demonstrate that for the general case, motion itself contains more information than previously exploited. This paper describes an unsupervised data driven Bayesian clustering algorithm which has detection of individual entities as its primary goal. We track simple image features and probabilistically group them into clusters representing independently moving entities. The numbers of clusters and the grouping of constituent features are determined without supervised learning or any subject-specific model. The new approach is instead, that space-time proximity and trajectory coherence through image space are used as the only probabilistic criteria for clustering. An important contribution of this work is how these criteria are used to perform a one-shot data association without iterating through combinatorial hypotheses of cluster assignments. Our proposed general detection algorithm can be augmented with subject-specific filtering, but is shown to already be effective at detecting individual entities in crowds of people, insects, and animals. This paper and the associated video examine the implementation and experiments of our motion clustering framework. Gabriel J. Brostow, Roberto Cipolla |
CVPR (1) | 2 |
| 2006 | Reconstruction in the Round Using Photometric Normals and SilhouettesabstractThis paper addresses the problem of obtaining complete, detailed reconstructions of shiny textureless objects. We present an algorithm which uses silhouettes of the object, as well as images obtained under varying illumination conditions. In contrast with previous photometric stereo techniques, ours is not limited to a single viewpoint and produces accurate reconstructions in full 3D. A number of images of the object are obtained from multiple viewpoints, under varying lighting conditions. Starting from the silhouettes, the algorithm recovers camera motion and constructs the object’s visual hull. This is then used to recover the illumination and initialise a multi-view photometric stereo scheme to obtain a closed surface reconstruction. The contributions of the paper are twofold: Firstly we describe a robust technique to estimate light directions and intensities and secondly, we introduce a novel formulation of photometric stereo which combines multiple viewpoints and hence allows closed surface reconstructions. The algorithm has been implemented as a practical model acquisition system. Here, a quantitative evaluation of the algorithm on synthetic data is presented together with a complete reconstruction of a challenging real object. George Vogiatzis, Carlos Hernández 0002, Roberto Cipolla |
CVPR (2) | 3 |
| 2006 | Sparse and Semi-supervised Visual Mapping with the S3PabstractThis paper is about mapping images to continuous output spaces using powerful Bayesian learning techniques. A sparse, semi-supervised Gaussian process regression model (S3GP) is introduced which learns a mapping using only partially labelled training data. We show that sparsity bestows efficiency on the S3GP which requires minimal CPU utilization for real-time operation; the predictions of uncertainty made by the S3GP are more accurate than those of other models leading to considerable performance improvements when combined with a probabilistic filter; and the ability to learn from semi-supervised data simplifies the process of collecting training data. The S3GP uses a mixture of different image features: this is also shown to improve the accuracy and consistency of the mapping. A major application of this work is its use as a gaze tracking system in which images of a human eye are mapped to screen coordinates: in this capacity our approach is efficient, accurate and versatile. Oliver Williams, Andrew Blake 0001, Roberto Cipolla |
CVPR (1) | 3 |
| 2006 | Face Recognition from Video Using the Generic Shape-Illumination Manifold
Ognjen Arandjelovic, Roberto Cipolla |
ECCV (4) | 2 |
| 2006 | Learning Discriminative Canonical Correlations for Object Recognition with Image Sets
Tae-Kyun Kim 0001, Josef Kittler, Roberto Cipolla |
ECCV (3) | 3 |
| 2006 | Multivariate Relevance Vector Machines for Tracking
Arasanathan Thayananthan, Ramanan Navaratnam, Björn Stenger, Philip Torr 0001, Roberto Cipolla |
ECCV (3) | 5 |
| 2006 | Semantic Photo SynthesisabstractAbstract Composite images are synthesized from existing photographs by artists who make concept art, e.g., storyboards for movies or architectural planning. Current techniques allow an artist to fabricate such an image by digitally splicing parts of stock photographs. While these images serve mainly to “quickly”convey how a scene should look, their production is laborious. We propose a technique that allows a person to design a new photograph with substantially less effort. This paper presents a method that generates a composite image when a user types in nouns, such as “boat”and “sand.”The artist can optionally design an intended image by specifying other constraints. Our algorithm formulates the constraints as queries to search an automatically annotated image database. The desired photograph, not a collage, is then synthesized using graph‐cut optimization, optionally allowing for further user interaction to edit or choose among alternative generated photos. An implementation of our approach, shown in the associated video, demonstrates our contributions of (1) a method for creating specific images with minimal human effort, and (2) a combined algorithm for automatically building an image library with semantic annotations from any photo collection. Matthew Johnson 0003, Gabriel J. Brostow, Jamie Shotton, Ognjen Arandjelovic, Vivek Kwatra, Roberto Cipolla |
Comput. Graph. Forum | 6 |
| 2006 | An information-theoretic approach to face recognition from face motion manifolds
Ognjen Arandjelovic, Roberto Cipolla |
Image Vis. Comput. | 2 |
| 2006 | Model-Based Hand Tracking Using a Hierarchical Bayesian FilterabstractThis paper sets out a tracking framework, which is applied to the recovery of three-dimensional hand motion from an image sequence. The method handles the issues of initialization, tracking, and recovery in a unified way. In a single input image with no prior information of the hand pose, the algorithm is equivalent to a hierarchical detection scheme, where unlikely pose candidates are rapidly discarded. In image sequences, a dynamic model is used to guide the search and approximate the optimal filtering equations. A dynamic model is given by transition probabilities between regions in parameter space and is learned from training data obtained by capturing articulated motion. The algorithm is evaluated on a number of image sequences, which include hand motion with self-occlusion in front of a cluttered background. Björn Stenger, Arasanathan Thayananthan, Philip Torr 0001, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2006 | Coarse-to-Fine Vision-Based Localization by Indexing Scale-Invariant FeaturesabstractThis paper presents a novel coarse-to-fine global localization approach inspired by object recognition and text retrieval techniques. Harris-Laplace interest points characterized by scale-invariant transformation feature descriptors are used as natural landmarks. They are indexed into two databases: a location vector space model (LVSM) and a location database. The localization process consists of two stages: coarse localization and fine localization. Coarse localization from the LVSM is fast, but not accurate enough, whereas localization from the location database using a voting algorithm is relatively slow, but more accurate. The integration of coarse and fine stages makes fast and reliable localization possible. If necessary, the localization result can be verified by epipolar geometry between the representative view in the database and the view to be localized. In addition, the localization system recovers the position of the camera by essential matrix decomposition. The localization system has been tested in indoor and outdoor environments. The results show that our approach is efficient and reliable. Junqiu Wang, Hongbin Zha, Roberto Cipolla |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2005 | Incremental Learning of Temporally-Coherent Gaussian Mixture ModelsabstractIn this paper we address the problem of learning Gaussian Mixture Models (GMMs) incrementally.Unlike previous approaches which universally assume that new data comes in blocks representable by GMMs which are then merged with the current model estimate, our method works for the case when novel data points arrive oneby-one, while requiring little additional memory.We keep only two GMMs in the memory and no historical data.The current fit is updated with the assumption that the number of components is fixed, which is increased (or reduced) when enough evidence for a new component is seen.This is deduced from the change from the oldest fit of the same complexity, termed the Historical GMM, the concept of which is central to our method.The performance of the proposed method is demonstrated qualitatively and quantitatively on several synthetic data sets and video sequences of faces acquired in realistic imaging conditions. Ognjen Arandjelovic, Roberto Cipolla |
BMVC | 2 |
| 2005 | Improved Image Annotation and Labelling through Multi-Label BoostingabstractThe majority of machine learning systems for object recognition is limited by their requirement of single labelled images for training, which are difficult to create or obtain in quantity. It is therefore impractical to use methods or techniques which require such data to build object recognizers for more than a relatively small subset of object classes. Instead, far more abundant multilabel data provides a ready means to create object recognition systems which are able to deal with large numbers of classes. In this paper we present a new object recognition system named MLBoost which learns from multi-label data through boosting and improves on state-of-the-art multi-label annotation and labelling systems. The system is trained on images with accompanying text and at no time is told which parts of each image correspond to which words, and as such the process is unsupervised. Having once been trained it is able to give segment labels and a list of descriptive words (an annotation) for any novel image. Matthew Johnson 0003, Roberto Cipolla |
BMVC | 2 |
| 2005 | Learning over Sets using Boosted Manifold Principal Angles (BoMPA)abstractIn this paper we address the problem of classifying vector sets. We motivate and introduce a novel method based on comparisons between corresponding vector subspaces. In particular, there are two main areas of novelty: (i) we extend the concept of principal angles between linear subspaces to manifolds with arbitrary nonlinearities; (ii) it is demonstrated how boosting can be used for application-optimal principal angle fusion. The strengths of the proposed method are empirically demonstrated on the task of automatic face recognition (AFR), in which it is shown to outperform state-of-the-art methods in the literature. Tae-Kyun Kim 0001, Ognjen Arandjelovic, Roberto Cipolla |
BMVC | 3 |
| 2005 | Hierarchical Part-Based Human Body Pose EstimationabstractThis paper addresses the problem of automatic detection and recovery of three-dimensional human body pose from monocular video sequences for HCI applications. We propose a new hierarchical part-based pose estimation method for the upper-body that efficiently searches the high dimensional articulation space. The body is treated as a collection of parts linked in a kinematic structure. Search for configurations of this collection is commenced from the most reliably detectable part. The rest of the parts are searched based on the detected locations of this anchor as they all are kinematically linked. Each part is represented by a set of 2D templates created from a 3D model, hence inherently encoding the 3D joint angles. The tree data structure is exploited to efficiently search through these templates. Multiple hypotheses are computed for each frame. By modelling these with a HMM, temporal coherence of body motion is exploited to find a smooth trajectory of articulation between frames using a modified Viterbi algorithm. Experimental results show that the proposed technique produces good estimates of the human 3D pose on a range of test videos in a cluttered environment. Ramanan Navaratnam, Arasanathan Thayananthan, Philip Torr 0001, Roberto Cipolla |
BMVC | 4 |
| 2005 | Hole Filling Through PhotomontageabstractTo fill holes in photographs of structured, man made environments, we propose a technique which automatically adjusts and clones large image patches that have similar structure. These source patches can come from elsewhere in the same image, or from other images shot from different perspectives. Two significant developments of this work are the ability to automatically detect and adjust source patches whose macrostructure is compatible with the hole region, and alternately, to interactively specify a user's desired search regions. In contrast to existing photomontage algorithms which either synthesize microstructure or require careful user interaction to fill holes, our approach handles macrostructure with an adjustable degree of automation. Marta Wilczkowiak, Gabriel J. Brostow, Ben Tordoff, Roberto Cipolla |
BMVC | 4 |
| 2005 | Real-time Interpretation of Hand Motions using a Sparse Bayesian Classifier on Motion Gradient Orientation ImagesabstractAn approach to recognise 10 elementary gestures is proposed and it can be applied to sign language recognition. In this work, a motion gradient orientation image is extracted directly from a raw video input and transformed to a motion feature vector. This feature vector is then classified into one of the 10 elementary gestures by a sparse Bayesian classifier. A training set of 628 samples and a testing set of over 1000 samples have been obtained to evaluate the proposed method. A real-time system was built and trained with the training set. From the experiment, the reported classification accuracy is 90% and the system can run in around 25 frames per second. Compared with other recently proposed methods that involve the use of hand tracking, the system can work reliably in real-time without relying on accurate tracking, and give a probabilistic output that is useful in complex motion analysis. Shu-Fai Wong, Roberto Cipolla |
BMVC | 2 |
| 2005 | Face Recognition with Image Sets Using Manifold Density DivergenceabstractIn many automatic face recognition applications, a set of a person's face images is available rather than a single image. In this paper, we describe a novel method for face recognition using image sets. We propose a flexible, semi-parametric model for learning probability densities confined to highly non-linear but intrinsically low-dimensional manifolds. The model leads to a statistical formulation of the recognition problem in terms of minimizing the divergence between densities estimated on these manifolds. The proposed method is evaluated on a large data set, acquired in realistic imaging conditions with severe illumination variation. Our algorithm is shown to match the best and outperform other state-of-the-art algorithms in the literature, achieving 94% recognition rate on average. Ognjen Arandjelovic, Gregory Shakhnarovich, John Fisher, Roberto Cipolla, Trevor Darrell |
CVPR (1) | 4 |
| 2005 | Visual Tracking in the Presence of Motion BlurabstractWe consider the problem of visual tracking of regions of interest in a sequence of motion blurred images. Traditional methods couple tracking with deblurring in order to correctly account for the effects of motion blur. Such coupling is usually appropriate, but computationally wasteful when visual tracking is the lone objective. Instead of deblurring images, we propose to match regions by blurring them. The matching score for two image regions is governed by a cost function that only involves the region deformation parameters and two motion blur vectors. We present an efficient algorithm to minimize the proposed cost function and demonstrate it on sequences of real blurred images. Hailin Jin, Paolo Favaro, Roberto Cipolla |
CVPR (2) | 3 |
| 2005 | Multi-View Stereo via Volumetric Graph-CutsabstractThis paper presents a novel formulation for the multi-view scene reconstruction problem. While this formulation benefits from a volumetric scene representation, it is amenable to a computationally tractable global optimisation using Graph-cuts. The algorithm proposed uses the visual hull of the scene to infer occlusions and as a constraint on the topology of the scene. A photo consistency-based surface cost functional is defined and discretised with a weighted graph. The optimal surface under this discretised functional is obtained as the minimum cut solution of the weighted graph. Our method provides a viewpoint independent surface regularisation, approximate handling of occlusions and a tractable optimisation scheme. Promising experimental results on real scenes as well as a quantitative evaluation on a synthetic scene are presented. George Vogiatzis, Philip Torr 0001, Roberto Cipolla |
CVPR (2) | 3 |
| 2005 | Contour-Based Learning for Object DetectionabstractWe present a novel categorical object detection scheme that uses only local contour-based features. A two-stage, partially supervised learning architecture is proposed: a rudimentary detector is learned from a very small set of segmented images and applied to a larger training set of un-segmented images; the second stage bootstraps these detections to learn an improved classifier while explicitly training against clutter. The detectors are learned with a boosting algorithm which creates a location-sensitive classifier using a discriminative set of features from a randomly chosen dictionary of contour fragments. We present results that are very competitive with other state-of-the-art object detection schemes and show robustness to object articulations, clutter, and occlusion. Our major contributions are the application of boosted local contour-based features for object detection in a partially supervised learning framework, and an efficient new boosting procedure for simultaneously selecting features and estimating per-feature parameters. Jamie Shotton, Andrew Blake 0001, Roberto Cipolla |
ICCV | 3 |
| 2005 | Using Frontier Points to Recover Shape, Reflectance and IllumunationabstractWe describe a method to recover the surface reflectance and the 3D shape of a non-Lambertian object as well as illumination, from a collection of images. It is based on the so-called frontier points, which are extracted from the outlines of an object. Frontier points provide 3D locations on the object surface where the surface normal is known. This information is exploited to infer the surface reflectance of the object and the light distribution of the scene both under varying illumination and fixed vantage point, and under varying vantage point and fixed illumination. We also show how to apply frontier points for shape recovery in photometric stereo. The effectiveness of frontier points for recovering reflectance, illumination and shape is confirmed by a number of experiments on both real and synthetic data. George Vogiatzis, Paolo Favaro, Roberto Cipolla |
ICCV | 3 |
| 2005 | Combining interest points and edges for content-based image retrievalabstractThis paper presents a novel approach using combined features to retrieve images containing specific objects, scenes or buildings. The content of an image is characterized by two kinds of features: Harris-Laplace interest points described by the SIFT descriptor and edges described by the edge color histogram. Edges and corners contain the maximal amount of information necessary for image retrieval. The feature detection in this work is an integrated process: edges are detected directly based on the Harris function; Harris interest points are detected at several scales and Harris-Laplace interest points are found using the Laplace function. The combination of edges and interest points brings efficient feature detection and high recognition ratio to the image retrieval system. Experimental results show this system has good performance. Junqiu Wang, Hongbin Zha, Roberto Cipolla |
ICIP (3) | 3 |
| 2005 | Vision-based Global Localization Using a Visual VocabularyabstractThis paper presents a novel coarse-to-fine global localization approach that is inspired by object recognition and text retrieval techniques. Harris-Laplace interest points characterized by SIFT descriptors are used as natural landmarks. These descriptors are indexed into two databases: an inverted index and a location database. The inverted index is built based on a visual vocabulary learned from the feature descriptors. In the location database, each location is directly represented by a set of scale invariant descriptors. The localization process consists of two stages: coarse localization and fine localization. Coarse localization from the inverted index is fast but not accurate enough; whereas localization from the location database using voting algorithm is relatively slow but more accurate. The combination of coarse and fine stages makes fast and reliable localization possible. In addition, if necessary, the localization result can be verified by epipolar geometry between the representative view in database and the view to be localized. Experimental results show that our approach is efficient and reliable. Junqiu Wang, Roberto Cipolla, Hongbin Zha |
ICRA | 2 |
| 2005 | Sparse Bayesian Learning for Efficient Visual TrackingabstractThis paper extends the use of statistical learning algorithms for object localization. It has been shown that object recognizers using kernel-SVMs can be elegantly adapted to localization by means of spatial perturbation of the SVM. While this SVM applies to each frame of a video independently of other frames, the benefits of temporal fusion of data are well-known. This is addressed here by using a fully probabilistic Relevance Vector Machine (RVM) to generate observations with Gaussian distributions that can be fused over time. Rather than adapting a recognizer, we build a displacement expert which directly estimates displacement from the target region. An object detector is used in tandem, for object verification, providing the capability for automatic initialization and recovery. This approach is demonstrated in real-time tracking systems where the sparsity of the RVM means that only a fraction of CPU time is required to track at frame rate. An experimental evaluation compares this approach to the state of the art showing it to be a viable method for long-term region tracking. Oliver Williams, Andrew Blake 0001, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2004 | An Illumination Invariant Face Recognition System for Access Control using VideoabstractIllumination and pose invariance are the most challenging aspects of face recognition. In this paper we describe a fully automatic face recognition system that uses video information to achieve illumination and pose robustness. In the proposed method, highly nonlinear manifolds of face motion are approximated using three Gaussian pose clusters. Pose robustness is achieved by comparing the corresponding pose clusters and probabilistically combining the results to derive a measure of similarity between two manifolds. Illumination is normalized on a per-pose basis. Region-based gamma intensity correction is used to correct for coarse illumination changes, while further refinement is achieved by combining a learnt linear manifold of illumination variation with constraints on face pattern distribution, derived from video. Comparative experimental evaluation is presented and the proposed method is shown to greatly outperform state-of-the-art algorithms. Consistent recognition rates of 94-100% are achieved across dramatic changes in illumination. 1 Ognjen Arandjelovic, Roberto Cipolla |
BMVC | 2 |
| 2004 | An Image-Based System for Urban NavigationabstractWe describe the prototype of a system intended to allow a userto navigate in an urban environment using a mobile telephone equipped wi th a camera. The system uses a database of views of building facades to det ermine the pose of a query view provided by the user. Our method is based o n a novel wide-baseline matching algorithm that can identify corres ponding building facades in two views despite significant changes of viewpoin t and lighting. We show that our system is capable of localising query views r eliably in a large part of Cambridge city centre. Duncan P. Robertson, Roberto Cipolla |
BMVC | 2 |
| 2004 | Likelihood Models For Template Matching using the PDF Projection TheoremabstractTemplate matching techniques are widely used in many computer vision tasks. Generally, a likelihood value is calculated from similarity measures, however the relation between these measures and the data likelihood is often incorrectly stated. It is clear that accurate likelihood estimation will improve the efficiency of the matching algorithms. This paper introduces a novel method for estimating the likelihood PDFS accurately based on the PDF Projection Theorem, which provides the correct relation between the feature likelihood and the data likelihood, permitting the use of different types of features for different types of objects and still estimating consistent likelihoods. The proposed method removes the normalization and bias problems that are usually associated with the likelihood calculations. We demonstrate that it significantly improves template matching in pose estimation problems. Qualitative and quantitative results are compared against traditional likelihood estimation schemes. Arasanathan Thayananthan, Ramanan Navaratnam, Philip Torr 0001, Roberto Cipolla |
BMVC | 4 |
| 2004 | Reconstructing Relief SurfacesabstractThis paper generalizes Markov Random Field (MRF) stereo methods to the generation of surface relief (height) fields rather than disparity or depth maps. This generalization enables the reconstruction of complete object models using the same algorithms that have been previously used to compute depth maps in binocular stereo. In contrast to traditional dense stereo where the parametrization is image based, here we advocate a parametrization by a height field over any base surface. In practice, the base surface is a coarse approximation to the true geometry, e.g., a bounding box, visual hull or triangulation of sparse correspondences, and is assigned or computed using other means. A dense set of sample points is defined on the base surface, each with a fixed normal direction and unknown height value. The estimation of heights for the sample points is achieved by a belief propagation technique. Our method provides a viewpoint independent smoothness constraint, a more compact parametrization and explicit handling of occlusions. We present experimental results on real scenes as well as a quantitative evaluation on an artificial scene. George Vogiatzis, Philip Torr 0001, Steven M. Seitz, Roberto Cipolla |
BMVC | 4 |
| 2004 | Sparse Finite Elements for Geodesic Contours with Level-Sets
Martin Weber 0001, Andrew Blake 0001, Roberto Cipolla |
ECCV (2) | 3 |
| 2004 | The Variational Ising Classifier (VIC) Algorithm for Coherently Contaminated DataabstractThere has been substantial progress in the past decade in the development of object classifiers for images, for example of faces, humans and vehi- cles. Here we address the problem of contaminations (e.g. occlusion, shadows) in test images which have not explicitly been encountered in training data. The Variational Ising Classifier (VIC) algorithm models contamination as a mask (a field of binary variables) with a strong spa- tial coherence prior. Variational inference is used to marginalize over contamination and obtain robust classification. In this way the VIC ap- proach can turn a kernel classifier for clean data into one that can tolerate contamination, without any specific training on contaminated positives. 1 Introduction Recent progress in discriminative object detection, especially for faces, has yielded good performance and efficiency [1, 2, 3, 4]. Such systems are capable of classifying those positives that can be generalized from positive training data. This is restrictive in practice in that test data may contain distortions that take it outside the strict ambit of the training positives. One example would be lighting changes (to a face) but this can be addressed reasonably effectively by a normalizing transformation applied to training and test images; doing so is common practice in face classification. Other sorts of disruption are not so easily factored out. A prime example is partial occlusion. The aim of this paper is to extend a classifier trained on clean positives to accept also partially occluded positives, without further training. The approach is to capture some of the regularity inherent in a typical pattern of contamination, namely its spatial coherence. This can be thought of as extending the generalizing capability of a classifier to tolerate the sorts of image distortion that occur as a result of contamination. As done previously in one-dimension, for image contours [5], the Variational Ising Classi- fier (VIC) models contamination explicitly as switches with a strong coherence prior in the form of an Ising model, but here over the full two-dimensional image array. In addition, the Ising model is loaded with a bias towards non-contamination. The aim is to incorporate these hidden contamination variables into a kernel classifier such as [1, 3]. In fact the Rel- evance Vector Machine (RVM) is particularly suitable [6] as it is explicitly probabilistic, so that contamination variables can be incorporated as a hidden layer of random variables. edge neighbours of i i Figure 1: The 2D Ising model is applied over a graph with edges e between neigh- bouring pixels (connected 4-wise). Classification is done by marginalization over all possible configurations of the hidden vari- able array, and this is made tractable by variational (mean field) inference. The inference scheme makes use of "hallucination" to fill in parts of the object that are unobserved due to occlusion. Results of VIC are given for face detection. First we show that the classifier performance is not significantly damaged by the inclusion of contamination variables. Then a contam- inated test set is generated using real test images and computer generated contaminations. Over this test data the VIC algorithm does indeed perform significantly better than a con- ventional classifier (similar to [4]). The hidden variable layer is shown to operate effec- tively, successfully inferring areas of contamination. Finally, inference of contamination is shown working on real images with real contaminations. 2 Bayesian modelling of contamination Classification requires P (F |I), the posterior for the proposition F that an object is present given the image data intensity array I. This can be computed in terms of likelihoods P (F | I) = P (I | F )P (F )/ P (I | F )P (F ) + P (I | F )P (F ) (1) so then the test P (F | I) > 1 becomes 2 log P (I | F ) - log P (I | F ) > t (2) where t is a prior-dependent threshold that controls the tradeoff between positive and neg- ative classification errors. Suppose we are given a likelihood P (I|, F ) for the presence of a face given contamination , an array of binary "observation" variables corresponding to each pixel Ij of I, such that j = 0 indicates contamination at that pixel, whereas j = 1 indicates a successfully observed pixel. Then, in principle, P (I|F ) = P (I|, F )P (), (3) (making the reasonable assumption P (|F ) = P (), that the pattern of contamination is object independent) and similarly for log P (I | F ). The marginalization itself is intractable, requiring a summation over all 2N possible configurations of , for images with N pixels. Approximating that marginalization is dealt with in the next section. In the meantime, there are two other problems to deal with: specifying the prior P (); and specifying the likeli- hood under contamination P (I|, F ) given only training data for the unoccluded object. 2.1 Prior over contaminations The prior contains two terms: the first expresses the belief that contamination will occur in coherent regions of a subimage. This takes the form of an Ising model [7] with energy UI() that penalizes adjacent pixels which differ in their labelling (see Figure 1); the second term UC biases generally against contamination a priori and its balance with the first term is mediated by the constant . The total prior energy is then U () = UI() + UC() = [1 - (e - )] + ( 1 e2 j ), (4) e j where (x) = 1 if x = 0 and 0 otherwise, and e1, e2 are the indices of the pixels at either end of edge e (figure 1). The prior energy determines a probability via a temperature constant 1/T0 [7]: P () e-U()/T0 = e-UI()/T0e-UC()/T0 (5) 2.2 Relevance vector machine An unoccluded classifier P (F |I, = 0) can be learned from training data using a Rele- vance Vector Machine (RVM) [6], trained on a database of frontal face and non-face im- ages [8] (see Section 4 for details). The probabilistic properties of the RVM make it a good choice when (later) it comes to marginalising over . For now we consider how to construct the likelihood itself. First the conventional, unoccluded case is considered for which the posterior P (F |I) is learned from positive and negative examples. Kernel functions [9] are computed between a candidate image I and a subset of relevance vectors {xk}, retained from the training set. Gaussian kernels are used here to compute y(I) = wk exp - (Ij - xkj)2 . (6) k j where wk are learned weights, and xkj is the jth pixel of the kth relevance vector. Then the posterior is computed via the logistic sigmoid function as 1 P (F |I, = 1) = (y(I)) = . (7) 1 + e-y(I) and finally the unoccluded data-likelihood would be P (I|F, = 1) (y(I))/P (F ). (8) 2.3 Hallucinating appearance The aim now is to derive the occluded likelihood from the unoccluded case, where the con- tamination mask is known, without any further training. To do this, (8) must be extended to give P (I|F, ) for arbitrary masks , despite the fact the pixels Ij from the object are not observed wherever j = 0. In principle one should take into account all possible (or at least probable) values for the occluded pixels. Here, for simplicity, a single fixed hallu- cination is substituted for occluded pixels, then we proceed as if those values had actually been observed. This gives P (I|F, ) (~ y(I, ))/P (F ) (9) Oliver Williams, Andrew Blake 0001, Roberto Cipolla |
NIPS | 3 |
| 2004 | Modelling and Interpretation of Architecture from Several Images
Anthony R. Dick, Philip Torr 0001, Roberto Cipolla |
Int. J. Comput. Vis. | 3 |
| 2004 | Towards a complete dense geometric and photometric reconstruction under varying pose and illumination
Martin Weber 0001, Andrew Blake 0001, Roberto Cipolla |
Image Vis. Comput. | 3 |
| 2004 | Reconstruction of surfaces of revolution from single uncalibrated views
Kwan-Yee Kenneth Wong, Paulo R. S. Mendonça, Roberto Cipolla |
Image Vis. Comput. | 3 |
| 2004 | Layered Motion Segmentation and Depth Ordering by Tracking EdgesabstractThis paper presents a new Bayesian framework for motion segmentation--dividing a frame from an image sequence into layers representing different moving objects--by tracking edges between frames. Edges are found using the Canny edge detector, and the Expectation-Maximization algorithm is then used to fit motion models to these edges and also to calculate the probabilities of the edges obeying each motion model. The edges are also used to segment the image into regions of similar color. The most likely labeling for these regions is then calculated by using the edge probabilities, in association with a Markov Random Field-style prior. The identification of the relative depth ordering of the different motion layers is also determined, as an integral part of the process. An efficient implementation of this framework is presented for segmenting two motions (foreground and background) using two frames. It is then demonstrated how, by tracking the edges into further frames, the probabilities may be accumulated to provide an even more accurate and robust estimate, and segment an entire sequence. Further extensions are then presented to address the segmentation of more than two motions. Here, a hierarchical method of initializing the Expectation-Maximization algorithm is described, and it is demonstrated that the Minimum Description Length principle may be used to automatically select the best number of motion layers. The results from over 30 sequences (demonstrating both two and three motions) are presented and discussed. Tom Drummond, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2004 | Reconstruction of sculpture from its profiles with unknown camera positionsabstractProfiles of a sculpture provide rich information about its geometry, and can be used for shape recovery under known camera motion. By exploiting correspondences induced by epipolar tangents on the profiles, a successful solution to motion estimation from profiles has been developed in the special case of circular motion. The main drawbacks of using circular motion alone, namely the difficulty in adding new views and part of the object always being invisible, can be overcome by incorporating arbitrary general views of the object and registering its new profiles with the set of profiles resulted from the circular motion. In this paper, we describe a complete and practical system for producing a three-dimensional (3-D) model from uncalibrated images of an arbitrary object using its profiles alone. Experimental results on various objects are presented, demonstrating the quality of the reconstructions using the estimated motion. Kwan-Yee Kenneth Wong, Roberto Cipolla |
IEEE Trans. Image Process. | 2 |
| 2003 | Learning a Kinematic Prior for Tree-Based FilteringabstractThe aim in this paper is to track articulated hand motion from monocular video. Bayesian filtering is implemented by using a tree-based representation of the posterior distribution. Each tree node corresponds to a partition of the state space with piecewise constant density. In a hierarchical search regions with low probability mass can be rapidly discarded, while the modes of the posterior can be approximated to high precision. Large sets of training data are captured using a data glove, and two techniques for constructing the tree are described: One method is to cluster the collected data points using a hierarchical clustering algorithm, and use the cluster centres as nodes. Alternatively, a lower dimensional eigenspace can be partitioned using a grid at multiple resolutions, and each partition centre corresponds to a node in the tree. The effectiveness of these techniques is demonstrated by using them for tracking 3D articulated hand motion in front of a cluttered background. 1 Arasanathan Thayananthan, Björn Stenger, Philip Torr 0001, Roberto Cipolla |
BMVC | 4 |
| 2003 | Bayesian Stochastic Mesh Optimization for 3D reconstructionabstractWe describe a mesh based approach to the problem of structure from motion. The input to the algorithm is a small set of images, sparse noisy feature correspondences (such as those provided by a Harris corner detector and cross correlation) and the camera geometry plus calibration. The output is a 3D mesh, that when projected onto each view, is visually consistent with the images. There are two contributions in this paper. The first is a Bayesian formulation in which simplicity and smoothness assumptions are encoded in the prior distribution. The resulting posterior is optimized by simulated annealing. The second and more important contribution is a way to make this optimization scheme more efficient. Generic simulated annealing has been long studied in computer vision and is thought to be highly inefficient. This is often because the proposal distribution searches regions of space which are far from the modes. In order to improve the performance of simulated annealing it has long been acknowledged that choice of the correct proposal distribution is of paramount importance to convergence. Taking inspiration from RANSAC andimportance sampling we craft a proposal distribution that is tailored to the problem of structure from motion. This makes our approach particularly robust to noise and ambiguity. We show results for an artificial object and an architectural scene. George Vogiatzis, Philip Torr 0001, Roberto Cipolla |
BMVC | 3 |
| 2003 | Shape Context and Chamfer Matching in Cluttered ScenesabstractThis paper compares two methods for object localization from contours: shape context and chamfer matching of templates. In the light of our experiments, we suggest improvements to the shape context: shape contexts are used to find corresponding features between model and image. In real images it is shown that the shape context is highly influenced by clutters; furthermore, even when the object is correctly localized, the feature correspondence may be poor. We show that the robustness of shape matching can be increased by including a figural continuity constraint. The combined shape and continuity cost is minimized using the Viterbi algorithm on features, resulting in improved localization and correspondence. Our algorithm can be generally applied to any feature based shape matching method. Chamfer matching correlates model templates with the distance transform of the edge image. This can be done efficiently using a coarse-to-fine search over the transformation parameters. The method is robust in clutter, however, multiple templates are needed to handle scale, rotation and shape variation. We compare both methods for locating hand shapes in cluttered images, and applied to word recognition in EZ-Gimpy images. Arasanathan Thayananthan, Björn Stenger, Philip Torr 0001, Roberto Cipolla |
CVPR (1) | 4 |
| 2003 | Filtering Using a Tree-Based EstimatorabstractWithin this paper a new framework for Bayesian tracking is presented, which approximates the posterior distribution at multiple resolutions. We propose a tree-based representation of the distribution, where the leaves define a partition of the state space with piecewise constant density. The advantage of this representation is that regions with low probability mass can be rapidly discarded in a hierarchical search, and the distribution can be approximated to arbitrary precision. We demonstrate the effectiveness of the technique by using it for tracking 3D articulated and nonrigid motion in front of cluttered background. More specifically, we are interested in estimating the joint angles, position and orientation of a 3D hand model in order to drive an avatar. Björn Stenger, Arasanathan Thayananthan, Philip Torr 0001, Roberto Cipolla |
ICCV | 4 |
| 2003 | A Sparse Probabilistic Learning Algorithm for Real-Time TrackingabstractWe address the problem of applying powerful pattern recognition algorithms based on kernels to efficient visual tracking. Recently S. Avidan, (2001) has shown that object recognizers using kernel-SVMs can be elegantly adapted to localization by means of spatial perturbation of the SVM, using optic flow. Whereas Avidan's SVM applies to each frame of a video independently of other frames, the benefits of temporal fusion of data are well known. Using a fully probabilistic 'relevance vector machine' (RVM) to generate observations with Gaussian distributions that can be fused over time is addressed. To improve performance further, rather than adapting a recognizer, we build a localizer directly using the regression form of the RVM. A classification SVM is used in tandem, for object verification, and this provides the capability of automatic initialization and recovery. The approach is demonstrated in real-time face and vehicle tracking systems. The 'sparsity' of the RVMs means that only a fraction of CPU time is required to track at frame rate. Tracker output is demonstrated in a camera management task in which zoom and pan are controlled in response to speaker/vehicle position and orientation, over an extended period. The advantages of temporal fusion in this system are demonstrated. Oliver Williams, Andrew Blake 0001, Roberto Cipolla |
ICCV | 3 |
| 2003 | Camera Calibration from Surfaces of RevolutionabstractThis paper addresses the problem of calibrating a pinhole camera from images of a surface of revolution. Camera calibration is the process of determining the intrinsic or internal parameters (i.e., aspect ratio, focal length, and principal point) of a camera, and it is important for both motion estimation and metric reconstruction of 3D models. In this paper, a novel and simple calibration technique is introduced, which is based on exploiting the symmetry of images of surfaces of revolution. Traditional techniques for camera calibration involve taking images of some precisely machined calibration pattern (such as a calibration grid). The use of surfaces of revolution, which are commonly found in daily life (e.g., bowls and vases), makes the process easier as a result of the reduced cost and increased accessibility of the calibration objects. In this paper, it is shown that two images of a surface of revolution will provide enough information for determining the aspect ratio, focal length, and principal point of a camera with fixed intrinsic parameters. The algorithms presented in this paper have been implemented and tested with both synthetic and real data. Experimental results show that the camera calibration method presented is both practical and accurate. Kwan-Yee Kenneth Wong, Paulo R. S. Mendonça, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2002 | Towards a Complete Dense Geometric and Photometric Reconstruction under Varying Pose and IlluminationabstractThis paper proposes a novel framework to construct a geometric and photometric model of a viewed object that can be used for visualisation in arbitrary pose and illumination. The method is solely based on images and does not require any specialised equipment. We assume that the object has a piece-wise smooth surface and that its reflectance can be modelled using a parametric bidirectional reflectance distribution function. Without assuming any prior knowledge on the object, geometry and reflectance have to be estimated simultaneously and occlusion and shadows have to be treated consistently. We exploit the geometric and photometric consistency using the fact that surface orientation and reflectance are local invariants. In a first implementation, we demonstrate the method using a Lambertian object placed on a turn-table and illuminated by a number of unknown point light-sources. A discrete voxel model is initialised to the visual hull and voxels identified as inconsistent with the invariants are removed iteratively. The resulting model is used to render images in novel pose and illumination. Martin Weber 0001, Andrew Blake 0001, Roberto Cipolla |
BMVC | 3 |
| 2002 | Reconstruction of Surfaces of Revolution from Single Uncalibrated ViewsabstractThis paper addresses the problem of recovering the 3D shape of a surface of revolution from a single uncalibrated perspective view.The algorithm introduced here makes use of the invariant properties of a surface of revolution and its silhouette to locate the image of the revolution axis, and to calibrate the focal length of the camera.The image is then normalized and rectified such that the resulting silhouette exhibits bilateral symmetry.Such a rectification leads to a simpler differential analysis of the silhouette, and yields a simple equation for depth recovery.Ambiguities in the reconstruction are analyzed and experimental results on real images are presented, which demonstrate the quality of the reconstruction. Kwan-Yee Kenneth Wong, Paulo R. S. Mendonça, Roberto Cipolla |
BMVC | 3 |
| 2002 | A Bayesian Estimation of Building Shape Using MCMC
Anthony R. Dick, Philip Torr 0001, Roberto Cipolla |
ECCV (2) | 3 |
| 2002 | Building Architectural Models from Many Views Using Map Constraints
Duncan P. Robertson, Roberto Cipolla |
ECCV (2) | 2 |
| 2002 | Real-time tracking of complex structures with on-line camera calibration
Tom Drummond, Roberto Cipolla |
Image Vis. Comput. | 2 |
| 2002 | Structure and motion estimation from apparent contours under circular motion
Kwan-Yee Kenneth Wong, Paulo R. S. Mendonça, Roberto Cipolla |
Image Vis. Comput. | 3 |
| 2002 | Estimating the Fundamental Matrix via Constrained Least-Squares: A Convex ApproachabstractIn this paper, a new method for the estimation of the fundamental matrix from point correspondences in stereo vision is presented. The minimization of the algebraic error is performed while taking explicitly into account the rank-two constraint on the fundamental matrix. It is shown how this nonconvex optimization problem can be solved avoiding local minima by using recently developed convexification techniques. The obtained estimate of the fundamental matrix turns out to be more accurate than the one provided by the linear criterion, where the rank constraint of the matrix is imposed after its computation by setting the smallest singular value to zero. This suggests that the proposed estimate can be used to initialize nonlinear criteria, such as the distance to epipolar lines and the gradient criterion, in order to obtain a more accurate estimate of the fundamental matrix. Graziano Chesi, Andrea Garulli, Antonio Vicino, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2002 | Real-Time Visual Tracking of Complex StructuresabstractPresents a framework for three-dimensional model-based tracking. Graphical rendering technology is combined with constrained active contour tracking to create a robust wire-frame tracking system. It operates in real time at video frame rate (25 Hz) on standard hardware. It is based on an internal CAD model of the object to be tracked which is rendered using a binary space partition tree to perform hidden line removal. A Lie group formalism is used to cast the motion computation problem into simple geometric terms so that tracking becomes a simple optimization problem solved by means of iterative reweighted least squares. A visual servoing system constructed using this framework is presented together with results showing the accuracy of the tracker. The paper then describes how this tracking system has been extended to provide a general framework for tracking in complex configurations. The adjoint representation of the group is used to transform measurements into common coordinate frames. The constraints are then imposed by means of Lagrange multipliers. Results from a number of experiments performed using this framework are presented and discussed. Tom Drummond, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Camera Self-Calibration from Unknown Planar Structures Enforcing the Multiview Constraints between CollineationsabstractIn this paper, we describe an efficient method to impose the constraints existing between the collineations between images which can be computed from a sequence of views of a planar structure. These constraints are usually not taken into account by multiview techniques in order not to increase the computational complexity of the algorithms. However, imposing the constraints is very useful since it allows a reduction of geometric errors in the reprojected features and provides a consistent set of collineations which can be used for several applications such as mosaicing, reconstruction, and self-calibration. In order to show the validity of our approach, this paper focus on self-calibration from unknown planar structures proposing a method exploiting the consistent set of collineations. Our method can deal with an arbitrary number of views and an arbitrary number of planes and varying camera internal parameters. However, for simplicity, this papers will only discuss the case with one plane in several views. The results obtained with synthetic and real data are very accurate and stable even when using only few images. Ezio Malis, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | Model-Based Hand Tracking Using an Unscented Kalman FilterabstractThis paper presents a novel method for hand tracking. It uses a 3D model built from quadrics which approximates the anatomy of a human hand. This approach allows for the use of results from projective geometry that yield an elegant technique to generate the projection of the model as a set of conics, as well as providing an efficient ray tracing algorithm to handle self-occlusion. Once the model is projected, an Unscented Kalman Filter is used to update its pose in order to minimise the geometric error between the model projection and a video sequence on the background. Results from experiments with real data show the accuracy of the technique. 1 Björn Stenger, Paulo R. S. Mendonça, Roberto Cipolla |
BMVC | 3 |
| 2001 | A Practical Method for Estimation of Point Light-SourcesabstractWe introduce a general model for point light-sources and show how the parameters of a source at finite distance can be estimated by shading on an object with known geometry and Lambertian reflectance. The parameters we estimate include not only the direction but also the location of the source. Furthermore, we argue that the system introduced here can be used in arbitrarily complex background illumination, provided one can switch on and off the light-source that is to be estimated. 1 Martin Weber 0001, Roberto Cipolla |
BMVC | 2 |
| 2001 | A Linear Iterative Method for Auto-Calibration using the DAC EquationabstractIn this paper, an iterative algorithm for auto-calibration is presented. The proposed algorithm switches between linearly estimating the dual of the absolute conic and the intrinsic parameters, while also incorporating the rank-3 constraint on the intrinsic parameters. The most important property of the algorithm is that it is completely general in the sense that any type of constraint on the intrinsic parameters might be used. The proposed algorithm locates in-between of a non-linear optimization and initial linear computation, and provides robust and sufficiently accurate initial values for a bundle adjustment routine. The performance of the algorithm is shown for both simulated and real data, especially in the important case of natural (zero skew and unit aspect ratio) cameras. Yongduek Seo, Anders Heyden, Roberto Cipolla |
CVPR (1) | 3 |
| 2001 | Model-Based 3D Tracking of an Articulated HandabstractThis paper presents a practical technique for model-based 3D hand tracking. An anatomically accurate hand model is built from truncated quadrics. This allows for the generation of 2D profiles of the model using elegant tools from projective geometry, and for an efficient method to handle self-occlusion. The pose of the hand model is estimated with an Unscented Kalman filter (UKF), which minimizes the geometric error between the profiles and edges extracted from the images. The use of the UKF permits higher frame rates than more sophisticated estimation methods such as particle filtering, whilst providing higher accuracy than the extended Kalman filter The system is easily scalable from single to multiple views, and from rigid to articulated models. First experiments on real data using one and two cameras demonstrate the quality of the proposed method for tracking a 7 DOF hand model. Björn Stenger, Paulo R. S. Mendonça, Roberto Cipolla |
CVPR (2) | 3 |
| 2001 | A Probabilistic Framework for Space Carving
Adrian Broadhurst, Tom Drummond, Roberto Cipolla |
ICCV | 3 |
| 2001 | Combining Single View Recognition and Multiple View Stereo for Architectural Scenes
Anthony R. Dick, Philip Torr 0001, Simon J. Ruffle, Roberto Cipolla |
ICCV | 4 |
| 2001 | Real-Time Tracking of Highly Articulated Structures in the Presence of Noisy MeasurementsabstractThis paper presents a novel approach for model-based real-time tracking of highly articulated structures such as humans. This approach is based on an algorithm which efficiently propagates statistics of probability distributions through a kinematic chain to obtain maximum a posteriori estimates of the motion of the entire structure. This algorithm yields the least squares solution in linear time (in the number of components of the model) and can also be applied to non-Gaussian statistics using a simple but powerful trick. The resulting implementation runs in real-time on standard hardware without any pre-processing of the video data and can thus operate on live video. Results from experiments performed using this system are presented and discussed. Tom Drummond, Roberto Cipolla |
ICCV | 2 |
| 2001 | Structure and Motion from SilhouettesabstractThis paper addresses the problem of recovering structure and motion from silhouettes. Silhouettes are projections of contour generators which are viewpoint dependent, and hence do not readily provide point correspondences for exploitation in motion estimation. Previous works have exploited correspondences induced by epipolar tangencies, and a successful solution has been developed in the special case of circular motion (turnable sequences). However, the main drawbacks are (1) new views cannot be added easily at a later time, and (2) part of the structure will always remain invisible under circular motion. In this paper we overcome the above problems by incorporating arbitrary general views and estimating the camera poses using silhouettes alone. We present a complete and practical system which produces high quality 3D models from 2D uncalibrated silhouettes. The 3D models thus obtained can be refined incrementally by adding new arbitrary views and estimating their poses. Experimental results on various objects are presented, demonstrating the quality of the reconstructions. Kwan-Yee Kenneth Wong, Roberto Cipolla |
ICCV | 2 |
| 2001 | Reconstruction of sculpture from uncalibrated image profilesabstractProfiles of a sculpture provide rich information about its geometry, and can be used for model reconstruction under known camera motion. By exploiting correspondences induced by epipolar tangents on the profiles, a successful solution to motion estimation has been developed for the case of circular motion. Arbitrary general views can then be incorporated to refine the model built from circular motion. Kwan-Yee Kenneth Wong, Roberto Cipolla |
ICIP (1) | 2 |
| 2001 | Epipolar Geometry from Profiles under Circular MotionabstractAddresses the problem of motion estimation from profiles (apparent contours) of an object rotating on a turntable in front of a single camera. A practical and accurate technique for solving this problem from profiles alone is developed. It is precise enough to reconstruct the shape of the object. No correspondences between points or lines are necessary. Symmetry of the surface of revolution swept out by the rotating object is exploited to obtain the image of the rotation axis and the homography relating epipolar lines in two views robustly and elegantly. These, together with geometric constraints for images of rotating objects, are used to obtain first the image of the horizon, which is the projection of the plane that contains the camera centers, and then the epipoles, thus fully determining the epipolar geometry of the image sequence. The estimation of this geometry by this sequential approach avoids many of the problems found in other algorithms. The search for the epipoles, by far the most critical step, is carried out as a simple 1D optimization. Parameter initialization is trivial and completely automatic at all stages. After the estimation of the epipolar geometry, the Euclidean motion is recovered using the fixed intrinsic parameters of the camera obtained either from a calibration grid or from self-calibration techniques. Finally, the spinning object is reconstructed from its profiles using the motion estimated in the previous stage. Results from real data are presented, demonstrating the efficiency and usefulness of the proposed methods. Paulo R. S. Mendonça, Kwan-Yee Kenneth Wong, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2000 | A Statistical Consistency Check for the Space Carving AlgorithmabstractThis paper investigates the use of the Space Carving algorithm with outdoor image sequences, using a lambertian lighting model. A new consistency function is proposed that uses a statistical comparison instead of the voxel centroid sampling that was initially proposed. This is important when there is more detail in the images than can be stored in a voxel representation. The new function is evaluated using synthetic data and real image sequences. 1 Introduction Many dierent techniques have been applied to the problem of reconstructing three-dimensional shape from image sequences. These techniques all work well for constrained problems, such as small camera baseline [6, 1, 2, 8], or smooth curved objects [4]. Recently voxel based algorithms [5, 9] have been demonstrated which can reconstruct very complex shapes, but at the cost of large memory requirements. One of these algorithms is called Space Carving [5] (see section 2.1 for a description). In this paper the eectiveness o... Adrian Broadhurst, Roberto Cipolla |
BMVC | 2 |
| 2000 | 3D Model Acquisition by Tracking 2D WireframesabstractThis paper presents a semi-automatic wireframe acquisition system. The sys-tem uses real-time (25Hz) tracking of a user specified 2D wireframe and in-termittent camera pose parameters to accumulate 3D position information. The 2D tracking framework enables the application of model-based con-straints in an intuitive way which interacts naturally with the Kalman filter formulation used. In particular, this is used to introduce feedback from the current 3D shape estimate to improve the robustness of the 2D tracking. The scheme allows wireframe models of simple edge based objects to be built in around 5 minutes. 1 Matthew A. Brown, Tom Drummond, Roberto Cipolla |
BMVC | 3 |
| 2000 | Automatic 3D Modelling of ArchitectureabstractThis paper describes a system which automatically derives 3D models of ar-chitectural scenes from multiple images. This system differs from previous structure from motion algorithms in that it explicitly makes use of strong ge-ometric constraints such as perpendicularity and verticality which are likely to be found in architecture. Structure is compactly represented as a piece-wise planar model which is initialised automatically by segmenting a feature-based reconstruction. An efficient technique for evaluation of model likeli-hood is also presented, which allows a rapid search through a large number of 3D models. 1 Anthony R. Dick, Philip Torr 0001, Roberto Cipolla |
BMVC | 3 |
| 2000 | An Interactive System for Constraint-Based Modelling
Duncan P. Robertson, Roberto Cipolla |
BMVC | 2 |
| 2000 | Segmentation of Multiple Motions by Edge Tracking between Two FramesabstractThis paper presents a method for segmenting multiple motions using edges. Recent work in this field has been constrained to the case of two motions, and this paper demonstrates that the approach can be extended to more than two motions. The image is first segmented into regions, and then the framework determines the motions present and labels the edges in the image. Initial-isation is particularly difficult, and a novel scheme is proposed which re-cursively splits motions to provide the Expectation-Maximisation algorithm with a reasonable guess, and a Minimum Description Length approach is used to determine the best number of models to use. The edge labels are then used to determine the the region labelling. A global optimisation is intro-duced to refine the motions and provide the most likely region labelling. 1 Tom Drummond, Roberto Cipolla |
BMVC | 3 |
| 2000 | On the Estimation of the Fundamental Matrix: A Convex Approach to Constrained Least-Squares
Graziano Chesi, Andrea Garulli, Antonio Vicino, Roberto Cipolla |
ECCV (1) | 4 |
| 2000 | Real-Time Tracking of Multiple Articulated Structures in Multiple Views
Tom Drummond, Roberto Cipolla |
ECCV (2) | 2 |
| 2000 | Multi-view Constraints between Collineations: Application to Self-Calibration from Unknown Planar Structures
Ezio Malis, Roberto Cipolla |
ECCV (2) | 2 |
| 2000 | Camera Pose Estimation and Reconstruction from Image Profiles under Circular Motion
Paulo R. S. Mendonça, Kwan-Yee Kenneth Wong, Roberto Cipolla |
ECCV (2) | 3 |
| 2000 | Motion Segmentation by Tracking Edge Information over Multiple Frames
Tom Drummond, Roberto Cipolla |
ECCV (2) | 3 |
| 2000 | Layer Extraction with a Bayesian Model of Shapes
Philip Torr 0001, Anthony R. Dick, Roberto Cipolla |
ECCV (2) | 3 |
| 2000 | Self-Calibration of Zooming Cameras Observing an Unknown Planar StructureabstractIn this paper, we propose a new self-calibration technique for cameras with changing zoom observing only a planar structure. The method does not need any metric or topologic knowledge about the structure since it is based on the estimation of the collineations existing between several views of a plane (thus only image correspondences are needed). The constraints existing between all the collineations are imposed using a very simple and efficient technique which does not need the solution of a complex optimisation problem. Finally, even if the structure of the plane is unknown it must be the same for all the images and this provides some constraints which allow the recovering of the varying focal length. Ezio Malis, Roberto Cipolla |
ICPR | 2 |
| 2000 | Automatic Segmentation and Matching of Planar Contours for Visual ServoingabstractWe present a complete system for segmenting, matching and tracking planar contours for use in visual servoing. Our system can be used with arbitrary contours of any shape and without any prior knowledge of their models. The system is first shown the target view. A selected contour is automatically extracted and its image shape is stored. The robot and object are then moved and the system automatically identifies the target. The matching step is done together with the estimation of the homography matrix between the two views of the contour. Then, a 2 1/2 D visual servoing technique is used to reposition the end-effector of a robot at the target position relative to the planar contour. The system has been successfully tested on several contours with very complex shapes such as leaves, keys and the coastal outlines of islands. Graziano Chesi, Ezio Malis, Roberto Cipolla |
ICRA | 3 |
| 2000 | Application of Lie Algebras to Visual Servoing
Tom Drummond, Roberto Cipolla |
Int. J. Comput. Vis. | 2 |
| 1999 | The Applications of Uncalibrated Occlusion JunctionsabstractWhen a scene is viewed from two different viewpoints there are often regions which are only visible in one of the two views. These occluded regions give important information about depth discontinuities in the image. When a background line is obscured by a foreground object, it forms a T-junction in the image, and these junction points can be used to detect occlusion. Two algorithms are presented in this paper. The first algorithm uses the trifocal tensor to automatically locate T-junctions visible in three views, whilst the second application uses junction points to obtain a constraint for the detection of planar surfaces. 1 Introduction The stereo correspondence problem is one of the oldest problems in computer vision, yet a general and robust solution remains elusive. One of the major difficulties for stereo matching is the presence of occlusion, although significant advances have been made by computing an occlusion map whilst computing the correspondence [2, 10, 5, 9]. Th... Adrian Broadhurst, Roberto Cipolla |
BMVC | 2 |
| 1999 | Collineation Estimation from Two Unmatched Views of an Unknown Planar Contour for Visual ServoingabstractIn this paper we describe a method to compute the collineation matrix between two unmatched images of an unknown planar contour described using a B-spline snake. The two images of the contour are matched and the collineation matrix is used to servo a camera mounted on the robot end-effector using a 2 1/2 D visual servoing technique. The experimental results, obtained using common planar objects, show that our method give very good results and allow the robot end-effector to be positioned with a great precision. 1 Introduction The visual servoing scheme of robot manipulators can be divided in three steps. In the first off-line learning step, the reference image of the object corresponding to a desired position of the robot is acquired and some image features are extracted. In general, objects are represented by free-form curves, i.e., arbitrary space curves of the type found in practice. A curve is usually described as a set of chained points. The reference image can be obtaine... Graziano Chesi, Ezio Malis, Roberto Cipolla |
BMVC | 3 |
| 1999 | Camera Calibration from Vanishing Points in Image of Architectural Scenes
Roberto Cipolla, Tom Drummond, Duncan P. Robertson |
BMVC | 1 |
| 1999 | Model Refinement from Planar ParallaxabstractThis paper presents a system for refining the accuracy and realism of coarse piecewise planar models from an uncalibrated sequence of images. First, dense depth maps are estimated by aligning a planar region of a scene in each image, approximating camera calibration, and generating dense planar paral-lax. These depth maps are then robustly fused to obtain incrementally refined surface estimates. It is envisaged that this system will extend the modelling capability of existing systems [3] which generate simple, piecewise planar architectural models. 1 Anthony R. Dick, Roberto Cipolla |
BMVC | 2 |
| 1999 | Real-time Tracking of Complex Structures with On-line Camera CalibrationabstractThis paper presents a novel three-dimensional model-based tracking system which has been incorporated into a visual servoing system. The tracking system combines modern graphical rendering technology with constrained active contour tracking techniques to create wireframe -snakes. It operates in real time at video frame rate (25 Hz) and is based on an internal CAD model of the object to be tracked. This model is rendered using a binary space partition tree to perform hidden line removal and the visible features are identified on-line at each frame and are tracked in the video feed. The tracking system has been extended to incorporate real-time on-line calibration and tracking of internal camera parameters. Results from on-line calibration and visual servoing experiments are presented. 1 Introduction The tracking of rigid three-dimensional objects is useful for numerous applications, including motion analysis, surveillance and robotic control tasks. This paper tackles two p... Tom Drummond, Roberto Cipolla |
BMVC | 2 |
| 1999 | Edge Tracking for Motion Segmentation and Depth OrderingabstractThis paper presents a new theoretical framework for motion segmentation based on the motion of tracked region edges. By considering the visible edges of an object, constraints may be placed on the motion labelling of edges. This enables the most likely region labelling and layer ordering to be established, thus producing a segmentation. An implementation is outlined and demonstrated on test sequences containing two motions. The image is divided into regions using a colour edgebased segmentation scheme and the normal motion of these edges is tracked. The EM algorithm is used to partition the edges and fit the best two motions accordingly. Hypothesising each motion in turn to be the foreground motion, the labelling constraints can be applied and the frame segmented. The hypothesis which best fits the observed edge motions indicates the layer ordering and leads to a very accurate segmentation. 1 Introduction The segmentation of a video sequence into moving objects is a first... Tom Drummond, Roberto Cipolla |
BMVC | 3 |
| 1999 | Reconstruction and Motion Estimation from Apparent Contours under Circular MotionabstractIn this paper we address the problem of recovering structure and motion from the contours of a smooth-curved surface. A novel and simpler technique for computing the structure of an object from its profiles is introduced. Experiments with real data show encouraging results, which are comparable to those obtained from much more sophisticated techniques. Furthermore, a new method for motion estimation from sequences of profiles is proposed. Preliminary results demonstrate the feasibility of the algorithm. 1 Introduction The recovering of structure and motion from sequences of images is a central problem in computer vision, and its solution has generated a rich pool of algorithms [8, 1]. Most of these algorithms rely on correspondences of points or lines between images, and work well when the scene being viewed is composed of polyhedral parts. However, for smooth surfaces without noticeable texture, point and line correspondences may not be easily established. In this case the profile of... Kwan-Yee Kenneth Wong, Paulo R. S. Mendonça, Roberto Cipolla |
BMVC | 3 |
| 1999 | Calibration of Image Sequences for Model VisualizationabstractThe object of this paper is to find a quick and accurate method for computing the projection matrices of an image sequence, so that the error is distributed evenly along the sequence. It assumes that a set of correspondences between points in the images is known, and that these points represent rigid points in the world. This paper extends the algebraic minimisation approach developed by Hartley so that it can be used for long image sequences. This is achieved by initially computing a trifocal tensor using the three most extreme views. The intermediate views are then computed linearly using the trifocal tensor. An iterative algorithm as presented which perturbs the twelve entries of one camera matrix so that the algebraic error along the whole sequence is minimised. Adrian Broadhurst, Roberto Cipolla |
CVPR | 2 |
| 1999 | Visual Tracking and Control using Lie AlgebrasabstractA novel approach to visual servoing is presented, which takes advantage of the structure of the Lie algebra of affine transformations. The aim of this project is to use feedback from a visual sensor to guide a robot arm to a target position. The sensor is placed in the end effector of the robot, the 'camera-in-hand' approach, and thus provides direct feedback of the robot motion relative to the target scene via observed transformations of the scene. These scene transformations are obtained by measuring the affine deformations of a target planar contour, captured by use of an active contour, or snake. Deformations of the snake are constrained using the Lie groups of affine and projective transformations. Properties of the Lie algebra of affine transformations are exploited to integrate observed deformations to the target contour which can be compensated with appropriate robot motion using a non-linear control structure. These techniques have been implemented using a video camera to control a 5 DoF robot arm. Experiments with this implementation are presented, together with a discussion of the results. Tom Drummond, Roberto Cipolla |
CVPR | 2 |
| 1999 | A Biprism-Stereo Camera SystemabstractIn this paper we propose a novel and practical stereo camera system that uses only one camera and a biprism placed in front of the camera. The equivalent of a stereo pair of images is formed as the left and right halves of a single CCD image using a biprism. The system is therefore cheap and extremely easy to calibrate since it requires only one CCD camera. An additional advantage of the geometrical set-up is that corresponding features lie on the same scanline automatically. The single camera and biprism have led to a simple stereo system for which correspondence is very easy and which is accurate for nearby objects in a small field of view. Since we we only, a single lens, calibration of the system is greatly simplified. This is due to the fact that we need to estimate only one focal length and one center of projection. Given the parameters in the biprism-stereo camera system, we can recover the depth of the object using only the disparity between the corresponding points. Doo Hyun Lee, In-So Kweon, Roberto Cipolla |
CVPR | 3 |
| 1999 | Estimation of Epipolar Geometry from Apparent Contours: Affine and Circular Motion CasesabstractThis paper addresses the problem of estimating the epipolar geometry from apparent contours in two special cases: under weak perspective and for circular motion. An appropriate parametrization of the fundamental matrix is introduced for both cases, as well as suitable cost functions for the estimation of the epipoles. The algorithm used in the affine approximation proven to be robust and accurate under several conditions. The circular motion case turned out to be much more difficult, but for a wide baseline the method introduced here is successful. For small viewing angles the technique is too sensitive to noise to be used in practice. Nevertheless, circular motion with small baseline can be well modeled by an affine camera system, and this approximation should be used in this circumstance. Paulo R. S. Mendonça, Roberto Cipolla |
CVPR | 2 |
| 1999 | A Simple Technique for Self-CalibrationabstractThis paper introduces an extension of Hartley's self-calibration technique based on properties of the essential matrix, allowing for the stable computation of varying focal lengths and principal point. It is well known that the three singular values of an essential must satisfy two conditions: one of them must be zero and the other two must be identical. An essential matrix is obtained from the fundamental matrix by a transformation involving the intrinsic parameters of the pair of cameras associated with the two views. Thus, constraints on the essential matrix can be translated into constraints on the intrinsic parameters of the pair of cameras. This allows for a search in the space of intrinsic parameters of the cameras in order to minimize a cost function related to the constraints. This approach is shown to be simpler than other methods, with comparable accuracy in the results. Another advantage of the technique is that it does not require as input a consistent set of weakly calibrated camera matrices (as defined by Harley) for the whole image sequence, i.e. a set of cameras consistent with the correspondences and known up to a projective transformation. Paulo R. S. Mendonça, Roberto Cipolla |
CVPR | 2 |
| 1999 | Extracting Group Transformations from Image Moments
Jun Sato, Roberto Cipolla |
Comput. Vis. Image Underst. | 2 |
| 1999 | Generalised Epipolar Constraints
Kalle Åström, Roberto Cipolla, Peter J. Giblin |
Int. J. Comput. Vis. | 2 |
| 1999 | Uncalibrated reconstruction of curved surfaces
Jun Sato, Roberto Cipolla |
Image Vis. Comput. | 2 |
| 1999 | Automated B-Spline Curve Representation Incorporating MDL and Error-Minimizing Control Point Insertion StrategiesabstractThe main issues of developing an automatic and reliable scheme for spline-fitting are discussed and addressed in this paper, which are not fully covered in previous papers or algorithms. The proposed method incorporates B-spline active contours, the minimum description length (MDL) principle, and a novel control point insertion strategy based on maximizing the potential for energy-reduction maximization (PERM). A comparison of test results shows that it outperforms one of the better existing methods. Tat-Jen Cham, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1999 | Affine Reconstruction of Curved Surfaces from Uncalibrated Views of Apparent ContoursabstractIn this paper, we consider uncalibrated reconstruction of curved surfaces from apparent contours. Since apparent contours are not fixed features (viewpoint independent), we cannot directly apply the recent results of the uncalibrated reconstruction from fixed features. We show that, nonetheless, curved surfaces can be reconstructed up to an affine ambiguity from their apparent contours viewed from uncalibrated cameras with unknown linear translations. Furthermore, we show that, even if the reconstruction is nonmetric (non-Euclidean), we can still extract useful information for many computer vision applications just from the apparent contours. We first show that if the camera motion is linear translation (but arbitrary direction and magnitude), the epipolar geometry can be recovered from the apparent contours without using any optimization process. The extracted epipolar geometry is next used for reconstructing curved surfaces from the deformations of the apparent contours viewed from uncalibrated cameras. The result is applied to distinguishing curved surfaces from fixed features in images. It is also shown that the time-to-contact to the curved surfaces can be computed from simple measurements of the apparent contours. Jun Sato, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1998 | Enhancing Human Face Detection Using Motion and Active Contours
Kin Choong Yow 0001, Roberto Cipolla |
ACCV (1) | 2 |
| 1998 | Applications of the "Creep-and-Merge" System: Corner DetectionabstractThe Creep–and–Merge (CAM) segmentation system (described in [3, 2]) is a novel architecture for region-based segmentation; it is designed to be efficient, insensitive to noise, scale and image geometry, capable of applying the widest range of statistical models, and to contain no adjustable parameters. This paper describes the continuing development of higher–level processing models utilising this system: we present a geometric (“true”) corner detector with good performance, and give qualitative and quantitative comparisons with other leading systems. 1 Antranig Basman, Joan Lasenby, Roberto Cipolla |
BMVC | 3 |
| 1998 | Analysis and Computation of an Affine Trifocal TensorabstractThis paper investigates the trifocal tensor for an affine trinocular rig and defines an affine trifocal tensor.The question of the degrees of freedom of the tensor entries will be addressed, and a novel algorithm to compute the tensor by a linear technique is presented.It will be shown that 4 point or 8 line correspondences along the three images are enough for a reliable computation of the trifocal tensor in the affine case (uncalibrated, weak perspective camera model), in contrast to the projective case, where at least 7 point or 13 line correspondences are needed for a linear computation (and usually with poor results).An analysis of the error in the approximation of a generic trifocal tensor by the affine trifocal tensor is carried out and preliminary experiments with synthetic and real data show the reliability and robustness of the approximation under a wide range of conditions. Paulo R. S. Mendonça, Roberto Cipolla |
BMVC | 2 |
| 1998 | Effective Corner MatchingabstractThis paper tackles the problem of obtaining a good initial set of corner matches between two images without resorting to any constraints from motion or structure models. Several different matching metrics, both traditional and statistical, are evaluated and the effect of matching using sub-pixel information is studied. It is found that, in most cases, the commonly-used cross-correlation does not perform as well as some other measures, such as the 2 test or the sum of squared differences, and that it is essential to use sub-pixel accuracy if mismatches are to be avoided. Further, a new technique, the Median Flow Filter, is introduced. This detects outliers by assuming that the image motion is locally similar. Any matches which are in gross disagreement with the local "median flow" are discarded. Experiments show this technique to be particularly effective, typically lowering the percentage of outliers from around 35% to less than 5%, permitting direct model fitting rathe... D. Sinclair, Roberto Cipolla, K. Wood |
BMVC | 3 |
| 1998 | A Statistical Framework for Long-Range Feature Matching in Uncalibrated Image MosaicingabstractThe problem considered is that of estimating the projective transformation between two images in situations where the image motion is large and feature matching is not aided by a proximity heuristic. The overall algorithm designed is based on a multiresolution, multihypothesis scheme, and similarities between tracking and matching through multiple resolution levels are exploited. Two major tools are developed in this paper: (i) a Bayesian framework for incorporating similarity measures of feature correspondences in regression to specify the different levels of confidence in the correspondences; and (ii) a Bayesian version of RANSAC, which is able to utilise prior estimates and matching probabilities. The algorithm is tested on a number of real images with large image motion and promising results were obtained. Tat-Jen Cham, Roberto Cipolla |
CVPR | 2 |
| 1998 | Affine Reconstruction of Curved Surfaces from Uncalibrated Views of Apparent ContoursabstractIn this paper, we show that even if the camera is uncalibrated, and its translational motion is unknown, curved surfaces can be reconstructed from their apparent contours up to a 3D affine ambiguity. Furthermore, we show that even if the reconstruction is nonmetric (non-Euclidean), we can still extract useful information for many computer vision applications just from the apparent contours. We first show that if the camera undergoes pure translation (unknown direction and magnitude), the epipolar geometry can be recovered from the apparent contours without using any search or optimisation process. The extracted epipolar geometry is next used for reconstructing curved surfaces from the deformations of the apparent contours viewed from uncalibrated cameras. The result is applied to distinguishing curved surfaces from fixed features in images. It is also shown that the time-to-contact to the curved surfaces can be computed from simple measurements of the apparent contours. The proposed method is implemented and tested on real images of curved surfaces. Jun Sato, Roberto Cipolla |
ICCV | 2 |
| 1998 | Quasi-Invariant Parameterisations and Matching of Curves in Images
Jun Sato, Roberto Cipolla |
Int. J. Comput. Vis. | 2 |
| 1997 | Uncalibrated Reconstruction of Curved Surfaces
Jun Sato, Roberto Cipolla |
BMVC | 2 |
| 1997 | Stereo Coupled Active ContoursabstractWe consider how tracking in stereo may be enhanced by coupling pairs of active contours in different views via affine epipolar geometry and various subsets of planar affine transformations, as well as by implementing temporal constraints imposed by curve rigidity. 3D curve tracking is achieved using a submanifold model, where it is shown how the coupling mechanisms can be decomposed to cater for fired and variable epipolar geometries. In the case of tracking planar curves, the canonical frame model is developed such that the various geometrical constraints needed in different situations may be efficiently selected. The results show that coupled active contours add consistency and robustness to tracking in stereo. Tat-Jen Cham, Roberto Cipolla |
CVPR | 2 |
| 1997 | Following Cusps
Roberto Cipolla, Gordon J. Fletcher, Peter J. Giblin |
Int. J. Comput. Vis. | 1 |
| 1997 | Affine integral invariants for extracting symmetry axes
Jun Sato, Roberto Cipolla |
Image Vis. Comput. | 2 |
| 1997 | Feature-based human face detection
Kin Choong Yow 0001, Roberto Cipolla |
Image Vis. Comput. | 2 |
| 1996 | Automated B-Spline Curve Representation with MDL-based Active ContoursabstractPresent spline-fitting methods used in computer vision do not fully address the main issues of developing an automatic and reliable algorithm, which are discussed in this paper. A paradigm for spline fitting is proposed, with features of the algorithm selected such that the main issues are resolved. This is achieved through the use of Bspline active contours, the minimum description length principle, and in conjunction with a control point insertion strategy based on the Potential for Energy-Reduction Maximisation (PERM). This strategy selects control points such that the formation of compatible collapse mechanisms for the splines is encouraged. An implementation of the algorithm is carried and tested on various images. A comparison with one of the better existing methods for spline fitting demonstrates that there is considerable potential for the algorithm to outperform current algorithms. 1 Introduction Representing curves by analytic functions instead of sets of data points has man... Tat-Jen Cham, Roberto Cipolla |
BMVC | 2 |
| 1996 | Affine Visual ServoingabstractSmall movements of a viewer relative to the surrounding scene induce deformations in the shape and detail of the projected image. This paper will consider the problem of using these deformations to provide visual feedback on the current position of the viewer relative to the scene. The implementation calculates the transformations that occur due to small movements around the current position. If a "target" transformation is specified, the equivalent motion can be interpolated. As a result, it is possible to position the viewer relying solely on visual feedback. All the calibrations required are performed within the algorithm, and the system is assumed to work using an uncalibrated camera. 1 Introduction Viewing a three dimensional world projected onto a two dimensional plane causes the image produced by a conventional camera to be both ambiguous and difficult to interpret automatically. The ability to navigate around obstacles with the information from a single viewpoint is... Geoffrey Cross, Roberto Cipolla |
BMVC | 2 |
| 1996 | Identifying Planar Regions in a Scene using Uncalibrated Stereo VisionabstractWe describe the use of well-known uncalibrated stereo algorithms for detecting planar regions in a scene from the transformation of feature locations between views. Simulations indicate that a typical set-up would have a resolution of the order of one centimetre. A fully operational system is not yet complete, but here we present a number of steps towards achieving this goal. Keywords: uncalibrated stereo vision, segmentation, planar. 1 Introduction We are developing a system which combines stereoscopic vision with a robotic manipulator to enable it to locate, reach and grasp unmodelled objects in an unstructured environment. Part of the system has been built. The algorithm for indicating the object of interest is described in [2] and the algorithm for visually guiding the robot arm to the object is described in [6]. Both these algorithms use uncalibrated stereo vision. The advantages of using uncalibrated stereo are that it is easier to set up, more robust to disturbances of t... Gabriel Hamid, Nicholas J. Hollinghurst, Roberto Cipolla |
BMVC | 3 |
| 1996 | Affine Integral Invariants for Extracting Symmetry AxesabstractIn this paper, we propose integral invariants based on group invariant parameterisation. The new invariants do not suffer from occlusion problems, do not require any correspondence of image features unlike existing algebraic invariants, and are less sensitive to noise than differential invariants. Our framework applies affine differential geometry to derive novel affine integral invariants. The new in-variants are exploited for extracting the symmetry axes of planar objects viewed under weak perspective. The proposed method is tested on natural leaves and is shown to extract symmetry axes reliably. Jun Sato, Roberto Cipolla |
BMVC | 2 |
| 1996 | Uncalibrated Visual ServoingabstractVisual servoing is a process to enable a robot to position a camera with respect to known landmarks using the visual data obtained by the camera itself to guide camera motion. A solution is described which requires very little a priori information freeing it from being specific to a particular configuration of robot and camera. The solution is based on closed loop control together with deliberate perturbations of the trajectory to provide calibration movements for refining that trajectory. Results from experiments in simulation and on a physical robot arm (camera-in-hand configuration) are presented. 1 Introduction Visual servoing is a process by which the appearance of landmarks is used to control the positioning of the camera with respect to the world. The camera is thus the sensor for a control scheme, in which the position of the camera itself is the object of control. The objective is for a robot to position the camera in a specific `target' pose (defined at initialisati... Michael W. Spratling, Roberto Cipolla |
BMVC | 2 |
| 1996 | Scale and Orientation Invariance in Human Face DetectionabstractHuman face detection has always been an important problem for face, expression and gesture recognition. Though numerous attempts have been made to detect and localize faces, these approaches have made assumptions that restrict their extension to more general cases. In this research, we propose a feature-based face detection algorithm that can be easily extended to detect faces under different scale and orientation. Feature points are detected from the image using spatial filters and grouped into face candidates using geometric and gray level constraints. A probabilistic framework is then used to evaluate the likelihood of the candidate as a face. We provide results to support the validity of the approach, and show that the algorithm can indeed cope efficiently with faces at different scale and orientation. 1 Introduction Human face recognition is a field that has important applications in our daily activities, such as for security verification, criminal identification, and... Kin Choong Yow 0001, Roberto Cipolla |
BMVC | 2 |
| 1996 | Generalised Epipolar Constraints
Kalle Åström, Roberto Cipolla, Peter J. Giblin |
ECCV (2) | 2 |
| 1996 | Geometric Saliency of Curve Correspondances and Grouping of Symmetric Comntours
Tat-Jen Cham, Roberto Cipolla |
ECCV (1) | 2 |
| 1996 | Reliable Extraction of the Camera Motion using Constraints on the Epipole
Jonathan M. Lawn, Roberto Cipolla |
ECCV (2) | 2 |
| 1996 | A probabilistic framework for perceptual grouping of features for human face detectionabstractPresent approaches to human face detection have made several assumptions that restrict their ability to be extended to general imaging conditions. We identify that the key factor in a generic and robust system is that of exploiting a large amount of evidence, related and reinforced by model knowledge through a probabilistic framework. In this paper, we propose a face detection framework that groups image features into meaningful entities-using perceptual organization, assigns probabilities to each of them, and reinforce there probabilities using Bayesian reasoning techniques. True hypotheses of faces will be reinforced to a high probability. The detection of faces under scale, orientation and viewpoint variations will be examined in a subsequent paper. Kin Choong Yow 0001, Roberto Cipolla |
FG | 2 |
| 1996 | Detection of human faces under scale, orientation and viewpoint variationsabstractMany current human face detection algorithms make implicit assumptions about the scale, orientation or viewpoint of faces in an image and exploit these constraints to detect and localize faces. The algorithm may be robust for the assumed conditions but it becomes very difficult to extend the results to general imaging conditions. In an earlier paper (Yow and Cipolla, 1996) we proposed a feature-based face detection algorithm to detect faces in a complex background. In this paper we examine its ability to detect faces under different scale, orientation and viewpoint. The results show that the algorithm can indeed cope with a good range of scale, orientation and viewpoint variations that is typical of a subject sitting in front of a computer terminal. Kin Choong Yow 0001, Roberto Cipolla |
FG | 2 |
| 1996 | Affine integral invariants and matching of curvesabstractWe propose integral invariants based on a group invariant parameterisation. These new invariants do not suffer from the occlusion problem, do not require any correspondence of image features unlike algebraic invariants, and are less sensitive to noise than differential invariants. Affine differential geometry is applied to this framework, and novel affine integral invariants are derived. A quasi-invariant parameterisation enables us to reduce the order of derivatives required. The proposed invariants are applied for extracting corresponding contour curves of natural images. The noise sensitivity of the proposed invariants is compared with that of differential invariants. Jun Sato, Roberto Cipolla |
ICPR | 2 |
| 1996 | Human-robot interface by pointing with uncalibrated stereo vision
Roberto Cipolla, Nicholas J. Hollinghurst |
Image Vis. Comput. | 1 |
| 1996 | Fast visual tracking by temporal consensus
Andrew H. Gee, Roberto Cipolla |
Image Vis. Comput. | 2 |
| 1995 | Towards an Automatic Human Face Localizations SystemabstractThis paper describes a method to detect and locate human faces in an image given no prior information about the size, orientation, and viewpoint of the faces in the image. This method uses a family of Gaussian derivative filters to search and extract human facial features from the image and then group them together into a set of partial faces using their geometric relationship. A belief network is then constructed for each possible face candidate and the belief values updated by evidences propagating through the network. Different instances of detected faces are then compared using their belief values and improbable face candidates discarded. The algorithm is tested on different instances of faces with varying sizes, orientation and viewpoint and the results indicate a 91% success rate in detection under viewpoint variation. Kin Choong Yow 0001, Roberto Cipolla |
BMVC | 2 |
| 1995 | Motion from the Frontier of Curved SurfacesabstractThe frontier of a curved surface is the envelope of contour generators showing the boundary, at least locally, of the visible region swept out under viewer motion. In general, the outlines of curved surfaces (apparent contours) from different viewpoints are generated by different contour generators on the surface and hence do not provide a constraint on viewer motion. We show that frontier points, however, have projections which correspond to a real point on the surface and can be used to constrain viewer motion by the epipolar constraint. We show how to recover viewer motion from frontier points for both continuous and discrete motion, calibrated and uncalibrated cameras. We present preliminary results of an iterative scheme to recover the epipolar line structure from real image sequences using only the outlines of curved surfaces. A statistical evaluation as also performed to estimate the stability of the solution.> Roberto Cipolla, Kalle Åström, Peter J. Giblin |
ICCV | 1 |
| 1995 | Surface Geometry from Cusps of Apparent ContoursabstractIt is known that the deformations of the apparent contours of a surface under perspective projection and viewer motion enable the recovery of the geometry of the surface, for example by utilising the epipolar parametrization. These methods break down with apparent contours that are singular i.e. with cusps. In this paper, we study this situation in detail and show how, nevertheless, the surface geometry (including the Gauss curvature and mean curvature of the surface) can be recovered by following the cusps. Indeed the formulae are much simpler in this case, and require lower spatio-temporal derivatives than in the general case of nonsingular apparent contours. We give a simulated example, and also show that following cusps does not by itself provide us with information on ego-motion.> Roberto Cipolla, Gordon J. Fletcher, Peter J. Giblin |
ICCV | 1 |
| 1995 | Symmetry detection through local skewed symmetries
Tat-Jen Cham, Roberto Cipolla |
Image Vis. Comput. | 2 |
| 1995 | Image registration using multi-scale texture moments
Jun Sato, Roberto Cipolla |
Image Vis. Comput. | 2 |
| 1995 | Skeletonization using an extended Euclidean distance transform
Mark W. Wright, Roberto Cipolla, Peter J. Giblin |
Image Vis. Comput. | 2 |
| 1994 | Skewed Symmetry Detection Through Local Skewed SymmetriesabstractWe explore how global symmetry can be detected prior to segmentation and under noise and occlusion. The definition of local symmetries is extended to affine geometries by considering the tangents and curvatures of local structures, and a quantitative measure of local symmetry known as symmetricity is introduced, which is based on Mahalanobis distances from the tangent-curvature states of local structures to the local skewed symmetry state-subspace. These symmetricity values, together with the associated local axes of symmetry, are spatially related in the local skewed symmetry field (LSSF). In the implementation, a fast, local symmetry detection algorithm allows initial hypotheses for the symmetry axis to be generated through the use of a modified Hough transform. This is then improved upon by maximising a global symmetry measure based on accumulated local support in the LSSF — a straight active contour model is used for this purpose. This produces useful estimates for the axis of symmetry and the angle of skew in the presence of contour fragmentation, artifacts and occlusion. Tat-Jen Cham, Roberto Cipolla |
BMVC | 2 |
| 1994 | Camera Motion Determination from Dymanic Perceptual Grouping of Line SegmentsabstractWe present here a discussion on the use of perceptual grouping to improve structure and egomotion recovery algorithms for monocular cameras. In particular we look at grouping lines to avoid the need for trinocular algorithms. We also present a number of methods for grouping lines, including a novel method that infers planar groups from those demonstrating a linear deformation. Jonathan M. Lawn, Roberto Cipolla |
BMVC | 2 |
| 1994 | Image Registration Using Multi-Scale Texture Moments
Jun Sato, Roberto Cipolla |
BMVC | 2 |
| 1994 | Skeletonisation using an Extended Euclidean Distance TransformationabstractA standard method to perform skeletonisation is to use a distance transform. Unfortunately such an approach has the drawback that only the Symmetric axis transform can be computed and not the more practical smoothed local symmetries or the more general symmetry set. Using singularity theory we introduce an extended distance transform which may be used to capture more of the symmetries of a shape. We describe the relationship of this extended distance transform to the skeletal shape descriptors themselves and other geometric phenomema related to the boundary of the curve. We then show how the extended distance transform can be used to derive skeletal descriptions of an object. Mark W. Wright, Roberto Cipolla, Peter J. Giblin |
BMVC | 2 |
| 1994 | Robust Egomotion Estimation from Affine Motion Parallax
Jonathan M. Lawn, Roberto Cipolla |
ECCV (1) | 2 |
| 1994 | Extracting the Affine Transformation from Texture Moments
Jun Sato, Roberto Cipolla |
ECCV (2) | 2 |
| 1994 | Active 3D Object Recognition using 3D Affine Invariants
Sven Vinther, Roberto Cipolla |
ECCV (2) | 2 |
| 1994 | A local approach to recovering global skewed symmetryabstractA local approach is adopted to allow recovery of global skewed symmetry in the presence of occlusion. Local skewed symmetries are established by extending the definition of local symmetries to affine geometries through the use of local derivatives. Symmetricity, a quantitative gauge of local symmetry based on Mahalanobis distances from the tangent-curvature states of local structures to the local skewed symmetry state-subspace, is also introduced to cope with noise. The symmetricity values and local symmetry axes for each pair of points are then spatially related in the local skewed symmetry field. The global symmetry detection algorithm implemented involves obtaining fast, initial estimates of the symmetry axis from a separate Hough transform technique, followed by maximising a global symmetry measure via a straight active contour model which is driven by effective symmetricity values. This produces useful estimates for the axis of symmetry and the angle of skew in the presence of contour fragmentation, artifacts and occlusion. Tat-Jen Cham, Roberto Cipolla |
ICPR (1) | 2 |
| 1994 | Estimating gaze from a single view of a faceabstractCurrent approaches to gaze tracking tend to be highly intrusive: the subject must either remain perfectly still, or wear cumbersome headgear to maintain a constant separation between the sensor and the eye. This paper describes a more flexible vision-based approach, which can estimate the direction of gaze from a single, monocular view of a face. The technique makes minimal assumptions about the structure of the face, requires very few image measurements, and produces a useful estimate of the facial orientation. The computational requirements are insignificant, so with automatic tracking of a few facial features it is possible to produce real-time gaze estimates. Andrew H. Gee, Roberto Cipolla |
ICPR (1) | 2 |
| 1994 | Determining the gaze of faces in images
Andrew H. Gee, Roberto Cipolla |
Image Vis. Comput. | 2 |
| 1994 | Uncalibrated stereo hand-eye coordination
Nicholas J. Hollinghurst, Roberto Cipolla |
Image Vis. Comput. | 2 |
| 1993 | Uncalibrated Stereo Hand-Eye CoordinationabstractThis paper describes a system that combines stereo vision with a 5-DOF robotic manipulator, to enable it to locate and reach for objects by sight. Our system uses an affine stereo algorithm, a simple but robust approximation to the geometry of stereo vision, to estimate positions and surface orientations. It can be calibrated very easily with just four reference points. These are defined by the robot itself, moving the gripper to four known positions (selfcalibration). The inevitable small errors are corrected by a feedback mechanism which implements image-based control of the gripper's position and orientation. Integral to this feedback mechanism is the use of affine active contour models which track the real-time motion of the gripper across the two images. Experiments show the system to be remarkably immune to unexpected translations and rotations of the cameras and changes of focal length — even after it has 'calibrated ' itself. A future goal is to use affine stereo to implement shape-based grasp planning, enabling the robot to pick up a wide range of unidentified objects left in its workspace. At present it can only pick up simple wooden blocks and tracks a single planar contour on its target object. 1 Nicholas J. Hollinghurst, Roberto Cipolla |
BMVC | 2 |
| 1993 | Epipole Estimation Using Affine Motion ParallaxabstractDetermining the motion of a camera from its image sequences has so far proved very difficult, and no practical algorithms have been found for freely moving cameras. This novel algorithm is based on motion parallax, but uses sparse visual motion estimates to extract the direction of translation of the camera directly, after which determination of the camera rotation and the depths of the image features follows easily. This method can also detect and reject independent motion, and provide a measure of the uncertainty of its estimates. 1 Jonathan M. Lawn, Roberto Cipolla |
BMVC | 2 |
| 1993 | Towards 3D Object Model Acquisition and Recognition using 3D Affine InvariantsabstractWe evaluate the power of 3D affine invariants in an object recognition scheme. These invariants are actively calculated by the real-time tracking of 2D image features (corners) over an image sequence. This is done optimally by using a Kalman filter. Object information is located in a hash table where it is stored and retrieved using the invariants as stable indices. Recognition takes place when significant evidence for a particular shape has been found from the table. Preliminary results with real data are presented, and some of the noise problems arising due to the weak perspective approximation and corner localisation errors are discussed. 1 Sven Vinther, Roberto Cipolla |
BMVC | 2 |
| 1993 | Robust structure from motion using motion parallaxabstractAn efficient and geometrically intuitive algorithm for reliably interpreting the image velocities of moving objects in 3-D is presented. It is well known that under a weak perspective the image motion of points on a plane can be characterized by an affine transformation. It is shown that the relative image motion of a nearby non-coplanar point and its projection on the plane is equivalent to motion parallax, and because it is independent of view rotations it is a reliable geometric cue to 3-D shape and viewer/object motion. The authors summarize why structure from motion algorithms are often very sensitive to errors in the measured image velocities and then show how to efficiently and reliably extract an incomplete qualitative solution. They also show how to augment this into a complete solution if additional constraints or views are available. A real-time example is presented in which the 3-D visual interpretation of hand gestures or a hand-held object is used as part of a man-machine interface. This is an alternative to the Polhemus coil instrumented Dataglove commonly used in sensing manual gestures.> Roberto Cipolla, Yasukazu Okamoto, Yoshinori Kuno |
ICCV | 1 |
| 1992 | Surface Orientation and Time to Contact from Image Divergence and Deformation
Roberto Cipolla, Andrew Blake 0001 |
ECCV | 1 |
| 1992 | Surface shape from the deformation of apparent contours
Roberto Cipolla, Andrew Blake 0001 |
Int. J. Comput. Vis. | 1 |
| 1991 | Parallel Implementation of Lagrangian Dynamics for Real-time Snakes
Rupert W. Curwen, Andrew Blake 0001, Roberto Cipolla |
BMVC | 3 |
| 1991 | Robust estimation of surface curvature from deformation of apparent contours
Andrew Blake 0001, Roberto Cipolla |
Image Vis. Comput. | 2 |
| 1990 | Towards qualitative vision: motion parallaxabstractA robot vehicle moving under visual guidance needs to compute approximate geometry of obstacles in its environment. It is unreasonable to assume that egomotion is known to the sort of precision that is available for a camera mounted on a high quality robot arm. Generally a nominal estimate for egomotion is available. One possibility is to refine this estimate using optic flow data (Harris, 1987). Alternatively, the problem can be turned on its head: what geometric information remains stable under perturbation of assumed egomotion? This question has been addressed by Koenderink and van Doom (1977), Nelson and Aloimonos (1988) and Verri et al. (1989), in the case of continuous motion fields and, in the domain of stereoscopic vision, by Weinshall (1990). Part of the answer, we claim, lies in the use of motion parallax as a geometric cue. Motion parallax, which is a relative measure of the positions of two points, can be very much more robust as a cue than the absolute position of a single point. This is true for computation of relative depth, curvature on specular surfaces and curvature on extremal boundaries. Andrew Blake 0001, Roberto Cipolla, Andrew Zisserman |
BMVC | 2 |
| 1990 | Robust Estimation of Surface Curvature from Deformation of Apparent Contours
Andrew Blake 0001, Roberto Cipolla |
ECCV | 2 |
| 1990 | The dynamic analysis of apparent contoursabstractThe authors develop previous theories of the analysis of deformation of apparent contours under viewer motion. Earlier results showing how surface curvature can be inferred from acceleration of image features are generalized for arbitrary viewer motion and perspective projection. It is shown that relative image acceleration, based on parallax measurements, is robust to uncertainties in robot motion. The theory has been implemented and extensively tested in a real-time (15 frames per second) tracking system based on deformable contours (snakes). It is shown that focusing attention by means of snakes allows rapid, robust computation of surface curvature, including discrimination of extremal and occluding contours.> Roberto Cipolla, Andrew Blake 0001 |
ICCV | 1 |
| 1990 | Stereoscopic tracking of bodies in motion
Roberto Cipolla, Masanobu Yamamoto |
Image Vis. Comput. | 1 |