VLDB 2026 Research / reviewers in the wild / expert
Andrew W. Fitzgibbon
dblp:f/AndrewWFitzgibbon · also Andrew William Fitzgibbon
· DBLP profile ↗
126ranked-venue papers
23as first author
6since 2021 · last 2025
0000-0002-9839-660XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 103 · 21 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 96 · 17 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 7Software engineering, systems software and programming languages · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On Stochastic Rounding with Few Random BitsabstractLarge-scale numerical computations make increasing use of low-precision (LP) floating point formats and mixed precision arithmetic, which can be enhanced by the technique of stochastic rounding (SR), that is, rounding an intermediate high- precision value up or down randomly as a function of the value's distance to the two rounding candidates. Stochastic rounding requires, in addition to the high-precision input value, a source of random bits. As the provision of high-quality random bits is an additional computational cost, it is of interest to require as few bits as possible while maintaining the desirable properties of SR in a given computation, or computational domain. This paper examines a number of possible implementations of few- bit stochastic rounding (FBSR), and shows how several natural implementations can introduce sometimes significant bias into the rounding process, which are not present in the case of infinite- bit, infinite-precision examinations of these implementations. The paper explores the impact of these biases in machine learning examples, and hence opens another class of configuration parameters of which practitioners should be aware when developing or adopting low-precision floating point. Andrew W. Fitzgibbon, Stephen Felix |
ARITH | 1 |
| 2024 | Towards Foundational Models for Molecular Learning on Large-Scale Multi-Task DatasetsabstractRecently, pre-trained foundation models have enabled significant advancements in multiple fields. In molecular machine learning, however, where datasets are often hand-curated, and hence typically small, the lack of datasets with labeled features, and codebases to manage those datasets, has hindered the development of foundation models. In this work, we present seven novel datasets categorized by size into three distinct categories: ToyMix, LargeMix and UltraLarge. These datasets push the boundaries in both the scale and the diversity of supervised labels for molecular learning. They cover nearly 100 million molecules and over 3000 sparsely defined tasks, totaling more than 13 billion individual labels of both quantum and biological nature. In comparison, our datasets contain 300 times more data points than the widely used OGB-LSC PCQM4Mv2 dataset, and 13 times more than the quantum-only QM1B dataset. In addition, to support the development of foundational models based on our proposed datasets, we present the Graphium graph machine learning library which simplifies the process of building and training molecular machine learning models for multi-task and multi-level molecular datasets. Finally, we present a range of baseline results as a starting point of multi-task and multi-level training on these datasets. Empirically, we observe that performance on low-resource biological datasets show improvement by also training on large amounts of quantum data. This indicates that there may be potential in multi-task and multi-level training of a foundation model and fine-tuning it to resource-constrained downstream tasks. The Graphium library is publicly available on Github and the dataset links are available in Part 1 and Part 2. Dominique Beaini, Shenyang Huang, Joao Alex Cunha, Gabriela Moisescu-Pareja, Oleksandr Dymov, Samuel Maddrell-Mander, Callum McLean, Frederik Wenkel, Luis Müller, Jama Hussein Mohamud, Ali Parviz, Michael Craig, Michal Koziarski, Jiarui Lu, Zhaocheng Zhu, Cristian Gabellini, Kerstin Kläser 0001, Josef Dean, Cas Wognum, Maciej Sypetkowski, Guillaume Rabusseau, Reihaneh Rabbany, Jian Tang 0005, Christopher Morris 0001, Mirco Ravanelli, Guy Wolf, Prudencio Tossou, Hadrien Mary, Therence Bois, Andrew W. Fitzgibbon, Blazej Banaszewski, Chad Martin, Dominic Masters |
ICLR | 31 |
| 2023 | Generating QM1B with PySCFIPU
Alexander Mathiasen, Hatem Helal, Kerstin Kläser 0001, Paul Balanca, Josef Dean, Carlo Luschi, Dominique Beaini, Andrew W. Fitzgibbon, Dominic Masters |
NeurIPS | 8 |
| 2022 | FLAG: Flow-based 3D Avatar Generation from Sparse ObservationsabstractTo represent people in mixed reality applications for collaboration and communication, we need to generate realistic and faithful avatar poses. However, the signal streams that can be applied for this task from head-mounted devices (HMDs) are typically limited to head pose and hand pose estimates. While these signals are valuable, they are an incomplete representation of the human body, making it challenging to generate a faithful full-body avatar. We address this challenge by developing a flow-based generative model of the 3D human body from sparse observations, wherein we learn not only a conditional distribution of 3D human pose, but also a probabilistic mapping from observations to the latent space from which we can generate a plausible pose along with uncertainty estimates for the joints. We show that our approach is not only a strong predictive model, but can also act as an efficient pose prior in different optimization settings where a good initial latent code plays a major role. Mohammad Sadegh Ali Akbarian, Pashmina Cameron, Federica Bogo, Andrew W. Fitzgibbon, Thomas J. Cashman 0001 |
CVPR | 4 |
| 2022 | Provably correct, asymptotically efficient, higher-order reverse-mode automatic differentiationabstractIn this paper, we give a simple and efficient implementation of reverse-mode automatic differentiation, which both extends easily to higher-order functions, and has run time and memory consumption linear in the run time of the original program. In addition to a formal description of the translation, we also describe an implementation of this algorithm, and prove its correctness by means of a logical relations argument. Faustyna Krawiec, Simon L. Peyton Jones, Neelakantan R. Krishnaswami, Tom Ellis, Richard A. Eisenberg, Andrew W. Fitzgibbon |
Proc. ACM Program. Lang. | 6 |
| 2021 | Hashing modulo alpha-equivalenceabstractIn many applications one wants to identify identical subtrees of a program syntax tree. This identification should ideally be robust to alpha-renaming of the program, but no existing technique has been shown to achieve this with good efficiency (better than O(n2) in expression size). We present a new, asymptotically efficient way to hash modulo alpha-equivalence. A key insight of our method is to use a weak (commutative) hash combiner at exactly one point in the construction, which admits an algorithm with O(n (logn)2) time complexity. We prove that the use of the commutative combiner nevertheless yields a strong hash with low collision probability. Numerical benchmarks attest to the asymptotic behaviour of the method. Krzysztof Maziarz, Tom Ellis, Alan Lawrence, Andrew W. Fitzgibbon, Simon L. Peyton Jones |
PLDI | 4 |
| 2020 | Who Left the Dogs Out? 3D Animal Reconstruction with Expectation Maximization in the Loop
Benjamin Biggs, Oliver Boyne, James Charles, Andrew W. Fitzgibbon, Roberto Cipolla |
ECCV (11) | 4 |
| 2020 | The Phong Surface: Efficient 3D Model Fitting Using Lifted Optimization
Jingjing Shen, Thomas J. Cashman 0001, Qi Ye 0001, Tim Hutton, Toby Sharp, Federica Bogo, Andrew W. Fitzgibbon, Jamie Shotton |
ECCV (1) | 7 |
| 2019 | Efficient differentiable programming in a functional array-processing languageabstractWe present a system for the automatic differentiation (AD) of a higher-order functional array-processing language. The core functional language underlying this system simultaneously supports both source-to-source forward-mode AD and global optimisations such as loop transformations. In combination, gradient computation with forward-mode AD can be as efficient as reverse mode, and that the Jacobian matrices required for numerical algorithms such as Gauss-Newton and Levenberg-Marquardt can be efficiently computed. Amir Shaikhha, Andrew W. Fitzgibbon, Dimitrios Vytiniotis, Simon L. Peyton Jones |
Proc. ACM Program. Lang. | 2 |
| 2018 | Creatures Great and SMAL: Recovering the Shape and Motion of Animals from Video
Benjamin Biggs, Thomas Roddick, Andrew W. Fitzgibbon, Roberto Cipolla |
ACCV (5) | 3 |
| 2018 | QRkit: Sparse, Composable QR Decompositions for Efficient and Stable Solutions to Problems in Computer VisionabstractEmbedded computer vision applications increasingly require the speed and power benefits of single-precision (32 bit) floating point. However, applications which make use of Levenberg-like optimization can lose significant accuracy when reducing to single precision, sometimes unrecoverably so. This accuracy can be regained using solvers based on QR rather than Cholesky decomposition, but the absence of sparse QR solvers for common sparsity patterns found in computer vision means that many applications cannot benefit. We introduce an open-source suite of solvers for Eigen, which efficiently compute the QR decomposition for matrices with some common sparsity patterns (block diagonal, horizontal and vertical concatenation, and banded). For problems with very particular sparsity structures, these elements can be composed together in 'kit' form, hence the name QRkit. We apply our methods to several computer vision problems, showing competitive performance and suitability especially in single precision arithmetic. Jan Svoboda, Thomas J. Cashman 0001, Andrew W. Fitzgibbon |
WACV | 3 |
| 2017 | On the Two-View Geometry of Unsynchronized CamerasabstractWe present new methods of simultaneously estimating camera geometry and time shift from video sequences from multiple unsynchronized cameras. Algorithms for simultaneous computation of a fundamental matrix or a homography with unknown time shift between images are developed. Our methods use minimal correspondence sets (eight for fundamental matrix and four and a half for homography) and therefore are suitable for robust estimation using RANSAC. Furthermore, we present an iterative algorithm that extends the applicability on sequences which are significantly unsynchronized, finding the correct time shift up to several seconds. We evaluated the methods on synthetic and wide range of real world datasets and the results show a broad applicability to the problem of camera synchronization. Cenek Albl, Zuzana Kukelova, Andrew W. Fitzgibbon, Jan Heller, Matej Smíd, Tomás Pajdla |
CVPR | 3 |
| 2017 | Revisiting the Variable Projection Method for Separable Nonlinear Least Squares ProblemsabstractVariable Projection (VarPro) is a framework to solve optimization problems efficiently by optimally eliminating a subset of the unknowns. It is in particular adapted for Separable Nonlinear Least Squares (SNLS) problems, a class of optimization problems including low-rank matrix factorization with missing data and affine bundle adjustment as instances. VarPro-based methods have received much attention over the last decade due to the experimentally observed large convergence basin for certain problem classes, where they have a clear advantage over standard methods based on Joint optimization over all unknowns. Yet no clear answers have been found in the literature as to why VarPro outperforms others and why Joint optimization, which has been successful in solving many computer vision tasks, fails on this type of problems. Also, the fact that VarPro has been mainly tested on small to medium-sized datasets has raised questions about its scalability. This paper intends to address these unsolved puzzles. Je Hyeong Hong, Christopher Zach, Andrew W. Fitzgibbon |
CVPR | 3 |
| 2017 | An Efficient Background Term for 3D Reconstruction and Tracking with Smooth Surface ModelsabstractWe present a novel strategy to shrink and constrain a 3D model, represented as a smooth spline-like surface, within the visual hull of an object observed from one or multiple views. This new background or silhouette term combines the efficiency of previous approaches based on an image-plane distance transform with the accuracy of formulations based on raycasting or ray potentials. The overall formulation is solved by alternating an inner nonlinear minization (raycasting) with a joint optimization of the surface geometry, the camera poses and the data correspondences. Experiments on 3D reconstruction and object tracking show that the new formulation corrects several deficiencies of existing approaches, for instance when modelling non-convex shapes. Moreover, our proposal is more robust against defects in the object segmentation and inherently handles the presence of uncertainty in the measurements (e.g. null depth values in images provided by RGB-D cameras). Mariano Jaimez, Thomas J. Cashman 0001, Andrew W. Fitzgibbon, Javier González 0001, Daniel Cremers |
CVPR | 3 |
| 2017 | Online generative model personalization for hand trackingabstractWe present a new algorithm for real-time hand tracking on commodity depth-sensing devices. Our method does not require a user-specific calibration session, but rather learns the geometry as the user performs live in front of the camera, thus enabling seamless virtual interaction at the consumer level. The key novelty in our approach is an online optimization algorithm that jointly estimates pose and shape in each frame, and determines the uncertainty in such estimates. This knowledge allows the algorithm to integrate per-frame estimates over time, and build a personalized geometric model of the captured user. Our approach can easily be integrated in state-of-the-art continuous generative motion tracking software. We provide a detailed evaluation that shows how our approach achieves accurate motion tracking for real-time applications, while significantly simplifying the workflow of accurate hand performance capture. We also provide quantitative evaluation datasets at http://gfx.uvic.ca/datasets/handy Anastasia Tkach, Andrea Tagliasacchi, Edoardo Remelli, Mark Pauly, Andrew W. Fitzgibbon |
ACM Trans. Graph. | 5 |
| 2016 | Better Together: Joint Reasoning for Non-rigid 3D Reconstruction with Specularities and Shading
Chris Russell 0001, Lourdes Agapito, Andrew W. Fitzgibbon, Liu-Yin Qi |
BMVC | 4 |
| 2016 | Efficient Intersection of Three Quadrics and Applications in Computer VisionabstractIn this paper, we present a new algorithm for finding all intersections of three quadrics. The proposed method is algebraic in nature and it is considerably more efficient than the Gröbner basis and resultant-based solutions previously used in computer vision applications. We identify several computer vision problems that are formulated and solved as systems of three quadratic equations and for which our algorithm readily delivers considerably faster results. Also, we propose new formulations of three important vision problems: absolute camera pose with unknown focal length, generalized pose-and-scale, and hand-eye calibration with known translation. These new formulations allow our algorithm to significantly outperform the state-of-the-art in speed. Zuzana Kukelova, Jan Heller, Andrew W. Fitzgibbon |
CVPR | 3 |
| 2016 | Fits Like a Glove: Rapid and Reliable Hand Shape PersonalizationabstractWe present a fast, practical method for personalizing a hand shape basis to an individual user's detailed hand shape using only a small set of depth images. To achieve this, we minimize an energy based on a sum of render-and-compare cost functions called the golden energy. However, this energy is only piecewise continuous, due to pixels crossing occlusion boundaries, and is therefore not obviously amenable to efficient gradient-based optimization. A key insight is that the energy is the combination of a smooth low-frequency function with a high-frequency, low-amplitude, piecewisecontinuous function. A central finite difference approximation with a suitable step size can therefore jump over the discontinuities to obtain a good approximation to the energy's low-frequency behavior, allowing efficient gradient-based optimization. Experimental results quantitatively demonstrate for the first time that detailed personalized models improve the accuracy of hand tracking and achieve competitive results in both tracking and model registration. David Joseph Tan, Thomas J. Cashman 0001, Jonathan Taylor 0001, Andrew W. Fitzgibbon, Daniel Tarlow, Sameh Khamis, Shahram Izadi, Jamie Shotton |
CVPR | 4 |
| 2016 | Projective Bundle Adjustment from Arbitrary Initialization Using the Variable Projection Method
Je Hyeong Hong, Christopher Zach, Andrew W. Fitzgibbon, Roberto Cipolla |
ECCV (1) | 3 |
| 2016 | ShadowHands: High-Fidelity Remote Hand Gesture Visualization using a Hand TrackerabstractThis paper presents ShadowHands - a novel technique for visualizing a remote user's hand gestures using a single depth sensor and hand tracking system. Previous work has shown that making distributed users better aware of each other's gestures facilitates remote collaboration. These systems presented virtual embodiments as a stream of raw 2D or 3D data -- this data is noisy, and requires high bandwidth and favorable camera positions. Instead, our work uses a hand tracker to capture gestures which we visualize with a high-fidelity hand model. Our system is practical, requiring only a single depth sensor placed below the screen, and can be used without per-user calibration. As we use a 3D model rather than raw data, we can augment the hand's appearance to improve saliency and aesthetics. We alpha-blend this visualization over a shared workspace, so the local user perceives the remote user's hand as if they were separated by a transparent display. We conducted an experiment to compare traditional hand embodiments with our new technique, showing a quantitative improvement in selection accuracy and qualitative improvements in feelings of mutual understanding and engagement. Erroll Wood, Jonathan Taylor 0001, John Fogarty, Andrew W. Fitzgibbon, Jamie Shotton |
ISS | 4 |
| 2016 | Efficient and precise interactive hand tracking through joint, continuous optimization of pose and correspondencesabstractFully articulated hand tracking promises to enable fundamentally new interactions with virtual and augmented worlds, but the limited accuracy and efficiency of current systems has prevented widespread adoption. Today's dominant paradigm uses machine learning for initialization and recovery followed by iterative model-fitting optimization to achieve a detailed pose fit. We follow this paradigm, but make several changes to the model-fitting, namely using: (1) a more discriminative objective function; (2) a smooth-surface model that provides gradients for non-linear optimization; and (3) joint optimization over both the model pose and the correspondences between observed data points and the model surface. While each of these changes may actually increase the cost per fitting iteration, we find a compensating decrease in the number of iterations. Further, the wide basin of convergence means that fewer starting points are needed for successful model fitting. Our system runs in real-time on CPU only, which frees up the commonly over-burdened GPU for experience designers. The hand tracker is efficient enough to run on low-power devices such as tablets. We can track up to several meters from the camera to provide a large working volume for interaction, even using the noisy data from current-generation depth cameras. Quantitative assessments on standard datasets show that the new approach exceeds the state of the art in accuracy. Qualitative results take the form of live recordings of a range of interactive experiences enabled by this new approach. Jonathan Taylor 0001, Lucas Bordeaux, Thomas J. Cashman 0001, Bob Corish, Cem Keskin, Toby Sharp, Eduardo Soto, David Sweeney, Julien P. C. Valentin, Benjamin Luff, Arran Topalian, Erroll Wood, Sameh Khamis, Pushmeet Kohli, Shahram Izadi, Richard Banks, Andrew W. Fitzgibbon, Jamie Shotton |
ACM Trans. Graph. | 17 |
| 2015 | Fitting models to data: Accuracy, Speed, Robustness
Andrew W. Fitzgibbon |
BMVC | 1 |
| 2015 | Accurate, Robust, and Flexible Real-time Hand TrackingabstractWe present a new real-time hand tracking system based on a single depth camera. The system can accurately reconstruct complex hand poses across a variety of subjects. It also allows for robust tracking, rapidly recovering from any temporary failures. Most uniquely, our tracker is highly flexible, dramatically improving upon previous approaches which have focused on front-facing close-range scenarios. This flexibility opens up new possibilities for human-computer interaction with examples including tracking at distances from tens of centimeters through to several meters (for controlling the TV at a distance), supporting tracking using a moving depth camera (for mobile scenarios), and arbitrary camera placements (for VR headsets). These features are achieved through a new pipeline that combines a multi-layered discriminative reinitialization strategy for per-frame pose estimation, followed by a generative model-fitting stage. We provide extensive technical details and a detailed qualitative and quantitative analysis. Toby Sharp, Cem Keskin, Duncan P. Robertson, Jonathan Taylor 0001, Jamie Shotton, David Kim 0002, Christoph Rhemann, Ido Leichter, Alon Vinnikov, Daniel Freedman, Pushmeet Kohli, Eyal Krupka, Andrew W. Fitzgibbon, Shahram Izadi |
CHI | 14 |
| 2015 | 3D scanning deformable objects with a single RGBD sensorabstractWe present a 3D scanning system for deformable objects that uses only a single Kinect sensor. Our work allows considerable amount of nonrigid deformations during scanning, and achieves high quality results without heavily constraining user or camera motion. We do not rely on any prior shape knowledge, enabling general object scanning with freeform deformations. To deal with the drift problem when nonrigidly aligning the input sequence, we automatically detect loop closures, distribute the alignment error over the loop, and finally use a bundle adjustment algorithm to optimize for the latent 3D shape and nonrigid deformation parameters simultaneously. We demonstrate high quality scanning results in some challenging sequences, comparing with state of art nonrigid techniques, as well as ground truth data. Mingsong Dou, Jonathan Taylor 0001, Henry Fuchs, Andrew W. Fitzgibbon, Shahram Izadi |
CVPR | 4 |
| 2015 | Large-scale and drift-free surface reconstruction using online subvolume registrationabstractDepth cameras have helped commoditize 3D digitization of the real-world. It is now feasible to use a single Kinect-like camera to scan in an entire building or other large-scale scenes. At large scales, however, there is an inherent chal-lenge of dealing with distortions and drift due to accumu-lated pose estimation errors. Existing techniques suffer from one or more of the following: a) requiring an expensive offline global optimization step taking hours to compute; b) needing a full second pass over the input depth frames to correct for accumulated errors; c) relying on RGB data alongside depth data to optimize poses; or d) requiring the user to create explicit loop closures to allow gross alignment errors to be resolved. In this paper, we present a method that addresses all of these issues. Our method supports online model correction, without needing to reprocess or store any input depth data. Even while performing global correction of a large 3D model, our method takes only minutes rather than hours to compute. Our model does not require any explicit loop closures to be detected and, finally, relies on depth data alone, allowing operation in low-lighting conditions. We show qualitative results on many large scale scenes, high-lighting the lack of error and drift in our reconstructions. We compare to state of the art techniques and demonstrate large-scale dense surface reconstruction “in the dark”, a capability not offered by RGB-D techniques. 1. Nicola Fioraio, Jonathan Taylor 0001, Andrew W. Fitzgibbon, Luigi Di Stefano, Shahram Izadi |
CVPR | 3 |
| 2015 | Learning an efficient model of hand shape variation from depth imagesabstractWe describe how to learn a compact and efficient model of the surface deformation of human hands. The model is built from a set of noisy depth images of a diverse set of subjects performing different poses with their hands. We represent the observed surface using Loop subdivision of a control mesh that is deformed by our learned parametric shape and pose model. The model simultaneously accounts for variation in subject-specific shape and subject-agnostic pose. Specifically, hand shape is parameterized as a linear combination of a mean mesh in a neutral pose with a small number of offset vectors. This mesh is then articulated using standard linear blend skinning (LBS) to generate the control mesh of a subdivision surface. We define an energy that encourages each depth pixel to be explained by our model, and the use of a smooth subdivision surface allows us to optimize for all parameters jointly from a rough initialization. The efficacy of our method is demonstrated using both synthetic and real data, where it is shown that hand shape variation can be represented using only a small number of basis components. We compare with other approaches including PCA and show a substantial improvement in the representational power of our model, while maintaining the efficiency of a linear shape basis. Sameh Khamis, Jonathan Taylor 0001, Jamie Shotton, Cem Keskin, Shahram Izadi, Andrew W. Fitzgibbon |
CVPR | 6 |
| 2015 | Exploiting uncertainty in regression forests for accurate camera relocalizationabstractRecent advances in camera relocalization use predictions from a regression forest to guide the camera pose optimization procedure. In these methods, each tree associates one pixel with a point in the scene's 3D world coordinate frame. In previous work, these predictions were point estimates and the subsequent camera pose optimization implicitly assumed an isotropic distribution of these estimates. In this paper, we train a regression forest to predict mixtures of anisotropic 3D Gaussians and show how the predicted uncertainties can be taken into account for continuous pose optimization. Experiments show that our proposed method is able to relocalize up to 40% more frames than the state of the art. Julien P. C. Valentin, Matthias Nießner, Jamie Shotton, Andrew W. Fitzgibbon, Shahram Izadi, Philip Torr 0001 |
CVPR | 4 |
| 2015 | Secrets of Matrix Factorization: Approximations, Numerics, Manifold Optimization and Random RestartsabstractMatrix factorization (or low-rank matrix completion) with missing data is a key computation in many computer vision and machine learning tasks, and is also related to a broader class of nonlinear optimization problems such as bundle adjustment. The problem has received much attention recently, with renewed interest in variable-projection approaches, yielding dramatic improvements in reliability and speed. However, on a wide class of problems, no one approach dominates, and because the various approaches have been derived in a multitude of different ways, it has been difficult to unify them. This paper provides a unified derivation of a number of recent approaches, so that similarities and differences are easily observed. We also present a simple meta-algorithm which wraps any existing algorithm, yielding 100% success rate on many standard datasets. Given 100% success, the focus of evaluation must turn to speed, as 100% success is trivially achieved if we do not care about speed. Again our unification allows a number of generic improvements applicable to all members of the family to be isolated, yielding a unified algorithm that outperforms our re-implementation of existing algorithms, which in some cases already outperform the original authors' publicly available codes. Je Hyeong Hong, Andrew W. Fitzgibbon |
ICCV | 2 |
| 2015 | Efficient Solution to the Epipolar Geometry for Radially Distorted CamerasabstractThe estimation of the epipolar geometry of two cameras from image matches is a fundamental problem of computer vision with many applications. While the closely related problem of estimating relative pose of two different uncalibrated cameras with radial distortion is of particular importance, none of the previously published methods is suitable for practical applications. These solutions are either numerically unstable, sensitive to noise, based on a large number of point correspondences, or simply too slow for real-time applications. In this paper, we present a new efficient solution to this problem that uses 10 image correspondences. By manipulating ten input polynomial equations, we derive a degree 10 polynomial equation in one variable. The solutions to this equation are efficiently found using the Sturm sequences method. In the experiments, we show that the proposed solution is stable, noise resistant, and fast, and as such efficiently usable in a practical Structure-from-Motion pipeline. Zuzana Kukelova, Jan Heller, Martin Bujnak, Andrew W. Fitzgibbon, Tomás Pajdla |
ICCV | 4 |
| 2015 | Reflection Modeling for Passive StereoabstractStereo reconstruction in presence of reality faces many challenges that still need to be addressed. This paper considers reflections, which introduce incorrect matches due to the observation violating the diffuse-world assumption underlying the majority of stereo techniques. Unlike most existing work, which employ regularization or robust data terms to suppress such errors, we derive two least squares models from first principles that generalize diffuse world stereo and explicitly take reflections into account. These models are parametrized by depth, orientation and material properties, resulting in a total of up to 5 parameters per pixel that have to be estimated. Additionally large non-local interactions between viewed and reflected surface have to be taken into account. These two properties make inference of the model appear prohibitive, but we present evidence that inference is actually possible using a variant of patch match stereo. Rahul Nair 0006, Andrew W. Fitzgibbon, Daniel Kondermann, Carsten Rother |
ICCV | 2 |
| 2015 | Towards Pointless Structure from Motion: 3D Reconstruction and Camera Parameters from General 3D CurvesabstractModern structure from motion (SfM) remains dependent on point features to recover camera positions, meaning that reconstruction is severely hampered in low-texture environments, for example scanning a plain coffee cup on an uncluttered table. We show how 3D curves can be used to refine camera position estimation in challenging low-texture scenes. In contrast to previous work, we allow the curves to be partially observed in all images, meaning that for the first time, curve-based SfM can be demonstrated in realistic scenes. The algorithm is based on bundle adjustment, so needs an initial estimate, but even a poor estimate from a few point correspondences can be substantially improved by including curves, suggesting that this method would benefit many existing systems. Irina Nurutdinova, Andrew W. Fitzgibbon |
ICCV | 2 |
| 2015 | Model-Based Tracking at 300Hz Using Raw Time-of-Flight ObservationsabstractConsumer depth cameras have dramatically improved our ability to track rigid, articulated, and deformable 3D objects in real-time. However, depth cameras have a limited temporal resolution (frame-rate) that restricts the accuracy and robustness of tracking, especially for fast or unpredictable motion. In this paper, we show how to perform model-based object tracking which allows to reconstruct the object's depth at an order of magnitude higher frame-rate through simple modifications to an off-the-shelf depth camera. We focus on phase-based time-of-flight (ToF) sensing, which reconstructs each low frame-rate depth image from a set of short exposure 'raw' infrared captures. These raw captures are taken in quick succession near the beginning of each depth frame, and differ in the modulation of their active illumination. We make two contributions. First, we detail how to perform model-based tracking against these raw captures. Second, we show that by reprogramming the camera to space the raw captures uniformly in time, we obtain a 10x higher frame-rate, and thereby improve the ability to track fast-moving objects. Jan Stühmer, Sebastian Nowozin, Andrew W. Fitzgibbon, Richard Szeliski, Travis Perry, Sunil Acharya, Daniel Cremers, Jamie Shotton |
ICCV | 3 |
| 2015 | Metric Regression Forests for Correspondence Estimation
Gerard Pons-Moll, Jonathan Taylor 0001, Jamie Shotton, Aaron Hertzmann, Andrew W. Fitzgibbon |
Int. J. Comput. Vis. | 5 |
| 2015 | What Can Pictures Tell Us About Web Pages? Improving Document Search Using ImagesabstractTraditional Web search engines do not use the images in the HTML pages to find relevant documents for a given query. Instead, they typically operate by computing a measure of agreement between the keywords provided by the user and only the text portion of each page. In this paper we study whether the content of the pictures appearing in a Web page can be used to enrich the semantic description of an HTML document and consequently boost the performance of a keyword-based search engine. We present a Web-scalable system that exploits a pure text-based search engine to find an initial set of candidate documents for a given query. Then, the candidate set is reranked using visual information extracted from the images contained in the pages. The resulting system retains the computational efficiency of traditional text-based search engines with only a small additional storage cost needed to encode the visual information. We test our approach on one of the TREC Million Query Track benchmarks where we show that the exploitation of visual content yields improvement in accuracies for two distinct text-based search engines, including the system with the best reported performance on this benchmark. We further validate our approach by collecting document relevance judgements on our search results using Amazon Mechanical Turk. The results of this experiment confirm the improvement in accuracy produced by our image-based reranker over a pure text-based system. Sergio Rodríguez-Vaamonde, Lorenzo Torresani, Andrew W. Fitzgibbon |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Multi-output Learning for Camera RelocalizationabstractWe address the problem of estimating the pose of a cam- era relative to a known 3D scene from a single RGB-D frame. We formulate this problem as inversion of the generative rendering procedure, i.e., we want to find the camera pose corresponding to a rendering of the 3D scene model that is most similar with the observed input. This is a non-convex optimization problem with many local optima. We propose a hybrid discriminative-generative learning architecture that consists of: (i) a set of M predictors which generate M camera pose hypotheses, and (ii) a 'selector' or 'aggregator' that infers the best pose from the multiple pose hypotheses based on a similarity function. We are interested in predictors that not only produce good hypotheses but also hypotheses that are different from each other. Thus, we propose and study methods for learning 'marginally relevant' predictors, and compare their performance when used with different selection procedures. We evaluate our method on a recently released 3D reconstruction dataset with challenging camera poses, and scene variability. Experiments show that our method learns to make multiple predictions that are marginally relevant and can effectively select an accurate prediction. Furthermore, our method outperforms the state-of-the-art discriminative approach for camera relocalization. Abner Guzmán-Rivera, Pushmeet Kohli, Ben Glocker, Jamie Shotton, Toby Sharp, Andrew W. Fitzgibbon, Shahram Izadi |
CVPR | 6 |
| 2014 | SphereFlow: 6 DoF Scene Flow from RGB-D PairsabstractWe take a new approach to computing dense scene flow between a pair of consecutive RGB-D frames. We exploit the availability of depth data by seeking correspondences with respect to patches specified not as the pixels inside square windows, but as the 3D points that are the inliers of spheres in world space. Our primary contribution is to show that by reasoning in terms of such patches under 6 DoF rigid body motions in 3D, we succeed in obtaining compelling results at displacements large and small without relying on either of two simplifying assumptions that pervade much of the earlier literature: brightness constancy or local surface planarity. As a consequence of our approach, our output is a dense field of 3D rigid body motions, in contrast to the 3D translations that are the norm in scene flow. Reasoning in our manner additionally allows us to carry out occlusion handling using a 6 DoF consistency check for the flow computed in both directions and a patchwise silhouette check to help reason about alignments in occlusion areas, and to promote smoothness of the flow fields using an intuitive local rigidity prior. We carry out our optimization in two steps, obtaining a first correspondence field using an adaptation of PatchMatch, and subsequently using alpha-expansion to jointly handle occlusions and perform regularization. We show attractive flow results on challenging synthetic and real-world scenes that push the practical limits of the aforementioned assumptions. Michael Hornacek, Andrew W. Fitzgibbon, Carsten Rother |
CVPR | 2 |
| 2014 | User-Specific Hand Modeling from Monocular Depth SequencesabstractThis paper presents a method for acquiring dense nonrigid shape and deformation from a single monocular depth sensor. We focus on modeling the human hand, and assume that a single rough template model is available. We combine and extend existing work on model-based tracking, subdivision surface fitting, and mesh deformation to acquire detailed hand models from as few as 15 frames of depth data. We propose an objective that measures the error of fit between each sampled data point and a continuous model surface defined by a rigged control mesh, and uses as-rigid-as-possible (ARAP) regularizers to cleanly separate the model and template geometries. A key contribution is our use of a smooth model based on subdivision surfaces that allows simultaneous optimization over both correspondences and model parameters. This avoids the use of iterated closest point (ICP) algorithms which often lead to slow convergence. Automatic initialization is obtained using a regression forest trained to infer approximate correspondences. Experiments show that the resulting meshes model the user's hand shape more accurately than just adapting the shape parameters of the skeleton, and that the retargeted skeleton accurately models the user's articulations. We investigate the effect of various modeling choices, and show the benefits of using subdivision surfaces and ARAP regularization. Jonathan Taylor 0001, Richard V. Stebbing, Varun Ramakrishna, Cem Keskin, Jamie Shotton, Shahram Izadi, Aaron Hertzmann, Andrew W. Fitzgibbon |
CVPR | 8 |
| 2014 | Highly Overparameterized Optical Flow Using PatchMatch Belief Propagation
Michael Hornacek, Frederic Besse, Jan Kautz, Andrew W. Fitzgibbon, Carsten Rother |
ECCV (3) | 4 |
| 2014 | PMBP: PatchMatch Belief Propagation for Correspondence Field Estimation
Frederic Besse, Carsten Rother, Andrew W. Fitzgibbon, Jan Kautz |
Int. J. Comput. Vis. | 3 |
| 2014 | Joint Demosaicing and Denoising via Learned Nonparametric Random FieldsabstractWe introduce a machine learning approach to demosaicing, the reconstruction of color images from incomplete color filter array samples. There are two challenges to overcome by a demosaicing method: 1) it needs to model and respect the statistics of natural images in order to reconstruct natural looking images and 2) it should be able to perform well in the presence of noise. To facilitate an objective assessment of current methods, we introduce a public ground truth data set of natural images suitable for research in image demosaicing and denoising. We then use this large data set to develop a machine learning approach to demosaicing. Our proposed method addresses both demosaicing challenges by learning a statistical model of images and noise from hundreds of natural images. The resulting model performs simultaneous demosaicing and denoising. We show that the machine learning approach has a number of benefits: 1) the model is trained to directly optimize a user-specified performance measure such as peak signal-to-noise ratio (PSNR) or structural similarity; 2) we can handle novel color filter array layouts by retraining the model on such layouts; and 3) it outperforms the previous state-of-the-art, in some setups by 0.7-dB PSNR, faithfully reconstructing edges, textures, and smooth areas. Our results demonstrate that in demosaicing and related imaging applications, discriminatively trained machine learning models have the potential for peak performance at comparatively low engineering effort. Daniel Khashabi, Sebastian Nowozin, Jeremy Jancsary, Andrew W. Fitzgibbon |
IEEE Trans. Image Process. | 4 |
| 2014 | Real-time non-rigid reconstruction using an RGB-D cameraabstractWe present a combined hardware and software solution for markerless reconstruction of non-rigidly deforming physical objects with arbitrary shape in real-time . Our system uses a single self-contained stereo camera unit built from off-the-shelf components and consumer graphics hardware to generate spatio-temporally coherent 3D models at 30 Hz. A new stereo matching algorithm estimates real-time RGB-D data. We start by scanning a smooth template model of the subject as they move rigidly. This geometric surface prior avoids strong scene assumptions, such as a kinematic human skeleton or a parametric shape model. Next, a novel GPU pipeline performs non-rigid registration of live RGB-D data to the smooth template using an extended non-linear as-rigid-as-possible (ARAP) framework. High-frequency details are fused onto the final mesh using a linear deformation model. The system is an order of magnitude faster than state-of-the-art methods, while matching the quality and robustness of many offline algorithms. We show precise real-time reconstructions of diverse scenes, including: large deformations of users' heads, hands, and upper bodies; fine-scale wrinkles and folds of skin and clothing; and non-rigid interactions performed by users on flexible objects such as toys. We demonstrate how acquired models can be used for many interactive scenarios, including re-texturing, online performance capture and preview, and real-time shape and motion re-targeting. Michael Zollhöfer, Matthias Nießner, Shahram Izadi, Christoph Rhemann, Christopher Zach, Matthew Fisher, Chenglei Wu, Andrew W. Fitzgibbon, Charles T. Loop, Christian Theobalt, Marc Stamminger |
ACM Trans. Graph. | 8 |
| 2014 | Kinectrack: 3D Pose Estimation Using a Projected Dense Dot PatternabstractKinectrack is a novel approach to six-DoF tracking that provides agile real-time pose estimation using only commodity hardware. The dot pattern emitter and IR camera components of the standard Kinect device are separated to allow the emitter to roam freely relative to a fixed camera. The six-DoF pose of the emitter component is recovered by matching the dense dot pattern observed by the camera to a pre-captured reference image. A novel matching technique is introduced to obtain the dense dot pattern correspondences efficiently in wide- and adaptive-baseline scenarios that requires only a small subset of the full dense dot pattern to fall within the field of view of the fixed camera. An auto-calibration process is proposed in order to obtain the intrinsic parameters of the fixed camera and the internal dot pattern reference image of the emitter. The system simultaneously recovers the six-DoF pose of the emitter device and the piecewise planar 3D scene structure. Kinectrack provides a low-cost method for tracking an object without any on-board computation, with small size and only simple electronics. This paper extends the original ISMAR 2012 submission, including a demonstration of robust pose tracking for AR and examples of matching in planar and non-planar scenes. Paul McIlroy, Shahram Izadi, Andrew W. Fitzgibbon |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2013 | Metric Regression Forests for Human Pose EstimationabstractTraditionally, human pose estimation algorithms could be classified into generative [2] and discriminative [4] approaches. Generative approaches model the likelihood of the observations given a pose estimate, however, they are susceptible to local minima and thus require good initial pose estimates. Discriminative approaches learn a direct mapping from image features to pose space from training data, however, they struggle to generalize to unseen poses. Building on previous work [3], Taylor et al. [5] bypass some of these limitations using a hybrid-approach that discriminatively predicts, for each pixel in a depth image, a corresponding point on the surface of a humanoid mesh model. This mesh model is then robustly fit to the resulting set of correspondences using local optimization. Surprisingly though, these correspondences are actually inferred using a random forest whose structure was trained using a classification objective that arbitrarily equates target model points belonging to the same predefined body part [3]. In this paper, we address Taylor et al.’s use of this proxy classification objective by proposing Metric Space Information Gain (MSIG), a replacement objective function for training a random forest to directly minimize the uncertainty over the target model points, naturally encoding the correlation between these points as a function of the geodesic distance. To this end, we view the surface of the model U as a metric space (U,dU) defined by the geodesic distance metric dU (see first panel of Figure 1). The natural objective function to minimize the uncertainty in the resulting true distributions that result from a split function s in such a space, is the information gain I(s) [1]. This is generally approximated using an empirical distribution Q = {ui} ⊆U drawn from the true unsplit distribution pU as I(s)≈ I(s;Q) = Ĥ(Q)− ∑ i∈{L,R} |Qi| |Q| Ĥ(Qi), (1) Gerard Pons-Moll, Jonathan Taylor 0001, Jamie Shotton, Aaron Hertzmann, Andrew W. Fitzgibbon |
BMVC | 5 |
| 2013 | Scene Coordinate Regression Forests for Camera Relocalization in RGB-D ImagesabstractWe address the problem of inferring the pose of an RGB-D camera relative to a known 3D scene, given only a single acquired image. Our approach employs a regression forest that is capable of inferring an estimate of each pixel's correspondence to 3D points in the scene's world coordinate frame. The forest uses only simple depth and RGB pixel comparison features, and does not require the computation of feature descriptors. The forest is trained to be capable of predicting correspondences at any pixel, so no interest point detectors are required. The camera pose is inferred using a robust optimization scheme. This starts with an initial set of hypothesized camera poses, constructed by applying the forest at a small fraction of image pixels. Preemptive RANSAC then iterates sampling more pixels at which to evaluate the forest, counting inliers, and refining the hypothesized poses. We evaluate on several varied scenes captured with an RGB-D camera and observe that the proposed technique achieves highly accurate relocalization and substantially out-performs two state of the art baselines. Jamie Shotton, Ben Glocker, Christopher Zach, Shahram Izadi, Antonio Criminisi, Andrew W. Fitzgibbon |
CVPR | 6 |
| 2013 | What can pictures tell us about web pages?: improving document search using imagesabstractTraditional Web search engines do not use the images in the HTML pages to find relevant documents for a given query. Instead, they typically operate by computing a measure of agreement between the keywords provided by the user and only the text portion of each page. In this paper we study whether the content of the pictures appearing in a Web page can be used to enrich the semantic description of an HTML document and consequently boost the performance of a keyword-based search engine. We present a Web-scalable system that exploits a pure text-based search engine to find an initial set of candidate documents for a given query. Then, the candidate set is reranked using semantic information extracted from the images contained in the pages. The resulting system retains the computational efficiency of traditional text-based search engines with only a small additional storage cost needed to encode the visual information. We test our approach on the TREC 2009 Million Query Track, where we show that our use of visual content yields improvement in accuracies for two distinct text-based search engines, including the system with the best reported performance on this benchmark. Sergio Rodríguez-Vaamonde, Lorenzo Torresani, Andrew W. Fitzgibbon |
SIGIR | 3 |
| 2013 | What Shape Are Dolphins? Building 3D Morphable Models from 2D Imagesabstract3D morphable models are low-dimensional parameterizations of 3D object classes which provide a powerful means of associating 3D geometry to 2D images. However, morphable models are currently generated from 3D scans, so for general object classes such as animals they are economically and practically infeasible. We show that, given a small amount of user interaction (little more than that required to build a conventional morphable model), there is enough information in a collection of 2D pictures of certain object classes to generate a full 3D morphable model, even in the absence of surface texture. The key restriction is that the object class should not be strongly articulated, and that a very rough rigid model should be provided as an initial estimate of the “mean shape.” The model representation is a linear combination of subdivision surfaces, which we fit to image silhouettes and any identifiable key points using a novel combined continuous-discrete optimization strategy. Results are demonstrated on several natural object classes, and show that models of rather high quality can be obtained from this limited information. Thomas J. Cashman 0001, Andrew W. Fitzgibbon |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Efficient Human Pose Estimation from Single Depth ImagesabstractWe describe two new approaches to human pose estimation. Both can quickly and accurately predict the 3D positions of body joints from a single depth image without using any temporal information. The key to both approaches is the use of a large, realistic, and highly varied synthetic set of training images. This allows us to learn models that are largely invariant to factors such as pose, body shape, field-of-view cropping, and clothing. Our first approach employs an intermediate body parts representation, designed so that an accurate per-pixel classification of the parts will localize the joints of the body. The second approach instead directly regresses the positions of body joints. By using simple depth pixel comparison features and parallelizable decision forests, both approaches can run super-real time on consumer hardware. Our evaluation investigates many aspects of our methods, and compares the approaches to each other and to the state of the art. Results on silhouettes suggest broader applicability to other imaging modalities. Jamie Shotton, Ross B. Girshick, Andrew W. Fitzgibbon, Toby Sharp, Mat Cook, Mark Finocchio, Richard Moore 0003, Pushmeet Kohli, Antonio Criminisi, Alex Kipman, Andrew Blake 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | PMBP: PatchMatch Belief Propagation for Correspondence Field EstimationabstractPatchMatch is a simple, yet very powerful and successful method for optimizing continuous labelling problems.The algorithm has two main ingredients: the update of the solution space by sampling and the use of the spatial neighbourhood to propagate samples.We show how these ingredients are related to steps in a specific form of belief propagation in the continuous space, called Particle Belief Propagation (PBP).However, PBP has thus far been too slow to allow complex state spaces.We show that unifying the two approaches yields a new algorithm, PMBP, which is more accurate than PatchMatch and orders of magnitude faster than PBP.To illustrate the benefits of our PMBP method we have built a new stereo matching algorithm with unary terms which are borrowed from the recent PatchMatch Stereo work and novel realistic pairwise terms that provide smoothness.We have experimentally verified that our method is an improvement over state-of-the-art techniques at sub-pixel accuracy level. IntroductionThis paper draws a new connection between two existing algorithms for estimation of correspondence fields between images: Belief Propagation [15,19] and PatchMatch [1, 2].Correspondence fields arise in problems such as dense stereo reconstruction, optical flow estimation, and a variety of computational photography applications such as recoloring, deblurring, high dynamic range imaging, and inpainting.By analysing the connection between the methods, we obtain a new algorithm which has performance superior to both its antecedents, and in the case of stereo matching, represents the current state of the art on the Middlebury benchmark at sub-pixel accuracy.The first contribution of our work is a detailed description of PatchMatch and belief propagation in terms that allow the connection between the two to be clearly described.This analysis is largely self-contained, and comprises the first major section of the paper.Our second contribution is in the use of this analysis to define a new algorithm: PatchMatch Belief Propagation (PMBP) which, despite its relative simplicity, is more accurate than PatchMatch and orders of magnitude faster than PBP. Frederic Besse, Carsten Rother, Andrew W. Fitzgibbon, Jan Kautz |
BMVC | 3 |
| 2012 | The Vitruvian manifold: Inferring dense correspondences for one-shot human pose estimationabstractFitting an articulated model to image data is often approached as an optimization over both model pose and model-to-image correspondence. For complex models such as humans, previous work has required a good initialization, or an alternating minimization between correspondence and pose. In this paper we investigate one-shot pose estimation: can we directly infer correspondences using a regression function trained to be invariant to body size and shape, and then optimize the model pose just once? We evaluate on several challenging single-frame data sets containing a wide variety of body poses, shapes, torso rotations, and image cropping. Our experiments demonstrate that one-shot pose estimation achieves state of the art results and runs in real-time. Jonathan Taylor 0001, Jamie Shotton, Toby Sharp, Andrew W. Fitzgibbon |
CVPR | 4 |
| 2012 | A unifying resolution-independent formulation for early visionabstractWe present a model for early vision tasks such as denoising, super-resolution, deblurring, and demosaicing. The model provides a resolution-independent representation of discrete images which admits a truly rotationally invariant prior. The model generalizes several existing approaches: variational methods, finite element methods, and discrete random fields. The primary contribution is a novel energy functional which has not previously been written down, which combines the discrete measurements from pixels with a continuous-domain world viewed through continous-domain point-spread functions. The value of the functional is that simple priors (such as total variation and generalizations) on the continous-domain world become realistic priors on the sampled images. We show that despite its apparent complexity, optimization of this model depends on just a few computational primitives, which although tedious to derive, can now be reused in many domains. We define a set of optimization algorithms which greatly overcome the apparent complexity of this model, and make possible its practical application. New experimental results include infinite-resolution upsampling, and a method for obtaining “subpixel superpixels”. Fabio Viola, Andrew W. Fitzgibbon, Roberto Cipolla |
CVPR | 2 |
| 2012 | Kinectrack: Agile 6-DoF tracking using a projected dot patternabstractWe present Kinectrack, a new six degree-of-freedom (6-DoF) tracker which allows real-time and low-cost pose estimation using only commodity hardware. We decouple the dot pattern emitter and IR camera of the Kinect. Keeping the camera fixed and moving the IR emitter in the environment, we recover the 6-DoF pose of the emitter by matching the observed dot pattern in the field-of-view of the camera to a pre-captured reference image. We propose a novel matching technique to obtain dot pattern correspondences efficiently in wide- and adaptive-baseline scenarios. We also propose an auto-calibration method to obtain the camera intrinsics and dot pattern reference image. The performance of Kinectrack is evaluated and the rotational and translational accuracy of the system is measured relative to ground truth for both planar and multi-planar scene geometry. Our system can simultaneously recover the 6-DoF pose of the device and also recover piecewise planar 3D scene structure, and can be used as a low-cost method for tracking a device without any on-board computation, with small size and only simple electronics. Paul McIlroy, Shahram Izadi, Andrew W. Fitzgibbon |
ISMAR | 3 |
| 2012 | KinÊtre: animating the world with the human bodyabstractKinÊtre allows novice users to scan arbitrary physical objects and bring them to life in seconds. The fully interactive system allows diverse static meshes to be animated using the entire human body. Traditionally, the process of mesh animation is laborious and requires domain expertise, with rigging specified manually by an artist when designing the character. KinÊtre makes creating animations a more playful activity, conducted by novice users interactively "at runtime". This paper describes the KinÊtre system in full, highlighting key technical contributions and demonstrating many examples of users animating meshes of varying shapes and sizes. These include non-humanoid meshes and incomplete surfaces produced by 3D scanning - two challenging scenarios for existing mesh animation systems. Rather than targeting professional CG animators, KinÊtre is intended to bring mesh animation to a new audience of novice users. We demonstrate potential uses of our system for interactive storytelling and new forms of physical gaming. Jiawen Chen 0001, Shahram Izadi, Andrew W. Fitzgibbon |
UIST | 3 |
| 2012 | In Memoriam: Mark EveringhamabstractRecounts the career and contributions pf Mark Everingham. Andrew Zisserman, John M. Winn, Andrew W. Fitzgibbon, Luc Van Gool, Josef Sivic, Christopher K. I. Williams, David C. Hogg |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Real-time human pose recognition in parts from single depth imagesabstractWe propose a new method to quickly and accurately predict 3D positions of body joints from a single depth image, using no temporal information. We take an object recognition approach, designing an intermediate body parts representation that maps the difficult pose estimation problem into a simpler per-pixel classification problem. Our large and highly varied training dataset allows the classifier to estimate body parts invariant to pose, body shape, clothing, etc. Finally we generate confidence-scored 3D proposals of several body joints by reprojecting the classification result and finding local modes. The system runs at 200 frames per second on consumer hardware. Our evaluation shows high accuracy on both synthetic and real test sets, and investigates the effect of several training parameters. We achieve state of the art accuracy in our comparison with related work and demonstrate improved generalization over exact whole-skeleton nearest neighbor matching. Jamie Shotton, Andrew W. Fitzgibbon, Mat Cook, Toby Sharp, Mark Finocchio, Richard Moore 0003, Alex Kipman, Andrew Blake 0001 |
CVPR | 2 |
| 2011 | Efficient regression of general-activity human poses from depth imagesabstractWe present a new approach to general-activity human pose estimation from depth images, building on Hough forests. We extend existing techniques in several ways: real time prediction of multiple 3D joints, explicit learning of voting weights, vote compression to allow larger training sets, and a comparison of several decision-tree training objectives. Key aspects of our work include: regression directly from the raw depth image, without the use of an arbitrary intermediate representation; applicability to general motions (not constrained to particular activities) and the ability to localize occluded as well as visible body joints. Experimental results demonstrate that our method produces state of the art results on several data sets including the challenging MSRC-5000 pose estimation test set, at a speed of about 200 frames per second. Results on silhouettes suggest broader applicability to other imaging modalities. Ross B. Girshick, Jamie Shotton, Pushmeet Kohli, Antonio Criminisi, Andrew W. Fitzgibbon |
ICCV | 5 |
| 2011 | KinectFusion: Real-time dense surface mapping and trackingabstractWe present a system for accurate real-time mapping of complex and arbitrary indoor scenes in variable lighting conditions, using only a moving low-cost depth camera and commodity graphics hardware. We fuse all of the depth data streamed from a Kinect sensor into a single global implicit surface model of the observed scene in real-time. The current sensor pose is simultaneously obtained by tracking the live depth frame relative to the global model using a coarse-to-fine iterative closest point (ICP) algorithm, which uses all of the observed depth data available. We demonstrate the advantages of tracking against the growing full surface model compared with frame-to-frame tracking, obtaining tracking and mapping results in constant time within room sized scenes with limited drift and high accuracy. We also show both qualitative and quantitative results relating to various aspects of our tracking and mapping system. Modelling of natural scenes, in real-time with only commodity sensor and GPU hardware, promises an exciting step forward in augmented reality (AR), in particular, it allows dense surfaces to be reconstructed in real-time, with a level of detail and robustness beyond any solution yet presented using passive computer vision. Richard A. Newcombe, Shahram Izadi, Otmar Hilliges, David Molyneaux, David Kim 0002, Andrew J. Davison, Pushmeet Kohli, Jamie Shotton, Steve Hodges 0001, Andrew W. Fitzgibbon |
ISMAR | 10 |
| 2011 | PiCoDes: Learning a Compact Code for Novel-Category RecognitionabstractWe introduce PiCoDes: a very compact image descriptor which nevertheless allows high performance on object category recognition. In particular, we address novel-category recognition: the task of defining indexing structures and image representations which enable a large collection of images to be searched for an object category that was not known when the index was built. Instead, the training images defining the category are supplied at query time. We explicitly learn descriptors of a given length (from as small as 16 bytes per image) which have good object-recognition performance. In contrast to previous work in the domain of object recognition, we do not choose an arbitrary intermediate representation, but explicitly learn short codes. In contrast to previous approaches to learn compact codes, we optimize explicitly for (an upper bound on) classification performance. Optimization directly for binary features is difficult and nonconvex, but we present an alternation scheme and convex upper bound which demonstrate excellent performance in practice. PiCoDes of 256 bytes match the accuracy of the current best known classifier for the Caltech256 benchmark, but they decrease the database storage size by a factor of 100 and speed-up the training and testing of novel classes by orders of magnitude. Alessandro Bergamo, Lorenzo Torresani, Andrew W. Fitzgibbon |
NIPS | 3 |
| 2011 | KinectFusion: real-time 3D reconstruction and interaction using a moving depth cameraabstractKinectFusion enables a user holding and moving a standard Kinect camera to rapidly create detailed 3D reconstructions of an indoor scene. Only the depth data from Kinect is used to track the 3D pose of the sensor and reconstruct, geometrically precise, 3D models of the physical scene in real-time. The capabilities of KinectFusion, as well as the novel GPU-based pipeline are described in full. Uses of the core system for low-cost handheld scanning, and geometry-aware augmented reality and physics-based interactions are shown. Novel extensions to the core GPU pipeline demonstrate object segmentation and user interaction directly in front of the sensor, without degrading camera tracking or reconstruction. These extensions are used to enable real-time multi-touch interactions anywhere, allowing any planar or non-planar reconstructed physical surface to be appropriated for touch. Shahram Izadi, David Kim 0002, Otmar Hilliges, David Molyneaux, Richard A. Newcombe, Pushmeet Kohli, Jamie Shotton, Steve Hodges 0001, Dustin Freeman, Andrew J. Davison, Andrew W. Fitzgibbon |
UIST | 11 |
| 2010 | Finding nemo: Deformable object class modelling using curve matchingabstractAn image search for “clownfish” yields many photos of clownfish, each of a different individual of a different 3D shape in a different pose. Yet, to the human observer, this set of images contains enough information to infer the underlying 3D deformable object class. Our goal is to recover such a deformable object class model directly from unordered images. For classes where feature-point correspondences can be found, this is a straightforward extension of non-rigid factorization, yielding a set of 3D basis shapes to explain the 2D data. However, when each image is of a different object instance, surface texture is generally unique to each individual, and does not give rise to usable image point correspondences. We overcome this sparsity using curve correspondences (crease-edge silhouettes or class-specific internal texture edges). Even rigid contour reconstruction is difficult due to the lack of reliable correspondences. We incorporate correspondence variation into the optimization, thereby extending contour-based reconstruction techniques to deformable object modelling. The notion of correspondence is extended to include mappings between 2D image curves and corresponding parts of the desired 3D object surface. Combined with class-specific priors, our method enables effective de-formable class reconstruction from unordered images, despite significant occlusion and the scarcity of shared 2D image features. Mukta Prasad, Andrew W. Fitzgibbon, Andrew Zisserman, Luc Van Gool |
CVPR | 2 |
| 2010 | Efficient Object Category Recognition Using Classemes
Lorenzo Torresani, Martin Szummer, Andrew W. Fitzgibbon |
ECCV (1) | 3 |
| 2010 | A fast natural Newton method
Nicolas Le Roux, Andrew W. Fitzgibbon |
ICML | 2 |
| 2009 | Learning query-dependent prefilters for scalable image retrievalabstractWe describe an algorithm for similar-image search which is designed to be efficient for extremely large collections of images. For each query, a small response set is selected by a fast prefilter, after which a more accurate ranker may be applied to each image in the response set. We consider a class of prefilters comprising disjunctions of conjunctions (“ORs of ANDs”) of Boolean features. AND filters can be implemented efficiently using skipped inverted files, a key component of Web-scale text search engines. These structures permit search in time proportional to the response set size. The prefilters are learned from training examples, and refined at query time to produce an approximately bounded response set. We cast prefiltering as an optimization problem: for each test query, select the OR-of-AND filter which maximizes training-set recall for an adjustable bound on response set size. This may be efficiently implemented by selecting from a large pool of candidate conjunctions of Boolean features using a linear program relaxation. Tests on object class recognition show that this relatively simple filter is nevertheless powerful enough to capture some semantic information. Lorenzo Torresani, Martin Szummer, Andrew W. Fitzgibbon |
CVPR | 3 |
| 2009 | Global Stereo Reconstruction under Second-Order Smoothness PriorsabstractSecond-order priors on the smoothness of 3D surfaces are a better model of typical scenes than first-order priors. However, stereo reconstruction using global inference algorithms, such as graph cuts, has not been able to incorporate second-order priors because the triple cliques needed to express them yield intractable (nonsubmodular) optimization problems. This paper shows that inference with triple cliques can be effectively performed. Our optimization strategy is a development of recent extensions to alpha -- expansion, based on the "QPBO" algorithm. The strategy is to repeatedly merge proposal depth maps using a novel extension of QPBO. Proposal depth maps can come from any source, for example, frontoparallel planes as in alpha-expansion, or indeed any existing stereo algorithm, with arbitrary parameter settings. Oliver J. Woodford, Philip Torr 0001, Ian D. Reid 0001, Andrew W. Fitzgibbon |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2008 | Global stereo reconstruction under second order smoothness priorsabstractSecond-order priors on the smoothness of 3D surfaces are a better model of typical scenes than first-order priors. However, stereo reconstruction using global inference algorithms, such as graph-cuts, has not been able to incorporate second-order priors because the triple cliques needed to express them yield intractable (non-submodular) optimization problems. This paper shows that inference with triple cliques can be effectively optimized. Our optimization strategy is a development of recent extensions to a-expansion, based on the "QPBO" algorithm [5, 14, 26]. The strategy is to repeatedly merge proposal depth maps using a novel extension of QPBO. Proposal depth maps can come from any source, for example fronto-parallel planes as in a-expansion, or indeed any existing stereo algorithm, with arbitrary parameter settings. Experimental results demonstrate the usefulness of the second-order prior and the efficacy of our optimization framework. An implementation of our stereo framework is available online [34]. Oliver J. Woodford, Philip Torr 0001, Ian D. Reid 0001, Andrew W. Fitzgibbon |
CVPR | 4 |
| 2008 | Editorial
Andrew W. Fitzgibbon, Camillo J. Taylor, Yann LeCun |
Int. J. Comput. Vis. | 1 |
| 2008 | Unwrap mosaics: a new representation for video editingabstractWe introduce a new representation for video which facilitates a number of common editing tasks. The representation has some of the power of a full reconstruction of 3D surface models from video, but is designed to be easy to recover from a priori unseen and uncalibrated footage. By modelling the image-formation process as a 2D-to-2D transformation from an object's texture map to the image, modulated by an object-space occlusion mask, we can recover a representation which we term the "unwrap mosaic". Many editing operations can be performed on the unwrap mosaic, and then re-composited into the original sequence, for example resizing objects, repainting textures, copying/cutting/pasting objects, and attaching effects layers to deforming objects. Alex Rav-Acha, Pushmeet Kohli, Carsten Rother, Andrew W. Fitzgibbon |
ACM Trans. Graph. | 4 |
| 2007 | On New View Synthesis Using Multiview StereoabstractWe show that application of modern multiview stereo techniques to the newview synthesis (NVS) problem introduces a number of non-trivial complexities. By simultaneously solving for the colour and depth of the new-view pixels we can eliminate the visual artefacts that conventional NVS-via-stereo suffers. The global occlusion reasoning which has led to considerable improvements in recent stereo algorithms can easily be included in the new algorithm, using a recently improved graph-cut-based optimizer for general multi-label conditional random fields (CRFs). However, the CRF priors that are important to success in stereo cannot be easily applied if the reconstruction is to be computed in the reference frame of the novel view. We address this problem by extending recent work on the fast optimization of texture priors in NVS to model the image edge structure, yielding a synthesis of the two approaches which yields good results on difficult image sequences. 1 Oliver J. Woodford, Ian D. Reid 0001, Philip Torr 0001, Andrew W. Fitzgibbon |
BMVC | 4 |
| 2007 | Combining local and global motion models for feature point trackingabstractAccurate feature point tracks through long sequences are a valuable substrate for many computer vision applications, e.g. non-rigid body tracking, video segmentation, video matching, and even object recognition. Existing algorithms may be arranged along an axis indicating how global the motion model used to constrain tracks is. Local methods, such as the KLT tracker, depend on local models of feature appearance, and are easily distracted by occlusions, repeated structure, and image noise. This leads to short tracks, many of which are incorrect. Alone, these require considerable postprocessing to obtain a useful result. In restricted scenes, for example a rigid scene through which a camera is moving, such postprocessing can make use of global motion models to allow "guided matching " which yields long high-quality feature tracks. However, many scenes of interest contain multiple motions or significant non-rigid deformations which mean that guided matching cannot be applied. In this paper we propose a general amalgam of local and global models to improve tracking even in these difficult cases. By viewing rank-constrained tracking as a probabilistic model of 2D tracks rather than 3D motion, we obtain a strong, robust motion prior, derived from the global motion in the scene. The result is a simple and powerful prior whose strength is easily tuned, enabling its use in any existing tracking algorithm. Aeron Buchanan, Andrew W. Fitzgibbon |
CVPR | 2 |
| 2007 | Efficient new-view synthesis using pairwise dictionary priorsabstractNew-view synthesis (NVS) using texture priors (as opposed to surface-smoothness priors) can yield high quality results, but the standard formulation is in terms of large-clique Markov random fields (MRFs). Only local optimization methods such as iterated conditional modes, which are prone to fall into local minima close to the initial estimate, are practical for solving these problems. In this paper we replace the large-clique energies with pairwise potentials, by restricting the patch dictionary for each clique to image regions suitable for that clique. This enables for the first time the use of a global optimization method, such as tree-reweighted message passing, to solve the NVS problem with image-based priors. We employ a robust, truncated quadratic kernel to reject outliers caused by occlusions, specularities and moving objects, within our global optimization. Because the MRF optimization is thus fast, computing the unary potentials becomes the new performance bottleneck. An additional contribution of this paper is a novel, fast method for enumerating color modes of the per-pixel unary potentials, despite the non-convex nature of our robust kernel. We compare the results of our technique with other rendering methods, and discuss the relative merits and flaws of regularizing color, and of local versus global dictionaries. Oliver J. Woodford, Ian D. Reid 0001, Andrew W. Fitzgibbon |
CVPR | 3 |
| 2007 | Reconstructing High Quality Face-Surfaces using Model Based StereoabstractWe present a novel model based stereo system, which accurately extracts the 3D shape and pose of faces from multiple images taken simultaneously. Extracting the 3D shape from images is important in areas such as pose-invariant face recognition and image manipulation. The method is based on a 3D morphable face model learned from a database of facial scans. The use of a strong face prior allows us to extract high precision surfaces from stereo data of faces, where traditional correlation based stereo methods fail because of the mostly textureless input images. The method uses two or more uncalibrated images of arbitrary baseline, estimating calibration and shape simultaneously. Results using two and three input images are presented. We replace the lighting and albedo estimation of a monocular method with the use of stereo information, making the system more accurate and robust. We evaluate the method using ground truth data and the standard PIE image dataset. A comparison with the state of the art monocular system shows that the new method has a significantly higher accuracy. Brian Amberg, Andrew Blake 0001, Andrew W. Fitzgibbon, Sami Romdhani, Thomas Vetter |
ICCV | 3 |
| 2007 | Learning priors for calibrating families of stereo camerasabstractOnline camera recalibration is necessary for long-term deployment of computer vision systems. Existing algorithms assume that the source of recalibration information is a set of features in a general 3D scene; and that enough features are observed that the calibration problem is well-constrained. However; these assumptions are frequently invalid outside the laboratory. Real-world scenes often lack texture, contain repeated texture, or are mostly planar, making calibration difficult or impossible. In this paper we consider the calibration of families of stereo cameras, where each camera is assumed to have parameters drawn from a common but unknown prior distribution. We show how estimation of this prior using a small-number of offline-calibrated cameras (e.g. from the same production line) allows online calibration of additional cameras using a small number of point correspondences; and that using the estimated prior significantly increases the accuracy and robustness of stereo camera calibration. Andrew W. Fitzgibbon, Duncan P. Robertson, Antonio Criminisi, Srikumar Ramalingam, Andrew Blake 0001 |
ICCV | 1 |
| 2007 | The Joint Manifold Model for Semi-supervised Multi-valued RegressionabstractMany computer vision tasks may be expressed as the problem of learning a mapping between image space and a parameter space. For example, in human body pose estimation, recent research has directly modelled the mapping from image features (z) to joint angles (θ). Fitting such models requires training data in the form of labelled (zθ) pairs, from which are learned the conditional densities p(zθ). Inference is then simple: given test image featuresz, the conditional (zθ) is immediately computed. However large amounts of training data are required to fit the models, particularly in the case where the spaces are high dimensional. We show how the use of unlabelled data—samples from the marginal distributions p(z) and p(θ)—may be used to improve fitting. This is valuable because it is often significantly easier to obtain unlabelled than labelled samples. We use a Gaussian process latent variable model to learn the mapping from a shared latent low-dimensional manifold to the feature and parameter spaces. This extends existing approaches to (a) use unlabelled data, and (b) represent one-to-many mappings. Experiments on synthetic and real problems demonstrate how the use of unlabelled data improves over existing techniques. In our comparisons, we include existing approaches that are explicitly semi-supervised as well as those which implicitly make use of unlabelled examples. Ramanan Navaratnam, Andrew W. Fitzgibbon, Roberto Cipolla |
ICCV | 2 |
| 2006 | Automatic Video Segmentation using Spatiotemporal T-JunctionsabstractThe problem of figure–ground segmentation is of great importance in both video editing and visual perception tasks. Classical video segmentation algorithms approach the problem from one of two perspectives. At one extreme, global approaches constrain the camera motion to simplify the image structure. At the other extreme, local approaches estimate motion in small image regions over a small number of frames and tend to produce noisy signals that are difficult to segment. With recent advances in image segmentation showing that sparse information is often sufficient for figure– ground segmentation it seems surprising then that with the extra temporal information of video, an unconstrained automatic figure–ground segmentation algorithm still eludes the research community. In this paper we present an automatic video segmentation algorithm that is intermediate between these two extremes and uses spatiotemporal features to regularize the segmentation. Detecting spatiotemporal T-junctions that indicate occlusion edges, we learn an occlusion edge model that is used within a colour contrast sensitive MRF to segment individual frames of a video sequence. T-junctions are learnt and classified using a support vector machine and a Gaussian mixture model is fitted to the (foreground, background) pixel pairs sampled from the detected T-junctions. Graph cut is then used to segment each frame of the video showing that sparse occlusion edge information can automatically initialize the video segmentation problem. 1 Nicholas Apostoloff, Andrew W. Fitzgibbon |
BMVC | 2 |
| 2006 | Semi-supervised Learning of Joint Density Models for Human Pose EstimationabstractLearning regression models (for example for body pose estimation, or BPE) currently requires large numbers of training examples—pairs of the form (image, pose parameters). These examples are difficult to obtain for many problems, demanding considerable effort in manual labelling. However it is easy to obtain unlabelled examples—in BPE, simply by collecting many images, and by sampling many poses using motion capture. We show how the use of unlabelled examples can improve the performance of such estimators, making better use of the difficult-to-obtain training examples. Because the distribution of parameters conditioned on a given image is often multimodal, conventional regression models must be extended to allow for multiple modes. Such extensions have to date had a pre-set number of modes, independent of the contents of the input image, and amount to fitting several regressors simultaneously. Our framework models instead the joint distribution of images and poses, so the conditional estimates are inherently multimodal, and the number of modes is a function of the joint-space complexity, rather than of the maximum number of output modes. We demonstrate the improvements obtainable by using unlabelled samples on synthetic examples and on a real pose estimation problem, and demonstrate in both cases the additional accuracy provided by the use of unlabelled data. 1 Ramanan Navaratnam, Andrew W. Fitzgibbon, Roberto Cipolla |
BMVC | 2 |
| 2006 | Fields of Experts for Image-based RenderingabstractImage priors for novel view synthesis have traditionally been non-parametric models based on large libraries of image patch exemplars, producing highquality results but making inference very slow. Recently a parametric framework, called Fields of Experts, has been proposed for image restoration that promises to speed up inference dramatically. In this paper we apply Fields of Experts for the first time to the problem of novel view synthesis, posed as a Markov random field labelling problem with very large cliques. Additionally, we introduce to computer vision for the first time a new optimization algorithm from statistical physics which reaches better minima than the ICM and simulated annealing algorithms to which such large-clique problems have previously been restricted. 1 Oliver J. Woodford, Ian D. Reid 0001, Philip Torr 0001, Andrew W. Fitzgibbon |
BMVC | 4 |
| 2006 | Interactive Feature Tracking using K-D Trees and Dynamic ProgrammingabstractA new approach to template tracking is presented, incorporating three distinct contributions. Firstly, an explicit definition for a feature track is given. Secondly, the advantages of an image preprocessing stage are demonstrated and, in particular, the effectiveness of highly compressed image patch data stored in k-d trees for fast and discriminatory image patch searches. Thirdly, the k-d trees are used to generate multiple track hypotheses which are efficiently merged to give the optimal solution using dynamic programming. The explicit separation of feature detection and trajectory determination creates the basis for the novel use of k-d trees and dynamic programming. Multiple appearances and occlusion handling are seamlessly integrated into this framework. Appearance variation through the sequence is robustly handled in an iterative process. The work presented is a significant foundation for a powerful off-line feature tracking system, particularly in the context of interactive applications. Aeron Buchanan, Andrew W. Fitzgibbon |
CVPR (1) | 2 |
| 2006 | Single View Reconstruction of Curved SurfacesabstractRecent advances in single-view reconstruction (SVR) have been in modelling power (curved 2.5D surfaces) and automation (automatic photo pop-up). We extend SVR along both of these directions. We increase modelling power in several ways: (i) We represent general 3D surfaces, rather than 2.5D Monge patches; (ii) We describe a closed-form method to reconstruct a smooth surface from its image apparent contour, including multilocal singularities ("kidney-bean" self-occlusions); (iii) We show how to incorporate user-specified data such as surface normals, interpolation and approximation constraints; (iv) We show how this algorithm can be adapted to deal with surfaces of arbitrary genus. We also show how the modelling process can be automated for simple object shapes and views, using a-priori object class information. We demonstrate these advances on natural images drawn from a number of object classes. Mukta Prasad, Andrew W. Fitzgibbon |
CVPR (2) | 2 |
| 2006 | Shift-Invariant Dynamic Texture Recognition
Franco Woolfe, Andrew W. Fitzgibbon |
ECCV (2) | 2 |
| 2006 | Editorial
Andrew W. Fitzgibbon, Marc Pollefeys, Luc Van Gool, Andrew Zisserman |
Int. J. Comput. Vis. | 1 |
| 2005 | A Plumbline Constraint for the Rational Function Lens Distortion ModelabstractThe recently introduced Rational Function (RF) model permits a linear solution of epipolar geometry and lens distortion from image correspondences for a very general class of lenses. In this paper we show that the model also permits a very elegant form of the plumbline constraint which allows uncalibrated correction of lens distortion from a single image containing lines known to be straight in the world. We show how this may be expressed as a factorization problem, and discuss its behaviour with noisy data. We also introduce a simple reduced parametrization of the RF model which again models a range of existing lenses. Because the RF model provides a very simple form for the distorted lines, nonlinear minimization can compute the Sampson distance to the distorted lines, allowing fast and accurate estimation of the RF model even from noisy images. David Claus, Andrew W. Fitzgibbon |
BMVC | 2 |
| 2005 | Fast Image-based Rendering using Hierarchical Texture Priors
Oliver J. Woodford, Andrew W. Fitzgibbon |
BMVC | 2 |
| 2005 | Learning Spatiotemporal T-Junctions for Occlusion DetectionabstractThe goal of motion segmentation and layer extraction can be viewed as the detection and localization of occluding surfaces. A feature that has been shown to be a particularly strong indicator of occlusion, in both computer vision and neuroscience, is the T-junction; however, little progress has been made in T-junction detection. One reason for this is the difficulty in distinguishing false T-junctions (i.e. those not on an occluding edge) and real T-junctions in cluttered images. In addition to this, their photometric profile alone is not enough for reliable detection. This paper overcomes the first problem by searching for T-junctions not in space, but in space-time. This removes many false T-junctions and creates a simpler image structure to explore. The second problem is mitigated by learning the appearance of T-junctions in these spatiotemporal images. An RVM T-junction classifier is learnt from hand-labelled data using SIFT to capture its redundancy. This detector is then demonstrated in a novel occlusion detector that fuses Canny edges and T-junctions in the spatiotemporal domain to detect occluding edges in the spatial domain. Nicholas Apostoloff, Andrew W. Fitzgibbon |
CVPR (2) | 2 |
| 2005 | Damped Newton Algorithms for Matrix Factorization with Missing DataabstractThe problem of low-rank matrix factorization in the presence of missing data has seen significant attention in recent computer vision research. The approach that dominates the literature is EM-like alternation of closed-form solutions for the two factors of the matrix. An obvious alternative is nonlinear optimization of both factors simultaneously, a strategy which has seen little published research. This paper provides a comprehensive comparison of the two strategies by evaluating previously published factorization algorithms as well as some second order methods not previously presented for this problem. We conclude that, although alternation approaches can be very quick, their propensity to glacial convergence in narrow valleys of the cost function means that average-case performance is worse than second-order strategies. Further, we demonstrate the importance of two main observations: one, that schemes based on closed-form solutions alone are not suitable and that non-linear optimization strategies are faster, more accurate and provide more flexible frameworks for continued progress; and two, that basic objective functions are not adequate and that regularization priors must be incorporated, a process that is easier with nonlinear methods. A. M. Buchanan, Andrew W. Fitzgibbon |
CVPR (2) | 2 |
| 2005 | A Rational Function Lens Distortion Model for General CamerasabstractWe introduce a new rational function (RF) model for radial lens distortion in wide-angle and catadioptric lenses, which allows the simultaneous linear estimation of motion and lens geometry from two uncalibrated views of a 3D scene. In contrast to existing models which admit such linear estimates, the new model is not specialized to any particular lens geometry, but is sufficiently general to model a variety of extreme distortions. The key step is to define the mapping between image (pixel) coordinates and 3D rays in camera coordinates as a linear combination of nonlinear functions of the image coordinates. Like a "kernel trick", this allows a linear algorithm to estimate nonlinear models, and in particular offers a simple solution to the estimation of nonlinear image distortion. The model also yields an explicit form for the epipolar curves, allowing correspondence search to be efficiently guided by the epipolar geometry. We show results of an implementation of the RF model in estimating the geometry of a real camera lens from uncalibrated footage, and compare the estimate to one obtained using a calibration grid. David Claus, Andrew W. Fitzgibbon |
CVPR (1) | 2 |
| 2005 | BRDF and geometry capture from extended inhomogeneous samples using flash photography
James A. Paterson, David Claus, Andrew W. Fitzgibbon |
Comput. Graph. Forum | 3 |
| 2005 | Image-Based Rendering Using Image-Based Priors
Andrew W. Fitzgibbon, Yonatan Wexler, Andrew Zisserman |
Int. J. Comput. Vis. | 1 |
| 2004 | Bayesian Video Matting Using Learnt Image Priors
Nicholas Apostoloff, Andrew W. Fitzgibbon |
CVPR (1) | 2 |
| 2004 | Reliable Fiducial Detection in Natural Scenes
David Claus, Andrew W. Fitzgibbon |
ECCV (4) | 2 |
| 2004 | Advanced Visual Tracking
Simon J. Julier, Andrew J. Davison, Andrew W. Fitzgibbon |
ISMAR | 3 |
| 2004 | Invariant Fitting of Two View GeometryabstractThis paper describes an extension of Bookstein's and Sampson's methods, for fitting conics, to the determination of epipolar geometry, both in the calibrated case, where the Essential matrix E is to be determined or in the uncalibrated case, where we seek the fundamental matrix F. We desire that the fitting of the relation be invariant to Euclidean transformations of the image, and show that there is only one suitable normalization of the coefficients and that this normalization gives rise to a quadratic form allowing eigenvector methods to be used to find E or F, or an arbitrary homography H. The resulting method has the advantage that it exhibits the improved stability of previous methods for estimating the epipolar geometry, such as the preconditioning method of Hartley, while also being invariant to equiform transformations. Philip Torr 0001, Andrew W. Fitzgibbon |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | Invariant Fitting of Two View Geometry
Andrew W. Fitzgibbon, Philip Torr 0001 |
BMVC | 1 |
| 2003 | 3D head tracking using non-linear optimizationabstractAccurate and reliable tracking of the 3D position of human heads is a continuing research problem in computer vision. This paper addresses the specific problem of model-based tracking with a generic deformable 3D head model. Following the work of Vetter and Blanz, a collection of head models is obtained from a 3D scanner, registered and parameterized to give a generic head model which is linearly parameterized by a small number of parameters. This is the 3D analogue of Cootes and Taylor’s active appearance models. We cast tracking as a parameter estimation problem, and note that many existing solutions to the problem—such as CONDENSATION and Kalman filtering—are analogous to nonlinear optimization strategies in numerical analysis. We show how careful analysis of the error function, parameterization of the model pose parameters, and choice of optimizer allows us to robustly track 3D head pose in digital video camera footage of quickly moving heads. 1 James A. Paterson, Andrew W. Fitzgibbon |
BMVC | 2 |
| 2003 | Joint Manifold Distance: a new approach to appearance based clusteringabstractWe wish to match sets of images to sets of images where both sets are undergoing various distortions such as viewpoint and lighting changes. To this end we have developed a joint manifold distance (JMD) which measures the distance between two subspaces, where each subspace is invariant to a desired group of transformations, for example affine warping of the image plane. The JMD may be seen as generalizing invariant distance metrics such as tangent distance in two important ways. First, formally representing priors on the image distribution avoids certain difficulties, which in previous work have required ad-hoc correction. The second contribution is the observation that previous distances have been computed using what amounted to "home-grown" nonlinear optimizers, and that more reliable results can be obtained by using generic optimizers which have been developed in the numerical analysis community, and which automatically set the parameters which home-grown methods must set by art. The JMD is used in this work to cluster faces in video. Sets of faces detected in contiguous frames define the subspaces, and distance between the subspaces is computed using JMD. In this way the principal cast of a movie can be 'discovered' as the principal clusters. We demonstrate the method on a feature-length movie. Andrew W. Fitzgibbon, Andrew Zisserman |
CVPR (1) | 1 |
| 2003 | Learning epipolar geometry from image sequencesabstractWe wish to determine the epipolar geometry of a stereo camera pair from image measurements alone. This paper describes a solution to this problem, which does not require a parametric model of the camera system, and consequently applies equally well to a wide class of stereo configurations. Examples in the paper range from a standard pinhole stereo configuration to more exotic systems combining curved mirrors and wide-angle lenses. The method described here allows epipolar curves to be learnt from multiple image pairs acquired by stereo cameras with fixed configuration. By aggregating information over the multiple image pairs, a dense map of the epipolar curves can be determined on the images. The algorithm requires a large number of images, but has the distinct benefit that the correspondence problem does not have to be explicitly solved. We show that for standard stereo configurations the results are comparable to those obtained from a state of the art parametric model method, despite the significantly weaker constraints on the non-parametric model. The new algorithm is simple to implement, so it may easily be employed on a new and possibly complex camera system. Yonatan Wexler, Andrew W. Fitzgibbon, Andrew Zisserman |
CVPR (2) | 2 |
| 2003 | Image-based rendering using image-based priorsabstractGiven a set of images acquired from known viewpoints, we describe a method for synthesizing the image which would be seen from a new viewpoint. In contrast to existing techniques, which explicitly reconstruct the 3D geometry of the scene, we transform the problem to the reconstruction of colour rather than depth. This retains the benefits of geometric constraints, but projects out the ambiguities in depth estimation which occur in textureless regions. On the other hand, regularization is still needed in order to generate high-quality images. The paper's second contribution is to constrain the generated views to lie in the space of images whose texture statistics are those of the input images. This amounts to an image-based prior on the reconstruction which regularizes the solution, yielding realistic synthetic views. Examples are given of new view generation for cameras interpolated between the acquisition viewpoints - which enables synthetic steadicam stabilization of a sequence with a high level of realism. Andrew W. Fitzgibbon, Yonatan Wexler, Andrew Zisserman |
ICCV | 1 |
| 2003 | Robust registration of 2D and 3D point sets
Andrew W. Fitzgibbon |
Image Vis. Comput. | 1 |
| 2002 | Real-time gesture recognition using deterministic boostingabstractA gesture recognition system which can reliably recognize single-hand gestures in real time on a 600Mhz notebook computer is described. The system has a vocabulary of 46 gestures including the American sign language letterspelling alphabet and digits. It includes mouse movements such as drag and drop, and is demonstrated controlling a windowed operating system, editing a document and performing file-system operations with extremely low error rates over long time periods. Real-time performance is provided by a novel combination of exemplar-based classification and a new “deterministic boosting ” algorithm which can allow for fast online retraining. Importantly, each frame of video is processed independently: no temporal Markov model is used to constrain gesture identity, and the search region is the entire image. This places stringent requirements on the accuracy and speed of recognition, which are met by our proposed architecture. 1 Raymond Lockton, Andrew W. Fitzgibbon |
BMVC | 2 |
| 2002 | On Affine Invariant Clustering and Automatic Cast Listing in Movies
Andrew W. Fitzgibbon, Andrew Zisserman |
ECCV (3) | 1 |
| 2002 | Bayesian Estimation of Layers from Multiple Images
Yonatan Wexler, Andrew W. Fitzgibbon, Andrew Zisserman |
ECCV (3) | 2 |
| 2002 | Automated reconstruction from multiple photographsabstractThis paper reviews two research themes. The first is wide baseline matching - establishing correspondences and cameras for images acquired from very different viewpoints. The second is automated surface reconstruction of a 3D scene from close-range images. These two themes are linked in an example application of automated architectural reconstruction from images. We demonstrate reconstructions of several university buildings from multiple photographs. Frederik Schaffalitzky, Andrew Zisserman, Tomás Werner, Andrew W. Fitzgibbon |
ICIP (3) | 4 |
| 2001 | Robust Registration of 2D and 3D Point Sets
Andrew W. Fitzgibbon |
BMVC | 1 |
| 2001 | Simultaneous linear estimation of multiple view geometry and lens distortionabstractA problem in uncalibrated stereo reconstruction is that cameras which deviate from the pinhole model have to be pre-calibrated in order to correct for nonlinear lens distortion. If they are not, and point correspondence is attempted using the uncorrected images, the matching constraints provided by the fundamental matrix must be set so loose that point matching is significantly hampered. This paper shows how linear estimation of the fundamental matrix from two-view point correspondences may be augmented to include one term of radial lens distortion. This is achieved by (1) changing from the standard radial-lens model to another which (as we show) has equivalent power, but which takes a simpler form in homogeneous coordinates, and (2) expressing fundamental matrix estimation as a quadratic eigenvalue problem (QEP), for which efficient algorithms are well known. I derive the new estimator, and compare its performance against bundle-adjusted calibration-grid data. The new estimator is fast enough to be included in a RANSAC-based matching loop, and we show cases of matching being rendered possible by its use. I show how the same lens can be calibrated in a natural scene where the lack of straight lines precludes most previous techniques. The modification when the multi-view relation is a planar homography or trifocal tensor is described. Andrew W. Fitzgibbon |
CVPR (1) | 1 |
| 2001 | Stochastic Rigidity: Image Registration for Nowhere-Static Scenes
Andrew W. Fitzgibbon |
ICCV | 1 |
| 2000 | Multibody Structure and Motion: 3-D Reconstruction of Independently Moving Objects
Andrew W. Fitzgibbon, Andrew Zisserman |
ECCV (1) | 1 |
| 1999 | Improving Augmented Reality using Image and Scene ConstraintsabstractThe goal of augmented reality is to insert virtual objects into real video sequences. This paper shows that by incorporating image-based geometric constraints over mul-tiple views, we improve on traditional techniques which use purely 3D information. The constraints imposed are chosen to directly target perceptual cues, important to the human visual system, by which errors in AR are most readily perceived. Imposi-tion of the constraints is achieved by constrained maximum-likelihood estimation, and blends projective, affine and Euclidean geometry as appropriate in different cases. We introduce a number of examples of augmented reality tasks, show how image-based constraints can be incorporated into current 3D-based systems, and demonstrate the improvements conferred. R. A. Smith, Andrew W. Fitzgibbon, Andrew Zisserman |
BMVC | 2 |
| 1999 | Parallax Geometry of Smooth Surfaces in Multiple ViewsabstractThis paper investigates the multiple view geometry of smooth surfaces and a plane, where the plane provides a planar homography mapping between the views. Innovations are made in three areas: first, new solutions are given for the computation of epipolar and trifocal geometry for this type of scene. In particular it is shown that the epipole may be determined from bitangents between the homography registered occluding contours, and a new minimal solution is given for computing the trifocal tensor: Second, algorithms are demonstrated for automatically estimating the fundamental matrix and trifocal tensor from images of such scenes. Third, a method is developed for estimating camera matrices for a sequence of images of these scenes. These three areas are combined in a "freehand scanner" application where 3D texture-mapped graphical models of smooth objects are acquired directly from a video sequence of the object and plane. Geoffrey Cross, Andrew W. Fitzgibbon, Andrew Zisserman |
ICCV | 2 |
| 1999 | The Problem of Degeneracy in Structure and Motion Recovery from Uncalibrated Image Sequences
Philip Torr 0001, Andrew W. Fitzgibbon, Andrew Zisserman |
Int. J. Comput. Vis. | 2 |
| 1999 | Direct Least Square Fitting of EllipsesabstractThis work presents a new efficient method for fitting ellipses to scattered data. Previous algorithms either fitted general conics or were computationally expensive. By minimizing the algebraic distance subject to the constraint 4ac-b/sup 2/=1, the new method incorporates the ellipticity constraint into the normalization factor. The proposed method combines several advantages: It is ellipse-specific, so that even bad data will always return an ellipse. It can be solved naturally by a generalized eigensystem. It is extremely robust, efficient, and easy to implement. Andrew W. Fitzgibbon, Maurizio Pilu, Robert B. Fisher |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1998 | Real-time Panoramic Mosaics and Augmented RealityabstractThis paper investigates estimating exact imaging transformations accurately, reliably and efficiently. It is shown that in certain common computer vision situations the transformation required can be defined by a small number of parameters. Search is only required over these parameters, and consequently the search algorithms to estimate the transformation can be run at frame rate, without sacrificing robustness or accuracy. Performance is superior to often used approximations to these transformations. Two examples are illustrated: planar panoramic mosaicing, and augmented reality. Both applications run at frame rate on standard desktop machines, such as an SGI Indy or a PC. 1 Introduction The idea explored in this paper is that by judicious modelling, exact transformations arising in computer vision applications can be specified by a reduced number of parameters. The great advantage of this is the reduction in the search space that must be explored when estimating these trans... Manish Jethwa, Andrew Zisserman, Andrew W. Fitzgibbon |
BMVC | 3 |
| 1998 | Automatic Camera Recovery for Closed or Open Image Sequences
Andrew W. Fitzgibbon, Andrew Zisserman |
ECCV (1) | 1 |
| 1998 | Maintaining Multiple Motion Model Hypotheses Through Many Views to Recover Matching and StructureabstractIn order to recover structure from images it is desirable to use many views to obtain the best possible estimates. However, whilst recovering projective structure and motion from such extended sequences problems arise that are not apparent from a general view-point/structure approach. Foremost amongst these are (a) maintaining image correspondences consistently through many images, and (b) identifying images, within the sequence, for which structure cannot be reliably recovery. Within this paper the use of multiple motion model hypotheses is explored as an aid to solve both of these problems. Philip Torr 0001, Andrew W. Fitzgibbon, Andrew Zisserman |
ICCV | 2 |
| 1998 | Simultaneous Registration of Multiple Range Views for Use in Reverse Engineering of CAD Models
David W. Eggert, Andrew W. Fitzgibbon, Robert B. Fisher |
Comput. Vis. Image Underst. | 2 |
| 1997 | High-level model acquisition from range images
Andrew W. Fitzgibbon, David W. Eggert, Robert B. Fisher |
Comput. Aided Des. | 1 |
| 1996 | Training PDMs on Models: The Case of Deformable SuperellipsesabstractThis paper addresses the following problem: How can we make a complicated mathematical shape model simpler while keeping a comparable level of representational power? The proposed solution is to use the original model itself -- which represents a class of shapes -- to train a Point Distribution Model. In this paper the idea is applied to the case of deformable superellipses. Maurizio Pilu, Andrew W. Fitzgibbon, Robert B. Fisher |
BMVC | 2 |
| 1996 | Ellipse-specific direct least-square fittingabstractEllipse fitting is one of the classic problems of pattern recognition and has been subject to considerable attention because of its many applications. This article presents the first direct method for specifically fitting ellipses in the least squares sense. Previous approaches used either generic conic fitting or relied on iterative methods to recover elliptic solutions. The proposed method is (i) ellipse-specific, (ii) directly solved by a generalised eigen-system, (iii) has a desirable low-eccentricity bias, and (iv) is robust to noise. We provide a theoretical demonstration, several examples and the Matlab coding of the algorithm. Maurizio Pilu, Andrew W. Fitzgibbon, Robert B. Fisher |
ICIP (3) | 2 |
| 1996 | Simultaneous registration of multiple range views for use in reverse engineeringabstractWhen reverse engineering a CAD model, it is necessary to integrate information from several views of an object into a common reference frame. Given a rough initial alignment, further pose refinement here uses an improved version of the interactive closes point algorithm. Incremental adjustments are computed simultaneously for all data sets, resulting in a more globally optimal set of transformations. Also, thresholds for removing outlier correspondences are not needed, as the merging data sets are considered as a whole. Motion updates are computed through force-based optimization, using implied springs between data sets. Experiments indicate that even for very rough initial positionings, registration accuracy approaches 25% of the interpoint sampling resolution of the images. David W. Eggert, Andrew W. Fitzgibbon, Robert B. Fisher |
ICPR | 2 |
| 1996 | Direct least squares fitting of ellipsesabstractThis paper presents a new efficient method for fitting ellipses to scattered data. Previous algorithms either fitted general conics or were computationally expensive. By minimizing the algebraic distance subject to the constraint 4ac-b/sup 2/=1 the new method incorporates the ellipticity constraint into the normalization factor. The new method combines several advantages: 1) it is ellipse-specific so that even bad data will always return an ellipse; 2) it can be solved naturally by a generalized eigensystem, and 3) it is extremely robust, efficient and easy to implement. We compare the proposed method to other approaches and show its robustness on several examples in which other nonellipse-specific approaches would fail or require computationally expensive iterative refinements. Andrew W. Fitzgibbon, Maurizio Pilu, Robert B. Fisher |
ICPR | 1 |
| 1996 | Convex hulls, occluding contours, aspect graphs and the Hough transform
Mark W. Wright, Andrew W. Fitzgibbon, Peter J. Giblin, Robert B. Fisher |
Image Vis. Comput. | 2 |
| 1996 | An Experimental Comparison of Range Image Segmentation AlgorithmsabstractA methodology for evaluating range image segmentation algorithms is proposed. This methodology involves (1) a common set of 40 laser range finder images and 40 structured light scanner images that have manually specified ground truth and (2) a set of defined performance metrics for instances of correctly segmented, missed, and noise regions, over- and under-segmentation, and accuracy of the recovered geometry. A tool is used to objectively compare a machine generated segmentation against the specified ground truth. Four research groups have contributed to evaluate their own algorithm for segmenting a range image into planar patches. Adam W. Hoover, Gillian Jean-Baptiste, Xiaoyi Jiang 0001, Patrick J. Flynn, Horst Bunke, Dmitry B. Goldgof, Kevin W. Bowyer, David W. Eggert, Andrew W. Fitzgibbon, Robert B. Fisher |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 1995 | A Buyer's Guide to Conic FittingabstractIn this paper we evaluate several methods of tting data to conic sec-tions. Conic tting is a commonly required task in machine vision, but many algorithms perform badly on incomplete or noisy data. We evalu-ate several algorithms under various noise and degeneracy conditions, identify the key parameters which aect sensitivity, and present the results of comparative experiments which emphasize the algorithms' behaviours under common examples of degenerate data. In addition, complexity analyses in terms of \nop counts are provided in order to further inform the choice of algorithm for a specic application. 1 Andrew W. Fitzgibbon, Robert B. Fisher |
BMVC | 1 |
| 1995 | Convex Hulls, Occluding Contours, Aspect Graphs and the Hough TransformabstractThe Hough transform is a standard technique for finding features such as lines in images.Typically edgels or other features are mapped into a partitioned parameter or Hough space as individual votes.The target image features are detected as peaks in the Hough space.In this paper we consider not just the peaks but the mapping of the entire shape boundary from image space to the Hough parameter space.We analyse this mapping and illustrate correspondences between features in Hough space and image space.Using this knowledge we present an algorithm to construct convex hulls of arbitrary 2D shapes with smooth and polygonal boundaries as well as isolated point sets.We also demonstrate its extension to the 3D case.We then show how this mapping changes as we move the origin in image space.The origin can be considered as a vantage point from which to view the object and the occluding contour can be extracted easily from Hough space as those points where R = 0. We demonstrate the potential for tracking of transitions in the mapping to be used to construct an aspect graph of arbitrary 2D and 3D shapes. Mark W. Wright, Andrew W. Fitzgibbon, Robert B. Fisher |
BMVC | 2 |
| 1994 | Direct Calibraction and Data Consistency in 3-D Laser ScanningabstractThis paper addresses two aspects of triangulation-based range sensors using structured laser light: calibration and measurements consistency. We present a direct calibration technique which does not require modelling any specific sensor component or phenomena, therefore is not limited in accuracy by the inability to model error sources. We also sketch some consistency tests based on two-camera geometry which make it possible to acquire satisfactory range images of highly reflective surfaces with holes. Experimental results indicating the validity of the methods are reported. Emanuele Trucco, Robert B. Fisher, Andrew W. Fitzgibbon |
BMVC | 3 |
| 1994 | Lack-of-fit Detection using the Run-distribution Test
Andrew W. Fitzgibbon, Robert B. Fisher |
ECCV (2) | 1 |
| 1993 | Visually Salient 3D Model Acquisition from Range DataabstractAutomatic model building is a crucial requirement of any model-based vision system which must work in unknown environments. Even in the specialised environments where CAD models are available for the small number of parts, these models often lack visual saliency, impacting on the robustness (more than the mean accuracy) of the system using them. Moreover, the time needed to make such models by hand may be prohibitive, particularly with freeform curved objects; hence automatic model acquisition is greatly desirable. We present a procedure which merges multiple range images of an unmodelled object to create a 3-D body-centred model of the part. By explicitly considering the specific problems of vision systems, we achieve a high-level, robust and accurate description of the unknown scene, where visual applicability of our generated model is guaranteed. The models are characterized by a pleasant `intuitive' feel which allows easy operator intervention if they must be altered, perhaps to g... Edvaldo M. Bispo, Andrew W. Fitzgibbon, Robert B. Fisher |
BMVC | 2 |
| 1993 | Invariant Fitting of Arbitrary Single-Extremum SurfacesabstractBesl and Jain's variable order surface fitting algorithm [1] is a useful method of constructing a noise-free reconstruction of 2jD range images with a small number of primitive regions. The use of bivariate polynomials as the approximation basis functions is linear, fast and easy to render robust. Seeding fits from regions classified by differential geometry is an important step towards a viewpoint invariant segmentation. However, in order to better approximate arbitrarily shaped surfaces, polynomials of high degree are needed. For a region-growing paradigm, the poor extrapolation power of high order polynomials slows convergence and generates "non-intuitive " segmentations when crossing curvature discontinuities. Such segmentations are difficult to match against traditional CAD-like models. Further, the instability of the segmentation makes invocation of the correct model from a large database extremely difficult. We show that these algorithms must of necessity trade representational richness for repeatability. In this paper we describe a new method of satisfying the requirement for high representational richness while retaining the ease of manipulation and recognition of single-extremum surface patches. By introducing a canonical reparameterised coordinate system, biquadratic patches can be made to approximate arbitrary single-extremum shapes in a viewpoint invariant manner. An iterative fitting algorithm is presented, which quickly converges to the appropriate description. Examples of the abilities of the new approach are supplied, and compared with alternative strategies. 1 Andrew W. Fitzgibbon, Robert B. Fisher |
BMVC | 1 |
| 1992 | Practical Aspect-graph Derivation Incorporating Feature Segmentation Performance
Andrew W. Fitzgibbon, Robert B. Fisher |
BMVC | 1 |