EDBT 2026 Demo / reviewers in the wild / expert
Thomas J. Cashman 0001
dblp:35/5640 · also Thomas Joseph Cashman 0001, Tom Cashman 0001
· DBLP profile ↗
21ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0001-7975-8567ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VoluMe - Authentic 3D Video Calls from Live Gaussian Splat Prediction
Martin de La Gorce, Charlie Hewitt, Robert Gerdisch, Zafiirah Hosenie, Givi Meishvili, Marek Kowalski, Thomas J. Cashman 0001, Antonio Criminisi |
ICCV | 8 |
| 2025 | DAViD: Data-Efficient and Accurate Vision Models from Synthetic Data DAViD also references Michelangelo's David - an iconic symbol of anatomical precision-and the David vs. Goliath story, reflecting our small yet powerful dataset and models
Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Charlie Hewitt, Lohit Petikam, Xian Xiao, Antonio Criminisi, Thomas J. Cashman 0001, Tadas Baltrusaitis |
ICCV | 7 |
| 2024 | Look Ma, no markers: holistic performance capture without the hassleabstractWe tackle the problem of highly-accurate, holistic performance capture for the face, body and hands simultaneously. Motion-capture technologies used in film and game production typically focus only on face, body or hand capture independently, involve complex and expensive hardware and a high degree of manual intervention from skilled operators. While machine-learning-based approaches exist to overcome these problems, they usually only support a single camera, often operate on a single part of the body, do not produce precise world-space results, and rarely generalize outside specific contexts. In this work, we introduce the first technique for markerfree, high-quality reconstruction of the complete human body, including eyes and tongue, without requiring any calibration, manual intervention or custom hardware. Our approach produces stable world-space results from arbitrary camera rigs as well as supporting varied capture environments and clothing. We achieve this through a hybrid approach that leverages machine learning models trained exclusively on synthetic data and powerful parametric models of human shape and motion. We evaluate our method on a number of body, face and hand reconstruction benchmarks and demonstrate state-of-the-art results that generalize on diverse datasets. Charlie Hewitt, Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Lohit Petikam, Shideh Rezaeifar, Louis Florentin, Zafiirah Hosenie, Thomas J. Cashman 0001, Julien Valentin, Darren Cosker, Tadas Baltrusaitis |
ACM Trans. Graph. | 8 |
| 2022 | FLAG: Flow-based 3D Avatar Generation from Sparse ObservationsabstractTo represent people in mixed reality applications for collaboration and communication, we need to generate realistic and faithful avatar poses. However, the signal streams that can be applied for this task from head-mounted devices (HMDs) are typically limited to head pose and hand pose estimates. While these signals are valuable, they are an incomplete representation of the human body, making it challenging to generate a faithful full-body avatar. We address this challenge by developing a flow-based generative model of the 3D human body from sparse observations, wherein we learn not only a conditional distribution of 3D human pose, but also a probabilistic mapping from observations to the latent space from which we can generate a plausible pose along with uncertainty estimates for the joints. We show that our approach is not only a strong predictive model, but can also act as an efficient pose prior in different optimization settings where a good initial latent code plays a major role. Mohammad Sadegh Ali Akbarian, Pashmina Cameron, Federica Bogo, Andrew W. Fitzgibbon, Thomas J. Cashman 0001 |
CVPR | 5 |
| 2022 | 3D Face Reconstruction with Dense Landmarks
Erroll Wood, Tadas Baltrusaitis, Charlie Hewitt, Matthew Johnson 0003, Jingjing Shen, Nikola Milosavljevic, Daniel Wilde, Stephan J. Garbin, Toby Sharp, Ivan Stojiljkovic, Thomas J. Cashman 0001, Julien P. C. Valentin |
ECCV (13) | 11 |
| 2021 | Full-Body Motion from a Single Head-Mounted Device: Generating SMPL Poses from Partial ObservationsabstractThe increased availability and maturity of head-mounted and wearable devices opens up opportunities for remote communication and collaboration. However, the signal streams provided by these devices (e.g., head pose, hand pose, and gaze direction) do not represent a whole person. One of the main open problems is therefore how to leverage these signals to build faithful representations of the user. In this paper, we propose a method based on variational autoencoders to generate articulated poses of a human skeleton based on noisy streams of head and hand pose. Our approach relies on a model of pose likelihood that is novel and theoretically well-grounded. We demonstrate on publicly available datasets that our method is effective even from very impoverished signals and investigate how pose prediction can be made more accurate and realistic. Andrea Dittadi, Sebastian Dziadzio, Darren Cosker, Ben Lundell, Thomas J. Cashman 0001, Jamie Shotton |
ICCV | 5 |
| 2021 | Fake it till you make it: face analysis in the wild using synthetic data aloneabstractWe demonstrate that it is possible to perform face-related computer vision in the wild using synthetic data alone. The community has long enjoyed the benefits of synthesizing training data with graphics, but the domain gap between real and synthetic data has remained a problem, especially for human faces. Researchers have tried to bridge this gap with data mixing, domain adaptation, and domain-adversarial training, but we show that it is possible to synthesize data with minimal domain gap, so that models trained on synthetic data generalize to real in-the-wild datasets. We describe how to combine a procedurally-generated parametric 3D face model with a comprehensive library of hand-crafted assets to render training images with unprecedented realism and diversity. We train machine learning systems for face-related tasks such as landmark localization and face parsing, showing that synthetic data can both match real data in accuracy as well as open up new approaches where manual labeling would be impossible. Erroll Wood, Tadas Baltrusaitis, Charlie Hewitt, Sebastian Dziadzio, Thomas J. Cashman 0001, Jamie Shotton |
ICCV | 5 |
| 2020 | The Phong Surface: Efficient 3D Model Fitting Using Lifted Optimization
Jingjing Shen, Thomas J. Cashman 0001, Qi Ye 0001, Tim Hutton, Toby Sharp, Federica Bogo, Andrew W. Fitzgibbon, Jamie Shotton |
ECCV (1) | 2 |
| 2018 | QRkit: Sparse, Composable QR Decompositions for Efficient and Stable Solutions to Problems in Computer VisionabstractEmbedded computer vision applications increasingly require the speed and power benefits of single-precision (32 bit) floating point. However, applications which make use of Levenberg-like optimization can lose significant accuracy when reducing to single precision, sometimes unrecoverably so. This accuracy can be regained using solvers based on QR rather than Cholesky decomposition, but the absence of sparse QR solvers for common sparsity patterns found in computer vision means that many applications cannot benefit. We introduce an open-source suite of solvers for Eigen, which efficiently compute the QR decomposition for matrices with some common sparsity patterns (block diagonal, horizontal and vertical concatenation, and banded). For problems with very particular sparsity structures, these elements can be composed together in 'kit' form, hence the name QRkit. We apply our methods to several computer vision problems, showing competitive performance and suitability especially in single precision arithmetic. Jan Svoboda, Thomas J. Cashman 0001, Andrew W. Fitzgibbon |
WACV | 2 |
| 2017 | An Efficient Background Term for 3D Reconstruction and Tracking with Smooth Surface ModelsabstractWe present a novel strategy to shrink and constrain a 3D model, represented as a smooth spline-like surface, within the visual hull of an object observed from one or multiple views. This new background or silhouette term combines the efficiency of previous approaches based on an image-plane distance transform with the accuracy of formulations based on raycasting or ray potentials. The overall formulation is solved by alternating an inner nonlinear minization (raycasting) with a joint optimization of the surface geometry, the camera poses and the data correspondences. Experiments on 3D reconstruction and object tracking show that the new formulation corrects several deficiencies of existing approaches, for instance when modelling non-convex shapes. Moreover, our proposal is more robust against defects in the object segmentation and inherently handles the presence of uncertainty in the measurements (e.g. null depth values in images provided by RGB-D cameras). Mariano Jaimez, Thomas J. Cashman 0001, Andrew W. Fitzgibbon, Javier González 0001, Daniel Cremers |
CVPR | 2 |
| 2016 | Fits Like a Glove: Rapid and Reliable Hand Shape PersonalizationabstractWe present a fast, practical method for personalizing a hand shape basis to an individual user's detailed hand shape using only a small set of depth images. To achieve this, we minimize an energy based on a sum of render-and-compare cost functions called the golden energy. However, this energy is only piecewise continuous, due to pixels crossing occlusion boundaries, and is therefore not obviously amenable to efficient gradient-based optimization. A key insight is that the energy is the combination of a smooth low-frequency function with a high-frequency, low-amplitude, piecewisecontinuous function. A central finite difference approximation with a suitable step size can therefore jump over the discontinuities to obtain a good approximation to the energy's low-frequency behavior, allowing efficient gradient-based optimization. Experimental results quantitatively demonstrate for the first time that detailed personalized models improve the accuracy of hand tracking and achieve competitive results in both tracking and model registration. David Joseph Tan, Thomas J. Cashman 0001, Jonathan Taylor 0001, Andrew W. Fitzgibbon, Daniel Tarlow, Sameh Khamis, Shahram Izadi, Jamie Shotton |
CVPR | 2 |
| 2016 | Efficient and precise interactive hand tracking through joint, continuous optimization of pose and correspondencesabstractFully articulated hand tracking promises to enable fundamentally new interactions with virtual and augmented worlds, but the limited accuracy and efficiency of current systems has prevented widespread adoption. Today's dominant paradigm uses machine learning for initialization and recovery followed by iterative model-fitting optimization to achieve a detailed pose fit. We follow this paradigm, but make several changes to the model-fitting, namely using: (1) a more discriminative objective function; (2) a smooth-surface model that provides gradients for non-linear optimization; and (3) joint optimization over both the model pose and the correspondences between observed data points and the model surface. While each of these changes may actually increase the cost per fitting iteration, we find a compensating decrease in the number of iterations. Further, the wide basin of convergence means that fewer starting points are needed for successful model fitting. Our system runs in real-time on CPU only, which frees up the commonly over-burdened GPU for experience designers. The hand tracker is efficient enough to run on low-power devices such as tablets. We can track up to several meters from the camera to provide a large working volume for interaction, even using the noisy data from current-generation depth cameras. Quantitative assessments on standard datasets show that the new approach exceeds the state of the art in accuracy. Qualitative results take the form of live recordings of a range of interactive experiences enabled by this new approach. Jonathan Taylor 0001, Lucas Bordeaux, Thomas J. Cashman 0001, Bob Corish, Cem Keskin, Toby Sharp, Eduardo Soto, David Sweeney, Julien P. C. Valentin, Benjamin Luff, Arran Topalian, Erroll Wood, Sameh Khamis, Pushmeet Kohli, Shahram Izadi, Richard Banks, Andrew W. Fitzgibbon, Jamie Shotton |
ACM Trans. Graph. | 3 |
| 2015 | Watertight conversion of trimmed CAD surfaces to Clough-Tocher splinesabstractThe boundary representations (B-reps) that are used to represent shape in Computer-Aided Design systems create unavoidable gaps at the face boundaries of a model. Although these inconsistencies can be kept below the scale that is important for visualisation and manufacture, they cause problems for many downstream tasks, making it difficult to use CAD models directly for simulation or advanced geometric analysis, for example. Motivated by this need for watertight models, we address the problem of converting B-rep models to a collection of cubic C1 Clough–Tocher splines. These splines allow a watertight join between B-rep faces, provide a homogeneous representation of shape, and also support local adaptivity. We perform a comparative study of the most prominent Clough–Tocher constructions and include some novel variants. Our criteria include visual fairness, invariance to affine reparameterisations, polynomial precision and approximation error. The constructions are tested on both synthetic data and CAD models that have been triangulated. Our results show that no construction is optimal in every scenario, with surface quality depending heavily on the triangulation and parameterisation that are used. Jirí Kosinka, Thomas J. Cashman 0001 |
Comput. Aided Geom. Des. | 2 |
| 2013 | Generalized Lane-Riesenfeld algorithms
Thomas J. Cashman 0001, Kai Hormann, Ulrich Reif |
Comput. Aided Geom. Des. | 1 |
| 2013 | Efficient Interpolation of Articulated Shapes Using Mixed Shape SpacesabstractAbstract Interpolation between compatible triangle meshes that represent different poses of some object is a fundamental operation in geometry processing. A common approach is to consider the static input shapes as points in a suitable shape space and then use simple linear interpolation in this space to find an interpolated shape. In this paper, we present a new interpolation technique that is particularly tailored for meshes that represent articulated shapes. It is up to an order of magnitude faster than state‐of‐the‐art methods and gives very similar results. To achieve this, our approach introduces a novel shape space that takes advantage of the underlying structure of articulated shapes and distinguishes between rigid parts and non‐rigid joints. This allows us to use fast vertex interpolation on the rigid parts and resort to comparatively slow edge‐based interpolation only for the joints. Stefano Marras, Thomas J. Cashman 0001, Kai Hormann |
Comput. Graph. Forum | 2 |
| 2013 | What Shape Are Dolphins? Building 3D Morphable Models from 2D Imagesabstract3D morphable models are low-dimensional parameterizations of 3D object classes which provide a powerful means of associating 3D geometry to 2D images. However, morphable models are currently generated from 3D scans, so for general object classes such as animals they are economically and practically infeasible. We show that, given a small amount of user interaction (little more than that required to build a conventional morphable model), there is enough information in a collection of 2D pictures of certain object classes to generate a full 3D morphable model, even in the absence of surface texture. The key restriction is that the object class should not be strongly articulated, and that a very rough rigid model should be provided as an initial estimate of the “mean shape.” The model representation is a linear combination of subdivision surfaces, which we fit to image silhouettes and any identifiable key points using a novel combined continuous-discrete optimization strategy. Results are demonstrated on several natural object classes, and show that models of rather high quality can be obtained from this limited information. Thomas J. Cashman 0001, Andrew W. Fitzgibbon |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Beyond Catmull-Clark? A Survey of Advances in Subdivision Surface MethodsabstractAbstract Subdivision surfaces allow smooth free‐form surface modelling without topological constraints. They have become a fundamental representation for smooth geometry, particularly in the animation and entertainment industries. This survey summarizes research on subdivision surfaces over the last 15 years in three major strands: analysis, integration into existing systems and the development of new schemes. We also examine the reason for the low adoption of new schemes with theoretical advantages, explain why Catmull–Clark surfaces have become a de facto standard in geometric modelling, and conclude by identifying directions for future research. Thomas J. Cashman 0001 |
Comput. Graph. Forum | 1 |
| 2012 | A continuous, editable representation for deforming mesh sequences with separate signals for time, pose and shapeabstractAbstract It is increasingly popular to represent non‐rigid motion using a deforming mesh sequence: a discrete sequence of frames, each of which is given as a mesh with a common graph structure. Such sequences have the flexibility to represent a wide range of mesh deformations used in practice, but they are also highly redundant, expensive to store, and difficult to edit in a time‐coherent manner. We address these limitations with a continuous representation that extracts redundancy in three separate phases, leading to separate editable signals in time, pose and shape. The representation can be applied to any deforming mesh sequence, in contrast to previous domain‐specific approaches. By modifying the three signal components, we demonstrate time‐coherent editing operations such as local repetition of part of a sequence, frame rate conversion and deformation transfer. We also show that our representation makes it possible to design new deforming sequences simply by sketching a curve in a 2D pose space. Thomas J. Cashman 0001, Kai Hormann |
Comput. Graph. Forum | 1 |
| 2009 | A symmetric, non-uniform, refine and smooth subdivision algorithm for general degree B-splines
Thomas J. Cashman 0001, Neil A. Dodgson, Malcolm A. Sabin |
Comput. Aided Geom. Des. | 1 |
| 2009 | Selective knot insertion for symmetric, non-uniform refine and smooth B-spline subdivision
Thomas J. Cashman 0001, Neil A. Dodgson, Malcolm A. Sabin |
Comput. Aided Geom. Des. | 1 |
| 2009 | NURBS with extraordinary points: high-degree, non-uniform, rational subdivision schemesabstractWe present a subdivision framework that adds extraordinary vertices to NURBS of arbitrarily high degree. The surfaces can represent any odd degree NURBS patch exactly. Our rules handle non-uniform knot vectors, and are not restricted to midpoint knot insertion. In the absence of multiple knots at extraordinary points, the limit surfaces have bounded curvature. Thomas J. Cashman 0001, Ursula H. Augsdörfer, Neil A. Dodgson, Malcolm A. Sabin |
ACM Trans. Graph. | 1 |