Tommaso Cavallari

dblp:153/7622 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0003-2490-5341ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 6 since 2021Systems, architecture and hardware · 4 · 1 since 2021
YearPublicationVenuePosition
2025 Scene Coordinate Reconstruction Priors
Wenjing Bian, Axel Barroso-Laguna, Tommaso Cavallari, Victor Adrian Prisacariu, Eric Brachmann
ICCV3
2025 ACE-G: Improving Generalization of Scene Coordinate Regression Through Query Pre-Training
abstract
Scene coordinate regression (SCR) has established itself as a promising learning-based approach to visual relocalization. After mere minutes of scene-specific training, SCR models estimate camera poses of query images with high accuracy. Still, SCR methods fall short of the generalization capabilities of more classical feature-matching approaches. When imaging conditions of query images, such as lighting or viewpoint, are too different from the training views, SCR models fail. Failing to generalize is an inherent limitation of previous SCR frameworks, since their training objective is to encode the training views in the weights of the coordinate regressor itself. The regressor essentially overfits to the training views, by design. We propose to separate the coordinate regressor and the map representation into a generic transformer and a scene-specific map code. This separation allows us to pre-train the transformer on tens of thousands of scenes. More importantly, it allows us to train the transformer to generalize from mapping images to unseen query images during pre-training. We demonstrate on multiple challenging relocalization datasets that our method, ACE-G, leads to significantly increased robustness while keeping the computational footprint attractive
Leonard Bruns, Axel Barroso-Laguna, Tommaso Cavallari, Áron Monszpart, Sowmya Munukutla, Victor Adrian Prisacariu, Eric Brachmann
ICCV3
2024 Map-Relative Pose Regression for Visual Re-Localization
abstract
Pose regression networks predict the camera pose of a query image relative to a known environment. Within this family of methods, absolute pose regression (APR) has recently shown promising accuracy in the range of a few centimeters in position error. APR networks encode the scene geometry implicitly in their weights. To achieve high accuracy, they require vast amounts of training data that, realistically, can only be created using novel view synthesis in a days-long process. This process has to be repeated for each new scene again and again. We present a new approach to pose regression, map-relative pose regression (marepo), that satisfies the data hunger of the pose regression network in a scene-agnostic fashion. We condition the pose regressor on a scene-specific map representation such that its pose predictions are relative to the scene map. This allows us to train the pose regressor across hundreds of scenes to learn the generic relation between a scene-specific map representation and the camera pose. Our map-relative pose regressor can be applied to new map representations immediately or after mere minutes of fine-tuning for the highest accuracy. Our approach outperforms previous pose regression methods by far on two public datasets, indoor and outdoor. Code is available: https://nianticlabs.github.io/marepo.
Tommaso Cavallari, Victor Adrian Prisacariu, Eric Brachmann
CVPR2
2024 Scene Coordinate Reconstruction: Posing of Image Collections via Incremental Learning of a Relocalizer
Eric Brachmann, Jamie Wynn, Tommaso Cavallari, Áron Monszpart, Daniyar Turmukhambetov, Victor Adrian Prisacariu
ECCV (56)4
2023 Accelerated Coordinate Encoding: Learning to Relocalize in Minutes Using RGB and Poses
abstract
Learning-based visual relocalizers exhibit leading pose accuracy, but require hours or days of training. Since training needs to happen on each new scene again, long training times make learning-based relocalization impractical for most applications, despite its promise of high accuracy. In this paper we show how such a system can actually achieve the same accuracy in less than 5 minutes. We start from the obvious: a relocalization network can be split in a scene-agnostic feature backbone, and a scene-specific prediction head. Less obvious: using an MLP prediction head allows us to optimize across thousands of view points simultaneously in each single training iteration. This leads to stable and extremely fast convergence. Furthermore, we substitute effective but slow end-to-end training using a robust pose solver with a curriculum over a reprojection loss. Our approach does not require privileged knowledge, such a depth maps or a 3D model, for speedy training. Overall, our approach is up to 300x faster in mapping than state-of-the-art scene coordinate regression, while keeping accuracy on par. Code is available: https://nianticlabs.github.io/ace
Eric Brachmann, Tommaso Cavallari, Victor Adrian Prisacariu
CVPR2
2023 Through Hawks' Eyes: Synthetically Reconstructing the Visual Field of a Bird in Flight
abstract
Birds of prey rely on vision to execute flight manoeuvres that are key to their survival, such as intercepting fast-moving targets or navigating through clutter. A better understanding of the role played by vision during these manoeuvres is not only relevant within the field of animal behaviour, but could also have applications for autonomous drones. In this paper, we present a novel method that uses computer vision tools to analyse the role of active vision in bird flight, and demonstrate its use to answer behavioural questions. Combining motion capture data from Harris' hawks with a hybrid 3D model of the environment, we render RGB images, semantic maps, depth information and optic flow outputs that characterise the visual experience of the bird in flight. In contrast with previous approaches, our method allows us to consider different camera models and alternative gaze strategies for the purposes of hypothesis testing, allows us to consider visual input over the complete visual field of the bird, and is not limited by the technical specifications and performance of a head-mounted camera light enough to attach to a bird's head in flight. We present pilot data from three sample flights: a pursuit flight, in which a hawk intercepts a moving target, and two obstacle avoidance flights. With this approach, we provide a reproducible method that facilitates the collection of large volumes of data across many individuals, opening up new avenues for data-driven models of animal behaviour. Supplementary Information: The online version contains supplementary material available at 10.1007/s11263-022-01733-2.
Sofía Miñano, Stuart Golodetz, Tommaso Cavallari, Graham K. Taylor
Int. J. Comput. Vis.3
2021 Scalable FPGA Median Filtering via a Directional Median Cascade
abstract
The 2-D median filter, one of the oldest and most well-established image-filtering techniques, still sees widespread use throughout computer vision. Despite its relative algorithmic simplicity, accelerating the 2-D median filter via a hardware implementation becomes increasingly challenging as the window size increases, since the resources required grow quartically with the window size. Previous works, in a non-FPGA context, have shown that separately applying several directional median filters to an image, and then taking the median of their results, yields performance that is competitive with, and in some cases even better than, that of a classic 2-D window median. Inspired by these approaches, we propose a novel way of substituting a 2-D median filter on an FPGA with a sequence of directional median filters, in our case arranged as a pipeline, in the pursuit of an FPGA implementation that achieves better scalability and hardware efficiency without sacrificing accuracy. We empirically show that the combination of three particular directional filters, in any order, achieves this, whilst requiring quadratically fewer resources on the FPGA and allowing for much higher throughput.
Oscar Rahnama, Stuart Golodetz, Tommaso Cavallari, Philip Torr 0001
FCCM3
2020 Beyond Controlled Environments: 3D Camera Re-localization in Changing Indoor Scenes
Johanna Wald, Torsten Sattler, Stuart Golodetz, Tommaso Cavallari, Federico Tombari
ECCV (7)4
2020 Scalable FPGA Median Filtering using Multiple Efficient Passes
abstract
The 2-D median filter, one of the oldest and most well-established image-filtering techniques, still sees widespread use throughout computer vision. Despite its relative algorithmic simplicity, accelerating the 2-D median filter via a hardware implementation becomes increasingly challenging as the window size increases, since the resources required grow quadratically with the window size. Previous works, in a non-FPGA context, have shown that applying a sequence of multiple directional median filters to an image yields results that are competitive with, and in some cases even better than, those of a classic 2-D window median. Inspired by these approaches, we propose a novel way of substituting a 2-D median filter on an FPGA with a sequence of directional median filters, in our case in the pursuit of an FPGA implementation that achieves better scalability and hardware efficiency without sacrificing accuracy. We empirically show that the combination of three particular directional filters, in any order, achieves this, whilst requiring quadratically fewer resources on the FPGA. Our approach allows for much higher throughput and is easier to implement as a pipeline.
Oscar Rahnama, Tommaso Cavallari, Philip Torr 0001, Stuart Golodetz
FPGA2
2020 Real-Time RGB-D Camera Pose Estimation in Novel Scenes Using a Relocalisation Cascade
abstract
Camera pose estimation is an important problem in computer vision, with applications as diverse as simultaneous localisation and mapping, virtual/augmented reality and navigation. Common techniques match the current image against keyframes with known poses coming from a tracker, directly regress the pose, or establish correspondences between keypoints in the current image and points in the scene in order to estimate the pose. In recent years, regression forests have become a popular alternative to establish such correspondences. They achieve accurate results, but have traditionally needed to be trained offline on the target scene, preventing relocalisation in new environments. Recently, we showed how to circumvent this limitation by adapting a pre-trained forest to a new scene on the fly. The adapted forests achieved relocalisation performance that was on par with that of offline forests, and our approach was able to estimate the camera pose in close to real time, which made it desirable for systems that require online relocalisation. In this paper, we present an extension of this work that achieves significantly better relocalisation performance whilst running fully in real time. To achieve this, we make several changes to the original approach: (i) instead of simply accepting the camera pose hypothesis produced by RANSAC without question, we make it possible to score the final few hypotheses it considers using a geometric approach and select the most promising one; (ii) we chain several instantiations of our relocaliser (with different parameter settings) together in a cascade, allowing us to try faster but less accurate relocalisation first, only falling back to slower, more accurate relocalisation as necessary; and (iii) we tune the parameters of our cascade, and the individual relocalisers it contains, to achieve effective overall performance. Taken together, these changes allow us to significantly improve upon the performance our original state-of-the-art method was able to achieve on the well-known 7-Scenes and Stanford 4 Scenes benchmarks. As additional contributions, we present a novel way of visualising the internal behaviour of our forests, and use the insights gleaned from this to show how to entirely circumvent the need to pre-train a forest on a generic scene.
Tommaso Cavallari, Stuart Golodetz, Nicholas A. Lord, Julien P. C. Valentin, Victor Adrian Prisacariu, Luigi Di Stefano, Philip Torr 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Let's Take This Online: Adapting Scene Coordinate Regression Network Predictions for Online RGB-D Camera Relocalisation
abstract
Many applications require a camera to be relocalised online, without expensive offline training on the target scene. Whilst both keyframe and sparse keypoint matching methods can be used online, the former often fail away from the training trajectory, and the latter can struggle in textureless regions. By contrast, scene coordinate regression (SCoRe) methods generalise to novel poses and can leverage dense correspondences to improve robustness, and recent work has shown how to adapt SCoRe forests between scenes, allowing their state-of-the-art performance to be leveraged online. However, because they use features hand-crafted for indoor use, they do not generalise well to harder outdoor scenes. Whilst replacing the forest with a neural network and learning suitable features for outdoor use is possible, the techniques used to adapt forests between scenes are unfortunately harder to transfer to a network context. In this paper, we address this by proposing a novel way of leveraging a network trained on one scene to predict points in another scene. Our approach replaces the appearance clustering performed by the branching structure of a regression forest with a two-step process that first uses the network to predict points in the original scene, and then uses these predicted points to look up clusters of points from the new scene. We show experimentally that our online approach achieves state-of-the-art performance on both the 7-Scenes and Cambridge Landmarks datasets, whilst running in under 300ms, making it highly effective in live scenarios.
Tommaso Cavallari, Luca Bertinetto, Jishnu Mukhoti, Philip Torr 0001, Stuart Golodetz
3DV1
2018 R3SGM: Real-Time Raster-Respecting Semi-Global Matching for Power-Constrained Systems
abstract
Stereo depth estimation is used for many computer vision applications. Though many popular methods strive solely for depth quality, for real-time mobile applications (e.g. prosthetic glasses or micro-UAVs), speed and power efficiency are equally, if not more, important. Many real-world systems rely on Semi-Global Matching (SGM) to achieve a good accuracy vs. speed balance, but power efficiency is hard to achieve with conventional hardware, making the use of embedded devices such as FPGAs attractive for low-power applications. However, the full SGM algorithm is ill-suited to deployment on FPGAs, and so most FPGA variants of it are partial, at the expense of accuracy. In a non-FPGA context, the accuracy of SGM has been improved by More Global Matching (MGM), which also helps tackle the streaking artifacts that afflict SGM. In this paper, we propose a novel, resource-efficient method that is inspired by MGM's techniques for improving depth quality, but which can be implemented to run in real time on a low-power FPGA. Through evaluation on multiple datasets (KITTI and Middlebury), we show that in comparison to other real-time capable stereo approaches, we can achieve a state-of-the-art balance between accuracy, power efficiency and speed, making our approach highly desirable for use in real-time systems with limited power.
Oscar Rahnama, Tommaso Cavallari, Stuart Golodetz, Simon Walker, Philip Torr 0001
FPT2
2018 Collaborative Large-Scale Dense 3D Reconstruction with Online Inter-Agent Pose Optimisation
abstract
Reconstructing dense, volumetric models of real-world 3D scenes is important for many tasks, but capturing large scenes can take significant time, and the risk of transient changes to the scene goes up as the capture time increases. These are good reasons to want instead to capture several smaller sub-scenes that can be joined to make the whole scene. Achieving this has traditionally been difficult: joining sub-scenes that may never have been viewed from the same angle requires a high-quality camera relocaliser that can cope with novel poses, and tracking drift in each sub-scene can prevent them from being joined to make a consistent overall scene. Recent advances, however, have significantly improved our ability to capture medium-sized sub-scenes with little to no tracking drift: real-time globally consistent reconstruction systems can close loops and re-integrate the scene surface on the fly, whilst new visual-inertial odometry approaches can significantly reduce tracking drift during live reconstruction. Moreover, high-quality regression forest-based relocalisers have recently been made more practical by the introduction of a method to allow them to be trained and used online. In this paper, we leverage these advances to present what to our knowledge is the first system to allow multiple users to collaborate interactively to reconstruct dense, voxel-based models of whole buildings using only consumer-grade hardware, a task that has traditionally been both time-consuming and dependent on the availability of specialised hardware. Using our system, an entire house or lab can be reconstructed in under half an hour and at a far lower cost than was previously possible.
Stuart Golodetz, Tommaso Cavallari, Nicholas A. Lord, Victor Adrian Prisacariu, David William Murray 0001, Philip Torr 0001
IEEE Trans. Vis. Comput. Graph.2
2017 Probabilistic Object Reconstruction with Online Global Model Correction
abstract
In recent years, major advances have been made in 3D scene reconstruction. However, much less progress has been made for objects, which can exhibit far fewer unambiguous geometric/texture cues than a full scene, and thus are much harder to track against. In this work we present a novel probabilistic object reconstruction framework that simultaneously allows for online, implicit deformation of the objects surface to reduce tracking drift and handle loop closure events. Coupled with our probabilistic formulation is the use of a multi subsegment representation of the object, used to enforce global consistency, with segmentation of the object built in to the formulation. Finally, we employ a CRF framework to refine the overall segmentation, defined by a probability field over the object. We present compelling results over the current state-of-the-art object reconstruction work and demonstrate robustness and consistency w.r.t. established dense SLAM frameworks.
Jack Hunt, Victor Adrian Prisacariu, Stuart Golodetz, Tommaso Cavallari, Nicholas A. Lord, Philip Torr 0001
3DV4
2017 On-the-Fly Adaptation of Regression Forests for Online Camera Relocalisation
abstract
Camera relocalisation is an important problem in computer vision, with applications in simultaneous localisation and mapping, virtual/augmented reality and navigation. Common techniques either match the current image against keyframes with known poses coming from a tracker, or establish 2D-to-3D correspondences between keypoints in the current image and points in the scene in order to estimate the camera pose. Recently, regression forests have become a popular alternative to establish such correspondences. They achieve accurate results, but must be trained offline on the target scene, preventing relocalisation in new environments. In this paper, we show how to circumvent this limitation by adapting a pre-trained forest to a new scene on the fly. Our adapted forests achieve relocalisation performance that is on par with that of offline forests, and our approach runs in under 150ms, making it desirable for real-time systems that require online relocalisation.
Tommaso Cavallari, Stuart Golodetz, Nicholas A. Lord, Julien P. C. Valentin, Luigi Di Stefano, Philip Torr 0001
CVPR1
2015 Volume-Based Semantic Labeling with Signed Distance Functions
Tommaso Cavallari, Luigi Di Stefano
PSIVT1
2014 Automatic detection of pole-like structures in 3D urban environments
abstract
This work aims at automatic detection of man-made pole-like structures in scans of urban environments acquired by a 3D sensor mounted on top a moving vehicle. Pole-like structures, such as e.g. road signs and streetlights, are widespread in these environments, and their reliable detection is relevant to applications dealing with autonomous navigation, facility damage detection, city planning and maintenance. Yet, due to the characteristic thin shape, detection of man-made pole-like structures is significantly prone to both noise as well as occlusions and clutter, the latter being pervasive nuisances when scanning urban environments. Our approach is based on a “local” stage, whereby local features are classified and clustered together, followed by a “global” stage aimed at further classification of candidate entities. The proposed pipeline turns out effective in experiments on a standard publicly available dataset as well as on a challenging dataset acquired during the project for validation purposes.
Federico Tombari, Nicola Fioraio, Tommaso Cavallari, Samuele Salti, Alioscia Petrelli, Luigi Di Stefano
IROS3