Matteo Toso

dblp:225/4596 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-8990-7156ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021
YearPublicationVenuePosition
2026 SplatFill: 3D Scene Inpainting via Depth-Guided Gaussian Splatting
Mahtab Dahaghin, Milind Gajanan Padalkar, Matteo Toso, Alessio Del Bue, Vittorio Murino
ICPR (16)3
2025 Maps from Motion (MfM): Generating 2D Semantic Maps from Sparse Multi-View Images
abstract
World-wide detailed 2D maps require enormous collective efforts. OpenStreetMap is the result of 11 million registered users manually annotating the GPS location of over 1.75 billion entries, including distinctive landmarks and common urban objects. At the same time, manual annotations can include errors and are slow to update, limiting the map's accuracy. Mapsfrom Motion (MfM) is a step for-ward to automatize such time-consuming map making procedure by computing 2D maps of semantic objects directly from a collection of uncalibrated multi-view images. From each image, we extract a set of object detections, and estimate their spatial arrangement in a top-down local map centered in the reference frame of the camera that captured the image. Aligning these local maps is not a trivial problem, since they provide incomplete, noisy fragments of the scene, and matching detections across them is unreliable because of the presence of repeated pattern and the limited appearance variability of urban objects. We address this with a novel graph-based framework, that encodes the spatial and semantic distribution of the objects detected in each image, and learns how to combine them to predict the objects' poses in a global reference system, while taking into account all possible detection matches and preserving the topology observed in each image. Despite the complexity of the problem, our best model achieves global2D registration with an average accuracy within 4 meters (i.e. below GPS accuracy) even on sparse sequences with strong view-point change, on which COLMAP has an 80% failure rate. We provide extensive evaluation on synthetic and real-world data, showing how the method obtains a solution even in scenarios where standard optimization techniques fail. Find more information at matteot90.github.io/MapsFromMotion.
Matteo Toso, Stefano Fiorini, Stuart Jamea, Alessio Del Bue
3DV1
2024 PRAGO: Differentiable Multi-View Pose Optimization From Objectness Detections
abstract
Robustly estimating camera poses from a set of images is a fundamental task which remains challenging for differentiable methods, especially in the case of small and sparse camera pose graphs. To overcome this challenge, we propose Pose-refined Rotation Averaging Graph Optimization (PRAGO). From a set of objectness detections on unordered images, our method reconstructs the rotational pose, and in turn, the absolute pose, in a differentiable manner benefiting from the optimization of a sequence of geometrical tasks. We show how our objectness pose-refinement module in PRAGO is able to refine the inherent ambiguities in pairwise relative pose estimation without removing edges and avoiding making early decisions on the viability of graph edges. PRAGO then refines the absolute rotations through iterative graph construction, reweighting the graph edges to compute the final rotational pose, which can be converted into absolute poses using translation averaging. We show that PRAGO is able to outperform non-differentiable solvers on small and sparse scenes extracted from 7-Scenes achieving a relative improvement of 21% for rotations while achieving similar translation estimates.
Matteo Taiana, Matteo Toso, Stuart James, Alessio Del Bue
3DV2
2024 Contrastive Gaussian Clustering for Weakly Supervised 3D Scene Segmentation
abstract
Abstract 3D scene segmentation is a crucial task in Computer Vision, with applications in autonomous driving, augmented reality, and robotics. Traditional methods often struggle to provide consistent and accurate segmentation across different viewpoints. To address this, we look at the growing field of novel view synthesis. Methods like NeRF and 3DGS take a set of images and implicitly learn a multi-view consistent representation of the geometry of the scene; the same strategy can be extended to learn a 3D segmentation of the scene that is consistent with the 2D segmentation of an initial training set of input images. We introduce Contrastive Gaussian Clustering, a novel approach for novel segmentation view synthesis and 3D scene segmentation. We extend 3D Gaussian Splatting to include a learnable 3D feature field, which allows us to cluster the 3D Gaussians into objects. Using a combination of contrastive learning and spatial regularization, our model can be trained on inconsistent 2D segmentation labels, and still learn to generate multi-view consistent masks. Moreover, the resulting model is extremely accurate, improving the IoU accuracy of the predicted masks by $$+8\%$$ + 8 % over the state of the art. Code and trained models are available at https://github.com/MyrnaCCS/contrastive-gaussian-clustering .
Myrna C. Silva, Mahtab Dahaghin, Matteo Toso, Alessio Del Bue
ICPR (23)3
2022 PoserNet: Refining Relative Camera Poses Exploiting Object Detections
Matteo Taiana, Matteo Toso, Stuart James, Alessio Del Bue
ECCV (33)2
2019 Fixing Implicit Derivatives: Trust-Region Based Learning of Continuous Energy Functions
abstract
We present a new technique for the learning of continuous energy functions that we refer to as Wibergian Learning. One common approach to inverse problems is to cast them as an energy minimisation problem, where the minimum cost solution found is used as an estimator of hidden parameters. Our new approach formally characterises the dependency between weights that control the shape of the energy function, and the location of minima, by describing minima as fixed points of optimisation methods. This allows for the use of gradient-based end-to- end training to integrate deep-learning and the classical inverse problem methods. We show how our approach can be applied to obtain state-of-the-art results in the diverse applications of tracker fusion and multiview 3D reconstruction.
Chris Russell 0001, Matteo Toso, Neill D. F. Campbell
NeurIPS2
2018 Rethinking Pose in 3D: Multi-stage Refinement and Recovery for Markerless Motion Capture
abstract
We propose a CNN-based approach for multi-camera markerless motion capture of the human body. Unlike existing methods that first perform pose estimation on individual cameras and generate 3D models as post-processing, our approach makes use of 3D reasoning throughout a multi-stage approach. This novelty allows us to use provisional 3D models of human pose to rethink where the joints should be located in the image and to recover from past mistakes. Our principled refinement of 3D human poses lets us make use of image cues, even from images where we previously misdetected joints, to refine our estimates as part of an end-to-end approach. Finally, we demonstrate how the high-quality output of our multi-camera setup can be used as an additional training source to improve the accuracy of existing single camera models.
Denis Tomè, Matteo Toso, Lourdes Agapito, Chris Russell 0001
3DV2