Silvano Galliani

dblp:47/1195 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-8575-2441ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2025 R-SCoRe: Revisiting Scene Coordinate Regression for Robust Large-Scale Visual Localization
abstract
Learning-based visual localization methods that use scene coordinate regression (SCR) offer the advantage of smaller map sizes. However, on datasets with complex illumination changes or image-level ambiguities, it remains a less robust alternative to feature matching methods. This work aims to close the gap. We introduce a covisibility graph-based global encoding learning and data augmentation strategy, along with a depth-adjusted reprojection loss to facilitate implicit triangulation. Additionally, we revisit the network architecture and local feature extraction module. Our method achieves state-of-the-art on challenging large-scale datasets without relying on network ensembles or 3D supervision. On Aachen Day-Night, we are 10× more accurate than previous SCR methods with similar map sizes and require at least 5× smaller map sizes than any other SCR method while still delivering superior accuracy. Code is available at: https://github.com/cvg/scrstudio.
Fangjinhua Wang, Silvano Galliani, Christoph Vogel, Marc Pollefeys
CVPR3
2024 GLACE: Global Local Accelerated Coordinate Encoding
abstract
Scene coordinate regression (SCR) methods are a family of visual localization methods that directly regress 2D-3D matches for camera pose estimation. They are effective in small-scale scenes but face significant challenges in large-scale scenes that are further amplified in the absence of ground truth 3D point clouds for supervision. Here, the model can only rely on reprojection constraints and needs to implicitly triangulate the points. The challenges stem from a fundamental dilemma: The network has to be invariant to observations of the same landmark at different viewpoints and lighting conditions, etc., but at the same time discriminate unrelated but similar observations. The latter becomes more relevant and severe in larger scenes. In this work, we tackle this problem by introducing the concept of co-visibility to the network. We propose GLACE, which integrates pre-trained global and local encodings and enables SCR to scale to large scenes with only a single small-sized network. Specifically, we propose a novel feature diffusion technique that implicitly groups the reprojection constraints with co-visibility and avoids overfitting to trivial solutions. Additionally, our position decoder parameterizes the output positions for large-scale scenes more effectively. Without using 3D models or depth maps for supervision, our method achieves state-of-the-art results on large-scale scenes with a low-map-size model. On Cambridge landmarks, with a single model, we achieve 17% lower median position error than Poker, the ensemble variant of the state-of-the-art SCR method ACE. Code is avail-able at: https://github.com/cvg/glace.
Fangjinhua Wang, Silvano Galliani, Christoph Vogel, Marc Pollefeys
CVPR3
2022 IterMVS: Iterative Probability Estimation for Efficient Multi-View Stereo
abstract
We present IterMVS, a new data-driven method for high-resolution multi-view stereo. We propose a novel GRU-based estimator that encodes pixel-wise probability distributions of depth in its hidden state. Ingesting multi-scale matching information, our model refines these distributions over multiple iterations and infers depth and confidence. To extract the depth maps, we combine traditional classification and regression in a novel manner. We verify the efficiency and effectiveness of our method on DTU, Tanks&Temples and ETH3D. While being the most efficient method in both memory and run-time, our model achieves competitive performance on DTU and better generalization ability on Tanks&Temples as well as ETH3D than most state-of-the-art methods. Code is available at https://github.com/FangjinhuaWang/IterMVS.
Fangjinhua Wang, Silvano Galliani, Christoph Vogel, Marc Pollefeys
CVPR2
2021 DeepVideoMVS: Multi-View Stereo on Video With Recurrent Spatio-Temporal Fusion
abstract
We propose an online multi-view depth prediction approach on posed video streams, where the scene geometry information computed in the previous time steps is propagated to the current time step in an efficient and geometrically plausible way. The backbone of our approach is a real-time capable, lightweight encoder-decoder that relies on cost volumes computed from pairs of images. We extend it by placing a ConvLSTM cell at the bottleneck layer, which compresses an arbitrary amount of past information in its states. The novelty lies in propagating the hidden state of the cell by accounting for the viewpoint changes between time steps. At a given time step, we warp the previous hidden state into the current camera plane using the previous depth prediction. Our extension brings only a small overhead of computation time and memory consumption, while improving the depth predictions significantly. As a result, we outperform the existing state-of-the-art multi-view stereo methods on most of the evaluated metrics in hundreds of indoor scenes while maintaining a real-time performance. Code available: https://github.com/ardaduz/deep-video-mvs
Arda Düzçeker, Silvano Galliani, Christoph Vogel, Pablo Speciale, Mihai Dusmanu, Marc Pollefeys
CVPR2
2021 PatchmatchNet: Learned Multi-View Patchmatch Stereo
abstract
We present PatchmatchNet, a novel and learnable cascade formulation of Patchmatch for high-resolution multi-view stereo. With high computation speed and low memory requirement, PatchmatchNet can process higher resolution imagery and is more suited to run on resource limited devices than competitors that employ 3D cost volume regularization. For the first time we introduce an iterative multi-scale Patchmatch in an end-to-end trainable architecture and improve the Patchmatch core algorithm with a novel and learned adaptive propagation and evaluation scheme for each iteration. Extensive experiments show a very competitive performance and generalization for our method on DTU, Tanks & Temples and ETH3D, but at a significantly higher efficiency than all existing top-performing models: at least two and a half times faster than state-of-the-art methods with twice less memory usage. Code is available at https://github.com/FangjinhuaWang/PatchmatchNet.
Fangjinhua Wang, Silvano Galliani, Christoph Vogel, Pablo Speciale, Marc Pollefeys
CVPR2
2020 Inference, Learning and Attention Mechanisms that Exploit and Preserve Sparsity in CNNs
Timo Hackel, Mikhail Usvyatsov, Silvano Galliani, Jan Dirk Wegner, Konrad Schindler
Int. J. Comput. Vis.3
2017 A Multi-view Stereo Benchmark with High-Resolution Images and Multi-camera Videos
abstract
Motivated by the limitations of existing multi-view stereo benchmarks, we present a novel dataset for this task. Towards this goal, we recorded a variety of indoor and outdoor scenes using a high-precision laser scanner and captured both high-resolution DSLR imagery as well as synchronized low-resolution stereo videos with varying fields-of-view. To align the images with the laser scans, we propose a robust technique which minimizes photometric errors conditioned on the geometry. In contrast to previous datasets, our benchmark provides novel challenges and covers a diverse set of viewpoints and scene types, ranging from natural scenes to man-made indoor and outdoor environments. Furthermore, we provide data at significantly higher temporal and spatial resolution. Our benchmark is the first to cover the important use case of hand-held mobile devices while also providing high-resolution DSLR camera images. We make our datasets and an online evaluation server available at http://www.eth3d.net.
Thomas Schöps, Johannes L. Schönberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, Andreas Geiger 0001
CVPR3
2017 Learned Multi-patch Similarity
abstract
Estimating a depth map from multiple views of a scene is a fundamental task in computer vision. As soon as more than two viewpoints are available, one faces the very basic question how to measure similarity across >2 image patches. Surprisingly, no direct solution exists, instead it is common to fall back to more or less robust averaging of two-view similarities. Encouraged by the success of machine learning, and in particular convolutional neural networks, we propose to learn a matching function which directly maps multiple image patches to a scalar similarity score. Experiments on several multi-view datasets demonstrate that this approach has advantages over methods based on pairwise patch similarity.
Wilfried Hartmann, Silvano Galliani, Michal Havlena, Luc Van Gool, Konrad Schindler
ICCV2
2016 Just Look at the Image: Viewpoint-Specific Surface Normal Prediction for Improved Multi-View Reconstruction
abstract
We present a multi-view reconstruction method that combines conventional multi-view stereo (MVS) with appearance-based normal prediction, to obtain dense and accurate 3D surface models. Reliable surface normals reconstructed from multi-view correspondence serve as training data for a convolutional neural network (CNN), which predicts continuous normal vectors from raw image patches. By training from known points in the same image, the prediction is specifically tailored to the materials and lighting conditions of the particular scene, as well as to the precise camera viewpoint. It is therefore a lot easier to learn than generic single-view normal estimation. The estimated normal maps, together with the known depth values from MVS, are integrated to dense depth maps, which in turn are fused into a 3D model. Experiments on the DTU dataset show that our method delivers 3D reconstructions with the same accuracy as MVS, but with significantly higher completeness.
Silvano Galliani, Konrad Schindler
CVPR1
2015 Massively Parallel Multiview Stereopsis by Surface Normal Diffusion
abstract
We present a new, massively parallel method for high-quality multiview matching. Our work builds on the Patchmatch idea: starting from randomly generated 3D planes in scene space, the best-fitting planes are iteratively propagated and refined to obtain a 3D depth and normal field per view, such that a robust photo-consistency measure over all images is maximized. Our main novelties are on the one hand to formulate Patchmatch in scene space, which makes it possible to aggregate image similarity across multiple views and obtain more accurate depth maps. And on the other hand a modified, diffusion-like propagation scheme that can be massively parallelized and delivers dense multiview correspondence over ten 1.9-Megapixel images in 3 seconds, on a consumer-grade GPU. Our method uses a slanted support window and thus has no fronto-parallel bias, it is completely local and parallel, such that computation time scales linearly with image size, and inversely proportional to the number of parallel threads. Furthermore, it has low memory footprint (four values per pixel, independent of the depth range). It therefore scales exceptionally well and can handle multiple large images at high depth resolution. Experiments on the DTU and Middlebury multiview datasets as well as oblique aerial images show that our method achieves very competitive results with high accuracy and completeness, across a range of different scenarios.
Silvano Galliani, Katrin Lasinger, Konrad Schindler
ICCV1
2012 Fast and Robust Surface Normal Integration by a Discrete Eikonal Equation
abstract
The integration of surface normals is a classic and fundamental task in computer vi-sion. In this paper we deal with a highly efficient fast marching (FM) method to perform the integration. In doing this we build upon a previous work of Ho and his coauthors. Their FM scheme is based on an analytic model that incorporates the eikonal equation. Our method is also built upon this equation, but it makes use of a complete discrete for-mulation for constructing the FM integrator (DEFM). We not only provide a theoretical justification of the proposed method, but also illustrate at hand of a simple example that our approach is much better suited to the task. Several more sophisticated tests confirm the robustness and higher accuracy of the DEFM model. Moreover, we present an ex-tension of DEFM that allows to integrate surface normals over non-trivial domains, e.g. featuring holes. Numerical results confirm desirable qualities of this method. 1
Silvano Galliani, Michael Breuß, Yong Chul Ju
BMVC1
2012 Shape from Shading for Rough Surfaces: Analysis of the Oren-Nayar Model
abstract
Due to their improved capability to handle realistic illumination scenarios, nonLambertian reflectance models are becoming increasingly more popular in the Shape from Shading (SfS) community. One of these advanced models is the Oren-Nayar model which is particularly suited to handle rough surfaces. However, not only the proper selection of the model is important, also the validation of stable and efficient algorithms plays a fundamental role when it comes to the practical applicability. While there are many works dealing with such algorithms in the case of Lambertian SfS, no such analysis has been performed so far for the Oren-Nayar model. In our paper we address this problem and present an in-depth study for such an advanced SfS model. To this end, we investigate under which conditions, i.e. model parameters, the Fast Marching (FM) method can be applied – a method that is known to be one of the most efficient algorithms for solving the underlying partial differential equations of Hamilton-Jacobi type. In this context, we do not only perform a general investigation of the model using Osher’s criterion for verifying the suitability of the FM method. We also conduct a parameter dependent analysis that shows, that FM can safely be used for the model for a wide range of settings relevant for practical applications. Thus, for the first time, it becomes possible to theoretically justify the use of the FM method as solver for the Oren-Nayar model which has been applied so far on a purely empirical basis only. Numerical experiments demonstrate the validity of our theoretical analysis. They show a stable behaviour of the FM method for the predicted range of model parameters.
Yong Chul Ju, Michael Breuß, Andrés Bruhn, Silvano Galliani
BMVC4
2007 The ball in the hole
abstract
"The Ball in The Hole" is an interactive video installation, running on GPL homemade software, which uses the principle of tele-presence.
Eleonora Maria Irene Oreggia, Silvano Galliani
ACM Multimedia2