Riccardo de Lutio

dblp:239/4210 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-2644-3876ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
3D vision · 86% Segmentation and scene understanding · 14%
Computer graphics and multimedia
4 papers
Image and video processing · 55% Rendering · 41% Computer animation and physical simulation · 5%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing › super-resolution › image super-resolution
guided super-resolution
1.022022
Learning Graph Regularisation for Guided Super-Resolution · CVPR 2022
Guided Super-Resolution As Pixel-to-Pixel Transformation · ICCV 2019
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.912025
OmniRe: Omni Urban Scene Reconstruction · ICLR 2025
Computer vision › 3D vision
3d scene reconstruction
0.912025
OmniRe: Omni Urban Scene Reconstruction · ICLR 2025
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
0.912025
OmniRe: Omni Urban Scene Reconstruction · ICLR 2025
Computer vision › 3D vision › point cloud processing
point cloud completion
0.912025
Towards Learning to Complete Anything in Lidar · ICML 2025
Computer vision › Segmentation and scene understanding
scene understanding
0.912025
Towards Learning to Complete Anything in Lidar · ICML 2025
Computer vision › 3D vision › 3d scene understanding
semantic scene completion
0.912025
Towards Learning to Complete Anything in Lidar · ICML 2025
Computer vision › 3D vision › 3d reconstruction
urban scene reconstruction
0.912025
OmniRe: Omni Urban Scene Reconstruction · ICLR 2025
Rendering
gaussian splatting
0.812024
3D Gaussian Ray Tracing: Fast Tracing of Particle Scenes · ACM Trans. Graph. 2024
Rendering › neural rendering
radiance field rendering
0.812024
3D Gaussian Ray Tracing: Fast Tracing of Particle Scenes · ACM Trans. Graph. 2024
Rendering
ray tracing
0.812024
3D Gaussian Ray Tracing: Fast Tracing of Particle Scenes · ACM Trans. Graph. 2024
Image and video processing › super-resolution › image super-resolution
depth super-resolution
0.612022
Learning Graph Regularisation for Guided Super-Resolution · CVPR 2022
Image and video processing
image restoration
0.612022
Learning Graph Regularisation for Guided Super-Resolution · CVPR 2022
Image and video processing › super-resolution
image super-resolution
0.612022
Learning Graph Regularisation for Guided Super-Resolution · CVPR 2022
Image and video processing
super-resolution
0.412019
Guided Super-Resolution As Pixel-to-Pixel Transformation · ICCV 2019

Methods — techniques the papers use, named apart from their topics

scene graph · 1.7gaussian splatting · 1.7zero-shot learning · 0.9multimodal sensor fusion · 0.9knowledge distillation · 0.9bounding volume hierarchy · 0.8GPU ray tracing · 0.8graph neural network · 0.6differentiable optimization · 0.6deep feature extraction · 0.6pixel-to-pixel mapping · 0.4multilayer perceptron · 0.4
YearPublicationVenuePosition
2025 OmniRe: Omni Urban Scene Reconstruction
abstract
We introduce OmniRe, a comprehensive system for efficiently creating high-fidelity digital twins of dynamic real-world scenes from on-device logs. Recent methods using neural fields or Gaussian Splatting primarily focus on vehicles, hindering a holistic framework for all dynamic foregrounds demanded by downstream applications, e.g., the simulation of human behavior. OmniRe extends beyond vehicle modeling to enable accurate, full-length reconstruction of diverse dynamic objects in urban scenes. Our approach builds scene graphs on 3DGS and constructs multiple Gaussian representations in canonical spaces that model various dynamic actors, including vehicles, pedestrians, cyclists, and others. OmniRe allows holistically reconstructing any dynamic object in the scene, enabling advanced simulations (~60 Hz) that include human-participated scenarios, such as pedestrian behavior simulation and human-vehicle interaction. This comprehensive simulation capability is unmatched by existing methods. Extensive evaluations on the Waymo dataset show that our approach outperforms prior state-of-the-art methods quantitatively and qualitatively by a large margin. We further extend our results to 5 additional popular driving datasets to demonstrate its generalizability on common urban scenes. Code and results are available at [omnire](https://ziyc.github.io/omnire/).
Jiawei Yang 0002, Riccardo de Lutio, Janick Martinez Esturo, Boris Ivanovic, Or Litany, Zan Gojcic, Sanja Fidler, Marco Pavone 0001, Yue Wang 0041
ICLR4
2025 Towards Learning to Complete Anything in Lidar
abstract
We propose CAL (Complete Anything in Lidar) for Lidar-based shape-completion in-the-wild. This is closely related to Lidar-based semantic/panoptic scene completion. However, contemporary methods can only complete and recognize objects from a closed vocabulary labeled in existing Lidar datasets. Different to that, our zero-shot approach leverages the temporal context from multi-modal sensor sequences to mine object shapes and semantic features of observed objects. These are then distilled into a Lidar-only instance-level completion and recognition model. Although we only mine partial shape completions, we find that our distilled model learns to infer full object shapes from multiple such partial observations across the dataset. We show that our model can be prompted on standard benchmarks for Semantic and Panoptic Scene Completion, localize objects as (amodal) 3D bounding boxes, and recognize objects beyond fixed class vocabularies.
Ayça Takmaz, Cristiano Saltori, Neehar Peri, Tim Meinhardt, Riccardo de Lutio, Laura Leal-Taixé, Aljosa Osep
ICML5
2024 3D Gaussian Ray Tracing: Fast Tracing of Particle Scenes
abstract
Particle-based representations of radiance fields such as 3D Gaussian Splatting have found great success for reconstructing and re-rendering of complex scenes. Most existing methods render particles via rasterization, projecting them to screen space tiles for processing in a sorted order. This work instead considers ray tracing the particles, building a bounding volume hierarchy and casting a ray for each pixel using high-performance GPU ray tracing hardware. To efficiently handle large numbers of semi-transparent particles, we describe a specialized rendering algorithm which encapsulates particles with bounding meshes to leverage fast ray-triangle intersections, and shades batches of intersections in depth-order. The benefits of ray tracing are well-known in computer graphics: processing incoherent rays for secondary lighting effects such as shadows and reflections, rendering from highly-distorted cameras common in robotics, stochastically sampling rays, and more. With our renderer, this flexibility comes at little cost compared to rasterization. Experiments demonstrate the speed and accuracy of our approach, as well as several applications in computer graphics and vision. We further propose related improvements to the basic Gaussian representation, including a simple use of generalized kernel functions which significantly reduces particle hit counts.
Nicolas Moënne-Loccoz, Ashkan Mirzaei, Or Perel, Riccardo de Lutio, Janick Martinez Esturo, Gavriel State, Sanja Fidler, Nicholas Sharp, Zan Gojcic
ACM Trans. Graph.4
2022 Learning Graph Regularisation for Guided Super-Resolution
abstract
We introduce a novel formulation for guided super-resolution. Its core is a differentiable optimisation layer that operates on a learned affinity graph. The learned graph potentials make it possible to leverage rich contextual information from the guide image, while the explicit graph optimisation within the architecture guarantees rigorous fidelity of the high-resolution target to the low-resolution source. With the decision to employ the source as a constraint rather than only as an input to the prediction, our method differs from state-of-the-art deep architectures for guided super-resolution, which produce targets that, when downsampled, will only approximately reproduce the source. This is not only theoretically appealing, but also produces crisper, more natural-looking images. A key property of our method is that, although the graph connectivity is restricted to the pixel lattice, the associated edge potentials are learned with a deep feature extractor and can encode rich context information over large receptive fields. By taking advantage of the sparse graph connectivity, it becomes possible to propagate gradients through the optimisation layer and learn the edge potentials from data. We extensively evaluate our method on several datasets, and consistently outperform recent baselines in terms of quantitative reconstruction errors, while also delivering visually sharper outputs. Moreover, we demonstrate that our method generalises particularly well to new datasets not seen during training.
Riccardo de Lutio, Alexander Becker 0002, Stefano D'Aronco, Stefania Russo, Jan Dirk Wegner, Konrad Schindler
CVPR1
2019 Guided Super-Resolution As Pixel-to-Pixel Transformation
abstract
Guided super-resolution is a unifying framework for several computer vision tasks where the inputs are a low-resolution source image of some target quantity (e.g., perspective depth acquired with a time-of-flight camera) and a high-resolution guide image from a different domain (e.g., a grey-scale image from a conventional camera); and the target output is a high-resolution version of the source (in our example, a high-res depth map). The standard way of looking at this problem is to formulate it as a super-resolution task, i.e., the source image is upsampled to the target resolution, while transferring the missing high-frequency details from the guide. Here, we propose to turn that interpretation on its head and instead see it as a pixel-to-pixel mapping of the guide image to the domain of the source image. The pixel-wise mapping is parametrised as a multi-layer perceptron, whose weights are learned by minimising the discrepancies between the source image and the downsampled target image. Importantly, our formulation makes it possible to regularise only the mapping function, while avoiding regularisation of the outputs; thus producing crisp, natural-looking images. The proposed method is unsupervised, using only the specific source and guide images to fit the mapping. We evaluate our method on two different tasks, super-resolution of depth maps and of tree height maps. In both cases, we clearly outperform recent baselines in quantitative comparisons, while delivering visually much sharper outputs.
Riccardo de Lutio, Stefano D'Aronco, Jan Dirk Wegner, Konrad Schindler
ICCV1