Rodrigo Marcuzzi

dblp:311/3640 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0001-8076-0293ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 50% Generative modeling · 24% Segmentation and scene understanding · 12%
Computer graphics and multimedia
2 papers
Geometric modeling and processing · 100%

Topics — the 19 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
point cloud
1.422024
Scaling Diffusion Models to Real-World 3D LiDAR Scene Completion · CVPR 2024
Temporal Consistent 3D LiDAR Representation Learning for Semantic Perception in Autonomous Driving · CVPR 2023
Machine learning › Generative modeling
3d generative model
1.012026
Toward Generating Realistic 3D Semantic Training Data for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › 3D vision
3d scene understanding
1.012026
Toward Generating Realistic 3D Semantic Training Data for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Segmentation and scene understanding
3d semantic segmentation
1.012026
Toward Generating Realistic 3D Semantic Training Data for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Geometric modeling and processing
point cloud processing
0.912025
Tree Skeletonization From 3D Point Clouds by Denoising Diffusion · ICCV 2025
Computer vision › 3D vision › 3d scene understanding
3d scene completion
0.812024
Scaling Diffusion Models to Real-World 3D LiDAR Scene Completion · CVPR 2024
Machine learning › Generative modeling
diffusion model
0.812024
Scaling Diffusion Models to Real-World 3D LiDAR Scene Completion · CVPR 2024
Machine learning › Generative modeling › diffusion model › 3d diffusion models
point cloud diffusion
0.812024
Scaling Diffusion Models to Real-World 3D LiDAR Scene Completion · CVPR 2024
Geometric modeling and processing › shape modeling
shape completion
0.812024
Efficient and Accurate Transformer-Based 3D Shape Completion and Reconstruction of Fruits for Agricultural Robots · ICRA 2024
Computer vision › 3D vision › point cloud analysis › point cloud learning › point cloud representation learning
LiDAR representation learning
0.712023
Temporal Consistent 3D LiDAR Representation Learning for Semantic Perception in Autonomous Driving · CVPR 2023
Robotics › Robot navigation and mapping
place recognition
0.612022
Retriever: Point Cloud Retrieval in Compressed 3D Maps · ICRA 2022
Computer vision › 3D vision
point cloud processing
0.612022
Retriever: Point Cloud Retrieval in Compressed 3D Maps · ICRA 2022
Computer vision › 3D vision › point cloud processing › 3d point cloud understanding
point cloud retrieval
0.612022
Retriever: Point Cloud Retrieval in Compressed 3D Maps · ICRA 2022
Robotics › Autonomous driving
perception
0.522026
Toward Generating Realistic 3D Semantic Training Data for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Scaling Diffusion Models to Real-World 3D LiDAR Scene Completion · CVPR 2024
Computer vision › 3D vision › 3d scene understanding › 3d scene completion
LiDAR scene completion
0.212024
Scaling Diffusion Models to Real-World 3D LiDAR Scene Completion · CVPR 2024
Computational science and engineering
agricultural robotics
0.212024
Efficient and Accurate Transformer-Based 3D Shape Completion and Reconstruction of Fruits for Agricultural Robots · ICRA 2024
Natural language and speech › Information extraction and text analysis › natural language semantics › semantic interpretation
semantic perception
0.212023
Temporal Consistent 3D LiDAR Representation Learning for Semantic Perception in Autonomous Driving · CVPR 2023
Computer vision › Segmentation and scene understanding
semantic segmentation
0.212023
Temporal Consistent 3D LiDAR Representation Learning for Semantic Perception in Autonomous Driving · CVPR 2023
Machine learning › Efficient and distributed learning
model compression
0.212022
Retriever: Point Cloud Retrieval in Compressed 3D Maps · ICRA 2022

Methods — techniques the papers use, named apart from their topics

diffusion model · 1.8transformer · 1.5template matching · 1.5neural network · 1.5point cloud generation · 1.0denoising diffusion · 0.9regularization loss · 0.8self-supervised pretraining · 0.7contrastive learning · 0.7deep neural network · 0.6attention mechanism · 0.6NetVLAD · 0.6
YearPublicationVenuePosition
2026 Toward Generating Realistic 3D Semantic Training Data for Autonomous Driving
abstract
Semantic scene understanding is crucial for robotics and computer vision applications. In autonomous driving, 3D semantic segmentation plays an important role for enabling safe navigation. Despite significant advances in the field, the complexity of collecting and annotating 3D data is a bottleneck in this developments. To overcome that data annotation limitation, synthetic simulated data has been used to generate annotated data on demand. There is still, however, a domain gap between real and simulated data. More recently, diffusion models have been in the spotlight, enabling close-to-real data synthesis. Those generative models have been recently applied to the 3D data domain for generating scene-scale data with semantic annotations. Still, those methods either rely on image projection or decoupled models trained with different resolutions in a coarse-to-fine manner. Such intermediary representations impact the generated data quality due to errors added in those transformations. In this work, we propose a novel approach able to generate 3D semantic scene-scale data without relying on any projection or decoupled trained multi-resolution models, achieving more realistic semantic scene data generation compared to previous state-of-the-art methods. Besides improving 3D semantic scene-scale data synthesis, we thoroughly evaluate the use of the synthetic scene samples as labeled data to train a semantic segmentation network. In our experiments, we show that using the synthetic annotated data generated by our method as training data together with the real semantic segmentation labels, leads to an improvement in the semantic segmentation model performance. Our results show the potential of generated scene-scale point clouds to generate more training data to extend existing datasets, reducing the data annotation effort.
Lucas Nunes, Rodrigo Marcuzzi, Jens Behley, Cyrill Stachniss
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Tree Skeletonization From 3D Point Clouds by Denoising Diffusion
Elias Marks, Lucas Nunes, Federico Magistri, Matteo Sodano, Rodrigo Marcuzzi, Lars Zimmermann, Jens Behley, Cyrill Stachniss
ICCV5
2025 3D Hierarchical Panoptic Segmentation in Real Orchard Environments Across Different Sensors
abstract
Crop yield estimation is a relevant problem in agriculture, because an accurate yield estimate can support farmers’ decisions on harvesting or precision intervention. Robots can help to automate this process. To do so, they need to be able to perceive the surrounding environment to identify target objects such as trees and plants. In this paper, we introduce a novel approach to address the problem of hierarchical panoptic segmentation of apple orchards on 3D data from different sensors. Our approach is able to simultaneously provide semantic segmentation, instance segmentation of trunks and fruits, and instance segmentation of trees (a trunk with its fruits). This allows us to identify relevant information such as individual plants, fruits, and trunks, and capture the relationship among them, such as precisely estimate the number of fruits associated to each tree in an orchard. To efficiently evaluate our approach for hierarchical panoptic segmentation, we provide a dataset designed specifically for this task. Our dataset is recorded in Bonn, Germany, in a real apple orchard with a variety of sensors, spanning from a terrestrial laser scanner to a RGB-D camera mounted on different robots platforms. The experiments show that our approach surpasses state-of-the-art approaches in 3D panoptic segmentation in the agricultural domain, while also providing full hierarchical panoptic segmentation. Our dataset is publicly available at https://www.ipb.uni-bonn.de/data/hops/. The open-source implementation of our approach is available at https://github.com/PRBonn/hapt3D.
Matteo Sodano, Federico Magistri, Elias Marks, Fares Hosn, Aibek Zurbayev, Rodrigo Marcuzzi, Meher V. R. Malladi, Jens Behley, Cyrill Stachniss
IROS6
2024 Scaling Diffusion Models to Real-World 3D LiDAR Scene Completion
abstract
Computer vision techniques play a central role in the perception stack of autonomous vehicles. Such methods are employed to perceive the vehicle surroundings given sensor data. 3D LiDAR sensors are commonly used to collect sparse 3D point clouds from the scene. However, compared to human perception, such systems struggle to deduce the unseen parts of the scene given those sparse point clouds. In this matter, the scene completion task aims at predicting the gaps in the LiDAR measurements to achieve a more complete scene representation. Given the promising results of recent diffusion models as generative models for images, we propose extending them to achieve scene completion from a single 3D LiDAR scan. Previous works used diffusion models over range images extracted from LiDAR data, directly applying image-based diffusion methods. Distinctly, we propose to directly operate on the points, reformulating the noising and denoising diffusion process such that it can efficiently work at scene scale. Together with our approach, we propose a regularization loss to stabilize the noise predicted during the denoising process. Our experimental evaluation shows that our method can complete the scene given a single LiDAR scan as input, producing a scene with more details compared to state-of-the-art scene completion methods. We believe that our proposed diffusion process formulation can support further research in diffusion models applied to scene-scale point cloud data.11Code: https://github.com/PRBonn/LiDiff
Lucas Nunes, Rodrigo Marcuzzi, Benedikt Mersch, Jens Behley, Cyrill Stachniss
CVPR2
2024 Efficient and Accurate Transformer-Based 3D Shape Completion and Reconstruction of Fruits for Agricultural Robots
abstract
Robots that operate in agricultural environments need a robust perception system that can deal with occlusions, which are naturally present in agricultural scenarios. In this paper, we address the problem of estimating 3D shapes of fruits when only partial observations are available. Generally speaking, such a shape completion can be realized by exploiting prior knowledge about the geometry of the fruit. This is typically done by template matching using traditional optimization algorithms, which are slow but accurate, or by encoding such knowledge into the weights of a neural network, leading to faster but often less accurate estimates. Our approach combines the best of both worlds. It exploits the benefit of having a template representing our object of interest with the advantages of using a neural network to learn how to deform a template. Our experimental evaluation demonstrates that our approach yields accurate estimation at a competitively low inference time in challenging greenhouse environments.
Federico Magistri, Rodrigo Marcuzzi, Elias Marks, Matteo Sodano, Jens Behley, Cyrill Stachniss
ICRA2
2023 Temporal Consistent 3D LiDAR Representation Learning for Semantic Perception in Autonomous Driving
abstract
Semantic perception is a core building block in autonomous driving, since it provides information about the drivable space and location of other traffic participants. For learning-based perception, often a large amount of diverse training data is necessary to achieve high performance. Data labeling is usually a bottleneck for developing such methods, especially for dense prediction tasks, e.g., semantic segmentation or panoptic segmentation. For 3D Li-DAR data, the annotation process demands even more effort than for images. Especially in autonomous driving, point clouds are sparse, and objects appearance depends on its distance from the sensor, making it harder to acquire large amounts of labeled training data. This paper aims at taking an alternative path proposing a self-supervised representation learning method for 3D LiDAR data. Our approach exploits the vehicle motion to match objects across time viewed in different scans. We then train a model to maximize the point-wise feature similarities from points of the associated object in different scans, which enables to learn a consistent representation across time. The experimental results show that our approach performs better than previous state-of-the-art self-supervised representation learning methods when fine-tuning to different downstream tasks. We furthermore show that with only 10% of labeled data, a network pre-trained with our approach can achieve better performance than the same network trained from scratch with all labels for semantic segmentation on SemanticKITTI.11Code: https://github.com/PRBonn/TARL
Lucas Nunes, Louis Wiesmann, Rodrigo Marcuzzi, Xieyuanli Chen, Jens Behley, Cyrill Stachniss
CVPR3
2022 Retriever: Point Cloud Retrieval in Compressed 3D Maps
abstract
Most autonomous driving and robotic applications require retrieving map data around the vehicle's current location. Those maps can cover large areas and are often stored in a compressed form to save memory and allow for efficient transmission. In this paper, we address the problem of place recognition in a compressed point cloud map. To this end, we propose a novel deep neural network architecture that directly operates on a compressed feature representation produced by a compression encoder. This enables us to bypass compute-heavy decompression of the map and exploits the compact as well as descriptive nature of the compressed features. Additionally, we propose an alternative to the commonly used NetVLAD layer to aggregate local descriptors. Here, we utilize an attention mechanism between local features and a latent code. Our experiments suggest that this produces a more descriptive feature representation of the point clouds for place recognition. We experimentally validate all architectural choices we made by our ablation studies and compare our performance to other state-of-the-art baselines on two commonly used datasets.
Louis Wiesmann, Rodrigo Marcuzzi, Cyrill Stachniss, Jens Behley
ICRA2