EDBT 2026 Demo / reviewers in the wild / expert
Alberto Pretto
dblp:01/3588
· DBLP profile ↗
25ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0003-1920-2887ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 6 since 2021Systems, architecture and hardware · 15 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revisiting Retentive Networks for Fast Range-View 3D LiDAR Semantic SegmentationabstractLiDAR semantic segmentation is a crucial task in autonomous driving and robotics, where real-time performance is essential for online decision-making. Recent trends exploit range images and Vision Transformers, using the self-attention mechanism. However, these approaches often lack explicit spatial priors and involve a large number of parameters. To tackle these limitations, we propose a novel method, adapting the Retentive Network architecture from the Natural Language Processing (NLP) field, for its efficient sequence modeling capabilities, directly operating on the range-view representation. Our approach incorporates a circular retention (CiR) mechanism that explicitly captures spatial relationships and continual circular property of the range image while modeling long-range dependencies and preserving the receptive field. In addition, we introduce a new set of range-view augmentations, adapted from 3D techniques, to improve generalization and mitigate class imbalance. Extensive experiments on three large-scale datasets, as SemanticKITTI, PandaSet and Semantic-POSS demonstrate that our method achieve state-of-the-art performance among range-view approaches on two out of three datasets, while satisfying real-time constraints. The code is available at https://github.com/SiMoM0/RangeRet. Simone Mosco, Daniel Fusaro, Wanmeng Li, Alberto Pretto |
WACV | 4 |
| 2025 | DPGLA: Bridging the Gap between Synthetic and Real Data for Unsupervised Domain Adaptation in 3D LiDAR Semantic SegmentationabstractAnnotating real-world LiDAR point clouds for use in intelligent autonomous systems is costly. To overcome this limitation, self-training-based Unsupervised Domain Adaptation (UDA) has been widely used to improve point cloud semantic segmentation by leveraging synthetic point cloud data. However, we argue that existing methods do not effectively utilize unlabeled data, as they either rely on predefined or fixed confidence thresholds, resulting in suboptimal performance. In this paper, we propose a Dynamic Pseudo-Label Filtering (DPLF) scheme to enhance real data utilization in point cloud UDA semantic segmentation. Additionally, we design a simple and efficient Prior-Guided Data Augmentation Pipeline (PG-DAP) to mitigate domain shift between synthetic and real-world point clouds. Finally, we utilize data mixing consistency loss to push the model to learn context-free representations. We implement and thoroughly evaluate our approach through extensive comparisons with state-of-the-art methods. Experiments on two challenging synthetic-to-real point cloud semantic segmentation tasks demonstrate that our approach achieves superior performance. Ablation studies confirm the effectiveness of the DPLF and PG-DAP modules. We release the code of our method in this paper. Wanmeng Li, Simone Mosco, Daniel Fusaro, Alberto Pretto |
IROS | 4 |
| 2024 | CombiNeRF: A Combination of Regularization Techniques for Few-Shot Neural Radiance Field View SynthesisabstractNeural Radiance Fields (NeRFs) have shown impressive results for novel view synthesis when a sufficiently large amount of views are available. When dealing with few-shot settings, i.e. with a small set of input views, the training could overfit those views, leading to artifacts and geometric and chromatic inconsistencies in the resulting rendering. Regularization is a valid solution that helps NeRF generalization. On the other hand, each of the most recent NeRF regularization techniques aim to mitigate a specific rendering problem. Starting from this observation, in this paper we propose CombiNeRF, a framework that synergically combines several regularization techniques, some of them novel, in order to unify the benefits of each. In particular, we regularize single and neighboring rays distributions and we add a smoothness term to regularize near geometries. After these geometric approaches, we propose to exploit Lipschitz regularization to both NeRF density and color networks and to use encoding masks for input features regularization. We show that CombiNeRF outperforms the state-of-the-art methods with few-shot settings in several publicly available datasets. We also present an ablation study on the LLFF and NeRF-Synthetic datasets that support the choices made. We release with this paper the open-source implementation of our framework. Matteo Bonotto, Luigi Sarrocco, Daniele Evangelista, Marco Imperoli, Alberto Pretto |
3DV | 5 |
| 2024 | PanNote: an Automatic Tool for Panoramic Image Annotation of People's PositionsabstractPanoramic cameras offer a 4π steradian field of view, which is desirable for tasks like people detection and tracking since nobody can exit the field of view. Despite the recent diffusion of low-cost panoramic cameras, their usage in robotics remains constrained by the limited availability of datasets featuring annotations in the robot space, including people’s 2D or 3D positions. To tackle this issue, we introduce PanNote, an automatic annotation tool for people’s positions in panoramic videos. Our tool is designed to be cost-effective and straightforward to use without requiring human intervention during the labeling process and enabling the training of machine learning models with low effort. The proposed method introduces a calibration model and a data association algorithm to fuse data from panoramic images and 2D LiDAR readings. We validate the capabilities of PanNote by collecting a real-world dataset. On these data, we compared manual labels, automatic labels and the predictions of a baseline deep neural network. Results clearly show the advantage of using our method, with a 15-fold speed up in labeling time and a considerable gain in performance while training deep neural models on automatically labelled data. Alberto Bacchin, Leonardo Barcellona, Sepideh Shamsizadeh, Emilio Olivastri, Alberto Pretto, Emanuele Menegatti |
ICRA | 5 |
| 2024 | IPC: Incremental Probabilistic Consensus-based Consistent Set Maximization for SLAM BackendsabstractIn SLAM (Simultaneous localization and mapping) problems, Pose Graph Optimization (PGO) is a technique to refine an initial estimate of a set of poses (positions and orientations) from a set of pairwise relative measurements. The optimization procedure can be negatively affected even by a single outlier measurement, with possible catastrophic and meaningless results. Although recent works on robust optimization aim to mitigate the presence of outlier measurements, robust solutions capable of handling large numbers of outliers are yet to come. This paper presents IPC, acronym for Incremental Probabilistic Consensus, a method that approximates the solution to the combinatorial problem of finding the maximally consistent set of measurements in an incremental fashion. It evaluates the consistency of each loop closure measurement through a consensus-based procedure, possibly applied to a subset of the global problem, where all previously integrated inlier measurements have veto power. We evaluated IPC on standard benchmarks against several state-of-the-art methods. Although it is simple and relatively easy to implement, IPC competes with or outperforms the other tested methods in handling outliers while providing online performances. We release with this paper an open-source implementation of the proposed method. Emilio Olivastri, Alberto Pretto |
ICRA | 2 |
| 2024 | Exploiting Local Features and Range Images for Small Data Real-Time Point Cloud Semantic SegmentationabstractSemantic segmentation of point clouds is an essential task for understanding the environment in autonomous driving and robotics. Recent range-based works achieve real-time efficiency, while point- and voxel-based methods produce better results but are affected by high computational complexity. Moreover, highly complex deep learning models are often not suited to efficiently learn from small datasets. Their generalization capabilities can easily be driven by the abundance of data rather than the architecture design. In this paper, we harness the information from the three-dimensional representation to proficiently capture local features, while introducing the range image representation to incorporate additional information and facilitate fast computation. A GPU-based KDTree allows for rapid building, querying, and enhancing projection with straightforward operations. Extensive experiments on SemanticKITTI and nuScenes datasets demonstrate the benefits of our modification in a "small data" setup, in which only one sequence of the dataset is used to train the models, but also in the conventional setup, where all sequences except one are used for training. We show that a reduced version of our model not only demonstrates strong competitiveness against full-scale state-of-the-art models but also operates in real-time, making it a viable choice for real-world case applications. The code of our method is available at https://github.com/Bender97/WaffleAndRange. Daniel Fusaro, Simone Mosco, Emanuele Menegatti, Alberto Pretto |
IROS | 4 |
| 2023 | A Graph-Based Optimization Framework for Hand-Eye Calibration for Multi-Camera SetupsabstractHand-eye calibration is the problem of estimating the spatial transformation between a reference frame, usually the base of a robot arm or its gripper, and the reference frame of one or multiple cameras. Generally, this calibration is solved as a non-linear optimization problem, what instead is rarely done is to exploit the underlying graph structure of the problem itself. Actually, the problem of hand-eye calibration can be seen as an instance of the Simultaneous Localization and Mapping (SLAM) problem. Inspired by this fact, in this work we present a pose-graph approach to the hand-eye calibration problem that extends a recent state-of-the-art solution in two different ways: i) by formulating the solution to eye-on-base setups with one camera; ii) by covering multi-camera robotic setups. The proposed approach has been validated in simulation against standard hand-eye calibration methods. Moreover, a real application is shown. In both scenarios, the proposed approach overcomes all alternative methods. We release with this paper an open-source implementation of our graph-based optimization framework for multi-camera setups. Daniele Evangelista, Emilio Olivastri, Davide Allegro, Emanuele Menegatti, Alberto Pretto |
ICRA | 5 |
| 2023 | Improving Generalization of Synthetically Trained Sonar Image Descriptors for Underwater Place Recognition
Ivano Donadi, Emilio Olivastri, Daniel Fusaro, Wanmeng Li, Daniele Evangelista, Alberto Pretto |
ICVS | 6 |
| 2022 | An Unified Iterative Hand-Eye Calibration Method for Eye-on-Base and Eye-in-Hand SetupsabstractThis paper presents an accurate and precise hand-eye calibration technique based on minimization of the reprojection error. Unlike traditional hand-eye calibration, the proposed method does not require an explicit estimate of the camera pose for each input image because it does not rely on mathematical description and problem formulation commonly used in standard hand-eye calibration algorithms. The proposed method is based on a nonlinear optimization problem, so that the estimation problem can be solved efficiently and robustly, and can be easily extended to different camera-robot setups (e.g., eye-on-base or eye-in-hand). An extensive evaluation based on simulated and real experiments has been performed, proving its good estimation accuracy in terms of reprojection error. The experimental results with real robots show that the proposed method is applicable to relevant industrial contexts and improves the quality and precision of the camera-robot transformation estimation with respect to state-of-the-art approaches. Daniele Evangelista, Davide Allegro, Matteo Terreran, Alberto Pretto, Stefano Ghidoni |
ETFA | 4 |
| 2022 | An Hybrid Approach to Improve the Performance of Encoder-Decoder Architectures for Traversability Analysis in Urban EnvironmentsabstractSelf-driving vehicles and autonomous ground robots require a reliable and accurate method to analyze the traversability of the surrounding environment for safe navigation. This paper proposes a hybrid approach that combines geometric and appearance features for training Deep Encoder-Decoder architectures to detect the traversability score in real urban contexts. The proposed approach has been tested with two Deep Learning architectures on a public dataset of outdoor driving scenarios. Thanks to our approach, we are able to reach high levels of accuracy in detecting the correct traversability score in environments of highly variable complexity. This demonstrates the effectiveness and robustness of the proposed method. Daniel Fusaro, Emilio Olivastri, Daniele Evangelista, Pietro Iob, Alberto Pretto |
IV | 5 |
| 2021 | Make It Easier: An Empirical Simplification of a Deep 3D Segmentation Network for Human Body Parts
Matteo Terreran, Daniele Evangelista, Jacopo Lazzaro, Alberto Pretto |
ICVS | 4 |
| 2020 | 3D Mapping of X-Ray Images in Inspections of Aerospace PartsabstractIn this work we present an industrial system for the inspection of composite parts in the aerospace industry, based on X-ray sensors and robotic manipulators. Such system is designed to identify any type of defects such as, missing gluing, core cell deformation, cracks or foreign objects, which may occur between layers of which these objects are composed. The inspection process involves back-projection of X-ray images onto the 3D CAD model of the inspected part, to directly locate the defects on the part itself. The complete system has been implemented in a real industrial workcell that involves two synchronized robots equipped with a X-ray source-detector system. The two robots move autonomously along a pre-computed trajectory without any human intervention, and the back-projection of the acquired images is efficiently performed at run-time using the proposed algorithm. The experiments demonstrate that the X-ray images back-projection is successful and can effectively replace standard manually guided inspections. This has a high impact on the factory automation cycle since it helps to reduce the effort and time needed for each inspection task. This work is part of a EU funded project called SPIRIT. Daniele Evangelista, Matteo Terreran, Alberto Pretto, Michele Moro, Carlo Ferrari, Emanuele Menegatti |
ETFA | 3 |
| 2018 | Robust Intrinsic and Extrinsic Calibration of RGB-D CamerasabstractColor-depth cameras (RGB-D cameras) have become the primary sensors in most robotics systems, from service robotics to industrial robotics applications. Typical consumer-grade RGB-D cameras are provided with a coarse intrinsic and extrinsic calibration that generally does not meet the accuracy requirements needed by many robotics applications [e.g., highly accurate three-dimensional (3-D) environment reconstruction and mapping, high precision object recognition, localization, etc.]. In this paper, we propose a human-friendly, reliable, and accurate calibration framework that enables to easily estimate both the intrinsic and extrinsic parameters of a general color-depth sensor couple. Our approach is based on a novel two components error model. This model unifies the error sources of RGB-D pairs based on different technologies, such as structured-light 3-D cameras and time-of-flight cameras. Our method provides some important advantages compared to other state-of-the-art systems: It is general (i.e., well suited for different types of sensors), based on an easy and stable calibration protocol, provides a greater calibration accuracy, and has been implemented within the robot operating system robotics framework. We report detailed experimental validations and performance comparisons to support our statements. Filippo Basso, Emanuele Menegatti, Alberto Pretto |
IEEE Trans. Robotics | 3 |
| 2017 | Automatic model based dataset generation for fast and accurate crop and weeds detectionabstractSelective weeding is one of the key challenges in the field of agriculture robotics. To accomplish this task, a farm robot should be able to accurately detect plants and to distinguish them between crop and weeds. Most of the promising state-of-the-art approaches make use of appearance-based models trained on large annotated datasets. Unfortunately, creating large agricultural datasets with pixel-level annotations is an extremely time consuming task, actually penalizing the usage of data-driven techniques. In this paper, we face this problem by proposing a novel and effective approach that aims to dramatically minimize the human intervention needed to train the detection and classification algorithms. The idea is to procedurally generate large synthetic training datasets randomizing the key features of the target environment (i.e., crop and weed species, type of soil, light conditions). More specifically, by tuning these model parameters, and exploiting a few real-world textures, it is possible to render a large amount of realistic views of an artificial agricultural scenario with no effort. The generated data can be directly used to train the model or to supplement real-world images. We validate the proposed methodology by using as testbed a modern deep learning based image segmentation architecture. We compare the classification results obtained using both real and synthetic images as training data. The reported results confirm the effectiveness and the potentiality of our approach. Maurilio Di Cicco, Ciro Potena, Giorgio Grisetti, Alberto Pretto |
IROS | 4 |
| 2015 | Plane Extraction for Indoor Place Recognition
Ciro Potena, Alberto Pretto, Domenico Daniele Bloisi, Daniele Nardi |
ACIVS | 2 |
| 2015 | D ^2 CO: Fast and Robust Registration of 3D Textureless Objects Using the Directional Chamfer Distance
Marco Imperoli, Alberto Pretto |
ICVS | 2 |
| 2014 | Unsupervised intrinsic and extrinsic calibration of a camera-depth sensor coupleabstractThe availability of affordable depth sensors in conjunction with common RGB cameras (even in the same device, e.g. the Microsoft Kinect) provides robots with a complete and instantaneous representation of both the appearance and the 3D structure of the current surrounding environment. This type of information enables robots to safely navigate, perceive and actively interact with other agents inside the working environment. It is clear that, in order to obtain a reliable and accurate representation, not only the intrinsic parameters of each sensors should be precisely calibrated, but also the extrinsic parameters relating the two sensors should be precisely known. In this paper, we propose a human-friendly and reliable calibration framework, that enables to easily estimate both the intrinsic and extrinsic parameters of a camera-depth sensor couple. Real world experiments using a Kinect show improvements for both the 3D structure estimation and the association tasks. Filippo Basso, Alberto Pretto, Emanuele Menegatti |
ICRA | 2 |
| 2014 | A robust and easy to implement method for IMU calibration without external equipmentsabstractMotion sensors as inertial measurement units (IMU) are widely used in robotics, for instance in the navigation and mapping tasks. Nowadays, many low cost micro electro mechanical systems (MEMS) based IMU are available off the shelf, while smartphones and similar devices are almost always equipped with low-cost embedded IMU sensors. Nevertheless, low cost IMUs are affected by systematic error given by imprecise scaling factors and axes misalignments that decrease accuracy in the position and attitudes estimation. In this paper, we propose a robust and easy to implement method to calibrate an IMU without any external equipment. The procedure is based on a multi-position scheme, providing scale and misalignments factors for both the accelerometers and gyroscopes triads, while estimating the sensor biases. Our method only requires the sensor to be moved by hand and placed in a set of different, static positions (attitudes). We describe a robust and quick calibration protocol that exploits an effective parameterless static filter to reliably detect the static intervals in the sensor measurements, where we assume local stability of the gravity's magnitude and stable temperature. We first calibrate the accelerometers triad taking measurement samples in the static intervals. We then exploit these results to calibrate the gyroscopes, employing a robust numerical integration technique. The performances of the proposed calibration technique has been successfully evaluated via extensive simulations and real experiments with a commercial IMU provided with a calibration certificate as reference data. David Tedaldi, Alberto Pretto, Emanuele Menegatti |
ICRA | 2 |
| 2011 | Omnidirectional dense large-scale mapping and navigation based on meaningful triangulationabstractIn this work, we propose a robust and efficient method to build dense 3D maps, using only the images grabbed by an omnidirectional camera. The map contains exhaustive information about both the structure and the appearance of the environment and it is well suited also for large scale environments. We start from the assumption that the surrounding environment (the scene) forms a piecewise smooth surface represented by a triangle mesh. Our system is able to infer, without any odometry information, the structure of the environment along with the ego-motion of the camera by performing a robust tracking of the projection of this surface in the omnidirectional image. The key idea is to use a guess of the triangle mesh subdivision based on a constrained Delaunay triangulation built according to a set of point features and edgelet features extracted from the image. In such a way, we take into account both the corners and the edges of the scene imaged by the camera, constrained by the topology of the triangulation in order to improve the stability of the tracking process. Both motion and structure parameters are estimated using a direct method inside an optimization framework, taking into account the topology of the subdivision in a robust and efficient way. We successfully tested our system in a challenging urban scenario along a large loop using an omnidirectional camera mounted on the roof of a car. Alberto Pretto, Emanuele Menegatti, Enrico Pagello |
ICRA | 1 |
| 2010 | Cooperative tracking of moving objects and face detection with a dual camera sensorabstractThis paper describes a sensor for autonomous surveillance capable of continuously monitoring the environment, while acquiring detailed images of specific areas. This is achieved by exploiting an omnidirectional camera and a PTZ camera, assembled together on a single mount. The two cameras form a single vision sensor, since data obtained processing the two images are used in a cooperative way. This system solves the problem affecting systems based on PTZ cameras only, since it does not exist a tracking system working reliably on PTZ images: the problem is solved here by performing the tracking in the omnidirectional image. This vision sensor is used in a surveillance application, that detects moving objects, and records all the faces of the people walking close to the sensor. It could be used to navigate or instruct a security or a service mobile robot. Stefano Ghidoni, Alberto Pretto, Emanuele Menegatti |
ICRA | 2 |
| 2009 | A visual odometry framework robust to motion blurabstractMotion blur is a severe problem in images grabbed by legged robots and, in particular, by small humanoid robots. Standard feature extraction and tracking approaches typically fail when applied to sequences of images strongly affected by motion blur. In this paper, we propose a new feature detection and tracking scheme that is robust even to non-uniform motion blur. Furthermore, we developed a framework for visual odometry based on features extracted out of and matched in monocular image sequences. To reliably extract and track the features, we estimate the point spread function (PSF) of the motion blur individually for image patches obtained via a clustering technique and only consider highly distinctive features during matching. We present experiments performed on standard datasets corrupted with motion blur and on images taken by a camera mounted on walking small humanoid robots to show the effectiveness of our approach. The experiments demonstrate that our technique is able to reliably extract and match features and that it is furthermore able to generate a correct visual odometry, even in presence of strong motion blur effects and without the aid of any inertial measurement sensor. Alberto Pretto, Emanuele Menegatti, Maren Bennewitz, Wolfram Burgard, Enrico Pagello |
ICRA | 1 |
| 2007 | A robotic sculpture speaking to peopleabstractThis video shows the interactive robotic sculpture conceived and realized by the artist Albano Guatti. The robotic part was totally developed by people at the IAS-lab and at IT+Robotics according to Guatti's concept. This work is the result of the meeting of robotics and art. The video shows that the robot is able to locate people in the environment, to navigate toward them avoiding the obstacles and to approach them as a polite waiter will do. In fact, the sculpture represents a waiter and a waitress (actually their suits, the statue do not have a body, they represent empty suits). The waiters chat among them (by play different pre-recorded voice files) while wandering in the environment. Once the robot locates a person and get close to her, it asks the customer one out of different pre-recorded questions. The more likely is: Would you like a drink?, that is the title of the sculpture. Emanuele Menegatti, Alberto Pretto, Stefano Tonello, Alvise Lastra, A. Guatti |
ICRA | 2 |
| 2006 | Omnidirectional vision scan matching for robot localization in dynamic environmentsabstractThe localization problem for an autonomous robot moving in a known environment is a well-studied problem which has seen many elegant solutions. Robot localization in a dynamic environment populated by several moving obstacles, however, is still a challenge for research. In this paper, we use an omnidirectional camera mounted on a mobile robot to perform a sort of scan matching. The omnidirectional vision system finds the distances of the closest color transitions in the environment, mimicking the way laser rangefinders detect the closest obstacles. The similarity of our sensor with classical rangefinders allows the use of practically unmodified Monte Carlo algorithms, with the additional advantage of being able to easily detect occlusions caused by moving obstacles. The proposed system was initially implemented in the RoboCup Middle-Size domain, but the experiments we present in this paper prove it to be valid in a general indoor environment with natural color transitions. We present localization experiments both in the RoboCup environment and in an unmodified office environment. In addition, we assessed the robustness of the system to sensor occlusions caused by other moving robots. The localization system runs in real-time on low-cost hardware. Emanuele Menegatti, Alberto Pretto, Alberto Scarpa, Enrico Pagello |
IEEE Trans. Robotics | 2 |
| 2004 | Testing omnidirectional vision-based Monte Carlo localization under occlusionabstractOne of the most challenging issues in mobile robot navigation is the localization problem in densely populated environments. In this paper, we present a new approach for vision-based localization able to solve this problem. The omnidirectional camera is used as a range finder sensitive to the distance of color transitions, whereas classical range finder;, like lasers or sonars, are sensitive to the distance of the nearest obstacles. The well-known Monte-Carlo localization technique was adapted for this new type of range sensor. The system runs in real time on a low-cost pc. In this paper we present experiments, performed in a crowded RoboCup middle-size field, proving the robustness of the approach to the occlusions of the vision sensor by moving obstacles (e.g other robots); occlusions that are very likely to occur in a real environment. Although, the system was implemented for the RoboCup environment, the system can be used in more general environments. Emanuele Menegatti, Alberto Pretto, Enrico Pagello |
IROS | 2 |
| 2004 | A New Omnidirectional Vision Sensor for Monte-Carlo Localization
Emanuele Menegatti, Alberto Pretto, Enrico Pagello |
RoboCup | 2 |