Stefano Ghidoni

dblp:77/5606 · DBLP profile ↗
← Back
32ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-3406-8719ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 2 first-author · 5 since 2021Systems, architecture and hardware · 12 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 4
YearPublicationVenuePosition
2026 SkelSplat: Robust Multi-view 3D Human Pose Estimation with Differentiable Gaussian Rendering
abstract
Accurate 3D human pose estimation is fundamental for applications such as augmented reality and human-robot interaction. State-of-the-art multi-view methods learn to fuse predictions across views by training on large annotated datasets, leading to poor generalization when the test scenario differs. To overcome these limitations, we propose SkelSplat, a novel framework for multi-view 3D human pose estimation based on differentiable Gaussian rendering. Human pose is modeled as a skeleton of 3D Gaussians, one per joint, optimized via differentiable rendering to enable seamless fusion of arbitrary camera views without 3D ground-truth supervision. Since Gaussian Splatting was originally designed for dense scene reconstruction, we propose a novel one-hot encoding scheme that enables independent optimization of human joints. SkelSplat outperforms approaches that do not rely on 3D ground truth in Human3.6M and CMU, while reducing the cross-dataset error up to 47.8% compared to learning-based methods. Experiments on Human3.6M-Occ and Occlusion-Person demonstrate robustness to occlusions, without scenario-specific fine-tuning. Our project page is available here: https://skelsplat.github.io.
Laura Bragagnolo, Leonardo Barcellona, Stefano Ghidoni
WACV3
2025 Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination
abstract
A world model provides an agent with a representation of its environment, enabling it to predict the causal consequences of its actions. Current world models typically cannot directly and explicitly imitate the actual environment in front of a robot, often resulting in unrealistic behaviors and hallucinations that make them unsuitable for real-world robotics applications. To overcome those challenges, we propose to rethink robot world models as learnable digital twins. We introduce DreMa, a new approach for constructing digital twins automatically using learned explicit representations of the real world and its dynamics, bridging the gap between traditional digital twins and world models. DreMa replicates the observed world and its structure by integrating Gaussian Splatting and physics simulators, allowing robots to imagine novel configurations of objects and to predict the future consequences of robot actions thanks to its compositionality. We leverage this capability to generate new data for imitation learning by applying equivariant transformations to a small set of demonstrations. Our evaluations across various settings demonstrate significant improvements in accuracy and robustness by incrementing actions and object distributions, reducing the data needed to learn a policy and improving the generalization of the agents. As a highlight, we show that a real Franka Emika Panda robot, powered by DreMa’s imagination, can successfully learn novel physical tasks from just a single example per task variation (one-shot policy learning). Our project page can be found in: https://dreamtomanipulate.github.io/.
Leonardo Barcellona, Andrii Zadaianchuk, Davide Allegro, Samuele Papa, Stefano Ghidoni, Efstratios Gavves
ICLR5
2024 Human-Robot Collaborative Transportation via Distance-based Role Allocation for Precise Positioning of Flexible Materials
abstract
Despite the importance of human-robot collaborative transportation of flexible material in many industrial scenarios, many works in the literature assume a passive role for the robot during the collaboration. The robot can only follow the human partner, without providing assistance in the more challenging phase of the collaboration such as precise material positioning. This work presents a framework for co-transportation, proposing a distance-based policy for dynamic leader role allocation through the task. For large distances from the target pose, the robot is mainly controlled by vision-based manual guidance exploiting haptic feedback and 3D human pose information; instead, close to the target material position, the robot acts as a leader guiding the human operator. The proposed framework is evaluated considering a carbon fiber draping task, which requires both co-transportation and precise positioning of flexible materials. Experimental results demonstrate how the robot leading the task in the final stage allows to achieve high task efficiency and alleviates human stress in the execution of the task.
Matteo Terreran, Alberto Gottardi, Emanuele Menegatti, Stefano Ghidoni
ETFA4
2024 MEMROC: Multi-Eye to Mobile RObot Calibration
abstract
This paper presents MEMROC (Multi-Eye to Mobile RObot Calibration), a novel motion-based calibration method that simplifies the process of accurately calibrating multiple cameras relative to a mobile robot’s reference frame. MEMROC utilizes a known calibration pattern to facilitate accurate calibration with a lower number of images during the optimization process. Additionally, it leverages robust ground plane detection for comprehensive 6-DoF extrinsic calibration, overcoming a critical limitation of many existing methods that struggle to estimate the complete camera pose. The proposed method addresses the need for frequent recalibration in dynamic environments, where cameras may shift slightly or alter their positions due to daily usage, operational adjustments, or vibrations from mobile robot movements. MEMROC exhibits remarkable robustness to noisy odometry data, requiring minimal calibration input data. This combination makes it highly suitable for daily operations involving mobile robots. A comprehensive set of experiments on both synthetic and real data proves MEMROC’s efficiency, surpassing existing state-of-the-art methods in terms of accuracy, robustness, and ease of use. To facilitate further research, we have made our code publicly available1.
Davide Allegro, Matteo Terreran, Stefano Ghidoni
IROS3
2024 WasteGAN: Data Augmentation for Robotic Waste Sorting through Generative Adversarial Networks
abstract
Robotic waste sorting poses significant challenges in both perception and manipulation, given the extreme variability of objects that should be recognized on a cluttered conveyor belt. While deep learning has proven effective in solving complex tasks, the necessity for extensive data collection and labeling limits its applicability in real-world scenarios like waste sorting. To tackle this issue, we introduce a data augmentation method based on a novel GAN architecture called wasteGAN. The proposed method allows to increase the performance of semantic segmentation models, starting from a very limited bunch of labeled examples, such as few as 100. The key innovations of wasteGAN include a novel loss function, a novel activation function, and a larger generator block. Overall, such innovations helps the network to learn from limited number of examples and synthesize data that better mirrors real-world distributions. We then leverage the higher-quality segmentation masks predicted from models trained on the wasteGAN synthetic data to compute semantic-aware grasp poses, enabling a robotic arm to effectively recognizing contaminants and separating waste in a real-world scenario. Through comprehensive evaluation encompassing dataset-based assessments and real-world experiments, our methodology demonstrated promising potential for robotic waste sorting, yielding performance gains of up to 5.8% in picking contaminants. The project page is available at https://github.com/bach05/wasteGAN.git.
Alberto Bacchin, Leonardo Barcellona, Matteo Terreran, Stefano Ghidoni, Emanuele Menegatti, Takuya Kiyokawa
IROS4
2023 FSG-Net: a Deep Learning model for Semantic Robot Grasping through Few-Shot Learning
abstract
Robot grasping has been widely studied in the last decade. Recently, Deep Learning made possible to achieve remarkable results in grasp pose estimation, using depth and RGB images. However, only few works consider the choice of the object to grasp. Moreover, they require a huge amount of data for generalizing to unseen object categories. For this reason, we introduce the Few-shot Semantic Grasping task where the objective is inferring a correct grasp given only five labelled images of a target unseen object. We propose a new deep learning architecture able to solve the aforementioned problem, leveraging on a Few-shot Semantic Segmentation module. We have evaluated the proposed model both in the Graspnet dataset and in a real scenario. In Graspnet, we achieve 40,95% accuracy in the Few-shot Semantic Grasping task, outperforming baseline approaches. In the real experiments, the results confirmed the generalization ability of the network.
Leonardo Barcellona, Alberto Bacchin, Alberto Gottardi, Emanuele Menegatti, Stefano Ghidoni
ICRA5
2022 An Unified Iterative Hand-Eye Calibration Method for Eye-on-Base and Eye-in-Hand Setups
abstract
This paper presents an accurate and precise hand-eye calibration technique based on minimization of the reprojection error. Unlike traditional hand-eye calibration, the proposed method does not require an explicit estimate of the camera pose for each input image because it does not rely on mathematical description and problem formulation commonly used in standard hand-eye calibration algorithms. The proposed method is based on a nonlinear optimization problem, so that the estimation problem can be solved efficiently and robustly, and can be easily extended to different camera-robot setups (e.g., eye-on-base or eye-in-hand). An extensive evaluation based on simulated and real experiments has been performed, proving its good estimation accuracy in terms of reprojection error. The experimental results with real robots show that the proposed method is applicable to relevant industrial contexts and improves the quality and precision of the camera-robot transformation estimation with respect to state-of-the-art approaches.
Daniele Evangelista, Davide Allegro, Matteo Terreran, Alberto Pretto, Stefano Ghidoni
ETFA5
2021 Deep Reinforcement Learning for Motion Planning in Human Robot cooperative Scenarios
abstract
In this paper we tackle motion planning in industrial human-robot cooperative scenarios modeled as a reinforcement learning problem solved in a simulated environment. The agent learns the most effective policy to reach the designated target position while avoiding collisions with a human, performing a pick and place task in the robot workspace, and with fixed obstacles. The policy acts as a feedback motion planner (or reactive motion planner), therefore at each time-step it senses the surrounding environment and computes the action to be performed. In this work a novel formulation of the action that guarantees the trajectory derivatives continuity is proposed to create smooth trajectories that are necessary for maximizing the human trust in the robot. The action is defined as the sub-trajectory the agent must follow for the duration of a time-step, therefore the complete trajectory is the concatenation of all the trajectories computed at each time-step. The proposed method does not require to infer the action the human is currently performing and/or foresee the space occupied by the human. Indeed, during the training phase in a simulated environment the agent experience how the human behaves in the specific scenario, therefore it learns the policy that best adapts to the human actions and movements. The proposed method is finally applied in a scenario of human-robot cooperative pick and place.
Giorgio Nicola, Stefano Ghidoni
ETFA2
2021 Ensemble of convolutional neural networks trained with different activation functions
Gianluca Maguolo, Loris Nanni, Stefano Ghidoni
Expert Syst. Appl.3
2021 Autonomous Learning of the Robot Kinematic Model
abstract
Robotics systems are becoming more and more autonomous and reconfigurable. In this context, the design of algorithms capable of deriving kinematics and dynamics models directly from data could be particularly useful. In this article, we present an algorithm that learns a forward kinematics model of a robot starting from a time series of visual observations. Our strategy can be applied to any robot with serial kinematics composed of revolute and prismatics joints. First, the algorithm identifies the robot kinematic structure, i.e., a high-level description of the robot geometry that defines the connections between the rigid-bodies composing the robot. Then, the algorithm derives the forward kinematics relying on a Gaussian process (GP) model. More precisely, the GP model is based on a polynomial kernel, defined exploiting the kinematic structure previously identified. The effectiveness of the proposed solution has been tested via extensive Monte Carlo simulations, as well as via experiments on a real UR10 robot.
Alberto Dalla Libera, Nicola Castaman, Stefano Ghidoni, Ruggero Carli
IEEE Trans. Robotics3
2020 Enhancing Deep Semantic Segmentation of RGB-D Data with Entangled Forests
abstract
Semantic segmentation is a problem which is getting more and more attention in the computer vision community. Nowadays, deep learning methods represent the state of the art to solve this problem, and the trend is to use deeper networks to get higher performance. The drawback with such models is a higher computational cost, which makes it difficult to integrate them on mobile robot platforms. In this work we want to explore how to obtain lighter deep learning models without compromising performance. To do so we will consider the features used in the 3D Entangled Forests algorithm and we will study the best strategies to integrate these within FuseNet deep network. Such new features allow us to shrink the network size without loosing performance, obtaining hence a lighter model which achieves state-of-the-art performance on the semantic segmentation task and represents an interesting alternative for mobile robotics applications, where computational power and energy are limited.
Matteo Terreran, Elia Bonetto, Stefano Ghidoni
ICPR3
2020 A Control Framework Definition to Overcome Position/Interaction Dynamics Uncertainties in Force-Controlled Tasks
abstract
Within the Industry 4.0 context, industrial robots need to show increasing autonomy. The manipulator has to be able to react to uncertainties/changes in the working environment, displaying a robust behavior. In this paper, a control framework is proposed to perform industrial interaction tasks in uncertain working scenes. The proposed methodology relies on two components: i) a 6D pose estimation algorithm aiming to recognize large and featureless parts; ii) a variable damping impedance controller (inner loop) enhanced by an adaptive saturation PI (outer loop) for high accuracy force control (i.e., zero steady-state force error and force overshoots avoidance). The proposed methodology allows to be robust w.r.t. task uncertainties (i.e. , positioning errors and interaction dynamics). The proposed approach has been evaluated in an assembly task of a side-wall panel to be installed inside the aircraft cabin. As a test platform, the KUKA iiwa 14 R820 has been used together with the Microsoft Kinect 2.0 as RGB-D sensor. Experiments show the reliability in the 6D pose estimation and the high-performance in the force-tracking task, avoiding force overshoots while achieving the tracking of the reference force.
Loris Roveda, Nicola Castaman, Paolo Franceschi, Stefano Ghidoni, Nicola Pedrocchi
ICRA4
2020 Real-time Object Detection using Deep Learning for helping People with Visual Impairments
abstract
Object detection plays a crucial role in the development of Electronic Travel Aids (ETAs), capable to guide a person with visual impairments towards a target object in an unknown indoor environment. In such a scenario, the object detector runs on a mobile device (e.g. smartphone) and needs to be fast, accurate, and, most importantly, lightweight. Nowadays, Deep Neural Networks (DNN) have become the state-of-the-art solution for object detection tasks, with many works improving speed and accuracy by proposing new architectures or extending existing ones. A common strategy is to use deeper networks to get higher performance, but that leads to a higher computational cost which makes it impractical to integrate them on mobile devices with limited computational power. In this work we compare different object detectors to find a suitable candidate to be implemented on ETAs, focusing on lightweight models capable of working in real-time on mobile devices with a good accuracy. In particular, we select two models: SSD Lite with Mobilenet V2 and Tiny-DSOD. Both models have been tested on the popular OpenImage dataset and a new dataset, named L-CAS Office dataset, collected to further test models' performance and robustness in a real scenario inspired by the actual perception challenges of a user with visual impairments.
Matteo Terreran, Andrea G. Tramontano, Jacobus Cornelius Lock, Stefano Ghidoni, Nicola Bellotto
IPAS4
2020 Robotic Object Sorting via Deep Reinforcement Learning: a generalized approach
abstract
This work proposes a general formulation for the Object Sorting problem, suitable to describe any non-deterministic environment characterized by friendly and adversarial interference. Such an approach, coupled with a Deep Reinforcement Learning algorithm, allows training policies to solve different sorting tasks without adjusting the architecture or modifying the learning method. Briefly, the environment is subdivided into a clutter, where objects are freely located, and a set of clusters, where objects should be placed according to predefined ordering and classification rules. A 3D grid discretizes such environment: the properties of an object within a cell depict its state. Such attributes include object category and order. A Markov Decision Process formulates the problem: at each time step, the state of the cells fully defines the environment's one. Users can custom-define object classes, ordering priorities, and failure rules. The latter by assigning a non-uniform risk probability to each cell. Performed experiments successfully trained and validated a Deep Reinforcement Learning model to solve several sorting tasks while minimizing the number of moves and failure probability. Obtained results demonstrate the capability of the system to handle non-deterministic events, like failures, and unpredictable external disturbances, like human user interventions.
Giorgio Nicola, Luca Tagliapietra, Elisa Tosello, Nicolò Navarin, Stefano Ghidoni, Emanuele Menegatti
RO-MAN5
2020 Enhancing semantic segmentation with detection priors and iterated graph cuts for robotics
Morris Antonello, Sabrina Chiesurin, Stefano Ghidoni
Eng. Appl. Artif. Intell.3
2019 Robot Task Planning via Deep Reinforcement Learning: a Tabletop Object Sorting Application
abstract
This paper proposes a Deep Reinforcement Learning powered approach for tabletop object sorting. Once perceived the environment, the system creates a semantic representation of the scene, describing the pose and category of each recognized object. This image is then provided as input to the trained Deep Neural Network in charge of choosing the correct action to be performed to successfully achieve the sorting task. Obtained results prove the capability of the proposed system, including its intrinsic robustness to failures and unpredictable interactions with humans or other environmental agents. Moreover, the use of semantic images makes the Deep Neural Network independent from the type of objects to be sorted and from their final placement location. Finally, the system is scalable, being capable of sorting as many known objects as recognized by the perception system. Currently, the system can sort objects belonging to two predefined categories while treating all the others as obstacles. Future works will extend the system making it capable of sorting potentially any type and number of object categories.
Federico Ceola, Elisa Tosello, Luca Tagliapietra, Giorgio Nicola, Stefano Ghidoni
SMC5
2019 Real-time Tracking-by-Detection of Human Motion in RGB-D Camera Networks
abstract
This paper presents a novel real-time tracking system capable of improving body pose estimation algorithms in distributed camera networks. The first stage of our approach introduces a linear Kalman filter operating at the body joints level, used to fuse single-view body poses coming from different detection nodes of the network and to ensure temporal consistency between them. The second stage, instead, refines the Kalman filter estimates by fitting a hierarchical model of the human body having constrained link sizes in order to ensure the physical consistency of the tracking. The effectiveness of the proposed approach is demonstrated through a broad experimental validation, performed on a set of sequences whose ground truth references are generated by a commercial marker-based motion capture system. The obtained results show how the proposed system outperforms the considered state-of-the-art approaches, granting accurate and reliable estimates. Moreover, the developed methodology constrains neither the number of persons to track, nor the number, position, synchronization, frame-rate, and manufacturer of the RGB-D cameras used. Finally, the real-time performances of the system are of paramount importance for a large number of real-world applications.
Alessandro Malaguti, Marco Carraro, Mattia Guidolin, Luca Tagliapietra, Emanuele Menegatti, Stefano Ghidoni
SMC6
2019 Bioimage Classification with Handcrafted and Learned Features
abstract
Bioimage classification is increasingly becoming more important in many biological studies including those that require accurate cell phenotype recognition, subcellular localization, and histopathological classification. In this paper, we present a new General Purpose (GenP) bioimage classification method that can be applied to a large range of classification problems. The GenP system we propose is an ensemble that combines multiple texture features (both handcrafted and learned descriptors) for superior and generalizable discriminative power. Our ensemble obtains a boosting of performance by combining local features, dense sampling features, and deep learning features. Each descriptor is used to train a different Support Vector Machine that is then combined by sum rule. We evaluate our method on a diverse set of bioimage classification tasks each represented by a benchmark database, including some of those available in the IICBU 2008 database. Each bioimage classification task represents a typical subcellular, cellular, and tissue level classification problem. Our evaluation on these datasets demonstrates that the proposed GenP bioimage ensemble obtains state-of-the-art performance without any ad-hoc dataset tuning of the parameters (thereby avoiding any risk of overfitting/overtraining). To reproduce the experiments reported in this paper, the MATLAB code of all the descriptors is available at https://github.com/LorisNanni and https://www.dropbox.com/s/bguw035yrqz0pwp/ElencoCode.docx?dl=0.
Loris Nanni, Sheryl Brahnam, Stefano Ghidoni, Alessandra Lumini
IEEE ACM Trans. Comput. Biol. Bioinform.3
2018 Multi-View 3D Entangled Forest for Semantic Segmentation and Mapping
abstract
Applications that provide location related services need to understand the environment in which humans live such that verbal references and human interaction are possible. We formulate this semantic labelling task as the problem of learning the semantic labels from the perceived 3D structure. In this contribution we propose a batch approach and a novel multi-view frame fusion technique to exploit multiple views for improving the semantic labelling results. The batch approach works offline and is the direct application of an existing single-view method to scene reconstructions with multiple views. The multi-view frame fusion works in an incremental fashion accumulating the single-view results, hence allowing the online multi-view semantic segmentation of single frames and the offline reconstruction of semantic maps. Our experiments show the superiority of the approaches based on our fusion scheme, which leads to a more accurate semantic labelling.
Morris Antonello, Daniel Wolf, Johann Prankl, Stefano Ghidoni, Emanuele Menegatti, Markus Vincze
ICRA4
2018 Human-Robot Cooperative Interaction Control for the Installation of Heavy and Bulky Components
abstract
The paper describes a human-robot cooperative installation methodology of heavy and bulky components based on marker-based visual servoing, force control, and human-robot cooperation. The main advance in the human-robot cooperation is achieved by a shared-control of the interaction during the installation task, relieving the human operator by the manipulated load and giving to the robot a partially autonomous behaviour in the force-tracking direction. Experimental results are shown in the context of the H2020 CleanSky 2 EURECA project in which a side-wall panel is installed in a 1:1 scale mock-up scenario of an A320 plane fuselage environment.
Loris Roveda, Nicola Castaman, Stefano Ghidoni, Paolo Franceschi, Nicoló Boscolo, Enrico Pagello, Nicola Pedrocchi
SMC3
2017 An ensemble of visual features for Gaussians of local descriptors and non-binary coding for texture descriptors
Loris Nanni, Michelangelo Paci, Sheryl Brahnam, Stefano Ghidoni
Expert Syst. Appl.4
2017 Handcrafted vs. non-handcrafted features for computer vision classification
Loris Nanni, Stefano Ghidoni, Sheryl Brahnam
Pattern Recognit.2
2017 How could a subcellular image, or a painting by Van Gogh, be similar to a great white shark or to a pizza?
Loris Nanni, Stefano Ghidoni
Pattern Recognit. Lett.2
2015 Performance evaluation of the 1st and 2nd generation Kinect for multimedia applications
abstract
Microsoft Kinect had a key role in the development of consumer depth sensors being the device that brought depth acquisition to the mass market. Despite the success of this sensor, with the introduction of the second generation, Microsoft has completely changed the technology behind the sensor from structured light to Time-Of-Flight. This paper presents a comparison of the data provided by the first and second generation Kinect in order to explain the achievements that have been obtained with the switch of technology. After an accurate analysis of the accuracy of the two sensors under different conditions, two sample applications, i.e., 3D reconstruction and people tracking, are presented and used to compare the performance of the two sensors.
Simone Zennaro, Matteo Munaro, Simone Milani, Pietro Zanuttigh, Andrea Bernardi, Stefano Ghidoni, Emanuele Menegatti
ICME6
2015 Improving the descriptors extracted from the co-occurrence matrix using preprocessing approaches
Loris Nanni, Sheryl Brahnam, Stefano Ghidoni, Emanuele Menegatti
Expert Syst. Appl.3
2015 Automatic Color Inspection for Colored Wires in Electric Cables
abstract
In this paper, an automatic optical inspection system for checking the sequence of colored wires in electric cable is presented. The system is able to inspect cables with flat connectors differing in the type and number of wires. This variability is managed in an automatic way by means of a self-learning subsystem and does not require manual input from the operator or loading new data to the machine. The system is coupled to a connector crimping machine and once the model of a correct cable is learned, it can automatically inspect each cable assembled by the machine. The main contributions of this paper are: (i) the self-learning system; (ii) a robust segmentation algorithm for extracting wires from images even if they are strongly bent and partially overlapped; and (iii) a color recognition algorithm able to cope with highlights and different finishing of the wire insulation. We report the system evaluation over a period of several months during the actual production of large batches of different cables; tests demonstrated a high level of accuracy and the absence of false negatives, which is a key point in order to guarantee defect-free productions.
Stefano Ghidoni, Matteo Finotto, Emanuele Menegatti
IEEE Trans Autom. Sci. Eng.1
2014 A feature-based approach to people re-identification using skeleton keypoints
abstract
In this paper we propose a novel methodology for people re-identification based on skeletal information. Features are evaluated on the skeleton joints and a highly distinctive and compact feature-based signature is generated for each user by concatenating descriptors of all visible joints. We compared a number of state-of-the-art 2D and 3D feature descriptors to be used with our signature on two newly acquired public datasets for people re-identification with RGB-D sensors. Moreover, we tested our approach against the best re-identification methods in the literature and on a widely used public video surveillance dataset. Our approach proved to be robust to strong illumination changes and occlusions. It achieved very high performance also on low resolution images, overcoming state-of-the-art methods in terms of recognition accuracy and efficiency. These features make our approach particularly suited for mobile robotics.
Matteo Munaro, Stefano Ghidoni, Deniz Tartaro Dizmen, Emanuele Menegatti
ICRA2
2013 A comparison of methods for extracting information from the co-occurrence matrix for subcellular classification
Loris Nanni, Sheryl Brahnam, Stefano Ghidoni, Emanuele Menegatti, Tonya Barrier
Expert Syst. Appl.3
2011 Self-learning visual inspection system for cable crimping machines
abstract
This paper presents a system for checking connectors while cables are being crimped to them. The system verifies that the wires color sequence is correct: an accurate color analysis technique has then been developed, in order to discriminate between similar colors, and filter noise factors.
Stefano Ghidoni, Matteo Finotto, Emanuele Menegatti
ICRA1
2010 Cooperative tracking of moving objects and face detection with a dual camera sensor
abstract
This paper describes a sensor for autonomous surveillance capable of continuously monitoring the environment, while acquiring detailed images of specific areas. This is achieved by exploiting an omnidirectional camera and a PTZ camera, assembled together on a single mount. The two cameras form a single vision sensor, since data obtained processing the two images are used in a cooperative way. This system solves the problem affecting systems based on PTZ cameras only, since it does not exist a tracking system working reliably on PTZ images: the problem is solved here by performing the tracking in the omnidirectional image. This vision sensor is used in a surveillance application, that detects moving objects, and records all the faces of the people walking close to the sensor. It could be used to navigate or instruct a security or a service mobile robot.
Stefano Ghidoni, Alberto Pretto, Emanuele Menegatti
ICRA1
2010 TerraMax Vision at the Urban Challenge 2007
abstract
This paper presents the TerraMax vision systems used during the 2007 DARPA Urban Challenge. First, a description of the different vision systems is provided, focusing on their hardware configuration, calibration method, and tasks. Then, each component is described in detail, focusing on the algorithms and sensor fusion opportunities: obstacle detection, road marking detection, and vehicle detection. The conclusions summarize the lesson learned from the developing of the passive sensing suite and its successful fielding in the Urban Challenge.
Alberto Broggi, Andrea Cappalunga, Claudio Caraffi, Stefano Cattani, Stefano Ghidoni, Paolo Grisleri, Pier Paolo Porta, Matteo Posterli, Paolo Zani
IEEE Trans. Intell. Transp. Syst.5
2009 A New Approach to Urban Pedestrian Detection for Automatic Braking
abstract
This paper presents an application of a pedestrian-detection system aimed at localizing potentially dangerous situations under specific urban scenarios. The approach used in this paper differs from those implemented in traditional pedestrian-detection systems, which are designed to localize all pedestrians in the area in front of the vehicle. Conversely, this approach searches for pedestrians in critical areas only. The environment is reconstructed with a standard laser scanner, whereas the following check for the presence of pedestrians is performed due to the fusion with a vision system. The great advantages of such an approach are that pedestrian recognition is performed on limited image areas, therefore boosting its timewise performance, and no assessment on the danger level is finally required before providing the result to either the driver or an onboard computer for automatic maneuvers. A further advantage is the drastic reduction of false alarms, making this system robust enough to control nonreversible safety systems.
Alberto Broggi, Pietro Cerri, Stefano Ghidoni, Paolo Grisleri, Ho Gi Jung
IEEE Trans. Intell. Transp. Syst.3