Gabriele Costante

dblp:123/6662 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-8417-9372ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 3 first-author · 8 since 2021Systems, architecture and hardware · 10 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
YearPublicationVenuePosition
2025 Active Illumination for Visual Ego-Motion Estimation in the Dark
abstract
Visual Odometry (VO) and Visual SLAM (VSLAM) systems often struggle in low-light and dark environments due to the lack of robust visual features. In this paper, we propose a novel active illumination framework to enhance the performance of VO and V-SLAM algorithms in these challenging conditions. The developed approach dynamically controls a moving light source to illuminate highly textured areas, thereby improving feature extraction and tracking. Specifically, a detector block, which incorporates a deep learning-based enhancing network, identifies regions with relevant features. Then, a pan-tilt controller is responsible for guiding the light beam toward these areas, so that to provide information-rich images to the ego-motion estimation algorithm. Experimental results on a real robotic platform demonstrate the effectiveness of the proposed method, showing a reduction in the pose estimation error up to 75 % with respect to a traditional fixed lighting technique.
Francesco Crocetti, Alberto Dionigi, Raffaele Brilli, Gabriele Costante, Paolo Valigi
ICRA4
2024 Infrastructure-less UWB-based Active Relative Localization
abstract
In multi-robot systems, relative localization between platforms plays a crucial role in many tasks, such as leader following, target tracking, or cooperative maneuvering. State of the Art (SotA) approaches either rely on infrastructure-based or on infrastructure-less setups. The former typically achieve high localization accuracy but require fixed external structures. The latter provide more flexibility, however, most of the works use cameras or lidars that require Line-of-Sight (LoS) to operate. Ultra Wide Band (UWB) devices are emerging as a viable alternative to build infrastructure-less solutions that do not require LoS. These approaches directly deploy the UWB sensors on the robots. However, they require that at least one of the platforms is static, limiting the advantages of an infrastructure-less setup. In this work, we remove this constraint and introduce an active method for infrastructureless relative localization. Our approach allows the robot to adapt its position to minimize the relative localization error of the other platform. To this aim, we first design a specialized anchor placement for the active localization task. Then, we propose a novel UWB Relative Localization Loss that adapts the Geometric Dilution Of Precision metric to the infrastructureless scenario. Lastly, we leverage this loss function to train an active Deep Reinforcement Learning-based controller for UWB relative localization. An extensive simulation campaign and real-world experiments validate our method, showing up to a 60% reduction of the localization error compared to current SotA approaches.
Valerio Brunacci, Alberto Dionigi, Alessio De Angelis, Gabriele Costante
IROS4
2024 The Power of Input: Benchmarking Zero-Shot Sim-to-Real Transfer of Reinforcement Learning Control Policies for Quadrotor Control
abstract
In the last decade, data-driven approaches have become popular choices for quadrotor control, thanks to their ability to facilitate the adaptation to unknown or uncertain flight conditions. Among the different data-driven paradigms, Deep Reinforcement Learning (DRL) is currently one of the most explored. However, the design of DRL agents for Micro Aerial Vehicles (MAVs) remains an open challenge. While some works have studied the output configuration of these agents (i.e., what kind of control to compute), there is no general consensus on the type of input data these approaches should employ. Multiple works simply provide the DRL agent with full state information, without questioning if this might be redundant and unnecessarily complicate the learning process, or pose superfluous constraints on the availability of such information in real platforms. In this work, we provide an in-depth benchmark analysis of different configurations of the observation space. We optimize multiple DRL agents in simulated environments with different input choices and study their robustness and their sim-to-real transfer capabilities with zero-shot adaptation. We believe that the outcomes and discussions presented in this work supported by extensive experimental results could be an important milestone in guiding future research on the development of DRL agents for aerial robot tasks.
Alberto Dionigi, Gabriele Costante, Giuseppe Loianno
IROS2
2024 LF2SLAM: Learning-based Features For visual SLAM
abstract
Autonomous robot navigation relies on the robot’s ability to understand its environment for localization, typically using a Visual Simultaneous Localization And Mapping (SLAM) algorithm that processes image sequences. While state-of-the-art methods have shown remarkable performance, they still have limitations. Geometric VO algorithms that leverage hand-crafted feature extractors require careful hyper-parameter tuning. Conversely, end-to-end data-driven VO algorithms suffer from limited generalization capabilities and require large datasets for their proper optimizations. Recently, promising results have been shown by hybrid approaches that integrate robust data-driven feature extraction with the geometric estimation pipeline. In this work, we follow these intuitions and propose a hybrid VO method, namely Learned Features For SLAM (LF2SLAM), that combines a deep neural network for feature extraction with a standard VO pipeline. The network is trained in a data-driven framework that includes a pose estimation component to learn feature extractors that are tailored for VO tasks. A novel loss function modification is introduced, using a binary mask that considers only the informative features. The experimental evaluation performed shows that our approach has remarkable generalization capabilities in scenarios that differ from those used for training. Furthermore, LF2SLAM exhibits robustness in more challenging scenarios, i.e., characterized by the presence of poor lighting and low amount of texture, with respect to the state-of-the-art ORB-SLAM3 algorithm.
Marco Legittimo, Francesco Crocetti, Mario Luca Fravolini, Giuseppe Mollica, Gabriele Costante
IROS5
2024 MA-VIED: A Multisensor Automotive Visual Inertial Event Dataset
abstract
Visual Inertial Odometry (VIO) and Simultaneous Localization and Mapping (SLAM) have experienced increasing interest in both the consumer and racing automotive sectors in recent decades. With the introduction of novel neuromorphic vision sensors, it is now possible to accurately localize a vehicle even under complex environmental conditions, leading to an improved and safer driving experience. In this paper, we propose MA-VIED, a large-scale driving dataset that collects race track-like loops, maneuvers, and standard driving scenarios, all bundled in a rich sensory dataset. MA-VIED provides highly accurate IMU data, standard and event camera streams, and RTK position data from a dual GPS antenna, both of which are hardware-synchronized with all cameras and IMU data. In addition, we collect accurate wheel odometry data and other data from the vehicle’s CAN bus. The dataset contains 13 sequences collected in urban, suburban, and racetrack-like environments with varying lighting conditions and driving dynamics. We provide ground-truth RTK data for algorithms evaluation and the calibration sequences for both IMU and cameras. We then present three tests to demonstrate how MA-VIED can be suitable for monocular VIO applications, using state-of-the-art VIO algorithms and an EKF-based sensor fusion solution. The experimental results show that MA-VIED can support the development and prototyping of novel automotive-oriented frame and event-based monocular VIO algorithms.
Giuseppe Mollica, Simone Felicioni, Marco Legittimo, Leonardo Meli, Gabriele Costante, Paolo Valigi
IEEE Trans. Intell. Transp. Syst.5
2023 Monocular Reactive Collision Avoidance for MAV Teleoperation with Deep Reinforcement Learning
abstract
Enabling Micro Aerial Vehicles (MAVs) with semi-autonomous capabilities to assist their teleoperation is crucial in several applications. Remote human operators do not have, in general, the situational awareness to perceive obstacles near the drone, nor the readiness to provide commands to avoid collisions. In this work, we devise a novel teleoperation setting that asks the operator to provide a simple high-level signal encoding the speed and the direction they expect the drone to follow. We then endow the MAV with an end-to-end Deep Reinforcement Learning (DRL) model that computes control commands to track the desired trajectory while performing collision avoidance. Differently from State-of-the-Art (SotA) works, it allows the robot to move freely in the 3D space, requires only the current RGB image captured by a monocular camera and the current robot position, and does not make any assumption about obstacle shape and size. We show the effectiveness and the generalization capabilities of our strategy by comparing it against a SotA baseline in photorealistic simulated environments.
Raffaele Brilli, Marco Legittimo, Francesco Crocetti, Mirko Leomanni, Mario Luca Fravolini, Gabriele Costante
ICRA6
2023 GaPT: Gaussian Process Toolkit for Online Regression with Application to Learning Quadrotor Dynamics
abstract
Gaussian Processes (GPs) are expressive models for capturing signal statistics and expressing prediction uncer-tainty. As a result, the robotics community has gathered interest in leveraging these methods for inference, planning, and control. Unfortunately, despite providing a closed-form inference solution, GPs are non-parametric models that typically scale cubically with the dataset size, hence making them difficult to be used especially on onboard Size, Weight, and Power (SWaP) constrained aerial robots. In addition, the integration of popular libraries with GPs for different kernels is not trivial. In this paper, we propose GaPT, a novel toolkit that converts GPs to their state space form and performs regression in linear time. GaPT is designed to be highly compatible with several optimizers popular in robotics. We thoroughly validate the proposed approach for learning quadrotor dynamics on both single and multiple input GP settings. GaPT accurately captures the system behavior in multiple flight regimes and operating conditions, including those producing highly nonlin-ear effects such as aerodynamic forces and rotor interactions. Moreover, the results demonstrate the superior computational performance of GaPT compared to a classical GP inference approach on both single and multi-input settings especially when considering large number of data points, enabling real-time regression speed on embedded platforms used on SWaP-constrained aerial robots.
Francesco Crocetti, Jeffrey Mao, Alessandro Saviolo, Gabriele Costante, Giuseppe Loianno
ICRA4
2023 Data-driven and uncertainty-aware robust airstrip surface estimation
abstract
Abstract The performances of aircraft braking control systems are strongly influenced by the tire friction force experienced during the braking phase. The availability of an accurate estimate of the current airstrip characteristics is a recognized issue for developing optimized braking control schemes. The study presented in this paper is focused on the robust online estimation of the airstrip characteristics from sensory data usually available on an aircraft. In order to capture the nonlinear dependency of the current best slip on sequential slip-friction measurements acquired during the braking maneuver, multilayer perceptron (MLP) approximators have been proposed. The MLP training is based on a synthetic data set derived from a widely used tire–road friction model. In order to achieve robust predictions, MLP architectures based on the drop-out mechanism have been applied not only in the offline training phase but also during the braking. This allowed to online compute a confidence interval measure for best friction estimate that has been exploited to refine the estimation via Kalman Filtering. Open loop and closed loop simulation studies in 15 representative airstrip scenarios (with multiple surface transitions) have been performed to evaluate the performance of the proposed robust estimation method in terms of estimation error, aircraft braking distance, and time, together with a quantitative comparison with a state-of-the-art benchmark approach.
Francesco Crocetti, Mario Luca Fravolini, Gabriele Costante, Paolo Valigi
Neural Comput. Appl.3
2022 A novel vision-based weakly supervised framework for autonomous yield estimation in agricultural applications
Enrico Bellocchio, Francesco Crocetti, Gabriele Costante, Mario Luca Fravolini, Paolo Valigi
Eng. Appl. Artif. Intell.3
2020 The Role of the Input in Natural Language Video Description
abstract
Natural language video description (NLVD) has recently received strong interest in the computer vision, natural language processing (NLP), multimedia, and autonomous robotics communities. The state-of-the-art (SotA) approaches obtained remarkable results when tested on the benchmark datasets. However, those approaches poorly generalize to new datasets. In addition, none of the existing works focus on the processing of the input to the NLVD systems, which is both visual and textual. In this paper, an extensive study is presented to deal with the role of the visual input, evaluated with respect to the overall NLP performance. This is achieved by performing data augmentation of the visual component, applying common transformations to model camera distortions, noise, lighting, and camera positioning that are typical in real-world operative scenarios. A t-SNE-based analysis is proposed to evaluate the effects of the considered transformations on the overall visual data distribution. For this study, the English subset of the Microsoft Research Video Description (MSVD) dataset is considered, which is used commonly for NLVD. It was observed that this dataset contains a relevant amount of syntactic and semantic errors. These errors have been amended manually, and the new version of the dataset (called MSVD-v2) is used in the experimentation. The MSVD-v2 dataset is released to help to gain insight into the NLVD problem.
Silvia Cascianelli, Gabriele Costante, Alessandro Devo, Thomas A. Ciarfuglia, Paolo Valigi, Mario Luca Fravolini
IEEE Trans. Multim.2
2020 Uncertainty Estimation for Data-Driven Visual Odometry
abstract
Over the past few years, we have witnessed a considerable diffusion of data-driven visual odometry (VO) approaches as viable alternatives to standard geometric-based strategies. Their success is mainly related to the improved robustness to image nonideal conditions (e.g., blur, high or low contrast, texture-poor scenarios). However, most of the data-driven State-of-the-Art (SotA) approaches do not provide any kind of information about the uncertainty of their estimates, which is crucial to effectively integrate them into robotic navigation systems. Inspired by this considerations, we propose uncertainty-aware VO (UA-VO), a novel deep neural network (DNN) architecture that computes relative pose predictions by processing sequence of images and, at the same time, provides uncertainty measures about those estimations. The confidence measure computed by UA-VO considers both epistemic and aleatoric uncertainties and accounts for heteroscedasticity, i.e., it is sample-dependent. We assess the benefits of UA-VO with different typology of experiments on three publicly available datasets and on a brand new set of sequences, we gathered to extend the evaluation.
Gabriele Costante, Michele Mancini
IEEE Trans. Robotics1
2020 Towards Generalization in Target-Driven Visual Navigation by Using Deep Reinforcement Learning
abstract
Among the main challenges in robotics, target-driven visual navigation has gained increasing interest in recent years. In this task, an agent has to navigate in an environment to reach a user specified target, only through vision. Recent fruitful approaches rely on deep reinforcement learning, which has proven to be an effective framework to learn navigation policies. However, current state-of-the-art methods require to retrain, or at least fine-tune, the model for every new environment and object. In real scenarios, this operation can be extremely challenging or even dangerous. For these reasons, we address generalization in target-driven visual navigation by proposing a novel architecture composed of two networks, both exclusively trained in simulation. The first one has the objective of exploring the environment, while the other one of locating the target. They are specifically designed to work together, while separately trained to help generalization. In this article, we test our agent in both simulated and real scenarios, and validate its generalization capabilities through extensive experiments with previously unseen goals and unknown mazes, even much larger than the ones used for training.
Alessandro Devo, Giacomo Mezzetti, Gabriele Costante, Mario Luca Fravolini, Paolo Valigi
IEEE Trans. Robotics3
2016 Fast robust monocular depth estimation for Obstacle Detection with fully convolutional networks
abstract
Obstacle Detection is a central problem for any robotic system, and critical for autonomous systems that travel at high speeds in unpredictable environment. This is often achieved through scene depth estimation, by various means. When fast motion is considered, the detection range must be longer enough to allow for safe avoidance and path planning. Current solutions often make assumption on the motion of the vehicle that limit their applicability, or work at very limited ranges due to intrinsic constraints. We propose a novel appearance-based Object Detection system that is able to detect obstacles at very long range and at a very high speed (~ 300Hz), without making assumptions on the type of motion. We achieve these results using a Deep Neural Network approach trained on real and synthetic images and trading some depth accuracy for fast, robust and consistent operation. We show how photo-realistic synthetic images are able to solve the problem of training set dimension and variety typical of machine learning approaches, and how our system is robust to massive blurring of test images.
Michele Mancini, Gabriele Costante, Paolo Valigi, Thomas A. Ciarfuglia
IROS2
2015 Exploiting Photometric Information for Planning Under Uncertainty
Gabriele Costante, Jeffrey A. Delmerico, Manuel Werlberger, Paolo Valigi, Davide Scaramuzza 0001
ISRR (1)1
2014 Exploiting transfer learning for personalized view invariant gesture recognition
abstract
A robust gesture recognition system is an essential component in many human-computer interaction applications. In particular, the widespread adoption of portable devices and the diffusion of autonomous systems with limited power and load capacity has increased the need of developing efficient recognition algorithms which operates on video streams recorded from low cost devices and which can cope with the challenging issue of point of view changes. A further challenge arises as different users tend to perform the same gesture with different styles and speeds. Thus a classifier trained with gestures data of certain set of users may work poorly when data from other users are being processed. However, as often a mobile device or a robot are intended to be used by a single or by a small group of people, it would be desirable to have a gesture recognition system designed specifically for these users. In this paper we introduce a novel approach to face the problems of view-invariance and user personalization in the context of gesture interaction systems. More specifically, we propose a domain adaptation framework based on a feature space augmentation approach operating on robust view-invariant Self Similarity Matrix descriptors. To prove the effectiveness of our method a dataset corresponding to 17 users performing 10 different gestures under 3 point of views is collected and an extensive experimental evaluation is performed.
Gabriele Costante, Valerio Galieni, Yan Yan 0002, Mario Luca Fravolini, Elisa Ricci 0001, Paolo Valigi
ICASSP1
2014 Personalizing vision-based gestural interfaces for HRI with UAVs: a transfer learning approach
abstract
Following recent works on HRI for UAVs, we present a gesture recognition system which operates on the video stream recorded from a passive monocular camera installed on a quadcopter. While many challenges must be addressed for building a real-time vision-based gestural interface, in this paper we specifically focus on the problem of user personalization. Different users tend to perform the same gesture with different styles and speed. Thus, a system trained on visual sequences depicting some users may work poorly when data from other people are available. On the other hand, collecting and annotating many user-specific data is time consuming. To avoid these issues, in this paper we propose a personalized gestural interface. We introduce a novel transfer learning algorithm which, exploiting both data downloaded from the web and gestures collected from other users, permits to learn a set of person-specific classifiers. We integrate the proposed gesture recognition module into a HRI system with a flying quadrotor robot. In our system first the UAV localizes a person and individuates her identity. Then, when a user performs a specific gesture, the system recognizes it adopting the associated user-specific classifier and the quadcopter executes the corresponding task. Our experimental evaluation demonstrates that the proposed personalized gesture recognition solution is advantageous with respect to generic ones.
Gabriele Costante, Enrico Bellocchio, Paolo Valigi, Elisa Ricci 0001
IROS1
2013 A transfer learning approach for multi-cue semantic place recognition
abstract
As researchers are striving for developing robotic systems able to move into the `the wild', the interest towards novel learning paradigms for domain adaptation has increased. In the specific application of semantic place recognition from cameras, supervised learning algorithms are typically adopted. However, once learning has been performed, if the robot is moved to another location, the acquired knowledge may be not useful, as the novel scenario can be very different from the old one. The obvious solution would be to retrain the model updating the robot internal representation of the environment. Unfortunately this procedure involves a very time consuming data-labeling effort at the human side. To avoid these issues, in this paper we propose a novel transfer learning approach for place categorization from visual cues. With our method the robot is able to decide automatically if and how much its internal knowledge is useful in the novel scenario. Differently from previous approaches, we consider the situation where the old and the novel scenario may differ significantly (not only the visual room appearance changes but also different room categories are present). Importantly, our approach does not require labeling from a human operator. We also propose a strategy for improving the performance of the proposed method by fusing two complementary visual cues. Our extensive experimental evaluation demonstrates the advantages of our approach on several sequences from publicly available datasets.
Gabriele Costante, Thomas A. Ciarfuglia, Paolo Valigi, Elisa Ricci 0001
IROS1
2012 A discriminative approach for appearance based loop closing
abstract
The place recognition module is a fundamental component in SLAM systems, as incorrect loop closures may result in severe errors in trajectory estimation. In the case of appearance-based methods the bag-of-words approach is typically employed for recognizing locations. This paper introduces a novel algorithm for improving loop closures detection performance by adopting a set of visual words weights, learned offline accordingly to a discriminative criterion. The proposed weights learning approach, based on the large margin paradigm, can be used for generic similarity functions and relies on an efficient online leaning algorithm in the training phase. As the computed weights are usually very sparse, a gain in terms of computational cost at recognition time is also obtained. Our experiments, conducted on publicly available datasets, demonstrate that the discriminative weights lead to loop closures detection results that are more accurate than the traditional bag-of-words method and that our place recognition approach is competitive with state-of-the-art methods.
Thomas A. Ciarfuglia, Gabriele Costante, Paolo Valigi, Elisa Ricci 0001
IROS2