Riccardo Monica

dblp:133/4772 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
6since 2021 · last 2025
0000-0002-1262-6348ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2025 Adaptive Complementary Filter for Hybrid Inside-Out Outside-In HMD Tracking With Smooth Transitions
abstract
Head-mounted displays (HMDs) in room-scale virtual reality are usually tracked using inside-out visual SLAM algorithms. Alternatively, to track the motion of the HMD with respect to a fixed real-world reference frame, an outside-in instrumentation like a motion capture system can be adopted. However, outside-in tracking systems may temporarily lose tracking as they suffer by occlusion and blind spots. A possible solution is to adopt a hybrid approach where the inside-out tracker of the HMD is augmented with an outside-in sensing system. On the other hand, when the tracking signal of the outside-in system is recovered after a loss of tracking the transition from inside-out tracking to hybrid tracking may generate a discontinuity, i.e a sudden change of the virtual viewpoint, that can be uncomfortable for the user. Therefore, hybrid tracking solutions for HMDs require advanced sensor fusion algorithms to obtain a smooth transition. This work proposes a method for hybrid tracking of a HMD with smooth transitions based on an adaptive complementary filter. The proposed approach can be configured with several parameters that determine a trade-off between user experience and tracking error. A user study was carried out in a room-scale virtual reality environment, where users carried out two different tasks while multiple signal tracking losses of the outside-in sensor system occurred. The results show that the proposed approach improves user experience compared to a standard Extended Kalman Filter, and that tracking error is lower compared to a state-of-the-art complementary filter when configured for the same quality of user experience.
Riccardo Monica, Dario Lodi Rizzini, Jacopo Aleotti
IEEE Trans. Vis. Comput. Graph.1
2024 Robot Manipulation of Tomato Fruits using a Commercial Soft Gripper
abstract
This paper presents a robot system for manipulation and picking of tomato fruits guided by computer vision. A CNN algorithm trained on a custom image dataset is used for fruit detection and the depth camera enables position estimation. The robot manipulator is equipped with a commercial soft gripper for picking fragile objects. Three planning procedures have been proposed to successfully reach and grasp the tomatoes. Experiments on simulated crops have compared the effectiveness of the proposed procedures.
Riccardo Monica, Dario Lodi Rizzini, Stefano Caselli
ETFA1
2024 Contact-Based in-Hand Package Pose Estimation Using a Collaborative Robot
abstract
In automated robot assembly and industrial palletization tasks it is crucial to ensure a good accuracy while placing objects given a planned target pose. To achieve this goal post-grasp strategies may be adopted that estimate or correct the displacement error between the expected and the actual grasp pose of an object. Standard in-hand post-grasp strategies require sensors like cameras to estimate the displacement error while the object is grasped. Other approaches are based on object re-grasping using special jigs and fixtures. In this paper a novel post-grasp strategy is proposed, where the displacement error is estimated in-hand by detecting collisions between the grasped object and a fixed peg. The proposed method estimates the displacement error after few collisions. The approach was evaluated on cardboard boxes thanks to the internal forcetorque sensor of a collaborative robot, achieving sub-millimeter and sub-degree residual placement errors.
Alessio Saccuti, Riccardo Monica, Jacopo Aleotti
ETFA2
2022 Detection of Unsorted Metal Components for Robot Bin Picking Using an Inexpensive RGB-D Sensor
abstract
This work investigates the problem of 6D pose estimation and robot bin picking of non-Lambertian reflecting objects based on a low-cost commercial 3D sensor. In particular, we address the task of estimating the pose of small metal hydraulic components of the same type, randomly placed in a bin. The system consists of a robot arm and an RGB-D sensor in eye-in-hand configuration. The proposed method works in two main phases. In the first phase a Convolutional Neural Network (CNN) extracts the bounding boxes of the objects contained in the bin from a single RGB image of the environment. In the second phase the 6D pose of the objects is estimated using a dense 3D reconstruction of the scene and by applying a template matching algorithm from multiple virtual views of the object CAD model. Experimental results have been carried out on a dataset containing both RGB and depth images. Preliminary experiments are also reported in the real setup.
Riccardo Monica, Alessio Saccuti, Jacopo Aleotti, Marco Lippi 0001
ETFA1
2022 Prediction of Depth Camera Missing Measurements Using Deep Learning for Next Best View Planning
abstract
Depth images usually contain pixels with invalid measurements. This paper presents a deep learning approach that receives as input a partially-known volumetric model of the environment and a camera pose, and it predicts the probability that a pixel would contain a valid depth measurement if a camera was placed at the given pose. The proposed network architecture consists of a 3D Convolutional Neural Network (CNN) module and a 2D CNN module, connected by a deep learning attention-based projection module. The method was integrated into a CNN-based probabilistic Next Best View plan-ner, resulting in a more realistic prediction of the information gain for each possible viewpoint with respect to state of the art approaches. Experiments were carried out in tabletop scenarios using a robot manipulator with an eye-in-hand depth camera.
Riccardo Monica, Jacopo Aleotti
ICRA1
2021 Multi-Robot Multiple Camera People Detection and Tracking in Automated Warehouses
abstract
In this work a multi-robot system is presented for people detection and tracking in automated warehouses. Each Automated Guided Vehicle (AGV) is equipped with multiple RGB cameras that can track the workers’ current locations on the floor thanks to a neural network that provides human pose estimation. Based on the local perception of the environment each AGV can exploit information about the tracked people for self-motion planning or collision avoidance.Additionally, data collected from each robot contributes to a global people detection and tracking system. A warehouse central management software fuses information received from all AGVs into a map of the current locations of workers. The estimated locations of workers are sent back to the AGVs to prevent potential collision. The proposed method is based on two-level hierarchy of Kalman filters. Experiments performed in a real warehouse show the viability of the proposed approach.
Michela Zaccaria, Mikhail Giorgini, Riccardo Monica, Jacopo Aleotti
INDIN3
2020 Integration of a Multi-Camera Vision System and Admittance Control for Robotic Industrial Depalletizing
abstract
This work addresses the task of robot depalletizing by means of a mobile manipulator, taking into account the problem of localizing the boxes to be removed from the pallet and a manipulation strategy that allows to pull the boxes without lifting them with the robot arm. The depalletizing task is of particular interest in the industrial scenario in order to increase efficiency, flexibility and economic affordability of automatic warehouses.The proposed solution makes use of a multi-sensor vision system and a force-controlled collaborative robot in order to detect the boxes on the pallet and to control the robot interaction with the boxes to be removed. The vision system comprises a fixed 3D Time-of-flight camera and an eye-in-hand 2D camera. Preliminary experimental results performed on a laboratory setup with a fixed-based robotic manipulator are reported to show the effectiveness of the perception and control system.
Davide Chiaravalli, Gianluca Palli, Riccardo Monica, Jacopo Aleotti, Dario Lodi Rizzini
ETFA3
2020 Surfel-Based Incremental Reconstruction of the Boundary Between Known and Unknown Space
abstract
This article presents the first surfel-based method for multi-view 3D reconstruction of the boundary between known and unknown space. The proposed approach integrates multiple views from a moving depth camera and it generates a set of surfels that encloses observed empty space, i.e., it models both the boundary between empty and occupied space, and the boundary between empty and unknown space. One novelty of the method is that it does not require a persistent voxel map of the environment to distinguish between unknown and empty space. The problem is solved thanks to an incremental algorithm that computes the Boolean union of two surfel bounded volumes: the known volume from previous frames and the space observed from the current depth image. A number of strategies were developed to cope with errors in surfel position and orientation. The method, implemented on CPU and GPU, was evaluated on real data acquired in indoor scenarios, and it was compared against state of the art approaches. Results show that the proposed method has a low number of false positive and false negatives, it is faster than a standard volumetric algorithm, it has a lower memory consumption, and it scales better in large environments.
Riccardo Monica, Jacopo Aleotti
IEEE Trans. Vis. Comput. Graph.1
2019 Humanoid Robot Next Best View Planning Under Occlusions Using Body Movement Primitives
abstract
This work presents an approach for humanoid Next Best View (NBV) planning that exploits full body motions to observe objects occluded by obstacles. The task is to explore a given region of interest in an initially unknown environment. The robot is equipped with a depth sensor, and it can perform both 2D and 3D mapping. As main contribution with respect to previous work, the proposed method does not rely on simple motions of the head and it was evaluated in real environments. The robot is guided by two behaviors: a target behavior that aims at observing the region of interest by exploiting body movements primitives, and an exploration behavior that aims at observing other unknown areas. Experiments show that the humanoid is able to peer around obstacles to reach a favourable point of view. Moreover, the proposed approach results in a more complete reconstruction of objects than a conventional algorithm that only changes the orientation of the head.
Riccardo Monica, Jacopo Aleotti
IROS1
2019 A Wave-Based Request-Response Protocol for Latency Minimization in WSNs
abstract
Transmission latency is a key performance metrics in most wireless sensor network (WSN) applications. Nodes in a WSN often keep their radio transceivers off, and turn them on periodically using a duty cycling mechanism. The latter is a major source of delay in the network, because transmissions must wait for the next receiver wake-up. In this paper, we present a cross-layer approach to minimize latency of a request-response (RR) protocol adopted in an IEEE 802.15.4-based WSN where the IPv6 routing protocol for low-power and lossy networks (RPLs) is used. Extra wake-ups are generated dynamically to match the predicted arrival time of the response packet, in order to reduce the duty cycling delay. The proposed approach is verified with the Cooja simulator, relying on the Contiki operating system (OS). The observed experimental results show a shorter RR delay with respect to a phase alignment (PA) approach.
Riccardo Monica, Luca Davoli, Gianluigi Ferrari 0001
IEEE Internet Things J.1
2019 Floorplan Generation of Indoor Environments From Large-Scale Terrestrial Laser Scanner Data
abstract
This letter presents a novel approach for automatic floorplan generation of indoor environments. The floorplan is computed from a large-scale point cloud obtained from registered terrestrial laser scans. In contrast to previous work, the proposed method does not assume either a flat ground, or flat ceiling, or planar walls. Moreover, the method exploits the detection of structural elements, i.e., parts having a constant section over the entire height of the building (such as walls and columns), which is beneficial in cluttered regions to compensate for the lack of information due to occlusions. The evaluation was performed in complex buildings, like industrial warehouses, that include machines and pallet racks, whose layout is included in the generated floorplan. The algorithm achieves a floorplan reconstruction with accuracy comparable to the resolution of the adopted sensor. Results are also compared to a ground truth acquired using a total station.
Mikhail Giorgini, Jacopo Aleotti, Riccardo Monica
IEEE Geosci. Remote. Sens. Lett.3
2019 Sensor-Based Optimization of Terrestrial Laser Scanning Measurement Setup on GPU
abstract
A novel formulation of the set cover problem is presented to find the optimal placement of the scan stations in a terrestrial laser scanning survey. The problem is formulated in 2-D by including sensor-based constraints such as coverage and overlap. The coverage constraint ensures a minimum density of horizontal scan lines on the ground. The overlap constraint enables automatic scan alignment and registration. The optimization problem takes into account both environment occlusions and a maximum allowed incidence angle of the laser beams. The adopted laser model includes fixed parameters such as laser height, angular resolution, field of view, and minimum and maximum sensor range. The sensor placement problem is solved using a numerical approach implemented on graphics processing unit (GPU). Thanks to the GPU acceleration, experiments have been performed in large-scale environments with internal structures.
Mikhail Giorgini, Stefano Marini, Riccardo Monica, Jacopo Aleotti
IEEE Geosci. Remote. Sens. Lett.3
2017 Multi-label Point Cloud Annotation by Selection of Sparse Control Points
abstract
This paper presents a user-friendly approach for multi-label point cloud annotation. The method requires the user to select sparse control points belonging to the objects through a mouse-based interface. Multiple control points may be assigned to the same label. The software utilizes the selected control points to perform a segmentation algorithm on the neighborhood graph, based on shortest path tree. The user is provided a real-time feedback about the result, and can correct segmentation errors. In contrast to previous work the method supports multi-label annotation of unorganized point clouds. The method has been evaluated by multiple users and compared with a standard rectangle-based selection technique. Results indicate that the proposed method is perceived as easier to use, and that it allows a faster segmentation even in complex scenarios with occlusions.
Riccardo Monica, Jacopo Aleotti, Michael Zillich, Markus Vincze
3DV1
2017 RGB-D fusion enhancement by mode filter for surfel cloud segmentation
abstract
This paper presents an algorithm for surfel color and position enhancement from RGB-D data acquired across multiple image frames. Surfel-based reconstruction algorithms associate each RGB-D frame pixel to a surfel in the model. As the reconstruction progresses, surfel color and position are the average of all observations. Our proposed algorithm is designed to enhance position discontinuities and to produce sharper colors, to facilitate subsequent segmentation steps on the 3D model. During reconstruction, several colors and positions are tracked for each surfel. Only at the end of reconstruction phase the most frequent value is chosen through a Winner Takes All policy. The result has been compared to the standard averaging policy of reconstruction algorithms. Experiments have been performed using both Flood Fill and Supervoxel-LCCP segmentation and by applying two segmentation evaluation metrics. Results show that the proposed method is suitable to enhance a surfel-based model for object segmentation purposes.
Riccardo Monica, Michael Zillich, Markus Vincze, Jacopo Aleotti
IROS1
2014 Global registration of mid-range 3D observations and short range next best views
abstract
This work proposes a method for autonomous robot exploration of unknown objects by sensor fusion of 3D range data. The approach aims at overcoming the physical limitation of the minimum sensing distance of range sensors. Two range sensors are used with complementary characteristics mounted in eye-in-hand configuration on a robot arm. The first sensor operates at mid-range and is used in the initial phase of exploration when the environment is unknown. The second sensor, which provides short-range data, is used in the following phase where the objects are explored at close distance through next best view planning. Next best view planning is performed using a volumetric representation of the environment. A complete point cloud model of each object is finally computed by global registration of all object observations including mid-range and short range views. In experiments performed in environments with multiple rigid objects the global registration algorithm has proven more accurate than a standard sequential registration approach.
Jacopo Aleotti, Dario Lodi Rizzini, Riccardo Monica, Stefano Caselli
IROS3
2013 Design and evaluation of a delay-efficient RPL routing metric
abstract
The Routing Protocol for Low power and Lossy Networks (RPL) is the IETF standard for IPv6 routing in low-power wireless sensor networks. It is a distance vector routing protocol that builds a Destination Oriented Directed Acyclic Graph (DODAG) rooted towards one sink (the DAG root), using an objective function and a set of metrics/constraints to compute the best path. In this paper, we propose a routing metric which minimizes the delay towards the DAG root, assuming that nodes run with very low duty cycles (e.g., under 1%) at the MAC layer. We evaluate the proposed routing metric with the Contiki operating system and compare its performance with that of the Expected Transmission Count (ETX) metric. Moreover, we propose some extensions to the ContikiMAC radio duty cycling protocol to support different sleeping periods of the nodes.
Pietro Gonizzi, Riccardo Monica, Gianluigi Ferrari 0001
IWCMC2