Primo Zingaretti

dblp:89/1175 · DBLP profile ↗
← Back
31ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0002-5709-2159ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5Systems, architecture and hardware · 3 · 1 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Immersive analytics with HMDs and CAVEs: A user study on 3D graph interaction
abstract
Integrating Virtual Reality (VR) and Human–Computer Interaction (HCI) has transformed user engagement with virtual environments, enhancing immersion and usability. Technologies like Cave Automatic Virtual Environment (CAVE) and Head-Mounted Displays (HMDs) have shown significant promise in visualizing data, especially for examining and comprehending intricate 3D datasets, such as graph visualizations. To explore the effectiveness of these technologies in data visualization, we conducted a user study comparing user experience and performance across these two systems when interacting with a large 3D graph. The virtual environment and interaction modalities were adapted to each platform: the HMD setup utilized dual 6-DOF controllers, while the CAVE configuration employed a Flystick2 controller and a trackball. Preliminary data on participants’ demographics, motion sickness sensitivity, and prior experience with graph theory were collected to provide context for the findings. Results show that users in the HMD condition reported significantly higher levels of perceived presence and involvement, as well as improved task performance in navigation and interaction tasks. While both systems were rated similarly for perceived usefulness and ease of use, the HMD environment offered a more immersive and emotionally positive experience overall. These findings contribute to immersive analytics research by demonstrating the comparative strengths of HMD-based systems for individual 3D graph exploration, while highlighting the potential advantages of CAVE for low-discomfort settings. The study underscores the importance of aligning system design with user profiles and task demands to optimize data exploration in virtual environments.
Nicola Capece, Marta Mondellini, Ugo Erra, Gabriele Gilio, Emanuele Balloni, Primo Zingaretti
Graph. Model.6
2026 TrackAid-DT: An egocentric vision dataset and benchmarks for lane detection in visually impaired athlete guidance
abstract
Running on marked tracks is central to athletic training, still athletes with visual impairment depend on external guides limiting independence in training and competition. Advancements in assistive and navigation technologies offer new possibilities for inclusive sports environments. Therefore, there is an urgent need to develop inclusive mobility solutions that empower visually impaired athletes. In this work, we focus on an autonomous lane guidance system designed to assist athletes in independent navigation. To address this investigation, we propose an end-to-end vision-based system built upon our TrackAid Dataset (TrackAid-DT). The system operates under real-world conditions and supports safe athletic running by delivering vibrotactile haptic feedback for real-time guidance, in contrast to prior works that primarily address walking or jogging. TrackAid-DT comprises egocentric, pixel-annotated monocular images acquired from chest-mounted cameras under diverse orientations, speeds, and lighting conditions, and is designed to train segmentation models that enable the proposed guidance system and its deployment on edge devices. We benchmark multiple advanced segmentation models using eight-fold cross validation, task specific pretraining with public KITTI dataset and report segmentation accuracy together with computational efficiency on both training and generalization sets. Among the considered models, pretrained TransUNet achieved the strongest segmentation quality (IoU 0.929 ± 0.064, F1 0.962 ± 0.036), while U-Net offered a more deployable balance (IoU 0.921 ± 0.064, 80.31 ms/frame, 12.45 FPS, and the lowest environmental impact at 16.49 g CO 2 e). We consider the trade-offs and deploy the system on multiple edge devices and validate it through outdoor prototype testing showing its practical effectiveness in real-world running scenarios. The system maintains stable lane detection and delivers timely haptic feedback at speeds of up to 10 km/h, demonstrating feasibility for autonomous vision based guidance in athletic running. To foster further development in this domain all code and data are publicly released (link: https://github.com/vrai-group/TrackAid-DT ).
Gagan Narang, Alessandro Galdelli, Oleksandr Kuznetsov, Primo Zingaretti, Adriano Mancini
Comput. Vis. Image Underst.4
2026 Orchestrating Generative AI Paradigms With Human-in-the-Loop for 3D Generation
abstract
Generative AI techniques are revolutionizing the creation of 3D and immersive content, yet challenges remain, such as achieving precise user control in 3D generation. Current text-to-3D and image-to-3D pipelines often produce outputs that deviate from user expectations, lacking the ability to refine or correct generated models effectively. To address these limitations, we propose Imagin3D, a novel human-in-the-loop (HITL) system that integrates Multimodal Large Language Models to enhance the controllability and adaptability of 3D content generation. Imagin3D leverages a Multi-View Question Answering module to evaluate the consistency of generated views with user-provided textual descriptions, enabling iterative refinement through guided inpainting while preserving multi-view consistency. This allows users to co-create 3D models, which are then synthesized into a final 3D asset using Neural Rendering. We validate Imagin3D through extensive quantitative evaluations and a comprehensive user study, demonstrating its effectiveness in improving usability, accuracy, and user satisfaction in interactive 3D generation tasks. Our results highlight the potential of HITL approaches to bridge the gap between AI-generated outputs and user intent, paving the way for more accessible and user-centered 3D generation workflows.
Emanuele Balloni, Lorenzo Stacchio, Marina Paolanti, Primo Zingaretti, Roberto Pierdicca
IEEE Trans. Vis. Comput. Graph.4
2025 COIGAN: Controllable Object Inpainting Through Generative Adversarial Network for Defect Synthesis in Data Augmentation
abstract
Predictive maintenance is a key aspect for the safety of critical infrastructure such as bridges, dams, and tunnels, where a failure can lead to catastrophic outcomes in terms of human lives and costs. The surge in Artificial Intelligence-driven visual robotic inspection methods necessitates high-quality datasets containing diverse defect classes with several instances on different conditions (e.g., material, illumination). In this context, we introduce a Controllable Object Inpainting Generative Adversarial Network (COIGAN) to synthetically generate realistic images that augment defect datasets. The effectiveness of the model is quantitatively validated by a Fréchet Inception Distance, which measures the similarity between the generated and training samples. To further evaluate the impact of COIGAN-generated images, a segmentation task was conducted, utilizing key performance metrics such as segmentation accuracy, mAP, mIoU, and F1 score, demonstrating that the synthetic images integrate seamlessly and produce results comparable to real defect images. Subsequently, COIGAN generability was successfully used for the segmentation of a defect-free dataset by inpainting defects. The results showcase COIGAN's ability to learn defect patterns and apply them in new contexts, preserving the original features of the base image and allowing the creation of new datasets with a desired multi-class distribution. Specifically, in the context of predictive maintenance, COIGAN enriches datasets, enabling deep learning models to more effectively identify potential infrastructure anomalies. Project page: https://bit.ly/4bzxwqf.
Massimiliano Biancucci, Alessandro Galdelli, Gagan Narang, Rocco Pietrini, Adriano Mancini, Primo Zingaretti
ICRA6
2025 A Neural Rendering system for fashion design process
Emanuele Balloni, Lorenzo Stacchio, Adriano Mancini, Emanuele Frontoni, Primo Zingaretti, Marina Paolanti
Eng. Appl. Artif. Intell.5
2025 RenderGAN: Enhancing Real-time Rendering Efficiency with Deep Learning
abstract
In the domain of computer graphics, achieving high visual quality in real-time rendering remains a formidable challenge due to the inherent time-quality tradeoff. Conventional real-time rendering engines sacrifice visual fidelity for interactive performance, while image generation using path-tracing techniques can be exceedingly time-consuming. In this article, we introduce RenderGAN, a deep learning-based solution designed to address this critical challenge in real-time rendering. RenderGAN uses G-Buffers and information from a real-time rendering engine as inputs to produce output images with exceptional visual fidelity. Its encoder–decoder architecture, trained using the Generative Adversarial Network (GAN) framework with perceptual loss, enhances image realism. To evaluate RenderGAN’s effectiveness, we quantitatively compare the generated images with those of a path-tracing engine, obtaining a remarkable Universal Image Quality Index (UIQI) value of 0.898. RenderGAN’s open source nature fosters collaboration, driving advancements in real-time computer graphics and rendering techniques. By bridging the gap between real-time and path-tracing rendering, RenderGAN opens new horizons for accelerated image generation, inspiring innovation and unlocking the full potential of real-time visual experiences. Project page: https://github.com/marcomameli1992/RenderNet
Marco Mameli, Marina Paolanti, Adriano Mancini, Primo Zingaretti, Roberto Pierdicca
ACM Trans. Multim. Comput. Commun. Appl.4
2025 MineVRA: Exploring the Role of Generative AI-Driven Content Development in XR Environments through a Context-Aware Approach
abstract
The convergence of Artificial Intelligence (AI), Computer Vision (CV), Computer Graphics (CG), and Extended Reality (XR) is driving innovation in immersive environments. A key challenge in these environments is the creation of personalized 3D assets, traditionally achieved through manual modeling, a time-consuming process that often fails to meet individual user needs. More recently, Generative AI (GenAI) has emerged as a promising solution for automated, context-aware content generation. In this paper, we present MineVRA (Multimodal generative artificial iNtelligence for contExt-aware Virtual Reality Assets), a novel Human-In-The-Loop (HITL) XR framework that integrates GenAI to facilitate coherent and adaptive 3D content generation in immersive scenarios. To evaluate the effectiveness of this approach, we conducted a comparative user study analyzing the performance and user satisfaction of GenAI-generated 3D objects compared to those generated by Sketchfab in different immersive contexts. The results suggest that GenAI can significantly complement traditional 3D asset libraries, with valuable design implications for the development of human-centered XR environments.
Lorenzo Stacchio, Emanuele Balloni, Emanuele Frontoni, Marina Paolanti, Primo Zingaretti, Roberto Pierdicca
IEEE Trans. Vis. Comput. Graph.5
2024 DeepReality: An open source framework to develop AI-based augmented reality applications
abstract
Augmented reality (AR) and Artificial Intelligence (AI) are technologies pioneers in innovation and alteration in several domains. AR allows the creation of an entirely new and interactive experience for users. However, there are several drawbacks in developing AR applications, such as the marker identification process and the creation of content itself. These are very time-consuming procedures and require ad-hoc development. The advantages of using AI to solve AR limitations have recently been explored in literature. Motivated by these findings, in this paper it is proposed DeepReality, a software toolkit plug-in for Unity 3D. It is conceived for allowing developers to integrate any Deep Learning (DL) models into Unity, through AR Foundation and Barracuda inference engine. DeepReality is aimed at simplifying and streamlining the usage of DL models in conjunction with AR. As such, users skilled in Unity and DL can easily create mobile applications (iOS and Android) to: extract visual features of real-world objects (framed with the device camera) via DL; Show on-screen content on top of those real-world objects, via AR. DeepReality performs object semantic processing within the scene, and extended semantic effects for incongruent objects, overcoming the environmental tracking, which is feature-based. In order to test DeepReality usability, experiments have been performed on the execution time and memory usage data, demonstrating the feasibility and possibility of integrating and using DNNs models in mobile applications for AR. The complexity analysis confirms that DeepReality can be completely executed on mobile devices. DeepReality is also open-source and it is freely available in the Unity asset store. By fostering accessible AI-AR integration, DeepReality addresses key shortcomings in existing approaches, encapsulating contributions such as versatile DL integration, open-source accessibility, operational validation, and comprehensive metrics analysis. DeepReality empowers developers to transcend boundaries, enriching AR applications with AI’s transformative potential. Our proposed framework fosters benchmarking, comparison, and a future harmonised by AR-AI synergy.
Roberto Pierdicca, Flavio Tonetto, Marina Paolanti, Marco Mameli, Riccardo Rosati 0002, Primo Zingaretti
Expert Syst. Appl.6
2024 Shelf Management: A deep learning-based system for shelf visual monitoring
abstract
Shelf monitoring plays a key role in optimizing retail shelf layout, enhancing the customer shopping experience and maximizing profit margins. The process of automating shelf audit involves the detection, localization and recognition of objects on store shelves, including diverse products with varying attributes in unconstrained environments. This facilitates the assessment of planogram compliance. Accurate product localization within shelves requires the identification of specific shelf rows. To address the current technological challenges, we introduce “Shelf Management”, a deep learning-based system that is carefully tailored to redesign shelf monitoring practices. Our system can navigate the complexities of shelf monitoring by using advanced deep learning techniques and object detection and recognition models. In addition, a complex semantic module enhances the accuracy of detecting and assigning products to their designated shelf rows and locations. In particular, we recognize the lack of finely annotated datasets at the SKU level. As a contribution to the field, we provide annotations for two novel datasets: SHARD (SHelf mAnagement Row Dataset) and SHAPE (SHelf mAnagement Product dataset). These datasets not only provide valuable resources, but also serve as benchmarks for further research in the field of retail. A complete pipeline is designed using a RetinaNet architecture for object detection with 0.752 mAP, followed by a Deep Hough transform to detect shelf rows as semantic lines with an F1 score of 97%, and a product recognition step using a MobileNetV3 architecture trained with triplet loss and used as a feature extractor together with FAISS for fast image retrieval with an accuracy of 93% on top-1 recognition. Localization is achieved using a deterministic approach based on product detection and shelf row detection. Source code and datasets are available at https://github.com/rokopi-byte/shelf_management.
Rocco Pietrini, Marina Paolanti, Adriano Mancini, Emanuele Frontoni, Primo Zingaretti
Expert Syst. Appl.5
2024 Generalizability and robustness evaluation of attribute-based zero-shot learning
abstract
In the field of deep learning, large quantities of data are typically required to effectively train models. This challenge has given rise to techniques like zero-shot learning (ZSL), which trains models on a set of "seen" classes and evaluates them on a set of "unseen" classes. Although ZSL has shown considerable potential, particularly with the employment of generative methods, its generalizability to real-world scenarios remains uncertain. The hypothesis of this work is that the performance of ZSL models is systematically influenced by the chosen "splits"; in particular, the statistical properties of the classes and attributes used in training. In this paper, we test this hypothesis by introducing the concepts of generalizability and robustness in attribute-based ZSL and carry out a variety of experiments to stress-test ZSL models against different splits. Our aim is to lay the groundwork for future research on ZSL models' generalizability, robustness, and practical applications. We evaluate the accuracy of state-of-the-art models on benchmark datasets and identify consistent trends in generalizability and robustness. We analyze how these properties vary based on the dataset type, differentiating between coarse- and fine-grained datasets, and our findings indicate significant room for improvement in both generalizability and robustness. Furthermore, our results demonstrate the effectiveness of dimensionality reduction techniques in improving the performance of state-of-the-art models in fine-grained datasets.
Luca Rossi 0007, Maria Chiara Fiorentino, Adriano Mancini, Marina Paolanti, Riccardo Rosati 0002, Primo Zingaretti
Neural Networks6
2023 Deep Reinforced Navigation of Agents in 2D Platform Video Games
Emanuele Balloni, Marco Mameli, Adriano Mancini, Primo Zingaretti
CGI (3)4
2023 Investigation on the Encoder-Decoder Application for Mesh Generation
Marco Mameli, Emanuele Balloni, Adriano Mancini, Emanuele Frontoni, Primo Zingaretti
CGI5
2020 Weight Estimation from an RGB-D camera in top-view configuration
abstract
The development of so-called soft-biometrics aims at providing information related to the physical and behavioural characteristics of a person. This paper focuses on body weight estimation based on the observation from a top-view RGB-D camera. In fact, the capability to estimate the weight of a person can be of help in many different applications, from health-related scenarios, to business intelligence and retail analytics. To deal with this issue, a TVWE (Top-View Weight Estimation) framework is proposed with the aim of predicting the weight. The approach relies on the adoption of Deep Neural Networks (DNNs) that have been trained on depth data. Each network has also been modified in their top section to replace classification with prediction inference. The performance of five state-of-art DNNs have been compared, namely VGG16, ResNet, Inception, DenseNet and Efficient-Net. In addition, a convolutional auto-encoder has also been included for completeness. Considering the limited literature in this domain, the TVWE framework has been evaluated on a new publicly available dataset: “VRAI Weight estimation Dataset”, which also collects, for each subject, labels related to weight, gender, and height. The experimental results have demonstrated that the proposed methods are suitable for this task, bringing different and significant insights for the application of the solution in different domains.
Marco Mameli, Marina Paolanti, Nicola Conci, Filippo Tessaro, Emanuele Frontoni, Primo Zingaretti
ICPR6
2020 Deep understanding of shopper behaviours and interactions using RGB-D vision
abstract
Abstract In retail environments, understanding how shoppers move about in a store’s spaces and interact with products is very valuable. While the retail environment has several favourable characteristics that support computer vision, such as reasonable lighting, the large number and diversity of products sold, as well as the potential ambiguity of shoppers’ movements, mean that accurately measuring shopper behaviour is still challenging. Over the past years, machine-learning and feature-based tools for people counting as well as interactions analytic and re-identification were developed with the aim of learning shopper skills based on occlusion-free RGB-D cameras in a top-view configuration. However, after moving into the era of multimedia big data, machine-learning approaches evolved into deep learning approaches, which are a more powerful and efficient way of dealing with the complexities of human behaviour. In this paper, a novel VRAI deep learning application that uses three convolutional neural networks to count the number of people passing or stopping in the camera area, perform top-view re-identification and measure shopper–shelf interactions from a single RGB-D video flow with near real-time performances has been introduced. The framework is evaluated on the following three new datasets that are publicly available: TVHeads for people counting, HaDa for shopper–shelf interactions and TVPR2 for people re-identification. The experimental results show that the proposed methods significantly outperform all competitive state-of-the-art methods (accuracy of 99.5% on people counting, 92.6% on interaction classification and 74.5% on re-id), bringing to different and significative insights for implicit and extensive shopper behaviour analysis for marketing applications.
Marina Paolanti, Rocco Pietrini, Adriano Mancini, Emanuele Frontoni, Primo Zingaretti
Mach. Vis. Appl.5
2019 Design, Large-Scale Usage Testing, and Important Metrics for Augmented Reality Gaming Applications
abstract
Augmented Reality (AR) offers the possibility to enrich the real world with digital mediated content, increasing in this way the quality of many everyday experiences. While in some research areas such as cultural heritage, tourism, or medicine there is a strong technological investment, AR for game purposes struggles to become a widespread commercial application. In this article, a novel framework for AR kid games is proposed, already developed by the authors for other AR applications such as Cultural Heritage and Arts. In particular, the framework includes different layers such as the development of a series of AR kid puzzle games in an intermediate structure which can be used as a standard for different applications development, the development of a smart configuration tool, together with general guidelines and long-life usage tests and metrics. The proposed application is designed for augmenting the puzzle experience, but can be easily extended to other AR gaming applications. Once the user has assembled the real puzzle, AR functionality within the mobile application can be unlocked, bringing to life puzzle characters, creating a seamless game that merges AR interactions with the puzzle reality. The main goals and benefits of this framework can be seen in the development of a novel set of AR tests and metrics in the pre-release phase (in order to help the commercial launch and developers), and in the release phase by introducing the measures for long-life app optimization, usage tests and hint on final users together with a measure to design policy, providing a method for automatic testing of quality and popularity improvements. Moreover, smart configuration tools, as part of the general framework, enabling multi-app and eventually also multi-user development, have been proposed, facilitating the serialization of the applications. Results were obtained from a large-scale user test with about 4 million users on a set of eight gaming applications, providing the scientific community a workflow for implicit quantitative analysis in AR gaming. Different data analytics developed on the data collected by the framework prove that the proposed approach is affordable and reliable for long-life testing and optimization.
Roberto Pierdicca, Emanuele Frontoni, Primo Zingaretti, Adriano Mancini, Jelena Loncarski, Marina Paolanti
ACM Trans. Multim. Comput. Commun. Appl.3
2018 Convolutional Networks for Semantic Heads Segmentation using Top-View Depth Data in Crowded Environment
abstract
Detecting and tracking people is a challenging task in a persistent crowded environment (i.e. retail, airport, station, etc.) for human behaviour analysis of security purposes. This paper introduces an approach to track and detect people in cases of heavy occlusions based on CNNs for semantic segmentation using top-view depth visual data. The purpose is the design of a novel U-Net architecture, U-Net3, that has been modified compared to the previous ones at the end of each layer. In particular, a batch normalization is added after the first ReLU activation function and after each max-pooling and up-sampling functions. The approach was applied and tested on a new and public available dataset, TVHeads Dataset, consisting of depth images of people recorded from an RGB-D camera installed in top-view configuration. Our variant outperforms baseline architectures while remaining computationally efficient at inference time. Results show high accuracy, demonstrating the effectiveness and suitability of our approach.
Daniele Liciotti, Marina Paolanti, Rocco Pietrini, Emanuele Frontoni, Primo Zingaretti
ICPR5
2018 Introduction to the Special Issue on Applications of Mechatronic and Embedded Systems (MESA) in ITS
abstract
Embedded systems result from the integration between mechanical and electronic components (hardware) and the information-driven functions (software). Embedded systems play a key role in the development of mechatronic systems, which involves finding an optimal balance between the basic mechanical structure, sensor and actuators, automatic digital information processing and control.
Massimo Bertozzi, Primo Zingaretti
IEEE Trans. Intell. Transp. Syst.3
2018 Mechatronic System to Help Visually Impaired Users During Walking and Running
abstract
Ambient assisted living and intelligent transportation systems are becoming strongly coupled. There is the necessity of improving the quality of life by developing inclusive mobility solutions for impaired people. In this paper, we focus on a monocular vision-based system to assist people during walking, jogging, and running in outdoor environments. The impaired user is guided along a path represented by a lane or line on a dedicated runway. We developed a set of image processing algorithms to extract lines/lanes to follow. The embedded system is based on a small camera and a board that is responsible for processing the images and communicating with the developed haptic device. The haptic device is formed by a set of two gloves equipped with vibration motors that drive the user to the right direction. The vibration sequences are generated according to a robotic-like controller, considering the user as a two wheel steering robot, where the rotational and translation velocity can be controlled. The results obtained show that the overall system is able to detect the right path and to provide the right stimuli to the user, by means of the gloves, up to a speed over 10 km/h.
Adriano Mancini, Emanuele Frontoni, Primo Zingaretti
IEEE Trans. Intell. Transp. Syst.3
2017 HMM-based Activity Recognition with a Ceiling RGB-D Camera
abstract
Automated recognition of Activities of Daily Living allows to identify possible health problems and apply corrective strategies in Ambient Assisted Living (AAL). Activities of Daily Living analysis can provide very useful information for elder care and long-term care services. This paper presents an automated RGB-D video analysis system that recognises human ADLs activities, related to classical daily actions. The main goal is to predict the probability of an analysed subject action. Thus, abnormal behaviour can be detected. The activity detection and recognition is performed using an affordable RGB-D camera. Human activities, despite their unstructured nature, tend to have a natural hierarchical structure; for instance, generally making a coffee involves a three-step process of turning on the coffee machine, putting sugar in cup and opening the fridge for milk. Action sequence recognition is then handled using a discriminative Hidden Markov Model (HMM). RADiaL, a dataset with RGB-D images and 3D position of each person for training as well as evaluating the HMM, has been built and made publicly available.
Daniele Liciotti, Emanuele Frontoni, Primo Zingaretti, Nicola Bellotto, Tom Duckett
ICPRAM3
2016 Robust and affordable retail customer profiling by vision and radio beacon sensor fusion
Mirco Sturari, Daniele Liciotti, Roberto Pierdicca, Emanuele Frontoni, Adriano Mancini, Marco Contigiani, Primo Zingaretti
Pattern Recognit. Lett.7
2015 Introduction to the Special Issue on Mechatronic and Embedded Systems and Applications in ITS
abstract
The papers in this special issue were presented at the The Mechatronic and Embedded Technologies in Intelligent Transportation Systems Symposium which includes contributions in technologies, methodologies and, in particular, applications of mechatronic and embedded systems in any aspect of ITSs, from automatic vehicle localization and monitoring to autonomous vehicles and environmental perception.
Massimo Bertozzi, Yanqing Gao, Primo Zingaretti
IEEE Trans. Intell. Transp. Syst.4
2015 Embedded Multisensor System for Safe Point-to-Point Navigation of Impaired Users
abstract
New smart objects to improve the quality of life in the ambient assisted living (AAL) scenario are capturing the interest of researchers and companies. In particular, novel assistive technologies are being developed to make accessible street navigation to impaired people. The solution that we propose in this new application domain of intelligent transportation systems is a framework for a safe point-to-point navigation, owing to high-detailed road graphs, including sidewalks, crosswalks, and generic “obstacles.” The system is based on a low-cost modular sensor box (embedded hardware) interfaced with a mobile/phone application that acts as an intelligent navigator. The main novelty is the capability to sense the surrounding area while being able to perform a fast path replanning, owing to a real-time link to a remote server, if an obstacle is detected. The sensing is performed using different sensors, such as ultrasound, lidar, and a 77-GHz mid-range automotive radar (absolutely novel in the AAL context), which are processed and fused in the well-established robot operating system (ROS). We tested the framework by analyzing its performance in two different configurations and environments by using, respectively, a sonar and a laser rangefinder in a building scenario and a radar in an urban environment. Even if in both cases results demonstrated a quite good robustness in the obstacle detection with a quasi-real-time route replanning, we were mainly interested and succeeded in demonstrating the high flexibility and extensibility of our framework.
Adriano Mancini, Emanuele Frontoni, Primo Zingaretti
IEEE Trans. Intell. Transp. Syst.3
2014 Feature Group Matching: a Novel Method to filter out Incorrect Local Feature Matchings
abstract
The importance of finding correct correspondences between two images is the major aspect in problems such as appearance-based robot localization and content-based image retrieval. Local feature matching has become a commonly used method to compare images, despite being highly probable that at least some of the matchings/correspondences it detects are incorrect. In this paper, we describe a novel approach to local feature matching, named Feature Group Matching (FGM), to select stable features and obtain a more reliable similarity value between two images. The proposed technique is demonstrated to be translational, rotational and scaling invariant. Experimental evaluation was performed on large and heterogeneous datasets of images using SIFT and SURF, the actual state-of-the-art feature extractors. Results show that FGM avoids almost 95% of incorrect matchings, reduces the visual aliasing (number of images considered similar) and increases both robotic localization and image retrieval accuracy on the average of 13%.
Emanuele Frontoni, Adriano Mancini, Primo Zingaretti
Int. J. Pattern Recognit. Artif. Intell.3
2011 Hybrid object-based approach for land use/land cover mapping using high spatial resolution imagery
abstract
Traditionally, remote sensing has employed pixel-based classification techniques to deal with land use/land cover (LULC) studies. Generally, pixel-based approaches have been proven to work well with low spatial resolution imagery (e.g. Landsat or System Pour L'Observation de la Terre sensors). Now, however, commercially available high spatial resolution images (e.g. aerial Leica ADS40 and Vexcel UltraCam sensors, and satellite IKONOS, Quickbird, GeoEye and WorldView sensors) can be problematic for pixel-based analysis due to their tendency to oversample the scene. This is driving research towards object-based approaches. This article proposes a hybrid classification method with the aim of incorporating the advantages of supervised pixel-based classification into object-based approaches. The method has been developed for medium-scale (1:10,000) LULC mapping using ADS40 imagery with 1 m ground sampling distance. First, spatial information is incorporated into a pixel-based classification (AdaBoost classifier) by means of additional texture features (Haralick, Gabor, Law features), which can be selected ‘ad hoc’ according to optimal training samples (‘Relief-F’ approach, Mahalanobis distances). Then a rule-based approach sorts segmented regions into thematic CORINE Land Cover classes in terms of membership class percentages (a modified Winner-Takes-All approach) and shape parameters. Finally, ancillary data (roads, rivers, etc.) are exploited to increase classification accuracy. The experimental results show that the proposed hybrid approach allows the extraction of more LULC classes than conventional pixel-based methods, while improving classification accuracy considerably. A second contribution of this article is the assessment of classification reliability by implementing a stability map, in addition to confusion matrices.
Eva Savina Malinverni, Anna Nora Tassetti, Adriano Mancini, Primo Zingaretti, Emanuele Frontoni, Annamaria Bernardini
Int. J. Geogr. Inf. Sci.4
2010 Road Change Detection from Multi-Spectral Aerial Data
abstract
The paper presents a novel approach to automate the Change Detection (CD) problem for the specific task of road extraction. Manual approaches to CD fail in terms of the time for releasing updated maps; in the contrary, automatic approaches, based on machine learning and image processing techniques, allow to update large areas in a short time with an accuracy and precision comparable to those obtained by human operators. This work is focused on the road-graph update starting from aerial, multi-spectral data. Georeferenced, ground data, acquired by a GPS and an inertial sensor, are integrated with aerial data to speed up the change detector. After roads extraction by means of a binary AdaBoost classifier, the old road-graph is updated exploiting a particle filter. In particular this filter results very useful to link (track) parts of roads not extracted by the classifier due to the presence of occlusions (e.g., shadows, trees).
Adriano Mancini, Emanuele Frontoni, Primo Zingaretti
ICPR3
2010 LCLU information system for object-oriented nomenclature
abstract
A Land Cover/Land Use (LCLU) Information System is proposed as a new dynamic and flexible approach to describe landscape objects. It is able to give a deeper and more realistic thematic description by storing membership land cover attributes for each polygon automatically extracted and classified by the T-MAP software. The proposed approach can overcome the traditional “hard” classification by taking directly into account “fuzzy” cover components and making the classification approach more bounded with the polygon characteristics and their changes. The LCLU Information System can be easily integrated with different databases, making it suitable for different nomenclatures and further analysis, regarding environmental indexes, class updating and classification stability assessment.
Eva Savina Malinverni, Anna Nora Tassetti, Primo Zingaretti
IGARSS3
2009 RoboBuntu: A Linux distribution for mobile robotics
abstract
During last years Linux started to climb the market of operating systems (OSs), and Ubuntu, derived by Debian OS, has become a good alternative to common OSs like Windows XP or Vista. The mobile robotics scientific community makes use of Linux based OSs to avoid the lack of stability that affects Microsoft OSs, especially when real time conditions must be satisfied. In this paper we present the Linux distribution RoboBuntu, acronym formed by the union of ROBOt and uBUNTU, to overcome the almost totally independent robotic software platforms existing today. The key idea behind RoboBuntu is the integration of different tools for mobile robotics into an embedded Ubuntu distribution. Another important characteristics of RoboBuntu is that every ldquohard steprdquo, like installation and configuration of OS and tools, is hidden to common users. In particular, RoboBuntu can be used either by students or researchers, as LiveCd, permanent installation on standard hard drive or, more interesting, on a USB storage flash disk.
Adriano Mancini, Emanuele Frontoni, Andrea Ascani, Primo Zingaretti
ICRA4
2008 Feature group matching for appearance-based localization
abstract
Local feature matching has become a commonly used method to compare images. For mobile robots, a reliable method for comparing images can constitute a key component for localization tasks. In this paper, we address the issues of appearance-based topological and metric localization by introducing a novel group matching approach to select less but more robust features to match the current robot view with reference images. Feature group matching is based on the consideration that feature descriptors together with spatial relations are more robust than classical approaches. Our datasets, each consisting of a large number of omnidirectional images, have been acquired over different day times (different lighting conditions) both in indoor and outdoor environments. The feature group matching outperforms the SIFT in indoor localization showing better performances both in the case of topological and metric localization. In outdoor SURF remains the best feature extraction method, as reported in literature.
Andrea Ascani, Emanuele Frontoni, Adriano Mancini, Primo Zingaretti
IROS4
2006 Aliasing Maps for Robot Global Localization
Emanuele Frontoni, Primo Zingaretti
ECAI2
2006 Complete classification of raw LIDAR data and 3D reconstruction of buildings
Gianfranco Forlani, Carla Nardinocchi, Marco Scaioni, Primo Zingaretti
Pattern Anal. Appl.4
1998 Fast Chain Coding of Region Boundaries
abstract
A fast single-pass algorithm to convert a multivalued image from a raster-based representation into chain codes is presented. All chain codes are obtained in linear time with respect to the number of chain segments that are generated at each raster according to a set of templates. A formal statement and the complexity and performance analysis of the algorithm are given.
Primo Zingaretti, Massimiliano Gasparroni, Lorenzo Vecci
IEEE Trans. Pattern Anal. Mach. Intell.1