VLDB 2026 Research / reviewers in the wild / expert
Adriano Mancini
dblp:52/1454
· DBLP profile ↗
29ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0001-5281-9200ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Freq2Clean: Enhancing Calcium Imaging Denoising via Frequency-Domain Fusion
Valerio Morelli, Daniele Berardini, Giorgio Letti, Sebastiano Curreli, Adriano Mancini, Tommaso Fellin, Vittorio Murino |
ICPR (12) | 5 |
| 2026 | TrackAid-DT: An egocentric vision dataset and benchmarks for lane detection in visually impaired athlete guidanceabstractRunning on marked tracks is central to athletic training, still athletes with visual impairment depend on external guides limiting independence in training and competition. Advancements in assistive and navigation technologies offer new possibilities for inclusive sports environments. Therefore, there is an urgent need to develop inclusive mobility solutions that empower visually impaired athletes. In this work, we focus on an autonomous lane guidance system designed to assist athletes in independent navigation. To address this investigation, we propose an end-to-end vision-based system built upon our TrackAid Dataset (TrackAid-DT). The system operates under real-world conditions and supports safe athletic running by delivering vibrotactile haptic feedback for real-time guidance, in contrast to prior works that primarily address walking or jogging. TrackAid-DT comprises egocentric, pixel-annotated monocular images acquired from chest-mounted cameras under diverse orientations, speeds, and lighting conditions, and is designed to train segmentation models that enable the proposed guidance system and its deployment on edge devices. We benchmark multiple advanced segmentation models using eight-fold cross validation, task specific pretraining with public KITTI dataset and report segmentation accuracy together with computational efficiency on both training and generalization sets. Among the considered models, pretrained TransUNet achieved the strongest segmentation quality (IoU 0.929 ± 0.064, F1 0.962 ± 0.036), while U-Net offered a more deployable balance (IoU 0.921 ± 0.064, 80.31 ms/frame, 12.45 FPS, and the lowest environmental impact at 16.49 g CO 2 e). We consider the trade-offs and deploy the system on multiple edge devices and validate it through outdoor prototype testing showing its practical effectiveness in real-world running scenarios. The system maintains stable lane detection and delivers timely haptic feedback at speeds of up to 10 km/h, demonstrating feasibility for autonomous vision based guidance in athletic running. To foster further development in this domain all code and data are publicly released (link: https://github.com/vrai-group/TrackAid-DT ). Gagan Narang, Alessandro Galdelli, Oleksandr Kuznetsov, Primo Zingaretti, Adriano Mancini |
Comput. Vis. Image Underst. | 5 |
| 2026 | Challenging DINOv3 foundation model under low inter-class variability: a case study on fetal brain ultrasoundabstractAbstract This study provides the first comprehensive evaluation of foundation models in fetal ultrasound (US) imaging under low inter-class variability conditions. While recent vision foundation models such as DINOv3 have shown remarkable transferability across medical domains, their ability to discriminate anatomically similar structures has not been systematically investigated. We address this gap by focusing on fetal brain standard planes–transthalamic (TT), transventricular (TV), and transcerebellar (TC)–which exhibit highly overlapping anatomical features and pose a critical challenge for reliable biometric assessment. To ensure a fair and reproducible evaluation, all publicly available fetal ultrasound datasets were curated and aggregated into a unified multicenter benchmark, FetalUS-188K, comprising more than 188,000 annotated images from heterogeneous acquisition settings. DINOv3 was pretrained in a self-supervised manner to learn ultrasound-aware representations. The learned features were then evaluated through standardized adaptation protocols, including linear probing with frozen backbone and full fine-tuning, under two initialization schemes: (i) pretraining on FetalUS-188K and (ii) initialization from natural-image DINOv3 weights. Models pretrained on fetal ultrasound data consistently outperformed those initialized on natural images, yielding weighted F1-score improvements of up to 21% (0.73 vs. 0.52 for ViT-B/16). This domain-adaptive pretraining proved critical for resolving low-margin class boundaries; while natural-image weights led to a representational collapse on the TV plane (F1-score $$\le$$ 16%), our approach preserved the subtle echogenic and structural cues necessary for its accurate discrimination. These results demonstrate that while generic foundation models fail to generalize under low inter-class variability, domain-specific pretraining is a technical prerequisite for achieving the robust and clinically reliable representations required for fetal brain biometric assessment. Edoardo Conti, Riccardo Rosati 0002, Lorenzo Federici, Adriano Mancini, Maria Chiara Fiorentino |
Neural Comput. Appl. | 4 |
| 2026 | From classification to landmark-aware identification: a multi-task multilabel approach for fetal brain ultrasound planesabstractAbstract Standard plane acquisition in fetal neurosonography is essential for prenatal diagnosis, yet remains highly operator-dependent due to the need to identify subtle anatomical landmarks that define each imaging plane. Current automated approaches either treat the problem as plane-level image classification—without explicitly verifying landmark visibility—or employ computationally expensive object detection methods that may not be suitable for real-time clinical integration. To address these limitations, we propose a novel multitask learning framework that combines detection-based spatial awareness during training with efficient multilabel classification at inference. Validated on a public dataset collected at BCNatal center, the method was compared against a standard ResNet-101 multilabel baseline. Results demonstrate that the proposed framework achieves an average F1-score of 0.92, significantly outperforming the baseline, particularly for the challenging Cavum Septi Pellucidi where precision improves from 0.76 to 0.93. Grad-CAM visualizations reveal substantially more focused and anatomically precise attention maps compared to baseline approaches, with activations concentrated on clinically relevant substructures such as the cerebellar vermis. These findings confirm that, by discarding the auxiliary detection head at inference, the model retains the spatial features learned during training while minimizing computational overhead. Maria Chiara Fiorentino, Riccardo Rosati 0002, Edoardo Conti, Costantino Tigano, Adriano Mancini |
Neural Comput. Appl. | 5 |
| 2025 | COIGAN: Controllable Object Inpainting Through Generative Adversarial Network for Defect Synthesis in Data AugmentationabstractPredictive maintenance is a key aspect for the safety of critical infrastructure such as bridges, dams, and tunnels, where a failure can lead to catastrophic outcomes in terms of human lives and costs. The surge in Artificial Intelligence-driven visual robotic inspection methods necessitates high-quality datasets containing diverse defect classes with several instances on different conditions (e.g., material, illumination). In this context, we introduce a Controllable Object Inpainting Generative Adversarial Network (COIGAN) to synthetically generate realistic images that augment defect datasets. The effectiveness of the model is quantitatively validated by a Fréchet Inception Distance, which measures the similarity between the generated and training samples. To further evaluate the impact of COIGAN-generated images, a segmentation task was conducted, utilizing key performance metrics such as segmentation accuracy, mAP, mIoU, and F1 score, demonstrating that the synthetic images integrate seamlessly and produce results comparable to real defect images. Subsequently, COIGAN generability was successfully used for the segmentation of a defect-free dataset by inpainting defects. The results showcase COIGAN's ability to learn defect patterns and apply them in new contexts, preserving the original features of the base image and allowing the creation of new datasets with a desired multi-class distribution. Specifically, in the context of predictive maintenance, COIGAN enriches datasets, enabling deep learning models to more effectively identify potential infrastructure anomalies. Project page: https://bit.ly/4bzxwqf. Massimiliano Biancucci, Alessandro Galdelli, Gagan Narang, Rocco Pietrini, Adriano Mancini, Primo Zingaretti |
ICRA | 5 |
| 2025 | A Neural Rendering system for fashion design process
Emanuele Balloni, Lorenzo Stacchio, Adriano Mancini, Emanuele Frontoni, Primo Zingaretti, Marina Paolanti |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Online Knowledge Distillation and Deep Supervision in HRNet: Green AI for Preterm Infants' Pose EstimationabstractThe current approach to deep-learning research is exemplified by the pursuit of red AI models—designs that show increasingly high performance but with increasingly high costs, in terms of economic requirements and environmental footprint. This approach is particularly detrimental in sectors like healthcare, which typically have limited resources. Meanwhile, green AI prioritizes efficiency and sustainability, by reducing the environmental footprint and making advanced technologies accessible. Following the green AI principles, this study focuses on the combined use of two techniques, namely Knowledge Distillation (KD) and Deep Supervision (DS), to reduce the costs of HRNet, a convolutional neural network designed for human pose estimation, here applied to support the diagnosis of neurological impairments in preterm infants. All the experiments are carried out on the BabyPose dataset, a collection of videos from a depth camera showing hospitalized preterm infants. By combining KD and DS, we can use a sub-network of HRNet that needs only 27.5% of its parameters and 61.7% of its FLOPs, without affecting performance (−0.59 percentage points in average precision). This achievement can have deep implications in the actual clinical practice, as it fosters democratization of high-quality technologies. Our codes are available at https://github.com/geronimaw/OnlineKD-HRNet-Human-Pose-Estimation.git . Alessandro Cacciatore, Daniele Berardini, Vito Scaraggi, Adriano Mancini, Sara Moccia, Lucia Migliorelli |
ACM Trans. Comput. Heal. | 4 |
| 2025 | OutfitAI: shop the outfit with a deep learning-based intelligent expert systemabstractAbstract In an age where consumer preferences are as diverse as they are dynamic, the ability to offer personalized fashion recommendations at scale remains a significant challenge for retailers. Consumers seek a shopping experience that not only understands their unique style preferences but also dynamically adapts to their evolving tastes. The fashion industry is at a crossroads, facing increasing consumer demand for personalization, sustainability and transparency in a rapidly evolving digital marketplace. Traditional retail practices, while rich in tradition and artistry, often struggle to up-to-date with the rapidly, ethically-conscious and technology-driven expectations of today’s consumers. “OutfitAI” is designed to address these challenges by leveraging the power of deep learning to revolutionize the fashion retail experience. By automating the process of background removal in fashion images, using advanced algorithms for personalized product matching, and integrating sustainability filters into the product discovery process, OutfitAI aims to deliver a shopping experience that is not only personalized and engaging, but also aligned with the ethical and environmental values of the contemporary consumer. Unlike existing solutions, OutfitAI uses state-of-the-art semantic segmentation for precise background removal, enabling detailed feature extraction from fashion images. This process enables accurate matching of user-uploaded images with similar fashion items from an extensive database of eco-friendly and ethically produced products sourced from leading e-tailers. Setting itself apart from the current state of the art, OutfitAI places a strong emphasis on ethical data use and privacy, implementing robust measures to ensure user privacy and transparency. It also pioneers the integration of sustainability into the digital fashion discovery process, promoting responsible consumption patterns among users. Through a comprehensive system architecture that combines technical innovation with a commitment to ethics and sustainability, OutfitAI not only addresses the technological needs of the fashion retail industry, but also responds to the growing demand for more responsible and transparent consumer technologies. Emanuele Balloni, Rocco Pietrini, Emanuele Frontoni, Adriano Mancini, Marina Paolanti |
Multim. Tools Appl. | 4 |
| 2025 | RenderGAN: Enhancing Real-time Rendering Efficiency with Deep LearningabstractIn the domain of computer graphics, achieving high visual quality in real-time rendering remains a formidable challenge due to the inherent time-quality tradeoff. Conventional real-time rendering engines sacrifice visual fidelity for interactive performance, while image generation using path-tracing techniques can be exceedingly time-consuming. In this article, we introduce RenderGAN, a deep learning-based solution designed to address this critical challenge in real-time rendering. RenderGAN uses G-Buffers and information from a real-time rendering engine as inputs to produce output images with exceptional visual fidelity. Its encoder–decoder architecture, trained using the Generative Adversarial Network (GAN) framework with perceptual loss, enhances image realism. To evaluate RenderGAN’s effectiveness, we quantitatively compare the generated images with those of a path-tracing engine, obtaining a remarkable Universal Image Quality Index (UIQI) value of 0.898. RenderGAN’s open source nature fosters collaboration, driving advancements in real-time computer graphics and rendering techniques. By bridging the gap between real-time and path-tracing rendering, RenderGAN opens new horizons for accelerated image generation, inspiring innovation and unlocking the full potential of real-time visual experiences. Project page: https://github.com/marcomameli1992/RenderNet Marco Mameli, Marina Paolanti, Adriano Mancini, Primo Zingaretti, Roberto Pierdicca |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | An automated CAD-to-XR framework based on generative AI and Shrinkwrap modelling for a User-Centred design approachabstract• Automated CAD to XR workflow for interactive Photorealistic Virtual Prototype (iPVP) • Unique texture generation module using a tailored approach based on GANs. • Shrinkwrap modelling for efficient 3D model simplification and texture UV mapping. • Significant time savings in virtual prototyping process. • Framework validated with a case study on sporting rifles, showing high-quality iPVP. CAD-to-XR is the workflow to generate interactive Photorealistic Virtual Prototypes (iPVPs) for Extended Reality (XR) apps from Computer-Aided Design (CAD) models. This process entails modelling, texturing, and XR programming. In the literature, no automatic CAD-to-XR frameworks simultaneously manage CAD simplification and texturing. There are no examples of their adoption for User-Centered Design (UCD). Moreover, such CAD-to-XR workflows do not seize the potentialities of generative algorithms to produce synthetic images (textures). The paper presents a framework for implementing the CAD-to-XR workflow. The solution consists of a module for texture generation based on Generative Adversarial Networks (GANs). The generated texture is then managed by another module (based on Shrinkwrap modelling) to develop the iPVP by simplifying the 3D model and UV mapping the generated texture. The geometric and material data is integrated into a graphic engine, which allows for programming an interactive experience with the iPVP in XR. The CAD-to-XR framework was validated on two components (rifle stock and forend) of a sporting rifle. The solution can automate the texturing process of different product versions in shorter times (compared to a manual procedure). After each product revision, it avoids tedious and manual activities required to generate a new iPVP. The image quality metrics highlight that images are generated in a “realistic” manner (the perceived quality of generated textures is highly comparable to real images). The quality of the iPVPs, generated through the proposed framework and visualised by users through a mixed reality head-mounted display, is equivalent to traditionally designed prototypes. Riccardo Rosati 0002, Paolo Senesi, Barbara Lonzi, Adriano Mancini, Marco Mandolini |
Adv. Eng. Informatics | 4 |
| 2024 | Social4Fashion: An intelligent expert system for forecasting fashion trends from social media contents
Emanuele Balloni, Rocco Pietrini, Matteo Fabiani, Emanuele Frontoni, Adriano Mancini, Marina Paolanti |
Expert Syst. Appl. | 5 |
| 2024 | Shelf Management: A deep learning-based system for shelf visual monitoringabstractShelf monitoring plays a key role in optimizing retail shelf layout, enhancing the customer shopping experience and maximizing profit margins. The process of automating shelf audit involves the detection, localization and recognition of objects on store shelves, including diverse products with varying attributes in unconstrained environments. This facilitates the assessment of planogram compliance. Accurate product localization within shelves requires the identification of specific shelf rows. To address the current technological challenges, we introduce “Shelf Management”, a deep learning-based system that is carefully tailored to redesign shelf monitoring practices. Our system can navigate the complexities of shelf monitoring by using advanced deep learning techniques and object detection and recognition models. In addition, a complex semantic module enhances the accuracy of detecting and assigning products to their designated shelf rows and locations. In particular, we recognize the lack of finely annotated datasets at the SKU level. As a contribution to the field, we provide annotations for two novel datasets: SHARD (SHelf mAnagement Row Dataset) and SHAPE (SHelf mAnagement Product dataset). These datasets not only provide valuable resources, but also serve as benchmarks for further research in the field of retail. A complete pipeline is designed using a RetinaNet architecture for object detection with 0.752 mAP, followed by a Deep Hough transform to detect shelf rows as semantic lines with an F1 score of 97%, and a product recognition step using a MobileNetV3 architecture trained with triplet loss and used as a feature extractor together with FAISS for fast image retrieval with an accuracy of 93% on top-1 recognition. Localization is achieved using a deterministic approach based on product detection and shelf row detection. Source code and datasets are available at https://github.com/rokopi-byte/shelf_management. Rocco Pietrini, Marina Paolanti, Adriano Mancini, Emanuele Frontoni, Primo Zingaretti |
Expert Syst. Appl. | 3 |
| 2024 | A deep-learning framework running on edge devices for handgun and knife detection from indoor video-surveillance camerasabstractAbstract The early detection of handguns and knives from surveillance videos is crucial to enhance people’s safety. Despite the increasing development of Deep Learning (DL) methods for general object detection, weapon detection from surveillance videos still presents open challenges. Among these, the most significant are: (i) the very small size of the weapons with respect to the camera field of view and (ii) the need of a real-time feedback, even when using low-cost edge devices for computation. Complex and recently-developed DL architectures could mitigate the former challenge but do not satisfy the latter one. To tackle such limitation, the proposed work addresses the weapon-detection task from an edge perspective. A double-step DL approach was developed and evaluated against other state-of-the-art methods on a custom indoor surveillance dataset. The approach is based on a first Convolutional Neural Network (CNN) for people detection which guides a second CNN to identify handguns and knives. To evaluate the performance in a real-world indoor environment, the approach was deployed on a NVIDIA Jetson Nano edge device which was connected to an IP camera. The system achieved near real-time performance without relying on expensive hardware. The results in terms of both COCO Average Precision (AP = 79.30) and Frames per Second (FPS = 5.10) on the low-power NVIDIA Jetson Nano pointed out the goodness of the proposed approach compared with the others, encouraging the spread of automated video surveillance systems affordable to everyone. Daniele Berardini, Lucia Migliorelli, Alessandro Galdelli, Emanuele Frontoni, Adriano Mancini, Sara Moccia |
Multim. Tools Appl. | 5 |
| 2024 | Generalizability and robustness evaluation of attribute-based zero-shot learningabstractIn the field of deep learning, large quantities of data are typically required to effectively train models. This challenge has given rise to techniques like zero-shot learning (ZSL), which trains models on a set of "seen" classes and evaluates them on a set of "unseen" classes. Although ZSL has shown considerable potential, particularly with the employment of generative methods, its generalizability to real-world scenarios remains uncertain. The hypothesis of this work is that the performance of ZSL models is systematically influenced by the chosen "splits"; in particular, the statistical properties of the classes and attributes used in training. In this paper, we test this hypothesis by introducing the concepts of generalizability and robustness in attribute-based ZSL and carry out a variety of experiments to stress-test ZSL models against different splits. Our aim is to lay the groundwork for future research on ZSL models' generalizability, robustness, and practical applications. We evaluate the accuracy of state-of-the-art models on benchmark datasets and identify consistent trends in generalizability and robustness. We analyze how these properties vary based on the dataset type, differentiating between coarse- and fine-grained datasets, and our findings indicate significant room for improvement in both generalizability and robustness. Furthermore, our results demonstrate the effectiveness of dimensionality reduction techniques in improving the performance of state-of-the-art models in fine-grained datasets. Luca Rossi 0007, Maria Chiara Fiorentino, Adriano Mancini, Marina Paolanti, Riccardo Rosati 0002, Primo Zingaretti |
Neural Networks | 3 |
| 2023 | Deep Reinforced Navigation of Agents in 2D Platform Video Games
Emanuele Balloni, Marco Mameli, Adriano Mancini, Primo Zingaretti |
CGI (3) | 3 |
| 2023 | Investigation on the Encoder-Decoder Application for Mesh Generation
Marco Mameli, Emanuele Balloni, Adriano Mancini, Emanuele Frontoni, Primo Zingaretti |
CGI | 3 |
| 2023 | A hybrid feature learning approach based on convolutional kernels for ATM fault prediction using event-log dataabstractPredictive Maintenance (PdM) methods aim to facilitate the scheduling of maintenance work before equipment failure. In this context, detecting early faults in automated teller machines (ATMs) has become increasingly important since these machines are susceptible to various types of unpredictable failures. ATMs track execution status by generating massive event-log data that collect system messages unrelated to the failure event. Predicting machine failure based on event logs poses additional challenges, mainly in extracting features that might represent sequences of events indicating impending failures. Accordingly, feature learning approaches are currently being used in PdM, where informative features are learned automatically from minimally processed sensor data. However, a gap remains to be seen on how these approaches can be exploited for deriving relevant features from event-log-based data. To fill this gap, we present a predictive model based on a convolutional kernel (MiniROCKET and HYDRA) to extract features from the original event-log data and a linear classifier to classify the sample based on the learned features. The proposed methodology is applied to a significant real-world collected dataset. Experimental results demonstrated how one of the proposed convolutional kernels (i.e. HYDRA) exhibited the best classification performance (accuracy of 0.759 and AUC of 0.693). In addition, statistical analysis revealed that the HYDRA and MiniROCKET models significantly overcome one of the established state-of-the-art approaches in time series classification (InceptionTime), and three non-temporal ML methods from the literature. The predictive model was integrated into a container-based decision support system to support operators in the timely maintenance of ATMs. Víctor Manuel Vargas Yun, Riccardo Rosati 0002, César Hervás-Martínez, Adriano Mancini, Luca Romeo, Pedro Antonio Gutiérrez |
Eng. Appl. Artif. Intell. | 4 |
| 2020 | Machine learning-based design support system for the prediction of heterogeneous machine parameters in industry 4.0
Luca Romeo, Jelena Loncarski, Marina Paolanti, Gianluca Bocchini, Adriano Mancini, Emanuele Frontoni |
Expert Syst. Appl. | 5 |
| 2020 | Deep understanding of shopper behaviours and interactions using RGB-D visionabstractAbstract In retail environments, understanding how shoppers move about in a store’s spaces and interact with products is very valuable. While the retail environment has several favourable characteristics that support computer vision, such as reasonable lighting, the large number and diversity of products sold, as well as the potential ambiguity of shoppers’ movements, mean that accurately measuring shopper behaviour is still challenging. Over the past years, machine-learning and feature-based tools for people counting as well as interactions analytic and re-identification were developed with the aim of learning shopper skills based on occlusion-free RGB-D cameras in a top-view configuration. However, after moving into the era of multimedia big data, machine-learning approaches evolved into deep learning approaches, which are a more powerful and efficient way of dealing with the complexities of human behaviour. In this paper, a novel VRAI deep learning application that uses three convolutional neural networks to count the number of people passing or stopping in the camera area, perform top-view re-identification and measure shopper–shelf interactions from a single RGB-D video flow with near real-time performances has been introduced. The framework is evaluated on the following three new datasets that are publicly available: TVHeads for people counting, HaDa for shopper–shelf interactions and TVPR2 for people re-identification. The experimental results show that the proposed methods significantly outperform all competitive state-of-the-art methods (accuracy of 99.5% on people counting, 92.6% on interaction classification and 74.5% on re-id), bringing to different and significative insights for implicit and extensive shopper behaviour analysis for marketing applications. Marina Paolanti, Rocco Pietrini, Adriano Mancini, Emanuele Frontoni, Primo Zingaretti |
Mach. Vis. Appl. | 3 |
| 2019 | Empowered Optical Inspection by Using Robotic Manipulator in Industrial ApplicationsabstractNowadays the inspection of products at the end of line represents a critical phase. At this stage, it is necessary to look for defects in order to prevent the quality check to fail, and to provide information for improving the production as well. This task can be performed by using several sensing technologies, and the contactless optical inspection plays a key role. In this regard, the use of advanced robotic manipulators offers the capability to change the viewpoint of a given object and to inspect its multiple faces. We propose an approach that combines the use of photometric stereo to derive a 3D model of objects, empowered by the super-resolution that is applied on the original dataset (upstream) or on the normal images (downstream) in order to increase the quality of the final 3D model. The vision system is mounted on a robotic manipulator, able to grasp and change the viewpoint, thus offering a more complete view of the object to be inspected. The obtained results show that the developed solution increases the quality of the derived 3D models used for inspection tasks on different faces of the objects; this is achieved by using the manipulation ability offered by the adopted robotic platform. Alessandro Galdelli, Daniele Proietti Pagnotta, Adriano Mancini, Alessandro Freddi, Andrea Monteriù, Emanuele Frontoni |
IROS | 3 |
| 2019 | Design, Large-Scale Usage Testing, and Important Metrics for Augmented Reality Gaming ApplicationsabstractAugmented Reality (AR) offers the possibility to enrich the real world with digital mediated content, increasing in this way the quality of many everyday experiences. While in some research areas such as cultural heritage, tourism, or medicine there is a strong technological investment, AR for game purposes struggles to become a widespread commercial application. In this article, a novel framework for AR kid games is proposed, already developed by the authors for other AR applications such as Cultural Heritage and Arts. In particular, the framework includes different layers such as the development of a series of AR kid puzzle games in an intermediate structure which can be used as a standard for different applications development, the development of a smart configuration tool, together with general guidelines and long-life usage tests and metrics. The proposed application is designed for augmenting the puzzle experience, but can be easily extended to other AR gaming applications. Once the user has assembled the real puzzle, AR functionality within the mobile application can be unlocked, bringing to life puzzle characters, creating a seamless game that merges AR interactions with the puzzle reality. The main goals and benefits of this framework can be seen in the development of a novel set of AR tests and metrics in the pre-release phase (in order to help the commercial launch and developers), and in the release phase by introducing the measures for long-life app optimization, usage tests and hint on final users together with a measure to design policy, providing a method for automatic testing of quality and popularity improvements. Moreover, smart configuration tools, as part of the general framework, enabling multi-app and eventually also multi-user development, have been proposed, facilitating the serialization of the applications. Results were obtained from a large-scale user test with about 4 million users on a set of eight gaming applications, providing the scientific community a workflow for implicit quantitative analysis in AR gaming. Different data analytics developed on the data collected by the framework prove that the proposed approach is affordable and reliable for long-life testing and optimization. Roberto Pierdicca, Emanuele Frontoni, Primo Zingaretti, Adriano Mancini, Jelena Loncarski, Marina Paolanti |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2018 | Mechatronic System to Help Visually Impaired Users During Walking and RunningabstractAmbient assisted living and intelligent transportation systems are becoming strongly coupled. There is the necessity of improving the quality of life by developing inclusive mobility solutions for impaired people. In this paper, we focus on a monocular vision-based system to assist people during walking, jogging, and running in outdoor environments. The impaired user is guided along a path represented by a lane or line on a dedicated runway. We developed a set of image processing algorithms to extract lines/lanes to follow. The embedded system is based on a small camera and a board that is responsible for processing the images and communicating with the developed haptic device. The haptic device is formed by a set of two gloves equipped with vibration motors that drive the user to the right direction. The vibration sequences are generated according to a robotic-like controller, considering the user as a two wheel steering robot, where the rotational and translation velocity can be controlled. The results obtained show that the overall system is able to detect the right path and to provide the right stimuli to the user, by means of the gloves, up to a speed over 10 km/h. Adriano Mancini, Emanuele Frontoni, Primo Zingaretti |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2016 | Robust and affordable retail customer profiling by vision and radio beacon sensor fusion
Mirco Sturari, Daniele Liciotti, Roberto Pierdicca, Emanuele Frontoni, Adriano Mancini, Marco Contigiani, Primo Zingaretti |
Pattern Recognit. Lett. | 5 |
| 2015 | Embedded Multisensor System for Safe Point-to-Point Navigation of Impaired UsersabstractNew smart objects to improve the quality of life in the ambient assisted living (AAL) scenario are capturing the interest of researchers and companies. In particular, novel assistive technologies are being developed to make accessible street navigation to impaired people. The solution that we propose in this new application domain of intelligent transportation systems is a framework for a safe point-to-point navigation, owing to high-detailed road graphs, including sidewalks, crosswalks, and generic “obstacles.” The system is based on a low-cost modular sensor box (embedded hardware) interfaced with a mobile/phone application that acts as an intelligent navigator. The main novelty is the capability to sense the surrounding area while being able to perform a fast path replanning, owing to a real-time link to a remote server, if an obstacle is detected. The sensing is performed using different sensors, such as ultrasound, lidar, and a 77-GHz mid-range automotive radar (absolutely novel in the AAL context), which are processed and fused in the well-established robot operating system (ROS). We tested the framework by analyzing its performance in two different configurations and environments by using, respectively, a sonar and a laser rangefinder in a building scenario and a radar in an urban environment. Even if in both cases results demonstrated a quite good robustness in the obstacle detection with a quasi-real-time route replanning, we were mainly interested and succeeded in demonstrating the high flexibility and extensibility of our framework. Adriano Mancini, Emanuele Frontoni, Primo Zingaretti |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2014 | Feature Group Matching: a Novel Method to filter out Incorrect Local Feature MatchingsabstractThe importance of finding correct correspondences between two images is the major aspect in problems such as appearance-based robot localization and content-based image retrieval. Local feature matching has become a commonly used method to compare images, despite being highly probable that at least some of the matchings/correspondences it detects are incorrect. In this paper, we describe a novel approach to local feature matching, named Feature Group Matching (FGM), to select stable features and obtain a more reliable similarity value between two images. The proposed technique is demonstrated to be translational, rotational and scaling invariant. Experimental evaluation was performed on large and heterogeneous datasets of images using SIFT and SURF, the actual state-of-the-art feature extractors. Results show that FGM avoids almost 95% of incorrect matchings, reduces the visual aliasing (number of images considered similar) and increases both robotic localization and image retrieval accuracy on the average of 13%. Emanuele Frontoni, Adriano Mancini, Primo Zingaretti |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2011 | Hybrid object-based approach for land use/land cover mapping using high spatial resolution imageryabstractTraditionally, remote sensing has employed pixel-based classification techniques to deal with land use/land cover (LULC) studies. Generally, pixel-based approaches have been proven to work well with low spatial resolution imagery (e.g. Landsat or System Pour L'Observation de la Terre sensors). Now, however, commercially available high spatial resolution images (e.g. aerial Leica ADS40 and Vexcel UltraCam sensors, and satellite IKONOS, Quickbird, GeoEye and WorldView sensors) can be problematic for pixel-based analysis due to their tendency to oversample the scene. This is driving research towards object-based approaches. This article proposes a hybrid classification method with the aim of incorporating the advantages of supervised pixel-based classification into object-based approaches. The method has been developed for medium-scale (1:10,000) LULC mapping using ADS40 imagery with 1 m ground sampling distance. First, spatial information is incorporated into a pixel-based classification (AdaBoost classifier) by means of additional texture features (Haralick, Gabor, Law features), which can be selected ‘ad hoc’ according to optimal training samples (‘Relief-F’ approach, Mahalanobis distances). Then a rule-based approach sorts segmented regions into thematic CORINE Land Cover classes in terms of membership class percentages (a modified Winner-Takes-All approach) and shape parameters. Finally, ancillary data (roads, rivers, etc.) are exploited to increase classification accuracy. The experimental results show that the proposed hybrid approach allows the extraction of more LULC classes than conventional pixel-based methods, while improving classification accuracy considerably. A second contribution of this article is the assessment of classification reliability by implementing a stability map, in addition to confusion matrices. Eva Savina Malinverni, Anna Nora Tassetti, Adriano Mancini, Primo Zingaretti, Emanuele Frontoni, Annamaria Bernardini |
Int. J. Geogr. Inf. Sci. | 3 |
| 2010 | Road Change Detection from Multi-Spectral Aerial DataabstractThe paper presents a novel approach to automate the Change Detection (CD) problem for the specific task of road extraction. Manual approaches to CD fail in terms of the time for releasing updated maps; in the contrary, automatic approaches, based on machine learning and image processing techniques, allow to update large areas in a short time with an accuracy and precision comparable to those obtained by human operators. This work is focused on the road-graph update starting from aerial, multi-spectral data. Georeferenced, ground data, acquired by a GPS and an inertial sensor, are integrated with aerial data to speed up the change detector. After roads extraction by means of a binary AdaBoost classifier, the old road-graph is updated exploiting a particle filter. In particular this filter results very useful to link (track) parts of roads not extracted by the classifier due to the presence of occlusions (e.g., shadows, trees). Adriano Mancini, Emanuele Frontoni, Primo Zingaretti |
ICPR | 1 |
| 2009 | RoboBuntu: A Linux distribution for mobile roboticsabstractDuring last years Linux started to climb the market of operating systems (OSs), and Ubuntu, derived by Debian OS, has become a good alternative to common OSs like Windows XP or Vista. The mobile robotics scientific community makes use of Linux based OSs to avoid the lack of stability that affects Microsoft OSs, especially when real time conditions must be satisfied. In this paper we present the Linux distribution RoboBuntu, acronym formed by the union of ROBOt and uBUNTU, to overcome the almost totally independent robotic software platforms existing today. The key idea behind RoboBuntu is the integration of different tools for mobile robotics into an embedded Ubuntu distribution. Another important characteristics of RoboBuntu is that every ldquohard steprdquo, like installation and configuration of OS and tools, is hidden to common users. In particular, RoboBuntu can be used either by students or researchers, as LiveCd, permanent installation on standard hard drive or, more interesting, on a USB storage flash disk. Adriano Mancini, Emanuele Frontoni, Andrea Ascani, Primo Zingaretti |
ICRA | 1 |
| 2008 | Feature group matching for appearance-based localizationabstractLocal feature matching has become a commonly used method to compare images. For mobile robots, a reliable method for comparing images can constitute a key component for localization tasks. In this paper, we address the issues of appearance-based topological and metric localization by introducing a novel group matching approach to select less but more robust features to match the current robot view with reference images. Feature group matching is based on the consideration that feature descriptors together with spatial relations are more robust than classical approaches. Our datasets, each consisting of a large number of omnidirectional images, have been acquired over different day times (different lighting conditions) both in indoor and outdoor environments. The feature group matching outperforms the SIFT in indoor localization showing better performances both in the case of topological and metric localization. In outdoor SURF remains the best feature extraction method, as reported in literature. Andrea Ascani, Emanuele Frontoni, Adriano Mancini, Primo Zingaretti |
IROS | 3 |