VLDB 2026 Research / reviewers in the wild / expert
Marina Paolanti
dblp:181/3045
· DBLP profile ↗
22ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-5523-7174ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 6 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Orchestrating Generative AI Paradigms With Human-in-the-Loop for 3D GenerationabstractGenerative AI techniques are revolutionizing the creation of 3D and immersive content, yet challenges remain, such as achieving precise user control in 3D generation. Current text-to-3D and image-to-3D pipelines often produce outputs that deviate from user expectations, lacking the ability to refine or correct generated models effectively. To address these limitations, we propose Imagin3D, a novel human-in-the-loop (HITL) system that integrates Multimodal Large Language Models to enhance the controllability and adaptability of 3D content generation. Imagin3D leverages a Multi-View Question Answering module to evaluate the consistency of generated views with user-provided textual descriptions, enabling iterative refinement through guided inpainting while preserving multi-view consistency. This allows users to co-create 3D models, which are then synthesized into a final 3D asset using Neural Rendering. We validate Imagin3D through extensive quantitative evaluations and a comprehensive user study, demonstrating its effectiveness in improving usability, accuracy, and user satisfaction in interactive 3D generation tasks. Our results highlight the potential of HITL approaches to bridge the gap between AI-generated outputs and user intent, paving the way for more accessible and user-centered 3D generation workflows. Emanuele Balloni, Lorenzo Stacchio, Marina Paolanti, Primo Zingaretti, Roberto Pierdicca |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Made-In: An immersive human-in-the-loop analytics platform for enhancing creative processes in fashionabstractThe fashion industry is undergoing a digital transformation, driven by growing demands for sustainability, personalization and immersive experiences. In this paper, we present Made-In (Multimodal and Collaborative Artificial Intelligence for the Design of Inclusive and Sustainable Fashion): an immersive, human-in-the-loop analytics system designed to support fashion professionals in exploring, comparing and contextualizing product data across digital and social platforms. Unlike generative or simulation-based approaches, Made-In provides creative decision support by aggregating real-world data from luxury brand websites and social media. This enables designers and merchandisers to make informed, context-aware choices. The system comprises three core modules: a 3D configurator for visualizing product assortments; a collection grid interface for the comparative analysis of e-commerce data; and a social media trend detector based on deep learning pipelines for image classification, object detection and color clustering. Two curated datasets, one derived from Instagram and the other from fashion e-tailers, provide the system with analytics. A user study with domain experts confirms the platform’s usability and relevance for trend forecasting, sustainability evaluation and visual merchandising strategy. The results demonstrate that Made-In effectively bridges the gap between data analytics and human creativity in fashion, offering a scalable solution that aligns with EU goals for digital sustainability and inclusivity. • Immersive AI system that supports sustainable digital fashion exploration. • Real-time trend detection from Instagram enables geo-localized style insights. • Interactive 3D and collection grids enhance visual merchandising decisions. • AI modules extract product data, dominant colors, and sustainability tags. • Usability study confirms system effectiveness for designers and retailers. Emanuele Balloni, Rocco Pietrini, Michele Sasso, Emanuele Frontoni, Marina Paolanti |
Comput. Vis. Image Underst. | 5 |
| 2025 | A Neural Rendering system for fashion design process
Emanuele Balloni, Lorenzo Stacchio, Adriano Mancini, Emanuele Frontoni, Primo Zingaretti, Marina Paolanti |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | OutfitAI: shop the outfit with a deep learning-based intelligent expert systemabstractAbstract In an age where consumer preferences are as diverse as they are dynamic, the ability to offer personalized fashion recommendations at scale remains a significant challenge for retailers. Consumers seek a shopping experience that not only understands their unique style preferences but also dynamically adapts to their evolving tastes. The fashion industry is at a crossroads, facing increasing consumer demand for personalization, sustainability and transparency in a rapidly evolving digital marketplace. Traditional retail practices, while rich in tradition and artistry, often struggle to up-to-date with the rapidly, ethically-conscious and technology-driven expectations of today’s consumers. “OutfitAI” is designed to address these challenges by leveraging the power of deep learning to revolutionize the fashion retail experience. By automating the process of background removal in fashion images, using advanced algorithms for personalized product matching, and integrating sustainability filters into the product discovery process, OutfitAI aims to deliver a shopping experience that is not only personalized and engaging, but also aligned with the ethical and environmental values of the contemporary consumer. Unlike existing solutions, OutfitAI uses state-of-the-art semantic segmentation for precise background removal, enabling detailed feature extraction from fashion images. This process enables accurate matching of user-uploaded images with similar fashion items from an extensive database of eco-friendly and ethically produced products sourced from leading e-tailers. Setting itself apart from the current state of the art, OutfitAI places a strong emphasis on ethical data use and privacy, implementing robust measures to ensure user privacy and transparency. It also pioneers the integration of sustainability into the digital fashion discovery process, promoting responsible consumption patterns among users. Through a comprehensive system architecture that combines technical innovation with a commitment to ethics and sustainability, OutfitAI not only addresses the technological needs of the fashion retail industry, but also responds to the growing demand for more responsible and transparent consumer technologies. Emanuele Balloni, Rocco Pietrini, Emanuele Frontoni, Adriano Mancini, Marina Paolanti |
Multim. Tools Appl. | 5 |
| 2025 | RenderGAN: Enhancing Real-time Rendering Efficiency with Deep LearningabstractIn the domain of computer graphics, achieving high visual quality in real-time rendering remains a formidable challenge due to the inherent time-quality tradeoff. Conventional real-time rendering engines sacrifice visual fidelity for interactive performance, while image generation using path-tracing techniques can be exceedingly time-consuming. In this article, we introduce RenderGAN, a deep learning-based solution designed to address this critical challenge in real-time rendering. RenderGAN uses G-Buffers and information from a real-time rendering engine as inputs to produce output images with exceptional visual fidelity. Its encoder–decoder architecture, trained using the Generative Adversarial Network (GAN) framework with perceptual loss, enhances image realism. To evaluate RenderGAN’s effectiveness, we quantitatively compare the generated images with those of a path-tracing engine, obtaining a remarkable Universal Image Quality Index (UIQI) value of 0.898. RenderGAN’s open source nature fosters collaboration, driving advancements in real-time computer graphics and rendering techniques. By bridging the gap between real-time and path-tracing rendering, RenderGAN opens new horizons for accelerated image generation, inspiring innovation and unlocking the full potential of real-time visual experiences. Project page: https://github.com/marcomameli1992/RenderNet Marco Mameli, Marina Paolanti, Adriano Mancini, Primo Zingaretti, Roberto Pierdicca |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | MineVRA: Exploring the Role of Generative AI-Driven Content Development in XR Environments through a Context-Aware ApproachabstractThe convergence of Artificial Intelligence (AI), Computer Vision (CV), Computer Graphics (CG), and Extended Reality (XR) is driving innovation in immersive environments. A key challenge in these environments is the creation of personalized 3D assets, traditionally achieved through manual modeling, a time-consuming process that often fails to meet individual user needs. More recently, Generative AI (GenAI) has emerged as a promising solution for automated, context-aware content generation. In this paper, we present MineVRA (Multimodal generative artificial iNtelligence for contExt-aware Virtual Reality Assets), a novel Human-In-The-Loop (HITL) XR framework that integrates GenAI to facilitate coherent and adaptive 3D content generation in immersive scenarios. To evaluate the effectiveness of this approach, we conducted a comparative user study analyzing the performance and user satisfaction of GenAI-generated 3D objects compared to those generated by Sketchfab in different immersive contexts. The results suggest that GenAI can significantly complement traditional 3D asset libraries, with valuable design implications for the development of human-centered XR environments. Lorenzo Stacchio, Emanuele Balloni, Emanuele Frontoni, Marina Paolanti, Primo Zingaretti, Roberto Pierdicca |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Embedding AI ethics into the design and use of computer vision technology for consumer's behaviour understandingabstractArtificial Intelligence (AI) techniques are becoming more and more sophisticated showing the potential to deeply understand and predict consumer behaviour in a way to boost the retail sector; however, retail-sensitive considerations underpinning their deployment have been poorly explored to date. This paper explores the application of AI technologies in the retail sector, focusing on their potential to enhance decision-making processes by preventing major ethical risks inherent to them, such as the propagation of bias and systems’ lack of explainability. Drawing on recent literature on AI ethics, this study proposes a methodological path for the design and the development of trustworthy, unbiased, and more explainable AI systems in the retail sector. Such framework grounds on European (EU) AI ethics principles and addresses the specific nuances of retail applications. To do this, we first examine the VRAI framework, a deep learning model used to analyse shopper interactions, people counting and re-identification, to highlight the critical need for transparency and fairness in AI operations. Second, the paper proposes actionable strategies for integrating high-level ethical guidelines into practical settings, and particularly, to mitigate biases leading to unfair outcomes in AI systems and improve their explainability. By doing so, the paper aims to show the key added value of embedding AI ethics requirements into AI practices and computer vision technology to truly promote technically and ethically robust AI in the retail domain. • Establishment of AI ethics guidelines for the retail domain. • Design of a detailed framework for AI ethics. • Evaluation of the VRAI framework on new explainability metrics. • Recommendations for AI transparency improvements. • Integration of ethics into the design of AI systems. Simona Tiribelli, Benedetta Giovanola, Rocco Pietrini, Emanuele Frontoni, Marina Paolanti |
Comput. Vis. Image Underst. | 5 |
| 2024 | Social4Fashion: An intelligent expert system for forecasting fashion trends from social media contents
Emanuele Balloni, Rocco Pietrini, Matteo Fabiani, Emanuele Frontoni, Adriano Mancini, Marina Paolanti |
Expert Syst. Appl. | 6 |
| 2024 | DeepReality: An open source framework to develop AI-based augmented reality applicationsabstractAugmented reality (AR) and Artificial Intelligence (AI) are technologies pioneers in innovation and alteration in several domains. AR allows the creation of an entirely new and interactive experience for users. However, there are several drawbacks in developing AR applications, such as the marker identification process and the creation of content itself. These are very time-consuming procedures and require ad-hoc development. The advantages of using AI to solve AR limitations have recently been explored in literature. Motivated by these findings, in this paper it is proposed DeepReality, a software toolkit plug-in for Unity 3D. It is conceived for allowing developers to integrate any Deep Learning (DL) models into Unity, through AR Foundation and Barracuda inference engine. DeepReality is aimed at simplifying and streamlining the usage of DL models in conjunction with AR. As such, users skilled in Unity and DL can easily create mobile applications (iOS and Android) to: extract visual features of real-world objects (framed with the device camera) via DL; Show on-screen content on top of those real-world objects, via AR. DeepReality performs object semantic processing within the scene, and extended semantic effects for incongruent objects, overcoming the environmental tracking, which is feature-based. In order to test DeepReality usability, experiments have been performed on the execution time and memory usage data, demonstrating the feasibility and possibility of integrating and using DNNs models in mobile applications for AR. The complexity analysis confirms that DeepReality can be completely executed on mobile devices. DeepReality is also open-source and it is freely available in the Unity asset store. By fostering accessible AI-AR integration, DeepReality addresses key shortcomings in existing approaches, encapsulating contributions such as versatile DL integration, open-source accessibility, operational validation, and comprehensive metrics analysis. DeepReality empowers developers to transcend boundaries, enriching AR applications with AI’s transformative potential. Our proposed framework fosters benchmarking, comparison, and a future harmonised by AR-AI synergy. Roberto Pierdicca, Flavio Tonetto, Marina Paolanti, Marco Mameli, Riccardo Rosati 0002, Primo Zingaretti |
Expert Syst. Appl. | 3 |
| 2024 | Shelf Management: A deep learning-based system for shelf visual monitoringabstractShelf monitoring plays a key role in optimizing retail shelf layout, enhancing the customer shopping experience and maximizing profit margins. The process of automating shelf audit involves the detection, localization and recognition of objects on store shelves, including diverse products with varying attributes in unconstrained environments. This facilitates the assessment of planogram compliance. Accurate product localization within shelves requires the identification of specific shelf rows. To address the current technological challenges, we introduce “Shelf Management”, a deep learning-based system that is carefully tailored to redesign shelf monitoring practices. Our system can navigate the complexities of shelf monitoring by using advanced deep learning techniques and object detection and recognition models. In addition, a complex semantic module enhances the accuracy of detecting and assigning products to their designated shelf rows and locations. In particular, we recognize the lack of finely annotated datasets at the SKU level. As a contribution to the field, we provide annotations for two novel datasets: SHARD (SHelf mAnagement Row Dataset) and SHAPE (SHelf mAnagement Product dataset). These datasets not only provide valuable resources, but also serve as benchmarks for further research in the field of retail. A complete pipeline is designed using a RetinaNet architecture for object detection with 0.752 mAP, followed by a Deep Hough transform to detect shelf rows as semantic lines with an F1 score of 97%, and a product recognition step using a MobileNetV3 architecture trained with triplet loss and used as a feature extractor together with FAISS for fast image retrieval with an accuracy of 93% on top-1 recognition. Localization is achieved using a deterministic approach based on product detection and shelf row detection. Source code and datasets are available at https://github.com/rokopi-byte/shelf_management. Rocco Pietrini, Marina Paolanti, Adriano Mancini, Emanuele Frontoni, Primo Zingaretti |
Expert Syst. Appl. | 2 |
| 2024 | GREEN PATH: an expert system for space planning and design by the generation of human trajectoriesabstractAbstract Public space is usually conceived as where people live, perceive, and interact with other people. The environment affects people in several different ways as well. The impact of environmental problems on humans is significant, affecting all human activities, including health and socio-economic development. Thus, there is a need to rethink how space is used. Dealing with the important needs raised by climate emergency, pandemic and digitization, the contributions of this paper consist in the creation of opportunities for developing generative approaches to space design and utilization. It is proposed GREEN PATH, an intelligent expert system for space planning. GREEN PATH uses human trajectories and deep learning methods to analyse and understand human behaviour for offering insights to layout designers. In particular, a Generative Adversarial Imitation Learning (GAIL) framework hybridised with classical reinforcement learning methods is proposed. An example of the classical reinforcement learning method used is continuous penalties, which allow us to model the shape of the trajectories and insert a bias, which is necessary for the generation, into the training. The structure of the framework and the formalisation of the problem to be solved allow for the evaluation of the results in terms of generation and prediction. The use case is a chosen retail domain that will serve as a demonstrator for optimising the layout environment and improving the shopping experience. Experiments were assessed on shoppers’ trajectories obtained from four different stores, considering two years. Marina Paolanti, Davide Manco, Rocco Pietrini, Emanuele Frontoni |
Multim. Tools Appl. | 1 |
| 2024 | Generalizability and robustness evaluation of attribute-based zero-shot learningabstractIn the field of deep learning, large quantities of data are typically required to effectively train models. This challenge has given rise to techniques like zero-shot learning (ZSL), which trains models on a set of "seen" classes and evaluates them on a set of "unseen" classes. Although ZSL has shown considerable potential, particularly with the employment of generative methods, its generalizability to real-world scenarios remains uncertain. The hypothesis of this work is that the performance of ZSL models is systematically influenced by the chosen "splits"; in particular, the statistical properties of the classes and attributes used in training. In this paper, we test this hypothesis by introducing the concepts of generalizability and robustness in attribute-based ZSL and carry out a variety of experiments to stress-test ZSL models against different splits. Our aim is to lay the groundwork for future research on ZSL models' generalizability, robustness, and practical applications. We evaluate the accuracy of state-of-the-art models on benchmark datasets and identify consistent trends in generalizability and robustness. We analyze how these properties vary based on the dataset type, differentiating between coarse- and fine-grained datasets, and our findings indicate significant room for improvement in both generalizability and robustness. Furthermore, our results demonstrate the effectiveness of dimensionality reduction techniques in improving the performance of state-of-the-art models in fine-grained datasets. Luca Rossi 0007, Maria Chiara Fiorentino, Adriano Mancini, Marina Paolanti, Riccardo Rosati 0002, Primo Zingaretti |
Neural Networks | 4 |
| 2022 | An offline parallel architecture for forensic multimedia classificationabstractAbstract Nowadays, the volume of the multimedia heterogeneous evidence presented for digital forensic analysis has significantly increased, thus requiring the application of big data technologies, cloud-based forensics services, as well as Machine Learning (ML) techniques. In digital forensics domain, ML algorithms have been applied for cybercrime investigation such as child abuse investigations, malware classification, and image forensics. This paper addresses this issues and deals with forensic analysis of digital images and videos. In particular, this work aims at proposing a multimedia classification tool with a parallel software architecture for a fast inspection, which is easy to use (to be used by officers during a search), requires limited hardware resources and it is built on an open-source software to limit its costs. Moreover, this tool must be able to quickly inspect multiple devices at a time. When positives are found in a device, such device will be seized for a deeper analysis later in the lab. It will not be seized otherwise, reducing the inconvenience for the suspect as well as the time required for the next analysis phase. As a case study, we focus on the identification of child pornography images. Experimental results show that the proposed architecture is capable of guaranteeing a high recall, a fast process and high performances in real scenarios. Luca Spalazzi, Marina Paolanti, Emanuele Frontoni |
Multim. Tools Appl. | 2 |
| 2022 | SeSAME: Re-identification-based ambient intelligence system for museum environment
Marina Paolanti, Roberto Pierdicca, Rocco Pietrini, Massimo Martini, Emanuele Frontoni |
Pattern Recognit. Lett. | 1 |
| 2021 | Human trajectory prediction and generation using LSTM models and GANs
Luca Rossi 0007, Marina Paolanti, Roberto Pierdicca, Emanuele Frontoni |
Pattern Recognit. | 2 |
| 2020 | Weight Estimation from an RGB-D camera in top-view configurationabstractThe development of so-called soft-biometrics aims at providing information related to the physical and behavioural characteristics of a person. This paper focuses on body weight estimation based on the observation from a top-view RGB-D camera. In fact, the capability to estimate the weight of a person can be of help in many different applications, from health-related scenarios, to business intelligence and retail analytics. To deal with this issue, a TVWE (Top-View Weight Estimation) framework is proposed with the aim of predicting the weight. The approach relies on the adoption of Deep Neural Networks (DNNs) that have been trained on depth data. Each network has also been modified in their top section to replace classification with prediction inference. The performance of five state-of-art DNNs have been compared, namely VGG16, ResNet, Inception, DenseNet and Efficient-Net. In addition, a convolutional auto-encoder has also been included for completeness. Considering the limited literature in this domain, the TVWE framework has been evaluated on a new publicly available dataset: “VRAI Weight estimation Dataset”, which also collects, for each subject, labels related to weight, gender, and height. The experimental results have demonstrated that the proposed methods are suitable for this task, bringing different and significant insights for the application of the solution in different domains. Marco Mameli, Marina Paolanti, Nicola Conci, Filippo Tessaro, Emanuele Frontoni, Primo Zingaretti |
ICPR | 2 |
| 2020 | Automatic Classification of Human Granulosa Cells in Assisted Reproductive Technology using vibrational spectroscopy imagingabstractIn the field of reproductive technology, the biochemical composition of female gametes has been successfully investigated with the use of vibrational spectroscopy. Currently, in assistive reproductive technology (ART), there are no shared criteria for the choice of oocyte, and automatic classification methods for the best quality oocytes have not yet been applied. In this paper, considering the lack of criteria in Assisted Reproductive Technology (ART), we use Machine Learning (ML) techniques to predict oocyte quality for a successful pregnancy. To improve the chances of successful implantation and minimize any complications during the pregnancy, Fourier transform infrared microspectroscopy (FTIRM) analysis has been applied on granulosa cells (GCs) collected along with the oocytes during oocyte aspiration, as it is routinely done in ART, and specific spectral biomarkers were selected by multivariate statistical analysis. A proprietary biological reference dataset (BRD) was successfully collected to predict the best oocyte for a successful pregnancy. Personal health information are stored, maintained and backed up using a cloud computing service. Using a user-friendly interface, the user will evaluate whether or not the selected oocyte will have a positive result. This interface includes a dashboard for retrospective analysis, reporting, real-time processing, and statistical analysis. The experimental results are promising and confirm the efficiency of the method in terms of classification metrics: precision, recall, and F1-score (F1) measures. Marina Paolanti, Marco Mameli, Emanuele Frontoni, Giorgia Gioacchini, Elisabetta Giorgini, Valentina Notarstefano, Carlotta Zacà, Oliana Carnevali, Andrea Bonarini |
ICPR | 1 |
| 2020 | Machine learning-based design support system for the prediction of heterogeneous machine parameters in industry 4.0
Luca Romeo, Jelena Loncarski, Marina Paolanti, Gianluca Bocchini, Adriano Mancini, Emanuele Frontoni |
Expert Syst. Appl. | 3 |
| 2020 | Deep understanding of shopper behaviours and interactions using RGB-D visionabstractAbstract In retail environments, understanding how shoppers move about in a store’s spaces and interact with products is very valuable. While the retail environment has several favourable characteristics that support computer vision, such as reasonable lighting, the large number and diversity of products sold, as well as the potential ambiguity of shoppers’ movements, mean that accurately measuring shopper behaviour is still challenging. Over the past years, machine-learning and feature-based tools for people counting as well as interactions analytic and re-identification were developed with the aim of learning shopper skills based on occlusion-free RGB-D cameras in a top-view configuration. However, after moving into the era of multimedia big data, machine-learning approaches evolved into deep learning approaches, which are a more powerful and efficient way of dealing with the complexities of human behaviour. In this paper, a novel VRAI deep learning application that uses three convolutional neural networks to count the number of people passing or stopping in the camera area, perform top-view re-identification and measure shopper–shelf interactions from a single RGB-D video flow with near real-time performances has been introduced. The framework is evaluated on the following three new datasets that are publicly available: TVHeads for people counting, HaDa for shopper–shelf interactions and TVPR2 for people re-identification. The experimental results show that the proposed methods significantly outperform all competitive state-of-the-art methods (accuracy of 99.5% on people counting, 92.6% on interaction classification and 74.5% on re-id), bringing to different and significative insights for implicit and extensive shopper behaviour analysis for marketing applications. Marina Paolanti, Rocco Pietrini, Adriano Mancini, Emanuele Frontoni, Primo Zingaretti |
Mach. Vis. Appl. | 1 |
| 2019 | Design, Large-Scale Usage Testing, and Important Metrics for Augmented Reality Gaming ApplicationsabstractAugmented Reality (AR) offers the possibility to enrich the real world with digital mediated content, increasing in this way the quality of many everyday experiences. While in some research areas such as cultural heritage, tourism, or medicine there is a strong technological investment, AR for game purposes struggles to become a widespread commercial application. In this article, a novel framework for AR kid games is proposed, already developed by the authors for other AR applications such as Cultural Heritage and Arts. In particular, the framework includes different layers such as the development of a series of AR kid puzzle games in an intermediate structure which can be used as a standard for different applications development, the development of a smart configuration tool, together with general guidelines and long-life usage tests and metrics. The proposed application is designed for augmenting the puzzle experience, but can be easily extended to other AR gaming applications. Once the user has assembled the real puzzle, AR functionality within the mobile application can be unlocked, bringing to life puzzle characters, creating a seamless game that merges AR interactions with the puzzle reality. The main goals and benefits of this framework can be seen in the development of a novel set of AR tests and metrics in the pre-release phase (in order to help the commercial launch and developers), and in the release phase by introducing the measures for long-life app optimization, usage tests and hint on final users together with a measure to design policy, providing a method for automatic testing of quality and popularity improvements. Moreover, smart configuration tools, as part of the general framework, enabling multi-app and eventually also multi-user development, have been proposed, facilitating the serialization of the applications. Results were obtained from a large-scale user test with about 4 million users on a set of eight gaming applications, providing the scientific community a workflow for implicit quantitative analysis in AR gaming. Different data analytics developed on the data collected by the framework prove that the proposed approach is affordable and reliable for long-life testing and optimization. Roberto Pierdicca, Emanuele Frontoni, Primo Zingaretti, Adriano Mancini, Jelena Loncarski, Marina Paolanti |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2018 | Convolutional Networks for Semantic Heads Segmentation using Top-View Depth Data in Crowded EnvironmentabstractDetecting and tracking people is a challenging task in a persistent crowded environment (i.e. retail, airport, station, etc.) for human behaviour analysis of security purposes. This paper introduces an approach to track and detect people in cases of heavy occlusions based on CNNs for semantic segmentation using top-view depth visual data. The purpose is the design of a novel U-Net architecture, U-Net3, that has been modified compared to the previous ones at the end of each layer. In particular, a batch normalization is added after the first ReLU activation function and after each max-pooling and up-sampling functions. The approach was applied and tested on a new and public available dataset, TVHeads Dataset, consisting of depth images of people recorded from an RGB-D camera installed in top-view configuration. Our variant outperforms baseline architectures while remaining computationally efficient at inference time. Results show high accuracy, demonstrating the effectiveness and suitability of our approach. Daniele Liciotti, Marina Paolanti, Rocco Pietrini, Emanuele Frontoni, Primo Zingaretti |
ICPR | 2 |
| 2017 | Customer Experience: A Design Approach and Supporting Platform
Maura Mengoni, Emanuele Frontoni, Luca Giraldi, Silvia Ceccacci, Roberto Pierdicca, Marina Paolanti |
PRO-VE | 6 |