Lorenzo Stacchio

dblp:280/0408 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-9341-7651ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RELD: Regularization by Latent Denoising
Pasquale Cascarano, Lorenzo Stacchio, Andrea Sebastiani, Alessandro Benfenati, Ulugbek Kamilov, Gustavo Marfia
IEEE Signal Process. Lett.2
2026 Orchestrating Generative AI Paradigms With Human-in-the-Loop for 3D Generation
abstract
Generative AI techniques are revolutionizing the creation of 3D and immersive content, yet challenges remain, such as achieving precise user control in 3D generation. Current text-to-3D and image-to-3D pipelines often produce outputs that deviate from user expectations, lacking the ability to refine or correct generated models effectively. To address these limitations, we propose Imagin3D, a novel human-in-the-loop (HITL) system that integrates Multimodal Large Language Models to enhance the controllability and adaptability of 3D content generation. Imagin3D leverages a Multi-View Question Answering module to evaluate the consistency of generated views with user-provided textual descriptions, enabling iterative refinement through guided inpainting while preserving multi-view consistency. This allows users to co-create 3D models, which are then synthesized into a final 3D asset using Neural Rendering. We validate Imagin3D through extensive quantitative evaluations and a comprehensive user study, demonstrating its effectiveness in improving usability, accuracy, and user satisfaction in interactive 3D generation tasks. Our results highlight the potential of HITL approaches to bridge the gap between AI-generated outputs and user intent, paving the way for more accessible and user-centered 3D generation workflows.
Emanuele Balloni, Lorenzo Stacchio, Marina Paolanti, Primo Zingaretti, Roberto Pierdicca
IEEE Trans. Vis. Comput. Graph.2
2025 A Neural Rendering system for fashion design process
Emanuele Balloni, Lorenzo Stacchio, Adriano Mancini, Emanuele Frontoni, Primo Zingaretti, Marina Paolanti
Eng. Appl. Artif. Intell.2
2025 MineVRA: Exploring the Role of Generative AI-Driven Content Development in XR Environments through a Context-Aware Approach
abstract
The convergence of Artificial Intelligence (AI), Computer Vision (CV), Computer Graphics (CG), and Extended Reality (XR) is driving innovation in immersive environments. A key challenge in these environments is the creation of personalized 3D assets, traditionally achieved through manual modeling, a time-consuming process that often fails to meet individual user needs. More recently, Generative AI (GenAI) has emerged as a promising solution for automated, context-aware content generation. In this paper, we present MineVRA (Multimodal generative artificial iNtelligence for contExt-aware Virtual Reality Assets), a novel Human-In-The-Loop (HITL) XR framework that integrates GenAI to facilitate coherent and adaptive 3D content generation in immersive scenarios. To evaluate the effectiveness of this approach, we conducted a comparative user study analyzing the performance and user satisfaction of GenAI-generated 3D objects compared to those generated by Sketchfab in different immersive contexts. The results suggest that GenAI can significantly complement traditional 3D asset libraries, with valuable design implications for the development of human-centered XR environments.
Lorenzo Stacchio, Emanuele Balloni, Emanuele Frontoni, Marina Paolanti, Primo Zingaretti, Roberto Pierdicca
IEEE Trans. Vis. Comput. Graph.1
2024 Investigating eXtended Reality-powered Digital Twins for Sequential Instruction Learning: the Case of the Rubik's Cube
abstract
Educational practices are increasingly experimenting with eXtended Reality (XR) paradigms to offer novel opportunities for boundaryless learning experiences with real-time interactions in immersive environments. Digital Twins (DT) are also gaining traction in this field to facilitate personalized learning experiences. However, a still unexplored space in learning frameworks amounts to the one where XR intersects with DTs. This work wants to move a step in such a direction with the design, implementation, and test of a DT-driven XR framework to learn procedural tasks. The framework offers three distinct learning modalities where virtual and physical interactions enhance learning retention by engaging users actively in digital and real-world environments. We contextualize such a framework for procedural task learning through one of its pivotal use cases: learning Rubik’s Cube notations. To evaluate and compare the effectiveness of these modalities, we perform an experimental user campaign evaluating short-term skill retention, performance accuracy, usability, and cognitive load of each of them. We then provide an extensive statistical analysis to compare each kind of guidance while analyzing correlations between the examined variables, offering insights into optimizing instructional methodologies within XR-based educational frameworks.
Shirin Hajahmadi, Lorenzo Stacchio, Alessandro Giacchè, Pasquale Cascarano, Gustavo Marfia
ISMAR2
2024 Analyzing cultural relationships visual cues through deep learning models in a cross-dataset setting
abstract
Abstract To study the evolution of specific cultures and times different kinds of pictures could be adopted. Family album photos may reveal socio-historical insights regarding those specific cultures and times. Along this path, this work addresses the problem of automatically dating an image by resorting to the analysis of an analog family album photo dataset. In particular, the IMAGO collection, which contains Italian photos shot in the 20th century, was considered. Thanks to the IMAGO dataset, it was possible to apply different deep learning-based architectures to date images belonging to photo albums without needing any other sources of information. In addition, we carried out cross-dataset experiments, which also involved models trained on American datasets, observing temporal shifts which may be due to known intercultural influences. We further explore such a possibility by qualitatively analyzing the cross-dataset interpretation of the trained deep-learning models with the Uniform Manifold Approximation and Projection (UMAP) algorithm. In conclusion, deep learning models revealed their potential in terms of possible applications to intercultural research, from different points of view.
Lorenzo Stacchio, Alessia Angeli, Giuseppe Lisanti, Gustavo Marfia
Neural Comput. Appl.1
2024 Making paper labels smart for augmented wine recognition
abstract
Abstract An invisible layer of knowledge is progressively growing with the emergence of situated visualizations and reality-based information retrieval systems. In essence, digital content will overlap with real-world entities, eventually providing insights into the surrounding environment and useful information for the user. The implementation of such a vision may appear close, but many subtle details separate us from its fulfillment. This kind of implementation, as the overlap between rendered virtual annotations and the camera’s real-world view, requires different computer vision paradigms for object recognition and tracking which often require high computing power and large-scale datasets of images. Nevertheless, these resources are not always available, and in some specific domains, the lack of an appropriate reference dataset could be disruptive for a considered task. In this particular scenario, we here consider the problem of wine recognition to support an augmented reading of their labels. In fact, images of wine bottle labels may not be available as wineries periodically change their designs, product information regulations may vary, and specific bottles may be rare, making the label recognition process hard or even impossible. In this work, we present augmented wine recognition, an augmented reality system that exploits optical character recognition paradigms to interpret and exploit the text within a wine label, without requiring any reference image. Our experiments show that such a framework can overcome the limitations posed by image retrieval-based systems while exhibiting a comparable performance.
Alessia Angeli, Lorenzo Stacchio, Lorenzo Donatiello, Alessandro Giacchè, Gustavo Marfia
Vis. Comput.2
2023 Unity-VRlines: Towards a Modular eXtended Reality Unity Flight Simulator
Giuseppe Di Maria, Lorenzo Stacchio, Gustavo Marfia
ICEC2
2022 Toward a Holistic Approach to the Socio-historical Analysis of Vernacular Photos
abstract
Although one of the most popular practices in photography since the end of the 19th century, an increase in scholarly interest in family photo albums dates back to the early 1980s. Such collections of photos may reveal sociological and historical insights regarding specific cultures and times. They are, however, in most cases scattered among private homes and only available on paper or photographic film, thus making their collection and analysis by historians, socio-cultural anthropologists, and cultural theorists very cumbersome. Computer-based methodologies could aid such a process in various ways, speeding up the cataloging step, for example, with the use of modern computer vision techniques. We here investigate such an approach, introducing the design and development of a multimedia application that may automatically catalog vernacular pictures drawn from family photo albums. To this aim, we introduce the IMAGO dataset, which is composed of photos belonging to family albums assembled at the University of Bologna’s Rimini campus since 2004. Exploiting the proposed application, IMAGO has offered the opportunity of experimenting with photos taken between the years 1845 and 2009. In particular, it has been possible to estimate their socio-historical content, i.e., the dates and contexts of the images, without resorting to any other sources of information. Exceeding our initial expectations, such an approach has revealed its merit not only in terms of performance but also in terms of the foreseeable implications for the benefit of socio-historical research. To the best of our knowledge, this contribution is among the few that move along this path at the intersection of socio-historical studies, multimedia computing, and artificial intelligence.
Lorenzo Stacchio, Alessia Angeli, Giuseppe Lisanti, Daniela Calanca, Gustavo Marfia
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Empowering Digital Twins with eXtended Reality Collaborations
abstract
The advancements of Artificial Intelligence, Big Data Analytics, and the Internet of Things paved the path to the emergence and use of Digital Twins (DTs) as technologies to “twin” the life of a physical entity in different fields, ranging from industry to healthcare. At the same time, the advent of eXtended Reality (XR) in industrial and consumer electronics has provided novel paradigms that may be put to good use to visualize and interact with DTs. XR technologies can support human-to-human interactions for training and remote assistance and could transform DTs into collaborative intelligence tools. We here present the Human Collaborative Intelligence empowered Digital Twin framework (HCLINT-DT) integrating human annotations (e.g., textual and vocal) to allow the creation of an all-in-one-place resource to preserve such knowledge. This framework could be adopted in many fields, supporting users to learn how to carry out an unknown process or explore others’ past experiences. The assessment of such a framework has involved implementing a DT supporting human annotations, reflected in both the physical world (Augmented Reality) and the virtual one (Virtual Reality). The outcomes of the interface design assessment confirm the interest in developing HCLINT-DT-based applications. Finally, we evaluated how the proposed framework could be translated into a manufacturing context.
Lorenzo Stacchio, Alessia Angeli, Gustavo Marfia
Virtual Real. Intell. Hardw.1