Davide Morelli

dblp:55/4866 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Sketch2Stitch: GANs for Abstract Sketch-Based Dress Synthesis
abstract
In the realm of creative expression, not everyone possesses the gift of effortlessly translating their imaginative visions into flawless sketches. More often than not, the outcome resembles an abstract, perhaps even slightly distorted representation. The art of producing impeccable sketches is not only challenging but also a time-consuming process. Our work is the first of this kind in transforming abstract, sometimes deformed garment sketches into photorealistic catalog images, to empower the everyday individual to become their own fashion designer. We create Sketch2Stitch, a dataset featuring over 65,000 abstract sketch images generated from garments of Dress Code [40] and VITON HD [6], two benchmark datasets in the virtual try-on task. Sketch2Stitch is the first dataset in the literature to provide abstract sketches in the fashion domain. We propose a StyleGAN-based generative framework that bridges freehand sketching with photorealistic garment synthesis. We demonstrate that our framework allows users to sketch rough outlines and optionally provide color hints, producing realistic designs in seconds. Experimental results demonstrate, both quantitatively and qualitatively, that the proposed framework achieves superior performance against various baselines and existing methods on both subsets of our dataset. Our work highlights a pathway toward AI-assisted fashion design tools, democratizing garment ideation for students, independent designers, and casual creators.
Faizan Farooq Khan, Eslam Abdelrahman, Davide Morelli, Marcella Cornia, Rita Cucchiara, Mohamed Elhoseiny 0001
WACV3
2026 Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing
abstract
Fashion illustration is a crucial medium for designers to convey their creative vision and transform design concepts into tangible representations that showcase the interplay between clothing and the human body. In the context of fashion design, computer vision techniques have the potential to enhance and streamline the design process. Departing from prior research primarily focused on virtual try-on, this article tackles the task of multimodal-conditioned fashion image editing. Our approach aims to generate human-centric fashion images guided by multimodal prompts, including text, human body poses, garment sketches, and fabric textures. To address this problem, we propose extending latent diffusion models to incorporate these multiple modalities and modifying the structure of the denoising network, taking multimodal prompts as input. To condition the proposed architecture on fabric textures, we employ textual inversion techniques and let diverse cross-attention layers of the denoising network attend to textual and texture information, thus incorporating different granularity conditioning details. Given the lack of datasets for the task, we extend two existing fashion datasets, Dress Code and VITON-HD, with multimodal annotations. Experimental evaluations demonstrate the effectiveness of our proposed approach in terms of realism and coherence concerning the provided multimodal inputs.
Alberto Baldrati, Davide Morelli, Marcella Cornia, Marco Bertini 0001, Rita Cucchiara
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation
abstract
In recent years, the fashion industry has increasingly adopted AI technologies to enhance customer experience, driven by the proliferation of e-commerce platforms and virtual applications. Among the various tasks, virtual try-on and multimodal fashion image editing–which utilizes diverse input modalities such as text, garment sketches, and body poses–have become a key area of research. Diffusion models have emerged as a leading approach for such generative tasks, offering superior image quality and diversity. However, most existing virtual try-on methods rely on having a specific garment input, which is often impractical in real-world scenarios where users may only provide textual specifications. To address this limitation, in this work we introduce Fashion Retrieval-Augmented Generation (Fashion-RAG), a novel method that enables the customization of fashion items based on user preferences provided in textual form. Our approach retrieves multiple garments that match the input specifications and generates a personalized image by incorporating attributes from the retrieved items. To achieve this, we employ textual inversion techniques, where retrieved garment images are projected into the textual embedding space of the Stable Diffusion text encoder, allowing seamless integration of retrieved elements into the generative process. Experimental results on the Dress Code dataset demonstrate that Fashion-RAG outperforms existing methods both qualitatively and quantitatively, effectively capturing fine-grained visual details from retrieved garments. To the best of our knowledge, this is the first work to introduce a retrieval-augmented generation approach specifically tailored for multimodal fashion image editing.
Fulvio Sanguigni, Davide Morelli, Marcella Cornia, Rita Cucchiara
IJCNN2
2025 Parents and Children: Distinguishing Multimodal Deepfakes from Natural Images
abstract
Recent advancements in diffusion models have enabled the generation of realistic deepfakes from textual prompts in natural language. While these models have numerous benefits across various sectors, they have also raised concerns about the potential misuse of fake images and cast new pressures on fake image detection. In this work, we pioneer a systematic study on deepfake detection generated by state-of-the-art diffusion models. Firstly, we conduct a comprehensive analysis of the performance of contrastive and classification-based visual features, respectively, extracted from CLIP-based models and ResNet or Vision Transformer (ViT)-based architectures trained on image classification datasets. Our results demonstrate that fake images share common low-level cues, which render them easily recognizable. Further, we devise a multimodal setting wherein fake images are synthesized by different textual captions, which are used as seeds for a generator. Under this setting, we quantify the performance of fake detection strategies and introduce a contrastive-based disentangling method that lets us analyze the role of the semantics of textual descriptions and low-level perceptual cues. Finally, we release a new dataset, called COCOFake, containing about 1.2 million images generated from the original COCO image–caption pairs using two recent text-to-image diffusion models, namely Stable Diffusion v1.4 and v2.0.
Roberto Amoroso, Davide Morelli, Marcella Cornia, Lorenzo Baraldi 0001, Alberto Del Bimbo, Rita Cucchiara
ACM Trans. Multim. Comput. Commun. Appl.2
2024 HEalthRecordBERT (HERBERT): Leveraging Transformers on Electronic Health Records for Chronic Kidney Disease Risk Stratification
abstract
Risk stratification is an essential tool in the fight against many diseases, including chronic kidney disease. Recent work has focused on applying techniques from machine learning and leveraging the information contained in a patient’s electronic health record (EHR). Irregular intervals between data entries and the large number of variables tracked in EHR datasets can make them challenging to work with. Many of the difficulties associated with these datasets can be overcome by using large language models, such as bidirectional encoder representations from transformers (BERT). Previous attempts to apply BERT to EHR for risk stratification have shown promise. In this work we propose HERBERT, a novel application of BERT to EHR data. We identify two key areas where BERT models must be modified to adapt them to EHR data, namely: the embedding layer and the pretraining task. We show how changes to these can lead to improved performance, relative to the previous state of the art. We evaluate our model by predicting the transition of chronic kidney disease patients to end stage renal disease. The strong performance of our model justifies our architectural changes and suggests that large language models could play an important role in future renal risk stratification.
Alex Moore, Bastien Orset, Arrash Yassaee, Benjamin Irving, Davide Morelli
ACM Trans. Comput. Heal.5
2024 conDENSE: Conditional Density Estimation for Time Series Anomaly Detection
abstract
In recent years deep learning methods, based on reconstruction errors, have facilitated huge improvements in unsupervised anomaly detection. These methods make the limiting assumption that the greater the distance between an observation and a prediction the lower the likelihood of that observation. In this paper we propose conDENSE, a novel anomaly detection algorithm, which does not use reconstruction errors but rather uses conditional density estimation in masked autoregressive flows. By directly estimating the likelihood of data, our model moves beyond approximating expected behaviour with a single point estimate, as is the case in reconstruction error models. We show how conditioning on a dense representation of the current trajectory, extracted from a variational autoencoder with a gated recurrent unit (GRU VAE), produces a model that is suitable for periodic datasets, while also improving performance on non-periodic datasets. Experiments on 31 time-series, including real-world anomaly detection benchmark datasets and synthetically generated data, show that the model can outperform state-of-the-art deep learning methods.
Alex Moore, Davide Morelli
J. Artif. Intell. Res.2
2023 Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing
abstract
Fashion illustration is used by designers to communicate their vision and to bring the design idea from conceptualization to realization, showing how clothes interact with the human body. In this context, computer vision can thus be used to improve the fashion design process. Differently from previous works that mainly focused on the virtual try-on of garments, we propose the task of multimodal-conditioned fashion image editing, guiding the generation of human-centric fashion images by following multimodal prompts, such as text, human body poses, and garment sketches. We tackle this problem by proposing a new architecture based on latent diffusion models, an approach that has not been used before in the fashion domain. Given the lack of existing datasets suitable for the task, we also extend two existing fashion datasets, namely Dress Code and VITON-HD, with multimodal annotations collected in a semi-automatic manner. Experimental results on these new datasets demonstrate the effectiveness of our proposal, both in terms of realism and coherence with the given multimodal inputs. Source code and collected multimodal annotations are publicly available at: https://github.com/aimagelab/multimodal-garment-designer.
Alberto Baldrati, Davide Morelli, Giuseppe Cartella, Marcella Cornia, Marco Bertini 0001, Rita Cucchiara
ICCV2
2023 LaDI-VTON: Latent Diffusion Textual-Inversion Enhanced Virtual Try-On
abstract
The rapidly evolving fields of e-commerce and metaverse continue to seek innovative approaches to enhance the consumer experience. At the same time, recent advancements in the development of diffusion models have enabled generative networks to create remarkably realistic images. In this context, image-based virtual try-on, which consists in generating a novel image of a target model wearing a given in-shop garment, has yet to capitalize on the potential of these powerful generative solutions. This work introduces LaDI-VTON, the first Latent Diffusion textual Inversion-enhanced model for the Virtual Try-ON task. The proposed architecture relies on a latent diffusion model extended with a novel additional autoencoder module that exploits learnable skip connections to enhance the generation process preserving the model's characteristics. To effectively maintain the texture and details of the in-shop garment, we propose a textual inversion component that can map the visual features of the garment to the CLIP token embedding space and thus generate a set of pseudo-word token embeddings capable of conditioning the generation process. Experimental results on Dress Code and VITON-HD datasets demonstrate that our approach outperforms the competitors by a consistent margin, achieving a significant milestone for the task. Source code and trained models are publicly available at: https://github.com/miccunifi/ladi-vton.
Davide Morelli, Alberto Baldrati, Giuseppe Cartella, Marcella Cornia, Marco Bertini 0001, Rita Cucchiara
ACM Multimedia1
2023 Modeling Mood Polarity and Declaration Occurrence by Neural Temporal Point Processes
abstract
Neural point processes provide the flexibility needed to deal with time series of heterogeneous nature within the robust framework of point processes. This aspect is of particular relevance when dealing with real-world data, mixing generative processes characterized by radically different distributions and sampling. This brief discusses a neural point process approach for health and behavioral data, comprising both sparse events coming from user subjective declarations as well as fast-flowing time series from wearable sensors. We propose and empirically validate different neural architectures and we assess the effect of including input sources of different nature. The empirical analysis is built on the top of a challenging original dataset, never published before, and collected as part of a real-world experiment in an uncontrolled setting. Results show the potential of neural point processes both in terms of predicting the next event type as well as in predicting the time to next user interaction.
Davide Bacciu, Davide Morelli, Vlad Pandelea
IEEE Trans. Neural Networks Learn. Syst.2
2022 Defect detection by a deep learning approach with active IR thermography
abstract
N owadays, non-destructive techniques (NDT) playa fundamental role in the production industry since early defects detection (EDD) can reduce possible costs and avoid catastrophic failures. Under these aspects, all methods for fast and reliable inspection deserve special attention. This paper proposes a method to detect manufacturing defects or other damage mechanisms without compromising the original condition of the material using active IR thermography and automatic semantic segmentation. The segmentation of defects in composite materials is achieved by using a deep learning algorithm on a high-variance dataset obtained performing lock-in thermography under five different heat source configurations. Experimental results on specimens with known defects have demonstrated that the proposed methodology provides satisfying performances in automatic defect detection.
Giovanna Guaragnella, Davide Morelli, Tiziana D'Orazio, Umberto Galietti, Bartolomeo Trentadue, Roberto Marani
CoDIT2
2022 Dress Code: High-Resolution Multi-category Virtual Try-On
Davide Morelli, Matteo Fincato, Marcella Cornia, Federico Landi, Fabio Cesari, Rita Cucchiara
ECCV (8)1
2022 Topographic mapping for quality inspection and intelligent filtering of smart-bracelet data
Davide Bacciu, Gioele Bertoncini, Davide Morelli
Neural Comput. Appl.3
2020 Design of Modern Supply Chain Networks Using Fuzzy Bargaining Game and Data Envelopment Analysis
abstract
This article proposes a novel methodology for multistage, multiproduct, multi-item, and closed-loop Supply Chain Network (SCN) design under uncertainty. The method considers that multiple products are manufactured by the SCN, each composed by multiple items, and that some of the sold products may require repair, refurbishing, or remanufacturing activities. We solve the two main decisions that take place in the medium-/short-term planning horizon, namely partners’ selection and allocation of the received orders among them. The partners’ selection problem is solved by a cross-efficiency fuzzy Data Envelopment Analysis technique, which allows evaluating the efficiency of each SCN member and ranking them against multiple conflicting objectives under uncertain data on their performance. Then, according to the estimated customers’ demand, the order allocation problem is solved by a fuzzy bargaining game problem, where each SCN actor behaves to simultaneously maximize both its own profit and the service level of the overall SCN in terms of efficiency, costs, and lead time. An illustrative example from the literature is finally presented.Note to Practitioners—We present a decision tool to address the optimal design, performance evaluation, and continuous improvement of modern cooperative SCNs. We propose an effective method to jointly solve the members’ selection and the orders’ allocation, considering the complex structure of modern SCNs, the multiobjective nature of the problems, and the uncertainty characterizing economic markets. Competition within SCNs stages and cooperation along the chain are considered, with the aim to improve both financial and environmental sustainability, while ensuring the highest service levels to customers.
Graziana Cavone, Mariagrazia Dotoli, Nicola Epicoco, Davide Morelli, Carla Seatzu
IEEE Trans Autom. Sci. Eng.4
2018 Randomized neural networks for preference learning with physiological data
Davide Bacciu, Michele Colombo, Davide Morelli, David Plans
Neurocomputing3
2017 ELM Preference Learning for Physiological Data
Davide Bacciu, Michele Colombo, Davide Morelli, David Plans
ESANN3
2017 DropIn: Making reservoir computing neural networks robust to missing inputs by dropout
abstract
The paper presents a novel, principled approach to train recurrent neural networks from the Reservoir Computing family that are robust to missing part of the input features at prediction time. By building on the ensembling properties of Dropout regularization, we propose a methodology, named DropIn, which efficiently trains a neural model as a committee machine of subnetworks, each capable of predicting with a subset of the original input features. We discuss the application of the DropIn methodology in the context of Reservoir Computing models and targeting applications characterized by input sources that are unreliable or prone to be disconnected, such as in pervasive wireless sensor networks and ambient intelligence. We provide an experimental assessment using real-world data from such application domains, showing how the Dropin methodology allows to maintain predictive performances comparable to those of a model without missing features, even when 20%–50% of the inputs are not available.
Davide Bacciu, Francesco Crecchi, Davide Morelli
IJCNN3
2016 A high-level and accurate energy model of parallel and concurrent workloads
abstract
Summary The ability to predict the energy needed by a system to perform a task, or several concurrent parallel tasks, allows the scheduler to enforce energy‐aware policies while providing acceptable performance. The approaches in literature to model energy consumption of tasks usually focus on low‐level descriptors and require invasive instrumentation of the computational environment. We developed an energy model and a methodology to automatically extract features that characterize the computational environment relying only on a single power meter that measures the energy consumption of the whole system. Once the model has been built, the energy consumption of concurrent tasks can be calculated, with a statistically insignificant error, even without any power meter. We show that our model can predict with high accuracy, even only using the utilization time of the cores in a high‐performance computing enclosure, without using performance counters. Hence, the model could be easily applicable to heterogeneous systems, where collecting representative performance counters can be problematic. Copyright © 2015 John Wiley & Sons, Ltd.
Davide Morelli, Andrea Canciani, Antonio Cisternino
Concurr. Comput. Pract. Exp.1
2012 Experience-Driven Procedural Music Generation for Games
abstract
As video games have grown from crude and simple circuit-based artefacts to a multibillion dollar worldwide industry, video-game music has become increasingly adaptive. Composers have had to use new techniques to avoid the traditional, event-based approach where music is composed mostly of looped audio tracks, which can lead to music that is too repetitive. In addition, these cannot scale well in the design of today's games, which have become increasingly complex and nonlinear in narrative. This paper outlines the use of experience-driven procedural music generation, to outline possible ways forward in the dynamic generation of music and audio according to user gameplay metrics.
David Plans, Davide Morelli
IEEE Trans. Comput. Intell. AI Games2
2006 Playing music: an installation based on Xenakis' musical games
abstract
Iannis Xenakis' works Duel and Strategie are two music games: sounds play the role of moves in a match where the players are the two orchestra conductors. They decide which part of the score is to be played in answer to the opposite conductor's choice, looking at a game matrix which contains the values of every couple of moves.Playing Music is an installation driven by software implementing the same logic as Duel.Each player makes his moves with simple physical actions, recognized by the software using a camera, and the score is projected on a screen so that the audience can easily understand the rules, after a few moves.
Marco Liuni, Davide Morelli
AVI2