Leonardo Rossi

dblp:49/10989 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Mcgm-styler: free-form styler for mask conditional text-to-image generative model
abstract
Abstract Generative models for text-to-image synthesis have made significant advancements in recent years, enabling the creation of highly detailed and stylistically diverse images. In this work, we introduce MCGM-Styler, an extension of our previous MCGM as reported (MCGM: Mask conditional text-to-image generative model, 2024) model, which generates images based on masked conditions to specify the action pose of subjects in a source image. Our key contribution is the addition of a new training step that enables the model to also perform style transfer, allowing it to generate images that not only force the pose by mask condition, but also adhere to both single and multiple target artistic styles. Unlike traditional approaches that require large datasets, MCGM-Styler is trained on a single image, making it highly efficient and adaptable. The model can handle scenes with one or multiple subjects, generating coherent and stylistically consistent outputs so the user can generate any subject(s) in any pose and style or generate an image with a mix of different styles. We evaluate our approach against existing works, particularly the DreamStyler model (Ahn et al. in Proc. AAAI Conf. Artif. Intell. 38:674–681), a state-of-the-art method for style transfer. Our results demonstrate that the MCGM-Styler achieves superior performance in preserving not only the pose of the concept, but also the style fidelity, highlighting its effectiveness in controllable image generation.
Rami Skaik, Leonardo Rossi, Tomaso Fontanini, Andrea Prati 0001
Vis. Comput.2
2025 Beyond Training: A Personalized Holistic Injury Prediction in Triathletes
abstract
Triathlon training combines swimming, cycling, and running, often at high volumes, to prepare athletes for long-distance events. The highly intense physical demand puts athletes at significant risk of overuse injuries. While wearable devices provide continuous, high-frequency insights into an athlete's physiological response to training, extracting meaningful, actionable patterns remains a challenge - especially for everyday users. Understanding these metrics and their relationship with injury risk is critical to optimizing training strategies and preventing injuries before they occur. This work proposes a learning model to identify patterns that indicate an increased risk of injury, allowing proactive adjustments to training loads. However, building a generalizable model in sports science and healthcare presents a key challenge: the need for large, high-quality labelled datasets, which are often limited by privacy concerns. To address this limitation, this work also explores the generation and application of a highly realistic synthetic dataset that ensures robust model training while mitigating privacy constraints.
Leonardo Rossi
SMARTCOMP1
2025 Mamba-ST: State Space Model for Efficient Style Transfer
abstract
The goal of style transfer is, given a content image and a style source, generating a new image preserving the content but with the artistic representation of the style source. Most of the state-of-the-art architectures use transformers or diffusion-based models to perform this task, despite the heavy computational burden that they require. In particular, transformers use self- and cross-attention layers which have large memory footprint, while diffusion models require high inference time. To overcome the above, this paper explores a novel design of Mamba, an emergent State-Space Model (SSM), called Mamba-ST, to perform style transfer. To do so, we adapt Mamba linear equation to simulate the behavior of cross-attention layers, which are able to combine two separate embeddings into a single output, but drastically reducing memory usage and time complexity. We modified the Mamba's inner equations so to accept inputs from, and combine, two separate data streams. To the best of our knowledge, this is the first attempt to adapt the equations of SSMs to a vision task like style transfer without requiring any other module like cross-attention or custom normalization layers. An extensive set of experiments demonstrates the superiority and efficiency of our method in performing style transfer compared to transformers and diffusion models. Results show improved quality in terms of both ArtFID and FID metrics. Code is available at https://github.com/FilippoBotti/MambaST.
Filippo Botti, Alex Ergasti, Leonardo Rossi, Tomaso Fontanini, Claudio Ferrari, Massimo Bertozzi, Andrea Prati 0001
WACV3
2025 Swin2-MoSE: A new single image supersolution model for remote sensing
abstract
Abstract Due to the limitations of current optical and sensor technologies and the high cost of updating them, the spectral and spatial resolution of satellites may not always meet desired requirements. For these reasons, Remote‐Sensing Single‐Image Super‐Resolution (RS‐SISR) techniques have gained significant interest. In this paper, Swin2‐MoSE model is proposed, an enhanced version of Swin2SR. The model introduces MoE‐SM, an enhanced Mixture‐of‐Experts (MoE) to replace the Feed‐Forward inside all Transformer block. MoE‐SM is designed with Smart‐Merger, and new layer for merging the output of individual experts, and with a new way to split the work between experts, defining a new per‐example strategy instead of the commonly used per‐token one. Furthermore, it is analyzed how positional encodings interact with each other, demonstrating that per‐channel bias and per‐head bias can positively cooperate. Finally, the authors propose to use a combination of Normalized‐Cross‐Correlation (NCC) and Structural Similarity Index Measure (SSIM) losses, to avoid typical MSE loss limitations. Experimental results demonstrate that Swin2‐MoSE outperforms any Swin derived models by up to 0.377–0.958 dB (PSNR) on task of , and resolution‐upscaling ( and OLI2MSI datasets). It also outperforms SOTA models by a good margin, proving to be competitive and with excellent potential, especially for complex tasks. Additionally, an analysis of computational costs is also performed. Finally, the efficacy of Swin2‐MoSE is shown, applying it to a semantic segmentation task (SeasoNet dataset). Code and pretrained are available on https://github.com/IMPLabUniPr/swin2‐mose/tree/official_code
Leonardo Rossi, Vittorio Bernuzzi, Tomaso Fontanini, Massimo Bertozzi, Andrea Prati 0001
IET Image Process.1
2024 Memory-Augmented Online Video Anomaly Detection
abstract
The ability to understand the surrounding scene is of paramount importance for Autonomous Vehicles (AVs). This paper presents a system capable to work in an online fashion, giving an immediate response to the arise of anomalies surrounding the AV, exploiting only the videos captured by a dash-mounted camera. Our architecture, called MOVAD, relies on two main modules: a Short-Term Memory Module to extract information related to the ongoing action, implemented by a Video Swin Transformer (VST), and a Long-Term Memory Module injected inside the classifier that considers also remote past information and action context thanks to the use of a Long-Short Term Memory (LSTM) network. The strengths of MOVAD are not only linked to its excellent performance, but also to its straightforward and modular architecture, trained in a end-to-end fashion with only RGB frames with as less assumptions as possible, which makes it easy to implement and play with. We evaluated the performance of our method on Detection of Traffic Anomaly (DoTA) dataset, a challenging collection of dash-mounted camera videos of accidents. After an extensive ablation study, MOVAD is able to reach an AUC score of 82.17%, surpassing the current state-of-the-art by +2.87 AUC. Our code and pretrained are available online on https://github.com/IMPLabUniPr/movad/tree/movad_vad
Leonardo Rossi, Vittorio Bernuzzi, Tomaso Fontanini, Massimo Bertozzi, Andrea Prati 0001
ICASSP1
2024 CFTS-GAN: Continual Few-Shot Teacher Student for Generative Adversarial Networks
Munsif Ali, Leonardo Rossi, Massimo Bertozzi
ICPR (25)2
2022 Self-Balanced R-CNN for instance segmentation
Leonardo Rossi, Akbar Karimi 0001, Andrea Prati 0001
J. Vis. Commun. Image Represent.1
2021 Recursively Refined R-CNN: Instance Segmentation with Self-RoI Rebalancing
Leonardo Rossi, Akbar Karimi 0001, Andrea Prati 0001
CAIP (1)1
2020 Adversarial Training for Aspect-Based Sentiment Analysis with BERT
abstract
Aspect-Based Sentiment Analysis (ABSA) studies the extraction of sentiments and their targets. Collecting labeled data for this task in order to help neural networks generalize better can be laborious and time-consuming. As an alternative, similar data to the real-world examples can be produced artificially through an adversarial process which is carried out in the embedding space. Although these examples are not real sentences, they have been shown to act as a regularization method which can make neural networks more robust. In this work, we fine- tune the general purpose BERT and domain specific post-trained BERT (BERT-PT) using adversarial training. After improving the results of post-trained BERT with different hyperparameters, we propose a novel architecture called BERT Adversarial Training (BAT) to utilize adversarial training for the two major tasks of Aspect Extraction and Aspect Sentiment Classification in sentiment analysis. The proposed model outperforms the general BERT as well as the in-domain post-trained BERT in both tasks. To the best of our knowledge, this is the first study on the application of adversarial training in ABSA. The code is publicly available on a GitHub repository at https://github.com/IMPLabUniPr/Adversarial-Training-for-ABSA.
Akbar Karimi 0001, Leonardo Rossi, Andrea Prati 0001
ICPR2
2020 A Novel Region of Interest Extraction Layer for Instance Segmentation
abstract
Given the wide diffusion of deep neural network architectures for computer vision tasks, several new applications are nowadays more and more feasible. Among them, a particular attention has been recently given to instance segmentation, by exploiting the results achievable by two-stage networks (such as Mask R-CNN or Faster R-CNN), derived from R-CNN. In these complex architectures, a crucial role is played by the Region of Interest (RoI) extraction layer, devoted to extracting a coherent subset of features from a single Feature Pyramid Network (FPN) layer attached on top of a backbone. This paper is motivated by the need to overcome the limitations of existing RoI extractors which select only one (the best) layer from FPN. Our intuition is that all the layers of FPN retain useful information. Therefore, the proposed layer (called Generic RoI Extractor - GRoIE) introduces non-local building blocks and attention mechanisms to boost the performance. A comprehensive ablation study at component level is conducted to find the best set of algorithms and parameters for the GRoIE layer. Moreover, GRoIE can be integrated seamlessly with every two-stage architecture for both object detection and instance segmentation tasks. Therefore, the improvements brought about by the use of GRoIE in different state-of-the-art architectures are also evaluated. The proposed layer leads up to gain a 1.1% AP improvement on bounding box detection and 1.7% AP improvement on instance segmentation. The code is publicly available on GitHub repository at https://github.com/IMPLabUniPr/mmdetection/tree/groie_dev.
Leonardo Rossi, Akbar Karimi 0001, Andrea Prati 0001
ICPR1
2012 Making turing machines accessible to blind students
abstract
In this paper we describe how we tried to make the well-known JFLAP Turing machine simulator accessible to blind students taking a theoretical computer science course. Software accessibility is an important topic for both legal and ethical reasons: in our case, however, we also wanted to make the accessible software usable by blind students in cooperation with the other students, in order to encourage the integration of the blind students within the rest of the class. For this reason, the accessible version of the JFLAP Turing machine simulator that we developed is as much similar as possible to and fully compatible with the original one. In the paper, we also report some very satisfactory preliminary validation results that indicate how the new software can really make Turing machines accessible to blind students.
Pierluigi Crescenzi, Leonardo Rossi, Gianluca Apollaro
SIGCSE2