VLDB 2026 Research / reviewers in the wild / expert
Tomaso Fontanini
dblp:194/5165
· DBLP profile ↗
9ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0001-6595-4874ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mcgm-styler: free-form styler for mask conditional text-to-image generative modelabstractAbstract Generative models for text-to-image synthesis have made significant advancements in recent years, enabling the creation of highly detailed and stylistically diverse images. In this work, we introduce MCGM-Styler, an extension of our previous MCGM as reported (MCGM: Mask conditional text-to-image generative model, 2024) model, which generates images based on masked conditions to specify the action pose of subjects in a source image. Our key contribution is the addition of a new training step that enables the model to also perform style transfer, allowing it to generate images that not only force the pose by mask condition, but also adhere to both single and multiple target artistic styles. Unlike traditional approaches that require large datasets, MCGM-Styler is trained on a single image, making it highly efficient and adaptable. The model can handle scenes with one or multiple subjects, generating coherent and stylistically consistent outputs so the user can generate any subject(s) in any pose and style or generate an image with a mix of different styles. We evaluate our approach against existing works, particularly the DreamStyler model (Ahn et al. in Proc. AAAI Conf. Artif. Intell. 38:674–681), a state-of-the-art method for style transfer. Our results demonstrate that the MCGM-Styler achieves superior performance in preserving not only the pose of the concept, but also the style fidelity, highlighting its effectiveness in controllable image generation. Rami Skaik, Leonardo Rossi, Tomaso Fontanini, Andrea Prati 0001 |
Vis. Comput. | 3 |
| 2025 | Mamba-ST: State Space Model for Efficient Style TransferabstractThe goal of style transfer is, given a content image and a style source, generating a new image preserving the content but with the artistic representation of the style source. Most of the state-of-the-art architectures use transformers or diffusion-based models to perform this task, despite the heavy computational burden that they require. In particular, transformers use self- and cross-attention layers which have large memory footprint, while diffusion models require high inference time. To overcome the above, this paper explores a novel design of Mamba, an emergent State-Space Model (SSM), called Mamba-ST, to perform style transfer. To do so, we adapt Mamba linear equation to simulate the behavior of cross-attention layers, which are able to combine two separate embeddings into a single output, but drastically reducing memory usage and time complexity. We modified the Mamba's inner equations so to accept inputs from, and combine, two separate data streams. To the best of our knowledge, this is the first attempt to adapt the equations of SSMs to a vision task like style transfer without requiring any other module like cross-attention or custom normalization layers. An extensive set of experiments demonstrates the superiority and efficiency of our method in performing style transfer compared to transformers and diffusion models. Results show improved quality in terms of both ArtFID and FID metrics. Code is available at https://github.com/FilippoBotti/MambaST. Filippo Botti, Alex Ergasti, Leonardo Rossi, Tomaso Fontanini, Claudio Ferrari, Massimo Bertozzi, Andrea Prati 0001 |
WACV | 4 |
| 2025 | Swin2-MoSE: A new single image supersolution model for remote sensingabstractAbstract Due to the limitations of current optical and sensor technologies and the high cost of updating them, the spectral and spatial resolution of satellites may not always meet desired requirements. For these reasons, Remote‐Sensing Single‐Image Super‐Resolution (RS‐SISR) techniques have gained significant interest. In this paper, Swin2‐MoSE model is proposed, an enhanced version of Swin2SR. The model introduces MoE‐SM, an enhanced Mixture‐of‐Experts (MoE) to replace the Feed‐Forward inside all Transformer block. MoE‐SM is designed with Smart‐Merger, and new layer for merging the output of individual experts, and with a new way to split the work between experts, defining a new per‐example strategy instead of the commonly used per‐token one. Furthermore, it is analyzed how positional encodings interact with each other, demonstrating that per‐channel bias and per‐head bias can positively cooperate. Finally, the authors propose to use a combination of Normalized‐Cross‐Correlation (NCC) and Structural Similarity Index Measure (SSIM) losses, to avoid typical MSE loss limitations. Experimental results demonstrate that Swin2‐MoSE outperforms any Swin derived models by up to 0.377–0.958 dB (PSNR) on task of , and resolution‐upscaling ( and OLI2MSI datasets). It also outperforms SOTA models by a good margin, proving to be competitive and with excellent potential, especially for complex tasks. Additionally, an analysis of computational costs is also performed. Finally, the efficacy of Swin2‐MoSE is shown, applying it to a semantic segmentation task (SeasoNet dataset). Code and pretrained are available on https://github.com/IMPLabUniPr/swin2‐mose/tree/official_code Leonardo Rossi, Vittorio Bernuzzi, Tomaso Fontanini, Massimo Bertozzi, Andrea Prati 0001 |
IET Image Process. | 3 |
| 2025 | MARS: Paying More Attention to Visual Attributes for Text-Based Person SearchabstractText-Based Person Search (TBPS) is a problem that gained significant interest within the research community. The task is that of retrieving one or more images of a specific individual based on a textual description. The multi-modal nature of the task requires learning representations that bridge text and image data within a shared latent space. Existing TBPS systems face two major challenges. One is defined as inter-identity noise that is due to the inherent vagueness and imprecision of text descriptions, and it indicates how descriptions of visual attributes can be generally associated to different people; the other is the intra-identity variations, which are all those nuisances, e.g., pose, illumination, that can alter the visual appearance of the same textual attributes for a given subject. To address these issues, this article presents a novel TBPS architecture named Mae-Attribute-Relation-Sensitive (MARS), which enhances current state-of-the-art models by introducing two key components: a Visual Reconstruction Loss and an Attribute Loss. The former employs a Masked AutoEncoder trained to reconstruct randomly masked image patches with the aid of the textual description. In doing so the model is encouraged to learn more expressive representations and textual–visual relations in the latent space. The attribute loss, instead, balances the contribution of different types of attributes, defined as adjective–noun chunks of text. This loss ensures that every attribute is taken into consideration in the person retrieval process. Extensive experiments on three commonly used datasets, namely CUHK-PEDES, ICFG-PEDES, and RSTPReid, report performance improvements, with significant gains in the Mean Average Precision (mAP) metric w.r.t. the current state of the art. Code will be available at https://github.com/ErgastiAlex/MARS . Alex Ergasti, Tomaso Fontanini, Claudio Ferrari, Massimo Bertozzi, Andrea Prati 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Memory-Augmented Online Video Anomaly DetectionabstractThe ability to understand the surrounding scene is of paramount importance for Autonomous Vehicles (AVs). This paper presents a system capable to work in an online fashion, giving an immediate response to the arise of anomalies surrounding the AV, exploiting only the videos captured by a dash-mounted camera. Our architecture, called MOVAD, relies on two main modules: a Short-Term Memory Module to extract information related to the ongoing action, implemented by a Video Swin Transformer (VST), and a Long-Term Memory Module injected inside the classifier that considers also remote past information and action context thanks to the use of a Long-Short Term Memory (LSTM) network. The strengths of MOVAD are not only linked to its excellent performance, but also to its straightforward and modular architecture, trained in a end-to-end fashion with only RGB frames with as less assumptions as possible, which makes it easy to implement and play with. We evaluated the performance of our method on Detection of Traffic Anomaly (DoTA) dataset, a challenging collection of dash-mounted camera videos of accidents. After an extensive ablation study, MOVAD is able to reach an AUC score of 82.17%, surpassing the current state-of-the-art by +2.87 AUC. Our code and pretrained are available online on https://github.com/IMPLabUniPr/movad/tree/movad_vad Leonardo Rossi, Vittorio Bernuzzi, Tomaso Fontanini, Massimo Bertozzi, Andrea Prati 0001 |
ICASSP | 3 |
| 2023 | FrankenMask: Manipulating semantic masks with transformers for face parts editingabstractIn this paper, we propose FrankenMask, a novel framework that allows swapping and rearranging face parts in semantic masks for automatic editing of shape-related facial attributes. This is a novel yet challenging task as substituting face parts in a semantic mask requires to account for possible spatial misalignment and the adaptation of surrounding regions. We obtain such a feature by combining a Transformer encoder to learn the spatial relationships of facial parts, with an encoder–decoder architecture, which reconstructs a complete mask from the composition of local parts. Reconstruction and attribute classification results demonstrate the effective synthesis of facial images, while showing the generation of accurate and plausible facial attributes. Code is available at https://github.com/TFonta/FrankenMask_semantic. Tomaso Fontanini, Claudio Ferrari, Giuseppe Lisanti, Leonardo Galteri, Stefano Berretti, Massimo Bertozzi, Andrea Prati 0001 |
Pattern Recognit. Lett. | 1 |
| 2023 | Unsupervised Discovery and Manipulation of Continuous Disentangled Factors of VariationabstractLearning a disentangled representation of a distribution in a completely unsupervised way is a challenging task that has drawn attention recently. In particular, much focus has been put in separating factors of variation (i.e., attributes) within the latent code of a Generative Adversarial Network (GAN). Achieving that permits control of the presence or absence of those factors in the generated samples by simply editing a small portion of the latent code. Nevertheless, existing methods that perform very well in a noise-to-image setting often fail when dealing with a real data distribution, i.e., when the discovered attributes need to be applied to real images. However, some methods are able to extract and apply a style to a sample but struggle to maintain its content and identity, while others are not able to locally apply attributes and end up achieving only a global manipulation of the original image. In this article, we propose a completely (i.e., truly ) unsupervised method that is able to extract a disentangled set of attributes from a data distribution and apply them to new samples from the same distribution by preserving their content. This is achieved by using an image-to-image GAN that maps an image and a random set of continuous attributes to a new image that includes those attributes. Indeed, these attributes are initially unknown and they are discovered during training by maximizing the mutual information between the generated samples and the attributes’ vector. Finally, the obtained disentangled set of continuous attributes can be used to freely manipulate the input samples. We prove the effectiveness of our method over a series of datasets and show its application on various tasks, such as attribute editing, data augmentation, and style transfer. Tomaso Fontanini, Luca Donati, Massimo Bertozzi, Andrea Prati 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Bag of indexes: a multi-index scheme for efficient approximate nearest neighbor search
Federico Magliani, Tomaso Fontanini, Andrea Prati 0001 |
Multim. Tools Appl. | 2 |
| 2020 | MetalGAN: Multi-domain label-less image synthesis using cGANs and meta-learning
Tomaso Fontanini, Eleonora Iotti, Luca Donati, Andrea Prati 0001 |
Neural Networks | 1 |