EDBT 2026 Demo / reviewers in the wild / expert
Leonardo Galteri
dblp:196/0941
· DBLP profile ↗
25ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0002-7247-9407ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 9 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating Diffusion Models with One-Step Distillation for Image and Video Super-Resolution
Alessio Bugetti, Leonardo Galteri, Marco Bertini 0001 |
MMSys | 2 |
| 2025 | Implicit and Explicit Attitudes Towards Virtual Agents in Education: An Experimental Study
Chiara Scuotto, Masiar Babazadeh, Leonardo Galteri, Pierpaolo Limone, Stefano Triberti |
CHIRA (2) | 3 |
| 2025 | Prompt-Engineered Detection of AI-Generated ImagesabstractThe proliferation of highly realistic AI-generated images presents significant challenges related to authenticity and misinformation. Although multimodal large language models (LLMs) possess advanced visual understanding capabilities, their effectiveness in distinguishing synthetic images from real ones requires systematic evaluation. This paper investigates the ability of four prominent LLMs—GPT-4V, Gemini, LLaVA, and Claude to detect AI-generated images. Using a case study methodology, twelve diverse synthetic images were presented to the models using ten distinct prompts, ranging from generic classification requests to detailed, forensically-guided queries. The study analyzes the accuracy and response patterns of each LLM, with a particular focus on the impact of prompt specificity and iterative refinement (prompt engineering) on detection performance. The results indicate that while GPT-4V demonstrated superior consistency, the performance of all tested LLMs, particularly Gemini, Claude, and LLaVA, was significantly influenced by prompt quality. Specific, detailed prompts markedly improved detection accuracy compared to generic ones. The findings underscore that effective prompt engineering is crucial for leveraging LLMs as tools for synthetic image detection and highlight the need for skilled human interaction to guide these systems to achieve reliable results. This research contributes to understanding the potential and current limitations of LLMs in addressing the challenges posed by synthetic media. Rita Ferrara, Leonardo Galteri |
SoMeT | 2 |
| 2024 | ARNIQA: Learning Distortion Manifold for Image Quality AssessmentabstractNo-Reference Image Quality Assessment (NR-IQA) aims to develop methods to measure image quality in alignment with human perception without the need for a high-quality reference image. In this work, we propose a self-supervised approach named ARNIQA (leArning distoRtion maNifold for Image Quality Assessment) for modeling the image distortion manifold to obtain quality representations in an intrinsic manner. First, we introduce an image degradation model that randomly composes ordered sequences of consecutively applied distortions. In this way, we can synthetically degrade images with a large variety of degradation patterns. Second, we propose to train our model by maximizing the similarity between the representations of patches of different images distorted equally, despite varying content. Therefore, images degraded in the same manner correspond to neighboring positions within the distortion manifold. Finally, we map the image representations to the quality scores with a simple linear regressor, thus without fine-tuning the encoder weights. The experiments show that our approach achieves state-of-the-art performance on several datasets. In addition, ARNIQA demonstrates improved data efficiency, generalization capabilities, and robustness compared to competing methods. The code and the model are publicly available at https://github.com/miccunifi/ARNIQA. Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini 0001, Alberto Del Bimbo |
WACV | 2 |
| 2024 | Reference-based Restoration of Digitized Analog VideotapesabstractAnalog magnetic tapes have been the main video data storage device for several decades. Videos stored on analog videotapes exhibit unique degradation patterns caused by tape aging and reader device malfunctioning that are different from those observed in film and digital video restoration tasks. In this work, we present a reference-based approach for the resToration of digitized Analog videotaPEs (TAPE). We leverage CLIP for zero-shot artifact detection to identify the cleanest frames of each video through textual prompts describing different artifacts. Then, we select the clean frames most similar to the input ones and employ them as references. We design a transformer-based Swin-UNet network that exploits both neighboring and reference frames via our Multi-Reference Spatial Feature Fusion (MRSFF) blocks. MRSFF blocks rely on cross-attention and attention pooling to take advantage of the most useful parts of each reference frame. To address the absence of ground truth in real-world videos, we create a synthetic dataset of videos exhibiting artifacts that closely resemble those commonly found in analog videotapes. Both quantitative and qualitative experiments show the effectiveness of our approach compared to other state-of-the-art methods. The code, the model, and the synthetic dataset are publicly available at https://github.com/miccunifi/TAPE. Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini 0001, Alberto Del Bimbo |
WACV | 2 |
| 2024 | Perceptual Quality Improvement in Videoconferencing Using Keyframes-Based GANabstractIn the latest years, videoconferencing has taken a fundamental role in interpersonal relations, both for personal and business purposes. Lossy video compression algorithms are the enabling technology for videoconferencing, as they reduce the bandwidth required for real-time video streaming. However, lossy video compression decreases the perceived visual quality. Thus, many techniques for reducing compression artifacts and improving video visual quality have been proposed in recent years. In this work, we propose a novel GAN-based method for compression artifacts reduction in videoconferencing. Given that, in this context, the speaker is typically in front of the camera and remains the same for the entire duration of the transmission, we can maintain a set of reference keyframes of the person from the higher-quality I-frames that are transmitted within the video stream and exploit them to guide the visual quality improvement; a novel aspect of this approach is the update policy that maintains and updates a compact and effective set of reference keyframes. First, we extract multi-scale features from the compressed and reference frames. Then, our architecture combines these features in a progressive manner according to facial landmarks. This allows the restoration of the high-frequency details lost after the video compression. Experiments show that the proposed approach improves visual quality and generates photo-realistic results even with high compression rates. Code and pre-trained networks are publicly available at https://github.com/LorenzoAgnolucci/Keyframes-GAN. Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini 0001, Alberto Del Bimbo |
IEEE Trans. Multim. | 2 |
| 2023 | Optimization Techniques of Deep Learning Models for Visual Quality ImprovementabstractVideo restoration is a widely studied task in the field of computer vision and image processing. The primary objective of video restoration is to improve the visual quality of degraded videos caused by various factors, such as noise, blur, compression artifacts, and other distortions. In this study, the integration of post-training quantization techniques was investigated to optimize deep learning models for super-resolution inference. The results indicate that reducing the precision of weights and activations in these models substantially decreases the computational complexity and memory requirements without compromising performance, rendering them more practical and cost-effective for real-world applications, where real-time inference is often required. When TensorRT was integrated with PyTorch, the efficiency of the model was further improved taking advantage of the INT8 computational capabilities of recent NVIDIA GPUs. Lorenzo Palloni, Leonardo Galteri, Marco Bertini 0001 |
SoMeT | 2 |
| 2023 | FrankenMask: Manipulating semantic masks with transformers for face parts editingabstractIn this paper, we propose FrankenMask, a novel framework that allows swapping and rearranging face parts in semantic masks for automatic editing of shape-related facial attributes. This is a novel yet challenging task as substituting face parts in a semantic mask requires to account for possible spatial misalignment and the adaptation of surrounding regions. We obtain such a feature by combining a Transformer encoder to learn the spatial relationships of facial parts, with an encoder–decoder architecture, which reconstructs a complete mask from the composition of local parts. Reconstruction and attribute classification results demonstrate the effective synthesis of facial images, while showing the generation of accurate and plausible facial attributes. Code is available at https://github.com/TFonta/FrankenMask_semantic. Tomaso Fontanini, Claudio Ferrari, Giuseppe Lisanti, Leonardo Galteri, Stefano Berretti, Massimo Bertozzi, Andrea Prati 0001 |
Pattern Recognit. Lett. | 4 |
| 2023 | (Compress and Restore)N: A Robust Defense Against Adversarial Attacks on Image ClassificationabstractModern image classification approaches often rely on deep neural networks, which have shown pronounced weakness to adversarial examples: images corrupted with specifically designed yet imperceptible noise that causes the network to misclassify. In this article, we propose a conceptually simple yet robust solution to tackle adversarial attacks on image classification. Our defense works by first applying a JPEG compression with a random quality factor; compression artifacts are subsequently removed by means of a generative model Artifact Restoration GAN. The process can be iterated ensuring the image is not degraded and hence the classification not compromised. We train different AR-GANs for different compression factors, so that we can change its parameters dynamically at each iteration depending on the current compression, making the gradient approximation difficult. We experiment with our defense against three white-box and two black-box attacks, with a particular focus on the state-of-the-art BPDA attack. Our method does not require any adversarial training, and is independent of both the classifier and the attack. Experiments demonstrate that dynamically changing the AR-GAN parameters is of fundamental importance to obtain significant robustness. Claudio Ferrari, Federico Becattini, Leonardo Galteri, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Restoration of Analog Videos Using Swin-UNetabstractIn this paper we present a system to restore analog videos of historical archives. These videos often contain severe visual degradation due to the deterioration of their tape supports that require costly and slow manual interventions to recover the original content. The proposed system uses a multi-frame approach and is able to deal also with severe tape mistracking, which results in completely scrambled frames. Tests on real-world videos from a major historical video archive show the effectiveness of our demo system. Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini 0001, Alberto Del Bimbo |
ACM Multimedia | 2 |
| 2022 | LANBIQUE: LANguage-based Blind Image QUality EvaluationabstractImage quality assessment is often performed with deep networks that are fine-tuned to regress a human provided quality score of a given image. Usually, this approach may lack generalization capabilities and, while being highly precise on similar image distribution, it may yield lower correlation on unseen distortions. In particular, they show poor performances, whereas images corrupted by noise, blur, or compression have been restored by generative models. As a matter of fact, evaluation of these generative models is often performed providing anecdotal results to the reader. In the case of image enhancement and restoration, reference images are usually available. Nevertheless, using signal based metrics often leads to counterintuitive results: Highly natural crisp images may obtain worse scores than blurry ones. However, blind reference image assessment may rank images reconstructed with GANs higher than the original undistorted images. To avoid time-consuming human-based image assessment, semantic computer vision tasks may be exploited instead. In this article, we advocate the use of language generation tasks to evaluate the quality of restored images. We refer to our assessment approach as LANguage-based Blind Image QUality Evaluation (LANBIQUE). We show experimentally that image captioning, used as a downstream task, may serve as a method to score image quality, independently of the distortion process that affects the data. Captioning scores are better aligned with human rankings with respect to classic signal based or No-reference image quality metrics. We show insights on how the corruption, by artefacts, of local image structure may steer image captions in the wrong direction. Leonardo Galteri, Lorenzo Seidenari, Pietro Bongini, Marco Bertini 0001, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Language Based Image Quality AssessmentabstractEvaluation of generative models, in the visual domain, is often performed providing anecdotal results to the reader. In the case of image enhancement, reference images are usually available. Nonetheless, using signal based metrics often leads to counterintuitive results: highly natural crisp images may obtain worse scores than blurry ones. On the other hand, blind reference image assessment may rank images reconstructed with GANs higher than the original undistorted images. To avoid time consuming human based image assessment, semantic computer vision tasks may be exploited instead [9, 25, 33]. In this paper we advocate the use of language generation tasks to evaluate the quality of restored images. We show experimentally that image captioning, used as a downstream task, may serve as a method to score image quality. Captioning scores are better aligned with human rankings with respect to signal based metrics or no-reference image quality metrics. We show insights on how the corruption, by artifacts, of local image structure may steer image captions in the wrong direction. Lorenzo Seidenari, Leonardo Galteri, Pietro Bongini, Marco Bertini 0001, Alberto Del Bimbo |
MMAsia | 2 |
| 2021 | Optical Flow based CNN for detection of unlearnt deepfake manipulationsabstractA new phenomenon named Deepfakes constitutes a serious threat in video manipulation. AI-based technologies have provided easy-to-use methods to create extremely realistic videos. On the side of multimedia forensics, being able to individuate this kind of fake contents becomes ever more crucial. In this work, a new forensic technique able to detect fake and original video sequences is proposed; it is based on the use of CNNs trained to distinguish possible motion dissimilarities in the temporal structure of a video sequence by exploiting optical flow fields. The results obtained highlight comparable performances with the state-of-the-art methods which, in general, only resort to single video frames. Furthermore, the proposed optical flow based detection scheme also provides a superior robustness in the more realistic cross-forgery operative scenario and can even be combined with frame-based approaches to improve their global effectiveness. Roberto Caldelli, Leonardo Galteri, Irene Amerini, Alberto Del Bimbo |
Pattern Recognit. Lett. | 2 |
| 2020 | Robust pedestrian detection in thermal imagery using synthesized imagesabstractIn this paper we propose a method for improving pedestrian detection in the thermal domain using two stages: first, a generative data augmentation approach is used, then a domain adaptation method using generated data adapts an RGB pedestrian detector. Our model, based on the Least-Squares Generative Adversarial Network, is trained to synthesize realistic thermal versions of input RGB images which are then used to augment the limited amount of labeled thermal pedestrian images available for training. We apply our generative data augmentation strategy in order to adapt a pretrained YOLOv3 pedestrian detector to detection in the thermal-only domain. Experimental results demonstrate the effectiveness of our approach: using less than 50% of available real thermal training data, and relying on synthesized data generated by our model in the domain adaptation phase, our detector achieves state-of-the-art results on the KAIST Multispectral Pedestrian Detection Benchmark; even if more real thermal data is available adding GAN generated images to the training data results in improved performance, thus showing that these images act as an effective form of data augmentation. To the best of our knowledge, our detector achieves the best single-modality detection results on KAIST with respect to the state-of-the-art. My Kieu, Lorenzo Berlincioni, Leonardo Galteri, Marco Bertini 0001, Andrew D. Bagdanov, Alberto Del Bimbo |
ICPR | 3 |
| 2020 | A NoGAN approach for image and video restoration and compression artifact removalabstractLossy image and video compression algorithms introduce several different types of visual artifacts that reduce the visual quality of the compressed media, and the higher the compression rate the higher is the strength of these artifacts. In this work, we describe an approach for visual quality improvement of compressed images and videos to be performed at presentation time, as to obtain the benefits of fast data transfer and reduced data storage, while enjoying a visual quality that could be obtained only reducing the compression rate. To obtain this result we propose to use a deep neural network trained using the NoGAN approach, adapting the popular DeOldify architecture used for colorization. We show how the proposed method can be applied both to image and video compression artifact removal and restoration. Filippo Mameli, Marco Bertini 0001, Leonardo Galteri, Alberto Del Bimbo |
ICPR | 3 |
| 2020 | Increasing Video Perceptual Quality with GANs and Semantic CodingabstractWe have seen a rise in video based user communication in the last year, unfortunately fueled by the spread of COVID-19 disease. Efficient low-latency delay of transmission of video is a challenging problem which must also deal with the segmented nature of network infrastructure not always allowing a high throughput. Lossy video compression is a basic requirement to enable such technology widely. While this may compromise the quality of the streamed video there are recent deep learning based solutions to restore quality of a lossy compressed video. Leonardo Galteri, Marco Bertini 0001, Lorenzo Seidenari, Tiberio Uricchio, Alberto Del Bimbo |
ACM Multimedia | 1 |
| 2020 | Image and Video Restoration and Compression Artefact Removal Using a NoGAN ApproachabstractLossy image and video compression algorithms introduce several types of visual artefacts that reduce the visual quality of the compressed media. In this work, we report results obtained using the NoGAN training approach and adapting the popular DeOldify architecture used for colorization, for image and video compression artefact removal and restoration. Filippo Mameli, Marco Bertini 0001, Leonardo Galteri, Alberto Del Bimbo |
ACM Multimedia | 3 |
| 2019 | Towards Real-Time Image Enhancement GANs
Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo |
CAIP (1) | 1 |
| 2019 | Coarse to Fine 3D Face Reconstruction from Single ImageabstractIn this demo we propose a coarse to fine reconstruction pipeline, which takes a single RGB image as input and outputs a detailed 3D model of the face. The pipeline is composed by two main blocks, the coarse reconstruction block, which is based on a 3D Morphable Model, and the refinement block, which instead grounds on a Generative Adversarial Network (GAN). Leonardo Galteri, Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo |
FG | 1 |
| 2019 | Fast Video Quality Enhancement using GANsabstractVideo compression algorithms result in a reduction of image quality, because of their lossy approach to reduce the required bandwidth. This affects commercial streaming services such as Netflix, or Amazon Prime Video, but affects also video conferencing and video surveillance systems. In all these cases it is possible to improve the video quality, both for human view and for automatic video analysis, without changing the compression pipeline, through a post-processing that eliminates the visual artifacts created by the compression algorithms. Generative Adversarial Networks have obtained extremely high quality results in image enhancement tasks; however, to obtain such results large generators are usually employed, resulting in high computational costs and processing time. In this work we present an architecture that can be used to reduce the computational cost and that has been implemented on mobile devices. A possible application is to improve video conferencing, or live streaming. In these cases there is no original uncompressed video stream available. Therefore, we report results using no-reference video quality metric showing high naturalness and quality even for efficient networks. Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
ACM Multimedia | 1 |
| 2019 | Deep 3D morphable model refinement via progressive growing of conditional Generative Adversarial Networks
Leonardo Galteri, Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo |
Comput. Vis. Image Underst. | 1 |
| 2019 | Deep Universal Generative Adversarial Compression Artifact RemovalabstractImage compression is a need that arises in many circumstances. Unfortunately, whenever a lossy compression algorithm is used, artifacts will manifest. Image artifacts, caused by compression tend to eliminate higher frequency details and, in certain cases, may add noise or small image structures. There are two main drawbacks of this phenomenon. First, images appear much less pleasant to the human eye. Second, computer vision algorithms, such as object detectors, may be hindered and their performance reduced. Removing such artifacts means recovering the original image from a perturbed version of it. This means that one ideally should invert the compression process through a complicated nonlinear image transformation. We propose an image transformation approach based on a feedforward fully convolutional residual network model. We show that this model can be optimized either traditionally, directly optimizing an image similarity loss (SSIM), or using a generative adversarial approach (GAN). Our GAN is able to produce images with more photorealistic details than SSIM-based networks. We describe a novel training procedure based on subpatches and devise a novel testing protocol to evaluate restored images quantitatively. We show that our approach can be used as a preprocessing step for different computer vision tasks in case images are degraded by compression to a point that state-of-the art algorithms fail. In this case, our GAN-based approach obtains better performance than MSE or SSIM trained networks. Different from previously proposed approaches, we are able to remove artifacts generated at any QF by inferring the image quality directly from data. Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo |
IEEE Trans. Multim. | 1 |
| 2018 | Video Compression for Object Detection AlgorithmsabstractVideo compression algorithms have been designed aiming at pleasing human viewers, and are driven by video quality metrics that are designed to account for the capabilities of the human visual system. However, thanks to the advances in computer vision systems more and more videos are going to be watched by algorithms, e.g. implementing video surveillance systems or performing automatic video tagging. This paper describes an adaptive video coding approach for computer vision-based systems. We show how to control the quality of video compression so that automatic object detectors can still process the resulting video, improving their detection performance, by preserving the elements of the scene that are more likely to contain meaningful content. Our approach is based on computation of saliency maps exploiting a fast objectness measure. The computational efficiency of this approach makes it usable in a real-time video coding pipeline. Experiments show that our technique outperforms standard H.265 in speed and coding efficiency, and can be applied to different types of video domains, from surveillance to web videos. Leonardo Galteri, Marco Bertini 0001, Lorenzo Seidenari, Alberto Del Bimbo |
ICPR | 1 |
| 2017 | Deep Generative Adversarial Compression Artifact RemovalabstractCompression artifacts arise in images whenever a lossy compression algorithm is applied. These artifacts eliminate details present in the original image, or add noise and small structures; because of these effects they make images less pleasant for the human eye, and may also lead to decreased performance of computer vision algorithms such as object detectors. To eliminate such artifacts, when decompressing an image, it is required to recover the original image from a disturbed version. To this end, we present a feed-forward fully convolutional residual network model trained using a generative adversarial framework. To provide a baseline, we show that our model can be also trained optimizing the Structural Similarity (SSIM), which is a better loss with respect to the simpler Mean Squared Error (MSE). Our GAN is able to produce images with more photorealistic details than MSE or SSIM based networks. Moreover we show that our approach can be used as a pre-processing step for object detection in case images are degraded by compression to a point that state-of-the art detectors fail. In this task, our GAN method obtains better performance than MSE or SSIM trained networks. Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo |
ICCV | 1 |
| 2017 | Spatio-Temporal Closed-Loop Object DetectionabstractObject detection is one of the most important tasks of computer vision. It is usually performed by evaluating a subset of the possible locations of an image, that are more likely to contain the object of interest. Exhaustive approaches have now been superseded by object proposal methods. The interplay of detectors and proposal algorithms has not been fully analyzed and exploited up to now, although this is a very relevant problem for object detection in video sequences. We propose to connect, in a closed-loop, detectors and object proposal generator functions exploiting the ordered and continuous nature of video sequences. Different from tracking we only require a previous frame to improve both proposal and detection: no prediction based on local motion is performed, thus avoiding tracking errors. We obtain three to four points of improvement in mAP and a detection time that is lower than Faster Regions with CNN features (R-CNN), which is the fastest Convolutional Neural Network (CNN) based generic object detector known at the moment. Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo |
IEEE Trans. Image Process. | 1 |