Fabrizio Guerrini

dblp:93/8761 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0001-5634-6615ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 2Security and privacy · 2 · 2 first-authorComputer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 Towards compression-aware iris presentation attack detection
abstract
With the growing integration of biometric recognition systems into high-security and large-scale deployment scenarios, it is becoming increasingly important to ensure their robustness under realistic operational constraints. This implies designing solutions able to withstand potential adversarial threats that could affect their integrity and accountability, and also taking into account requirements of real-world operating systems such as limited availability of bandwidth and memory for data transmission and storage. Hence, the proposed study deals with presentation attack detection (PAD) for iris recognition, evaluating the effectiveness of Transformer-based frameworks at detecting spoofing attacks relying on fabricated or artificial biometric evidences. More specifically, we focus on the effects of image compression on the quality of iris images and on the resulting PAD performance, considering both traditional techniques such as JPEG as well as next-generation learning-based image codecs such as JPEG AI. We then examine the feasibility of mitigating compression-induced performance degradation by fine-tuning the adopted models on compressed images, achieving improvements in terms of half total error rate between 5% and 10% for images compressed at the worst JPEG and JPEG AI qualities. We also evaluate the generalizability of the developed solutions by testing them on learning-based codecs not considered during training, to check whether similar PAD-relevant artifacts are introduced by different compressions. Furthermore, we investigate the redundancy within the embeddings generated by the employed detectors, and demonstrated it is possible to significantly compress them while preserving the achievable PAD performance. Overall, the study provides a systematic analysis of iris PAD under compression constraints, offering insights into model adaptation, cross-codec robustness, and representation efficiency in scenarios where visual data coding plays a central role.
Rocco Albano, Filippo Battaglia, Alessandro Gnutti, Emanuele Maiorana, Fabrizio Guerrini, Giuseppe Campobello, Pierangelo Migliorati, Patrizio Campisi
Signal Process. Image Commun.5
2026 JPEG AI Compressed Domain Face Detection: A Multi-Scale Bridging Perspective
abstract
Learning-based image coding is showing improved compression efficiency, while also offering a novel advantage in enabling computer vision tasks directly within the compressed domain. The latent representation created by deep learning methods inherently contains all visual features, without a computationally expensive synthesis process at the decoder. This paper is an invited extension of a previous solution for JPEG AI compressed domain face detection that adapts a RetinaFace-based detector to operate directly on the latent tensor. In addition to a former single-scale bridging solution, this work provides a novel multi-scale bridging architecture to enable a more effective multi-scale compressed domain face detection. The results show a significant performance gain, improving accuracy up to 20% for detection of tiny faces on the WIDER FACE dataset compared to single-scale bridging, and further narrowing the gap when compared to detection on uncompressed or JPEG AI decoded images. Furthermore, since the computationally expensive decoding step is bypassed and since the bridges consist of lower-complexity networks, the overall processing cost is significantly reduced. Single and multi-scale bridging, respectively, have about 10% and 32% the complexity of applying pixel domain face detection on decoded images. The proposed architecture is expected to be extended to other multiscale sensitive vision tasks, as JPEG AI is not specifically designed for any single downstream application.
Ayman Alkhateeb, Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, João Ascenso, Fernando Pereira 0001
IEEE Trans. Multim.3
2025 Variable-Size Symmetry-Based Graph Fourier Transforms for Image Compression
Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, Antonio Ortega
IEEE Trans. Circuits Syst. Video Technol.2
2024 JPEG AI Compressed Domain Face Detection
abstract
Learning-based image coding has achieved competitive performance in terms of compression efficiency, while also gaining a key advantage in the ability to carry out computer vision tasks directly in the compressed domain. In fact, the latent representation which is generated using deep learning techniques may natively encapsulate all visual features needed for processing tasks, thereby eliminating the need to perform the expensive synthesis transform process at the decoder side. In this paper, it is proposed to perform face detection using the latent code present in the JPEG AI architecture. First, some experiments show how decoded images can be efficiently processed for face detection without retraining, albeit with some performance degradation. Then, for the first time a compressed domain RetinaFace-based detector applied to JPEG AI latent representations is competitively proposed. The performance achieved is comparable to the performance of the original RetinaFace applied to the reconstructed JPEG AI images, while reducing computational complexity since it bypasses the image decoding process. It is expected that this approach might be extended to other vision tasks since the JPEG AI representation format is not tailored specifically for any computer vision task.
Ayman Alkhateeb, Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, João Ascenso, Fernando Pereira 0001
MMSP3
2023 Image forgery detection: a survey of recent deep-learning approaches
abstract
Abstract In the last years, due to the availability and easy of use of image editing tools, a large amount of fake and altered images have been produced and spread through the media and the Web. A lot of different approaches have been proposed in order to assess the authenticity of an image and in some cases to localize the altered (forged) areas. In this paper, we conduct a survey of some of the most recent image forgery detection methods that are specifically designed upon Deep Learning (DL) techniques, focusing on commonly found copy-move and splicing attacks. DeepFake generated content is also addressed insofar as its application is aimed at images, achieving the same effect as splicing. This survey is especially timely because deep learning powered techniques appear to be the most relevant right now, since they give the best overall performances on the available benchmark datasets. We discuss the key-aspects of these methods, while also describing the datasets on which they are trained and validated. We also discuss and compare (where possible) their performance. Building upon this analysis, we conclude by addressing possible future research trends and directions, in both deep learning architectural and evaluation approaches, and dataset building for easy methods comparison.
Marcello Zanardelli, Fabrizio Guerrini, Riccardo Leonardi, Nicola Adami
Multim. Tools Appl.2
2021 Symmetry-Based Graph Fourier Transforms: Are They Optimal For Image Compression?
abstract
Traditional block-based transforms are based on applying a single transform to all blocks. As an alternative, better performance in image and video processing and representation can be achieved by choosing one among a discrete set of transforms for each block. As an example, our recently proposed set of multiple transforms called Symmetry-Based Graph Fourier Transforms (SBGFTs) have shown good performance in terms of energy compaction, improving HEVC intra coding performance when used to replace the Discrete Cosine Transform (DCT). This paper further explores the performance of the SBGFTs in a multiple transforms, non-linear approximation perspective, by comparing them with two alternative sets of orthogonal transforms, namely, the Karhunen-Loève Transform (KLT) and the Sparse Orthonormal Transform (SOT). Experimental results confirm that SBGFTs achieve superior representation ability in this context as well, suggesting that they could assume a central role in image compression.
Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, Antonio Ortega
ICIP2
2021 Combining Appearance and Gradient Information for Image Symmetry Detection
abstract
This work addresses the challenging problem of reflection symmetry detection in unconstrained environments. Starting from the understanding on how the visual cortex manages planar symmetry detection, it is proposed to treat the problem in two stages: i) the design of a stable metric that extracts subsets of consistently oriented candidate segments, whenever the underlying 2D signal appearance exhibits definite near symmetric correspondences; ii) the ranking of such segments on the basis of the surrounding gradient orientation specularity, in order to reflect real symmetric object boundaries. Since these operations are related to the way the human brain performs planar symmetry detection, a better correspondence can be established between the outcomes of the proposed algorithm and a human-constructed ground truth. When compared to the testing sets used in recent symmetry detection competitions, a remarkable performance gain can be observed. In additional, further validation has been achieved by conducting perceptual validation experiments with users on a newly built dataset.
Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi
IEEE Trans. Image Process.2
2020 2D Discrete Mirror Transform for Image Non-Linear Approximation
abstract
In this paper, a new 2D transform named Discrete Mirror Transform (DMT) is presented. The DMT is computed by decomposing a signal into its even and odd parts around an optimal location in a given direction so that the signal energy is maximally split between the two components. After minimizing the information required to regenerate the original signal by removing redundant structures, the process is iterated leading the signal energy to distribute into a continuously smaller set of coefficients. The DMT can be displayed as a binary tree, where each node represents the single (even or odd) signal derived from the decomposition in the previous level. An optimized version of the DMT (ODMT) is also introduced, by exploiting the possibility to choose different directions at which performing the decomposition. Experimental simulations have been carried out in order to test the sparsity properties of the DMT and ODMT when applied on images: referring to both transforms, the results show a superior performance with respect to the popular Discrete Cosine Transform (DCT) and Discrete Wavelet Transform (DWT) in terms of non-linear approximation.
Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi
ICPR2
2020 Minimal Information Exchange for Secure Image Hash-Based Geometric Transformations Estimation
abstract
Signal processing applications dealing with secure transmission are enjoying increasing attention lately. This paper provides some theoretical insights as well as a practical solution for transmitting a hash of an image to a central server to be compared with a reference image. The proposed solution employs a rigid image registration technique viewed in a distributed source coding perspective. In essence, it embodies a phase encoding framework to let the decoder estimate the transformation parameters using a very modest amount of information about the original image. The problem is first cast in an ideal setting and then it is solved in a realistic scenario, giving more prominence to low computational complexity in both the transmitter and receiver, minimal hash size, and hash security. Satisfactory experimental results are reported on a standard images set.
Fabrizio Guerrini, Marco Dalai, Riccardo Leonardi
IEEE Trans. Inf. Forensics Secur.1
2019 Iterative Mirror Decomposition for Signal Representation
abstract
In this paper it is shown how to describe any finite-energy continuous or discrete signal through an ordered set of positions to uniquely represent it. This is obtained by designing an iterative decomposition through a series of mirror operations around those positions. The purpose is to find at any step of the decomposition the location that provides for the maximum decoupling between the even and odd components of the signal with respect to it. The algorithm can then be iterated at infinity determining a sequence of positions. The per location information determines the optimal energy decoupling strategy at each stage providing remarkable sparsity in the representation. Thanks to the sparsity of the resulting representation, experimental simulations demonstrate superior approximation capabilities of this proposed non-linear mirror transform.
Fabrizio Guerrini, Alessandro Gnutti, Riccardo Leonardi
ICASSP1
2019 Coding of Image Intra Prediction Residuals Using Symmetric Graphs
abstract
The Discrete Cosine Transform (DCT) is widely deployed by modern image and video coding standards such as JPEG and H.26x. In most cases, the DCT is applied in a separable manner to rows and columns, which limits its ability to represent signals with diagonal orientation. As an alternative, non-separable transforms can represent signals with different orientations, but are significantly more computationally complex. To address this problem, in this paper we propose a set of non-separable Symmetry-Based Graph Fourier Transforms (SBGFTs), whose symmetric structures lead to a faster implementation. We study a practical image coding scenario that exploits the proposed SBGFTs, where for each intra predicted image residual block the optimal graph is chosen by solving a graph-based Rate-Distortion (R-D) problem. Experimental results indicate a coding efficiency higher than JPEG and JPEG2000.
Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, Antonio Ortega
ICIP2
2018 Symmetry-Based Graph Fourier Transforms for Image Representation
abstract
It is well-known that the application of the Discrete Cosine Transform (DCT) in transform coding schemes is justified by the fact that it belongs to a family of transforms asymptotically equivalent to the Karhunen-Loeve Transform (KLT) of a first order Markov process. However, when the pixel-to-pixel correlation is low the DCT does not provide a compression performance comparable with the KLT. In this paper, we propose a set of symmetry-based Graph Fourier Transforms (GFT) whose associated graphs present a totally or partially symmetric grid. We show that this family of transforms well represents both natural images and residual signals outperforming the DCT in terms of energy compaction. We also investigate how to reduce the cardinality of the set of transforms through an analysis that studies the relation between efficient symmetry-based GFTs and the directional modes used in H.265 standard. Experimental results indicate that coding efficiency is high.
Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, Antonio Ortega
ICIP2
2018 Combining the use of CNN classification and strength-driven compression for the robust identification of bacterial species on hyperspectral culture plate images
abstract
Huge streams of diagnostic images are expected to be produced daily in the emerging field of digital microbiology imaging because of the ongoing worldwide spread of Full Laboratory Automation systems. This is redefining the way microbiologists execute diagnostic tasks. In this context, the authors want to assess the suitability and effectiveness of a deep learning approach to solve the diagnostically relevant but visually challenging task of directly identifying pathogens on bacterial growing plates. In particular, starting from hyperspectral acquisitions in the VNIR range and spatial‐spectral processing of cultured plates, they approach the identification problem as the classification of computed spectral signatures of the bacterial colonies. In a highly relevant clinical context (urinary tract infections) and on a database of acquired hyperspectral images, they designed and trained a convolutional neural network for pathogen identification, assessing its performance and comparing it against conventional classification solutions. At the same time, given the expected data flow and possible conservation and transmission needs, they are interested in evaluating the combined use of classification and lossy data compression. To this end, after selecting a suitable wavelet‐based compression technology, they test coding strength‐driven operating points looking for configurations able to provably prevent any classification performance degradation.
Alberto Signoroni, Mattia Savardi, Mario Pezzoni, Fabrizio Guerrini, Simone Arrigoni, Giovanni Turra
IET Comput. Vis.4
2017 Even/odd decomposition made sparse: A fingerprint to hidden patterns
Fabrizio Guerrini, Alessandro Gnutti, Riccardo Leonardi
Signal Process.1
2017 Interactive Film Recombination
abstract
In this article, we discuss an innovative media entertainment application called Interactive Movietelling. As an offspring of Interactive Storytelling applied to movies, we propose to integrate narrative generation through artificial intelligence (AI) planning with video processing and modeling to construct filmic variants starting from the baseline content. The integration is possible thanks to content description using semantic attributes pertaining to intermediate-level concepts shared between video processing and planning levels. The output is a recombination of segments taken from the input movie performed so as to convey an alternative plot. User tests on the prototype proved how promising Interactive Movietelling might be, even if it was designed at a proof of concept level. Possible improvements that are suggested here lead to many challenging research issues.
Fabrizio Guerrini, Nicola Adami, Sergio Benini, Alberto Piacenza, Julie Porteous, Marc Cavazza, Riccardo Leonardi
ACM Trans. Multim. Comput. Commun. Appl.1
2013 Tracking characters in movies within logical story units
abstract
In this paper, we propose a methodology to allow movie character recognition and tracking within movie scenes. In detail, we present a combination of a tracking algorithm robust against the problems of the currently available face detection algorithms and a face recognition process. We test how effective the system is in terms of both face tracking effectiveness and precision-recall results obtained for the recognition of the main characters present in an input movie.
Alberto Piacenza, Fabrizio Guerrini, Nicola Adami, Riccardo Leonardi
MMSP2
2011 Generating story variants with constrained video recombination
abstract
We present a novel approach to the automatic generation of filmic variants within an implemented Video-Based Storytelling (VBS) system that successfully integrates video segmentation with stochastically controlled re-ordering techniques and narrative generation via AI planning. We have introduced flexibility into the video recombination process by sequencing video shots in a way that maintains local video consistency and this is combined with exploitation of shot polysemy to enable shot reuse in a range of valid semantic contexts. Results of evaluations on output narratives using a shared set of video data show consistency in terms of local video sequences and global causality with no loss of generative power.
Alberto Piacenza, Fabrizio Guerrini, Nicola Adami, Riccardo Leonardi, Julie Porteous, Jonathan Teutenberg, Marc Cavazza
ACM Multimedia2
2011 Changing video arrangement for constructing alternative stories
abstract
Currently, automatic generation of filmic variants faces a number of key technical issues and thus it usually resorts to the shooting of multiple versions of alternative scenes. However, recent advancements in video analysis has made this objective feasible, though semantic consistency must be somehow preserved. This demo presents a video-based storytelling (VBS) system that successfully integrates video processing with narrative generation by means of a shared semantic description. The novel filmic variants are constructed through a flexible video recombination process that takes advantage of the polysemy of baseline video segments. The short output video clips shown in this demo prove how the generated narratives are semantically consistent while keeping generative power intact.
Alberto Piacenza, Fabrizio Guerrini, Nicola Adami, Riccardo Leonardi, Jonathan Teutenberg, Julie Porteous, Marc Cavazza
ACM Multimedia2
2011 High Dynamic Range Image Watermarking Robust Against Tone-Mapping Operators
abstract
High dynamic range (HDR) images represent the future format for digital images since they allow accurate rendering of a wider range of luminance values. However, today special types of preprocessing, collectively known as tone-mapping (TM) operators, are needed to adapt HDR images to currently existing displays. Tone-mapped images, although of reduced dynamic range, have nonetheless high quality and hence retain some commercial value. In this paper, we propose a solution to the problem of HDR image watermarking, e.g., for copyright embedding, that should survive TM. Therefore, the requirements imposed on the watermark encompass imperceptibility, a certain degree of security, and robustness to TM operators. The proposed watermarking system belongs to the blind, detectable category; it is based on the quantization index modulation (QIM) paradigm and employs higher order statistics as a feature. Experimental analysis shows positive results and demonstrates the system effectiveness with current state-of-art TM algorithms.
Fabrizio Guerrini, Masahiro Okuda, Nicola Adami, Riccardo Leonardi
IEEE Trans. Inf. Forensics Secur.1
2010 CBCD based on color features and landmark MDS-assisted distance estimation
abstract
Content-Based Copy Detection (CBCD) of digital videos is an important research field that aims at the identification of modified copies of an original clip, e.g., on the Internet. In this application, the video content is uniquely identified by the content itself, by extracting some compact features that are robust to a certain set of video transformations. Given the huge amount of data present in online video databases, the computational complexity of the feature extraction and comparison is a very important issue. In this paper, a landmark based multi-dimensional scaling technique is proposed to speed up the detection procedure which is based on exhaustive search and the MPEG-7 Dominant Color Descriptor. The method is evaluated under the MPEG Video Signature Core Experiment conditions, and simulation results show impressive time savings at the cost of a slightly reduced detection performance.
Marzia Corvaglia, Fabrizio Guerrini, Riccardo Leonardi, Pierangelo Migliorati, Eliana Rossi
ICASSP2
2010 Toward a multi-feature approach to Content-Based Copy Detection
abstract
Video Content-Based Copy Detection (CBCD) is an emergent research field which is targeted to the identification of modified copies of an original clip in a given dataset, e.g., on the Internet. As opposed to digital watermarking, the content itself is used to uniquely identify the video through the extraction of features that need to be robust against a certain set of predetermined video attacks. This paper advocates the use of multiple features together with detection performance estimation to construct a flexible video signature instead of a fixed, single feature based one. To combine diverse features, a normalized linear combination is also proposed. The system performance boost is evaluated through the MPEG Video Signature Core Experiment dataset and experimental results show how the proposed signature scheme can achieve impressive improvements with respect to the single feature approach.
Marzia Corvaglia, Fabrizio Guerrini, Riccardo Leonardi, Pierangelo Migliorati, Eliana Rossi
ICIP2
2006 Image Watermarking Robust Against Non-Linear Value-Metric Scaling Based on Higher Order Statistics
abstract
A new QIM-based image watermarking system for still images is proposed. The new system is expressly designed to cope with non-linear value-metric scaling attacks such as histogram stretching and gamma correction. By recognizing that any value-metric scaling attack must not change the global appearance of the image, we argue that the watermark should be inserted into high level visual features. We move a first step into this direction by proposing a system embedding the watermark into the kurtosis of selected image blocks. Though the kurtosis is not strictly invariant against non-linear gain, its value tends to remain constant whenever the image content is not altered significantly. The experiments we carried out confirm the validity of the new system, though some problems still need to be solved to make it suitable for real applications
Fabrizio Guerrini, Riccardo Leonardi, Mauro Barni
ICASSP (5)1