Sebastiano Battiato

dblp:b/SBattiato · DBLP profile ↗
← Back
117ranked-venue papers
39as first author
40since 2021 · last 2026
0000-0001-6127-2470ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 85 · 31 first-author · 24 since 2021Artificial intelligence and machine learning · 35 · 8 first-author · 15 since 2021Security and privacy · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Theory of computation · 2 · 2 first-authorComputer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 TRUST: Leveraging Text Robustness for Unsupervised Domain Adaptation
abstract
Recent unsupervised domain adaptation (UDA) methods have shown great success in addressing classical domain shifts (e.g., synthetic-to-real), but they still suffer under complex shifts (e.g. geographical shift), where both the background and object appearances differ significantly across domains. Prior works showed that the language modality can help in the adaptation process, exhibiting more robustness to such complex shifts. In this paper, we introduce TRUST, a novel UDA approach that exploits the robustness of the language modality to guide the adaptation of a vision model. TRUST generates pseudo-labels for target samples from their captions and introduces a novel uncertainty estimation strategy that uses normalised CLIP similarity scores to estimate the uncertainty of the generated pseudo-labels. Such estimated uncertainty is then used to reweight the classification loss, mitigating the adverse effects of wrong pseudo-labels obtained from low-quality captions. To further increase the robustness of the vision model, we propose a multimodal soft-contrastive learning loss that aligns the vision and language feature spaces, by leveraging captions to guide the contrastive training of the vision model on target images. In our contrastive loss, each pair of images acts as both a positive and a negative pair and their feature representations are attracted and repulsed with a strength proportional to the similarity of their captions. This solution avoids the need for hardly determining positive and negative pairs, which is critical in the UDA setting. Our approach outperforms previous methods, setting the new state-of-the-art on classical (DomainNet) and complex (GeoNet) domain shifts. The code is available at https://github.com/MattiaLitrico/TRUST-Leveraging-Text-Robustness-for-Unsupervised-Domain-Adaptation.
Mattia Litrico, Mario Valerio Giuffrida, Sebastiano Battiato, Devis Tuia
AAAI3
2026 HypDeformNet: Edge-Deployable Deep Architecture with Jacobian-Stable Hyperbolic Deformation and Lipschitz Distillation for Immunotherapy Response Prediction
Francesco Rundo, Massimo Orazio Spata, Giuseppe L. Banna, Sebastiano Battiato
ICPR (8)4
2026 A Novel Metric for Detecting Memorization in Generative Models for Brain MRI Synthesis
abstract
Deep generative models have emerged as a transformative tool in medical imaging, offering substantial potential for synthetic data generation. However, recent empirical studies highlight a critical vulnerability: these models can memorize sensitive training data, posing significant risks of unauthorized patient information disclosure. Detecting memorization in generative models remains particularly challenging, necessitating scalable methods capable of identifying training data leakage across large sets of generated samples. In this work, we propose DeepSSIM, a novel self-supervised metric for quantifying memorization in generative models. DeepSSIM is trained to: i) project images into a learned embedding space and ii) force the cosine similarity between embeddings to match the ground-truth Structural Similarity Index (SSIM) scores computed in the image space. To capture domain-specific anatomical features, training incorporates structure-preserving augmentations, allowing DeepSSIM to estimate similarity reliably without requiring precise spatial alignment. We evaluate DeepSSIM in two case studies using synthetic brain MRI and chest X-ray data generated by a Latent Diffusion Model (LDM) trained under memorization-prone conditions. Compared to state-of-the-art memorization metrics, DeepSSIM achieves superior performance, improving F1 scores by an average of +52.03% over the best existing method. Code and data are publicly available at https://github.com/brAIn-science/DeepSSIM.
Antonio Scardace, Lemuel Puglisi, Francesco Guarnera, Sebastiano Battiato, Daniele Ravì
WACV4
2026 Proto-LeakNet: Towards signal-leak aware attribution in synthetic human face imagery
abstract
The growing sophistication of synthetic image and deepfake generation models has turned source attribution and authenticity verification into a critical challenge for modern computer vision systems. Recent studies suggest that diffusion pipelines unintentionally imprint persistent statistical traces, known as signal-leaks, within their outputs, particularly in latent representations. Building on this observation, we propose Proto-LeakNet, a signal-leak-aware and interpretable attribution framework that integrates Closed-set classification with a density-based Open-set evaluation on the learned embeddings, enabling analysis of unseen generators without retraining. Acting in the latent domain of diffusion models, our method re-simulates partial forward diffusion to expose residual generator-specific cues. A temporal attention encoder aggregates multi-step latent features, while a feature-weighted prototype head structures the embedding space and enables transparent attribution. Trained solely on closed data and achieving a Macro AUC of 98.13%, Proto-LeakNet learns a latent geometry that remains robust under post-processing, surpassing state-of-the-art methods, and achieves strong separability both between real images and known generators, and between known and unseen ones. The codebase will be available after acceptance.
Claudio Giusti, Luca Guarnera, Sebastiano Battiato
Comput. Vis. Image Underst.3
2026 Correction: CNNMC: a convolutional neural network with Monte Carlo dropout for speaker recognition
Massimo Orazio Spata, Alessandro Ortis, Georgia Fargetta, Sebastiano Battiato
J. Inf. Secur.4
2026 Stability-plasticity inspired knowledge distillation expert system with Lipschitz-regularized neuro-modulation for silicon-carbide power modules health monitoring in next-generation electric vehicles
abstract
Edge-side Prognostics and Health Management (PHM) for Electric Vehicles (EVs) demands vehicle sub-systems Remaining Useful Life (RUL) predictors running on automotive-grade Microcontroller Units (MCUs) under tight memory, latency, and energy budgets. Classical Knowledge Distillation (KD) often degrades long-horizon accuracy and fails on the complex forecasting-to-decision pipeline for Silicon-Carbide (SiC) traction-inverter power modules. We propose Neuro-Modulated Knowledge Distillation (NM-KD), distilling a Multi-Scale Temporal Fusion Transformer (TFT-MS) into a lightweight model deployable on MCUs. NM-KD adapts loss weights via learned gates, bounds student sensitivity with a Lipschitz regularizer, and leverages Stochastic Weight Averaging (SWA) modulation. The TFT-MS teacher is benchmarked against several architectures spanning six design families, while the student selection is validated against five MCU-deployable alternatives. On SiC power devices, the student matches short-horizon performance and outperforms the teacher at long horizon (50-step accuracy 95.13% vs. 94.32%). On-target deployment reduces non-volatile memory by 83.12%, latency by 81.73%, and energy by 60.57%. Cross-batch generalization on 21 modules from two independent manufacturing lots confirms robustness to process variability. A lightweight temporal remapping converts accelerated-cycle predictions into field-relevant RUL estimates. NM-KD enables fast, low-power, accurate PHM on automotive Electronic Control Units (ECUs).
Francesco Rundo, Massimo Orazio Spata, Carmelo Pino, Michele Calabretta, Angelo Alberto Messina, Michael Rundo, Sebastiano Battiato
Expert Syst. Appl.7
2026 Uncertainty-guided Open-Set Source-Free Unsupervised Domain Adaptation with Target-private Class Segregation
abstract
Standard Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target, requiring simultaneous access to both source and target data. Moreover, UDA approaches commonly assume that source and target domains share the same labels space. Yet, these two assumptions are hardly satisfied in real-world scenarios. This paper considers the more challenging Source-Free Open-set Domain Adaptation (SF-OSDA) setting, where both assumptions do not hold. We propose a novel approach for SF-OSDA that takes advantage of the granularity of target-private categories by segregating their samples into multiple unknown classes. Starting from an initial clustering-based pseudo-labels initialisation, our method progressively improves the segregation of target-private samples by refining their pseudo-labels with the guide of an uncertainty-based sample selection module. Additionally, we propose a novel contrastive loss, named NL-InfoNCELoss, that, integrating negative learning into self-supervised contrastive learning, enhances the model robustness to noisy pseudo-labels. Extensive experiments on benchmark datasets demonstrate the superiority of our proposed approach over competing methods, establishing new state-of-the-art performance. Notably, additional analyses show that our method is able to learn the underlying semantics of novel classes, opening the possibility to perform novel class discovery.
Mattia Litrico, Davide Talon, Sebastiano Battiato, Alessio Del Bue, Mario Valerio Giuffrida, Pietro Morerio
Int. J. Comput. Vis.3
2026 Fraud is not just rarity: A causal prototype attention approach to realistic synthetic oversampling
abstract
Detecting fraudulent credit card transactions remains a significant challenge, due to the extreme class imbalance in real-world data and the often subtle patterns that separate fraud from legitimate activity. Existing research commonly attempts to address this by generating synthetic samples for the minority class using approaches such as GANs, VAEs (Variational Autoencoders), or hybrid generative models. However, these techniques, particularly when applied only to minority-class data, tend to result in overconfident classifiers and poor latent cluster separation, ultimately limiting real-world detection performance. In this study, we propose the Causal Prototype Attention Classifier (CPAC), an interpretable architecture that promotes class-aware clustering and improved latent space structure through prototype-based attention mechanisms and we couple it with the encoder of a Variational Autoencoder–Generative Adversarial Network (VAE-GAN) in order to achieve improved latent cluster separation moving beyond post-hoc sample augmentation. We compared CPAC-augmented models to traditional oversamplers, such as SMOTE, as well as to state-of-the-art generative models, both with and without CPAC-based latent classifiers. Our results show that classifier-guided latent shaping with CPAC delivers superior performance, achieving an F1-score of 93.74% and recall of 92.85%, along with improved latent cluster separation. Further ablation studies and visualizations provide deeper insight into the benefits and limitations of classifier-driven representation learning for fraud detection. The codebase for this work will be available at final submission.
Claudio Giusti, Luca Guarnera, Mirko Casu, Sebastiano Battiato
Knowl. Based Syst.4
2026 Count2Density: Crowd density estimation without location-level annotations
abstract
Crowd density estimation is a well-known computer vision task aimed at estimating the density distribution of people in an image. The primary challenge in this domain is the reliance on fine-grained location-level annotations — i.e., points placed on top of each individual — to train deep networks. Collecting such detailed annotations is both tedious, time-consuming, and poses a significant barrier to scalability for real-world applications. To alleviate this burden, we present Count2Density : a novel pipeline designed to predict meaningful density maps containing quantitative spatial information using only count-level annotations (i.e., the total number of people) during training. To achieve this, Count2Density generates pseudo-density maps leveraging past predictions stored in a Historical Map Bank, thereby reducing confirmation bias. This bank is initialised using an unsupervised saliency estimator to provide an initial spatial prior and is iteratively updated with an Exponential Moving Average of predicted density maps. These pseudo-density maps are obtained by sampling locations from estimated crowd areas using a hypergeometric distribution, with the number of samplings determined by the count-level annotations. To further enhance the spatial awareness of the model and promote robust feature learning, we add a self-supervised contrastive spatial regulariser to encourage similar feature representations within crowded regions while maximising dissimilarity with background regions. Experimental results demonstrate that our approach significantly outperforms cross-domain adaptation methods and achieves better results than recent state-of-the-art approaches in semi-supervised settings across several datasets. Additional analyses validate the effectiveness of each individual component of our pipeline, including the self-supervised contrastive regulariser, confirming the ability of Count2Density to effectively retrieve spatial information from count-level annotations and enabling accurate subregion counting.
Mattia Litrico, Michael P. Pound, Sotirios A. Tsaftaris, Sebastiano Battiato, Mario Valerio Giuffrida
Pattern Recognit.5
2025 WILD: a new in-the-Wild Image Linkage Dataset for synthetic image attribution
abstract
Synthetic image source attribution is an open challenge, with an increasing number of image generators being released yearly. The complexity and the sheer number of available generative techniques, as well as the scarcity of high-quality open source datasets of diverse nature for this task, make training and benchmarking synthetic image source attribution models very challenging. WILD1is a new in-the-Wild Image Linkage Dataset designed to provide a powerful training and benchmarking tool for synthetic image attribution models. The dataset is built out of a closed set of 10 popular commercial generators, which constitutes the training base of attribution models, and an open set of 10 additional generators, simulating a real-world in-the-wild scenario. Each generator is represented by 1,000 images, for a total of 10,000 images in the closed set and 10,000 images in the open set. Half of the images are post-processed with a wide range of operators. WILD allows benchmarking attribution models in a wide range of tasks, including closed and open set identification and verification, and robust attribution with respect to post-processing and adversarial attacks. Models trained on WILD are expected to benefit from the challenging scenario represented by the dataset itself. Moreover, an assessment of seven baseline methodologies on closed and open set attribution is presented, including robustness tests with respect to post-processing.
Pietro Bongini, Sara Mandelli, Andrea Montibeller, Mirko Casu, Orazio Pontorno, Claudio Vittorio Ragaglia, Luca Zanchetta, Mattia Aquilina, Taiba Majid Wani, Luca Guarnera, Benedetta Tondi, Giulia Boato, Paolo Bestagini, Irene Amerini, Francesco G. B. De Natale, Sebastiano Battiato, Mauro Barni
IJCNN16
2025 End-to-end Audio Deepfake Detection from RAW Waveforms: a RawNet-Based Approach with Cross-Dataset Evaluation
abstract
Audio deepfakes represent a growing threat to digital security and trust, leveraging advanced generative models to produce synthetic speech that closely mimics real human voices. Detecting such manipulations is especially challenging under open-world conditions, where spoofing methods encountered during testing may differ from those seen during training. In this work, we propose an end-to-end deep learning framework for audio deepfake detection that operates directly on raw waveforms. Our model, RawNetLite, is a lightweight convolutional-recurrent architecture designed to capture both spectral and temporal features without handcrafted preprocessing. To enhance robustness, we introduce a training strategy that combines data from multiple domains and adopts Focal Loss to emphasize difficult or ambiguous samples. We further demonstrate that incorporating codec-based manipulations and applying waveform-level audio augmentations (e.g., pitch shifting, noise, and time stretching) leads to significant generalization improvements under realistic acoustic conditions. The proposed model achieves over 99.7% F1 and 0.25% EER on in-domain data (FakeOrReal), and up to 83.4% F1 with 16.4% EER on a challenging out-of-distribution test set (AVSpoof2021 + CodecFake). These findings highlight the importance of diverse training data, tailored objective functions and audio augmentations in building resilient and generalizable audio forgery detectors. Code and pretrained models are available at https://iplab.dmi.unict.it/mfs/Deepfakes/PaperRawNet2025/.
Andrea Di Pierno, Luca Guarnera, Dario Allegra, Sebastiano Battiato
IJCNN4
2025 Attention-Enhanced Convolutional Deep Architecture with Adaptive Jacobian Regularization for Joint Multi-Modal Optical and X-Ray Screening of Silicon-Carbide Power Devices
abstract
The automotive development of Electric Vehicles (EVs) is accelerating due to environmental concerns and technological advances. The complexity and performance demand of EVs require the use of Silicon-Carbide (SiC) technologies for the related efficiency, high-temperature resistance, and robust dynamic switching. To ensure the high performance of electric vehicles, the rigorous screening of the delivered SiC devices plays a pivotal role. Defects in SiC production can cause failures that directly affect the performance of electric vehicle engines, particularly in the traction inverter subsystem. To identify these defects early, advanced visual screening through optical microscopy and X-ray approaches has been proposed by leading car makers. These screening methodologies require advanced technical expertise and are still prone to significant errors, even when performed manually by experienced operators. To address these inefficiencies, the authors propose an ensemble deep learning pipeline for performing a robust automated visual inspection of SiC devices. By combining Multi-Head Attention blocks with adaptive input data distortion compensation through Jacobian regularization within convolutional architectures, the proposed model leverages global context and local feature maps, enhancing accuracy and robustness in multimodal defects detection. The proposed combined approach, tested on ACEPACK/TPACK DRIVE SiC power modules provided by STMicroelectronics, achieved an average accuracy of approximately 93% in both methodologies.
Francesco Rundo, Carmelo Pino, Giulia Castagnolo, Angelo Alberto Messina, Michele Calabretta, Sebastiano Battiato
IJCNN6
2025 Adversarial Attacks on Deepfake Detectors: A Challenge in the Era of AI-Generated Media (AADD-2025)
abstract
The rapid proliferation of AI-generated media, particularly hyper-realistic deepfakes, has underscored the critical need for robust detection systems to mitigate risks such as misinformation and identity theft. However, state-of-the-art deepfake detectors remain vulnerable to adversarial attacks-subtle perturbations designed to evade classification. To address this gap, we organized the Adversarial Attacks on Deepfake Detectors (AADD-2025) challenge, a competitive evaluation aimed at advancing methodologies to expose and strengthen weaknesses in deepfake detection models. The challenge tasked participants with generating adversarial examples capable of evading four diverse classifiers (including ResNet, DenseNet, and two blind models) while preserving structural similarity to original deepfakes. A dataset comprising 16 subsets of high- and low-quality deepfake images generated by GAN-based and diffusion models (e.g., StableDiffusion, StyleGAN3) was provided. Participants were evaluated using a weighted combination of Structural Similarity Index (SSIM) and attack success rates across all classifiers. Thirteen teams proposed innovative solutions leveraging techniques such as latent-space manipulation, ensemble gradient optimization, surrogate modeling, and frequency-domain perturbation. Top-performing approaches, including MR-CAS (1st place), Safe AI (2nd place), and RoMa (3rd place), achieved high SSIM scores (0.74-0.93) while successfully misleading classifiers. Notably, MR-CAS's latent diffusion model inversion strategy and Safe AI's consensus-orthogonal gradient weighting framework demonstrated superior transferability across architectures, including Vision Transformers. The challenge revealed critical insights: latent-space attacks outperformed pixel-level methods, ensemble-based strategies enhanced cross-model robustness, and adversarial perturbations optimized for both CNNs and transformers proved most effective. However, gaps persist in generalizing attacks across heterogeneous models and maintaining perceptual fidelity, highlighting the urgency of developing adaptive defenses and hybrid detection mechanisms. By fostering collaboration and innovation, AADD-2025 provides a benchmark for evaluating adversarial robustness in deepfake detection and underscores the need for resilient systems in the era of AI-generated media.
Sebastiano Battiato, Mirko Casu, Francesco Guarnera, Luca Guarnera, Giovanni Puglisi, Orazio Pontorno, Claudio Vittorio Ragaglia, Zahid Akhtar
ACM Multimedia1
2025 (DFF '25) 1st Deepfake Forensics Workshop: Detection, Attribution, Recognition, and Adversarial Challenges in the Era of AI-Generated Media
abstract
The proliferation of generative models, particularly Generative Adversarial Networks (GANs) and Diffusion Models, has reshaped multimedia content creation. Alongside creative and commercial opportunities, they have introduced unprecedented risks through the production of highly realistic synthetic content, or deepfakes. These artifacts challenge visual and auditory trust, with major implications for media, security, politics, and law. This workshop provides a forum to examine deepfake technology from forensic, technical, legal, and social perspectives. It will bring together experts to advance robust and explainable detection methods, define benchmarking practices, and address ethical and regulatory frameworks. Topics include detection and attribution, adversarial countermeasures, multimodal analysis, model traceability, legal admissibility of synthetic content, as well as real-world deployment challenges and dataset creation. Further information about the workshop is available at https://iplab.dmi.unict.it/mfs/acm-dff-ws-2025/
Sebastiano Battiato, Mirko Casu, Francesco Guarnera, Luca Guarnera, Giovanni Puglisi, Orazio Pontorno, Claudio Vittorio Ragaglia, Zahid Akhtar
ACM Multimedia1
2025 CNNMC: a convolutional neural network with Monte Carlo dropout for speaker recognition
abstract
Speaker recognition is the task of identifying or verifying a person’s identity using their voice. This problem involves challenges like variations in speech due to emotional states, health conditions, heterogeneity of microphone models, different environments and background noise. Accurate speaker recognition is critical for security, personalization, and forensic applications. Applying a CNN with Monte Carlo dropout can enhance Speaker Recognition by enabling robust uncertainty-aware predictions, making the presented architecture particularly effective for smaller, noisy datasets without the need for large-scale pre-training. This approach helps mitigate overfitting and improves generalization, making it effective in handling diverse speech patterns. The designed deep learning model showcases superior performance in multiple dimensions, achieving a peak validation accuracy of 93.27% for speaker recognition on a specific dataset recorded in the wild by phone, and 0.030 of EER, showing competitive performance with respect to state-of-the-art baselines.
Massimo Orazio Spata, Alessandro Ortis, Georgia Fargetta, Sebastiano Battiato
EURASIP J. Inf. Secur.4
2025 An open source framework for video streaming in cloud gaming
abstract
Abstract Digital games often play the role of vectors for innovations in computer graphics and multimedia. Contemporary digital games feature realistic graphics and complex mechanics, which increase the computational burden of a proper gaming experience; consequently, the costs for customer equipment raises. Cloud gaming is a technique in which a high-performance server running a videogame receives the player’s input and streams the video back to a lightweight client. Currently, available open-source frameworks designed in this scope suffer from deprecation, limited capabilities or strict hardware and operating system dependencies. In this paper, we present a novel open-source framework to define and benchmark architectures for remote rendering of screen content, to provide a platform-agnostic tool for researchers that want to contribute to this area. We run a substantial set of experiments about streaming gaming sessions with different combinations of network conditions, codec parameters and transmission policies, thus collecting network statistics and dumping transmitted and received video frames, finally demonstrating the usability of the provided tools. We show the various features of the framework to profile and perform both an analysis of network statistics and of the stream’s visual quality.
Lorenzo Catania, Oliver Giudice, Sebastiano Battiato, Filippo Stanco, Dario Allegra
Multim. Tools Appl.3
2025 DeepFeatureX-SN: Generalization of deepfake detection via contrastive learning
abstract
Abstract The rapid advancement of generative artificial intelligence, particularly in the domains of Generative Adversarial Networks (GANs) and Diffusion Models (DMs), has led to the creation of increasingly sophisticated deepfakes. These synthetic images pose significant challenges for detection systems and present growing concerns in the realm of Cybersecurity. The potential misuse of deepfakes for disinformation, fraud, and identity theft underscores the critical need for robust detection methods. This paper introduces DeepFeatureX-SN (‘Deep Features eXtractors based Siamese Network’), an innovative deep learning model designed to address the complex task of not only distinguishing between real and synthetic images but also identifying the specific employed generative technique (GAN or DM). Our approach makes use of a tripartite structure of specialized base models, each trained using Siamese networks and contrastive learning techniques, to extract discriminative features unique to real, GAN-generated, and DM-generated images. These features are then combined through a CNN-based classifier for final categorization. Extensive experiments demonstrate the model’s superior performance, with a detection accuracy of 97.29%, strong generalization to unseen generative architectures (achieving an average accuracy of 67.40%, which surpasses most existing approaches by over 10%) and robustness against various image manipulations, all of which are crucial for real-world Cybersecurity applications. DeepFeatureX-SN achieves state-of-the-art results across multiple datasets, showing particular strength in detecting images from novel GAN and DM implementations. Furthermore, a comprehensive ablation study validates the effectiveness of each component in our proposed architecture. This research contributes significantly to the field, offering a more nuanced and accurate approach to identifying and categorizing synthetic images. The results obtained in the different configurations in the generalization tests demonstrate the good capabilities of the model, outperforming methods found in the literature. Codes and models are available at https://iplab.dmi.unict.it/mfs/Deepfakes/DeepFeatureX-SN/ .
Orazio Pontorno, Luca Guarnera, Sebastiano Battiato
Multim. Tools Appl.3
2025 Smoking Detection and Cessation: An Updated Scoping Review of Digital and Mobile Health Technologies
abstract
Digital and mobile health technologies offer promising solutions for smoking detection and cessation. This scoping review examines the current state of research and development in this field, encompassing smartphone applications, wearable devices, and sensor-based systems. We analyzed 49 studies published between 2019 and 2023 from PubMed and ACM Digital Library, focusing on technology features, outcomes, and evaluation methods. Wearable sensors and smartphone apps show potential in combating smoking addiction and improving quit rates. Motion sensors for hand-to-mouth gesture detection achieve high accuracy in controlled settings but face challenges in real-world applications. Machine learning models and wireless signal detection techniques yield encouraging results but require further refinement. Smartphone apps provide personalized plans and progress tracking, though most rely on manual logging and lack rigorous scientific evaluation. Our findings suggest that digital health technologies could significantly enhance smoking cessation efforts. However, more robust evaluation methods and integration of sensor data with machine learning are needed to improve usability and effectiveness. Continued research and innovation in this field are crucial for developing reliable, practical solutions and integrating these technologies into clinical programs.
Mirko Casu, Francesco Guarnera, Giusy Rita Maria La Rosa, Sebastiano Battiato, Pasquale Caponnetto, Riccardo Polosa, Rosalia Emma
IEEE J. Biomed. Health Informatics4
2025 Benchmarking computer vision architectures for cloud detection from lidar ceilometer backscatter data
abstract
Abstract Cloud detection is fundamental for accurate weather monitoring, often achieved through remote sensing technology, such as satellite imagery or radar. This study explores the use of lidar ceilometer backscatter data, a rich but noisy source of atmospheric information, to enhance cloud detection. Leveraging data acquired from a Lufft CHM 15k ceilometer over three months near Mount Etna, Italy, we gathered a novel dataset comprising time-height plots derived from backscatter profiles. The Weather Research and Forecasting (WRF) model was used for ground-truth data labeling, ensuring reliable model validation. We benchmarked state-of-the-art deep learning architectures, including CNN-based models (e.g., ResNet50, VGG16, InceptionV3, EfficientNet) and the Vision Transformer (ViT), on our collected dataset. Among these, ResNet50 achieved the highest accuracy ( $$89.57\%$$ 89.57 % ), closely followed by ViT ( $$89.36\%$$ 89.36 % ), showcasing the efficacy of residual learning and transformer-based approaches in extracting complex patterns from atmospheric data. Our results highlight the potential of lidar-based systems for accurate cloud detection, complementing other remote sensing technologies. Our work contributes to the field by introducing a publicly available dataset and providing comprehensive benchmarking results that establish a baseline for future research. This study also opens avenues for broader applications of ceilometer data, such as the detection of pollutants and other atmospheric phenomena. Our dataset is publicly available at https://zenodo.org/records/10616434 .
Alessio Barbaro Chisari, Luca Guarnera, Alessandro Ortis, Wladimiro Carlo Patatu, Sebastiano Battiato, Mario Valerio Giuffrida
Vis. Comput.5
2024 Advantages of brain parcellation in Multiple Sclerosis Lesion Segmentation
abstract
Segmentation of multiple sclerosis lesions plays an important role in understanding disease status. In this work, we focus on the effectiveness of brain parcellation in enhancing the performance of segmentation for multiple sclerosis lesions in Magnetic Resonance Imaging. Brain parcellation does not improve the segmentation performance, but make the results more robust in terms of overall variability (e.g. standard deviation), by dividing the brain into physically significant sub-regions that the model can concentrate on. Our approach combines parcellation with the existing diffusion-based model to increase sensitivity, particularly in regions with small anomalies. We conducted a thorough evaluation of a reference dataset on the field using all available modalities. Our results show how the parcellation of the brain when integrated into a diffusion-based pipeline, makes the segmentation of MS more stable, lowering deviations from the average, and improving some of the results w.r.t. state-of-the-art. This method achieves good segmentation capabilities even with small datasets, providing promising indications for further research.
Dario Samuele Pishvai, Alessia Rondinella, Francesco Guarnera, Sebastiano Battiato
BIBM4
2024 Federated Learning in a Semi-Supervised Environment for Earth Observation Data
abstract
We propose FedRec, a federated learning workflow taking advantage of unlabelled data in a semi-supervised environment to assist in the training of a supervised aggregated model.In our proposed method, an encoder architecture extracting features from unlabelled data is aggregated with the feature extractor of a classification model via weight averaging.The fully connected layers of the supervised models are also averaged in a federated fashion.We show the effectiveness of our approach by comparing it with the state-of-the-art federated algorithm, an isolated and a centralised baseline, on novel cloud detection datasets.Our code is available at https://github.com/CasellaJr/FedRec.* This work has been partly supported by the Spoke "FutureHPC & BigData" of the ICSC -Centro Nazionale di Ricerca in "High Performance Computing, Big Data and Quantum Computing", funded by European Union -
Bruno Casella, Alessio Barbaro Chisari, Marco Aldinucci, Sebastiano Battiato, Mario Valerio Giuffrida
ESANN4
2024 Innovative Methods for Non-Destructive Inspection of Handwritten Documents
abstract
Handwritten document analysis is an area of forensic science, with the goal of establishing authorship of documents through examination of inherent characteristics. Law enforcement agencies use standard protocols based on manual processing of handwritten documents. This method is time-consuming, is often subjective in its evaluation, and is not replicable. To overcome these limitations, in this paper we present a framework capable of extracting and analyzing intrinsic measures of manuscript documents related to text line heights, space between words, and character sizes using image processing and deep learning techniques. The final feature vector for each document involved consists of the mean (η) and standard deviation (σ) for every type of measure collected. By quantifying the Euclidean distance between the feature vectors of the documents to be compared, authorship can be discerned. Our study pioneered the comparison between traditionally handwritten documents and those produced with digital tools (e.g., tablets). Experimental results demonstrate the ability of our method to objectively determine authorship in different writing media, outperforming the state of the art.
Eleonora Breci, Luca Guarnera, Sebastiano Battiato
ICASSP3
2024 On the Cloud Detection from Backscattered Images Generated from a Lidar-Based Ceilometer: Current State and Opportunities
abstract
Accurate weather monitoring depends significantly on cloud detection, a crucial process achievable through remote sensing tools such as satellite imagery and radar or through the analysis of data obtained from ceilometers. A ceilometer is a lidar-based device allowing to analyse the atmosphere and detect the presence of particles within clouds. The data retrieved from ceilometers involve analysis of the backscatter of the lidar signal returning to the surface. Given the inherent noise in this data, we leverage deep learning models to detect the presence of clouds in the data. To label the data, we take advantage of a Weather Research & Forecasting (WRF) model, which provided us with ground-truth used for validation purposes. We performed a comparative analysis with current state-of-the-art deep learning architectures on this specialist domain. This comparative analysis shows that the best model is ResNet 50, but also a transformer-based model, such as ViT, achieves great results. These preliminary results pave the scenario for future works aimed at detecting other particles composing the atmosphere, such as polluting agents that can be detected from the ceilometer backscatter data.
Alessio Barbaro Chisari, Alessandro Ortis, Luca Guarnera, Wladimiro Carlo Patatu, Rosaria Ausilia Giandolfo, Emanuele Spampinato, Sebastiano Battiato, Mario Valerio Giuffrida
ICIP7
2024 On the Exploitation of DCT-Traces in the Generative-AI Domain
abstract
Deepfakes represent one of the toughest challenges in the world of Cybersecurity and Digital Forensics, especially considering the high-quality results obtained with recent generative AI-based solutions. Almost all generative models leave unique traces in synthetic data that, if analyzed and identified in detail, can be exploited to improve the generalization limitations of existing deepfake detectors. In this paper we analyzed deepfake images in the frequency domain generated by both GAN and Diffusion Model engines, examining in detail the underlying statistical distribution of Discrete Cosine Transform (DCT) coefficients. Recognizing that not all coefficients contribute equally to image detection, we hypothesize the existence of a unique “discriminative fingerprint”, embedded in specific combinations of coefficients. To identify them, Machine Learning classifiers were trained on various combinations of coefficients. In addition, the Explainable AI (XAI) LIME algorithm was used to search for intrinsic discriminative combinations of coefficients. Finally, we performed a robustness test to analyze the persistence of traces by applying JPEG compression. The experimental results reveal the existence of traces left by the generative models that are more discriminative and persistent at JPEG attacks. Code and dataset are available at github/opontorno/dcts_analysis_deepfakes.
Orazio Pontorno, Luca Guarnera, Sebastiano Battiato
ICIP3
2024 DeepFeatureX Net: Deep Features eXtractors Based Network for Discriminating Synthetic from Real Images
Orazio Pontorno, Luca Guarnera, Sebastiano Battiato
ICPR (21)3
2024 ICPR 2024 Competition on Multiple Sclerosis Lesion Segmentation - Methods and Results
Alessia Rondinella, Francesco Guarnera, Elena Crispino, Giulia Russo, Clara Di Lorenzo, Davide Maimone, Francesco Pappalardo 0001, Sebastiano Battiato
ICPR (34)8
2024 TADM: Temporally-Aware Diffusion Model for Neurodegenerative Progression on Brain MRI
Mattia Litrico, Francesco Guarnera, Mario Valerio Giuffrida, Daniele Ravì, Sebastiano Battiato
MICCAI (2)5
2024 Mastering Deepfake Detection: A Cutting-edge Approach to Distinguish GAN and Diffusion-model Images
abstract
Detecting and recognizing deepfakes is a pressing issue in the digital age. In this study, we first collected a dataset of pristine images and fake ones properly generated by nine different Generative Adversarial Network (GAN) architectures and four Diffusion Models (DM). The dataset contained a total of 83,000 images, with equal distribution between the real and deepfake data. Then, to address different deepfake detection and recognition tasks, we proposed a hierarchical multi-level approach. At the first level, we classified real images from AI-generated ones. At the second level, we distinguished between images generated by GANs and DMs. At the third level (composed of two additional sub-levels), we recognized the specific GAN and DM architectures used to generate the synthetic data. Experimental results demonstrated that our approach achieved more than 97% classification accuracy, outperforming existing state-of-the-art methods. The models obtained in the different levels turn out to be robust to various attacks such as JPEG compression (with different quality factor values) and resize (and others), demonstrating that the framework can be used and applied in real-world contexts (such as the analysis of multimedia data shared in the various social platforms) for support even in forensic investigations to counter the illicit use of these powerful and modern generative models. We are able to identify the specific GAN and DM architecture used to generate the image, which is critical in tracking down the source of the deepfake. Our hierarchical multi-level approach to deepfake detection and recognition shows promising results in identifying deepfakes allowing focus on underlying task by improving (about 2% on the average) standard multiclass flat detection systems. The proposed method has the potential to enhance the performance of deepfake detection systems, aid in the fight against the spread of fake images, and safeguard the authenticity of digital media.
Luca Guarnera, Oliver Giudice, Sebastiano Battiato
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Enhancing Multiple Sclerosis Lesion Segmentation in Multimodal MRI Scans with Diffusion Models
abstract
Accurate segmentation of Multiple Sclerosis (MS) lesions from Magnetic Resonance Imaging (MRI) scans is crucial for clinical diagnosis and effective treatment planning. In this work, we investigate the effectiveness of Diffusion Models (DM) in achieving pixel-wise segmentation of MS lesions. DM significantly improves segmentation sensitivity, especially in regions with subtle abnormalities. We conducted extensive experiments using the magnetic resonance volumes from a public dataset, encompassing various imaging modalities. Our analysis demonstrated how DM can achieve performance levels that are on par with state-of-the-art techniques, as evidenced by a mean Dice coefficient comparable to the best existing methods. Furthermore, some variants of standard DM exhibits robustness across various imaging modalities, showcasing its versatility in clinical settings.
Alessia Rondinella, Francesco Guarnera, Oliver Giudice, Alessandro Ortis, Giulia Russo, Elena Crispino, Francesco Pappalardo 0001, Sebastiano Battiato
BIBM8
2023 Assessing forensic ballistics three-dimensionally through graphical reconstruction and immersive VR observation
abstract
Abstract A crime scene can provide valuable evidence critical to explain reason and modality of the occurred crime, and it can also lead to the arrest of criminals. The type of evidence collected by crime scene investigators or by law enforcement may accordingly effective involved cases. Bullets and cartridge cases examination is of paramount importance in forensic science because they may contain traces of microscopic striations, impressions and markings, which are unique and reproducible as “ballistic fingerprints”. The analysis of bullets and cartridge cases is a complicated and challenging process, typically based on optical comparison, leading to the identification of the employed firearm. New methods have recently been proposed for more accurate comparisons, which rely on three-dimensionally reconstructed data. This paper aims at further advancing recent methods by introducing a novel immersive technique for ballistics comparison by means of Virtual Reality. Users can three-dimensionally examine the cartridge cases shapes through intuitive natural gestures, from any vantage viewpoint (including internal iper-magnified views), while having at their disposal sets of visual aids which could not be easily implemented in desktop-based applications. A user study was conducted to assess viability and performance of our solution, which involved fourteen individuals acquainted with the standard procedures used by law enforcement agencies. Results clearly indicated that our approach lead to faster adaptation of users to the UI/UX and more accurate and explainable ballistics examination results.
Luca Guarnera, Oliver Giudice, Salvatore Livatino, Antonino Barbaro Paratore, Angelo Salici, Sebastiano Battiato
Multim. Tools Appl.6
2022 A Robust Misalignment Estimation Approach in Non-Aligned Double JPEG Compression Scenario
abstract
The estimation of the misalignment between consecutive JPEG compressions is a really important task that can be useful in forensics investigation for first quantization matrix estimation and forgery localization. Based on the analysis of DCT histograms obtained applying a third compression, an effective shift estimation solution has been designed. More-over, to increase the overall robustness in challenging conditions (e.g., small patches) several fusion strategies combining all the DCT coefficient information have been investigated. Finally, the effectiveness of the proposed solution has been demonstrated considering several scenarios (i.e., different patch sizes and quantization matrices) and comparisons with state-of-the-art solutions.
Giovanni Puglisi, Sebastiano Battiato
ICIP2
2022 Semantic food segmentation for health monitoring
abstract
This paper presents semantic food segmentation to detect individual food items in an image. The presented approach has been developed in the context of the FoodRec project, which aims to study and develop an automatic framework to track and monitor the dietary habits of people, during their smoke quitting protocol. The goal of food segmentation is to train a model that can look at the images of food items and infer semantic information to recognize individual food items present in an image. In this contribution, we propose a novel Convolutional Deconvolutional Pyramid Network for food segmentation to understand the semantic information of an image at a pixel level. This network employs convolution and deconvolution layers to build a feature pyramid and achieves high-level semantic feature map representation. As a consequence, the novel semantic segmentation network generates a dense and precise segmentation map of the input food image. Furthermore, the proposed method demonstrated significant improvements on a well-known public benchmark dataset.
Mazhar Hussain, Alessandro Ortis, Riccardo Polosa, Sebastiano Battiato
ICMV4
2022 Natural Gas Leakage Detection: a Deep Learning Framework on IR Video Data
abstract
Undetected gas leakages may result in serious fire and explosion accidents with consequences like injuries among workers and financial losses. Automated leak detectors aimed to catch in time the gas emissions could reduce the incident risks. Several monitoring techniques have been developed over the years, among them the Optical Gas Imaging (OGI) is a widely-used method but it typically requires manual analysis (slow and error-prone). This paper introduces an automated gas leakage detection framework exploiting Infrared video data. A novel Recurrent Neural Network architecture was designed and trained on an ad-hoc collected large-scale dataset. Experimental results demonstrated the effectiveness of the proposed framework outperforming the state-of-the-art approaches with an average accuracy of 98%. The robustness of the technique was also validated in different scenarios and with different camera settings.
Maria Ausilia Napoli Spatafora, Dario Allegra, Oliver Giudice, Filippo Stanco, Sebastiano Battiato
ICPR5
2022 CNN-based first quantization estimation of double compressed JPEG images
abstract
Multiple JPEG compressions leave artifacts in digital images: residual traces that could be exploited in forensics investigations to recover information about the device employed for acquisition or image editing software. In this paper, a novel First Quantization Estimation (FQE) algorithm based on convolutional neural networks (CNNs) is proposed. In particular, a solution based on an ensemble of CNNs was developed in conjunction with specific regularization strategies exploiting assumptions about neighboring element values of the quantization matrix to be inferred. Mostly designed to work in the aligned case, the solution was tested in challenging scenarios involving different input patch sizes, quantization matrices (both standard and custom) and datasets (i.e., RAISE and UCID collections). Comparisons with state-of-the-art solutions confirmed the effectiveness of the presented solution demonstrating for the first time to cover the widest combinations of parameters of double JPEG compressions.
Sebastiano Battiato, Oliver Giudice, Francesco Guarnera, Giovanni Puglisi
J. Vis. Commun. Image Represent.1
2021 Fine-Grained Image Classification for Pollen Grain Microscope Images
Francesca Trenta, Alessandro Ortis, Sebastiano Battiato
CAIP (1)3
2021 Advanced Deep Network with Attention and Genetic-Driven Reinforcement Learning Layer for an Efficient Cancer Treatment Outcome Prediction
abstract
In the last few years, medical researchers have investigated promising approaches for cancer treatment, leading to a major interest in the immunotherapeutic approach. The target of immunotherapy is to boost a subject’s immune system in order to fight cancer. However, scientific studies confirmed that not all patients have a positive response to immunotherapy treatment. Medical research has long been engaged in the search for predictive immunotherapeutic-response biomarkers. Based on these considerations, we developed a non-invasive advanced pipeline with a downstream 3D deep classifier with attention and reinforcement learning for early prediction of patients responsive to immunotherapeutic treatment from related chest-abdomen CT-scan imaging. We have tested the proposed pipeline within a clinical trial that recruited patients with metastatic bladder cancer. Our experiment results achieved accuracy close to 93%.
Francesco Rundo, Giuseppe L. Banna, Francesca Trenta, Sebastiano Battiato
ICIP4
2021 Advanced Densely Connected System with Embedded Spatio-Temporal Features Augmentation for Immunotherapy Assessment
abstract
In medical field, the term “immunotherapy” refers to a form of cancer treatment that uses the ability of body's immune system to prevent and destroy cancer cells. In the last few years, immunotherapy has demonstrated to be a very effective treatment in fighting cancer diseases. However, immunotherapy does not work for every patients and moreover, certain types of immunotherapy drugs could have side effects. With this regard, scientific researchers are investigating for effective ways to select the patients who are more likely to respond to the treatment. Hence, pre-clinical data confirmed that, sometimes, the composition of immune system cells infiltrating the tumor micro-environment may interfere with the efficacy of immunotherapy treatments. In this work, we developed a 3D Deep Network with a downstream classifier for selecting and properly augmenting features from chest-abdomen CT images toward improving cancer outcome prediction. In our work, we proposed an effective solution to a specific type of aggressive bladder cancer, called Metastatic Urothelial Carcinoma (mUC). Our experiment results achieved high accuracy confirming the effectiveness of the proposed pipeline.
Francesco Rundo, Giuseppe L. Banna, Francesca Trenta, Sebastiano Battiato
IJCNN4
2021 Estimating Previous Quantization Factors on Multiple JPEG Compressed Images
abstract
Abstract The JPEG compression algorithm has proven to be efficient in saving storage and preserving image quality thus becoming extremely popular. On the other hand, the overall process leaves traces into encoded signals which are typically exploited for forensic purposes: for instance, the compression parameters of the acquisition device (or editing software) could be inferred. To this aim, in this paper a novel technique to estimate “previous” JPEG quantization factors on images compressed multiple times, in the aligned case by analyzing statistical traces hidden on Discrete Cosine Transform (DCT) histograms is exploited. Experimental results on double, triple and quadruple compressed images, demonstrate the effectiveness of the proposed technique while unveiling further interesting insights.
Sebastiano Battiato, Oliver Giudice, Francesco Guarnera, Giovanni Puglisi
EURASIP J. Inf. Secur.1
2021 Exploiting objective text description of images for visual sentiment analysis
Alessandro Ortis, Giovanni Maria Farinella, Giovanni Torrisi, Sebastiano Battiato
Multim. Tools Appl.4
2021 EgoCart: A Benchmark Dataset for Large-Scale Indoor Image-Based Localization in Retail Stores
abstract
We consider the task of localizing shopping carts in a retail store from egocentric images. Addressing this task allows to infer information on the behavior of the customers to understand how they move in the store and what they pay more attention to. To study the problem, we propose a large dataset of images collected in a real retail store. The dataset comprises 19, 531 RGB images along with depth maps, ground truth camera poses, as well as class labels specifying the areas of the store in which each image has been acquired. We release the dataset to the public to encourage research in large-scale image-based indoor localization and to address the scarcity of large datasets to tackle the problem. We hence perform a benchmark of several image-based localization techniques exploiting images and depth information on the proposed dataset. In our study, both localization performances and space/time requirements are compared. The results show that, while state-of-the-art approaches allow to achieve good results, there is space for improvement.
Emiliano Spera, Antonino Furnari, Sebastiano Battiato, Giovanni Maria Farinella
IEEE Trans. Circuits Syst. Video Technol.3
2020 Visual Saliency Detection guided by Neural Signals
abstract
Saliency detection is a fundamental process of human visual perception, since it allows us to identify the most important parts of a scene, directing our analysis and interpretation capabilities on a reduced set of information and reducing reaction times. However, current approaches for automatic saliency detection either attempt to mimic human capabilities by building attention maps from hand-crafted feature analysis, or employ convolutional neural networks trained as black boxes, without any architectural or information prior from human biology. In this paper, we present an approach for saliency detection that combines the success of deep learning in identifying representations for visual data with a training paradigm aimed at matching neural activity provided directly by brain signals recorded while subjects look at images. We show that our approach is able to capture correspondences between visual elements and neural activities, successfully generalizing to unseen images to identify their most salient regions.
Simone Palazzo, Francesco Rundo, Sebastiano Battiato, Daniela Giordano, Concetto Spampinato
FG3
2020 POLLEN13K: A Large Scale Microscope Pollen Grain Image Dataset
abstract
Pollen grain classification has a remarkable role in many fields from medicine to biology and agronomy. Indeed, automatic pollen grain classification is an important task for all related applications and areas. This work presents the first large-scale pollen grain image dataset, including more than 13 thousands objects. After an introduction to the problem of pollen grain classification and its motivations, the paper focuses on the employed data acquisition steps, which include aerobiological sampling, microscope image acquisition, object detection, segmentation and labelling. Furthermore, a baseline experimental assessment for the task of pollen classification on the built dataset, together with discussion on the achieved results, is presented.
Sebastiano Battiato, Alessandro Ortis, Francesca Trenta, Lorenzo Ascari, Mara Politi, Consolata Siniscalco
ICIP1
2020 Animated Gif Optimization By Adaptive Color Local Table Management
abstract
After thirty years of the GIF file format, today is becoming more popular than ever: being a great way of communication for friends and communities on Instant Messengers and Social Networks. While being so popular, the original compression method to encode GIF images have not changed a bit. On the other hand popularity means that storage saving becomes an issue for hosting platforms. In this paper a parametric optimization technique for animated GIFs will be presented. The proposed technique is based on Local Color Table selection and color remapping in order to create optimized animated GIFs while preserving the original format. The technique achieves good results in terms of byte reduction with limited or no loss of perceived color quality. Tests carried out on 1000 GIF files demonstrate the effectiveness of the proposed optimization strategy.
Oliver Giudice, Dario Allegra, Francesco Guarnera, Filippo Stanco, Sebastiano Battiato
ICIP5
2020 Single Architecture and Multiple task deep Neural Network for Altered Fingerprint Analysis
abstract
Fingerprints are one of the most copious evidence in a crime scene and, for this reason, they are frequently used by law enforcement for identification of individuals. But fingerprints can be altered. “Altered fingerprints” refers to intentionally damage of the friction ridge pattern and they are often used by smart criminals in hope to evade law enforcement. We use a deep neural network approach training an Inception-v3 architecture. This paper proposes a method for detection of altered fingerprints, identification of types of alterations and recognition of gender, hand and fingers. We also produce activation maps that show which part of a fingerprint the neural network has focused on, in order to detect where alterations are positioned. The proposed approach achieves an accuracy of 98.21%, 98.46%, 92.52%, 97.53% and 92,18% for the classification of fakeness, alterations, gender, hand and fingers, respectively on the SO.CO.FING. dataset.
Oliver Giudice, Mattia Litrico, Sebastiano Battiato
ICIP3
2020 Computational Data Analysis for First Quantization Estimation on JPEG Double Compressed Images
abstract
Multimedia Forensics experts work consists in providing answers about integrity of a specific media content and from where it comes from. Exploitation of any traces from JPEG double compressed images is often one of the main investigative path to be used for these purposes. Thus it is fundamental to have tools and algorithms able to safely estimate the first quantization matrix to further proceed with camera model identification and related tasks. In this paper, a technique based on extensive simulation is proposed, with the aim to infer the first quantization for a certain numbers of Discrete Cosine Transform (DCT) coefficients exploiting local image statistics without using any a-priori knowledge. The method provides also a reliable confidence value for the estimation which is of great importance for forensic purposes. Experimental results w.r.t. the state-of-the-art demonstrate the effectiveness of the proposed technique both in terms of precision and overall reliability.
Sebastiano Battiato, Oliver Giudice, Francesco Guarnera, Giovanni Puglisi
ICPR1
2020 Anticipating Activity from Multimodal Signals
abstract
Images, videos, audio signals, sensor data, can be easily collected in huge quantity by different devices and processed in order to emulate the human capability of elaborating a variety of different stimuli. Are multimodal signals useful to understand and anticipate human actions if acquired from the user viewpoint? This paper proposes to build an embedding space where inputs of different nature, but semantically correlated, are projected in a new representation space and properly exploited to anticipate the future user activity. To this purpose, we built a new multimodal dataset comprising video, audio, tri-axial acceleration, angular velocity, tri-axial magnetic field, pressure and temperature. To benchmark the proposed multimodal anticipation challenge, we consider classic classifiers on top of deep learning methods used to build the embedding space representing multimodal signals. The achieved results show that the exploitation of different modalities is useful to improve the anticipation of the future activity.
Tiziana Rotondo, Giovanni Maria Farinella, Davide Giacalone, Sebastiano Mauro Strano, Valeria Tomaselli, Sebastiano Battiato
ICPR6
2020 Survey on visual sentiment analysis
abstract
Visual Sentiment Analysis aims to understand how images affect people, in terms of evoked emotions. Although this field is rather new, a broad range of techniques have been developed for various data sources and problems, resulting in a large body of research. This paper reviews pertinent publications and tries to present an exhaustive overview of the field. After a description of the task and the related applications, the subject is tackled under different main headings. The paper also describes principles of design of general Visual Sentiment Analysis systems from three main points of view: emotional models, dataset definition, feature design. A formalization of the problem is discussed, considering different levels of granularity, as well as the components that can affect the sentiment toward an image in different ways. To this aim, this paper considers a structured formalization of the problem which is usually used for the analysis of text, and discusses it's suitability in the context of Visual Sentiment Analysis. The paper also includes a description of new challenges, the evaluation from the viewpoint of progress toward more sophisticated systems and related practical applications, as well as a summary of the insights resulting from this study.
Alessandro Ortis, Giovanni Maria Farinella, Sebastiano Battiato
IET Image Process.3
2020 SceneAdapt: Scene-based domain adaptation for semantic segmentation using adversarial learning
Daniele Di Mauro, Antonino Furnari, Giuseppe Patanè 0002, Sebastiano Battiato, Giovanni Maria Farinella
Pattern Recognit. Lett.4
2020 EGO-CH: Dataset and fundamental tasks for visitors behavioral understanding using egocentric vision
Francesco Ragusa, Antonino Furnari, Sebastiano Battiato, Giovanni Signorello, Giovanni Maria Farinella
Pattern Recognit. Lett.3
2019 Advanced Motion-Tracking System with Multi-Layers Deep Learning Framework for Innovative Car-Driver Drowsiness Monitoring
abstract
Recently, the ability to monitor driver drowsiness has attracted a great deal of attention in the automotive industry, in order to prevent the risk due to an inadequate driver psycho-physical state. Specifically, the research effort has focused on the study of the physiological signals to assess the attention level. The main idea consists in verifying the drowsiness level through analyzing the Heart Rate Variability (HRV). The HRV allows to understand the activity of the autonomic nervous system that regulates a series of unconscious and involuntary activities (e.g. the heartbeat, the blood pressure). The HRV is traditionally obtained from electrocardiography (ECG) even though the photoplethysmography (PPG) signal has been proposed as valid alternative to ECG in order to overcome some limitations derived from it. For the above reasons, we analyzed the skin micro-movements and changes in facial color due to blood circulation quite indistinguishable with naked eye in order to extract facial landmarks and to reconstruct PPG signal. The results we obtained by validation confirmed the correlation between the PPG signal detected by sensors and the reconstructed PPG signal from facial landmarks.
Francesca Trenta, Sabrina Conoci, Francesco Rundo, Sebastiano Battiato
FG4
2019 Siamese Ballistics Neural Network
abstract
Firearm identification is crucial in many investigative scenario. The crime scene often contains traces left by firearms in terms of bullets and cartridges. Traces analysis is a fundamental step in the Forensics Ballistics Analysis Process to identify which firearm fired a specific cartridge. In this paper we present a fully automated technique to compare cartridges represented as a set of 3D point-clouds. The overall approach is based on Siamese Neural Network learning paradigm that we use to build a suitable embedding space where the 3D point-cloud of the cartridges are compared. The proposed approach has been assessed by considering the NBTRD dataset. Obtained results support the exploitation of the proposed technique in ballistic analysis.
Oliver Giudice, Luca Guarnera, Antonino Barbaro Paratore, Giovanni Maria Farinella, Sebastiano Battiato
ICIP5
2019 A New Study On Wood Fibers Textures: Documents Authentication Through LBP Fingerprint
abstract
The authentication of printed material based on textures is a critical and challenging problem for many security agencies in many contexts: valuable documents, banknotes, tickets or rare collectible cards are often targets for forgery. This motivates the study of low-cost, fast and reliable approaches for documents authenticity analysis. In this paper, we present a new approach based on the extraction of translucent patterns from paper sheet by means of a specific-built framework. A fingerprint is obtained by computing a Local Binary Pattern descriptor on the digital image. To validate the robustness of the proposed method for authentication analysis, we introduce a novel dataset and perform retrieval tests under both, ideal and noisy conditions. Experimental results prove the validity of the proposed strategy.
Francesco Guarnera, Dario Allegra, Oliver Giudice, Filippo Stanco, Sebastiano Battiato
ICIP5
2019 A New Framework for Studying Tubes Rearrangement Strategies in Surveillance Video Synopsis
abstract
The manual review of raw surveillance video is a time consuming task which can be optimized by using a Video Synopsis (VS) algorithm. The aim of such approaches is to condense a long video into shorter one to allow a quicker review of surveillance data. However, VS is a complex problem. A typical object-based VS algorithm requires three main modules to perform the following tasks: object detection and tracking, tubes rearrangement, condensed video generation. Although the aforementioned three steps are equally critical, we realized that the core of Video Synopsis lies in the tubes rearrangement. This led us to propose an original approach to tackle the problem of tubes rearrangement. To this aim, we first introduce a new toolbox to generate a proper testing dataset, which allows to bypass the lack of public databases including proper annotated videos for testing synopsis approaches. Additionally, we propose an improvement of a tubes arrangement algorithm based on graph colouring and we prove its validity on our generated dataset. For a proper comparison, we show that our algorithm also outperforms the original one on UA-DETRAC public dataset.
Giovanna Pappalardo, Dario Allegra, Filippo Stanco, Sebastiano Battiato
ICIP4
2019 Estimating the occupancy status of parking areas by counting cars and non-empty stalls
Daniele Di Mauro, Antonino Furnari, Giuseppe Patanè 0002, Sebastiano Battiato, Giovanni Maria Farinella
J. Vis. Commun. Image Represent.4
2019 Egocentric visitors localization in natural sites
Filippo L. M. Milotta, Antonino Furnari, Sebastiano Battiato, Giovanni Signorello, Giovanni Maria Farinella
J. Vis. Commun. Image Represent.3
2018 Scene Adaptation for Semantic Segmentation using Adversarial Learning
abstract
Semantic Segmentation algorithms based on the deep learning paradigm have reached outstanding performances. However, in order to achieve good results in a new domain, it is generally demanded to fine-tune a pre-trained deep architecture using new labeled data coming from the target application domain. The fine-tuning procedure is also required when the domain application settings change, e. g., when a camera is moved, or a new camera is installed. This implies the collection and pixel-wise la-beling of images to be used for training, which slows down the deployment of semantic segmentation systems in real industrial scenarios and increases the industrial costs. Taking into account the aforementioned issues, in this paper we propose an approach based on Adversarial Learning to perform scene adaptation for semantic segmentation. We frame scene adaptation as the task of predicting semantic segmentation masks for images belonging to a Target Scene Context given labeled images coming from a Source Scene Context and unlabeled images coming from the Target Scene Context. Experiments highlight that the proposed method achieves promising performances both when the two scenes contain similar content (i.e., they are related to two different points of view of the same scene) and when the observed scenes contain unrelated content (i.e., they account to completely different scenes).
Daniele Di Mauro, Antonino Furnari, Giuseppe Patanè 0002, Sebastiano Battiato, Giovanni Maria Farinella
AVSS4
2018 Visual Sentiment Analysis Based on on Objective Text Description of Images
abstract
Visual Sentiment Analysis aims to estimate the polarity of the sentiment evoked by images in terms of positive or negative sentiment. To this aim, most of the state of the art works exploit the text associated to a social post provided by the user. However, such textual data is typically noisy due to the subjectivity of the user which usually includes text useful to maximize the diffusion of the social post. In this paper we extract and employ an Objective Text description of images automatically extracted from the visual content rather than the classic Subjective Text provided by the users. The proposed method defines a multimodal embedding space based on the contribute of both visual and textual features. The sentiment polarity is then inferred by a supervised Support Vector Machine trained on the representations of the obtained embedding space. Experiments performed on a representative dataset of 47235 labelled samples demonstrate that the exploitation of the proposed Objective Text helps to outperform state-of-the-art for sentiment polarity estimation.
Alessandro Ortis, Giovanni Maria Farinella, Giovanni Torrisi, Sebastiano Battiato
CBMI4
2018 A Fast Palette Reordering Technique Based on GPU-Optimized Genetic Algorithms
abstract
Color re-indexing is one of main approaches for improving the loss-less compression of color indexed images. Zero-order entropy reduction of indexes matrix is the key to obtain high compression ratio. However, obtaining the optimal re-indexed palette is a challenging problem that cannot be solved by brute-force approaches. In this paper we propose a novel re-indexing approach where the Travelling Salesman Problem is solved through Ant Colony Optimization. Our method is proved to achieve high quality results by outperforming state-of-art ones in term of compression gain. Additionally, we exploit clustering and GPU computing to make our solution extremely fast.
Oliver Giudice, Dario Allegra, Filippo Stanco, Giorgio Mario Grasso, Sebastiano Battiato
ICIP5
2018 Egocentric Shopping Cart Localization
abstract
This work investigates the new problem of image-based egocentric shopping cart localization in retail stores. The contribution of our work is two-fold. First, we propose a novel large-scale dataset for image-based egocentric shopping cart localization. The dataset has been collected using cameras placed on shopping carts in a large retail store. It contains a total of 19,531 image frames, each labelled with its six Degrees Of Freedom pose. We study the localization problem by analysing how cart locations should be represented and estimated, and how to assess the localization results. Second, we benchmark different image-based techniques to address the task. Specifically, we investigate two families of algorithms: classic methods based on image retrieval and emerging methods based on regression. Experimental results show that methods based on image retrieval largely outperform regression-based approaches. We also point out that deep metric learning techniques allow to learn better visual representations w.r.t. other architectures, and are useful to improve the localization results of both retrieval-based and regression-based approaches. Our findings suggest that deep metric learning techniques can help bridge the gap between retrieval-based and regression-based methods.
Emiliano Spera, Antonino Furnari, Sebastiano Battiato, Giovanni Maria Farinella
ICPR3
2018 Evaluation of Levenberg-Marquardt neural networks and stacked autoencoders clustering for skin lesion analysis, screening and follow-up
abstract
Traditional methods for early detection of melanoma rely on the visual analysis of the skin lesions performed by a dermatologist. The analysis is based on the so‐called ABCDE (Asymmetry, Border irregularity, Colour variegation, Diameter, Evolution) criteria, although confirmation is obtained through biopsy performed by a pathologist. The proposed method exploits an automatic pipeline based on morphological analysis and evaluation of skin lesion dermoscopy images. Preliminary segmentation and pre‐processing of dermoscopy image by SC‐cellular neural networks is performed, in order to obtain ad‐hoc grey‐level skin lesion image that is further exploited to extract analytic innovative hand‐crafted image features for oncological risks assessment. In the end, a pre‐trained Levenberg–Marquardt neural network is used to perform ad‐hoc clustering of such features in order to achieve an efficient nevus discrimination (benign against melanoma), as well as a numerical array to be used for follow‐up rate definition and assessment. Moreover, the authors further evaluated a combination of stacked autoencoders in lieu of the Levenberg–Marquardt neural network for the clustering step.
Francesco Rundo, Sabrina Conoci, Giuseppe L. Banna, Alessandro Ortis, Filippo Stanco, Sebastiano Battiato
IET Comput. Vis.6
2018 Preface
Sebastiano Battiato, Patrizia Daniele, Giovanni Maria Farinella, Sofia Giuffrè, Laura Scrimali 0001
J. Glob. Optim.1
2018 Personal-location-based temporal segmentation of egocentric videos for lifelogging applications
Antonino Furnari, Sebastiano Battiato, Giovanni Maria Farinella
J. Vis. Commun. Image Represent.2
2018 Market basket analysis from egocentric videos
Vito Santarcangelo, Giovanni Maria Farinella, Antonino Furnari, Sebastiano Battiato
Pattern Recognit. Lett.4
2017 Park Smart
abstract
The paper presents Park Smart, a solution which aim is to solve the pain of finding a free parking space in public and private areas (e.g. cities, malls, etc.), and hence to optimize parking stalls allocation as well as to increase revenues for the companies which manage them. The proposed solution exploits cutting edge technologies such as IoT, Cloud Computing and Deep Learning.
Daniele Di Mauro, Marco Moltisanti, Giuseppe Patanè 0002, Sebastiano Battiato, Giovanni Maria Farinella
AVSS4
2017 Next-active-object prediction from egocentric videos
Antonino Furnari, Sebastiano Battiato, Kristen Grauman, Giovanni Maria Farinella
J. Vis. Commun. Image Represent.2
2017 Distortion adaptive Sobel filters for the gradient estimation of wide angle images
Antonino Furnari, Giovanni Maria Farinella, Arcangelo Bruna, Sebastiano Battiato
J. Vis. Commun. Image Represent.4
2017 Organizing egocentric videos of daily living activities
Alessandro Ortis, Giovanni Maria Farinella, Valeria D'Amico, Luca Addesso, Giovanni Torrisi, Sebastiano Battiato
Pattern Recognit.6
2017 Recognizing Personal Locations From Egocentric Videos
abstract
Contextual awareness in wearable computing allows for construction of intelligent systems, which are able to interact with the user in a more natural way. In this paper, we study how personal locations arising from the user's daily activities can be recognized from egocentric videos. We assume that few training samples are available for learning purposes. Considering the diversity of the devices available on the market, we introduce a benchmark dataset containing egocentric videos of eight personal locations acquired by a user with four different wearable cameras. To make our analysis useful in real-world scenarios, we propose a method to reject negative locations, i.e., those not belonging to any of the categories of interest for the end-user. We assess the performances of the main state-of-the-art representations for scene and object classification on the considered task, as well as the influence of device-specific factors such as the field of view and the wearing modality. Concerning the different device-specific factors, experiments revealed that the best results are obtained using a head-mounted wide-angular device. Our analysis shows the effectiveness of using representations based on convolutional neural networks, employing basic transfer learning techniques and an entropy-based rejection algorithm.
Antonino Furnari, Giovanni Maria Farinella, Sebastiano Battiato
IEEE Trans. Hum. Mach. Syst.3
2017 Affine Covariant Features for Fisheye Distortion Local Modeling
abstract
Perspective cameras are the most popular imaging sensors used in computer vision. However, many application fields, including automotive, surveillance, and robotics, require the use of wide angle cameras (e.g., fisheye), which allow to acquire a larger portion of the scene using a single device at the cost of the introduction of noticeable radial distortion in the images. Affine covariant feature detectors have proved successful in a variety of computer vision applications, including object recognition, image registration, and visual search. Moreover, their robustness to a series of variabilities related to both the scene and the image acquisition process has been thoroughly studied in the literature. In this paper, we investigate their effectiveness on fisheye images providing both theoretical and experimental analyses. As theoretical outcome, we show that the inherently non-linear radial distortion can be locally approximated by linear functions with a reasonably small error. The experimental analysis builds on Mikolajczyk's benchmark to assess the robustness of three popular affine region detectors (i.e., maximally stable extremal regions, and Harris and Hessian affine region detectors), with respect to different variabilities as well as to radial distortion. To support the evaluations, we rely on the Oxford data set and introduce a novel benchmark data set comprising 50 images depicting different scene categories. Experiments are carried out on rectilinear images to which radial distortion is artificially added, and on real-world images acquired using fisheye lenses. Our analysis points out that affine region detectors can be effectively employed directly on fisheye images and that the radial distortion is locally modeled as an additional affine variability.
Antonino Furnari, Giovanni Maria Farinella, Arcangelo Bruna, Sebastiano Battiato
IEEE Trans. Image Process.4
2016 Learning Approaches for Parking Lots Classification
Daniele Di Mauro, Sebastiano Battiato, Giuseppe Patanè 0002, Marco Leotta, Daniele Maio, Giovanni Maria Farinella
ACIVS2
2016 The Social Picture
abstract
We present The Social Picture, a framework to collect and explore huge amount of crowdsourced social images about public events, cultural heritage sites and other customized private events.The Social Picture aims to create social communities of users that contribute to the creation of image collections about common interests. The collections can be explored through a number of advanced Computer Vision and Machine Learning algorithms, able to capture the visual content of images in order to organize them in a semantic way. The interfaces of The Social Picture allow the users to create customized collections by exploiting semantic filters based on visual features, social network tags, geolocation, and other information related to the images.
Sebastiano Battiato, Giovanni Maria Farinella, Filippo L. M. Milotta, Alessandro Ortis, Luca Addesso, Antonino Casella, Valeria D'Amico, Giovanni Torrisi
ICMR1
2016 Aligning shapes for symbol classification and retrieval
Sebastiano Battiato, Giovanni Maria Farinella, Oliver Giudice, Giovanni Puglisi
Multim. Tools Appl.1
2016 Semantic segmentation of images exploiting DCT based features and random forest
Daniele Ravì, M. Bober, Giovanni Maria Farinella, Mirko Guarnera, Sebastiano Battiato
Pattern Recognit.5
2016 Special issue on "Video analytics for audience measurement in retail and digital signage"
Sebastiano Battiato, Andrea Cavallaro, Cosimo Distante
Pattern Recognit. Lett.1
2015 On Blind Source Camera Identification
Giovanni Maria Farinella, Mario Valerio Giuffrida, V. Digiacomo, Sebastiano Battiato
ACIVS4
2015 Fast and Low Power Consumption Outliers Removal for Motion Vector Estimation
Giuseppe Spampinato, Arcangelo Bruna, Giovanni Maria Farinella, Sebastiano Battiato, Giovanni Puglisi
ACIVS4
2015 Generalized Sobel Filters for gradient estimation of distorted images
abstract
In this paper we tackle the problem of correctly estimating the gradient of distorted images. The proper estimation of the gradient in the presence of distortion is of great interest due to the large number of applications relying on wide angle cameras (e.g., in surveillance, automotive, robotics). To this aim we propose the Generalized Sobel Filters (GSF), a family of adaptive Sobel filters able to correctly estimate the gradient of distorted images. To assess the performances of the proposed method, we acquired a benchmark dataset of high resolution images belonging to different categories which are relevant to application domains where the gradient estimation is usually employed. We build an objective evaluation pipeline and perform experiments which show that our method outperforms the state-of-the-art.
Antonino Furnari, Giovanni Maria Farinella, Arcangelo Bruna, Sebastiano Battiato
ICIP4
2015 RECfusion: Automatic Video Curation Driven by Visual Content Popularity
abstract
The proliferation of mobile devices and the diffusion of social media have changed the communication paradigm of people that share multimedia data by allowing new interaction models (e.g., social networks). In social events (e.g., concerts), the automatic video understanding goal includes the interpretation of which visual contents are the most popular. The popularity of a visual content depends on how many people are looking at that scene, and therefore it could be obtained through the "visual consensus" among multiple video streams acquired by the different users devices. In this work we present RECfusion, a system able to automatically create a single video from multiple video sources by taking into account the popularity of the acquired scenes. The frames composing the final popular video are selected from the different video streams by considering those visual scenes which are pointed and recorded by the highest number of users' devices. Results on two benchmark datasets confirm the effectiveness of the proposed system.
Alessandro Ortis, Giovanni Maria Farinella, Valeria D'Amico, Luca Addesso, Giovanni Torrisi, Sebastiano Battiato
ACM Multimedia6
2015 An integrated system for vehicle tracking and classification
Sebastiano Battiato, Giovanni Maria Farinella, Antonino Furnari, Giovanni Puglisi, Anique Snijders, Jelmer Spiekstra
Expert Syst. Appl.1
2015 Representing scenes for real-time context classification on mobile devices
Giovanni Maria Farinella, Daniele Ravì, Valeria Tomaselli, Mirko Guarnera, Sebastiano Battiato
Pattern Recognit.5
2014 Classifying food images represented as Bag of Textons
abstract
The classification of food images is an interesting and challenging problem since the high variability of the image content which makes the task difficult for current state-of-the-art classification methods. The image representation to be employed in the classification engine plays an important role. We believe that texture features have been not properly considered in this application domain. This paper points out, through a set of experiments, that textures are fundamental to properly recognize different food items. For this purpose the bag of visual words model (BoW) is employed. Images are processed with a bank of rotation and scale invariant filters and then a small codebook of Textons is built for each food class. The learned class-based Textons are hence collected in a single visual dictionary. The food images are represented as visual words distributions (Bag of Textons) and a Support Vector Machine is used for the classification stage. The experiments demonstrate that the image representation based on Bag of Textons is more accurate than existing (and more complex) approaches in classifying the 61 classes of the Pittsburgh Fast-Food Image Dataset.
Giovanni Maria Farinella, Marco Moltisanti, Sebastiano Battiato
ICIP3
2014 Affine region detectors on the fisheye domain
abstract
Feature extractors play an important role in different Computer Vision application domains such as registration, recognition and visual search. Different detectors have been proposed and evaluated so far assuming images taken with classic cameras. However, many operating cameras (e.g., in surveillance and automotive) are built considering a fisheye model and a preprocessing step is performed to remove the distortion of the images before running a detector. The following question arises: are the current detectors suitable to work directly in the fisheye domain? To answer this question, in this paper a benchmark dataset and objective evaluation measures are considered to evaluate the performances of the state-of-the-art detectors in the fisheye domain. Test images are properly generated starting from benchmark rectilinear images and considering different fisheye focal lengths. The experiments evaluate the performances of the detectors against both increasing fisheye distortion and the combination of the fisheye distortion with photometric and geometric variability of the image content. The experiments demonstrate that affine covariant detectors can be employed directly in the fisheye domain. Furthermore, although the transformation between the rectilinear and the fisheye coordinates is not affine, we show that the mapping can be locally approximated by linear functions with a small error.
Antonino Furnari, Giovanni Maria Farinella, Giovanni Puglisi, Arcangelo Bruna, Sebastiano Battiato
ICIP5
2014 Aligning codebooks for near duplicate image detection
Sebastiano Battiato, Giovanni Maria Farinella, Giovanni Puglisi, Daniele Ravì
Multim. Tools Appl.1
2014 First Quantization Matrix Estimation From Double Compressed JPEG Images
abstract
One of the most common problems in the image forensics field is the reconstruction of the history of an image or a video. The data related to the characteristics of the camera that carried out the shooting, together with the reconstruction of the (possible) further processing, allow us to have some useful hints about the originality of the visual document under analysis. For example, if an image has been subjected to more than one JPEG compression, we can state that the considered image is not the exact bitstream generated by the camera at the time of shooting. It is then useful to estimate the quantization steps of the first compression, which, in case of JPEG images edited and then saved again in the same format, are no more available in the embedded metadata. In this paper, we present a novel algorithm to achieve this goal in case of double JPEG compressed images. The proposed approach copes with the case when the second quantization step is lower than the first one, exploiting the effects of successive quantizations followed by dequantizations. To improve the results of the estimation, a proper filtering strategy together with a function devoted to find the first quantization step, have been designed. Experimental results and comparisons with the state-of-the-art methods, confirm the effectiveness of the proposed approach.
Fausto Galvan, Giovanni Puglisi, Arcangelo Bruna, Sebastiano Battiato
IEEE Trans. Inf. Forensics Secur.4
2014 Saliency-Based Selection of Gradient Vector Flow Paths for Content Aware Image Resizing
abstract
Content-aware image resizing techniques allow to take into account the visual content of images during the resizing process. The basic idea beyond these algorithms is the removal of vertical and/or horizontal paths of pixels (i.e., seams) containing low salient information. In this paper, we present a method which exploits the gradient vector flow (GVF) of the image to establish the paths to be considered during the resizing. The relevance of each GVF path is straightforward derived from an energy map related to the magnitude of the GVF associated to the image to be resized. To make more relevant, the visual content of the images during the content-aware resizing, we also propose to select the generated GVF paths based on their visual saliency properties. In this way, visually important image regions are better preserved in the final resized image. The proposed technique has been tested, both qualitatively and quantitatively, by considering a representative data set of 1000 images labeled with corresponding salient objects (i.e., ground-truth maps). Experimental results demonstrate that our method preserves crucial salient regions better than other state-of-the-art algorithms.
Sebastiano Battiato, Giovanni Maria Farinella, Giovanni Puglisi, Daniele Ravì
IEEE Trans. Image Process.1
2013 First JPEG quantization matrix estimation based on histogram analysis
abstract
To assess if a digital image has been (or not) doubly compressed is a challenging issue especially in forensics domain where could be fundamental clarify if, in addition to the compression at the time of shooting, the picture was decompressed (in some way) and then resaved. This is not a clear indication of forgery, but it guarantees that the image, probably, is not the original one. In this paper we propose a novel technique able to recover the coefficients of the first compression in a double compressed JPEG image under some assumptions. The proposed approach exploits how successive quantizations followed by dequantizations introduce some regularities (e.g., sequence of zero and not zero values) on the histograms of coefficient distributions that could be analyzed to recover the original compression parameters. Experimental results and comparisons with state of the art methods confirm the effectiveness of the proposed approach.
Giovanni Puglisi, Arcangelo Bruna, Fausto Galvan, Sebastiano Battiato
ICIP4
2012 Content-aware image resizing with seam selection based on Gradient Vector Flow
abstract
Content-aware image resizing is an effective technique that allows to take into account the visual content of images during the resizing process. The basic idea beyond these algorithms is the resizing of an image by considering vertical and/or horizontal paths of pixels (i.e., seams) which contain low salient information. In this paper we exploit the Gradient Vector Flow (GVF) of the image to establish the paths to be considered during the resizing. The relevance of each path is derived from a saliency map obtained by considering the magnitude of the GVF associated to the image under consideration. The proposed technique has been tested, both qualitatively and quantitatively, by considering a representative set of images labeled with corresponding salient objects (i.e., ground-truth maps). Experimental results demonstrate that our method preserves crucial salient regions better than other state-of-the-art algorithms.
Sebastiano Battiato, Giovanni Maria Farinella, Giovanni Puglisi, Daniele Ravì
ICIP1
2012 Aligning Bags of Shape Contexts for Blurred Shape Model based symbol classification
Sebastiano Battiato, Giovanni Maria Farinella, Oliver Giudice, Giovanni Puglisi
ICPR1
2012 Robust Image Alignment for Tampering Detection
abstract
The widespread use of classic and newest technologies available on Internet (e.g., emails, social networks, digital repositories) has induced a growing interest on systems able to protect the visual content against malicious manipulations that could be performed during their transmission. One of the main problems addressed in this context is the authentication of the image received in a communication. This task is usually performed by localizing the regions of the image which have been tampered. To this aim the aligned image should be first registered with the one at the sender by exploiting the information provided by a specific component of the forensic hash associated to the image. In this paper we propose a robust alignment method which makes use of an image hash component based on the Bag of Features paradigm. The proposed signature is attached to the image before transmission and then analyzed at destination to recover the geometric transformations which have been applied to the received image. The estimator is based on a voting procedure in the parameter space of the model used to recover the geometric transformation occurred into the manipulated image. The proposed image hash encodes the spatial distribution of the image features to deal with highly textured and contrasted tampering patterns. A block-wise tampering detection which exploits an histograms of oriented gradients representation is also proposed. A non-uniform quantization of the histogram of oriented gradient space is used to build the signature of each image block for tampering purposes. Experiments show that the proposed approach obtains good margin of performances with respect to state-of-the art methods.
Sebastiano Battiato, Giovanni Maria Farinella, Enrico Messina, Giovanni Puglisi
IEEE Trans. Inf. Forensics Secur.1
2011 Robust video stabilization approach based on a voting strategy
abstract
Today many people in the world without any (or with little) knowledge about video recording, thanks to the widespread use of mobile devices (PDAs, mobile phones, etc.) take videos. However the unwanted movements of their hands typically blur and introduce disturbing jerkiness in the recorded sequences. A fundamental issue is the overall robustness with respect to different scene contents (indoor, outdoor, etc.) and conditions (illumination changes, moving objects, etc.). In this paper we propose an accurate and robust image alignment algorithm for video stabilization purposes based on a voting strategy. Experimental results confirm the effectiveness of the proposed approach.
Giovanni Puglisi, Sebastiano Battiato
ICIP2
2011 Understanding geometric manipulations of images through bovw-based hashing
abstract
The increasing use of low cost imaging devices and the innovations in terms of media distribution technologies induce a growing interest on technologies able to protect digital visual media against malicious manipulations of the visual contents. One of the main problems addressed in this research area is the blind detection of traces of forgery on an image obtained through the internet. Specifically, in this paper we consider the context of communications, where malicious image manipulations should be detected by a receiver. In the proposed method, an image hash based on the Bag of Visual Words paradigm is attached as signature to the image before trans mission. The forensic hash is then analyzed at destination to detect the geometric transformations which have been applied to the received image. This task is fundamental for further processing which usually assumes that the received image is aligned with the original one, as in the case of tampering detection systems. Experiments show that the proposed approach outperforms state-of-the art methods by obtaining a good margin in terms of performances.
Sebastiano Battiato, Giovanni Maria Farinella, Enrico Messina, Giovanni Puglisi
ICME1
2011 Third ACM international workshop on multimedia in forensics and intelligence (MiFor 2011)
abstract
This paper introduces the context of the workshop and the associated papers.
Sebastiano Battiato, Sabu Emmanuel, Adrian Ulges, Marcel Worring
ACM Multimedia1
2011 Robust image registration and tampering localization exploiting bag of features based forensic signature
abstract
The distribution of digital images with the classic and newest technologies available on Internet (e.g., emails, social networks, digital repositories) has induced a growing interest on systems able to protect the visual content against malicious manipulations that could be performed during their transmission. One of the main problems addressed in this context is the authentication of the image received in a communication. This task is usually performed by localizing the regions of the image which have been tampered. To this aim the received image should be first registered with the one at the sender by exploiting the information provided by a specific component of the forensic hash associated with the image. In this paper we propose a robust alignment method which makes use of an image signature based on the Bag of Features paradigm. The alignment is based on a voting procedure in the parameter space of the model used to recover the geometric transformation occurred into the manipulated image. Experiments show that the proposed approach obtains good margin in terms of performances with respect to state-of-the art methods.
Sebastiano Battiato, Giovanni Maria Farinella, Enrico Messina, Giovanni Puglisi
ACM Multimedia1
2011 A Robust Image Alignment Algorithm for Video Stabilization Purposes
abstract
Today, many people in the world without any (or with little) knowledge about video recording, thanks to the widespread use of mobile devices (personal digital assistants, mobile phones, etc.), take videos. However, the unwanted movements of their hands typically blur and introduce disturbing jerkiness in the recorded sequences. Many video stabilization techniques have been hence developed with different performances but only fast strategies can be implemented on embedded devices. A fundamental issue is the overall robustness with respect to different scene contents (indoor, outdoor, etc.) and conditions (illumination changes, moving objects, etc.). In this paper, we propose a fast and robust image alignment algorithm for video stabilization purposes. Our contribution is twofold: a fast and accurate block-based local motion estimator together with a robust alignment algorithm based on voting. Experimental results confirm the effectiveness of both local and global motion estimators.
Giovanni Puglisi, Sebastiano Battiato
IEEE Trans. Circuits Syst. Video Technol.2
2010 Red-eyes removal through cluster based Linear Discriminant Analysis
abstract
Red-eye artifact is a well-known problem in digital photography. Since the large diffusion of mobile devices with embedded camera and flashgun, automatic detection and correction of red-eyes have become an important task. In this paper we describe a technique that makes use of three steps to identify and correct red-eyes. First, red-eye candidates are extracted from the input image by using simple color segmentation coupled with geometrical constraints. A set of linear discriminant classifiers is then learned on the clustered patches space, and hence employed to distinguish between eyes and non-eyes patches. The proposed cluster-based Linear Discriminant Analysis is used to deal with the multi-modally nature of the input space. The third step of the pipeline is devoted to artifacts correction through de-saturation and brightness reduction. Experimental results on a large dataset of images demonstrate the effectiveness of the pro- posed pipeline that outperforms other existing solutions in terms of hit rates maximization, false positives reduction and ad-hoc quality measure.
Sebastiano Battiato, Giovanni Maria Farinella, Mirko Guarnera, Giuseppe Messina, Daniele Ravì
ICIP1
2010 Characterization of signal perturbation using voting based curve fitting for multispectral images
abstract
Signal degradation impacts the final quality of images acquired using remote sensing radiometer. The effectiveness of a restoration algorithm strongly depends on two main factors: an accurate model of the disturbs introduced by the acquisition device and adaptation of the filtering method to image content. In this paper we target the first factor, by providing a solution for characterizing multispectral image signal degradation. A framework for estimating signal disturbs from heterogeneous sets of multispectral images is presented jointly with a voting-based technique for determining the best coefficients of the fitting equation. Tests conducted on multispectral images confirm the effectiveness of the proposed approach.
Sebastiano Battiato, Giovanni Puglisi, Rosetta Rizzo
ICIP1
2010 Boosting Gray Codes for Red Eyes Removal
abstract
Since the large diffusion of digital camera and mobile devices with embedded camera and flashgun, the red-eyes artifacts have de-facto become a critical problem. The technique herein described makes use of three main steps to identify and remove red-eyes. First, red eyes candidates are extracted from the input image by using an image filtering pipeline. A set of classifiers is then learned on gray code features extracted in the clustered patches space, and hence employed to distinguish between eyes and non-eyes patches. Once red-eyes are detected, artifacts are removed through desaturation and brightness reduction. The proposed method has been tested on large dataset of images achieving effective results in terms of hit rates maximization, false positives reduction and quality measure.
Sebastiano Battiato, Giovanni Maria Farinella, Mirko Guarnera, Giuseppe Messina, Daniele Ravì
ICPR1
2010 Second ACM international workshop on multimedia in forensics, security and intelligence (MiFor 2010)
abstract
This paper introduces the context of the workshop and the associated papers.
Sebastiano Battiato, Sabu Emmanuel, Adrian Ulges, Marcel Worring
ACM Multimedia1
2010 3D ancient mosaics
abstract
Digital 3D mosaics generation is a current trend of NPR (Non Photorealistic Rendering) field; in this demo we present an interactive system realized in JAVA where the user can simulate ancient mosaic in a 3D environment starting for any input image. Different simulation engines able to render the so-called "Opus Musivum"and "Opus Vermiculatum" are employed. Different parameters can be dynamically adjusted to obtain very impressive results.
Sebastiano Battiato, Giovanni Puglisi
ACM Multimedia1
2010 Exploiting visual and text features for direct marketing learning in time and space constrained domains
Sebastiano Battiato, Giovanni Maria Farinella, Giovanni Giuffrida, Catarina Sismeiro, Giuseppe Tribulato
Pattern Anal. Appl.1
2010 A Robust Block-Based Image/Video Registration Approach for Mobile Imaging Devices
abstract
Digital video stabilization enables to acquire video sequences without disturbing jerkiness by compensating unwanted camera movements. In this paper, we propose a novel fast image registration algorithm based on block matching. Unreliable motion vectors (i.e., not related with jitter movements) are properly filtered out by making use of ad-hoc rules taking into account local similarity, local “activity,” and matching effectiveness. Moreover, a temporal analysis of the relative error computed at each frame has been performed. Reliable information is then used to retrieve inter-frame transformation parameters. Experiments on real cases confirm the effectiveness of the proposed approach even in critical conditions.
Sebastiano Battiato, Arcangelo Bruna, Giovanni Puglisi
IEEE Trans. Multim.1
2009 A bio-inspired CNN with re-indexing engine for lossless DNA microarray compression and segmentation
abstract
The DNA microarray images allow to analyze the natural gene expressions. In this paper we propose an advanced method to efficiently address the imaging storage as well as the performance of the algorithm used to retrieve information from DNA images. The cellular neural networks (CNNs) based core is able to provide a method to extract foreground (the DNA gene expression information) from DNA images. It is also proposed an innovative method to compress the DNA image by re-organizing the signal data belonging to the background by making use of a novel way to apply the re-indexing techniques to almost ¿uncorrelated¿ signal. Experiments confirm how the proposed method outperform previous solution in almost all cases.
Sebastiano Battiato, Francesco Rundo
ICIP1
2009 Spatial Hierarchy of Textons Distributions for Scene Classification
Sebastiano Battiato, Giovanni Maria Farinella, Giovanni Gallo, Daniele Ravì
MMM1
2009 Using visual and text features for direct marketing on multimedia messaging services domain
Sebastiano Battiato, Giovanni Maria Farinella, Giovanni Giuffrida, Catarina Sismeiro, Giuseppe Tribulato
Multim. Tools Appl.1
2008 Scene categorization using bag of Textons on spatial hierarchy
abstract
This paper proposes a method to recognize scene categories using bags of visual words obtained hierarchically partitioning into subregion the input images. Specifically, for each subregions the texton histogram and the extension of the sub-region is taken into account. The bags of visual words, obtained in this way, are weighted and used in a similarity measure during the categorization. Experimental tests using ten different scene categories show that the proposed approach achieves good performances with respect to the state of the art methods.
Sebastiano Battiato, Giovanni Maria Farinella, Giovanni Gallo, Daniele Ravì
ICIP1
2008 A robust video stabilization system by adaptive motion vectors filtering
abstract
Digital video stabilization allows to acquire video sequences without disturbing jerkiness, removing unwanted camera movements. In this paper we propose a novel fast video stabilization algorithm based on block matching of local motion vectors. Some of these vectors are properly filtered out by making use of ad-hoc rules taking into account local similarity, local ldquoactivityrdquo and matching effectiveness. Also a temporal analysis of the relative error computed at each frame has been achieved. Reliable information are then used to retrieve inter-frame transformation parameters. Experiments on real cases confirm the effectiveness of the proposed approach even in critical conditions.
Sebastiano Battiato, Giovanni Puglisi, Arcangelo Bruna
ICME1
2008 Regular texture removal for video stabilization
abstract
In this paper we propose a novel fast fuzzy classifier able to find regular and low distorted near regular texture taking into account the constraints of video stabilization applications. Digital video stabilization allows to acquire video sequences without disturbing jerkiness, removing unwanted camera movements. In presence of regular or near regular texture, video stabilization approaches typically fail. These kind of patterns, due to their periodicity, create multiple matching that degrade motion estimation performances. The proposed classifier has been used as a filtering module in a block based video stabilization approach. Experiments on real sequences with (and without) regular texture confirm the effectiveness of the proposed approach.
Sebastiano Battiato, Giovanni Puglisi, Arcangelo Bruna
ICPR1
2007 A Novel Image Re-Indexing by Self Organizing Motor Maps
abstract
Palette re-ordering is a well known and very effective approach for improving the compression of color indexed images. If the spatial distribution of the indexes in the image is smooth, greater compression ratios may be obtained. As known, obtaining an optimal re-indexing scheme is not a trivial task. In this paper we provide a novel algorithm for palette re-ordering problem making use of a motor map neural network. Experimental results show the real effectiveness of the proposed method both in terms of compression ratio and zero-order entropy of local differences. Also its computational complexity is competitive with previous works in the field.
Sebastiano Battiato, Francesco Rundo, Filippo Stanco
ICIP (6)1
2007 Digital Mosaic Frameworks - An Overview
abstract
Abstract Art often provides valuable hints for technological innovations especially in the field of Image Processing and Computer Graphics. In this paper we survey in a unified framework several methods to transform raster input images into good quality mosaics. For each of the major different approaches in literature the paper reports a short description and a discussion of the most relevant issues. To complete the survey comparisons among the different techniques both in terms of visual quality and computational complexity are provided.
Sebastiano Battiato, Gianpiero di Blasi, Giovanni Maria Farinella, Giovanni Gallo
Comput. Graph. Forum1
2007 Self Organizing Motor Maps for Color-Mapped Image Re-Indexing
abstract
Palette re-ordering is an effective approach for improving the compression of color-indexed images. If the spatial distribution of the indexes in the image is smooth, greater compression ratios may be obtained. As is already known, obtaining an optimal re-indexing scheme is not a trivial task. In this paper, we provide a novel algorithm for palette re-ordering problem making use of a motor map neural network. Experimental results show the real effectiveness of the proposed method both in terms of compression ratio and zero-order entropy of local differences. Also, its computational complexity is competitive with previous works in the field.
Sebastiano Battiato, Francesco Rundo, Filippo Stanco
IEEE Trans. Image Process.1
2004 Adaptive image data fusion for consumer devices application
abstract
This paper presents a complete system for building an improved picture with high dynamic range by using different pictures of the same scene acquired under different exposure settings. The image data fusion is achieved by merging the original data by weighting each single contribution on a pixel basis by suitable data function. Experiments confirm the effectiveness of such approach.
Alessandro Capra, Alfio Castorina, Paolo Vivirito, Sebastiano Battiato
MMSP4
2004 An efficient Re-indexing algorithm for color-mapped images
abstract
The efficiency of lossless compression algorithms for fixed-palette images (indexed images) may change if a different indexing scheme is adopted. Many lossless compression algorithms adopt a differential-predictive approach. Hence, if the spatial distribution of the indexes over the image is smooth, greater compression ratios may be obtained. Because of this, finding an indexing scheme that realizes such a smooth distribution is a relevant issue. Obtaining an optimal re-indexing scheme is suspected to be a hard problem and only approximate solutions have been provided in literature. In this paper, we restate the re-indexing problem as a graph optimization problem: an optimal re-indexing corresponds to the heaviest Hamiltonian path in a weighted graph. It follows that any algorithm which finds a good approximate solution to this graph-theoretical problem also provides a good re-indexing. We propose a simple and easy-to-implement approximation algorithm to find such a path. The proposed technique compares favorably with most of the algorithms proposed in literature, both in terms of computational complexity and of compression ratio.
Sebastiano Battiato, Giovanni Gallo, Gaetano Impoco, Filippo Stanco
IEEE Trans. Image Process.1
2003 A light viewfinder pipeline for consumer devices application
abstract
The paper describes an image generation pipeline able to realize a "viewfinder". The viewfinder of a typical handset device allows user to track in real time the scene under detection. The pipeline is composed by a set of blocks implementing demosaicing and image enhancement algorithms with new and efficient techniques able to reduce considerably the computational overhead. Experiments show how modest computational resources can be coupled with acceptable perceived image quality for this particular target.
Sebastiano Battiato, Alfio Castorina, Mirko Guarnera, Filippo Vella
ICME1
2003 Image quality improvement by adaptive exposure correction techniques
abstract
The proposed paper concerns the processing of images in digital format and, more specifically, particular techniques that can be advantageously used in digital still cameras for improving the quality of images acquired with a non-optimal exposure. The proposed approach analyses the CCD/CMOS sensor Bayer data or the corresponding color generated image and, after identifying specific features, it adjusts the exposure level according to a 'camera response' like function.
Giuseppe Messina, Alfio Castorina, Sebastiano Battiato, Angelo Bosco
ICME3
2002 Temporal noise reduction of Bayer matrixed video data
abstract
This paper describes a new approach for noise reduction of video sequences that directly processes raw data frames acquired by an image sensor. The proposed noise reduction filter operates on Bayer matrixed video sequences instead of the canonical YUV format allowing saving of resources in terms of time and space; this is particularly relevant for real time processing. Noise level is constantly monitored in order to change the filter strength adaptively. Experiments show the effectiveness of the proposed approach.
Angelo Bosco, Massimo Mancuso, Sebastiano Battiato, Giuseppe Spampinato
ICME (1)3
2002 A locally adaptive zooming algorithm for digital images
Sebastiano Battiato, Giovanni Gallo, Filippo Stanco
Image Vis. Comput.1
2000 An Efficient Algorithm for the Approximate Median Selection Problem
Sebastiano Battiato, Domenico Cantone, Dario Catalano, Gianluca Cincotti, Micha Hofri
CIAC1