VLDB 2026 Research / reviewers in the wild / expert
Simone Palazzo
dblp:46/11431
· DBLP profile ↗
67ranked-venue papers
9as first author
40since 2021 · last 2026
0000-0002-2441-0982ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 7 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 4 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Global-Local Feature Decoding with Adapter-Guided SAMv2 for Salient Object Detection
Morteza Moradi 0001, Mohammad Moradi 0001, Simone Palazzo, Ali Borji, Concetto Spampinato |
ICPR (10) | 3 |
| 2026 | Learning long- and short-term dynamics for human attention prediction using large video models
Morteza Moradi 0001, Mohammad Moradi 0001, Ali Borji, Federica Proietto Salanitri, Giovanni Bellitto, Francesco Rundo, Simone Palazzo, Concetto Spampinato |
Comput. Vis. Image Underst. | 7 |
| 2026 | Knowledge distillation meets video foundation models: A video saliency prediction case study
Morteza Moradi 0001, Mohammad Moradi 0001, Concetto Spampinato, Ali Borji, Simone Palazzo |
J. Vis. Commun. Image Represent. | 5 |
| 2025 | Automated MoCA Score Estimation Using Eye-Gaze Data and Vision TransformersabstractCognitive impairment is a growing public health concern, with early detection playing a crucial role in improving patient outcomes. The Montreal Cognitive Assessment (MoCA) is widely used for screening mild cognitive impairment (MCI) and early-stage dementia. However, traditional MoCA assessments require manual scoring by trained professionals, making the process labor-intensive, time-consuming, and susceptible to human error. To overcome these limitations, we propose an automated pipeline for MoCA score estimation using eye-gaze data and Vision Transformers (ViTs). Our approach leverages gaze-tracking technology to capture spatial and temporal eyemovement patterns during structured cognitive tasks, identifying subtle cognitive impairments that may otherwise go unnoticed. The raw gaze data is preprocessed and mapped onto taskrelevant image regions, where a pretrained ViT extracts highdimensional feature representations. To address inconsistencies in gaze sampling and improve temporal modeling, we introduce a time-aware positional embedding mechanism that enhances the model's ability to infer cognitive performance. These extracted features are then processed by a transformer-based classification model to predict MoCA scores with high accuracy. We validate our approach using a dataset collected from seven cognitive gaming sessions, demonstrating its effectiveness in automated cognitive assessment. The experimental results indicate that our method provides a reliable and efficient alternative to traditional MoCA evaluations, reducing dependency on human intervention while maintaining diagnostic accuracy. Raffaele Mineo, Isaak Kavasidis, Federica Proietto Salanitri, Lisa Passarello, Giovanni Piccininno, Vincenzo Masciale, Alessandro Anselmo, Cristiano Convertino, Domenico Rotondi, Nicola Laurieri, Simone Palazzo, Concetto Spampinato, Manuela Pennisi, Daniela Giordano |
CBMS | 11 |
| 2025 | Performance Evaluation Of Reinforcement Learning Algorithms For Navigation In Unstructured EnvironmentsabstractAutonomous navigation in unstructured outdoor environments presents significant challenges due to the complexity and variability of terrain and obstacles. Traditional navigation methods struggle with adaptability, while learning-based approaches, particularly deep Reinforcement Learning (RL), have shown promise but faces difficulties in generalization and data efficiency. To address these challenges, we tested the performance of various deep RL algorithms for point-goal navigation, on MIDGARD, our photorealistic simulation platform built on Unreal Engine. We compare PPO, A2C, SAC, and TD3, evaluating their effectiveness based on success rates and reward progression. Our results indicate that PPO outperforms other algorithms, achieving the highest success rate, while off-policy methods struggle due to inefficient exploration and policy updates. Francesco Cancelliere, Giuseppe Sutera, Dario C. Guastella, Simone Palazzo, Giovanni Muscato, Concetto Spampinato |
ECMS | 4 |
| 2025 | EEG-Music Emotion Recognition: Challenge OverviewabstractAs our understanding of emotions continues to evolve, the ability of machines to accurately interpret and respond to emotional cues is more important than ever. Traditional methods of emotion recognition often fall short, particularly when it comes to the subtle and complex responses elicited by music. The EEG-Music Emotion Recognition Challenge aims to leverage electroencephalography (EEG) to decode emotional states from brain signals while subjects listen to music. This initiative seeks to uncover the intricate relationship between neural activity and emotional responses, offering insights for advancing adaptive user interfaces. We propose two tracks: (1) Person Identification aims to identify the subject from whom the EEG was recorded, while (2) Emotion Recognition targets the decoding of emotional state of the subject while listening to a musical stimulus. Salvatore Calcagno 0002, Simone Carnemolla, Isaak Kavasidis, Simone Palazzo, Daniela Giordano, Concetto Spampinato |
ICASSP | 4 |
| 2025 | Distilling Knowledge from Large Video Models for Driver Visual Attention PredictionabstractDriver attention prediction has gained significant attention recently due to its role in developing advanced driver assistance systems (ADAS) and intelligent vehicles. The emergence of video foundation models (VFMs) has opened up new possibilities for improving video understanding tasks like video saliency prediction (VSP). However, these large models are often not cost-effective for ADAS and intelligent vehicles due to their size and resource demands. To address this, we present an early effort to use knowledge distillation for predicting driver visual attention, employing the first VFM-based VSP model, SalFoM, as the teacher network. Given that driver attention prediction datasets are smaller than those used for large models, fine-tuning such models is challenging due to their high parameter count. To overcome this, we designed a VFM-based driver attention prediction network with fewer parameters than the teacher network. Experimental results show our model’s effectiveness on benchmark datasets. Morteza Moradi 0001, Mohammad Moradi 0001, Concetto Spampinato, Ali Borji, Simone Palazzo |
ICASSP | 5 |
| 2025 | Radar-Based Imaging for Sign Language Recognition in Medical Communication
Raffaele Mineo, Gaia Caligiore, Federica Proietto Salanitri, Isaak Kavasidis, Senya Polikovsky, Sabina Fontana, Egidio Ragonese, Concetto Spampinato, Simone Palazzo |
MICCAI (6) | 9 |
| 2025 | SeeingSounds: Learning Audio-to-Visual Alignment via TextabstractWe introduce SeeingSounds, a lightweight and modular framework for audio-to-image generation that leverages the interplay between audio, language, and vision—without requiring any paired audio-visual data or training on visual generative models. Rather than treating audio as a substitute for text or relying solely on audio-to-text mappings, our method performs dual alignment: audio is projected into a semantic language space via a frozen language encoder, and, contextually grounded into the visual domain using a vision-language model. This approach, inspired by cognitive neuroscience, reflects the natural cross-modal associations observed in human perception. The model operates on frozen diffusion backbones and trains only lightweight adapters, enabling efficient and scalable learning. Moreover, it supports fine-grained and interpretable control through procedural text prompt generation, where audio transformations (e.g., volume or pitch shifts) translate into descriptive prompts (e.g., “a distant thunder”) that guide visual outputs. Extensive experiments across standard benchmarks confirm that SeeingSounds outperforms existing methods in both zero-shot and supervised settings, establishing a new state of the art in controllable audio-to-visual generation. Simone Carnemolla, Matteo Pennisi, Chiara Maria Russo, Simone Palazzo, Daniela Giordano, Concetto Spampinato |
MMAsia | 4 |
| 2025 | DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision ModelsabstractUnderstanding and explaining the behavior of machine learning models is essential for building transparent and trustworthy AI systems. We introduce DEXTER, a data-free framework that employs
diffusion models and large language models to generate global, textual explanations of visual classifiers. DEXTER operates by optimizing text prompts to synthesize class-conditional images that strongly activate a target classifier. These synthetic samples are then used to elicit detailed natural language reports that describe class-specific decision patterns and biases. Unlike prior work, DEXTER enables natural language explanation
about a classifier's decision process without access to training data or ground-truth labels. We demonstrate DEXTER's flexibility across three tasks—activation maximization, slice discovery and debiasing, and bias explanation—each illustrating its ability to uncover the internal mechanisms of visual classifiers. Quantitative and qualitative evaluations, including a user study, show that DEXTER produces accurate, interpretable outputs. Experiments on ImageNet, Waterbirds, CelebA, and FairFaces confirm that DEXTER outperforms existing approaches in global model explanation and class-level bias reporting. Code is available at https://github.com/perceivelab/dexter. Simone Carnemolla, Matteo Pennisi, Sarinda Samarasinghe, Giovanni Bellitto, Simone Palazzo, Daniela Giordano, Mubarak Shah, Concetto Spampinato |
NeurIPS | 5 |
| 2025 | DiffExplainer: Towards cross-modal global explanations with diffusion modelsabstractWe present DiffExplainer , a novel framework that, leveraging language-vision models, enables multimodal global explainability. DiffExplainer employs diffusion models conditioned on optimized text prompts, synthesizing images that maximize class outputs and hidden features of a classifier, thus providing a visual tool for explaining decisions. Moreover, the analysis of generated visual descriptions allows for automatic identification of biases and spurious features, as opposed to traditional methods that often rely on manual intervention. The cross-modal transferability of language-vision models also enables the possibility to describe decisions in a more human-interpretable way, i.e., through text. We conduct comprehensive experiments demonstrating the effectiveness of DiffExplainer on (1) the generation of high-quality images explaining model decisions, surpassing existing activation maximization methods, and (2) the automated identification of biases and spurious features. • Leverages diffusion models to generate images explaining classifier decisions. • Enables bias and spurious feature detection without manual intervention. • Outperforms activation maximization methods in image quality and feature analysis. • Enables specific model analysis by the use of fixed prompts. Matteo Pennisi, Giovanni Bellitto, Simone Palazzo, Isaak Kavasidis, Mubarak Shah, Concetto Spampinato |
Comput. Vis. Image Underst. | 3 |
| 2025 | Recent advancements in driver's attention prediction
Morteza Moradi 0001, Simone Palazzo, Francesco Rundo, Concetto Spampinato |
Multim. Tools Appl. | 2 |
| 2025 | Wake-Sleep Consolidated LearningabstractWe propose wake-sleep consolidated learning (WSCL), a learning strategy leveraging complementary learning system (CLS) theory and the wake-sleep phases of the human brain to improve the performance of deep neural networks (DNNs) for visual classification tasks in continual learning (CL) settings. Our method learns continually via the synchronization between distinct wake and sleep phases. During the wake phase, the model is exposed to sensory input and adapts its representations, ensuring stability through a dynamic parameter freezing mechanism and storing episodic memories in a short-term temporary memory (similar to what happens in the hippocampus). During the sleep phase, the training process is split into nonrapid eye movement (NREM) and rapid eye movement (REM) stages. In the NREM stage, the model's synaptic weights are consolidated using replayed samples from the short-term and long-term memory and the synaptic plasticity mechanism is activated, strengthening important connections and weakening unimportant ones. In the REM stage, the model is exposed to previously-unseen realistic visual sensory experience, and the dreaming process is activated, which enables the model to explore the potential feature space, thus preparing synapses for future knowledge. We evaluate the effectiveness of our approach on four benchmark datasets: CIFAR-10, CIFAR-100, Tiny-ImageNet, and FG-ImageNet. In all cases, our method outperforms the baselines and prior work, yielding a significant performance gain on continual visual classification tasks. Furthermore, we demonstrate the usefulness of all processing stages and the importance of dreaming to enable positive forward transfer (FWT). The code is available at: https://github.com/perceivelab/wscl. Amelia Sorrenti, Giovanni Bellitto, Federica Proietto Salanitri, Matteo Pennisi, Simone Palazzo, Concetto Spampinato |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | MatFuse: Controllable Material Generation with Diffusion ModelsabstractCreating high-quality materials in computer graphics is a challenging and time-consuming task, which requires great expertise. To simplify this process, we introduce MatFuse, a unified approach that harnesses the gener-ative power of diffusion models for creation and editing of 3D materials. Our method integrates multiple sources of conditioning, including color palettes, sketches, text, and pictures, enhancing creative possibilities and granting fine-grained control over material synthesis. Additionally, MatFuse enables map-level material editing capabilities through latent manipulation by means of a multi-encoder compression model which learns a disentangled latent rep-resentation for each map. We demonstrate the effectiveness of MatFuse under multiple conditioning settings and ex-plore the potential of material editing. Finally, we assess the quality of the generated materials both quantitatively in terms of CLIP-IQA and FID scores and qualitatively by conducting a user study. Source code for training MatFuse and supplemental mate-rials are publicly available at https: //gvecchio.com/ matfuse. Giuseppe Vecchio, Renato Sortino, Simone Palazzo, Concetto Spampinato |
CVPR | 3 |
| 2024 | Vito: Vision Transformer Optimization Via Knowledge Distillation On DecodersabstractIn this paper, we propose ViTO, a novel knowledge distillation strategy that aims to convert a CNN model into a transformer-based counterpart that incorporates the advantages of transformers while retaining or improving its inductive bias. Our approach is based on a two-level transformer architecture that includes an inner model for learning visual representations and an outer model that aims to match the teacher’s predictions through autoregression. Specifically, given an image in a batch, the outer model classifies the image by using, in addition to the image’s visual properties, also the predictions it has made on images previously seen within the same batch. The effect of this strategy is to allow the transformer to estimate self- and cross-attention across all input batch images to learn autoregressively intra-class and inter-class correlations.We experimentally validate ViTO on several standard benchmarks obtaining better performance than existing knowledge distillation strategies on transformers. Furthermore, our distilled transformer-based model shows better robustness properties than standard vision transformers, demonstrating the effectiveness of our proposed distillation strategy. Giovanni Bellitto, Renato Sortino, Paolo Spadaro, Simone Palazzo, Federica Proietto Salanitri, Giuseppe Fiameni, Efstratios Gavves, Concetto Spampinato |
ICIP | 4 |
| 2024 | Back to Supervision: Boosting Word Boundary Detection Through Frame Classification
Simone Carnemolla, Salvatore Calcagno 0002, Simone Palazzo, Daniela Giordano |
ICPR (33) | 3 |
| 2024 | Evidential Federated Learning for Skin Lesion Image Classification
Rutger Hendrix, Federica Proietto Salanitri, Concetto Spampinato, Simone Palazzo, Ulas Bagci |
ICPR (29) | 4 |
| 2024 | SalFoM: Dynamic Saliency Prediction with Video Foundation Models
Morteza Moradi 0001, Mohammad Moradi 0001, Francesco Rundo, Concetto Spampinato, Ali Borji, Simone Palazzo |
ICPR (22) | 6 |
| 2024 | FedRewind: Rewinding Continual Model Exchange for Decentralized Federated Learning
Luca Palazzo, Matteo Pennisi, Federica Proietto Salanitri, Giovanni Bellitto, Simone Palazzo, Concetto Spampinato |
ICPR (25) | 5 |
| 2024 | Incremental Object 6D Pose Estimation
Amelia Sorrenti, Yik Lung Pang, Giovanni Bellitto, Simone Palazzo, Concetto Spampinato, Changjae Oh |
ICPR (26) | 5 |
| 2024 | Saliency-driven Experience Replay for Continual LearningabstractWe present Saliency-driven Experience Replay - SER - a biologically-plausible approach based on replicating human visual saliency to enhance classification models in continual learning settings. Inspired by neurophysiological evidence that the primary visual cortex does not contribute to object manifold untangling for categorization and that primordial saliency biases are still embedded in the modern brain, we propose to employ auxiliary saliency prediction features as a modulation signal to drive and stabilize the learning of a sequence of non-i.i.d. classification tasks. Experimental results confirm that SER effectively enhances the performance (in some cases up to about twenty percent points) of state-of-the-art continual learning methods, both in class-incremental and task-incremental settings. Moreover, we show that saliency-based modulation successfully encourages the learning of features that are more robust to the presence of spurious features and to adversarial attacks than baseline methods. Code is available at: https://github.com/perceivelab/SER Giovanni Bellitto, Federica Proietto Salanitri, Matteo Pennisi, Matteo Boschini, Lorenzo Bonicelli, Angelo Porrello, Simone Calderara, Simone Palazzo, Concetto Spampinato |
NeurIPS | 8 |
| 2024 | MultiVD: A Transformer-based Multitask Approach for Software Vulnerability Detection
Claudio Curto, Daniela Giordano, Simone Palazzo, Daniel Gustav Indelicato |
SECRYPT | 3 |
| 2024 | FedER: Federated Learning through Experience Replay and privacy-preserving data synthesisabstractIn the medical field, multi-center collaborations are often sought to yield more generalizable findings by leveraging the heterogeneity of patient and clinical data. However, recent privacy regulations hinder the possibility to share data, and consequently, to come up with machine learning-based solutions that support diagnosis and prognosis. Federated learning (FL) aims at sidestepping this limitation by bringing AI-based solutions to data owners and only sharing local AI models, or parts thereof, that need then to be aggregated. However, most of the existing federated learning solutions are still at their infancy and show several shortcomings, from the lack of a reliable and effective aggregation scheme able to retain the knowledge learned locally to weak privacy preservation as real data may be reconstructed from model updates. Furthermore, the majority of these approaches, especially those dealing with medical data, relies on a centralized distributed learning strategy that poses robustness, scalability and trust issues. In this paper we present a federated learning strategy, FedER, that, exploiting experience replay and generative adversarial concepts, effectively integrates features from local nodes, providing models able to generalize across multiple datasets while maintaining privacy. FedER is tested on two tasks — tuberculosis and melanoma classification — using multiple datasets in order to simulate realistic non-i.i.d. medical data scenarios. Results show that our approach achieves performance comparable to standard (non-federated) learning and significantly outperforms state-of-the-art federated methods. Remarkably, we also observe that FedER enables any node model to be used as a global federation model. Indeed, the experience replay strategy with privacy-preserving synthetic data allows all node models to converge to reach the same optimum without the need of a single shared model. Code is available at https://github.com/perceivelab/FedER. Matteo Pennisi, Federica Proietto Salanitri, Giovanni Bellitto, Bruno Casella, Marco Aldinucci, Simone Palazzo, Concetto Spampinato |
Comput. Vis. Image Underst. | 6 |
| 2024 | Rebuttal to "Comments on 'Decoding Brain Representations by Multimodal Learning of Neural Activity and Visual Features' "abstractBharadwaj et al. (2023) present a comments paper evaluating the classification accuracy of several state-of-the-art methods using EEG data averaged over random class samples. According to the results, some of the methods achieve above-chance accuracy, while the method proposed in (Palazzo et al. 2020), that is the target of their analysis, does not. In this rebuttal, we address these claims and explain why they are not grounded in the cognitive neuroscience literature, and why the evaluation procedure is ineffective and unfair. Simone Palazzo, Concetto Spampinato, Isaak Kavasidis, Daniela Giordano, Joseph Schmidt, Mubarak Shah |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | A Convolutional-Transformer Model for FFR and iFR Assessment From Coronary AngiographyabstractThe quantification of stenosis severity from X-ray catheter angiography is a challenging task. Indeed, this requires to fully understand the lesion's geometry by analyzing dynamics of the contrast material, only relying on visual observation by clinicians. To support decision making for cardiac intervention, we propose a hybrid CNN-Transformer model for the assessment of angiography-based non-invasive fractional flow-reserve (FFR) and instantaneous wave-free ratio (iFR) of intermediate coronary stenosis. Our approach predicts whether a coronary artery stenosis is hemodynamically significant and provides direct FFR and iFR estimates. This is achieved through a combination of regression and classification branches that forces the model to focus on the cut-off region of FFR (around 0.8 FFR value), which is highly critical for decision-making. We also propose a spatio-temporal factorization mechanisms that redesigns the transformer's self-attention mechanism to capture both local spatial and temporal interactions between vessel geometry, blood flow dynamics, and lesion morphology. The proposed method achieves state-of-the-art performance on a dataset of 778 exams from 389 patients. Unlike existing methods, our approach employs a single angiography view and does not require knowledge of the key frame; supervision at training time is provided by a classification loss (based on a threshold of the FFR/iFR values) and a regression loss for direct estimation. Finally, the analysis of model interpretability and calibration shows that, in spite of the complexity of angiographic imaging data, our method can robustly identify the location of the stenosis and correlate prediction uncertainty to the provided output scores. Raffaele Mineo, Federica Proietto Salanitri, Giovanni Bellitto, Isaak Kavasidis, Ovidio De Filippo, M. Millesimo, Gaetano Maria de Ferrari, Marco Aldinucci, Daniela Giordano, Simone Palazzo, Fabrizio D'Ascenzo, Concetto Spampinato |
IEEE Trans. Medical Imaging | 10 |
| 2023 | Dynamic Graph Attention: Unraveling Spatio-Temporal Synchrony in EEG DataabstractIn this paper, we propose a deep model based on graph convolutional networks for emotion recognition using EEG data. The model encodes spatial and temporal features of EEG channels and learns relationships between nodes through a self-attention mechanism, capturing spatio-temporal synchrony in brain regions. Experimental results show that our model outperforms existing approaches, with the attention mechanism contributing significantly to classification accuracy. In particular, the attention scores provide insights into how EEG channels influence each other at different times, revealing spatio-temporal patterns of brain connectivity related to emotions. Federica Proietto Salanitri, Giovanni Bellitto, Raffaele Mineo, Matteo Pennisi, Amelia Sorrenti, Salvatore Calcagno 0002, Daniela Giordano, Simone Palazzo, Concetto Spampinato |
BIBM | 8 |
| 2023 | Collective Driver Attention: Towards a Comprehensive Visual Understanding of Traffic ScenesabstractThanks to state-of-the-art deep learning-based methods for driver's attention prediction, it becomes possible to estimate where drivers look at in different traffic scenes. However, such estimation only takes into account visual information of the front view from a single vehicle. To remedy the lack of comprehensiveness of this approach, modern advanced driver-assistance systems (ADAS) further incorporate individual-specific features, including blood pressure and heart rate, to provide more precise safety advice. Nonetheless, there is still room for the improvement of safety-related recommendations by means of predicting collective drivers' attention. Specifically, the conceptual idea presented in this work is based on integrating visual understanding of the surrounding environment of a vehicle, driver-specific information and estimated attention in order to create a collective knowledge of the road, obstacles, distraction points, pedestrians and drivers' consciousness level from viewpoints of several drivers to predict a holistic attention map. Then, such a 360-degree attention map enables drivers to being aware not only of their front view, but also of back view and around (left and right) sides to help them prevent accidents and keep away from obstacles. The proposed framework takes advantages of edge and cloud computing for processing real-time information and large-scale computing, respectively. This work is intended to open a broader window towards the development of next generation of networked ADAS systems by employing several heterogeneous sources of information, implicit participation of drivers and their visual understanding and reasoning. Morteza Moradi 0001, Simone Palazzo, Concetto Spampinato |
CoDIT | 2 |
| 2023 | A Baseline on Continual Learning Methods for Video Action RecognitionabstractContinual learning has recently attracted attention from the research community, as it aims to solve long-standing limitations of classic supervised-trained models. However, most research on this subject has tackled continual learning in simple image classification scenarios. In this paper, we present a benchmark of state-of-the-art continual learning methods on video action recognition. Besides the increased complexity due to the temporal dimension, the video setting imposes stronger requirements on computing resources for top-performing rehearsal methods. To counteract the increased memory requirements, we present two method-agnostic variants for rehearsal methods, exploiting measures of either model confidence or data information to select memorable samples. Our experiments show that, as expected from the literature, rehearsal methods outperform other approaches; moreover, the proposed memory-efficient variants are shown to be effective at retaining a certain level of performance with a smaller buffer size. Giulia Castagnolo, Concetto Spampinato, Francesco Rundo, Daniela Giordano, Simone Palazzo |
ICIP | 5 |
| 2023 | A Privacy-Preserving Walk in the Latent Space of Generative Models for Medical Applications
Matteo Pennisi, Federica Proietto Salanitri, Giovanni Bellitto, Simone Palazzo, Ulas Bagci, Concetto Spampinato |
MICCAI (3) | 4 |
| 2023 | TinyHD: Efficient Video Saliency Prediction with Heterogeneous Decoders using Hierarchical Maps DistillationabstractVideo saliency prediction has recently attracted attention of the research community, as it is an upstream task for several practical applications. However, current solutions are particurly computationally demanding, especially due to the wide usage of spatio-temporal 3D convolutions. We observe that, while different model architectures achieve similar performance on benchmarks, visual variations between predicted saliency maps are still significant. Inspired by this intuition, we propose a lightweight model that employs multiple simple heterogeneous decoders and adopts several practical approaches to improve accuracy while keeping computational costs low, such as hierarchical multi-map knowledge distillation, multi-output saliency prediction, unlabeled auxiliary datasets and channel reduction with teacher assistant supervision. Our approach achieves saliency prediction accuracy on par or better than state-of-the-art methods on DFH1K, UCF-Sports and Hollywood2 benchmarks, while enhancing significantly the efficiency of the model. Feiyan Hu, Simone Palazzo, Federica Proietto Salanitri, Giovanni Bellitto, Morteza Moradi 0001, Concetto Spampinato, Kevin McGuinness |
WACV | 2 |
| 2023 | Transformer-based image generation from scene graphsabstractGraph-structured scene descriptions can be efficiently used in generative models to control the composition of the generated image. Previous approaches are based on the combination of graph convolutional networks and adversarial methods for layout prediction and image generation, respectively. In this work, we show how employing multi-head attention to encode the graph information, as well as using a transformer-based model in the latent space for image generation can improve the quality of the sampled data, without the need to employ adversarial models with the subsequent advantage in terms of training stability. The proposed approach, specifically, is entirely based on transformer architectures both for encoding scene graphs into intermediate object layouts and for decoding these layouts into images, passing through a lower dimensional space learned by a vector-quantized variational autoencoder. Our approach shows an improved image quality with respect to state-of-the-art methods as well as a higher degree of diversity among multiple generations from the same scene graph. We evaluate our approach on three public datasets: Visual Genome, COCO, and CLEVR. We achieve an Inception Score of 13.7 and 12.8, and an FID of 52.3 and 60.3, on COCO and Visual Genome, respectively. We perform ablation studies on our contributions to assess the impact of each component. Code is available at https://github.com/perceivelab/trf-sg2im. Renato Sortino, Simone Palazzo, Francesco Rundo, Concetto Spampinato |
Comput. Vis. Image Underst. | 2 |
| 2023 | MeT: A graph transformer for semantic segmentation of 3D meshesabstractPolygonal meshes have become the standard for discretely approximating 3D shapes, thanks to their efficiency and high flexibility in capturing non-uniform shapes. This non-uniformity, however, leads to irregularity in the mesh structure, making tasks like segmentation of 3D meshes particularly challenging. Semantic segmentation of 3D mesh has been typically addressed through CNN-based approaches, leading to good accuracy. Recently, transformers have gained enough momentum both in NLP and computer vision fields, achieving performance at least on par with CNN models, supporting the long-sought architecture universalism. Following this trend, we propose a transformer-based method for semantic segmentation of 3D mesh motivated by a better modeling of the graph structure of meshes, by means of global attention mechanisms. In order to address the limitations of standard transformer architectures in modeling relative positions of non-sequential data, as in the case of 3D meshes, as well as in capturing the local context, we perform positional encoding by means the Laplacian eigenvectors of the adjacency matrix, replacing the traditional sinusoidal positional encodings, and by introducing clustering-based features into the self-attention and cross-attention operators. Experimental results, carried out on three sets of the Shape COSEG Dataset (Wang et al., 2012), on the human segmentation dataset proposed in Maron et al. (2017) and on the ShapeNet benchmark (Chang et al., 2015), show how the proposed approach yields state-of-the-art performance on semantic segmentation of 3D meshes. Giuseppe Vecchio, Luca Prezzavento, Carmelo Pino, Francesco Rundo, Simone Palazzo, Concetto Spampinato |
Comput. Vis. Image Underst. | 5 |
| 2022 | Transfer Without Forgetting
Matteo Boschini, Lorenzo Bonicelli, Angelo Porrello, Giovanni Bellitto, Matteo Pennisi, Simone Palazzo, Concetto Spampinato, Simone Calderara |
ECCV (23) | 6 |
| 2022 | Effects of Auxiliary Knowledge on Continual LearningabstractIn Continual Learning (CL), a neural network is trained on a stream of data whose distribution changes over time. In this context, the main problem is how to learn new information without forgetting old knowledge (i.e., Catastrophic Forgetting). Most existing CL approaches focus on finding solutions to preserve acquired knowledge, so working on the "past" of the model. However, we argue that as the model has to continually learn new tasks, it is also important to put focus on the "present" knowledge that could improve following tasks learning. In this paper we propose a new, simple, CL algorithm that focuses on solving the current task in a way that might facilitate the learning of the next ones. More specifically, our approach combines the main data stream with a secondary, diverse and uncorrelated stream, from which the network can draw auxiliary knowledge. This helps the model from different perspectives, since auxiliary data may contain useful features for the current and the next tasks and incoming task classes can be mapped onto auxiliary classes. Furthermore, the addition of data to the current task is implicitly making the classifier more robust as we are forcing the extraction of more discriminative features. Our method can outperform existing state-of-the-art models on the most common CL Image Classification benchmarks. Giovanni Bellitto, Matteo Pennisi, Simone Palazzo, Lorenzo Bonicelli, Matteo Boschini, Simone Calderara |
ICPR | 3 |
| 2022 | Transforming Image Generation from Scene GraphsabstractGenerating images from semantic visual knowledge is a challenging task, that can be useful to condition the synthesis process in complex, subtle, and unambiguous ways, compared to alternatives such as class labels or text descriptions. Although generative methods conditioned by semantic representations exist, they do not provide a way to control the generation process aside from the specification of constraints between objects. As an example, the possibility to iteratively generate or modify images by manually adding specific items is a desired property that, to our knowledge, has not been fully investigated in the literature. In this work we propose a transformer-based approach conditioned by scene graphs that, conversely to recent transformer-based methods, also employs a decoder to autoregressively compose images, making the synthesis process more effective and controllable. The proposed architecture is composed by three modules: 1) a graph convolutional network, to encode the relationships of the input graph; 2) an encoder-decoder transformer, which autoregressively composes the output image; 3) an auto-encoder, employed to generate representations used as input/output of each generation step by the transformer. Results obtained on CIFAR10 and MNIST images show that our model is able to satisfy semantic constraints defined by a scene graph and to model relations between visual objects in the scene by taking into account a user-provided partial rendering of the desired target. Renato Sortino, Simone Palazzo, Concetto Spampinato |
ICPR | 2 |
| 2021 | SurfaceNet: Adversarial SVBRDF Estimation from a Single ImageabstractIn this paper we present SurfaceNet, an approach for estimating spatially-varying bidirectional reflectance distribution function (SVBRDF) material properties from a single image. We pose the problem as an image translation task and propose a novel patch-based generative adversarial network (GAN) that is able to produce high-quality, high-resolution surface reflectance maps. The employment of the GAN paradigm has a twofold objective: 1) allowing the model to recover finer details than standard translation models; 2) reducing the domain shift between synthetic and real data distributions in an unsupervised way.An extensive evaluation, carried out on a public benchmark of synthetic and real images under different illumination conditions, shows that SurfaceNet largely outperforms existing SVBRDF reconstruction methods, both quantitatively and qualitatively. Furthermore, SurfaceNet exhibits a remarkable ability in generating high-quality maps from real samples without any supervision at training time.Source code available at https://github.com/perceivelab/surfacenet. Giuseppe Vecchio, Simone Palazzo, Concetto Spampinato |
ICCV | 2 |
| 2021 | An explainable AI system for automated COVID-19 assessment and lesion categorization from CT-scans
Matteo Pennisi, Isaak Kavasidis, Concetto Spampinato, Vincenzo Schininà, Simone Palazzo, Federica Proietto Salanitri, Giovanni Bellitto, Francesco Rundo, Marco Aldinucci, Massimo Cristofaro, Paolo Campioni, Elisa Pianura, Federica Di Stefano 0002, Ada Petrone, Fabrizio Albarello, Giuseppe Ippolito, Salvatore Cuzzocrea, Sabrina Conoci |
Artif. Intell. Medicine | 5 |
| 2021 | Hierarchical Domain-Adapted Feature Learning for Video Saliency PredictionabstractAbstract In this work, we propose a 3D fully convolutional architecture for video saliency prediction that employs hierarchical supervision on intermediate maps (referred to as conspicuity maps) generated using features extracted at different abstraction levels. We provide the base hierarchical learning mechanism with two techniques for domain adaptation and domain-specific learning. For the former, we encourage the model to unsupervisedly learn hierarchical general features using gradient reversal at multiple scales, to enhance generalization capabilities on datasets for which no annotations are provided during training. As for domain specialization, we employ domain-specific operations (namely, priors, smoothing and batch normalization) by specializing the learned features on individual datasets in order to maximize performance. The results of our experiments show that the proposed model yields state-of-the-art accuracy on supervised saliency prediction. When the base hierarchical model is empowered with domain-specific modules, performance improves, outperforming state-of-the-art models on three out of five metrics on the DHF1K benchmark and reaching the second-best results on the other two. When, instead, we test it in an unsupervised domain adaptation setting, by enabling hierarchical gradient reversal layers, we obtain performance comparable to supervised state-of-the-art. Source code, trained models and example outputs are publicly available at https://github.com/perceivelab/hd2s . Giovanni Bellitto, Federica Proietto Salanitri, Simone Palazzo, Francesco Rundo, Daniela Giordano, Concetto Spampinato |
Int. J. Comput. Vis. | 3 |
| 2021 | Decoding Brain Representations by Multimodal Learning of Neural Activity and Visual FeaturesabstractThis work presents a novel method of exploring human brain-visual representations, with a view towards replicating these processes in machines. The core idea is to learn plausible computational and biological representations by correlating human neural activity and natural images. Thus, we first propose a model, EEG-ChannelNet, to learn a brain manifold for EEG classification. After verifying that visual information can be extracted from EEG data, we introduce a multimodal approach that uses deep image and EEG encoders, trained in a siamese configuration, for learning a joint manifold that maximizes a compatibility measure between visual features and brain representations. We then carry out image classification and saliency detection on the learned manifold. Performance analyses show that our approach satisfactorily decodes visual information from neural signals. This, in turn, can be used to effectively supervise the training of deep learning models, as demonstrated by the high performance of image classification and saliency detection on out-of-training classes. The obtained results show that the learned brain-visual features lead to improved performance and simultaneously bring deep models more in line with cognitive neuroscience work related to visual perception and attention. Simone Palazzo, Concetto Spampinato, Isaak Kavasidis, Daniela Giordano, Joseph Schmidt, Mubarak Shah |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Exploiting structured high-level knowledge for domain-specific visual classification
Simone Palazzo, Francesca Murabito, Carmelo Pino, Francesco Rundo, Daniela Giordano, Mubarak Shah, Concetto Spampinato |
Pattern Recognit. | 1 |
| 2020 | Visual Saliency Detection guided by Neural SignalsabstractSaliency detection is a fundamental process of human visual perception, since it allows us to identify the most important parts of a scene, directing our analysis and interpretation capabilities on a reduced set of information and reducing reaction times. However, current approaches for automatic saliency detection either attempt to mimic human capabilities by building attention maps from hand-crafted feature analysis, or employ convolutional neural networks trained as black boxes, without any architectural or information prior from human biology. In this paper, we present an approach for saliency detection that combines the success of deep learning in identifying representations for visual data with a training paradigm aimed at matching neural activity provided directly by brain signals recorded while subjects look at images. We show that our approach is able to capture correspondences between visual elements and neural activities, successfully generalizing to unseen images to identify their most salient regions. Simone Palazzo, Francesco Rundo, Sebastiano Battiato, Daniela Giordano, Concetto Spampinato |
FG | 1 |
| 2020 | Deep Recurrent-Convolutional Model for Automated Segmentation of Craniomaxillofacial CT ScansabstractIn this paper we define a deep learning architecture for automated segmentation of anatomical structures in Craniomaxillofacial (CMF) CT scans that leverages the recent success of encoder-decoder models for semantic segmentation of natural images. In particular, we propose a fully convolutional deep network that combines the advantages of recent fully convolutional models, such as Tiramisu, with squeeze-and-excitation blocks for feature recalibration, integrated with convolutional LSTMs to model spatio-temporal correlations between consecutive slices. The proposed segmentation network shows superior performance and generalization capabilities (to different structures and imaging modalities) than state of the art methods on automated segmentation of CMF structures (e.g., mandibles and airways) in several standard benchmarks (e.g., MICCAI datasets) and on new datasets proposed herein, effectively facing shape variability. Francesca Murabito, Simone Palazzo, Federica Proietto Salanitri, Francesco Rundo, Ulas Bagci, Daniela Giordano, Rosalia Leonardi, Concetto Spampinato |
ICPR | 2 |
| 2020 | Deep Multi-stage Model for Automated Landmarking of Craniomaxillofacial CT ScansabstractIn this paper we define a deep multi-stage architecture for automated landmarking of craniomaxillofacial (CMF) CT images. Our model is composed of three subnetworks that first localize, on reduced-resolution images, areas where landmarks may be found and then refine the search, at full-resolution scale, through a hierarchical structure aiming at increasing the granularity of the investigated region. The multi-stage pipeline is designed to deal with full resolution data and does not require any additional pre-processing step to reduce search space, as opposed to existing methods that can be only adopted for searching landmarks located in well-defined anatomical structures (e.g., mandibles). The automated landmarking system is tested on identifying landmarks located in several CMF regions, achieving an average error of 0.8 mm, significantly lower than expert readings. The proposed model also outperforms baselines and is on par with existing models that employ additional upstream segmentation, on state-of-the-art benchmarks. Simone Palazzo, Giovanni Bellitto, Luca Prezzavento, Francesco Rundo, Ulas Bagci, Daniela Giordano, Rosalia Leonardi, Concetto Spampinato |
ICPR | 1 |
| 2020 | Domain Adaptation for Outdoor Robot Traversability Estimation from RGB data with Safety-Preserving LossabstractBeing able to estimate the traversability of the area surrounding a mobile robot is a fundamental task in the design of a navigation algorithm. However, the task is often complex, since it requires evaluating distances from obstacles, type and slope of terrain, and dealing with non-obvious discontinuities in detected distances due to perspective. In this paper, we present an approach based on deep learning to estimate and anticipate the traversing score of different routes in the field of view of an on-board RGB camera. The backbone of the proposed model is based on a state-of-the-art deep segmentation model, which is fine-tuned on the task of predicting route traversability. We then enhance the model's capabilities by a) addressing domain shifts through gradient-reversal unsupervised adaptation, and b) accounting for the specific safety requirements of a mobile robot, by encouraging the model to err on the safe side, i.e., penalizing errors that would cause collisions with obstacles more than those that would cause the robot to stop in advance. Experimental results show that our approach is able to satisfactorily identify traversable areas and to generalize to unseen locations. Simone Palazzo, Dario C. Guastella, Luciano Cantelli, Paolo Spadaro, Francesco Rundo, Giovanni Muscato, Daniela Giordano, Concetto Spampinato |
IROS | 1 |
| 2020 | Adversarial Framework for Unsupervised Learning of Motion Dynamics in Videos
Concetto Spampinato, Simone Palazzo, P. D'Oro, Daniela Giordano, Mubarak Shah |
Int. J. Comput. Vis. | 2 |
| 2020 | MASK-RL: Multiagent Video Object Segmentation Framework Through Reinforcement LearningabstractIntegrating human-provided location priors into video object segmentation has been shown to be an effective strategy to enhance performance, but their application at large scale is unfeasible. Gamification can help reduce the annotation burden, but it still requires user involvement. We propose a video object segmentation framework that leverages the combined advantages of user feedback for segmentation and gamification strategy by simulating multiple game players through a reinforcement learning (RL) model that reproduces human ability to pinpoint moving objects and using the simulated feedback to drive the decisions of a fully convolutional deep segmentation network. Experimental results on the DAVIS-17 benchmark show that: 1) including user-provided prior, even if not precise, yields high performance; 2) our RL agent replicates satisfactorily the same variability of humans in identifying spatiotemporal salient objects; and 3) employing artificially generated priors in an unsupervised video object segmentation model reaches state-of-the-art performance. Giuseppe Vecchio, Simone Palazzo, Daniela Giordano, Francesco Rundo, Concetto Spampinato |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Top-down saliency detection driven by visual classification
Francesca Murabito, Concetto Spampinato, Simone Palazzo, Daniela Giordano, Konstantin Pogorelov, Michael Riegler 0001 |
Comput. Vis. Image Underst. | 3 |
| 2017 | Deep Learning Human Mind for Automated Visual ClassificationabstractxWhat if we could effectively read the mind and transfer human visual capabilities to computer vision methods? In this paper, we aim at addressing this question by developing the first visual object classifier driven by human brain signals. In particular, we employ EEG data evoked by visual object stimuli combined with Recurrent Neural Networks (RNN) to learn a discriminative brain activity manifold of visual categories in a reading the mind effort. Afterward, we transfer the learned capabilities to machines by training a Convolutional Neural Network (CNN)-based regressor to project images onto the learned manifold, thus allowing machines to employ human brain-based features for automated visual classification. We use a 128-channel EEG with active electrodes to record brain activity of several subjects while looking at images of 40 ImageNet object classes. The proposed RNN-based approach for discriminating object classes using brain signals reaches an average accuracy of about 83%, which greatly outperforms existing methods attempting to learn EEG visual object representations. As for automated object categorization, our human brain-driven approach obtains competitive performance, comparable to those achieved by powerful CNN models and it is also able to generalize over different visual datasets. Concetto Spampinato, Simone Palazzo, Isaak Kavasidis, Daniela Giordano, Nasim Souly, Mubarak Shah |
CVPR | 2 |
| 2017 | Generative Adversarial Networks Conditioned by Brain SignalsabstractRecent advancements in generative adversarial networks (GANs), using deep convolutional models, have supported the development of image generation techniques able to reach satisfactory levels of realism. Further improvements have been proposed to condition GANs to generate images matching a specific object category or a short text description. In this work, we build on the latter class of approaches and investigate the possibility of driving and conditioning the image generation process by means of brain signals recorded, through an electroencephalograph (EEG), while users look at images from a set of 40 ImageNet object categories with the objective of generating the seen images. To accomplish this task, we first demonstrate that brain activity EEG signals encode visually-related information that allows us to accurately discriminate between visual object categories and, accordingly, we extract a more compact class-dependent representation of EEG data using recurrent neural networks. Afterwards, we use the learned EEG manifold to condition image generation employing GANs, which, during inference, will read EEG signals and convert them into images. We tested our generative approach using EEG signals recorded from six subjects while looking at images of the aforementioned 40 visual classes. The results show that for classes represented by well-defined visual patterns (e.g., pandas, airplane, etc.), the generated images are realistic and highly resemble those evoking the EEG signals used for conditioning GANs, resulting in an actual reading-the-mind process. Simone Palazzo, Concetto Spampinato, Isaak Kavasidis, Daniela Giordano, Mubarak Shah |
ICCV | 1 |
| 2017 | Brain2Image: Converting Brain Signals into ImagesabstractReading the human mind has been a hot topic in the last decades, and recent research in neuroscience has found evidence on the possibility of decoding, from neuroimaging data, how the human brain works. At the same time, the recent rediscovery of deep learning combined to the large interest of scientific community on generative methods has enabled the generation of realistic images by learning a data distribution from noise. The quality of generated images increases when the input data conveys information on visual content of images. Leveraging on these recent trends, in this paper we present an approach for generating images using visually-evoked brain signals recorded through an electroencephalograph (EEG). More specifically, we recorded EEG data from several subjects while observing images on a screen and tried to regenerate the seen images. To achieve this goal, we developed a deep-learning framework consisting of an LSTM stacked with a generative method, which learns a more compact and noise-free representation of EEG data and employs it to generate the visual stimuli evoking specific brain responses. Isaak Kavasidis, Simone Palazzo, Concetto Spampinato, Daniela Giordano, Mubarak Shah |
ACM Multimedia | 2 |
| 2017 | Deep learning for automated skeletal bone age assessment in X-ray images
Concetto Spampinato, Simone Palazzo, Daniela Giordano, Marco Aldinucci, Rosalia Leonardi |
Medical Image Anal. | 2 |
| 2017 | Gamifying Video Object SegmentationabstractVideo object segmentation can be considered as one of the most challenging computer vision problems. Indeed, so far, no existing solution is able to effectively deal with the peculiarities of real-world videos, especially in cases of articulated motion and object occlusions; limitations that appear more evident when we compare the performance of automated methods with the human one. However, manually segmenting objects in videos is largely impractical as it requires a lot of time and concentration. To address this problem, in this paper we propose an interactive video object segmentation method, which exploits, on one hand, the capability of humans to identify correctly objects in visual scenes, and on the other hand, the collective human brainpower to solve challenging and large-scale tasks. In particular, our method relies on a game with a purpose to collect human inputs on object locations, followed by an accurate segmentation phase achieved by optimizing an energy function encoding spatial and temporal constraints between object regions as well as human-provided location priors. Performance analysis carried out on complex video benchmarks, and exploiting data provided by over 60 users, demonstrated that our method shows a better trade-off between annotation times and segmentation accuracy than interactive video annotation and automated video object segmentation approaches. Concetto Spampinato, Simone Palazzo, Daniela Giordano |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | A diversity-based search approach to support annotation of a large fish image dataset
Daniela Giordano, Simone Palazzo, Concetto Spampinato |
Multim. Syst. | 2 |
| 2016 | Fine-grained object recognition in underwater visual data
Concetto Spampinato, Simone Palazzo, Pierre-Hugues Joalland, Sébastien Paris, Hervé Glotin, Katy Blanc, Diane Lingrand, Frédéric Precioso |
Multim. Tools Appl. | 2 |
| 2015 | Rejecting False Positives in Video Object Segmentation
Daniela Giordano, Isaak Kavasidis, Simone Palazzo, Concetto Spampinato |
CAIP (1) | 3 |
| 2015 | Superpixel-based video object segmentation using perceptual organization and location priorabstractIn this paper we present an approach for segmenting objects in videos taken in complex scenes with multiple and different targets. The method does not make any specific assumptions about the videos and relies on how objects are perceived by humans according to Gestalt laws. Initially, we rapidly generate a coarse foreground segmentation, which provides predictions about motion regions by analyzing how superpixel segmentation changes in consecutive frames. We then exploit these location priors to refine the initial segmentation by optimizing an energy function based on appearance and perceptual organization, only on regions where motion is observed. We evaluated our method on complex and challenging video sequences and it showed significant performance improvements over recent state-of-the-art methods, being also fast enough to be used for “on-the-fly” processing. Daniela Giordano, Francesca Murabito, Simone Palazzo, Concetto Spampinato |
CVPR | 3 |
| 2015 | Using the Eyes to "See" the ObjectsabstractThis paper investigates how to exploit eye gaze data for understanding visual content. In particular, we propose a human-in-the-loop approach for object segmentation in videos, where humans provide significant cues on spatiotemporal relations between object parts (i.e. superpixels in our approach) by simply looking at video sequences. Such constraints, together with object appearance properties, are encoded into an energy function so as to tackle the segmentation problem as a labeling one. The proposed method uses gaze data from only two people and was tested on two challenging visual benchmarks: 1) SegTrack v2 and 2) FBMS-59. The achieved performance showed how our method outperformed more complex video object segmentation approaches, while reducing the effort needed for collecting human feedback Concetto Spampinato, Simone Palazzo, Francesca Murabito, Daniela Giordano |
ACM Multimedia | 2 |
| 2015 | Nonparametric label propagation using mutual local similarity in nearest neighbors
Daniela Giordano, Isaak Kavasidis, Simone Palazzo, Concetto Spampinato |
Comput. Vis. Image Underst. | 3 |
| 2014 | Kernel Density Estimation Using Joint Spatial-Color-Depth Data for Background ModelingabstractThe use of low-cost devices for depth estimation, such as Microsoft Kinect, is becoming more and more popular in computer vision research. In this paper, we propose an algorithm for background modeling which exploits this kind of devices to make the background and foreground models more robust to effects such as camouflage and illumination changes. Our algorithm, after a preprocessing stage for aligning color and depth data and for filtering/filling noisy depth measurements, explicitly models the scene's background and foreground with a Kernel Density Estimation approach in a quantized x-y-hue-saturation-depth space. The results in three different indoor environments, with different lighting conditions, showed that our approach is able to achieve an accuracy in foreground segmentation over 90% and that the combination of depth data and illumination-independent color space proved to be very robust against noise and illumination changes. Daniela Giordano, Simone Palazzo, Concetto Spampinato |
ICPR | 2 |
| 2014 | Large Scale Data Processing in Ecology: A Case Study on Long-Term Underwater Video MonitoringabstractEcology is, nowadays, an interdisciplinary, collabo- rative and data-intensive science, therefore, discovering, integrat- ing and analysing daily-produced data is necessary to support researchers to investigate complex questions, ranging from single particles to animals to the biosphere [1]. As a consequence, ecology-related multimedia content has been produced massively in recent years: for example, the Xeno-canto project1 and the Pl@ntNet project2 respectively collected 140,000 audio records of 8,700 bird species and about 60,000 thousand images covering thousand of plant species, to be used by scientists or professionals. Unfortunately, a manual analysis of such amount of generated data is impossible: automatic analysis tools combined with high- performance computing (HPC) solutions are therefore heavily demanded for making sense of such big ecological data. In this paper we present a case study of large-scale video processing on HPC facilities for underwater fish monitoring in the context of the Fish4Knowledge project 3, where a system to analyse long-term underwater camera footage has been developed. The paper is meant to report on the employed hardware/software architecture, the design and deployment of the parallel job manager, and the problems encountered during the whole process, from load balancing to job submission policies to bottlenecks. Simone Palazzo, Concetto Spampinato, Daniela Giordano |
PDP | 1 |
| 2014 | A texton-based kernel density estimation approach for background modeling under extreme conditions
Concetto Spampinato, Simone Palazzo, Isaak Kavasidis |
Comput. Vis. Image Underst. | 2 |
| 2014 | An innovative web-based collaborative platform for video annotation
Isaak Kavasidis, Simone Palazzo, Roberto Di Salvo, Daniela Giordano, Concetto Spampinato |
Multim. Tools Appl. | 2 |
| 2014 | Understanding fish behavior during typhoon events in real-life underwater environments
Concetto Spampinato, Simone Palazzo, Bas Boom, Jacco van Ossenbruggen, Isaak Kavasidis, Roberto Di Salvo, Fang-Pang Lin, Daniela Giordano, Lynda Hardman, Robert B. Fisher |
Multim. Tools Appl. | 2 |
| 2014 | A rule-based event detection system for real-life underwater domain
Concetto Spampinato, Emma Beauxis-Aussalet, Simone Palazzo, Cigdem Beyan, Jacco van Ossenbruggen, Jiyin He, Bas Boom |
Mach. Vis. Appl. | 3 |
| 2013 | Covariance based modeling of underwater scenes for fish detectionabstractIn this paper we present an algorithm for visual object detection in a underwater real-life context which explicitly models both the background and the foreground for each frame - thus helping to avoid foreground absorption into similar background -, and integrates both colour and texture features (which have proved effective in overcoming the limitations of colour-only appearance descriptors) into a covariance-based model, which provides an elegant way to merge multiple features together and enforce structural relationships. A joint domain-range model combined to a post-processing approach based on Markov Random Field takes into account the spatial dependency between pixels in the classification process, unlike the classical pixel-oriented modeling techniques. Our results show the effectiveness of this approach in the underwater environment, which presents a lot of variety in scene conditions, objects' motion patterns, shapes and colouring, and background activity. Simone Palazzo, Isaak Kavasidis, Concetto Spampinato |
ICIP | 1 |
| 2012 | Evaluation of tracking algorithm performance without ground-truth dataabstractVisual tracking is a topic on which a lot of scientific work has been carried out in the last years. An important aspect of tracking algorithms is the performance evaluation, which has been carried out typically through hand-labeled ground-truth data. Since the manual generation of ground truth is a time-consuming, error-prone and tedious task, recently many researchers have focused their attention on self-evaluation techniques for performance analysis. In this paper we propose a novel tool that enables image processing researchers to test the performance of tracking algorithms without resorting to hand-labeled ground truth data. The proposed approach consists of computing a set of features describing shape, appearance and motion of the tracked objects and combining them through a naive Bayesian classifier, in order to obtain a probability score representing the overall evaluation of each tracking decision. The method was tested on three different targets (vehicles, humans and fish) with three different tracking algorithms and the results show how this approach is able to reflect the quality of the performed tracking. Concetto Spampinato, Simone Palazzo, Daniela Giordano |
ICIP | 2 |
| 2012 | Enhancing object detection performance by integrating motion objectness and perceptual organization
Concetto Spampinato, Simone Palazzo |
ICPR | 2 |