EDBT 2026 Demo / reviewers in the wild / expert
Concetto Spampinato
dblp:32/5543
· DBLP profile ↗
114ranked-venue papers
20as first author
49since 2021 · last 2026
0000-0001-6653-2577ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 57 · 11 first-author · 23 since 2021Artificial intelligence and machine learning · 55 · 8 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 2 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 9 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Global-Local Feature Decoding with Adapter-Guided SAMv2 for Salient Object Detection
Morteza Moradi 0001, Mohammad Moradi 0001, Simone Palazzo, Ali Borji, Concetto Spampinato |
ICPR (10) | 5 |
| 2026 | Learning long- and short-term dynamics for human attention prediction using large video models
Morteza Moradi 0001, Mohammad Moradi 0001, Ali Borji, Federica Proietto Salanitri, Giovanni Bellitto, Francesco Rundo, Simone Palazzo, Concetto Spampinato |
Comput. Vis. Image Underst. | 8 |
| 2026 | Knowledge distillation meets video foundation models: A video saliency prediction case study
Morteza Moradi 0001, Mohammad Moradi 0001, Concetto Spampinato, Ali Borji, Simone Palazzo |
J. Vis. Commun. Image Represent. | 3 |
| 2026 | AdverIN: Monotonic adversarial intensity attack for domain generalization in medical image segmentation
Zheyuan Zhang 0001, Bin Wang 0068, Lanhong Yao, Elif Keles, Debesh Jha, Matthew Antalek, Gorkem Durak, Alpay Medetalibeyoglu, Concetto Spampinato, Baris Turkbey, Boqing Gong, Ulas Bagci |
Medical Image Anal. | 9 |
| 2026 | Decoding attention from the visual cortex: fMRI-based prediction of human saliency maps
Salvatore Calcagno 0002, Marco Finocchiaro, Giovanni Bellitto, Concetto Spampinato, Federica Proietto Salanitri |
Pattern Recognit. Lett. | 4 |
| 2026 | SAM-guided prompt learning for Multiple Sclerosis lesion segmentationabstractAccurate segmentation of Multiple Sclerosis (MS) lesions remains a critical challenge in medical image analysis due to their small size, irregular shape, and sparse distribution. Despite recent progress in vision foundation models — such as SAM and its medical variant MedSAM — these models have not yet been explored in the context of MS lesion segmentation. Moreover, their reliance on manually crafted prompts and high inference-time computational cost limits their applicability in clinical workflows, especially in resource-constrained environments. In this work, we introduce a novel training-time framework for effective and efficient MS lesion segmentation. Our method leverages SAM solely during training to guide a prompt learner that automatically discovers task-specific embeddings. At inference, SAM is replaced by a lightweight convolutional aggregator that maps the learned embeddings directly into segmentation masks—enabling fully automated, low-cost deployment. We show that our approach significantly outperforms existing specialized methods on the public MSLesSeg dataset, establishing new performance benchmarks in a domain where foundation models had not previously been applied. To assess generalizability, we also evaluate our method on pancreas and prostate segmentation tasks, where it achieves competitive accuracy while requiring an order of magnitude fewer parameters and computational resources compared to SAM-based pipelines. By eliminating the need for foundation models at inference time, our framework enables efficient segmentation without sacrificing accuracy. This design bridges the gap between large-scale pretraining and real-world clinical deployment, offering a scalable and practical solution for MS lesion segmentation and beyond. Code is available at https://github.com/perceivelab/MS-SAM-LESS . Federica Proietto Salanitri, Giovanni Bellitto, Salvatore Calcagno 0002, Ulas Bagci, Concetto Spampinato, Manuela Pennisi |
Pattern Recognit. Lett. | 5 |
| 2025 | Automated MoCA Score Estimation Using Eye-Gaze Data and Vision TransformersabstractCognitive impairment is a growing public health concern, with early detection playing a crucial role in improving patient outcomes. The Montreal Cognitive Assessment (MoCA) is widely used for screening mild cognitive impairment (MCI) and early-stage dementia. However, traditional MoCA assessments require manual scoring by trained professionals, making the process labor-intensive, time-consuming, and susceptible to human error. To overcome these limitations, we propose an automated pipeline for MoCA score estimation using eye-gaze data and Vision Transformers (ViTs). Our approach leverages gaze-tracking technology to capture spatial and temporal eyemovement patterns during structured cognitive tasks, identifying subtle cognitive impairments that may otherwise go unnoticed. The raw gaze data is preprocessed and mapped onto taskrelevant image regions, where a pretrained ViT extracts highdimensional feature representations. To address inconsistencies in gaze sampling and improve temporal modeling, we introduce a time-aware positional embedding mechanism that enhances the model's ability to infer cognitive performance. These extracted features are then processed by a transformer-based classification model to predict MoCA scores with high accuracy. We validate our approach using a dataset collected from seven cognitive gaming sessions, demonstrating its effectiveness in automated cognitive assessment. The experimental results indicate that our method provides a reliable and efficient alternative to traditional MoCA evaluations, reducing dependency on human intervention while maintaining diagnostic accuracy. Raffaele Mineo, Isaak Kavasidis, Federica Proietto Salanitri, Lisa Passarello, Giovanni Piccininno, Vincenzo Masciale, Alessandro Anselmo, Cristiano Convertino, Domenico Rotondi, Nicola Laurieri, Simone Palazzo, Concetto Spampinato, Manuela Pennisi, Daniela Giordano |
CBMS | 12 |
| 2025 | Performance Evaluation Of Reinforcement Learning Algorithms For Navigation In Unstructured EnvironmentsabstractAutonomous navigation in unstructured outdoor environments presents significant challenges due to the complexity and variability of terrain and obstacles. Traditional navigation methods struggle with adaptability, while learning-based approaches, particularly deep Reinforcement Learning (RL), have shown promise but faces difficulties in generalization and data efficiency. To address these challenges, we tested the performance of various deep RL algorithms for point-goal navigation, on MIDGARD, our photorealistic simulation platform built on Unreal Engine. We compare PPO, A2C, SAC, and TD3, evaluating their effectiveness based on success rates and reward progression. Our results indicate that PPO outperforms other algorithms, achieving the highest success rate, while off-policy methods struggle due to inefficient exploration and policy updates. Francesco Cancelliere, Giuseppe Sutera, Dario C. Guastella, Simone Palazzo, Giovanni Muscato, Concetto Spampinato |
ECMS | 6 |
| 2025 | EEG-Music Emotion Recognition: Challenge OverviewabstractAs our understanding of emotions continues to evolve, the ability of machines to accurately interpret and respond to emotional cues is more important than ever. Traditional methods of emotion recognition often fall short, particularly when it comes to the subtle and complex responses elicited by music. The EEG-Music Emotion Recognition Challenge aims to leverage electroencephalography (EEG) to decode emotional states from brain signals while subjects listen to music. This initiative seeks to uncover the intricate relationship between neural activity and emotional responses, offering insights for advancing adaptive user interfaces. We propose two tracks: (1) Person Identification aims to identify the subject from whom the EEG was recorded, while (2) Emotion Recognition targets the decoding of emotional state of the subject while listening to a musical stimulus. Salvatore Calcagno 0002, Simone Carnemolla, Isaak Kavasidis, Simone Palazzo, Daniela Giordano, Concetto Spampinato |
ICASSP | 6 |
| 2025 | Distilling Knowledge from Large Video Models for Driver Visual Attention PredictionabstractDriver attention prediction has gained significant attention recently due to its role in developing advanced driver assistance systems (ADAS) and intelligent vehicles. The emergence of video foundation models (VFMs) has opened up new possibilities for improving video understanding tasks like video saliency prediction (VSP). However, these large models are often not cost-effective for ADAS and intelligent vehicles due to their size and resource demands. To address this, we present an early effort to use knowledge distillation for predicting driver visual attention, employing the first VFM-based VSP model, SalFoM, as the teacher network. Given that driver attention prediction datasets are smaller than those used for large models, fine-tuning such models is challenging due to their high parameter count. To overcome this, we designed a VFM-based driver attention prediction network with fewer parameters than the teacher network. Experimental results show our model’s effectiveness on benchmark datasets. Morteza Moradi 0001, Mohammad Moradi 0001, Concetto Spampinato, Ali Borji, Simone Palazzo |
ICASSP | 3 |
| 2025 | Zero-shot Decentralized Federated LearningabstractCLIP has revolutionized zero-shot learning by enabling task generalization without fine-tuning. While prompting techniques like CoOp and CoCoOp enhance CLIP’s adaptability, their effectiveness in Federated Learning (FL) remains an open challenge. Existing federated prompt learning approaches, such as FedCoOp and FedTPG, improve performance but face generalization issues, high communication costs, and reliance on a central server, limiting scalability and privacy.We propose Zero-shot Decentralized Federated Learning (ZeroDFL), a fully decentralized framework that enables zero-shot adaptation across distributed clients without a central coordinator. ZeroDFL employs an iterative prompt-sharing mechanism, allowing clients to optimize and exchange textual prompts to enhance generalization while drastically reducing communication overhead.We validate ZeroDFL on nine diverse image classification datasets, demonstrating that it consistently outperforms—or remains on par with—state-of-the-art federated prompt learning methods. More importantly, ZeroDFL achieves this performance in a fully decentralized setting while reducing communication overhead by 118× compared to FedTPG. These results highlight that our approach not only enhances generalization in federated zero-shot learning but also improves scalability, efficiency, and privacy preservation—paving the way for decentralized adaptation of large vision-language models in real-world applications. Code is available at: https://github.com/perceivelab/ZeroDFL Alessio Masano, Matteo Pennisi, Federica Proietto Salanitri, Concetto Spampinato, Giovanni Bellitto |
IJCNN | 4 |
| 2025 | Radar-Based Imaging for Sign Language Recognition in Medical Communication
Raffaele Mineo, Gaia Caligiore, Federica Proietto Salanitri, Isaak Kavasidis, Senya Polikovsky, Sabina Fontana, Egidio Ragonese, Concetto Spampinato, Simone Palazzo |
MICCAI (6) | 8 |
| 2025 | Pre-Forgettable Models: Prompt Learning as a Native Mechanism for UnlearningabstractFoundation models have transformed multimedia analysis by enabling robust and transferable representations across diverse modalities and tasks. However, their static deployment conflicts with growing societal and regulatory demands-particularly the need to unlearn specific data upon request, as mandated by privacy frameworks such as the GDPR. Traditional unlearning approaches, including retraining, activation editing, or distillation, are often computationally expensive, fragile, and ill-suited for real-time or continuously evolving systems. In this paper, we propose a paradigm shift: rethinking unlearning not as a retroactive intervention but as a built-in capability. We introduce a prompt-based learning framework that unifies knowledge acquisition and removal within a single training phase. Rather than encoding information in model weights, our approach binds class-level semantics to dedicated prompt tokens. This design enables instant unlearning simply by removing the corresponding prompt-without retraining, model modification, or access to original data. Experiments demonstrate that our framework preserves predictive performance on retained classes while effectively erasing forgotten ones. Beyond utility, our method exhibits strong privacy and security guarantees: it is resistant to membership inference attacks, and prompt removal prevents any residual knowledge extraction, even under adversarial conditions. This ensures compliance with data protection principles and safeguards against unauthorized access to forgotten information, making the framework suitable for deployment in sensitive and regulated environments. Overall, by embedding removability into the architecture itself, this work establishes a new foundation for designing modular, scalable and ethically responsive AI models. Rutger Hendrix, Giovanni Patanè, Leonardo G. Russo, Simone Carnemolla, Federica Proietto Salanitri, Giovanni Bellitto, Concetto Spampinato, Matteo Pennisi |
ACM Multimedia | 7 |
| 2025 | SeeingSounds: Learning Audio-to-Visual Alignment via TextabstractWe introduce SeeingSounds, a lightweight and modular framework for audio-to-image generation that leverages the interplay between audio, language, and vision—without requiring any paired audio-visual data or training on visual generative models. Rather than treating audio as a substitute for text or relying solely on audio-to-text mappings, our method performs dual alignment: audio is projected into a semantic language space via a frozen language encoder, and, contextually grounded into the visual domain using a vision-language model. This approach, inspired by cognitive neuroscience, reflects the natural cross-modal associations observed in human perception. The model operates on frozen diffusion backbones and trains only lightweight adapters, enabling efficient and scalable learning. Moreover, it supports fine-grained and interpretable control through procedural text prompt generation, where audio transformations (e.g., volume or pitch shifts) translate into descriptive prompts (e.g., “a distant thunder”) that guide visual outputs. Extensive experiments across standard benchmarks confirm that SeeingSounds outperforms existing methods in both zero-shot and supervised settings, establishing a new state of the art in controllable audio-to-visual generation. Simone Carnemolla, Matteo Pennisi, Chiara Maria Russo, Simone Palazzo, Daniela Giordano, Concetto Spampinato |
MMAsia | 6 |
| 2025 | DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision ModelsabstractUnderstanding and explaining the behavior of machine learning models is essential for building transparent and trustworthy AI systems. We introduce DEXTER, a data-free framework that employs
diffusion models and large language models to generate global, textual explanations of visual classifiers. DEXTER operates by optimizing text prompts to synthesize class-conditional images that strongly activate a target classifier. These synthetic samples are then used to elicit detailed natural language reports that describe class-specific decision patterns and biases. Unlike prior work, DEXTER enables natural language explanation
about a classifier's decision process without access to training data or ground-truth labels. We demonstrate DEXTER's flexibility across three tasks—activation maximization, slice discovery and debiasing, and bias explanation—each illustrating its ability to uncover the internal mechanisms of visual classifiers. Quantitative and qualitative evaluations, including a user study, show that DEXTER produces accurate, interpretable outputs. Experiments on ImageNet, Waterbirds, CelebA, and FairFaces confirm that DEXTER outperforms existing approaches in global model explanation and class-level bias reporting. Code is available at https://github.com/perceivelab/dexter. Simone Carnemolla, Matteo Pennisi, Sarinda Samarasinghe, Giovanni Bellitto, Simone Palazzo, Daniela Giordano, Mubarak Shah, Concetto Spampinato |
NeurIPS | 8 |
| 2025 | DiffExplainer: Towards cross-modal global explanations with diffusion modelsabstractWe present DiffExplainer , a novel framework that, leveraging language-vision models, enables multimodal global explainability. DiffExplainer employs diffusion models conditioned on optimized text prompts, synthesizing images that maximize class outputs and hidden features of a classifier, thus providing a visual tool for explaining decisions. Moreover, the analysis of generated visual descriptions allows for automatic identification of biases and spurious features, as opposed to traditional methods that often rely on manual intervention. The cross-modal transferability of language-vision models also enables the possibility to describe decisions in a more human-interpretable way, i.e., through text. We conduct comprehensive experiments demonstrating the effectiveness of DiffExplainer on (1) the generation of high-quality images explaining model decisions, surpassing existing activation maximization methods, and (2) the automated identification of biases and spurious features. • Leverages diffusion models to generate images explaining classifier decisions. • Enables bias and spurious feature detection without manual intervention. • Outperforms activation maximization methods in image quality and feature analysis. • Enables specific model analysis by the use of fixed prompts. Matteo Pennisi, Giovanni Bellitto, Simone Palazzo, Isaak Kavasidis, Mubarak Shah, Concetto Spampinato |
Comput. Vis. Image Underst. | 6 |
| 2025 | Large-scale multi-center CT and MRI segmentation of pancreas with deep learningabstractAutomated volumetric segmentation of the pancreas on cross-sectional imaging is needed for diagnosis and follow-up of pancreatic diseases. While CT-based pancreatic segmentation is more established, MRI-based segmentation methods are understudied, largely due to a lack of publicly available datasets, benchmarking research efforts, and domain-specific deep learning methods. In this retrospective study, we collected a large dataset (767 scans from 499 participants) of T1-weighted (T1 W) and T2-weighted (T2 W) abdominal MRI series from five centers between March 2004 and November 2022. We also collected CT scans of 1,350 patients from publicly available sources for benchmarking purposes. We introduced a new pancreas segmentation method, called PanSegNet , combining the strengths of nnUNet and a Transformer network with a new linear attention module enabling volumetric computation. We tested PanSegNet ’s accuracy in cross-modality (a total of 2,117 scans) and cross-center settings with Dice and Hausdorff distance (HD95) evaluation metrics. We used Cohen’s kappa statistics for intra and inter-rater agreement evaluation and paired t-tests for volume and Dice comparisons, respectively. For segmentation accuracy, we achieved Dice coefficients of 88.3% (±7.2%, at case level) with CT, 85.0% (±7.9%) with T1 W MRI, and 86.3% (±6.4%) with T2 W MRI. There was a high correlation for pancreas volume prediction with R 2 of 0.91, 0.84, and 0.85 for CT, T1 W, and T2 W, respectively. We found moderate inter-observer (0.624 and 0.638 for T1 W and T2 W MRI, respectively) and high intra-observer agreement scores. All MRI data is made available at https://osf.io/kysnj/ . Our source code is available at https://github.com/NUBagciLab/PaNSegNet . • We develop a first-ever cross-platform compatible (T1 W, T2 W, and CT) pancreas segmentation tool, named PanSegNet . • PaNSegNet has innovative “linear self-attention” blocks to reduce computational cost significantly while operating on 3D. • We shared our both source code and multi-center multi-contrast MRI datasets with ground truths. • PaNSegNet underwent rigorous validation, including cross-domain and multi-center comparisons between CT and MRI scans. Zheyuan Zhang 0001, Elif Keles, Gorkem Durak, Yavuz Taktak, Onkar Susladkar, Vandan Gorade, Debesh Jha, Asli C. Ormeci, Alpay Medetalibeyoglu, Lanhong Yao, Bin Wang 0068, Ilkin Isler, Linkai Peng, Hongyi Pan, Camila Lopes Vendrami, Amir Bourhani, Yury Velichko, Boqing Gong, Concetto Spampinato, Ayis Pyrros, Pallavi Tiwari, Derk C. F. Klatte, Megan Engels, Sanne Hoogenboom, Candice W. Bolan, Emil Agarunov, Nassier Harfouch, Chenchan Huang, Marco J. Bruno, Ivo Schoots, Rajesh Keswani, Frank H. Miller, Tamas Gonda, Cemal Yazici, Temel Tirkes, Baris Turkbey, Michael B. Wallace, Ulas Bagci |
Medical Image Anal. | 19 |
| 2025 | Recent advancements in driver's attention prediction
Morteza Moradi 0001, Simone Palazzo, Francesco Rundo, Concetto Spampinato |
Multim. Tools Appl. | 4 |
| 2025 | Wake-Sleep Consolidated LearningabstractWe propose wake-sleep consolidated learning (WSCL), a learning strategy leveraging complementary learning system (CLS) theory and the wake-sleep phases of the human brain to improve the performance of deep neural networks (DNNs) for visual classification tasks in continual learning (CL) settings. Our method learns continually via the synchronization between distinct wake and sleep phases. During the wake phase, the model is exposed to sensory input and adapts its representations, ensuring stability through a dynamic parameter freezing mechanism and storing episodic memories in a short-term temporary memory (similar to what happens in the hippocampus). During the sleep phase, the training process is split into nonrapid eye movement (NREM) and rapid eye movement (REM) stages. In the NREM stage, the model's synaptic weights are consolidated using replayed samples from the short-term and long-term memory and the synaptic plasticity mechanism is activated, strengthening important connections and weakening unimportant ones. In the REM stage, the model is exposed to previously-unseen realistic visual sensory experience, and the dreaming process is activated, which enables the model to explore the potential feature space, thus preparing synapses for future knowledge. We evaluate the effectiveness of our approach on four benchmark datasets: CIFAR-10, CIFAR-100, Tiny-ImageNet, and FG-ImageNet. In all cases, our method outperforms the baselines and prior work, yielding a significant performance gain on continual visual classification tasks. Furthermore, we demonstrate the usefulness of all processing stages and the importance of dreaming to enable positive forward transfer (FWT). The code is available at: https://github.com/perceivelab/wscl. Amelia Sorrenti, Giovanni Bellitto, Federica Proietto Salanitri, Matteo Pennisi, Simone Palazzo, Concetto Spampinato |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | MatFuse: Controllable Material Generation with Diffusion ModelsabstractCreating high-quality materials in computer graphics is a challenging and time-consuming task, which requires great expertise. To simplify this process, we introduce MatFuse, a unified approach that harnesses the gener-ative power of diffusion models for creation and editing of 3D materials. Our method integrates multiple sources of conditioning, including color palettes, sketches, text, and pictures, enhancing creative possibilities and granting fine-grained control over material synthesis. Additionally, MatFuse enables map-level material editing capabilities through latent manipulation by means of a multi-encoder compression model which learns a disentangled latent rep-resentation for each map. We demonstrate the effectiveness of MatFuse under multiple conditioning settings and ex-plore the potential of material editing. Finally, we assess the quality of the generated materials both quantitatively in terms of CLIP-IQA and FID scores and qualitatively by conducting a user study. Source code for training MatFuse and supplemental mate-rials are publicly available at https: //gvecchio.com/ matfuse. Giuseppe Vecchio, Renato Sortino, Simone Palazzo, Concetto Spampinato |
CVPR | 4 |
| 2024 | Domain Generalization with fourier Transform and soft thresholdingabstractDomain generalization aims to train models on multiple source domains so that they can generalize well to unseen target domains. Among many domain generalization methods, Fourier-transformbased domain generalization methods have gained popularity primarily because they exploit the power of Fourier transformation to capture essential patterns and regularities in the data, making the model more robust to domain shifts. The mainstream Fouriertransform-based domain generalization swaps the Fourier amplitude spectrum while preserving the phase spectrum between the source and the target images. However, it neglects background interference in the amplitude spectrum. To overcome this limitation, we introduce a soft-thresholding function in the Fourier domain. We apply this newly designed algorithm to retinal fundus image segmentation, which is important for diagnosing ocular diseases but the neural network’s performance can degrade across different sources due to domain shifts. The proposed technique basically enhances fundus image augmentation by eliminating small values in the Fourier domain and providing better generalization. The innovative nature of the soft thresholding fused with Fourier-transform-based domain generalization improves neural network models’ performance by reducing the target images’ background interference significantly. Experiments on public data validate our approach’s effectiveness over conventional and state-of-the-art methods with superior segmentation metrics. Hongyi Pan, Bin Wang 0068, Zheyuan Zhang 0001, Xin Zhu 0005, Debesh Jha, A. Enis Çetin, Concetto Spampinato, Ulas Bagci |
ICASSP | 7 |
| 2024 | Vito: Vision Transformer Optimization Via Knowledge Distillation On DecodersabstractIn this paper, we propose ViTO, a novel knowledge distillation strategy that aims to convert a CNN model into a transformer-based counterpart that incorporates the advantages of transformers while retaining or improving its inductive bias. Our approach is based on a two-level transformer architecture that includes an inner model for learning visual representations and an outer model that aims to match the teacher’s predictions through autoregression. Specifically, given an image in a batch, the outer model classifies the image by using, in addition to the image’s visual properties, also the predictions it has made on images previously seen within the same batch. The effect of this strategy is to allow the transformer to estimate self- and cross-attention across all input batch images to learn autoregressively intra-class and inter-class correlations.We experimentally validate ViTO on several standard benchmarks obtaining better performance than existing knowledge distillation strategies on transformers. Furthermore, our distilled transformer-based model shows better robustness properties than standard vision transformers, demonstrating the effectiveness of our proposed distillation strategy. Giovanni Bellitto, Renato Sortino, Paolo Spadaro, Simone Palazzo, Federica Proietto Salanitri, Giuseppe Fiameni, Efstratios Gavves, Concetto Spampinato |
ICIP | 8 |
| 2024 | Evidential Federated Learning for Skin Lesion Image Classification
Rutger Hendrix, Federica Proietto Salanitri, Concetto Spampinato, Simone Palazzo, Ulas Bagci |
ICPR (29) | 3 |
| 2024 | SalFoM: Dynamic Saliency Prediction with Video Foundation Models
Morteza Moradi 0001, Mohammad Moradi 0001, Francesco Rundo, Concetto Spampinato, Ali Borji, Simone Palazzo |
ICPR (22) | 4 |
| 2024 | FedRewind: Rewinding Continual Model Exchange for Decentralized Federated Learning
Luca Palazzo, Matteo Pennisi, Federica Proietto Salanitri, Giovanni Bellitto, Simone Palazzo, Concetto Spampinato |
ICPR (25) | 6 |
| 2024 | Incremental Object 6D Pose Estimation
Amelia Sorrenti, Yik Lung Pang, Giovanni Bellitto, Simone Palazzo, Concetto Spampinato, Changjae Oh |
ICPR (26) | 6 |
| 2024 | Saliency-driven Experience Replay for Continual LearningabstractWe present Saliency-driven Experience Replay - SER - a biologically-plausible approach based on replicating human visual saliency to enhance classification models in continual learning settings. Inspired by neurophysiological evidence that the primary visual cortex does not contribute to object manifold untangling for categorization and that primordial saliency biases are still embedded in the modern brain, we propose to employ auxiliary saliency prediction features as a modulation signal to drive and stabilize the learning of a sequence of non-i.i.d. classification tasks. Experimental results confirm that SER effectively enhances the performance (in some cases up to about twenty percent points) of state-of-the-art continual learning methods, both in class-incremental and task-incremental settings. Moreover, we show that saliency-based modulation successfully encourages the learning of features that are more robust to the presence of spurious features and to adversarial attacks than baseline methods. Code is available at: https://github.com/perceivelab/SER Giovanni Bellitto, Federica Proietto Salanitri, Matteo Pennisi, Matteo Boschini, Lorenzo Bonicelli, Angelo Porrello, Simone Calderara, Simone Palazzo, Concetto Spampinato |
NeurIPS | 9 |
| 2024 | FedER: Federated Learning through Experience Replay and privacy-preserving data synthesisabstractIn the medical field, multi-center collaborations are often sought to yield more generalizable findings by leveraging the heterogeneity of patient and clinical data. However, recent privacy regulations hinder the possibility to share data, and consequently, to come up with machine learning-based solutions that support diagnosis and prognosis. Federated learning (FL) aims at sidestepping this limitation by bringing AI-based solutions to data owners and only sharing local AI models, or parts thereof, that need then to be aggregated. However, most of the existing federated learning solutions are still at their infancy and show several shortcomings, from the lack of a reliable and effective aggregation scheme able to retain the knowledge learned locally to weak privacy preservation as real data may be reconstructed from model updates. Furthermore, the majority of these approaches, especially those dealing with medical data, relies on a centralized distributed learning strategy that poses robustness, scalability and trust issues. In this paper we present a federated learning strategy, FedER, that, exploiting experience replay and generative adversarial concepts, effectively integrates features from local nodes, providing models able to generalize across multiple datasets while maintaining privacy. FedER is tested on two tasks — tuberculosis and melanoma classification — using multiple datasets in order to simulate realistic non-i.i.d. medical data scenarios. Results show that our approach achieves performance comparable to standard (non-federated) learning and significantly outperforms state-of-the-art federated methods. Remarkably, we also observe that FedER enables any node model to be used as a global federation model. Indeed, the experience replay strategy with privacy-preserving synthetic data allows all node models to converge to reach the same optimum without the need of a single shared model. Code is available at https://github.com/perceivelab/FedER. Matteo Pennisi, Federica Proietto Salanitri, Giovanni Bellitto, Bruno Casella, Marco Aldinucci, Simone Palazzo, Concetto Spampinato |
Comput. Vis. Image Underst. | 7 |
| 2024 | Rebuttal to "Comments on 'Decoding Brain Representations by Multimodal Learning of Neural Activity and Visual Features' "abstractBharadwaj et al. (2023) present a comments paper evaluating the classification accuracy of several state-of-the-art methods using EEG data averaged over random class samples. According to the results, some of the methods achieve above-chance accuracy, while the method proposed in (Palazzo et al. 2020), that is the target of their analysis, does not. In this rebuttal, we address these claims and explain why they are not grounded in the cognitive neuroscience literature, and why the evaluation procedure is ineffective and unfair. Simone Palazzo, Concetto Spampinato, Isaak Kavasidis, Daniela Giordano, Joseph Schmidt, Mubarak Shah |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | A Convolutional-Transformer Model for FFR and iFR Assessment From Coronary AngiographyabstractThe quantification of stenosis severity from X-ray catheter angiography is a challenging task. Indeed, this requires to fully understand the lesion's geometry by analyzing dynamics of the contrast material, only relying on visual observation by clinicians. To support decision making for cardiac intervention, we propose a hybrid CNN-Transformer model for the assessment of angiography-based non-invasive fractional flow-reserve (FFR) and instantaneous wave-free ratio (iFR) of intermediate coronary stenosis. Our approach predicts whether a coronary artery stenosis is hemodynamically significant and provides direct FFR and iFR estimates. This is achieved through a combination of regression and classification branches that forces the model to focus on the cut-off region of FFR (around 0.8 FFR value), which is highly critical for decision-making. We also propose a spatio-temporal factorization mechanisms that redesigns the transformer's self-attention mechanism to capture both local spatial and temporal interactions between vessel geometry, blood flow dynamics, and lesion morphology. The proposed method achieves state-of-the-art performance on a dataset of 778 exams from 389 patients. Unlike existing methods, our approach employs a single angiography view and does not require knowledge of the key frame; supervision at training time is provided by a classification loss (based on a threshold of the FFR/iFR values) and a regression loss for direct estimation. Finally, the analysis of model interpretability and calibration shows that, in spite of the complexity of angiographic imaging data, our method can robustly identify the location of the stenosis and correlate prediction uncertainty to the provided output scores. Raffaele Mineo, Federica Proietto Salanitri, Giovanni Bellitto, Isaak Kavasidis, Ovidio De Filippo, M. Millesimo, Gaetano Maria de Ferrari, Marco Aldinucci, Daniela Giordano, Simone Palazzo, Fabrizio D'Ascenzo, Concetto Spampinato |
IEEE Trans. Medical Imaging | 12 |
| 2024 | Guest Editorial Introduction to the Issue on Pre-Trained Models for Multi-Modality UnderstandingabstractIn the ever-evolving domain of multimedia, the significance of multi-modality understanding cannot be overstated. As multimedia content becomes increasingly sophisticated and ubiquitous, the ability to effectively combine and analyze the diverse information from different types of data, such as text, audio, image, video and point clouds, will be paramount in pushing the boundaries of what technology can achieve in understanding and interacting with the world around us. Accordingly, multi-modality understanding has attracted a tremendous amount of research, establishing itself as an emerging topic. Pre-trained models, in particular, have revolutionized this field, providing a way to leverage vast amounts of data without task-specific annotation to facilitate various downstream tasks. Wengang Zhou 0001, Jiajun Deng, Nicu Sebe, Qi Tian 0001, Alan L. Yuille, Concetto Spampinato, Zakia Hammal |
IEEE Trans. Multim. | 6 |
| 2023 | Dynamic Graph Attention: Unraveling Spatio-Temporal Synchrony in EEG DataabstractIn this paper, we propose a deep model based on graph convolutional networks for emotion recognition using EEG data. The model encodes spatial and temporal features of EEG channels and learns relationships between nodes through a self-attention mechanism, capturing spatio-temporal synchrony in brain regions. Experimental results show that our model outperforms existing approaches, with the attention mechanism contributing significantly to classification accuracy. In particular, the attention scores provide insights into how EEG channels influence each other at different times, revealing spatio-temporal patterns of brain connectivity related to emotions. Federica Proietto Salanitri, Giovanni Bellitto, Raffaele Mineo, Matteo Pennisi, Amelia Sorrenti, Salvatore Calcagno 0002, Daniela Giordano, Simone Palazzo, Concetto Spampinato |
BIBM | 9 |
| 2023 | Collective Driver Attention: Towards a Comprehensive Visual Understanding of Traffic ScenesabstractThanks to state-of-the-art deep learning-based methods for driver's attention prediction, it becomes possible to estimate where drivers look at in different traffic scenes. However, such estimation only takes into account visual information of the front view from a single vehicle. To remedy the lack of comprehensiveness of this approach, modern advanced driver-assistance systems (ADAS) further incorporate individual-specific features, including blood pressure and heart rate, to provide more precise safety advice. Nonetheless, there is still room for the improvement of safety-related recommendations by means of predicting collective drivers' attention. Specifically, the conceptual idea presented in this work is based on integrating visual understanding of the surrounding environment of a vehicle, driver-specific information and estimated attention in order to create a collective knowledge of the road, obstacles, distraction points, pedestrians and drivers' consciousness level from viewpoints of several drivers to predict a holistic attention map. Then, such a 360-degree attention map enables drivers to being aware not only of their front view, but also of back view and around (left and right) sides to help them prevent accidents and keep away from obstacles. The proposed framework takes advantages of edge and cloud computing for processing real-time information and large-scale computing, respectively. This work is intended to open a broader window towards the development of next generation of networked ADAS systems by employing several heterogeneous sources of information, implicit participation of drivers and their visual understanding and reasoning. Morteza Moradi 0001, Simone Palazzo, Concetto Spampinato |
CoDIT | 3 |
| 2023 | Ensemble and Personalized Transformer Models for Subject Identification and Relapse Detection in E-Prevention ChallengeabstractIn this short paper, we present the devised solutions for the subject identification and relapse detection tasks, which are part of the e-Prevention Challenge hosted at the ICASSP 2023 conference [1] [2] [3]. We specifically design an ensemble scheme of six models - five transformer-based ones and a CNN model - for the identification of subjects from wearable devices, while a personalized - one for each subject - scheme is used for relapse detection in psychotic disorder. Our final submitted solutions yield top performance on both tracks of the challenge: we ranked 2ndon the subject identification task (with an accuracy of 93.85%) and 1ston the relapse detection task (with a ROC-AUC and PR-AUC of about 0.65). Code and details are available at https://github.com/perceivelab/e-prevention-icassp-2023. Salvatore Calcagno 0002, Raffaele Mineo, Daniela Giordano, Concetto Spampinato |
ICASSP | 4 |
| 2023 | A Baseline on Continual Learning Methods for Video Action RecognitionabstractContinual learning has recently attracted attention from the research community, as it aims to solve long-standing limitations of classic supervised-trained models. However, most research on this subject has tackled continual learning in simple image classification scenarios. In this paper, we present a benchmark of state-of-the-art continual learning methods on video action recognition. Besides the increased complexity due to the temporal dimension, the video setting imposes stronger requirements on computing resources for top-performing rehearsal methods. To counteract the increased memory requirements, we present two method-agnostic variants for rehearsal methods, exploiting measures of either model confidence or data information to select memorable samples. Our experiments show that, as expected from the literature, rehearsal methods outperform other approaches; moreover, the proposed memory-efficient variants are shown to be effective at retaining a certain level of performance with a smaller buffer size. Giulia Castagnolo, Concetto Spampinato, Francesco Rundo, Daniela Giordano, Simone Palazzo |
ICIP | 2 |
| 2023 | A Privacy-Preserving Walk in the Latent Space of Generative Models for Medical Applications
Matteo Pennisi, Federica Proietto Salanitri, Giovanni Bellitto, Simone Palazzo, Ulas Bagci, Concetto Spampinato |
MICCAI (3) | 6 |
| 2023 | TinyHD: Efficient Video Saliency Prediction with Heterogeneous Decoders using Hierarchical Maps DistillationabstractVideo saliency prediction has recently attracted attention of the research community, as it is an upstream task for several practical applications. However, current solutions are particurly computationally demanding, especially due to the wide usage of spatio-temporal 3D convolutions. We observe that, while different model architectures achieve similar performance on benchmarks, visual variations between predicted saliency maps are still significant. Inspired by this intuition, we propose a lightweight model that employs multiple simple heterogeneous decoders and adopts several practical approaches to improve accuracy while keeping computational costs low, such as hierarchical multi-map knowledge distillation, multi-output saliency prediction, unlabeled auxiliary datasets and channel reduction with teacher assistant supervision. Our approach achieves saliency prediction accuracy on par or better than state-of-the-art methods on DFH1K, UCF-Sports and Hollywood2 benchmarks, while enhancing significantly the efficiency of the model. Feiyan Hu, Simone Palazzo, Federica Proietto Salanitri, Giovanni Bellitto, Morteza Moradi 0001, Concetto Spampinato, Kevin McGuinness |
WACV | 6 |
| 2023 | Transformer-based image generation from scene graphsabstractGraph-structured scene descriptions can be efficiently used in generative models to control the composition of the generated image. Previous approaches are based on the combination of graph convolutional networks and adversarial methods for layout prediction and image generation, respectively. In this work, we show how employing multi-head attention to encode the graph information, as well as using a transformer-based model in the latent space for image generation can improve the quality of the sampled data, without the need to employ adversarial models with the subsequent advantage in terms of training stability. The proposed approach, specifically, is entirely based on transformer architectures both for encoding scene graphs into intermediate object layouts and for decoding these layouts into images, passing through a lower dimensional space learned by a vector-quantized variational autoencoder. Our approach shows an improved image quality with respect to state-of-the-art methods as well as a higher degree of diversity among multiple generations from the same scene graph. We evaluate our approach on three public datasets: Visual Genome, COCO, and CLEVR. We achieve an Inception Score of 13.7 and 12.8, and an FID of 52.3 and 60.3, on COCO and Visual Genome, respectively. We perform ablation studies on our contributions to assess the impact of each component. Code is available at https://github.com/perceivelab/trf-sg2im. Renato Sortino, Simone Palazzo, Francesco Rundo, Concetto Spampinato |
Comput. Vis. Image Underst. | 4 |
| 2023 | MeT: A graph transformer for semantic segmentation of 3D meshesabstractPolygonal meshes have become the standard for discretely approximating 3D shapes, thanks to their efficiency and high flexibility in capturing non-uniform shapes. This non-uniformity, however, leads to irregularity in the mesh structure, making tasks like segmentation of 3D meshes particularly challenging. Semantic segmentation of 3D mesh has been typically addressed through CNN-based approaches, leading to good accuracy. Recently, transformers have gained enough momentum both in NLP and computer vision fields, achieving performance at least on par with CNN models, supporting the long-sought architecture universalism. Following this trend, we propose a transformer-based method for semantic segmentation of 3D mesh motivated by a better modeling of the graph structure of meshes, by means of global attention mechanisms. In order to address the limitations of standard transformer architectures in modeling relative positions of non-sequential data, as in the case of 3D meshes, as well as in capturing the local context, we perform positional encoding by means the Laplacian eigenvectors of the adjacency matrix, replacing the traditional sinusoidal positional encodings, and by introducing clustering-based features into the self-attention and cross-attention operators. Experimental results, carried out on three sets of the Shape COSEG Dataset (Wang et al., 2012), on the human segmentation dataset proposed in Maron et al. (2017) and on the ShapeNet benchmark (Chang et al., 2015), show how the proposed approach yields state-of-the-art performance on semantic segmentation of 3D meshes. Giuseppe Vecchio, Luca Prezzavento, Carmelo Pino, Francesco Rundo, Simone Palazzo, Concetto Spampinato |
Comput. Vis. Image Underst. | 6 |
| 2022 | Transfer Without Forgetting
Matteo Boschini, Lorenzo Bonicelli, Angelo Porrello, Giovanni Bellitto, Matteo Pennisi, Simone Palazzo, Concetto Spampinato, Simone Calderara |
ECCV (23) | 7 |
| 2022 | Transforming Image Generation from Scene GraphsabstractGenerating images from semantic visual knowledge is a challenging task, that can be useful to condition the synthesis process in complex, subtle, and unambiguous ways, compared to alternatives such as class labels or text descriptions. Although generative methods conditioned by semantic representations exist, they do not provide a way to control the generation process aside from the specification of constraints between objects. As an example, the possibility to iteratively generate or modify images by manually adding specific items is a desired property that, to our knowledge, has not been fully investigated in the literature. In this work we propose a transformer-based approach conditioned by scene graphs that, conversely to recent transformer-based methods, also employs a decoder to autoregressively compose images, making the synthesis process more effective and controllable. The proposed architecture is composed by three modules: 1) a graph convolutional network, to encode the relationships of the input graph; 2) an encoder-decoder transformer, which autoregressively composes the output image; 3) an auto-encoder, employed to generate representations used as input/output of each generation step by the transformer. Results obtained on CIFAR10 and MNIST images show that our model is able to satisfy semantic constraints defined by a scene graph and to model relations between visual objects in the scene by taking into account a user-provided partial rendering of the desired target. Renato Sortino, Simone Palazzo, Concetto Spampinato |
ICPR | 3 |
| 2022 | Deep Learning Car Driver Motion Magnified Saccadic Eye Movements for Advanced Driving Assistance SystemabstractAutomotive industry is making rapid progress in the development of next generation cars with higher levels of autonomy and intelligent assistance. Although the general advanced driver assistance system (ADAS) architecture is widely discussed, limited interaction between driver and these intelligent solutions sometimes make these approaches inefficient. For these reasons, the authors triggered an investigation about driver's feedback in relation to the assistance inputs provided by the ADAS technologies. In this context, the goal of this proposal is the design of an intelligent system that learns from the analysis of the car driver eyes saccadic movements, the correlated level of attention towards the salient driving scene. With this approach, we enabled a visual-feedback system which learns the driver eye's fixing dynamic associated to the analyzed driving scene. Through ad-hoc enhanced motion magnification technique, the authors were able to amplify the mentioned saccadic dynamics in order to allow a downstream deep classifier to associate this physiological behavior with the corresponding level of the driver attention. The collected performances (over 97%) confirmed the effectiveness of the proposed method. Francesco Rundo, Angelo Alberto Messina, Michele Calabretta, Matteo Dilonardo, Salvatore Coffa, Concetto Spampinato |
IJCNN | 6 |
| 2022 | On the Effectiveness of Lipschitz-Driven Rehearsal in Continual LearningabstractRehearsal approaches enjoy immense popularity with Continual Learning (CL) practitioners. These methods collect samples from previously encountered data distributions in a small memory buffer; subsequently, they repeatedly optimize on the latter to prevent catastrophic forgetting. This work draws attention to a hidden pitfall of this widespread practice: repeated optimization on a small pool of data inevitably leads to tight and unstable decision boundaries, which are a major hindrance to generalization. To address this issue, we propose Lipschitz-DrivEn Rehearsal (LiDER), a surrogate objective that induces smoothness in the backbone network by constraining its layer-wise Lipschitz constants w.r.t. replay examples. By means of extensive experiments, we show that applying LiDER delivers a stable performance gain to several state-of-the-art rehearsal CL methods across multiple datasets, both in the presence and absence of pre-training. Through additional ablative experiments, we highlight peculiar aspects of buffer overfitting in CL and better characterize the effect produced by LiDER. Code is available at https://github.com/aimagelab/LiDER. Lorenzo Bonicelli, Matteo Boschini, Angelo Porrello, Concetto Spampinato, Simone Calderara |
NeurIPS | 4 |
| 2022 | Distributed workflows with JupyterabstractThe designers of a new coordination interface enacting complex workflows have to tackle a dichotomy: choosing a language-independent or language-dependent approach. Language-independent approaches decouple workflow models from the host code’s business logic and advocate portability. Language-dependent approaches foster flexibility and performance by adopting the same host language for business and coordination code. Jupyter Notebooks, with their capability to describe both imperative and declarative code in a unique format, allow taking the best of the two approaches, maintaining a clear separation between application and coordination layers but still providing a unified interface to both aspects. We advocate the Jupyter Notebooks’ potential to express complex distributed workflows, identifying the general requirements for a Jupyter-based Workflow Management System (WMS) and introducing a proof-of-concept portable implementation working on hybrid Cloud-HPC infrastructures. As a byproduct, we extended the vanilla IPython kernel with workflow-based parallel and distributed execution capabilities. The proposed Jupyter-workflow (Jw) system is evaluated on common scenarios for High Performance Computing (HPC) and Cloud, showing its potential in lowering the barriers between prototypical Notebooks and production-ready implementations. Iacopo Colonnelli, Marco Aldinucci, Barbara Cantalupo, Luca Padovani, Sergio Rabellino, Concetto Spampinato, Roberto Morelli, Rosario Di Carlo, Nicolò Magini, Carlo Cavazzoni |
Future Gener. Comput. Syst. | 6 |
| 2021 | SurfaceNet: Adversarial SVBRDF Estimation from a Single ImageabstractIn this paper we present SurfaceNet, an approach for estimating spatially-varying bidirectional reflectance distribution function (SVBRDF) material properties from a single image. We pose the problem as an image translation task and propose a novel patch-based generative adversarial network (GAN) that is able to produce high-quality, high-resolution surface reflectance maps. The employment of the GAN paradigm has a twofold objective: 1) allowing the model to recover finer details than standard translation models; 2) reducing the domain shift between synthetic and real data distributions in an unsupervised way.An extensive evaluation, carried out on a public benchmark of synthetic and real images under different illumination conditions, shows that SurfaceNet largely outperforms existing SVBRDF reconstruction methods, both quantitatively and qualitatively. Furthermore, SurfaceNet exhibits a remarkable ability in generating high-quality maps from real samples without any supervision at training time.Source code available at https://github.com/perceivelab/surfacenet. Giuseppe Vecchio, Simone Palazzo, Concetto Spampinato |
ICCV | 3 |
| 2021 | An explainable AI system for automated COVID-19 assessment and lesion categorization from CT-scans
Matteo Pennisi, Isaak Kavasidis, Concetto Spampinato, Vincenzo Schininà, Simone Palazzo, Federica Proietto Salanitri, Giovanni Bellitto, Francesco Rundo, Marco Aldinucci, Massimo Cristofaro, Paolo Campioni, Elisa Pianura, Federica Di Stefano 0002, Ada Petrone, Fabrizio Albarello, Giuseppe Ippolito, Salvatore Cuzzocrea, Sabrina Conoci |
Artif. Intell. Medicine | 3 |
| 2021 | Hierarchical Domain-Adapted Feature Learning for Video Saliency PredictionabstractAbstract In this work, we propose a 3D fully convolutional architecture for video saliency prediction that employs hierarchical supervision on intermediate maps (referred to as conspicuity maps) generated using features extracted at different abstraction levels. We provide the base hierarchical learning mechanism with two techniques for domain adaptation and domain-specific learning. For the former, we encourage the model to unsupervisedly learn hierarchical general features using gradient reversal at multiple scales, to enhance generalization capabilities on datasets for which no annotations are provided during training. As for domain specialization, we employ domain-specific operations (namely, priors, smoothing and batch normalization) by specializing the learned features on individual datasets in order to maximize performance. The results of our experiments show that the proposed model yields state-of-the-art accuracy on supervised saliency prediction. When the base hierarchical model is empowered with domain-specific modules, performance improves, outperforming state-of-the-art models on three out of five metrics on the DHF1K benchmark and reaching the second-best results on the other two. When, instead, we test it in an unsupervised domain adaptation setting, by enabling hierarchical gradient reversal layers, we obtain performance comparable to supervised state-of-the-art. Source code, trained models and example outputs are publicly available at https://github.com/perceivelab/hd2s . Giovanni Bellitto, Federica Proietto Salanitri, Simone Palazzo, Francesco Rundo, Daniela Giordano, Concetto Spampinato |
Int. J. Comput. Vis. | 6 |
| 2021 | Decoding Brain Representations by Multimodal Learning of Neural Activity and Visual FeaturesabstractThis work presents a novel method of exploring human brain-visual representations, with a view towards replicating these processes in machines. The core idea is to learn plausible computational and biological representations by correlating human neural activity and natural images. Thus, we first propose a model, EEG-ChannelNet, to learn a brain manifold for EEG classification. After verifying that visual information can be extracted from EEG data, we introduce a multimodal approach that uses deep image and EEG encoders, trained in a siamese configuration, for learning a joint manifold that maximizes a compatibility measure between visual features and brain representations. We then carry out image classification and saliency detection on the learned manifold. Performance analyses show that our approach satisfactorily decodes visual information from neural signals. This, in turn, can be used to effectively supervise the training of deep learning models, as demonstrated by the high performance of image classification and saliency detection on out-of-training classes. The obtained results show that the learned brain-visual features lead to improved performance and simultaneously bring deep models more in line with cognitive neuroscience work related to visual perception and attention. Simone Palazzo, Concetto Spampinato, Isaak Kavasidis, Daniela Giordano, Joseph Schmidt, Mubarak Shah |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Exploiting structured high-level knowledge for domain-specific visual classification
Simone Palazzo, Francesca Murabito, Carmelo Pino, Francesco Rundo, Daniela Giordano, Mubarak Shah, Concetto Spampinato |
Pattern Recognit. | 7 |
| 2020 | Visual Saliency Detection guided by Neural SignalsabstractSaliency detection is a fundamental process of human visual perception, since it allows us to identify the most important parts of a scene, directing our analysis and interpretation capabilities on a reduced set of information and reducing reaction times. However, current approaches for automatic saliency detection either attempt to mimic human capabilities by building attention maps from hand-crafted feature analysis, or employ convolutional neural networks trained as black boxes, without any architectural or information prior from human biology. In this paper, we present an approach for saliency detection that combines the success of deep learning in identifying representations for visual data with a training paradigm aimed at matching neural activity provided directly by brain signals recorded while subjects look at images. We show that our approach is able to capture correspondences between visual elements and neural activities, successfully generalizing to unseen images to identify their most salient regions. Simone Palazzo, Francesco Rundo, Sebastiano Battiato, Daniela Giordano, Concetto Spampinato |
FG | 5 |
| 2020 | Deep Recurrent-Convolutional Model for Automated Segmentation of Craniomaxillofacial CT ScansabstractIn this paper we define a deep learning architecture for automated segmentation of anatomical structures in Craniomaxillofacial (CMF) CT scans that leverages the recent success of encoder-decoder models for semantic segmentation of natural images. In particular, we propose a fully convolutional deep network that combines the advantages of recent fully convolutional models, such as Tiramisu, with squeeze-and-excitation blocks for feature recalibration, integrated with convolutional LSTMs to model spatio-temporal correlations between consecutive slices. The proposed segmentation network shows superior performance and generalization capabilities (to different structures and imaging modalities) than state of the art methods on automated segmentation of CMF structures (e.g., mandibles and airways) in several standard benchmarks (e.g., MICCAI datasets) and on new datasets proposed herein, effectively facing shape variability. Francesca Murabito, Simone Palazzo, Federica Proietto Salanitri, Francesco Rundo, Ulas Bagci, Daniela Giordano, Rosalia Leonardi, Concetto Spampinato |
ICPR | 8 |
| 2020 | Deep Multi-stage Model for Automated Landmarking of Craniomaxillofacial CT ScansabstractIn this paper we define a deep multi-stage architecture for automated landmarking of craniomaxillofacial (CMF) CT images. Our model is composed of three subnetworks that first localize, on reduced-resolution images, areas where landmarks may be found and then refine the search, at full-resolution scale, through a hierarchical structure aiming at increasing the granularity of the investigated region. The multi-stage pipeline is designed to deal with full resolution data and does not require any additional pre-processing step to reduce search space, as opposed to existing methods that can be only adopted for searching landmarks located in well-defined anatomical structures (e.g., mandibles). The automated landmarking system is tested on identifying landmarks located in several CMF regions, achieving an average error of 0.8 mm, significantly lower than expert readings. The proposed model also outperforms baselines and is on par with existing models that employ additional upstream segmentation, on state-of-the-art benchmarks. Simone Palazzo, Giovanni Bellitto, Luca Prezzavento, Francesco Rundo, Ulas Bagci, Daniela Giordano, Rosalia Leonardi, Concetto Spampinato |
ICPR | 8 |
| 2020 | Domain Adaptation for Outdoor Robot Traversability Estimation from RGB data with Safety-Preserving LossabstractBeing able to estimate the traversability of the area surrounding a mobile robot is a fundamental task in the design of a navigation algorithm. However, the task is often complex, since it requires evaluating distances from obstacles, type and slope of terrain, and dealing with non-obvious discontinuities in detected distances due to perspective. In this paper, we present an approach based on deep learning to estimate and anticipate the traversing score of different routes in the field of view of an on-board RGB camera. The backbone of the proposed model is based on a state-of-the-art deep segmentation model, which is fine-tuned on the task of predicting route traversability. We then enhance the model's capabilities by a) addressing domain shifts through gradient-reversal unsupervised adaptation, and b) accounting for the specific safety requirements of a mobile robot, by encouraging the model to err on the safe side, i.e., penalizing errors that would cause collisions with obstacles more than those that would cause the robot to stop in advance. Experimental results show that our approach is able to satisfactorily identify traversable areas and to generalize to unseen locations. Simone Palazzo, Dario C. Guastella, Luciano Cantelli, Paolo Spadaro, Francesco Rundo, Giovanni Muscato, Daniela Giordano, Concetto Spampinato |
IROS | 8 |
| 2020 | Adversarial Framework for Unsupervised Learning of Motion Dynamics in Videos
Concetto Spampinato, Simone Palazzo, P. D'Oro, Daniela Giordano, Mubarak Shah |
Int. J. Comput. Vis. | 1 |
| 2020 | MASK-RL: Multiagent Video Object Segmentation Framework Through Reinforcement LearningabstractIntegrating human-provided location priors into video object segmentation has been shown to be an effective strategy to enhance performance, but their application at large scale is unfeasible. Gamification can help reduce the annotation burden, but it still requires user involvement. We propose a video object segmentation framework that leverages the combined advantages of user feedback for segmentation and gamification strategy by simulating multiple game players through a reinforcement learning (RL) model that reproduces human ability to pinpoint moving objects and using the simulated feedback to drive the decisions of a fully convolutional deep segmentation network. Experimental results on the DAVIS-17 benchmark show that: 1) including user-provided prior, even if not precise, yields high performance; 2) our RL agent replicates satisfactorily the same variability of humans in identifying spatiotemporal salient objects; and 3) employing artificially generated priors in an unsupervised video object segmentation model reaches state-of-the-art performance. Giuseppe Vecchio, Simone Palazzo, Daniela Giordano, Francesco Rundo, Concetto Spampinato |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2018 | ThoughtViz: Visualizing Human Thoughts Using Generative Adversarial NetworkabstractStudying human brain signals has always gathered great attention from the scientific community. In Brain Computer Interface (BCI) research, for example, changes of brain signals in relation to specific tasks (e.g., thinking something) are detected and used to control machines. While extracting spatio-temporal cues from brain signals for classifying state of human mind is an explored path, decoding and visualizing brain states is new and futuristic. Following this latter direction, in this paper, we propose an approach that is able not only to read the mind, but also to decode and visualize human thoughts. More specifically, we analyze brain activity, recorded by an ElectroEncephaloGram (EEG), of a subject while thinking about a digit, character or an object and synthesize visually the thought item. To accomplish this, we leverage the recent progress of adversarial learning by devising a conditional Generative Adversarial Network (GAN), which takes, as input, encoded EEG signals and generates corresponding images. In addition, since collecting large EEG signals in not trivial, our GAN model allows for learning distributions with limited training data. Performance analysis carried out on three different datasets -- brain signals of multiple subjects thinking digits, characters, and objects -- show that our approach is able to effectively generate images from thoughts of a person. They also demonstrate that EEG signals encode explicitly cues from thoughts which can be effectively used for generating semantically relevant visualizations. Praveen Tirupattur, Yogesh S. Rawat, Concetto Spampinato, Mubarak Shah |
ACM Multimedia | 3 |
| 2018 | Top-down saliency detection driven by visual classification
Francesca Murabito, Concetto Spampinato, Simone Palazzo, Daniela Giordano, Konstantin Pogorelov, Michael Riegler 0001 |
Comput. Vis. Image Underst. | 2 |
| 2017 | An Eye Tracker based Computer System to Support Oculomotor and Attention Deficit InvestigationsabstractEye tracking is a non-invasive procedure to acquire eye-gaze data. The accuracy offered by the new eye tracking technologies gives to physicians and scientists a great opportunity to employ eye trackers to perform quantitative assessment of eye movements for diagnostic and rehabilitation purposes. However, eye trackers do not support physicians in their analysis, as they typically lack specific software solutions tailored to the diseases under investigation. For instance, ophthalmologists need to use eye trackers in tests such as visually guided saccades or smooth pursuit eye movements, while neuropsychiatrists need them in psychological tests to measure attention deficit. The development of a software tool for a specific test cannot be done by physicians (even using the high-level programming frameworks bundled with eye tracker devices) and needs to be carried out by expert computer programmers. Thus, in this paper we propose a general computer system for eye tracker - based investigation, which allows physicians to easily customize experiments in different scenarios without the need of expert computer programmers. To demonstrate the generalization capabilities of the proposed system, we show how it was employed to support the investigation of eye movements for patients affected by a) glycogen storage disease, b) idiopathic congenital nystagmus and c) neurodevelopmental diseases such as autism spectrum disorders. Daniela Giordano, Carmelo Pino, Isaak Kavasidis, Concetto Spampinato, Massimo Di Pietro, Renata Rizzo, Anna Scuderi, Rita Barone |
CBMS | 4 |
| 2017 | Deep Learning Human Mind for Automated Visual ClassificationabstractxWhat if we could effectively read the mind and transfer human visual capabilities to computer vision methods? In this paper, we aim at addressing this question by developing the first visual object classifier driven by human brain signals. In particular, we employ EEG data evoked by visual object stimuli combined with Recurrent Neural Networks (RNN) to learn a discriminative brain activity manifold of visual categories in a reading the mind effort. Afterward, we transfer the learned capabilities to machines by training a Convolutional Neural Network (CNN)-based regressor to project images onto the learned manifold, thus allowing machines to employ human brain-based features for automated visual classification. We use a 128-channel EEG with active electrodes to record brain activity of several subjects while looking at images of 40 ImageNet object classes. The proposed RNN-based approach for discriminating object classes using brain signals reaches an average accuracy of about 83%, which greatly outperforms existing methods attempting to learn EEG visual object representations. As for automated object categorization, our human brain-driven approach obtains competitive performance, comparable to those achieved by powerful CNN models and it is also able to generalize over different visual datasets. Concetto Spampinato, Simone Palazzo, Isaak Kavasidis, Daniela Giordano, Nasim Souly, Mubarak Shah |
CVPR | 1 |
| 2017 | Generative Adversarial Networks Conditioned by Brain SignalsabstractRecent advancements in generative adversarial networks (GANs), using deep convolutional models, have supported the development of image generation techniques able to reach satisfactory levels of realism. Further improvements have been proposed to condition GANs to generate images matching a specific object category or a short text description. In this work, we build on the latter class of approaches and investigate the possibility of driving and conditioning the image generation process by means of brain signals recorded, through an electroencephalograph (EEG), while users look at images from a set of 40 ImageNet object categories with the objective of generating the seen images. To accomplish this task, we first demonstrate that brain activity EEG signals encode visually-related information that allows us to accurately discriminate between visual object categories and, accordingly, we extract a more compact class-dependent representation of EEG data using recurrent neural networks. Afterwards, we use the learned EEG manifold to condition image generation employing GANs, which, during inference, will read EEG signals and convert them into images. We tested our generative approach using EEG signals recorded from six subjects while looking at images of the aforementioned 40 visual classes. The results show that for classes represented by well-defined visual patterns (e.g., pandas, airplane, etc.), the generated images are realistic and highly resemble those evoking the EEG signals used for conditioning GANs, resulting in an actual reading-the-mind process. Simone Palazzo, Concetto Spampinato, Isaak Kavasidis, Daniela Giordano, Mubarak Shah |
ICCV | 2 |
| 2017 | Semi Supervised Semantic Segmentation Using Generative Adversarial NetworkabstractSemantic segmentation has been a long standing challenging task in computer vision. It aims at assigning a label to each image pixel and needs a significant number of pixel-level annotated data, which is often unavailable. To address this lack of annotations, in this paper, we leverage, on one hand, a massive amount of available unlabeled or weakly labeled data, and on the other hand, non-real images created through Generative Adversarial Networks. In particular, we propose a semi-supervised framework - based on Generative Adversarial Networks (GANs) - which consists of a generator network to provide extra training examples to a multi-class classifier, acting as discriminator in the GAN framework, that assigns sample a label y from the K possible classes or marks it as a fake sample (extra class). The underlying idea is that adding large fake visual data forces real samples to be close in the feature space, which, in turn, improves multiclass pixel classification. To ensure a higher quality of generated images by GANs with consequently improved pixel classification, we extend the above framework by adding weakly annotated data, i.e., we provide class level information to the generator. We test our approaches on several challenging benchmarking visual datasets, i.e. PASCAL, SiftFLow, Stanford and CamVid, achieving competitive performance compared to state-of-the-art semantic segmentation methods. Nasim Souly, Concetto Spampinato, Mubarak Shah |
ICCV | 2 |
| 2017 | Brain2Image: Converting Brain Signals into ImagesabstractReading the human mind has been a hot topic in the last decades, and recent research in neuroscience has found evidence on the possibility of decoding, from neuroimaging data, how the human brain works. At the same time, the recent rediscovery of deep learning combined to the large interest of scientific community on generative methods has enabled the generation of realistic images by learning a data distribution from noise. The quality of generated images increases when the input data conveys information on visual content of images. Leveraging on these recent trends, in this paper we present an approach for generating images using visually-evoked brain signals recorded through an electroencephalograph (EEG). More specifically, we recorded EEG data from several subjects while observing images on a screen and tried to regenerate the seen images. To achieve this goal, we developed a deep-learning framework consisting of an LSTM stacked with a generative method, which learns a more compact and noise-free representation of EEG data and employs it to generate the visual stimuli evoking specific brain responses. Isaak Kavasidis, Simone Palazzo, Concetto Spampinato, Daniela Giordano, Mubarak Shah |
ACM Multimedia | 3 |
| 2017 | A Holistic Multimedia System for Gastrointestinal Tract Disease DetectionabstractAnalysis of medical videos for detection of abnormalities and diseases requires both high precision and recall, but also real-time processing for live feedback and scalability for massive screening of entire populations. Existing work on this field does not provide the necessary combination of retrieval accuracy and performance.; [email protected] this paper, a multimedia system is presented where the aim is to tackle automatic analysis of videos from the human gastrointestinal (GI) tract. The system includes the whole pipeline from data collection, processing and analysis, to visualization. The system combines filters using machine learning, image recognition and extraction of global and local image features. Furthermore, it is built in a modular way so that it can easily be extended. At the same time, it is developed for efficient processing in order to provide real-time feedback to the doctors. Our experimental evaluation proves that our system has detection and localisation accuracy at least as good as existing systems for polyp detection, it is capable of detecting a wider range of diseases, it can analyze video in real-time, and it has a low resource consumption for scalability. Konstantin Pogorelov, Sigrun Losada Eskeland, Thomas de Lange, Carsten Griwodz, Kristin Ranheim Randel, Håkon Kvale Stensland, Duc-Tien Dang-Nguyen, Concetto Spampinato, Dag Johansen, Michael Riegler 0001, Pål Halvorsen |
MMSys | 8 |
| 2017 | KVASIR: A Multi-Class Image Dataset for Computer Aided Gastrointestinal Disease DetectionabstractAutomatic detection of diseases by use of computers is an important, but still unexplored field of research. Such innovations may improve medical practice and refine health care systems all over the world. However, datasets containing medical images are hardly available, making reproducibility and comparison of approaches almost impossible. In this paper, we present KVASIR, a dataset containing images from inside the gastrointestinal (GI) tract. The collection of images are classified into three important anatomical landmarks and three clinically significant findings. In addition, it contains two categories of images related to endoscopic polyp removal. Sorting and annotation of the dataset is performed by medical doctors (experienced endoscopists). In this respect, KVASIR is important for research on both single- and multi-disease computer aided detection. By providing it, we invite and enable multimedia researcher into the medical domain of detection and retrieval. Konstantin Pogorelov, Kristin Ranheim Randel, Carsten Griwodz, Sigrun Losada Eskeland, Thomas de Lange, Dag Johansen, Concetto Spampinato, Duc-Tien Dang-Nguyen, Mathias Lux, Peter Thelin Schmidt, Michael Riegler 0001, Pål Halvorsen |
MMSys | 7 |
| 2017 | Nerthus: A Bowel Preparation Quality Video DatasetabstractBowel preparation (cleansing) is considered to be a key precondition for successful colonoscopy (endoscopic examination of the bowel). The degree of bowel cleansing directly affects the possibility to detect diseases and may influence decisions on screening and follow-up examination intervals. An accurate assessment of bowel preparation quality is therefore important. Despite the use of reliable and validated bowel preparation scales, the grading may vary from one doctor to another. An objective and automated assessment of bowel cleansing would contribute to reduce such inequalities and optimize use of medical resources. This would also be a valuable feature for automatic endoscopy reporting in the future. In this paper, we present Nerthus, a dataset containing videos from inside the gastrointestinal (GI) tract, showing different degrees of bowel cleansing. By providing this dataset, we invite multimedia researchers to contribute in the medical field by making systems automatically evaluate the quality of bowel cleansing for colonoscopy. Such innovations would probably contribute to improve the medical field of GI endoscopy. Konstantin Pogorelov, Kristin Ranheim Randel, Thomas de Lange, Sigrun Losada Eskeland, Carsten Griwodz, Dag Johansen, Concetto Spampinato, Mario Taschwer, Mathias Lux, Peter Thelin Schmidt, Michael Riegler 0001, Pål Halvorsen |
MMSys | 7 |
| 2017 | Deep learning for automated skeletal bone age assessment in X-ray images
Concetto Spampinato, Simone Palazzo, Daniela Giordano, Marco Aldinucci, Rosalia Leonardi |
Medical Image Anal. | 1 |
| 2017 | Gamifying Video Object SegmentationabstractVideo object segmentation can be considered as one of the most challenging computer vision problems. Indeed, so far, no existing solution is able to effectively deal with the peculiarities of real-world videos, especially in cases of articulated motion and object occlusions; limitations that appear more evident when we compare the performance of automated methods with the human one. However, manually segmenting objects in videos is largely impractical as it requires a lot of time and concentration. To address this problem, in this paper we propose an interactive video object segmentation method, which exploits, on one hand, the capability of humans to identify correctly objects in visual scenes, and on the other hand, the collective human brainpower to solve challenging and large-scale tasks. In particular, our method relies on a game with a purpose to collect human inputs on object locations, followed by an accurate segmentation phase achieved by optimizing an energy function encoding spatial and temporal constraints between object regions as well as human-provided location priors. Performance analysis carried out on complex video benchmarks, and exploiting data provided by over 60 users, demonstrated that our method shows a better trade-off between annotation times and segmentation accuracy than interactive video annotation and automated video object segmentation approaches. Concetto Spampinato, Simone Palazzo, Daniela Giordano |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | An ontological ubiquitous city information platform provided with Cyber-Physical-Social-SystemsabstractAim of the paper is to illustrate the methodology used to implement an ubiquitous city platform called Wi-City-Plus provided with mobile and centralized Decision Support Systems (DSSs) taking advantage of all the data of city interest, including social data, and those sensed by networked ambient Cyber-Physical-Systems (CPSs). The paper proposes to model such data by a suitable RDF ontology specifically designed to support the main user scenarios in smart cities and to implement both the technical and the service interoperability envisaged by the CPS ideal model. This ontology is obtained by reusing existing ontologies and allows the DSSs to access by SPARQL queries all the relevant data independent of the proprietary technology of the CPSs and of the DBMS adopted by the data owners. Examples derived from the pilot version of the proposed platform allows us to point out advantages and perspectives of our approach. Alfio Costanzo, Alberto Faro, Daniela Giordano, Concetto Spampinato |
CCNC | 4 |
| 2016 | Implementing Cyber Physical social Systems for smart cities: A semantic web perspectiveabstractAim of the paper is to give a demonstration of the main Cyber-Physical-Systems (CPSs) and social networking facilities used to implement an ubiquitous city platform called Wi-City-Plus provided with mobile and centralized decision support systems (DSSs) taking advantage of all the data of city interest, including social data, and those sensed by networked monitoring devices. In particular, the paper adopts a semantic web approach that allow us to model such data by a suitable ontology specifically designed to support the main user scenarios in smart cities and to implement both the technical interoperability and the service integration envisaged by the CPS ideal model. Alfio Costanzo, Alberto Faro, Daniela Giordano, Concetto Spampinato |
CCNC | 4 |
| 2016 | GeoSentiment: A tool for analyzing geographically distributed event-related sentimentsabstractIn this demo paper we present GeoSentiment, a tool for effective assessment and visualization of event-related sentiments in geographically confined populations. GeoSentiment is developed as a web tool application and provides to stakeholders an easy-to-use and powerful means to investigate how events are perceived by people and which factors may influence such perception. GeoSentiment relies on different services for a) retrieving and mining official statistical information as well as the most-common social networks and b) performing sentiment analysis. It is provided with an interactive interface, which enables rendering and deep exploration of all the processed data and results. Carmelo Pino, Isaak Kavasidis, Concetto Spampinato |
CCNC | 3 |
| 2016 | Assessment and visualization of geographically distributed event-related sentiments by mining social networks and newsabstractIn this paper we present a method/tool for integrating heterogeneous data, assessing and visualizing sentiments related to big impact events in geographically confined populations. The data employed are official statistical information provided by governments, news web sites and user submitted georeferenced comments retrieved from various social networks. Sentiment analysis is applied to the retrieved comments and results are visualized on interactive maps, thus providing an effective tool for decision makers and analysts to evaluate how socioeconomic factors can influence the mood and the opinion of a specific area. The method is designed to assist city planners, business managers or social scientists for strategic planning and decision making. Furthermore, the experimental evaluation showed that the proposed method is robust, achieves state-of-the-art performance and allows an easy-exploration of big, distributed and heterogeneous information. Carmelo Pino, Isaak Kavasidis, Concetto Spampinato |
CCNC | 3 |
| 2016 | Multimedia and Medicine: Teammates for Better Disease Detection and SurvivalabstractHealth care has a long history of adopting technology to save lives and improve the quality of living. Visual information is frequently applied for disease detection and assessment, and the established fields of computer vision and medical imaging provide essential tools. It is, however, a misconception that disease detection and assessment are provided exclusively by these fields and that they provide the solution for all challenges. Integration and analysis of data from several sources, real-time processing, and the assessment of usefulness for end-users are core competences of the multimedia community and are required for the successful improvement of health care systems. We have conducted initial investigations into two use cases surrounding diseases of the gastrointestinal (GI) tract, where the detection of abnormalities provides the largest chance of successful treatment if the initial observation of disease indicators occurs before the patient notices any symptoms. Although such detection is typically provided visually by applying an endoscope, we are facing a multitude of new multimedia challenges that differ between use cases. In real-time assistance for colonoscopy, we combine sensor information about camera position and direction to aid in detecting, investigate means for providing support to doctors in unobtrusive ways, and assist in reporting. In the area of large-scale capsular endoscopy, we investigate questions of scalability, performance and energy efficiency for the recording phase, and combine video summarization and retrieval questions for analysis. Michael Riegler 0001, Mathias Lux, Carsten Griwodz, Concetto Spampinato, Thomas de Lange, Sigrun Losada Eskeland, Konstantin Pogorelov, Wallapak Tavanapong, Peter Thelin Schmidt, Cathal Gurrin, Dag Johansen, Håvard D. Johansen, Pål Halvorsen |
ACM Multimedia | 4 |
| 2016 | Right inflight?: a dataset for exploring the automatic prediction of movies suitable for a watching situationabstractIn this paper, we present the dataset Right Inflight developed to support the exploration of the match between video content and the situation in which that content is watched. Specifically, we look at videos that are suitable to be watched on an airplane, where the main assumption is that that viewers watch movies with the intent of relaxing themselves and letting time pass quickly, despite the inconvenience and discomfort of flight. The aim of the dataset is to support the development of recommender systems, as well as computer vision and multimedia retrieval algorithms capable of automatically predicting which videos are suitable for inflight consumption. Our ultimate goal is to promote a deeper understanding of how people experience video content, and of how technology can support people in finding or selecting video content that supports them in regulating their internal states in certain situations. Right Inflight consists of 318 human-annotated movies, for which we provide links to trailers, a set of pre-computed low-level visual, audio and text features as well as user ratings. The annotation was performed by crowdsourcing workers, who were asked to judge the appropriateness of movies for inflight consumption. Michael Riegler 0001, Martha A. Larson, Concetto Spampinato, Pål Halvorsen, Mathias Lux, Jonas Markussen, Konstantin Pogorelov, Carsten Griwodz, Håkon Kvale Stensland |
MMSys | 3 |
| 2016 | Generating reliable video annotations by exploiting the crowdabstractIn computer vision and machine learning, the availability of annotated datasets is of crucial importance for both learning and performance evaluation. However, annotating visual datasets is a tedious and error-prone task and computer vision researchers usually dedicate a large amount of their time for collecting and generating annotations, which most of the time cannot be re-used in other scenarios. In this paper, we propose a simple, but effective, interactive video object segmentation method exploiting large noisy data gathered from crowd of users while playing a web game. Experimental results, carried out on two challenging video benchmarks, show how it is possible to generate reliable object segmentations in videos with a small human effort, achieving an accuracy comparable to the one obtained with manually-labeled annotations and also outperforming state-of-the-art video object segmentation approaches. Roberto Di Salvo, Concetto Spampinato, Daniela Giordano |
WACV | 2 |
| 2016 | A diversity-based search approach to support annotation of a large fish image dataset
Daniela Giordano, Simone Palazzo, Concetto Spampinato |
Multim. Syst. | 3 |
| 2016 | Special issue on multimedia in ecology
Concetto Spampinato, Vasileios Mezaris, Jacco van Ossenbruggen |
Multim. Syst. | 1 |
| 2016 | Fine-grained object recognition in underwater visual data
Concetto Spampinato, Simone Palazzo, Pierre-Hugues Joalland, Sébastien Paris, Hervé Glotin, Katy Blanc, Diane Lingrand, Frédéric Precioso |
Multim. Tools Appl. | 1 |
| 2016 | Special issue on "Fine-grained categorization in ecological multimedia"
Concetto Spampinato, Vasileios Mezaris, Marco Cristani |
Pattern Recognit. Lett. | 1 |
| 2015 | Rejecting False Positives in Video Object Segmentation
Daniela Giordano, Isaak Kavasidis, Simone Palazzo, Concetto Spampinato |
CAIP (1) | 4 |
| 2015 | Automatic Summary Creation by Applying Natural Language Processing on Unstructured Medical Records
Daniela Giordano, Isaak Kavasidis, Concetto Spampinato |
CAIP (2) | 3 |
| 2015 | Superpixel-based video object segmentation using perceptual organization and location priorabstractIn this paper we present an approach for segmenting objects in videos taken in complex scenes with multiple and different targets. The method does not make any specific assumptions about the videos and relies on how objects are perceived by humans according to Gestalt laws. Initially, we rapidly generate a coarse foreground segmentation, which provides predictions about motion regions by analyzing how superpixel segmentation changes in consecutive frames. We then exploit these location priors to refine the initial segmentation by optimizing an energy function based on appearance and perceptual organization, only on regions where motion is observed. We evaluated our method on complex and challenging video sequences and it showed significant performance improvements over recent state-of-the-art methods, being also fast enough to be used for “on-the-fly” processing. Daniela Giordano, Francesca Murabito, Simone Palazzo, Concetto Spampinato |
CVPR | 4 |
| 2015 | Using the Eyes to "See" the ObjectsabstractThis paper investigates how to exploit eye gaze data for understanding visual content. In particular, we propose a human-in-the-loop approach for object segmentation in videos, where humans provide significant cues on spatiotemporal relations between object parts (i.e. superpixels in our approach) by simply looking at video sequences. Such constraints, together with object appearance properties, are encoded into an energy function so as to tackle the segmentation problem as a labeling one. The proposed method uses gaze data from only two people and was tested on two challenging visual benchmarks: 1) SegTrack v2 and 2) FBMS-59. The achieved performance showed how our method outperformed more complex video object segmentation approaches, while reducing the effort needed for collecting human feedback Concetto Spampinato, Simone Palazzo, Francesca Murabito, Daniela Giordano |
ACM Multimedia | 1 |
| 2015 | Nonparametric label propagation using mutual local similarity in nearest neighbors
Daniela Giordano, Isaak Kavasidis, Simone Palazzo, Concetto Spampinato |
Comput. Vis. Image Underst. | 4 |
| 2014 | Kernel Density Estimation Using Joint Spatial-Color-Depth Data for Background ModelingabstractThe use of low-cost devices for depth estimation, such as Microsoft Kinect, is becoming more and more popular in computer vision research. In this paper, we propose an algorithm for background modeling which exploits this kind of devices to make the background and foreground models more robust to effects such as camouflage and illumination changes. Our algorithm, after a preprocessing stage for aligning color and depth data and for filtering/filling noisy depth measurements, explicitly models the scene's background and foreground with a Kernel Density Estimation approach in a quantized x-y-hue-saturation-depth space. The results in three different indoor environments, with different lighting conditions, showed that our approach is able to achieve an accuracy in foreground segmentation over 90% and that the combination of depth data and illumination-independent color space proved to be very robust against noise and illumination changes. Daniela Giordano, Simone Palazzo, Concetto Spampinato |
ICPR | 3 |
| 2014 | Summary Abstract for the 3rd ACM International Workshop on Multimedia Analysis for Ecological DataabstractThe 3rd ACM International Workshop on Multimedia Anal- ysis for Ecological Data (MAED'14) is held as part of ACM Multimedia 2014. Concetto Spampinato, Vasileios Mezaris, Marco Cristani |
ACM Multimedia | 1 |
| 2014 | Large Scale Data Processing in Ecology: A Case Study on Long-Term Underwater Video MonitoringabstractEcology is, nowadays, an interdisciplinary, collabo- rative and data-intensive science, therefore, discovering, integrat- ing and analysing daily-produced data is necessary to support researchers to investigate complex questions, ranging from single particles to animals to the biosphere [1]. As a consequence, ecology-related multimedia content has been produced massively in recent years: for example, the Xeno-canto project1 and the Pl@ntNet project2 respectively collected 140,000 audio records of 8,700 bird species and about 60,000 thousand images covering thousand of plant species, to be used by scientists or professionals. Unfortunately, a manual analysis of such amount of generated data is impossible: automatic analysis tools combined with high- performance computing (HPC) solutions are therefore heavily demanded for making sense of such big ecological data. In this paper we present a case study of large-scale video processing on HPC facilities for underwater fish monitoring in the context of the Fish4Knowledge project 3, where a system to analyse long-term underwater camera footage has been developed. The paper is meant to report on the employed hardware/software architecture, the design and deployment of the parallel job manager, and the problems encountered during the whole process, from load balancing to job submission policies to bottlenecks. Simone Palazzo, Concetto Spampinato, Daniela Giordano |
PDP | 2 |
| 2014 | Parallel stochastic systems biology in the cloudabstractThe stochastic modelling of biological systems, coupled with Monte Carlo simulation of models, is an increasingly popular technique in bioinformatics. The simulation-analysis workflow may result computationally expensive reducing the interactivity required in the model tuning. In this work, we advocate the high-level software design as a vehicle for building efficient and portable parallel simulators for the cloud. In particular, the Calculus of Wrapped Components (CWC) simulator for systems biology, which is designed according to the FastFlow pattern-based approach, is presented and discussed. Thanks to the FastFlow framework, the CWC simulator is designed as a high-level workflow that can simulate CWC models, merge simulation results and statistically analyse them in a single parallel workflow in the cloud. To improve interactivity, successive phases are pipelined in such a way that the workflow begins to output a stream of analysis results immediately after simulation is started. Performance and effectiveness of the CWC simulator are validated on the Amazon Elastic Compute Cloud. Marco Aldinucci, Massimo Torquati, Concetto Spampinato, Maurizio Drocco, Claudia Misale, Cristina Calcagno, Mario Coppo |
Briefings Bioinform. | 3 |
| 2014 | Discovering biological knowledge by integrating high-throughput data and scientific literature on the cloudabstractSUMMARY In this paper, we present a bioinformatics knowledge discovery tool for extracting and validating associations between biological entities. By mining specialized scientific literature, the tool not only generates biological hypotheses in the form of associations between genes, proteins, miRNA and diseases but also validates the plausibility of such associations against high‐throughput biological data (e.g. microarray) and annotated databases (e.g. Gene Ontology). Both the knowledge discovery system and its validation are carried out by exploiting the advantages and the potentialities of the Cloud, which allowed us to derive and check the validity of thousands of biological associations in a reasonable amount of time. The system was tested on a dataset containing more than 1000 gene–disease associations achieving an average recall of about 71%, outperforming existing approaches. The results also showed that porting a data‐intensive application in an Infrastructure as a Service cloud environment boosts significantly the application's efficiency. Copyright © 2013 John Wiley & Sons, Ltd. Concetto Spampinato, Isaak Kavasidis, Marco Aldinucci, Carmelo Pino, Daniela Giordano, Alberto Faro |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | A texton-based kernel density estimation approach for background modeling under extreme conditions
Concetto Spampinato, Simone Palazzo, Isaak Kavasidis |
Comput. Vis. Image Underst. | 1 |
| 2014 | An innovative web-based collaborative platform for video annotation
Isaak Kavasidis, Simone Palazzo, Roberto Di Salvo, Daniela Giordano, Concetto Spampinato |
Multim. Tools Appl. | 5 |
| 2014 | MTAP special issue on methods and tools for ground truth collection in multimedia applications
Concetto Spampinato, Bas Boom, Jiyin He |
Multim. Tools Appl. | 1 |
| 2014 | Understanding fish behavior during typhoon events in real-life underwater environments
Concetto Spampinato, Simone Palazzo, Bas Boom, Jacco van Ossenbruggen, Isaak Kavasidis, Roberto Di Salvo, Fang-Pang Lin, Daniela Giordano, Lynda Hardman, Robert B. Fisher |
Multim. Tools Appl. | 1 |
| 2014 | A rule-based event detection system for real-life underwater domain
Concetto Spampinato, Emma Beauxis-Aussalet, Simone Palazzo, Cigdem Beyan, Jacco van Ossenbruggen, Jiyin He, Bas Boom |
Mach. Vis. Appl. | 1 |
| 2013 | Covariance based modeling of underwater scenes for fish detectionabstractIn this paper we present an algorithm for visual object detection in a underwater real-life context which explicitly models both the background and the foreground for each frame - thus helping to avoid foreground absorption into similar background -, and integrates both colour and texture features (which have proved effective in overcoming the limitations of colour-only appearance descriptors) into a covariance-based model, which provides an elegant way to merge multiple features together and enforce structural relationships. A joint domain-range model combined to a post-processing approach based on Markov Random Field takes into account the spatial dependency between pixels in the classification process, unlike the classical pixel-oriented modeling techniques. Our results show the effectiveness of this approach in the underwater environment, which presents a lot of variety in scene conditions, objects' motion patterns, shapes and colouring, and background activity. Simone Palazzo, Isaak Kavasidis, Concetto Spampinato |
ICIP | 3 |
| 2013 | Summary abstract for the 2nd ACM international workshop on multimedia analysis for ecological dataabstractThe 2nd ACM International Workshop on Multimedia Analysis for Ecological Data (MAED'13) is held as part of ACM Multimedia 2013. MAED'13, following the first workshop of the MAED series (MAED'12) that was held as part of ACM Multimedia 2012, is concerned with the processing, interpretation, and visualization of ecology-related multimedia content with the aim to support biologists in their investigations for analyzing and monitoring natural environments. Concetto Spampinato, Vasileios Mezaris, Jacco van Ossenbruggen |
ACM Multimedia | 1 |
| 2012 | First International Workshop on Visual Interfaces for Ground Truth Collection in Computer Vision ApplicationsabstractThe goal of the First International Workshop on Visual Interfaces for Ground Truth Collection in Computer Vision Applications is to bring together practitioners and researchers in computer vision and in HCI to share ideas and experiences in designing and implementing visual interfaces for ground truth data generation. Concetto Spampinato, Bas Boom, Jiyin He |
AVI | 1 |
| 2012 | BioWizard: Discovering and validating associations between biological entities by integrated analysis of scientific literature and experimental dataabstractIn this paper, we present BioWizard, a bioinformatics knowledge discovery tool for extracting and validating implicit associations between biological entities. By mining specialized scientific literature, BioWizard not only generates biological hypotheses in the form of associations between genes, proteins and diseases, but also validates the plausibility of such associations against high-throughput biological data (microarrays) and annotated databases. The main novelties of the proposed approach are that: (1) it infers associations between biological entities by mining full text papers instead of only abstracts as usually performed by the existing tools, (2) a named entity recognition that improves the precision of the derived associations by enriching the vocabularies used in the mining loop with terms extracted directly from the text and, (3) the inferred associations are filtered according to their evidence in experimental data. We tested the precision and the recall of our system in retrieving known-associations (which did not appear in the same document) from gold standards and the results shown the ability of BioWizard in retrieving valid associations, thus providing a valuable tool for the use of biomedical researchers to speed up scientific progress. Concetto Spampinato, Daniela Giordano, Isaak Kavasidis, Sebastiano Milardo |
CBMS | 1 |
| 2012 | Content based recommender system by using eye gaze dataabstractIn this work, we present a proactive content based recommender system that employs web document clustering performed by using eye gaze data. Generally, recommender systems are used in commercial applications, where information about the user's habits and interests are of crucial importance in order to plan marketing strategies, or in information retrieval systems in order to suggest similar resources a user is interested in. Commonly, these systems use explicit relevance feedback techniques (e.g. mouse or keyboard) to improve their performance and to recommend products. In contrast, the proposed system permits to capture user's interest by using implicit relevance feedback, based on data acquired by an eye tracker Tobii T60. The purpose of the system is to collect eye gaze data during web navigation and, by employing clustering techniques, to suggest web documents similar to those that the user, implicitly, expressed greater interest. Performance evaluation was carried out on 30 users and the results show that the proposed system enhanced navigation experience in about 73% of the cases. Daniela Giordano, Isaak Kavasidis, Carmelo Pino, Concetto Spampinato |
ETRA | 4 |
| 2012 | Implementing Ubiquitous Services with Ontologies: Methodology and Case Study
Alfio Costanzo, Alberto Faro, Daniela Giordano, Concetto Spampinato |
FedCSIS | 4 |
| 2012 | Context Aware Services for Mobile Users: JQMobile vs Flash Builder Implementations
Alfio Costanzo, Alberto Faro, Daniela Giordano, Concetto Spampinato |
FedCSIS | 4 |
| 2012 | A Feature Model Configuration for Multimedia Applications by an OWL-based Approach
Giuseppe Santoro, Carmelo Pino, Concetto Spampinato |
FedCSIS | 3 |
| 2012 | Evaluation of tracking algorithm performance without ground-truth dataabstractVisual tracking is a topic on which a lot of scientific work has been carried out in the last years. An important aspect of tracking algorithms is the performance evaluation, which has been carried out typically through hand-labeled ground-truth data. Since the manual generation of ground truth is a time-consuming, error-prone and tedious task, recently many researchers have focused their attention on self-evaluation techniques for performance analysis. In this paper we propose a novel tool that enables image processing researchers to test the performance of tracking algorithms without resorting to hand-labeled ground truth data. The proposed approach consists of computing a set of features describing shape, appearance and motion of the tracked objects and combining them through a naive Bayesian classifier, in order to obtain a probability score representing the overall evaluation of each tracking decision. The method was tested on three different targets (vehicles, humans and fish) with three different tracking algorithms and the results show how this approach is able to reflect the quality of the performed tracking. Concetto Spampinato, Simone Palazzo, Daniela Giordano |
ICIP | 1 |
| 2012 | Enhancing object detection performance by integrating motion objectness and perceptual organization
Concetto Spampinato, Simone Palazzo |
ICPR | 1 |
| 2012 | Multimedia analysis for ecological dataabstractThe ACM International Workshop on Multimedia Analysis for Ecological Data (MAED'12) is held as part of ACM Multimedia 2012. MAED'12 is concerned with the processing, interpretation, and visualization of ecology-related multimedia content with the aim to support biologists in their investigations for analyzing and monitoring natural environments, with particular attention to living organisms and pollution effects. Concetto Spampinato, Vasileios Mezaris, Jacco van Ossenbruggen |
ACM Multimedia | 1 |
| 2012 | Location Intelligence Services for Mobiles using Ruby on Rails and JQueryMobile
Alfio Costanzo, Alberto Faro, Concetto Spampinato |
WEBIST | 3 |
| 2012 | Combining literature text mining with microarray data: advances for system biology modelingabstractA huge amount of important biomedical information is hidden in the bulk of research articles in biomedical fields. At the same time, the publication of databases of biological information and of experimental datasets generated by high-throughput methods is in great expansion, and a wealth of annotated gene databases, chemical, genomic (including microarray datasets), clinical and other types of data repositories are now available on the Web. Thus a current challenge of bioinformatics is to develop targeted methods and tools that integrate scientific literature, biological databases and experimental data for reducing the time of database curation and for accessing evidence, either in the literature or in the datasets, useful for the analysis at hand. Under this scenario, this article reviews the knowledge discovery systems that fuse information from the literature, gathered by text mining, with microarray data for enriching the lists of down and upregulated genes with elements for biological understanding and for generating and validating new biological hypothesis. Finally, an easy to use and freely accessible tool, GeneWizard, that exploits text mining and microarray data fusion for supporting researchers in discovering gene-disease relationships is described. Alberto Faro, Daniela Giordano, Concetto Spampinato |
Briefings Bioinform. | 3 |
| 2011 | A learning tool for assessing skeletal bone age in radiologyabstractThe assessment of skeletal bone age is an important step both in diagnostic and in therapeutic investigations of en-docrinological problems and growth disorders of children. Currently, there are two main approaches for skeletal bone age estimation that use X-Ray images: 1) the Greulich and Pyle (G&P) method and 2) the Tanner and Whitehouse (TW2 or TW3) methods. Although the G&P method is the most widely used for its simplicity, the TW2/TW3 method is the most accurate one, but it requires intensive training, especially for novice radiologists. The method is complex since it involves the simultaneous assessment of shapes and relative distances among several bones of the hand and with the further complication of the interaction between sex, age and race. In this paper we propose a computer-based system for assisting radiologists in the training with the Tan-ner&Whitehouse method; the system is based on an engine for automated assessment of skeletal bone age from X-Rays. In detail, after evaluating the user level, the system automatically proposes personalized training sessions by drawing from a wide repository of hand X-rays, and it suggests remedial sessions upon identification the features of the cases where the users demonstrated misconceptions. Finally, all the X-Rays that show combinations of features provenly difficult, are treated as teaching cases and are converted into Medical Imaging Resource Center (MIRC) format to ease their integration into more comprehensive digital educational resources. Alberto Faro, Daniela Giordano, Concetto Spampinato, Rosalia Leonardi |
CBMS | 3 |
| 2011 | Eye tracker based method for quantitative analysis of pathological nystagmusabstractIn this paper we propose a method for quantitative assessment of pathological nystagmus by using eye gaze data recorded with an eye tracker (Tobii T60). In detail, we use data acquired while patients perform two tests, the smooth pursuit and the saccadic movement test (implemented on the Tobii T60 using its API), that may indicate altered ophthalmic functions typical of the pathological nystagmus. Afterwards, data is analyzed by using statistical, spectral and chaotic methods in order to provide the physicians with a useful tool for assessing quantitatively the severity (or the presence) of the pathological nystagmus. A pilot study carried out on a set of 15 patients is here reported for a preliminary investigation of the suitability of the proposed eye tracker method. Daniela Giordano, Carmelo Pino, Concetto Spampinato, Massimo Di Pietro, Alfredo Reibaldi |
CBMS | 3 |
| 2011 | Adaptive Background Modeling Integrated With Luminosity Sensors and Occlusion Processing for Reliable Vehicle DetectionabstractThis paper presents a novel vehicle detection and tracking system with stationary camera that relies on a recursive background-modeling approach, i.e., the adaptive Poisson mixture model, which is integrated with a hardware module consisting of luminosity sensors. The luminosity information side channel allows the system to effectively handle rapid changes in illumination, which is typical of outdoor applications and bottleneck of the existing background pixel classification methods. A novel algorithm for detecting and removing partial and full occlusions among blobs is also proposed. Partial occlusions are detected by evaluating the ratio between the area of the vehicle and the area of the vehicle's convex hull and are suppressed by identifying a cutting line using curvature analysis. A predictive model of the shape and motion features of the vehicles over consecutive frames instead corrects the error of the previous levels when full occlusions or background-vehicle occlusions occur in the scene. Quantitative evaluation and comparisons on some real-world scenarios demonstrate that the proposed approach outperforms state-of-the-art methods in terms of both vehicle detection and processing time, particularly due to the robustness and the efficiency of the background-modeling algorithm. Alberto Faro, Daniela Giordano, Concetto Spampinato |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2010 | Visual attention for implicit relevance feedback in a content based image retrievalabstractIn this paper we propose an implicit relevance feedback method with the aim to improve the performance of known Content Based Image Retrieval (CBIR) systems by re-ranking the retrieved images according to users' eye gaze data. This represents a new mechanism for implicit relevance feedback, in fact usually the sources taken into account for image retrieval are based on the natural behavior of the user in his/her environment estimated by analyzing mouse and keyboard interactions. In detail, after the retrieval of the images by querying CBIRs with a keyword, our system computes the most salient regions (where users look with a greater interest) of the retrieved images by gathering data from an unobtrusive eye tracker, such as Tobii T60. According to the features, in terms of color, texture, of these relevant regions our system is able to re-rank the images, initially, retrieved by the CBIR. Performance evaluation, carried out on a set of 30 users by using Google Images and "pyramid" like keyword, shows that about the 87% of the users is more satisfied of the output images when the re-raking is applied. Alberto Faro, Daniela Giordano, Carmelo Pino, Concetto Spampinato |
ETRA | 4 |
| 2010 | An interactive interface for remote administration of clinical tests based on eye trackingabstractA challenging goal today is the use of computer networking and advanced monitoring technologies to extend human intellectual capabilities in medical decision making. Modern commercial eye trackers are used in many of research fields, but the improvement of eye tracking technology, in terms of precision on the eye movements capture, has led to consider the eye tracker as a tool for vision analysis, so that its application in medical research, e.g. in ophthalmology, cognitive psychology and in neuroscience has grown considerably. The improvements of the human eye tracker interface become more and more important to allow medical doctors to increase their diagnosis capacity, especially if the interface allows them to remotely administer the clinical tests more appropriate for the problem at hand. In this paper, we propose a client/server eye tracking system that provides an interactive system for monitoring patients eye movements depending on the clinical test administered by the medical doctors. The system supports the retrieval of the gaze information and provides statistics to both medical research and disease diagnosis. Alberto Faro, Daniela Giordano, Concetto Spampinato, Davide De Tommaso, Simona Ullo |
ETRA | 3 |
| 2009 | Discovering Genes-Diseases Associations From Specialized Literature Using the GridabstractThis paper proposes a novel method for text mining on the Grid, aimed at pointing out hidden relationships for hypothesis generation and suitable for semi-interactive querying. The method is based on unsupervised clustering and the outputs are visualized with contextual information. Grid implementation is crucial for feasibility. We demonstrate it with a mining run for discovering genes-diseases associations from bibliographic sources and annotated databases. The proposed methodology is in view of a Grid architecture specialized in bioinformatics mining tasks. Some performance considerations are provided. Alberto Faro, Daniela Giordano, Francesco Maiorana, Concetto Spampinato |
IEEE Trans. Inf. Technol. Biomed. | 4 |
| 2008 | Evaluation of the Traffic Parameters in a Metropolitan Area by Fusing Visual Perceptions and CNN Processing of Webcam ImagesabstractThis paper proposes a traffic monitoring architecture based on a high-speed communication network whose nodes are equipped with fuzzy processors and cellular neural network (CNN) embedded systems. It implements a real-time mobility information system where visual human perceptions sent by people working on the territory and video-sequences of traffic taken from webcams are jointly processed to evaluate the fundamental traffic parameters for every street of a metropolitan area. This paper presents the whole methodology for data collection and analysis and compares the accuracy and the processing time of the proposed soft computing techniques with other existing algorithms. Moreover, this paper discusses when and why it is recommended to fuse the visual perceptions of the traffic with the automated measurements taken from the webcams to compute the maximum traveling time that is likely needed to reach any destination in the traffic network. Alberto Faro, Daniela Giordano, Concetto Spampinato |
IEEE Trans. Neural Networks | 3 |
| 2005 | Transcranial Magnetic Stimulation (TMS) to Evaluate and Classify Mental Diseases Using Neural Networks
Alberto Faro, Daniela Giordano, Manuela Pennisi, Giacomo Scarciofalo, Concetto Spampinato, Francesco Tramontana |
AIME | 5 |