Nicola Strisciuglio

dblp:136/0221 · DBLP profile ↗
← Back
40ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0002-7478-3509ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 22 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BEAM: Exact Benchmarking of Explainable AI Attribution Methods
abstract
The rationale behind deep neural network (DNN) predictions is often difficult to understand by humans. While EXplainable AI (XAI) methods aim at improving the explainability of DNNs, their explanations need to be reliably evaluated and compared to ground truth (GT) explanations. Existing evaluation protocols often suffer from unreliability due to the black-box nature of DNNs and resulting lack of GTs. Consequently, many existing works resorted to guessed GTs. In this paper, we shift from guessed GTs and propose a novel evaluation framework for benchmarking XAI attribution methods, which consists of a carefully designed synthetic convolutional or vision transformer image classification model accompanied by synthetic GTs. It enables precise representation of input node contributions. We also propose high-fidelity metrics to quantify the alignment between explanations of the investigated XAI method and GTs. We investigate this approach by constructing synthetic image classification models and benchmarking several widely used XAI attribution methods. Our results provide essential insights into the performance of XAI methods including 1) the imbalance in explanation fidelity between positively and negatively contributing pixels, 2) all instances of a concept being highlighted even when only a subset contributes to model output, 3) non-linear relationships between input and output being incorrectly explained, and 4) XAI methods not explaining that an input belongs to a class due to absence of a feature. Finally, our framework allows for exact evaluation of XAI methods that can be subsequently deployed for real-world task explanations. We release code and materials at https://github.com/rbrandt1/BEAM .
Rafaël Brandt, Nicola Strisciuglio, Daan Raatjes, Georgi Gaydadjiev
ICPR (14)2
2026 MedHyCLIP: Hyperbolic CLIP adaptation for universal medical anomaly detection
Keyu Guo, Hongkai Wei, Yongle Huang, Shijie Sun 0001, Yueming Shi, Huansheng Song, Nicola Strisciuglio
Pattern Recognit.9
2026 Animating Faces With Emotions Through a Generative Adversarial Network Preserving Identity
abstract
Artificially applying specific emotions to videos of people faces with a neutral expression, while preserving the identity of the subject is a challenging task. When parts of the face are synthetically moved to generate an emotion, it typically results in spatio-temporal artifacts in the generated videos, or inconsistency to preserve the identity of subjects. Existing methods that deploy spatio-temporal convolutions and de-convolutions to generate consecutive frames in a single step are not able to ensure proper motion dynamics, in the sense that the emotion may be not visible on the face or the facial features are distorted in the video. At the same time, approaches that generate motion and identity in two separate steps are not able to ensure the consistency of the subject identity after the generation of the emotion. In this paper we propose a novel method, Video Identity-Consistent Emotion GAN (VICEGAN), that improves the video generative capabilities of two-step methods. We decouple motion and content generation, thus ensuring the consistency of subject identity in the generated videos by using an encoder-decoder generator and a new identity-preserving loss in an adversarial framework. The proposed neural network architecture also guarantees the generation of proper motion of the target expressions, mitigating the presence of artifacts. We evaluated VICEGAN on the MUG dataset and compared it with a method based on a GAN, ImaGINator, demonstrating superior performance both quantitatively and qualitatively, and with a popular method based on a diffusion model, LFDM, showing a better capability to generate recognizable emotions.
Antonio Greco 0001, Nicola Strisciuglio, Mario Vento
IEEE Trans. Affect. Comput.2
2025 Vision on the Move: Automated Hazardous Material Plate Detection in Freight Transport
Melissa Tijink, Stanislav Levendeev, Ewaldo Nieuwenhuis, Luuk J. Spreeuwers, Nicola Strisciuglio, Estefanía Talavera
CAIP (1)5
2025 Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
abstract
Vision-Language Models (VLMs) learn a shared feature space for text and images, enabling the comparison of inputs of different modalities. While prior works demonstrated that VLMs organize natural language representations into regular structures encoding composite meanings, it remains unclear if compositional patterns also emerge in the visual embedding space. In this work, we investigate compositionality in the image domain, where the analysis of compositional properties is challenged by noise and sparsity of visual data. We address these problems and propose a framework, called Geodesically Decomposable Embeddings (GDE), that approximates image representations with geometry-aware compositional structures in the latent space. We demonstrate that visual embeddings of pre-trained VLMs exhibit a compositional arrangement, and evaluate the effectiveness of this property in the tasks of compositional classification and group robustness. GDE achieves stronger performance in compositional classification compared to its counterpart method that assumes linear geometry of the latent space. Notably, it is particularly effective for group robustness, where we achieve higher results than task-specific solutions. Our results indicate that VLMs can automatically develop a human-like form of compositional reasoning in the visual domain, making their underlying processes more interpretable. Code is available at https://github.com/BerasiDavide/vlm_image_compositionality.
Davide Berasi, Matteo Farina, Massimiliano Mancini, Elisa Ricci 0001, Nicola Strisciuglio
CVPR5
2025 Do ImageNet-trained Models Learn Shortcuts? The Impact of Frequency Shortcuts on Generalization
abstract
Frequency shortcuts refer to specific frequency patterns that models heavily rely on for correct classification. Previous studies have shown that models trained on small image datasets often exploit such shortcuts, potentially impairing their generalization performance. However, existing methods for identifying frequency shortcuts require expensive computations and become impractical for analyzing models trained on large datasets. In this work, we propose the first approach to more efficiently analyze frequency shortcuts at a large scale. We show that both CNN and transformer models learn frequency shortcuts on ImageNet. We also expose that frequency shortcut solutions can yield good performance on out-of-distribution (OOD) test sets which largely retain texture information. However, these shortcuts, mostly aligned with texture patterns, hinder model generalization on rendition-based OOD test sets. These observations suggest that current OOD evaluations often overlook the impact of frequency shortcuts on model generalization. Future benchmarks could thus benefit from explicitly assessing and accounting for these shortcuts to build models that generalize across a broader range of OOD scenarios. Codes are available at https://github.com/nis-research/hfss.
Shunxin Wang, Raymond N. J. Veldhuis, Nicola Strisciuglio
CVPR3
2025 Dynamic Sparse Training versus Dense Training: The Unexpected Winner in Image Corruption Robustness
abstract
It is generally perceived that Dynamic Sparse Training opens the door to a new era of scalability and efficiency for artificial neural networks at, perhaps, some costs in accuracy performance for the classification task. At the same time, Dense Training is widely accepted as being the "de facto" approach to train artificial neural networks if one would like to maximize their robustness against image corruption. In this paper, we question this general practice. Consequently, \textit{we claim that}, contrary to what is commonly thought, the Dynamic Sparse Training methods can consistently outperform Dense Training in terms of robustness accuracy, particularly if the efficiency aspect is not considered as a main objective (i.e., sparsity levels between 10\% and up to 50\%), without adding (or even reducing) resource cost. We validate our claim on two types of data, images and videos, using several traditional and modern deep learning architectures for computer vision and three widely studied Dynamic Sparse Training algorithms. Our findings reveal a new yet-unknown benefit of Dynamic Sparse Training and open new possibilities in improving deep learning robustness beyond the current state of the art.
Boqian Wu, Qiao Xiao, Shunxin Wang, Nicola Strisciuglio, Mykola Pechenizkiy, Maurice van Keulen, Decebal Constantin Mocanu, Elena Mocanu
ICLR4
2025 Multilayer perceptron ensembles in a truly sparse training context
abstract
Abstract Ensemble learning for artificial neural networks (ANNs) is an effective method to enhance predictive performance. However, ANNs are computationally and memory intensive, and naively training multiple networks can lead to excessive training times and costs. An effective tool for improving ensemble efficiency is introducing topological sparsity. Even though several implementations of efficient ensembles have been proposed, none of them can provide actual benefits in terms of computational overhead as the sparsity is simulated using binary masks. In this paper, we address this issue by introducing a Truly Sparse Ensemble without binary masks and directly incorporate native sparsity. We also propose two algorithms for initializing new subnetworks within the ensemble, leveraging this native topological sparsity to enhance subnetwork diversity. We demonstrate the performance of the resulting models at high levels of sparsity on several datasets in terms of classification accuracy, floating point operations (FLOPs), and actual running time. The proposed methods outperform all baseline dense and truly sparse models on tabular data, successfully diversify the training trajectory of the subnetworks, and increase the topological distance between subnetworks after re-initialization.
Peter R. D. van der Wal, Nicola Strisciuglio, George Azzopardi, Decebal Constatin Mocanu
Neural Comput. Appl.2
2024 Fourier-Basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
abstract
Computer vision models normally witness degraded performance when deployed in real-world scenarios, due to unexpected changes in inputs that were not accounted for during training. Data augmentation is commonly used to address this issue, as it aims to increase data variety and reduce the distribution gap between training and test data. However, common visual augmentations might not guaran-tee extensive robustness of computer vision models. In this paper, we propose Auxiliary Fourier-basis Augmentation (AFA), a complementary technique targeting augmentation in the frequency domain and filling the robustness gap left by visual augmentations. We demonstrate the utility of augmentation via Fourier-basis additive noise in a straightforward and efficient adversarial setting. Our results show that AFA benefits the robustness of models against common corruptions, OOD generalization, and consistency of performance of models against increasing perturbations, with negligible deficit to the standard performance of models. It can be seamlessly integrated with other augmentation techniques to further boost performance. Codes and models are available at https://github.com/nis-research/afa-augment.
Puru Vaish, Shunxin Wang, Nicola Strisciuglio
CVPR3
2024 Controllable Privacy in Face Recognition: A Filter-based Approach
abstract
Recent advancements in deep learning for face recognition have led to concerns regarding privacy and algorithmic bias, particularly in inferring demographic attributes from facial templates. Existing methods often struggle to balance privacy preservation with utility and operational efficiency. In this paper, we propose a filter-based privacy-enhancing method inspired by the information bottleneck concept. Our approach involves training a filter estimator that assigns scores to intermediate-layer elements based on their sensitivity to target attributes. By selectively replacing sensitive elements with noise while allowing less sensitive ones to pass through, our method aims to enhance privacy while managing the trade-off with verification performance. Most importantly, our approach allows for post-training tuning of the privacy-utility trade-off, providing flexibility for different operational requirements. Evaluation across multiple face recognition networks and datasets demonstrates that our approach can achieve substantial gains in the gender obfuscation task while maintaining adequate verification performance and computational efficiency suitable for real-time applications.
Zohra Rezgui, Nicola Strisciuglio, Raymond N. J. Veldhuis
IJCB2
2024 PushPull-Net: Inhibition-Driven ResNet Robust to Image Corruptions
Guru Swaroop Bennabhaktula, Enrique Alegre, Nicola Strisciuglio, George Azzopardi
ICPR (8)3
2024 Regressing Transformers for Data-efficient Visual Place Recognition
abstract
Visual place recognition is a critical task in computer vision, especially for localization and navigation systems. Existing methods often rely on contrastive learning: image descriptors are trained to have small distance for similar images and larger distance for dissimilar ones in a latent space. However, this approach struggles to ensure accurate distance-based image similarity representation, particularly when training with binary pairwise labels, and complex re-ranking strategies are required. This work introduces a fresh perspective by framing place recognition as a regression problem, using camera field-of-view overlap as similarity ground truth for learning. By optimizing image descriptors to align directly with graded similarity labels, this approach enhances ranking capabilities without expensive re-ranking, offering data-efficient training and strong generalization across several benchmark datasets.
Maria Leyva-Vallina, Nicola Strisciuglio, Nicolai Petkov
ICRA2
2024 CAST: Clustering self-Attention using Surrogate Tokens for efficient transformers
abstract
The Transformer architecture has shown to be a powerful tool for a wide range of tasks. It is based on the self-attention mechanism, which is an inherently computationally expensive operation with quadratic computational complexity: memory usage and compute time increase quadratically with the length of the input sequences, thus limiting the application of Transformers. In this work, we propose a novel Clustering self-Attention mechanism using Surrogate Tokens (CAST), to optimize the attention computation and achieve efficient transformers. CAST utilizes learnable surrogate tokens to construct a cluster affinity matrix, used to cluster the input sequence and generate novel cluster summaries. The self-attention from within each cluster is then combined with the cluster summaries of other clusters, enabling information flow across the entire input sequence. CAST improves efficiency by reducing the complexity from O ( N 2 ) to O ( α N ) where N is the sequence length, and α is constant according to the number of clusters and samples per cluster. We show that CAST performs better than or comparable to the baseline Transformers on long-range sequence modeling tasks, while also achieving higher results on time and memory efficiency than other efficient transformers. • Computation of self-attention Transformers is limited by the input sequence length. • We propose CAST, an efficient self-attention mechanism with clustered attention. • We propose the use of surrogate tokens to optimize self-attention in transformers. • We observe that cluster summaries enhance training efficiency and results. • CAST reduces complexity of self-attention computation from O ( N 2 ) to O ( α N ) .
Adjorn van Engelenhoven, Nicola Strisciuglio, Estefanía Talavera
Pattern Recognit. Lett.2
2024 RDA-INR: Riemannian Diffeomorphic Autoencoding via Implicit Neural Representations
abstract
Abstract. Diffeomorphic registration frameworks such as large deformation diffeomorphic metric mapping (LDDMM) are used in computer graphics and the medical domain for atlas building, statistical latent modeling, and pairwise and groupwise registration. In recent years, researchers have developed neural network–based approaches regarding diffeomorphic registration to improve the accuracy and computational efficiency of traditional methods. In this work, we focus on a limitation of neural network–based atlas building and statistical latent modeling methods, namely that they either (i) are resolution dependent or (ii) disregard any data- or problem-specific geometry needed for proper mean-variance analysis. In particular, we overcome this limitation by designing a novel encoder based on resolution-independent implicit neural representations. The encoder achieves resolution invariance for LDDMM-based statistical latent modeling. Additionally, the encoder adds LDDMM Riemannian geometry to resolution-independent deep learning models for statistical latent modeling. We investigate how the Riemannian geometry improves latent modeling and is required for a proper mean-variance analysis. To highlight the benefit of resolution independence for LDDMM-based data variability modeling, we show that our approach outperforms current neural network–based LDDMM latent code models. Our work paves the way for more research into how Riemannian geometry; shape, respectively, image analysis; and deep learning can be combined.
Sven Dummer, Nicola Strisciuglio, Christoph Brune
SIAM J. Imaging Sci.2
2023 Defocus Blur Synthesis and Deblurring via Interpolation and Extrapolation in Latent Space
Ioana Mazilu, Shunxin Wang, Sven Dummer, Raymond N. J. Veldhuis, Christoph Brune, Nicola Strisciuglio
CAIP (2)6
2023 Data-Efficient Large Scale Place Recognition with Graded Similarity Supervision
abstract
Visual place recognition (VPR) is a fundamental task of computer vision for visual localization. Existing methods are trained using image pairs that either depict the same place or not. Such a binary indication does not consider continuous relations of similarity between images of the same place taken from different positions, determined by the continuous nature of camera pose. The binary similarity induces a noisy supervision signal into the training of VPR methods, which stall in local minima and require expensive hard mining algorithms to guarantee convergence. Motivated by the fact that two images of the same place only partially share visual cues due to camera pose differences, we deploy an automatic re-annotation strategy to re-label VPR datasets. We compute graded similarity labels for image pairs based on available localization metadata. Furthermore, we propose a new Generalized Contrastive Loss (GCL) that uses graded similarity labels for training contrastive networks. We demonstrate that the use of the new labels and GCL allow to dispense from hard-pair mining, and to train image descriptors that perform better in VPR by nearest neighbor search, obtaining superior or comparable results than methods that require expensive hard-pair mining and re-ranking techniques.
Maria Leyva-Vallina, Nicola Strisciuglio, Nicolai Petkov
CVPR2
2023 What do neural networks learn in image classification? A frequency shortcut perspective
abstract
Frequency analysis is useful for understanding the mechanisms of representation learning in neural networks (NNs). Most research in this area focuses on the learning dynamics of NNs for regression tasks, while little for classification. This study empirically investigates the latter and expands the understanding of frequency shortcuts. First, we perform experiments on synthetic datasets, designed to have a bias in different frequency bands. Our results demonstrate that NNs tend to find simple solutions for classification, and what they learn first during training depends on the most distinctive frequency characteristics, which can be either low- or high-frequencies. Second, we confirm this phenomenon on natural images. We propose a metric to measure class-wise frequency characteristics and a method to identify frequency shortcuts. The results show that frequency shortcuts can be texture-based or shape-based, depending on what best simplifies the objective. Third, we validate the transferability of frequency shortcuts on out-of-distribution (OOD) test sets. Our results suggest that frequency shortcuts can be transferred across datasets and cannot be fully avoided by larger model capacity and data augmentation. We recommend that future research should focus on effective training schemes mitigating frequency shortcut learning. Codes and data are available at https://github.com/nis-research/nn-frequency-shortcuts.
Shunxin Wang, Raymond N. J. Veldhuis, Christoph Brune, Nicola Strisciuglio
ICCV4
2023 Benchmarking deep networks for facial emotion recognition in the wild
abstract
Abstract Emotion recognition from face images is a challenging task that gained interest in recent years for its applications to business intelligence and social robotics. Researchers in computer vision and affective computing focused on optimizing the classification error on benchmark data sets, which do not extensively cover possible variations that face images may undergo in real environments. Following on investigations carried out in the field of object recognition, we evaluated the robustness of existing methods for emotion recognition when their input is subjected to corruptions caused by factors present in real-world scenarios. We constructed two data sets on top of the RAF-DB test set, named RAF-DB-C and RAF-DB-P, that contain images modified with 18 types of corruption and 10 of perturbation. We benchmarked existing networks (VGG, DenseNet, SENet and Xception) trained on the original images of RAF-DB and compared them with ARM, the current state-of-the-art method on the RAF-DB test set. We carried out an extensive study on the effects that modifications to the training data or network architecture have on the classification of corrupted and perturbed data. We observed a drop of recognition performance of ARM, with the classification error raising up to 200% of that achieved on the original RAF-DB test set. We demonstrate that the use of the AutoAugment data augmentation and an anti-aliasing filter within down-sampling layers provide existing networks with increased robustness to out-of-distribution variations, substantially reducing the error on corrupted inputs and outperforming ARM. We provide insights about the resilience of existing emotion recognition methods and an estimation of their performance in real scenarios. The processing time required by the modifications we investigated (35 ms in the worst case) supports their suitability for application in real-world scenarios. The RAF-DB-C and RAF-DB-P test sets, trained models and evaluation framework are available at https://github.com/MiviaLab/emotion-robustness .
Antonio Greco 0001, Nicola Strisciuglio, Mario Vento, Vincenzo Vigilante
Multim. Tools Appl.2
2021 MTStereo 2.0: Accurate Stereo Depth Estimation via Max-Tree Matching
Rafaël Brandt, Nicola Strisciuglio, Nicolai Petkov
CAIP (1)2
2020 A robust contour detection operator with combined push-pull inhibition and surround suppression
abstract
Contour detection is a salient operation in many computer vision applications as it extracts features that are important for distinguishing objects in scenes. It is believed to be a primary role of simple cells in visual cortex of the mammalian brain. Many of such cells receive push-pull inhibition or surround suppression. We propose a computational model that exhibits a combination of these two phenomena. It is based on two existing models, which have been proven to be very effective for contour detection. In particular, we introduce a brain-inspired contour operator that combines push-pull and surround inhibition. It turns out that this combination results in a more effective contour detector, which suppresses texture while keeping the strongest responses to lines and edges, when compared to existing models. The proposed model consists of a Combination of Receptive Field (or CORF) model with push-pull inhibition, extended with surround suppression. We demonstrate the effectiveness of the proposed approach on the RuG and Berkeley benchmark data sets of 40 and 500 images, respectively. The proposed push-pull CORF operator with surround suppression outperforms the one without suppression with high statistical significance.
Damiano Melotti, Kevin Heimbach, Antonio Jose Rodríguez-Sánchez, Nicola Strisciuglio, George Azzopardi
Inf. Sci.4
2020 U-COSFIRE filters for vessel tortuosity quantification with application to automated diagnosis of retinopathy of prematurity
Sivakumar Ramachandran, Nicola Strisciuglio, Anand Vinekar, Renu John, George Azzopardi
Neural Comput. Appl.2
2020 Enhanced robustness of convolutional networks with a push-pull inhibition layer
abstract
Abstract Convolutional neural networks (CNNs) lack robustness to test image corruptions that are not seen during training. In this paper, we propose a new layer for CNNs that increases their robustness to several types of corruptions of the input images. We call it a ‘push–pull’ layer and compute its response as the combination of two half-wave rectified convolutions, with kernels of different size and opposite polarity. Its implementation is based on a biologically motivated model of certain neurons in the visual system that exhibit response suppression, known as push–pull inhibition. We validate our method by replacing the first convolutional layer of the LeNet, ResNet and DenseNet architectures with our push–pull layer. We train the networks on original training images from the MNIST and CIFAR data sets and test them on images with several corruptions, of different types and severities, that are unseen by the training process. We experiment with various configurations of the ResNet and DenseNet models on a benchmark test set with typical image corruptions constructed on the CIFAR test images. We demonstrate that our push–pull layer contributes to a considerable improvement in robustness of classification of corrupted images, while maintaining state-of-the-art performance on the original image classification task. We released the code and trained models at the url http://github.com/nicstrisc/Push-Pull-CNN-layer .
Nicola Strisciuglio, Manuel Lopez-Antequera, Nicolai Petkov
Neural Comput. Appl.1
2020 Efficient binocular stereo correspondence matching with 1-D Max-Trees
abstract
Extraction of depth from images is of great importance for various computer vision applications. Methods based on convolutional neural networks are very accurate but have high computation requirements, which can be achieved with GPUs. However, GPUs are difficult to use on devices with low power requirements like robots and embedded systems. In this light, we propose a stereo matching method appropriate for applications in which limited computational and energy resources are available. The algorithm is based on a hierarchical representation of image pairs which is used to restrict disparity search range. We propose a cost function that takes into account region contextual information and a cost aggregation method that preserves disparity borders. We tested the proposed method on the Middlebury and KITTI benchmark data sets and on the TrimBot2020 synthetic data. We achieved accuracy and time efficiency results that show that the method is suitable to be deployed on embedded and robotics systems.
Rafaël Brandt, Nicola Strisciuglio, Nicolai Petkov, Michael H. F. Wilkinson
Pattern Recognit. Lett.2
2019 Place Recognition in Gardens by Learning Visual Representations: Data Set and Benchmark Analysis
Maria Leyva-Vallina, Nicola Strisciuglio, Nicolai Petkov
CAIP (1)2
2019 Trainable COPE Features for Sound Event Detection
Nicola Strisciuglio, Nicolai Petkov
CIARP1
2019 Learning representations of sound using trainable COPE feature extractors
abstract
Sound analysis research has mainly been focused on speech and music processing. The deployed methodologies are not suitable for analysis of sounds with varying background noise, in many cases with very low signal-to-noise ratio (SNR). In this paper, we present a method for the detection of patterns of interest in audio signals . We propose novel trainable feature extractors, which we call COPE (Combination of Peaks of Energy). The structure of a COPE feature extractor is determined using a single prototype sound pattern in an automatic configuration process , which is a type of representation learning. We construct a set of COPE feature extractors, configured on a number of training patterns. Then we take their responses to build feature vectors that we use in combination with a classifier to detect and classify patterns of interest in audio signals . We carried out experiments on four public data sets: MIVIA audio events, MIVIA road events, ESC-10 and TU Dortmund data sets. The results that we achieved (recognition rate equal to 91.71% on the MIVIA audio events, 94% on the MIVIA road events, 81.25% on the ESC-10 and 94.27% on the TU Dortmund) demonstrate the effectiveness of the proposed method and are higher than the ones obtained by other existing approaches. The COPE feature extractors have high robustness to variations of SNR. Real-time performance is achieved even when the value of a large number of features is computed.
Nicola Strisciuglio, Mario Vento, Nicolai Petkov
Pattern Recognit.1
2019 Learning skeleton representations for human action recognition
abstract
Automatic interpretation of human actions gained strong interest among researchers in patter recognition and computer vision because of its wide range of applications, such as in social and home robotics, elderly people health care, surveillance, among others. In this paper, we propose a method for recognition of human actions by analysis of skeleton poses. The method that we propose is based on novel trainable feature extractors, which can learn the representation of prototype skeleton examples and can be employed to recognize skeleton poses of interest. We combine the proposed feature extractors with an approach for classification of pose sequences based on string kernels. We carried out experiments on three benchmark data sets (MIVIA-S, MSRSDA and MHAD) and the results that we achieved are comparable or higher than the ones obtained by other existing methods. A further important contribution of this work is the MIVIA-S dataset, that we collected and made publicly available.
Alessia Saggese, Nicola Strisciuglio, Mario Vento, Nicolai Petkov
Pattern Recognit. Lett.2
2019 Robust Inhibition-Augmented Operator for Delineation of Curvilinear Structures
abstract
Delineation of curvilinear structures in images is an important basic step of several image processing applications, such as segmentation of roads or rivers in aerial images, vessels or staining membranes in medical images, and cracks in pavements and roads, among others. Existing methods suffer from insufficient robustness to noise. In this paper, we propose a novel operator for the detection of curvilinear structures in images, which we demonstrate to be robust to various types of noise and effective in several applications. We call it RUSTICO, which stands for RobUST Inhibition-augmented Curvilinear Operator. It is inspired by the push-pull inhibition in visual cortex and takes as input the responses of two trainable B-COSFIRE filters of opposite polarity. The output of RUSTICO consists of a magnitude map and an orientation map. We carried out experiments on a data set of synthetic stimuli with noise drawn from different distributions, as well as on several benchmark data sets of retinal fundus images, crack pavements, and aerial images and a new data set of rose bushes used for automatic gardening. We evaluated the performance of RUSTICO by a metric that considers the structural properties of line networks (connectivity, area, and length) and demonstrated that RUSTICO outperforms many existing methods with high statistical significance. RUSTICO exhibits high robustness to noise and texture.
Nicola Strisciuglio, George Azzopardi, Nicolai Petkov
IEEE Trans. Image Process.1
2017 A real-time system for audio source localization with cheap sensor device
abstract
We propose an architecture for real-time audio source localization based on the integration of localization methodologies within a framework that employs a cheap acquisition sensor. The architecture that we present takes as input the audio signals from two calibrated microphones. Then, it computes biological-inspired features of the sound signal and estimates its direction by means of a Gaussian Mixture Model estimator. We carried out an extensive experimental analysis on four data sets, one of which we realized and made publicly available. We evaluated several characteristics of the sound localization architecture and its use in real scenarios.
Alessia Saggese, Nicola Strisciuglio, Mario Vento, Nicolai Petkov
AVSS2
2017 Detection of Curved Lines with B-COSFIRE Filters: A Case Study on Crack Delineation
Nicola Strisciuglio, George Azzopardi, Nicolai Petkov
CAIP (1)1
2016 Time-frequency analysis for audio event detection in real scenarios
abstract
We propose a sound analysis system for the detection of audio events in surveillance applications. The method that we propose combines short- and long-time analysis in order to increase the reliability of the detection. The basic idea is that a sound is composed of small, atomic audio units and some of them are distinctive of a particular class of sounds. Similarly to the words in a text, we count the occurrence of audio units for the construction of a feature vector that describes a given time interval. A classifier is then used to learn which audio units are distinctive for the different classes of sound. We compare the performance of different sets of short-time features by carrying out experiments on the MIVIA audio event data set. We study the performance and the stability of the proposed system when it is employed in live scenarios, so as to characterize its expected behavior when used in real applications.
Alessia Saggese, Nicola Strisciuglio, Mario Vento, Nicolai Petkov
AVSS2
2016 Supervised vessel delineation in retinal fundus images with the automatic selection of B-COSFIRE filters
abstract
The inspection of retinal fundus images allows medical doctors to diagnose various pathologies. Computer-aided diagnosis systems can be used to assist in this process. As a first step, such systems delineate the vessel tree from the background. We propose a method for the delineation of blood vessels in retinal images that is effective for vessels of different thickness. In the proposed method, we employ a set of B -COSFIRE filters selective for vessels and vessel-endings. Such a set is determined in an automatic selection process and can adapt to different applications. We compare the performance of different selection methods based upon machine learning and information theory. The results that we achieve by performing experiments on two public benchmark data sets, namely DRIVE and STARE, demonstrate the effectiveness of the proposed approach.
Nicola Strisciuglio, George Azzopardi, Mario Vento, Nicolai Petkov
Mach. Vis. Appl.1
2016 Audio Surveillance of Roads: A System for Detecting Anomalous Sounds
abstract
In the last decades, several systems based on video analysis have been proposed for automatically detecting accidents on roads to ensure a quick intervention of emergency teams. However, in some situations, the visual information is not sufficient or sufficiently reliable, whereas the use of microphones and audio event detectors can significantly improve the overall reliability of surveillance systems. In this paper, we propose a novel method for detecting road accidents by analyzing audio streams to identify hazardous situations such as tire skidding and car crashes. Our method is based on a two-layer representation of an audio stream: at a low level, the system extracts a set of features that is able to capture the discriminant properties of the events of interest, and at a high level, a representation based on a bag-of-words approach is then exploited in order to detect both short and sustained events. The deployment architecture for using the system in real environments is discussed, together with an experimental analysis carried out on a data set made publicly available for benchmarking purposes. The obtained results confirm the effectiveness of the proposed approach.
Pasquale Foggia, Nicolai Petkov, Alessia Saggese, Nicola Strisciuglio, Mario Vento
IEEE Trans. Intell. Transp. Syst.4
2015 Car crashes detection by audio analysis in crowded roads
abstract
In the last years, video surveillance has been employed for roads monitoring in order to detect abnormal events and improve the safety procedures in case of emergency. Certain events, such as car crashes or tire skidding, are difficult or impossible to detect when only the visual information is considered. In this paper we describe a preliminary system to detect events in roads by means of audio analysis. The system that we propose combines short- and long-time analysis of the audio signal in order to detect both impulsive and sustained events. We present the preliminary results achieved by the proposed system on a data set specifically made for roads surveillance, which we made publicly available. We also discuss the architectural deployment of such system in real environments with respect to a model of the noise of road traffic. The achieved results are promising and confirm the effectiveness of the system.
Pasquale Foggia, Alessia Saggese, Nicola Strisciuglio, Mario Vento, Nicolai Petkov
AVSS3
2015 Multiscale Blood Vessel Delineation Using B-COSFIRE Filters
Nicola Strisciuglio, George Azzopardi, Mario Vento, Nicolai Petkov
CAIP (2)1
2015 Trainable COSFIRE filters for vessel delineation with application to retinal images
George Azzopardi, Nicola Strisciuglio, Mario Vento, Nicolai Petkov
Medical Image Anal.2
2015 Reliable detection of audio events in highly noisy environments
Pasquale Foggia, Nicolai Petkov, Alessia Saggese, Nicola Strisciuglio, Mario Vento
Pattern Recognit. Lett.4
2014 Cascade classifiers trained on gammatonegrams for reliably detecting audio events
abstract
In this paper we propose a novel method for the detection of events of interest through audio analysis. The system that we propose is based on the representation of the audio streams through a Gammatone image, which describes the time-frequency distribution of the energy of the signal; this representation is inspired by the functioning of the human auditory system. A pool of AdaBoost cascade classifiers, one for each class of events of interest, is involved in the event detection stage. The performance of the proposed system has been evaluated on a large data set of audio events for surveillance applications and the achieved results, compared with two state of the art approaches, confirm its effectiveness.
Pasquale Foggia, Alessia Saggese, Nicola Strisciuglio, Mario Vento
AVSS3
2014 Exploiting the deep learning paradigm for recognizing human actions
abstract
In this paper we propose a novel method for recognizing human actions by exploiting a multi-layer representation based on a deep learning based architecture. A first level feature vector is extracted and then a high level representation is obtained by taking advantage of a Deep Belief Network trained using a Restricted Boltzmann Machine. The classification is finally performed by a feed-forward neural network. The main advantage behind the proposed approach lies in the fact that the high level representation is automatically built by the system exploiting the regularities in the dataset; given a suitably large dataset, it can be expected that such a representation can outperform a hand-design description scheme. The proposed approach has been tested on two standard datasets and the achieved results, compared with state of the art algorithms, confirm its effectiveness.
Pasquale Foggia, Alessia Saggese, Nicola Strisciuglio, Mario Vento
AVSS3
2013 Audio surveillance using a bag of aural words classifier
abstract
In this paper we propose a novel approach for the audio-based detection of events. The approach adopts the bag of words paradigm, and has two main advantages over other techniques present in the literature: the ability to automatically adapt (through a learning phase) to both short, impulsive sounds and long, sustained ones, and the ability to work in noisy environments where the sounds of interest are superimposed to background sounds possibly having similar characteristics. The proposed method has been experimentally validated on a large database of sounds, including several kinds of background noise, which are superimposed to the sounds to be recognized. The obtained performance has been compared with the results of another audio event detection algorithm from the literature, showing a significant improvement.
Vincenzo Carletti, Pasquale Foggia, Gennaro Percannella, Alessia Saggese, Nicola Strisciuglio, Mario Vento
AVSS5