Alessia Saggese

dblp:64/11434 · DBLP profile ↗
← Back
63ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0003-4687-7994ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 25 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Deep learning based empty shelf detection based on autonomous mobile robot
abstract
The issue of out-of-stock (OOS) represents a substantial challenge for retailers, often resulting in significant sales losses. To address this problem, this paper introduces an autonomous mobile robotic platform built on the Robot Operating System (ROS) framework, designed to accelerate the restocking process in supermarkets. The platform autonomously detects empty shelves and notifies human operators, streamlining inventory management. Equipped with advanced navigation capabilities, the proposed system employs a deep learning-based, two-stage architecture that identifies shelving areas and subsequently detects empty shelves. To validate the performance of the proposed two-stage artificial vision algorithm, two datasets were used: the first comprises approximately 2000 images (900 of them collected by our team from three different supermarkets), while the second dataset consists of around 5600 manually annotated images extracted from videos recorded in a supermarket by the robotic platform itself. Additionally, in order to validate the entire robotic system, an extensive experimental evaluation was conducted in a supermarket during regular business hours. The results demonstrate that the proposed platform substantially outperforms human operators, identifying OOS items eight times faster than traditional human operator based methods. This advancement provides valuable assistance to supermarket staff, significantly enhancing operational efficiency.
Giuseppe De Simone, Alessia Saggese, Pasquale Foggia, Mario Vento
Comput. Vis. Image Underst.2
2026 Simultaneous person attribute recognition using task-specific attention network on embedded devices
abstract
Pedestrian attribute recognition has become an important task in computer vision, particularly for retail and marketing and critical applications in security and surveillance. Despite its potential, achieving real-time performance on embedded devices while maintaining accuracy has been a significant challenge. In this paper, we propose a novel multi-task method for simultaneous person attribute recognition using task-specific attention network, which shares low-level representations across related tasks, reducing computational and memory requirements without compromising accuracy. In particular, we employ a spatial-channel attention mechanism to selectively focus on relevant regions without increasing the computational complexity of the backbone. Furthermore, we use a knowledge distillation technique to deal with missing labels and gradient normalization for dealing with task imbalances, since varying task difficulties lead to disproportionate gradient magnitudes during training. The experimental results demonstrate the effectiveness of our approach, achieving a mean accuracy of 0.889 while maintaining real-time performance at 114 frames per second on an embedded board with limited resources. These results highlight the practical viability and novelty of our system as a robust and scalable solution for pedestrian attribute recognition on embedded devices in real-world scenarios.
George Azzopardi, Antonio Greco 0001, Alessia Saggese, Bruno Vento
Eng. Appl. Artif. Intell.3
2026 Joint underwater image enhancement and multi-scale fish detection
abstract
Automatic fish detection in images plays a crucial role in marine biodiversity monitoring and environmental conservation. The advent of deep learning has significantly improved the results achievable in fish detection; however, underwater challenges such as scale variability, optical distortions, and color inconsistencies reveal some limitations of conventional deep learning-based detectors. To address these issues, we propose an improved fish detection model designed for underwater environments trained to perform joint underwater image enhancement and multi-scale fish detection. While incorporating structural enhancements for multi-scale robustness, such as a dedicated tiny object detection head and attention mechanisms, the core of our approach resides in the joint training of an underwater image enhancement module with the detector. We evaluate our method on two challenging test sets. The experimental results demonstrate that our model significantly outperforms state-of-the-art methods, achieving an average precision between 0.84 and 0.89. Despite its superior accuracy, the proposed solution maintains a lightweight architecture with only 24 millions of parameters and ensures real-time processing capabilities (13 frames per second on edge hardware), highlighting its potential for effective deployment in marine research and autonomous fisheries monitoring. • YOLO-JUICE boosts small-object detection with an extra tiny-object head. • YOLO-JUICE preserves details without extra cost with a space to depth module. • YOLO-JUICE refines spatial and channel features with a spatial-channel attention. • YOLO-JUICE is jointly trained for image classification and enhancement.
Vincenzo Carletti, Antonio Greco 0001, Andrea Vincenzo Ricciardi, Alessia Saggese, Bruno Vento
Eng. Appl. Artif. Intell.4
2026 DRIVE: Distributed Robotic Intelligence for Vision-based Exploration for retail shelf monitoring
abstract
Out-of-stock (OOS) detection in retail environments is essential to ensure efficient inventory management and maintain high levels of customer satisfaction. Within this context, mobile robotic platforms, equipped with a camera and an empty shelf object detector, have emerged as a promising solution. However, detector-based approaches suffer from a fundamental trade-off between false positives and missed detections, with limited generalization capabilities due to small available training datasets, and high false positive rate in cluttered retail scenes. To overcome these challenges, we propose DRIVE (Distributed Robotic Intelligence for Vision-based Exploration), a novel distributed architecture that combines a lightweight on-board object detector with a cloud-based transformer-powered semantic validation stage. This two-tier design mitigates the precision–recall trade-off of traditional detectors, reducing false positives without sacrificing recall, while ensuring real-time feasibility on resource-constrained platforms. Furthermore, to enable robust domain adaptation under low-data regimes without catastrophic forgetting, we fine-tune the vision transformer backbone using Parameter-Efficient Fine-Tuning (PEFT) via Low-Rank Adaptation (LoRA), thus injecting less than 1% additional parameters while preserving pretrained knowledge. Extensive experiments in a real supermarket environments demonstrate that DRIVE achieves impressive robustness and accuracy compared to state-of-the-art detection-based solutions, paving the way for scalable, autonomous OOS detection in dynamic retail scenarios.
Alessia Saggese, Mario Vento
J. Syst. Archit.1
2025 Multi-modal Human-Robot Collaboration in Production Lines Through Speech Commands and Gestures
Vincenzo Carletti, Antonio Greco 0001, Domenico Longobardi, Pierluigi Ritrovato, Alessia Saggese, Mario Vento
CAIP (2)5
2025 Leveraging Vision-Language Models for Improving Detection of Obstacles on Railway Tracks
Vincenzo Carletti, Antonio Greco 0001, Alessia Saggese, Camilla Spingola, Bruno Vento
CAIP (2)3
2025 Multimodal Audio-Visual Emotion Recognition for Social Robotics
Giuseppe De Simone, Luca Greco 0001, Alessia Saggese, Mario Vento
CAIP (2)3
2024 Empowering Human Interaction: A Socially Assistive Robot for Support in Trade Shows
abstract
Social robots are increasingly finding applications in sectors such as industry, retail, and healthcare. They employ a combination of verbal and non-verbal cues in order to allow an empathetic and efficient interaction with the humans. In contrast to traditional robots, social robots possess contextual awareness, enabling intelligent responses to human interactions. Within this context, in this paper we propose a framework for social robots, integrating advanced audio and video analytics capabilities with a novel engagement algorithm, designed and developed to facilitate effective communication in multi-user settings. Furthermore, other than conversation skills, the proposed social robot is enhanced with the capability to interactively play games with the human. The effectiveness of the proposed system was tested in real-world scenarios, during two trade shows in Italy, providing valuable insights into its performance and adaptability to different contexts and different audiences.
Giuseppe De Simone, Alessia Saggese, Mario Vento
RO-MAN2
2024 Robust speech command recognition in challenging industrial environments
Stefano Bini, Vincenzo Carletti, Alessia Saggese, Mario Vento
Comput. Commun.3
2024 Facial Soft-biometrics Obfuscation through Adversarial Attacks
abstract
Sharing facial pictures through online services, especially on social networks, has become a common habit for thousands of users. This practice hides a possible threat to privacy: the owners of such services, as well as malicious users, could automatically extract information from faces using modern and effective neural networks. In this article, we propose the harmless use of adversarial attacks, i.e., variations of images that are almost imperceptible to the human eye and that are typically generated with the malicious purpose to mislead Convolutional Neural Networks (CNNs). Such attacks have been instead adopted to (1) obfuscate soft biometrics (gender, age, ethnicity) but (2) without degrading the quality of the face images posted online. We achieve the above-mentioned two conflicting goals by modifying the implementations of four of the most popular adversarial attacks, namely FGSM, PGD, DeepFool, and C&W, in order to constrain the average amount of noise they generate on the image and the maximum perturbation they add on the single pixel. We demonstrate, in an experimental framework including three popular CNNs, namely VGG16, SENet, and MobileNetV3, that the considered obfuscation method, which requires at most 4 seconds for each image, is effective not only when we have a complete knowledge of the neural network that extracts the soft biometrics (white box attacks) but also when the adversarial attacks are generated in a more realistic black box scenario. Finally, we prove that an opponent can implement defense techniques to partially reduce the effect of the obfuscation, but substantially paying in terms of accuracy over clean images; this result, confirmed by the experiments carried out with three popular defense methods, namely adversarial training, denoising autoencoder, and Kullback-Leibler autoencoder, shows that it is not convenient for the opponent to defend himself and that the proposed approach is robust to defenses.
Vincenzo Carletti, Pasquale Foggia, Antonio Greco 0001, Alessia Saggese, Mario Vento
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Highly Crowd Detection and Counting Based on Curriculum Learning
Lidia Fotia, Gennaro Percannella, Alessia Saggese, Mario Vento
CAIP (2)3
2023 Multi-task learning on the edge for effective gender, age, ethnicity and emotion recognition
Pasquale Foggia, Antonio Greco 0001, Alessia Saggese, Mario Vento
Eng. Appl. Artif. Intell.3
2023 A Social Robot Architecture for Personalized Real-Time Human-Robot Interaction
abstract
In the age of the Internet of Things (IoT), the combination of robotics and artificial intelligence has paved the way for the development of social robots able to undertake realistic conversations with humans, making them the perfect human interface in applications like for instance, robotic house assistants and hotel concierges. Despite the several solutions developed in recent years, the definition of social robot requirements and the software modules needed to meet such requirements has not yet been formalized. In this article, we define the requirements of a social robot and propose a software architecture that includes all the necessary modules to meet them. The proposed architecture, implemented using robot operating system (ROS) nodes, is hardware-independent, enabling its reuse across different robotic platforms with interchangeable types of sensors and actuators. We deployed a social robot based on this architecture to interact with attendees of a real exhibition context and validated the reliability of our proposed solution through a survey of 161 users. The results of the user study indicated a high-quality user experience with the social robot, with scores ranging from 4 to 5 (being 5 the maximum score). Sharing these design choices and evaluation results could significantly benefit the development of future social robotics applications in the context of IoT.
Pasquale Foggia, Antonio Greco 0001, Antonio Roberto, Alessia Saggese, Mario Vento
IEEE Internet Things J.4
2023 Degramnet: effective audio analysis based on a fully learnable time-frequency representation
abstract
Abstract Current state-of-the-art audio analysis algorithms based on deep learning rely on hand-crafted Spectrogram-like audio representations, that are more compact than descriptors obtained from the raw waveform; the latter are, in turn, far from achieving good generalization capabilities when few data are available for the training. However, Spectrogram-like representations have two main limitations: (1) The parameters of the filters are defined a priori, regardless of the specific audio analysis task; (2) such representations do not perform any denoising operation on the audio signal, neither in the time domain nor in the frequency domain. To overcome these limitations, we propose a new general-purpose convolutional architecture for audio analysis tasks that we call DEGramNet, which is trained with audio samples described with a novel, compact and learnable time–frequency representation that we call DEGram. The proposed representation is fully trainable: Indeed, it is able to learn the frequencies of interest for the specific audio analysis task; in addition, it performs denoising through a custom time–frequency attention module, which amplifies the frequency and time components in which the sound is actually located. It implies that the proposed representation can be easily adapted to the specific problem at hands, for instance giving more importance to the voice frequencies when the network needs to be used for speaker recognition. DEGramNet achieved state-of-the-art performance on the VGGSound dataset (for Sound Event Classification) and comparable accuracy with a complex and special-purpose approach based on network architecture search over the VoxCeleb dataset (for Speaker Identification). Moreover, we demonstrate that DEGram allows to achieve high accuracy with lightweight neural networks that can be used in real-time on embedded systems, making the solution suitable for Cognitive Robotics applications.
Pasquale Foggia, Antonio Greco 0001, Antonio Roberto, Alessia Saggese, Mario Vento
Neural Comput. Appl.4
2023 A deep learning based system for handwashing procedure evaluation
abstract
Abstract Hand washing preparation can be considered as one of the main strategies for reducing the risk of surgical site contamination and thus the infections risks. Within this context, in this paper we propose an embedded system able to automatically analyze, in real-time, the sequence of images acquired by a depth camera to evaluate the quality of the handwashing procedure. In particular, the designed system runs on an NVIDIA Jetson Nano $$^{{\mathrm{TM}}}$$ TM computing platform. We adopt a convolutional neural network, followed by a majority voting scheme, to classify the movement of the worker according to one of the ten gestures defined by the World Health Organization. To test the proposed system, we collect a dataset built by 74 different video sequences. The results achieved on this dataset confirm the effectiveness of the proposed approach.
Antonio Greco 0001, Gennaro Percannella, Pierluigi Ritrovato, Alessia Saggese, Mario Vento
Neural Comput. Appl.4
2023 A multi-task network for speaker and command recognition in industrial environments
abstract
In industrial environments, it is crucial to establish a strong collaboration between humans and robots to enhance productivity. However, the nature of the work demands that workers have the authority to provide specific instructions to the robots. The scientific community has extensively investigated these dual requirements, aiming to develop advanced systems capable of recognizing voice commands and implementing speaker authentication. Nevertheless, in the industrial context, these tasks should be executed simultaneously on low-cost and low-power embedded devices that can be mounted on board the robotic platform. To overcome this challenge, we propose a multi-task network for Speech-Command Recognition and Speaker Identification. Additionally, we employ the GradNorm adaptive algorithm to address the issue of task imbalance. To evaluate the proposed system, we introduce a new dataset, MIVIA-ISC, consisting of 20,857 samples uttered by 562 speakers for 31 distinct commands. Our approach significantly reduces the network size by 47% and its execution time by 48% compared to the commonly used methodology, which employs one network for each task. Furthermore, our approach demonstrates a significant improvement in the accuracy of the Speaker Identification task, achieving an 11% increase compared to the corresponding single-task network. Importantly, this enhancement is achieved without compromising the accuracy of the Speech-Command Recognition task, which experiences only a minimal 3% decrease in performance.
Stefano Bini, Gennaro Percannella, Alessia Saggese, Mario Vento
Pattern Recognit. Lett.3
2022 Benchmarking deep neural networks for gesture recognition on embedded devices
abstract
The gesture is one of the most used forms of communication between humans; in recent years, given the new trend of factories to be adapted to Industry 4.0 paradigm, the scientific community has shown a growing interest towards the design of Gesture Recognition (GR) algorithms for Human-Robot Interaction (HRI) applications. Within this context, the GR algorithm needs to work in real time and over embedded platforms, with limited resources. Anyway, when looking at the available scientific literature, the aim of the different proposed neural networks (i.e. 2D and 3D) and of the different modalities used for feeding the network (i.e. RGB, RGB-D, optical flow) is typically the optimization of the accuracy, without strongly paying attention to the feasibility over low power hardware devices. Anyway, the analysis related to the trade-off between accuracy and computational burden (for both networks and modalities) becomes important so as to allow GR algorithms to work in industrial robotics applications. In this paper, we perform a wide benchmarking focusing not only on the accuracy but also on the computational burden, involving two different architectures (2D and 3D), with two different backbones (MobileNet, ResNeXt) and four types of input modalities (RGB, Depth, Optical Flow, Motion History Image) and their combinations.
Stefano Bini, Antonio Greco 0001, Alessia Saggese, Mario Vento
RO-MAN3
2022 Effective training of convolutional neural networks for age estimation based on knowledge distillation
abstract
Abstract Age estimation from face images can be profitably employed in several applications, ranging from digital signage to social robotics, from business intelligence to access control. Only in recent years, the advent of deep learning allowed for the design of extremely accurate methods based on convolutional neural networks (CNNs) that achieve a remarkable performance in various face analysis tasks. However, these networks are not always applicable in real scenarios, due to both time and resource constraints that the most accurate approaches often do not meet. Moreover, in case of age estimation, there is the lack of a large and reliably annotated dataset for training deep neural networks. Within this context, we propose in this paper an effective training procedure of CNNs for age estimation based on knowledge distillation, able to allow smaller and simpler “student” models to be trained to match the predictions of a larger “teacher” model. We experimentally show that such student models are able to almost reach the performance of the teacher, obtaining high accuracy over the LFW+, LAP 2016 and Adience datasets, but being up to 15 times faster. Furthermore, we evaluate the performance of the student models in the presence of image corruptions, and we demonstrate that some of them are even more resilient to these corruptions than the teacher model.
Antonio Greco 0001, Alessia Saggese, Mario Vento, Vincenzo Vigilante
Neural Comput. Appl.2
2022 Vehicles Detection for Smart Roads Applications on Board of Smart Cameras: A Comparative Analysis
abstract
Video analytics can be profitably adopted in smart roads environments to automatically detect abnormal situations. Within this context, vehicle detection is the first and foremost stage, and its accuracy is crucial, since any detection error will affect the performance of any subsequent step. Furthermore, in smart road environments it is often preferred to perform the video analysis directly on board of smart surveillance cameras, in order to reduce bandwidth usage and eliminate the cost of setup and maintenance of powerful processing servers; on the flip side, processing on board of smart cameras implies the detection algorithm to be fast and slim, since the resources available on this kind of embedded device are limited. In the era of deep learning, it seems that the questionwhat is the best method for vehicle detection?may have a trivial answer, since this class of methods includes some very accurate ones. Anyway, according to the above consideration, the best suited method for this application is not necessarily the most accurate one, but for sure the most accurate one running on the available hardware at a given resolution and frame rate. Starting from the above considerations, in this paper we perform an analysis of the methods available in the literature for vehicle detection, by comparing them in terms of accuracy and computational burden, with the aim to answer the following question:what is the best method for vehicles detection when working with smart cameras?
Antonio Greco 0001, Alessia Saggese, Mario Vento, Vincenzo Vigilante
IEEE Trans. Intell. Transp. Syst.2
2021 DENet: a deep architecture for audio surveillance applications
Antonio Greco 0001, Antonio Roberto, Alessia Saggese, Mario Vento
Neural Comput. Appl.3
2020 Which are the factors affecting the performance of audio surveillance systems?
abstract
Sound event recognition systems are rapidly becoming part of our life, since they can be profitably used in several vertical markets, ranging from audio security applications to scene classification and multi-modal analysis in social robotics. In the last years, a not negligible part of the scientific community started to apply Convolutional Neural Networks (CNNs) to image-based representations of the audio stream, due to their successful adoption in almost all the computer vision tasks. In this paper, we carry out a detailed benchmark of various widely used CNN architectures and visual representations on a popular dataset, namely the MIVIA Audio Events database. Our analysis is aimed at understanding how these factors affect the sound event recognition performance with a particular focus on the false positive rate, very relevant in audio surveillance solutions. In fact, although most of the proposed solutions achieve a high recognition rate, the capability of distinguishing the events-of-interest from the background is often not yet sufficient for real systems, and prevent its usage in real applications. Our comprehensive experimental analysis investigates this aspect and allows to identify useful design guidelines for increasing the specificity of sound event recognition systems.
Antonio Greco 0001, Antonio Roberto, Alessia Saggese, Mario Vento
ICPR3
2020 Comparing performance of graph matching algorithms on huge graphs
Vincenzo Carletti, Pasquale Foggia, Antonio Greco 0001, Alessia Saggese, Mario Vento
Pattern Recognit. Lett.4
2020 AReN: A Deep Learning Approach for Sound Event Recognition Using a Brain Inspired Representation
abstract
Audio surveillance is gaining in the last years wide interest. This is due to the large number of situations in which this kind of systems can be used, either alone or combined with video-based algorithms. In this paper we propose a deep learning method to automatically recognize events of interest in the context of audio surveillance (namely screams, broken glasses and gun shots). The audio stream is represented by a gammatonegram image. We propose a 21-layer CNN to which we feed sections of the gammatonegram representation. At the output of this CNN there are units that correspond to the classes. We trained the CNN, called AReN, by taking advantage of a problem-driven data augmentation, which extends the training dataset with gammatonegram images extracted by sounds acquired with different signal to noise ratios. We experimented it with three datasets freely available, namely SESA, MIVIA Audio Events and MIVIA Road Events and we achieved 91.43%, 99.62% and 100% recognition rate, respectively. We compared our method with other state of the art methodologies based both on traditional machine learning methodologies and deep learning. The comparison confirms the effectiveness of the proposed approach, which outperforms the existing methods in terms of recognition rate. We experimentally prove that the proposed network is resilient to the noise, has the capability to significantly reduce the false positive rate and is able to generalize in different scenarios. Furthermore, AReN is able to process 5 audio frames per second on a standard CPU and, consequently, it is suitable for real audio surveillance applications.
Antonio Greco 0001, Nicolai Petkov, Alessia Saggese, Mario Vento
IEEE Trans. Inf. Forensics Secur.3
2019 A System for Controlling How Carefully Surgeons Are Cleaning Their Hands
Luca Greco 0001, Gennaro Percannella, Pierluigi Ritrovato, Alessia Saggese, Mario Vento
CAIP (2)4
2019 A Challenging Voice Dataset for Robotic Applications in Noisy Environments
Antonio Roberto, Alessia Saggese, Mario Vento
CAIP (2)2
2019 MIVIABot: A Cognitive Robot for Smart Museum
Alessia Saggese, Mario Vento, Vincenzo Vigilante
CAIP (1)1
2019 Emotion analysis from faces for social robotics
abstract
A social robot is able to perceive the information about the environment (both in terms of persons and objects populating the scene), to reason about the acquired information and to interact with the human in a proper way. Among the information required for the interaction with a human, the capability of analysing the emotion is surely among the most important ones. Another relevant feature of social robots is the possibility to interact with the human in real time, without any latency that could introduce a delay in the talk between the human and the robot. It means that all the processing of the information (for instance the analysis of the sequence of images for the detection of the persons and the consequent emotion analysis) needs to be performed directly on board of the robot, without any possibility to use high performance servers (for instance services on the cloud) but only small devices that can be installed directly on board of the robotic platform. In this paper we propose MIVIAbot, a robotic platform based on the Pepper humanoid robot and equipped with a small, low cost and low energy consumption embedded device installed directly on board of Pepper, without any requirements for a Wi-Fi connection that could introduce any latency for the transmission of the data. Furthermore, we also propose a method for the analysis of the emotion of the persons by face analysis, able to run on the considered hardware platform in real time without paying in terms of accuracy. The experimentation conducted over various widely adopted datasets of images and videos for emotion analysis confirms both the efficiency and the effectiveness of the proposed approach.
Antonio Greco 0001, Antonio Roberto, Alessia Saggese, Mario Vento, Vincenzo Vigilante
SMC3
2019 SoReNet: a novel deep network for audio surveillance applications
abstract
In the era of third generation surveillance systems, it becomes more and more useful to have available a solution able to automatically detect abnormal events. The interest for audio analysis is thus growing in the last years, due to the large amount of situations where a microphone and an audio surveillance system can be profitably used by the human operator in charge of control. In this paper, we propose a method for automatically analyzing the audio stream for surveillance purposes: it is able to detect the presence of abnormal events such as screams, gun shots and broken glasses. Instead than processing directly raw data (the audio signal), the stream is represented by means of an image, namely the spectrogram, a time-frequency representation of the audio stream. In this way, we formulate the problem of audio analysis as a problem of image classification. Thus, we propose to use a Convolutional Neural Network with the following two main properties: inspired by VGG network, we employed very small kernels in convolutional layers; furthermore, we adopted a pyramidal structure in fully connected layers. These choices allow to have good generalization capabilities of the network even in presence of a not so wide dataset. The performance, computed over a standard dataset already used for benchmarking purposes in the field of audio surveillance, confirms the effectiveness of the proposed approach.
Antonio Greco 0001, Alessia Saggese, Mario Vento, Vincenzo Vigilante
SMC2
2019 A human-like description of scene events for a proper UAV-based video content analysis
Danilo Cavaliere, Vincenzo Loia, Alessia Saggese, Sabrina Senatore, Mario Vento
Knowl. Based Syst.3
2019 Learning skeleton representations for human action recognition
abstract
Automatic interpretation of human actions gained strong interest among researchers in patter recognition and computer vision because of its wide range of applications, such as in social and home robotics, elderly people health care, surveillance, among others. In this paper, we propose a method for recognition of human actions by analysis of skeleton poses. The method that we propose is based on novel trainable feature extractors, which can learn the representation of prototype skeleton examples and can be employed to recognize skeleton poses of interest. We combine the proposed feature extractors with an approach for classification of pose sequences based on string kernels. We carried out experiments on three benchmark data sets (MIVIA-S, MSRSDA and MHAD) and the results that we achieved are comparable or higher than the ones obtained by other existing methods. A further important contribution of this work is the MIVIA-S dataset, that we collected and made publicly available.
Alessia Saggese, Nicola Strisciuglio, Mario Vento, Nicolai Petkov
Pattern Recognit. Lett.1
2019 Semantically Enhanced UAVs to Increase the Aerial Scene Understanding
abstract
Visual tracking supported by unmanned aerial vehicles (UAVs) has generated a lot of interest in recent years, especially in application domains such as surveillance, search for missing persons and traffic monitoring. The major challenges in visual tracking with small UAVs arise in the form of target representation, target appearance change, target detection and localization in real time computation. Reliable target detection depends on factors such as occlusions, image noise, illumination and pose changes, or image blur that may compromise the object labeling. To mitigate these issues, this paper proposes a hybrid solution: along with the tracked objects, scenes are completely depicted by adding contextual information, i.e., data describing places, natural features, or in general points of interest. Each scenario indeed is semantically described by ontological statements that define the context and then, by inference, support the object tracking task in the object identification and labeling. The synergy between the tracking methods and semantic modeling can bridge the object labeling gap, enhancing the scene understanding and awareness when alarming situations are discovered. Experimental results are promising and confirm the applicability of the proposed framework in supporting drones in object identification and critical situation detection tasks.
Danilo Cavaliere, Vincenzo Loia, Alessia Saggese, Sabrina Senatore, Mario Vento
IEEE Trans. Syst. Man Cybern. Syst.3
2018 Gender recognition from face images using trainable shape and color features
abstract
Gender recognition from face images is an important application and it is still an open computer vision problem, even though it is something trivial from the human visual system. Variations in pose, lighting, and expression are few of the problems that make such an application challenging for a computer system. Neurophysiological studies demonstrate that the human brain is able to distinguish men and women also in absence of external cues, by analyzing the shape of specific parts of the face. In this paper, we describe an automatic procedure that combines trainable shape and color features for gender classification. In particular the proposed method fuses edge-based and color-blob-based features by means of trainable COSFIRE filters. The former types of feature are able to extract information about the shape of a face whereas the latter extract information about shades of colors in different parts of the face. We use these two sets of features to create a stacked classification SVM model and demonstrate its effectiveness on the GENDER-COLOR-FERET dataset, where we achieve an accuracy of 96.4%.
George Azzopardi, Pasquale Foggia, Antonio Greco 0001, Alessia Saggese, Mario Vento
ICPR4
2018 Challenging the Time Complexity of Exact Subgraph Isomorphism for Huge and Dense Graphs with VF3
abstract
Graph matching is essential in several fields that use structured information, such as biology, chemistry, social networks, knowledge management, document analysis and others. Except for special classes of graphs, graph matching has in the worst-case an exponential complexity; however, there are algorithms that show an acceptable execution time, as long as the graphs are not too large and not too dense. In this paper we introduce a novel subgraph isomorphism algorithm, VF3, particularly efficient in the challenging case of graphs with thousands of nodes and a high edge density. Its performance, both in terms of time and memory, has been assessed on a large dataset of 12,700 random graphs with a size up to 10,000 nodes, made publicly available. VF3 has been compared with four other state-of-the-art algorithms, and the huge experimentation required more than two years of processing time. The results confirm that VF3 definitely outperforms the other algorithms when the graphs become huge and dense, but also has a very good performance on smaller or sparser graphs.
Vincenzo Carletti, Pasquale Foggia, Alessia Saggese, Mario Vento
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 Fast gender recognition in videos using a novel descriptor based on the gradient magnitudes of facial landmarks
abstract
The growing interest in recent years for gender recognition from face images is mainly attributable to the wide range of possible applications that can be used for commercial and marketing purposes. It is desirable that such algorithms process high resolution video frames acquired by using surveillance cameras in real-time. To the best of our knowledge, however, there are no studies which analyze the computational impact of the methods and the difficulties related to the processing of faces extracted from videos captured in the wild. We propose a novel face descriptor based on the gradient magnitudes of facial landmarks, which are points automatically extracted from the face contour, eyes, eyebrows, nose, mouth and chin. We evaluate the effectiveness and efficiency of the proposed approach on two new datasets, which we made available online and that consist of color face images and color video sequences acquired in real scenarios. The proposed approach is more efficient and effective than three commercial libraries.
George Azzopardi, Antonio Greco 0001, Alessia Saggese, Mario Vento
AVSS3
2017 A real-time system for audio source localization with cheap sensor device
abstract
We propose an architecture for real-time audio source localization based on the integration of localization methodologies within a framework that employs a cheap acquisition sensor. The architecture that we present takes as input the audio signals from two calibrated microphones. Then, it computes biological-inspired features of the sound signal and estimates its direction by means of a Gaussian Mixture Model estimator. We carried out an extensive experimental analysis on four data sets, one of which we realized and made publicly available. We evaluated several characteristics of the sound localization architecture and its use in real scenarios.
Alessia Saggese, Nicola Strisciuglio, Mario Vento, Nicolai Petkov
AVSS1
2016 Improving reliability of people tracking by adding semantic reasoning
abstract
Even the best performing object tracking algorithm on well known datasets, commits several errors that prevent a concrete adoption in real case scenarios unless you do not accept some compromise about tracking quality and reliability. The aim of this paper is to demonstrate that adding to a traditional object tracking solution a knowledge based reasoner build on top of semantic web technologies, it is possible to identify and properly manage common tracking problems. The proposed approach has been evaluated using View 001 and View 003 of the PETS2009 dataset with interesting results.
Luca Greco 0001, Pierluigi Ritrovato, Alessia Saggese, Mario Vento
AVSS3
2016 Time-frequency analysis for audio event detection in real scenarios
abstract
We propose a sound analysis system for the detection of audio events in surveillance applications. The method that we propose combines short- and long-time analysis in order to increase the reliability of the detection. The basic idea is that a sound is composed of small, atomic audio units and some of them are distinctive of a particular class of sounds. Similarly to the words in a text, we count the occurrence of audio units for the construction of a feature vector that describes a given time interval. A classifier is then used to learn which audio units are distinctive for the different classes of sound. We compare the performance of different sets of short-time features by carrying out experiments on the MIVIA audio event data set. We study the performance and the stability of the proposed system when it is employed in live scenarios, so as to characterize its expected behavior when used in real applications.
Alessia Saggese, Nicola Strisciuglio, Mario Vento, Nicolai Petkov
AVSS1
2016 International Contest on Pattern Recognition techniques for indirect immunofluorescence images analysis
abstract
This contest is a joint initiative organized by the University of Salerno (Italy) and the University of Queensland (Australia) with the support of the Sullivan Nicolaides Pathology (SNP), Australia. The contest primarily aims to provide a platform for scientists and practitioners for performing research to develop Computer Aided Diagnosis (CAD) systems for pathology tests utilizing indirect immunofluorescence protocol. In particular, the contest considers the Antinuclear Antibodies (ANA) test using Human Epithelial type 2 (HEp-2) cells. The competition is divided into four tasks that address specific problems: (1) HEp-2 cell classification; (2) Patient specimen classification; (3) HEp-2 mitotic cell identification and (4) Cell segmentation.
Brian C. Lovell, Gennaro Percannella, Alessia Saggese, Mario Vento, Arnold Wiliem
ICPR3
2016 Action recognition by using kernels on aclets sequences
Luc Brun, Gennaro Percannella, Alessia Saggese, Mario Vento
Comput. Vis. Image Underst.3
2016 Executable thematic special issue on pattern recognition techniques for indirect immunofluorescence images analysis
Mehrtash Harandi, Brian C. Lovell, Gennaro Percannella, Alessia Saggese, Mario Vento, Arnold Wiliem
Pattern Recognit. Lett.4
2016 Computer Aided Diagnosis for Anti-Nuclear Antibodies HEp-2 images: Progress and challenges
Peter Hobson, Brian C. Lovell, Gennaro Percannella, Alessia Saggese, Mario Vento, Arnold Wiliem
Pattern Recognit. Lett.4
2016 HEp-2 staining pattern recognition at cell and specimen levels: Datasets, algorithms and results
Peter Hobson, Brian C. Lovell, Gennaro Percannella, Alessia Saggese, Mario Vento, Arnold Wiliem
Pattern Recognit. Lett.4
2016 Audio Surveillance of Roads: A System for Detecting Anomalous Sounds
abstract
In the last decades, several systems based on video analysis have been proposed for automatically detecting accidents on roads to ensure a quick intervention of emergency teams. However, in some situations, the visual information is not sufficient or sufficiently reliable, whereas the use of microphones and audio event detectors can significantly improve the overall reliability of surveillance systems. In this paper, we propose a novel method for detecting road accidents by analyzing audio streams to identify hazardous situations such as tire skidding and car crashes. Our method is based on a two-layer representation of an audio stream: at a low level, the system extracts a set of features that is able to capture the discriminant properties of the events of interest, and at a high level, a representation based on a bag-of-words approach is then exploited in order to detect both short and sustained events. The deployment architecture for using the system in real environments is discussed, together with an experimental analysis carried out on a data set made publicly available for benchmarking purposes. The obtained results confirm the effectiveness of the proposed approach.
Pasquale Foggia, Nicolai Petkov, Alessia Saggese, Nicola Strisciuglio, Mario Vento
IEEE Trans. Intell. Transp. Syst.3
2015 Automatic detection of long term parked cars
abstract
The detection of illegal roadside parking is becoming more and more interesting in the field of intelligent transportation systems, since it may cause traffic congestion or accidents. In this paper we propose a method able to analyze videos acquired by traditional surveillance cameras and to automatically detect the vehicles stopped in a forbidden area. Two main contributions have been introduced: first, spatio temporal information related to the stopped vehicles are encoded by a heat map; second, the background is not updated by evaluating the movement of the vehicle in a single time instant, but instead the whole movement of the vehicles, encoded into the heat map, is taken into account. Two widely adopted datasets, namely the iLids and the PETS 2000, have been used to experimentally evaluate the proposed approach and the results achieved, compared with state of the art methodologies, confirm its effectiveness.
Vincenzo Carletti, Pasquale Foggia, Antonio Greco 0001, Alessia Saggese, Mario Vento
AVSS4
2015 Human action recognition using an improved string edit distance
abstract
In this paper we propose an improvement of a human action recognition method that uses a string-based representation and a string edit distance to compare the observed action with reference actions in the training set. In particular, the original improvement is based on a specific formulation of the string edit distance that is more suited to take into account the problems related to noise and to different execution speeds that are observed in an action recognition system. The experimentation has been carried out on two widely adopted datasets, namely the MIVIA and the MHAD datasets, and the obtained results, compared with both the original method and other state of the art approaches, confirm the significance of the proposed improvement and the effectiveness of the method.
Pasquale Foggia, Benoit Gaüzère, Alessia Saggese, Mario Vento
AVSS3
2015 Car crashes detection by audio analysis in crowded roads
abstract
In the last years, video surveillance has been employed for roads monitoring in order to detect abnormal events and improve the safety procedures in case of emergency. Certain events, such as car crashes or tire skidding, are difficult or impossible to detect when only the visual information is considered. In this paper we describe a preliminary system to detect events in roads by means of audio analysis. The system that we propose combines short- and long-time analysis of the audio signal in order to detect both impulsive and sustained events. We present the preliminary results achieved by the proposed system on a data set specifically made for roads surveillance, which we made publicly available. We also discuss the architectural deployment of such system in real environments with respect to a model of the noise of road traffic. The achieved results are promising and confirm the effectiveness of the system.
Pasquale Foggia, Alessia Saggese, Nicola Strisciuglio, Mario Vento, Nicolai Petkov
AVSS2
2015 Locally Adapted Gain Control for Reliable Foreground Detection
Duber Martinez, Alessia Saggese, Mario Vento, Humberto Loaiza, Eduardo F. Caicedo
CAIP (1)2
2015 A hierarchical neuro-fuzzy architecture for human behavior analysis
Giovanni Acampora, Pasquale Foggia, Alessia Saggese, Mario Vento
Inf. Sci.3
2015 Reliable detection of audio events in highly noisy environments
Pasquale Foggia, Nicolai Petkov, Alessia Saggese, Nicola Strisciuglio, Mario Vento
Pattern Recognit. Lett.3
2015 Real-Time Fire Detection for Video-Surveillance Applications Using a Combination of Experts Based on Color, Shape, and Motion
abstract
In this paper, we propose a method that is able to detect fires by analyzing videos acquired by surveillance cameras. Two main novelties have been introduced. First, complementary information, based on color, shape variation, and motion analysis, is combined by a multiexpert system. The main advantage deriving from this approach lies in the fact that the overall performance of the system significantly increases with a relatively small effort made by the designer. Second, a novel descriptor based on a bag-of-words approach has been proposed for representing motion. The proposed method has been tested on a very large dataset of fire videos acquired both in real environments and from the web. The obtained results confirm a consistent reduction in the number of false positives, without paying in terms of accuracy or renouncing the possibility to run the system on embedded platforms.
Pasquale Foggia, Alessia Saggese, Mario Vento
IEEE Trans. Circuits Syst. Video Technol.2
2015 Designing Huge Repositories of Moving Vehicles Trajectories for Efficient Extraction of Semantic Data
abstract
The rapid development of digital cameras equipped with video analytics software is providing the availability of large amount of traffic data describing the trajectories traced by each vehicle and person within a scene. These data offer enormous potential when coupled with a querying system able to extract synthetic but meaningful information as those obtained by spatiotemporal queries; the latter allow, for instance, to select all those trajectories passing through some parts of the scene, even in given sequences, and adding restrictions on the properties of the objects (the category of the vehicles, their color and size, and so on). In this paper we propose a novel system for efficiently storing and querying large amounts of 3D data (trajectories over time), specifically designed for making possible the formulation of a wide variety of spatio-temporal 3D queries. The method is based on a novel 3D data schema which is reconducted to a set of 2D schemata, being the latter the only ones available in currently ready-to-use database environments. An implementation of the system over PostGIS is presented in this paper, together with a performance assessment on a huge trajectory database. The obtained results confirm the effectiveness of the proposed approach and its applicability to real applications.
Antonio d'Acierno, Alessia Saggese, Mario Vento
IEEE Trans. Intell. Transp. Syst.2
2014 Detection of anomalous driving behaviors by unsupervised learning of graphs
abstract
In this paper we propose a graph based approach for detecting abnormal behaviors starting from the analysis of vehicles' trajectories. The scene is partitioned into zones and is dynamically represented as a graph by evaluating the distribution of trajectories belonging to the training set. Furthermore, four different strategies are proposed in order to verify if a test trajectory belongs to the scene and then can be considered normal by evaluating the probability that this trajectory belongs to the graph. Our algorithms have been tested on the standard MIT Trajectories dataset and the obtained results confirm the effectiveness of the proposed approach.
Luc Brun, Benito Cappellania, Alessia Saggese, Mario Vento
AVSS3
2014 HAck: A system for the recognition of human actions by kernels of visual strings
abstract
In this paper we propose HAcK, a novel method for recognizing Human Actions by string Kernel; the main idea is to represent each action through a sequence of visual characters, namely a string, able to model the temporal evolution of the events. Visual characters are extracted by analyzing global descriptors of the scene and by taking advantage on the depth information provided by a Kinect sensor. The similarity between actions is evaluated with a fast global alignment kernel, which allows to deal with actions of different length as well as with the noise introduced during the features extraction step. HAcK has been evaluated over two standard datasets and the obtained results, compared with state of the art approaches, confirm its effectiveness and its applicability in real environments.
Luc Brun, Gennaro Percannella, Alessia Saggese, Mario Vento
AVSS3
2014 A reliable string kernel based approach for solving queries by sketch
abstract
In this paper we propose a novel and efficient method for solving queries by sketch in traffic scenarios, aiming to find the k nearest neighbor trajectories to the one hand drawn by the human operator. Each trajectory is represented as a sequence of symbols, namely a string, and it is stored into a k-d tree by taking into account the similarity between trajectories, evaluated by a global fast alignment kernel. The experimentation has been conducted over the standard MIT trajectories dataset and results confirm the effectiveness and the robustness of the proposed approach.
Luc Brun, Alessia Saggese, Mario Vento
AVSS2
2014 Cascade classifiers trained on gammatonegrams for reliably detecting audio events
abstract
In this paper we propose a novel method for the detection of events of interest through audio analysis. The system that we propose is based on the representation of the audio streams through a Gammatone image, which describes the time-frequency distribution of the energy of the signal; this representation is inspired by the functioning of the human auditory system. A pool of AdaBoost cascade classifiers, one for each class of events of interest, is involved in the event detection stage. The performance of the proposed system has been evaluated on a large data set of audio events for surveillance applications and the achieved results, compared with two state of the art approaches, confirm its effectiveness.
Pasquale Foggia, Alessia Saggese, Nicola Strisciuglio, Mario Vento
AVSS2
2014 Exploiting the deep learning paradigm for recognizing human actions
abstract
In this paper we propose a novel method for recognizing human actions by exploiting a multi-layer representation based on a deep learning based architecture. A first level feature vector is extracted and then a high level representation is obtained by taking advantage of a Deep Belief Network trained using a Restricted Boltzmann Machine. The classification is finally performed by a feed-forward neural network. The main advantage behind the proposed approach lies in the fact that the high level representation is automatically built by the system exploiting the regularities in the dataset; given a suitably large dataset, it can be expected that such a representation can outperform a hand-design description scheme. The proposed approach has been tested on two standard datasets and the achieved results, compared with state of the art algorithms, confirm its effectiveness.
Pasquale Foggia, Alessia Saggese, Nicola Strisciuglio, Mario Vento
AVSS2
2014 Pattern recognition in stained HEp-2 cells: Where are we now?
Pasquale Foggia, Gennaro Percannella, Alessia Saggese, Mario Vento
Pattern Recognit.3
2014 Dynamic Scene Understanding for Behavior Analysis Based on String Kernels
abstract
This paper aims at dynamically understanding the properties of a scene from the analysis of moving object trajectories. Two different applications are proposed: the former is devoted to identify abnormal behaviors, while the latter allows to extract the k, most of the similar trajectories to the one hand-drawn by an human operator. A set of normal trajectories' models is extracted using a novel unsupervised learning technique: the scene is adaptively partitioned into zones using the distribution of the training set and each trajectory is represented as a sequence of symbols by considering positional information (the zones crossed in the scene), speed, and shape. The main novelty is the use of a kernel-based approach for evaluating the similarity between the trajectories. Furthermore, we define a novel and efficient kernel-based clustering algorithm, aimed at obtaining groups of normal trajectories. Experimentations, conducted over three standard data sets, confirm the effectiveness of the proposed approach.
Luc Brun, Alessia Saggese, Mario Vento
IEEE Trans. Circuits Syst. Video Technol.2
2013 Audio surveillance using a bag of aural words classifier
abstract
In this paper we propose a novel approach for the audio-based detection of events. The approach adopts the bag of words paradigm, and has two main advantages over other techniques present in the literature: the ability to automatically adapt (through a learning phase) to both short, impulsive sounds and long, sustained ones, and the ability to work in noisy environments where the sounds of interest are superimposed to background sounds possibly having similar characteristics. The proposed method has been experimentally validated on a large database of sounds, including several kinds of background noise, which are superimposed to the sounds to be recognized. The obtained performance has been compared with the results of another audio event detection algorithm from the literature, showing a significant improvement.
Vincenzo Carletti, Pasquale Foggia, Gennaro Percannella, Alessia Saggese, Nicola Strisciuglio, Mario Vento
AVSS4
2013 Recognizing Human Actions by a Bag of Visual Words
abstract
In this paper a novel method for action recognition based on the bag of visual words approach is proposed. The main contribution is to model each action through a high level features vector computed as the histogram of the visual words: the visual words are extracted by analyzing global descriptors of the scene and their occurrences are evaluated according to a codebook, a kind of dictionary, which encodes the typical visual words, automatically extracted during the learning phase. The classification is performed by using an SVM classifier, trained only by using high level features vectors, in order to increase the overall reliability of the system. The experimentation has been conducted over two recently proposed datasets, the MIVIA and the MHAD, the promising results confirm the robustness and the stability of the proposed approach.
Pasquale Foggia, Gennaro Percannella, Alessia Saggese, Mario Vento
SMC3
2013 A real time algorithm for people tracking using contextual reasoning
Rosario Di Lascio, Pasquale Foggia, Gennaro Percannella, Alessia Saggese, Mario Vento
Comput. Vis. Image Underst.4
2012 Combining Neural Networks and Fuzzy Systems for Human Behavior Understanding
abstract
The psychological overcharge issue related to human inadequacy to maintain a constant level of attention in simultaneously monitoring multiple visual information sources makes necessary to develop enhanced video surveillance systems that automatically understand human behaviors and identify dangerous situations. This paper introduces a semantic human behavioral analysis (HBA) system based on a neuro-fuzzy approach that, independently from the specific application, translates tracking kinematic data into a collection of semantic labels characterizing the behavior of different actors in a scene in order to appropriately classify the current situation. Different from other HBA approaches, the proposed system shows high level of scalability, robustness and tolerance for tracking imprecision and, for this reason, it could represent a valid choice for improving the performance of current systems.
Giovanni Acampora, Pasquale Foggia, Alessia Saggese, Mario Vento
AVSS3
2012 An Ensemble of Rejecting Classifiers for Anomaly Detection of Audio Events
abstract
Audio analytic systems are receiving an increasing interest in the scientific community, not only as stand alone systems for the automatic detection of abnormal events by the interpretation of the audio track, but also in conjunction with video analytics tools for enforcing the evidence of anomaly detection. In this paper we present an automatic recognizer of a set of abnormal audio events that works by extracting suitable features from the signals obtained by microphones installed into a surveilled area, and by classifying them using two classifiers that operate at different time resolutions. An original aspect of the proposed system is the estimation of the reliability of each response of the individual classifiers. In this way, each classifier is able to reject the samples having an overall reliability below a threshold. This approach allows our system to combine only reliable decisions, so increasing the overall performance of the method. The system has been tested on a large dataset of samples acquired from real world scenarios, the audio classes of interests are represented by gunshot, scream and glass breaking in addition to the background sounds. The preliminary results obtained encourage further research in this direction.
Donatello Conte, Pasquale Foggia, Gennaro Percannella, Alessia Saggese, Mario Vento
AVSS4