Antonio Greco 0001

dblp:55/5022-1 · DBLP profile ↗
← Back
45ranked-venue papers
18as first author
32since 2021 · last 2026
0000-0002-5495-2432ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 10 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 9 since 2021Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 1 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 ShapeBlend: Boosting out-of-distribution robustness in image classification via shape-based blending augmentation
abstract
Deep Neural Networks (DNNs) often struggle to generalize beyond their training distributions, making them vulnerable to domain shifts. To enhance robustness, various approaches have been developed, particularly focusing on data-centric methods. Modifying the training data can increase diversity and improve generalization, while also introducing bias that positively guides the model’s decision-making. Prior research suggests that DNNs tend to overemphasize texture-based patterns, at the expense of more robust shape-based representations. We introduce ShapeBlend, a novel data augmentation technique that emphasizes image contours, and hence shape features. It blends a contour map from the push-pull CORF operator with the original image at varying strengths. ShapeBlend consistently outperforms state-of-the-art methods across major Out-of-Distribution (OOD) benchmarks (ImageNet-A, ImageNet-R, ImageNet-C, and ImageNet- C ¯ ), setting new records in robustness. Moreover, ShapeBlend’s versatility allows its application during inference. To fully leverage ShapeBlend, we propose Shape-Enhanced Voting (SEV), an inference strategy that aggregates predictions from multiple ShapeBlend-processed images. The combination of ShapeBlend and SEV further enhances domain robustness, with performance gains varying based on the chosen configuration. • We introduce ShapeBlend: a shape-based augmentation for OOD robustness. • ShapeBlend improves performance across multiple robustness benchmarks. • Inference performance is boosted with Shape-enhanced Voting (SEV). • ShapeBlend is compatible with existing pipelines like AugMix and DeepAugment. • A simple, explainable method with strong theoretical motivation.
George Azzopardi, Sabatino Esposito, Antonio Greco 0001, Mario Vento
Comput. Vis. Image Underst.3
2026 Simultaneous person attribute recognition using task-specific attention network on embedded devices
abstract
Pedestrian attribute recognition has become an important task in computer vision, particularly for retail and marketing and critical applications in security and surveillance. Despite its potential, achieving real-time performance on embedded devices while maintaining accuracy has been a significant challenge. In this paper, we propose a novel multi-task method for simultaneous person attribute recognition using task-specific attention network, which shares low-level representations across related tasks, reducing computational and memory requirements without compromising accuracy. In particular, we employ a spatial-channel attention mechanism to selectively focus on relevant regions without increasing the computational complexity of the backbone. Furthermore, we use a knowledge distillation technique to deal with missing labels and gradient normalization for dealing with task imbalances, since varying task difficulties lead to disproportionate gradient magnitudes during training. The experimental results demonstrate the effectiveness of our approach, achieving a mean accuracy of 0.889 while maintaining real-time performance at 114 frames per second on an embedded board with limited resources. These results highlight the practical viability and novelty of our system as a robust and scalable solution for pedestrian attribute recognition on embedded devices in real-world scenarios.
George Azzopardi, Antonio Greco 0001, Alessia Saggese, Bruno Vento
Eng. Appl. Artif. Intell.2
2026 Joint underwater image enhancement and multi-scale fish detection
abstract
Automatic fish detection in images plays a crucial role in marine biodiversity monitoring and environmental conservation. The advent of deep learning has significantly improved the results achievable in fish detection; however, underwater challenges such as scale variability, optical distortions, and color inconsistencies reveal some limitations of conventional deep learning-based detectors. To address these issues, we propose an improved fish detection model designed for underwater environments trained to perform joint underwater image enhancement and multi-scale fish detection. While incorporating structural enhancements for multi-scale robustness, such as a dedicated tiny object detection head and attention mechanisms, the core of our approach resides in the joint training of an underwater image enhancement module with the detector. We evaluate our method on two challenging test sets. The experimental results demonstrate that our model significantly outperforms state-of-the-art methods, achieving an average precision between 0.84 and 0.89. Despite its superior accuracy, the proposed solution maintains a lightweight architecture with only 24 millions of parameters and ensures real-time processing capabilities (13 frames per second on edge hardware), highlighting its potential for effective deployment in marine research and autonomous fisheries monitoring. • YOLO-JUICE boosts small-object detection with an extra tiny-object head. • YOLO-JUICE preserves details without extra cost with a space to depth module. • YOLO-JUICE refines spatial and channel features with a spatial-channel attention. • YOLO-JUICE is jointly trained for image classification and enhancement.
Vincenzo Carletti, Antonio Greco 0001, Andrea Vincenzo Ricciardi, Alessia Saggese, Bruno Vento
Eng. Appl. Artif. Intell.2
2026 Multi-task road surface condition recognition through a self-conditioned multi-label approach
abstract
Autonomous vehicle applications are gaining growing attention for their increasing reliability and capability to reduce the ecological impact of transportation. However, most prototypes fail to perceive the conditions of the road surface, which are crucial to adapt the driving controllers and avoid risks, e.g., due to hydroplaning. We propose a method to identify the major causes of grip reduction, namely friction, unevenness, and type of material, using a sequential architecture based on an LSTM module. As the multiple combination of these conditions alters the appearance of the ground surface, our model is trained to identify the most evident conditions first, then to use the predicted information to infer the less evident conditions. After the sequential prediction, a final classification stage allows for surpassing the limits of the sequential approach and recovering errors that occurred in early iterations, thus reducing their impact on subsequent ones. Experimental results show that the method outperforms previous ones on a challenging 27-class and 3-task dataset, by about 2.5 and 5.0 percentage points in terms of overall and balanced accuracy, respectively, while marginally increasing the inference time. Moreover, the conducted thorough error analysis demonstrates that our approach can implicitly learn the application constraints, like the absence of the material label if the surface is covered by snow, and provide useful information to interpret the sequence of predictions produced at each iteration of the sequential architecture.
Diego Gragnaniello, Antonio Greco 0001, Mattia Marseglia, Carlo Sansone, Bruno Vento
Expert Syst. Appl.2
2026 Benchmarking illegal waste dumping detection: A public video dataset and a reference baseline
abstract
Illegal waste dumping represents one of the most pervasive threats to environmental health and the well-being of terrestrial ecosystems. As widely documented in the literature, this phenomenon is deeply rooted in cultural factors and civic behavior, making early detection a crucial component in effective prevention strategies. Recent advances in artificial intelligence offer promising avenues for automatically identifying illicit dumping actions; however, progress in this direction is severely hindered by the scarcity of publicly available video datasets tailored to the Illegal Waste Dumping Detection (IWDD) task. To address this critical gap, we introduce Mivia-IWDD-500 , a novel, fully balanced dataset comprising 500 videos: 250 positive samples and 250 negative samples. The positive class is further evenly divided into 125 videos depicting static disposal events, intentional and spatially localized acts such as depositing garbage bags or bulky items, and 125 videos capturing dynamic disposal behaviors, which involve brief, spontaneous, and spatially dispersed actions such as discarding waste while walking or from a moving vehicle. Alongside the dataset, we present a baseline model designed to serve as a reference point for future research. The proposed baseline model achieves an F 1 -score of 0.79, demonstrating the viability of the dataset and establishing a solid foundation for subsequent advancements.
Antonio Greco 0001, Andrea Vincenzo Ricciardi, Carlo Sansone, Bruno Vento
Image Vis. Comput.1
2026 STEP-FACE: Sequential TExtual-Visual Prompting for multi-task face analysis
abstract
Jointly analyzing stable facial attributes (gender, age) and transient affect (emotion) is a useful functionality for applications such as adaptive human–computer interaction, wellbeing screening and security. However, the design of such a solution remains challenging due to data scarcity, label imbalance, missing annotations and task interference in multi-task learning. In this paper, we propose STEP-FACE, a parameter-efficient framework that adapts a Vision-Language Model (VLM) to simultaneous gender, emotion and age recognition. It is based on Sequential TExtual-visual Prompting (STEP), the proposed parameter-efficient procedure which first learns task-specific textual prompts and then uses them to optimize visual prompts, injecting relevant cues for face analysis into the visual encoder. The inference is performed with the visual branch only, reducing memory and latency with respect to standard VLMs. The multi-task learning procedure designed for STEP-FACE further enforces task-representative batching, dealing with missing labels, balancing the task-specific losses and treating age estimation as a classification problem by adopting an ordinal-ranking formulation. The effectiveness of the proposed solution is confirmed by the experimental results obtained on widely used face analysis benchmarks: STEP-FACE achieves state-of-the-art or competitive performance, with 97.8% and 98.3% accuracy for gender recognition on FairFace and VGGFace2, 86.1% for emotion recognition on RAF-DB, and 63.8% and 63.5% for age classification on FairFace and UTKFace, while consistently improving over the considered baselines and remaining competitive with existing single-task and multi-task solutions in a parameter-efficient multi-task VLM setting.
Antonio Greco 0001, Camilla Spingola, Mario Vento
Knowl. Based Syst.1
2026 Parvelous: pedestrian attribute recognition using a vision encoder for multi-task learning on unbalanced sample distributions
abstract
Abstract Pedestrian attribute recognition (PAR) is a critical task for real-time video surveillance and person re-identification in-the-wild. While modern vision–language models pre-trained on billions of image–text pairs have recently achieved outstanding accuracy, their substantial latency and high memory requirements make them impractical for real-world deployments. To overcome these limitations, we present Parvelous , an efficient and versatile multi-task framework built on an optimized vision encoder pre-trained for image–text matching and specifically adapted to the PAR domain. Through targeted architectural refinements and a selective layer-wise fine-tuning strategy, our framework ensures both efficiency and strong task specialization. Modular task-specific branches equipped with channel-wise attention are tailored in depth to address the varying complexity of binary and multi-class attribute recognition tasks. Additionally, a multi-task loss based on asymmetric loss functions mitigates the severe class imbalance inherent in standard PAR datasets, fostering robust learning across diverse attributes. Extensive evaluations on the public MIVIA PAR KD benchmark demonstrate that Parvelous achieves a 0.957 accuracy rate, setting a new state of the art, while delivering up to an 80-fold inference speedup compared to competing vision language models. This combination of accuracy and computational efficiency positions Parvelous as a practical and deployable solution for real-world PAR applications.
Antonio Greco 0001, Andrea Vincenzo Ricciardi, Bruno Vento, Antonio Vitale
Neural Comput. Appl.1
2026 Guest editorial: special issue "from bench to the wild: recent advances in computer vision methods (WILD-VISION)"
George Azzopardi, Laura Fernández-Robles, Antonio Greco 0001, Bruno Vento
Pattern Recognit.3
2026 Animating Faces With Emotions Through a Generative Adversarial Network Preserving Identity
abstract
Artificially applying specific emotions to videos of people faces with a neutral expression, while preserving the identity of the subject is a challenging task. When parts of the face are synthetically moved to generate an emotion, it typically results in spatio-temporal artifacts in the generated videos, or inconsistency to preserve the identity of subjects. Existing methods that deploy spatio-temporal convolutions and de-convolutions to generate consecutive frames in a single step are not able to ensure proper motion dynamics, in the sense that the emotion may be not visible on the face or the facial features are distorted in the video. At the same time, approaches that generate motion and identity in two separate steps are not able to ensure the consistency of the subject identity after the generation of the emotion. In this paper we propose a novel method, Video Identity-Consistent Emotion GAN (VICEGAN), that improves the video generative capabilities of two-step methods. We decouple motion and content generation, thus ensuring the consistency of subject identity in the generated videos by using an encoder-decoder generator and a new identity-preserving loss in an adversarial framework. The proposed neural network architecture also guarantees the generation of proper motion of the target expressions, mitigating the presence of artifacts. We evaluated VICEGAN on the MUG dataset and compared it with a method based on a GAN, ImaGINator, demonstrating superior performance both quantitatively and qualitatively, and with a popular method based on a diffusion model, LFDM, showing a better capability to generate recognizable emotions.
Antonio Greco 0001, Nicola Strisciuglio, Mario Vento
IEEE Trans. Affect. Comput.1
2026 FAN-TAST-IC: Fast Alarm Notification with Task-Aware Spatio-Temporal Image Classification
abstract
Video-based alarm notification systems play a critical role in safety-critical applications such as fire detection and pedestrian monitoring, because they allow for prompt intervention in the event of dangerous situations. However, existing approaches often struggle to balance high sensitivity and specificity with real-time performance, particularly in complex scenes. To address these limitations, we propose a novel method enabling Fast Alarm Notification with Task-Aware Spatio-Temporal Image Classification (FAN-TAST-IC). It is an innovative and efficient framework that combines a lightweight task-specific object detector with a pre-trained Vision-Language Model (VLM) encoder and a binary classifier. Unlike end-to-end multimodal systems, FAN-TAST-IC leverages the VLM solely as a frozen visual feature extractor, preserving its rich semantic knowledge while ensuring computational efficiency. The object detector performs a rough but real-time filtering of the temporal frames and spatial positions where objects can be. This filtering is then refined in two distinct ways, depending on the time constraints of the specific application, either to discard temporally incoherent detections or to confirm those with high confidence. The selection of such candidates drastically reduces the input space for the VLM and classifier to dubious detections only. Thus, the latter is trained on detector-guided positive and negative samples, enabling precise alarm validation, improving specificity, and preserving sensitivity without requiring extensive labeled datasets or fine-tuning the VLM. Experiments on fire and pedestrian detection tasks demonstrate the effectiveness of our method since FAN-TAST-IC consistently outperforms all the compared approaches, achieving superior F-scores and precision on challenging benchmarks while maintaining real-time capabilities.
Diego Gragnaniello, Antonio Greco 0001, Carlo Sansone, Bruno Vento
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Multi-modal Human-Robot Collaboration in Production Lines Through Speech Commands and Gestures
Vincenzo Carletti, Antonio Greco 0001, Domenico Longobardi, Pierluigi Ritrovato, Alessia Saggese, Mario Vento
CAIP (2)2
2025 Leveraging Vision-Language Models for Improving Detection of Obstacles on Railway Tracks
Vincenzo Carletti, Antonio Greco 0001, Alessia Saggese, Camilla Spingola, Bruno Vento
CAIP (2)2
2025 An Extended Dataset and a Baseline for Pedestrian Attribute Recognition with Advanced Neural Networks
Antonio Greco 0001, Bruno Vento
CAIP (1)1
2025 Real-time joint recognition of weather and ground surface conditions by a multi-task deep network
abstract
Climate change and the occurrence of intense and unexpected weather events highlighted the need for real-time weather warning systems, especially in smart roads and isolated scenarios like rural areas. In this work, we propose to jointly recognize the weather and the ground surface conditions using existing video surveillance systems. Previous works separately tackled these two tasks even if they are correlated to each other. We propose a convolutional neural network with shared weights in the lower layers and two separate classification branches on top to exploit the correlation between the tasks and, at the same time, learn diverse high-level features for each task. Moreover, the network architecture implements attention mechanisms allowing the classification branches to focus on diverse image regions. The method is versatile and allows us to train the network on partially labeled data. The experimental analysis on real data demonstrate the effectiveness of the proposed method on both tasks, confirmed by the accuracy comparison with existing methods for the recognition of weather and ground surface conditions. The multi-task solution improves the inference speed (50 frames per second) and reduces the required memory (less than 1 GB) with respect to a system with two different single-task approaches; these results confirm that the proposed solution is ready for video surveillance applications to support smart cities. • A novel multi-task neural network for real-time recognition of weather and ground surface conditions is proposed. • A task-specific attention mechanism focuses the network’s receptive field on relevant image parts. • A masked asymmetric loss is proposed to deal with unbalanced and partially labeled datasets • We adaptively balance the backpropagated gradients of the two tasks.
Diego Gragnaniello, Antonio Greco 0001, Carlo Sansone, Bruno Vento
Eng. Appl. Artif. Intell.2
2025 FaiResGAN: Fair and robust blind face restoration with biometrics preservation
abstract
Modern computer vision technologies enable systems to detect, recognize, and analyze facial features, but challenges arise when images are noisy, blurred, or low quality. Blind face restoration, which aims to recover high-quality facial images without prior knowledge of degradation, addresses this issue. In this paper, we introduce Fair Restoration GAN (FaiResGAN), a novel Generative Adversarial Network (GAN) designed to balance face restoration with the preservation of soft biometrics (identity, ethnicity, age, and gender). Our model incorporates a pseudo-random batch composition algorithm to promote fairness and mitigate bias, alongside a realistic degradation model simulating corruptions typical in surveillance images. Experimental results show that FaiResGAN outperforms state-of-the-art blind face restoration methods, both quantitatively and qualitatively. A user study involving 40 participants showed that FaiResGAN-restored images were preferred by 70% of users. Additionally, tests on VGGFace2, UTKFace, and FairFace datasets demonstrate FaiResGAN’s superior performance in preserving soft biometric attributes and ensuring fair restoration across different genders and ethnicities.
George Azzopardi, Antonio Greco 0001, Mario Vento
Image Vis. Comput.2
2025 FOCUS: Improving fire detection on videos by scenario adaptation
abstract
Fire detection from video is effective for most video surveillance applications. The algorithm that processes the video acquired by cameras in real time has a twofold goal: detect as many fires as possible and keep the number of false alarms low. While existing approaches obtain the first goal, they often produce many false alarms due to their inability to account for the specific characteristics of diverse application environments. This paper introduces Fire Observation and Control Using Scenarios (FOCUS), a novel configurable fire detection method designed to bridge the gap between the literature methods and the application needs by exploiting scenario-specific knowledge. FOCUS leverages scalable and configurable modules for robust fire detection, incorporating three key steps: (1) fire detection, which identifies potential fire regions using visual cues; (2) fire candidate filtering through motion analysis, to eliminate false positives by analyzing the dynamic behavior of the identified fire candidates; (3) a vision-language model, which evaluates and confirms fire alarms by correlating visual evidence with contextual knowledge. By tailoring the configuration to the scenario and integrating the advanced filtering mechanisms according to the complexity of the environment, FOCUS improves performance in all the considered application scenarios. The analysis of the results shows that the proposed approach outperforms existing methods demonstrating higher resilience on real data, which enables its usage in real-world applications.
Diego Gragnaniello, Antonio Greco 0001, Carlo Sansone, Bruno Vento
Image Vis. Comput.2
2025 FLAME: fire detection in videos combining a deep neural network with a model-based motion analysis
abstract
Abstract Among the catastrophic natural events posing hazards to human lives and infrastructures, fire is the phenomenon causing more frequent damages. Thanks to the spread of smart cameras, video fire detection is gaining more attention as a solution to monitor wide outdoor areas where no specific sensors for smoke detection are available. However, state-of-the-art fire detectors assure a satisfactory Recall but exhibit a high false-positive rate that renders the application practically unusable. In this paper, we propose FLAME, an efficient and adaptive classification framework to address fire detection from videos. The framework integrates a state-of-the-art deep neural network for frame-wise object detection, in an automatic video analysis tool. The advantages of our approach are twofold. On the one side, we exploit advances in image detector technology to ensure a high Recall. On the other side, we design a model-based motion analysis that improves the system’s Precision by filtering out fire candidates occurring in the scene’s background or whose movements differ from those of the fire. The proposed technique, able to be executed in real-time on embedded systems, has proven to surpass the methods considered for comparison on a recent literature dataset representing several scenarios. The code and the dataset used for designing the system have been made publicly available by the authors at ( https://mivia.unisa.it/large-fire-dataset-with-negative-samples-lfdn/ ).
Diego Gragnaniello, Antonio Greco 0001, Carlo Sansone, Bruno Vento
Neural Comput. Appl.2
2025 Guest editorial: special issue on pedestrian attribute recognition and person re-identification
Antonio Greco 0001, Modesto Castrillón-Santana, Bruno Vento
Pattern Anal. Appl.1
2025 Video Fire Recognition Using Zero-Shot Vision-Language Models Guided by a Task-Aware Object Detector
abstract
Fire detection from images or videos has gained a growing interest in recent years due to the criticality of the application. Both reliable real-time detectors and efficient retrieval techniques, able to process large databases acquired by sensor networks, are needed. Even if the reliability of artificial vision methods improved in the last years, some issues are still open problems. In particular, literature methods often reveal a low generalization capability when employed in scenarios different from the training ones in terms of framing distance, surrounding environment, or weather conditions. This can be addressed by considering contextual information and, more specifically, using vision-language models capable of interpreting and describing the framed scene. In this work, we propose FIRE-TASTIC: Fire Recognition with Task-Aware Spatio-Temporal Image Captioning, a novel framework to use object detectors in conjunction with vision-language models for fire detection and information retrieval. The localization capability of the former makes it able to detect even tiny fire traces but expose the system to false alarms. These are strongly reduced by the impressive zero-shot generalization capability of the latter, which can recognize and describe fire-like objects without prior fine-tuning. We also present a variant of the FIRE-TASTIC framework based on visual question answering instead of image captioning, which allows one to customize the retrieved information with personalized questions. To integrate the high-level information provided by both neural networks, we propose a novel method to query the vision-language models using the temporal and spatial localization information provided by the object detector. The proposal can improve the retrieval performance, as evidenced by the experiments conducted on two recent fire detection datasets, showing the effectiveness and the generalization capabilities of FIRE-TASTIC, which surpasses the state of the art. Moreover, the vision-language model, which is unsuitable for video processing due to its high computational load, is executed only on suspicious frames, allowing for real-time processing. This makes FIRE-TASTIC suitable for both real-time processing and information retrieval on large datasets.
Diego Gragnaniello, Antonio Greco 0001, Carlo Sansone, Bruno Vento
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Fire and smoke detection from videos: A literature review under a novel taxonomy
abstract
The recent development of deep learning based fire detection techniques and the availability of smart cameras able to execute these algorithms on the edge paved the way for sophisticated and efficient video-based firefighting systems. However, the limited available data to train these algorithms cast shadows on their robustness and generalization capability. In this survey, we review 153 papers published in the literature and 17 publicly available fire detection datasets with the aim of identifying application scenarios that better describe real-world fire detection challenges. In the proposed taxonomy, these are characterized by two features: i) the fire size in the framed scene that depends on several parameters, foremost the distance from the fire but also the camera optic; ii) the background activity, due to the presence of moving objects that may mislead the detector. On this basis, we analyzed the existing methods under a common scheme according to this new taxonomy and matched the solutions with the needs of specific application scenarios. Similarly, for 9 interesting video datasets acquired from cameras, we labeled 536 videos according to the proposed taxonomy and shared these annotations with the community. The aim of this fire detection review is two-fold: on one hand, we classify the existing scientific works according to the real application scenarios, determining the features that are promising in specific operative conditions; on the other hand, we provide a detailed analysis and annotation of available datasets to promote the development of more reliable validation protocols and the collection of data from missing scenarios.
Diego Gragnaniello, Antonio Greco 0001, Carlo Sansone, Bruno Vento
Expert Syst. Appl.2
2024 Facial Soft-biometrics Obfuscation through Adversarial Attacks
abstract
Sharing facial pictures through online services, especially on social networks, has become a common habit for thousands of users. This practice hides a possible threat to privacy: the owners of such services, as well as malicious users, could automatically extract information from faces using modern and effective neural networks. In this article, we propose the harmless use of adversarial attacks, i.e., variations of images that are almost imperceptible to the human eye and that are typically generated with the malicious purpose to mislead Convolutional Neural Networks (CNNs). Such attacks have been instead adopted to (1) obfuscate soft biometrics (gender, age, ethnicity) but (2) without degrading the quality of the face images posted online. We achieve the above-mentioned two conflicting goals by modifying the implementations of four of the most popular adversarial attacks, namely FGSM, PGD, DeepFool, and C&W, in order to constrain the average amount of noise they generate on the image and the maximum perturbation they add on the single pixel. We demonstrate, in an experimental framework including three popular CNNs, namely VGG16, SENet, and MobileNetV3, that the considered obfuscation method, which requires at most 4 seconds for each image, is effective not only when we have a complete knowledge of the neural network that extracts the soft biometrics (white box attacks) but also when the adversarial attacks are generated in a more realistic black box scenario. Finally, we prove that an opponent can implement defense techniques to partially reduce the effect of the obfuscation, but substantially paying in terms of accuracy over clean images; this result, confirmed by the experiments carried out with three popular defense methods, namely adversarial training, denoising autoencoder, and Kullback-Leibler autoencoder, shows that it is not convenient for the opponent to defend himself and that the proposed approach is robust to defenses.
Vincenzo Carletti, Pasquale Foggia, Antonio Greco 0001, Alessia Saggese, Mario Vento
ACM Trans. Multim. Comput. Commun. Appl.3
2023 PAR Contest 2023: Pedestrian Attributes Recognition with Multi-task Learning
Antonio Greco 0001, Bruno Vento
CAIP (1)1
2023 Multi-task learning on the edge for effective gender, age, ethnicity and emotion recognition
Pasquale Foggia, Antonio Greco 0001, Alessia Saggese, Mario Vento
Eng. Appl. Artif. Intell.2
2023 A Social Robot Architecture for Personalized Real-Time Human-Robot Interaction
abstract
In the age of the Internet of Things (IoT), the combination of robotics and artificial intelligence has paved the way for the development of social robots able to undertake realistic conversations with humans, making them the perfect human interface in applications like for instance, robotic house assistants and hotel concierges. Despite the several solutions developed in recent years, the definition of social robot requirements and the software modules needed to meet such requirements has not yet been formalized. In this article, we define the requirements of a social robot and propose a software architecture that includes all the necessary modules to meet them. The proposed architecture, implemented using robot operating system (ROS) nodes, is hardware-independent, enabling its reuse across different robotic platforms with interchangeable types of sensors and actuators. We deployed a social robot based on this architecture to interact with attendees of a real exhibition context and validated the reliability of our proposed solution through a survey of 161 users. The results of the user study indicated a high-quality user experience with the social robot, with scores ranging from 4 to 5 (being 5 the maximum score). Sharing these design choices and evaluation results could significantly benefit the development of future social robotics applications in the context of IoT.
Pasquale Foggia, Antonio Greco 0001, Antonio Roberto, Alessia Saggese, Mario Vento
IEEE Internet Things J.2
2023 Benchmarking deep networks for facial emotion recognition in the wild
abstract
Abstract Emotion recognition from face images is a challenging task that gained interest in recent years for its applications to business intelligence and social robotics. Researchers in computer vision and affective computing focused on optimizing the classification error on benchmark data sets, which do not extensively cover possible variations that face images may undergo in real environments. Following on investigations carried out in the field of object recognition, we evaluated the robustness of existing methods for emotion recognition when their input is subjected to corruptions caused by factors present in real-world scenarios. We constructed two data sets on top of the RAF-DB test set, named RAF-DB-C and RAF-DB-P, that contain images modified with 18 types of corruption and 10 of perturbation. We benchmarked existing networks (VGG, DenseNet, SENet and Xception) trained on the original images of RAF-DB and compared them with ARM, the current state-of-the-art method on the RAF-DB test set. We carried out an extensive study on the effects that modifications to the training data or network architecture have on the classification of corrupted and perturbed data. We observed a drop of recognition performance of ARM, with the classification error raising up to 200% of that achieved on the original RAF-DB test set. We demonstrate that the use of the AutoAugment data augmentation and an anti-aliasing filter within down-sampling layers provide existing networks with increased robustness to out-of-distribution variations, substantially reducing the error on corrupted inputs and outperforming ARM. We provide insights about the resilience of existing emotion recognition methods and an estimation of their performance in real scenarios. The processing time required by the modifications we investigated (35 ms in the worst case) supports their suitability for application in real-world scenarios. The RAF-DB-C and RAF-DB-P test sets, trained models and evaluation framework are available at https://github.com/MiviaLab/emotion-robustness .
Antonio Greco 0001, Nicola Strisciuglio, Mario Vento, Vincenzo Vigilante
Multim. Tools Appl.1
2023 Degramnet: effective audio analysis based on a fully learnable time-frequency representation
abstract
Abstract Current state-of-the-art audio analysis algorithms based on deep learning rely on hand-crafted Spectrogram-like audio representations, that are more compact than descriptors obtained from the raw waveform; the latter are, in turn, far from achieving good generalization capabilities when few data are available for the training. However, Spectrogram-like representations have two main limitations: (1) The parameters of the filters are defined a priori, regardless of the specific audio analysis task; (2) such representations do not perform any denoising operation on the audio signal, neither in the time domain nor in the frequency domain. To overcome these limitations, we propose a new general-purpose convolutional architecture for audio analysis tasks that we call DEGramNet, which is trained with audio samples described with a novel, compact and learnable time–frequency representation that we call DEGram. The proposed representation is fully trainable: Indeed, it is able to learn the frequencies of interest for the specific audio analysis task; in addition, it performs denoising through a custom time–frequency attention module, which amplifies the frequency and time components in which the sound is actually located. It implies that the proposed representation can be easily adapted to the specific problem at hands, for instance giving more importance to the voice frequencies when the network needs to be used for speaker recognition. DEGramNet achieved state-of-the-art performance on the VGGSound dataset (for Sound Event Classification) and comparable accuracy with a complex and special-purpose approach based on network architecture search over the VoxCeleb dataset (for Speaker Identification). Moreover, we demonstrate that DEGram allows to achieve high accuracy with lightweight neural networks that can be used in real-time on embedded systems, making the solution suitable for Cognitive Robotics applications.
Pasquale Foggia, Antonio Greco 0001, Antonio Roberto, Alessia Saggese, Mario Vento
Neural Comput. Appl.2
2023 A deep learning based system for handwashing procedure evaluation
abstract
Abstract Hand washing preparation can be considered as one of the main strategies for reducing the risk of surgical site contamination and thus the infections risks. Within this context, in this paper we propose an embedded system able to automatically analyze, in real-time, the sequence of images acquired by a depth camera to evaluate the quality of the handwashing procedure. In particular, the designed system runs on an NVIDIA Jetson Nano $$^{{\mathrm{TM}}}$$ TM computing platform. We adopt a convolutional neural network, followed by a majority voting scheme, to classify the movement of the worker according to one of the ten gestures defined by the World Health Organization. To test the proposed system, we collect a dataset built by 74 different video sequences. The results achieved on this dataset confirm the effectiveness of the proposed approach.
Antonio Greco 0001, Gennaro Percannella, Pierluigi Ritrovato, Alessia Saggese, Mario Vento
Neural Comput. Appl.1
2022 Benchmarking deep neural networks for gesture recognition on embedded devices
abstract
The gesture is one of the most used forms of communication between humans; in recent years, given the new trend of factories to be adapted to Industry 4.0 paradigm, the scientific community has shown a growing interest towards the design of Gesture Recognition (GR) algorithms for Human-Robot Interaction (HRI) applications. Within this context, the GR algorithm needs to work in real time and over embedded platforms, with limited resources. Anyway, when looking at the available scientific literature, the aim of the different proposed neural networks (i.e. 2D and 3D) and of the different modalities used for feeding the network (i.e. RGB, RGB-D, optical flow) is typically the optimization of the accuracy, without strongly paying attention to the feasibility over low power hardware devices. Anyway, the analysis related to the trade-off between accuracy and computational burden (for both networks and modalities) becomes important so as to allow GR algorithms to work in industrial robotics applications. In this paper, we perform a wide benchmarking focusing not only on the accuracy but also on the computational burden, involving two different architectures (2D and 3D), with two different backbones (MobileNet, ResNeXt) and four types of input modalities (RGB, Depth, Optical Flow, Motion History Image) and their combinations.
Stefano Bini, Antonio Greco 0001, Alessia Saggese, Mario Vento
RO-MAN2
2022 Effective training of convolutional neural networks for age estimation based on knowledge distillation
abstract
Abstract Age estimation from face images can be profitably employed in several applications, ranging from digital signage to social robotics, from business intelligence to access control. Only in recent years, the advent of deep learning allowed for the design of extremely accurate methods based on convolutional neural networks (CNNs) that achieve a remarkable performance in various face analysis tasks. However, these networks are not always applicable in real scenarios, due to both time and resource constraints that the most accurate approaches often do not meet. Moreover, in case of age estimation, there is the lack of a large and reliably annotated dataset for training deep neural networks. Within this context, we propose in this paper an effective training procedure of CNNs for age estimation based on knowledge distillation, able to allow smaller and simpler “student” models to be trained to match the predictions of a larger “teacher” model. We experimentally show that such student models are able to almost reach the performance of the teacher, obtaining high accuracy over the LFW+, LAP 2016 and Adience datasets, but being up to 15 times faster. Furthermore, we evaluate the performance of the student models in the presence of image corruptions, and we demonstrate that some of them are even more resilient to these corruptions than the teacher model.
Antonio Greco 0001, Alessia Saggese, Mario Vento, Vincenzo Vigilante
Neural Comput. Appl.1
2022 Vehicles Detection for Smart Roads Applications on Board of Smart Cameras: A Comparative Analysis
abstract
Video analytics can be profitably adopted in smart roads environments to automatically detect abnormal situations. Within this context, vehicle detection is the first and foremost stage, and its accuracy is crucial, since any detection error will affect the performance of any subsequent step. Furthermore, in smart road environments it is often preferred to perform the video analysis directly on board of smart surveillance cameras, in order to reduce bandwidth usage and eliminate the cost of setup and maintenance of powerful processing servers; on the flip side, processing on board of smart cameras implies the detection algorithm to be fast and slim, since the resources available on this kind of embedded device are limited. In the era of deep learning, it seems that the questionwhat is the best method for vehicle detection?may have a trivial answer, since this class of methods includes some very accurate ones. Anyway, according to the above consideration, the best suited method for this application is not necessarily the most accurate one, but for sure the most accurate one running on the available hardware at a given resolution and frame rate. Starting from the above considerations, in this paper we perform an analysis of the methods available in the literature for vehicle detection, by comparing them in terms of accuracy and computational burden, with the aim to answer the following question:what is the best method for vehicles detection when working with smart cameras?
Antonio Greco 0001, Alessia Saggese, Mario Vento, Vincenzo Vigilante
IEEE Trans. Intell. Transp. Syst.1
2021 Guess the Age 2021: Age Estimation from Facial Images with Deep Convolutional Neural Networks
Antonio Greco 0001
CAIP (2)1
2021 DENet: a deep architecture for audio surveillance applications
Antonio Greco 0001, Antonio Roberto, Alessia Saggese, Mario Vento
Neural Comput. Appl.1
2020 Which are the factors affecting the performance of audio surveillance systems?
abstract
Sound event recognition systems are rapidly becoming part of our life, since they can be profitably used in several vertical markets, ranging from audio security applications to scene classification and multi-modal analysis in social robotics. In the last years, a not negligible part of the scientific community started to apply Convolutional Neural Networks (CNNs) to image-based representations of the audio stream, due to their successful adoption in almost all the computer vision tasks. In this paper, we carry out a detailed benchmark of various widely used CNN architectures and visual representations on a popular dataset, namely the MIVIA Audio Events database. Our analysis is aimed at understanding how these factors affect the sound event recognition performance with a particular focus on the false positive rate, very relevant in audio surveillance solutions. In fact, although most of the proposed solutions achieve a high recognition rate, the capability of distinguishing the events-of-interest from the background is often not yet sufficient for real systems, and prevent its usage in real applications. Our comprehensive experimental analysis investigates this aspect and allows to identify useful design guidelines for increasing the specificity of sound event recognition systems.
Antonio Greco 0001, Antonio Roberto, Alessia Saggese, Mario Vento
ICPR1
2020 Benchmarking deep network architectures for ethnicity recognition using a new large face dataset
abstract
Abstract Although in recent years we have witnessed an explosion of the scientific research in the recognition of facial soft biometrics such as gender, age and expression with deep neural networks, the recognition of ethnicity has not received the same attention from the scientific community. The growth of this field is hindered by two related factors: on the one hand, the absence of a dataset sufficiently large and representative does not allow an effective training of convolutional neural networks for the recognition of ethnicity; on the other hand, the collection of new ethnicity datasets is far from simple and must be carried out manually by humans trained to recognize the basic ethnicity groups using the somatic facial features. To fill this gap in the facial soft biometrics analysis, we propose the VGGFace2 Mivia Ethnicity Recognition (VMER) dataset, composed by more than 3,000,000 face images annotated with 4 ethnicity categories, namely African American, East Asian, Caucasian Latin and Asian Indian. The final annotations are obtained with a protocol which requires the opinion of three people belonging to different ethnicities, in order to avoid the bias introduced by the well-known other race effect. In addition, we carry out a comprehensive performance analysis of popular deep network architectures, namely VGG-16, VGG-Face, ResNet-50 and MobileNet v2. Finally, we perform a cross-dataset evaluation to demonstrate that the deep network architectures trained with VMER generalize on different test sets better than the same models trained on the largest ethnicity dataset available so far. The ethnicity labels of the VMER dataset and the code used for the experiments are available upon request at https://mivia.unisa.it .
Antonio Greco 0001, Gennaro Percannella, Mario Vento, Vincenzo Vigilante
Mach. Vis. Appl.1
2020 Age from Faces in the Deep Learning Revolution
abstract
Face analysis includes a variety of specific problems as face detection, person identification, gender and ethnicity recognition, just to name the most common ones; in the last two decades, significant research efforts have been devoted to the challenging task of age estimation from faces, as witnessed by the high number of published papers. The explosion of the deep learning paradigm, that is determining a spectacular increasing of the performance, is in the public eye; consequently, the number of approaches based on deep learning is impressively growing and this also happened for age estimation. The exciting results obtained have been recently surveyed on almost all the specific face analysis problems; the only exception stands for age estimation, whose last survey dates back to 2010 and does not include any deep learning based approach to the problem. This paper provides an analysis of the deep methods proposed in the last six years; these are analysed from different points of view: the network architecture together with the learning procedure, the used datasets, data preprocessing and augmentation, and the exploitation of additional data coming from gender, race and face expression. The review is completed by discussing the results obtained on public datasets, so as the impact of different aspects on system performance, together with still open issues.
Vincenzo Carletti, Antonio Greco 0001, Gennaro Percannella, Mario Vento
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Comparing performance of graph matching algorithms on huge graphs
Vincenzo Carletti, Pasquale Foggia, Antonio Greco 0001, Alessia Saggese, Mario Vento
Pattern Recognit. Lett.3
2020 AReN: A Deep Learning Approach for Sound Event Recognition Using a Brain Inspired Representation
abstract
Audio surveillance is gaining in the last years wide interest. This is due to the large number of situations in which this kind of systems can be used, either alone or combined with video-based algorithms. In this paper we propose a deep learning method to automatically recognize events of interest in the context of audio surveillance (namely screams, broken glasses and gun shots). The audio stream is represented by a gammatonegram image. We propose a 21-layer CNN to which we feed sections of the gammatonegram representation. At the output of this CNN there are units that correspond to the classes. We trained the CNN, called AReN, by taking advantage of a problem-driven data augmentation, which extends the training dataset with gammatonegram images extracted by sounds acquired with different signal to noise ratios. We experimented it with three datasets freely available, namely SESA, MIVIA Audio Events and MIVIA Road Events and we achieved 91.43%, 99.62% and 100% recognition rate, respectively. We compared our method with other state of the art methodologies based both on traditional machine learning methodologies and deep learning. The comparison confirms the effectiveness of the proposed approach, which outperforms the existing methods in terms of recognition rate. We experimentally prove that the proposed network is resilient to the noise, has the capability to significantly reduce the false positive rate and is able to generalize in different scenarios. Furthermore, AReN is able to process 5 audio frames per second on a standard CPU and, consequently, it is suitable for real audio surveillance applications.
Antonio Greco 0001, Nicolai Petkov, Alessia Saggese, Mario Vento
IEEE Trans. Inf. Forensics Secur.1
2019 Emotion analysis from faces for social robotics
abstract
A social robot is able to perceive the information about the environment (both in terms of persons and objects populating the scene), to reason about the acquired information and to interact with the human in a proper way. Among the information required for the interaction with a human, the capability of analysing the emotion is surely among the most important ones. Another relevant feature of social robots is the possibility to interact with the human in real time, without any latency that could introduce a delay in the talk between the human and the robot. It means that all the processing of the information (for instance the analysis of the sequence of images for the detection of the persons and the consequent emotion analysis) needs to be performed directly on board of the robot, without any possibility to use high performance servers (for instance services on the cloud) but only small devices that can be installed directly on board of the robotic platform. In this paper we propose MIVIAbot, a robotic platform based on the Pepper humanoid robot and equipped with a small, low cost and low energy consumption embedded device installed directly on board of Pepper, without any requirements for a Wi-Fi connection that could introduce any latency for the transmission of the data. Furthermore, we also propose a method for the analysis of the emotion of the persons by face analysis, able to run on the considered hardware platform in real time without paying in terms of accuracy. The experimentation conducted over various widely adopted datasets of images and videos for emotion analysis confirms both the efficiency and the effectiveness of the proposed approach.
Antonio Greco 0001, Antonio Roberto, Alessia Saggese, Mario Vento, Vincenzo Vigilante
SMC1
2019 SoReNet: a novel deep network for audio surveillance applications
abstract
In the era of third generation surveillance systems, it becomes more and more useful to have available a solution able to automatically detect abnormal events. The interest for audio analysis is thus growing in the last years, due to the large amount of situations where a microphone and an audio surveillance system can be profitably used by the human operator in charge of control. In this paper, we propose a method for automatically analyzing the audio stream for surveillance purposes: it is able to detect the presence of abnormal events such as screams, gun shots and broken glasses. Instead than processing directly raw data (the audio signal), the stream is represented by means of an image, namely the spectrogram, a time-frequency representation of the audio stream. In this way, we formulate the problem of audio analysis as a problem of image classification. Thus, we propose to use a Convolutional Neural Network with the following two main properties: inspired by VGG network, we employed very small kernels in convolutional layers; furthermore, we adopted a pyramidal structure in fully connected layers. These choices allow to have good generalization capabilities of the network even in presence of a not so wide dataset. The performance, computed over a standard dataset already used for benchmarking purposes in the field of audio surveillance, confirms the effectiveness of the proposed approach.
Antonio Greco 0001, Alessia Saggese, Mario Vento, Vincenzo Vigilante
SMC1
2019 VF3-Light: A lightweight subgraph isomorphism algorithm and its experimental evaluation
Vincenzo Carletti, Pasquale Foggia, Antonio Greco 0001, Mario Vento, Vincenzo Vigilante
Pattern Recognit. Lett.3
2018 Gender recognition from face images using trainable shape and color features
abstract
Gender recognition from face images is an important application and it is still an open computer vision problem, even though it is something trivial from the human visual system. Variations in pose, lighting, and expression are few of the problems that make such an application challenging for a computer system. Neurophysiological studies demonstrate that the human brain is able to distinguish men and women also in absence of external cues, by analyzing the shape of specific parts of the face. In this paper, we describe an automatic procedure that combines trainable shape and color features for gender classification. In particular the proposed method fuses edge-based and color-blob-based features by means of trainable COSFIRE filters. The former types of feature are able to extract information about the shape of a face whereas the latter extract information about shades of colors in different parts of the face. We use these two sets of features to create a stacked classification SVM model and demonstrate its effectiveness on the GENDER-COLOR-FERET dataset, where we achieve an accuracy of 96.4%.
George Azzopardi, Pasquale Foggia, Antonio Greco 0001, Alessia Saggese, Mario Vento
ICPR3
2017 Fast gender recognition in videos using a novel descriptor based on the gradient magnitudes of facial landmarks
abstract
The growing interest in recent years for gender recognition from face images is mainly attributable to the wide range of possible applications that can be used for commercial and marketing purposes. It is desirable that such algorithms process high resolution video frames acquired by using surveillance cameras in real-time. To the best of our knowledge, however, there are no studies which analyze the computational impact of the methods and the difficulties related to the processing of faces extracted from videos captured in the wild. We propose a novel face descriptor based on the gradient magnitudes of facial landmarks, which are points automatically extracted from the face contour, eyes, eyebrows, nose, mouth and chin. We evaluate the effectiveness and efficiency of the proposed approach on two new datasets, which we made available online and that consist of color face images and color video sequences acquired in real scenarios. The proposed approach is more efficient and effective than three commercial libraries.
George Azzopardi, Antonio Greco 0001, Alessia Saggese, Mario Vento
AVSS2
2016 Gender recognition from face images with trainable COSFIRE filters
abstract
Gender recognition from face images is an important application in the fields of security, retail advertising and marketing. We propose a novel descriptor based on COSFIRE filters for gender recognition. A COSFIRE filter is trainable, in that its selectivity is determined in an automatic configuration process that analyses a given prototype pattern of interest. We demonstrate the effectiveness of the proposed approach on a new dataset called GENDER-FERET with 474 training and 472 test samples and achieve an accuracy rate of 93.7%. It also outperforms an approach that relies on handcrafted features and an ensemble of classifiers. Furthermore, we perform another experiment by using the images of the Labeled Faces in the Wild (LFW) dataset to train our classifier and the test images of the GENDER-FERET dataset for evaluation. This experiment demonstrates the generalization ability of the proposed approach and it also outperforms two commercial libraries, namely Face++ and Luxand.
George Azzopardi, Antonio Greco 0001, Mario Vento
AVSS2
2016 Counting people by RGB or depth overhead cameras
Luca Del Pizzo, Pasquale Foggia, Antonio Greco 0001, Gennaro Percannella, Mario Vento
Pattern Recognit. Lett.3
2015 Automatic detection of long term parked cars
abstract
The detection of illegal roadside parking is becoming more and more interesting in the field of intelligent transportation systems, since it may cause traffic congestion or accidents. In this paper we propose a method able to analyze videos acquired by traditional surveillance cameras and to automatically detect the vehicles stopped in a forbidden area. Two main contributions have been introduced: first, spatio temporal information related to the stopped vehicles are encoded by a heat map; second, the background is not updated by evaluating the movement of the vehicle in a single time instant, but instead the whole movement of the vehicles, encoded into the heat map, is taken into account. Two widely adopted datasets, namely the iLids and the PETS 2000, have been used to experimentally evaluate the proposed approach and the results achieved, compared with state of the art methodologies, confirm its effectiveness.
Vincenzo Carletti, Pasquale Foggia, Antonio Greco 0001, Alessia Saggese, Mario Vento
AVSS3