Mario Vento

dblp:46/113 · DBLP profile ↗
← Back
158ranked-venue papers
1as first author
33since 2021 · last 2026
0000-0002-2948-741XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 96 · 1 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 58 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 6 since 2021Human-computer interaction and ubiquitous computing · 13 · 5 since 2021Databases, data management, data science and information retrieval · 5Computer networks · 3 · 3 since 2021Security and privacy · 3 · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Improving Linear Algebra Problem Solving in Higher Education with Retrieval-Augmented Large Language Models
Giovannina Albano, Nicola Capuano, Giovanni Casella, Mario Vento
CSEDU (1)4
2026 On the Detectability of Active Gradient Inversion Attacks in Federated Learning
abstract
One of the key advantages of Federated Learning (FL) is its ability to collaboratively train a Machine Learning (ML) model while keeping clients' data on-site. However, this can create a false sense of security. Despite not sharing private data increases the overall privacy, prior studies have shown that gradients exchanged during the FL training remain vulnerable to Gradient Inversion Attacks (GIAs). These attacks allow reconstructing the clients' local data, breaking the privacy promise of FL. GIAs can be launched by either a passive or an active server. In the latter case, a malicious server manipulates the global model to facilitate data reconstruction. While effective, earlier attacks falling under this category have been demonstrated to be detectable by clients, limiting their real-world applicability. Recently, novel active GIAs have emerged, claiming to be far stealthier than previous approaches. This work provides the first comprehensive analysis of these claims, investigating four state-of-the-art GIAs. We propose novel lightweight client-side detection techniques, based on statistically improbable weight structures and anomalous loss and gradient dynamics. Extensive evaluation across several configurations demonstrates that our methods enable clients to effectively detect active GIAs without any modifications to the FL training protocol.
Vincenzo Carletti, Pasquale Foggia, Carlo Mazzocca, Giuseppe Parrella, Mario Vento
SP5
2026 ShapeBlend: Boosting out-of-distribution robustness in image classification via shape-based blending augmentation
abstract
Deep Neural Networks (DNNs) often struggle to generalize beyond their training distributions, making them vulnerable to domain shifts. To enhance robustness, various approaches have been developed, particularly focusing on data-centric methods. Modifying the training data can increase diversity and improve generalization, while also introducing bias that positively guides the model’s decision-making. Prior research suggests that DNNs tend to overemphasize texture-based patterns, at the expense of more robust shape-based representations. We introduce ShapeBlend, a novel data augmentation technique that emphasizes image contours, and hence shape features. It blends a contour map from the push-pull CORF operator with the original image at varying strengths. ShapeBlend consistently outperforms state-of-the-art methods across major Out-of-Distribution (OOD) benchmarks (ImageNet-A, ImageNet-R, ImageNet-C, and ImageNet- C ¯ ), setting new records in robustness. Moreover, ShapeBlend’s versatility allows its application during inference. To fully leverage ShapeBlend, we propose Shape-Enhanced Voting (SEV), an inference strategy that aggregates predictions from multiple ShapeBlend-processed images. The combination of ShapeBlend and SEV further enhances domain robustness, with performance gains varying based on the chosen configuration. • We introduce ShapeBlend: a shape-based augmentation for OOD robustness. • ShapeBlend improves performance across multiple robustness benchmarks. • Inference performance is boosted with Shape-enhanced Voting (SEV). • ShapeBlend is compatible with existing pipelines like AugMix and DeepAugment. • A simple, explainable method with strong theoretical motivation.
George Azzopardi, Sabatino Esposito, Antonio Greco 0001, Mario Vento
Comput. Vis. Image Underst.4
2026 Deep learning based empty shelf detection based on autonomous mobile robot
abstract
The issue of out-of-stock (OOS) represents a substantial challenge for retailers, often resulting in significant sales losses. To address this problem, this paper introduces an autonomous mobile robotic platform built on the Robot Operating System (ROS) framework, designed to accelerate the restocking process in supermarkets. The platform autonomously detects empty shelves and notifies human operators, streamlining inventory management. Equipped with advanced navigation capabilities, the proposed system employs a deep learning-based, two-stage architecture that identifies shelving areas and subsequently detects empty shelves. To validate the performance of the proposed two-stage artificial vision algorithm, two datasets were used: the first comprises approximately 2000 images (900 of them collected by our team from three different supermarkets), while the second dataset consists of around 5600 manually annotated images extracted from videos recorded in a supermarket by the robotic platform itself. Additionally, in order to validate the entire robotic system, an extensive experimental evaluation was conducted in a supermarket during regular business hours. The results demonstrate that the proposed platform substantially outperforms human operators, identifying OOS items eight times faster than traditional human operator based methods. This advancement provides valuable assistance to supermarket staff, significantly enhancing operational efficiency.
Giuseppe De Simone, Alessia Saggese, Pasquale Foggia, Mario Vento
Comput. Vis. Image Underst.4
2026 DRIVE: Distributed Robotic Intelligence for Vision-based Exploration for retail shelf monitoring
abstract
Out-of-stock (OOS) detection in retail environments is essential to ensure efficient inventory management and maintain high levels of customer satisfaction. Within this context, mobile robotic platforms, equipped with a camera and an empty shelf object detector, have emerged as a promising solution. However, detector-based approaches suffer from a fundamental trade-off between false positives and missed detections, with limited generalization capabilities due to small available training datasets, and high false positive rate in cluttered retail scenes. To overcome these challenges, we propose DRIVE (Distributed Robotic Intelligence for Vision-based Exploration), a novel distributed architecture that combines a lightweight on-board object detector with a cloud-based transformer-powered semantic validation stage. This two-tier design mitigates the precision–recall trade-off of traditional detectors, reducing false positives without sacrificing recall, while ensuring real-time feasibility on resource-constrained platforms. Furthermore, to enable robust domain adaptation under low-data regimes without catastrophic forgetting, we fine-tune the vision transformer backbone using Parameter-Efficient Fine-Tuning (PEFT) via Low-Rank Adaptation (LoRA), thus injecting less than 1% additional parameters while preserving pretrained knowledge. Extensive experiments in a real supermarket environments demonstrate that DRIVE achieves impressive robustness and accuracy compared to state-of-the-art detection-based solutions, paving the way for scalable, autonomous OOS detection in dynamic retail scenarios.
Alessia Saggese, Mario Vento
J. Syst. Archit.2
2026 STEP-FACE: Sequential TExtual-Visual Prompting for multi-task face analysis
abstract
Jointly analyzing stable facial attributes (gender, age) and transient affect (emotion) is a useful functionality for applications such as adaptive human–computer interaction, wellbeing screening and security. However, the design of such a solution remains challenging due to data scarcity, label imbalance, missing annotations and task interference in multi-task learning. In this paper, we propose STEP-FACE, a parameter-efficient framework that adapts a Vision-Language Model (VLM) to simultaneous gender, emotion and age recognition. It is based on Sequential TExtual-visual Prompting (STEP), the proposed parameter-efficient procedure which first learns task-specific textual prompts and then uses them to optimize visual prompts, injecting relevant cues for face analysis into the visual encoder. The inference is performed with the visual branch only, reducing memory and latency with respect to standard VLMs. The multi-task learning procedure designed for STEP-FACE further enforces task-representative batching, dealing with missing labels, balancing the task-specific losses and treating age estimation as a classification problem by adopting an ordinal-ranking formulation. The effectiveness of the proposed solution is confirmed by the experimental results obtained on widely used face analysis benchmarks: STEP-FACE achieves state-of-the-art or competitive performance, with 97.8% and 98.3% accuracy for gender recognition on FairFace and VGGFace2, 86.1% for emotion recognition on RAF-DB, and 63.8% and 63.5% for age classification on FairFace and UTKFace, while consistently improving over the considered baselines and remaining competitive with existing single-task and multi-task solutions in a parameter-efficient multi-task VLM setting.
Antonio Greco 0001, Camilla Spingola, Mario Vento
Knowl. Based Syst.3
2026 Animating Faces With Emotions Through a Generative Adversarial Network Preserving Identity
abstract
Artificially applying specific emotions to videos of people faces with a neutral expression, while preserving the identity of the subject is a challenging task. When parts of the face are synthetically moved to generate an emotion, it typically results in spatio-temporal artifacts in the generated videos, or inconsistency to preserve the identity of subjects. Existing methods that deploy spatio-temporal convolutions and de-convolutions to generate consecutive frames in a single step are not able to ensure proper motion dynamics, in the sense that the emotion may be not visible on the face or the facial features are distorted in the video. At the same time, approaches that generate motion and identity in two separate steps are not able to ensure the consistency of the subject identity after the generation of the emotion. In this paper we propose a novel method, Video Identity-Consistent Emotion GAN (VICEGAN), that improves the video generative capabilities of two-step methods. We decouple motion and content generation, thus ensuring the consistency of subject identity in the generated videos by using an encoder-decoder generator and a new identity-preserving loss in an adversarial framework. The proposed neural network architecture also guarantees the generation of proper motion of the target expressions, mitigating the presence of artifacts. We evaluated VICEGAN on the MUG dataset and compared it with a method based on a GAN, ImaGINator, demonstrating superior performance both quantitatively and qualitatively, and with a popular method based on a diffusion model, LFDM, showing a better capability to generate recognizable emotions.
Antonio Greco 0001, Nicola Strisciuglio, Mario Vento
IEEE Trans. Affect. Comput.3
2025 Leveraging Open-Source LLMs for Zero-Shot Vulnerability Detection: A Comparative Analysis
Nicola Capuano, Vincenzo Carletti, Pasquale Foggia, Giuseppe Parrella, Mario Vento
AINA (5)5
2025 Multi-modal Human-Robot Collaboration in Production Lines Through Speech Commands and Gestures
Vincenzo Carletti, Antonio Greco 0001, Domenico Longobardi, Pierluigi Ritrovato, Alessia Saggese, Mario Vento
CAIP (2)6
2025 Multimodal Audio-Visual Emotion Recognition for Social Robotics
Giuseppe De Simone, Luca Greco 0001, Alessia Saggese, Mario Vento
CAIP (2)4
2025 SoK: Gradient Inversion Attacks in Federated Learning
Vincenzo Carletti, Pasquale Foggia, Carlo Mazzocca, Giuseppe Parrella, Mario Vento
USENIX Security Symposium5
2025 A Multi-task learning U-Net model for end-to-end HEp-2 cell image analysis
abstract
Antinuclear Antibody ( ANA ) testing is pivotal to help diagnose patients with a suspected autoimmune disease. The Indirect Immunofluorescence ( IIF ) microscopy performed with human epithelial type 2 (HEp-2) cells as the substrate is the reference method for ANA screening. It allows for the detection of antibodies binding to specific intracellular targets, resulting in various staining patterns that should be identified for diagnosis purposes. In recent years, there has been an increasing interest in devising deep learning methods for automated cell segmentation and classification of staining patterns, as well as for other tasks related to this diagnostic technique (such as intensity classification). However, little attention has been devoted to architectures aimed at simultaneously managing multiple interrelated tasks, via a shared representation. In this paper, we propose a deep neural network model that extends U-Net in a Multi-Task Learning (MTL) fashion, thus offering an end-to-end approach to tackle three fundamental tasks of the diagnostic procedure, i.e., HEp-2 cell specimen intensity classification, specimen segmentation, and pattern classification. The experiments were conducted on one of the largest publicly available datasets of HEp-2 images. The results showed that the proposed approach significantly outperformed the competing state-of-the-art methods for all the considered tasks. • IIF on Hep-2 cells is crucial for diagnosing autoimmune diseases. • It involves one segmentation task and two classification tasks. • A Multi-Task Learning approach could help but hasn’t been explored till now. • We propose a Multi-Task, end-to-end U-Net architecture for performing all three tasks. • We achieve significant improvements over methods designed for individual tasks.
Gennaro Percannella, Umberto Petruzzello, Francesco Tortorella, Mario Vento
Artif. Intell. Medicine4
2025 FaiResGAN: Fair and robust blind face restoration with biometrics preservation
abstract
Modern computer vision technologies enable systems to detect, recognize, and analyze facial features, but challenges arise when images are noisy, blurred, or low quality. Blind face restoration, which aims to recover high-quality facial images without prior knowledge of degradation, addresses this issue. In this paper, we introduce Fair Restoration GAN (FaiResGAN), a novel Generative Adversarial Network (GAN) designed to balance face restoration with the preservation of soft biometrics (identity, ethnicity, age, and gender). Our model incorporates a pseudo-random batch composition algorithm to promote fairness and mitigate bias, alongside a realistic degradation model simulating corruptions typical in surveillance images. Experimental results show that FaiResGAN outperforms state-of-the-art blind face restoration methods, both quantitatively and qualitatively. A user study involving 40 participants showed that FaiResGAN-restored images were preferred by 70% of users. Additionally, tests on VGGFace2, UTKFace, and FairFace datasets demonstrate FaiResGAN’s superior performance in preserving soft biometric attributes and ensuring fair restoration across different genders and ethnicities.
George Azzopardi, Antonio Greco 0001, Mario Vento
Image Vis. Comput.3
2025 Detecting malicious IoT network communication through Graph Neural Networks in real-world conditions
abstract
Internet of Things (IoT) devices are increasingly permeating homes, industries, and many other environments. The need for robust security measures in IoT networks has never been more critical, since they are becoming the preferred target for cyberattacks. In this paper, we address the challenge of detecting abnormal communication patterns in IoT networks using Graph Neural Networks (GNNs). To this end, we have conducted a comprehensive and fair comparison of machine learning approaches and GNNs, for both static and dynamic graphs, across three recent datasets, IoT23, IoTID20, IoT Traces, that contain recordings of network communications among IoT devices in real environments. Differently from the state-of-the-art, we face the problem as a node anomaly detection task under the realistic assumption of only having normal traffic samples for training the GNNs. Furthermore, we have also restricted the false positive rate below 1% to make the system practical for human operators willing to use it as an Anomaly-based IDS (A-IDS). Finally, the experimental results highlight the relevance of structural information to effectively address the task in real-world conditions. • The security of IoT devices is becoming a relevant issue in modern networks. • Host-based analysis is unfeasible in many IoT devices. • Graph Neural Networks are becoming a promising tool for network traffic analysis. • Anomaly detection is the most suitable approach scenarios where attacks are unknow. • A comparison among ML methods and GNNs for anomaly detection is proposed.
Vincenzo Carletti, Pasquale Foggia, Francesco Rosa, Mario Vento
Pattern Recognit. Lett.4
2024 Improving Learning from Visual Demonstration Methods by Target Localization
abstract
This paper presents a novel approach to multi-task visual-guided imitation learning. Upon evaluating the current state-of-the-art method, we observed its capability to replicate the intent of the demonstrator, but with the flaw of manipulating incorrect objects. To address this issue, our study introduces a new approach based on the assumption that explicitly addressing task-relevant problems, such as target object localization, can enhance system performance. Our validation shows that the proposal overtakes the leading method thanks to its ability to drive the robot motion towards the target object.
Pasquale Foggia, Francesco Rosa, Mario Vento
RO-MAN3
2024 Empowering Human Interaction: A Socially Assistive Robot for Support in Trade Shows
abstract
Social robots are increasingly finding applications in sectors such as industry, retail, and healthcare. They employ a combination of verbal and non-verbal cues in order to allow an empathetic and efficient interaction with the humans. In contrast to traditional robots, social robots possess contextual awareness, enabling intelligent responses to human interactions. Within this context, in this paper we propose a framework for social robots, integrating advanced audio and video analytics capabilities with a novel engagement algorithm, designed and developed to facilitate effective communication in multi-user settings. Furthermore, other than conversation skills, the proposed social robot is enhanced with the capability to interactively play games with the human. The effectiveness of the proposed system was tested in real-world scenarios, during two trade shows in Italy, providing valuable insights into its performance and adaptability to different contexts and different audiences.
Giuseppe De Simone, Alessia Saggese, Mario Vento
RO-MAN3
2024 Robust speech command recognition in challenging industrial environments
Stefano Bini, Vincenzo Carletti, Alessia Saggese, Mario Vento
Comput. Commun.4
2024 Facial Soft-biometrics Obfuscation through Adversarial Attacks
abstract
Sharing facial pictures through online services, especially on social networks, has become a common habit for thousands of users. This practice hides a possible threat to privacy: the owners of such services, as well as malicious users, could automatically extract information from faces using modern and effective neural networks. In this article, we propose the harmless use of adversarial attacks, i.e., variations of images that are almost imperceptible to the human eye and that are typically generated with the malicious purpose to mislead Convolutional Neural Networks (CNNs). Such attacks have been instead adopted to (1) obfuscate soft biometrics (gender, age, ethnicity) but (2) without degrading the quality of the face images posted online. We achieve the above-mentioned two conflicting goals by modifying the implementations of four of the most popular adversarial attacks, namely FGSM, PGD, DeepFool, and C&W, in order to constrain the average amount of noise they generate on the image and the maximum perturbation they add on the single pixel. We demonstrate, in an experimental framework including three popular CNNs, namely VGG16, SENet, and MobileNetV3, that the considered obfuscation method, which requires at most 4 seconds for each image, is effective not only when we have a complete knowledge of the neural network that extracts the soft biometrics (white box attacks) but also when the adversarial attacks are generated in a more realistic black box scenario. Finally, we prove that an opponent can implement defense techniques to partially reduce the effect of the obfuscation, but substantially paying in terms of accuracy over clean images; this result, confirmed by the experiments carried out with three popular defense methods, namely adversarial training, denoising autoencoder, and Kullback-Leibler autoencoder, shows that it is not convenient for the opponent to defend himself and that the proposed approach is robust to defenses.
Vincenzo Carletti, Pasquale Foggia, Antonio Greco 0001, Alessia Saggese, Mario Vento
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Highly Crowd Detection and Counting Based on Curriculum Learning
Lidia Fotia, Gennaro Percannella, Alessia Saggese, Mario Vento
CAIP (2)4
2023 Assistive force control in collaborative human-robot transportation
abstract
Collaborative robotics has gained significant traction in the industrial scenario due to its ability to merge human cognitive abilities with robot strength and dexterity. One specific area where this technology is promising is the transportation of heavy and/or bulky objects. In the scenarios where the human leads, physical human-robot interaction triggers cognitive human-robot interaction, by which the robot is called to adapt its behavior to the collaborator’s intention. Based on this principle, this paper introduces a novel control architecture, namely assistive force control (AFC), by which the robot’s purpose is to alleviate the human collaborator’s effort during transportation. Instead of acting on the robot’s motion, the AFC acts on its causes, by intuitively defining assistive forces, which are input to a lower-level direct force controller. We validate the proposed architecture on two real-case transportation scenarios involving an industrial robot collaboratively carrying objects with different subjects. Our preliminary results show that low effort is required for human operators to manipulate heavy objects, confirming that the proposed architecture is well-suited for collaborative transportation in real-world scenarios.
Bruno G. C. Lima, Enrico Ferrentino, Pasquale Chiacchio, Mario Vento
RO-MAN4
2023 Multi-task learning on the edge for effective gender, age, ethnicity and emotion recognition
Pasquale Foggia, Antonio Greco 0001, Alessia Saggese, Mario Vento
Eng. Appl. Artif. Intell.4
2023 A Social Robot Architecture for Personalized Real-Time Human-Robot Interaction
abstract
In the age of the Internet of Things (IoT), the combination of robotics and artificial intelligence has paved the way for the development of social robots able to undertake realistic conversations with humans, making them the perfect human interface in applications like for instance, robotic house assistants and hotel concierges. Despite the several solutions developed in recent years, the definition of social robot requirements and the software modules needed to meet such requirements has not yet been formalized. In this article, we define the requirements of a social robot and propose a software architecture that includes all the necessary modules to meet them. The proposed architecture, implemented using robot operating system (ROS) nodes, is hardware-independent, enabling its reuse across different robotic platforms with interchangeable types of sensors and actuators. We deployed a social robot based on this architecture to interact with attendees of a real exhibition context and validated the reliability of our proposed solution through a survey of 161 users. The results of the user study indicated a high-quality user experience with the social robot, with scores ranging from 4 to 5 (being 5 the maximum score). Sharing these design choices and evaluation results could significantly benefit the development of future social robotics applications in the context of IoT.
Pasquale Foggia, Antonio Greco 0001, Antonio Roberto, Alessia Saggese, Mario Vento
IEEE Internet Things J.5
2023 Benchmarking deep networks for facial emotion recognition in the wild
abstract
Abstract Emotion recognition from face images is a challenging task that gained interest in recent years for its applications to business intelligence and social robotics. Researchers in computer vision and affective computing focused on optimizing the classification error on benchmark data sets, which do not extensively cover possible variations that face images may undergo in real environments. Following on investigations carried out in the field of object recognition, we evaluated the robustness of existing methods for emotion recognition when their input is subjected to corruptions caused by factors present in real-world scenarios. We constructed two data sets on top of the RAF-DB test set, named RAF-DB-C and RAF-DB-P, that contain images modified with 18 types of corruption and 10 of perturbation. We benchmarked existing networks (VGG, DenseNet, SENet and Xception) trained on the original images of RAF-DB and compared them with ARM, the current state-of-the-art method on the RAF-DB test set. We carried out an extensive study on the effects that modifications to the training data or network architecture have on the classification of corrupted and perturbed data. We observed a drop of recognition performance of ARM, with the classification error raising up to 200% of that achieved on the original RAF-DB test set. We demonstrate that the use of the AutoAugment data augmentation and an anti-aliasing filter within down-sampling layers provide existing networks with increased robustness to out-of-distribution variations, substantially reducing the error on corrupted inputs and outperforming ARM. We provide insights about the resilience of existing emotion recognition methods and an estimation of their performance in real scenarios. The processing time required by the modifications we investigated (35 ms in the worst case) supports their suitability for application in real-world scenarios. The RAF-DB-C and RAF-DB-P test sets, trained models and evaluation framework are available at https://github.com/MiviaLab/emotion-robustness .
Antonio Greco 0001, Nicola Strisciuglio, Mario Vento, Vincenzo Vigilante
Multim. Tools Appl.3
2023 Degramnet: effective audio analysis based on a fully learnable time-frequency representation
abstract
Abstract Current state-of-the-art audio analysis algorithms based on deep learning rely on hand-crafted Spectrogram-like audio representations, that are more compact than descriptors obtained from the raw waveform; the latter are, in turn, far from achieving good generalization capabilities when few data are available for the training. However, Spectrogram-like representations have two main limitations: (1) The parameters of the filters are defined a priori, regardless of the specific audio analysis task; (2) such representations do not perform any denoising operation on the audio signal, neither in the time domain nor in the frequency domain. To overcome these limitations, we propose a new general-purpose convolutional architecture for audio analysis tasks that we call DEGramNet, which is trained with audio samples described with a novel, compact and learnable time–frequency representation that we call DEGram. The proposed representation is fully trainable: Indeed, it is able to learn the frequencies of interest for the specific audio analysis task; in addition, it performs denoising through a custom time–frequency attention module, which amplifies the frequency and time components in which the sound is actually located. It implies that the proposed representation can be easily adapted to the specific problem at hands, for instance giving more importance to the voice frequencies when the network needs to be used for speaker recognition. DEGramNet achieved state-of-the-art performance on the VGGSound dataset (for Sound Event Classification) and comparable accuracy with a complex and special-purpose approach based on network architecture search over the VoxCeleb dataset (for Speaker Identification). Moreover, we demonstrate that DEGram allows to achieve high accuracy with lightweight neural networks that can be used in real-time on embedded systems, making the solution suitable for Cognitive Robotics applications.
Pasquale Foggia, Antonio Greco 0001, Antonio Roberto, Alessia Saggese, Mario Vento
Neural Comput. Appl.5
2023 A deep learning based system for handwashing procedure evaluation
abstract
Abstract Hand washing preparation can be considered as one of the main strategies for reducing the risk of surgical site contamination and thus the infections risks. Within this context, in this paper we propose an embedded system able to automatically analyze, in real-time, the sequence of images acquired by a depth camera to evaluate the quality of the handwashing procedure. In particular, the designed system runs on an NVIDIA Jetson Nano $$^{{\mathrm{TM}}}$$ TM computing platform. We adopt a convolutional neural network, followed by a majority voting scheme, to classify the movement of the worker according to one of the ten gestures defined by the World Health Organization. To test the proposed system, we collect a dataset built by 74 different video sequences. The results achieved on this dataset confirm the effectiveness of the proposed approach.
Antonio Greco 0001, Gennaro Percannella, Pierluigi Ritrovato, Alessia Saggese, Mario Vento
Neural Comput. Appl.5
2023 A multi-task network for speaker and command recognition in industrial environments
abstract
In industrial environments, it is crucial to establish a strong collaboration between humans and robots to enhance productivity. However, the nature of the work demands that workers have the authority to provide specific instructions to the robots. The scientific community has extensively investigated these dual requirements, aiming to develop advanced systems capable of recognizing voice commands and implementing speaker authentication. Nevertheless, in the industrial context, these tasks should be executed simultaneously on low-cost and low-power embedded devices that can be mounted on board the robotic platform. To overcome this challenge, we propose a multi-task network for Speech-Command Recognition and Speaker Identification. Additionally, we employ the GradNorm adaptive algorithm to address the issue of task imbalance. To evaluate the proposed system, we introduce a new dataset, MIVIA-ISC, consisting of 20,857 samples uttered by 562 speakers for 31 distinct commands. Our approach significantly reduces the network size by 47% and its execution time by 48% compared to the commonly used methodology, which employs one network for each task. Furthermore, our approach demonstrates a significant improvement in the accuracy of the Speaker Identification task, achieving an 11% increase compared to the corresponding single-task network. Importantly, this enhancement is achieved without compromising the accuracy of the Speech-Command Recognition task, which experiences only a minimal 3% decrease in performance.
Stefano Bini, Gennaro Percannella, Alessia Saggese, Mario Vento
Pattern Recognit. Lett.4
2022 Joint Intensity Classification and Specimen Segmentation on HEp-2 Images: a Deep Learning Approach
abstract
Antinuclear antibody (ANA) testing is performed to help diagnose patients with possible autoimmune diseases. The indirect immunofluorescence (IIF) microscopy performed with human epithelial type 2 (HEp-2) cells as the substrate is the reference method for ANA screening. It allows for the detection of antibody binding to specific intracellular targets, resulting in distinct staining patterns whose recognition may be helpful for diagnosis purposes. In recent years, there has been an increasing interest in devising deep learning methods for automatic segmentation and classification of staining patterns, as well as for other tasks related to this diagnostic technique. However, little attention has been paid to architectures aimed at managing more functions related to ANA testing by simultaneously training tasks that use a shared representation. In this paper, we propose a deep neural network model based on U-Net that exploits an end-to-end approach for joint intensity classification and specimen segmentation on HEp-2 cell images. To the best of our knowledge, this is the first work proposing the adoption of a single framework tailored to address these two mandatory tasks in the HEp-2 clinical workflow. The experiments were conducted on I3A, the largest publicly available dataset of HEp-2 images. The results showed that the proposed approach outperformed the competing state-of-the-art methods for both the considered tasks, achieving (in 5-fold cross validation) 93.83% and 90.21% for intensity classification accuracy and segmentation accuracy, respectively.
Gennaro Percannella, Umberto Petruzzello, Pierluigi Ritrovato, Leonardo Rundo, Francesco Tortorella, Mario Vento
ICPR6
2022 Benchmarking deep neural networks for gesture recognition on embedded devices
abstract
The gesture is one of the most used forms of communication between humans; in recent years, given the new trend of factories to be adapted to Industry 4.0 paradigm, the scientific community has shown a growing interest towards the design of Gesture Recognition (GR) algorithms for Human-Robot Interaction (HRI) applications. Within this context, the GR algorithm needs to work in real time and over embedded platforms, with limited resources. Anyway, when looking at the available scientific literature, the aim of the different proposed neural networks (i.e. 2D and 3D) and of the different modalities used for feeding the network (i.e. RGB, RGB-D, optical flow) is typically the optimization of the accuracy, without strongly paying attention to the feasibility over low power hardware devices. Anyway, the analysis related to the trade-off between accuracy and computational burden (for both networks and modalities) becomes important so as to allow GR algorithms to work in industrial robotics applications. In this paper, we perform a wide benchmarking focusing not only on the accuracy but also on the computational burden, involving two different architectures (2D and 3D), with two different backbones (MobileNet, ResNeXt) and four types of input modalities (RGB, Depth, Optical Flow, Motion History Image) and their combinations.
Stefano Bini, Antonio Greco 0001, Alessia Saggese, Mario Vento
RO-MAN4
2022 Effective training of convolutional neural networks for age estimation based on knowledge distillation
abstract
Abstract Age estimation from face images can be profitably employed in several applications, ranging from digital signage to social robotics, from business intelligence to access control. Only in recent years, the advent of deep learning allowed for the design of extremely accurate methods based on convolutional neural networks (CNNs) that achieve a remarkable performance in various face analysis tasks. However, these networks are not always applicable in real scenarios, due to both time and resource constraints that the most accurate approaches often do not meet. Moreover, in case of age estimation, there is the lack of a large and reliably annotated dataset for training deep neural networks. Within this context, we propose in this paper an effective training procedure of CNNs for age estimation based on knowledge distillation, able to allow smaller and simpler “student” models to be trained to match the predictions of a larger “teacher” model. We experimentally show that such student models are able to almost reach the performance of the teacher, obtaining high accuracy over the LFW+, LAP 2016 and Adience datasets, but being up to 15 times faster. Furthermore, we evaluate the performance of the student models in the presence of image corruptions, and we demonstrate that some of them are even more resilient to these corruptions than the teacher model.
Antonio Greco 0001, Alessia Saggese, Mario Vento, Vincenzo Vigilante
Neural Comput. Appl.3
2022 Vehicles Detection for Smart Roads Applications on Board of Smart Cameras: A Comparative Analysis
abstract
Video analytics can be profitably adopted in smart roads environments to automatically detect abnormal situations. Within this context, vehicle detection is the first and foremost stage, and its accuracy is crucial, since any detection error will affect the performance of any subsequent step. Furthermore, in smart road environments it is often preferred to perform the video analysis directly on board of smart surveillance cameras, in order to reduce bandwidth usage and eliminate the cost of setup and maintenance of powerful processing servers; on the flip side, processing on board of smart cameras implies the detection algorithm to be fast and slim, since the resources available on this kind of embedded device are limited. In the era of deep learning, it seems that the questionwhat is the best method for vehicle detection?may have a trivial answer, since this class of methods includes some very accurate ones. Anyway, according to the above consideration, the best suited method for this application is not necessarily the most accurate one, but for sure the most accurate one running on the available hardware at a given resolution and frame rate. Starting from the above considerations, in this paper we perform an analysis of the methods available in the literature for vehicle detection, by comparing them in terms of accuracy and computational burden, with the aim to answer the following question:what is the best method for vehicles detection when working with smart cameras?
Antonio Greco 0001, Alessia Saggese, Mario Vento, Vincenzo Vigilante
IEEE Trans. Intell. Transp. Syst.3
2021 Sentiment analysis for customer relationship management: an incremental learning approach
abstract
Abstract In recent years there has been a significant rethinking of corporate management, which is increasingly based on customer orientation principles. As a matter of fact, customer relationship management processes and systems are ever more popular and crucial to facing today’s business challenges. However, the large number of available customer communication stimuli coming from different (direct and indirect) channels, require automatic language processing techniques to help filter and qualify such stimuli, determine priorities, facilitate the routing of requests and reduce the response times. In this scenario, sentiment analysis plays an important role in measuring customer satisfaction, tracking consumer opinion, interacting with consumers and building customer loyalty. The research described in this paper proposes an approach based on Hierarchical Attention Networks for detecting the sentiment polarity of customer communications. Unlike other existing approaches, after initial training, the defined model can improve over time during system operation using the feedback provided by CRM operators thanks to an integrated incremental learning mechanism. The paper also describes the developed prototype as well as the dataset used for training the model which includes over 30.000 annotated items. The results of two experiments aimed at measuring classifier performance and validating the retraining mechanism are also presented and discussed. In particular, the classifier accuracy turned out to be better than that of other algorithms for the supported languages (macro-averaged f1-score of 0.89 and 0.79 for Italian and English respectively) and the retraining mechanism was able to improve the classification accuracy on new samples without degrading the overall system performance.
Nicola Capuano, Luca Greco 0001, Pierluigi Ritrovato, Mario Vento
Appl. Intell.4
2021 DENet: a deep architecture for audio surveillance applications
Antonio Greco 0001, Antonio Roberto, Alessia Saggese, Mario Vento
Neural Comput. Appl.4
2021 Two parallel versions of VF3: Performance analysis on a wide database of graphs
Vincenzo Carletti, Pasquale Foggia, Gennaro Percannella, Pierluigi Ritrovato, Mario Vento
Pattern Recognit. Lett.5
2020 Which are the factors affecting the performance of audio surveillance systems?
abstract
Sound event recognition systems are rapidly becoming part of our life, since they can be profitably used in several vertical markets, ranging from audio security applications to scene classification and multi-modal analysis in social robotics. In the last years, a not negligible part of the scientific community started to apply Convolutional Neural Networks (CNNs) to image-based representations of the audio stream, due to their successful adoption in almost all the computer vision tasks. In this paper, we carry out a detailed benchmark of various widely used CNN architectures and visual representations on a popular dataset, namely the MIVIA Audio Events database. Our analysis is aimed at understanding how these factors affect the sound event recognition performance with a particular focus on the false positive rate, very relevant in audio surveillance solutions. In fact, although most of the proposed solutions achieve a high recognition rate, the capability of distinguishing the events-of-interest from the background is often not yet sufficient for real systems, and prevent its usage in real applications. Our comprehensive experimental analysis investigates this aspect and allows to identify useful design guidelines for increasing the specificity of sound event recognition systems.
Antonio Greco 0001, Antonio Roberto, Alessia Saggese, Mario Vento
ICPR4
2020 Benchmarking deep network architectures for ethnicity recognition using a new large face dataset
abstract
Abstract Although in recent years we have witnessed an explosion of the scientific research in the recognition of facial soft biometrics such as gender, age and expression with deep neural networks, the recognition of ethnicity has not received the same attention from the scientific community. The growth of this field is hindered by two related factors: on the one hand, the absence of a dataset sufficiently large and representative does not allow an effective training of convolutional neural networks for the recognition of ethnicity; on the other hand, the collection of new ethnicity datasets is far from simple and must be carried out manually by humans trained to recognize the basic ethnicity groups using the somatic facial features. To fill this gap in the facial soft biometrics analysis, we propose the VGGFace2 Mivia Ethnicity Recognition (VMER) dataset, composed by more than 3,000,000 face images annotated with 4 ethnicity categories, namely African American, East Asian, Caucasian Latin and Asian Indian. The final annotations are obtained with a protocol which requires the opinion of three people belonging to different ethnicities, in order to avoid the bias introduced by the well-known other race effect. In addition, we carry out a comprehensive performance analysis of popular deep network architectures, namely VGG-16, VGG-Face, ResNet-50 and MobileNet v2. Finally, we perform a cross-dataset evaluation to demonstrate that the deep network architectures trained with VMER generalize on different test sets better than the same models trained on the largest ethnicity dataset available so far. The ethnicity labels of the VMER dataset and the code used for the experiments are available upon request at https://mivia.unisa.it .
Antonio Greco 0001, Gennaro Percannella, Mario Vento, Vincenzo Vigilante
Mach. Vis. Appl.3
2020 Age from Faces in the Deep Learning Revolution
abstract
Face analysis includes a variety of specific problems as face detection, person identification, gender and ethnicity recognition, just to name the most common ones; in the last two decades, significant research efforts have been devoted to the challenging task of age estimation from faces, as witnessed by the high number of published papers. The explosion of the deep learning paradigm, that is determining a spectacular increasing of the performance, is in the public eye; consequently, the number of approaches based on deep learning is impressively growing and this also happened for age estimation. The exciting results obtained have been recently surveyed on almost all the specific face analysis problems; the only exception stands for age estimation, whose last survey dates back to 2010 and does not include any deep learning based approach to the problem. This paper provides an analysis of the deep methods proposed in the last six years; these are analysed from different points of view: the network architecture together with the learning procedure, the used datasets, data preprocessing and augmentation, and the exploitation of additional data coming from gender, race and face expression. The review is completed by discussing the results obtained on public datasets, so as the impact of different aspects on system performance, together with still open issues.
Vincenzo Carletti, Antonio Greco 0001, Gennaro Percannella, Mario Vento
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Trends in graph-based representations for Pattern Recognition
Luc Brun, Pasquale Foggia, Mario Vento
Pattern Recognit. Lett.3
2020 Comparing performance of graph matching algorithms on huge graphs
Vincenzo Carletti, Pasquale Foggia, Antonio Greco 0001, Alessia Saggese, Mario Vento
Pattern Recognit. Lett.5
2020 Special issue on "Applications of graph-based techniques to pattern recognition"
Pasquale Foggia, Mario Vento, Cheng-Lin Liu 0001
Pattern Recognit. Lett.2
2020 Trends in IoT based solutions for health care: Moving AI to the edge
Luca Greco 0001, Gennaro Percannella, Pierluigi Ritrovato, Francesco Tortorella, Mario Vento
Pattern Recognit. Lett.5
2020 AReN: A Deep Learning Approach for Sound Event Recognition Using a Brain Inspired Representation
abstract
Audio surveillance is gaining in the last years wide interest. This is due to the large number of situations in which this kind of systems can be used, either alone or combined with video-based algorithms. In this paper we propose a deep learning method to automatically recognize events of interest in the context of audio surveillance (namely screams, broken glasses and gun shots). The audio stream is represented by a gammatonegram image. We propose a 21-layer CNN to which we feed sections of the gammatonegram representation. At the output of this CNN there are units that correspond to the classes. We trained the CNN, called AReN, by taking advantage of a problem-driven data augmentation, which extends the training dataset with gammatonegram images extracted by sounds acquired with different signal to noise ratios. We experimented it with three datasets freely available, namely SESA, MIVIA Audio Events and MIVIA Road Events and we achieved 91.43%, 99.62% and 100% recognition rate, respectively. We compared our method with other state of the art methodologies based both on traditional machine learning methodologies and deep learning. The comparison confirms the effectiveness of the proposed approach, which outperforms the existing methods in terms of recognition rate. We experimentally prove that the proposed network is resilient to the noise, has the capability to significantly reduce the false positive rate and is able to generalize in different scenarios. Furthermore, AReN is able to process 5 audio frames per second on a standard CPU and, consequently, it is suitable for real audio surveillance applications.
Antonio Greco 0001, Nicolai Petkov, Alessia Saggese, Mario Vento
IEEE Trans. Inf. Forensics Secur.4
2019 A System for Controlling How Carefully Surgeons Are Cleaning Their Hands
Luca Greco 0001, Gennaro Percannella, Pierluigi Ritrovato, Alessia Saggese, Mario Vento
CAIP (2)5
2019 A Challenging Voice Dataset for Robotic Applications in Noisy Environments
Antonio Roberto, Alessia Saggese, Mario Vento
CAIP (2)3
2019 MIVIABot: A Cognitive Robot for Smart Museum
Alessia Saggese, Mario Vento, Vincenzo Vigilante
CAIP (1)2
2019 Emotion analysis from faces for social robotics
abstract
A social robot is able to perceive the information about the environment (both in terms of persons and objects populating the scene), to reason about the acquired information and to interact with the human in a proper way. Among the information required for the interaction with a human, the capability of analysing the emotion is surely among the most important ones. Another relevant feature of social robots is the possibility to interact with the human in real time, without any latency that could introduce a delay in the talk between the human and the robot. It means that all the processing of the information (for instance the analysis of the sequence of images for the detection of the persons and the consequent emotion analysis) needs to be performed directly on board of the robot, without any possibility to use high performance servers (for instance services on the cloud) but only small devices that can be installed directly on board of the robotic platform. In this paper we propose MIVIAbot, a robotic platform based on the Pepper humanoid robot and equipped with a small, low cost and low energy consumption embedded device installed directly on board of Pepper, without any requirements for a Wi-Fi connection that could introduce any latency for the transmission of the data. Furthermore, we also propose a method for the analysis of the emotion of the persons by face analysis, able to run on the considered hardware platform in real time without paying in terms of accuracy. The experimentation conducted over various widely adopted datasets of images and videos for emotion analysis confirms both the efficiency and the effectiveness of the proposed approach.
Antonio Greco 0001, Antonio Roberto, Alessia Saggese, Mario Vento, Vincenzo Vigilante
SMC4
2019 SoReNet: a novel deep network for audio surveillance applications
abstract
In the era of third generation surveillance systems, it becomes more and more useful to have available a solution able to automatically detect abnormal events. The interest for audio analysis is thus growing in the last years, due to the large amount of situations where a microphone and an audio surveillance system can be profitably used by the human operator in charge of control. In this paper, we propose a method for automatically analyzing the audio stream for surveillance purposes: it is able to detect the presence of abnormal events such as screams, gun shots and broken glasses. Instead than processing directly raw data (the audio signal), the stream is represented by means of an image, namely the spectrogram, a time-frequency representation of the audio stream. In this way, we formulate the problem of audio analysis as a problem of image classification. Thus, we propose to use a Convolutional Neural Network with the following two main properties: inspired by VGG network, we employed very small kernels in convolutional layers; furthermore, we adopted a pyramidal structure in fully connected layers. These choices allow to have good generalization capabilities of the network even in presence of a not so wide dataset. The performance, computed over a standard dataset already used for benchmarking purposes in the field of audio surveillance, confirms the effectiveness of the proposed approach.
Antonio Greco 0001, Alessia Saggese, Mario Vento, Vincenzo Vigilante
SMC3
2019 A human-like description of scene events for a proper UAV-based video content analysis
Danilo Cavaliere, Vincenzo Loia, Alessia Saggese, Sabrina Senatore, Mario Vento
Knowl. Based Syst.5
2019 Learning representations of sound using trainable COPE feature extractors
abstract
Sound analysis research has mainly been focused on speech and music processing. The deployed methodologies are not suitable for analysis of sounds with varying background noise, in many cases with very low signal-to-noise ratio (SNR). In this paper, we present a method for the detection of patterns of interest in audio signals . We propose novel trainable feature extractors, which we call COPE (Combination of Peaks of Energy). The structure of a COPE feature extractor is determined using a single prototype sound pattern in an automatic configuration process , which is a type of representation learning. We construct a set of COPE feature extractors, configured on a number of training patterns. Then we take their responses to build feature vectors that we use in combination with a classifier to detect and classify patterns of interest in audio signals . We carried out experiments on four public data sets: MIVIA audio events, MIVIA road events, ESC-10 and TU Dortmund data sets. The results that we achieved (recognition rate equal to 91.71% on the MIVIA audio events, 94% on the MIVIA road events, 81.25% on the ESC-10 and 94.27% on the TU Dortmund) demonstrate the effectiveness of the proposed method and are higher than the ones obtained by other existing approaches. The COPE feature extractors have high robustness to variations of SNR. Real-time performance is achieved even when the value of a large number of features is computed.
Nicola Strisciuglio, Mario Vento, Nicolai Petkov
Pattern Recognit.2
2019 VF3-Light: A lightweight subgraph isomorphism algorithm and its experimental evaluation
Vincenzo Carletti, Pasquale Foggia, Antonio Greco 0001, Mario Vento, Vincenzo Vigilante
Pattern Recognit. Lett.4
2019 Learning skeleton representations for human action recognition
abstract
Automatic interpretation of human actions gained strong interest among researchers in patter recognition and computer vision because of its wide range of applications, such as in social and home robotics, elderly people health care, surveillance, among others. In this paper, we propose a method for recognition of human actions by analysis of skeleton poses. The method that we propose is based on novel trainable feature extractors, which can learn the representation of prototype skeleton examples and can be employed to recognize skeleton poses of interest. We combine the proposed feature extractors with an approach for classification of pose sequences based on string kernels. We carried out experiments on three benchmark data sets (MIVIA-S, MSRSDA and MHAD) and the results that we achieved are comparable or higher than the ones obtained by other existing methods. A further important contribution of this work is the MIVIA-S dataset, that we collected and made publicly available.
Alessia Saggese, Nicola Strisciuglio, Mario Vento, Nicolai Petkov
Pattern Recognit. Lett.3
2019 Semantically Enhanced UAVs to Increase the Aerial Scene Understanding
abstract
Visual tracking supported by unmanned aerial vehicles (UAVs) has generated a lot of interest in recent years, especially in application domains such as surveillance, search for missing persons and traffic monitoring. The major challenges in visual tracking with small UAVs arise in the form of target representation, target appearance change, target detection and localization in real time computation. Reliable target detection depends on factors such as occlusions, image noise, illumination and pose changes, or image blur that may compromise the object labeling. To mitigate these issues, this paper proposes a hybrid solution: along with the tracked objects, scenes are completely depicted by adding contextual information, i.e., data describing places, natural features, or in general points of interest. Each scenario indeed is semantically described by ontological statements that define the context and then, by inference, support the object tracking task in the object identification and labeling. The synergy between the tracking methods and semantic modeling can bridge the object labeling gap, enhancing the scene understanding and awareness when alarming situations are discovered. Experimental results are promising and confirm the applicability of the proposed framework in supporting drones in object identification and critical situation detection tasks.
Danilo Cavaliere, Vincenzo Loia, Alessia Saggese, Sabrina Senatore, Mario Vento
IEEE Trans. Syst. Man Cybern. Syst.5
2018 Gender recognition from face images using trainable shape and color features
abstract
Gender recognition from face images is an important application and it is still an open computer vision problem, even though it is something trivial from the human visual system. Variations in pose, lighting, and expression are few of the problems that make such an application challenging for a computer system. Neurophysiological studies demonstrate that the human brain is able to distinguish men and women also in absence of external cues, by analyzing the shape of specific parts of the face. In this paper, we describe an automatic procedure that combines trainable shape and color features for gender classification. In particular the proposed method fuses edge-based and color-blob-based features by means of trainable COSFIRE filters. The former types of feature are able to extract information about the shape of a face whereas the latter extract information about shades of colors in different parts of the face. We use these two sets of features to create a stacked classification SVM model and demonstrate its effectiveness on the GENDER-COLOR-FERET dataset, where we achieve an accuracy of 96.4%.
George Azzopardi, Pasquale Foggia, Antonio Greco 0001, Alessia Saggese, Mario Vento
ICPR5
2018 Challenging the Time Complexity of Exact Subgraph Isomorphism for Huge and Dense Graphs with VF3
abstract
Graph matching is essential in several fields that use structured information, such as biology, chemistry, social networks, knowledge management, document analysis and others. Except for special classes of graphs, graph matching has in the worst-case an exponential complexity; however, there are algorithms that show an acceptable execution time, as long as the graphs are not too large and not too dense. In this paper we introduce a novel subgraph isomorphism algorithm, VF3, particularly efficient in the challenging case of graphs with thousands of nodes and a high edge density. Its performance, both in terms of time and memory, has been assessed on a large dataset of 12,700 random graphs with a size up to 10,000 nodes, made publicly available. VF3 has been compared with four other state-of-the-art algorithms, and the huge experimentation required more than two years of processing time. The results confirm that VF3 definitely outperforms the other algorithms when the graphs become huge and dense, but also has a very good performance on smaller or sparser graphs.
Vincenzo Carletti, Pasquale Foggia, Alessia Saggese, Mario Vento
IEEE Trans. Pattern Anal. Mach. Intell.4
2017 Fast gender recognition in videos using a novel descriptor based on the gradient magnitudes of facial landmarks
abstract
The growing interest in recent years for gender recognition from face images is mainly attributable to the wide range of possible applications that can be used for commercial and marketing purposes. It is desirable that such algorithms process high resolution video frames acquired by using surveillance cameras in real-time. To the best of our knowledge, however, there are no studies which analyze the computational impact of the methods and the difficulties related to the processing of faces extracted from videos captured in the wild. We propose a novel face descriptor based on the gradient magnitudes of facial landmarks, which are points automatically extracted from the face contour, eyes, eyebrows, nose, mouth and chin. We evaluate the effectiveness and efficiency of the proposed approach on two new datasets, which we made available online and that consist of color face images and color video sequences acquired in real scenarios. The proposed approach is more efficient and effective than three commercial libraries.
George Azzopardi, Antonio Greco 0001, Alessia Saggese, Mario Vento
AVSS4
2017 An efficient and effective method for people detection from top-view depth cameras
abstract
The detection of persons from videos is particularly important in many computer vision contexts being an enabling technology for several relevant applications either for security and safety or for business intelligence purposes. The adoption of a depth sensor mounted in a top-view position is often used to achieve high person detection accuracy as it allows to cope effectively with occlusions and difficult lighting conditions. In this paper, we propose a new method for people detection from depth maps produced by sensors mounted in a zenithal position. The method is designed with the aim of providing an optimal trade off between the detection accuracy and the computational complexity. The proposed approach adopts a dynamic background modeling strategy in order to find the objects of interest into the scene; then a lightweight algorithm is used to filter out the noise from the foreground image and to determine the position of the persons into the scene. The experimental analysis carried out on a public and large dataset allowed to demonstrate that the method is fast and accurate. The method has been compared with respect to two different approaches available in the literature for people detection from a depth camera mounted in a zenithal position: an unsupervised method that is fast although not highly accurate, and a supervised one that conversely is very accurate but less computationally efficient. The proposed method allows to achieve comparable accuracy of the supervised approach using very few computational resources, with a reduction of an order of magnitude of the processing times.
Vincenzo Carletti, Luca Del Pizzo, Gennaro Percannella, Mario Vento
AVSS4
2017 A real-time system for audio source localization with cheap sensor device
abstract
We propose an architecture for real-time audio source localization based on the integration of localization methodologies within a framework that employs a cheap acquisition sensor. The architecture that we present takes as input the audio signals from two calibrated microphones. Then, it computes biological-inspired features of the sound signal and estimates its direction by means of a Gaussian Mixture Model estimator. We carried out an extensive experimental analysis on four data sets, one of which we realized and made publicly available. We evaluated several characteristics of the sound localization architecture and its use in real scenarios.
Alessia Saggese, Nicola Strisciuglio, Mario Vento, Nicolai Petkov
AVSS3
2017 Graph edit distance as a quadratic assignment problem
Sébastien Bougleux, Luc Brun, Vincenzo Carletti, Pasquale Foggia, Benoit Gaüzère, Mario Vento
Pattern Recognit. Lett.6
2016 Gender recognition from face images with trainable COSFIRE filters
abstract
Gender recognition from face images is an important application in the fields of security, retail advertising and marketing. We propose a novel descriptor based on COSFIRE filters for gender recognition. A COSFIRE filter is trainable, in that its selectivity is determined in an automatic configuration process that analyses a given prototype pattern of interest. We demonstrate the effectiveness of the proposed approach on a new dataset called GENDER-FERET with 474 training and 472 test samples and achieve an accuracy rate of 93.7%. It also outperforms an approach that relies on handcrafted features and an ensemble of classifiers. Furthermore, we perform another experiment by using the images of the Labeled Faces in the Wild (LFW) dataset to train our classifier and the test images of the GENDER-FERET dataset for evaluation. This experiment demonstrates the generalization ability of the proposed approach and it also outperforms two commercial libraries, namely Face++ and Luxand.
George Azzopardi, Antonio Greco 0001, Mario Vento
AVSS3
2016 Towards semantic context-aware drones for aerial scenes understanding
abstract
Visual object tracking with unmanned aerial vehicles (UAVs) plays a central role in the aerial surveillance. Reliable object detection depends on many factors such as large displacements, occlusions, image noise, illumination and pose changes or image blur that may compromise the object labeling. The paper presents a proposal for a hybrid solution that adds semantic information to the video tracking processing: along with the tracked objects, the scene is completely depicted by data from places, natural features, or in general Points of Interest (POIs). Each scene from a video sequence is semantically described by ontological statements which, by inference, support the object identification which often suffers from some weakness in the object tracking methods. The synergy between the tracking methods and semantic technologies seems to bridge the object labeling gap, enhance the understanding of the situation awareness, as well as critical alarming situations.
Danilo Cavaliere, Sabrina Senatore, Mario Vento, Vincenzo Loia
AVSS3
2016 Improving reliability of people tracking by adding semantic reasoning
abstract
Even the best performing object tracking algorithm on well known datasets, commits several errors that prevent a concrete adoption in real case scenarios unless you do not accept some compromise about tracking quality and reliability. The aim of this paper is to demonstrate that adding to a traditional object tracking solution a knowledge based reasoner build on top of semantic web technologies, it is possible to identify and properly manage common tracking problems. The proposed approach has been evaluated using View 001 and View 003 of the PETS2009 dataset with interesting results.
Luca Greco 0001, Pierluigi Ritrovato, Alessia Saggese, Mario Vento
AVSS4
2016 Time-frequency analysis for audio event detection in real scenarios
abstract
We propose a sound analysis system for the detection of audio events in surveillance applications. The method that we propose combines short- and long-time analysis in order to increase the reliability of the detection. The basic idea is that a sound is composed of small, atomic audio units and some of them are distinctive of a particular class of sounds. Similarly to the words in a text, we count the occurrence of audio units for the construction of a feature vector that describes a given time interval. A classifier is then used to learn which audio units are distinctive for the different classes of sound. We compare the performance of different sets of short-time features by carrying out experiments on the MIVIA audio event data set. We study the performance and the stability of the proposed system when it is employed in live scenarios, so as to characterize its expected behavior when used in real applications.
Alessia Saggese, Nicola Strisciuglio, Mario Vento, Nicolai Petkov
AVSS3
2016 Interactive pose calibration of a set of cameras for video surveillance
abstract
There has been an increase of video surveillance systems in operation in public areas. The classical systems simply send the images to monitors. Nevertheless, there is a demand on giving more intelligence on these systems and asking them to automatically track objects or recognise people. One of the basic low-level tasks that these systems have to face with is the accurate deduction of the cameras' poses. We present a method that deducts these poses in an interactive way when the automatic method fails or generates a large error. The user is asked for mapping some points between the images from these cameras when the alignment between them fails in a completely automatic way. Experimental validation has demonstrated that with really few interactions, the reduction of the pose error is considerable.
Gaetano Manzo, Francesc Serratosa, Mario Vento
ETFA3
2016 International Contest on Pattern Recognition techniques for indirect immunofluorescence images analysis
abstract
This contest is a joint initiative organized by the University of Salerno (Italy) and the University of Queensland (Australia) with the support of the Sullivan Nicolaides Pathology (SNP), Australia. The contest primarily aims to provide a platform for scientists and practitioners for performing research to develop Computer Aided Diagnosis (CAD) systems for pathology tests utilizing indirect immunofluorescence protocol. In particular, the contest considers the Antinuclear Antibodies (ANA) test using Human Epithelial type 2 (HEp-2) cells. The competition is divided into four tasks that address specific problems: (1) HEp-2 cell classification; (2) Patient specimen classification; (3) HEp-2 mitotic cell identification and (4) Cell segmentation.
Brian C. Lovell, Gennaro Percannella, Alessia Saggese, Mario Vento, Arnold Wiliem
ICPR4
2016 Action recognition by using kernels on aclets sequences
Luc Brun, Gennaro Percannella, Alessia Saggese, Mario Vento
Comput. Vis. Image Underst.4
2016 Online human assisted and cooperative pose estimation of 2D cameras
Gaetano Manzo, Francesc Serratosa, Mario Vento
Expert Syst. Appl.3
2016 Supervised vessel delineation in retinal fundus images with the automatic selection of B-COSFIRE filters
abstract
The inspection of retinal fundus images allows medical doctors to diagnose various pathologies. Computer-aided diagnosis systems can be used to assist in this process. As a first step, such systems delineate the vessel tree from the background. We propose a method for the delineation of blood vessels in retinal images that is effective for vessels of different thickness. In the proposed method, we employ a set of B -COSFIRE filters selective for vessels and vessel-endings. Such a set is determined in an automatic selection process and can adapt to different applications. We compare the performance of different selection methods based upon machine learning and information theory. The results that we achieve by performing experiments on two public benchmark data sets, namely DRIVE and STARE, demonstrate the effectiveness of the proposed approach.
Nicola Strisciuglio, George Azzopardi, Mario Vento, Nicolai Petkov
Mach. Vis. Appl.3
2016 Executable thematic special issue on pattern recognition techniques for indirect immunofluorescence images analysis
Mehrtash Harandi, Brian C. Lovell, Gennaro Percannella, Alessia Saggese, Mario Vento, Arnold Wiliem
Pattern Recognit. Lett.5
2016 Computer Aided Diagnosis for Anti-Nuclear Antibodies HEp-2 images: Progress and challenges
Peter Hobson, Brian C. Lovell, Gennaro Percannella, Alessia Saggese, Mario Vento, Arnold Wiliem
Pattern Recognit. Lett.5
2016 HEp-2 staining pattern recognition at cell and specimen levels: Datasets, algorithms and results
Peter Hobson, Brian C. Lovell, Gennaro Percannella, Alessia Saggese, Mario Vento, Arnold Wiliem
Pattern Recognit. Lett.5
2016 Counting people by RGB or depth overhead cameras
Luca Del Pizzo, Pasquale Foggia, Antonio Greco 0001, Gennaro Percannella, Mario Vento
Pattern Recognit. Lett.5
2016 Audio Surveillance of Roads: A System for Detecting Anomalous Sounds
abstract
In the last decades, several systems based on video analysis have been proposed for automatically detecting accidents on roads to ensure a quick intervention of emergency teams. However, in some situations, the visual information is not sufficient or sufficiently reliable, whereas the use of microphones and audio event detectors can significantly improve the overall reliability of surveillance systems. In this paper, we propose a novel method for detecting road accidents by analyzing audio streams to identify hazardous situations such as tire skidding and car crashes. Our method is based on a two-layer representation of an audio stream: at a low level, the system extracts a set of features that is able to capture the discriminant properties of the events of interest, and at a high level, a representation based on a bag-of-words approach is then exploited in order to detect both short and sustained events. The deployment architecture for using the system in real environments is discussed, together with an experimental analysis carried out on a data set made publicly available for benchmarking purposes. The obtained results confirm the effectiveness of the proposed approach.
Pasquale Foggia, Nicolai Petkov, Alessia Saggese, Nicola Strisciuglio, Mario Vento
IEEE Trans. Intell. Transp. Syst.5
2015 Automatic detection of long term parked cars
abstract
The detection of illegal roadside parking is becoming more and more interesting in the field of intelligent transportation systems, since it may cause traffic congestion or accidents. In this paper we propose a method able to analyze videos acquired by traditional surveillance cameras and to automatically detect the vehicles stopped in a forbidden area. Two main contributions have been introduced: first, spatio temporal information related to the stopped vehicles are encoded by a heat map; second, the background is not updated by evaluating the movement of the vehicle in a single time instant, but instead the whole movement of the vehicles, encoded into the heat map, is taken into account. Two widely adopted datasets, namely the iLids and the PETS 2000, have been used to experimentally evaluate the proposed approach and the results achieved, compared with state of the art methodologies, confirm its effectiveness.
Vincenzo Carletti, Pasquale Foggia, Antonio Greco 0001, Alessia Saggese, Mario Vento
AVSS5
2015 Human action recognition using an improved string edit distance
abstract
In this paper we propose an improvement of a human action recognition method that uses a string-based representation and a string edit distance to compare the observed action with reference actions in the training set. In particular, the original improvement is based on a specific formulation of the string edit distance that is more suited to take into account the problems related to noise and to different execution speeds that are observed in an action recognition system. The experimentation has been carried out on two widely adopted datasets, namely the MIVIA and the MHAD datasets, and the obtained results, compared with both the original method and other state of the art approaches, confirm the significance of the proposed improvement and the effectiveness of the method.
Pasquale Foggia, Benoit Gaüzère, Alessia Saggese, Mario Vento
AVSS4
2015 Car crashes detection by audio analysis in crowded roads
abstract
In the last years, video surveillance has been employed for roads monitoring in order to detect abnormal events and improve the safety procedures in case of emergency. Certain events, such as car crashes or tire skidding, are difficult or impossible to detect when only the visual information is considered. In this paper we describe a preliminary system to detect events in roads by means of audio analysis. The system that we propose combines short- and long-time analysis of the audio signal in order to detect both impulsive and sustained events. We present the preliminary results achieved by the proposed system on a data set specifically made for roads surveillance, which we made publicly available. We also discuss the architectural deployment of such system in real environments with respect to a model of the noise of road traffic. The achieved results are promising and confirm the effectiveness of the system.
Pasquale Foggia, Alessia Saggese, Nicola Strisciuglio, Mario Vento, Nicolai Petkov
AVSS4
2015 A Verification-Based Multithreshold Probing Approach to HEp-2 Cell Segmentation
Xiaoyi Jiang 0001, Gennaro Percannella, Mario Vento
CAIP (2)3
2015 Locally Adapted Gain Control for Reliable Foreground Detection
Duber Martinez, Alessia Saggese, Mario Vento, Humberto Loaiza, Eduardo F. Caicedo
CAIP (1)3
2015 Multiscale Blood Vessel Delineation Using B-COSFIRE Filters
Nicola Strisciuglio, George Azzopardi, Mario Vento, Nicolai Petkov
CAIP (2)3
2015 Benchmarking human epithelial type 2 interphase cells classification methods on a very large dataset
Peter Hobson, Brian C. Lovell, Gennaro Percannella, Mario Vento, Arnold Wiliem
Artif. Intell. Medicine4
2015 A hierarchical neuro-fuzzy architecture for human behavior analysis
Giovanni Acampora, Pasquale Foggia, Alessia Saggese, Mario Vento
Inf. Sci.4
2015 Trainable COSFIRE filters for vessel delineation with application to retinal images
George Azzopardi, Nicola Strisciuglio, Mario Vento, Nicolai Petkov
Medical Image Anal.3
2015 A long trip in the charming world of graphs for Pattern Recognition
Mario Vento
Pattern Recognit.1
2015 Reliable detection of audio events in highly noisy environments
Pasquale Foggia, Nicolai Petkov, Alessia Saggese, Nicola Strisciuglio, Mario Vento
Pattern Recognit. Lett.5
2015 Real-Time Fire Detection for Video-Surveillance Applications Using a Combination of Experts Based on Color, Shape, and Motion
abstract
In this paper, we propose a method that is able to detect fires by analyzing videos acquired by surveillance cameras. Two main novelties have been introduced. First, complementary information, based on color, shape variation, and motion analysis, is combined by a multiexpert system. The main advantage deriving from this approach lies in the fact that the overall performance of the system significantly increases with a relatively small effort made by the designer. Second, a novel descriptor based on a bag-of-words approach has been proposed for representing motion. The proposed method has been tested on a very large dataset of fire videos acquired both in real environments and from the web. The obtained results confirm a consistent reduction in the number of false positives, without paying in terms of accuracy or renouncing the possibility to run the system on embedded platforms.
Pasquale Foggia, Alessia Saggese, Mario Vento
IEEE Trans. Circuits Syst. Video Technol.3
2015 Designing Huge Repositories of Moving Vehicles Trajectories for Efficient Extraction of Semantic Data
abstract
The rapid development of digital cameras equipped with video analytics software is providing the availability of large amount of traffic data describing the trajectories traced by each vehicle and person within a scene. These data offer enormous potential when coupled with a querying system able to extract synthetic but meaningful information as those obtained by spatiotemporal queries; the latter allow, for instance, to select all those trajectories passing through some parts of the scene, even in given sequences, and adding restrictions on the properties of the objects (the category of the vehicles, their color and size, and so on). In this paper we propose a novel system for efficiently storing and querying large amounts of 3D data (trajectories over time), specifically designed for making possible the formulation of a wide variety of spatio-temporal 3D queries. The method is based on a novel 3D data schema which is reconducted to a set of 2D schemata, being the latter the only ones available in currently ready-to-use database environments. An implementation of the system over PostGIS is presented in this paper, together with a performance assessment on a huge trajectory database. The obtained results confirm the effectiveness of the proposed approach and its applicability to real applications.
Antonio d'Acierno, Alessia Saggese, Mario Vento
IEEE Trans. Intell. Transp. Syst.3
2014 Detection of anomalous driving behaviors by unsupervised learning of graphs
abstract
In this paper we propose a graph based approach for detecting abnormal behaviors starting from the analysis of vehicles' trajectories. The scene is partitioned into zones and is dynamically represented as a graph by evaluating the distribution of trajectories belonging to the training set. Furthermore, four different strategies are proposed in order to verify if a test trajectory belongs to the scene and then can be considered normal by evaluating the probability that this trajectory belongs to the graph. Our algorithms have been tested on the standard MIT Trajectories dataset and the obtained results confirm the effectiveness of the proposed approach.
Luc Brun, Benito Cappellania, Alessia Saggese, Mario Vento
AVSS4
2014 HAck: A system for the recognition of human actions by kernels of visual strings
abstract
In this paper we propose HAcK, a novel method for recognizing Human Actions by string Kernel; the main idea is to represent each action through a sequence of visual characters, namely a string, able to model the temporal evolution of the events. Visual characters are extracted by analyzing global descriptors of the scene and by taking advantage on the depth information provided by a Kinect sensor. The similarity between actions is evaluated with a fast global alignment kernel, which allows to deal with actions of different length as well as with the noise introduced during the features extraction step. HAcK has been evaluated over two standard datasets and the obtained results, compared with state of the art approaches, confirm its effectiveness and its applicability in real environments.
Luc Brun, Gennaro Percannella, Alessia Saggese, Mario Vento
AVSS4
2014 A reliable string kernel based approach for solving queries by sketch
abstract
In this paper we propose a novel and efficient method for solving queries by sketch in traffic scenarios, aiming to find the k nearest neighbor trajectories to the one hand drawn by the human operator. Each trajectory is represented as a sequence of symbols, namely a string, and it is stored into a k-d tree by taking into account the similarity between trajectories, evaluated by a global fast alignment kernel. The experimentation has been conducted over the standard MIT trajectories dataset and results confirm the effectiveness and the robustness of the proposed approach.
Luc Brun, Alessia Saggese, Mario Vento
AVSS3
2014 Cascade classifiers trained on gammatonegrams for reliably detecting audio events
abstract
In this paper we propose a novel method for the detection of events of interest through audio analysis. The system that we propose is based on the representation of the audio streams through a Gammatone image, which describes the time-frequency distribution of the energy of the signal; this representation is inspired by the functioning of the human auditory system. A pool of AdaBoost cascade classifiers, one for each class of events of interest, is involved in the event detection stage. The performance of the proposed system has been evaluated on a large data set of audio events for surveillance applications and the achieved results, compared with two state of the art approaches, confirm its effectiveness.
Pasquale Foggia, Alessia Saggese, Nicola Strisciuglio, Mario Vento
AVSS4
2014 Exploiting the deep learning paradigm for recognizing human actions
abstract
In this paper we propose a novel method for recognizing human actions by exploiting a multi-layer representation based on a deep learning based architecture. A first level feature vector is extracted and then a high level representation is obtained by taking advantage of a Deep Belief Network trained using a Restricted Boltzmann Machine. The classification is finally performed by a feed-forward neural network. The main advantage behind the proposed approach lies in the fact that the high level representation is automatically built by the system exploiting the regularities in the dataset; given a suitably large dataset, it can be expected that such a representation can outperform a hand-design description scheme. The proposed approach has been tested on two standard datasets and the achieved results, compared with state of the art algorithms, confirm its effectiveness.
Pasquale Foggia, Alessia Saggese, Nicola Strisciuglio, Mario Vento
AVSS4
2014 Classifying Anti-nuclear Antibodies HEp-2 Images: A Benchmarking Platform
abstract
There has been an ongoing effort in improving reliability and consistency of pathology test results due to their critical role in making an accurate diagnosis. One way to do this is by applying image-based Computer Aided Diagnosis (CAD) systems. This paper proposes a comprehensive benchmarking platform comprising over 1,000 images to evaluate CAD systems for the Anti-Nuclear Antibody (ANA) test via the Indirect Immunofluorescence (IIF) protocol applied on Human Epithelial Type 2 (HEp-2) cells. While prior works in this domain have primarily focussed on classifying individual cell images derived from ANA IIF HEp-2 images, our proposed benchmarking platform goes beyond this by considering the ANA IIF HEp-2 image classification problem. Generally the existing works derive an ANA IIF HEp-2 image label from the dominant pattern of the cell images (we call this approach baseline). In this work, we argue that this approach cannot be used to achieve an acceptable performance, thus, the problem of classifying ANA IIF HEp-2 images (or ANA images in short) is still largely unexplored. To demonstrate that, we propose a simple-yet-effective CAD system which is inspired from the recent success of object bank representation in the object classification domain. We evaluate the proposed system, the baseline and a recent CAD system and show that our proposed system considerably outperforms the others.
Peter Hobson, Brian C. Lovell, Gennaro Percannella, Mario Vento, Arnold Wiliem
ICPR4
2014 Graph Matching and Learning in Pattern Recognition in the Last 10 Years
abstract
In this paper, we examine the main advances registered in the last ten years in Pattern Recognition methodologies based on graph matching and related techniques, analyzing more than 180 papers; the aim is to provide a systematic framework presenting the recent history and the current developments. This is made by introducing a categorization of graph-based techniques and reporting, for each class, the main contributions and the most outstanding research results.
Pasquale Foggia, Gennaro Percannella, Mario Vento
Int. J. Pattern Recognit. Artif. Intell.3
2014 Special issue on the analysis and recognition of indirect immuno-fluorescence images
Pasquale Foggia, Gennaro Percannella, Paolo Soda, Mario Vento
Pattern Recognit.4
2014 Pattern recognition in stained HEp-2 cells: Where are we now?
Pasquale Foggia, Gennaro Percannella, Alessia Saggese, Mario Vento
Pattern Recognit.4
2014 Mitotic cells recognition in HEp-2 images
Giulio Iannello, Gennaro Percannella, Paolo Soda, Mario Vento
Pattern Recognit. Lett.4
2014 Dynamic Scene Understanding for Behavior Analysis Based on String Kernels
abstract
This paper aims at dynamically understanding the properties of a scene from the analysis of moving object trajectories. Two different applications are proposed: the former is devoted to identify abnormal behaviors, while the latter allows to extract the k, most of the similar trajectories to the one hand-drawn by an human operator. A set of normal trajectories' models is extracted using a novel unsupervised learning technique: the scene is adaptively partitioned into zones using the distribution of the training set and each trajectory is represented as a sequence of symbols by considering positional information (the zones crossed in the scene), speed, and shape. The main novelty is the use of a kernel-based approach for evaluating the similarity between the trajectories. Furthermore, we define a novel and efficient kernel-based clustering algorithm, aimed at obtaining groups of normal trajectories. Experimentations, conducted over three standard data sets, confirm the effectiveness of the proposed approach.
Luc Brun, Alessia Saggese, Mario Vento
IEEE Trans. Circuits Syst. Video Technol.3
2013 Audio surveillance using a bag of aural words classifier
abstract
In this paper we propose a novel approach for the audio-based detection of events. The approach adopts the bag of words paradigm, and has two main advantages over other techniques present in the literature: the ability to automatically adapt (through a learning phase) to both short, impulsive sounds and long, sustained ones, and the ability to work in noisy environments where the sounds of interest are superimposed to background sounds possibly having similar characteristics. The proposed method has been experimentally validated on a large database of sounds, including several kinds of background noise, which are superimposed to the sounds to be recognized. The obtained performance has been compared with the results of another audio event detection algorithm from the literature, showing a significant improvement.
Vincenzo Carletti, Pasquale Foggia, Gennaro Percannella, Alessia Saggese, Nicola Strisciuglio, Mario Vento
AVSS6
2013 Tracking System with Re-identification Using a Graph Kernels Approach
Amal Mahboubi, Luc Brun, Donatello Conte, Pasquale Foggia, Mario Vento
CAIP (1)5
2013 Recognizing Human Actions by a Bag of Visual Words
abstract
In this paper a novel method for action recognition based on the bag of visual words approach is proposed. The main contribution is to model each action through a high level features vector computed as the histogram of the visual words: the visual words are extracted by analyzing global descriptors of the scene and their occurrences are evaluated according to a codebook, a kind of dictionary, which encodes the typical visual words, automatically extracted during the learning phase. The classification is performed by using an SVM classifier, trained only by using high level features vectors, in order to increase the overall reliability of the system. The experimentation has been conducted over two recently proposed datasets, the MIVIA and the MHAD, the promising results confirm the robustness and the stability of the proposed approach.
Pasquale Foggia, Gennaro Percannella, Alessia Saggese, Mario Vento
SMC4
2013 A real time algorithm for people tracking using contextual reasoning
Rosario Di Lascio, Pasquale Foggia, Gennaro Percannella, Alessia Saggese, Mario Vento
Comput. Vis. Image Underst.5
2013 Counting moving persons in crowded scenes
Donatello Conte, Pasquale Foggia, Gennaro Percannella, Mario Vento
Mach. Vis. Appl.4
2013 Benchmarking HEp-2 Cells Classification Methods
abstract
In this paper, we report on the first edition of the HEp-2 Cells Classification contest, held at the 2012 edition of the International Conference on Pattern Recognition, and focused on indirect immunofluorescence (IIF) image analysis. The IIF methodology is used to detect autoimmune diseases by searching for antibodies in the patient serum but, unfortunately, it is still a subjective method that depends too heavily on the experience and expertise of the physician. This has been the motivation behind the recent initial developments of computer aided diagnosis systems in this field. The contest aimed to bring together researchers interested in the performance evaluation of algorithms for IIF image analysis: 28 different recognition systems able to automatically recognize the staining pattern of cells within IIF images were tested on the same undisclosed dataset. In particular, the dataset takes into account the six staining patterns that occur most frequently in the daily diagnostic practice: centromere, nucleolar, homogeneous, fine speckled, coarse speckled, and cytoplasmic. In the paper, we briefly describe all the submitted methods, analyze the obtained results, and discuss the design choices conditioning the performance of each method.
Pasquale Foggia, Gennaro Percannella, Paolo Soda, Mario Vento
IEEE Trans. Medical Imaging4
2012 Combining Neural Networks and Fuzzy Systems for Human Behavior Understanding
abstract
The psychological overcharge issue related to human inadequacy to maintain a constant level of attention in simultaneously monitoring multiple visual information sources makes necessary to develop enhanced video surveillance systems that automatically understand human behaviors and identify dangerous situations. This paper introduces a semantic human behavioral analysis (HBA) system based on a neuro-fuzzy approach that, independently from the specific application, translates tracking kinematic data into a collection of semantic labels characterizing the behavior of different actors in a scene in order to appropriately classify the current situation. Different from other HBA approaches, the proposed system shows high level of scalability, robustness and tolerance for tracking imprecision and, for this reason, it could represent a valid choice for improving the performance of current systems.
Giovanni Acampora, Pasquale Foggia, Alessia Saggese, Mario Vento
AVSS4
2012 An Ensemble of Rejecting Classifiers for Anomaly Detection of Audio Events
abstract
Audio analytic systems are receiving an increasing interest in the scientific community, not only as stand alone systems for the automatic detection of abnormal events by the interpretation of the audio track, but also in conjunction with video analytics tools for enforcing the evidence of anomaly detection. In this paper we present an automatic recognizer of a set of abnormal audio events that works by extracting suitable features from the signals obtained by microphones installed into a surveilled area, and by classifying them using two classifiers that operate at different time resolutions. An original aspect of the proposed system is the estimation of the reliability of each response of the individual classifiers. In this way, each classifier is able to reject the samples having an overall reliability below a threshold. This approach allows our system to combine only reliable decisions, so increasing the overall performance of the method. The system has been tested on a large dataset of samples acquired from real world scenarios, the audio classes of interests are represented by gunshot, scream and glass breaking in addition to the background sounds. The preliminary results obtained encourage further research in this direction.
Donatello Conte, Pasquale Foggia, Gennaro Percannella, Alessia Saggese, Mario Vento
AVSS5
2012 A classification-based approach to segment HEp-2 cells
abstract
In this paper we propose a new method for cells segmentation in HEp-2 images addressing and overcoming the main limitations of the existing approaches. The proposed method adopts image reconstruction for a preliminary image segmentation and, then, it employs a sort of classifier-controlled dilation for better determining the structure of the cells, where the classifier is trained using data of the image at hand. We compare the performance of the proposed method with the most representative approaches from the scientific literature on a common and publicly available dataset of HEp-2 images.
Gennaro Percannella, Paolo Soda, Mario Vento
CBMS3
2012 Removing Object Reflections in Videos by Global Optimization
abstract
This paper presents a novel algorithm for the removal of reflections generated by objects on reflecting floors. The algorithm uses both chromatic properties of the reflections and geometrical constraints on their positions; however, it does not make use of a model of the reflected objects, and so it can be applied to scenes containing several kinds of objects (e.g., people, baggage, animals, vehicles, etc.). The proposed method has been validated by an extensive set of experiments on a large video database. In these experiments, the method has been compared to two other recent reflection removal algorithms. The experimental results show that the proposed method is fast and effective, both in absolute terms and in comparison with the other algorithms.
Donatello Conte, Pasquale Foggia, Gennaro Percannella, Mario Vento
IEEE Trans. Circuits Syst. Video Technol.4
2011 A MultiView Appearance Model for people re-identification
abstract
Tracking moving objects across events that break the continuity of the trajectory, such as occlusions or temporary exits from the scene, usually requires that a model of the objects of interest is created and maintained. Commonly used model representations are prone to errors when the objects can change the direction of their motion. In this paper we introduce a novel model representation, the MultiView Appearance Model, specifically devised to deal with the issue. The algorithms that create and update the model also take into account the problem of object apparent size changes due to perspective. An experimental evaluation of the proposed model representation has been performed on the PETS2010 dataset. The results show a consistent improvement of the performances in comparison with another well-known appearance model.
Donatello Conte, Pasquale Foggia, Gennaro Percannella, Mario Vento
AVSS4
2011 Trainable estimators for indirect people counting: A comparative study
abstract
Estimating the number of people in a scene is a very relevant issue due to the possibility of using it in a large number of contexts where it is necessary to automatically monitor an area for security/safety reasons, for economic purposes, etc. The large number of people counting approaches available in the literature can be roughly abscribed to two categories: direct approaches and indirect ones. In the first category there are methods that first detect people and then count them; differently, the indirect methods face the counting problem by establishing a relation between some scene features and the estimated number of people. Some recent comparative evaluations carried out in the framework of the PETS initiative have demonstrated that the indirect methods tends to be more robust than direct ones, above all when they are used in very crowded conditions. In this paper, we analyze the behavior of an indirect approach that is based on a trainable estimator that does not require an explicit formulation of a priori knowledge about the perspective and density effects present in the scene at hand. In particular, we investigate on the way the counting accuracy in different crowding conditions is affected by the choice of the trainable estimator.
Giovanni Acampora, Vincenzo Loia, Gennaro Percannella, Mario Vento
FUZZ-IEEE4
2011 A middleware platform for real-time processing of multiple video streams based on the data-flow paradigm
abstract
In this paper we introduce a new software platform for the realization of intelligent video-surveillance applications and, more generally, of real-time video stream processing systems. The platform is implemented as a middleware, providing general purpose services, and a collection of dynamically loaded modules carrying out domain-specific tasks. The architecture of the platform follows a data-flow paradigm, where the application is organized as a processing network whose nodes are activated by the middleware as soon as their inputs are available and a processor is ready. This architecture is beneficial both with respect to the development process, simplifying the module implementation and favoring the reuse of software components, and with respect to the performance, since the middleware can automatically parallelize the processing using the available processors or cores. The platform has been validated by converting an existing video surveillance application, demonstrating both the improvement in the development process and the performance increment.
Pasquale Foggia, Mario Vento
ICME2
2010 A Method for Counting People in Crowded Scenes
abstract
This paper presents a novel method to count people for video surveillance applications. Methods in the literature either follow a direct approach, by first detecting people and then counting them, or an indirect approach, by establishing a relation between some easily detectable scene features and the estimated number of people. The indirect approach is considerably more robust, but it is not easy to take into account such factors as perspective or people groups with different densities. The proposed technique, while based on the indirect approach, specifically addresses these problems; furthermore it is based on a trainable estimator that does not require an explicit formulation of a priori knowledge about the perspective and density effects present in the scene at hand. In the experimental evaluation, the method has been extensively compared with the algorithm by Albiol et al, which provided the highest performance at the PETS 2009 contest on people counting. The experimentation has used the public PETS 2009 datasets. The results confirm that the proposed method improves the accuracy, while retaining the robustness of the indirect approach.
Donatello Conte, Pasquale Foggia, Gennaro Percannella, Francesco Tufano, Mario Vento
AVSS5
2010 A Method Based on the Indirect Approach for Counting People in Crowded Scenes
abstract
This paper presents a method for counting people in a scene by establishing a mapping between some scene features and the number of people avoiding the complex foreground detection problem. The method is based on the use of SURF features and of an ∈-SVR regressor to provide an estimate of this count. The algorithm takes specifically into account problems due to partial occlusions and to perspective.
Donatello Conte, Pasquale Foggia, Gennaro Percannella, Mario Vento
AVSS4
2010 Performance Evaluation of a People Tracking System on PETS2009 Database
abstract
In this paper a system for autonomous video surveillance in relatively unconstrained environments is described. The system consists of two principal phases: object detection and object tracking. An adaptive background subtraction, together with a set of corrective algorithms, is used to cope with variable lighting, dynamic and articulate scenes, etc. The tracking algorithm is based on a matrix representation of the problem, and is used to face splitting and occlusion problems. When the tracking algorithm fails in following actual object trajectories, an appearance-based module is used to restore object identities. An experimental evaluation, carried out on the PETS2009 dataset for tracking, shows promising results.
Donatello Conte, Pasquale Foggia, Gennaro Percannella, Mario Vento
AVSS4
2010 Early experiences in mitotic cells recognition on HEp-2 slides
abstract
Indirect immunofluorescence (IIF) imaging is the recommended laboratory technique to detect autoantibodies in patient serum, but it suffers from several issues limiting its reliability and reproducibility. IIF slides are observed by specialists at the fluorescence microscope, reporting fluorescence intensity and staining pattern and looking for mitotic cells. Indeed, the presence of such cells is a key factor to assess the correctness of slide preparation process and the reported staining pattern. Therefore, the ability to detect mitotic cells is needed to develop a complete computer-aided-diagnosis system in IIF, which can support the specialists from image acquisition up to image classification. Although recent research in IIF has been directed to image acquisition, image segmentation, fluorescence intensity classification and staining pattern recognition, no works presented methods suited to classify such cells. Hence, this paper presents an heterogeneous set of features used to describe the peculiarities of mitotic cells and then tests five classifiers, belonging to different classification paradigms. The approach has been evaluated on an annotated dataset of mitotic cells. The measured performances are promising, achieving a classification accuracy of 86.5 %.
Pasquale Foggia, Gennaro Percannella, Paolo Soda, Mario Vento
CBMS4
2010 Counting Moving People in Videos by Salient Points Detection
abstract
This paper presents a novel method to count people for video surveillance applications. The problem is faced by establishing a mapping between some scene features and the number of people. Moreover, the proposed technique takes specifically into account problems due to perspective. In the experimental evaluation, the method has been compared with respect to the algorithm by Albiol et al., which provided the highest performance at the PETS 2009 contest on people counting, using the same datasets. The results confirm that the proposed method improves the accuracy, while retaining the robustness of Albiol's algorithm.
Donatello Conte, Pasquale Foggia, Gennaro Percannella, Francesco Tufano, Mario Vento
ICPR5
2010 Reflection Removal in Color Videos
abstract
This paper presents a novel method for reflection removal in the context of an object detection system. The method is based on chromatic properties of the reflections and does not require a geometric model of the objects. An experimental evaluation of the proposed method has been performed on a large database, showing its effectiveness.
Donatello Conte, Pasquale Foggia, Gennaro Percannella, Francesco Tufano, Mario Vento
ICPR5
2009 An Algorithm for Detection of Partially Camouflaged People
abstract
Several video analysis applications perform object detection using a background subtraction approach. Camouflage can be a serious problem for these applications, since the objects of interest may appear fragmented into small,disconnected pieces, with a dramatic negative impact on later processing phases such as classification or tracking. Nevertheless, this problem is largely underestimated in the literature. In this paper an effective, model-based solution is presented for the case of people detection. The proposed method acts as a post-processing phase, grouping together the fragmented blocks to restore the original object. A quantitative evaluation of the effectiveness of this method has been performed on real world videos from a video-surveillance application. The videos used for the experiments (with metadata) have been made publicly available on the Internet.
Donatello Conte, Pasquale Foggia, Gennaro Percannella, Francesco Tufano, Mario Vento
AVSS5
2009 Benchmarking graph-based clustering algorithms
Pasquale Foggia, Gennaro Percannella, Carlo Sansone, Mario Vento
Image Vis. Comput.4
2009 A multiple expert system for classifying fluorescent intensity in antinuclear autoantibodies analysis
Paolo Soda, Giulio Iannello, Mario Vento
Pattern Anal. Appl.3
2008 A Graph-Based Algorithm for Cluster Detection
abstract
In some Computer Vision applications there is the need for grouping, in one or more clusters, only a part of the whole dataset. This happens, for example, when samples of interest for the application at hand are present together with several noisy samples. In this paper we present a graph-based algorithm for cluster detection that is particularly suited for detecting clusters of any size and shape, without the need of specifying either the actual number of clusters or the other parameters. The algorithm has been tested on data coming from two different computer vision applications. A comparison with other four state-of-the-art graph-based algorithms was also provided, demonstrating the effectiveness of the proposed approach.
Pasquale Foggia, Gennaro Percannella, Carlo Sansone, Mario Vento
Int. J. Pattern Recognit. Artif. Intell.4
2007 A New Approach for Stereo Matching in Autonomous Mobile Robot Applications
Pasquale Foggia, Jean-Michel Jolion, Alessandro Limongiello, Mario Vento
IJCAI4
2007 Segmentation of news videos based on audio-video information
Massimo De Santo, Gennaro Percannella, Carlo Sansone, Mario Vento
Pattern Anal. Appl.4
2006 Learning Graphs from Examples: an Application to the Prediction of the Toxicity of Chemical Compounds
abstract
A common problem encountered in structural pattern recognition is the difficulty of constructing classification models or rules from a set of examples, due to the complexity of the structures needed to represent the patterns. In this paper, we present an extension of a method for structural learning. The goal of the method is to find descriptions which are general (in other words, are successfully applicable to recognize objects different from the ones in the training set), preserving at the same time their discrimination ability. This method has been applied to predictive toxicology evaluation, that is the inference of the cancerogenic characteristics of chemical compounds.
Pasquale Foggia, Alessandro Limongiello, Francesco Tufano, Mario Vento
Int. J. Pattern Recognit. Artif. Intell.4
2006 Preface
Luc Brun, Mario Vento
Pattern Recognit.2
2006 A graph-based, multi-resolution algorithm for tracking objects in presence of occlusions
Donatello Conte, Pasquale Foggia, Jean-Michel Jolion, Mario Vento
Pattern Recognit.4
2005 A Graph-Theoretical Clustering Method for Detecting Clusters of Micro-Calcifications in Mammographic Images
abstract
In this paper we propose a method based on a graph-theoretical cluster analysis for automatically finding cluster of micro-calcifications in mammographic images. It is applied to the image after a micro-calcification detection phase and is able to cope with the unavoidable false positives that each automatic detection algorithm produces. The proposed approach has been tested on a standard database of 40 mammographic images and revealed to be very effective even when the detection phase gives rise to several false positives.
Luigi P. Cordella, Gennaro Percannella, Carlo Sansone, Mario Vento
CBMS4
2004 Thirty Years Of Graph Matching In Pattern Recognition
abstract
A recent paper posed the question: "Graph Matching: What are we really talking about?". Far from providing a definite answer to that question, in this paper we will try to characterize the role that graphs play within the Pattern Recognition field. To this aim two taxonomies are presented and discussed. The first includes almost all the graph matching algorithms proposed from the late seventies, and describes the different classes of algorithms. The second taxonomy considers the types of common applications of graph-based techniques in the Pattern Recognition and Machine Vision field.
Donatello Conte, Pasquale Foggia, Carlo Sansone, Mario Vento
Int. J. Pattern Recognit. Artif. Intell.4
2004 A Multi-Expert System For Shot Change Detection In Mpeg Movies
abstract
Shot Change Detection (SCD) in MPEG coded videos is a complex and still open research problem whose interest is growing up more and more due to the diffusion of Video Databases and Digital Libraries. Techniques providing fully satisfactory performances on complex video domains are not yet available even if a number of proposals exist; such proposals show very often to be complementary in their results. In this context, the Authors investigated the use of Multi-Expert Systems (MES) for approaching the SCD problem. In the present paper, we propose and discuss a strategy to select the SCD techniques to be combined and a method for choosing an effective combining rule. In order to assess the performance of the proposed MES, we set up a database that is significantly wider than the ones commonly used in the field. Experimental results demonstrate that the proposed system performs better than each of the single SCD technique considered.
Massimo De Santo, Gennaro Percannella, Carlo Sansone, Mario Vento
Int. J. Pattern Recognit. Artif. Intell.4
2004 Combining experts for anchorperson shot detection in news videos
Massimo De Santo, Gennaro Percannella, Carlo Sansone, Mario Vento
Pattern Anal. Appl.4
2004 A (Sub)Graph Isomorphism Algorithm for Matching Large Graphs
abstract
We present an algorithm for graph isomorphism and subgraph isomorphism suited for dealing with large graphs. A first version of the algorithm has been presented in a previous paper, where we examined its performance for the isomorphism of small and medium size graphs. The algorithm is improved here to reduce its spatial complexity and to achieve a better performance on large graphs; its features are analyzed in detail with special reference to time and memory requirements. The results of a testing performed on a publicly available database of synthetically generated graphs and on graphs relative to a real application dealing with technical drawings are presented, confirming the effectiveness of the approach, especially when working with large graphs.
Luigi P. Cordella, Pasquale Foggia, Carlo Sansone, Mario Vento
IEEE Trans. Pattern Anal. Mach. Intell.4
2003 Weighted minimum common supergraph for cluster representation
abstract
Graphs are a powerful and versatile tool useful for representing patterns in various fields of science and engineering. In many applications, for example, in image processing and pattern recognition, it is required to measure the similarity of objects for clustering similar patterns. In this paper a new structural method for representing a cluster of graphs is proposed. Using this method it becomes easy to extract the common information shared in the patterns of a cluster, make evident this information and separate it from noise and distortions that usually affect graph representation of real images.
Horst Bunke, Corrado Guidobaldi, Mario Vento
ICIP (2)3
2003 Graph matching applications in pattern recognition and image processing
abstract
In this paper we will try to characterize the role that graphs are conquering within the pattern recognition field. To this aim, a taxonomy built considering the most common applications of graph based techniques in the pattern recognition and image processing field is presented and discussed.
Donatello Conte, Pasquale Foggia, Carlo Sansone, Mario Vento
ICIP (2)4
2003 Automatic classification of clustered microcalcifications by a multiple expert system
Massimo De Santo, Mario Molinara, Francesco Tortorella, Mario Vento
Pattern Recognit.4
2003 Preface
Pasquale Foggia, Carlo Sansone, Mario Vento
Pattern Recognit. Lett.3
2003 A large database of graphs and its use for benchmarking graph isomorphism algorithms
Massimo De Santo, Pasquale Foggia, Carlo Sansone, Mario Vento
Pattern Recognit. Lett.4
2002 Learning structural shape descriptions from examples
Luigi P. Cordella, Pasquale Foggia, Carlo Sansone, Mario Vento
Pattern Recognit. Lett.4
2001 A Classification Reliability Driven Reject Rule for Multi-Expert Systems
abstract
In this paper we propose a reject rule applicable to a Multi-Expert System (MES). The rule is adaptive to the given domain and allows the achievement of the best trade-off between reject and error rates as a function of the costs attributed to errors and rejects in the considered application. The results of the method are particularly effective since the method does not rely on particular statistical assumptions, as other reject rules. An experimental analysis carried out on publicly available databases is reported together with a comparison with other methods present in the literature.
Carlo Sansone, Francesco Tortorella, Mario Vento
Int. J. Pattern Recognit. Artif. Intell.3
2001 Symbolic vs. Connectionist Learning: An Experimental Comparison in a Structured Domain
abstract
During the last two decades, the attempts to find effective solutions to the problem of learning any kind of structured information have been splitting the scientific community. A "holy war" has been fought between the advocates of a symbolic approach to learning and the advocates of a connectionist approach. One of the most repeated claims of the symbolic party has been that symbolic methods are able to cope with structured information while connectionist ones are not. However, in the last few years, the possibility of employing connectionist methods for structured data has been widely investigated and several approaches have been proposed. A novel algorithm for learning structured descriptions, ascribable to the category of symbolic techniques, is proposed. It faces the problem directly in the space of graphs by defining the proper inference operators, as graph generalization and graph specialization, and obtains general and consistent prototypes with a low computational cost with respect to other symbolic learning systems. The proposed algorithm is compared with a recent connectionist method for learning structured data (P. Frasconi et al., 1998), with reference to a problem of handwritten character recognition from a standard database on the Web. The orthogonality of the two approaches strongly suggests their combination in a multiclassifier system so as to retain the strengths of both of them, while overcoming their weaknesses. The results on an experimental case study demonstrated that the adoption of a parallel combination scheme of the two algorithms could improve the recognition performance by about 10 percent. A truce or an alliance between the symbolic and the connectionist worlds?.
Pasquale Foggia, Roberto Genna, Mario Vento
IEEE Trans. Knowl. Data Eng.3
2000 Fast Graph Matching for Detecting CAD Image Components
abstract
The performance of an attributed relational graph (ARG) matching algorithm, tailored for dealing with large graphs, is evaluated in the context of a real application. The detection of component parts in CAD images of mechanical drawings. The matching problem is a graph-subgraph isomorphism and the algorithm exploits semantic information about nodes while does not require information about the topology of the graphs to be matched. Experimental results, compared with those obtained with a different method, show the overall efficiency of the algorithm and the matching time reduction obtainable by exploiting the semantic information held by ARGs.
Luigi P. Cordella, Pasquale Foggia, Carlo Sansone, Mario Vento
ICPR4
2000 Combining Experts with Different Features for Classifying Clustered Microcalcifications in Mammograms
abstract
At present, mammography is the only non-invasive diagnostic technique of breast cancer at a very early stage. A visual clue of such disease particularly significant is the presence of clusters of microcalcifications. Reliable methods for an automatic recognition of malignant clusters are very difficult to accomplish because of the small size of the microcalcifications and the poor quality of the mammographic images. In this paper we propose a novel approach for automating the recognition of malignant clusters, based on the adoption of a multiple expert system. The approach has been successfully tested on a standard database of 40 mammographic images.
Luigi P. Cordella, Francesco Tortorella, Mario Vento
ICPR3
2000 Symbol recognition in documents: a collection of techniques?
Luigi P. Cordella, Mario Vento
Int. J. Document Anal. Recognit.2
2000 Signature Verification: Increasing Performance by a Multi-Stage System
Carlo Sansone, Mario Vento
Pattern Anal. Appl.2
2000 To reject or not to reject: that is the question-an answer in case of neural classifiers
abstract
A method defining a reject option that is applicable to a given 0-reject classifier is proposed. The reject option is based on an estimate of the classification reliability, measured by a reliability evaluator /spl Psi/. Trivially, once a reject threshold /spl sigma/ has been fixed, a sample is rejected if the corresponding value of /spl Psi/ is below /spl sigma/. Obviously, as /spl sigma/ represents the least tolerable classification reliability level, when its value varies the reject option becomes more or less severe. In order to adapt the behavior of the reject option to the requirements of the considered application domain, a function P characterizing the reject option's adequacy to the domain has been introduced. It is shown that P can be expressed as a function of /spl sigma/ and, consequently, the optimal value for /spl sigma/ is defined as the one which maximizes the function P. The method for determining the optimal threshold value is independent of the specific 0-reject classifier, while the definition of the reliability evaluators is related to the classifier's architecture. General criteria for defining appropriate reliability evaluators within a classification paradigm are illustrated in the paper and are based on the localization, in the feature space, of the samples that could be classified with a low reliability. The definition of the reliability evaluators for three popular architectures of neural networks (backpropagation, learning vector quantization and probabilistic network) is presented. Finally, the method has been tested with reference to a complex classification problem with data generated according to a distribution-of-distributions model.
Claudio De Stefano, Carlo Sansone, Mario Vento
IEEE Trans. Syst. Man Cybern. Part C3
1999 Document Validation by Signature: A Serial Multi-Expert Approach
abstract
A three-stage serial multi-expert system for facing the problem of signature verification is proposed. The first two stages, respectively devoted to the recognition of random and simple forgeries and of skilled forgeries, employ suitable criteria for estimating the reliability of the performed classification. In case of uncertainty the signature is forwarded to the successive stage which takes the final decision, taking into account the decisions of the previous stages together with their reliability estimations. Criteria for selecting the features to be used at each stage and for computing classification reliability and reliability thresholds are discussed The obtained results are compared with those achieved, on the same database of 49 writers, by other existing systems.
Luigi P. Cordella, Pasquale Foggia, Carlo Sansone, Mario Vento
ICDAR4
1999 Combining statistical and structural approaches for handwritten character description
Pasquale Foggia, Carlo Sansone, Francesco Tortorella, Mario Vento
Image Vis. Comput.4
1999 Reliability Parameters to Improve Combination Strategies in Multi-Expert Systems
Luigi P. Cordella, Pasquale Foggia, Carlo Sansone, Mario Vento
Pattern Anal. Appl.4
1999 Definition and Validation of a Distance Measure Between Structural Primitives
Pasquale Foggia, Carlo Sansone, Francesco Tortorella, Mario Vento
Pattern Anal. Appl.4
1999 Multiclassification: reject criteria for the Bayesian combiner
Pasquale Foggia, Carlo Sansone, Francesco Tortorella, Mario Vento
Pattern Recognit.4
1998 Graph matching: a fast algorithm and its evaluation
abstract
A graph matching algorithm is illustrated and its performance compared with that of a well known algorithm performing the same task. According to the proposed algorithm the matching process is carried out by using a state space representation: a state represents a partial solution of the matching between two graphs, and a transition between states corresponds to the addition of a new pair of matched nodes. A set of feasibility rules is introduced for pruning states corresponding to partial matching solutions not satisfying the required graph morphism. Results outlining the computational cost reduction achieved by the method are given with reference to a set of randomly generated graphs.
Luigi P. Cordella, Pasquale Foggia, Carlo Sansone, Francesco Tortorella, Mario Vento
ICPR5
1997 Character Recognition by Geometrical Moments on Structural Decompositions
abstract
A novel description method is presented. It is based on the combination of structural and statistical approaches, and is applied to the problem of unconstrained isolated handwritten character recognition. Characters are preliminarily decomposed in terms of structural primitives (circular arcs) and successively described in terms of statistical features (geometrical moments suitably normalized). A multilayer perceptron is adopted in the classification stage. Results of the method on the digits of the ETL Database are reported.
Pasquale Foggia, Carlo Sansone, Francesco Tortorella, Mario Vento
ICDAR4
1996 An efficient algorithm for the inexact matching of ARG graphs using a contextual transformational model
abstract
The paper illustrates an algorithm for the inexact matching of attributed relational graphs. A sample graph is considered matchable with one of the prototypes if, by using a defined set of syntactic and semantic transformations, it can be made isomorphic to the graph of the prototype. The applicability of a transformation is contextually defined, i.e. each transformation can be defined with reference to a prototype, and can be applied only when the sample graph is being matched with that prototype. The reduction of the computational complexity with respect to a brute-force approach is given with reference to an OCR application.
Luigi P. Cordella, Pasquale Foggia, Carlo Sansone, Mario Vento
ICPR4
1996 A distance measure for structural descriptions using circular arcs as primitives
abstract
This paper proposes a structural description scheme using circular arcs as primitives. On this scheme, a metric for defining a distance between pairs of circular arcs and relations among them, is introduced and its main properties are discussed. This metric is based on a set of perceptive criteria which allow to increase its effectiveness in application domains characterized by high variability in the shape of the visual patterns. The whole approach is general enough to be satisfactorily used in a wide class of applications. The metric has been validated by employing it in a nearest neighbour classifier, which has been used for automatic recognition of handwritten digits extracted from a standard character database.
Claudio De Stefano, Pasquale Foggia, Francesco Tortorella, Mario Vento
ICPR4
1995 A neural network classifier for OCR using structural descriptions
Luigi P. Cordella, Claudio De Stefano, Mario Vento
Mach. Vis. Appl.3
1995 An entropy based method for extracting robust binary-templates
Claudio De Stefano, Francesco Tortorella, Mario Vento
Mach. Vis. Appl.3
1995 A method for improving classification reliability of multilayer perceptrons
abstract
Criteria for evaluating the classification reliability of a neural classifier and for accordingly making a reject option are proposed. Such an option, implemented by means of two rules which can be applied independently of topology, size, and training algorithms of the neural classifier, allows one to improve the classification reliability. It is assumed that a performance function P is defined which, taking into account the requirements of the particular application, evaluates the quality of the classification in terms of recognition, misclassification, and reject rates. Under this assumption the optimal reject threshold value, determining the best trade-off between reject rate and misclassification rate, is the one for which the function P reaches its absolute maximum. No constraints are imposed on the form of P, but the ones necessary in order that P actually measures the quality of the classification process. The reject threshold is evaluated on the basis of some statistical distributions characterizing the behavior of the classifier when operating without reject option; these distributions are computed once the training phase of the net has been completed. The method has been tested with a neural classifier devised for handprinted and multifont printed characters, by using a database of about 300000 samples. Experimental results are discussed.
Luigi P. Cordella, Claudio De Stefano, Francesco Tortorella, Mario Vento
IEEE Trans. Neural Networks4
1994 Can a sequential thinning algorithm be parallelized?
abstract
It is commonly presumed that only parallel thinning algorithms can be efficiently implemented on a parallel machine. In this paper it is shown that also a sequential thinning algorithm can have parallel features which can be made explicit and successfully used for a parallel implementation. To this end, the main phases of a fully sequential algorithm are reformulated in such a way that each phase can be carried out by using parallel operators. Experimental results, obtained on a general purpose SIMD machine, are finally discussed.
Antonio d'Acierno, Claudio De Stefano, Francesco Tortorella, Mario Vento
ICPR (3)4
1993 Using entropy for drawing reliable templates
abstract
The method presented allows the drawing of templates that are reliable in locating distorted occurrences of the symbol to recognize and robust against the noise and the false alarms present on the background. A learning by showing technique is used to draw the template by considering samples of the symbol coming from a training set. The design of the template involves the evaluation of the reliablity of each pixel, obtained by estimating the entropy which characterizes the pixel. To be considered reliable, a pixel must show an entropy less than a threshold H/sub o/ set so as to maximize a match performance figure: the obtained template is constituted only by the pixels meeting this requirement.>
Claudio De Stefano, Francesco Tortorella, Mario Vento
ICDAR3
1992 Improving character recognition rate by a multi-net neural classifier
abstract
A neural classifier for isolated omnifont characters is discussed. A method for characterizing a given training set of characters, based on the definition of some statistical parameters is introduced; on the basis of such characterization an architecture is defined made of a set of neural networks properly connected. Depending on the value of the parameters characterizing the training set, both sizing and training of each network are separately carried out according to a suitable methodology. It is shown that higher recognition rates can be achieved than those obtained by using a single neural network as classifier.>
Luigi P. Cordella, Claudio De Stefano, Francesco Tortorella, Mario Vento
ICPR (2)4
1992 A method for the recognition of symbols on geographic maps
abstract
Presents a method for the recognition of symbols on binary images, especially tailored for geographic maps. The prototyping stage is performed by means of a geometric approach: using a suitable distance definition, the prototype of a class is obtained as the pattern minimizing the sum of distances among itself and all the samples of the considered class. In this phase some statistical parameters characterizing the distortions occurring on the symbols are also obtained, so allowing the estimate of the recognition reliability. On the basis of the misclassification and reject costs, some tuning parameters, affecting the classification criteria, are also evaluated in order to maximize the classification performances. Experimental results for several test maps are finally presented.>
Claudio De Stefano, Francesco Tortorella, Mario Vento
ICPR (1)3
1988 A preliminary approach to the design and evaluation of a reconfigurable architecture for computer vision
abstract
An analytical model describing the performance of a mesh-connected array of state-of-the-art VLSI microprocessors embedded in a CSP-like environment, is presented. The mode has been developed in the framework of a study aimed at the design and evaluation of a reconfigurable architecture especially suitable for the execution of both low and high level vision tasks. The validity of the model is demonstrated and its use illustrated.>
Angelo Chianese, Luigi P. Cordella, Massimo De Santo, Angelo Marcelli, Mario Vento
ICPR5