EDBT 2026 Demo / reviewers in the wild / expert
Danilo Avola
dblp:91/3593
· DBLP profile ↗
44ranked-venue papers
37as first author
19since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 17 first-author · 7 since 2021Artificial intelligence and machine learning · 18 · 13 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-authorSecurity and privacy · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SATEER: Subject-Aware Transformer for EEG-Based Emotion RecognitionabstractThis study presents a Subject-Aware Transformer-based neural network designed for the Electroencephalogram (EEG) Emotion Recognition task (SATEER), which entails the analysis of EEG signals to classify and interpret human emotional states. SATEER processes the EEG waveforms by transforming them into Mel spectrograms, which can be seen as particular cases of images with the number of channels equal to the number of electrodes used during the recording process; this type of data can thus be processed using a Computer Vision pipeline. Distinct from preceding approaches, this model addresses the variability in individual responses to identical stimuli by incorporating a User Embedder module. This module enables the association of individual profiles with their EEGs, thereby enhancing classification accuracy. The efficacy of the model was rigorously evaluated using four publicly available datasets, demonstrating superior performance over existing methods in all conducted benchmarks. For instance, on the AMIGOS dataset (A dataset for Multimodal research of affect, personality traits, and mood on Individuals and GrOupS), SATEER's accuracy exceeds 99.8% accuracy across all labels and showcases an improvement of 0.47% over the state of the art. Furthermore, an exhaustive ablation study underscores the pivotal role of the User Embedder module and each other component of the presented model in achieving these advancements. Romeo Lanzino, Danilo Avola, Federico Fontana, Luigi Cinque, Francesco Scarcello, Gian Luca Foresti |
Int. J. Neural Syst. | 2 |
| 2024 | Semantically Guided Representation Learning For Action Anticipation
Anxhelo Diko, Danilo Avola, Bardh Prenkaj, Federico Fontana, Luigi Cinque |
ECCV (28) | 2 |
| 2024 | FaceVision-GAN: A 3D Model Face Reconstruction Method from a Single Image Using GANsabstractGenerative algorithms have been very successful in recent years. This phenomenon derives from the strong computational power that even consumer computers can provide. Moreover, a huge amount of data is available today for feeding deep learning algorithms. In this context, human 3D face mesh reconstruction is becoming an important but challenging topic in computer vision and computer graphics. It could be exploited in different application areas, from security to avatarization. This paper provides a 3D face reconstruction pipeline based on Generative Adversarial Networks (GANs). It can generate high-quality depth and correspondence maps from 2D images, which are exploited for producing a 3D model of the subject’s face. Danilo Avola, Luigi Cinque, Gian Luca Foresti, Marco Raoul Marini |
ICPRAM | 1 |
| 2024 | Signal enhancement and efficient DTW-based comparison for wearable gait recognitionabstractThe popularity of biometrics-based user identification has significantly increased over the last few years. User identification based on the face, fingerprints, and iris, usually achieves very high accuracy only in controlled setups and can be vulnerable to presentation attacks, spoofing, and forgeries. To overcome these issues, this work proposes a novel strategy based on a relatively less explored biometric trait, i.e., gait, collected by a smartphone accelerometer, which can be more robust to the attacks mentioned above. According to the wearable sensor-based gait recognition state-of-the-art, two main classes of approaches exist: 1) those based on machine and deep learning; 2) those exploiting hand-crafted features. While the former approaches can reach a higher accuracy, they suffer from problems like, e.g., performing poorly outside the training data, i.e., lack of generalizability. This paper proposes an algorithm based on hand-crafted features for gait recognition that can outperform the existing machine and deep learning approaches. It leverages a modified Majority Voting scheme applied to Fast Window Dynamic Time Warping, a modified version of the Dynamic Time Warping (DTW) algorithm with relaxed constraints and majority voting, to recognize gait patterns. We tested our approach named MV-FWDTW on the ZJU-gaitacc, one of the most extensive datasets for the number of subjects, but especially for the number of walks per subject and walk lengths. Results set a new state-of-the-art gait recognition rate of 98.82% in a cross-session experimental setup. We also confirm the quality of the proposed method using a subset of the OU-ISIR dataset, another large state-of-the-art benchmark with more subjects but much shorter walk signals. Danilo Avola, Luigi Cinque, Maria De Marsico, Alessio Fagioli 0001, Gian Luca Foresti, Maurizio Mancini, Alessio Mecca |
Comput. Secur. | 1 |
| 2024 | Spatio-Temporal Image-Based Encoded Atlases for EEG Emotion RecognitionabstractEmotion recognition plays an essential role in human-human interaction since it is a key to understanding the emotional states and reactions of human beings when they are subject to events and engagements in everyday life. Moving towards human-computer interaction, the study of emotions becomes fundamental because it is at the basis of the design of advanced systems to support a broad spectrum of application areas, including forensic, rehabilitative, educational, and many others. An effective method for discriminating emotions is based on ElectroEncephaloGraphy (EEG) data analysis, which is used as input for classification systems. Collecting brain signals on several channels and for a wide range of emotions produces cumbersome datasets that are hard to manage, transmit, and use in varied applications. In this context, the paper introduces the Empátheia system, which explores a different EEG representation by encoding EEG signals into images prior to their classification. In particular, the proposed system extracts spatio-temporal image encodings, or atlases, from EEG data through the Processing and transfeR of Interaction States and Mappings through Image-based eNcoding (PRISMIN) framework, thus obtaining a compact representation of the input signals. The atlases are then classified through the Empátheia architecture, which comprises branches based on convolutional, recurrent, and transformer models designed and tuned to capture the spatial and temporal aspects of emotions. Extensive experiments were conducted on the Shanghai Jiao Tong University (SJTU) Emotion EEG Dataset (SEED) public dataset, where the proposed system significantly reduced its size while retaining high performance. The results obtained highlight the effectiveness of the proposed approach and suggest new avenues for data representation in emotion recognition from EEG signals. Danilo Avola, Luigi Cinque, Angelo Di Mambro, Alessio Fagioli 0001, Marco Raoul Marini, Daniele Pannone, Bruno Fanini, Gian Luca Foresti |
Int. J. Neural Syst. | 1 |
| 2024 | ReViT: Enhancing vision transformers feature diversity with attention residual connections
Anxhelo Diko, Danilo Avola, Marco Cascio, Luigi Cinque |
Pattern Recognit. | 2 |
| 2023 | Keyrtual: A Lightweight Virtual Musical Keyboard Based on RGB-D and Sensors Fusion
Danilo Avola, Luigi Cinque, Marco Raoul Marini, Andrea Princic, Valerio Venanzi |
CAIP (2) | 1 |
| 2022 | Human Silhouette and Skeleton Video Synthesis Through Wi-Fi SignalsabstractThe increasing availability of wireless access points (APs) is leading toward human sensing applications based on Wi-Fi signals as support or alternative tools to the widespread visual sensors, where the signals enable to address well-known vision-related problems such as illumination changes or occlusions. Indeed, using image synthesis techniques to translate radio frequencies to the visible spectrum can become essential to obtain otherwise unavailable visual data. This domain-to-domain translation is feasible because both objects and people affect electromagnetic waves, causing radio and optical frequencies variations. In the literature, models capable of inferring radio-to-visual features mappings have gained momentum in the last few years since frequency changes can be observed in the radio domain through the channel state information (CSI) of Wi-Fi APs, enabling signal-based feature extraction, e.g. amplitude. On this account, this paper presents a novel two-branch generative neural network that effectively maps radio data into visual features, following a teacher-student design that exploits a cross-modality supervision strategy. The latter conditions signal-based features in the visual domain to completely replace visual data. Once trained, the proposed method synthesizes human silhouette and skeleton videos using exclusively Wi-Fi signals. The approach is evaluated on publicly available data, where it obtains remarkable results for both silhouette and skeleton videos generation, demonstrating the effectiveness of the proposed cross-modality supervision strategy. Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti |
Int. J. Neural Syst. | 1 |
| 2022 | Affective Action and Interaction Recognition by Multi-View Representation Learning from Handcrafted Low-Level Skeleton FeaturesabstractHuman feelings expressed through verbal (e.g. voice) and non-verbal communication channels (e.g. face or body) can influence either human actions or interactions. In the literature, most of the attention was given to facial expressions for the analysis of emotions conveyed through non-verbal behaviors. Despite this, psychology highlights that the body is an important indicator of the human affective state in performing daily life activities. Therefore, this paper presents a novel method for affective action and interaction recognition from videos, exploiting multi-view representation learning and only full-body handcrafted characteristics selected following psychological and proxemic studies. Specifically, 2D skeletal data are extracted from RGB video sequences to derive diverse low-level skeleton features, i.e. multi-views, modeled through the bag-of-visual-words clustering approach generating a condition-related codebook. In this way, each affective action and interaction within a video can be represented as a frequency histogram of codewords. During the learning phase, for each affective class, training samples are used to compute its global histogram of codewords stored in a database and later used for the recognition task. In the recognition phase, the video frequency histogram representation is matched against the database of class histograms and classified as the closest affective class in terms of Euclidean distance. The effectiveness of the proposed system is evaluated on a specifically collected dataset containing 6 emotion for both actions and interactions, on which the proposed system obtains 93.64% and 90.83% accuracy, respectively. In addition, the devised strategy also achieves in line performances with other literature works based on deep learning when tested on a public collection containing 6 emotions plus a neutral state, demonstrating the effectiveness of the presented approach and confirming the findings in psychological and proxemic studies. Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti |
Int. J. Neural Syst. | 1 |
| 2022 | SIRe-Networks: Convolutional neural networks architectural extension for information preservation via skip/residual connections and interlaced auto-encoders
Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti |
Neural Networks | 1 |
| 2022 | 3D hand pose and shape estimation from RGB images for keypoint-based hand gesture recognition
Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Adriano Fragomeni, Daniele Pannone |
Pattern Recognit. | 1 |
| 2022 | Deep Temporal Analysis for Non-Acted Body Affect RecognitionabstractIn the field of body affect recognition, the majority of literature is based on experiments performed on datasets where trained actors simulate emotional reactions. These acted and unnatural expressions differ from the more challenging genuine emotions, thus leading to less valuable results. In this article, a solution for basic non-acted emotion recognition based on 3D skeleton and Deep Neural Networks (DNNs) is provided. The proposed work introduces three majors contributions. First, temporal local movements performed by subjects are examined frame-by-frame, unlike the current state-of-the-art in non-acted body affect recognition where only static or global body features are considered. Second, an original set of global and time-dependent features for body movement description is provided. Third, this is one of the first works to use deep learning methods in the current non-acted body affect recognition literature. Due to the novelty of the topic, only the UCLIC dataset is currently considered the benchmark for comparative tests. On the latter, the proposed method outperforms all the competitors. Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Cristiano Massaroni |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Multimodal Feature Fusion and Knowledge-Driven Learning via Experts Consult for Thyroid Nodule ClassificationabstractComputer-aided diagnosis (CAD) is becoming a prominent approach to assist clinicians spanning across multiple fields. These automated systems take advantage of various computer vision (CV) procedures, as well as artificial intelligence (AI) techniques, to formulate a diagnosis of a given image, e.g., computed tomography and ultrasound. Advances in both areas (CV and AI) are enabling ever increasing performances of CAD systems, which can ultimately avoid performing invasive procedures such as fine-needle aspiration. In this study, a novel end-to-end knowledge-driven classification framework is presented. The system focuses on multimodal data generated by thyroid ultrasonography, and acts as a CAD system by providing a thyroid nodule classification into the benign and malignant categories. Specifically, the proposed system leverages cues provided by an ensemble of experts to guide the learning phase of a densely connected convolutional network (DenseNet). The ensemble is composed by various networks pretrained on ImageNet, including AlexNet, ResNet, VGG, and others. The previously computed multimodal feature parameters are used to create ultrasonography domain experts via transfer learning, decreasing, moreover, the number of samples required for training. To validate the proposed method, extensive experiments were performed, providing detailed performances for both the experts ensemble and the knowledge-driven DenseNet. As demonstrated by the results, the proposed system achieves relevant performances in terms of qualitative metrics for the thyroid nodule classification task, thus resulting in a great asset when formulating a diagnosis. Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Sebastiano Filetti, Giorgio Grani, Emanuele Rodolà |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Person Re-Identification Through Wi-Fi Extracted Radio Biometric SignaturesabstractPerson re-identification (Re-ID) is a challenging task that tries to recognize a person across different cameras, and that can prove useful in video surveillance as well as in forensics and security applications. However, traditional Re-ID systems analyzing image or video sequences suffer from well-known issues such as illumination changes, occlusions, background clutter, and long-term re-identification. To simultaneously address all these difficult problems, we explore a Re-ID solution based on an alternative medium that is inherently not affected by them, i.e., the Wi-Fi technology. The latter, due to the widespread use of wireless communications, has grown rapidly and is already enabling the development of Wi-Fi sensing applications, such as human localization or counting. These sensing procedures generally exploit Wi-Fi signals variations that are a direct consequence, among other things, of human presence, and which can be observed through the channel state information (CSI) of Wi-Fi access points. Following this rationale, in this paper, for the first time in literature, we show how the pervasive Wi-Fi technology can also be directly exploited for person Re-ID. More accurately, Wi-Fi signals amplitude and phase are extracted from CSI measurements and analyzed through a two-branch deep neural network working in a siamese-like fashion. The designed pipeline can extract meaningful features from signals, i.e., radio biometric signatures, that ultimately allow the person Re-ID. The effectiveness of the proposed system is evaluated on a specifically collected dataset, where remarkable performances are obtained; suggesting that Wi-Fi signal variations differ between different people and can consequently be used for their re-identification. Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Chiara Petrioli |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2021 | LieToMe: An Ensemble Approach for Deception Detection from Facial CuesabstractDeception detection is a relevant ability in high stakes situations such as police interrogatories or court trials, where the outcome is highly influenced by the interviewed person behavior. With the use of specific devices, e.g. polygraph or magnetic resonance, the subject is aware of being monitored and can change his behavior, thus compromising the interrogation result. For this reason, video analysis-based methods for automatic deception detection are receiving ever increasing interest. In this paper, a deception detection approach based on RGB videos, leveraging both facial features and stacked generalization ensemble, is proposed. First, a face, which is well-known to present several meaningful cues for deception detection, is identified, aligned, and masked to build video signatures. These signatures are constructed starting from five different descriptors, which allow the system to capture both static and dynamic facial characteristics. Then, video signatures are given as input to four base-level algorithms, which are subsequently fused applying the stacked generalization technique, resulting in a more robust meta-level classifier used to predict deception. By exploiting relevant cues via specific features, the proposed system achieves improved performances on a public dataset of famous court trials, with respect to other state-of-the-art methods based on facial features, highlighting the effectiveness of the proposed method. Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti |
Int. J. Neural Syst. | 1 |
| 2021 | Forward-looking sonar image compression by integrating keypoint clustering and morphological skeleton
Danilo Avola, Marco Bernardi, Luigi Cinque, Gian Luca Foresti, Daniele Pannone, Chiara Petrioli |
Multim. Tools Appl. | 1 |
| 2021 | Automatic estimation of optimal UAV flight parameters for real-time wide areas monitoring
Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Daniele Pannone, Claudio Piciarelli |
Multim. Tools Appl. | 1 |
| 2021 | Data integration by two-sensors in a LEAP-based Virtual Glove for human-system interactionabstractAbstract Virtual Glove (VG) is a low-cost computer vision system that utilizes two orthogonal LEAP motion sensors to provide detailed 4D hand tracking in real–time. VG can find many applications in the field of human-system interaction, such as remote control of machines or tele-rehabilitation. An innovative and efficient data-integration strategy, based on the velocity calculation, for selecting data from one of the LEAPs at each time, is proposed for VG. The position of each joint of the hand model, when obscured to a LEAP, is guessed and tends to flicker. Since VG uses two LEAP sensors, two spatial representations are available each moment for each joint: the method consists of the selection of the one with the lower velocity at each time instant. Choosing the smoother trajectory leads to VG stabilization and precision optimization, reduces occlusions (parts of the hand or handling objects obscuring other hand parts) and/or, when both sensors are seeing the same joint, reduces the number of outliers produced by hardware instabilities. The strategy is experimentally evaluated, in terms of reduction of outliers with respect to a previously used data selection strategy on VG, and results are reported and discussed. In the future, an objective test set has to be imagined, designed, and realized, also with the help of an external precise positioning equipment, to allow also quantitative and objective evaluation of the gain in precision and, maybe, of the intrinsic limitations of the proposed strategy. Moreover, advanced Artificial Intelligence-based (AI-based) real-time data integration strategies, specific for VG, will be designed and tested on the resulting dataset. Giuseppe Placidi, Danilo Avola, Luigi Cinque, Matteo Polsinelli, Eleni Theodoridou, João Manuel R. S. Tavares |
Multim. Tools Appl. | 2 |
| 2021 | R-SigNet: Reduced space writer-independent feature learning for offline writer-dependent signature verification
Danilo Avola, Manoochehr Joodi Bigdello, Luigi Cinque, Alessio Fagioli 0001, Marco Raoul Marini |
Pattern Recognit. Lett. | 1 |
| 2020 | Fusing Self-Organized Neural Network and Keypoint Clustering for Localized Real-Time Background SubtractionabstractMoving object detection in video streams plays a key role in many computer vision applications. In particular, separation between background and foreground items represents a main prerequisite to carry out more complex tasks, such as object classification, vehicle tracking, and person re-identification. Despite the progress made in recent years, a main challenge of moving object detection still regards the management of dynamic aspects, including bootstrapping and illumination changes. In addition, the recent widespread of Pan-Tilt-Zoom (PTZ) cameras has made the management of these aspects even more complex in terms of performance due to their mixed movements (i.e. pan, tilt, and zoom). In this paper, a combined keypoint clustering and neural background subtraction method, based on Self-Organized Neural Network (SONN), for real-time moving object detection in video sequences acquired by PTZ cameras is proposed. Initially, the method performs a spatio-temporal tracking of the sets of moving keypoints to recognize the foreground areas and to establish the background. Then, it adopts a neural background subtraction, localized in these areas, to accomplish a foreground detection able to manage bootstrapping and gradual illumination changes. Experimental results on three well-known public datasets, and comparisons with different key works of the current literature, show the efficiency of the proposed method in terms of modeling and background subtraction. Danilo Avola, Marco Bernardi, Luigi Cinque, Cristiano Massaroni, Gian Luca Foresti |
Int. J. Neural Syst. | 1 |
| 2020 | Online separation of handwriting from freehand drawing using extreme learning machines
Danilo Avola, Marco Bernardi, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni |
Multim. Tools Appl. | 1 |
| 2020 | MIFTel: a multimodal interactive framework based on temporal logic rules
Danilo Avola, Luigi Cinque, Alberto Del Bimbo, Marco Raoul Marini |
Multim. Tools Appl. | 1 |
| 2020 | Homography vs similarity transformation in aerial mosaicking: which is the best at different altitudes?
Danilo Avola, Luigi Cinque, Gian Luca Foresti, Daniele Pannone |
Multim. Tools Appl. | 1 |
| 2020 | LieToMe: Preliminary study on hand gestures for deception detection via Fisher-LSTM
Danilo Avola, Luigi Cinque, Maria De Marsico, Alessio Fagioli 0001, Gian Luca Foresti |
Pattern Recognit. Lett. | 1 |
| 2020 | 2-D Skeleton-Based Action Recognition via Two-Branch Stacked LSTM-RNNsabstractAction recognition in video sequences is an interesting field for many computer vision applications, including behavior analysis, event recognition, and video surveillance. In this article, a method based on 2D skeleton and two-branch stacked Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM) cells is proposed. Unlike 3D skeletons, usually generated by RGB-D cameras, the 2D skeletons adopted in this article are reconstructed starting from RGB video streams, therefore allowing the use of the proposed approach in both indoor and outdoor environments. Moreover, any case of missing skeletal data is managed by exploiting 3D-Convolutional Neural Networks (3D-CNNs). Comparative experiments with several key works on KTH and Weizmann datasets show that the method described in this paper outperforms the current state-of-the-art. Additional experiments on UCF Sports and IXMAS datasets demonstrate the effectiveness of our method in the presence of noisy data and perspective changes, respectively. Further investigations on UCF Sports, HMDB51, UCF101, and Kinetics400 highlight how the combination between the proposed two-branch stacked LSTM and the 3D-CNN-based network can manage missing skeleton information, greatly improving the overall accuracy. Moreover, additional tests on KTH and UCF Sports datasets also show the robustness of our approach in the presence of partial body occlusions. Finally, comparisons on UT-Kinect and NTU-RGB+D datasets show that the accuracy of the proposed method is fully comparable to that of works based on 3D skeletons. Danilo Avola, Marco Cascio, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni, Emanuele Rodolà |
IEEE Trans. Multim. | 1 |
| 2020 | A UAV Video Dataset for Mosaicking and Change Detection From Low-Altitude FlightsabstractIn recent years, the technology of small-scale unmanned aerial vehicles (UAVs) has steadily improved in terms of flight time, automatic control, and image acquisition. This has lead to the development of several applications for low-altitude tasks, such as vehicle tracking, person identification, and object recognition. These applications often require to stitch together several video frames to get a comprehensive view of large areas (mosaicking), or to detect differences between images or mosaics acquired at different times (change detection). However, the datasets used to test mosaicking and change detection algorithms are typically acquired at high-altitudes, thus ignoring the specific challenges of low-altitude scenarios. The purpose of this paper is to fill this gap by providing the UAV mosaicking and change detection dataset. It consists of 50 challenging aerial video sequences acquired at low-altitude in different environments with and without the presence of vehicles, persons, and objects, plus metadata and telemetry. In addition, this paper provides some performance metrics to evaluate both the quality of the obtained mosaics and the correctness of the detected changes. Finally, the results achieved by two baseline algorithms, one for mosaicking and one for detection, are presented. The aim is to provide a shared performance reference that can be used for comparison with future algorithms that will be tested on the dataset. Danilo Avola, Luigi Cinque, Gian Luca Foresti, Niki Martinel, Daniele Pannone, Claudio Piciarelli |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2019 | Master and Rookie Networks for Person Re-identification
Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Cristiano Massaroni |
CAIP (2) | 1 |
| 2019 | A Shape Comparison Reinforcement Method Based on Feature Extractors and F1-ScoreabstractEvaluating object segmentation is a topic of great interest for shape comparison techniques. In this work, ad-hoc metrics for a detailed segmentation analysis and a novel keypoint based method for comparing pairs of shapes are presented. As references, two different segmentation approaches were used: a handmade segmentation and an automatic one based on a Convolutional Neural Network (CNN). The proposed comparison approach consists of a combination between a keypoint extractor and an invariant scale shape identifier. The overall validation process is established according to different steps, which allow to measure the similarity between shapes. First, Reinforced Matched (RM) and Reinforced Ratio (RR) strategies are implemented. Moreover, five different state-of-the-art keypoint extractors are compared, i.e., SIFT, SURF, ORB, A-KAZE, and BRISK. Experimental tests were performed on a popular collection of images, i.e., the Berkeley Segmentation Dataset and Benchmark 300 (BSDS300), which contains shapes segmented both manually and automatically. The experimental results have shown the effectiveness of the proposed method. Danilo Avola, Luigi Cinque, Gian Luca Foresti, Francesco Lamacchia, Marco Raoul Marini, Luca Perini, Kristjana Qorraj, Gabriele Telesca |
SMC | 1 |
| 2019 | An interactive and low-cost full body rehabilitation framework based on 3D immersive serious games
Danilo Avola, Luigi Cinque, Gian Luca Foresti, Marco Raoul Marini |
J. Biomed. Informatics | 1 |
| 2019 | Fusing depth and colour information for human action recognition
Danilo Avola, Marco Bernardi, Gian Luca Foresti |
Multim. Tools Appl. | 1 |
| 2019 | A Vision-Based System for Internal Pipeline InspectionabstractThe internal inspection of large pipeline infrastructures, such as sewers and waterworks, is a fundamental task for the prevention of possible failures. In particular, visual inspection is typically performed by human operators on the basis of video sequences either acquired on-line or recorded for further off-line analysis. In this work, we propose a vision-based software approach to assist the human operator by conveniently showing the acquired data and by automatically detecting and highlighting the pipeline sections where relevant anomalies could occur. Claudio Piciarelli, Danilo Avola, Daniele Pannone, Gian Luca Foresti |
IEEE Trans. Ind. Informatics | 2 |
| 2019 | Exploiting Recurrent Neural Networks and Leap Motion Controller for the Recognition of Sign Language and Semaphoric Hand GesturesabstractHand gesture recognition is still a topic of great interest for the computer vision community. In particular, sign language and semaphoric hand gestures are two foremost areas of interest due to their importance in human-human communication and human-computer interaction, respectively. Any hand gesture can be represented by sets of feature vectors that change over time. Recurrent neural networks (RNNs) are suited to analyze this type of set thanks to their ability to model the long-term contextual information of temporal sequences. In this paper, an RNN is trained by using as features the angles formed by the finger bones of the human hands. The selected features, acquired by a leap motion controller sensor, are chosen because the majority of human hand gestures produce joint movements that generate truly characteristic corners. The proposed method, including the effectiveness of the selected angles, was initially tested by creating a very challenging dataset composed by a large number of gestures defined by the American sign language. On the latter, an accuracy of over 96% was achieved. Afterwards, by using the Shape Retrieval Contest (SHREC) dataset, a wide collection of semaphoric hand gestures, the method was also proven to outperform in accuracy competing approaches of the current literature. Danilo Avola, Marco Bernardi, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni |
IEEE Trans. Multim. | 1 |
| 2018 | Combining Keypoint Clustering and Neural Background Subtraction for Real-time Moving Object Detection by PTZ CamerasabstractDetection of moving objects is a topic of great interest in computer vision. This task represents a prerequisite for more complex duties, such as classification and re-identification. One of the main challenges regards the management of dynamic factors, with particular reference to bootstrapping and illumination change issues. The recent widespread of PTZ cameras has made these issues even more complex in terms of performance due to their composite movements (i.e., pan, tilt, and zoom). This paper proposes a combined keypoint clustering and neural background subtraction method for real-time moving object detection in video sequences acquired by PTZ cameras. Initially, the method performs a spatio-temporal tracking of the sets of moving keypoints to recognize the foreground areas and to establish the background. Subsequently, it adopts a neural background subtraction to accomplish a foreground detection, in these areas, able to manage bootstrapping and gradual illumination changes. Experimental results on two well-known public datasets and comparisons with different key works of the current state-of-the-art demonstrate the remarkable results of the proposed method. Danilo Avola, Marco Bernardi, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni |
ICPRAM | 1 |
| 2018 | A Rover-based System for Searching Encrypted Targets in Unknown EnvironmentsabstractIn the last decade, there has been a widespread use of autonomous robots in several application fields, such as border controls, precision agriculture, and military operations. Usually, in the latter, there is the need to encrypt the acquired data, or to mark as relevant some positions or areas. In this paper, we present a client-server rover-based system able to search encrypted targets within an unknown environment. The system uses a rover to explore an unknown environment through a Simultaneous Localization And Mapping (SLAM) algorithm and acquires the scene with a standard RGB camera. Then, by using visual cryptography, it is possible to encrypt the acquired RGB data and to send it to a server, which decrypts the data and checks if it contains a target object. The experiments performed on several objects show the effectiveness of the proposed system. Danilo Avola, Luigi Cinque, Gian Luca Foresti, Marco Raoul Marini, Daniele Pannone |
ICPRAM | 1 |
| 2018 | VRheab: a fully immersive motor rehabilitation system based on recurrent neural network
Danilo Avola, Luigi Cinque, Gian Luca Foresti, Marco Raoul Marini, Daniele Pannone |
Multim. Tools Appl. | 1 |
| 2017 | Aerial video surveillance system for small-scale UAV environment monitoringabstractChange detection algorithms are commonly used to detect novelties for surveillance purposes in public and private places equipped by static or Pan-Tilt-Zoom (PTZ) cameras. Often, these techniques are also used as prerequisite to support more complex algorithms, including event recognition, object classification, person re-identification, and many others. With regard to small-scale Unmanned Aerial Vehicles (UAVs) at low-altitude, the change detection techniques require further investigation. In fact, most of the works currently available in the literature process video sequences acquired at very high-altitude for large-scale operations, such as vegetation monitoring, mapping of buildings, and so on. In a wide range of application contexts that require, for example, frequent monitoring or high spatial resolution for detecting small objects, video sequences acquired at high-altitude are not suitable. This paper presents a change detection system based on histogram equalization and RGB-Local Binary Pattern (RGB-LBP) operator for monitoring of wide areas by small-scale UAVs at low-altitude. Extensive experimental results show the robustness of the proposed pipeline. These latter were performed by using challenging video sequences of the public UAV Mosaicking and Change Detection (UMCD) dataset and measured a set of well-known statistical metrics. Finally, a performance analysis of the proposed algorithm is also provided. Danilo Avola, Gian Luca Foresti, Niki Martinel, Christian Micheloni, Daniele Pannone, Claudio Piciarelli |
AVSS | 1 |
| 2017 | Adaptive bootstrapping management by keypoint clustering for background initialization
Danilo Avola, Marco Bernardi, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni |
Pattern Recognit. Lett. | 1 |
| 2017 | A keypoint-based method for background modeling and foreground detection using a PTZ camera
Danilo Avola, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni, Daniele Pannone |
Pattern Recognit. Lett. | 1 |
| 2016 | A Practical Framework for the Development of Augmented Reality Applications by using ArUco MarkersabstractThe Augmented Reality (AR) is an expanding field of the Computer Graphics (CG) that merges items of the real-world environment (e.g., places, objects) with digital information (e.g., multimedia files, virtual objects) to provide users with an enhanced interactive multi-sensorial experience of the real-world that surrounding them. Currently, a wide range of devices is used to vehicular AR systems. Common devices (e.g., cameras equipped on smartphones) enable users to receive multimedia information about target objects (non-immersive AR). Advanced devices (e.g., virtual windscreens) provide users with a set of virtual information about points of interest (POIs) or places (semi-immersive AR). Finally, an ever-increasing number of new devices (e.g., HeadMounted Display, HMD) support users to interact with mixed reality environments (immersive AR). This paper presents a practical framework for the development of non-immersive augmented reality applications through which target objects are enriched with multimedia information. On each target object is applied a different ArUco marker. When a specific application hosted inside a device recognizes, via camera, one of these markers, then the related multimedia information are loaded and added to the target object. The paper also reports a complete case study together with some considerations on the framework and future work. Danilo Avola, Luigi Cinque, Gian Luca Foresti, Cristina Mercuri, Daniele Pannone |
ICPRAM | 1 |
| 2016 | A multipurpose autonomous robot for target recognition in unknown environmentsabstractIn recent years, the technological improvements of consumer robots, in terms of processing capacity and sensors, are enabling an ever-increasing number of researchers to quickly develop both scale prototypes and alternative low cost solutions. In these contexts, a critical aspect is the design of ad-hoc algorithms according to the features of the available hardware. This paper proposes a prototype of an autonomous robot for mapping unknown environments and recognizing target objects. During the setup phase one or more target objects are shown to the RGB camera of the robot which, for each of them, extracts and stores a set of A-KAZE features. Afterwards, the robot adopts the ultrasonic distance measurement and the RGB stream to map the whole environment and search a set of A-KAZE features matchable with those previously acquired. The paper also reports both preliminary tests carried out on a reference indoor environment and a case study performed in an outdoor one that validate the proposed system. Danilo Avola, Gian Luca Foresti, Luigi Cinque, Cristiano Massaroni, Gabriele Vitale, Luca Lombardi |
INDIN | 1 |
| 2015 | Basis for the implementation of an EEG-based single-trial binary brain computer interface through the disgust produced by remembering unpleasant odors
Giuseppe Placidi, Danilo Avola, Andrea Petracca, Fiorella Sgallari, Matteo Spezialetti |
Neurocomputing | 2 |
| 2010 | Interacting annotations in MADCOW 2.0abstractMADCOW 2.0 is a system for annotation of Web content, supporting the production and exploration of personal and public annotations on text, images and videos in a Web page. Its design starts from the main requirement that the annotation activity does not have to disrupt the normal browsing of Web pages by a user. MADCOW 2.0 allows interaction with the annotated portions of the page to provide access to the annotation content. Conversely, the representation of the existing notes supports different forms of exploration of the Web page, and can become the starting point for further navigation over the Web. A uniform style of interaction has been adopted for creating and accessing annotations on text, images and videos, and some novel solutions have been introduced to cope with overlaps between the annotated portions. The annotation user experience is facilitated by enabling forms of in-place annotation and manipulation of both the annotated portion and the annotation content. Danilo Avola, Paolo Bottoni, Stefano Levialdi, Emanuele Panizzi |
AVI | 1 |
| 2008 | A Novel Approach for Practical Semantic Web Data Management
Giorgio Gianforme, Roberto De Virgilio, Stefano Paolozzi, Pierluigi Del Nostro, Danilo Avola |
KES (2) | 5 |
| 2007 | Sketch Style Recognition in Human Computer Interaction
Danilo Avola, Fernando Ferri, Patrizia Grifoni |
SEKE | 1 |