EDBT 2026 Demo / reviewers in the wild / expert
Luigi Cinque
dblp:64/4176
· DBLP profile ↗
97ranked-venue papers
30as first author
29since 2021 · last 2026
0000-0001-9149-2175ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 20 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 14 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 8 · 4 since 2021Databases, data management, data science and information retrieval · 7 · 7 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Systems, architecture and hardware · 2Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Thyroid Nodule Classification via Weak Self-Supervision and Transfer Learning
Alessio Fagioli 0001, Marco Cascio, Gian Luca Foresti, Luigi Cinque |
ICPRAM | 4 |
| 2026 | MusiKeyrtual: A Framework to Play a Musical Keyboard in Augmented RealityabstractThe advances in Machine and Deep Learning (ML and DL, respectively) contribute to the evolution of modern Augmented Reality (AR) systems, adapting software to complex and interesting applications. Moreover, lower-end systems, such as smartphones, are now capable of running AR applications, albeit often requiring smaller and lighter DL and ML models to accommodate hardware limitations. In this context, we propose MusiKeyrtual, a lightweight application that allows users to play a musical keyboard drawn on paper; an improvement over the previous version, Keyrtual. The application requires only a smartphone to run. The pipeline proposed addresses the hardware limitations of smartphones, both in terms of limited computational capabilities and in terms of using a single RGB camera, which cannot detect depth. Quantitative and qualitative results highlight the effectiveness of the proposed pipeline in exploiting the capabilities of modern smartphones. Valerio Venanzi, Andrea Princic, Marco Raoul Marini, Gian Luca Foresti, Luigi Cinque |
Int. J. Hum. Comput. Interact. | 5 |
| 2026 | Data-related Ablation for Reinforcing Deep Learning in Explaining Complex PhenomenaabstractDeep Learning (DL) models excel at automatically learning intricate patterns within complex data, but their black box nature undermines human trust. To address this, current validation strategies typically focus on the model itself, modifying its architecture to assess the role and importance of the components. However, this model-centric view overlooks the critical learning substrate, which is represented by the data, implicitly assuming that it accurately represents the target phenomenon. This implicit trust in data means that evaluation may fail to detect whether high performance stems from exploiting biases or data quirks rather than learning relevant patterns. We present a novel data-related ablation as a complement to the traditional architectural ablation. Using this framework for Electroencephalography (EEG) signals of Emotional Recognition (ER) and Motor Execution (ME) as a case study, we show that seemingly high-accuracy models often rely heavily on process-irrelevant features, maintaining performance even when key information is eliminated. This shows that a standard, data-independent evaluation can be misleading about whether a model truly captured the intended process; the proposed approach helps distinguish robust learning from leaning on incidental characteristics. Therefore, incorporating data-related ablation is essential for developing reliable and generalizable DL models in fields that rely on data derived from complex and often not completely known phenomena. Romeo Lanzino, Luigi Cinque, Gian Luca Foresti, Giuseppe Placidi |
Int. J. Neural Syst. | 2 |
| 2025 | Hand Gesture Recognition Using MediaPipe Landmarks and Deep Learning NetworksabstractAdvanced Human Computer Interaction techniques are commonly used in multiple application areas, from entertainment to rehabilitation. In this context, this paper proposes a framework to recognize hand gestures using a limited number of landmarks from the video images. This hand gesture recognition system comprises an image processing module that extracts and processes the coordinates of 21 hand points called landmarks, and a deep neural network module that models and classifies the hand gestures. These landmarks are extracted automatically through MediaPipe software. The experiments were carried out over the IPN Hand dataset in an independent-user scenario using a Subject-Wise Cross Validation. They cover the use of different landmark-based formats, normalizations, lengths of the gesture representations, and number of landmarks used as inputs. The system obtains significantly better accuracy when using the raw coordinates of the 21 landmarks through 125 timesteps and a light Recurre nt Neural Network architecture (80.56 ± 1.19 %) or the hand anthropometric measures (82.20 ± 1.15 %) compared to using the speed of the hand landmarks through the gesture (72.93 ± 1.34 %). The proposed framework studied the effect of different landmark-based normalizations over the raw coordinates, obtaining an accuracy of 83.67 ± 1.12 % when using as reference the wrist landmark from each frame, and an accuracy of 84.66 ± 1.09 % when using as reference the wrist landmark from the first video frame of the current gesture. In addition, the proposed solution provided high recognition performance even when only using the coordinates from 6 (82.15 ± 1.16 %) or 4 (81.46 ± 1.17 %) specific hand landmarks using as reference the wrist landmark from the first video frame of the current gesture. Manuel Gil-Martín, Marco Raoul Marini, Iván Martín-Fernández, Sergio Esteban Romero, Luigi Cinque |
ICAART (3) | 5 |
| 2025 | An Optimized and Accelerated Object Instance Segmentation Model for Low-Power Edge DevicesabstractDeep learning, for sustainable applications or in cases of energy scarcity, requires using available, cost-effective, and energy-efficient accelerators together with efficient models. We explore using the Yolact model, for instance, segmentation, running on a low power consumption device (e.g., Intel Neural Computing Stick 2 (NCS2)), to detect and segment-specific objects. We have changed the Feature Pyramid Network (FPN) and pruning techniques to make the model usable for this application. The final model achieves a noticeable result in Frames Per Second (FPS) on the edge device while achieving a consistent mean Average Precision (mAP). Diego Bellani, Valerio Venanzi, Shadi Andishmand, Luigi Cinque, Marco Raoul Marini |
ICPRAM | 4 |
| 2025 | SATEER: Subject-Aware Transformer for EEG-Based Emotion RecognitionabstractThis study presents a Subject-Aware Transformer-based neural network designed for the Electroencephalogram (EEG) Emotion Recognition task (SATEER), which entails the analysis of EEG signals to classify and interpret human emotional states. SATEER processes the EEG waveforms by transforming them into Mel spectrograms, which can be seen as particular cases of images with the number of channels equal to the number of electrodes used during the recording process; this type of data can thus be processed using a Computer Vision pipeline. Distinct from preceding approaches, this model addresses the variability in individual responses to identical stimuli by incorporating a User Embedder module. This module enables the association of individual profiles with their EEGs, thereby enhancing classification accuracy. The efficacy of the model was rigorously evaluated using four publicly available datasets, demonstrating superior performance over existing methods in all conducted benchmarks. For instance, on the AMIGOS dataset (A dataset for Multimodal research of affect, personality traits, and mood on Individuals and GrOupS), SATEER's accuracy exceeds 99.8% accuracy across all labels and showcases an improvement of 0.47% over the state of the art. Furthermore, an exhaustive ablation study underscores the pivotal role of the User Embedder module and each other component of the presented model in achieving these advancements. Romeo Lanzino, Danilo Avola, Federico Fontana, Luigi Cinque, Francesco Scarcello, Gian Luca Foresti |
Int. J. Neural Syst. | 4 |
| 2025 | A Context-Dependent CNN-Based Framework for Multiple Sclerosis Segmentation in MRIabstractDespite several automated strategies for identification/segmentation of Multiple Sclerosis (MS) lesions in Magnetic Resonance Imaging (MRI) being developed, they consistently fall short when compared to the performance of human experts. This emphasizes the unique skills and expertise of human professionals in dealing with the uncertainty resulting from the vagueness and variability of MS, the lack of specificity of MRI concerning MS, and the inherent instabilities of MRI. Physicians manage this uncertainty in part by relying on their radiological, clinical, and anatomical experience. We have developed an automated framework for identifying and segmenting MS lesions in MRI scans by introducing a novel approach to replicating human diagnosis, a significant advancement in the field. This framework has the potential to revolutionize the way MS lesions are identified and segmented, being based on three main concepts: (1) Modeling the uncertainty; (2) Use of separately trained Convolutional Neural Networks (CNNs) optimized for detecting lesions, also considering their context in the brain, and to ensure spatial continuity; (3) Implementing an ensemble classifier to combine information from these CNNs. The proposed framework has been trained, validated, and tested on a single MRI modality, the FLuid-Attenuated Inversion Recovery (FLAIR) of the MSSEG benchmark public data set containing annotated data from seven expert radiologists and one ground truth. The comparison with the ground truth and each of the seven human raters demonstrates that it operates similarly to human raters. At the same time, the proposed model demonstrates more stability, effectiveness and robustness to biases than any other state-of-the-art model though using just the FLAIR modality. Giuseppe Placidi, Luigi Cinque, Gian Luca Foresti, Francesca Galassi, Filippo Mignosi, Michele Nappi, Matteo Polsinelli |
Int. J. Neural Syst. | 2 |
| 2025 | Leveraging spatial-channel attention in U-Net for enhanced segmentation of martian dust storms
Daniele Venturini, Marco Raoul Marini, Luigi Cinque, Gian Luca Foresti |
Image Vis. Comput. | 3 |
| 2024 | A Natural Interaction System for Medical Training through VR TechnologyabstractVirtual Reality (VR) technology is rapidly gaining traction as a pivotal tool in medical education, offering immersive and interactive learning environments that show considerable promise, especially in anatomy training. Its ability to simulate complex anatomical structures in a three-dimensional space allows for a deeper understanding and visualization that is difficult to achieve through traditional two-dimensional methods. This study evaluates a VR-based training system that enhances anatomical learning through principles of Human-Computer Interaction (HCI), e.g., hand-tracking technology to avoid the need for traditional controllers. The effectiveness and usability of this system were assessed using the System Usability Scale (SUS), with additional analysis of whether demographic factors such as age, gender, and prior VR experience influence the outcomes. The high achieved results reflect user-friendliness and potential educational effectiveness across diverse user groups. The intuitive nature of the proposed natural interactions significantly enhances the accessibility and engagement of learners, demonstrating that this technology could make advanced medical training more inclusive and broadly accessible. This suggests promising avenues for further research into its application in more complex anatomical and procedural training, aiming to exploit VR’s potential in medical education as a future standard. Marco Raoul Marini, Alessio Mecca, Gian Luca Foresti, Luigi Cinque |
CBMS | 4 |
| 2024 | Semantically Guided Representation Learning For Action Anticipation
Anxhelo Diko, Danilo Avola, Bardh Prenkaj, Federico Fontana, Luigi Cinque |
ECCV (28) | 5 |
| 2024 | FaceVision-GAN: A 3D Model Face Reconstruction Method from a Single Image Using GANsabstractGenerative algorithms have been very successful in recent years. This phenomenon derives from the strong computational power that even consumer computers can provide. Moreover, a huge amount of data is available today for feeding deep learning algorithms. In this context, human 3D face mesh reconstruction is becoming an important but challenging topic in computer vision and computer graphics. It could be exploited in different application areas, from security to avatarization. This paper provides a 3D face reconstruction pipeline based on Generative Adversarial Networks (GANs). It can generate high-quality depth and correspondence maps from 2D images, which are exploited for producing a 3D model of the subject’s face. Danilo Avola, Luigi Cinque, Gian Luca Foresti, Marco Raoul Marini |
ICPRAM | 2 |
| 2024 | Signal enhancement and efficient DTW-based comparison for wearable gait recognitionabstractThe popularity of biometrics-based user identification has significantly increased over the last few years. User identification based on the face, fingerprints, and iris, usually achieves very high accuracy only in controlled setups and can be vulnerable to presentation attacks, spoofing, and forgeries. To overcome these issues, this work proposes a novel strategy based on a relatively less explored biometric trait, i.e., gait, collected by a smartphone accelerometer, which can be more robust to the attacks mentioned above. According to the wearable sensor-based gait recognition state-of-the-art, two main classes of approaches exist: 1) those based on machine and deep learning; 2) those exploiting hand-crafted features. While the former approaches can reach a higher accuracy, they suffer from problems like, e.g., performing poorly outside the training data, i.e., lack of generalizability. This paper proposes an algorithm based on hand-crafted features for gait recognition that can outperform the existing machine and deep learning approaches. It leverages a modified Majority Voting scheme applied to Fast Window Dynamic Time Warping, a modified version of the Dynamic Time Warping (DTW) algorithm with relaxed constraints and majority voting, to recognize gait patterns. We tested our approach named MV-FWDTW on the ZJU-gaitacc, one of the most extensive datasets for the number of subjects, but especially for the number of walks per subject and walk lengths. Results set a new state-of-the-art gait recognition rate of 98.82% in a cross-session experimental setup. We also confirm the quality of the proposed method using a subset of the OU-ISIR dataset, another large state-of-the-art benchmark with more subjects but much shorter walk signals. Danilo Avola, Luigi Cinque, Maria De Marsico, Alessio Fagioli 0001, Gian Luca Foresti, Maurizio Mancini, Alessio Mecca |
Comput. Secur. | 2 |
| 2024 | Spatio-Temporal Image-Based Encoded Atlases for EEG Emotion RecognitionabstractEmotion recognition plays an essential role in human-human interaction since it is a key to understanding the emotional states and reactions of human beings when they are subject to events and engagements in everyday life. Moving towards human-computer interaction, the study of emotions becomes fundamental because it is at the basis of the design of advanced systems to support a broad spectrum of application areas, including forensic, rehabilitative, educational, and many others. An effective method for discriminating emotions is based on ElectroEncephaloGraphy (EEG) data analysis, which is used as input for classification systems. Collecting brain signals on several channels and for a wide range of emotions produces cumbersome datasets that are hard to manage, transmit, and use in varied applications. In this context, the paper introduces the Empátheia system, which explores a different EEG representation by encoding EEG signals into images prior to their classification. In particular, the proposed system extracts spatio-temporal image encodings, or atlases, from EEG data through the Processing and transfeR of Interaction States and Mappings through Image-based eNcoding (PRISMIN) framework, thus obtaining a compact representation of the input signals. The atlases are then classified through the Empátheia architecture, which comprises branches based on convolutional, recurrent, and transformer models designed and tuned to capture the spatial and temporal aspects of emotions. Extensive experiments were conducted on the Shanghai Jiao Tong University (SJTU) Emotion EEG Dataset (SEED) public dataset, where the proposed system significantly reduced its size while retaining high performance. The results obtained highlight the effectiveness of the proposed approach and suggest new avenues for data representation in emotion recognition from EEG signals. Danilo Avola, Luigi Cinque, Angelo Di Mambro, Alessio Fagioli 0001, Marco Raoul Marini, Daniele Pannone, Bruno Fanini, Gian Luca Foresti |
Int. J. Neural Syst. | 2 |
| 2024 | ReViT: Enhancing vision transformers feature diversity with attention residual connections
Anxhelo Diko, Danilo Avola, Marco Cascio, Luigi Cinque |
Pattern Recognit. | 4 |
| 2023 | Keyrtual: A Lightweight Virtual Musical Keyboard Based on RGB-D and Sensors Fusion
Danilo Avola, Luigi Cinque, Marco Raoul Marini, Andrea Princic, Valerio Venanzi |
CAIP (2) | 2 |
| 2023 | Siamese Network to Investigate Scanner-Dependency in MRIabstractMagnetic resonance imaging (MRI) is an effective imaging tool that, due to its non-invasiveness and multiple-parameter nature, is frequently used in medicine. In particular, the MRI's inherent flexibility deriving from the usage of multiple parameters allows to obtain images of variable contrast and quality. However, intrinsic MRI contrast variability often comes with drawbacks in terms of differences in different scanners, thus resulting in the impossibility of standardizing the image contrast. In particular, this variability could negatively affect the automatic analysis of Deep Learning (DL) methods, both in the training phase and in the test phase. In this work, we present several results on how images collected from different MRI scanners are handled by DL methods. To this end, we trained a Siamese network (SNN), based on the EfficientNet-B0 Convolutional Neural Network (EN-CNN), to learn how to recognize the scanner that has generated a given image. The output encoding features of the SNN have been projected into a 2D space with Uniform Manifold Approximation and Projection (UMAP) and have been discussed. Regarding the training phase, the UMAP projects show that the network is capable of separating MR images encoded features from different MRI scanners. Moreover, even if the MR images of different subjects are acquired with the same scanner, the results suggest that there are considerable differences in how the SNN encoded those features. The test phase confirmed that the SNN architecture is capable of recognizing images from different MRI scanners. Matteo Polsinelli, Luigi Cinque, Filippo Mignosi, Giuseppe Placidi, Genny Tortora |
CBMS | 2 |
| 2023 | Hand Tracking and Gesture Recognition by Multiple Contactless Sensors: A SurveyabstractHand tracking and gesture recognition are fundamental in a multitude of applications. Various sensors have been used for this purpose, however, all monocular vision systems face limitations caused by occlusions. Wearable equipment overcome said limitations, although deemed impractical in some cases. Using more than one sensor provides a way to overcome this problem, but necessitates more complicated designs. In this work, we aim to highlight contemporary methods used for hand tracking and gesture recognition by collecting publications of systems developed in the last decade, that employ contactless devices as RGB cameras, IR, and depth sensors, along with some preceding pillar works. Additionally, we briefly present common steps, techniques, and basic algorithms used during the process of developing modern hand tracking and gesture recognition systems and, finally, we derive the trend for the next future. Eleni Theodoridou, Luigi Cinque, Filippo Mignosi, Giuseppe Placidi, Matteo Polsinelli, João Manuel R. S. Tavares, Matteo Spezialetti |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2022 | Human Silhouette and Skeleton Video Synthesis Through Wi-Fi SignalsabstractThe increasing availability of wireless access points (APs) is leading toward human sensing applications based on Wi-Fi signals as support or alternative tools to the widespread visual sensors, where the signals enable to address well-known vision-related problems such as illumination changes or occlusions. Indeed, using image synthesis techniques to translate radio frequencies to the visible spectrum can become essential to obtain otherwise unavailable visual data. This domain-to-domain translation is feasible because both objects and people affect electromagnetic waves, causing radio and optical frequencies variations. In the literature, models capable of inferring radio-to-visual features mappings have gained momentum in the last few years since frequency changes can be observed in the radio domain through the channel state information (CSI) of Wi-Fi APs, enabling signal-based feature extraction, e.g. amplitude. On this account, this paper presents a novel two-branch generative neural network that effectively maps radio data into visual features, following a teacher-student design that exploits a cross-modality supervision strategy. The latter conditions signal-based features in the visual domain to completely replace visual data. Once trained, the proposed method synthesizes human silhouette and skeleton videos using exclusively Wi-Fi signals. The approach is evaluated on publicly available data, where it obtains remarkable results for both silhouette and skeleton videos generation, demonstrating the effectiveness of the proposed cross-modality supervision strategy. Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti |
Int. J. Neural Syst. | 3 |
| 2022 | Affective Action and Interaction Recognition by Multi-View Representation Learning from Handcrafted Low-Level Skeleton FeaturesabstractHuman feelings expressed through verbal (e.g. voice) and non-verbal communication channels (e.g. face or body) can influence either human actions or interactions. In the literature, most of the attention was given to facial expressions for the analysis of emotions conveyed through non-verbal behaviors. Despite this, psychology highlights that the body is an important indicator of the human affective state in performing daily life activities. Therefore, this paper presents a novel method for affective action and interaction recognition from videos, exploiting multi-view representation learning and only full-body handcrafted characteristics selected following psychological and proxemic studies. Specifically, 2D skeletal data are extracted from RGB video sequences to derive diverse low-level skeleton features, i.e. multi-views, modeled through the bag-of-visual-words clustering approach generating a condition-related codebook. In this way, each affective action and interaction within a video can be represented as a frequency histogram of codewords. During the learning phase, for each affective class, training samples are used to compute its global histogram of codewords stored in a database and later used for the recognition task. In the recognition phase, the video frequency histogram representation is matched against the database of class histograms and classified as the closest affective class in terms of Euclidean distance. The effectiveness of the proposed system is evaluated on a specifically collected dataset containing 6 emotion for both actions and interactions, on which the proposed system obtains 93.64% and 90.83% accuracy, respectively. In addition, the devised strategy also achieves in line performances with other literature works based on deep learning when tested on a public collection containing 6 emotions plus a neutral state, demonstrating the effectiveness of the presented approach and confirming the findings in psychological and proxemic studies. Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti |
Int. J. Neural Syst. | 3 |
| 2022 | SIRe-Networks: Convolutional neural networks architectural extension for information preservation via skip/residual connections and interlaced auto-encoders
Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti |
Neural Networks | 2 |
| 2022 | 3D hand pose and shape estimation from RGB images for keypoint-based hand gesture recognition
Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Adriano Fragomeni, Daniele Pannone |
Pattern Recognit. | 2 |
| 2022 | Deep Temporal Analysis for Non-Acted Body Affect RecognitionabstractIn the field of body affect recognition, the majority of literature is based on experiments performed on datasets where trained actors simulate emotional reactions. These acted and unnatural expressions differ from the more challenging genuine emotions, thus leading to less valuable results. In this article, a solution for basic non-acted emotion recognition based on 3D skeleton and Deep Neural Networks (DNNs) is provided. The proposed work introduces three majors contributions. First, temporal local movements performed by subjects are examined frame-by-frame, unlike the current state-of-the-art in non-acted body affect recognition where only static or global body features are considered. Second, an original set of global and time-dependent features for body movement description is provided. Third, this is one of the first works to use deep learning methods in the current non-acted body affect recognition literature. Due to the novelty of the topic, only the UCLIC dataset is currently considered the benchmark for comparative tests. On the latter, the proposed method outperforms all the competitors. Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Cristiano Massaroni |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Multimodal Feature Fusion and Knowledge-Driven Learning via Experts Consult for Thyroid Nodule ClassificationabstractComputer-aided diagnosis (CAD) is becoming a prominent approach to assist clinicians spanning across multiple fields. These automated systems take advantage of various computer vision (CV) procedures, as well as artificial intelligence (AI) techniques, to formulate a diagnosis of a given image, e.g., computed tomography and ultrasound. Advances in both areas (CV and AI) are enabling ever increasing performances of CAD systems, which can ultimately avoid performing invasive procedures such as fine-needle aspiration. In this study, a novel end-to-end knowledge-driven classification framework is presented. The system focuses on multimodal data generated by thyroid ultrasonography, and acts as a CAD system by providing a thyroid nodule classification into the benign and malignant categories. Specifically, the proposed system leverages cues provided by an ensemble of experts to guide the learning phase of a densely connected convolutional network (DenseNet). The ensemble is composed by various networks pretrained on ImageNet, including AlexNet, ResNet, VGG, and others. The previously computed multimodal feature parameters are used to create ultrasonography domain experts via transfer learning, decreasing, moreover, the number of samples required for training. To validate the proposed method, extensive experiments were performed, providing detailed performances for both the experts ensemble and the knowledge-driven DenseNet. As demonstrated by the results, the proposed system achieves relevant performances in terms of qualitative metrics for the thyroid nodule classification task, thus resulting in a great asset when formulating a diagnosis. Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Sebastiano Filetti, Giorgio Grani, Emanuele Rodolà |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Person Re-Identification Through Wi-Fi Extracted Radio Biometric SignaturesabstractPerson re-identification (Re-ID) is a challenging task that tries to recognize a person across different cameras, and that can prove useful in video surveillance as well as in forensics and security applications. However, traditional Re-ID systems analyzing image or video sequences suffer from well-known issues such as illumination changes, occlusions, background clutter, and long-term re-identification. To simultaneously address all these difficult problems, we explore a Re-ID solution based on an alternative medium that is inherently not affected by them, i.e., the Wi-Fi technology. The latter, due to the widespread use of wireless communications, has grown rapidly and is already enabling the development of Wi-Fi sensing applications, such as human localization or counting. These sensing procedures generally exploit Wi-Fi signals variations that are a direct consequence, among other things, of human presence, and which can be observed through the channel state information (CSI) of Wi-Fi access points. Following this rationale, in this paper, for the first time in literature, we show how the pervasive Wi-Fi technology can also be directly exploited for person Re-ID. More accurately, Wi-Fi signals amplitude and phase are extracted from CSI measurements and analyzed through a two-branch deep neural network working in a siamese-like fashion. The designed pipeline can extract meaningful features from signals, i.e., radio biometric signatures, that ultimately allow the person Re-ID. The effectiveness of the proposed system is evaluated on a specifically collected dataset, where remarkable performances are obtained; suggesting that Wi-Fi signal variations differ between different people and can consequently be used for their re-identification. Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Chiara Petrioli |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | LieToMe: An Ensemble Approach for Deception Detection from Facial CuesabstractDeception detection is a relevant ability in high stakes situations such as police interrogatories or court trials, where the outcome is highly influenced by the interviewed person behavior. With the use of specific devices, e.g. polygraph or magnetic resonance, the subject is aware of being monitored and can change his behavior, thus compromising the interrogation result. For this reason, video analysis-based methods for automatic deception detection are receiving ever increasing interest. In this paper, a deception detection approach based on RGB videos, leveraging both facial features and stacked generalization ensemble, is proposed. First, a face, which is well-known to present several meaningful cues for deception detection, is identified, aligned, and masked to build video signatures. These signatures are constructed starting from five different descriptors, which allow the system to capture both static and dynamic facial characteristics. Then, video signatures are given as input to four base-level algorithms, which are subsequently fused applying the stacked generalization technique, resulting in a more robust meta-level classifier used to predict deception. By exploiting relevant cues via specific features, the proposed system achieves improved performances on a public dataset of famous court trials, with respect to other state-of-the-art methods based on facial features, highlighting the effectiveness of the proposed method. Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti |
Int. J. Neural Syst. | 3 |
| 2021 | Forward-looking sonar image compression by integrating keypoint clustering and morphological skeleton
Danilo Avola, Marco Bernardi, Luigi Cinque, Gian Luca Foresti, Daniele Pannone, Chiara Petrioli |
Multim. Tools Appl. | 3 |
| 2021 | Automatic estimation of optimal UAV flight parameters for real-time wide areas monitoring
Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Daniele Pannone, Claudio Piciarelli |
Multim. Tools Appl. | 2 |
| 2021 | Data integration by two-sensors in a LEAP-based Virtual Glove for human-system interactionabstractAbstract Virtual Glove (VG) is a low-cost computer vision system that utilizes two orthogonal LEAP motion sensors to provide detailed 4D hand tracking in real–time. VG can find many applications in the field of human-system interaction, such as remote control of machines or tele-rehabilitation. An innovative and efficient data-integration strategy, based on the velocity calculation, for selecting data from one of the LEAPs at each time, is proposed for VG. The position of each joint of the hand model, when obscured to a LEAP, is guessed and tends to flicker. Since VG uses two LEAP sensors, two spatial representations are available each moment for each joint: the method consists of the selection of the one with the lower velocity at each time instant. Choosing the smoother trajectory leads to VG stabilization and precision optimization, reduces occlusions (parts of the hand or handling objects obscuring other hand parts) and/or, when both sensors are seeing the same joint, reduces the number of outliers produced by hardware instabilities. The strategy is experimentally evaluated, in terms of reduction of outliers with respect to a previously used data selection strategy on VG, and results are reported and discussed. In the future, an objective test set has to be imagined, designed, and realized, also with the help of an external precise positioning equipment, to allow also quantitative and objective evaluation of the gain in precision and, maybe, of the intrinsic limitations of the proposed strategy. Moreover, advanced Artificial Intelligence-based (AI-based) real-time data integration strategies, specific for VG, will be designed and tested on the resulting dataset. Giuseppe Placidi, Danilo Avola, Luigi Cinque, Matteo Polsinelli, Eleni Theodoridou, João Manuel R. S. Tavares |
Multim. Tools Appl. | 3 |
| 2021 | R-SigNet: Reduced space writer-independent feature learning for offline writer-dependent signature verification
Danilo Avola, Manoochehr Joodi Bigdello, Luigi Cinque, Alessio Fagioli 0001, Marco Raoul Marini |
Pattern Recognit. Lett. | 3 |
| 2020 | Guidelines for Effective Automatic Multiple Sclerosis Lesion Segmentation by Magnetic Resonance Imaging
Giuseppe Placidi, Luigi Cinque, Matteo Polsinelli |
ICPRAM | 2 |
| 2020 | Fusing Self-Organized Neural Network and Keypoint Clustering for Localized Real-Time Background SubtractionabstractMoving object detection in video streams plays a key role in many computer vision applications. In particular, separation between background and foreground items represents a main prerequisite to carry out more complex tasks, such as object classification, vehicle tracking, and person re-identification. Despite the progress made in recent years, a main challenge of moving object detection still regards the management of dynamic aspects, including bootstrapping and illumination changes. In addition, the recent widespread of Pan-Tilt-Zoom (PTZ) cameras has made the management of these aspects even more complex in terms of performance due to their mixed movements (i.e. pan, tilt, and zoom). In this paper, a combined keypoint clustering and neural background subtraction method, based on Self-Organized Neural Network (SONN), for real-time moving object detection in video sequences acquired by PTZ cameras is proposed. Initially, the method performs a spatio-temporal tracking of the sets of moving keypoints to recognize the foreground areas and to establish the background. Then, it adopts a neural background subtraction, localized in these areas, to accomplish a foreground detection able to manage bootstrapping and gradual illumination changes. Experimental results on three well-known public datasets, and comparisons with different key works of the current literature, show the efficiency of the proposed method in terms of modeling and background subtraction. Danilo Avola, Marco Bernardi, Luigi Cinque, Cristiano Massaroni, Gian Luca Foresti |
Int. J. Neural Syst. | 3 |
| 2020 | Online separation of handwriting from freehand drawing using extreme learning machines
Danilo Avola, Marco Bernardi, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni |
Multim. Tools Appl. | 3 |
| 2020 | MIFTel: a multimodal interactive framework based on temporal logic rules
Danilo Avola, Luigi Cinque, Alberto Del Bimbo, Marco Raoul Marini |
Multim. Tools Appl. | 2 |
| 2020 | Homography vs similarity transformation in aerial mosaicking: which is the best at different altitudes?
Danilo Avola, Luigi Cinque, Gian Luca Foresti, Daniele Pannone |
Multim. Tools Appl. | 2 |
| 2020 | LieToMe: Preliminary study on hand gestures for deception detection via Fisher-LSTM
Danilo Avola, Luigi Cinque, Maria De Marsico, Alessio Fagioli 0001, Gian Luca Foresti |
Pattern Recognit. Lett. | 2 |
| 2020 | A light CNN for detecting COVID-19 from CT scans of the chest
Matteo Polsinelli, Luigi Cinque, Giuseppe Placidi |
Pattern Recognit. Lett. | 2 |
| 2020 | 2-D Skeleton-Based Action Recognition via Two-Branch Stacked LSTM-RNNsabstractAction recognition in video sequences is an interesting field for many computer vision applications, including behavior analysis, event recognition, and video surveillance. In this article, a method based on 2D skeleton and two-branch stacked Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM) cells is proposed. Unlike 3D skeletons, usually generated by RGB-D cameras, the 2D skeletons adopted in this article are reconstructed starting from RGB video streams, therefore allowing the use of the proposed approach in both indoor and outdoor environments. Moreover, any case of missing skeletal data is managed by exploiting 3D-Convolutional Neural Networks (3D-CNNs). Comparative experiments with several key works on KTH and Weizmann datasets show that the method described in this paper outperforms the current state-of-the-art. Additional experiments on UCF Sports and IXMAS datasets demonstrate the effectiveness of our method in the presence of noisy data and perspective changes, respectively. Further investigations on UCF Sports, HMDB51, UCF101, and Kinetics400 highlight how the combination between the proposed two-branch stacked LSTM and the 3D-CNN-based network can manage missing skeleton information, greatly improving the overall accuracy. Moreover, additional tests on KTH and UCF Sports datasets also show the robustness of our approach in the presence of partial body occlusions. Finally, comparisons on UT-Kinect and NTU-RGB+D datasets show that the accuracy of the proposed method is fully comparable to that of works based on 3D skeletons. Danilo Avola, Marco Cascio, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni, Emanuele Rodolà |
IEEE Trans. Multim. | 3 |
| 2020 | A UAV Video Dataset for Mosaicking and Change Detection From Low-Altitude FlightsabstractIn recent years, the technology of small-scale unmanned aerial vehicles (UAVs) has steadily improved in terms of flight time, automatic control, and image acquisition. This has lead to the development of several applications for low-altitude tasks, such as vehicle tracking, person identification, and object recognition. These applications often require to stitch together several video frames to get a comprehensive view of large areas (mosaicking), or to detect differences between images or mosaics acquired at different times (change detection). However, the datasets used to test mosaicking and change detection algorithms are typically acquired at high-altitudes, thus ignoring the specific challenges of low-altitude scenarios. The purpose of this paper is to fill this gap by providing the UAV mosaicking and change detection dataset. It consists of 50 challenging aerial video sequences acquired at low-altitude in different environments with and without the presence of vehicles, persons, and objects, plus metadata and telemetry. In addition, this paper provides some performance metrics to evaluate both the quality of the obtained mosaics and the correctness of the detected changes. Finally, the results achieved by two baseline algorithms, one for mosaicking and one for detection, are presented. The aim is to provide a shared performance reference that can be used for comparison with future algorithms that will be tested on the dataset. Danilo Avola, Luigi Cinque, Gian Luca Foresti, Niki Martinel, Daniele Pannone, Claudio Piciarelli |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2019 | Master and Rookie Networks for Person Re-identification
Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Cristiano Massaroni |
CAIP (2) | 3 |
| 2019 | An Image-Based Encoding to Record and Track Immersive VR Sessions
Bruno Fanini, Luigi Cinque |
ICCSA (2) | 2 |
| 2019 | A Shape Comparison Reinforcement Method Based on Feature Extractors and F1-ScoreabstractEvaluating object segmentation is a topic of great interest for shape comparison techniques. In this work, ad-hoc metrics for a detailed segmentation analysis and a novel keypoint based method for comparing pairs of shapes are presented. As references, two different segmentation approaches were used: a handmade segmentation and an automatic one based on a Convolutional Neural Network (CNN). The proposed comparison approach consists of a combination between a keypoint extractor and an invariant scale shape identifier. The overall validation process is established according to different steps, which allow to measure the similarity between shapes. First, Reinforced Matched (RM) and Reinforced Ratio (RR) strategies are implemented. Moreover, five different state-of-the-art keypoint extractors are compared, i.e., SIFT, SURF, ORB, A-KAZE, and BRISK. Experimental tests were performed on a popular collection of images, i.e., the Berkeley Segmentation Dataset and Benchmark 300 (BSDS300), which contains shapes segmented both manually and automatically. The experimental results have shown the effectiveness of the proposed method. Danilo Avola, Luigi Cinque, Gian Luca Foresti, Francesco Lamacchia, Marco Raoul Marini, Luca Perini, Kristjana Qorraj, Gabriele Telesca |
SMC | 2 |
| 2019 | An interactive and low-cost full body rehabilitation framework based on 3D immersive serious games
Danilo Avola, Luigi Cinque, Gian Luca Foresti, Marco Raoul Marini |
J. Biomed. Informatics | 2 |
| 2019 | Exploiting Recurrent Neural Networks and Leap Motion Controller for the Recognition of Sign Language and Semaphoric Hand GesturesabstractHand gesture recognition is still a topic of great interest for the computer vision community. In particular, sign language and semaphoric hand gestures are two foremost areas of interest due to their importance in human-human communication and human-computer interaction, respectively. Any hand gesture can be represented by sets of feature vectors that change over time. Recurrent neural networks (RNNs) are suited to analyze this type of set thanks to their ability to model the long-term contextual information of temporal sequences. In this paper, an RNN is trained by using as features the angles formed by the finger bones of the human hands. The selected features, acquired by a leap motion controller sensor, are chosen because the majority of human hand gestures produce joint movements that generate truly characteristic corners. The proposed method, including the effectiveness of the selected angles, was initially tested by creating a very challenging dataset composed by a large number of gestures defined by the American sign language. On the latter, an accuracy of over 96% was achieved. Afterwards, by using the Shape Retrieval Contest (SHREC) dataset, a wide collection of semaphoric hand gestures, the method was also proven to outperform in accuracy competing approaches of the current literature. Danilo Avola, Marco Bernardi, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni |
IEEE Trans. Multim. | 3 |
| 2018 | Combining Keypoint Clustering and Neural Background Subtraction for Real-time Moving Object Detection by PTZ CamerasabstractDetection of moving objects is a topic of great interest in computer vision. This task represents a prerequisite for more complex duties, such as classification and re-identification. One of the main challenges regards the management of dynamic factors, with particular reference to bootstrapping and illumination change issues. The recent widespread of PTZ cameras has made these issues even more complex in terms of performance due to their composite movements (i.e., pan, tilt, and zoom). This paper proposes a combined keypoint clustering and neural background subtraction method for real-time moving object detection in video sequences acquired by PTZ cameras. Initially, the method performs a spatio-temporal tracking of the sets of moving keypoints to recognize the foreground areas and to establish the background. Subsequently, it adopts a neural background subtraction to accomplish a foreground detection, in these areas, able to manage bootstrapping and gradual illumination changes. Experimental results on two well-known public datasets and comparisons with different key works of the current state-of-the-art demonstrate the remarkable results of the proposed method. Danilo Avola, Marco Bernardi, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni |
ICPRAM | 3 |
| 2018 | A Rover-based System for Searching Encrypted Targets in Unknown EnvironmentsabstractIn the last decade, there has been a widespread use of autonomous robots in several application fields, such as border controls, precision agriculture, and military operations. Usually, in the latter, there is the need to encrypt the acquired data, or to mark as relevant some positions or areas. In this paper, we present a client-server rover-based system able to search encrypted targets within an unknown environment. The system uses a rover to explore an unknown environment through a Simultaneous Localization And Mapping (SLAM) algorithm and acquires the scene with a standard RGB camera. Then, by using visual cryptography, it is possible to encrypt the acquired RGB data and to send it to a server, which decrypts the data and checks if it contains a target object. The experiments performed on several objects show the effectiveness of the proposed system. Danilo Avola, Luigi Cinque, Gian Luca Foresti, Marco Raoul Marini, Daniele Pannone |
ICPRAM | 2 |
| 2018 | Towards EEG-based BCI driven by emotions for addressing BCI-Illiteracy: a meta-analytic reviewabstractMany critical aspects affect the correct operation of a Brain Computer Interface. The term ‘BCI-illiteracy’ describes the impossibility of using a BCI paradigm. At present, a universal solution does not exist and seeking innovative protocols to drive a BCI is mandatory. This work presents a meta-analytic review on recent advances in emotions recognition with the perspective of using emotions as voluntary, stimulus-independent, commands for BCIs. 60 papers, based on electroencephalography measurements, were selected to evaluate what emotions have been most recognised and what brain regions were activated by them. It was found that happiness, sadness, anger and calm were the most recognised emotions. Relevant discriminant locations for emotions recognition and for the particular case of discrete emotions recognition were identified in the temporal, frontal and parietal areas. The meta-analysis was mainly performed on stimulus-elicited emotions, due to the limited amount of literature about self-induced emotions. The obtained results represent a good starting point for the development of BCI driven by emotions and allow to: (1) ascertain that emotions are measurable and recognisable one from another (2) select a subset of most recognisable emotions and the corresponding active brain regions. Matteo Spezialetti, Luigi Cinque, João Manuel R. S. Tavares, Giuseppe Placidi |
Behav. Inf. Technol. | 2 |
| 2018 | VRheab: a fully immersive motor rehabilitation system based on recurrent neural network
Danilo Avola, Luigi Cinque, Gian Luca Foresti, Marco Raoul Marini, Daniele Pannone |
Multim. Tools Appl. | 2 |
| 2017 | A Virtual Glove System for the Hand Rehabilitation based on Two Orthogonal LEAP Motion ControllersabstractHand rehabilitation therapy is fundamental in the recovery process for patients suffering from post-stroke or post-surgery impairments.Traditional approaches require the presence of therapist during the sessions, involving high costs and subjective measurements of the patients' abilities and progresses.Recently, several alternative approaches have been proposed.Mechanical devices are often expensive, cumbersome and patient specific, while virtual devices are not subject to this limitations, but, especially if based on a single sensor, could suffer from occlusions.In this paper a novel multi-sensor approach, based on the simultaneous use of two LEAP motion controllers, is proposed.The hardware and software design is illustrated and the measurements error induced by the mutual infrared interference is discussed.Finally, a calibration procedure, a tracking model prototype based on the sensors turnover and preliminary experimental results are presented. Giuseppe Placidi, Luigi Cinque, Andrea Petracca, Matteo Polsinelli, Matteo Spezialetti |
ICPRAM | 2 |
| 2017 | Iterative Adaptive Sparse Sampling Method for Magnetic Resonance Imaging
Giuseppe Placidi, Luigi Cinque, Andrea Petracca, Matteo Polsinelli, Matteo Spezialetti |
ICPRAM | 2 |
| 2017 | Adaptive bootstrapping management by keypoint clustering for background initialization
Danilo Avola, Marco Bernardi, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni |
Pattern Recognit. Lett. | 3 |
| 2017 | A keypoint-based method for background modeling and foreground detection using a PTZ camera
Danilo Avola, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni, Daniele Pannone |
Pattern Recognit. Lett. | 2 |
| 2016 | A Practical Framework for the Development of Augmented Reality Applications by using ArUco MarkersabstractThe Augmented Reality (AR) is an expanding field of the Computer Graphics (CG) that merges items of the real-world environment (e.g., places, objects) with digital information (e.g., multimedia files, virtual objects) to provide users with an enhanced interactive multi-sensorial experience of the real-world that surrounding them. Currently, a wide range of devices is used to vehicular AR systems. Common devices (e.g., cameras equipped on smartphones) enable users to receive multimedia information about target objects (non-immersive AR). Advanced devices (e.g., virtual windscreens) provide users with a set of virtual information about points of interest (POIs) or places (semi-immersive AR). Finally, an ever-increasing number of new devices (e.g., HeadMounted Display, HMD) support users to interact with mixed reality environments (immersive AR). This paper presents a practical framework for the development of non-immersive augmented reality applications through which target objects are enriched with multimedia information. On each target object is applied a different ArUco marker. When a specific application hosted inside a device recognizes, via camera, one of these markers, then the related multimedia information are loaded and added to the target object. The paper also reports a complete case study together with some considerations on the framework and future work. Danilo Avola, Luigi Cinque, Gian Luca Foresti, Cristina Mercuri, Daniele Pannone |
ICPRAM | 2 |
| 2016 | A multipurpose autonomous robot for target recognition in unknown environmentsabstractIn recent years, the technological improvements of consumer robots, in terms of processing capacity and sensors, are enabling an ever-increasing number of researchers to quickly develop both scale prototypes and alternative low cost solutions. In these contexts, a critical aspect is the design of ad-hoc algorithms according to the features of the available hardware. This paper proposes a prototype of an autonomous robot for mapping unknown environments and recognizing target objects. During the setup phase one or more target objects are shown to the RGB camera of the robot which, for each of them, extracts and stores a set of A-KAZE features. Afterwards, the robot adopts the ultrasonic distance measurement and the RGB stream to map the whole environment and search a set of A-KAZE features matchable with those previously acquired. The paper also reports both preliminary tests carried out on a reference indoor environment and a case study performed in an outdoor one that validate the proposed system. Danilo Avola, Gian Luca Foresti, Luigi Cinque, Cristiano Massaroni, Gabriele Vitale, Luca Lombardi |
INDIN | 3 |
| 2016 | Tuning of level-set speed function for speckled image segmentation
Luigi Cinque, Rossella Cossu, Daniela Mansutti, Rosa Maria Spitaleri, Malgorzata Blaszczyk |
Pattern Anal. Appl. | 1 |
| 2015 | Speed Parameters in the Level-Set Segmentation
Luigi Cinque, Rossella Cossu |
CAIP (2) | 1 |
| 2015 | Speckled Images Segmentation and Algorithm Comparison
Luigi Cinque, Rossella Cossu, Rosa Maria Spitaleri |
ICPRAM (2) | 1 |
| 2009 | Practical Parallel Algorithms for Dictionary Data CompressionabstractPRAM CREW parallel algorithms requiring logarithmic time and a linear number of processors exist for sliding (LZ1) and static dictionary compression. On the other hand, LZ2 compression seems hard to parallelize. Both adaptive methods work with prefix dictionaries, that is, all prefixes of a dictionary element are dictionary elements.Therefore, it is reasonable to use prefix dictionaries also for the static method. A left to right semi-greedy approach exists to compute an optimal parsing of a string with a prefix static dictionary. The left to right greedy approach is enough to achieve optimal compression with a sliding dictionary since such dictionary is both prefix and suffix. We assume the window is bounded by a constant. With the practical assumption that the dictionary elements have constant length we present PRAM EREW algorithms for sliding and static dictionary compression still requiring logarithmic time and a linear number of processors. A PRAM EREW decoder for static dictionary compression can be easily designed with a linear number of processors and logarithmic time. A work-optimal logarithmic time PRAM EREW decoder exists for sliding dictionary compression when the window has constant length. The simplest model for parallel computation is an array of processors with distributed memory and no interconnections, therefore, no communication cost. An approximation scheme to optimal compression with prefix static dictionaries was designed running with the same complexity of the previous algorithms on such model. It was presented for a massively parallel architecture but in virtue of its scalability it can be implemented on a small scale system as well. We describe such approach and extend it to the sliding dictionary method. The approximation scheme for sliding dictionaries is suitable for small scale systems but due to its adaptiveness it is practical for a large scale system when the file size is large. A two-dimensional extension of the sliding dictionary method to lossless compression of bi-level images, called block matching, is also discussed. We designed a parallel implementation of such heuristic on a constant size array of processors and experimented it with up to 32 processors of a 256 Intel Xeon 3.06 GHz processors machine (avogadro.cilea.it) on a test set of large topographic images. We achieved the expected speed-up, obtaining parallel compression and decompression about twenty-five times faster than the sequential ones. Luigi Cinque, Sergio De Agostino, Luca Lombardi |
DCC | 1 |
| 2009 | Decomposition of two-dimensional shapes for efficient retrieval
Cecilia Di Ruberto, Luigi Cinque |
Image Vis. Comput. | 2 |
| 2008 | Fast viewpoint-invariant articulated hand detection combining curve and graph matchingabstractWe present an approach to viewpoint invariant hand detection which merges model based representation of shape and exact curve matching with graph search in order to achieve a very low false alarm rate system able to work in real time. The method proposed makes few assumptions on the articulated object nature and can be applied to recognize other articulated objects as well. Luigi Cinque, Marco Cupelli, Enver Sangineto |
FG | 1 |
| 2008 | Identifying elephant photos by multi-curve matching
Alessandro Ardovini, Luigi Cinque, Enver Sangineto |
Pattern Recognit. | 2 |
| 2007 | A Parallel Decoder for Lossless Image Compression by Block MatchingabstractA work-optimal O(lognlogM) time PRAM-EREW algorithm for lossless image compression by block matching was shown in L. Cinque et al., (2003), where n is the size of the image and M is the maximum size of the match. The design of a parallel decoder was left as an open problem. By slightly modifying the parallel encoder, in this paper we show how to implement the decoder in O(lognlogM) time with O(n/logn) processors on the PRAM-EREW. With the realistic assumption that the size of the compressed image is O(n1/2), the parallel decoder requires O(log2n) time and O(n/logn) processors on the mesh of trees Luigi Cinque, Sergio De Agostino |
DCC | 1 |
| 2007 | Deformation tolerant generalized Hough transform for sketch-based image retrieval in complex scenes
Marco Anelli, Luigi Cinque, Enver Sangineto |
Image Vis. Comput. | 2 |
| 2006 | Articulated Object Recognition: A General Framework and a Case StudyabstractWe present in this paper a general-purpose approach for articulated object recognition. We split the recognition process in two distinct phases. In the former we use standard model-based techniques in order to recognize and localize in the input image the rigid components the articulated object is composed of. In the second phase the spatial configurations formed by the recognized components are analyzed and compared with the valid configurations of the object we are searching. The comparison is based on a constraint satisfaction method which can deal with both missing components and false positives. The proposed method is based on a redundant set of constraints which represent the valid spatial configurations of the object's components. Such constraints are not embedded in the system nor are domain-specific but they are learned during a suitable training phase. We show how this approach can be used in different scenarios with different kinds of articulated objects and we present a case study concerning a robotic application. Luigi Cinque, Enver Sangineto, Steven L. Tanimoto |
AVSS | 1 |
| 2006 | 3-D Virtual Environments on Mobile Devices for Remote SurveillanceabstractIn this paper we present a distributed videosurveillance framework. Our end is the remote monitoring of the behavior of people moving in a scene exploiting a virtual reconstruction on low capabilities devices, like PDAs and cell phones. The main novelty of this system is the effective integration of the computer vision and computer graphics modules. The first, using a probabilistic frameworks, can detect the position, the trajectory and the posture of peoples moving in the scene. The second exploits the new possibility of both standard 3D graphics libraries on mobile (namely JSR184 and M3G graphic format) and new PDAs processing capability in order to reconstruct the remote surveillance data in real-time. Roberto Vezzani, Rita Cucchiara, Alessio Malizia, Luigi Cinque |
AVSS | 4 |
| 2004 | A Semantic-Based System for Querying Personal Digital Libraries
Luigi Cinque, Alessio Malizia, Roberto Navigli |
Document Analysis Systems | 1 |
| 2004 | A Simple Lossless Compression Heuristic for RGB ImagesabstractIn this paper, we show a simple lossless compression heuristic for color images in RGB format. The main advantage of this approach is that it provides a highly parallelizable compressor and decompressor. The lossless image compression methods often consist of two distinct and independent components (context-based methods): modeling and coding phase. It can be applied independently to each block of 8/spl times/8 pixels achieving 70 to 80 percent of the compression obtained with LOCO-I (JPEG-LS). Luigi Cinque, Franco Liberati, Sergio De Agostino |
Data Compression Conference | 1 |
| 2004 | Image retrieval using resegmentation driven by query rectangles
Luigi Cinque, Fabio De Rosa, Fabio Lecca, Stefano Levialdi |
Image Vis. Comput. | 1 |
| 2004 | A clustering fuzzy approach for image segmentation
Luigi Cinque, Gian Luca Foresti, Luca Lombardi |
Pattern Recognit. | 1 |
| 2002 | DAN: An Automatic Segmentation and Classification Engine for Paper Documents
Luigi Cinque, Stefano Levialdi, Alessio Malizia, Fabio De Rosa |
Document Analysis Systems | 1 |
| 2002 | A Parallel Algorithm for Lossless Image Compression by Block MatchingabstractSummary form only given. We show a parallel algorithm using a rectangle greedy matching technique which requires a linear number of processors and O(log(M)log(n)) time on the PRAM EREW model. The algorithm is suitable for practical parallel architectures as a mesh of trees, a pyramid or a multigrid. We implement a sequential procedure which simulates the compression performed by the parallel algorithm and it achieves 95 to 97 percent of the compression of a previous sequential heuristic. To achieve logarithmic time we partition an m/spl times/n image, I, in x/spl times/y rectangular areas where x and y are /spl Theta/(log/sup 1/2 / mn). In parallel for each area, one processor applies the sequential parsing algorithm, so that, in logarithmic time, each area is parsed in rectangles, some of which are monochromatic. Before encoding, we compute larger monochromatic rectangles by merging the ones adjacent on the horizontal boundaries and then on the vertical boundaries, doubling in this way the length and width of each area at each step. Luigi Cinque, Sergio De Agostino, Franco Liberati |
DCC | 1 |
| 2002 | Improvements to image magnification
Alberto M. Biancardi, Luigi Cinque, Luca Lombardi |
Pattern Recognit. | 2 |
| 2002 | Segmentation of page images having artifacts of photocopying and scanning
Luigi Cinque, Stefano Levialdi, Luca Lombardi, Steven L. Tanimoto |
Pattern Recognit. | 1 |
| 2001 | LZ1 Compression of Binary Images Using a Simple Rectangle Greedy Matching Technique
Luigi Cinque, Eernesto Grande, Sergio De Agostino |
Data Compression Conference | 1 |
| 2001 | Color-based image retrieval using spatial-chromatic histograms
Luigi Cinque, Gianluigi Ciocca, Stefano Levialdi, A. Pellicanò, Raimondo Schettini |
Image Vis. Comput. | 1 |
| 2001 | A BSP realisation of Jarvis' algorithm
Luigi Cinque, C. di Maggio |
Pattern Recognit. Lett. | 1 |
| 2000 | Optimal Range Segmentation Parameters through Genetic AlgorithmsabstractA wide number of algorithms for surface segmentation in range images have been recently proposed characterized by different approaches (edge filling, region growing,...), different surface types (either for planar or curved surfaces) and different parameters involved. Optimization of the parameter set is a particularly critical task since the range of parameter variability is often quite large: parameter selection depends on surface type, sensors and the required speed which strongly of affect performance. A framework for parameter optimization is proposed based on genetic algorithms. Such algorithms allow a general approach that has been successfully applied on different state-of-the-art segmenters and different range image databases. Luigi Cinque, Stefano Levialdi, Gianluca Pignalberi, Rita Cucchiara, Stefano Martinz |
ICPR | 1 |
| 1998 | Retrieval of images using rich region descriptionsabstractWe present a combination of techniques which can improve recall in image retrieval systems: a fast image segmentation algorithm suitable for color images which enables one to describe images in a database using multiple segmentations. A simple example shows that the use of a "rich description" satisfies a broader range of queries. Luigi Cinque, Fabio Lecca, Stefano Levialdi, Steven L. Tanimoto |
ICPR | 1 |
| 1998 | Self-organizing map for segmenting 3D biological imagesabstractAn image processing method for features extraction and segmentation from three-dimensional (3D) image datasets is presented. Kohonen's self-organizing map (SOM) is used to perform segmentation. Previously, the segmentation method worked on a 2D dataset based on a projection of the three-dimensional dataset (Nguyen et al., 1998). Our 3D approach to segment biological images preserves the 3D object orientations with respect to the surrounding cell volume. A few examples from genetics and brain analysis are provided in order to demonstrate the performance of the proposed method. Luigi Cinque, Raniero Romagnoli, Stefano Levialdi, P. T. A. Nguyen, Ling Guan |
ICPR | 1 |
| 1998 | Matching the resolution level to salient image features
Paolo Bottoni, Luigi Cinque, Stefano Levialdi, Piero Mussio |
Pattern Recognit. | 2 |
| 1998 | 2-D Object Recognition By Multiscale Tree Matching
Virginio Cantoni, Luigi Cinque, Concettina Guerra, Stefano Levialdi, Luca Lombardi |
Pattern Recognit. | 2 |
| 1998 | A multiresolution approach for page segmentation
Luigi Cinque, Luca Lombardi, Gianmarco Manzini |
Pattern Recognit. Lett. | 1 |
| 1998 | Shape description using cubic polynomial Bezier curves
Luigi Cinque, Stefano Levialdi, Alessio Malizia |
Pattern Recognit. Lett. | 1 |
| 1998 | Image thresholding using fuzzy entropiesabstractAn image can be regarded as a fuzzy subset of a plane. A fuzzy entropy measuring the blur in an image is a functional which increases when the sharpness of its argument image decreases. We generalize and extend the relation "sharper than" between fuzzy sets in view of implementing the properties of a relation "sharper than" between images. We show that there are infinitely many implementations of this relation into an ordering between fuzzy sets (equivalently, images). Relying upon these orderings, we construct classes of fuzzy entropies which are useful for image thresholding by cost minimization. Assuming the image to be a degraded version of an ideal two level image (object/background), a fuzzy entropy can be introduced in a cost functional to force the fitting function to be as close as possible to a two-valued function. The minimization problem is numerically solved, and the results obtained on a synthetic image are reported. Silvano Di Zenzo, Luigi Cinque, Stefano Levialdi |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 1997 | Indexing pictorial documents by their content: a survey of current techniques
Maria De Marsico, Luigi Cinque, Stefano Levialdi |
Image Vis. Comput. | 2 |
| 1997 | A Parallel Algorithm for Graph Matching and Its MasPar ImplementationabstractSearch of discrete spaces is important in combinatorial optimization. Such problems arise in artificial intelligence, computer vision, operations research, and other areas. For realistic problems, the search spaces to be processed are usually huge, necessitating long computation times, pruning heuristics, or massively parallel processing. We present an algorithm that reduces the computation time for graph matching by employing both branch-and-bound pruning of the search tree and massively-parallel search of the as-yet-unpruned portions of the space. Most research on parallel search has assumed that a multiple-instruction-stream/multiple-data-stream (MIMD) parallel computer is available. Since massively parallel stream (SIMD) computers are much less expensive than MIMD systems with equal numbers of processors, the question arises as to whether SIMD systems can efficiently handle state-space search problems. We demonstrate that the answer is yes, and in particular, that graph matching has a natural and efficient implementation on SIMD machines. Luigi Cinque, Steven L. Tanimoto, Linda G. Shapiro, Dean Yasuda |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1996 | A parallel partial-sums computation on a pyramid machineabstractA computationally efficient solution to the partial-sums problem is given on a pyramid computer. The algorithm is optimal, since for n elements it requires O(n/lg n) processors and O(lg n) time (excluding the time to load the image). Luigi Cinque |
ICPR | 1 |
| 1996 | Shape Description by a Syntactic Pyramidal ApproachabstractIn the syntactic approach to pattern recognition, patterns are represented as strings, where each pattern is expressed as a composition of its component patterns, called subpatterns and pattern primitives. This approach draws an analogy between the structure of patterns and the syntax of a language. The patterns are considered at a single resolution level, and the recognition of each pattern is usually made by parsing the pattern structure according to a given set of syntax rules, obtained in the first stage of the analysis. This paper describes a pyramidal approach to 2-D object representation which relies on a linguistic description of the object contour at different resolution levels. Our approach, which is parallel and context-sensitive, implies the use of higher dimensional grammars for defining production rules which describe the evolution of the contour between levels. This method is computationally efficient (due to the fast coarse-fine search), rotationally invariant, as well as rather insensitive to noise. Stefano Levialdi, Luigi Cinque |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 1996 | Run-Based Algorithms for Binary Image Analysis and ProcessingabstractIn this paper we suggest a variant of a binary image representation based on run length encoding. This variant allows one to build a "graph representation" for a number of computing tasks like component labeling, computations of Euler number, diameter and convex hull, and the detection of local extrema and multiple points. Finally, a running application in the raster-to-vector conversion of digital maps is provide. Silvano Di Zenzo, Luigi Cinque, Stefano Levialdi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1996 | An improved algorithm for relational distance graph matching
Luigi Cinque, Dean Yasuda, Linda G. Shapiro, Steven L. Tanimoto |
Pattern Recognit. | 1 |
| 1995 | Shape description and recognition by a multiresolution approach
Luigi Cinque, Luca Lombardi |
Image Vis. Comput. | 1 |
| 1995 | Fast pyramidal algorithms for image thresholding
Luigi Cinque, Stefano Levialdi, Azriel Rosenfeld |
Pattern Recognit. | 1 |
| 1995 | Parallel prefix computation on a pyramid computer
Luigi Cinque, Giancarlo Bongiovanni |
Pattern Recognit. Lett. | 1 |
| 1995 | Evaluating digital angles by a parallel diffusion process
Luigi Cinque, Luca Lombardi, Azriel Rosenfeld |
Pattern Recognit. Lett. | 1 |
| 1994 | Recognizing 2D objects by a multi-resolution approachabstractThe method proposed in this paper for two-dimensional object recognition relies on the idea of a structural coding of an object at varying levels of resolution. A tree structure describes the evolution of the contour at increasing levels of detail. Each tree node represents a contour segment through a set of attributes to provide a richer description of the image shape. In addition to the attributes, a weight is associated to each descriptor indicating the significance of the corresponding primitives. An approach for 2D object recognition based on a matching of the attributed tree representation of the candidates with that of the model is proposed. Virginio Cantoni, Luca Lombardi, Luigi Cinque, Stefano Levialdi, Concettina Guerra |
ICPR (3) | 3 |
| 1993 | Image segmentation by a multiresolution approach
Giancarlo Bongiovanni, Luigi Cinque, Stefano Levialdi, Azriel Rosenfeld |
Pattern Recognit. | 2 |
| 1992 | Computing shape description transforms on a multiresolution architecture
Luigi Cinque, Concettina Guerra, Stefano Levialdi |
CVGIP Image Underst. | 1 |
| 1990 | Optimal parallel computation of the quadtree medial axis transform on a multi-layered architectureabstractThe quadtree medial axis is a compact image representation that can be used to derive a number of geometrical properties of an image component. A parallel algorithm for computing the quadtree medial axis transform is described. For an n*n image, the algorithm takes O(log n) time on an n*n pyramid.> Luigi Cinque, Concettina Guerra, Stefano Levialdi |
ICPR (2) | 1 |