VLDB 2026 Research / reviewers in the wild / expert
Gian Luca Foresti
dblp:93/5522
· DBLP profile ↗
203ranked-venue papers
25as first author
50since 2021 · last 2026
0000-0002-8425-6892ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 93 · 12 first-author · 11 since 2021Artificial intelligence and machine learning · 90 · 9 first-author · 36 since 2021Databases, data management, data science and information retrieval · 17 · 3 since 2021Human-computer interaction and ubiquitous computing · 13 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021Systems, architecture and hardware · 2Computer networks · 1Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging the Gaps: Learning to Estimate Missing Text in Fragmentary Greek Inscriptions
Silvia Zottin, Axel De Nardin, Maddalena Zunino, Valentina Mignosa, Gian Luca Foresti |
ICDAR (3) | 5 |
| 2026 | Thyroid Nodule Classification via Weak Self-Supervision and Transfer Learning
Alessio Fagioli 0001, Marco Cascio, Gian Luca Foresti, Luigi Cinque |
ICPRAM | 3 |
| 2026 | A Glossary-Based Learning Platform to Support Novice Learners in IT Literacy
Erica Perseghin, Gian Luca Foresti |
ITiCSE (2) | 2 |
| 2026 | MOSAIC: Maximizing out-of-distribution sensitivity via aligned image classificationabstractOut-of-Distribution (OOD) classification is a domain generalization task in computer vision. Deep learning models are typically developed and tested under the implicit assumption that training and test data are drawn independently and identically distributed (IID) from the same distribution. Overlooking OOD images can lead to poor performance under unseen or adverse viewing conditions, which are common in real-world scenarios. In this work, the proposed solution can be described as a data-driven approach to solve the OOD classification task in computer vision. The proposed approach consists of three stages, a training stage for exploiting labeled source data with different data augmentation strategies using powerful pretrained vision transformer models, an intermediate stage for weighted model ensemble and post-processing strategies, and finally an inference stage for exploiting unlabeled target data by using test-time learning. The proposed data-driven approach enhances the OOD generalization ability of deep models that withstand shifts in nuisances such as shape, pose, context, texture, occlusion, and weather in OOD or rare scenarios. Extensive data-augmentation strategies are used to improve the OOD generalization of deep models across various nuisances. The effectiveness of the proposed approach is evaluated using two standard computer vision benchmarks: ROBIN and a test set provided by the OOD-CV Challenge 2023. The experimental results show that the proposed approach demonstrates a performance improvement of 2.73% the ROBIN test set and achieves accuracy of 94.04% for the Challenge test set in terms of OOD robustness evaluation with classification accuracy. Furthermore, the proposed solution has secured a position within the top three OOD-based rankings on the OOD-CV Challenge Image Classification Leaderboard, 2023. Hussain Ahmad Madni, Rao Muhammad Umer, Carsten Marr, Gian Luca Foresti |
Comput. Vis. Image Underst. | 4 |
| 2026 | U-DIADS-TL: a novel dataset for text line segmentation in historical manuscriptsabstractAbstract Text line segmentation in historical documents remains a significant challenge due to degraded manuscripts, complex layouts, and diverse handwriting styles. Developing robust computational methods is hindered by the scarcity of high-quality ground truth annotations, which require expert knowledge and are time-intensive to produce. Few-shot learning has emerged as a promising solution by enabling model training with minimal annotated data, yet its application to historical document analysis is still largely unexplored. To address this limitation, we introduce U-DIADS-TL (Uniud - Document Image Analysis DataSet - Text Line), a dataset specifically designed for text line segmentation in ancient manuscripts. U-DIADS-TL provides noise-free annotations with non-overlapping text elements and accommodates diverse document structures, including multi-column layouts. To encourage few-shot learning approaches, we offer only three training images, allowing researchers to develop segmentation models that can generalize from limited supervision. Our dataset serves as a critical bridge between deep learning and historical document analysis, fostering the creation of efficient, adaptable segmentation models for real-world applications. Silvia Zottin, Axel De Nardin, Claudio Piciarelli, Gian Luca Foresti |
Int. J. Document Anal. Recognit. | 4 |
| 2026 | MusiKeyrtual: A Framework to Play a Musical Keyboard in Augmented RealityabstractThe advances in Machine and Deep Learning (ML and DL, respectively) contribute to the evolution of modern Augmented Reality (AR) systems, adapting software to complex and interesting applications. Moreover, lower-end systems, such as smartphones, are now capable of running AR applications, albeit often requiring smaller and lighter DL and ML models to accommodate hardware limitations. In this context, we propose MusiKeyrtual, a lightweight application that allows users to play a musical keyboard drawn on paper; an improvement over the previous version, Keyrtual. The application requires only a smartphone to run. The pipeline proposed addresses the hardware limitations of smartphones, both in terms of limited computational capabilities and in terms of using a single RGB camera, which cannot detect depth. Quantitative and qualitative results highlight the effectiveness of the proposed pipeline in exploiting the capabilities of modern smartphones. Valerio Venanzi, Andrea Princic, Marco Raoul Marini, Gian Luca Foresti, Luigi Cinque |
Int. J. Hum. Comput. Interact. | 4 |
| 2026 | Data-related Ablation for Reinforcing Deep Learning in Explaining Complex PhenomenaabstractDeep Learning (DL) models excel at automatically learning intricate patterns within complex data, but their black box nature undermines human trust. To address this, current validation strategies typically focus on the model itself, modifying its architecture to assess the role and importance of the components. However, this model-centric view overlooks the critical learning substrate, which is represented by the data, implicitly assuming that it accurately represents the target phenomenon. This implicit trust in data means that evaluation may fail to detect whether high performance stems from exploiting biases or data quirks rather than learning relevant patterns. We present a novel data-related ablation as a complement to the traditional architectural ablation. Using this framework for Electroencephalography (EEG) signals of Emotional Recognition (ER) and Motor Execution (ME) as a case study, we show that seemingly high-accuracy models often rely heavily on process-irrelevant features, maintaining performance even when key information is eliminated. This shows that a standard, data-independent evaluation can be misleading about whether a model truly captured the intended process; the proposed approach helps distinguish robust learning from leaning on incidental characteristics. Therefore, incorporating data-related ablation is essential for developing reliable and generalizable DL models in fields that rely on data derived from complex and often not completely known phenomena. Romeo Lanzino, Luigi Cinque, Gian Luca Foresti, Giuseppe Placidi |
Int. J. Neural Syst. | 3 |
| 2026 | SAGE-networks: Shape-aware geometric embeddings for writer re-identification in historical manuscripts
Alessio Fagioli 0001, Nicola Follador, Marco Cascio, Emanuela Colombi, Gian Luca Foresti |
Pattern Recognit. | 5 |
| 2026 | FsBAD: Data-efficient feature reconstruction for few-shot brain anomaly detectionabstractData efficiency remains a central challenge in brain anomaly detection, where annotated datasets are often scarce. Most existing methods are tailored to single-class settings and show limited ability to generalize. We introduce FsBAD, a feature reconstruction-based approach designed for few-shot brain anomaly detection with minimal supervision. FsBAD reconstructs a nominal version of an anomalous brain scan by leveraging a small set of aligned reference samples. To enhance reconstruction quality, we propose a novel feature alignment strategy that integrates regression with distribution regularization, promoting both semantic accuracy and nominal consistency. While FsBAD is optimized for brain imaging, we evaluate its generalization capabilities on liver and retina datasets. Experiments across all three domains show that FsBAD consistently outperforms state-of-the-art methods in both image-wise classification and pixel-wise anomaly localization, even in extremely low-shot (2- to 15-shot) settings. This demonstrates FsBAD’s potential as a scalable, data-efficient solution for brain anomaly detection and its robustness across medical imaging tasks. Hussain Ahmad Madni, Hafsa Shujat, Axel De Nardin, Silvia Zottin, Gian Luca Foresti |
Pattern Recognit. Lett. | 5 |
| 2025 | ICDAR 2025 Competition on FEw-Shot Text Line Segmentation of Ancient Handwritten Documents (FEST)
Silvia Zottin, Axel De Nardin, Giuseppe Branca, Claudio Piciarelli, Gian Luca Foresti |
ICDAR (5) | 5 |
| 2025 | In-domain versus out-of-domain transfer learning for document layout analysisabstractAbstract Data availability is a big concern in the field of document analysis, especially when working on tasks that require a high degree of precision when it comes to the definition of the ground truths on which to train deep learning models. A notable example is represented by the task of document layout analysis in handwritten documents, which requires pixel-precise segmentation maps to highlight the different layout components of each document page. These segmentation maps are typically very time-consuming and require a high degree of domain knowledge to be defined, as they are intrinsically characterized by the content of the text. For this reason in the present work, we explore the effects of different initialization strategies for deep learning models employed for this type of task by relying on both in-domain and cross-domain datasets for their pre-training. To test the employed models we use two publicly available datasets with heterogeneous characteristics both regarding their structure as well as the languages of the contained documents. We show how a combination of cross-domain and in-domain transfer learning approaches leads to the best overall performance of the models, as well as speeding up their convergence process. Axel De Nardin, Silvia Zottin, Claudio Piciarelli, Gian Luca Foresti, Emanuela Colombi |
Int. J. Document Anal. Recognit. | 4 |
| 2025 | SATEER: Subject-Aware Transformer for EEG-Based Emotion RecognitionabstractThis study presents a Subject-Aware Transformer-based neural network designed for the Electroencephalogram (EEG) Emotion Recognition task (SATEER), which entails the analysis of EEG signals to classify and interpret human emotional states. SATEER processes the EEG waveforms by transforming them into Mel spectrograms, which can be seen as particular cases of images with the number of channels equal to the number of electrodes used during the recording process; this type of data can thus be processed using a Computer Vision pipeline. Distinct from preceding approaches, this model addresses the variability in individual responses to identical stimuli by incorporating a User Embedder module. This module enables the association of individual profiles with their EEGs, thereby enhancing classification accuracy. The efficacy of the model was rigorously evaluated using four publicly available datasets, demonstrating superior performance over existing methods in all conducted benchmarks. For instance, on the AMIGOS dataset (A dataset for Multimodal research of affect, personality traits, and mood on Individuals and GrOupS), SATEER's accuracy exceeds 99.8% accuracy across all labels and showcases an improvement of 0.47% over the state of the art. Furthermore, an exhaustive ablation study underscores the pivotal role of the User Embedder module and each other component of the presented model in achieving these advancements. Romeo Lanzino, Danilo Avola, Federico Fontana, Luigi Cinque, Francesco Scarcello, Gian Luca Foresti |
Int. J. Neural Syst. | 6 |
| 2025 | Unsupervised Brain MRI Anomaly Detection via Inter-Realization ChannelsabstractAccurate anomaly detection in brain Magnetic Resonance Imaging (MRI) is crucial for early diagnosis of neurological disorders, yet remains a significant challenge due to the high heterogeneity of brain abnormalities and the scarcity of annotated data. Traditional one-class classification models require extensive training on normal samples, limiting their adaptability to diverse clinical cases. In this work, we introduce MadIRC, an unsupervised anomaly detection framework that leverages Inter-Realization Channels (IRC) to construct a robust nominal model without any reliance on labeled data. We extensively evaluate MadIRC on brain MRI as the primary application domain, achieving a localization AUROC of 0.96 outperforming state-of-the-art supervised anomaly detection methods. Additionally, we further validate our approach on liver CT and retinal images to assess its generalizability across medical imaging modalities. Our results demonstrate that MadIRC provides a scalable, label-free solution for brain MRI anomaly detection, offering a promising avenue for integration into real-world clinical workflows. Hussain Ahmad Madni, Hafsa Shujat, Axel De Nardin, Silvia Zottin, Gian Luca Foresti |
Int. J. Neural Syst. | 5 |
| 2025 | A Context-Dependent CNN-Based Framework for Multiple Sclerosis Segmentation in MRIabstractDespite several automated strategies for identification/segmentation of Multiple Sclerosis (MS) lesions in Magnetic Resonance Imaging (MRI) being developed, they consistently fall short when compared to the performance of human experts. This emphasizes the unique skills and expertise of human professionals in dealing with the uncertainty resulting from the vagueness and variability of MS, the lack of specificity of MRI concerning MS, and the inherent instabilities of MRI. Physicians manage this uncertainty in part by relying on their radiological, clinical, and anatomical experience. We have developed an automated framework for identifying and segmenting MS lesions in MRI scans by introducing a novel approach to replicating human diagnosis, a significant advancement in the field. This framework has the potential to revolutionize the way MS lesions are identified and segmented, being based on three main concepts: (1) Modeling the uncertainty; (2) Use of separately trained Convolutional Neural Networks (CNNs) optimized for detecting lesions, also considering their context in the brain, and to ensure spatial continuity; (3) Implementing an ensemble classifier to combine information from these CNNs. The proposed framework has been trained, validated, and tested on a single MRI modality, the FLuid-Attenuated Inversion Recovery (FLAIR) of the MSSEG benchmark public data set containing annotated data from seven expert radiologists and one ground truth. The comparison with the ground truth and each of the seven human raters demonstrates that it operates similarly to human raters. At the same time, the proposed model demonstrates more stability, effectiveness and robustness to biases than any other state-of-the-art model though using just the FLAIR modality. Giuseppe Placidi, Luigi Cinque, Gian Luca Foresti, Francesca Galassi, Filippo Mignosi, Michele Nappi, Matteo Polsinelli |
Int. J. Neural Syst. | 3 |
| 2025 | Leveraging spatial-channel attention in U-Net for enhanced segmentation of martian dust storms
Daniele Venturini, Marco Raoul Marini, Luigi Cinque, Gian Luca Foresti |
Image Vis. Comput. | 4 |
| 2024 | A Natural Interaction System for Medical Training through VR TechnologyabstractVirtual Reality (VR) technology is rapidly gaining traction as a pivotal tool in medical education, offering immersive and interactive learning environments that show considerable promise, especially in anatomy training. Its ability to simulate complex anatomical structures in a three-dimensional space allows for a deeper understanding and visualization that is difficult to achieve through traditional two-dimensional methods. This study evaluates a VR-based training system that enhances anatomical learning through principles of Human-Computer Interaction (HCI), e.g., hand-tracking technology to avoid the need for traditional controllers. The effectiveness and usability of this system were assessed using the System Usability Scale (SUS), with additional analysis of whether demographic factors such as age, gender, and prior VR experience influence the outcomes. The high achieved results reflect user-friendliness and potential educational effectiveness across diverse user groups. The intuitive nature of the proposed natural interactions significantly enhances the accessibility and engagement of learners, demonstrating that this technology could make advanced medical training more inclusive and broadly accessible. This suggests promising avenues for further research into its application in more complex anatomical and procedural training, aiming to exploit VR’s potential in medical education as a future standard. Marco Raoul Marini, Alessio Mecca, Gian Luca Foresti, Luigi Cinque |
CBMS | 3 |
| 2024 | ICDAR 2024 Competition on Few-Shot and Many-Shot Layout Segmentation of Ancient Manuscripts (SAM)
Silvia Zottin, Axel De Nardin, Gian Luca Foresti, Emanuela Colombi, Claudio Piciarelli |
ICDAR (6) | 3 |
| 2024 | FaceVision-GAN: A 3D Model Face Reconstruction Method from a Single Image Using GANsabstractGenerative algorithms have been very successful in recent years. This phenomenon derives from the strong computational power that even consumer computers can provide. Moreover, a huge amount of data is available today for feeding deep learning algorithms. In this context, human 3D face mesh reconstruction is becoming an important but challenging topic in computer vision and computer graphics. It could be exploited in different application areas, from security to avatarization. This paper provides a 3D face reconstruction pipeline based on Generative Adversarial Networks (GANs). It can generate high-quality depth and correspondence maps from 2D images, which are exploited for producing a 3D model of the subject’s face. Danilo Avola, Luigi Cinque, Gian Luca Foresti, Marco Raoul Marini |
ICPRAM | 3 |
| 2024 | A One-Shot Learning Approach to Document Layout Segmentation of Ancient Arabic ManuscriptsabstractDocument layout segmentation is a challenging task due to the variability and complexity of document layouts. Ancient manuscripts in particular are often damaged by age, have very irregular layouts, and are characterized by progressive editing from different authors over a large time window. All these factors make the semantic segmentation process of specific areas, such as main text and side text, very difficult. However, the study of these manuscripts turns out to be fundamental for historians and humanists, so much so that in recent years the demand for machine learning approaches aimed at simplifying the extraction of information from these documents has consistently increased, leading document layout analysis to become an increasingly important research area. In order for machine learning techniques to be applied effectively to this task, however, a large amount of correctly and precisely labeled images is required for their training. This is obviously a limitation for this field of research as ground truth must be precisely and manually crafted by expert humanists, making it a very time-consuming process. In this paper, with the aim of overcoming this limitation, we present an efficient document layout segmentation framework, which while being trained on only one labeled page per manuscript still achieves state-of-the-art performance compared to other popular approaches trained on all the available data when tested on a challenging dataset of ancient Arabic manuscripts. Axel De Nardin, Silvia Zottin, Claudio Piciarelli, Emanuela Colombi, Gian Luca Foresti |
WACV | 5 |
| 2024 | Signal enhancement and efficient DTW-based comparison for wearable gait recognitionabstractThe popularity of biometrics-based user identification has significantly increased over the last few years. User identification based on the face, fingerprints, and iris, usually achieves very high accuracy only in controlled setups and can be vulnerable to presentation attacks, spoofing, and forgeries. To overcome these issues, this work proposes a novel strategy based on a relatively less explored biometric trait, i.e., gait, collected by a smartphone accelerometer, which can be more robust to the attacks mentioned above. According to the wearable sensor-based gait recognition state-of-the-art, two main classes of approaches exist: 1) those based on machine and deep learning; 2) those exploiting hand-crafted features. While the former approaches can reach a higher accuracy, they suffer from problems like, e.g., performing poorly outside the training data, i.e., lack of generalizability. This paper proposes an algorithm based on hand-crafted features for gait recognition that can outperform the existing machine and deep learning approaches. It leverages a modified Majority Voting scheme applied to Fast Window Dynamic Time Warping, a modified version of the Dynamic Time Warping (DTW) algorithm with relaxed constraints and majority voting, to recognize gait patterns. We tested our approach named MV-FWDTW on the ZJU-gaitacc, one of the most extensive datasets for the number of subjects, but especially for the number of walks per subject and walk lengths. Results set a new state-of-the-art gait recognition rate of 98.82% in a cross-session experimental setup. We also confirm the quality of the proposed method using a subset of the OU-ISIR dataset, another large state-of-the-art benchmark with more subjects but much shorter walk signals. Danilo Avola, Luigi Cinque, Maria De Marsico, Alessio Fagioli 0001, Gian Luca Foresti, Maurizio Mancini, Alessio Mecca |
Comput. Secur. | 5 |
| 2024 | Spatio-Temporal Image-Based Encoded Atlases for EEG Emotion RecognitionabstractEmotion recognition plays an essential role in human-human interaction since it is a key to understanding the emotional states and reactions of human beings when they are subject to events and engagements in everyday life. Moving towards human-computer interaction, the study of emotions becomes fundamental because it is at the basis of the design of advanced systems to support a broad spectrum of application areas, including forensic, rehabilitative, educational, and many others. An effective method for discriminating emotions is based on ElectroEncephaloGraphy (EEG) data analysis, which is used as input for classification systems. Collecting brain signals on several channels and for a wide range of emotions produces cumbersome datasets that are hard to manage, transmit, and use in varied applications. In this context, the paper introduces the Empátheia system, which explores a different EEG representation by encoding EEG signals into images prior to their classification. In particular, the proposed system extracts spatio-temporal image encodings, or atlases, from EEG data through the Processing and transfeR of Interaction States and Mappings through Image-based eNcoding (PRISMIN) framework, thus obtaining a compact representation of the input signals. The atlases are then classified through the Empátheia architecture, which comprises branches based on convolutional, recurrent, and transformer models designed and tuned to capture the spatial and temporal aspects of emotions. Extensive experiments were conducted on the Shanghai Jiao Tong University (SJTU) Emotion EEG Dataset (SEED) public dataset, where the proposed system significantly reduced its size while retaining high performance. The results obtained highlight the effectiveness of the proposed approach and suggest new avenues for data representation in emotion recognition from EEG signals. Danilo Avola, Luigi Cinque, Angelo Di Mambro, Alessio Fagioli 0001, Marco Raoul Marini, Daniele Pannone, Bruno Fanini, Gian Luca Foresti |
Int. J. Neural Syst. | 8 |
| 2024 | Robust Federated Learning for Heterogeneous Model and DataabstractData privacy and security is an essential challenge in medical clinical settings, where individual hospital has its own sensitive patients data. Due to recent advances in decentralized machine learning in Federated Learning (FL), each hospital has its own private data and learning models to collaborate with other trusted participating hospitals. Heterogeneous data and models among different hospitals raise major challenges in robust FL, such as gradient leakage, where participants can exploit model weights to infer data. Here, we proposed a robust FL method to efficiently tackle data and model heterogeneity, where we train our model using knowledge distillation and a novel weighted client confidence score on hematological cytomorphology data in clinical settings. In the knowledge distillation, each participant learns from other participants by a weighted confidence score so that knowledge from clean models is distributed other than the noisy clients possessing noisy data. Moreover, we use symmetric loss to reduce the negative impact of data heterogeneity and label diversity by reducing overfitting the model to noisy labels. In comparison to the current approaches, our proposed method performs the best, and this is the first demonstration of addressing both data and model heterogeneity in end-to-end FL that lays the foundation for robust FL in laboratories and clinical applications. Hussain Ahmad Madni, Rao Muhammad Umer, Gian Luca Foresti |
Int. J. Neural Syst. | 3 |
| 2024 | U-DIADS-Bib: a full and few-shot pixel-precise dataset for document layout analysis of ancient manuscripts
Silvia Zottin, Axel De Nardin, Emanuela Colombi, Claudio Piciarelli, Filippo Pavan, Gian Luca Foresti |
Neural Comput. Appl. | 6 |
| 2023 | Efficient few-shot learning for pixel-precise handwritten document layout analysisabstractLayout analysis is a task of uttermost importance in ancient handwritten document analysis and represents a fundamental step toward the simplification of subsequent tasks such as optical character recognition and automatic transcription. However, many of the approaches adopted to solve this problem rely on a fully supervised learning paradigm. While these systems achieve very good performance on this task, the drawback is that pixel-precise text labeling of the entire training set is a very time-consuming process, which makes this type of information rarely available in a real-world scenario. In the present paper, we address this problem by proposing an efficient few-shot learning framework that achieves performances comparable to current state-of-the-art fully supervised methods on the publicly available DIVA-HisDB dataset. Axel De Nardin, Silvia Zottin, Matteo Paier, Gian Luca Foresti, Emanuela Colombi, Claudio Piciarelli |
WACV | 4 |
| 2023 | A late fusion deep neural network for robust speaker identification using raw waveforms and gammatone cepstral coefficientsabstractSpeaker identification aims at determining the speaker identity by analyzing his voice characteristics, and relies typically on statistical models or machine learning techniques. Frequency-domain features are by far the most used choice to encode the audio input in sound recognition. Recently, some studies have also analyzed the use of time-domain raw waveform (RW) with deep neural network (DNN) architectures. In this paper, we hypothesize that both time-domain and frequency-domain features can be used to increase the robustness of speaker identification task in adverse noisy and reverberation conditions, and we present a method based on a late fusion DNN using RWs and gammatone cepstral coefficients (GTCCs). We analyze the characteristics of RW and spectrum-based short-time features, reporting advantages and limitations, and we show that the joint use can increase the identification accuracy. The proposed late fusion DNN model consists of two independent DNN branches made primarily by convolutional neural networks (CNN) and fully connected neural networks (NN) layers. The two DNN branches have as input short-time RW audio fragments and GTCCs, respectively. The late fusion is computed on the predicted scores of the DNN branches. Since the method is based on short segments, it has the advantage of being independent from the size of the input audio signal, and the identification task can be computed by summing the predicted scores over several short-time frames. Analysis of speaker identification performance computed with simulations show that the late fusion DNN model improves the accuracy rate in adverse noise and reverberation conditions in comparison to the RW, the GTCC, and the mel-frequency cepstral coefficients (MFCCs) features. Experiments with real-world speech datasets confirm the efficiency of the proposed method, especially with small-size audio samples. Daniele Salvati, Carlo Drioli, Gian Luca Foresti |
Expert Syst. Appl. | 3 |
| 2023 | Swarm-FHE: Fully Homomorphic Encryption-based Swarm Learning for Malicious ClientsabstractSwarm Learning (SL) is a promising approach to perform the distributed and collaborative model training without any central server. However, data sensitivity is the main concern for privacy when collaborative training requires data sharing. A neural network, especially Generative Adversarial Network (GAN), is able to reproduce the original data from model parameters, i.e. gradient leakage problem. To solve this problem, SL provides a framework for secure aggregation using blockchain methods. In this paper, we consider the scenario of compromised and malicious participants in the SL environment, where a participant can manipulate the privacy of other participant in collaborative training. We propose a method, Swarm-FHE, Swarm Learning with Fully Homomorphic Encryption (FHE), to encrypt the model parameters before sharing with the participants which are registered and authenticated by blockchain technology. Each participant shares the encrypted parameters (i.e. ciphertexts) with other participants in SL training. We evaluate our method with training of the convolutional neural networks on the CIFAR-10 and MNIST datasets. On the basis of a considerable number of experiments and results with different hyperparameter settings, our method performs better as compared to other existing methods. Hussain Ahmad Madni, Rao Muhammad Umer, Gian Luca Foresti |
Int. J. Neural Syst. | 3 |
| 2023 | A Sentiment Analysis Anomaly Detection System for Cyber IntelligenceabstractConsidering the 2030 United Nations intent of world connection, Cyber Intelligence becomes the main area of the human dimension able of inflicting changes in geopolitical dynamics. In cyberspace, the new battlefield is the mind of people including new weapons like abuse of social media with information manipulation, deception by activists and misinformation. In this paper, a Sentiment Analysis system with Anomaly Detection (SAAD) capability is proposed. The system, scalable and modular, uses an OSINT-Deep Learning approach to investigate on social media sentiment in order to predict suspicious anomaly trend in Twitter posts. Anomaly detection is investigated with a new semi-supervised process that is able to detect potentially dangerous situations in critical areas. The main contributions of the paper are the system suitability for working in different areas and domains, the anomaly detection procedure in sentiment context and a time-dependent confusion matrix to address model evaluation with unbalanced dataset. Real experiments and tests were performed on Sahel Region. The detected anomalies in negative sentiment have been checked by experts of Sahel area, proving true links between the models results and real situations observable from the tweets. Roberta Maisano, Gian Luca Foresti |
Int. J. Neural Syst. | 2 |
| 2023 | Few-Shot Pixel-Precise Document Layout Segmentation via Dynamic Instance Generation and Local ThresholdingabstractOver the years, the humanities community has increasingly requested the creation of artificial intelligence frameworks to help the study of cultural heritage. Document Layout segmentation, which aims at identifying the different structural components of a document page, is a particularly interesting task connected to this trend, specifically when it comes to handwritten texts. While there are many effective approaches to this problem, they all rely on large amounts of data for the training of the underlying models, which is rarely possible in a real-world scenario, as the process of producing the ground truth segmentation task with the required precision to the pixel level is a very time-consuming task and often requires a certain degree of domain knowledge regarding the documents at hand. For this reason, in this paper, we propose an effective few-shot learning framework for document layout segmentation relying on two novel components, namely a dynamic instance generation and a segmentation refinement module. This approach is able of achieving performances comparable to the current state of the art on the popular Diva-HisDB dataset, while relying on just a fraction of the available data. Axel De Nardin, Silvia Zottin, Claudio Piciarelli, Emanuela Colombi, Gian Luca Foresti |
Int. J. Neural Syst. | 5 |
| 2023 | Performance evaluation of a Wi-Fi-based multi-node network for distributed audio-visual sensorsabstractAbstract The experimental research described in this manuscript proposes a complete network system for distributed multimedia acquisition by mobile remote nodes, streaming to a central unit, and centralized real-time processing of the collected signals. Particular attention is placed on the hardware structure of the system and on the research of the best network performances for an efficient and secure streaming. Specifically, these acoustic and video sensors, microphone arrays and video cameras respectively, can be employed in any robotic vehicles and systems, both mobile and fixed. The main objective is to intercept unidentified sources, like any kind of vehicles or robotic vehicles, drones, or people whose identity is not a-priory known whose instantaneous location and trajectory are also unknown. The proposed multimedia network infrastructure is analysed and studied in terms of efficiency and robustness, and experiments are conducted on the field to validate it. The hardware and software components of the system were developed using suitable technologies and multimedia transmission protocols to meet the requirements and constraints of computation performance, energy efficiency, and data transmission security. Niccolò Cecchinato, Andrea Toma, Carlo Drioli, Giovanni Ferrin, Gian Luca Foresti |
Multim. Tools Appl. | 5 |
| 2022 | Efficient Detection and Localization of Acoustic Sources with a low complexity CNN network and the Diagonal Unloading BeamformingabstractDetection of acoustic events and direction of arrival estimation of acoustic sources are nowadays central topics in the field of acoustic array signal processing, providing both theoretical and practical relevant perspectives. Reconnaissance and surveillance against intrusions, search and rescue in hostile environments, speaker detection and localization are examples of real applications in which an accurate and efficient analysis of the acoustic scene is required. In this context, we have been investigating an efficient CNN neural network-based method capable of learning from the multi-channel signal an intrinsic function between a specific choice of audio-related features and both the nature of the acoustic event and its spatial location. In this work, we investigate an extended CNN network and compare its performance with the accuracy of a reduced complexity CNN network. The main novelty introduced in this research with respect to the state-of-the-art is that we propose a Diagonal Unloading (DU) Beamforming-based method that produces acoustic maps of azimuth and elevation angles to generate the feature representation of the acoustic signal. A comparative study with the Log-Mel Spectrogram feature representation is also conduced along with a method with fusion of the two feature representations that has been experimented in this work. The dataset for both training and validation of the CNN network belongs to the DCASE challenge. The experiments demonstrated the benefits introduced by the DU-Acoustic Map feature representation that provides additional information about the position of acoustic sources, in terms of angles of azimuth and elevation, through the acoustic maps. The accuracy and efficiency of the proposed deep learning-based method are confirmed by the results. Andrea Toma, Daniele Salvati, Carlo Drioli, Gian Luca Foresti |
IJCNN | 4 |
| 2022 | Pro-CCaps: Progressively Teaching Colourisation to CapsulesabstractAutomatic image colourisation studies how to colourise greyscale images. Existing approaches exploit convolutional layers that extract image-level features learning the colourisation on the entire image, but miss entities-level ones due to pooling strategies. We believe that entity-level features are of paramount importance to deal with the intrinsic multimodality of the problem (i.e., the same object can have different colours, and the same colour can have different properties). Models based on capsule layers aim to identify entity-level features in the image from different points of view, but they do not keep track of global features.Our network architecture integrates entity-level features into the image-level features to generate a plausible image colourisation. We observed that results obtained with direct integration of such two representations are largely dominated by the image-level features, thus resulting in unsaturated colours for the entities. To limit such an issue, we propose a gradual growth of the reconstruction phase of the model while training. By advantaging of prior knowledge from each growing step, we obtain a stable collaboration between image-level and entity-level features that ultimately generates stable and vibrant colourisations. Experimental results on three benchmark datasets, and a user study, demonstrate that our approach has competitive performance with respect to the state-of-the-art and provides more consistent colourisation. Rita Pucci, Christian Micheloni, Gian Luca Foresti, Niki Martinel |
WACV | 3 |
| 2022 | Human Silhouette and Skeleton Video Synthesis Through Wi-Fi SignalsabstractThe increasing availability of wireless access points (APs) is leading toward human sensing applications based on Wi-Fi signals as support or alternative tools to the widespread visual sensors, where the signals enable to address well-known vision-related problems such as illumination changes or occlusions. Indeed, using image synthesis techniques to translate radio frequencies to the visible spectrum can become essential to obtain otherwise unavailable visual data. This domain-to-domain translation is feasible because both objects and people affect electromagnetic waves, causing radio and optical frequencies variations. In the literature, models capable of inferring radio-to-visual features mappings have gained momentum in the last few years since frequency changes can be observed in the radio domain through the channel state information (CSI) of Wi-Fi APs, enabling signal-based feature extraction, e.g. amplitude. On this account, this paper presents a novel two-branch generative neural network that effectively maps radio data into visual features, following a teacher-student design that exploits a cross-modality supervision strategy. The latter conditions signal-based features in the visual domain to completely replace visual data. Once trained, the proposed method synthesizes human silhouette and skeleton videos using exclusively Wi-Fi signals. The approach is evaluated on publicly available data, where it obtains remarkable results for both silhouette and skeleton videos generation, demonstrating the effectiveness of the proposed cross-modality supervision strategy. Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti |
Int. J. Neural Syst. | 5 |
| 2022 | Affective Action and Interaction Recognition by Multi-View Representation Learning from Handcrafted Low-Level Skeleton FeaturesabstractHuman feelings expressed through verbal (e.g. voice) and non-verbal communication channels (e.g. face or body) can influence either human actions or interactions. In the literature, most of the attention was given to facial expressions for the analysis of emotions conveyed through non-verbal behaviors. Despite this, psychology highlights that the body is an important indicator of the human affective state in performing daily life activities. Therefore, this paper presents a novel method for affective action and interaction recognition from videos, exploiting multi-view representation learning and only full-body handcrafted characteristics selected following psychological and proxemic studies. Specifically, 2D skeletal data are extracted from RGB video sequences to derive diverse low-level skeleton features, i.e. multi-views, modeled through the bag-of-visual-words clustering approach generating a condition-related codebook. In this way, each affective action and interaction within a video can be represented as a frequency histogram of codewords. During the learning phase, for each affective class, training samples are used to compute its global histogram of codewords stored in a database and later used for the recognition task. In the recognition phase, the video frequency histogram representation is matched against the database of class histograms and classified as the closest affective class in terms of Euclidean distance. The effectiveness of the proposed system is evaluated on a specifically collected dataset containing 6 emotion for both actions and interactions, on which the proposed system obtains 93.64% and 90.83% accuracy, respectively. In addition, the devised strategy also achieves in line performances with other literature works based on deep learning when tested on a public collection containing 6 emotions plus a neutral state, demonstrating the effectiveness of the presented approach and confirming the findings in psychological and proxemic studies. Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti |
Int. J. Neural Syst. | 5 |
| 2022 | Masked Transformer for Image Anomaly LocalizationabstractImage anomaly detection consists in detecting images or image portions that are visually different from the majority of the samples in a dataset. The task is of practical importance for various real-life applications like biomedical image analysis, visual inspection in industrial production, banking, traffic management, etc. Most of the current deep learning approaches rely on image reconstruction: the input image is projected in some latent space and then reconstructed, assuming that the network (mostly trained on normal data) will not be able to reconstruct the anomalous portions. However, this assumption does not always hold. We thus propose a new model based on the Vision Transformer architecture with patch masking: the input image is split in several patches, and each patch is reconstructed only from the surrounding data, thus ignoring the potentially anomalous information contained in the patch itself. We then show that multi-resolution patches and their collective embeddings provide a large improvement in the model's performance compared to the exclusive use of the traditional square patches. The proposed model has been tested on popular anomaly detection datasets such as MVTec and head CT and achieved good results when compared to other state-of-the-art approaches. Axel De Nardin, Pankaj Mishra, Gian Luca Foresti, Claudio Piciarelli |
Int. J. Neural Syst. | 3 |
| 2022 | SIRe-Networks: Convolutional neural networks architectural extension for information preservation via skip/residual connections and interlaced auto-encoders
Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti |
Neural Networks | 4 |
| 2022 | 3D hand pose and shape estimation from RGB images for keypoint-based hand gesture recognition
Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Adriano Fragomeni, Daniele Pannone |
Pattern Recognit. | 4 |
| 2022 | Acoustic Source Localization Using a Geometrically Sampled Grid SRP-PHAT Algorithm With Max-Pooling OperationabstractThe steered response power phase transform (SRP-PHAT) is a well-known algorithm for acoustic source localization using microphone arrays. It consists in the computation of the generalized cross-correlation (GCC) between each microphone pair, and in the coherent summation of the GCC values in the grid search space. Several improvements based on the volumetric grid have been proposed in order to achieve spatial resolution scalability and to reduce the computational cost by using a coarser grid. In general, the problem of the volumetric based methods is that the noise and the reverberation are projected into the search space since all GCC information is used to build the acoustic map. It is hence proposed a volumetric grid SRP-PHAT algorithm based on the geometrically sampled grid (GSG) that incorporates a max-pooling (MP) operation in the volume accumulation of the GCC values in order to improve the localization performance. The MP is the solution of a minimization-maximization problem that aims at minimizing the deleterious effect of noise and reverberation and at maximizing the accuracy of the GCC values related to the target sound source. Simulations and real-world experiments demonstrate the efficiency of the proposed SRP-GSG-MP algorithm in adverse conditions. Daniele Salvati, Carlo Drioli, Gian Luca Foresti |
IEEE Signal Process. Lett. | 3 |
| 2022 | Deep Temporal Analysis for Non-Acted Body Affect RecognitionabstractIn the field of body affect recognition, the majority of literature is based on experiments performed on datasets where trained actors simulate emotional reactions. These acted and unnatural expressions differ from the more challenging genuine emotions, thus leading to less valuable results. In this article, a solution for basic non-acted emotion recognition based on 3D skeleton and Deep Neural Networks (DNNs) is provided. The proposed work introduces three majors contributions. First, temporal local movements performed by subjects are examined frame-by-frame, unlike the current state-of-the-art in non-acted body affect recognition where only static or global body features are considered. Second, an original set of global and time-dependent features for body movement description is provided. Third, this is one of the first works to use deep learning methods in the current non-acted body affect recognition literature. Due to the novelty of the topic, only the UCLIC dataset is currently considered the benchmark for comparative tests. On the latter, the proposed method outperforms all the competitors. Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Cristiano Massaroni |
IEEE Trans. Affect. Comput. | 4 |
| 2022 | Lord of the Rings: Hanoi Pooling and Self-Knowledge Distillation for Fast and Accurate Vehicle ReidentificationabstractVehicle reidentification has seen increasing interest, thanks to its fundamental impact on intelligent surveillance systems and smart transportation. The visual data acquired from monitoring camera networks come with severe challenges, including occlusions, color and illumination changes, as well as orientation issues (a vehicle can be seen from the side/front/rear due to different camera viewpoints). To deal with such challenges, the community has spent much effort in learning robust feature representations that hinge on additional visual attributes and part-driven methods, but with the side effects of requiring extensive human annotation labor as well as increasing computational complexity. In this article, we propose an approach that learns a feature representation robust to vehicle orientation issues without the need for extra-labeled data and adding negligible computational overheads. The former objective is achieved through the introduction of a Hanoi pooling layer exploiting ring regions and the image pyramid approach yielding a multiscale representation of vehicle appearance. The latter is tackled by transferring the accuracy of a deep network to its first layers, thus reducing the inference effort by the early stop of a test example. This is obtained by means of a self-knowledge distillation framework encouraging multiexit network decisions to agree with each other. Results demonstrate that the proposed approach significantly improves the accuracy of early (i.e., very fast) exits while maintaining the same accuracy of a deep (slow) baseline. Moreover, our solution obtains the best existing performance on three benchmark datasets.11[Online]. Available:https://github.com/iN1k1/. Niki Martinel, Matteo Dunnhofer, Rita Pucci, Gian Luca Foresti, Christian Micheloni |
IEEE Trans. Ind. Informatics | 4 |
| 2021 | Experiencing Contemporary Art at a Distance
Barbara Rita Barricelli, Antonella Varesano, Giuliana Carbi, Torkil Clemmensen, Gian Luca Foresti, José L. Abdelnour-Nocera, Maja Ciric, Gerrit C. van der Veer, Fabio Pittarello, Nuno Nunes 0001, Letizia Bollini, Alexandra Verdeil |
INTERACT (5) | 5 |
| 2021 | Time Delay Estimation for Speaker Localization Using CNN-Based Parametrized GCC-PHAT Features
Daniele Salvati, Carlo Drioli, Gian Luca Foresti |
Interspeech | 3 |
| 2021 | CNN-Based Processing of Acoustic and Radio Frequency Signals for Speaker Localization from MAVsabstractA novel speaker localization algorithm from micro aerial vehicles (MAVs) is investigated. It introduces a joint direction of arrival (DOA) and distance prediction method based on processing and fusion of the multi-channel speech data with radio frequency (RF) measurements of the received signal strength. Possible applications include unmanned aerial vehicles (UAVs)based reconnaissance and surveillance against intrusions and search and rescue in hostile environments. A 3-stages convolutional neural network (CNN) with a fusion layer is proposed to perform this task with the objective of augmenting the source localization from multi-channel speech signals. Two parallel CNNs process the speech and RF data, and the regression network produces predictions of the angle and distance from the source after the fusion layer. To show the performance and effectiveness of this RF-assisted method, the experimental scenario and datasets are presented and experiments are then discussed along with the results that have been obtained. Andrea Toma, Daniele Salvati, Carlo Drioli, Gian Luca Foresti |
Interspeech | 4 |
| 2021 | LieToMe: An Ensemble Approach for Deception Detection from Facial CuesabstractDeception detection is a relevant ability in high stakes situations such as police interrogatories or court trials, where the outcome is highly influenced by the interviewed person behavior. With the use of specific devices, e.g. polygraph or magnetic resonance, the subject is aware of being monitored and can change his behavior, thus compromising the interrogation result. For this reason, video analysis-based methods for automatic deception detection are receiving ever increasing interest. In this paper, a deception detection approach based on RGB videos, leveraging both facial features and stacked generalization ensemble, is proposed. First, a face, which is well-known to present several meaningful cues for deception detection, is identified, aligned, and masked to build video signatures. These signatures are constructed starting from five different descriptors, which allow the system to capture both static and dynamic facial characteristics. Then, video signatures are given as input to four base-level algorithms, which are subsequently fused applying the stacked generalization technique, resulting in a more robust meta-level classifier used to predict deception. By exploiting relevant cues via specific features, the proposed system achieves improved performances on a public dataset of famous court trials, with respect to other state-of-the-art methods based on facial features, highlighting the effectiveness of the proposed method. Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti |
Int. J. Neural Syst. | 5 |
| 2021 | Editorial: From Pioneering Artificial Neural Networks to Deep Learning and Beyond
Gian Luca Foresti |
Int. J. Neural Syst. | 1 |
| 2021 | A Continuous Learning Approach for Real-Time Network Intrusion DetectionabstractNetwork intrusion detection is becoming a challenging task with cyberattacks that are becoming more and more sophisticated. Failing the prevention or detection of such intrusions might have serious consequences. Machine learning approaches try to recognize network connection patterns to classify unseen and known intrusions but also require periodic re-training to keep the performances at a high level. In this paper, a novel continuous learning intrusion detection system, called Soft-Forgetting Self-Organizing Incremental Neural Network (SF-SOINN), is introduced. SF-SOINN, besides providing continuous learning capabilities, is able to perform fast classification, is robust to noise, and it obtains good performances with respect to the existing approaches. The main characteristic of SF-SOINN is the ability to remove nodes from the neural network based on their utility estimate. SF-SOINN has been validated on the well-known NSL-KDD and CIC-IDS-2017 intrusion detection datasets as well as on some artificial data to show the classification capability on more general tasks. Marcello Rinaldo Martina, Gian Luca Foresti |
Int. J. Neural Syst. | 2 |
| 2021 | Supervised Anomaly Detection with Highly Imbalanced Datasets Using Capsule NetworksabstractDetecting anomalous patterns in data is a relevant task in many practical applications, such as defective items detection in industrial inspection systems, cancer identification in medical images, or attacker detection in network intrusion detection systems. This paper focuses on detection of anomalous images, this is images that visually deviate from a reference set of regular data. While anomaly detection has been widely studied in the context of classical machine learning, the application of modern deep learning techniques in this field is still limited. We here propose a capsule-based network for anomaly detection in an extremely imbalanced fully supervised context: we assume that anomaly samples are available, but their amount is limited if compared to regular data. By using a variant of the standard CapsNet architecture, we achieved state-of-the-art results on the MNIST, F-MNIST and K-MNIST datasets. Claudio Piciarelli, Pankaj Mishra, Gian Luca Foresti |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2021 | Forward-looking sonar image compression by integrating keypoint clustering and morphological skeleton
Danilo Avola, Marco Bernardi, Luigi Cinque, Gian Luca Foresti, Daniele Pannone, Chiara Petrioli |
Multim. Tools Appl. | 4 |
| 2021 | Automatic estimation of optimal UAV flight parameters for real-time wide areas monitoring
Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Daniele Pannone, Claudio Piciarelli |
Multim. Tools Appl. | 4 |
| 2021 | Editorial for the special issue on the DAFNE project (DigitalAnastylosis of Frescoes challeNgE)
Virginio Cantoni, Gian Luca Foresti, Nicu Sebe |
Pattern Recognit. Lett. | 2 |
| 2021 | Acoustic Target Tracking Through a Cluster of Mobile AgentsabstractThis paper discusses the problem of tracking a moving target by means of a cluster of mobile agents that is able to sense the acoustic emissions of the target, with the aim of improving the target localization and tracking performance with respect to conventional fixed-array acoustic localization. We handle the acoustic part of the problem by modeling the cluster as a sensor network, and we propose a centralized control strategy for the agents that exploits the spatial sensitivity pattern of the sensor network to estimate the best possible cluster configuration with respect to the expected target position. In order to take into account the position estimation delay due to the frame-based nature of the processing, the possible positions of the acoustic target in a given future time interval are represented in terms of a compatible set, that is, the set of all possible future positions of the target, given its dynamics and its present state. A frame-by-frame cluster reconfiguration algorithm is presented, which adapts the position of each sensing agent with the goal of pursuing the maximum overlap between the region of high acoustic sensitivity of the entire cluster and the compatible set of the sound-emitting target. The tracking scheme iterates, at each observation frame, the computation of the target compatible set, the reconfiguration of the cluster, and the target acoustic localization. The reconfiguration step makes use of an opportune cost function proportional to the difference of the compatibility set and the acoustic sensitivity spatial pattern determined by the mobile agent positions. Simulations under different geometric configurations and positioning constraints demonstrate the ability of the proposed approach to effectively localize and track a moving target based on its acoustic emission. The Doppler effect related to moving sources and sensors is taken into account, and its impact on performance is analyzed. We compare the localization results with conventional static-array localization and positioning of acoustic sensors through genetic algorithm optimization, and results demonstrate the sensible improvements in terms of localization and tracking performance. Although the method is discussed here with respect to acoustic target tracking, it can be effectively adapted to video-based localization and tracking, or to multimodal information settings (e.g., audio and video). Carlo Drioli, Giulia Giordano, Daniele Salvati, Franco Blanchini, Gian Luca Foresti |
IEEE Trans. Cybern. | 5 |
| 2020 | Impairments in decoding facial and vocal emotional expressions in high functioning autistic adults and adolescentsabstractThe present investigation shows that gender of stimuli, age, and emotional categories affects the ability of adults and adolescent with Autistic Spectrum Conditions (ASC) to decode facial and vocal emotional expressions. A total of 60 subjects participated to the research: 15 ASC and 15 control adolescents aged between 10-14 years; and 15 ASC and 15 control young adults aged between 20-24 years. Their tasks consisted in decoding: a) 24 adults and 24 children contemporary facial emotional expressions of happiness, sadness, anger, fear, surprise, and disgust; and b) 20 adult's vocal emotional expressions of the same abovementioned emotions (except disgust). Significant differences were observed between ASC and typically developed peers. The data suggest that gender, type (voices or faces) of stimuli, and participants' age affect the emotion recognition process making difficult the definition of a common and shared pattern of emotional expression's recognition compliance among autistic and control groups. These results suggest that efficient and effective e-health technologies need to be able to learn and adapt to user individual traits and subjective needs to offer personalized assistance and support. Anna Esposito, Italia Cirillo, Antonietta Maria Esposito, Leopoldina Fortunati, Gian Luca Foresti, Sergio Escalera, Nikolaos G. Bourbakis |
FG | 5 |
| 2020 | Fixed simplex coordinates for angular margin loss in CapsNetabstractA more stationary and discriminative embedding is necessary for robust classification of images. We focus our attention on the newel CapsNet model and we propose the angular margin loss function in composition with margin loss. We define a fixed classifier implemented with fixed weights vectors obtained by the vertex coordinates of a simplex polytope. The advantage of using simplex polytope is that we obtain the maximal symmetry for stationary features angularly centred. Each weight vector is to be considered as the centroid of a class in the dataset. The embedding of an image is obtained through the capsule network encoding phase, that is identified as digitcaps matrix. Based on the centroids from the simplex coordinates and the embedding from the model, we compute the angular distance between the image embedding and the centroid of the correspondent class of the image. We take this angular distance as angular margin loss. We keep the computation proposed for margin loss in the original architecture of CapsNet. We train the model to minimise the angular between the embedding and the centroid of the class and maximise the magnitude of the embedding for the predicted class. The experiments on different datasets demonstrate that the angular margin loss improves the capability of capsule networks with complex datasets. Rita Pucci, Christian Micheloni, Gian Luca Foresti, Niki Martinel |
ICPR | 3 |
| 2020 | Deep Iterative Residual Convolutional Network for Single Image Super-ResolutionabstractDeep convolutional neural networks (CNNs) have recently achieved great success for single image super-resolution (SISR) task due to their powerful feature representation capabilities. The most recent deep learning based SISR methods focus on designing deeper / wider models to learn the non-linear mapping between low-resolution (LR) inputs and high-resolution (HR) outputs. These existing SR methods do not take into account the image observation (physical) model and thus require a large number of network's trainable parameters with a great volume of training data. To address these issues, we propose a deep Iterative Super-Resolution Residual Convolutional Network (ISRResCNet) that exploits the powerful image regularization and large-scale optimization techniques by training the deep network in an iterative manner with a residual learning approach. Extensive experimental results on various super-resolution benchmarks demonstrate that our method with a few trainable parameters improves the results for different scaling factors in comparison with the state-of-art methods. Rao Muhammad Umer, Gian Luca Foresti, Christian Micheloni |
ICPR | 2 |
| 2020 | Two-Microphone End-to-End Speaker Joint Identification and Localization Via Convolutional Neural NetworksabstractWe present an end-to-end scheme based on convolutional neural networks (CNNs) for speaker joint identification and localization. We investigate the possibility to estimate both the direction of arrival (DOA) and the identity of the speaker in far-field noisy and reverberant conditions using a two-channel microphone array. The proposed CNN network is designed to map the raw waveform of the two channels into the speaker identity and into the DOA of its speech signal. We analyze the identification and localization performance with simulated experiments in noisy and reverberation conditions. Daniele Salvati, Carlo Drioli, Gian Luca Foresti |
IJCNN | 3 |
| 2020 | Drone swarm patrolling with uneven coverage requirementsabstractSwarms of drones are being more and more used in many practical scenarios, such as surveillance, environmental monitoring, search and rescue in hardly‐accessible areas and so on. While a single drone can be guided by a human operator, the deployment of a swarm of multiple drones requires proper algorithms for automatic task‐oriented control. In this study, the authors focus on visual coverage optimisation with drone‐mounted camera sensors. In particular, they consider the specific case in which the coverage requirements are uneven, meaning that different parts of the environment have different coverage priorities. They model these coverage requirements with relevance maps and propose a deep reinforcement learning algorithm to guide the swarm. This study first defines a proper learning model for a single drone, and then extends it to the case of multiple drones both with greedy and cooperative strategies. Experimental results show the performance of the proposed method, also compared with a standard patrolling algorithm. Claudio Piciarelli, Gian Luca Foresti |
IET Comput. Vis. | 2 |
| 2020 | Fusing Self-Organized Neural Network and Keypoint Clustering for Localized Real-Time Background SubtractionabstractMoving object detection in video streams plays a key role in many computer vision applications. In particular, separation between background and foreground items represents a main prerequisite to carry out more complex tasks, such as object classification, vehicle tracking, and person re-identification. Despite the progress made in recent years, a main challenge of moving object detection still regards the management of dynamic aspects, including bootstrapping and illumination changes. In addition, the recent widespread of Pan-Tilt-Zoom (PTZ) cameras has made the management of these aspects even more complex in terms of performance due to their mixed movements (i.e. pan, tilt, and zoom). In this paper, a combined keypoint clustering and neural background subtraction method, based on Self-Organized Neural Network (SONN), for real-time moving object detection in video sequences acquired by PTZ cameras is proposed. Initially, the method performs a spatio-temporal tracking of the sets of moving keypoints to recognize the foreground areas and to establish the background. Then, it adopts a neural background subtraction, localized in these areas, to accomplish a foreground detection able to manage bootstrapping and gradual illumination changes. Experimental results on three well-known public datasets, and comparisons with different key works of the current literature, show the efficiency of the proposed method in terms of modeling and background subtraction. Danilo Avola, Marco Bernardi, Luigi Cinque, Cristiano Massaroni, Gian Luca Foresti |
Int. J. Neural Syst. | 5 |
| 2020 | A Neural Network for Image Anomaly Detection with Deep Pyramidal Representations and Dynamic RoutingabstractImage anomaly detection is an application-driven problem where the aim is to identify novel samples, which differ significantly from the normal ones. We here propose Pyramidal Image Anomaly DEtector (PIADE), a deep reconstruction-based pyramidal approach, in which image features are extracted at different scale levels to better catch the peculiarities that could help to discriminate between normal and anomalous data. The features are dynamically routed to a reconstruction layer and anomalies can be identified by comparing the input image with its reconstruction. Unlike similar approaches, the comparison is done by using structural similarity and perceptual loss rather than trivial pixel-by-pixel comparison. The proposed method performed at par or better than the state-of-the-art methods when tested on publicly available datasets such as CIFAR10, COIL-100 and MVTec. Pankaj Mishra, Claudio Piciarelli, Gian Luca Foresti |
Int. J. Neural Syst. | 3 |
| 2020 | Online separation of handwriting from freehand drawing using extreme learning machines
Danilo Avola, Marco Bernardi, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni |
Multim. Tools Appl. | 4 |
| 2020 | Homography vs similarity transformation in aerial mosaicking: which is the best at different altitudes?
Danilo Avola, Luigi Cinque, Gian Luca Foresti, Daniele Pannone |
Multim. Tools Appl. | 3 |
| 2020 | A scalable system for the monitoring of video transmission components in delay-sensitive networked applications
Andrea Bulfone, Carlo Drioli, Giovanni Ferrin, Gian Luca Foresti |
Multim. Tools Appl. | 4 |
| 2020 | Deep interactive encoding with capsule networks for image classification
Rita Pucci, Christian Micheloni, Gian Luca Foresti, Niki Martinel |
Multim. Tools Appl. | 3 |
| 2020 | Adaptive neural tree exploiting expert nodes to classify high-dimensional data
Shadi Abpeikar, Mehdi Ghatee, Gian Luca Foresti, Christian Micheloni |
Neural Networks | 3 |
| 2020 | LieToMe: Preliminary study on hand gestures for deception detection via Fisher-LSTM
Danilo Avola, Luigi Cinque, Maria De Marsico, Alessio Fagioli 0001, Gian Luca Foresti |
Pattern Recognit. Lett. | 5 |
| 2020 | Diagonal Unloading Beamforming in the Spherical Harmonic Domain for Acoustic Source Localization in Reverberant EnvironmentsabstractSpherical microphone arrays allow the sound field analysis in three dimensions with the advantage of having the same resolution in all directions. By considering the frequency-independent character of the steering vectors in the spherical harmonic (SH) domain, we propose a very low-complexity SH diagonal unloading (DU) beamforming with a novel frequency smoothing power transform (FSPT) of the covariance matrices. We consider the direction of arrival (DOA) estimation problem of acoustic sources in reverberant conditions. The DU beamforming provides high resolution directional response since it exploits the subspace orthogonality property of the covariance matrix by the removal or the attenuation of the signal subspaces, obtained through the subtraction of an opportune diagonal matrix from the covariance matrix. The FSPT aims at smoothing the narrowband covariance matrices of the entire set of frequency domain components, and it pursues this goal by minimizing the narrowband error contributions due to reverberation in the broadband frequency smoothing covariance matrix. We analyze the DOA estimation performance using speech signals with simulations and real acoustic data in reverberant conditions. The results show that the proposed SH-DU-FSPT has a DOA estimation performance comparable to that of high resolution state-of-the-art methods with a significant reduction of the computational cost, since the steering directional responses are computed on the broadband frequency smoothing covariance matrix. Daniele Salvati, Carlo Drioli, Gian Luca Foresti |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Deep Pyramidal Pooling With Attention for Person Re-IdentificationabstractLearning discriminative, view-invariant and multi-scale representations of object appearance with different semantic levels is of paramount importance for person Re-Identification (ReID). Recently, the community has focused on learning deep Re-ID models to capture a single holistic representation. To improve the achieved results, additional visual attributes and object part-driven models have been considered, inevitably introducing additional human annotation labor or computational efforts. In this paper, we argue that pyramid-inspired methods capturing multi-scale information may overcome such requirements. Precisely, multi-scale pooled regions representing visual information of an object are integrated within a novel deep architecture factorizing them into discriminative features at multiple semantic levels. These are exploited through an attention mechanism later considered in an identification-similarity multi-task loss, trained by means of a curriculum learning strategy. Extensive results on three person ReID benchmarks demonstrate that better performance than existing methods are achieved. Code is available at https://github.com/iN1k1. Niki Martinel, Gian Luca Foresti, Christian Micheloni |
IEEE Trans. Image Process. | 2 |
| 2020 | 2-D Skeleton-Based Action Recognition via Two-Branch Stacked LSTM-RNNsabstractAction recognition in video sequences is an interesting field for many computer vision applications, including behavior analysis, event recognition, and video surveillance. In this article, a method based on 2D skeleton and two-branch stacked Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM) cells is proposed. Unlike 3D skeletons, usually generated by RGB-D cameras, the 2D skeletons adopted in this article are reconstructed starting from RGB video streams, therefore allowing the use of the proposed approach in both indoor and outdoor environments. Moreover, any case of missing skeletal data is managed by exploiting 3D-Convolutional Neural Networks (3D-CNNs). Comparative experiments with several key works on KTH and Weizmann datasets show that the method described in this paper outperforms the current state-of-the-art. Additional experiments on UCF Sports and IXMAS datasets demonstrate the effectiveness of our method in the presence of noisy data and perspective changes, respectively. Further investigations on UCF Sports, HMDB51, UCF101, and Kinetics400 highlight how the combination between the proposed two-branch stacked LSTM and the 3D-CNN-based network can manage missing skeleton information, greatly improving the overall accuracy. Moreover, additional tests on KTH and UCF Sports datasets also show the robustness of our approach in the presence of partial body occlusions. Finally, comparisons on UT-Kinect and NTU-RGB+D datasets show that the accuracy of the proposed method is fully comparable to that of works based on 3D skeletons. Danilo Avola, Marco Cascio, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni, Emanuele Rodolà |
IEEE Trans. Multim. | 4 |
| 2020 | A UAV Video Dataset for Mosaicking and Change Detection From Low-Altitude FlightsabstractIn recent years, the technology of small-scale unmanned aerial vehicles (UAVs) has steadily improved in terms of flight time, automatic control, and image acquisition. This has lead to the development of several applications for low-altitude tasks, such as vehicle tracking, person identification, and object recognition. These applications often require to stitch together several video frames to get a comprehensive view of large areas (mosaicking), or to detect differences between images or mosaics acquired at different times (change detection). However, the datasets used to test mosaicking and change detection algorithms are typically acquired at high-altitudes, thus ignoring the specific challenges of low-altitude scenarios. The purpose of this paper is to fill this gap by providing the UAV mosaicking and change detection dataset. It consists of 50 challenging aerial video sequences acquired at low-altitude in different environments with and without the presence of vehicles, persons, and objects, plus metadata and telemetry. In addition, this paper provides some performance metrics to evaluate both the quality of the obtained mosaics and the correctness of the detected changes. Finally, the results achieved by two baseline algorithms, one for mosaicking and one for detection, are presented. The aim is to provide a shared performance reference that can be used for comparison with future algorithms that will be tested on the dataset. Danilo Avola, Luigi Cinque, Gian Luca Foresti, Niki Martinel, Daniele Pannone, Claudio Piciarelli |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2019 | Master and Rookie Networks for Person Re-identification
Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Cristiano Massaroni |
CAIP (2) | 5 |
| 2019 | Distributional memory explainable word embeddings in continuous space
Lauro Snidaro, Giovanni Ferrin, Gian Luca Foresti |
FUSION | 3 |
| 2019 | End-to-End Speaker Identification in Noisy and Reverberant Environments Using Raw Waveform Convolutional Neural Networks
Daniele Salvati, Carlo Drioli, Gian Luca Foresti |
INTERSPEECH | 3 |
| 2019 | A Shape Comparison Reinforcement Method Based on Feature Extractors and F1-ScoreabstractEvaluating object segmentation is a topic of great interest for shape comparison techniques. In this work, ad-hoc metrics for a detailed segmentation analysis and a novel keypoint based method for comparing pairs of shapes are presented. As references, two different segmentation approaches were used: a handmade segmentation and an automatic one based on a Convolutional Neural Network (CNN). The proposed comparison approach consists of a combination between a keypoint extractor and an invariant scale shape identifier. The overall validation process is established according to different steps, which allow to measure the similarity between shapes. First, Reinforced Matched (RM) and Reinforced Ratio (RR) strategies are implemented. Moreover, five different state-of-the-art keypoint extractors are compared, i.e., SIFT, SURF, ORB, A-KAZE, and BRISK. Experimental tests were performed on a popular collection of images, i.e., the Berkeley Segmentation Dataset and Benchmark 300 (BSDS300), which contains shapes segmented both manually and automatically. The experimental results have shown the effectiveness of the proposed method. Danilo Avola, Luigi Cinque, Gian Luca Foresti, Francesco Lamacchia, Marco Raoul Marini, Luca Perini, Kristjana Qorraj, Gabriele Telesca |
SMC | 3 |
| 2019 | From person to group re-identification via unsupervised transfer of sparse features
Giuseppe Lisanti, Niki Martinel, Christian Micheloni, Alberto Del Bimbo, Gian Luca Foresti |
Image Vis. Comput. | 5 |
| 2019 | An interactive and low-cost full body rehabilitation framework based on 3D immersive serious games
Danilo Avola, Luigi Cinque, Gian Luca Foresti, Marco Raoul Marini |
J. Biomed. Informatics | 3 |
| 2019 | Fusing depth and colour information for human action recognition
Danilo Avola, Marco Bernardi, Gian Luca Foresti |
Multim. Tools Appl. | 3 |
| 2019 | Distributed person re-identification through network-wise rank fusion consensus
Niki Martinel, Gian Luca Foresti, Christian Micheloni |
Pattern Recognit. Lett. | 2 |
| 2019 | Power Method for Robust Diagonal Unloading Localization BeamformingabstractWe propose a robust version of the diagonal unloading (DU) beamforming for the acoustic source localization problem in high noise conditions. The DU beamformer exploits the subspace orthogonality property by the removal or the attenuation of the signal subspaces, obtained through the subtraction of an opportune diagonal matrix from the covariance matrix. As a result, it provides high-resolution directional response with low computational complexity. We show that a robust DU beamformer can be implemented by subtracting the largest eigenvalue of the estimated covariance matrix from the diagonal elements, and that this implementation is valid in general (i.e., for both the single-source and the multiple-source case). We propose the use of the power method for the estimation of the largest eigenvalue in the DU procedure. We show with numerical simulations that the proposed method improves the localization performance in high noise conditions without substantial increment of the computational cost. Applications for this method include a number of scenarios involving multirotor aerial systems due to its robustness to the noise and its low computational complexity. Daniele Salvati, Carlo Drioli, Gian Luca Foresti |
IEEE Signal Process. Lett. | 3 |
| 2019 | A Vision-Based System for Internal Pipeline InspectionabstractThe internal inspection of large pipeline infrastructures, such as sewers and waterworks, is a fundamental task for the prevention of possible failures. In particular, visual inspection is typically performed by human operators on the basis of video sequences either acquired on-line or recorded for further off-line analysis. In this work, we propose a vision-based software approach to assist the human operator by conveniently showing the acquired data and by automatically detecting and highlighting the pipeline sections where relevant anomalies could occur. Claudio Piciarelli, Danilo Avola, Daniele Pannone, Gian Luca Foresti |
IEEE Trans. Ind. Informatics | 4 |
| 2019 | Exploiting Recurrent Neural Networks and Leap Motion Controller for the Recognition of Sign Language and Semaphoric Hand GesturesabstractHand gesture recognition is still a topic of great interest for the computer vision community. In particular, sign language and semaphoric hand gestures are two foremost areas of interest due to their importance in human-human communication and human-computer interaction, respectively. Any hand gesture can be represented by sets of feature vectors that change over time. Recurrent neural networks (RNNs) are suited to analyze this type of set thanks to their ability to model the long-term contextual information of temporal sequences. In this paper, an RNN is trained by using as features the angles formed by the finger bones of the human hands. The selected features, acquired by a leap motion controller sensor, are chosen because the majority of human hand gestures produce joint movements that generate truly characteristic corners. The proposed method, including the effectiveness of the selected angles, was initially tested by creating a very challenging dataset composed by a large number of gestures defined by the American sign language. On the latter, an accuracy of over 96% was achieved. Afterwards, by using the Shape Retrieval Contest (SHREC) dataset, a wide collection of semaphoric hand gestures, the method was also proven to outperform in accuracy competing approaches of the current literature. Danilo Avola, Marco Bernardi, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni |
IEEE Trans. Multim. | 4 |
| 2018 | Describing Capability Through Lexical Semantics Exploitation: Foundational ArgumentsabstractIn everyday life as well as in asymmetric warfare domain, to achieve the intended goals, agents often do not make use of the designed and purpose-built tools, but some other tools whose features simply fit for the purpose. The present paper discusses the possibility of capturing and integrating relations and features from context that could drive the retrieval of possible candidate substitutes for the properly designed artifacts through Lexical Semantics Exploitation. Generative Lexicon theory assumes a structure (QualiaStructure) organizing the semantic content carried by lexical items through roles. Among them, theTelicrole exposes the function or purpose of the predicated entity and theConstitutiverole exposes its component parts. We argue that the typical function an entity has been thought for is related to its internal constituents. We also argue that a knowledge base and a proper metrics can be conveniently built extractingQualiaelements from suitable text corpora. Giovanni Ferrin, Lauro Snidaro, Gian Luca Foresti |
FUSION | 3 |
| 2018 | Context-Based Goal-Driven Reasoning for Improved Target TrackingabstractTracking objects in complex dynamic environments can be less challenging once their behavior is recognized. Inferring on targets' future actions based on their past can be addressed via probabilistic reasoning. Context information plays a crucial role in the reasoning process as it provides additional clues about targets' behavior. Combining context reasoning with target tracking continues to increase with the availability of supporting information. The framework here discussed views target's actions as a Hidden Markov Model (HMM) with relevant context associated with each node. Context is at each time step selected based on immediate and goal driven sets of actions. Inference in the HMM is conditioned on prior target's measurements and the belief state conditioned on context. This posterior is then compared with the target's state estimate in order to adjust the switching probability in the Interactive Multiple Models (IMM) tracking process. Lubos Vaci, Lauro Snidaro, Gian Luca Foresti |
FUSION | 3 |
| 2018 | Combining Keypoint Clustering and Neural Background Subtraction for Real-time Moving Object Detection by PTZ CamerasabstractDetection of moving objects is a topic of great interest in computer vision. This task represents a prerequisite for more complex duties, such as classification and re-identification. One of the main challenges regards the management of dynamic factors, with particular reference to bootstrapping and illumination change issues. The recent widespread of PTZ cameras has made these issues even more complex in terms of performance due to their composite movements (i.e., pan, tilt, and zoom). This paper proposes a combined keypoint clustering and neural background subtraction method for real-time moving object detection in video sequences acquired by PTZ cameras. Initially, the method performs a spatio-temporal tracking of the sets of moving keypoints to recognize the foreground areas and to establish the background. Subsequently, it adopts a neural background subtraction to accomplish a foreground detection, in these areas, able to manage bootstrapping and gradual illumination changes. Experimental results on two well-known public datasets and comparisons with different key works of the current state-of-the-art demonstrate the remarkable results of the proposed method. Danilo Avola, Marco Bernardi, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni |
ICPRAM | 4 |
| 2018 | A Rover-based System for Searching Encrypted Targets in Unknown EnvironmentsabstractIn the last decade, there has been a widespread use of autonomous robots in several application fields, such as border controls, precision agriculture, and military operations. Usually, in the latter, there is the need to encrypt the acquired data, or to mark as relevant some positions or areas. In this paper, we present a client-server rover-based system able to search encrypted targets within an unknown environment. The system uses a rover to explore an unknown environment through a Simultaneous Localization And Mapping (SLAM) algorithm and acquires the scene with a standard RGB camera. Then, by using visual cryptography, it is possible to encrypt the acquired RGB data and to send it to a server, which decrypts the data and checks if it contains a target object. The experiments performed on several objects show the effectiveness of the proposed system. Danilo Avola, Luigi Cinque, Gian Luca Foresti, Marco Raoul Marini, Daniele Pannone |
ICPRAM | 3 |
| 2018 | An augmented reality system for technical staff trainingabstractAugmented reality (AR) systems are getting more and more popular in several application fields, such as medicine, education, cultural heritage, etc. The recent hardware development of AR-oriented smartglasses extends these applications to hands-free contexts, in which the user cannot hold a tablet or a smartphone device. An example of this scenario is the ARsupported technical training of human operators. In this paper, we propose an AR smartglasses-based system for the training of technical staff working on ship engine parts. The size of components and the impossibility to use fiducial markers led to the development of an ad-hoc detection and tracking algorithm to align virtual parts over the real-world images. The system is currently under evaluation by a large multinational company leader in the field of marine market. Claudio Piciarelli, Marco Vernier, Mattia Zanier, Gian Luca Foresti |
INDIN | 4 |
| 2018 | Wide-Slice Residual Networks for Food RecognitionabstractImage-based food recognition pose new challenges for mainstream computer vision algorithms. Recent works in the field focused either on hand-crafted representations or on learning these by exploiting deep neural networks (DNN). Despite the success of DNN-based works, these exploit off-the-shelf deep architectures which are not cast to the specific food classification problem. We believe that better results can be obtained if the architecture is defined with respect to an analysis of the food composition. Following such an intuition, this work introduces a new deep scheme that is designed to handle the food structure. In particular, we focus on the vertical food traits that are common to a large number of categories (i.e., 15% of the whole data in current datasets). Towards the final objective, we first introduce a slice convolution block to capture such specific information. Then, we leverage on the recent success of deep residual blocks and combine those with the sliced convolution to produce the classification score. Extensive evaluations on three benchmark datasets demonstrated that our solution has better performance than existing approaches (e.g., a top-1 accuracy of 90.27% on the Food-101 dataset). Niki Martinel, Gian Luca Foresti, Christian Micheloni |
WACV | 2 |
| 2018 | VRheab: a fully immersive motor rehabilitation system based on recurrent neural network
Danilo Avola, Luigi Cinque, Gian Luca Foresti, Marco Raoul Marini, Daniele Pannone |
Multim. Tools Appl. | 3 |
| 2018 | Sensitivity-based region selection in the steered response power algorithm
Daniele Salvati, Carlo Drioli, Gian Luca Foresti |
Signal Process. | 3 |
| 2018 | A Low-Complexity Robust Beamforming Using Diagonal Unloading for Acoustic Source LocalizationabstractIn acoustic array processing, beamforming is a class of algorithms commonly used to estimate the position of a radiating sound source. This paper presents a diagonal unloading (DU) transformation method for the conventional response power beamforming to achieve robust localization with low computational complexity. The transformation is obtained by subtracting an opportune diagonal matrix from the covariance matrix of the array output vector. Specifically, the DU beamformer aims at subtracting the signal subspace from the noisy signal space. It is, hence, a data-dependent covariance matrix conditioning method. We show how to calculate precisely the unloading parameters, and we present a comparison of the proposed DU beamforming, the robust minimum variance distortionless response (MVDR) filter, and the multiple signal classification (MUSIC) method, in terms of their respective eigenanalyses. Theoretical analysis and experiments conducted on both simulated and real acoustic data demonstrate that the DU beamformer localization performance is comparable to that of robust MVDR and MUSIC. Since its computational cost is equivalent to that of a conventional beamformer, the proposed DU beamformer method can, thus, be very attractive due to its effectiveness and computational efficiency. Daniele Salvati, Carlo Drioli, Gian Luca Foresti |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Aerial video surveillance system for small-scale UAV environment monitoringabstractChange detection algorithms are commonly used to detect novelties for surveillance purposes in public and private places equipped by static or Pan-Tilt-Zoom (PTZ) cameras. Often, these techniques are also used as prerequisite to support more complex algorithms, including event recognition, object classification, person re-identification, and many others. With regard to small-scale Unmanned Aerial Vehicles (UAVs) at low-altitude, the change detection techniques require further investigation. In fact, most of the works currently available in the literature process video sequences acquired at very high-altitude for large-scale operations, such as vegetation monitoring, mapping of buildings, and so on. In a wide range of application contexts that require, for example, frequent monitoring or high spatial resolution for detecting small objects, video sequences acquired at high-altitude are not suitable. This paper presents a change detection system based on histogram equalization and RGB-Local Binary Pattern (RGB-LBP) operator for monitoring of wide areas by small-scale UAVs at low-altitude. Extensive experimental results show the robustness of the proposed pipeline. These latter were performed by using challenging video sequences of the public UAV Mosaicking and Change Detection (UMCD) dataset and measured a set of well-known statistical metrics. Finally, a performance analysis of the proposed algorithm is also provided. Danilo Avola, Gian Luca Foresti, Niki Martinel, Christian Micheloni, Daniele Pannone, Claudio Piciarelli |
AVSS | 2 |
| 2017 | Group Re-identification via Unsupervised Transfer of Sparse Features EncodingabstractPerson re-identification is best known as the problem of associating a single person that is observed from one or more disjoint cameras. The existing literature has mainly addressed such an issue, neglecting the fact that people usually move in groups, like in crowded scenarios. We believe that the additional information carried by neighboring individuals provides a relevant visual context that can be exploited to obtain a more robust match of single persons within the group. Despite this, re-identifying groups of people compound the common single person re-identification problems by introducing changes in the relative position of persons within the group and severe self-occlusions. In this paper, we propose a solution for group re-identification that grounds on transferring knowledge from single person reidentification to group re-identification by exploiting sparse dictionary learning. First, a dictionary of sparse atoms is learned using patches extracted from single person images. Then, the learned dictionary is exploited to obtain a sparsity-driven residual group representation, which is finally matched to perform the re-identification. Extensive experiments on the i-LIDS groups and two newly collected datasets show that the proposed solution outperforms stateof-the-art approaches. Giuseppe Lisanti, Niki Martinel, Alberto Del Bimbo, Gian Luca Foresti |
ICCV | 4 |
| 2017 | An ADAS Design based on IoT V2X Communications to Improve Safety - Case Study and IoT Architecture Reference Model
Yakusheva Nadezda, Gian Luca Foresti, Christian Micheloni |
VEHITS | 2 |
| 2017 | Adaptive bootstrapping management by keypoint clustering for background initialization
Danilo Avola, Marco Bernardi, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni |
Pattern Recognit. Lett. | 4 |
| 2017 | A keypoint-based method for background modeling and foreground detection using a PTZ camera
Danilo Avola, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni, Daniele Pannone |
Pattern Recognit. Lett. | 3 |
| 2017 | Person Reidentification in a Distributed Camera Network FrameworkabstractPlenty of research has been conducted to obtain the best reidentification performance between a single camera-pairs. None of the current approaches has addressed the reidentification in a camera network by considering the network topology (i.e., the structure of the monitored environment). We introduce a distributed network person reidentification framework which introduces the following contributions. 1) a camera matching cost to measure the reidentification performance between nodes of the network and 2) a derivation of the distance vector algorithm which allows to learn the network topology thus to prioritize and limit the cameras inquired for the matching of the probe. Results on three benchmark datasets show that the network topology can be learned in an unsupervised fashion and network-wise reidentification performance improves. As a side effect, we obtain that the communication bandwidth usage is reduced. Niki Martinel, Gian Luca Foresti, Christian Micheloni |
IEEE Trans. Cybern. | 2 |
| 2017 | Discriminant Context Information Analysis for Post-Ranking Person Re-IdentificationabstractExisting approaches for person re-identification are mainly based on creating distinctive representations or on learning optimal metrics. The achieved results are then provided in the form of a list of ranked matching persons. It often happens that the true match is not ranked first but it is in the first positions. This is mostly due to the visual ambiguities shared between the true match and other "similar" persons. At the current state, there is a lack of a study of such visual ambiguities which limit the re-identification performance within the first ranks. We believe that an analysis of the similar appearances of the first ranks can be helpful in detecting, hence removing, such visual ambiguities. We propose to achieve such a goal by introducing an unsupervised post-ranking framework. Once the initial ranking is available, content and context sets are extracted. Then, these are exploited to remove the visual ambiguities and to obtain the discriminant feature space which is finally exploited to compute the new ranking. An in-depth analysis of the performance achieved on three public benchmark data sets support our believes. For every data set, the proposed method remarkably improves the first ranks results and outperforms the state-of-the-art approaches. Jorge García 0002, Niki Martinel, Alfredo Gardel Vicente, Ignacio Bravo Muñoz, Gian Luca Foresti, Christian Micheloni |
IEEE Trans. Image Process. | 5 |
| 2016 | Mobile ocular biometrics in visible spectrum using local image descriptors: A preliminary studyabstractOcular biometrics refers to personal identification using iris, conjunctival vasculature, periocular or eye movements. Contrary to most of other biometric traits, ocular biometrics does not require high user cooperation and close capture distance. Biometrics is now adopted ubiquitously as an alternative to passwords on mobile devices. Especially, ocular biometrics in the visible spectrum has attracted a lot of attention owing to the fact that it can be acquired using the regular RGB cameras already available in all mobile devices. The use of local image descriptors (i.e., analysis of microtextural features) for ocular biometrics is gaining more and more popularity because of their compactness, computationally inexpensiveness, excellent performance and flexibility. In this work, we explore the possibility of performing large scale mobile ocular biometric recognition in the visible spectrum using local image descriptors. We design a weighted fusion scheme to combine the information originating from four different local descriptors. The experimental analysis of the devised scheme, on newly collected and publicly available large scale database using three different mobile devices, shows promising results. Zahid Akhtar, Christian Micheloni, Gian Luca Foresti |
ICIP | 3 |
| 2016 | Distributed and Unsupervised Cost-Driven Person Re-IdentificationabstractThe problem of re-identify persons across single disjoint camera-pairs has received great attention from the community. Despite this, when the re-identification process has to be carried out on a large camera network a different approach has to be considered. In particular, existing approaches have neglected the importance of the network topology (i.e., the structure of the monitored environment) in such a process. To try filling such a gap, we propose a Distributed and Unsupervised Cost-Driven Person Re-Identification framework (DUPRe) which introduces the following contributions: (i) a camera matching cost to measure the re-identification performance between nodes of the network; (ii) a derivation of the distance vector algorithm which allows to learn the network topology hence to prioritize and limit the cameras inquired for the re-identification. Results on two benchmark datasets show that our solution brings to significant network-wise re-identification improvements. Niki Martinel, Gian Luca Foresti, Christian Micheloni |
ICPR | 2 |
| 2016 | A Practical Framework for the Development of Augmented Reality Applications by using ArUco MarkersabstractThe Augmented Reality (AR) is an expanding field of the Computer Graphics (CG) that merges items of the real-world environment (e.g., places, objects) with digital information (e.g., multimedia files, virtual objects) to provide users with an enhanced interactive multi-sensorial experience of the real-world that surrounding them. Currently, a wide range of devices is used to vehicular AR systems. Common devices (e.g., cameras equipped on smartphones) enable users to receive multimedia information about target objects (non-immersive AR). Advanced devices (e.g., virtual windscreens) provide users with a set of virtual information about points of interest (POIs) or places (semi-immersive AR). Finally, an ever-increasing number of new devices (e.g., HeadMounted Display, HMD) support users to interact with mixed reality environments (immersive AR). This paper presents a practical framework for the development of non-immersive augmented reality applications through which target objects are enriched with multimedia information. On each target object is applied a different ArUco marker. When a specific application hosted inside a device recognizes, via camera, one of these markers, then the related multimedia information are loaded and added to the target object. The paper also reports a complete case study together with some considerations on the framework and future work. Danilo Avola, Luigi Cinque, Gian Luca Foresti, Cristina Mercuri, Daniele Pannone |
ICPRAM | 3 |
| 2016 | A multipurpose autonomous robot for target recognition in unknown environmentsabstractIn recent years, the technological improvements of consumer robots, in terms of processing capacity and sensors, are enabling an ever-increasing number of researchers to quickly develop both scale prototypes and alternative low cost solutions. In these contexts, a critical aspect is the design of ad-hoc algorithms according to the features of the available hardware. This paper proposes a prototype of an autonomous robot for mapping unknown environments and recognizing target objects. During the setup phase one or more target objects are shown to the RGB camera of the robot which, for each of them, extracts and stores a set of A-KAZE features. Afterwards, the robot adopts the ultrasonic distance measurement and the RGB stream to map the whole environment and search a set of A-KAZE features matchable with those previously acquired. The paper also reports both preliminary tests carried out on a reference indoor environment and a case study performed in an outdoor one that validate the proposed system. Danilo Avola, Gian Luca Foresti, Luigi Cinque, Cristiano Massaroni, Gabriele Vitale, Luca Lombardi |
INDIN | 2 |
| 2016 | Modeling feature distances by orientation driven classifiers for person re-identification
Jorge García 0002, Niki Martinel, Alfredo Gardel Vicente, Ignacio Bravo Muñoz, Gian Luca Foresti, Christian Micheloni |
J. Vis. Commun. Image Represent. | 5 |
| 2016 | A pool of multiple person re-identification expertsabstractThe person re-identification problem, i.e. recognizing a person across non-overlapping cameras at different times and locations, is of fundamental importance for video surveillance applications. Due to pose variations, illumination conditions, background clutter, and occlusions, re-identify a person is an inherently difficult problem which is still far from being solved. In this work, inspired by the recent police lineup innovations, we propose a re-identification approach where Multiple Re-identification Experts (MuRE) are trained to reliably match new probes. The answers from all the experts are then combined to achieve a final decision. The proposed method has been evaluated on three datasets showing significant improvements over state-of-the-art approaches. Niki Martinel, Christian Micheloni, Gian Luca Foresti |
Pattern Recognit. Lett. | 3 |
| 2016 | A weighted MVDR beamformer based on SVM learning for sound source localization
Daniele Salvati, Carlo Drioli, Gian Luca Foresti |
Pattern Recognit. Lett. | 3 |
| 2016 | Sound Source and Microphone Localization From Acoustic Impulse ResponsesabstractThis letter proposes a new method for source and microphone localization in reverberant environments using a randomly arranged sensor array, under the hypothesis that the position of one reference sensor and the geometry of the environment are known, and the other microphone positions are unknown. A minimum mean square error (MMSE) estimator that exploits early reflections is proposed. The MMSE estimator is solved by a grid search method that combines the information on early reflections estimated using a multichannel blind system identification and the time difference of arrivals between the reflections and the direct-path calculated with the image-source model. Simulations under different reverberant scenarios demonstrate the ability of the proposed approaches in localizing source and microphone. Daniele Salvati, Carlo Drioli, Gian Luca Foresti |
IEEE Signal Process. Lett. | 3 |
| 2016 | Introduction of New Associate EditorsabstractPresents a listing of the new Associate Editors for this issue of the publication. Nikolaos V. Boulgouris, David Bull 0001, Marco Cagnazzo, Andrea Cavallaro, Gene Cheung, Amit K. Roy-Chowdhury, Pedro Comesaña Alfaro, Sarp Ertürk, Markus Flierl, Gian Luca Foresti, Gang Hua 0001, Zhu Li 0001, Weisi Lin, Siwei Ma 0001, Pramod Kumar Meher, Debargha Mukherjee, Aleksandra Pizurica, Andrea Prati 0001, Paolo Remagnino, Arun Ross, Shin'ichi Satoh 0001, Andreas E. Savakis, Heiko Schwarz, Ling Shao 0001, Shervin Shirmohammadi, Giuseppe Valenzise, Meng Wang 0001, Zhou Wang 0001, Yonggang Wen 0001, Dong Xu 0001, Junsong Yuan 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 11 |
| 2016 | Dynamic Reconfiguration in Camera Networks: A Short SurveyabstractThere is a clear trend in camera networks toward enhanced functionality and flexibility, and a fixed static deployment is typically not sufficient to fulfill these increased requirements. Dynamic network reconfiguration helps to optimize the network performance to the currently required specific tasks while considering the available resources. Although several reconfiguration methods have been recently proposed, e.g., for maximizing the global scene coverage or maximizing the image quality of specific targets, there is a lack of a general framework highlighting the key components shared by all these systems. In this paper, we propose a reference framework for network reconfiguration and present a short survey of some of the most relevant state-of-the-art works in this field, showing how they can be reformulated in our framework. Finally, we discuss the main open research challenges in camera network reconfiguration. Claudio Piciarelli, Lukas Esterle, Asif Khan 0003, Bernhard Rinner, Gian Luca Foresti |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2015 | Artifact "Metaphors": Gaining capability using "Wrong" tools
Giovanni Ferrin, Lauro Snidaro, Gian Luca Foresti |
FUSION | 3 |
| 2015 | Data-driven vocal folds models for the representation of both acoustic and high speed video dataabstractThe aim of this paper is to evaluate the effectiveness of a class of data-driven physical models to represent both acoustic and high-speed video data of the voice production process. Voice production analysis through numerical models of the phonation process is nowday a mature research field, and reliable dynamical glottal models of different accuracy and complexity are available. Although they are traditionally used to represent the acoustic emission during phonation, the biomechanical nature of the modeling makes them well suited to also represent high speed video recordings of the vocal folds oscillations. We discuss here a data-driven, numerically simulated model of the folds motion within an audio-video data analysis context. A model structure is proposed which is based on physical knowledge and data-driven machine learning components. A model inversion algorithm is designed that exploits acoustic data related to the glottal excitation and high speed video data of the folds, to estimate the parameters of the model and to represent the phonation characteristics. It is shown here how machine learning techniques can be effectively used in combination to biomechanical modeling, in order to fit match the osbserved data. The method is assessed on data from different subjects uttering sustained vowels. Carlo Drioli, Gian Luca Foresti |
IJCNN | 2 |
| 2015 | Enhanced videokymographic data analysis based on vocal folds dynamics modelingabstractThe automatic analysis of temporal patterns of vocal folds motion and the tracking of glottal cues such as folds edge position or glottal area, has recently become a topic of interest in the field of laryngeal video imaging. We discuss here the use of a numerically simulated model of the folds motion within a video analysis context, for the analysis of videokymographic data and glottal cues segmentation. The proposed algorithm exploits both visual and acoustic data related to the glottal excitation, to estimate the parameters of the model. The trained model is then used to enhance the analysis and segmentation of visual glottal cues, i.e. The folds edge displacement and glottal area. Objective measures are reported of the accuracy with which the visual glottal cues and the acoustic voice emission are represented by the model. The method is illustrated and assessed on data from different subjects Carlo Drioli, Gian Luca Foresti |
INTERSPEECH | 2 |
| 2015 | Frequency map selection using a RBFN-based classifier in the MVDR beamformer for speaker localization in reverberant rooms
Daniele Salvati, Carlo Drioli, Gian Luca Foresti |
INTERSPEECH | 3 |
| 2015 | A neural tree for classification using convex objective function
Asha Rani 0005, Gian Luca Foresti, Christian Micheloni |
Pattern Recognit. Lett. | 2 |
| 2015 | Time-varying delay measurement of video capture-to-display components with application to visual servoing
Carlo Drioli, Gian Luca Foresti |
Signal Process. Image Commun. | 2 |
| 2015 | Kernelized Saliency-Based Person Re-Identification Through Multiple Metric LearningabstractPerson re-identification in a non-overlapping multi-camera scenario is an open and interesting challenge. While the task can hardly be completed by machines, we, as humans, are inherently able to sample those relevant persons' details that allow us to correctly solve the problem in a fraction of a second. Thus, knowing where a human might fixate to recognize a person is of paramount interest for re-identification. Inspired by the human gazing capabilities, we want to identify the salient regions of a person appearance to tackle the problem. Toward this objective, we introduce the following main contributions. A kernelized graph-based approach is used to detect the salient regions of a person appearance, later used as a weighting tool in the feature extraction process. The proposed person representation combines visual features either considering or not the saliency. These are then exploited in a pairwise-based multiple metric learning framework. Finally, the non-Euclidean metrics that have been separately learned for each feature are fused to re-identify a person. The proposed kernelized saliency-based person re-identification through multiple metric learning has been evaluated on four publicly available benchmark data sets to show its superior performance over the state-of-the-art approaches (e.g., it achieves a rank 1 correct recognition rate of 42.41% on the VIPeR data set). Niki Martinel, Christian Micheloni, Gian Luca Foresti |
IEEE Trans. Image Process. | 3 |
| 2014 | MoBio_LivDet: Mobile biometric liveness detectionabstractBiometric authentication is now being used ubiquitously as an alternative to passwords on mobile devices. However, current biometric systems are vulnerable to simple spoofing attacks. Several liveness detection methods have been proposed to determine whether there is a live person or an artificial replica in front of the biometric sensor. Yet, the problem is unsolved due to hardship in finding discriminative and computationally inexpensive features for spoofing attacks. Moreover, previous liveness detection approaches are not explicitly aimed for mobile biometric, thus principally unsuited for portable devices. Therefore, we build a software-based multi-biometric prototype that detects face, iris and fingerprint spoofing attacks on mobile devices. We present MoBio_LivDet (Mobile Biometric Liveness Detection), a novel approach that analyzes local features and global structures of the biometric images using a set of low-level feature descriptors and decision level fusion. The system allows user to balance the security level (robustness against spoofing) and convenience that they want. The proposed method is highly fast, simple, efficient, robust and does not require user-cooperation, thus making it extremely apt for mobile devices. Experimental analysis on publicly available face, iris and fingerprint data sets with real spoofing attacks show promising results. Zahid Akhtar, Christian Micheloni, Claudio Piciarelli, Gian Luca Foresti |
AVSS | 4 |
| 2014 | Person Orientation and Feature Distances Boost Re-identificationabstractMost of the open challenges in person re-identification arise from the large variations of human appearance and from the different camera views that may be involved, making pure feature matching an unreliable solution. To tackle these challenges state-of-the-art methods assume that a unique inter-camera transformation of features undergoes between two cameras. However, the combination of view points, scene illumination and photometric settings, etc., together with the appearance, pose and orientation of a person make the inter-camera transformation of features multi-modal. To address these challenges we introduce three main contributions. We propose a method to extract multiple frames of the same person with different orientation. We learn the pair wise feature dissimilarities space (PFDS) formed by the subspace of pair wise feature dissimilarities computed between images of persons with similar orientation and the subspace of pair wise feature dissimilarities computed between images of persons non-similar orientations. Finally, a classifier is trained to capture the multi-modal inter-camera transformation of pair wise images for each subspace. To validate the proposed approach we show the superior performance of our approach to state-of-the-art methods using two publicly available benchmark datasets. Jorge García 0002, Niki Martinel, Gian Luca Foresti, Alfredo Gardel Vicente, Christian Micheloni |
ICPR | 3 |
| 2014 | Incoherent Frequency Fusion for Broadband Steered Response Power Algorithms in Noisy EnvironmentsabstractThe steered response power (SRP) algorithms have been shown to be among the most effective and robust ones in noisy environments for direction of arrival (DOA) estimation. In broadband signal applications, the SRP methods typically perform their computations in the frequency-domain by applying a fast Fourier transform (FFT) on a signal portion, calculating the response power on each frequency bin, and subsequently fusing these estimates to obtain the final result. We introduce a frequency response incoherent fusion method based on a normalized arithmetic mean (NAM). Experiments are presented that rely on the SRP algorithms for the localization of motor vehicles in a noisy outdoor environment, focusing our discussion on performance differences with respect to different signal-to-noise ratios (SNR), and on spatial resolution issues for closely spaced sources. We demonstrate that the proposed fusion method provides higher resolution for the delay-and-sum SRP, and improved performances for minimum variance distortionless response (MVDR) and multiple signal classification (MUSIC). Daniele Salvati, Carlo Drioli, Gian Luca Foresti |
IEEE Signal Process. Lett. | 3 |
| 2014 | Camera Selection for Adaptive Human-Computer InterfaceabstractVideo analytics has become a very important topic in computer vision. This paper introduces advanced video analytics human-computer interfaces for a video surveillance system to ease the tasks of security operators. The visualization of the most relevant views is provided by the human-computer interface module that preemptively activates cameras that will probably cover the motion of interesting objects. Human-computer interaction principles have been considered to develop the novel user interface. Four prototypes have been designed and usability performance has been evaluated, exploiting standard methods. Results obtained from such evaluations show the efficiency of the novel information visualization technique. Niki Martinel, Christian Micheloni, Claudio Piciarelli, Gian Luca Foresti |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2013 | Context in fusion: Some considerations in a JDL perspective
Lauro Snidaro, Ingrid Visentini, James Llinas, Gian Luca Foresti |
FUSION | 4 |
| 2013 | Robust Painting Recognition and Registration for Mobile Augmented RealityabstractIn this work we introduce a novel approach for painting recognition and registration for mobile Augmented Reality applications. To address the challenges of real-time painting recognition and registration we introduce three main contributions: i) A relevant painting region detector extracts the painting region from the given image. ii) Two local and global features are extracted from the relevant region to robustly match a painting database. iii) A RANSAC homography estimation method is used to overlay the additional content in an AR framework. Experiments have been carried out on a dataset built with publicly available images. Niki Martinel, Christian Micheloni, Gian Luca Foresti |
IEEE Signal Process. Lett. | 3 |
| 2012 | Markov Logic Networks for context integration and situation assessment in maritime domain
Lauro Snidaro, Ingrid Visentini, Karna Bryan, Gian Luca Foresti |
FUSION | 4 |
| 2012 | A balanced neural tree for pattern classification
Christian Micheloni, Asha Rani 0005, Sanjeev Kumar 0001, Gian Luca Foresti |
Neural Networks | 4 |
| 2012 | Fusing multiple video sensors for surveillanceabstractReal-time detection, tracking, recognition, and activity understanding of moving objects from multiple sensors represent fundamental issues to be solved in order to develop surveillance systems that are able to autonomously monitor wide and complex environments. The algorithms that are needed span therefore from image processing to event detection and behaviour understanding, and each of them requires dedicated study and research. In this context, sensor fusion plays a pivotal role in managing the information and improving system performance. Here we present a novel fusion framework for combining the data coming from multiple and possibly heterogeneous sensors observing a surveillance area. Lauro Snidaro, Ingrid Visentini, Gian Luca Foresti |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2011 | PTZ network configuration for optimal 3D coverageabstractDuring the last years, the need for security-oriented surveillance systems has grown higher and higher. Nowadays many public environments, such as airports, train stations, etc. are monitored by some sort of video-surveillance system in order to detect or prevent security issues. The involved technology ranges from the use of plain closed-circuit cameras (CCTV) to sophisticated computer-based video processing systems. The CCTV approach has been the only feasible choice in the past, and it is still widely used, however its limits are more and more evident: the increase of the number of sensors (modern surveillance systems can use hundreds of cameras) is often not matched by an adequate number of human operators, whose attention is spread on many different tasks and quickly decreases over time. Modern computer-based systems try to face these problems using automatic video analysis and understanding techniques, in order to cover wide areas and simultaneously highlight only the potential security issues and thus requiring the attention of a human operator only in a limited number of cases (e.g. [6, 5]). The research in this field has been very active and produced many techniques for video analysis and interpretation, but many works are limited to the use of static cameras. Only recently the research community started focusing on more sophisticated sensors like Pan-Tilt-Zoom (PTZ) cameras, and the research on dynamic, active networks of PTZ cameras is still limited (for an example of some recent works in this field, see [1]). Many of these works focus on exploiting the dynamic features of a network of PTZ cameras to improve tracking performance [3, 4, 13, 10, 12], while relatively few works address the problem of optimizing the camera coverage of the monitored area according to specific criteria. Angella et al. [2] propose a method to maximize the area coverage by using a 3D model of the observed zone, but their work only aims at finding a good initial camera displacement, which cannot be dynamically modified according to the observed data. Mittal and Davis [8, 7] also consider the presence of dynamic occluding objects in order to evaluate the visibility of the scene. Piciarelli et al. [11] propose a method to automatically and dynamically reconfigure the camera orientations and zoom levels using an Expectation-Maximization-based approach. Claudio Piciarelli, Gian Luca Foresti |
AVSS | 2 |
| 2011 | Multiple acoustic sources localization using incident Signal Power comparisonabstractWe present a novel approach to locate multiple acoustic sources in far-field environments, in order to solve an interesting problem in different application domain, such as: audio surveillance systems and soundscape analysis frameworks. This approach aims at finding a solution to the ambiguities in Direction Of Arrivals (DOAs) combination caused by simultaneous multiple sources. The algorithm is based on two steps: the separation of the sources by means of beamforming techniques and the comparison of the Incident Signal Power (ISP) spectrum by means of a spectral distance measure. We implemented a prototype, composed by two linear arrays, that has been successfully tested in a real noisy environment. Daniele Salvati, Antonio Rodà, Sergio Canazza, Gian Luca Foresti |
AVSS | 4 |
| 2011 | Contexts, co-texts and situations in fusion domain
Giovanni Ferrin, Lauro Snidaro, Gian Luca Foresti |
FUSION | 3 |
| 2010 | Human Action Recognition using a Hybrid NTLD ClassifierabstractThis work proposes a hybrid classifier to recognize human actions in different contexts. In particular, the proposed hybrid classifier (a neural tree with linear discriminant nodes NTLD), is a neural tree whose nodes can be either simple preceptrons or recursive fisher linear discriminant (RFLD) classifiers. A novel technique to substitute bad trained perceptron with more performant linear discriminators is introduced. For a given frame, geometrical features are extracted from the skeleton of the human blob (silhouette). These geometrical features are collected for a fixed number of consecutive frames to recognize the corresponding activity. The resulting feature vector is adopted as input to the NTLD classifier. The performance of the proposed classifier has been evaluated on two available databases. Asha Rani 0005, Sanjeev Kumar 0001, Christian Micheloni, Gian Luca Foresti |
AVSS | 4 |
| 2010 | Revisiting the role of abductive inference in fusion domain
Giovanni Ferrin, Lauro Snidaro, Gian Luca Foresti |
FUSION | 3 |
| 2010 | Selecting classifiers by F-score for real-time video tracking
Ingrid Visentini, Lauro Snidaro, Gian Luca Foresti |
FUSION | 3 |
| 2010 | Stereo rectification of uncalibrated and heterogeneous images
Sanjeev Kumar 0001, Christian Micheloni, Claudio Piciarelli, Gian Luca Foresti |
Pattern Recognit. Lett. | 4 |
| 2009 | Stereo Localization Based on Network's Uncalibrated Camera PairsabstractIn this paper, a stereo framework for a robust real time localization of objects using networkpsilas camera pairs is presented. The stereo system contains a combination of static and pan-tilt-zoom (PTZ) cameras instead of traditional dual head mounted cameras. The proposed novelty consists in applying stereo vision to heterogeneous cameras belonging to a video-surveillance network. First, a look-up-table (LUT) is built with the rectification transformations computed for some predefined pan and tilt values. Then, the LUT is used to compute rectification transformations by means of neural networks for any arbitrary pan and tilt settings. Different zoom levels are compensated by resizing images according to their focal ratio and by applying zero padding. Localization of any object is made using its 3D position information obtained by a modified stereo concept. Experimental results are presented for the localization of moving objects in a parking lot scenario. Sanjeev Kumar 0001, Christian Micheloni, Claudio Piciarelli, Gian Luca Foresti |
AVSS | 4 |
| 2009 | Modelling and Managing Domain Context for Automatic Surveillance SystemsabstractIn this paper we propose an architecture for a surveillance application aimed to the automatic recognition of complex events. The main novelty of this work consists in the design of an effective and viable solution for the actual implementation of a complex automatic surveillance system that explicitly separates signal processing routines from the reasoning modules. We describe our ongoing efforts in the representation of the domain knowledge and in the development of a framework that allows the operator to easily check and update the systempsilas knowledge base. The taxonomical knowledge is expressed through ontologies (OWL), the event classification logic is expressed using a dedicated rule language (Jess), and the implementation is based on the java language. Lauro Snidaro, Massimo Belluz, Gian Luca Foresti |
AVSS | 3 |
| 2009 | Multi-sensor Multi-cue Fusion for Object Detection in Video SurveillanceabstractWe here present a multi-sensor data fusion architecture that takes into account the performance of video sensors in detecting moving targets for video surveillance purposes. Target detection and tracking is performed via classification by an ensemble of classifiers learned online using heterogeneous features for each target. A novel approach is then used to estimate the position of the target on the ground plane map by temporally fusing likelihood maps, then by approximating likelihoods analytically by a Gaussian function, and eventually projecting and fusing the likelihood functions. Experimental results are shown on real-world video sequences. Lauro Snidaro, Ingrid Visentini, Gian Luca Foresti |
AVSS | 3 |
| 2009 | Structuring relations for fusion in intelligence
Giovanni Ferrin, Lauro Snidaro, Gian Luca Foresti |
FUSION | 3 |
| 2009 | Exploiting temporal statistics for events analysis and understanding
Christian Micheloni, Lauro Snidaro, Gian Luca Foresti |
Image Vis. Comput. | 3 |
| 2009 | Computer vision methods for ambient intelligence
Paolo Remagnino, Gian Luca Foresti |
Image Vis. Comput. | 2 |
| 2009 | Active Tuning of Intrinsic Camera ParametersabstractIn the last years, the research effort of the scientific community to study systems for ambient intelligence has been really strong. Usually, the systems developed so far base their analysis on images acquired by automatic cameras. In this paper, we propose a way to develop new smart systems that are able to actively decide both what to see and how to see it. In particular, the main idea is to tune the acquisition parameters on the basis of what the system desires to acquire. The regulation strategy is based on two camera parameters, focus and iris. It aims to identify an optimal sequence of steps to enhance the acquisition quality of an object of interest. To this end, a hierarchy of neural networks has been employed first to select which parameter must be regulated then to adjust it. The proposed solution can be applied to both static and moving cameras. The results show how the proposed technique can be applied to images acquired by a moving camera with zoom capabilities for surveillance purposes. Christian Micheloni, Gian Luca Foresti |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2008 | Dynamic Models for People Detection and TrackingabstractIn this paper we propose a real-time algorithm for detecting and tracking moving objects in a video sequence. Based on the on-line boosting framework, our algorithm is able to detect an object as a member of a class, e.g. pedestrian, then a specific model for each instance of the class can be built on-line allowing at the same time robust tracking and recognition of the particular instance as it leaves and re-enters the scene. Promising experimental results have been performed on standard video sequences. Lauro Snidaro, Ingrid Visentini, Gian Luca Foresti |
AVSS | 3 |
| 2008 | A security assistance system combining person tracking with chemical attributes and video event analysis
Christopher Becher, Gian Luca Foresti, Peter Kaul, Wolfgang Koch 0001, Frank P. Lorenz, Daniel Lubczyk, Christian Micheloni, Claudio Piciarelli, Konstantin Safenreiter, Carsten Siering, Macarena Varela, Siegfried R. Waldvogel, Monika Wieneke |
FUSION | 2 |
| 2008 | Soft data issues in fusion of video surveillance
Giovanni Ferrin, Lauro Snidaro, Sergio Canazza, Gian Luca Foresti |
FUSION | 4 |
| 2008 | Support vector machines for robust trajectory clusteringabstractMany event analysis systems are based on the detection of uncommon feature patterns that could be associated to anomalous events; the uncommon patterns are identified by comparison with a "normality model" describing the previously acquired data. In this work we propose an anomaly detection system based on trajectory clustering with single-class support vector machines. However, SVM parameter tuning would require an a-priori estimate of the number of outlier trajectories in the training data, which is unknown. We here propose a technique for automatic estimation of the number of outliers, thus avoiding the arbitrary choice of constant tuning parameters. Claudio Piciarelli, Christian Micheloni, Gian Luca Foresti |
ICIP | 3 |
| 2008 | Anomalous trajectory patterns detectionabstractIn the field of event analysis, the detection of anomalous events has often been based on the creation of a model representing the most common patterns of activity detected within a monitored scene. This way, anomalous events can be identified by comparison with the model as patterns differing from typical events. In particular, trajectories of moving objects have often been used as a feature for anomalous event detection. In this paper we propose a combination of clustering and SVM techniques in order to automatically detect anomalous trajectories. Claudio Piciarelli, Christian Micheloni, Gian Luca Foresti |
ICPR | 3 |
| 2008 | On-line boosted cascade for object detectionabstractOn-line boosting is a recent advancement in the field of machine learning that has opened a new spectrum of possibilities in many diverse fields. With respect to a static strong classifier, the on-line algorithm updates the ensemble using new incoming samples. This idea has been successfully exploited in tasks such as detection and tracking as a classification problem with good results. Our purpose is to provide an efficient and robust framework to build a cascade of on-line updated classifiers that, speeding up the application time, allows the employment of a higher number of features, thus achieving better detection performance. Ingrid Visentini, Lauro Snidaro, Gian Luca Foresti |
ICPR | 3 |
| 2008 | Adaptive video communication for an intelligent distributed system: Tuning sensors parameters for surveillance purposes
Christian Micheloni, Marco Lestuzzi, Gian Luca Foresti |
Mach. Vis. Appl. | 3 |
| 2008 | Trajectory-Based Anomalous Event DetectionabstractDuring the last years, the task of automatic event analysis in video sequences has gained an increasing attention among the research community. The application domains are disparate, ranging from video surveillance to automatic video annotation for sport videos or TV shots. Whatever the application field, most of the works in event analysis are based on two main approaches: the former based on explicit event recognition, focused on finding high-level, semantic interpretations of video sequences, and the latter based on anomaly detection. This paper deals with the second approach, where the final goal is not the explicit labeling of recognized events, but the detection of anomalous events differing from typical patterns. In particular, the proposed work addresses anomaly detection by means of trajectory analysis, an approach with several application fields, most notably video surveillance and traffic monitoring. The proposed approach is based on single-class support vector machine (SVM) clustering, where the novelty detection SVM capabilities are used for the identification of anomalous trajectories. Particular attention is given to trajectory classification in absence ofaprioriinformation on the distribution of outliers. Experimental results prove the validity of the proposed approach. Claudio Piciarelli, Christian Micheloni, Gian Luca Foresti |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | Anomalous trajectory detection using support vector machinesabstractOne of the most promising approaches to event analysis in video sequences is based on the automatic modelling of common patterns of activity for later detection of anomalous events. This approach is especially useful in those applications that do not necessarily require the exact identification of the events, but need only the detection of anomalies that should be reported to a human operator (e.g. video surveillance or traffic monitoring applications). In this paper we propose a trajectory analysis method based on Support Vector Machines; the SVM model is trained on a given set of trajectories and can subsequently detect trajectories substantially differing from the training ones. Particular emphasis is placed on a novel method for estimating the parameter v, since it heavily influences the performances of the system but cannot be easily estimated a-priori. Experimental results are given both on synthetic and real-world data. Claudio Piciarelli, Gian Luca Foresti |
AVSS | 2 |
| 2007 | Representing and recognizing complex events in surveillance applicationsabstractIn this paper, we investigate the problem of representing and maintaining rule knowledge for a video surveillance application. We focus on complex events representation which cannot be straightforwardly represented by canonical means. In particular, we highlight the ongoing efforts for a unifying framework for computable rule and taxonomical knowledge representation. Lauro Snidaro, Massimo Belluz, Gian Luca Foresti |
AVSS | 3 |
| 2007 | Domain knowledge for surveillance applicationsabstractIn this paper, we address the problem of representing domain knowledge for situation awareness in a security application. While ontologies are appropriate for describing taxonomical knowledge, they cannot express more complex knowledge such as entailments. In this paper, we describe how domain knowledge can be encoded through OWL ontologies and SWRL rules in order to reason about the entities and their interactions in a surveillance application. We describe how events can be described through ontologies and how video sequences can be annotated using the MPEG-7 standard. Lauro Snidaro, Massimo Belluz, Gian Luca Foresti |
FUSION | 3 |
| 2007 | Tuning Asymboost Cascades Improves Face DetectionabstractThe face detection problem is certainly one of the most studied topics in artificial vision. This interest raises from the conscience that this is a crucial step for every system that uses biometric information. Video surveillance and security systems, biometrics, HCI and multimedia applications are some examples of systems that exploit face localization to improve their robustness. AdaBoost and AsymBoost based classifiers are widely used to achieve high performances saving computational time. In this paper, a new reactive strategy to build a strong classifier cascade is provided; at each stage of the cascade a different tradeoff between accuracy and computational complexity is explored. The results will show that this method is effective, and propose a way to construct a rapid and robust multipose detector. Ingrid Visentini, Christian Micheloni, Gian Luca Foresti |
ICIP (4) | 3 |
| 2007 | Expert environments: machine intelligence methods for ambient intelligenceabstractThe ideas put forward by Donald A. Norman (1999) in his monograph entitled The Invisible Computer can be considered the main source of inspiration of a new research area, called ambient intelligence. Ambient intelligence, commonly abbreviated as AmI, is primarily concerned with human–environment interactions. An environment is seen anthropomorphically, as an intelligent agent able to interact with users, creating for them processes to interpret, inform, communicate and dialogue (Abowd & Mynatt, 2000; Remagnino & Foresti, 2005). The history of ambient intelligence starts in Europe in 2001 with the Fifth European Framework Program. At that time, the IST Program Advisory Group (ISTAG) of the European Commission (Directorate General on Information Society and the Media) introduced the concept of ambient intelligence by publishing the report Scenarios for Ambient Intelligence in 2010 (Ducatel et al., 2001). Since then, ambient intelligence has been recognized in Europe as one of the key concepts related to the information and communication technology society. An updated version of the report was published in 2003 under the title Ambient Intelligence: from Vision to Reality (Ducatel et al., 2003). Ambient intelligence's emphasis is on support to human interactions with the environment, user-friendliness, ubiquitous accessibility etc. and requires competences from many research areas, ranging from computer vision, machine learning, distributed computing and middleware, context awareness systems, sensor networks etc. Enticing illustrative scenarios have been published, in which the user wears technology that communicates with systems and devices present in the environment in order to provide information and receive services. Since its inception, ambient intelligence has inspired the design and implementation of system prototypes for intelligent spaces, developed and tested in controlled environments. The scope of this special issue is to publish innovative ideas on a selection of topics related to ambient intelligence. Cameras and computer vision algorithms, for instance, can be used to unobtrusively acquire awareness of the environment. In the paper entitled ‘A multi-camera vision system for fall detection and alarm generation’, Cucchiara et al. propose a system for using cameras to detect people's falls. The use of cameras is preferred in ambient intelligence to accelerometers or other devices because they are not intrusive, people tend to get used to their presence and people's actions are less affected. In this paper, a system of multiple cameras is used to monitor people's movements in house environments and detect changes in people's posture with the specific goal to identify falls. In the paper entitled ‘Understanding intention of movement from electroencephalograms’, Lakany and Conway analyse people's intentions, using electroencephalogram waves. Their method is non-intrusive and it uses a brain–computer interface. Support vector machines are used to select features and perform classification of intentions, and tests are performed to detect the direction of users' movements. While the previous papers address ambient intelligence from the point of view of sensing technologies and algorithms, the following two papers are mainly focused on knowledge representation for context awareness. The paper entitled ‘Knowledge representation for ambient security’ by Snidaro and Foresti proposes an ontology-based methodology for representing the interaction between the user and the environment with specific reference to security scenarios. Similarly, in the paper entitled ‘Context-aware environments: from specification to implementation’ Reigner et al. are concerned with the problem of implementing a context model for a smart environment. Their paper proposes interesting approaches based on ‘networks of situations’, introducing a comparison of the use of Petri nets and hidden Markov models. Finally, in the paper entitled ‘Collection, storage and application of human knowledge in expert system development’ Balch et al. propose an analysis of the knowledge engineering flow, encompassing knowledge acquisition, representation and inference. This flow analysis is presented and applied to the petroleum industry application domain, and specific software tools that use fuzzy logic are utilized. Paolo Remagnino, Andrea Prati 0001, Gian Luca Foresti, Rita Cucchiara |
Expert Syst. J. Knowl. Eng. | 3 |
| 2007 | Knowledge representation for ambient securityabstractAbstract: Ambient intelligence envisages an articulated, though transparent, interaction between the user and the environment. According to this grand vision, appliances and systems embedded in the environment have to react to the user's presence and provide services in a customized fashion. Therefore, ambient intelligence systems should be endowed with context awareness capabilities in order to provide the proper responses for each user. This paper specifically shows how the system can be instructed to recognize events occurring in the observed environment for security purposes. Lauro Snidaro, Gian Luca Foresti |
Expert Syst. J. Knowl. Eng. | 2 |
| 2007 | Novel concepts and challenges for the next generation of video surveillance systems
Paolo Remagnino, Sergio A. Velastin, Gian Luca Foresti, Mohan M. Trivedi |
Mach. Vis. Appl. | 3 |
| 2007 | Quality-Based Fusion of Multiple Video Sensors for Video SurveillanceabstractIn this correspondence, we address the problem of fusing data for object tracking for video surveillance. The fusion process is dynamically regulated to take into account the performance of the sensors in detecting and tracking the targets. This is performed through a function that adjusts the measurement error covariance associated with the position information of each target according to the quality of its segmentation. In this manner, localization errors due to incorrect segmentation of the blobs are reduced thus improving tracking accuracy. Experimental results on video sequences of outdoor environments show the effectiveness of the proposed approach. Lauro Snidaro, Ruixin Niu, Gian Luca Foresti, Pramod K. Varshney |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2006 | Sensor Bandwidth Assignment through Video AnnotationabstractThe state of the art of surveillance systems include a large set of techniques for both low level and high level tasks. In particular, the research community has witnessed in the last decade a high proliferation of techniques that span from object detection and tracking to object recogni- tion and event understanding. Although some techniques have been proven to be very effective those tasks cannot be considered solved. Although more effort is needed in the event analysis field, a new problem arises from the develop- ment of large scale networked surveillance systems: infor- mation sharing. The way information is shared between the nodes of the surveillance network today represents a key- point issue. To provide a first and novel solution to such a problem, we propose an innovative system architecture for a video surveillance system with distributed processing over multiple processing units and with distributed communica- tion over multiple heterogeneous channels (wireless, satel- lite, local IP networks, etc.). In particular, a new real-time technique for changing the video transmission parameters (e.g., frame rate, spatial/color resolution, etc.) according to the bandwidth available will be here presented. Christian Micheloni, Lauro Snidaro, Ingrid Visentini, Gian Luca Foresti |
AVSS | 4 |
| 2006 | Fusion of trajectory clusters for situation assessmentabstractIn this paper, we address the problem of identifying anomalous events in the context of a multi sensor surveillance system. Targets' trajectories are analyzed and compared to common patterns of activity represented as clusters of trajectories. Here we extend our previous work to cater for observations provided by multiple cameras observing the same scene. Data fusion is performed within the Dempster-Shafer theory of evidence framework. The proposed approach is validated through experimental results performed in the context of an automatic road traffic monitoring application Lauro Snidaro, Claudio Piciarelli, Gian Luca Foresti |
FUSION | 3 |
| 2006 | Activity Analysis for Video Security SystemsabstractVideo security systems are required to extract semantic information on the ongoing activities in a given environment, in order to detect suspicious or potentially dangerous behaviours. To achieve this goal we proposed a real-time trajectory clustering algorithm able to detect common patterns of activity; the output of this algorithm can be used to facilitate high level behaviour analysis and situation assessment. In this paper we extend our previous work using the Dempster-Shafer theory of evidence framework. This allows an explicit handling of evidence accrual and fusion that gives more meaningful data for high-level activity analysis modules. Lauro Snidaro, Claudio Piciarelli, Gian Luca Foresti |
ICIP | 3 |
| 2006 | Real-time image processing for active monitoring of wide areas
Christian Micheloni, Gian Luca Foresti |
J. Vis. Commun. Image Represent. | 2 |
| 2006 | Growing Hierarchical Tree SOM: An unsupervised neural network with dynamic topology
Alberto Forti, Gian Luca Foresti |
Neural Networks | 2 |
| 2006 | On-line trajectory clustering for anomalous events detection
Claudio Piciarelli, Gian Luca Foresti |
Pattern Recognit. Lett. | 2 |
| 2005 | An integrated surveillance system for outdoor securityabstractAn integrated system for the detection, active tracking and recognition of people in wide outdoor environments is hereafter discussed. Specifically, a static sensor with a wide view is used to detect people inside the environment and to classify their behaviours. As outcome of anomalous activities, an active camera is selected to focus its attention on a particular person. Here, techniques for face detection are employed to determine a region of interest where to extract features used by a tracking algorithm for an autonomous gaze of a PTZ camera. Finally, a face recognition phase is considered to recognize the person of interest. Results show how the integrated system is able to detect, track and recognise people inside a tough environment such as a parking lot. Christian Micheloni, Elena Salvador, Flavio Bigaran, Gian Luca Foresti |
AVSS | 4 |
| 2005 | Trajectory clustering and its applications for video surveillanceabstractIn this paper we present a trajectory clustering method suited for video surveillance and monitoring systems. The clusters are dynamic and built in real-time as the trajectory data is acquired, without the need of an off-line processing step. We show how the obtained clusters can be successfully used both to give proper feedback to the low-level tracking system and to collect valuable information for the high-level event analysis modules. Claudio Piciarelli, Gian Luca Foresti, Lauro Snidaro |
AVSS | 2 |
| 2005 | Zoom on target while trackingabstractIn this paper, the problem of continuous tracking of moving objects with a PTZ camera is addressed. In particular, the problem of tracking moving objects during zoom phases is solved by using a feature clustering technique. In order to adopt such a method, we need, first, a step where during tracking with a pan&tilt camera we can identify the mobile objects in the monitored scene. Therefore, a set of good trackable features belonging to the selected target is extracted. In this research, we adopt a feature clustering method that is able to discriminate between features associated with the background and features associated with different moving objects. As a result, for each moving object, we have a set of correctly tracked features that is used to track the objects. Experiments have been performed on outdoor environments where either people or vehicles have been tracked. The results highlight how such a technique can be included in a more complex system able to maintain targets in the field of view of the camera, and to zoom on an object of interest when desired. Christian Micheloni, Gian Luca Foresti |
ICIP (3) | 2 |
| 2005 | A multi-camera approach to sensor evaluation in video surveillanceabstractIn this paper, we provide an objective sensor performance evaluation approach to automatically carry out the sensor evaluation task in a multi-sensor system for video surveillance. That is, given a set of sensors monitoring the same area, the proposed technique estimates the accuracy of each sensor in detecting a given target. This measure depends on the conditions under which the sensor is operating and is computed per sensor and per target at every time instant. In this way, data coming from the different sensors may be weighted accordingly and fused. Promising experimental results are presented on real video sequences. Lauro Snidaro, Gian Luca Foresti |
ICIP (1) | 2 |
| 2005 | Detecting moving people in video streams
Gian Luca Foresti, Christian Micheloni, Claudio Piciarelli |
Pattern Recognit. Lett. | 1 |
| 2004 | A new feature clustering method for object detection with an active cameraabstractFeature based methods for ego-motion estimation are widely used in computer vision but they must deal with errors in feature tracking. In this paper, we propose a robust real-time method for ego-motion estimation by assuming an affine motion of the background from the previous to the current frame. A new clustering technique is applied on image's subareas to select in a fast and reliable way three features for the affine transform computation. The previous frame after being warped according to the computed affine transform is processed with the current frame by a change detection method in order to detect mobile objects. Results are presented in the context of a visual-based surveillance system for monitoring outdoor environments. Christian Micheloni, Gian Luca Foresti, Flavio Alberti |
ICIP | 2 |
| 2004 | A clustering fuzzy approach for image segmentation
Luigi Cinque, Gian Luca Foresti, Luca Lombardi |
Pattern Recognit. | 2 |
| 2004 | An adaptive high-order neural tree for pattern recognitionabstractA new neural tree model, called adaptive high-order neural tree (AHNT), is proposed for classifying large sets of multidimensional patterns. The AHNT is built by recursively dividing the training set into subsets and by assigning each subset to a different child node. Each node is composed of a high-order perceptron (HOP) whose order is automatically tuned taking into account the complexity of the pattern set reaching that node. First-order nodes divide the input space with hyperplanes, while HOPs divide the input space arbitrarily, but at the expense of increased complexity. Experimental results demonstrate that the AHNT generalizes better than trees with homogeneous nodes, produces small trees and avoids the use of complex comparative statistical tests and/or a priori selection of large parameter sets. Gian Luca Foresti, T. Dolso |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2004 | Automatic visual recognition of deformable objects for grasping and manipulationabstractThis paper describes a vision-based system that is able to automatically recognize deformable objects, to estimate their pose, and to select suitable picking points. A hierarchical self-organized neural network is used to segment color images based on texture information. A morphological analysis allows the recognition of the objects and the picking points extraction. The proposed approach is useful in all of the situations where texture properties are significant for detecting regions of interest on deformable objects. Several tests on a large number of images, acquired in real operative working conditions, demonstrate the effectiveness of the system. Gian Luca Foresti, Felice Andrea Pellegrino |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2003 | Fast Good Features Selection for Wide Area MonitoringabstractRecently the surveillance of wide areas has pointed the interest of the research community. The use of active vision seems to be the most effective solutions for these needs. Against the better acquiring resolution there is the problem of the apparent motion inducted by the camera motion known as ego-motion. Feature based methods for ego-motion estimation are widely used in computer vision but they deal with feature recovery and with errors in feature tracking. In this paper, we propose a fast method to extract and select new features during camera motion. This is achieved by adopting a reference map containing well trackable features that is updated at each frame by introducing new good features related to regions appearing in the current image. A new procedure is applied to reject badly tracked features. The current frame and the background after compensation are processed by a change detection method in order to locate mobile objects. Results are presented in the context of a visual-based surveillance system for monitoring outdoor environments. Christian Micheloni, Gian Luca Foresti |
AVSS | 2 |
| 2003 | Automatic Camera Selection and Fusion for Outdoor Surveillance under Changing Weather ConditionsabstractAn outdoor multi-camera video surveillance system operating under changing weather conditions is presented. A new confidence measure, appearance ratio (AR), is defined to evaluate automatically the sensors' performance for each time instant. By comparing their ARs, the system can select the most appropriate cameras to perform specific tasks. When redundant measurements are available for a target, the AR measures are used to perform a weighted fusion of them. Experimental results are presented on outdoor scenes under different weather conditions. Lauro Snidaro, Ruixin Niu, Pramod K. Varshney, Gian Luca Foresti |
AVSS | 4 |
| 2003 | A robust face detection system for real environmentsabstractIn this paper, a robust real-time face detection system based on the integration of different location methods is proposed. A hierarchical architecture composed of three levels is designed. At the first level, a change detection method is applied to detect blobs of moving objects (i.e., humans) in the scene. Then, the silhouette of each blob is analyzed to focalize the attention of the system on small image areas where the probability of finding human heads is high. At the second level, two different methods, i.e., the skin color and the principal component analysis, are applied to locate human faces. Finally, the higher level fuses the obtained location data to improve the face detection reliability. The system is tested in outdoor environments in the context of a video-based surveillance system. Gian Luca Foresti, Christian Micheloni, Lauro Snidaro |
ICIP (3) | 1 |
| 2003 | Real-time thresholding with Euler numbers
Lauro Snidaro, Gian Luca Foresti |
Pattern Recognit. Lett. | 2 |
| 2002 | A distributed sensor network for video surveillance of outdoor environmentsabstractA distributed sensor network (DSN) for video surveillance is presented. The system is able to manage heterogeneous sensors (e.g. optical, infrared, radar, etc.) to operate during night and day and in the presence of different weather conditions (e.g. fog, rain, etc.). Data fusion is therefore mandatory and exploited at different levels to integrate the information produced by the sensors. The architecture presented has low network requirements, is easily scalable, maintainable and allows an easy distribution of the system on a wide outdoor area. Gian Luca Foresti, Lauro Snidaro |
ICIP (1) | 1 |
| 2002 | Automatic detection and indexing of video-event shots for surveillance applicationsabstractIncreased communication capabilities and automatic scene understanding allow human operators to simultaneously monitor multiple environments. Due to the amount of data to be processed in new surveillance systems, the human operator must be helped by automatic processing tools in the work of inspecting video sequences. In this paper, a novel approach allowing layered content-based retrieval of video-event shots referring to potentially interesting situations is presented. Interpretation of events is used for defining new video-event shot detection and indexing criteria. Interesting events refer to potentially dangerous situations: abandoned objects and predefined human events are considered in this paper. Video-event shot detection and indexing capabilities are used for online and offline content-based retrieval of scenes to be detected. Gian Luca Foresti, Lucio Marcenaro, Carlo S. Regazzoni |
IEEE Trans. Multim. | 1 |
| 2002 | Generalized neural trees for pattern classificationabstractIn this paper, a new neural tree (NT) model, the generalized NT (GNT), is presented. The main novelty of the GNT consists in the definition of a new training rule that performs an overall optimization of the tree. Each time the tree is increased by a new level, the whole tree is reevaluated. The training rule uses a weight correction strategy that takes into account the entire tree structure, and it applies a normalization procedure to the activation values of each node such that these values can be interpreted as a probability. The weight connection updating is calculated by minimizing a cost function, which represents a measure of the overall probability of correct classification. Significant results on both synthetic and real data have been obtained by comparing the classification performances among multilayer perceptrons (MLPs), NTs, and GNTs. In particular, the GNT model displays good classification performances for training sets having complex distributions. Moreover, its particular structure provides an easily probabilistic interpretation of the pattern classification task and allows growing small neural trees with good generalization properties. Gian Luca Foresti, Christian Micheloni |
IEEE Trans. Neural Networks | 1 |
| 2002 | Invariant feature extraction and neural trees for range surface classificationabstractIn this paper, a neural tree-based approach for classifying range images into a set of nonoverlapping regions is presented. An innovative procedure is applied to extract invariant surface features from each pixel of the range image. These features are: 1) robust to noise, and 2) invariant to scale, shift, rotations, curvature variations, and direction of the normal. Then, a generalized neural tree is used to classify each image point as belonging to one of the six surface models of differential geometry, i.e., peak, ridge, valley, saddle, pit, and flat. Comparisons with other methods and experiments on both synthetic and real three-dimensional range images are proposed. Gian Luca Foresti |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2001 | Use of neural networks for behaviour understanding in railway transport monitoring applicationsabstractInterest for advanced video-based surveillance applications has been growing rapidly. This is especially true in the field of railway urban transport where video-based surveillance can be exploited to face many relevant security aspects (e.g. vandal acts, overcrowding situations, abandoned object detection, etc.). This paper investigates an open problem in the implementation of video-based surveillance systems for transport applications, i.e.: the implementation of reliable image understanding modules in order to recognize dangerous situations with reduced false alarm and misdetection rates. We consider the use of a neural network-based classifier for detecting the behavior of vandals in metro stations. The achieved results show that the classifier choice mentioned above allows one to achieve very good performances also in the presence of high scene complexity. Claudio Sacchi, Carlo S. Regazzoni, Gianluca Gera, Gian Luca Foresti |
ICIP (1) | 4 |
| 2001 | Distributed architectures and logical-task decomposition in multimedia surveillance systemsabstractIn the past few years, the development of complex surveillance systems has captured the interest of both the research and industrial worlds. Strong and challenging requirements of modern society are involved in this problem, which aims to increase safety and security in several application domains such as transport, tourism, home and bank security, military applications, etc. At the same time, fast improvements in microelectronics, telecommunications, and computer science make it necessary to consider new perspectives in this field. The main objective of this paper is to investigate, discuss, and evaluate the impact of distributed processing and new communication techniques on multimedia surveillance systems, which represent the so-called third-generation surveillance systems (3 GSSs). In particular, aspects related to the distribution of intelligence among multiple-processing and wide-bandwidth resources are discussed in detail. It is shown how distribution of intelligence can be obtained by a hierarchical architecture that partitions, in a dynamic way, the main logical processing tasks (i.e., representation, recognition, and communication) performed in a 3 GSS physical architecture made up of intelligent cameras, hubs, and central control rooms. The advantages of this solution are pointed out in terms of 1) increased flexibility and reconfigurability and 2) optimal allocation of available processing and bandwidth resources. Finally, a case study is analyzed that allows one to gain a deeper insight into a distributed surveillance system. Lucio Marcenaro, Franco Oberti, Gian Luca Foresti, Carlo S. Regazzoni |
Proc. IEEE | 3 |
| 2001 | Visual inspection of sea bottom structures by an autonomous underwater vehicleabstractThis paper describes a vision-based system for inspections of underwater structures, e.g., pipelines, cables, etc., by an autonomous underwater vehicle (AUV). Usually underwater inspections are performed by remote operated vehicles (ROVs) driven by human operators placed in a support vessel. However, this task is often challenging, especially in conditions of poor visibility or in presence of strong currents. The system proposed allows the AUV to accomplish the task in autonomy. Moreover, the use of a three-dimensional (3-D) model of the environment and of an extended Kalman filter (EKF) allows the guidance and the control of the vehicle in real time. Experiments done on real underwater images have demonstrated the validity of the proposed method and its efficiency in the case of critical and complex situations. Gian Luca Foresti |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2000 | Generalized Neural Trees for Outdoor Scene UnderstandingabstractA new model of a neural tree, called generalized neural tree (GNT), is presented. In the GNT learning process, the whole tree structure is considered at each learning step, and the entire training set is used to update each node. The main novelty of the proposed approach is that the output obtained when a pattern is presented to the network has a probabilistic interpretation. Experimental tests have been performed by applying the GNT in the context of a visual-based surveillance system for outdoor scenes. In particular, objects moving in the observed scene are firstly classified into 5 different categories. Then, the trajectory of such objects, together with the class information is provided to a second GNT which gives a final interpretation of the scene in terms of presence of dangerous situations. Gian Luca Foresti, Walter Vanzella |
ICIP | 1 |
| 2000 | A Vision Based System for Object Detection in Underwater ImagesabstractIn this paper, a vision-based system for underwater object detection is presented. The system is able to detect automatically a pipeline placed on the sea bottom, and some objects, e.g. trestles and anodes, placed in its neighborhoods. A color compensation procedure has been introduced in order to reduce problems connected with the light attenuation in the water. Artificial neural networks are then applied in order to classify in real-time the pixels of the input image into different classes, corresponding e.g. to different objects present in the observed scene. Geometric reasoning is applied to reduce the detection of false objects and to improve the accuracy of true detected objects. The results on real underwater images representing a pipeline structure in different scenarios are shown. The presence of seaweed and sand, different illumination conditions and water depth, different pipeline diameter and small variations of the camera tilt angle are considered to evaluate the algorithm performances. Gian Luca Foresti, Stefania Gentili |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2000 | A hierarchical approach to feature extraction and groupingabstractIn this paper, the problem of extracting and grouping image features from complex scenes is solved by a hierarchical approach based on two main processes: voting and clustering. Voting is performed for assigning a score to both global and local features. The score represents the evidential support provided by input data for the presence of a feature. Clustering aims at individuating a minimal set of significant local features by grouping together simpler correlated observations. It is based on a spatial relation between simple observations on a fixed level, i.e., the definition of a distance in an appropriate space. As the multilevel structure of the system implies that input data for an intermediate level are outputs of the lower level, voting can be seen as a functional representation of the "part-of" relation between features at different abstraction levels. The proposed approach has been tested on both synthetic and real images and compared with other existing feature grouping methods. Gian Luca Foresti, Carlo S. Regazzoni |
IEEE Trans. Image Process. | 1 |
| 1999 | A Probabilistic Approach to Object Classification by Neural TreesabstractIn this paper, a probabilistic approach is followed to improve the classification performances of neural trees. A neural tree whose nodes are generalized perceptrons without hidden layers and with activation function characterized by a sigmoidal behaviour is considered. The standard classification method may sometimes end up with wrong conclusions, e.g., the pattern is close to one of the decision hyperplanes. This situation occurs when the activation vector in one or more internal nodes (doubt nodes) of the NT is characterized by some values close to the highest value. To this end, multiple paths are followed by appropriately backtracking into the tree. A gain function is assigned to each node and a probabilistic technique to search for the path which maximizes this gain is proposed to improve the classification performances of the standard NT. Gian Luca Foresti |
ICIP (1) | 1 |
| 1999 | Noise-Robust and Invariant Object Classification by the High-Order Statistical Pattern SpectrumabstractA new shape descriptor, the high order statistical pattern spectrum (HSP), able to extract from real images a set of descriptive features which can be used to classify objects regardless of their positions, sizes, orientations and the presence of noise, has been developed. The HSP is an internal, noise-robust, noninformation-preserving operator which combines the properties of invariance of the high order pattern spectrum and the properties of noise robustness of the statistical pattern spectrum. A neural network trained by a back-propagation algorithm has been used to test the method on different classification problems. Experimental results are presented on both synthetic and real images corrupted by various levels of noise and containing an object in different positions. Comparisons with other existing shape descriptor operators have been also performed. Gian Luca Foresti, Stefania Gentili |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1999 | Outdoor Scene Classification by a Neural Tree-Based Approach
Gian Luca Foresti |
Pattern Anal. Appl. | 1 |
| 1999 | Object recognition and tracking for remote video surveillanceabstractA system for real-time object recognition and tracking for remote video surveillance is presented. In order to meet real-time requirements, a unique feature, i.e., the statistical morphological skeleton, which achieves low computational complexity, accuracy of localization, and noise robustness has been considered for both object recognition and tracking. Recognition is obtained by comparing an analytical approximation of the skeleton function extracted from the analyzed image with that obtained from model objects stored into a database. Tracking is performed by applying an extended Kalman filter to a set of observable quantities derived from the detected skeleton and other geometric characteristics of the moving object. Several experiments are shown to illustrate the validity of the proposed method and to demonstrate its usefulness in video-based applications. Gian Luca Foresti |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1998 | A Line Segment Based Approach for 3D Motion Estimation and Tracking of Multiple ObjectsabstractA line segment based approach for 3D motion estimation and tracking of multiple objects from a monocular image sequence is presented. Objects are described by means of 3D line segments, and their presence in the scene is associated with the detection of 2D line segments on the image plane. A change detection algorithm is applied to detect moving objects on the image plane and a Hough-based algorithm is used to individuate 2D line segments. 3D parameters of each line segment are estimated, at each time instant, by means of an extended Kalman filter (EKF), whose observations are the displacements of 2D line segment endpoints on the image plane. Results on both synthetic and real scenes are presented. Gian Luca Foresti |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1998 | Exploiting neural trees in range image understanding
Gian Luca Foresti, Goffredo G. Pieroni |
Pattern Recognit. Lett. | 1 |
| 1998 | A real-time system for video surveillance of unattended outdoor environmentsabstractThis paper describes a visual surveillance system for remote monitoring of unattended outdoor environments. The system, which works in real time, is able to detect, localize, track, and classify multiple objects moving in a surveilled area. The object classification task is based on a statistical morphological operator, the statistical pecstrum (called specstrum), which is invariant to translations, rotations, and scale variations, and it is robust to noise. Classification is performed by matching the specstrum extracted from each detected object with the specstra extracted from multiple views of different real object models contained in a large database. Outdoor images are used to test the system in real functioning conditions. Performances about good classification percentage, false and missed alarms, viewpoint invariance, noise robustness, and processing time are evaluated. Gian Luca Foresti |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1997 | 3D Object Recognition by Neural TreesabstractIn this paper, a two stage method for 3D object recognition from range images is presented. The first stage extracts local surface features from the input range images. These features are used in the second stage to group image pixels into different surface patches according to the six surface classes proposed by the differential geometry. A neural tree architecture whose nodes are perceptrons without hidden layers and with sigmoidal activation functions is used. A new strategy is proposed to split the training set when it is not linearly separable in order to assure the convergence of the tree learning process. This method has been successfully applied to a large number of synthetic and real images, some of which are presented in the result section. Gian Luca Foresti, Goffredo G. Pieroni |
ICIP (3) | 1 |
| 1997 | Object Pose Estimation in Underwater Acoustic ImagesabstractWe address the problem of the recognition of man-made objects and the estimation of the related orientation in 2D acoustic images acquired with a forward looking sonar or an acoustic camera. A voting-based approach is described that is able to recognize objects and to estimate their two-dimensional pose by using information coming from boundary segments and their angular relations. The method is directly applied to the edge discontinuities of underwater acoustic images, whose quality is usually affected by some undesired effects such as object blurring, speckle noise, and geometrical distortions degrading the edge detection. The voting approach is robust with respect to these effects, so that good results are obtained even with images of poor quality. The sequences of simulated and real acoustic images are presented in order to test the validity of the proposed method in terms of the average estimation error and computational load. Vittorio Murino, Gian Luca Foresti, Andrea Trucco |
ICIP (1) | 2 |
| 1997 | A Belief-Based Approach for Adaptive Image ProcessingabstractThis paper proposes a new approach to the problem of intelligently regulating image-processing parameters of a distributed network. The proposed approach is based on two-step probabilistic process: (a) belief updating, which consists in computing a functional cost at each node of the network and, (b) belief maximization, which depends on maximizing this functional cost by using a stochastic optimization algorithm. The architecture of an image processing system, consisting of three modules connected in a chain-like structure, is presented as an example showing the capabilities of the proposed approach. Each module is provided with a priori information about the set of parameters that manage a particular data transformation, and with evaluation criteria to judge data quality and to decide on the parameters to be adjusted. Experimental results obtained by using a digitally controlled camera and lens objective, are presented to show the validity of the proposed approach. Vittorio Murino, Gian Luca Foresti, Carlo S. Regazzoni |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 1997 | 2D into 3D Hough-space mapping for planar object pose estimation
Vittorio Murino, Gian Luca Foresti |
Image Vis. Comput. | 2 |
| 1997 | A real-time model-based method for 3-D object orientation estimation in outdoor scenesabstractThis letter presents a new method for real-time determination of the three-dimensional (3-D) orientation, i.e., the rotation angle with respect to one of the principal object axes, of a moving known object from a monocular image sequence. The method is composed by three steps: 1) extraction of the morphological skeleton from binary images, 2) projection of the skeleton function on two planes and its analytical approximation by nonuniform rational B-splines, and 3) comparison with a set of data stored into a model database. Several experiments performed on real images prove the method's validity. Gian Luca Foresti, Carlo S. Regazzoni |
IEEE Signal Process. Lett. | 1 |
| 1996 | Properties of binary statistical morphologyabstractThe properties and applications of a class of statistical morphological operators, i.e. binary statistical morphology (BSM) operators, for binary image processing are described. The proposed operators are based on quantization of the output of a statistical morphological operator, modeled as a binary probabilistic hypothesis-testing step. The operator obtained is shown to be equivalent to a rank-order filter. Relationships are established between the quantization threshold, rank of the equivalent rank-order filter and parameters of the model. It is also shown that basic BSM operators, i.e. binary statistical dilation and binary statistical erosion can be used as the basis for defining more complex filters. In this paper, attention is paid to describe specific properties of BSM operators which are useful for different applications, e.g. shape description. Carlo S. Regazzoni, Gian Luca Foresti |
ICPR | 2 |
| 1996 | Grouping as a Searching Process for Minimum-Energy Configurations of Labelled Random Fields
Vittorio Murino, Carlo S. Regazzoni, Gian Luca Foresti |
Comput. Vis. Image Underst. | 3 |
| 1996 | A Hough-Based Matching of 2D Line Segments in a Monocular Image SequenceabstractThe paper describes a method for detecting 2D straight segments and their correspondences in successive frames of an image sequence by means of a Hough-based matching approach. The main advantage of this method is the possibility of extracting and matching 2D straight segments directly in the feature space, without the need for complex matching operations and time-consuming inverse transformations. An additional advantage is that only four attributes of 2D straight segments are required to perform an efficient matching process: position, orientation, length, and midpoint. Tests were performed on both synthetic and real images containing complex man-made objects moving in a scene. A comparison with a well-known 2D line matching algorithm is also made. Gian Luca Foresti, Carlo S. Regazzoni |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1996 | A distributed probabilistic system for adaptive regulation of image processing parametersabstractA distributed optimization framework and its application to the regulation of the behavior of a network of interacting image processing algorithms are presented. The algorithm parameters used to regulate information extraction are explicitly represented as state variables associated with all network nodes. Nodes are also provided with message-passing procedures to represent dependences between parameter settings at adjacent levels. The regulation problem is defined as a joint-probability maximization of a conditional probabilistic measure evaluated over the space of possible configurations of the whole set of state variables (i.e., parameters). The global optimization problem is partitioned and solved in a distributed way, by considering local probabilistic measures for selecting and estimating the parameters related to specific algorithms used within the network. The problem representation allows a spatially varying tuning of parameters, depending on the different informative contents of the subareas of an image. An application of the proposed approach to an image processing problem is described. The processing chain chosen as an example consists of four modules. The first three algorithms correspond to network nodes. The topmost node is devoted to integrating information derived from applying different parameter settings to the algorithms of the chain. The nodes associated with data-transformation processes to be regulated are represented by an optical sensor and two filtering units (for edge-preserving and edge-extracting filterings), and a straight-segment detection module is used as an integration site. Vittorio Murino, Gian Luca Foresti, Carlo S. Regazzoni |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 1995 | 3D pose estimation and shape coding of moving objects based on statistical morphological skeletonabstractRecognition-based tracking and image coding methods are often based on different techniques. However, these techniques share the necessity of reliable and fast methods for the representation of the information content of the scene. In this paper, a method based on a lossy shape descriptor is presented which can be used for both recognition and coding purposes. The statistical morphological skeleton provides a noise-robust shape descriptor on which a further approximation phase is performed in order to improve the compression ratio. Then, it is possible to estimate the object's pose through a comparison of the shape descriptor with a set of object models stored in a database. An application to surveillance is presented where the obtained description is used to transmit shape information to a remote control center. Carlo S. Regazzoni, Gian Luca Foresti, Anastasios N. Venetsanopoulos |
ICIP (3) | 2 |
| 1995 | A Multilevel Fusion Approach to Object Identification in Outdoor Road ScenesabstractThe task of object identification is fundamental to the operations of an autonomous vehicle. It can be accomplished by using techniques based on a Multisensor Fusion framework, which allows the integration of data coming from different sensors. In this paper, an approach to the synergic interpretation of data provided by thermal and visual sensors is proposed. Such integration is justified by the necessity for solving the ambiguities that may arise from separate data interpretations. The architecture of a distributed Knowledge-Based system is described. It performs an Intelligent Data Fusion process by integrating, in an opportunistic way, data acquired with a thermal and a video (b/w) camera. Data integration is performed at various architecture levels in order to increase the robustness of the whole recognition process. A priori models allow the system to obtain interesting data from both sensors; to transform such data into intermediate symbolic objects; and, finally, to recognize environmental situations on which to perform further processing. Some results are reported for different environmental conditions (i.e. a road scene by day and by night, with and without the presence of obstacles). Vittorio Murino, Carlo S. Regazzoni, Gian Luca Foresti, Gianni Vernazza |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 1995 | Circular arc extraction by direct clustering in a 3D Hough parameter space
Gian Luca Foresti, Carlo S. Regazzoni, Gianni Vernazza |
Signal Process. | 1 |
| 1995 | A Gibbs Markov random field model for active imaging at microwave frequencies
Carlo S. Regazzoni, Gian Luca Foresti |
Signal Process. | 2 |
| 1994 | Statistical morphological filters for binary image processingabstractA new class of statistical morphological operators for binary image processing is introduced. These operators are based on a digitized version of the mean field approximation. The main advantage of the new operators is provided by the capability of taking into account both noise and shape information. Binary statistical dilation (BSD) and binary statistical erosion (BSE) are considered as a case study. Extensivity properties of BSD and BSE are also discussed.> Carlo S. Regazzoni, Anastasios N. Venetsanopoulos, Gian Luca Foresti, Gianni Vernazza |
ICASSP (5) | 3 |
| 1994 | Shape Representation from Image Sequences by using Binary Statistical MorphologyabstractA real-time visual surveillance system is based on three main image processing phases, devoted to extract information about the observed scene: change detection, focus of attention, feature-extraction. In this paper attention is paid to a theory (i.e., binary statistical morphology) which provides a common framework for designing fast and noise-robust methods for the three tasks of interest. The main theoretical novelty is to establish a link between binary statistical morphology and voting methods. An application is presented which deals with intruder detection in a railway-crossing area.> Carlo S. Regazzoni, Gian Luca Foresti, Anastasios N. Venetsanopoulos |
ICIP (2) | 2 |
| 1993 | Distributed spatial reasoning for multisensory image interpretation
Gian Luca Foresti, Vittorio Murino, Carlo S. Regazzoni, Gianni Vernazza |
Signal Process. | 1 |
| 1991 | A numerical and symbolic fusion method for interpretation of image sequenceabstractThe problem of analyzing a sequence of images by taking into account the symbolic and numerical content of the signal is considered. An algorithm for segmentation and tracking of regions among images of a sequence is presented. The method is based on a Gibbs Markov random field (GMRF) model which couples the image process to a spatial-temporal region process. Optical flow field is used to adaptively decide the temporal clique to be used during annealing of energy. Displacement vectors of pixels belonging to recognized regions are predicted by using prior knowledge about object behavior.> Fabio Arduini, R. Cabri, Gian Luca Foresti, Vittorio Murino, Carlo S. Regazzoni |
ICASSP | 3 |