EDBT 2026 Demo / reviewers in the wild / expert
Francesca Odone
dblp:73/2633
· DBLP profile ↗
68ranked-venue papers
2as first author
20since 2021 · last 2026
0000-0002-3463-2263ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 2 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | P4: Place with Purpose - Pose and Prompt-Guided Human Synthesis in Real Scenes
Dadan Khan, Mohammad Zohaib, Radu Timofte, Francesca Odone |
ICPR (2) | 4 |
| 2026 | FiloAnalyzer: a deep learning approach for cell filopodia segmentationabstractPURPOSE: Filopodia are thin, finger-like extensions that project from the surface of cells. Present in various cell types, analyzing these structures can yield significant insights into cellular behavior and function. However, a manual analysis is impractical, resulting in a bottleneck in processing large-scale datasets. METHODS: This paper introduces FiloAnalyzer, an open-source toolbox for automatically analyzing cell filopodia. The toolbox is powered with a user-friendly GUI, allowing the segmentation of batches of microscopy images and extracting a set of quantitative parameters, incorporating both a deep learning-based and an image-processing method. As data and annotation scarcity are typical of this domain, we compare popular ImageNet supervised to self-supervised pre-training, adopting state-of-the-art contrastive, self-distillation, and generative approaches based on diffusion models in an attempt to maximize segmentation performance in a low-data regime. RESULTS: We validate our toolbox on a dataset of annotated fluorescent images of neural crest cells, benchmarking our results with three popular filopodia segmentation toolboxes. Furthermore, we show how self-supervision can be a promising tool to deal with the limited availability of annotated data in this domain, obtaining a significant improvement over ImageNet supervised pre-training, with as few as ten annotated images. CONCLUSION: Our experiments show that the deep learning pipeline outperforms image processing-based available alternatives. We release both an annotated dataset and the toolbox as open-source to foster further research on this topic. The toolbox code and dataset are available at https://github.com/Malga-Vision/FiloAnalyzerToolbox. Vito Paolo Pastore, Riccardo Rorato, Larbi Touijer, Roberto Di Via, Francesca Odone, Lisa M. Galli, Laura W. Burrus, Simone Bianco 0002 |
BMC Bioinform. | 5 |
| 2025 | Disentangled representations of microscopy imagesabstractMicroscopy image analysis is fundamental for different applications, from diagnosis to synthetic engineering and environmental monitoring. Modern acquisition systems have granted the possibility to acquire an escalating amount of images, requiring a consequent development of a large collection of deep learning-based automatic image analysis methods. Although deep neural networks have demonstrated great performance in this field, interpretability — an essential requirement for microscopy image analysis — remains an open challenge.This work proposes a Disentangled Representation Learning (DRL) methodology to enhance model interpretability for microscopy image classification. Exploiting benchmark datasets from three different microscopic image domains (plankton, yeast vacuoles, and human cells), we show how a DRL framework, based on transferring a representation learnt from synthetic data, can provide a good trade-off between accuracy and interpretability in this domain. Jacopo Dapueto, Vito Paolo Pastore, Nicoletta Noceti, Francesca Odone |
IJCNN | 4 |
| 2025 | Diffusing DeBias: Synthetic Bias Amplification for Model DebiasingabstractThe effectiveness of deep learning models in classification tasks is often challenged by the quality and quantity of training data whenever they are affected by strong spurious correlations between specific attributes and target labels. This results in a form of bias affecting training data, which typically leads to unrecoverable weak generalization in prediction. This paper addresses this problem by leveraging bias amplification with generated synthetic data only: we introduce Diffusing DeBias (DDB), a novel approach acting as a plug-in for common methods of unsupervised model debiasing, exploiting the inherent bias-learning tendency of diffusion models in data generation. Specifically, our approach adopts conditional diffusion models to generate synthetic bias-aligned images, which fully replace the original training set for learning an effective bias amplifier model to be subsequently incorporated into an end-to-end and a two-step unsupervised debiasing approach. By tackling the fundamental issue of bias-conflicting training samples’ memorization in learning auxiliary models, typical of this type of technique, our proposed method outperforms the current state-of-the-art in multiple benchmark datasets, demonstrating its potential as a versatile and effective tool for tackling bias in deep learning models. Code is available at https://github.com/Malga-Vision/DiffusingDeBias Massimiliano Ciranni, Vito Paolo Pastore, Roberto Di Via, Enzo Tartaglione, Francesca Odone, Vittorio Murino |
NeurIPS | 5 |
| 2025 | Looking at Model Debiasing through the Lens of Anomaly DetectionabstractDeep neural networks are likely to learn unintended spurious correlations between training data and labels when dealing with biased data, potentially limiting the generalization to unseen samples not presenting the same bias. In this context, model debiasing approaches can be de-vised aiming at reducing the model's dependency on such unwanted correlations, either leveraging the knowledge of bias information or not. In this work, we focus on the latter and more realistic scenario, showing the importance of accurately predicting the bias-conflicting and bias-aligned samples to obtain compelling performance in bias mitigation. On this ground, we propose to conceive the problem of model bias from an out-of-distribution perspective, intro-ducing a new bias identification method based on anomaly detection. We claim that when data is mostly biased, bias-conflicting samples can be regarded as outliers with respect to the bias-aligned distribution in the feature space of a bi-ased model, thus allowing for precisely detecting them with an anomaly detection method. Coupling the proposed bias identification approach with bias-conflicting data upsampling and augmentation in a two-step strategy, we reach state-of-the-art performance on synthetic and real benchmark datasets. Ultimately, our proposed approach shows that the data bias issue does not necessarily require complex debiasing methods, given that an accurate bias identification procedure is defined. Source code is available at https://github.com/Malga-Vision/MoDAD Vito Paolo Pastore, Massimiliano Ciranni, Davide Marinelli, Francesca Odone, Vittorio Murino |
WACV | 4 |
| 2025 | Self-Supervised Pre-Training with Diffusion Model for Few-Shot Landmark Detection in X-Ray ImagesabstractDeep neural networks have been extensively applied in the medical domain for various tasks, including image classification, segmentation, and landmark detection. However, their application is often hindered by data scarcity, both in terms of available annotations and images. This study introduces a novel application of denoising diffusion probabilistic models (DDPMs) to the landmark detection task, specifically addressing the challenge of limited annotated data in x-ray imaging. Our key innovation lies in leveraging DDPMs for self-supervised pre-training in landmark detection, a previously unexplored approach in this domain. This method enables accurate landmark detection with minimal annotated training data (as few as 50 images), surpassing both ImageNet supervised pretraining and traditional self-supervised techniques across three popular x-ray benchmark datasets. To our knowledge, this work represents the first application of diffusion models for self-supervised learning in landmark detection, which may offer a valuable pre-training approach in few-shot regimes, for mitigating data scarcity. To bolster further development and reproducibility, we provide open access to our code and pre-trained models for a variety of x-ray related applications: https://github.com/Malga-Vision/DiffusionXray-FewShot-LandmarkDetection Roberto Di Via, Francesca Odone, Vito Paolo Pastore |
WACV | 2 |
| 2025 | View-to-label: Multi-view consistency for self-supervised monocular 3D object detectionabstractFor autonomous vehicles, driving safely is highly dependent on the capability to correctly perceive the environment in the 3D space, hence the task of 3D object detection represents a fundamental aspect of perception. While 3D sensors deliver accurate metric perception, monocular approaches enjoy cost and availability advantages that are valuable in a wide range of applications. Unfortunately, training monocular methods requires a vast amount of annotated data. To compensate for this need, we propose a novel approach to self-supervise 3D object detection purely from RGB video sequences, leveraging geometric constraints and weak labels. Unlike other approaches that exploit additional sensors during training, our method relies on the temporal continuity of video sequences. A supervised pre-training on synthetic data produces initial plausible 3D boxes, then our geometric and photometrically grounded losses provide a strong self-supervision signal that allows the model to be fine-tuned on real data without labels. Our experiments on Autonomous Driving benchmark datasets showcase the effectiveness and generality of our approach and the competitive performance compared to other self-supervised approaches. • Self-supervised 3D object detection from RGB videos using geometric constraints. • Temporal continuity repaces additional sensors in training. • Self-supervision enables fine-tuning on real data without labels. Issa Mouawad, Nikolas Brasch, Fabian Manhardt, Federico Tombari, Francesca Odone |
Comput. Vis. Image Underst. | 5 |
| 2025 | Predicting Engagement of Older People's Virtual Teams from Video Call AnalysisabstractThis study examines seniors’ creative engagement in group activities using synchronous communication tools and explores automatic assessment methods through behavioral and psychophysiological measurements. Working with a small senior group on collaborative creative tasks, we implemented a comprehensive data collection approach using audio-visual and physiological measurements. Machine learning models were used to evaluate group creative engagement levels using various data subsets. Results show that engagement assessment can be effective with different feature combinations, allowing flexibility across contexts and constraints. The multimodal approach, combining facial, audio, and body analysis, achieved optimal performance and is recommended when conditions permit. Our research provides insights into seniors’ online creative participation and presents an automated system for detecting creative engagement in virtual teams, supporting active participation strategies. Nicoletta Noceti, Simone Campisi, Alice Chirico, Vittorio Cuculo, Giuliano Grossi, Monica Michelotto, Francesca Odone, Andrea Gaggioli, Raffaella Lanzarotti |
Int. J. Hum. Comput. Interact. | 7 |
| 2024 | Transferring disentangled representations: bridging the gap between synthetic and real imagesabstractDeveloping meaningful and efficient representations that separate the fundamental structure of the data generation mechanism is crucial in representation learning. However, Disentangled Representation Learning has not fully shown its potential on real images, because of correlated generative factors, their resolution and limited access to ground truth labels. Specifically on the latter, we investigate the possibility of leveraging synthetic data to learn general-purpose disentangled representations applicable to real data, discussing the effect of fine-tuning and what properties of disentanglement are preserved after the transfer. We provide an extensive empirical study to address these issues. In addition, we propose a new interpretable intervention-based metric, to measure the quality of factors encoding in the representation. Our results indicate that some level of disentanglement, transferring a representation from synthetic to real data, is possible and effective. Jacopo Dapueto, Nicoletta Noceti, Francesca Odone |
NeurIPS | 3 |
| 2024 | Head pose estimation with uncertainty and an application to dyadic interaction detectionabstractDetermining the visual focus of attention of people in a scene is a fundamental cue to understand social interactions from videos. Gaze direction is ideal for determining eye contact, a basic cue of non-verbal communication, but it is not always easy to recognise. Head direction is a well-known proxy of gaze direction, more robust to the variability of the scene, thus offering a valuable alternative. In this work, we consider HHP-net, a method for estimating the head direction from single frames based on a heteroscedastic neural network to estimate people’s head pose from a minimal set of head key points. We formulate the problem as a multi-task regression, to predict the pose as a triplet of Euler angles from the output of a 2D pose estimator. HHP-net also provides a measure of the aleatoric heteroscedastic uncertainties associated with the angles, through an ad-hoc loss function we introduce. In a thorough experimental analysis, we show that our model is efficient and effective compared with the state of the art, with only ∼2 degrees of degradation in the worst case counterbalanced by a space occupation ∼12 times smaller. We also show the beneficial effects of uncertainty on interpretability. Finally, we discuss the robustness of our method to input variability, showing that it can be seen as a plug-in to different pose estimators. As a proof-of-concept, we address social interaction analysis, with an algorithm to detect dyadic interactions in images. Federico Figari Tomenotti, Nicoletta Noceti, Francesca Odone |
Comput. Vis. Image Underst. | 3 |
| 2024 | Top-tuning: A study on transfer learning for an efficient alternative to fine tuning for image classification with fast kernel methodsabstractThe impressive performance of deep learning architectures is associated with a massive increase in model complexity. Millions of parameters need to be tuned, with training and inference time scaling accordingly, together with energy consumption. But is massive fine-tuning always necessary? In this paper, focusing on image classification, we consider a simple transfer learning approach exploiting pre-trained convolutional features as input for a fast-to-train kernel method. We refer to this approach as top-tuning since only the kernel classifier is trained on the target dataset. In our study, we perform more than 3000 training processes focusing on 32 small to medium-sized target datasets, a typical situation where transfer learning is necessary. We show that the top-tuning approach provides comparable accuracy with respect to fine-tuning, with a training time between one and two orders of magnitude smaller. These results suggest that top-tuning is an effective alternative to fine-tuning in small/medium datasets, being especially useful when training time efficiency and computational resources saving are crucial. Paolo Didier Alfano, Vito Paolo Pastore, Lorenzo Rosasco, Francesca Odone |
Image Vis. Comput. | 4 |
| 2024 | Computer vision and deep learning meet plankton: Milestones and future directionsabstractPlanktonic organisms play a pivotal role within aquatic ecosystems, serving as the foundation of the aquatic food chain while also playing a critical role in climate regulation and the production of oxygen. In recent years, the advent of automated systems for capturing in-situ images has led to a huge influx of plankton images, making manual classification impractical. This, at the same time, has opened up opportunities for the application of machine learning and deep learning solutions. This paper undertakes an extensive analysis of the broad range of computer vision techniques and methodologies that have emerged to facilitate the automatic analysis of small- to large-scale datasets containing plankton images. By focusing on different computer vision tasks, we present findings and limitations in order to offer a comprehensive overview of the current state-of-the-art, while also pinpointing the open challenges that demand further research and attention. Massimiliano Ciranni, Vittorio Murino, Francesca Odone, Vito Paolo Pastore |
Image Vis. Comput. | 3 |
| 2023 | An Automatic Tool Performing Functional Analysis in MR Urography in ChildrenabstractMagnetic Resonance urography (MRU) can be used to evaluate abnormalities of the urinary tract in children, with the advantage of being a non-invasive technique and allowing both morphologic and functional assessments. Today, MRU analysis is usually performed using semi-automatic software that typically requires manual segmentation of kidney and pelvis and other time-consuming interactions. In this work, we propose a deep learning approach to automatize the functional MRU analysis. Our pipeline first employs an Attention U-Net for kidney and pelvis segmentation on morphological magnetic resonance and then an image registration process to align the segmentations on the functional MR. The automatic segmentation of morphological MR has been tested on 107 patients using cross validation, achieving a dice score of$\mathbf{0.87}\pm \mathbf{0.15}$and$\mathbf{0.91}\pm \mathbf{0.11}$for left and right kidney, and a dice score of$\mathbf{0.75}\pm \mathbf{0.24}$and$\mathbf{0.71}\pm \mathbf{0.25}$for the left and right pelvis respectively. These segmentations are used to extract morphological and functional parameters to assess the urinary-tract function in children undergoing analysis. The proposed approach has been integrated into a commercial web viewer (DicomVision 0.18.3) so that it can be used easily by clinical staff. Our tests demonstrate that this automated tool allows for rapid and comprehensive analysis of children MRU, thus laying the pave for its wider exploitation in clinical routines. Elena Vincenzi, Alice Fantazzini, Martina Gulino, Simone Manini, Alessandro Verri, Francesca Odone, Luca Basso, Maria Beatrice Damasio, Curzio Basso |
CBMS | 6 |
| 2023 | Efficient unsupervised learning of biological images with compressed deep featuresabstractMachine learning has significantly impacted the analysis of biological images and is now an important part of many biological data analysis pipelines. A variety of biological and biomedical domain-related tasks is gaining benefit from image analysis and pattern recognition tools developed currently. Applications include diagnostic histopathology, environmental monitoring, synthetic biology, genomics, and proteomics. Particularly in the last decade, several deep learning and advanced computer vision methods such as convolutional neural networks (CNNs), typically trained in a supervised fashion, have started to be largely employed in biological image classification. Moreover, the advancement of automatic acquisition systems has been generating a massive amount of biological data, which requires to be analyzed by domain experts. However, the cost of manual annotation of such data has become a bottleneck, impairing the application of supervised machine learning algorithms. Biological images generally have an intrinsic high variability, whose identity is sometimes hard to assign and strongly dependent on the annotator’s expertise. In this context, a limited number of annotation-free (i.e., unsupervised) learning solutions have been proposed, typically based on hand-crafted features, specifically tailored for a certain biological domain. Nonetheless, a successful unsupervised learning approach must be accurate, and sufficiently robust to deal with different biological domains. This paper aims at providing a viable solution to these issues, proposing an unsupervised learning algorithm based on compressed deep features for image classification. We exploit features extracted from ImageNet pre-trained transformers and CNNs, further compressed with a customized β-Variational AutoEncoder (β-VAE), that we call reconstruction VAE (R-VAE). We test our algorithm on biological images coming from diverse domains characterized by high variability in shape and texture information and acquired with widely differing imaging platforms. Considered image datasets range from multi-cellular organisms (plankton, coral) to sub-cellular organelles (budding yeast vacuoles, human cells’ nuclei, etc.). Our results show that the compressed deep features extracted from different pre-trained vision models establish new unsupervised learning state-of-the-art performances for the investigated datasets. Vito Paolo Pastore, Massimiliano Ciranni, Simone Bianco 0002, Jennifer Carol Fung, Vittorio Murino, Francesca Odone |
Image Vis. Comput. | 6 |
| 2023 | Uncertainty-Aware Gaze Tracking for Assisted Living EnvironmentsabstractEffective assisted living environments must be able to infer how their occupants interact in a variety of scenarios. Gaze direction provides strong indications of how a person engages with the environment and its occupants. In this paper, we investigate the problem of gaze tracking in multi-camera assisted living environments. We propose a gaze tracking method based on predictions generated by a neural network regressor that relies only on the relative positions of facial keypoints to estimate gaze. For each gaze prediction, our regressor also provides an estimate of its own uncertainty, which is used to weigh the contribution of previously estimated gazes within a tracking framework based on an angular Kalman filter. Our gaze estimation neural network uses confidence gated units to alleviate keypoint prediction uncertainties in scenarios involving partial occlusions or unfavorable views of the subjects. We evaluate our method using videos from the MoDiPro dataset, which we acquired in a real assisted living facility, and on the publicly available MPIIFaceGaze, GazeFollow, and Gaze360 datasets. Experimental results show that our gaze estimation network outperforms sophisticated state-of-the-art methods, while additionally providing uncertainty predictions that are highly correlated with the actual angular error of the corresponding estimates. Finally, an analysis of the temporal integration performance of our method demonstrates that it generates accurate and temporally stable gaze predictions. Paris Her, Logan Manderle, Philipe A. Dias, Henry Medeiros 0001, Francesca Odone |
IEEE Trans. Image Process. | 5 |
| 2022 | Real time Vehicle Color Recognition on a budget: an investigation on the usage of CNN architecturesabstractIn this work, we consider the problem of vehicle color recognition and target scenarios with limited computational resources. Indeed, in real traffic monitoring systems running on the field, algorithms must be light in terms of inference time and memory, but also accurate and robust to the scene variability. We employ end-to-end Convolutional Neural Networks to investigate under which conditions the use of such methodologies– that are state-of-the-art in a multitude of vision-based tasks but often lead to a significant computational burden– can provide us a good compromise between efficiency and effectiveness. We reason on the structure and size of the networks, while monitoring the performance in terms of color classification accuracy and computational effort. We provide an extensive experimental analysis comparing the methods using a benchmark and a private dataset acquired on the field, with almost 20K of images covering a variety of scene conditions. Simone Campisi, Luca Colombini, Alberto Lovato, Francesca Odone, Nicoletta Noceti |
AVSS | 4 |
| 2022 | Efficient Unsupervised Learning for Plankton ImagesabstractMonitoring plankton populations in situ is fundamental to preserve the aquatic ecosystem. Plankton microorganisms are in fact susceptible of minor environmental perturbations, that can reflect into consequent morphological and dynamical modifications. Nowadays, the availability of advanced automatic or semi-automatic acquisition systems has been allowing the production of an increasingly large amount of plankton image data. The adoption of machine learning algorithms to classify such data may be affected by the significant cost of manual annotation, due to both the huge quantity of acquired data and the numerosity of plankton species. To address these challenges, we propose an efficient unsupervised learning pipeline to provide accurate classification of plankton microorganisms. We build a set of image descriptors exploiting a two-step procedure. First, a Variational Autoencoder (VAE) is trained on features extracted by a pre-trained neural network. We then use the learnt latent space as image descriptor for clustering. We compare our method with state-of-the-art unsupervised approaches, where a set of pre-defined hand-crafted features is used for clustering of plankton images. The proposed pipeline outperforms the benchmark algorithms for all the plankton datasets included in our analysis, providing better image embedding properties. Paolo Didier Alfano, Marco Rando, Marco Letizia, Francesca Odone, Lorenzo Rosasco, Vito Paolo Pastore |
ICPR | 4 |
| 2022 | HHP-Net: A light Heteroscedastic neural network for Head Pose estimation with uncertaintyabstractIn this paper we introduce a novel method to estimate the head pose of people in single images starting from a small set of head keypoints. To this purpose, we propose a regression model that exploits keypoints computed automatically by 2D pose estimation algorithms and outputs the head pose represented by yaw, pitch, and roll. Our model is simple to implement and more efficient with respect to the state of the art –faster in inference and smaller in terms of memory occupancy –with comparable accuracy.Our method also provides a measure of the heteroscedastic uncertainties associated with the three angles, through an appropriately designed loss function; we show there is a correlation between error and uncertainty values, thus this extra source of information may be used in subsequent computational steps. As an example application, we address social interaction analysis in images: we propose an algorithm for a quantitative estimation of the level of interaction between people, starting from their head poses and reasoning on their mutual positions. Giorgio Cantarini, Federico Figari Tomenotti, Nicoletta Noceti, Francesca Odone |
WACV | 4 |
| 2022 | Cross-view action recognition with small-scale datasets
Gaurvi Goyal, Nicoletta Noceti, Francesca Odone |
Image Vis. Comput. | 3 |
| 2021 | On The Precision Of Markerless 3d Semantic Features: An Experimental Study On Violin PlayingabstractHuman motion analysis is an essential task in several domains and, depending on the application field, it requires different level of accuracy. In the motor control field it is commonly performed with motion capture systems and infrared markers that guarantee a high accuracy. However, these systems are expensive, cumbersome, and may induce bias. An alternative to marker-based technologies are image-based marker-less systems, that are cheaper and do not affect the naturalness of the motion. Although their accuracy level seems to limit their use in motor control field, a thorough quantitative comparison with marker-based techniques does not appear to be available yet. We compare the estimates of a 3D image-based marker-less pipeline we propose, with a standard marker-based system; the analysis is carried out on a multi-sensor dataset acquired to study the motion of violin players. The results we obtain on the precision level are suggesting that marker-less systems may successfully track performances in real-world settings. Matteo Moro, Maura Casadio, Leigh A. Mrotek, Rajiv Ranganathan, Robert A. Scheidt, Francesca Odone |
ICIP | 6 |
| 2020 | Single View Learning in Action RecognitionabstractViewpoint is an essential aspect of how an action is visually perceived, with the motion appearing substantially different for some viewpoint pairs. Data driven action recognition algorithms compensate for this by including a variety of viewpoints in their training data, adding to the cost of data acquisition as well as training. We propose a novel methodology that leverages deeply pretrained features to learn actions from a single viewpoint using domain adaptation for knowledge transfer. We demonstrate the effectiveness of this pipeline on 3 different datasets: IXMAS, MoCA and NTU RGBD+, and compare with both classical and deep learning methods. Our method requires low training data and demonstrates unparalleled cross-view action recognition accuracies for single view learning. Gaurvi Goyal, Nicoletta Noceti, Francesca Odone |
ICPR | 3 |
| 2020 | Learning dictionaries of kinematic primitives for action classificationabstractThis paper proposes a method based on visual motion primitives to address the problem of action understanding. The approach builds in an unsupervised way a dictionary of kinematic primitives from a set of sub-movements obtained by segmenting the velocity profile of an action on the basis of local minima derived directly from the optical flow. The dictionary is then used to describe each sub-movement as a linear combination of atoms using sparse coding. The descriptive capability of the proposed motion representation is experimentally validated on the MoCA dataset, a collection of synchronized multi-view videos and motion capture data of cooking activities. The results show that the approach, despite its simplicity, has a good performance in action classification, especially when the motion primitives are combined over time. Also, the method is proved to be tolerant to view point changes, and can thus support cross-view action recognition. Overall, the method may be seen as a backbone of a general approach to action understanding, with potential applications in robotics. Alessia Vignolo, Nicoletta Noceti, Alessandra Sciutti, Francesca Odone, Giulio Sandini |
ICPR | 4 |
| 2020 | Stairway to Elders: Bridging Space, Time and Emotions in Their Social Environment for Wellbeing
Giuseppe Boccignone, Claudio de'Sperati, Marco Granato, Giuliano Grossi, Raffaella Lanzarotti, Nicoletta Noceti, Francesca Odone |
ICPRAM | 7 |
| 2020 | Boosting car plate recognition systems performances with agile re-trainingabstractIn this work, we report an experimental study on an Automatic Licence Plate Recognition system developed and commercialized by a partner company, with the main goals of critically analysing the original system and of devising effective but minimally invasive design changes. From a scientific point of view, ours is an attempt of reducing the gap between the different experimental approaches in academia and industry. The system is organized in layers, with an initial car plate proposal step followed by a OCR step. To cope with the drawbacks of the pre-existing system, we inserted an intermediate CNN binary classification step to discriminate between plates and non plates independently from the OCR module. Our solution incorporates new data available from working installations, in a closed refinement loop. We evaluate the modified system on 8 different installations. With respect to the original performances, we obtained significant improvements with an impact on both false positive (-9.8%) and false negatives (-5%). Giorgio Cantarini, Nicoletta Noceti, Francesca Odone |
IPAS | 3 |
| 2020 | Gaze Estimation for Assisted Living EnvironmentsabstractEffective assisted living environments must be able to perform inferences on how their occupants interact with one another as well as with surrounding objects. To accomplish this goal using a vision-based automated approach, multiple tasks such as pose estimation, object segmentation and gaze estimation must be addressed. Gaze direction provides some of the strongest indications of how a person interacts with the environment. In this paper, we propose a simple neural network regressor that estimates the gaze direction of individuals in a multi-camera assisted living scenario, relying only on the relative positions of facial keypoints collected from a single pose estimation model. To handle cases of keypoint occlusion, our model exploits a novel confidence gated unit in its input layer. In addition to the gaze direction, our model also outputs an estimation of its own prediction uncertainty. Experimental results on a public benchmark demonstrate that our approach performs on par with a complex, dataset-specific baseline, while its uncertainty predictions are highly correlated to the actual angular error of corresponding estimations. Finally, experiments on images from a real assisted living environment demonstrate that our model has a higher suitability for its final application. Philipe A. Dias, Damiano Malafronte, Henry Medeiros 0001, Francesca Odone |
WACV | 4 |
| 2020 | Positive technology for elderly well-being: A review
Giuliano Grossi, Raffaella Lanzarotti, Paolo Napoletano, Nicoletta Noceti, Francesca Odone |
Pattern Recognit. Lett. | 5 |
| 2019 | Visual Tracking with Autoencoder-Based Maximum A Posteriori Data FusionabstractIn this paper, a novel method for tracker fusion is proposed and evaluated for vision-based object tracking. This work combines three distinct popular techniques into a recursive Bayesian estimation algorithm. First, a semi-supervised learning approach is used to train deep neural networks capable of detecting anomalous visual tracking behavior. Next, the network output is used to compute maximum a posteriori scores. Finally, these scores are integrated into the observation weighing mechanism of an existing data fusion algorithm. We evaluated the proposed algorithm on the OTB-100 benchmark dataset and compared its performance to the performance of the baseline fusion approach. Yevgeniy Reznichenko, Enrico Prampolini, Abubakar Siddique 0003, Henry Medeiros 0001, Francesca Odone |
COMPSAC (1) | 5 |
| 2017 | Exploring Biological Motion Regularities of Human Actions: A New Perspective on Video AnalysisabstractThe ability to detect potentially interacting agents in the surrounding environment is acknowledged to be one of the first perceptual tasks developed by humans, supported by the ability to recognise biological motion. The precocity of this ability suggests that it might be based on rather simple motion properties, and it can be interpreted as an atomic building block of more complex perception tasks typical of interacting scenarios, as the understanding of non-verbal communication cues based on motion or the anticipation of others’ action goals. In this article, we propose a novel perspective for video analysis, bridging cognitive science and machine vision, which leverages the use of computational models of the perceptual primitives that are at the basis of biological motion perception in humans. Our work offers different contributions. In a first part, we propose an empirical formulation for the Two-Thirds Power Law , a well-known invariant law of human movement, and thoroughly discuss its readability in experimental settings of increasing complexity. In particular, we consider unconstrained video analysis scenarios, where, to the best of our knowledge, the invariant law has not found application so far. The achievements of this analysis pave the way for the second part of the work, in which we propose and evaluate a general representation scheme for biological motion characterisation to discriminate biological movements with respect to non-biological dynamic events in video sequences. The method is proposed as the first layer of a more complex architecture for behaviour analysis and human-machine interaction, providing in particular a new way to approach the problem of human action understanding. Nicoletta Noceti, Francesca Odone, Alessandra Sciutti, Giulio Sandini |
ACM Trans. Appl. Percept. | 2 |
| 2017 | Scale Invariant and Noise Robust Interest Points With ShearletsabstractShearlets are a relatively new directional multi-scale framework for signal analysis, which have been shown effective to enhance signal discontinuities, such as edges and corners at multiple scales even in the presence of a large quantity of noise. In this paper, we consider blob-like features in the shearlets framework. We derive a measure, which is very effective for blob detection, and, based on this measure, we propose a blob detector and a keypoint description, whose combination outperforms the state-of-the-art algorithms with noisy and compressed images. We also demonstrate that the measure satisfies the perfect scale invariance property in the continuous case. We evaluate the robustness of our algorithm to different types of noise, including blur, compression artifacts, and Gaussian noise. Furthermore, we carry on a comparative analysis on benchmark data, referring, in particular, to tolerance to noise and image compression. Miguel A. Duval, Nicoletta Noceti, Francesca Odone, Ernesto De Vito |
IEEE Trans. Image Process. | 3 |
| 2016 | An integrated artificial vision framework for assisting visually impaired users
Manuela Chessa, Nicoletta Noceti, Francesca Odone, Fabio Solari, Joan Sosa-García, Luca Zini |
Comput. Vis. Image Underst. | 3 |
| 2016 | Portable and fast text detection
Luca Zini, Francesca Odone |
Mach. Vis. Appl. | 2 |
| 2016 | Adaptive Body Gesture Representation for Automatic Emotion RecognitionabstractWe present a computational model and a system for the automated recognition of emotions starting from full-body movement. Three-dimensional motion data of full-body movements are obtained either from professional optical motion-capture systems (Qualisys) or from low-cost RGB-D sensors (Kinect and Kinect2). A number of features are then automatically extracted at different levels, from kinematics of a single joint to more global expressive features inspired by psychology and humanistic theories (e.g., contraction index, fluidity, and impulsiveness). An abstraction layer based on dictionary learning further processes these movement features to increase the model generality and to deal with intraclass variability, noise, and incomplete information characterizing emotion expression in human movement. The resulting feature vector is the input for a classifier performing real-time automatic emotion recognition based on linear support vector machines. The recognition performance of the proposed model is presented and discussed, including the tradeoff between precision of the tracking measures (we compare the Kinect RGB-D sensor and the Qualisys motion-capture system) versus dimension of the training dataset. The resulting model and system have been successfully applied in the development of serious games for helping autistic children learn to recognize and express emotions by means of their full-body movement. Stefano Piana, Alessandra Staglianò, Francesca Odone, Antonio Camurri |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2015 | Banknote Recognition as a CBIR Problem
Joan Sosa-García, Francesca Odone |
SISAP | 2 |
| 2015 | Structured multi-class feature selection with an application to face recognition
Luca Zini, Nicoletta Noceti, Giovanni Fusco 0004, Francesca Odone |
Pattern Recognit. Lett. | 4 |
| 2015 | Edges and Corners With ShearletsabstractShearlets are a relatively new and very effective multi-scale framework for signal analysis. Contrary to the traditional wavelets, shearlets are capable to efficiently capture the anisotropic information in multivariate problem classes. Therefore, shearlets can be seen as the valid choice for multi-scale analysis and detection of directional sensitive visual features like edges and corners. In this paper, we start by reviewing the main properties of shearlets that are important for edge and corner detection. Then, we study algorithms for multi-scale edge and corner detection based on the shearlet representation. We provide an extensive experimental assessment on benchmark data sets which empirically confirms the potential of shearlets feature detection. Miguel A. Duval, Francesca Odone, Ernesto De Vito |
IEEE Trans. Image Process. | 2 |
| 2015 | Online Space-Variant Background Modeling With Sparse CodingabstractIn this paper, we propose a sparse coding approach to background modeling. The obtained model is based on dictionaries which we learn and keep up to date as new data are provided by a video camera. We observe that, without dynamic events, video frames may be seen as noisy data belonging to the background. Over time, such background is subject to local and global changes due to variable illumination conditions, camera jitter, stable scene changes, and intermittent motion of background objects. To capture the locality of some changes, we propose a space-variant analysis where we learn a dictionary of atoms for each image patch, the size of which depends on the background variability. At run time, each patch is represented by a linear combination of the atoms learnt online. A change is detected when the atoms are not sufficient to provide an appropriate representation, and stable changes over time trigger an update of the current dictionary. Even if the overall procedure is carried out at a coarse level, a pixel-wise segmentation can be obtained by comparing the atoms with the patch corresponding to the dynamic event. Experiments on benchmarks indicate that the proposed method achieves very good performances on a variety of scenarios. An assessment on long video streams confirms our method incorporates periodical changes, as the ones caused by variations in natural illumination. The model, fully data driven, is suitable as a main component of a change detection system. Alessandra Staglianò, Nicoletta Noceti, Alessandro Verri, Francesca Odone |
IEEE Trans. Image Process. | 4 |
| 2014 | Ask the Image: Supervised Pooling to Preserve Feature LocalityabstractIn this paper we propose a weighted supervised pooling method for visual recognition systems. We combine a standard Spatial Pyramid Representation which is commonly adopted to encode spatial information, with an appropriate Feature Space Representation favoring semantic information in an appropriate feature space. For the latter, we propose a weighted pooling strategy exploiting data supervision to weigh each local descriptor coherently with its likelihood to belong to a given object class. The two representations are then combined adaptively with Multiple Kernel Learning. Experiments on common benchmarks (Caltech-256 and PASCAL VOC-2007) show that our image representation improves the current visual recognition pipeline and it is competitive with similar state-of-art pooling methods. We also evaluate our method on a real Human-Robot Interaction setting, where the pure Spatial Pyramid Representation does not provide sufficient discriminative power, obtaining a remarkable improvement. Sean Ryan Fanello, Nicoletta Noceti, Carlo Ciliberto, Giorgio Metta, Francesca Odone |
CVPR | 5 |
| 2014 | Semi-supervised learning of sparse representations to recognize people spatial orientationabstractIn this paper we consider the problem of classifying people spatial orientation with respect to the camera viewpoint from 2D images. Structured multi-class feature selection allows us to control the amount of redundancy of our input data, while semi-supervised learning helps us coping with the intrinsic ambiguity of output labels. We model the multi-class classification problem with an all-pairs strategy based on the use of a coding matrix. A thorough experimental evaluation on the TUD Multiview Pedestrian benchmark dataset demonstrates the superiority of our approach w.r.t. state-of-the-art. Nicoletta Noceti, Francesca Odone |
ICIP | 2 |
| 2014 | Emotional CharadesabstractThis is a short description of the Emotional Charades serious game demo. Our goal is to focus on emotion expression through body gestures, making the players aware of the amount of affective information their bodies convey. The whole framework aims at helping children with autism to understand and express emotions. We also want to compare the performances of our automatic recognition system and the ones achieved by humans. Stefano Piana, Alessandra Staglianò, Francesca Odone, Antonio Camurri |
ICMI | 3 |
| 2014 | A Spectral Graph Kernel and Its Application to Collective Activities ClassificationabstractIn this work we consider a machine learning setting where data are represented as graphs. First, we derive a kernel function which evaluates the similarity between graphs, while capturing pair-wise constraints between graph nodes. Second, we apply it to the problem of classifying collective activities: on this respect we first represent groups of people located in a spatial neighborhood as graphs, and then train a multi-class classifier able to capture the behavior of the groups. We evaluate our approach on a benchmark dataset and report a comparative analysis with other state-of-art methods which highlights the benefits of our approach. Nicoletta Noceti, Francesca Odone |
ICPR | 2 |
| 2014 | Humans in groups: The importance of contextual information for understanding collective activities
Nicoletta Noceti, Francesca Odone |
Pattern Recognit. | 2 |
| 2014 | Geometrical and computational aspects of Spectral Support Estimation for novelty detection
Alessandro Rudi, Francesca Odone, Ernesto De Vito |
Pattern Recognit. Lett. | 2 |
| 2014 | Multiview Matching of Articulated ObjectsabstractWe address the problem of multiview association of articulated objects observed using possibly moving and hand-held cameras. Starting from trajectory data, we encode the temporal evolution of the objects and perform matching without making assumptions on scene geometry and with only weak assumptions on the field-of-view overlaps. After generating a viewpoint invariant representation using self-similarity matrices, we put in correspondence the spatio-temporal object descriptions using spectral methods on the resulting matching graph. We validate the proposed method on three publicly available real-world datasets and compare it with alternative approaches. Moreover, we present an extensive analysis of the accuracy of the proposed method in different contexts, with varying noise levels on the input data, varying amount of overlap between the fields of view, and varying duration of the available observations. Luca Zini, Francesca Odone, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Background modeling through dictionary learningabstractIn this work we build a model of the background based on dictionary learning. The image is divided into patches of equal size and a background model is obtained as a sparse linear combination of patch prototypes learnt from the image stream and updated when necessary to take into account stable variations. By enforcing sparsity, the obtained reconstruction can be computed and maintained effectively. The proposed method is stable with respect to illumination changes, correctly incorporates stable background changes in the model, and cancels out moving objects. Experiments on benchmark data indicate that the proposed method reaches very good pixel-wise performances even if relatively large patches are used. Alessandra Staglianò, Nicoletta Noceti, Alessandro Verri, Francesca Odone |
ICIP | 4 |
| 2013 | Precise people counting in real timeabstractIn this paper we propose a motion-based people counting algorithm that relies on a weak camera calibration and produces a smooth estimate of the number of people in the scene. The method performs an analysis of the severity of possible occlusions and the integration of instantaneous observations over time. The key features of the algorithm are a simple pipeline, a small computational cost, the use of a model-free approach that does not need complex training procedures and its ability to work in different types of scenarios. We report results on both benchmark and acquired in-house datasets of different degrees of complexity, showing how our solution achieves comparable or superior performances with respect to state-of-art methods, while providing real-time performances. Luca Zini, Nicoletta Noceti, Francesca Odone |
ICIP | 3 |
| 2013 | Keep it simple and sparse: real-time action recognition
Sean Ryan Fanello, Ilaria Gori, Giorgio Metta, Francesca Odone |
J. Mach. Learn. Res. | 4 |
| 2012 | Combining Retrieval and Classification for Real-Time Face RecognitionabstractIn this paper we propose a real time face recognition method that combines face matching and identity verification modules in a feedback loop, exploiting the temporal efficiency of matching and the performances of SVM classifiers. Our approach represents an ad-hoc solution for settings characterized by variable quantity, quality and distribution of labeled data among the identities. We assess the procedure on two data sets of different complexities, showing the effectiveness of our solution. For its intrinsic peculiarities and its limited computational cost the method finds application in real time systems, and will be implemented on a wearable device for supporting visually impaired people to localize known faces. Giovanni Fusco 0004, Nicoletta Noceti, Francesca Odone |
AVSS | 3 |
| 2012 | Learning common behaviors from large sets of unlabeled temporal series
Nicoletta Noceti, Francesca Odone |
Image Vis. Comput. | 2 |
| 2010 | Learning how to grasp objects
Annalisa Barla, Luca Baldassarre, Nicoletta Noceti, Francesca Odone |
ESANN | 4 |
| 2009 | Combined Motion and Appearance Models for Robust Object Tracking in Real-TimeabstractThis paper proposes a tracking architecture that finds a trade-off between accuracy and efficiency, via a combined solution of motion and appearance information. We explore the use of color features into a tracking pipeline based on Kalman filtering. The devised architecture is made of simple modules, combined to reach a robust final result, while keeping the computation cost low (we perform 20 fps). The method has been evaluated on three benchmark datasets and is currently under use on real video-surveillance systems, reporting very good tracking results. Nicoletta Noceti, Augusto Destrero, Alberto Lovato, Francesca Odone |
AVSS | 4 |
| 2009 | A Classification Architecture Based on Connected Components for Text Detection in Unconstrained EnvironmentsabstractThe paper presents a method for efficient text detection in unconstrained environments, based on image features derived from connected components and on a classification architecture implementing a focus of attention approach.The main application motivating the work is container code detection with the final goal of checking freight trains composition. Although the method is strongly influenced by the application experimental evidence speaks in favour of its generality: we present results on container codes, car plates images and on the benchmark dataset ICDAR. Luca Zini, Augusto Destrero, Francesca Odone |
AVSS | 3 |
| 2009 | Spatio-temporal constraints for on-line 3D object recognition in videos
Nicoletta Noceti, Elisabetta Delponte, Francesca Odone |
Comput. Vis. Image Underst. | 3 |
| 2009 | A Regularized Framework for Feature Selection in Face Detection and Authentication
Augusto Destrero, Christine De Mol, Francesca Odone, Alessandro Verri |
Int. J. Comput. Vis. | 3 |
| 2009 | A Sparsity-Enforcing Method for Learning Face FeaturesabstractIn this paper, we propose a new trainable system for selecting face features from over-complete dictionaries of image measurements. The starting point is an iterative thresholding algorithm which provides sparse solutions to linear systems of equations. Although the proposed methodology is quite general and could be applied to various image classification tasks, we focus here on the case study of face and eyes detection. For our initial representation, we adopt rectangular features in order to allow straightforward comparisons with existing techniques. For computational efficiency and memory saving requirements, instead of implementing the full optimization scheme on tenths of thousands of features, we propose a three-stage architecture which consists of finding first intermediate solutions to smaller size optimization problems, then merging the obtained results, and next applying further selection procedures. The devised system requires the solution of a number of independent problems, and, hence, the necessary computations could be implemented in parallel. Experimental results obtained on both benchmark and newly acquired face and eyes images indicate that our method is a serious competitor to other feature selection schemes recently popularized in computer vision for dealing with problems of real-time object detection. A major advantage of the proposed system is that it performs well even with relatively small training sets. Augusto Destrero, Christine De Mol, Francesca Odone, Alessandro Verri |
IEEE Trans. Image Process. | 3 |
| 2008 | Spectral Algorithms for Supervised LearningabstractWe discuss how a large class of regularization methods, collectively known as spectral regularization and originally designed for solving ill-posed inverse problems, gives rise to regularized learning algorithms. All of these algorithms are consistent kernel methods that can be easily implemented. The intuition behind their derivation is that the same principle allowing for the numerical stabilization of a matrix inversion problem is crucial to avoid overfitting. The various methods have a common derivation but different computational and theoretical properties. We describe examples of such algorithms, analyze their classification performance on several data sets and discuss their applicability to real-world problems. L. Lo Gerfo, Lorenzo Rosasco, Francesca Odone, Ernesto De Vito, Alessandro Verri |
Neural Comput. | 3 |
| 2007 | A Regularized Approach to Feature Selection for Face Detection
Augusto Destrero, Christine De Mol, Francesca Odone, Alessandro Verri |
ACCV (2) | 3 |
| 2007 | A system for face detection and tracking in unconstrained environmentsabstractWe describe a trainable system for face detection and tracking. The structure of the system is based on multiple cues that discard non face areas as soon as possible: we combine motion, skin, and face detection. The latter is the core of our system and consists of a hierarchy of small SVM classifiers built on the output of an automatic feature selection procedure. Our feature selection is entirely data-driven and allows us to obtain powerful descriptions from a relatively small set of data. Finally, a Kalman tracking on the face region optimizes detection results over time. We present an experimental analysis of the face detection module and results obtained with the whole system on the specific task of counting people entering the scene. Augusto Destrero, Francesca Odone, Alessandro Verri |
AVSS | 2 |
| 2006 | SVD-matching using SIFT features
Elisabetta Delponte, Francesco Isgrò, Francesca Odone, Alessandro Verri |
Graph. Model. | 3 |
| 2005 | Support vector algorithms as regularization networks
Andrea Caponnetto, Lorenzo Rosasco, Francesca Odone, Alessandro Verri |
ESANN | 3 |
| 2005 | Feature selection with nonparametric statisticsabstractIn this paper we discuss a general framework for feature selection based on nonparametric statistics. The three stage approach we propose is based on the assumption that the available data set is representative of a certain concept and aims at learning from the data the selection of a subset of descriptive features out of a large pool of measurements. The first stage requires the computation of a large number of image features. Simple significance tests and the maximum likelihood principle are at the basis of the second stage in which a saliency measure is used to reject the features which do not appear to be descriptive of the given data set. The third and final stage, by using the Spearman independence rank test, selects a maximal number of pairwise independent features. We report experiments on a face dataset (the MIT-CBCL database) which confirm the quality and the potential of the approach. Emanuele Franceschi, Francesca Odone, Fabrizio Smeraldi, Alessandro Verri |
ICIP (1) | 2 |
| 2005 | Learning from Examples as an Inverse ProblemabstractMany works related learning from examples to regularization techniques for inverse problems, emphasizing the strong algorithmic and conceptual analogy of certain learning algorithms with regularization algorithms. In particular it is well known that regularization schemes such as Tikhonov regularization can be effectively used in the context of learning and are closely related to algorithms such as support vector machines. Nevertheless the connection with inverse problem was considered only for the discrete (finite sample) problem and the probabilistic aspects of learning from examples were not taken into account. In this paper we provide a natural extension of such analysis to the continuous (population) case and study the interplay between the discrete and continuous problems. From a theoretical point of view, this allows to draw a clear connection between the consistency approach in learning theory and the stability convergence property in ill-posed inverse problems. The main mathematical result of the paper is a new probabilistic bound for the regularized least-squares algorithm. By means of standard results on the approximation term, the consistency of the algorithm easily follows. Ernesto De Vito, Lorenzo Rosasco, Andrea Caponnetto, Umberto De Giovannini, Francesca Odone |
J. Mach. Learn. Res. | 5 |
| 2005 | Building kernels from binary strings for image matchingabstractIn the statistical learning framework, the use of appropriate kernels may be the key for substantial improvement in solving a given problem. In essence, a kernel is a similarity measure between input points satisfying some mathematical requirements and possibly capturing the domain knowledge. In this paper, we focus on kernels for images: we represent the image information content with binary strings and discuss various bitwise manipulations obtained using logical operators and convolution with nonbinary stencils. In the theoretical contribution of our work, we show that histogram intersection is a Mercer's kernel and we determine the modifications under which a similarity measure based on the notion of Hausdorff distance is also a Mercer's kernel. In both cases, we determine explicitly the mapping from input to feature space. The presented experimental results support the relevance of our analysis for developing effective trainable systems. Francesca Odone, Annalisa Barla, Alessandro Verri |
IEEE Trans. Image Process. | 1 |
| 2004 | Learning, Regularization and Ill-Posed Inverse ProblemsabstractMany works have shown that strong connections relate learning from ex- amples to regularization techniques for ill-posed inverse problems. Nev- ertheless by now there was no formal evidence neither that learning from examples could be seen as an inverse problem nor that theoretical results in learning theory could be independently derived using tools from reg- ularization theory. In this paper we provide a positive answer to both questions. Indeed, considering the square loss, we translate the learning problem in the language of regularization theory and show that consis- tency results and optimal regularization parameter choice can be derived by the discretization of the corresponding inverse problem. 1 Introduction The main goal of learning from examples is to infer an estimator, given a finite sample of data drawn according to a fixed but unknown probabilistic input-output relation. The desired property of the selected estimator is to perform well on new data, i.e. it should gen- eralize. The fundamental works of Vapnik and further developments [16], [8], [5], show that the key to obtain a meaningful solution to the above problem is to control the complex- ity of the solution space. Interestingly, as noted by [12], [8], [2], this is the idea underlying regularization techniques for ill-posed inverse problems [15], [7]. In such a context to avoid undesired oscillating behavior of the solution we have to restrict the solution space. Not surprisingly the form of the algorithms proposed in both theories is strikingly similar. Anyway a careful analysis shows that a rigorous connection between learning and regular- ization for inverse problem is not straightforward. In this paper we consider the square loss and show that the problem of learning can be translated into a convenient inverse problem and consistency results can be derived in a general setting. When a generic loss is consid- ered the analysis becomes immediately more complicated. Some previous works on this subject considered the special case in which the elements of the input space are fixed and not probabilistically drawn [11], [9]. Some weaker results in the same spirit of those presented in this paper can be found in [13] where anyway the connections with inverse problems is not discussed. Finally, our analysis is close to the idea of stochastic inverse problems discussed in [16]. It follows the plan of the paper. Af- ter recalling the main concepts and notation of learning and inverse problems, in section 4 we develop a formal connection between the two theories. In section 5 the main results are stated and discussed. Finally in section 6 we conclude with some remarks and open problems. 2 Learning from examples We briefly recall some basic concepts of learning theory [16], [8]. In the framework of learning, there are two sets of variables: the input space X, compact subset of Rn, and the output space Y , compact subset of R. The relation between the input x X and the output y Y is described by a probability distribution (x, y) = (x)(y|x) on X Y . The distribution is known only through a sample z = (x, y) = ((x1, y1), . . . , (x , y )), called training set, drawn i.i.d. according to . The goal of learning is, given the sample z, to find a function fz : X R such that fz(x) is an estimate of the output y when the new input x is given. The function fz is called estimator and the rule that, given a sample z, provides us with fz is called learning algorithm. Given a measurable function f : X R, the ability of f to describe the distribution is measured by its expected risk defined as I[f ] = (f (x) - y)2 d(x,y). XY The regression function g(x) = y d(y|x), Y is the minimizer of the expected risk over the set of all measurable functions and always exists since Y is compact. Usually, the regression function cannot be reconstructed exactly since we are given only a finite, possibly small, set of examples z. To overcome this problem, in the regularized least squares algorithm an hypothesis space H is fixed, and, given > 0, an estimator f z is defined as the solution of the regularized least squares problem, 1 min{ (f (xi) - yi)2 + f 2H}. (1) f H i=1 The regularization parameter has to be chosen depending on the available data, = ( , z), in such a way that, for every > 0 lim P I[f ( ,z) z ] - inf I[f] = 0. (2) + f H We note that in general inffH I[f ] is larger that I[g] and represents a sort of irreducible error associated with the choice of the space H. The above convergence in probability is usually called consistency of the algorithm [16] [14]. 3 Ill-Posed Inverse Problems and Regularization In this section we give a very brief account of linear inverse problems and regularization theory [15], [7]. Let H and K be two Hilbert spaces and A : H K a linear bounded operator. Consider the equation Af = g (3) where g, g K and g - g K . Here g represents the exact, unknown data and g the available, noisy data. Finding the function f satisfying the above equation, given A and g, is the linear inverse problem associated to Eq. (3). The above problem is, in general, ill- posed, that is, the Uniqueness can be restored introducing the Moore-Penrose generalized inverse f = Ag defined as the minimum norm solution of the problem min Af - g 2 . (4) K f H However the operator A is usually not bounded so, in order to ensure a continuous de- pendence of the solution on the data, the following Tikhonov regularization scheme can be considered1 min{ Af - g 2 + f 2 K H}, (5) f H whose unique minimizer is given by f = (AA + I )-1Ag , (6) where A denotes the adjoint of A. A crucial step in the above algorithm is the choice of the regularization parameter = (, g), as a function of the noise level and the data g, in such a way that lim f (,g) = 0, (7) - f 0 H that is, the regularized solution f (,g) converges to the generalized solution f = Ag (f exists if and only if P g Range(A), where P is the projection on the closure of the range of A and, in that case, Af = P g) when the noise goes to zero. The similarity between regularized least squares algorithm (1) and Tikhonov regulariza- tion (5) is apparent. However, several difficulties emerge. First, to treat the problem of learning in the setting of ill-posed inverse problems we have to define a direct problem by means of a suitable operator A. Second, in the context of learning, it is not clear the nature of the noise . Finally we have to clarify the relation between consistency (2) and the kind of convergence expressed by (7). In the following sections we will show a possible way to tackle these problems. 4 Learning as an Inverse Problem We can now show how the problem of learning can be rephrased in a framework close to the one presented in the previous section. We assume that hypothesis space H is a reproducing kernel Hilbert space [1] with a contin- uous kernel K : X X R. If x X, we let Kx(s) = K(s, x), and, if is the marginal distribution of on X, we define the bounded linear operator A : H L2(X, ) as (Af )(x) = f, Kx = f (x), H 1In the framework of inverse problems, many other regularization procedures are introduced [7]. For simplicity we only treat the Tikhonov regularization. that is, A is the canonical injection of H in L2(X, ). In particular, for all f H, the expected risk becomes, I[f ] = Af - g 2 + I[g], L2(X,) where g is the regression function [2]. The above equation clarifies that if the expected risk admits a minimizer fH on the hypothesis space H, then it is exactly the generalized solution2 f = Ag of the problem Af = g. (8) Moreover, given a training set z = (x, y), we get a discretized version Ax : H E of A, that is (Axf)i = f, Kx = f (x i H i), where E = R is the finite dimensional euclidean space endowed with the scalar product 1 y, y = y E iyi. i=1 It is straightforward to check that 1 (f (xi) - yi)2 = Axf - y 2 , E i=1 so that the estimator f z given by the regularized least squares algorithm is the regularized solution of the discrete problem Axf = y. (9) At this point it is useful to remark the following two facts. First, in learning from examples we are not interested into finding an approximation of the generalized solution of the dis- cretized problem (9), but we want to find a stable approximation of the solution of the exact problem (8) (compare with [9]). Second, we notice that in learning theory the consistency property (2) involves the control of the quantity I[f z ] - inf I[f] = Af - g 2 Af . (10) L2(X,) - inf - g 2L2(X,) f H f H If P is the projection on the closure of the range of A, the definition of P gives 2 I[f z ] - inf I[f] = Afz - P g (11) f H L2(X,) (the above equality stronlgy depends on the fact that the loss function is the square loss). In the inverse problem setting, the square root of the above quantity is called the residue of the solution f z . Hence, consistency is controlled by the residue of the estimator, instead of the reconstruction error f z - f (as in inverse problems). In particular, consistency H is a weaker condition than the one required by (7) and does not require the existence of the generalized solution fH. 5 Regularization, Stochastic Noise and Consistency To apply the framework of ill-posed inverse problems of Section 3 to the formulation of learning proposed above, we note that the operator Ax in the discretized problem (9) differs from the operator A in the exact problem (8) and a measure of the difference between Ax and A is required. Moreover, the noisy data y E and the exact data g L2(X, ) belong to different spaces, so that the notion of noise has to be modified. Given the above premise our derivation of consistency results is developed in two steps: we first study the residue of the solution by means of a measure of the noise due to discretization and then we show a possible way to give a probabilistic evaluation of the noise previously introduced. 2The fact that fH is the minimal norm solution of (4) is ensured by the assumption that the support of the measure is X, since in this case the operator A is injective. 5.1 Bounding the Residue of the Regularized Solution We recall that the regularized solutions of problems (9) and (8) are given by f = (A A y, z x x + I )-1A x f = (AA + I)-1Ag. The above equations show that f and f depend only on A A z x x and AA which are operators from H into H and on Ay and Ag which are elements of x H, so that the space E disappears. This observation suggests that noise levels could be A A x x - AA L(H) and A y , where is the uniform operator norm. To this purpose, for x - Ag H L(H) every = (1, 2) R2+ we define the collection of training sets. U = {z (X Y ) | Ay A x - Ag H 1, Ax x - AA L(H) 2, N} and we let M = sup{|y| | y Y }. The next theorem is the central result of the paper. Theorem 1 If > 0, the following inequalities hold 1. for any training set z U M Af 2 + 1 z - P g L2(X,) - Af - P g L2(X,) 4 2 2. if P g Range(A), for any training set z U, M f 2 + 1 z - f H - f - f H 3 2 2 Moreover if we choose = (, z) in such a way that lim0 sup (,z) = 0 zU 2 lim 1 0 sup = 0 zU (12) (,z) then lim 2 0 sup = 0 zU (,z) lim sup Af (,z) = 0. (13) z - Pg 0 zU L2(X,) We omit the complete proof and refer to [3]. Briefly, the idea is to note that Af z - P g L2(X,) - Af - P g L2(X,) 1 Af = (AA) 2 (f z - Af L2(X,) z - f) H where the last equation follows by polar decomposition of the operator A. Moreover a simple algebraic computation gives f A A y+(AA+I)-1(A y z -f = (AA+I)-1(AA-Ax x)(Ax x+I)-1Ax x -Ag) where the relevant quantities for definition of the noise appear. The first item in the above proposition quantifies the difference between the residues of the regularized solutions of the exact and discretized problems in terms of the noise level = (1, 2). As mentioned before this is exactly the kind of result needed to derive consistency. On the other hand the last part of the proposition gives sufficient conditions on the parameter to ensure convergence of the residue to zero as the level noise decreases. The above results were obtained introducing the collection U of training sets compatible with a certain noise level . It is left to quantify the noise level corresponding to a training set of cardinality . This will be achieved in a probabilistic setting in the next section. 5.2 Stochastic Evaluation of the Noise In this section we estimate the discretization noise = (1, 2). Theorem 2 Let 1, 2 > 0 and = supxX K(x, x), then M 2 P Ag - A x y A H + 1, AA - Ax x L(H) + 2 2 2 1 2 1 - e-22M2 - e-24 (14) The proof is given in [3] and it is based on McDiarmid inequality [10] applied to the random variables F (z) = A x y - Ag G(z) = A A . H x x - AA L(H) Other estimates of the noise can be given using, for example, union bounds and Hoeffd- ing's inequality. Anyway rather then providing a tight analysis our concern was to find an natural, explicit and easy to prove estimate of . 5.3 Consistency and Regularization Parameter Choice Combining Theorems 1 and 2, we easily derive the following corollary. Corollary 1 Given 0 < < 1, with probability greater that 1 - , Af z - Pg - Af - Pg L2(X,) L2(X,) M 1 4 + 1 + log (15) 2 2 Lorenzo Rosasco, Andrea Caponnetto, Ernesto De Vito, Francesca Odone, Umberto De Giovannini |
NIPS | 4 |
| 2003 | Histogram intersection kernel for image classificationabstractIn this paper we address the problem of classifying images, by exploiting global features that describe color and illumination properties, and by using the statistical learning paradigm. The contribution of this paper is twofold. First, we show that histogram intersection has the required mathematical properties to be used as a kernel function for support vector machines (SVMs). Second, we give two examples of how a SVM, equipped with such a kernel, can achieve very promising results on image classification based on color information. Annalisa Barla, Francesca Odone, Alessandro Verri |
ICIP (3) | 2 |
| 2002 | Hausdorff Kernel for 3D Object Acquisition and Detection
Annalisa Barla, Francesca Odone, Alessandro Verri |
ECCV (4) | 2 |
| 2002 | Kernel-Based 3D Object Representation
Annalisa Barla, Francesca Odone |
ICANN | 2 |
| 2002 | Layered Representation of a Video Shot with Mosaicing
Emanuele Trucco, Francesca Odone, Andrea Fusiello |
Pattern Anal. Appl. | 2 |
| 1998 | Visual Learning of Weight from Shape Using Support Vector MachinesabstractWe investigate the automatic estimation of fish weight from sets of morphometric measurements. Our solution combines a vision system with a robust regression method, the Support Vector Machine (SVM). Measurements are taken automatically from two binarised views of each fish in a training sample, then fed to a quadratic SVM along with approximate weight estimates. The SVM learns the law linking weight to shape directly (without computing volume) and compensates for several inaccuracies in the training measurements. We suggest a methodology identifying optimal shape measurements for the task, and report results obtained with a sample of 99 trouts between 300 and 600g, showing good accuracy and reliability, and better performance with respect to length-weight relations adopted commonly in fisheries science. 1 Introduction This work explores a new way of estimating fish weight from shape using computer vision. The relation between weight and shape is important both for fish biology [4, 5... Francesca Odone, Emanuele Trucco, Alessandro Verri |
BMVC | 1 |