Fabio Cuzzolin

dblp:60/2919 · DBLP profile ↗
← Back
66ranked-venue papers
24as first author
22since 2021 · last 2026
0000-0002-9271-2130ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 16 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Credal Ensemble Distillation for Uncertainty Quantification
abstract
Deep ensembles (DE) have emerged as a powerful approach for quantifying predictive uncertainty and distinguishing its aleatoric and epistemic components, thereby enhancing model robustness and reliability. However, their high computational and memory costs during inference pose significant challenges for wide practical deployment. To overcome this issue, we propose credal ensemble distillation (CED), a novel framework that compresses a DE into a single model, CREDIT, for classification tasks. Instead of a single softmax probability distribution, CREDIT predicts class-wise probability intervals that define a credal set, a convex set of probability distributions, for uncertainty quantification. Empirical results on out-of-distribution detection benchmarks demonstrate that CED achieves superior or comparable uncertainty estimation compared to several existing baselines, while substantially reducing inference overhead compared to DE.
Fabio Cuzzolin, David Moens, Hans Hallez
AAAI2
2026 Direct interval propagation methods using neural-network surrogates for uncertainty quantification in physical systems surrogate model
Ghifari Adam Faza, Jolan Wauters, Fabio Cuzzolin, Hans Hallez, David Moens
Knowl. Based Syst.3
2026 A Review of Uncertainty Representation and Quantification in Neural Networks
abstract
Effectively estimating the uncertainty attached to neural network predictions thus becomes essential to improve robustness, reliability, and trustworthiness. This paper provides an overview of various methodologies for representing, quantifying, and distinguishing two major types of uncertainties (namely, 'aleatoric' and 'epistemic' uncertainty) in neural networks. The review covers classical probabilistic techniques such as Bayesian neural networks and deep ensembles, methods from generalized probability that leverage uncertainty representations such as Dirichlet distributions, belief functions, random sets, probability intervals, and credal sets, among others. Additionally, interval-based approaches employing interval models are also examined. We discuss the strengths and limitations of various methodologies and identify promising research directions for potential future exploration.
Fabio Cuzzolin, Keivan Shariatmadar, David Moens, Hans Hallez
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 TDCC: A Trustworthy Deep Credal Clustering Method for Uncertain Data
abstract
Deep clustering has achieved remarkable success in handling various types of real-world data, but often suffers from overconfidence, forcing ambiguous samples into specific clusters even when the evidence is insufficient. To address this limitation, we propose trustworthy deep credal clustering, a novel framework for uncertainty that integrates deep neural networks with the Dempster-Shafer Theory of evidence (DST). This method leverages credal cluster structures to enhance the model's robustness against uncertain data. Our model can refrain from assigning uncertain samples to a specific cluster, thereby reducing errors and enhancing the model's trustworthiness. Theoretically, we derive closed-form solutions for updating cluster memberships and prototypes, employing a coordinate descent strategy to rigorously optimize the objective function. Experiments on various datasets confirm that our proposed trustworthy clustering method leads to enhanced overall clustering effectiveness. Code is available at https://github.com/H1nkik/Trustworthy-Clustering.
Kuang Zhou, Fabio Cuzzolin
IEEE Trans. Cybern.3
2025 A Unified Evaluation Framework for Epistemic Predictions
abstract
Predictions of uncertainty-aware models are diverse, ranging from single point estimates (often averaged over prediction samples) to predictive distributions, to set-valued or credal-set representations. We propose a novel unified evaluation framework for uncertainty-aware classifiers, applicable to a wide range of model classes, which allows users to tailor the trade-off between accuracy and precision of predictions via a suitably designed performance metric. This makes possible the selection of the most suitable model for a particular real-world application as a function of the desired trade-off. Our experiments, concerning Bayesian, ensemble, evidential, deterministic, credal and belief function classifiers on the CIFAR-10, MNIST and CIFAR-100 datasets, show that the metric behaves as desired.
Shireen Kudukkil Manchingal, Muhammad Mubashar, Fabio Cuzzolin
AISTATS4
2025 Random-Set Neural Networks
abstract
Machine learning is increasingly deployed in safety-critical domains where erroneous predictions may lead to potentially catastrophic consequences, highlighting the need for learning systems to be aware of how confident they are in their own predictions: in other words, 'to know when they do not know’. In this paper, we propose a novel Random-Set Neural Network (RS-NN) approach to classification which predicts *belief functions* (rather than classical probability vectors) over the class list using the mathematics of *random sets*, i.e., distributions over the collection of *sets* of classes. RS-NN encodes the 'epistemic' uncertainty induced by training sets that are insufficiently representative or limited in size via the size of the convex set of probability vectors associated with a predicted belief function. Our approach outperforms state-of-the-art Bayesian and Ensemble methods in terms of accuracy, uncertainty estimation and out-of-distribution (OoD) detection on multiple benchmarks (CIFAR-10 vs SVHN/Intel-Image, MNIST vs FMNIST/KMNIST, ImageNet vs ImageNet-O). RS-NN also scales up effectively to large-scale architectures (e.g. WideResNet-28-10, VGG16, Inception V3, EfficientNetB2 and ViT-Base-16), exhibits remarkable robustness to adversarial attacks and can provide statistical guarantees in a conformal learning setting.
Shireen Kudukkil Manchingal, Muhammad Mubashar, Keivan Shariatmadar, Fabio Cuzzolin
ICLR5
2025 Credal Wrapper of Model Averaging for Uncertainty Estimation in Classification
abstract
This paper presents an innovative approach, called credal wrapper, to formulating a credal set representation of model averaging for Bayesian neural networks (BNNs) and deep ensembles (DEs), capable of improving uncertainty estimation in classification tasks. Given a finite collection of single predictive distributions derived from BNNs or DEs, the proposed credal wrapper approach extracts an upper and a lower probability bound per class, acknowledging the epistemic uncertainty due to the availability of a limited amount of distributions. Such probability intervals over classes can be mapped on a convex set of probabilities (a credal set) from which, in turn, a unique prediction can be obtained using a transformation called intersection probability transformation. In this article, we conduct extensive experiments on several out-of-distribution (OOD) detection benchmarks, encompassing various dataset pairs (CIFAR10/100 vs SVHN/Tiny-ImageNet, CIFAR10 vs CIFAR10-C, CIFAR100 vs CIFAR100-C and ImageNet vs ImageNet-O) and using different network architectures (such as VGG16, ResNet-18/50, EfficientNet B2, and ViT Base). Compared to the BNN and DE baselines, the proposed credal wrapper method exhibits superior performance in uncertainty estimation and achieves a lower expected calibration error on corrupted data.
Fabio Cuzzolin, Keivan Shariatmadar, David Moens, Hans Hallez
ICLR2
2025 Future themes in regulating artificial intelligence in investment management
abstract
We are witnessing the emergence of the “first generation” of AI and AI-adjacent soft and hard laws such as the EU AI Act or South Korea's Basic Act on AI. In parallel, existing industry regulations, such as GDPR, MIFID II or SM&CR, are being “retrofitted” and reinterpreted from the perspective of AI. In this paper we identify and analyze ten novel, “second generation” themes which are likely to become regulatory considerations in the near future: non-personal data, managerial accountability, robo-advisory, generative AI, privacy enhancing techniques (PETs), profiling, emergent behaviours, smart contracts, ESG and algorithm management. The themes have been identified on the basis of ongoing developments in AI, existing regulations and industry discussions. Prior to making any new regulatory recommendations we explore whether novel issues can be solved by existing regulations. The contribution of this paper is a comprehensive picture of emerging regulatory considerations for AI in investment management, as well as broader financial services, and the ways they might be addressed by regulations – future or existing ones.
Wojtek Buczynski, Felix Steffek, Mateja Jamnik, Fabio Cuzzolin, Barbara J. Sahakian
Comput. Law Secur. Rev.4
2025 CreINNs: Credal-Set Interval Neural Networks for Uncertainty Estimation in Classification Tasks
abstract
Effective uncertainty estimation is becoming increasingly attractive for enhancing the reliability of neural networks. This work presents a novel approach, termed Credal-Set Interval Neural Networks (CreINNs), for classification. CreINNs retain the fundamental structure of traditional Interval Neural Networks, capturing weight uncertainty through deterministic intervals. CreINNs are designed to predict an upper and a lower probability bound for each class, rather than a single probability value. The probability intervals can define a credal set, facilitating estimating different types of uncertainties associated with predictions. Experiments on standard multiclass and binary classification tasks demonstrate that the proposed CreINNs can achieve superior or comparable quality of uncertainty estimation compared to variational Bayesian Neural Networks (BNNs) and Deep Ensembles. Furthermore, CreINNs significantly reduce the computational complexity of variational BNNs during inference. Moreover, the effective uncertainty quantification of CreINNs is also verified when the input data are intervals.
Keivan Shariatmadar, Shireen Kudukkil Manchingal, Fabio Cuzzolin, David Moens, Hans Hallez
Neural Networks4
2025 Guest Editorial Special Issue on Effective Feature Fusion in Deep Neural Networks
Yanwei Pang, Fahad Shahbaz Khan, Fabio Cuzzolin
IEEE Trans. Neural Networks Learn. Syst.4
2024 Credal Learning Theory
abstract
Statistical learning theory is the foundation of machine learning, providing theoretical bounds for the risk of models learned from a (single) training set, assumed to issue from an unknown probability distribution. In actual deployment, however, the data distribution may (and often does) vary, causing domain adaptation/generalization issues. In this paper we lay the foundations for a `credal' theory of learning, using convex sets of probabilities (credal sets) to model the variability in the data-generating distribution. Such credal sets, we argue, may be inferred from a finite sample of training sets. Bounds are derived for the case of finite hypotheses spaces (both assuming realizability or not), as well as infinite model spaces, which directly generalize classical results.
Michele Caprio, Maryam Sultana, Eleni Elia, Fabio Cuzzolin
NeurIPS4
2024 Credal Deep Ensembles for Uncertainty Quantification
abstract
This paper introduces an innovative approach to classification called Credal Deep Ensembles (CreDEs), namely, ensembles of novel Credal-Set Neural Networks (CreNets). CreNets are trained to predict a lower and an upper probability bound for each class, which, in turn, determine a convex set of probabilities (credal set) on the class set. The training employs a loss inspired by distributionally robust optimization which simulates the potential divergence of the test distribution from the training distribution, in such a way that the width of the predicted probability interval reflects the epistemic uncertainty about the future data distribution. Ensembles can be constructed by training multiple CreNets, each associated with a different random seed, and averaging the outputted intervals. Extensive experiments are conducted on various out-of-distributions (OOD) detection benchmarks (CIFAR10/100 vs SVHN/Tiny-ImageNet, CIFAR10 vs CIFAR10-C, ImageNet vs ImageNet-O) and using different network architectures (ResNet50, VGG16, and ViT Base). Compared to Deep Ensemble baselines, CreDEs demonstrate higher test accuracy, lower expected calibration error, and significantly improved epistemic uncertainty estimation.
Fabio Cuzzolin, Shireen Kudukkil Manchingal, Keivan Shariatmadar, David Moens, Hans Hallez
NeurIPS2
2024 A Hybrid Graph Network for Complex Activity Detection in Video
abstract
Interpretation and understanding of video presents a challenging computer vision task in numerous fields - e.g. autonomous driving and sports analytics. Existing approaches to interpreting the actions taking place within a video clip are based upon Temporal Action Localisation (TAL), which typically identifies short-term actions. The emerging field of Complex Activity Detection (CompAD) extends this analysis to long-term activities, with a deeper understanding obtained by modelling the internal structure of a complex activity taking place within the video.We address the CompAD problem using a hybrid graph neural network which combines attention applied to a graph encoding the local (short-term) dynamic scene with a temporal graph modelling the overall long-duration activity. Our approach is as follows: i) Firstly, we propose a novel feature extraction technique which, for each video snippet, generates spatiotemporal ‘tubes’ for the active elements (‘agents’) in the (local) scene by detecting individual objects, tracking them and then extracting 3D features from all the agent tubes as well as the overall scene. ii) Next, we construct a local scene graph where each node (representing either an agent tube or the scene) is connected to all other nodes. Attention is then applied to this graph to obtain an overall representation of the local dynamic scene. iii) Finally, all local scene graph representations are interconnected via a temporal graph, to estimate the complex activity class together with its start and end time.The proposed framework outperforms all previous state-of-the-art methods on all three datasets including ActivityNet-1.3, Thumos-14, and ROAD.
Salman Khan 0004, Izzeddin Teeti, Andrew Bradley, Mohamed Elhoseiny 0001, Fabio Cuzzolin
WACV5
2024 Feature boosting with efficient attention for scene parsing
Vivek Singh Bawa, Shailza Sharma, Fabio Cuzzolin
Neurocomputing3
2023 ROAD-R: the autonomous driving dataset with logical requirements
abstract
Abstract Neural networks have proven to be very powerful at computer vision tasks. However, they often exhibit unexpected behaviors, acting against background knowledge about the problem at hand. This calls for models (i) able to learn from requirements expressing such background knowledge, and (ii) guaranteed to be compliant with the requirements themselves. Unfortunately, the development of such models is hampered by the lack of real-world datasets equipped with formally specified requirements. In this paper, we introduce the ROad event Awareness Dataset with logical Requirements (ROAD-R), the first publicly available dataset for autonomous driving with requirements expressed as logical constraints. Given ROAD-R, we show that current state-of-the-art models often violate its logical constraints, and that it is possible to exploit them to create models that (i) have a better performance, and (ii) are guaranteed to be compliant with the requirements themselves.
Eleonora Giunchiglia, Mihaela Catalina Stoian, Salman Khan 0004, Fabio Cuzzolin, Thomas Lukasiewicz
Mach. Learn.4
2023 ROAD: The Road Event Awareness Dataset for Autonomous Driving
abstract
Humans drive in a holistic fashion which entails, in particular, understanding dynamic road events and their evolution. Injecting these capabilities in autonomous vehicles can thus take situational awareness and decision making closer to human-level performance. To this purpose, we introduce the ROad event Awareness Dataset (ROAD) for Autonomous Driving, to our knowledge the first of its kind. ROAD is designed to test an autonomous vehicle's ability to detect road events, defined as triplets composed by an active agent, the action(s) it performs and the corresponding scene locations. ROAD comprises videos originally from the Oxford RobotCar Dataset, annotated with bounding boxes showing the location in the image plane of each road event. We benchmark various detection tasks, proposing as a baseline a new incremental algorithm for online road event awareness termed 3D-RetinaNet. We also report the performance on the ROAD tasks of Slowfast and YOLOv5 detectors, as well as that of the winners of the ICCV2021 ROAD challenge, which highlight the challenges faced by situation awareness in autonomous driving. ROAD is designed to allow scholars to investigate exciting tasks such as complex (road) activity detection, future event anticipation and continual learning. The dataset is available at https://github.com/gurkirt/road-dataset; the baseline can be found at https://github.com/gurkirt/3D-RetinaNet.
Gurkirt Singh, Stephen Akrigg, Manuele Di Maio, Valentina Fontana, Reza Javanmard Alitappeh, Salman Khan 0004, Suman Saha 0001, Kossar Jeddi Saravi, Farzad Yousefi, Jacob Culley, Tom Nicholson, Jordan Omokeowa, Stanislao Grazioso, Andrew Bradley, Giuseppe Di Gironimo, Fabio Cuzzolin
IEEE Trans. Pattern Anal. Mach. Intell.16
2022 Vision-based Intention and Trajectory Prediction in Autonomous Vehicles: A Survey
abstract
This survey targets intention and trajectory prediction in Autonomous Vehicles (AV), as AV companies compete to create dedicated prediction pipelines to avoid collisions. The survey starts with a formal definition of the prediction problem and highlights its challenges, to then critically compare the models proposed in the last 2-3 years in terms of how they overcome these challenges. Further, it lists the latest methodological and technical trends in the field and comments on the efficacy of different machine learning blocks in modelling various aspects of the prediction problem. It also summarises the popular datasets and metrics used to evaluate prediction models, before concluding with the possible research gaps and future directions.
Izzeddin Teeti, Salman Khan 0004, Ajmal Shahbaz, Andrew Bradley, Fabio Cuzzolin
IJCAI5
2022 Semantics-Driven Generative Replay for Few-Shot Class Incremental Learning
abstract
We deal with the problem of few-shot class incremental learning (FSCIL), which requires a model to continuously recognize new categories for which limited training data are available. Existing FSCIL methods depend on prior knowledge to regularize the model parameters for combating catastrophic forgetting. Devising an effective prior in a low-data regime, however, is not trivial. The memory-replay based approaches from the fully-supervised class incremental learning (CIL) literature cannot be used directly for FSCIL as the generative memory-replay modules of CIL are hard to train from few training samples. However, generative replay can tackle both the stability and plasticity of the models simultaneously by generating a large number of class-conditional samples. Convinced by this fact, we propose a generative modeling-based FSCIL framework using the paradigm of memory-replay in which a novel conditional few-shot generative adversarial network (GAN) is incrementally trained to produce visual features while ensuring the stability-plasticity trade-off through novel loss functions and combating the mode-collapse problem effectively. Furthermore, the class-specific synthesized visual features from the few-shot GAN are constrained to match the respective latent semantic prototypes obtained from a well-defined semantic space. We find that the advantages of this semantic restriction is two-fold, in dealing with forgetting, while making the features class-discernible. The model requires a single per-class prototype vector to be maintained in a dynamic memory buffer. Experimental results on the benchmark and large-scale CiFAR-100, CUB-200, and Mini-ImageNet confirm the superiority of our model over the current FSCIL state of the art.
Aishwarya Agarwal, Biplab Banerjee, Fabio Cuzzolin, Subhasis Chaudhuri
ACM Multimedia3
2022 An intelligent system for complex violence pattern analysis and detection
abstract
Video surveillance has shown encouraging outcomes to monitor human activities and prevent crimes in real time. To this extent, violence detection (VD) has received substantial attention from the research community due to its vast applications, such as ensuring security over public areas and industrial settings through smart machine intelligence. However, because of changing illumination, complex background and low resolution, the analysis of violence patterns remains challenging in the industrial video surveillance domain. In this paper, we propose a computationally intelligent VD approach to precisely detect violent scenes through deep analysis of surveillance video sequential patterns. First, the video stream acquired through the vision sensor is processed by a lightweight convolutional neural network (CNN) for the segmentation of important shots. Next, temporal optical flow features are extracted from the informative shots via a residential optical flow CNN. These are concatenated with appearance-invariant features extracted from a Darknet CNN model. Finally, a multilayer long short-term memory network is plugged to generate the final feature map for learning the violence patterns in a sequence of frames. In addition, we contribute to the existing surveillance VD data set by considering its indoor and outdoor scenarios separately for the proposed method's evaluation, achieving a 2% increase in accuracy over surveillance fight data set. Experiments also show encouraging results over the state of the art on other challenging benchmark data sets.
Fath U Min Ullah, Mohammad S. Obaidat, Khan Muhammad 0001, Amin Ullah, Sung Wook Baik, Fabio Cuzzolin, Joel J. P. C. Rodrigues, Victor Hugo C. de Albuquerque
Int. J. Intell. Syst.6
2021 Spatiotemporal Deformable Scene Graphs for Complex Activity Detection
Salman Khan 0001, Fabio Cuzzolin
BMVC2
2021 SVD-GAN for Real-Time Unsupervised Video Anomaly Detection
R. Dinesh Jackson Samuel, Fabio Cuzzolin
BMVC2
2021 DeepSmoke: Deep learning model for smoke detection and segmentation in outdoor environments
Salman Khan 0004, Khan Muhammad 0001, Tanveer Hussain 0001, Javier Del Ser, Fabio Cuzzolin, Siddhartha Bhattacharyya 0001, Zahid Akhtar, Victor Hugo C. de Albuquerque
Expert Syst. Appl.5
2020 Video-Based Crowd Counting Using a Multi-scale Optical Flow Pyramid Network
Mohammad Asiful Hossain, Kevin Cannons, Daesik Jang, Fabio Cuzzolin
ACCV (5)4
2020 Generalized Zero-Shot Learning using Generated Proxy Unseen Samples and Entropy Separation
abstract
The recent generative model-driven Generalized Zero-shot Learning (GZSL) techniques overcome the prevailing issue of the model bias towards the seen classes by synthesizing the visual samples of the unseen classes through leveraging the corresponding semantic prototypes. Although such approaches significantly improve the GZSL performance due to data augmentation, they violate the principal assumption of GZSL regarding the unavailability of semantic information of unseen classes during training. In this work, we propose to use a generative model (GAN) for synthesizing the visual proxy samples while strictly adhering to the standard assumptions of the GZSL. The aforementioned proxy samples are generated by exploring the early training regime of the GAN. We hypothesize that such proxy samples can effectively be used to characterize the average entropy of the label distribution of the samples from the unseen classes. Further, we train a classifier on the visual samples from the seen classes and proxy samples using entropy separation criterion such that an average entropy of the label distribution is low and high, respectively, for the visual samples from the seen classes and the proxy samples. Such entropy separation criterion generalizes well during testing where the samples from the unseen classes exhibit higher entropy than the entropy of the samples from the seen classes. Subsequently, low and high entropy samples are classified using supervised learning and ZSL rather than GZSL. We show the superiority of the proposed method by experimenting on AWA1, CUB, HMDB51, and UCF101 datasets.
Omkar Gune, Biplab Banerjee, Subhasis Chaudhuri, Fabio Cuzzolin
ACM Multimedia4
2020 Evidence Combination Based on Credal Belief Redistribution for Pattern Classification
abstract
Evidence theory, also called belief function theory, provides an efficient tool to represent and combine uncertain information for pattern classification. Evidence combination can be interpreted, in some applications, as classifier fusion. The sources of evidence corresponding to multiple classifiers usually exhibit different classification qualities, and they are often discounted using different weights before combination. In order to achieve the best possible fusion performance, a new credal belief redistribution (CBR) method is proposed to revise such evidence. The rationale of CBR consists of transferring belief from one class not just to other classes, but also to the associated disjunctions of classes (i.e., meta-classes). As classification accuracy for different objects in a given classifier can also vary, the evidence is revised according to prior knowledge mined from its training neighbors. If the selected neighbors are relatively close to the evidence, a large amount of belief will be discounted for redistribution. Otherwise, only a small fraction of belief will enter the redistribution procedure. An imprecision matrix estimated based on these neighbors is employed to specifically redistribute the discounted beliefs. This matrix expresses the likelihood of misclassification (i.e., the probability of a test pattern belonging to a class different from the one assigned to it by the classifier). In CBR, the discounted beliefs are divided into two parts. One part is transferred between singleton classes, whereas the other is cautiously committed to the associated meta-classes. By doing this, one can efficiently reduce the chance of misclassification by modeling partial imprecision. The multiple revised pieces of evidence are finally combined by the Dempster-Shafer rule to reduce uncertainty and further improve classification accuracy. The effectiveness of CBR is extensively validated on several real datasets from the UCI repository and critically compared with that of other related fusion methods.
Zhunga Liu, Yu Liu 0005, Jean Dezert, Fabio Cuzzolin
IEEE Trans. Fuzzy Syst.4
2018 TraMNet - Transition Matrix Network for Efficient Action Tube Proposals
Gurkirt Singh, Suman Saha 0001, Fabio Cuzzolin
ACCV (6)3
2018 Incremental Tube Construction for Human Action Detection
Harkirat S. Behl, Michael Sapienza, Gurkirt Singh, Suman Saha 0001, Fabio Cuzzolin, Philip Torr 0001
BMVC5
2018 A Belief-Theoretical Approach to Example-Based Pose Estimation
abstract
In example-based human pose estimation, the configuration of an evolving object is sought given visual evidence, having to rely uniquely on a set of sample images. We assume here that, at each time instant of a training session, a number of feature measurements is extracted from the available images, while ground truth is provided in the form of the true object pose. In this scenario, a sensible approach consists in learning maps from features to poses, using the information provided by the training set. In particular, multivalued mappings linking feature values to set of training poses can be constructed. To this purpose we propose a belief modeling regression (BMR) approach in which a probability measure on any individual feature space maps to a convex set of probabilities on the set of training poses, in a form of a belief function. Given a test image, its feature measurements translate into a collection of belief functions on the set of training poses which, when combined, yield there an entire family of probability distributions. From the latter either a single central pose estimate or a set of extremal ones can be computed, together with a measure of how reliable the estimate is. Contrarily to other competing models, in BMR the sparsity of the training samples can be taken into account to model the level of uncertainty associated with these estimates. We illustrate BMR's performance in an application to human pose recovery, showing how it outperforms our implementation of both relevant vector machine and Gaussian process regression. Finally, we discuss motivation and advantages of the proposed approach with respect to its most direct competitors.
Wenjuan Gong, Fabio Cuzzolin
IEEE Trans. Fuzzy Syst.2
2017 AMTnet: Action-Micro-Tube Regression by End-to-end Trainable Deep Architecture
abstract
Dominant approaches to action detection can only provide sub-optimal solutions to the problem, as they rely on seeking frame-level detections, to later compose them into ‘action tubes’ in a post-processing step. With this paper we radically depart from current practice, and take a first step towards the design and implementation of a deep network architecture able to classify and regress whole video subsets, so providing a truly optimal solution of the action detection problem. In this work, in particular, we propose a novel deep net framework able to regress and classify 3D region proposals spanning two successive video frames, whose core is an evolution of classical region proposal networks (RPNs). As such, our 3D-RPN net is able to effectively encode the temporal aspect of actions by purely exploiting appearance, as opposed to methods which heavily rely on expensive flow maps. The proposed model is end-to-end trainable and can be jointly optimised for action localisation and classification in a single step. At test time the network predicts ‘micro-tubes’ encompassing two successive frames, which are linked up into complete action tubes via a new algorithm which exploits the temporal encoding learned by the network and cuts computation time by 50%. Promising results on the J-HMDB-21 and UCF-101 action detection datasets show that our model does outperform the state-of-the-art when relying purely on appearance.
Suman Saha 0001, Gurkirt Singh, Fabio Cuzzolin
ICCV3
2017 Online Real-Time Multiple Spatiotemporal Action Localisation and Prediction
abstract
We present a deep-learning framework for real-time multiple spatio-temporal (S/T) action localisation and classification. Current state-of-the-art approaches work offline, and are too slow to be useful in real-world settings. To overcome their limitations we introduce two major developments. Firstly, we adopt real-time SSD (Single Shot Multi-Box Detector) CNNs to regress and classify detection boxes in each video frame potentially containing an action of interest. Secondly, we design an original and efficient online algorithm to incrementally construct and label ‘action tubes’ from the SSD frame level detections. As a result, our system is not only capable of performing S/T detection in real time, but can also perform early action prediction in an online fashion. We achieve new state-of-the-art results in both S/T action localisation and early action prediction on the challenging UCF101-24 and J-HMDB-21 benchmarks, even when compared to the top offline competitors. To the best of our knowledge, ours is the first real-time (up to 40fps) system able to perform online S/T action localisation on the untrimmed videos of UCF101-24.
Gurkirt Singh, Suman Saha 0001, Michael Sapienza, Philip Torr 0001, Fabio Cuzzolin
ICCV5
2017 The Total Belief Theorem
Chunlai Zhou, Fabio Cuzzolin
UAI2
2017 Active Incremental Recognition of Human Activities in a Streaming Context
Rocco De Rosa, Ilaria Gori, Fabio Cuzzolin, Nicolò Cesa-Bianchi
Pattern Recognit. Lett.3
2016 Deep Learning for Detecting Multiple Space-Time Action Tubes in Videos
Suman Saha 0001, Gurkirt Singh, Michael Sapienza, Philip Torr 0001, Fabio Cuzzolin
BMVC5
2016 Belief functions: Theory and applications (BELIEF 2014)
Fabio Cuzzolin
Int. J. Approx. Reason.1
2015 Robust classification of multivariate time series by imprecise hidden Markov models
Alessandro Antonucci 0001, Rocco De Rosa, Alessandro Giusti, Fabio Cuzzolin
Int. J. Approx. Reason.4
2015 Robust Temporally Coherent Laplacian Protrusion Segmentation of 3D Articulated Bodies
Fabio Cuzzolin, Diana Mateus, Radu Horaud
Int. J. Comput. Vis.1
2014 Online Action Recognition via Nonparametric Incremental Learning
Rocco De Rosa, Nicolò Cesa-Bianchi, Ilaria Gori, Fabio Cuzzolin
BMVC4
2014 Learning Discriminative Space-Time Action Parts from Weakly Labelled Videos
Michael Sapienza, Fabio Cuzzolin, Philip Torr 0001
Int. J. Comput. Vis.2
2014 Learning Pullback HMM Distances
abstract
Recent work in action recognition has exposed the limitations of methods which directly classify local features extracted from spatio-temporal video volumes. In opposition, encoding the actions' dynamics via generative dynamical models has a number of attractive features: however, using all-purpose distances for their classification does not necessarily deliver good results. We propose a general framework for learning distance functions for generative dynamical models, given a training set of labelled videos. The optimal distance function is selected among a family of pullback ones, induced by a parametrised automorphism of the space of models. We focus here on hidden Markov models and their model space, and design an appropriate automorphism there. Experimental results are presented which show how pullback learning greatly improves action recognition performances with respect to base distances.
Fabio Cuzzolin, Michael Sapienza
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 $L_{p}$ Consonant Approximations of Belief Functions
abstract
In this paper, we solve the problem of approximating a belief measure with a necessity measure or “consonant belief function” in a geometric framework. Consonant belief functions form a simplicial complex in both the space of all belief functions and the space of all mass vectors: Partial approximations are first sought in each component of the complex, while global solutions are selected among them. As a first step in this line of study, we seek here approximations that minimize Lpnorms. Approximations in the mass space can be interpreted in terms of mass redistribution, while approximations in the belief space generalize the maximal outer consonant approximation. We compare them with each other and with other classical approximations and illustrate them with the help of a running example.
Fabio Cuzzolin
IEEE Trans. Fuzzy Syst.1
2013 Belief modeling regression for pose estimation
Fabio Cuzzolin, Wenjuan Gong
FUSION1
2012 Learning discriminative space-time actions from weakly labelled videos
abstract
Current state-of-the-art action classification methods extract feature representations from the entire video clip in which the action unfolds, however this representation may include irrelevant scene context and movements which are shared amongst multiple action classes. For example, a waving action may be performed whilst walking, however if the walking movement and scene context appear in other action classes, then they should not be included in a waving movement classifier. In this work, we propose an action classification framework in which more discriminative action subvolumes are learned in a weakly supervised setting, owing to the difficulty of manually labelling massive video datasets. The learned models are used to simultaneously classify video clips and to localise actions to a given space-time subvolume. Each subvolume is cast as a bag-offeatures (BoF) instance in a multiple-instance-learning framework, which in turn is used to learn its class membership. We demonstrate quantitatively that even with single fixedsized subvolumes, the classification performance of our proposed algorithm is superior to the state-of-the-art BoF baseline on the majority of performance measures, and shows promise for space-time action localisation on the most challenging video datasets.
Michael Sapienza, Fabio Cuzzolin, Philip Torr 0001
BMVC2
2012 On the relative belief transform
Fabio Cuzzolin
Int. J. Approx. Reason.1
2011 On Consistent Approximations of Belief Functions in the Mass Space
Fabio Cuzzolin
ECSQARU1
2010 Credal Sets Approximation by Lower Probabilities: Application to Credal Networks
Alessandro Antonucci 0001, Fabio Cuzzolin
IPMU2
2010 The geometry of consonant belief functions: Simplicial complexes of necessity measures
Fabio Cuzzolin
Fuzzy Sets Syst.1
2010 Three alternative combinatorial formulations of the theory of evidence
abstract
In this paper we introduce three alternative combinatorial formulations of the theory of evidence (ToE), by proving that both plausibility and commonality functions share the structure of “sum function” with belief functions. We compute their Moebius inverses, which we call basic plausibility and c ommonality assignments. In the framework of the geometric approach to uncertainty measures the equivalence of the associated formulations of the ToE is mirrored by the geometric congruence of the related simplices. We can therefore describe the point-wise geometry of these sum functions in terms of rigid transformations mapping them onto each other. Combination rules can be applied to plausibility and commonality functions through their Moebius inverses, leading to interesting applications of such inverses to the probabilistic transformation problem.
Fabio Cuzzolin
Intell. Data Anal.1
2010 Credal Semantics of Bayesian Transformations in Terms of Probability Intervals
abstract
In this paper, we propose a credal representation of the interval probability associated with a belief function (b.f.) and show how it relates to several classical Bayesian transformations of b.f.'s through the notion of ¿focus¿ of a pair of simplices. While a b.f. corresponds to a polytope of probabilities consistent with it, the related interval probability is geometrically represented by a pair of upper and lower simplices. Starting from the interpretation of the pignistic function as the center of mass of the credal set of consistent probabilities, we prove that the relative belief of singletons, the relative plausibility of singletons, and the intersection probability can all be described as the foci of different pairs of simplices in the region of all probability measures. The formulation of frameworks similar to the transferable belief model for such Bayesian transformations appears then at hand.
Fabio Cuzzolin
IEEE Trans. Syst. Man Cybern. Part B1
2009 Complexes of Outer Consonant Approximations
Fabio Cuzzolin
ECSQARU1
2009 The Intersection Probability and Its Properties
Fabio Cuzzolin
ECSQARU1
2008 Coherent Laplacian 3-D protrusion segmentation
abstract
In this paper, an analysis of locally linear embedding (LLE) in the context of clustering is developed. As LLE conserves the local affine coordinates of points, shape protrusions as high-curvature regions of the surface are preserved. Also, LLEpsilas covariance constraint acts as a force stretching those protrusions and making them wider separated and lower dimensional. A novel scheme for unsupervised body-part segmentation along time sequences is thus proposed in which 3-D shapes are clustered after embedding. Clusters are propagated in time, and merged or split in an unsupervised fashion to accommodate changes of the body topology. Comparisons on synthetic, and real data with ground truth, are run with direct segmentation in 3-D by EM clustering and ISOMAP-based clustering. Robustness and the effects of topology transitions are discussed.
Fabio Cuzzolin, Diana Mateus, David Knossow, Edmond Boyer, Radu Horaud
CVPR1
2008 Articulated shape matching using Laplacian eigenfunctions and unsupervised point registration
abstract
Matching articulated shapes represented by voxel-sets reduces to maximal sub-graph isomorphism when each set is described by a weighted graph. Spectral graph theory can be used to map these graphs onto lower dimensional spaces and match shapes by aligning their embeddings in virtue of their invariance to change of pose. Classical graph isomorphism schemes relying on the ordering of the eigenvalues to align the eigenspaces fail when handling large data-sets or noisy data. We derive a new formulation that finds the best alignment between two congruent K-dimensional sets of points by selecting the best subset of eigenfunctions of the Laplacian matrix. The selection is done by matching eigenfunction signatures built with histograms, and the retained set provides a smart initialization for the alignment problem with a considerable impact on the overall performance. Dense shape matching casted into graph matching reduces then, to point registration of embeddings under orthogonal transformations; the registration is solved using the framework of unsupervised clustering and the EM algorithm. Maximal subset matching of non identical shapes is handled by defining an appropriate outlier class. Experimental results on challenging examples show how the algorithm naturally treats changes of topology, shape variations and different sampling densities.
Diana Mateus, Radu Horaud, David Knossow, Fabio Cuzzolin, Edmond Boyer
CVPR4
2008 On the Credal Structure of Consistent Probabilities
Fabio Cuzzolin
JELIA1
2008 Dual Properties of the Relative Belief of Singletons
Fabio Cuzzolin
PRICAI1
2008 Alternative Formulations of the Theory of Evidence Based on Basic Plausibility and Commonality Assignments
Fabio Cuzzolin
PRICAI1
2008 A Geometric Approach to the Theory of Evidence
abstract
In this paper, we propose a geometric approach to the theory of evidence based on convex geometric interpretations of its two key notions of belief function (b.f.) and Dempster's sum. On one side, we analyze the geometry of b.f.'s as points of a polytope in the Cartesian space called belief space, and discuss the intimate relationship between basic probability assignment and convex combination. On the other side, we study the global geometry of Dempster's rule by describing its action on those convex combinations. By proving that Dempster's sum and convex closure commute, we are able to depict the geometric structure of conditional subspaces, i.e., sets of b.f.'s conditioned by a given functionb. Natural applications of these geometric methods to classical problems such as probabilistic approximation and canonical decomposition are outlined.
Fabio Cuzzolin
IEEE Trans. Syst. Man Cybern. Part C1
2007 On the Orthogonal Projection of a Belief Function
Fabio Cuzzolin
ECSQARU1
2007 Articulated Shape Matching by Robust Alignment of Embedded Representations
abstract
In this paper we propose a general framework to solve the articulated shape matching problem, formulated as finding point-to-point correspondences between two shapes represented by 2-D or 3-D point clouds. The original point- sets are embedded in a spectral representation and the actual matching is carried out in the embedded space. We analyze the advantages of this choice as well as the reasons for which the task remains a difficult one. In particular, we show that although embedded-space matching still has intrinsic combinatorial difficulties, it can be solved by searching for an optimal orthogonal transformation that aligns the two shape embeddings. Relying on the model based clustering formalism, we propose a probabilistic formulation which casts the matching into an EM algorithm. Outliers are properly handled by the algorithm and a simple strategy is adopted to initialize it. Experiments are performed with three embedding methods (Isomap, LLE, and Laplacian embedding) and with 3-D voxelsets representing a human-motion sequence.
Diana Mateus, Fabio Cuzzolin, Radu Horaud, Edmond Boyer
ICCV2
2007 Articulated Shape Matching Using Locally Linear Embedding and Orthogonal Alignment
abstract
In this paper we propose a method for matching articulated shapes represented as large sets of 3D points by aligning the corresponding embedded clouds generated by locally linear embedding. In particular we show that the problem is equivalent to aligning two sets of points under an orthogonal transformation acting onto the d-dimensional embeddings. The method may well be viewed as belonging to the model-based clustering framework and is implemented as an EM algorithm that alternates between the estimation of correspondences between data-points and the estimation of an optimal alignment transformation. Correspondences are initialized by embedding one set of data- points onto the other one through out-of-sample extension. Results for pairs of voxelsets representing moving persons are presented. Empirical evidence on the influence of the dimension of the embedding space is provided, suggesting that working with higher-dimensional spaces helps matching in challenging real-world scenarios, without collateral effects on the convergence.
Diana Mateus, Fabio Cuzzolin, Radu Horaud, Edmond Boyer
ICCV2
2007 Two New Bayesian Approximations of Belief Functions Based on Convex Geometry
abstract
In this paper, we analyze from a geometric perspective the meaningful relations taking place between belief and probability functions in the framework of the geometric approach to the theory of evidence. Starting from the case of binary domains, we identify and study three major geometric entities relating a generic belief function (b.f.) to the set of probabilities P: 1) the dual line connecting belief and plausibility functions; 2) the orthogonal complement of P; and 3) the simplex of consistent probabilities. Each of them is in turn associated with a different probability measure that depends on the original b.f. We focus in particular on the geometry and properties of the orthogonal projection of a b.f. onto P and its intersection probability, provide their interpretations in terms of degrees of belief, and discuss their behavior with respect to affine combination.
Fabio Cuzzolin
IEEE Trans. Syst. Man Cybern. Part B1
2006 Using Bilinear Models for View-invariant Action and Identity Recognition
abstract
Human identification from gait is a challenging task in realistic surveillance scenarios in which people walking along arbitrary directions are imaged by a single camera. In this paper, motivated by the view-invariance issue in the human ID from gait problem, we address the general problem of classifying the "content" of human motions of unknown "style". Given a dataset of sequences in which different people walking normally are seen from several wellseparated views, we propose a three-layer scheme based on bilinear models, in which image sequences are mapped to observation vectors of fixed dimension using Markov modeling. We test our approach on the CMU Mobo database, showing how bilinear separation outperforms other approaches, opening the way to view- and action-invariant identity recognition, as well as subject- and view-invariant action recognition.
Fabio Cuzzolin
CVPR (2)1
2005 On the properties of relative plausibilities
abstract
In this paper we investigate the properties of the relative plausibility function, the probability built by normalizing the plausibilities of singletons associated with a belief function. On one side, we stress how this probability is a perfect representative of the original belief function when combined with any arbitrary probability through Dempster's rule. This leads to conjecture that this function should also be the solution of the probabilistic approximation problem, formulated naturally in terms of Dempster's rule. On the other side, the geometric properties of relative plausibilities are studied in the context of the geometric approach to the theory of evidence, yielding a description of the representation property which suggests a sketch for the general proof of our conjecture.
Fabio Cuzzolin
SMC1
2004 Action modeling with volumetric data
abstract
In this paper we propose and test an action recognition algorithm in which the images of the scene captured by a significant number of cameras are first used to generate a volumetric representation of a moving human body in terms of voxsets by means of volumetric intersection. The recognition stage is then performed directly on 3D data, allowing the system to avoid critical problems like viewpoint dependence and motion trajectory variability. Suitable features are extracted from the voxset representing the body and fed to a classical hidden Markov model to produce a finite-state description of the motion.
Fabio Cuzzolin, Augusto Sarti, Stefano Tubaro
ICIP1
2004 Invariant action classification with volumetric data
abstract
We propose an action recognition algorithm in which the image sequences capturing a moving human body produced by a significant number of cameras are first used to generate a volumetric representation of the body by means of volumetric intersection. Classification is then performed directly on 3D data, making the system inherently insensitive to viewpoint dependence and motion trajectory variability. Suitable features are extracted from the voxset approximating the body, and fed to a hidden Markov model to produce a finite-state description of the motion. The Kullback-Leibler distance is finally used to classify new sequences.
Fabio Cuzzolin, Augusto Sarti, Stefano Tubaro
MMSP1
2004 Geometry of Dempster's rule of combination
abstract
In this paper, we analyze Shafer's belief functions (BFs) as geometric entities, focusing in particular on the geometric behavior of Dempster's rule of combination in the belief space, i.e., the set Stheta of all the admissible BFs defined over a given finite domain theta. The study of the orthogonal sums of affine subspaces allows us to unveil a convex decomposition of Dempster's rule of combination in terms of Bayes' rule of conditioning and prove that under specific conditions orthogonal sum and affine closure commute. A direct consequence of these results is the simplicial shape of the conditional subspaces , i.e., the sets of all the possible combinations of a given BF s. We show how Dempster's rule exhibits a rather elegant behavior when applied to BFs assigning the same mass to a fixed subset (constant mass loci). The resulting affine spaces have a common intersection that is characteristic of the conditional subspace, called focus. The affine geometry of these foci eventually suggests an interesting geometric construction of the orthogonal sum of two BFs.
Fabio Cuzzolin
IEEE Trans. Syst. Man Cybern. Part B1
1997 Using Hidden Markov Models and Dynamic Size Functions for Gesture Recognition
Andrea Sorrentino, Fabio Cuzzolin, Ruggero Frezza
BMVC2