Marco Cristani

dblp:58/2811 · DBLP profile ↗
← Back
139ranked-venue papers
11as first author
36since 2021 · last 2026
0000-0002-0523-6042ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 84 · 5 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 84 · 9 first-author · 15 since 2021Systems, architecture and hardware · 11 · 10 since 2021Software engineering, systems software and programming languages · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 A Comprehensive Survey on Deep Learning-based Predictive Maintenance
abstract
With the advent of Industrial 4.0 and the push toward Industry 5.0, the data generated by the industries have become surprisingly large. This abundance of data significantly boosts machine and deep learning models for Predictive Maintenance (PdM). The PdM plays a vital role in extending the lifespan of industrial equipment and machines while also helping to reduce the risk of unscheduled downtime. Given its multidisciplinary nature, the field of PdM has been approached from many different angles: this comprehensive survey aims at providing an up-to-date overview focused on all the learning-based industrial PdM strategies, discussing weaknesses and strengths. The survey is based on the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) methodological flow, allowing a systematic and complete review of the literature. In particular, firstly, we explore the main learning models used for PdM, mainly Convolutional Neural Networks (ConvNets), Autoencoders (AEs), Generative Adversarial Networks (GANs), and Transformers, also giving an overview of the newest models such as diffusion models and foundation models. Then, we discuss the main learning paradigms applied to PdM, i.e., supervised, unsupervised, ensemble, transfer, federated, and reinforcement learning. Furthermore, this work discusses the pipeline of the data-driven PdM and its benefits, practical applications, datasets, and benchmarks. In addition, the evaluation metrics for each PdM stage and the state-of-the-art hardware devices used are discussed. Finally, the challenges and future work are presented.
Dong Seon Cheng, Francesco Setti, Franco Fummi, Marco Cristani, Luigi Capogrosso
ACM Trans. Embed. Comput. Syst.5
2026 WP-FSCIL: A Well-Prepared Few-Shot Class-Incremental Learning Framework for Pill Recognition
abstract
Few-shot Class-incremental Pill Recognition (FSCIPR) aims to develop an automatic pill recognition system that requires only a few training data and can continuously adapt to new classes, providing technical support for applications in hospitals, portable apps, and assistance for visually impaired individuals. This task faces three core challenges: overfitting, fine-grained classification problems, and catastrophic forgetting. We propose the Well-Prepared Few-shot Class-incremental Learning (WP-FSCIL) framework, which addresses overfitting through a parameter-freezing strategy, enhances the robustness and discriminative power of backbone features with Center-Triplet (CT) loss and supervised contrastive loss for fine-grained classification, and alleviates catastrophic forgetting using a multi-dimensional Knowledge Distillation (KD) strategy based on flexible Pseudo-feature Synthesis (PFS). By flexibly synthesizing any number of old-class features, the PFS strategy resolves the issue of insufficient samples in the KD process, enabling Response-based KD (KD1) and Relation-based KD (KD2) to comprehensively preserve old knowledge. The effectiveness of WP-FSCIL has been validated through experiments conducted on two publicly available pill datasets. These experiments show that WP-FSCIL outperforms existing state-of-the-art methods, demonstrating its superior performance.
Chen Li 0022, Marco Cristani, Hongzan Sun, Marcin Grzegorzek, Huiling Chen 0001
IEEE J. Biomed. Health Informatics3
2025 Seeing the Abstract: Translating the Abstract Language for Vision Language Models
abstract
Natural language goes beyond dryly describing visual content. It contains rich abstract concepts to express feeling, creativity and properties that cannot be directly perceived. Yet, current research in Vision Language Models (VLMs) has not shed light on abstract-oriented language. Our research breaks new ground by uncovering its wide presence and under-estimated value, with extensive analysis. Particularly, we focus our investigation on the fashion domain, a highly-representative field with abstract expressions. By analyzing recent large-scale multimodal fashion datasets, we find that abstract terms have a dominant presence, rivaling the concrete ones, providing novel information, and being useful in the retrieval task. However, a critical challenge emerges: current general-purpose or fashion-specific VLMs are pre-trained with databases that lack sufficient abstract words in their text corpora, thus hindering their ability to effectively represent abstract-oriented language. We propose a training-free and model-agnostic method, Abstract-to-Concrete Translator (ACT), to shift abstract representations towards well-represented concrete ones in the VLM latent space, using pre-trained models and existing multi-modal databases. On the text-to-image retrieval task, despite being training-free, ACT outperforms the fine-tuned VLMs in both same- and cross-dataset settings, exhibiting its effectiveness with a strong generalization capability. Moreover, the improvement introduced by ACT is consistent with various VLMs, making it a plug-and-play solution.
Davide Talon, Federico Girella, Marco Cristani, Yiming Wang 0002
CVPR4
2025 Human-Centered Digital Twin for Industry 5.0
abstract
Moving beyond the automation-driven paradigm of Industry 4.0, Industry 5.0 emphasizes human-centric industrial systems where human creativity and instincts complement precise and advanced machines. With this new paradigm, there is a growing need for resource-efficient and user-preferred manufac-turing solutions that integrate humans into industrial processes. Unfortunately, methodologies for incorporating human elements into industrial processes remain underdeveloped. In this work, we present the first pipeline for the creation of a human-centered Digital Twin (DT), leveraging Unreal Engine's MetaHuman technology to track worker alertness in real-time. Our findings demonstrate the potential of integrating Artificial Intelligence (AI) and human-centered design within Industry 5.0 to enhance both worker safety and industrial efficiency.
Francesco Biondani, Luigi Capogrosso, Nicola Dall'Ora, Enrico Fraccaroli, Marco Cristani, Franco Fummi
DATE5
2025 Towards Real Unsupervised Anomaly Detection Via Confident Meta-Learning
abstract
So-called unsupervised anomaly detection is better described as semi-supervised, as it assumes all training data are nominal. This assumption simplifies training but requires manual data curation, introducing bias and limiting adaptability. We propose Confident Meta-learning (CoMet), a novel training strategy that enables deep anomaly detection models to learn from uncurated datasets where nominal and anomalous samples coexist, eliminating the need for explicit filtering. Our approach integrates Soft Confident Learning, which assigns lower weights to low-confidence samples, and Meta-Learning, which stabilizes training by regularizing updates based on training validation loss covariance. This prevents overfitting and enhances robustness to noisy data. CoMet is model-agnostic and can be applied to any anomaly detection method trainable via gradient descent. Experiments on MVTec-AD, VIADUCT, and KSDD2 with two state-of-the-art models demonstrate the effectiveness of our approach, consistently improving over the baseline methods, remaining insensitive to anomalies in the training set, and setting a new state-of-the-art across all datasets. Code is available at https://github.com/aqeeelmirza/CoMet
Muhammad Aqeel, Shakiba Sharifi, Marco Cristani, Francesco Setti
ICCV3
2025 LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing
Federico Girella, Davide Talon, Zanxi Ruan, Yiming Wang 0002, Marco Cristani
ICCV6
2025 Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues
Francesco Taioli, Edoardo Zorzi, Gianni Franchi, Alberto Castellini, Alessandro Farinelli, Marco Cristani, Yiming Wang 0002
ICCV6
2025 Pre-trained Multiple Latent Variable Generative Models are Good Defenders Against Adversarial Attacks
abstract
Attackers can deliberately perturb classifiers' input with subtle noise, altering final predictions. Among proposed countermeasures, adversarial purification employs generative networks to preprocess input images, filtering out adversarial noise. In this study, we propose specific generators, defined Multiple Latent Variable Generative Models (MLVGMs), for adversarial purification. These models possess multiple latent variables that naturally disentangle coarse from fine features. Taking advantage of these properties, we autoencode images to maintain class-relevant information, while discarding and re-sampling any detail, including adversarial noise. The procedure is completely training-free, exploring the generalization abilities of pretrained MLVGMs on the adversarial purification down-stream task. Despite the lack of large models, trained on billions of samples, we show that smaller MLVGMs are already competitive with traditional methods, and can be used as foundation models. Official code released at https://github.com/SerezD/gen_adversarial.
Dario Serez, Marco Cristani, Alessio Del Bue, Vittorio Murino, Pietro Morerio
WACV2
2024 Leveraging Latent Diffusion Models for Training-Free in-Distribution Data Augmentation for Surface Defect Detection
abstract
Defect detection is the task of identifying defects in production samples. Usually, defect detection classifiers are trained on ground-truth data formed by normal samples (negative data) and samples with defects (positive data), where the latter are consistently fewer than normal samples. State-of-the-art data augmentation procedures add synthetic defect data by superimposing artifacts to normal samples to mitigate problems related to unbalanced training data. These techniques often produce out-of-distribution images, resulting in systems that learn what is not a normal sample but cannot accurately identify what a defect looks like. In this work, we introduce DIAG, a training-free Diffusion-based In-distribution Anomaly Generation pipeline for data augmentation. Unlike conventional image generation techniques, we implement a human-in-the-loop pipeline, where domain experts provide multimodal guidance to the model through text descriptions and region localization of the possible anomalies. This strategic shift enhances the interpretability of results and fosters a more robust human feedback loop, facilitating iterative improvements of the generated outputs. Remarkably, our approach operates in a zero-shot manner, avoiding time-consuming fine-tuning procedures while achieving superior performance. We demonstrate the efficacy and versatility of DIAG with respect to state-of-the-art data augmentation approaches on the challenging KSDD2 dataset, with an improvement in AP of approximately 18 % when positive samples are available and 28 % when they are missing. The source code is available at https://github.com/intelligolabs/DIAG.
Federico Girella, Franco Fummi, Francesco Setti, Marco Cristani, Luigi Capogrosso
CBMI5
2024 MTL-Split: Multi-Task Learning for Edge Devices using Split Computing
abstract
Split Computing (SC), where a Deep Neural Network (DNN) is intelligently split with a part of it deployed on an edge device and the rest on a remote server is emerging as a promising approach. It allows the power of DNNs to be leveraged for latency-sensitive applications that do not allow the entire DNN to be deployed remotely, while not having sufficient computation bandwidth available locally. In many such embedded systems scenarios, such as those in the automotive domain, computational resource constraints also necessitate Multi-Task Learning (MTL), where the same DNN is used for multiple inference tasks instead of having dedicated DNNs for each task, which would need more computing bandwidth. However, how to partition such a multi-tasking DNN to be deployed within a SC framework has not been sufficiently studied. This paper studies this problem, and MTL-Split, our novel proposed architecture, shows encouraging results on both synthetic and real-world data. The source code is available at https://github.com/intelligolabs/MTL-Split.
Luigi Capogrosso, Enrico Fraccaroli, Samarjit Chakraborty, Franco Fummi, Marco Cristani
DAC5
2024 Enhancing Split Computing and Early Exit Applications through Predefined Sparsity
abstract
In the past decade, Deep Neural Networks (DNNs) achieved state-of-the-art performance in a broad range of problems, spanning from object classification and action recognition to smart building and healthcare. The flexibility that makes DNNs such a pervasive technology comes at a price: the computational requirements preclude their deployment on most of the resource-constrained edge devices available today to solve real-time and real-world tasks. This paper introduces a novel approach to address this challenge by combining the concept of predefined sparsity with Split Computing (SC) and Early Exit (EE). In particular, SC aims at splitting a DNN with a part of it deployed on an edge device and the rest on a remote server. Instead, EE allows the system to stop using the remote server and rely solely on the edge device’s computation if the answer is already good enough. Specifically, how to apply such a predefined sparsity to a SC and EE paradigm has never been studied. This paper studies this problem and shows how predefined sparsity significantly reduces the computational, storage, and energy burdens during the training and inference phases, regardless of the hardware platform. This makes it a valuable approach for enhancing the performance of SC and EE applications. Experimental results showcase reductions exceeding 4× in storage and computational complexity without compromising performance. The source code is available at https://github.com/intelligolabs/sparsity_sc_ee.
Luigi Capogrosso, Enrico Fraccaroli, Giulio Petrozziello, Francesco Setti, Samarjit Chakraborty, Franco Fummi, Marco Cristani
FDL7
2024 Dif4FF: Leveraging Multimodal Diffusion Models and Graph Neural Networks for Accurate New Fashion Product Performance Forecasting
Andrea Avogaro, Luigi Capogrosso, Franco Fummi, Marco Cristani
ICPR (8)4
2024 SITUATE: Indoor Human Trajectory Prediction Through Geometric Features and Self-supervised Vision Representation
Luigi Capogrosso, Andrea Toaiari, Andrea Avogaro, Aditya Jivoji, Franco Fummi, Marco Cristani
ICPR (16)7
2024 Forecast the Forecasting
Marco Cristani
ICPRAM1
2024 Exploring 3D Human Pose Estimation and Forecasting from the Robot's Perspective: The HARPER Dataset
abstract
We introduce HARPER, a novel dataset for 3D body pose estimation and forecasting in dyadic interactions between users and Spot, the quadruped robot manufactured by Boston Dynamics. The key-novelty of HARPER is its focus on the robot’s perspective, i.e., on the data captured by the robot’s sensors. This makes 3D body pose analysis challenging, as being close to the ground results in only partial captures of humans. The scenario underlying HARPER includes 15 actions, of which 10 involve physical contact between the robot and users. The corpus contains recordings not only from Spot’s built-in stereo cameras but also from a 6-camera OptiTrack system, with all recordings synchronized. This setup leads to ground-truth skeletal representations with a precision of less than a millimeter. Additionally, the corpus includes reproducible benchmarks for 3D Human Pose Estimation, Human Pose Forecasting, and Collision Prediction, all based on publicly available baseline approaches. This enables future HARPER users to rigorously compare their results with those provided in this work.
Andrea Avogaro, Andrea Toaiari, Federico Cunico, Xiangmin Xu 0003, Haralambos Dafas, Alessandro Vinciarelli, Liying Li 0001, Marco Cristani
IROS8
2024 Mind the Error! Detection and Localization of Instruction Errors in Vision-and-Language Navigation
abstract
Vision-and-Language Navigation in Continuous Environments (VLN-CE) is one of the most intuitive yet challenging embodied AI tasks. Agents are tasked to navigate towards a target goal by executing a set of low-level actions, following a series of natural language instructions. All VLN-CE methods in the literature assume that language instructions are exact. However, in practice, instructions given by humans can contain errors when describing a spatial environment due to inaccurate memory or confusion. Current VLN-CE benchmarks do not address this scenario, making the state-of-the-art methods in VLN-CE fragile in the presence of erroneous instructions from human users. For the first time, we propose a novel benchmark dataset that introduces various types of instruction errors considering potential human causes. This benchmark provides valuable insight into the robustness of VLN systems in continuous environments. We observe a noticeable performance drop (up to −25%) in Success Rate when evaluating the state-of-the-art VLN-CE methods on our benchmark. Moreover, we formally define the task of Instruction Error Detection and Localization, and establish an evaluation protocol on top of our benchmark dataset. We also propose an effective method, based on a cross-modal transformer architecture, that achieves the best performance in error detection and localization, compared to baselines. Surprisingly, our proposed method has revealed errors in the validation set of the two commonly used datasets for VLN-CE, i.e., R2R-CE and RxR-CE, demonstrating the utility of our technique in other tasks.
Francesco Taioli, Stefano Rosa, Alberto Castellini, Lorenzo Natale, Alessio Del Bue, Alessandro Farinelli, Marco Cristani, Yiming Wang 0002
IROS7
2024 I2EDL: Interactive Instruction Error Detection and Localization
abstract
In the Vision-and-Language Navigation in Continuous Environments (VLN-CE) task, the human user guides an autonomous agent to reach a target goal via a series of low-level actions following a textual instruction in natural language. However, most existing methods do not address the likely case where users may make mistakes when providing such instruction (e.g., "turn left" instead of "turn right"). In this work, we address a novel task of Interactive VLN in Continuous Environments (IVLN-CE), which allows the agent to interact with the user during the VLN-CE navigation to verify any doubts regarding the instruction errors. We propose an Interactive Instruction Error Detector and Localizer (I2EDL) that triggers the user-agent interaction upon the detection of instruction errors during the navigation. We leverage a pre-trained module to detect instruction errors and pinpoint them in the instruction by cross-referencing the textual input and past observations. In such way, the agent is able to query the user for a timely correction, without demanding the user's cognitive load, as we locate the probable errors to a precise part of the instruction. We evaluate the proposed I2EDL on a dataset of instructions containing errors, and further devise a novel metric, the Success weighted by Interaction Number (SIN), to reflect both the navigation performance and the interaction effectiveness. We show how the proposed method can ask focused requests for corrections to the user, which in turn increases the navigation success, while minimizing the interactions.
Francesco Taioli, Stefano Rosa, Alberto Castellini, Lorenzo Natale, Alessio Del Bue, Alessandro Farinelli, Marco Cristani, Yiming Wang 0002
RO-MAN7
2024 Unsupervised Active Visual Search With Monte Carlo Planning Under Uncertain Detections
abstract
We propose a solution for Active Visual Search of objects in an environment, whose 2D floor map is the only known information. Our solution has three key features that make it more plausible and robust to detector failures compared to state-of-the-art methods: i) it is unsupervised as it does not need any training sessions. ii) During the exploration, a probability distribution on the 2D floor map is updated according to an intuitive mechanism, while an improved belief update increases the effectiveness of the agent's exploration. iii) We incorporate the awareness that an object detector may fail into the aforementioned probability modelling by exploiting the success statistics of a specific detector. Our solution is dubbed POMP-BE-PD (Pomcp-based Online Motion Planning with Belief by Exploration and Probabilistic Detection). It uses the current pose of an agent and an RGB-D observation to learn an optimal search policy, exploiting a POMDP solved by a Monte-Carlo planning approach. On the Active Vision Dataset Benchmark, we increase the average success rate over all the environments by a significant 35 % while decreasing the average path length by 4 % with respect to competing methods. Thus, our results are state-of-the-art, even without any training procedure.
Francesco Taioli, Francesco Giuliari, Yiming Wang 0002, Riccardo Berra, Alberto Castellini, Alessio Del Bue, Alessandro Farinelli, Marco Cristani, Francesco Setti
IEEE Trans. Pattern Anal. Mach. Intell.8
2023 The Post-pandemic Effects on IoT for Safety: The Safe Place Project
abstract
COVID-19 had substantial effects on the IoT community which designs systems for safety: the urge to face masks worn by everyone, the analysis of crowds to avoid the spread of the disease, and the sanitization of public environments has led to exceptional research acceleration and fast engineering of the related solutions. Now that the pandemic is losing power, some applications are becoming less important, while others are proving to be useful regardless of the criticality of COVID-19. The Safe Place project is a prime example of this situation (DATE23 MPP category: final stage). Safe Place is an Italian 3M euro regional industrial/academic project, financed by European funds, created to ensure a multidisciplinary choral reaction to COVID-19 in critical environments such as rest homes and public places. Safe Place consortium was able to understand what is no longer useful in this post-pandemic period, and what instead is potentially attractive for the market. For example, the detection of face masks has little importance, while sanitization does have much. This paper shares such analysis, which emerged through a co-design process of three public Safe Place project demonstrators, involving heterogeneous figures spanning from scientists to lawyers.
Federico Cunico, Luigi Capogrosso, Alberto Castellini, Francesco Setti, Patrik Pluchino, Filippo Zordan, Valeria Santus, Anna Spagnolli, Stefano Cordibella, Giambattista Gennari, Mauro Borgo, Alberto Sozza, Stefano Troiano, Roberto Flor, Andrea Zanella, Alessandro Farinelli, Luciano Gamberini, Marco Cristani
DATE18
2023 Towards Deep Learning-based Occupancy Detection Via WiFi Sensing in Unconstrained Environments
abstract
In the context of smart buildings and smart cities, the design of low-cost and privacy-aware solutions for recognizing the presence of humans and their activities is becoming of great interest. Existing solutions exploiting wearables and video-based systems have several drawbacks, such as high cost, low usability, poor portability, and privacy-related issues. Consequently, more ubiquitous and accessible solutions, such as WiFi sensing, became the focus of attention. However, at the current state-of-the-art, WiFi sensing is subject to low accuracy and poor generalization, primarily affected by environmental factors, such as humidity and temperature variations, and furniture position changes. Such is-sues are partially solved at the cost of complex data preprocessing pipelines. In this paper, we present a highly accurate, resource-efficient deep learning-based occupancy detection solution, which is resilient to variations in humidity and temperature. The approach is tested on an extensive benchmark, where people are free to move and the furniture layout does change. In addition, based on a consolidated algorithm of explainable AI, we quantify the importance of the WiFi signal w.r.t. humidity and temperature for the proposed approach. Notably, humidity and temperature can indeed be predicted based on WiFi signals; this promotes the expressivity of the WiFi signal and at the same time the need for a non-linear model to properly deal with it.
Cristian Turetta, Geri Skenderi, Luigi Capogrosso, Florenc Demrozi, Philipp H. Kindt, Alejandro Masrur, Franco Fummi, Marco Cristani, Graziano Pravadelli
DATE8
2023 Split-Et-Impera: A Framework for the Design of Distributed Deep Learning Applications
abstract
Many recent pattern recognition applications rely on complex distributed architectures in which sensing and computational nodes interact together through a communication network. Deep neural networks (DNNs) play an important role in this scenario, furnishing powerful decision mechanisms, at the price of a high computational effort. Consequently, powerful state-of-the-art DNNs are frequently split over various computational nodes, e.g., a first part stays on an embedded device and the rest on a server. Deciding where to split a DNN is a challenge in itself, making the design of deep learning applications even more complicated. Therefore, we propose Split-Et-Impera, a novel and practical framework that i) determines the set of the best-split points of a neural network based on deep network interpretability principles without performing a tedious try-and-test approach, ii) performs a communication-aware simulation for the rapid evaluation of different neural network rearrangements, and iii) suggests the best match between the quality of service requirements of the application and the performance in terms of accuracy and latency time.
Luigi Capogrosso, Federico Cunico, Michele Lora, Marco Cristani, Franco Fummi, Davide Quaglia
DDECS4
2023 HermesBDD: A Multi-Core and Multi-Platform Binary Decision Diagram Package
abstract
BDDs are representations of a Boolean expression in the form of a directed acyclic graph. BDDs are widely used in several fields, particularly in model checking and hardware verification. There are several implementations for BDD manipulation, where each package differs depending on the application. This paper presents HermesBDD: a novel multi-core and multi-platform binary decision diagram package focused on high performance and usability. HermesBDD supports a static and dynamic memory management mechanism, the possibility to exploit lock-free hash tables, and a simple parallel implementation of the IF-THEN-ELSE procedure based on a higher-level wrapper for threads and futures. HermesBDD is completely written in C++ with no need to rely on external libraries and is developed according to software engineering principles for reliability and easy maintenance over time. We provide experimental results on the n-Queens problem, the de-facto SAT solver benchmark for BDDs, demonstrating a significant speedup of 18.73× over our non-parallel baselines, and a remarkable performance boost w.r.t. other state-of-the-art BDDs packages.
Luigi Capogrosso, Luca Geretti, Marco Cristani, Franco Fummi, Tiziano Villa
DDECS3
2023 Neuro-Symbolic Empowered Denoising Diffusion Probabilistic Models for Real-Time Anomaly Detection in Industry 4.0: Wild-and-Crazy-Idea Paper
abstract
Industry 4.0 involves the integration of digital technologies, such as IoT, Big Data, and AI, into manufacturing and industrial processes to increase efficiency and productivity. As these technologies become more interconnected and interdependent, Industry 4.0 systems become more complex, which brings the difficulty of identifying and stopping anomalies that may cause disturbances in the manufacturing process. This paper aims to propose a diffusion-based model for real-time anomaly prediction in Industry 4.0 processes. Using a neuro-symbolic approach, we integrate industrial ontologies in the model, thereby adding formal knowledge on smart manufacturing. Finally, we propose a simple yet effective way of distilling diffusion models through Random Fourier Features for deployment on an embedded system for direct integration into the manufacturing process. To the best of our knowledge, this approach has never been explored before.
Luigi Capogrosso, Alessio Mascolini, Federico Girella, Geri Skenderi, Sebastiano Gaiardelli, Nicola Dall'Ora, Francesco Ponzio, Enrico Fraccaroli, Santa Di Cataldo, Sara Vinco, Enrico Macii, Franco Fummi, Marco Cristani
FDL13
2023 Leveraging Commonsense for Object Localisation in Partial Scenes
abstract
We propose an end-to-end solution to address the problem of object localisation in partial scenes, where we aim to estimate the position of an object in an unknown area given only a partial 3D scan of the scene. We propose a novel scene representation to facilitate the geometric reasoning, Directed Spatial Commonsense Graph (D-SCG), a spatial scene graph that is enriched with additional concept nodes from a commonsense knowledge base. Specifically, the nodes of D-SCG represent the scene objects and the edges are their relative positions. Each object node is then connected via different commonsense relationships to a set of concept nodes. With the proposed graph-based scene representation, we estimate the unknown position of the target object using a Graph Neural Network that implements a sparse attentional message passing mechanism. The network first predicts the relative positions between the target object and each visible object by learning a rich representation of the objects via aggregating both the object nodes and the concept nodes in D-SCG. These relative positions then are merged to obtain the final position. We evaluate our method using Partial ScanNet, improving the state-of-the-art by 5.9% in terms of the localisation accuracy at a 8x faster training speed.
Francesco Giuliari, Geri Skenderi, Marco Cristani, Alessio Del Bue, Yiming Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Under the hood of transformer networks for trajectory forecasting
abstract
Transformer Networks have established themselves as the de-facto state-of-the-art for trajectory forecasting but there is currently no systematic study on their capability to model the motion patterns of people, without interactions with other individuals nor the social context. There is abundant literature on LSTMs, CNNs and GANs on this subject. However methods adopting Transformer techniques achieve great performances by complex models and a clear analysis of their adoption as plain sequence models is missing. This paper proposes the first in-depth study of Transformer Networks (TF) and the Bidirectional Transformers (BERT) for the forecasting of the individual motion of people, without bells and whistles. We conduct an exhaustive evaluation of the input/output representations, problem formulations and sequence modelling, including a novel analysis of their capability to predict multi-modal futures. Out of comparative evaluation on the ETH+UCY benchmark, both TF and BERT are top performers in predicting individual motions and remain within a narrow margin wrt more complex techniques, including both social interactions and scene contexts. Source code will be released for all conducted experiments.
Luca Franco, Leonardo Placidi, Francesco Giuliari, Irtiza Hasan, Marco Cristani, Fabio Galasso
Pattern Recognit.5
2022 Spatial Commonsense Graph for Object Localisation in Partial Scenes
abstract
We solve object localisation in partial scenes, a new problem of estimating the unknown position of an object (e.g. where is the bag?) given a partial 3D scan of a scene. The proposed solution is based on a novel scene graph model, the Spatial Commonsense Graph (SCG), where objects are the nodes and edges define pairwise distances between them, enriched by concept nodes and relationships from a commonsense knowledge base. This allows SCG to better generalise its spatial inference over unknown 3D scenes. The SCG is used to estimate the unknown position of the target object in two steps: first, we feed the SCG into a novel Proximity Prediction Network, a graph neural network that uses attention to perform distance prediction between the node representing the target object and the nodes representing the observed objects in the SCG; second, we propose a Localisation Module based on circular intersection to estimate the object position using all the predicted pairwise distances in order to be independent of any reference system. We create a new dataset of partially reconstructed scenes to benchmark our method and baselines for object localisation in partial scenes, where our proposed method achieves the best localisation performance.
Francesco Giuliari, Geri Skenderi, Marco Cristani, Yiming Wang 0002, Alessio Del Bue
CVPR3
2022 POP: Mining POtential Performance of New Fashion Products via Webly Cross-modal Query Expansion
Christian Joppi, Geri Skenderi, Marco Cristani
ECCV (38)3
2022 Pose Forecasting in Industrial Human-Robot Collaboration
Alessio Sampieri, Guido Maria D'Amely di Melendugno, Andrea Avogaro, Federico Cunico, Francesco Setti, Geri Skenderi, Marco Cristani, Fabio Galasso
ECCV (38)7
2022 I-SPLIT: Deep Network Interpretability for Split Computing
abstract
This work makes a substantial step in the field of split computing, i.e., how to split a deep neural network to host its early part on an embedded device and the rest on a server. So far, potential split locations have been identified exploiting uniquely architectural aspects, i.e., based on the layer sizes. Under this paradigm, the efficacy of the split in terms of accuracy can be evaluated only after having performed the split and retrained the entire pipeline, making an exhaustive evaluation of all the plausible splitting points prohibitive in terms of time. Here we show that not only the architecture of the layers does matter, but the importance of the neurons contained therein too. A neuron is important if its gradient with respect to the correct class decision is high. It follows that a split should be applied right after a layer with a high density of important neurons, in order to preserve the information flowing until then. Upon this idea, we propose Interpretable Split (I-SPLIT): a procedure that identifies the most suitable splitting points by providing a reliable prediction on how well this split will perform in terms of classification accuracy, beforehand of its effective implementation. As a further major contribution of I-SPLIT, we show that the best choice for the splitting point on a multiclass categorization problem depends also on which specific classes the network has to deal with. Exhaustive experiments have been carried out on two networks, VGG16 and ResNet-50, and three datasets, Tiny-Imagenet-200, notMNIST, and Chest X-Ray Pneumonia. The source code is available at https://github.com/vips4/I-Split.
Federico Cunico, Luigi Capogrosso, Francesco Setti, Damiano Carra, Franco Fummi, Marco Cristani
ICPR6
2022 MovingFashion: a Benchmark for the Video-to-Shop Challenge
abstract
Retrieving clothes that are worn in social media videos (Instagram, TikTok) is the latest frontier of e-fashion, referred to as "video-to-shop" in the computer vision literature. In this paper, we present MovingFashion, the first publicly available dataset to cope with this challenge. MovingFashion is composed of 14855 social videos, each one of them associated with e-commerce "shop" images where the corresponding clothing items are clearly portrayed. In addition, we present a novel baseline for this scenario, dubbed SEAM Match-RCNN. The model is trained by image-tovideo domain adaptation, allowing the use of video sequences where only their association with a shop image is given, eliminating the need for millions of annotated bounding boxes. SEAM Match-RCNN builds an embedding, where an attention-based weighted sum of few frames (10) of a social video is enough to individuate the correct product within the first 5 retrieved items in a 14K+ shop element gallery with an accuracy of 80%. This provides the best performance on MovingFashion, comparing exhaustively against the related state-of-the-art approaches and alternative baselines1.
Marco Godi, Christian Joppi, Geri Skenderi, Marco Cristani
WACV4
2022 SHREC 2022 track on online detection of heterogeneous gestures
Marco Emporio, Ariel Caputo, Andrea Giachetti 0001, Marco Cristani, Guido Borghi, Andrea D'Eusanio, Minh-Quan Le, Hai-Dang Nguyen, Minh-Triet Tran, Felix Ambellan, Martin Hanik, Esfandiar Nava-Yazdani, Christoph von Tycowicz
Comput. Graph.4
2022 Special Issue on Conformal and Probabilistic Prediction with Applications: Preface
Alex Gammerman, Vladimir Vovk, Marco Cristani
Pattern Recognit.3
2021 DOHMO: Embedded Computer Vision in Co-Housing Scenarios
abstract
This paper presents DOHMO, an embedded computer vision system where multiple sensors, including intelligent cameras, are connected to actuators that regulate illumination and doors. The system aims at assisting elderly and impaired people in co-housing scenarios, in accordance with privacy design principles. The paper provides details of two core elements of the system: The first one is the BOX-IO controller, a fully scalable and customizable hardware and software IoT ecosystem that can collect, control, and monitor data, operational flows and business scenarios, whether indoor or outdoor. The second one is the embedded 3DEverywhere intelligent camera, a device composed of an embedded system that receives input data provided by a 3D/2D camera, analyzes it, and returns the metadata of this analysis. We illustrate how they can be connected and how simple decision mechanisms can be implemented in such a framework. In particular, illumination can be triggered on and off by the detected presence of people, overcoming the limitations of typical sensors, while doors can be opened or closed based on person trajectories in an intelligent manner. To substantiate the proposed system, numerous experiments are performed in a lab and a co-housina scenario.
Geri Skenderi, Alessia Bozzini, Luigi Capogrosso, Enrico Carlo Agrillo, Giovanni Perbellini, Franco Fummi, Marco Cristani
FDL7
2021 POMP++: Pomcp-based Active Visual Search in unknown indoor environments
abstract
In this paper, we focus on the problem of learning online an optimal policy for Active Visual Search (AVS) of objects in unknown indoor environments. We propose POMP++, a planning strategy that introduces a novel formulation on top of the classic Partially Observable Monte Carlo Planning (POMCP) framework, to allow training-free online policy learning in unknown environments. We present a new belief reinvigoration strategy that enables the use of POMCP with a dynamically growing state space to address the online generation of the floor map. We evaluate our method on two public benchmark datasets, AVD that is acquired by real robotic platforms and Habitat ObjectNav that is rendered from real 3D scene scans, achieving the best success rate with an improvement of >10% over the state-of-the-art methods.
Francesco Giuliari, Alberto Castellini, Riccardo Berra, Alessio Del Bue, Alessandro Farinelli, Marco Cristani, Francesco Setti, Yiming Wang 0002
IROS6
2021 Forecasting People Trajectories and Head Poses by Jointly Reasoning on Tracklets and Vislets
abstract
In this article, we explore the correlation between people trajectories and their head orientations. We argue that people trajectory and head pose forecasting can be modelled as a joint problem. Recent approaches on trajectory forecasting leverage short-term trajectories (aka tracklets) of pedestrians to predict their future paths. In addition, sociological cues, such as expected destination or pedestrian interaction, are often combined with tracklets. In this article, we propose MiXing-LSTM (MX-LSTM) to capture the interplay between positions and head orientations (vislets) thanks to a joint unconstrained optimization of full covariance matrices during the LSTM backpropagation. We additionally exploit the head orientations as a proxy for the visual attention, when modeling social interactions. MX-LSTM predicts future pedestrians location and head pose, increasing the standard capabilities of the current approaches on long-term trajectory forecasting. Compared to the state-of-the-art, our approach shows better performances on an extensive set of public benchmarks. MX-LSTM is particularly effective when people move slowly, i.e., the most challenging scenario for all other models. The proposed approach also allows for accurate predictions on a longer time horizon.
Irtiza Hasan, Francesco Setti, Theodore Tsesmelis, Vasileios Belagiannis, Sikandar Amin, Alessio Del Bue, Marco Cristani, Fabio Galasso
IEEE Trans. Pattern Anal. Mach. Intell.7
2021 Infinite Feature Selection: A Graph-based Feature Filtering Approach
abstract
We propose a filtering feature selection framework that considers subsets of features as paths in a graph, where a node is a feature and an edge indicates pairwise (customizable) relations among features, dealing with relevance and redundancy principles. By two different interpretations (exploiting properties of power series of matrices and relying on Markov chains fundamentals) we can evaluate the values of paths (i.e., feature subsets) of arbitrary lengths, eventually go to infinite, from which we dub our framework Infinite Feature Selection (Inf-FS). Going to infinite allows to constrain the computational complexity of the selection process, and to rank the features in an elegant way, that is, considering the value of any path (subset) containing a particular feature. We also propose a simple unsupervised strategy to cut the ranking, so providing the subset of features to keep. In the experiments, we analyze diverse settings with heterogeneous features, for a total of 11 benchmarks, comparing against 18 widely-known comparative approaches. The results show that Inf-FS behaves better in almost any situation, that is, when the number of features to keep are fixed a priori, or when the decision of the subset cardinality is part of the process.
Giorgio Roffo, Simone Melzi, Umberto Castellani, Alessandro Vinciarelli, Marco Cristani
IEEE Trans. Pattern Anal. Mach. Intell.5
2020 POMP: Pomcp-based Online Motion Planning for active visual search in indoor environments
Yiming Wang 0002, Francesco Giuliari, Riccardo Berra, Alberto Castellini, Alessio Del Bue, Alessandro Farinelli, Marco Cristani, Francesco Setti
BMVC7
2020 Leveraging Acoustic Images for Effective Self-supervised Audio Representation Learning
Valentina Sanguineti, Pietro Morerio, Niccolò Pozzetti, Danilo Greco, Marco Cristani, Vittorio Murino
ECCV (22)5
2020 Transformer Networks for Trajectory Forecasting
abstract
Most recent successes on forecasting the people motion are based on LSTM models and all most recent progress has been achieved by modelling the social interaction among people and the people interaction with the scene. We question the use of the LSTM models and propose the novel use of Transformer Networks for trajectory forecasting. This is a fundamental switch from the sequential step-by-step processing of LSTMs to the only-attention-based memory mechanisms of Transformers. In particular, we consider both the original Transformer Network (TF) and the larger Bidirectional Transformer (BERT), state-of-the-art on all natural language processing tasks. Our proposed Transformers predict the trajectories of the individual people in the scene. These are “simple” models because each person is modelled separately without any complex human-human nor scene interaction terms. In particular, the TF modelwithoutbellsandwhistlesyields the best score on the largest and most challenging trajectory forecasting benchmark of TrajNet [1]. Additionally, its extension which predicts multiple plausible future trajectories performs on par with more engineered techniques on the 5 datasets of ETH [2]+UCY [3]. Finally, we show that Transformers may deal with missing observations, as it may be the case with real sensor data. Code is available at github.com/FGiuliari/Trajectory-Transformer.
Francesco Giuliari, Irtiza Hasan, Marco Cristani, Fabio Galasso
ICPR3
2020 SIMCO: SIMilarity-based object COunting
abstract
We present SIMCO, a completely agnostic multiclass object counting approach. SIMCO starts by detecting foreground objects through a novel Mask RCNN-based architecture trained beforehand (just once) on a brand-new synthetic 2D shape dataset, InShape; the idea is to highlight every object resembling a primitive 2D shape (circle, square, rectangle, etc.), Each object detected is described by a low-dimensional embedding, obtained from a novel similarity-based head branch; this latter implements a triplet loss, encouraging similar objects (same 2D shape + color and scale) to map close. Subsequently, SIMCO uses this embedding for clustering, so that different “classes” of similar objects can emerge and be counted, making SIMCO the very first multi-class unsupervised counter. The only required assumption is that repeated objects are present in the image. Experiments show that SIMCO provides state-of-the-art scores on counting benchmarks and that it can also help in many challenging image understanding tasks.
Marco Godi, Christian Joppi, Andrea Giachetti 0001, Marco Cristani
ICPR4
2020 Toward a Wearable System for Predicting Freezing of Gait in People Affected by Parkinson's Disease
abstract
Some wearable solutions exploiting on-body acceleration sensors have been proposed to recognize Freezing of Gait (FoG) in people affected by Parkinson Disease (PD). Once a FoG event is detected, these systems generate a sequence of rhythmic stimuli to allow the patient restarting the gait. While these solutions are effective in detecting FoG events, they are unable to predict FoG to prevent its occurrence. This paper fills in the gap by presenting a machine learning-based approach that classifies accelerometer data from PD patients, recognizing a pre-FOG phase to further anticipate FoG occurrence in advance. Gait was monitored by three tri-axial accelerometer sensors worn on the back, hip and ankle. Gait features were then extracted from the accelerometer's raw data through data windowing and non-linear dimensionality reduction. A k-nearest neighbor algorithm (k-NN) was used to classify gait in three classes of events: pre-FoG, no-FoG and FoG. The accuracy of the proposed solution was compared to state-of-the-art approaches. Our study showed that: (i) we achieved performances overcoming the state-of-the-art approaches in terms of FoG detection, (ii) we were able, for the very first time in the literature, to predict FoG by identifying the pre-FoG events with an average sensitivity and specificity of, respectively, 94.1% and 97.1%, and (iii) our algorithm can be executed on resource-constrained devices. Future applications include the implementation on a mobile device, and the administration of rhythmic stimuli by a wearable device to help the patient overcome the FoG.
Florenc Demrozi, Ruggero Angelo Bacchin, Stefano Tamburin, Marco Cristani, Graziano Pravadelli
IEEE J. Biomed. Health Informatics4
2019 Texel-Att: Representing and Classifying Element-Based Textures by Attributes
Marco Godi, Christian Joppi, Andrea Giachetti 0001, Fabio Pellacini, Marco Cristani
BMVC5
2019 Berrick: a low-cost robotic head platform for human-robot interaction
abstract
We propose a low-cost, open-source platform to encourage large-scale study and research on human-robot social interaction. The paramount of social interaction lies on face-to-face dynamics, so we focus on realizing an anthropomorphic robotic head, Berrick, with multimodal sensing and acting capabilities. Taking from the InMoov robot [1], we isolate its head and redesign its electronics, exploiting the implementation of a new board able to combine an autonomous vision system with the already present motorized facial platform, with Wi-Fi and Bluetooth connectivity for multiagent communication. At the present moment, Berrick is capable of running fundamental yet sophisticated tasks of face detection and gazing, which are crucial to trigger and drive situated social exchanges. In particular, we present here a novel social cue embedded into Berrick, dubbed light-based gazing, functional to communicate the internal state of the robot during a dyadic interaction, thus facilitating situated social exchanges. With a 250€ cost, a 3 days time-to-build (with fully 3D printable parts) and a publicly available documentation, Berrick can become a reference for approaching human robot social interaction at a large-scale at the universities as well as at lower education degrees.
Riccardo Berra, Francesco Setti, Marco Cristani
SMC3
2019 Human-Centric Light Sensing and Estimation From RGBD Images: The Invisible Light Switch
abstract
Lighting design in indoor environments is of primary importance for at least two reasons: 1) people should perceive an adequate light; 2) an effective lighting design means consistent energy saving. We present the Invisible Light Switch (ILS) to address both aspects. ILS dynamically adjusts the room illumination level to save energy while maintaining constant the light level perception of the users. So the energy saving is invisible to them. Our proposed ILS leverages a radiosity model to estimate the light level which is perceived by a person within an indoor environment, taking into account the person position and her/his viewing frustum (head pose). ILS may therefore dim those luminaires, which are not seen by the user, resulting in an effective energy saving, especially in large open offices (where light may otherwise be ON everywhere for a single person). To quantify the system performance, we have collected a new dataset where people wear luxmeter devices while working in office rooms. The luxmeters measure the amount of light (in Lux) reaching the people gaze, which we consider a proxy to their illumination level perception. Our initial results are promising: in a room with 8 LED luminaires, the energy consumption in a day may be reduced from 18585 to 6206 watts with ILS (currently needing 1560 watts for operations). While doing so, the drop in perceived lighting decreases by just 200 lux, a value considered negligible when the original illumination level is above 1200 lux, as is normally the case in offices.
Theodore Tsesmelis, Irtiza Hasan, Marco Cristani, Alessio Del Bue, Fabio Galasso
WACV3
2019 RGBD2lux: Dense Light Intensity Estimation With an RGBD Sensor
abstract
Lighting design and modelling or industrial applications like luminaire planning and commissioning rely heavily on time-consuming manual measurements or on physically coherent computational simulations. Regarding the latter, standard approaches are based on CAD modeling simulations and offline rendering, with long processing times and therefore inflexible workflows. Thus, in this paper we propose a computer vision based system to measure lighting with just a single RGBD camera. The proposed method uses both depth data and images from the sensor to provide a dense measure of light intensity in the field of view of the camera. We evaluate our system on novel ground truth data and compare it to state-of-the-art commercial light planning software. Our system provides improved performance, while being completely automated, given that the CAD model is extracted from the depth and the albedo estimated with the support of RGB images. To the best of our knowledge, this is the first automatic framework for the estimation of lighting in general indoor scenarios from RGBD input.
Theodore Tsesmelis, Irtiza Hasan, Marco Cristani, Fabio Galasso, Alessio Del Bue
WACV3
2019 Evaluating the Group Detection Performance: The GRODE Metrics
abstract
The detection of groups of individuals is attracting the attention of many researchers in diverse fields, from automated surveillance to human-computer interaction, with a growing number of approaches published every year. Unexpectedly, the evaluation metrics for this problem are not consolidated, with some measures inherited from the people detection field, other from clustering, other designed specifically for a particular approach, thus lacking in generalization and making the comparisons between different approaches hard to be carried out. Moreover, most of the existent metrics are scarcely expressive, addressing groups as they are atomic entities, ignoring that they may have different cardinalities, and that group detection approaches may fail in capturing the exact number of individuals that compose it. This paper fills this gap presenting the GROup DEtection (GRODE) metrics, which formally define precision and recall on the groups, including the group cardinality as a variable. This gives the possibility to investigate aspects never considered so far, such as the tendency of a method of over- or under-segmenting, or of better dealing with specific group cardinalities. The GRODE metrics have been evaluated first on controlled scenarios, where the differences with alternative metrics are evident. Then, the metrics have been applied to eight approaches of group detection, on eight public datasets, providing a fresh-new panorama of the state-of-the-art, discovering interesting strengths and pitfalls of the recent approaches.
Francesco Setti, Marco Cristani
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 Analyzing Body Fat from Depth Images
abstract
We present a novel framework to directly estimate body fat percentage from depth images of human subjects and to visually evaluate salient points of the body shape related to the fat distribution. For this purpose, we created a novel, publicly available dataset including front and back depth images of a set of subjects with specific features (active young men or professional sportsmen) with associated ground truth fat values estimated with dual-energy x-ray absorptiometry (DXA) scanning. These depth images were obtained with depth rendering of an available dataset of whole body scans, simulating low-cost depth sensor acquisitions. We customized a ResNet-50 regressor to estimate fat percentage values directly from the front/back scans, achieving promising accuracy (standard errors of estimate SEE less than 2.1 on the depth renderings and 2.5 on a small set of real depth scans). We also demonstrate that, using a custom perturbation-based procedure for analyzing deep networks, it is possible to highlight, on subjects' depth images, the specific body areas related to fat accumulation (typically neck, shoulders, hip, and abdomen) and those characterizing skinny subjects (chest and abdomen).
Marco Carletti, Marco Cristani, Valentina Cavedon, Chiara Milanese, Carlo Zancanaro, Andrea Giachetti 0001
3DV2
2018 Understanding Deep Architectures by Visual Summaries
Marco Godi, Marco Carletti, Maya Aghaei, Francesco Giuliari, Marco Cristani
BMVC5
2018 Recognition self-awareness for active object recognition on depth images
Andrea Roberti, Marco Carletti, Francesco Setti, Umberto Castellani, Paolo Fiorini, Marco Cristani
BMVC6
2018 MX-LSTM: Mixing Tracklets and Vislets to Jointly Forecast Trajectories and Head Poses
abstract
Recent approaches on trajectory forecasting use tracklets to predict the future positions of pedestrians exploiting Long Short Term Memory (LSTM) architectures. This paper shows that adding vislets, that is, short sequences of head pose estimations, allows to increase significantly the trajectory forecasting performance. We then propose to use vislets in a novel framework called MX-LSTM, capturing the interplay between tracklets and vislets thanks to a joint unconstrained optimization of full covariance matrices during the LSTM backpropagation. At the same time, MX-LSTM predicts the future head poses, increasing the standard capabilities of the long-term trajectory forecasting approaches. With standard head pose estimators and an attentional-based social pooling, MX-LSTM scores the new trajectory forecasting state-of-the-art in all the considered datasets (Zara01, Zara02, UCY, and TownCentre) with a dramatic margin when the pedestrians slow down, a case where most of the forecasting approaches struggle to provide an accurate solution.
Irtiza Hasan, Francesco Setti, Theodore Tsesmelis, Alessio Del Bue, Fabio Galasso, Marco Cristani
CVPR6
2018 An Energy Saving Approach to Active Object Recognition and Localization
abstract
We propose an Active Object Recognition (AOR) strategy explicitly suited to work with robotic arms in human-robot cooperation scenarios. So far, AOR policies on robotic arms have focused on heterogeneous constraints, most of them related to classification accuracy, classification confidence, number of moves etc., discarding physical and energetic constraints a real robot has to fulfill. Our strategy overcomes this weakness by exploiting a POMDP-based AOR algorithm that explicitly considers manipulability and energetic terms in the planning optimization. The manipulability term avoids the robotic arm to get close to singularities, which require expensive and straining backtracking steps; the energetic term deals with the arm gravity compensation when in static conditions, which is crucial in AOR policies where time is spent in the classifier belief update, before doing the next movement. Several experiments have been carried out on a redundant, 7-DoF Panda arm manipulator, on a multi-object recognition task. This allows to appreciate the improvement of our solution with respect to other competitors evaluated on simulations only.
Andrea Roberti, Riccardo Muradore, Paolo Fiorini, Marco Cristani, Francesco Setti
IECON4
2018 "Seeing is Believing": Pedestrian Trajectory Forecasting Using Visual Frustum of Attention
abstract
In this paper we show the importance of the head pose estimation in the task of trajectory forecasting. This cue, when produced by an oracle and injected in a novel socially-based energy minimization approach, allows to get state-of-the-art performances on four different forecasting benchmarks, without relying on additional information such as expected destination and desired speed, which are supposed to be know beforehand for most of the current forecasting techniques. Our approach uses the head pose estimation for two aims: 1) to define a view frustum of attention, highlighting the people a given subject is more interested about, in order to avoid collisions; 2) to give a shorttime estimation of what would be the desired destination point. Moreover, we show that when the head pose estimation is given by a real detector, though the performance decreases, it still remains at the level of the top score forecasting systems.
Irtiza Hasan, Francesco Setti, Theodore Tsesmelis, Alessio Del Bue, Marco Cristani, Fabio Galasso
WACV5
2018 Looking beyond appearances: Synthetic training data for deep CNNs in re-identification
Igor Barros Barbosa, Marco Cristani, Barbara Caputo, Aleksander Rognhaugen, Theoharis Theoharis
Comput. Vis. Image Underst.2
2018 Count on Me: Learning to Count on a Single Image
abstract
Individuating and locating repetitive patterns in still images is a fundamental task in image processing, typically achieved by means of correlation strategies. In this paper, we provide a solid solution to this task using a differential geometry approach, operating on Lie algebra, and exploiting a mixture of templates. The proposed method asks the user to locate a few instances of the target patterns (seeds) that become visual templates used to explore the image. We propose an iterative algorithm to locate patches similar to the seeds working in three steps: first, clustering the detected patches to generate templates of different classes, then looking for the affine transformations, living on a Lie algebra that best links the templates and the detected patches, and finally detecting new patches with a convolutional strategy. The process ends when no new patches are found. We will show how our method is able to process heterogeneous unstructured images with multiple visual motifs and extremely crowded scenarios with high precision and recall, outperforming all the state-of-the-art methods.
Francesco Setti, Davide Conigliaro, Michele Tobanelli, Marco Cristani
IEEE Trans. Circuits Syst. Video Technol.4
2018 Discrete Time Evolution Process Descriptor for Shape Analysis and Matching
abstract
In shape analysis and matching, it is often important to encode information about the relation between a given point and other points on a shape, namely, its context . To this aim, we propose a theoretically sound and efficient approach for the simulation of a discrete time evolution process that runs through all possible paths between pairs of points on a surface represented as a triangle mesh in the discrete setting. We demonstrate how this construction can be used to efficiently construct a multiscale point descriptor, called the Discrete Time Evolution Process Descriptor , which robustly encodes the structure of neighborhoods of a point across multiple scales. Our work is similar in spirit to the methods based on diffusion geometry, and derived signatures such as the HKS or the WKS, but provides information that is complementary to these descriptors and can be computed without solving an eigenvalue problem. We demonstrate through extensive experimental evaluation that our descriptor can be used to obtain accurate results in shape matching in different scenarios. Our approach outperforms similar methods and is especially robust in the presence of large nonisometric deformations, including missing parts.
Simone Melzi, Maks Ovsjanikov, Giorgio Roffo, Marco Cristani, Umberto Castellani
ACM Trans. Graph.4
2017 Clothing and People - A Social Signal Processing Perspective
abstract
In our society and century, clothing is not anymore used only as a means for body protection. Our paper builds upon the evidence, studied within the social sciences, that clothing brings a clear communicative message in terms of social signals, influencing the impression and behaviour of others towards a person. In fact, clothing correlates with personality traits, both in terms of self-assessment and assessments that unacquainted people give to an individual. The consequences of these facts are important: the influence of clothing on the decision making of individuals has been investigated in the literature, showing that it represents a discriminative factor to differentiate among diverse groups of people. Unfortunately, this has been observed after cumbersome and expensive manual annotations, on very restricted populations, limiting the scope of the resulting claims. With this position paper, we want to sketch the main steps of the very first systematic analysis, driven by social signal processing techniques, of the relationship between clothing and social signals, both sent and perceived. Thanks to human parsing technologies, which exhibit high robustness owing to deep learning architectures, we are now capable to isolate visual patterns characterising a large types of garments. These algorithms will be used to capture statistical relations on a large corpus of evidence to confirm the sociological findings and to go beyond the state of the art.
Maedeh Aghaei, Federico Parezzan, Mariella Dimiccoli, Petia Radeva, Marco Cristani
FG5
2017 Tiny head pose classification by bodily cues
abstract
The head pose is an important cue for computer vision. Traditionally considered in human computer interaction applications, it becomes very hard to model in surveillance scenarios, due to the tiny head size. Additionally, no public dataset contains continuous head pose annotations in open scenery, making the challenge even harder to face. Here we present a framework based on Faster RCNN, which introduces a branch in the network architecture related to the head pose estimation. The key idea is to leverage the presence of the people body to better infer the head pose, through a joint optimization process. Additionally, we enrich the Town Center dataset with head pose labels, promoting further study on this topic. Results on this novel benchmark and ablation studies on other task-specific datasets promote our idea and confirm the importance of the body cues to contextualize the head pose estimation.
Irtiza Hasan, Theodore Tsesmelis, Fabio Galasso, Alessio Del Bue, Marco Cristani
ICIP5
2017 What your Facebook Profile Picture Reveals about your Personality
abstract
People spend considerable effort managing the impressions they give others. Social psychologists have shown that people manage these impressions differently depending upon their personality. Facebook and other social media provide a new forum for this fundamental process; hence, understanding people's behaviour on social media could provide interesting insights on their personality. In this paper we investigate automatic personality recognition from Facebook profile pictures. We analyze the effectiveness of four families of visual features and we discuss some human interpretable patterns that explain the personality traits of the individuals. For example, extroverts and agreeable individuals tend to have warm colored pictures and to exhibit many faces in their portraits, mirroring their inclination to socialize; while neurotic ones have a prevalence of pictures of indoor places. Then, we propose a classification approach to automatically recognize personality traits from these visual features. Finally, we compare the performance of our classification approach to the one obtained by human raters and we show that computer-based classifications are significantly more accurate than averaged human-based classifications for Extraversion and Neuroticism.
Cristina Segalin, Fabio Celli, Luca Polonio, Michal Kosinski, David Stillwell, Nicu Sebe, Marco Cristani, Bruno Lepri
ACM Multimedia7
2017 Social profiling through image understanding: Personality inference using convolutional neural networks
Cristina Segalin, Dong Seon Cheng, Marco Cristani
Comput. Vis. Image Underst.3
2017 The S-Hock dataset: A new benchmark for spectator crowd analysis
Francesco Setti, Davide Conigliaro, Paolo Rota, Chiara Bassetti, Nicola Conci, Nicu Sebe, Marco Cristani
Comput. Vis. Image Underst.7
2017 The Pictures We Like Are Our Image: Continuous Mapping of Favorite Pictures into Self-Assessed and Attributed Personality Traits
abstract
Flickr allows its users to tag the pictures they like as “favorite”. As a result, many users of the popular photo-sharing platform produce galleries of favorite pictures. This article proposes new approaches, based on Computational Aesthetics, capable to infer the personality traits of Flickr users from the galleries above. In particular, the approaches map low-level features extracted from the pictures into numerical scores corresponding to the Big-Five Traits, both self-assessed and attributed. The experiments were performed over 60,000 pictures tagged as favorite by 300 users (the PsychoFlickr Corpus). The results show that it is possible to predict beyond chance both self-assessed and attributed traits. In line with the state-of-the-art of Personality Computing, these latter are predicted with higher effectiveness (correlation up to 0.68 between actual and predicted traits).
Cristina Segalin, Alessandro Perina, Marco Cristani, Alessandro Vinciarelli
IEEE Trans. Affect. Comput.3
2017 Soft Ngram Representation and Modeling for Protein Remote Homology Detection
abstract
Remote homology detection represents a central problem in bioinformatics, where the challenge is to detect functionally related proteins when their sequence similarity is low. Recent solutions employ representations derived from the sequence profile, obtained by replacing each amino acid of the sequence by the corresponding most probable amino acid in the profile. However, the information contained in the profile could be exploited more deeply, provided that there is a representation able to capture and properly model such crucial evolutionary information. In this paper, we propose a novel profile-based representation for sequences, called soft Ngram. This representation, which extends the traditional Ngram scheme (obtained by grouping N consecutive amino acids), permits considering all of the evolutionary information in the profile: this is achieved by extracting Ngrams from the whole profile, equipping them with a weight directly computed from the corresponding evolutionary frequencies. We illustrate two different approaches to model the proposed representation and to derive a feature vector, which can be effectively used for classification using a support vector machine (SVM). A thorough evaluation on three benchmarks demonstrates that the new approach outperforms other Ngram-based methods, and shows very promising results also in comparison with a broader spectrum of techniques.
Pietro Lovato, Marco Cristani, Manuele Bicego
IEEE ACM Trans. Comput. Biol. Bioinform.2
2016 Detecting conversational groups in images and sequences: A robust game-theoretic approach
Sebastiano Vascon, Eyasu Zemene Mequanint, Marco Cristani, Hayley Hung, Marcello Pelillo, Vittorio Murino
Comput. Vis. Image Underst.3
2016 Special issue on "Fine-grained categorization in ecological multimedia"
Concetto Spampinato, Vasileios Mezaris, Marco Cristani
Pattern Recognit. Lett.3
2015 The S-HOCK dataset: Analyzing crowds at the stadium
abstract
The topic of crowd modeling in computer vision usually assumes a single generic typology of crowd, which is very simplistic. In this paper we adopt a taxonomy that is widely accepted in sociology, focusing on a particular category, the spectator crowd, which is formed by people “interested in watching something specific that they came to see” [6]. This can be found at the stadiums, amphitheaters, cinema, etc. In particular, we propose a novel dataset, the Spectators Hockey (S-HOCK), which deals with 4 hockey matches during an international tournament. In the dataset, a massive annotation has been carried out, focusing on the spectators at different levels of details: at a higher level, people have been labeled depending on the team they are supporting and the fact that they know the people close to them; going to the lower levels, standard pose information has been considered (regarding the head, the body) but also fine grained actions such as hands on hips, clapping hands etc. The labeling focused on the game field also, permitting to relate what is going on in the match with the crowd behavior. This brought to more than 100 millions of annotations, useful for standard applications as people counting and head pose estimation but also for novel tasks as spectator categorization. For all of these we provide protocols and baseline results, encouraging further research.
Davide Conigliaro, Paolo Rota, Francesco Setti, Chiara Bassetti, Nicola Conci, Nicu Sebe, Marco Cristani
CVPR7
2015 Infinite Feature Selection
abstract
Filter-based feature selection has become crucial in many classification settings, especially object recognition, recently faced with feature learning strategies that originate thousands of cues. In this paper, we propose a feature selection method exploiting the convergence properties of power series of matrices, and introducing the concept of infinite feature selection (Inf-FS). Considering a selection of features as a path among feature distributions and letting these paths tend to an infinite number permits the investigation of the importance (relevance and redundancy) of a feature when injected into an arbitrary set of cues. Ranking the importance individuates candidate features, which turn out to be effective from a classification point of view, as proved by a thoroughly experimental section. The Inf-FS has been tested on thirteen diverse benchmarks, comparing against filters, embedded methods, and wrappers, in all the cases we achieve top performances, notably on the classification tasks of PASCAL VOC 2007-2012.
Giorgio Roffo, Simone Melzi, Marco Cristani
ICCV3
2015 Semantically-driven automatic creation of training sets for object recognition
Dong Seon Cheng, Francesco Setti, Nicola Zeni, Roberta Ferrario, Marco Cristani
Comput. Vis. Image Underst.5
2015 Non-myopic information theoretic sensor management of a single pan-tilt-zoom camera for multiple object detection and tracking
Pietro Salvagnini, Federico Pernici, Marco Cristani, Giuseppe Lisanti, Alberto Del Bimbo, Vittorio Murino
Comput. Vis. Image Underst.3
2015 Joint Individual-Group Modeling for Tracking
abstract
We present a novel probabilistic framework that jointly models individuals and groups for tracking. Managing groups is challenging, primarily because of their nonlinear dynamics and complex layout which lead to repeated splitting and merging events. The proposed approach assumes a tight relation of mutual support between the modeling of individuals and groups, promoting the idea that groups are better modeled if individuals are considered and vice versa. This concept is translated in a mathematical model using a decentralized particle filtering framework which deals with a joint individual-group state space. The model factorizes the joint space into two dependent subspaces, where individuals and groups share the knowledge of the joint individual-group distribution. The assignment of people to the different groups (and thus group initialization, split and merge) is implemented by two alternative strategies: using classifiers trained beforehand on statistics of group configurations, and through online learning of a Dirichlet process mixture model, assuming that no training data is available before tracking. These strategies lead to two different methods that can be used on top of any person detector (simulated using the ground truth in our experiments). We provide convincing results on two recent challenging tracking benchmarks.
Loris Bazzani, Matteo Zanotto, Marco Cristani, Vittorio Murino
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Recognizing People by Their Personal Aesthetics: A Statistical Multi-level Approach
Cristina Segalin, Alessandro Perina, Marco Cristani
ACCV (3)3
2014 A Game-Theoretic Probabilistic Approach for Detecting Conversational Groups
Sebastiano Vascon, Eyasu Zemene Mequanint, Marco Cristani, Hayley Hung, Marcello Pelillo, Vittorio Murino
ACCV (5)3
2014 Weighted bag of visual words for object recognition
abstract
Bag of Visual words (BoV) is one of the most successful strategy for object recognition, used to represent an image as a vector of counts using a learned vocabulary. This strategy assumes that the representation is built using patches that are either densely extracted or sampled from the images using feature detectors. However, the dense strategy captures also the noisy background information, whereas the feature detection strategy can lose important parts of the objects. In this paper we propose a solution in-between these two strategies, by densely extracting patches from the image, and weighting them accordingly to their salience. Intuitively, highly salient patches have an important role in describing an object, while those with low saliency are still taken with low emphasis, instead of discarding them. We embed this idea in the word encoding mechanism adopted in the BoV approaches. The technique is successfully applied to vector quantization and Fisher vector, on Caltech-101 and Caltech-256.
Marco San-Biagio, Loris Bazzani, Marco Cristani, Vittorio Murino
ICIP3
2014 Biometrics on visual preferences: A "pump and distill" regression approach
abstract
We present a statistical behavioural biometric approach for recognizing people by their aesthetic preferences, using colour images. In the enrollment phase, a model is learnt for each user, using a training set of preferred images. In the recognition/authentication phase, such model is tested with an unseen set of pictures preferred by a probe subject. The approach is dubbed “pump and distill”, since the training set of each user is pumped by bagging, producing a set of image ensembles. In the distill step, each ensemble is reduced into a set of surrogates, that is, aggregates of images sharing a similar visual content. Finally, LASSO regression is performed on these surrogates; the resulting regressor, employed as a classifier, takes test images belonging to a single user, predicting his identity. The approach improves the state-of-the-art on recognition and authentication tasks in average, on a dataset of 40000 Flickr images and 200 users. In practice, given a pool of 20 preferred images of a user, the approach recognizes his identity with an accuracy of 92%, and sets an authentication accuracy of 91% in terms of normalized Area Under the Curve of the CMC and ROC curve, respectively.
Cristina Segalin, Alessandro Perina, Marco Cristani
ICIP3
2014 Statistical Analysis of Personality and Identity in Chats Using a Keylogging Platform
abstract
Interacting via text chats can be considered as a hybrid type of communication, in which textual information delivery follows turn-taking dynamics, resembling spoken interactions. An interesting research question is whether personality can be observed in chats, similarly as happening in face-to-face exchanges. After an encouraging preliminary work on Skype, in this study we have set up our own chat service in which key-logging functionalities have been activated, so that the timings of each key pressing can be measured. Using this framework, we organized semi-structured chats between 50 subjects, whose personality traits have been analyzed through psychometric tests, and a single operator, for a total of 16 hours of conversation. On this data, we have observed that some personality traits are linked with the way we are chatting (measured by stylometric cues), by means of statistically significant correlations and regression studies. Finally, we have assessed that some of the stylometric cues are very discriminative for the recognition of a user in a identification scenario. These facts taken together could underlie that some personality traits drive us in chatting in a particular fashion, which turns out to be very recognizable.
Giorgio Roffo, Cinzia Giorgetta, Roberta Ferrario, Walter Riviera, Marco Cristani
ICMI5
2014 Personal Aesthetics for Soft Biometrics: A Generative Multi-resolution Approach
abstract
Are we recognizable by our image preferences? This paper answers affirmatively the question, presenting a soft biometric approach where the preferred images of an individual are used as his personal signature in identification tasks. The approach builds a multi-resolution latent space, formed by multiple Counting Grids, where similar images are mapped nearby. On this space, a set of preferred images of a user produces an ensemble of intensity maps, highlighting in an intuitive way his personal aesthetic preferences. These maps are then used for learning a battery of discriminative classifiers (one for each resolution), which characterizes the user and serves to perform identification. Results are promising: on a dataset of 200 users, and 40K images, using 20 preferred images as biometric template gives 66% of probability of guessing the correct user. This makes the "personal aesthetics" a very hot topic for soft biometrics, while its usage in standard biometric applications seems to be far from being effective, as we show in a simple user study.
Cristina Segalin, Alessandro Perina, Marco Cristani
ICMI3
2014 Summary Abstract for the 3rd ACM International Workshop on Multimedia Analysis for Ecological Data
abstract
The 3rd ACM International Workshop on Multimedia Anal- ysis for Ecological Data (MAED'14) is held as part of ACM Multimedia 2014.
Concetto Spampinato, Vasileios Mezaris, Marco Cristani
ACM Multimedia3
2014 Information theoretic sensor management for multi-target tracking with a single pan-tilt-zoom camera
abstract
Automatic multiple target tracking with pan-tilt-zoom (PTZ) cameras is a hard task, with few approaches in the literature, most of them proposing simplistic scenarios. In this paper, we present a PTZ camera management framework which lies on information theoretic principles: at each time step, the next camera pose (pan, tilt, focal length) is chosen, according to a policy which ensures maximum information gain. The formulation takes into account occlusions, physical extension of targets, realistic pedestrian detectors and the mechanical constraints of the camera. Convincing comparative results on synthetic data, realistic simulations and the implementation on a real video surveillance camera validate the effectiveness of the proposed method.
Pietro Salvagnini, Federico Pernici, Marco Cristani, Giuseppe Lisanti, Iacopo Masi, Alberto Del Bimbo, Vittorio Murino
WACV3
2014 Encoding Structural Similarity by Cross-covariance Tensors for Image Classification
abstract
In computer vision, an object can be modeled in two main ways: by explicitly measuring its characteristics in terms of feature vectors, and by capturing the relations which link an object with some exemplars, that is, in terms of similarities. In this paper, we propose a new similarity-based descriptor, dubbed structural similarity cross-covariance tensor (SS-CCT), where self-similarities come into play: Here the entity to be measured and the exemplar are regions of the same object, and their similarities are encoded in terms of cross-covariance matrices. These matrices are computed from a set of low-level feature vectors extracted from pairs of regions that cover the entire image. SS-CCT shares some similarities with the widely used covariance matrix descriptor, but extends its power focusing on structural similarities across multiple parts of an image, instead of capturing local similarities in a single region. The effectiveness of SS-CCT is tested on many diverse classification scenarios, considering objects and scenes on widely known benchmarks (Caltech-101, Caltech-256, PASCAL VOC 2007 and SenseCam). In all the cases, the results obtained demonstrate the superiority of our new descriptor against diverse competitors. Furthermore, we also reported an analysis on the reduced computational burden achieved by using and efficient implementation that takes advantage from the integral image representation.
Marco San-Biagio, Samuele Martelli, Marco Crocco, Marco Cristani, Vittorio Murino
Int. J. Pattern Recognit. Artif. Intell.4
2014 Faved! Biometrics: Tell Me Which Image You Like and I'll Tell You Who You Are
abstract
This paper builds upon the belief that every human being has a built-in image aesthetic evaluation system. This sort of personal aesthetics mostly follows certain aesthetic rules widely studied in image aesthetics (e.g., rules of thirds, colorfulness, etc.), though it likely contains some innate, unique preferences. This paper is a proof of concept of this intuition, presenting personal aesthetics as a novel behavioral biometrical trait. In our scenario, personal aesthetics activate when an individual is presented with a set of photos he may like or dislike. The goal is to distill and encode the uniqueness of his visual preferences into a compact template. To this aim, we extract a pool of low- and high-level state-of-the-art image features from a set of Flickr images preferred by a user, feeding them successively into a LASSO regressor. LASSO highlights the most discriminant cues for the individual, allowing authentication and recognition tasks. The results are surprising given only 1 image as test. We can match the user identity against a gallery of 200 individuals definitely much better than chance. Using 20 images (all preferred by a single user) as a biometrical trait, we reach an AUC of 96%, considering the cumulative matching characteristic curve. Extensive experiments also support the interpretability of our approach, effectively modeling what is the “what we like” that distinguishes us from others.
Pietro Lovato, Manuele Bicego, Cristina Segalin, Alessandro Perina, Nicu Sebe, Marco Cristani
IEEE Trans. Inf. Forensics Secur.6
2013 Semi-supervised multi-feature learning for person re-identification
abstract
Person re-identification is probably the open challenge for low-level video surveillance in the presence of a camera network with non-overlapped fields of view. A large number of direct approaches has emerged in the last five years, often proposing novel visual features specifically designed to highlight the most discriminant aspects of people, which are invariant to pose, scale and illumination. On the other hand, learning-based methods are usually based on simpler features, and are trained on pairs of cameras to discriminate between individuals. In this paper, we present a method that joins these two ideas: given an arbitrary state-of-the-art set of features, no matter their number, dimensionality or descriptor, the proposed multi-class learning approach learns how to fuse them, ensuring that the features agree on the classification result. The approach consists of a semi-supervised multi-feature learning strategy, that requires at least a single image per person as training data. To validate our approach, we present results on different datasets, using several heterogeneous features, that set a new level of performance in the person re-identification problem.
Dario Figueira, Loris Bazzani, Hà Quang Minh, Marco Cristani, Alexandre Bernardino, Vittorio Murino
AVSS4
2013 Reading between the turns: Statistical modeling for identity recognition and verification in chats
abstract
Identity safekeeping has recently become an important problem for the social web: as a case study, we focus here on instant messaging platforms, proposing novel soft-biometric cues for user recognition and verification. Specifically, we design a set of features encoding effectively how a person converses: since chats are crossbreeds of written text and face-to-face verbal communication, the features inherit equally from textual authorship attribution and conversational analysis of speech. Importantly, our cues ignore completely the semantics of the chat, relying solely on non-verbal aspects, taking care of possible privacy and ethical issues. We apply our approach on a novel dataset of 94 different individuals, whose chat conversations have been recorded for an average period of five months; recognition rate, intended as normalized AUC on CMC curve, is 95.73%, while verification rate amounts to 95.66%, as normalized AUC on ROC curve.
Giorgio Roffo, Cristina Segalin, Alessandro Vinciarelli, Vittorio Murino, Marco Cristani
AVSS5
2013 SDALF+C: Augmenting the SDALF Descriptor by Relation-Based Information for Multi-shot Re-identification
Sylvie Jasmine Poletti, Vittorio Murino, Marco Cristani
CIARP (2)3
2013 Statistical Analysis of Visual Attentional Patterns for Video Surveillance
Giorgio Roffo, Marco Cristani, Frank E. Pollick, Cristina Segalin, Vittorio Murino
CIARP (2)2
2013 Encoding Classes of Unaligned Objects Using Structural Similarity Cross-Covariance Tensors
Marco San-Biagio, Samuele Martelli, Marco Crocco, Marco Cristani, Vittorio Murino
CIARP (1)4
2013 Heterogeneous Auto-similarities of Characteristics (HASC): Exploiting Relational Information for Classification
abstract
Capturing the essential characteristics of visual objects by considering how their features are inter-related is a recent philosophy of object classification. In this paper, we embed this principle in a novel image descriptor, dubbed Heterogeneous Auto-Similarities of Characteristics (HASC). HASC is applied to heterogeneous dense features maps, encoding linear relations by co variances and nonlinear associations through information-theoretic measures such as mutual information and entropy. In this way, highly complex structural information can be expressed in a compact, scale invariant and robust manner. The effectiveness of HASC is tested on many diverse detection and classification scenarios, considering objects, textures and pedestrians, on widely known benchmarks (Caltech-101, Brodatz, Daimler Multi-Cue). In all the cases, the results obtained with standard classifiers demonstrate the superiority of HASC with respect to the most adopted local feature descriptors nowadays, such as SIFT, HOG, LBP and feature co variances. In addition, HASC sets the state-of-the-art on the Brodatz texture dataset and the Daimler Multi-Cue pedestrian dataset, without exploiting ad-hoc sophisticated classifiers.
Marco San-Biagio, Marco Crocco, Marco Cristani, Samuele Martelli, Vittorio Murino
ICCV3
2013 We like it! Mapping image preferences on the counting grid
abstract
Modeling user preferences in photographic images is often reduced to analyzing intermediate explicit representations (e.g. textual tags) as means of capturing the objective and subjective properties of image perception, trying to distill the essence of what gives pleasure. We propose an alternative approach that bypasses the necessity to build an explicit conceptual coding of image preferences, operating directly on the raw properties of the images, extracted with heterogeneous feature descriptors. This is achieved through the counting grid model, which fuses together content-based and aesthetics themes into a 2D map in an unsupervised way. We show that certain locations in this map correspond to perceptually intuitive image classes, even without relying on tags or other user-defined information. Moreover, we show that users' individual preferences can be represented as distributions over the map, allowing us to evaluate the affinity between different users' appreciations. We experiment on a large Flickr dataset, clustering users by affinity, and validating these clusters by checking users that belong to the same Flickr photo groups.
Pietro Lovato, Alessandro Perina, Dong Seon Cheng, Cristina Segalin, Nicu Sebe, Marco Cristani
ICIP6
2013 Person re-identification with a PTZ camera: An introductory study
abstract
We present an introductory study that paves the way for a new kind of person re-identification, by exploiting a single Pan-Tilt-Zoom (PTZ) camera. PTZ devices allow to zoom on body regions, acquiring discriminative visual patterns that enrich the appearance description of an individual. This intuition has been translated into a statistical direct reidentification scheme, which collects two images for each probe subject: the first image captures the probe individual, focusing on the whole body; the second can be a zoomed body part (head, torso or legs) or another whole body image, and is the outcome of an action-selection mechanism, driven by feature selection principles. The validation of this technique is also explored: in order to allow repeatability, two novel multi-resolution benchmarks have been created. On these data, we demonstrate that our approach selects effective actions, by focusing on body portions which discriminate each subject. Moreover, we show that the proposed compound of two images overwhelms standard multi-shot descriptions, composed by many more pictures.
Pietro Salvagnini, Loris Bazzani, Marco Cristani, Vittorio Murino
ICIP3
2013 Multi-scale f-formation discovery for group detection
abstract
We present an unsupervised approach for the automatic detection of static interactive groups. The approach builds upon a novel multi-scale Hough voting policy, which incorporates in a flexible way the sociological notion of group as F-formation; the goal is to model at the same time small arrangements of close friends and aggregations of many individuals spread over a large area. Our technique is based on a competition of different voting sessions, each one specialized for a particular group cardinality; all the votes are then evaluated using information theoretic criteria, producing the final set of groups. The proposed technique has been applied on public benchmark sequences and a novel cocktail party dataset, evaluating new group detection metrics and obtaining state-of-the-art performances.
Francesco Setti, Oswald Lanz, Roberta Ferrario, Vittorio Murino, Marco Cristani
ICIP5
2013 Unveiling the multimedia unconscious: implicit cognitive processes and multimedia content analysis
abstract
One of the main findings of cognitive sciences is that automatic processes of which we are unaware shape, to a significant extent, our perception of the environment. The phenomenon applies not only to the real world, but also to multimedia data we consume every day. Whenever we look at pictures, watch a video or listen to audio recordings, our conscious attention efforts focus on the observable content, but our cognition spontaneously perceives intentions, beliefs, values, attitudes and other constructs that, while being outside of our conscious awareness, still shape our reactions and behavior. So far, multimedia technologies have neglected such a phenomenon to a large extent. This paper argues that taking into account cognitive effects is possible and it can also improve multimedia approaches. As a supporting proof-of-concept, the paper shows not only that there are visual patterns correlated with the personality traits of 300 Flickr users to a statistically significant extent, but also that the personality traits (both self-assessed and attributed by others) of those users can be inferred from the images these latter post as "favourite".
Marco Cristani, Alessandro Vinciarelli, Cristina Segalin, Alessandro Perina
ACM Multimedia1
2013 Symmetry-driven accumulation of local features for human characterization and re-identification
Loris Bazzani, Marco Cristani, Vittorio Murino
Comput. Vis. Image Underst.2
2013 Social interactions by visual focus of attention in a three-dimensional environment
abstract
Abstract In human behaviour analysis, the visual focus of attention (VFOA) of a person is a very important cue. VFOA detection is difficult, though, especially in a unconstrained and crowded environment, typical of video surveillance scenarios. In this paper, we estimate the VFOA by defining the Subjective View Frustum, which approximates the visual field of a person in a three‐dimensional representation of the scene. This opens up to several intriguing behavioural investigations. In particular, we propose the Inter‐Relation Pattern Matrix, which suggests possible social interactions between the people present in a scene. Theoretical justifications and experimental results substantiate the validity and the goodness of the analysis performed.
Loris Bazzani, Marco Cristani, Diego Tosato, Michela Farenzena, Giulia Paggetti, Gloria Menegaz, Vittorio Murino
Expert Syst. J. Knowl. Eng.2
2013 Human behavior analysis in video surveillance: A Social Signal Processing perspective
Marco Cristani, Ramachandra Raghavendra, Alessio Del Bue, Vittorio Murino
Neurocomputing1
2013 Characterizing Humans on Riemannian Manifolds
abstract
In surveillance applications, head and body orientation of people is of primary importance for assessing many behavioral traits. Unfortunately, in this context people are often encoded by a few, noisy pixels so that their characterization is difficult. We face this issue, proposing a computational framework which is based on an expressive descriptor, the covariance of features. Covariances have been employed for pedestrian detection purposes, actually a binary classification problem on Riemannian manifolds. In this paper, we show how to extend to the multiclassification case, presenting a novel descriptor, named weighted array of covariances, especially suited for dealing with tiny image representations. The extension requires a novel differential geometry approach in which covariances are projected on a unique tangent space where standard machine learning techniques can be applied. In particular, we adopt the Campbell-Baker-Hausdorff expansion as a means to approximate on the tangent space the genuine (geodesic) distances on the manifold in a very efficient way. We test our methodology on multiple benchmark datasets, and also propose new testing sets, getting convincing results in all the cases.
Diego Tosato, Mauro Spera, Marco Cristani, Vittorio Murino
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 Tell Me What You Like and I'll Tell You What You Are: Discriminating Visual Preferences on Flickr Data
Pietro Lovato, Alessandro Perina, Nicu Sebe, Omar Zandonà, Alessio Montagnini, Manuele Bicego, Marco Cristani
ACCV (1)7
2012 Stereo-Based Framework for Pedestrian Detection with Partial Occlusion Handling
abstract
The pedestrian detection literature has been recently extended by the availability of large-scale multisensory datasets, able to capture complementary aspects of the objects of interest, namely, appearance, motion, and depth. In this paper, we exploit this multimodal scenario to propose a new set of composite descriptors dubbed CO2, CO-variances of visual features and CO-occurrences of depth fields. Covariances of visual features allow us to integrate at low-level heterogeneous visual cues related to intensity and texture. Co-occurrences of depth fields are brand new descriptors, which use range information for characterizing the global shape of a pedestrian while being also able to identify its occluded parts. This paper illustrates how these descriptors can be instantiated and combined together, improving detection capabilities taking also benefit from the proper handling of occlusions. Experimental results show that CO2, fed into a standard discriminative classification system, set state-of-the-art performances on recent multi-modal intensity- and stereo-based pedestrian datasets.
Samuele Martelli, Marco Cristani, Vittorio Murino
AVSS2
2012 Online Bayesian Non-parametrics for Social Group Detection
Matteo Zanotto, Loris Bazzani, Marco Cristani, Vittorio Murino
BMVC3
2012 Decentralized particle filter for joint individual-group tracking
abstract
In this paper, we address the task of tracking groups of people in surveillance scenarios. This is a major challenge in computer vision, since groups are structured entities, subjected to repeated split and merge events. Our solution is a joint individual-group tracking framework, inspired by a recent technique dubbed decentralized particle filtering. The proposed strategy factorizes the joint individual-group state space in two dependent subspaces where individuals and groups share the knowledge of the joint individual-group distribution. In practice, we establish a tight relation of mutual support between the modeling of individuals and that of groups, promoting the idea that groups are better tracked if individuals are considered, and viceversa. Extensive experiments on a published and novel dataset validate our intuition, opening up to many future developments.
Loris Bazzani, Marco Cristani, Vittorio Murino
CVPR2
2012 A regularized spectral algorithm for Hidden Markov Models with applications in computer vision
abstract
Hidden Markov Models (HMMs) are among the most important and widely used techniques to deal with sequential or temporal data. Their application in computer vision ranges from action/gesture recognition to videosurveillance through shape analysis. Although HMMs are often embedded in complex frameworks, this paper focuses on theoretical aspects of HMM learning. We propose a regularized algorithm for learning HMMs in the spectral framework, whose computations have no local minima. Compared with recently proposed spectral algorithms for HMMs, our method is guaranteed to produce probability values which are always physically meaningful and which, on synthetic mathematical models, give very good approximations to true probability values. Furthermore, we place no restriction on the number of symbols and the number of states. On various pattern recognition data sets, our algorithm consistently outperforms classical HMMs, both in accuracy and computational speed. This and the fact that HMMs are used in vision as building blocks for more powerful classification approaches, such as generative embedding approaches or more complex generative models, strongly support spectral HMMs (SHMMs) as a new basic tool for pattern recognition.
Hà Quang Minh, Marco Cristani, Alessandro Perina, Vittorio Murino
CVPR2
2012 Learning Discriminative Spatial Relations for Detector Dictionaries: An Application to Pedestrian Detection
Enver Sangineto, Marco Cristani, Alessio Del Bue, Vittorio Murino
ECCV (2)2
2012 Low-level multimodal integration on Riemannian manifolds for automatic pedestrian detection
Marco San-Biagio, Marco Crocco, Marco Cristani, Samuele Martelli, Vittorio Murino
FUSION3
2012 Joining feature-based and similarity-based pattern description paradigms for object detection
Samuele Martelli, Marco Cristani, Loris Bazzani, Diego Tosato, Vittorio Murino
ICPR2
2012 A multiple kernel learning approach to multi-modal pedestrian classification
Marco San-Biagio, Aydin Ulas, Marco Crocco, Marco Cristani, Umberto Castellani, Vittorio Murino
ICPR4
2012 Conversationally-inspired stylometric features for authorship attribution in instant messaging
abstract
Authorship attribution (AA) aims at recognizing automatically the author of a given text sample. Traditionally applied to literary texts, AA faces now the new challenge of recognizing the identity of people involved in chat conversations. These share many aspects with spoken conversations, but AA approaches did not take it into account so far. Hence, this paper tries to fill the gap and proposes two novelties that improve the effectiveness of traditional AA approaches for this type of data: the first is to adopt features inspired by Conversation Analysis (in particular for turn-taking), the second is to extract the features from individual turns rather than from entire conversations. The experiments have been performed over a corpus of dyadic chat conversations (77 individuals in total). The performance in identifying the persons involved in each exchange, measured in terms of area under the Cumulative Match Characteristic curve, is 89.5%.
Marco Cristani, Giorgio Roffo, Cristina Segalin, Loris Bazzani, Alessandro Vinciarelli, Vittorio Murino
ACM Multimedia1
2012 Stel Component Analysis: Joint Segmentation, Modeling and Recognition of Objects Classes
abstract
Models that captures the common structure of an object class have appeared few years ago in the literature (Jojic and Caspi in Proceedings of IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), pp. 212–219, 2004 ; Winn and Jojic in Proceedings of International Conference on Computer Vision (ICCV), pp. 756–763, 2005 ); they are often referred as “stel models.” Their main characteristic is to segment objects in clear, often semantic, parts as a consequence of the modeling constraint which forces the regions belonging to a single segment to have a tight distribution over local measurements, such as color or texture. This self-similarity within a region in a single image is typical of many meaningful image parts, even when across different images of similar objects, the corresponding parts may not have similar local measurements. Moreover, the segmentation itself is expected to be consistent within a class, although still flexible. These models have been applied mostly to segmentation scenarios. In this paper, we extent those ideas (1) proposing to capture correlations that exist in structural elements of an image class due to global effects, (2) exploiting the segmentations to capture feature co-occurrences and (3) allowing the use of multiple, eventually sparse, observation of different nature. In this way we obtain richer models more suitable to recognition tasks. We accomplish these requirements using a novel approach we dubbed stel component analysis . Experimental results show the flexibility of the model as it can deal successfully with image/video segmentation and object recognition where, in particular, it can be used as an alternative of, or in conjunction with, bag-of-features and related classifiers, where stel inference provides a meaningful spatial partition of features.
Alessandro Perina, Nebojsa Jojic, Marco Cristani, Vittorio Murino
Int. J. Comput. Vis.3
2012 Free Energy Score Spaces: Using Generative Information in Discriminative Classifiers
abstract
A score function induced by a generative model of the data can provide a feature vector of a fixed dimension for each data sample. Data samples themselves may be of differing lengths (e.g., speech segments or other sequential data), but as a score function is based on the properties of the data generation process, it produces a fixed-length vector in a highly informative space, typically referred to as "score space." Discriminative classifiers have been shown to achieve higher performances in appropriately chosen score spaces with respect to what is achievable by either the corresponding generative likelihood-based classifiers or the discriminative classifiers using standard feature extractors. In this paper, we present a novel score space that exploits the free energy associated with a generative model. The resulting free energy score space (FESS) takes into account the latent structure of the data at various levels and can be shown to lead to classification performance that at least matches the performance of the free energy classifier based on the same generative model and the same factorization of the posterior. We also show that in several typical computer vision and computational biology applications the classifiers optimized in FESS outperform the corresponding pure generative approaches, as well as a number of previous approaches combining discriminating and generative models.
Alessandro Perina, Marco Cristani, Umberto Castellani, Vittorio Murino, Nebojsa Jojic
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 Multiple-shot person re-identification by chromatic and epitomic analyses
Loris Bazzani, Marco Cristani, Alessandro Perina, Vittorio Murino
Pattern Recognit. Lett.2
2011 Custom Pictorial Structures for Re-identification
abstract
We propose a novel methodology for re-identification, based on Pictorial Structures (PS). Whenever face or other biometric information is missing, humans recognize an individual by selectively focusing on the body parts, looking for part-to-part correspondences. We want to take inspiration from this strategy in a re-identification context, using PS to achieve this objective. For single image re-identification, we adopt PS to localize the parts, extract and match their descriptors. When multiple images of a single individual are available, we propose a new algorithm to customize the fit of PS on that specific person, leading to what we call a Custom Pictorial Structure (CPS). CPS learns the appearance of an individual, improving the localization of its parts, thus obtaining more reliable visual characteristics for re-identification. It is based on the statistical learning of pixel attributes collected through spatio-temporal reasoning. The use of PS and CPS leads to state-of-the-art results on all the available public benchmarks, and opens a fresh new direction for research on re-identification.
Dong Seon Cheng, Marco Cristani, Michele Stoppa, Loris Bazzani, Vittorio Murino
BMVC2
2011 Social interaction discovery by statistical analysis of F-formations
abstract
We present a novel approach for detecting social interactions in a crowded scene by employing solely visual cues. The detection of social interactions in unconstrained scenarios is a valuable and important task, especially for surveillance purposes. Our proposal is inspired by the social signaling literature, and in particular it considers the sociological notion of F-formation. An F-formation is a set of possible configurations in space that people may assume while participating in a social interaction. Our system takes as input the positions of the people in a scene and their (head) orientations; then, employing a voting strategy based on the Hough transform, it recognizes F-formations and the individuals associated with them. Experiments on simulations and real data promote our idea.
Marco Cristani, Loris Bazzani, Giulia Paggetti, Andrea Fossati, Diego Tosato, Alessio Del Bue, Gloria Menegaz, Vittorio Murino
BMVC1
2011 Fast FPGA-based architecture for pedestrian detection based on covariance matrices
abstract
Pedestrian detection is a crucial task in several video surveillance and automotive scenarios, but only a few detection systems are designed to be realized on an embedded architecture, allowing to increase the processing speed which is one of the key requirements in real applications. In this paper, we propose a novel SoC (System on Chip) architecture for fast pedestrian detection in video. Our implementation is based on a linear SVM (Support Vector Machine) classification frame- work, learned on a set of overlapped image patches. Each patch is described by a covariance matrix of a set of image features. Exploiting the inner parallelism of the FPGA (Field Programmable Gate Array) boards, we dramatically accelerate the covariance matrices computation that plays a crucial role in the framework. In the experiments, we show the effectiveness and the efficiency of our pedestrian detection system, reaching a detection speed of 132 fps at VGA resolution.
Samuele Martelli, Diego Tosato, Marco Cristani, Vittorio Murino
ICIP3
2011 An Experimental Framework for Evaluating PTZ Tracking Algorithms
Pietro Salvagnini, Marco Cristani, Alessio Del Bue, Vittorio Murino
ICVS2
2011 Statistical 3D Shape Analysis by Local Generative Descriptors
abstract
In this paper, we propose a new approach for surface representation. Generative models are exploited for encoding the variations of local geometric properties of 3D shapes. Surfaces are locally modeled as a stochastic process which spans a neighborhood area through a set of circular geodesic pathways, captured by a modified version of a Hidden Markov Model (HMM) named multicircular HMM (MC-HMM). The approach proposed consists of two main phases: 1) local geometric feature collection and 2) MC-HMM parameter estimation. The effectiveness of our proposal is demonstrated by several applicative scenarios, all using well-known benchmark data sets, such as multiple view registration, matching of deformable shapes, and object recognition on cluttered scenes. The results achieved are very promising and open up the use of generative models as geometric descriptors in an extensive range of applications.
Umberto Castellani, Marco Cristani, Vittorio Murino
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Generative modeling and classification of dialogs by a low-level turn-taking feature
Marco Cristani, Anna Pesarin, Carlo Drioli, Alessandro Tavano, Alessandro Perina, Vittorio Murino
Pattern Recognit.1
2010 Person re-identification by symmetry-driven accumulation of local features
abstract
In this paper, we present an appearance-based method for person re-identification. It consists in the extraction of features that model three complementary aspects of the human appearance: the overall chromatic content, the spatial arrangement of colors into stable regions, and the presence of recurrent local motifs with high entropy. All this information is derived from different body parts, and weighted opportunely by exploiting symmetry and asymmetry perceptual principles. In this way, robustness against very low resolution, occlusions and pose, viewpoint and illumination changes is achieved. The approach applies to situations where the number of candidates varies continuously, considering single images or bunch of frames for each individual. It has been tested on several public benchmark datasets (ViPER, iLIDS, ETHZ), gaining new state-of-the-art performances.
Michela Farenzena, Loris Bazzani, Alessandro Perina, Vittorio Murino, Marco Cristani
CVPR5
2010 Object Recognition with Hierarchical Stel Models
Alessandro Perina, Nebojsa Jojic, Umberto Castellani, Marco Cristani, Vittorio Murino
ECCV (6)4
2010 Multi-class Classification on Riemannian Manifolds for Video Surveillance
Diego Tosato, Michela Farenzena, Mauro Spera, Vittorio Murino, Marco Cristani
ECCV (2)5
2010 Collaborative particle filters for group tracking
abstract
Tracking groups of people is a highly informative task in surveillance, and it represents a still open and little explored issue. In this paper, we propose a brand new framework for group tracking, that consists in two separate particle filters, one focusing on groups as atomic entities (the multi-group tracker), and the other modeling each individual separately (the multi-object tracker). The latter helps the multi-group tracker in better defining the nature of a group, evaluating the membership of each individual with respect to different groups, and allowing a robust management of the occlusions. The coupling of the two processes is theoretically founded due to the revision of the posterior distribution of the multi-group tracker with the statistics accumulated by the multi-object tracker. Experimental comparative results certify the goodness of the proposed technique.
Loris Bazzani, Marco Cristani, Vittorio Murino
ICIP2
2010 Part-based human detection on Riemannian manifolds
abstract
In this paper we propose a novel part-based framework for pedestrian detection. We model a human as a hierarchy of fixed overlapped parts, each of which described by covariances of features. Each part is modeled by a boosted classifier, learnt using Logitboost on Riemannian manifolds. All the classifiers are then linked to form a high-level classifier, through weighted summation, whose weights are estimated during the learning. The final classifier is simple, light and robust. The experimental results show that we outperform the state-of-the-art human detection performances on the INRIA person dataset.
Diego Tosato, Michela Farenzena, Marco Cristani, Vittorio Murino
ICIP3
2010 Multiple-Shot Person Re-identification by HPE Signature
abstract
In this paper, we propose a novel appearance-based method for person re-identification, that condenses a set of frames of the same individual into a highly informative signature, called Histogram Plus Epitome, HPE. It incorporates complementary global and local statistical descriptions of the human appearance, focusing on the overall chromatic content, via histograms representation, and on the presence of recurrent local patches, via epitome estimation. The matching of HPEs provides optimal performances against low resolution, occlusions, pose and illumination variations, defining novel state-of-the-art results on all the datasets considered.
Loris Bazzani, Marco Cristani, Alessandro Perina, Michela Farenzena, Vittorio Murino
ICPR2
2010 2LDA: Segmentation for Recognition
abstract
Following the trend of “segmentation for recognition”, we present 2LDA, a novel generative model to automatically segment an image in 2 segments, background and foreground, while inferring a latent Dirichlet allocation (LDA) topic distribution on both segments. The idea is to merge two separate modules, LDA and the segmentation module, explicitly considering (and exchanging) the uncertainty between them. The resulting model adds spatial relationships to LDA, which in turn helps in using the topics to segment an image. The experimental results show that, unlike LDA, our model can be used to recognize objects, and also outperforms the state of the art algorithms.
Alessandro Perina, Marco Cristani, Vittorio Murino
ICPR2
2010 A Re-evaluation of Pedestrian Detection on Riemannian Manifolds
abstract
Boosting covariance data on Riemannian manifolds has proven to be a convenient strategy in a pedestrian detection context. In this paper we show that the detection performances of the state-of-the-art approach of Tuzel et al. can be greatly improved, from both a computational and a qualitative point of view, by considering practical and theoretical issues, and allowing also the estimation of occlusions in a fine way. The resulting detection system reaches the best performance on the INRIA dataset, setting novel state-of-the art results.
Diego Tosato, Michela Farenzena, Marco Cristani, Vittorio Murino
ICPR3
2010 Pervasive video analysis: workshop overview
abstract
This workshop aims at tackling the novel challenging scenarios in pervasive video analysis which require not only to address specific problems (e.g., tracking, recognition) on a single view, but to deal with a set of distributed observations, eventually integrated with subjective mobile video streams. Accepted papers cover a wide range of subjects going from the joint analysis of video sequences, taken from fixed location and mobile cameras, to situation awareness and understanding.
Hamid K. Aghajan, Marco Cristani, Vittorio Murino, Nicu Sebe
ACM Multimedia2
2010 Toward an automatically generated soundtrack from low-level cross-modal correlations for automotive scenarios
abstract
In this paper, we propose a novel recommendation policy for driving scenarios. While driving a car, listening to an audio track may enrich the atmosphere, conveying emotions that let the driver sense a more arousing experience. Here, we are introducing a recommendation policy that, given a video sequence taken by a camera mounted onboard a car, chooses the most suitable audio piece from a predetermined set of melodies. The mixing mechanism takes inspiration from a set of generic qualitative aesthetical rules for cross-modal linking, realized by associating audio and video features. The contribution of this paper is to translate such qualitative rules into quantitative terms, learning from an extensive training dataset cross-modal statistical correlations, and validating them in a thoroughly way. In this way, we are able to define what are the audio and video features that correlate at best (i.e., promoting or rejecting some aesthetical rules), and what are their correlation intensities. This knowledge is then employed for the realization of the recommendation policy. A set of user studies illustrate and validate the policy, thus encouraging further developments toward a real implementation in an automotive application.
Marco Cristani, Anna Pesarin, Carlo Drioli, Vittorio Murino, Antonio Rodà, Michele Grapulin, Nicu Sebe
ACM Multimedia1
2010 Learning natural scene categories by selective multi-scale feature extraction
Alessandro Perina, Marco Cristani, Vittorio Murino
Image Vis. Comput.2
2009 Learning Approach to Analyze Tumour Heterogeneity in DCE-MRI Data During Anti-cancer Treatment
Alessandro Daducci, Umberto Castellani, Marco Cristani, Paolo Farace, Pasquina Marzola, Andrea Sbarbati, Vittorio Murino
AIME3
2009 Stel component analysis: Modeling spatial correlations in image class structure
abstract
As a useful concept in the study of the low level image class structure, we introduce the notion of a structure element - `stel.' The notion is related to the notions of a pixel, superpixel, segment or a part, but instead of referring to an element or a region of a single image, stel is a probabilistic element of an entire image class. Stels often define clear object or scene parts as a consequence of the modeling constraint which forces the regions belonging to a single stel to have a tight distribution over local measurements, such as color or texture. This self-similarity within a region in a single image is typical of most meaningful image parts, even when in different images of similar objects the corresponding parts may not have similar local measurements. The stel itself is expected to be consistent within a class, yet flexible, which we accomplish using a novel approach we dubbed stel component analysis. Experimental results show how stel component analysis can assist in image/video segmentation and object recognition where, in particular, it can be used as an alternative of, or in conjunction with, bag-of-features and related classifiers, where stel inference provides a meaningful spatial partition of features.
Nebojsa Jojic, Alessandro Perina, Marco Cristani, Vittorio Murino, Brendan J. Frey
CVPR3
2009 A hybrid generative/discriminative classification framework based on free-energy terms
abstract
Hybrid, generative-discriminative, techniques have proven to be valuable approaches in tackling difficult object or scene recognition problems. In general, a generative model over the available data for each image class is first learned providing a relatively comprehensive statistical multi-level representation. In this way, new meaningful image features become available, which encode the degree of fitness of the data with respect to the model at different representation levels. Such features are then fed into a discriminative classifier which can exploit the intrinsic data separability. In this paper, we propose the use of variational free energy terms as feature vectors, so that the degree of fitness of the data and the uncertainty over the generative process are explicitly included in the data description. The proposed method is automatically superior to a pure generative classification, and we also experimentally validate it on a wide selection of generative models applied to challenging benchmarks in hard computer vision tasks such as scene, object, and shape recognition. In several instances, the proposed approach outperforms the current state-of-the-art techniques as for classification results, while also showing to be computationally inexpensive.
Alessandro Perina, Marco Cristani, Umberto Castellani, Vittorio Murino, Nebojsa Jojic
ICCV2
2009 Online subjective feature selection for occlusion management in tracking applications
abstract
Most of the state-of-the-art tracking algorithms are prone to error when dealing with occlusions, especially when the involved moving objects are hardly discernible in appearance. In this paper, we propose a multi-object particle filtering tracking framework particularly suited to manage the occlusion problem. The presented solution consists in the introduction of a online subjective feature selection mechanism, which highlights and employs the most discriminant features characterizing a single object with respect to the neighbouring objects. The policy adopted fits formally in the observation step of the particle filtering process, it is effective and not computationally costly. Trials carried out on illustrative synthetic data and on recent challenging benchmark sequences report compelling performances and encourage further development of the technique.
Loris Bazzani, Marco Cristani, Manuele Bicego, Vittorio Murino
ICIP2
2009 Free energy score space
abstract
Score functions induced by generative models extract fixed-dimension feature vectors from different-length data observations by subsuming the process of data generation, projecting them in highly informative spaces called score spaces. In this way, standard discriminative classifiers are proved to achieve higher performances than a solely generative or discriminative approach. In this paper, we present a novel score space that exploits the free energy associated to a generative model through a score function. This function aims at capturing both the uncertainty of the model learning and ``local compliance of data observations with respect to the generative process. Theoretical justifications and convincing comparative classification results on various generative models prove the goodness of the proposed strategy.
Alessandro Perina, Marco Cristani, Umberto Castellani, Vittorio Murino, Nebojsa Jojic
NIPS2
2009 Fully non-homogeneous hidden Markov model double net: A generative model for haplotype reconstruction and block discovery
Alessandro Perina, Marco Cristani, Luciano Xumerle, Vittorio Murino, Pier Franco Pignatti, Giovanni Malerba
Artif. Intell. Medicine2
2008 Unsupervised Learning of Saliency Concepts for Natural Image Classification and Retrieval
Alessandro Perina, Marco Cristani, Vittorio Murino
CIARP2
2008 Geo-located image analysis using latent representations
abstract
Image categorization is undoubtedly one of the most challenging open problems faced in computer vision, far from being solved by employing pure visual cues. Recently, additional textual ldquotagsrdquo can be associated to images, enriching their semantic interpretation beyond the pure visual aspect, and helping to bridge the so-called semantic gap. One of the latest class of tags consists in geo-location data, containing information about the geographical site where an image has been captured. Such data motivate, if not require, novel strategies to categorize images, and pose new problems to focus on. In this paper, we present a statistical method for geo-located image categorization, in which categories are formed by clustering geographically proximal images with similar visual appearance. The proposed strategy permits also to deal with the geo-recognition problem, i.e., to infer the geographical area depicted by images with no available location information. The method lies in the wide literature on statistical latent representations, in particular, the probabilistic latent semantic analysis (pLSA) paradigm has been extended, introducing a latent aspect which characterizes peculiar visual features of different geographical zones. Experiments on categorization and georecognition have been carried out employing a well-known geographical image repository: results are actually very promising, opening new interesting challenges and applications in this research field.
Marco Cristani, Alessandro Perina, Umberto Castellani, Vittorio Murino
CVPR1
2008 A statistical signature for automatic dialogue classification
abstract
In the last few years, there has been a certain attention to the problem of human-human communication, trying to devise artificial systems able to mediate a conversational setting between two or more people. In this paper, we designed an automatic system based on a generative structure able to classify hard dialog acts. The generative model is composed by integrating a hierarchical Gaussian mixture model and the Influence Model, originating a brand new method able to deal with such difficult scenarios. The method has been tested on a set of conversational settings involving dialogues between adults and children and adults, in flat and arguing discussions, proving very accurate classification results.
Anna Pesarin, Marco Cristani, Vittorio Murino, Carlo Drioli, Alessandro Perina, Alessandro Tavano
ICPR2
2008 Geo-located Image Grouping Using Latent Descriptions
Marco Cristani, Alessandro Perina, Vittorio Murino
ICVS1
2008 Visual MRI: Merging information visualization and non-parametric clustering techniques for MRI dataset analysis
Umberto Castellani, Marco Cristani, Carlo Combi, Vittorio Murino, Andrea Sbarbati, Pasquina Marzola
Artif. Intell. Medicine2
2008 Sparse points matching by combining 3D mesh saliency with statistical descriptors
abstract
Abstract This paper proposes new methodology for the detection and matching of salient points over several views of an object. The process is composed by three main phases. In the first step, detection is carried out by adopting a new perceptually‐inspired 3D saliency measure. Such measure allows the detection of few sparse salient points that characterize distinctive portions of the surface. In the second step, a statistical learning approach is considered to describe salient points across different views. Each salient point is modelled by a Hidden Markov Model (HMM), which is trained in an unsupervised way by using contextual 3D neighborhood information, thus providing a robust and invariant point signature. Finally, in the third step, matching among points of different views is performed by evaluating a pairwise similarity measure among HMMs. An extensive and comparative experimental session has been carried out, considering real objects acquired by a 3D scanner from different points of view, where objects come from standard 3D databases. Results are promising, as the detection of salient points is reliable, and the matching is robust and accurate.
Umberto Castellani, Marco Cristani, Simone Fantoni, Vittorio Murino
Comput. Graph. Forum2
2007 Audio-Visual Event Recognition in Surveillance Video Sequences
abstract
In the context of the automated surveillance field, automatic scene analysis and understanding systems typically consider only visual information, whereas other modalities, such as audio, are typically disregarded. This paper presents a new method able to integrate audio and visual information for scene analysis in a typical surveillance scenario, using only one camera and one monaural microphone. Visual information is analyzed by a standard visual background/foreground (BG/FG) modelling module, enhanced with a novelty detection stage and coupled with an audio BG/FG modelling scheme. These processes permit one to detect separate audio and visual patterns representing unusual unimodal events in a scene. The integration of audio and visual data is subsequently performed by exploiting the concept of synchrony between such events. The audio-visual (AV) association is carried out online and without need for training sequences, and is actually based on the computation of a characteristic feature called audio-video concurrence matrix, allowing one to detect and segment AV events, as well as to discriminate between them. Experimental tests involving classification and clustering of events show all the potentialities of the proposed approach, also in comparison with the results obtained by employing the single modalities and without considering the synchrony issue
Marco Cristani, Manuele Bicego, Vittorio Murino
IEEE Trans. Multim.1
2006 Acoustic Range Image Segmentation by Effective Mean Shift
abstract
Image perception in underwater environment is a difficult task for a human operator, and data segmentation becomes a crucial step toward an higher level interpretation and recognition of the observing scenarios. This paper contributes to the related state of the art, by fitting the mean shift clustering paradigm to the segmentation of acoustical range images, providing a segmentation approach in which whatever parameter tuning is absent. Moreover, the method exploits actively the connectivity information provided by the range map, by using reverse projection as acceleration technique. Therefore, the method is able to produce, starting from raw range data, meaningful segmented clouds of points in a fully automatic and efficient fashion.
Umberto Castellani, Marco Cristani, Vittorio Murino
ICIP2
2006 Unsupervised scene analysis: A hidden Markov model approach
Manuele Bicego, Marco Cristani, Vittorio Murino
Comput. Vis. Image Underst.2
2004 Audio-Video Integration for Background Modelling
Marco Cristani, Manuele Bicego, Vittorio Murino
ECCV (2)1