Pietro Morerio

dblp:119/8523 · DBLP profile ↗
← Back
60ranked-venue papers
10as first author
30since 2021 · last 2026
0000-0001-5259-1496ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 7 first-author · 21 since 2021Artificial intelligence and machine learning · 31 · 2 first-author · 18 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Uncertainty-guided Open-Set Source-Free Unsupervised Domain Adaptation with Target-private Class Segregation
abstract
Standard Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target, requiring simultaneous access to both source and target data. Moreover, UDA approaches commonly assume that source and target domains share the same labels space. Yet, these two assumptions are hardly satisfied in real-world scenarios. This paper considers the more challenging Source-Free Open-set Domain Adaptation (SF-OSDA) setting, where both assumptions do not hold. We propose a novel approach for SF-OSDA that takes advantage of the granularity of target-private categories by segregating their samples into multiple unknown classes. Starting from an initial clustering-based pseudo-labels initialisation, our method progressively improves the segregation of target-private samples by refining their pseudo-labels with the guide of an uncertainty-based sample selection module. Additionally, we propose a novel contrastive loss, named NL-InfoNCELoss, that, integrating negative learning into self-supervised contrastive learning, enhances the model robustness to noisy pseudo-labels. Extensive experiments on benchmark datasets demonstrate the superiority of our proposed approach over competing methods, establishing new state-of-the-art performance. Notably, additional analyses show that our method is able to learn the underlying semantics of novel classes, opening the possibility to perform novel class discovery.
Mattia Litrico, Davide Talon, Sebastiano Battiato, Alessio Del Bue, Mario Valerio Giuffrida, Pietro Morerio
Int. J. Comput. Vis.6
2025 Embodied Image Captioning: Self-Supervised Learning Agents for Spatially Coherent Image Descriptions
abstract
We present a self-supervised method to improve an agent's abilities in describing arbitrary objects while actively exploring a generic environment. This is a challenging problem, as current models struggle to obtain coherent image captions due to different camera viewpoints and clutter. We propose a three-phase framework to fine-tune existing captioning models that enhances caption accuracy and consistency across views via a consensus mechanism. First, an agent explores the environment, collecting noisy image-caption pairs. Then, a consistent pseudo-caption for each object instance is distilled via consensus using a large language model. Finally, these pseudo-captions are used to fine-tune an off-the-shelf captioning model, with the addition of contrastive learning. We analyse the performance of the combination of captioning models, exploration policies, pseudo-labeling methods, and fine-tuning strategies, on our manually labeled test set. Results show that a policy can be trained to mine samples with higher disagreement compared to classical baselines. Our pseudo-captioning method, in combination with all policies, has a higher semantic similarity compared to other existing methods, and fine-tuning improves caption accuracy and consistency by a significant margin. Code and test set annotations available at https://hsp-iit.github.io/embodied-captioning/
Tommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio, Alessio Del Bue, Lorenzo Natale
ICCV4
2025 ReassembleNet: Learnable Keypoints and Diffusion for 2D Fresco Reconstruction
abstract
The task of reassembly is a significant challenge across multiple domains, including archaeology, genomics, and molecular docking, requiring the precise placement and orientation of elements to reconstruct an original structure. In this work, we address key limitations in state-of-the-art Deep Learning methods for reassembly, namely i) scalability; ii) multimodality; and iii) real-world applicability: beyond square or simple geometric shapes, realistic and complex erosion, or other real-world problems. We propose ReassembleNet, a method that reduces complexity by representing each input piece as a set of contour keypoints and learning to select the most informative ones by Graph Neural Networks pooling inspired techniques. ReassembleNet effectively lowers computational complexity while enabling the integration of features from multiple modalities, including both geometric and texture data. Further enhanced through pretraining on a semi-synthetic dataset. We then apply diffusion-based pose estimation to recover the original structure. We improve on prior methods by 57% and 87% for RMSE Rotation and Translation, respectively.
Adeela Islam, Stefano Fiorini, Stuart James, Pietro Morerio, Alessio Del Bue
ICCV4
2025 BillBoard Splatting (BBSplat): Learnable Textured Primitives for Novel View Synthesis
abstract
We present billboard Splatting (BBSplat) - a novel approach for novel view synthesis based on textured geometric primitives. BBSplat represents the scene as a set of optimizable textured planar primitives with learnable RGB textures and alpha-maps to control their shape. BBSplat primitives can be used in any Gaussian Splatting pipeline as drop-in replacements for Gaussians. The proposed primitives close the rendering quality gap between 2D and 3D Gaussian Splatting (GS), enabling the accurate extraction of 3D mesh as in the 2DGS framework. Additionally, the explicit nature of planar primitives enables the use of the ray-tracing effects in rasterization. Our novel regularization term encourages textures to have a sparser structure, enabling an efficient compression that leads to a reduction in the storage space of the model up to x17 times compared to 3DGS. Our experiments show the efficiency of BBSplat on standard datasets of real indoor and outdoor scenes such as Tanks&Temples, DTU, and Mip-NeRF-360. Namely, we achieve a state-of-the-art PSNR of 29.72 for DTU at Full HD resolution.
David Svitov, Pietro Morerio, Lourdes Agapito, Alessio Del Bue
ICCV2
2025 Direction-Aware Room Impulse Response Estimation for Immersive Audio Rendering in Real Environments
abstract
Evolving multimedia systems are increasingly being adopted in virtual reality and gaming applications. Such systems emphasize immersion to engage users by bridging the gap between real and virtual content. In this context, visual and acoustic stimuli are the two key media that dictate such immersion. While visual 3D rendering is advancing rapidly, the same is not true for audio, where most research is limited to the reconstruction of the room impulse response (RIR) using omnidirectional audio or, at best, binaural. Such methods do not adequately account for the directions and orientations of the acoustic signals with respect to either the source or the listener, thereby compromising immersion quality. In this work, we explore the effect of adding such "directionality" to the training data to improve the estimation of the room’s acoustic parameters. A more accurate set of such parameters implies in fact a more realistic predicted RIR, leading to a more immersive experience of the acoustic scene. Specifically, we propose a novel framework driven by a suitable loss function to account for directionality in ambisonic microphones, and novel variants of loss functions for both omnidirectional and ambisonic cases. We also propose to account for microphone characteristics and their contribution to the predicted RIRs. Experiments were performed using two datasets of real recordings and the results established the efficacy of the proposed methods
Giovanni Zanin, Ritujoy Biswas, Pietro Morerio, Sylvio Barbon Junior, Alberto Carini, Alessio Del Bue, Vittorio Murino
ACM Multimedia3
2025 Pre-trained Multiple Latent Variable Generative Models are Good Defenders Against Adversarial Attacks
abstract
Attackers can deliberately perturb classifiers' input with subtle noise, altering final predictions. Among proposed countermeasures, adversarial purification employs generative networks to preprocess input images, filtering out adversarial noise. In this study, we propose specific generators, defined Multiple Latent Variable Generative Models (MLVGMs), for adversarial purification. These models possess multiple latent variables that naturally disentangle coarse from fine features. Taking advantage of these properties, we autoencode images to maintain class-relevant information, while discarding and re-sampling any detail, including adversarial noise. The procedure is completely training-free, exploring the generalization abilities of pretrained MLVGMs on the adversarial purification down-stream task. Despite the lack of large models, trained on billions of samples, we show that smaller MLVGMs are already competitive with traditional methods, and can be used as foundation models. Official code released at https://github.com/SerezD/gen_adversarial.
Dario Serez, Marco Cristani, Alessio Del Bue, Vittorio Murino, Pietro Morerio
WACV5
2025 Guest Editorial: Special Issue on Multimodal Learning
Michael Ying Yang, Paolo Rota, Massimiliano Mancini, Pietro Morerio, Bodo Rosenhahn, Vittorio Murino
Int. J. Comput. Vis.4
2025 CDHN: Cross-domain hallucination network for 3D keypoints estimation
abstract
This paper presents a novel method to estimate sparse 3D keypoints from single-view RGB images . Our network is trained in two steps using a knowledge distillation framework. In the first step, the teacher is trained to extract 3D features from point cloud data, which are used in combination with 2D features to estimate the 3D keypoints. In the second step, the teacher teaches the student module to hallucinate the 3D features from RGB images that are similar to those extracted from the point clouds. This procedure helps the network during inference to extract 2D and 3D features directly from images, without requiring point clouds as input. Moreover, the network also predicts a confidence score for every keypoint, which is used to select the valid ones from a set of N predicted keypoints. This allows the prediction of different number of keypoints depending on the object’s geometry. We use the estimated keypoints for computing the relative pose between two views of an object. The results are compared with those of KP-Net and StarMap , which are the state-of-the-art for estimating 3D keypoints from a single-view RGB image. The average angular distance error of our approach (5.94°) is 8.46° and 55.26° lower than that of KP-Net (14.40°) and StarMap (61.20°), respectively.
Mohammad Zohaib, Milind Gajanan Padalkar, Pietro Morerio, Matteo Taiana, Alessio Del Bue
Pattern Recognit.3
2024 HAHA: Highly Articulated Gaussian Human Avatars with Textured Mesh Prior
David Svitov, Pietro Morerio, Lourdes Agapito, Alessio Del Bue
ACCV (9)2
2024 DiffAssemble: A Unified Graph-Diffusion Model for 2D and 3D Reassembly
abstract
Reassembly tasks play a fundamental role in many fields and multiple approaches exist to solve specific reassembly problems. In this context, we posit that a general unified model can effectively address them all, irrespective of the input data type (images, 3D, etc.). We introduce DiffAssemble, a Graph Neural Network (GNN)-based architecture that learns to solve reassembly tasks using a diffusion model formulation. Our method treats the elements of a set, whether pieces of 2D patch or 3D object fragments, as nodes of a spatial graph. Training is performed by introducing noise into the position and rotation of the elements and iteratively denoising them to reconstruct the coherent initial pose. DiffAssemble achieves state-of-the-art (SOTA) results in most 2D and 3D reassembly tasks and is the first learning-based approach that solves 2D puzzles for both rotation and translation. Furthermore, we highlight its remarkable reduction in run-time, performing 11 times faster than the quickest optimization-based method for puzzle solving. Code available at https://github.com/IIT-PAVIS/DiffAssemble.
Gianluca Scarpellini, Stefano Fiorini, Francesco Giuliari, Pietro Morerio, Alessio Del Bue
CVPR4
2024 Look Around and Learn: Self-training Object Detection by Exploration
Gianluca Scarpellini, Stefano Rosa, Pietro Morerio, Lorenzo Natale, Alessio Del Bue
ECCV (56)3
2024 Re-assembling the past: The RePAIR dataset and benchmark for real world 2D and 3D puzzle solving
abstract
This paper proposes the RePAIR dataset that represents a challenging benchmark to test modern computational and data driven methods for puzzle-solving and reassembly tasks. Our dataset has unique properties that are uncommon to current benchmarks for 2D and 3D puzzle solving. The fragments and fractures are realistic, caused by a collapse of a fresco during a World War II bombing at the Pompeii archaeological park. The fragments are also eroded and have missing pieces with irregular shapes and different dimensions, challenging further the reassembly algorithms. The dataset is multi-modal providing high resolution images with characteristic pictorial elements, detailed 3D scans of the fragments and meta-data annotated by the archaeologists. Ground truth has been generated through several years of unceasing fieldwork, including the excavation and cleaning of each fragment, followed by manual puzzle solving by archaeologists of a subset of approx. 1000 pieces among the 16000 available. After digitizing all the fragments in 3D, a benchmark was prepared to challenge current reassembly and puzzle-solving methods that often solve more simplistic synthetic scenarios. The tested baselines show that there clearly exists a gap to fill in solving this computationally complex problem.
Theodore Tsesmelis, Luca Palmieri 0002, Marina Khoroshiltseva, Adeela Islam, Gur Elkin, Ofir Itzhak Shahar, Gianluca Scarpellini, Stefano Fiorini, Yaniv Ohayon, Nadav Alali, Sinem Aslan, Pietro Morerio, Sebastiano Vascon, Elena Gravina, Maria Cristina Napolitano, Giuseppe Scarpati, Gabriel Zuchtriegel, Alexandra Spühler, Michel E. Fuchs, Stuart James, Ohad Ben-Shahar, Marcello Pelillo, Alessio Del Bue
NeurIPS12
2024 Leveraging Next-Active Objects for Context-Aware Anticipation in Egocentric Videos
abstract
Objects are crucial for understanding human-object interactions. By identifying the relevant objects, one can also predict potential future interactions or actions that may occur with these objects. In this paper, we study the problem of Short-Term Object interaction anticipation (STA) and propose NAOGAT (Next-Active-Object Guided Anticipation Transformer), a multi-modal end-to-end transformer network, that attends to objects in observed frames in order to anticipate the next-active-object (NAO) and, eventually, to guide the model to predict context-aware future actions. The task is challenging since it requires anticipating future action along with the object with which the action occurs and the time after which the interaction will begin, a.k.a. the time to contact (TTC). Compared to existing video modeling architectures for action anticipation, NAOGAT captures the relationship between objects and the global scene context in order to predict detections for the next active object and anticipate relevant future actions given these detections, leveraging the objects’ dynamics to improve accuracy. One of the key strengths of our approach, in fact, is its ability to exploit the motion dynamics of objects within a given clip , which is often ignored by other models, and separately decoding the object-centric and motion-centric information. Through our experiments, we show that our model outperforms existing methods on two separate datasets, Ego4D and EpicKitchens-100 ("Unseen Set"), as measured by several additional metrics, such as time to contact, and next-active-object localization. The code can be found on project page : sanketsans.github.io/wacv24
Sanket Kumar Thakur, Cigdem Beyan, Pietro Morerio, Vittorio Murino, Alessio Del Bue
WACV3
2024 Positional diffusion: Graph-based diffusion models for set ordering
abstract
Positional reasoning is the process of ordering an unsorted set of parts into a consistent structure. To address this problem, we present Positional Diffusion , a plug-and-play graph formulation with Diffusion Probabilistic Models. Using a diffusion process, we add Gaussian noise to the set elements’ position and map them to a random position in a continuous space. Positional Diffusion learns to reverse the noising process and recover the original positions through an Attention-based Graph Neural Network. To evaluate our method, we conduct extensive experiments on three different tasks and seven datasets, comparing our approach against the state-of-the-art methods for visual puzzle-solving, sentence ordering, and room arrangement, demonstrating that our method outperforms long-lasting research on puzzle solving with up to + 17 % compared to the second-best deep learning method, and performs on par against the state-of-the-art methods on sentence ordering and room rearrangement. Our work highlights the suitability of diffusion models for ordering problems and proposes a novel formulation and method for solving various ordering tasks. We release our code at https://github.com/IIT-PAVIS/Positional_Diffusion . • The article presents a novel method for Ordering Elements of a Set in 1D and 2D space. • We propose a task-agnostic method, Positional Diffusion for different ordering tasks • Our approach combines Graph Neural Networks with Diffusion Probabilistic Models. • Without any task-specific modes, our method can outperform task-specific approaches. • We test our approach on Sentence ordering, Visual Puzzles, and Furniture Arrangement.
Francesco Giuliari, Gianluca Scarpellini, Stefano Fiorini, Stuart James, Pietro Morerio, Yiming Wang 0002, Alessio Del Bue
Pattern Recognit. Lett.5
2023 Learnable Data Augmentation for One-Shot Unsupervised Domain Adaptation
Julio Ivan Davila Carrazco, Pietro Morerio, Alessio Del Bue, Vittorio Murino
BMVC2
2023 Guiding Pseudo-labels with Uncertainty Estimation for Source-free Unsupervised Domain Adaptation
abstract
Standard Unsupervised Domain Adaptation (UDA) methods assume the availability of both source and target data during the adaptation. In this work, we investigate Source-free Unsupervised Domain Adaptation (SF-UDA), a specific case of UDA where a model is adapted to a target domain without access to source data. We propose a novel approach for the SF-UDA setting based on a loss reweighting strategy that brings robustness against the noise that inevitably affects the pseudo-labels. The classification loss is reweighted based on the reliability of the pseudo-labels that is measured by estimating their uncertainty. Guided by such reweighting strategy, the pseudo-labels are progressively refined by aggregating knowledge from neighbouring samples. Furthermore, a self-supervised contrastive framework is leveraged as a target space regulariser to enhance such knowledge aggregation. A novel negative pairs exclusion strategy is proposed to identify and exclude negative pairs made of samples sharing the same class, even in presence of some noise in the pseudo-labels. Our method outperforms previous methods on three major benchmarks by a large margin. We set the new SF-UDA state-of-the-art on VisDA-C and DomainNet with a performance gain of + 1.8% on both benchmarks and on PACS with + 12.3% in the single-source setting and +6.6% in multi-target adaptation. Additional analyses demonstrate that the proposed approach is robust to the noise, which results in significantly more accurate pseudo-labels compared to state-of-the-art approaches.
Mattia Litrico, Alessio Del Bue, Pietro Morerio
CVPR3
2023 Audio-Visual Inpainting: Reconstructing Missing Visual Information with Sound
abstract
We tackle audio-visual inpainting, the problem of completing an image in such a way to be consistent with the sound associated to the scene. To this end, we propose a multimodal, audio-visual inpainting method (AVIN), and show how to leverage sound to reconstruct semantically consistent images. AVIN is a 2-stage algorithm, which first learns the scene semantics and reconstructs low resolution images based on a conditional probability distribution of pixels in the space conditioned to audio, and then refines such result with a GAN-based network to increase the resolution of the reconstructed image. We show that AVIN is able to recover the original content, especially in the hard cases where the missing area heavily degrades the scene semantics: it can perform cross-modal generation whenever no visual context is observed at all, reconstructing visual data from sound only. Code will be made available upon acceptance.
Valentina Sanguineti, Sanket Kumar Thakur, Pietro Morerio, Alessio Del Bue, Vittorio Murino
ICASSP3
2023 Person Re-Identification without Identification via Event Anonymization
abstract
Wide-scale use of visual surveillance in public spaces puts individual privacy at stake while increasing resource consumption (energy, bandwidth, and computation). Neuromorphic vision sensors (event-cameras) have been recently considered a valid solution to the privacy issue because they do not capture detailed RGB visual information of the subjects in the scene. However, recent deep learning architectures have been able to reconstruct images from event cameras with high fidelity, reintroducing a potential threat to privacy for event-based vision applications. In this paper, we aim to anonymize event-streams to protect the identity of human subjects against such image reconstruction attacks. To achieve this, we propose an end-to-end network architecture jointly optimized for the twofold objective of preserving privacy and performing a downstream task such as person ReId. Our network learns to scramble events, enforcing the degradation of images recovered from the privacy attacker. In this work, we also bring to the community the first ever event-based person ReId dataset gathered to evaluate the performance of our approach. We validate our approach with extensive experiments and report results on the synthetic event data simulated from the publicly available SoftBio dataset and our proposed Event-ReId dataset. The code is available at https://github.com/IIT-PAVIS/ReId_without_Id
Shafiq Ahmad, Pietro Morerio, Alessio Del Bue
ICCV2
2023 Enhancing Next Active Object-Based Egocentric Action Anticipation with Guided Attention
abstract
Short-term action anticipation (STA) in first-person videos is a challenging task that involves understanding the next active object interactions and predicting future actions. Existing action anticipation methods have primarily focused on utilizing features extracted from video clips, but often overlooked the importance of objects and their interactions. To this end, we propose a novel approach that applies a guided attention mechanism between the objects, and the spatiotemporal features extracted from video clips, enhancing the motion and contextual information, and further decoding the object-centric and motion-centric information to address the problem of STA in egocentric videos. Our method, GANO (Guided Attention for Next active Objects) is a multi-modal, end-to-end, single transformer-based network. The experimental results performed on the largest egocentric dataset demonstrate that GANO outperforms the existing state-of-the-art methods for the prediction of the next active object label, its bounding box location, the corresponding future action, and the time to contact the object. The ablation study shows the positive contribution of the guided attention mechanism compared to other fusion methods. Moreover, it is possible to improve the next active object location and class label prediction results of GANO by just appending the learnable object tokens with the region of interest embeddings. Related implementations are available at: sanketsans.github.io/guided-attention-egocentric.html
Sanket Kumar Thakur, Cigdem Beyan, Pietro Morerio, Vittorio Murino, Alessio Del Bue
ICIP3
2023 3DSGrasp: 3D Shape-Completion for Robotic Grasp
abstract
Real-world robotic grasping can be done robustly if a complete 3D Point Cloud Data (PCD) of an object is available. However, in practice, PCDs are often incomplete when objects are viewed from few and sparse viewpoints before the grasping action, leading to the generation of wrong or inaccurate grasp poses. We propose a novel grasping strategy, named 3DSGrasp, that predicts the missing geometry from the partial PCD to produce reliable grasp poses. Our proposed PCD completion network is a Transformer-based encoder-decoder network with an Offset-Attention layer. Our network is inherently invariant to the object pose and point's permutation, which generates PCDs that are geometrically consistent and completed properly. Experiments on a wide range of partial PCD show that 3DSGrasp outperforms the best state-of-the-art method on PCD completion tasks and largely improves the grasping success rate in real-world scenarios. The code and dataset are available at: https://github.com/NunoDuarte/3DSGrasp.
Seyed Saber Mohammadi, Nuno Ferreira Duarte, Dimitrios Dimou, Yiming Wang 0002, Matteo Taiana, Pietro Morerio, Atabak Dehban, Plinio Moreno, Alexandre Bernardino, Alessio Del Bue, José Santos-Victor
ICRA6
2023 Towards Equivariant Optical Flow Estimation with Deep Learning
abstract
Methods for Optical Flow (OF) estimation based on Deep Learning have considerably improved traditional approaches in challenging and realistic conditions. However, data-driven approaches can inherently be biased, leading to unexpected under-performance in real application scenarios. In this paper, we first observe that the OF estimation accuracy varies with motion direction, and name this phenomenon ‘OF sign imbalance’. The sign imbalance cannot be assessed by means of the endpoint-error (EPE), the typical training and evaluation metric for Deep Optical Flow estimators. This paper tackles this issue by proposing a new metric to assess the sign imbalance, which is compared to the endpoint-error. We provide an extensive evaluation of the sign imbalance for the state-of-the-art optical flow estimators. Based on the evaluation, we propose two strategies to mitigate the phenomenon, i) by constraining the model estimations during inference, and, ii) by constraining the loss function during training. Testing and training code is available at: www.github.com/stsavian/equivariant_of_estimation.
Stefano Savian, Pietro Morerio, Alessio Del Bue, Andrea Janes, Tammam Tillo
WACV2
2022 Cleaning Noisy Labels by Negative Ensemble Learning for Source-Free Unsupervised Domain Adaptation
abstract
Conventional Unsupervised Domain Adaptation (UDA) methods presume source and target domain data to be simultaneously available during training. Such an assumption may not hold in practice, as source data is often inaccessible (e.g., due to privacy reasons). On the contrary, a pre-trained source model is usually available, which performs poorly on target due to the well-known domain shift problem. This translates into a significant amount of misclassifications, which can be interpreted as structured noise affecting the inferred target pseudo-labels. In this work, we cast UDA as a pseudo-label refinery problem in the challenging source-free scenario. We propose Negative Ensemble Learning (NEL) technique, a unified method for adaptive noise filtering and progressive pseudo-label refinement. NEL is devised to tackle noisy pseudo-labels by enhancing diversity in ensemble members with different stochastic (i) input augmentation and (ii) feedback. The latter is achieved by leveraging the novel concept of Disjoint Residual Labels, which allow propagating diverse information to the different members. Eventually, a single model is trained with the refined pseudo-labels, which leads to a robust performance on the target domain. Extensive experiments show that the proposed method achieves state-of-the-art performance on major UDA benchmarks, such as Digit5, PACS, Visda-C, and DomainNet, without using source data samples at all.
Pietro Morerio, Vittorio Murino
WACV2
2022 Unsupervised Synthetic Acoustic Image Generation for Audio-Visual Scene Understanding
abstract
Acoustic images are an emergent data modality for multimodal scene understanding. Such images have the peculiarity of distinguishing the spectral signature of the sound coming from different directions in space, thus providing a richer information as compared to that derived from single or binaural microphones. However, acoustic images are typically generated by cumbersome and costly microphone arrays which are not as widespread as ordinary microphones. This paper shows that it is still possible to generate acoustic images from off-the-shelf cameras equipped with only a single microphone and how they can be exploited for audio-visual scene understanding. We propose three architectures inspired by Variational Autoencoder, U-Net and adversarial models, and we assess their advantages and drawbacks. Such models are trained to generate spatialized audio by conditioning them to the associated video sequence and its corresponding monaural audio track. Our models are trained using the data collected by a microphone array as ground truth. Thus they learn to mimic the output of an array of microphones in the very same conditions. We assess the quality of the generated acoustic images considering standard generation metrics and different downstream tasks (classification, cross-modal retrieval and sound localization). We also evaluate our proposed models by considering multimodal datasets containing acoustic images, as well as datasets containing just monaural audio signals and RGB video frames. In all of the addressed downstream tasks we obtain notable performances using the generated acoustic data, when compared to the state of the art and to the results obtained using real acoustic images as input.
Valentina Sanguineti, Pietro Morerio, Alessio Del Bue, Vittorio Murino
IEEE Trans. Image Process.2
2021 Audio-Visual Localization by Synthetic Acoustic Image Generation
abstract
Acoustic images constitute an emergent data modality for multimodal scene understanding. Such images have the peculiarity to distinguish the spectral signature of sounds coming from different directions in space, thus providing richer information than the one derived from mono and binaural microphones. However, acoustic images are typically generated by cumbersome microphone arrays, which are not as widespread as ordinary microphones mounted on optical cameras. To exploit this empowered modality while using standard microphones and cameras we propose to leverage the generation of synthetic acoustic images from common audio-video data for the task of audio-visual localization. The generation of synthetic acoustic images is obtained by a novel deep architecture, based on Variational Autoencoder and U-Net models, which is trained to reconstruct the ground truth spatialized audio data collected by a microphone array, from the associated video and its corresponding monaural audio signal. Namely, the model learns how to mimic what an array of microphones can produce in the same conditions. We assess the quality of the generated synthetic acoustic images on the task of unsupervised sound source localization in a qualitative and quantitative manner, while also considering standard generation metrics. Our model is evaluated by considering both multimodal datasets containing acoustic images, used for the training, and unseen datasets containing just monaural audio signals and RGB frames, showing to reach more accurate localization results as compared to the state of the art.
Valentina Sanguineti, Pietro Morerio, Alessio Del Bue, Vittorio Murino
AAAI2
2021 End-To-End Pairwise Human Proxemics from Uncalibrated Single Images
abstract
In this work, we address the ill-posed problem of estimating pairwise metric distances between people using only a single uncalibrated image. We propose an end-to-end model, DeepProx, that takes as inputs two skeletal joints as a set of 2D image coordinates and outputs the metric distance between them. We show that an increased performance is achieved by a geometrical loss over simplified camera parameters provided at training time. Further, DeepProx achieves a remarkable generalisation over novel viewpoints through domain generalisation techniques. We validate our proposed method quantitatively and qualitatively against baselines on public datasets for which we provided groundtruth on interpersonal distances.
Pietro Morerio, Matteo Bustreo, Yiming Wang 0002, Alessio Del Bue
ICIP1
2021 Predicting Gaze from Egocentric Social Interaction Videos and IMU Data
abstract
Gaze prediction in egocentric videos is a fairly new research topic, which might have several applications for assistive technology (e.g., supporting blind people in their daily interactions), security (e.g., attention tracking in risky work environments), education (e.g., augmented / mixed reality training simulators, immersive games) and so forth. Egocentric gaze is typically estimated from video while few works attempt to use inertial measurement unit (IMU) data, a sensor modality often available in wearable devices (e.g., augmented reality headsets). Instead, in this paper, we examine whether joint learning of egocentric video and corresponding IMU data can improve the first-person gaze prediction compared to using these modalities separately. In this respect, we propose a multimodal network and evaluate it on several unconstrained social interaction scenarios captured by a first-person perspective. The proposed multimodal network achieves better results compared to unimodal methods as well as several (multimodal) baselines, showing that using egocentric video together with IMU data can boost the first-person gaze estimation performance.
Sanket Kumar Thakur, Cigdem Beyan, Pietro Morerio, Alessio Del Bue
ICMI3
2021 Single Image Human Proxemics Estimation for Visual Social Distancing
abstract
In this work, we address the problem of estimating the so-called "Social Distancing" given a single uncalibrated image in unconstrained scenarios. Our approach proposes a semi-automatic solution to approximate the homography matrix between the scene ground and image plane. With the estimated homography, we then leverage an off-the-shelf pose detector to detect body poses on the image and to reason upon their inter-personal distances using the length of their body-parts. Inter-personal distances are further locally inspected to detect possible violations of the social distancing rules. We validate our proposed method quantitatively and qualitatively against baselines on public domain datasets for which we provided groundtruth on interpersonal distances. Besides, we demonstrate the application of our method deployed in a real testing scenario where statistics on the inter-personal distances are currently used to improve the safety in a critical environment.
Maya Aghaei, Matteo Bustreo, Yiming Wang 0002, Gian Luca Bailo, Pietro Morerio, Alessio Del Bue
WACV5
2021 Distillation Multiple Choice Learning for Multimodal Action Recognition
abstract
In this work, we address the problem of learning an ensemble of specialist networks using multimodal data, while considering the realistic and challenging scenario of possible missing modalities at test time. Our goal is to leverage the complementary information of multiple modalities to the benefit of the ensemble and each individual network. We introduce a novel Distillation Multiple Choice Learning framework for multimodal data, where different modality networks learn in a cooperative setting from scratch, strengthening one another. The modality networks learned using our method achieve significantly higher accuracy than if trained separately, due to the guidance of other modalities. We evaluate this approach on three video action recognition benchmark datasets. We obtain state-of-the-art results in comparison to other approaches that work with missing modalities at test time.
Nuno C. Garcia, Sarah Adel Bargal, Vitaly Ablavsky, Pietro Morerio, Vittorio Murino, Stan Sclaroff
WACV4
2021 Intra-Camera Supervised Person Re-Identification
abstract
Abstract Existing person re-identification (re-id) methods mostly exploit a large set of cross-camera identity labelled training data. This requires a tedious data collection and annotation process, leading to poor scalability in practical re-id applications. On the other hand unsupervised re-id methods do not need identity label information, but they usually suffer from much inferior and insufficient model performance. To overcome these fundamental limitations, we propose a novel person re-identification paradigm based on an idea ofindependentper-camera identity annotation. This eliminates the most time-consuming and tedious inter-camera identity labelling process, significantly reducing the amount of human annotation efforts. Consequently, it gives rise to a more scalable and more feasible setting, which we callIntra-Camera Supervised (ICS)person re-id, for which we formulate a Multi-tAsk mulTi-labEl (MATE) deep learning method. Specifically, MATE is designed for self-discovering the cross-camera identity correspondence in a per-camera multi-task inference framework. Extensive experiments demonstrate the cost-effectiveness superiority of our method over the alternative approaches on three large person re-id datasets. For example, MATE yields 88.7% rank-1 score on Market-1501 in the proposed ICS person re-id setting, significantly outperforming unsupervised learning models and closely approaching conventional fully supervised learning competitors.
Xiangping Zhu, Xiatian Zhu, Minxian Li, Pietro Morerio, Vittorio Murino, Shaogang Gong
Int. J. Comput. Vis.4
2021 Excitation Dropout: Encouraging Plasticity in Deep Neural Networks
Andrea Zunino, Sarah Adel Bargal, Pietro Morerio, Jianming Zhang 0001, Stan Sclaroff, Vittorio Murino
Int. J. Comput. Vis.3
2020 Leveraging Acoustic Images for Effective Self-supervised Audio Representation Learning
Valentina Sanguineti, Pietro Morerio, Niccolò Pozzetti, Danilo Greco, Marco Cristani, Vittorio Murino
ECCV (22)2
2020 Complex-Object Visual Inspection: Empirical Studies on A Multiple Lighting Solution
abstract
The design of an automatic visual inspection system is usually performed in two stages. While the first stage consists in selecting the most suitable hardware setup for highlighting most effectively the defects on the surface to be inspected, the second stage concerns the development of algorithmic solutions to exploit the potentials offered by the collected data. In this paper, first, we present a novel illumination setup embedding four illumination configurations to resemble diffused, dark-field, and front lighting techniques. Second, we analyze the contributions brought by deploying the proposed setup in the training phase only, mimicking the scenario in which an already developed visual inspection system cannot be modified on the customer site. Along with an exhaustive set of experiments, in this paper, we demonstrate the suitability of the proposed setup for effective illumination of complex-objects, defined as manufactured items with variable surface characteristics that cannot be determined a priori. Eventually, we provide insights into the importance of multiple light configurations availability during training and their natural boosting effect which, without the need to modify the system design in the evaluation phase, lead to improvements in the overall system performance.
Maya Aghaei, Matteo Bustreo, Pietro Morerio, Nicolò Carissimi, Alessio Del Bue, Vittorio Murino
ICPR3
2020 Compact CNN Structure Learning by Knowledge Distillation
abstract
The concept of compressing deep Convolutional Neural Networks (CNNs) is essential to use limited computation, power, and memory resources on embedded devices. However, existing methods achieve this objective at the cost of a drop in inference accuracy in computer vision tasks. To address such a drawback, we propose a framework that leverages knowledge distillation along with customizable block-wise optimization to learn a lightweight CNN structure while preserving better control over the compression-performance tradeoff. Considering specific resource constraints, e.g., floating-point operations per inference (FLOPs) or model-parameters, our method results in a state of the art network compression while being capable of achieving better inference accuracy. In a comprehensive evaluation, we demonstrate that our method is effective, robust, and consistent with results over a variety of network architectures and datasets, at negligible training overhead. In particular, for the already compact network MobileNet_v2, our method offers up to 2× and 5.2× better model compression in terms of FLOPs and model-parameters, respectively, while getting 1.05% better model performance than the baseline network.
Andrea Zunino, Pietro Morerio, Vittorio Murino
ICPR3
2020 Generative Pseudo-label Refinement for Unsupervised Domain Adaptation
abstract
We investigate and characterize the inherent resilience of conditional Generative Adversarial Networks (cGANs) against noise in their conditioning labels, and exploit this fact in the context of Unsupervised Domain Adaptation (UDA). In UDA, a classifier trained on the labelled source set can be used to infer pseudo-labels on the unlabelled target set. However, this will result in a significant amount of misclassified examples (due to the well-known domain shift issue), which can be interpreted as noise injection in the ground-truth labels for the target set. We show that cGANs are, to some extent, robust against such "shift noise". Indeed, cGANs trained with noisy pseudo-labels, are able to filter such noise and generate cleaner target samples. We exploit this finding in an iterative procedure where a generative model and a classifier are jointly trained: in turn, the generator allows to sample cleaner data from the target distribution, and the classifier allows to associate better labels to target samples, progressively refining target pseudo-labels. Results on common benchmarks show that our method performs better or comparably with the unsupervised domain adaptation state of the art.
Pietro Morerio, Riccardo Volpi, Ruggero Ragonesi, Vittorio Murino
WACV1
2020 Audio-Visual Model Distillation Using Acoustic Images
abstract
In this paper, we investigate how to learn rich and robust feature representations for audio classification from visual data and acoustic images, a novel audio data modality. Former models learn audio representations from raw signals or spectral data acquired by a single microphone, with remarkable results in classification and retrieval. However, such representations are not so robust towards variable environmental sound conditions. We tackle this drawback by exploiting a new multimodal labeled action recognition dataset acquired by a hybrid audio-visual sensor that provides RGB video, raw audio signals, and spatialized acoustic data, also known as acoustic images, where the visual and acoustic images are aligned in space and synchronized in time. Using this richer information, we train audio deep learning models in a teacher-student fashion. In particular, we distill knowledge into audio networks from both visual and acoustic image teachers. Our experiments suggest that the learned representations are more powerful and have better generalization capabilities than the features learned from models trained using just single-microphone audio data.
Andrés F. Pérez, Valentina Sanguineti, Pietro Morerio, Vittorio Murino
WACV3
2020 Predicting Intentions from Motion: The Subject-Adversarial Adaptation Approach
abstract
Abstract This paper aims at investigating the action prediction problem from a pure kinematic perspective. Specifically, we address the problem of recognizing future actions, indeed human intentions, underlying a same initial (and apparently unrelated) motor act. This study is inspired by neuroscientific findings asserting that motor acts at the very onset are embedding information about the intention with which are performed, even when different intentions originate from a same class of movements. To demonstrate this claim in computational and empirical terms, we designed an ad hoc experiment and built a new 3D and 2D dataset where, in both training and testing, we analyze a same class of grasping movements underlying different intentions. We investigate how much the intention discriminants generalize across subjects, discovering that each subject tends to affect the prediction by his/her own bias. Inspired by the domain adaptation problem, we propose to interpret each subject as a domain, leading to a novel subject adversarial paradigm. The proposed approach favorably copes with our new problem, boosting the considered baseline features encoding 2D and 3D information and which do not exploit the subject information.
Andrea Zunino, Jacopo Cavazza, Riccardo Volpi, Pietro Morerio, Andrea Cavallo, Cristina Becchio, Vittorio Murino
Int. J. Comput. Vis.4
2020 Learning with Privileged Information via Adversarial Discriminative Modality Distillation
abstract
Heterogeneous data modalities can provide complementary cues for several tasks, usually leading to more robust algorithms and better performance. However, while training data can be accurately collected to include a variety of sensory modalities, it is often the case that not all of them are available in real life (testing) scenarios, where a model has to be deployed. This raises the challenge of how to extract information from multimodal data in the training stage, in a form that can be exploited at test time, considering limitations such as noisy or missing modalities. This paper presents a new approach in this direction for RGB-D vision tasks, developed within the adversarial learning and privileged information frameworks. We consider the practical case of learning representations from depth and RGB videos, while relying only on RGB data at test time. We propose a new approach to train a hallucination network that learns to distill depth information via adversarial learning, resulting in a clean approach without several losses to balance or hyperparameters. We report state-of-the-art results for object classification on the NYUD dataset, and video action recognition on the largest multimodal dataset available for this task, the NTU RGB+D, as well as on the Northwestern-UCLA.
Nuno C. Garcia, Pietro Morerio, Vittorio Murino
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Unsupervised Domain-Adaptive Person Re-Identification Based on Attributes
abstract
Pedestrian attributes, e. g., hair length, clothes type and color, locally describe the semantic appearance of a person. Training person re-identification (ReID) algorithms under the supervision of such attributes have proven to be effective in extracting local features which are important for ReID. Unlike person identity, attributes are consistent across different domains (or datasets). However, most of ReID datasets lack attribute annotations. On the other hand, there are several datasets labeled with sufficient attributes for the case of pedestrian attribute recognition. Exploiting such data for ReID purpose can be a way to alleviate the shortage of attribute annotations in ReID case. In this work, an unsupervised domain adaptive ReID feature learning framework is proposed to make full use of attribute annotations. We propose to transfer attribute-related features from their original domain to the ReID one: to this end, we introduce an adversarial discriminative domain adaptation method in order to learn domain invariant features for encoding semantic attributes. Experiments on three large-scale datasets validate the effectiveness of the proposed ReID framework.
Xiangping Zhu, Pietro Morerio, Vittorio Murino
ICIP2
2019 Scalable and compact 3D action recognition with approximated RBF kernel machines
Jacopo Cavazza, Pietro Morerio, Vittorio Murino
Pattern Recognit.2
2018 Dropout as a Low-Rank Regularizer for Matrix Factorization
abstract
Regularization for matrix factorization (MF) and approximation problems has been carried out in many different ways. Due to its popularity in deep learning, dropout has been applied also for this class of problems. Despite its solid empirical performance, the theoretical properties of dropout as a regularizer remain quite elusive for this class of problems. In this paper, we present a theoretical analysis of dropout for MF, where Bernoulli random variables are used to drop columns of the factors. We demonstrate the equivalence between dropout and a fully deterministic model for MF in which the factors are regularized by the sum of the product of squared Euclidean norms of the columns. Additionally, we inspect the case of a variable sized factorization and we prove that dropout achieves the global minimum of a convex approximation problem with (squared) nuclear norm regularization. As a result, we conclude that dropout can be used as a low-rank regularizer with data dependent singular-value thresholding.
Jacopo Cavazza, Pietro Morerio, Benjamin D. Haeffele, Connor Lane, Vittorio Murino, René Vidal
AISTATS2
2018 Adversarial Feature Augmentation for Unsupervised Domain Adaptation
abstract
Recent works showed that Generative Adversarial Networks (GANs) can be successfully applied in unsupervised domain adaptation, where, given a labeled source dataset and an unlabeled target dataset, the goal is to train powerful classifiers for the target samples. In particular, it was shown that a GAN objective function can be used to learn target features indistinguishable from the source ones. In this work, we extend this framework by (i) forcing the learned feature extractor to be domain-invariant, and (ii) training it through data augmentation in the feature space, namely performing feature augmentation. While data augmentation in the image space is a well established technique in deep learning, feature augmentation has not yet received the same level of attention. We accomplish it by means of a feature generator trained by playing the GAN minimax game against source features. Results show that both enforcing domain-invariance and performing feature augmentation lead to superior or comparable performance to state-of-the-art results in several unsupervised domain adaptation benchmarks.
Riccardo Volpi, Pietro Morerio, Silvio Savarese, Vittorio Murino
CVPR2
2018 Modality Distillation with Multiple Stream Networks for Action Recognition
Nuno C. Garcia, Pietro Morerio, Vittorio Murino
ECCV (8)2
2018 Minimal-Entropy Correlation Alignment for Unsupervised Deep Domain Adaptation
Pietro Morerio, Jacopo Cavazza, Vittorio Murino
ICLR (Poster)1
2018 Video Gesture Analysis for Autism Spectrum Disorder Detection
abstract
Autism is a behavioral neurological disorder affecting a significant percentage of worldwide population. It especially starts manifesting at very low ages, but it is difficult to early diagnose it since there is not a specific exam or trial that is able to spot it safely. Its detection is in fact mainly dependent from the medical expertise used to assess the patient behavior during direct interviews. This work aims at providing an automatic objective support to the doctor for the assessment of (early) diagnosis of possible autistic subjects by only using video sequences. The underlying idea and rationale come from the psychological and neuroscience studies claiming that the executions of simple motor acts are different between pathological and healthy subjects, and this can be sufficient to discriminate between them. To this end, we devised an experiment in which we recorded, using a standard video camera, patient and healthy children performing the same simple gesture of grasping a bottle. By only processing the video clips depicting the grasping action using a recurrent deep neural network, we are able to discriminate with a good accuracy between the 2 classes of subjects. The designed deep model is also able to provide a sort of attention map in which the zones in the video of major interest are identified in space and time: this “explains” in a certain way which areas the model deems more relevant to the classification purpose, which could also be used by the doctor to make the diagnosis. In the end, this work constitutes a first step towards the development of an automatic computational system devoted to the early diagnosis of autistic subjects, providing the medical expert of a supportive objective method, potentially simple to use in clinical and also more open settings.
Andrea Zunino, Pietro Morerio, Andrea Cavallo, Caterina Ansuini, Jessica Podda, Francesca Battaglia, Edvige Veneselli, Cristina Becchio, Vittorio Murino
ICPR2
2017 Hand pose recognition in First Person Vision through graph spectral analysis
abstract
With the growing availability of wearable technology, video recording devices have become so intimately tied to individuals, that they are able to record the movements of users' hands, making hand-based applications one the most explored area in First Person Vision (FPV). In particular, hand pose recognition plays a fundamental role in tasks such as gesture and activity recognition, which in turn represent the base for developing human-machine interfaces or augmented reality applications. In this work we propose a graph-based representation of hands seen from the point of view of the user, obtained through the shape-fitting capability of a modified Instantaneous Topological Map. Spectral analysis of the graph Laplacian allows to arrange eigenvalues in vectors of features, which prove to be discriminative in classifying the considered hand poses.
Mohamad Baydoun, Alejandro Betancourt, Pietro Morerio, Lucio Marcenaro, Matthias Rauterberg, Carlo S. Regazzoni
ICASSP3
2017 Curriculum Dropout
abstract
Dropout is a very effective way of regularizing neural networks. Stochastically “dropping out” units with a certain probability discourages over-specific co-adaptations of feature detectors, preventing overfitting and improving network generalization. Besides, Dropout can be interpreted as an approximate model aggregation technique, where an exponential number of smaller networks are averaged in order to get a more powerful ensemble. In this paper, we show that using a fixed dropout probability during training is a suboptimal choice. We thus propose a time scheduling for the probability of retaining neurons in the network. This induces an adaptive regularization scheme that smoothly increases the difficulty of the optimization problem. This idea of “starting easy” and adaptively increasing the difficulty of the learning problem has its roots in curriculum learning and allows one to train better models. Indeed, we prove that our optimization strategy implements a very general curriculum scheme, by gradually adding noise to both the input and intermediate feature representations within the network architecture. Experiments on seven image classification datasets and different network architectures show that our method, named Curriculum Dropout, frequently yields to better generalization and, at worst, performs just as well as the standard Dropout method.
Pietro Morerio, Jacopo Cavazza, Riccardo Volpi, René Vidal, Vittorio Murino
ICCV1
2017 Left/right hand segmentation in egocentric videos
Alejandro Betancourt, Pietro Morerio, Emilia I. Barakova, Lucio Marcenaro, Matthias Rauterberg, Carlo S. Regazzoni
Comput. Vis. Image Underst.2
2016 A Cognitive Control-Inspired Approach to Object Tracking
abstract
Under a tracking framework, the definition of the target state is the basic step for automatic understanding of dynamic scenes. More specifically, far object tracking raises challenges related to the potentially abrupt size changes of the targets as they approach the sensor. If not handled, size changes can introduce heavy issues in data association and position estimation. This is why adaptability and self-awareness of a tracking module are desirable features. The paradigm of cognitive dynamic systems (CDSs) can provide a framework under which a continuously learning cognitive module can be designed. In particular, CDS theory describes a basic vocabulary of components that can be used as the founding blocks of a module capable to learn behavioral rules from continuous active interactions with the environment. This quality is the fundamental to deal with dynamic situations. In this paper we propose a general CDS-based approach to tracking. We show that such a CDS-inspired design can lead to the self-adaptability of a Bayesian tracker in fusing heterogeneous object features, overcoming size change issues. The experimental results on infrared sequences show how the proposed framework is able to outperform other existing far object tracking methods.
Andrea Mazzù, Pietro Morerio, Lucio Marcenaro, Carlo S. Regazzoni
IEEE Trans. Image Process.2
2015 A Dynamic Approach and a New Dataset for Hand-detection in First Person Vision
Alejandro Betancourt, Pietro Morerio, Emilia I. Barakova, Lucio Marcenaro, Matthias Rauterberg, Carlo S. Regazzoni
CAIP (1)2
2015 Filtering SVM frame-by-frame binary classification in a detection framework
abstract
Classifying frames, or parts of them, is a common way of carrying out detection tasks in computer vision. However, frame by frame classification suffers from sudden significant variations in image texture, colour and luminosity, resulting in noise in the extracted features and consequently in the decisions taken. Support Vector Machines have been widely validated as powerful tools for frame by frame detection of non-separable datasets, but are extremely sensitive to these variations between adjacent frames, creating as consequence sudden flickering in the classification results. This work proposes a Dynamic Bayesian Network to smooth the classification results of Support Vector Machines (SVM) in detection tasks. The method is evaluated in First Person Vision (FPV) videos, where a SVM is used to decide whether or not the user's hands are in his field of view.
Alejandro Betancourt, Pietro Morerio, Lucio Marcenaro, Matthias Rauterberg, Carlo S. Regazzoni
ICIP2
2015 Optimizing Superpixel Clustering for Real-Time Egocentric-Vision Applications
abstract
In this work, we propose a strategy for optimizing a superpixel algorithm for video signals, in order to get closer to real time performances which are on the one hand needed for egocentric vision applications and on the other must be bearable by wearable technologies. Instead of applying the algorithm frame by frame, we propose a technique inspired to Bayesian filtering and to video coding which allows to re-initialize superpixels using the information from the previous frame. This results in faster convergence and demonstrates how performances improve with respect to the standard application of the algorithm from scratch at each frame.
Pietro Morerio, Gabriel Claudiu Georgiu, Lucio Marcenaro, Carlo S. Regazzoni
IEEE Signal Process. Lett.1
2015 The Evolution of First Person Vision Methods: A Survey
abstract
The emergence of new wearable technologies, such as action cameras and smart glasses, has increased the interest of computer vision scientists in the first person perspective. Nowadays, this field is attracting attention and investments of companies aiming to develop commercial devices with first person vision (FPV) recording capabilities. Due to this interest, an increasing demand of methods to process these videos, possibly in real time, is expected. The current approaches present a particular combinations of different image features and quantitative methods to accomplish specific objectives like object detection, activity recognition, user-machine interaction, and so on. This paper summarizes the evolution of the state of the art in FPV video analysis between 1997 and 2014, highlighting, among others, the most commonly used features, methods, challenges, and opportunities within the field.
Alejandro Betancourt, Pietro Morerio, Carlo S. Regazzoni, Matthias Rauterberg
IEEE Trans. Circuits Syst. Video Technol.2
2014 A generative superpixel method
Pietro Morerio, Lucio Marcenaro, Carlo S. Regazzoni
FUSION1
2014 Exploiting an event based state estimator in presence of sparse measurements in video analytics
abstract
Recently, a Bayesian estimator with a hybrid update was developed [1], based on a mathematical formulation of sampling. Such an Event Based State Estimator (EBSE) allows for a stable synchronous state estimate, relying on asynchronous measurements. Usefulness of such a filter comes with its approximate analytic formulation, which is attainable given a send-on-delta sampling strategy. We argue that such a formulation can be extended to cope with a failing detector in case the filter is used for tracking. The basic idea is to approach the issue as a package loss problem, where a missed target is assimilated to a lost package. More in detail, we propose that this approach can be exploited in video tracking, where faulty detectors are commonplace. We show how tracking performance with a poor pedestrian detector, failing to recognize its target, can improve with respect to standard Kalman filter.
Pietro Morerio, Mattia Pompei, Lucio Marcenaro, Carlo S. Regazzoni
ICASSP1
2013 A bio-inspired knowledge representation method for anomaly detection in cognitive Video Surveillance systems
Simone Chiappino, Pietro Morerio, Lucio Marcenaro, Carlo S. Regazzoni
FUSION2
2013 Hand detection in First Person Vision
Pietro Morerio, Lucio Marcenaro, Carlo S. Regazzoni
FUSION1
2012 Performance Evaluation of Multi-camera Visual Tracking
abstract
Main drawbacks in single-camera multi-target visual tracking can be partially removed by increasing the amount of information gathered on the scene, i.e. by adding cameras. By adopting such a multi-camera approach, multiple sensors cooperate for overall scene understanding. However, new issues arise such as data association and data fusion. This work addresses the issue of evaluating the performance of a multi-camera tracking algorithm based on Rao-Blackwellized Monte Carlo data association (RBMCDA) on real data. For this purpose, a new metric based on three performance indexes is developed.
Lucio Marcenaro, Pietro Morerio, Carlo S. Regazzoni
AVSS2
2012 People Count Estimation In Small Crowds
abstract
This work addresses the problem of people counting in crowded situations, such as urban environments, in computer vision. As crowding density increases in a scene, it might become impossible to count people as single individuals: a global group-based approach is then preferable and in fact often necessary. A simple method for estimating the count of people in such tight crowds is here proposed, relying on accurate camera calibration. A training phase is also needed by the algorithm in order to learn the parameters needed for estimation.
Pietro Morerio, Lucio Marcenaro, Carlo S. Regazzoni
AVSS1
2012 A multi-sensor cognitive approach for active security monitoring of abnormal overcrowding situations
Simone Chiappino, Pietro Morerio, Lucio Marcenaro, Elisabetta Fuiano, Giulia Repetto, Carlo S. Regazzoni
FUSION2
2012 Early fire and smoke detection based on colour features and motion analysis
abstract
This work addresses the issue of fire and smoke detection in a scene within a video surveillance framework. Detection of fire and smoke pixels is at first achieved by means of a motion detection algorithm. In addition, separation of smoke and fire pixels using colour information (within appropriate spaces, specifically chosen in order to enhance specific chromatic features) is performed. In parallel, a pixel selection based on the dynamics of the area is carried out in order to reduce false detection. The output of the three parallel algorithms are eventually fused by means of a MLP.
Pietro Morerio, Lucio Marcenaro, Carlo S. Regazzoni, Gianluca Gera
ICIP1