VLDB 2026 Research / reviewers in the wild / expert
Roberto Vezzani
dblp:v/RobertoVezzani
· DBLP profile ↗
50ranked-venue papers
10as first author
11since 2021 · last 2026
0000-0002-1046-6870ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 8 first-author · 5 since 2021Artificial intelligence and machine learning · 28 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GazeD: Context-Aware Diffusion for Accurate 3D Gaze EstimationabstractWe introduce GazeD, a new 3D gaze estimation method that jointly provides 3D gaze and human pose from a single RGB image. Leveraging the ability of diffusion models to deal with uncertainty, it generates multiple plausible 3D gaze and pose hypotheses based on the 2D context information extracted from the input image. Specifically, we condition the denoising process on the 2D pose, the surroundings of the subject, and the context of the scene. With GazeD we also introduce a novel way of representing the 3D gaze by positioning it as an additional body joint at a fixed distance from the eyes. The rationale is that the gaze is usually closely related to the pose, and thus it can benefit from being jointly denoised during the diffusion process. Evaluations across three benchmark datasets demonstrate that GazeD achieves state-of-the-art performance in 3D gaze estimation, even surpassing methods that rely on temporal information. Project details will be available at https://aimagelab.ing.unimore.it/go/gazed Riccardo Catalini, Davide Di Nucci, Guido Borghi, Davide Davoli 0002, Lorenzo Garattoni, Giampiero Francesca, Yuki Kawana, Roberto Vezzani |
3DV | 8 |
| 2026 | RI-PIENO - Revised and Improved Petrol-filling Itinerary Estimation aNd OptimizationabstractEfficient energy provisioning is a fundamental requirement for modern transportation systems, making refueling path optimization a critical challenge. Existing solutions often focus either on inter-vehicle communication or intra-vehicle monitoring, leveraging Intelligent Transportation Systems, Digital Twins, and Software-Defined Internet of Vehicles with Cloud/Fog/Edge infrastructures. However, integrated frameworks that adapt dynamically to driver mobility patterns are still underdeveloped. Building on our previous PIENO framework, we present RI-PIENO (Revised and Improved Petrol-filling Itinerary Estimation aNd Optimization), a system that combines intra-vehicle sensor data with external geospatial and fuel price information, processed via IoT-enabled Cloud/Fog services. RI-PIENO models refueling as a dynamic, time-evolving directed acyclic graph that reflects both habitual daily trips and real-time vehicular inputs, transforming the system from a static recommendation tool into a continuously adaptive decision engine. We validate RI-PIENO in a daily-commute use case through realistic multi-driver, multi-week simulations, showing that it achieves significant cost savings and more efficient routing compared to previous approaches. The framework is designed to leverage emerging roadside infrastructure and V2X communication, supporting scalable deployment within next-generation IoT and vehicular networking ecosystems. Marco Savarese, Antonio de Blasi, Carmine Zaccagnino, Giacomo Salici, Silvia Cascianelli, Roberto Vezzani, Carlo Augusto Grazia |
CCNC | 6 |
| 2026 | A Workflow for Cost- and Time-Aware Refueling Itinerary OptimizationabstractThe complete workflow of the RI-PIENO framework is presented, a system for refueling itinerary optimization that extends the original PIENO design. While prior work introduced the conceptual modules of RI-PIENO, their operational pipeline was not described in detail. This study makes the workflow explicit, covering the end-to-end process from CAN Bus data acquisition and stop detection to the construction of daily trip graphs, refueling optimization, and mileage prediction. By clarifying the sequence of operations, the contribution provides a reproducible and extensible foundation for future research and development. Marco Savarese, Carmine Zaccagnino, Antonio de Blasi, Giacomo Salici, Silvia Cascianelli, Roberto Vezzani, Carlo Augusto Grazia |
CCNC | 6 |
| 2026 | Fake3DGS: A Benchmark for 3D Manipulation Detection in Neural Rendering
Davide Di Nucci, Riccardo Catalini, Guido Borghi, Roberto Vezzani |
ICPR (5) | 4 |
| 2026 | SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses
Alessandro Simoni, Riccardo Catalini, Davide Di Nucci, Guido Borghi, Davide Davoli 0002, Lorenzo Garattoni, Gianpiero Francesca, Yuki Kawana, Roberto Vezzani |
ICPR (3) | 9 |
| 2025 | BRUM: Robust 3D Vehicle Reconstruction from 360° Sparse ImagesabstractAccurate 3D reconstruction of vehicles is vital for applications such as vehicle inspection, predictive maintenance, and urban planning. Existing methods like Neural Radiance Fields and Gaussian Splatting have shown impressive results but remain limited by their reliance on dense input views, which hinders real-world applicability. This paper addresses the challenge of reconstructing vehicles from sparse-view inputs, leveraging depth maps and a robust pose estimation architecture to synthesize novel views and augment training data. Specifically, we enhance Gaussian Splatting by integrating a selective photometric loss, applied only to high-confidence pixels, and replacing standard Structure-from-Motion pipelines with the DUSt3R architecture to improve camera pose estimation. Furthermore, we present a novel dataset featuring both synthetic and real-world public transportation vehicles, enabling extensive evaluation of our approach. Experimental results demonstrate state-of-the-art performance across multiple benchmarks, showcasing the method's ability to achieve high-quality reconstructions even under constrained input conditions. Code and data are publicly available at https://aimagelab.ing.unimore.it/go/brum. Davide Di Nucci, Matteo Tomei, Guido Borghi, Luca Ciuffreda, Roberto Vezzani, Rita Cucchiara |
IV | 5 |
| 2025 | 3D Pose Nowcasting: Forecast the future to improve the presentabstractTechnologies to enable safe and effective collaboration and coexistence between humans and robots have gained significant importance in the last few years. A critical component useful for realizing this collaborative paradigm is the understanding of human and robot 3D poses using non-invasive systems. Therefore, in this paper, we propose a novel vision-based system leveraging depth data to accurately establish the 3D locations of skeleton joints. Specifically, we introduce the concept of Pose Nowcasting, denoting the capability of the proposed system to enhance its current pose estimation accuracy by jointly learning to forecast future poses. The experimental evaluation is conducted on two different datasets, providing accurate and real-time performance and confirming the validity of the proposed method on both the robotic and human scenarios. • We introduce the novel task of 3D Pose Nowcasting. • Our Pose Nowcasting system is based on both 3D Pose Estimation and Forecasting. • We show that knowledge about pose forecasting improves the accuracy of pose estimation. • We apply the proposed system both to human and robots. • Result on different dataset show state-of-the-art performance and robustness. Alessandro Simoni, Francesco Marchetti, Guido Borghi, Federico Becattini, Lorenzo Seidenari, Roberto Vezzani, Alberto Del Bimbo |
Comput. Vis. Image Underst. | 6 |
| 2023 | Depth-based 3D human pose refinement: Evaluating the refinet frameworkabstractIn recent years, Human Pose Estimation has achieved impressive results on RGB images. The advent of deep learning architectures and large annotated datasets have contributed to these achievements. However, little has been done towards estimating the human pose using depth maps, and especially towards obtaining a precise 3D body joint localization. To fill this gap, this paper presents RefiNet, a depth-based 3D human pose refinement framework. Given a depth map and an initial coarse 2D human pose, RefiNet regresses a fine 3D pose. The framework is composed of three modules, based on different data representations, i.e. 2D depth patches, 3D human skeletons, and point clouds. An extensive experimental evaluation is carried out to investigate the impact of the model hyper-parameters and to compare RefiNet with off-the-shelf 2D methods and literature approaches. Results confirm the effectiveness of the proposed framework and its limited computational requirements. Andrea D'Eusanio, Alessandro Simoni, Stefano Pini, Guido Borghi, Roberto Vezzani, Rita Cucchiara |
Pattern Recognit. Lett. | 5 |
| 2021 | Multi-Category Mesh Reconstruction From Image CollectionsabstractRecently, learning frameworks have shown the capability of inferring the accurate shape, pose, and texture of an object from a single RGB image. However, current methods are trained on image collections of a single category in order to exploit specific priors, and they often make use of category-specific 3D templates. In this paper, we present an alternative approach that infers the textured mesh of objects combining a series of deformable 3D models and a set of instance-specific deformation, pose, and texture. Differently from previous works, our method is trained with images of multiple object categories using only foreground masks and rough camera poses as supervision. Without specific 3D templates, the framework learns category-level models which are deformed to recover the 3D shape of the depicted object. The instance-specific deformations are predicted independently for each vertex of the learned 3D mesh, enabling the dynamic subdivision of the mesh during the training process. Experiments show that the proposed framework can distinguish between different object categories and learn category-specific shape priors in an unsupervised manner. Predicted shapes are smooth and can leverage from multiple steps of subdivision during the training process, obtaining comparable or state-of-the-art results on two public datasets. Models and code are publicly released1. Alessandro Simoni, Stefano Pini, Roberto Vezzani, Rita Cucchiara |
3DV | 3 |
| 2021 | SHREC 2021: Skeleton-based hand gesture recognition in the wild
Ariel Caputo, Andrea Giachetti 0001, Simone Soso, Deborah Pintani, Andrea D'Eusanio, Stefano Pini, Guido Borghi, Alessandro Simoni, Roberto Vezzani, Rita Cucchiara, Andrea Ranieri, Franca Giannini, Katia Lupinetti, Marina Monti, Mehran Maghoumi, Joseph J. LaViola Jr., Minh-Quan Le, Hai-Dang Nguyen, Minh-Triet Tran |
Comput. Graph. | 9 |
| 2021 | Video Frame Synthesis Combining Conventional and Event CamerasabstractEvent cameras are biologically-inspired sensors that gather the temporal evolution of the scene. They capture pixel-wise brightness variations and output a corresponding stream of asynchronous events. Despite having multiple advantages with respect to conventional cameras, their use is limited due to the scarce compatibility of asynchronous event streams with traditional data processing and vision algorithms. In this regard, we present a framework that synthesizes RGB frames from the output stream of an event camera and an initial or a periodic set of color key-frames. The deep learning-based frame synthesis framework consists of an adversarial image-to-image architecture and a recurrent module. Two public event-based datasets, DDD17 and MVSEC, are used to obtain qualitative and quantitative per-pixel and perceptual results. In addition, we converted into event frames two additional well-known datasets, namely Kitti and Cityscapes, in order to present semantic results, in terms of object detection and semantic segmentation accuracy. Extensive experimental evaluation confirms the quality and the capability of the proposed approach of synthesizing frame sequences from color key-frames and sequences of intermediate events. Stefano Pini, Guido Borghi, Roberto Vezzani |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2020 | A Transformer-Based Network for Dynamic Hand Gesture RecognitionabstractTransformer-based neural networks represent a successful self-attention mechanism that achieves state-of-the-art results in language understanding and sequence modeling. However, their application to visual data and, in particular, to the dynamic hand gesture recognition task has not yet been deeply investigated. In this paper, we propose a transformer-based architecture for the dynamic hand gesture recognition task. We show that the employment of a single active depth sensor, specifically the usage of depth maps and the surface normals estimated from them, achieves state-of-the-art results, overcoming all the methods available in the literature on two automotive datasets, namely NVidia Dynamic Hand Gesture and Briareo. Moreover, we test the method with other data types available with common RGB-D devices, such as infrared and color data. We also assess the performance in terms of inference time and number of parameters, showing that the proposed framework is suitable for an online in-car infotainment system. Andrea D'Eusanio, Alessandro Simoni, Stefano Pini, Guido Borghi, Roberto Vezzani, Rita Cucchiara |
3DV | 5 |
| 2020 | Baracca: a Multimodal Dataset for Anthropometric Measurements in AutomotiveabstractThe recent spread of depth sensors has enabled new methods to automatically estimate anthropometric measurements, in place of manual procedures or expensive 3D scanners. Generally, the use of depth data is limited by the lack of depth-based public datasets containing accurate anthropometric annotations. Therefore, in this paper we propose a new dataset, called Baracca, specifically designed for the automotive context, including in-car and outside views. The dataset is multimodal: it has been acquired with synchronized depth, infrared, thermal and RGB cameras in order to deal with the requirements imposed by the automotive context. In addition, we propose several baselines to test the challenges of the presented dataset and provide considerations for future work. Stefano Pini, Andrea D'Eusanio, Guido Borghi, Roberto Vezzani, Rita Cucchiara |
IJCB | 4 |
| 2020 | RefiNet: 3D Human Pose Refinement with Depth MapsabstractHuman Pose Estimation is a fundamental task for many applications in the Computer Vision community and it has been widely investigated in the 2D domain, i.e. intensity images. Therefore, most of the available methods for this task are mainly based on 2D Convolutional Neural Networks and huge manually-annotated RGB datasets, achieving stunning results. In this paper, we propose RefiNet, a multi-stage framework that regresses an extremely-precise 3D human pose estimation from a given 2D pose and a depth map. The framework consists of three different modules, each one specialized in a particular refinement and data representation, i.e. depth patches, 3D skeleton and point clouds. Moreover, we present a new dataset, called Baracca, acquired with RGB, depth and thermal cameras and specifically created for the automotive context. Experimental results confirm the quality of the refinement procedure that largely improves the human pose estimations of off-the-shelf 2D methods. Andrea D'Eusanio, Stefano Pini, Guido Borghi, Roberto Vezzani, Rita Cucchiara |
ICPR | 4 |
| 2020 | Face-from-Depth for Head Pose Estimation on Depth ImagesabstractDepth cameras allow to set up reliable solutions for people monitoring and behavior understanding, especially when unstable or poor illumination conditions make unusable common RGB sensors. Therefore, we propose a complete framework for the estimation of the head and shoulder pose based on depth images only. A head detection and localization module is also included, in order to develop a complete end-to-end system. The core element of the framework is a Convolutional Neural Network, called POSEidon+, that receives as input three types of images and provides the 3D angles of the pose as output. Moreover, a Face-from-Depth component based on a Deterministic Conditional GAN model is able to hallucinate a face from the corresponding depth image. We empirically demonstrate that this positively impacts the system performances. We test the proposed framework on two public datasets, namely Biwi Kinect Head Pose and ICT-3DHP, and on Pandora, a new challenging dataset mainly inspired by the automotive setup. Experimental results show that our method overcomes several recent state-of-art works based on both intensity and depth input data, running in real-time at more than 30 frames per second. Guido Borghi, Matteo Fabbri, Roberto Vezzani, Simone Calderara, Rita Cucchiara |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Learning to Generate Facial Depth MapsabstractIn this paper, an adversarial architecture for facial depth map estimation from monocular intensity images is presented. By following an image-to-image approach, we combine the advantages of supervised learning and adversarial training, proposing a conditional Generative Adversarial Network that effectively learns to translate intensity face images into the corresponding depth maps. Two public datasets, namely Biwi database and Pandora dataset, are exploited to demonstrate that the proposed model generates high-quality synthetic depth images, both in terms of visual appearance and informative content. Furthermore, we show that the model is capable of predicting distinctive facial details by testing the generated depth maps through a deep model trained on authentic depth maps for the face verification task. Stefano Pini, Filippo Grazioli, Guido Borghi, Roberto Vezzani, Rita Cucchiara |
3DV | 4 |
| 2018 | Face Verification from Depth using Privileged Information
Guido Borghi, Stefano Pini, Filippo Grazioli, Roberto Vezzani, Rita Cucchiara |
BMVC | 4 |
| 2018 | Learning to Detect and Track Visible and Occluded Body Joints in a Virtual World
Matteo Fabbri, Fabio Lanzi, Simone Calderara, Andrea Palazzi, Roberto Vezzani, Rita Cucchiara |
ECCV (4) | 5 |
| 2018 | Hands on the wheel: A Dataset for Driver Hand Detection and TrackingabstractThe ability to detect, localize and track the hands is crucial in many applications requiring the understanding of the person behavior, attitude and interactions. In particular, this is true for the automotive context, in which hand analysis allows to predict preparatory movements for maneuvers or to investigate the driver's attention level. Moreover, due to the recent diffusion of cameras inside new car cockpits, it is feasible to use hand gestures to develop new Human-Car Interaction systems, more user-friendly and safe. In this paper, we propose a new dataset, called Turms, that consists of infrared images of driver's hands, collected from the back of the steering wheel, an innovative point of view. The Leap Motion device has been selected for the recordings, thanks to its stereo capabilities and the wide view-angle. Besides, we introduce a method to detect the presence and the location of driver's hands on the steering wheel, during driving activity tasks. Guido Borghi, Elia Frigieri, Roberto Vezzani, Rita Cucchiara |
FG | 3 |
| 2018 | Fully Convolutional Network for Head Detection with Depth ImagesabstractHead detection and localization are one of the most investigated and demanding tasks of the Computer Vision community. These are also a key element for many disciplines, like Human Computer Interaction, Human Behavior Understanding, Face Analysis and Video Surveillance. In last decades, many efforts have been conducted to develop accurate and reliable head or face detectors on standard RGB images, but only few solutions concern other types of images, such as depth maps. In this paper, we propose a novel method for head detection on depth images, based on a deep learning approach. In particular, the presented system overcomes the classic sliding-window approach, that is often the main computational bottleneck of many object detectors, through a Fully Convolutional Network. Two public datasets, namely Pandora and Watch-n-Patch, are exploited to train and test the proposed network. Experimental results confirm the effectiveness of the method, that is able to exceed all the state-of-art works based on depth images and to run with real time performance. Diego Ballotta, Guido Borghi, Roberto Vezzani, Rita Cucchiara |
ICPR | 3 |
| 2018 | Domain Translation with Conditional GANs: from Depth to RGB Face-to-FaceabstractCan faces acquired by low-cost depth sensors be useful to catch some characteristic details of the face? Typically the answer is no. However, new deep architectures can generate RGB images from data acquired in a different modality, such as depth data. In this paper, we propose a new Deterministic Conditional GAN, trained on annotated RGB-D face datasets, effective for a face-to-face translation from depth to RGB. Although the network cannot reconstruct the exact somatic features for unknown individual faces, it is capable to reconstruct plausible faces; their appearance is accurate enough to be used in many pattern recognition tasks. In fact, we test the network capability to hallucinate with some Perceptual Probes, as for instance face aspect classification or landmark detection. Depth face can be used in spite of the correspondent RGB images, that often are not available due to difficult luminance conditions. Experimental results are very promising and are as far as better than previously proposed approaches: this domain translation can constitute a new way to exploit depth data in new future applications. Matteo Fabbri, Guido Borghi, Fabio Lanzi, Roberto Vezzani, Simone Calderara, Rita Cucchiara |
ICPR | 4 |
| 2017 | POSEidon: Face-from-Depth for Driver Pose EstimationabstractFast and accurate upper-body and head pose estimation is a key task for automatic monitoring of driver attention, a challenging context characterized by severe illumination changes, occlusions and extreme poses. In this work, we present a new deep learning framework for head localization and pose estimation on depth images. The core of the proposal is a regressive neural network, called POSEidon, which is composed of three independent convolutional nets followed by a fusion layer, specially conceived for understanding the pose by depth. In addition, to recover the intrinsic value of face appearance for understanding head position and orientation, we propose a new Face-from-Depth model for learning image faces from depth. Results in face reconstruction are qualitatively impressive. We test the proposed framework on two public datasets, namely Biwi Kinect Head Pose and ICT-3DHP, and on Pandora, a new challenging dataset mainly inspired by the automotive setup. Results show that our method overcomes all recent state-of-art works, running in real time at more than 30 frames per second. Guido Borghi, Marco Venturelli, Roberto Vezzani, Rita Cucchiara |
CVPR | 3 |
| 2017 | Embedded recurrent network for head pose estimation in carabstractAn accurate and fast driver's head pose estimation is a rich source of information, in particular in the automotive context. Head pose is a key element for driver's behavior investigation, pose analysis, attention monitoring and also a useful component to improve the efficacy of Human-Car Interaction systems. In this paper, a Recurrent Neural Network is exploited to tackle the problem of driver head pose estimation, directly and only working on depth images to be more reliable in presence of varying or insufficient illumination. Experimental results, obtained from two public dataset, namely Biwi Kinect Head Pose and ICT-3DHP Database, prove the efficacy of the proposed method that overcomes state-of-art works. Besides, the entire system is implemented and tested on two embedded boards with real time performance. Guido Borghi, Riccardo Gasparini, Roberto Vezzani, Rita Cucchiara |
Intelligent Vehicles Symposium | 3 |
| 2016 | Fast gesture recognition with Multiple Stream Discrete HMMs on 3D skeletonsabstractHMMs are widely used in action and gesture recognition due to their implementation simplicity, low computational requirement, scalability and high parallelism. They have worth performance even with a limited training set. All these characteristics are hard to find together in other even more accurate methods. In this paper, we propose a novel double-stage classification approach, based on Multiple Stream Discrete Hidden Markov Models (MSD-HMM) and 3D skeleton joint data, able to reach high performances maintaining all advantages listed above. The approach allows both to quickly classify pre-segmented gestures (offline classification), and to perform temporal segmentation on streams of gestures (online classification) faster than real time. We test our system on three public datasets, MSRAction3D, UTKinect-Action and MSRDailyAction, and on a new dataset, Kinteract Dataset, explicitly created for Human Computer Interaction (HCI). We obtain state of the art performances on all of them. Guido Borghi, Roberto Vezzani, Rita Cucchiara |
ICPR | 2 |
| 2016 | YACCLAB - Yet Another Connected Components Labeling BenchmarkabstractThe problem of labeling the connected components (CCL) of a binary image is well-defined and several proposals have been presented in the past. Since an exact solution to the problem exists and should be mandatory provided as output, algorithms mainly differ on their execution speed. In this paper, we propose and describe YACCLAB, Yet Another Connected Components Labeling Benchmark. Together with a rich and varied dataset, YACCLAB contains an open source platform to test new proposals and to compare them with publicly available competitors. Textual and graphical outputs are automatically generated for three kinds of test, which analyze the methods from different perspectives. The fairness of the comparisons is guaranteed by running on the same system and over the same datasets. Examples of usage and the corresponding comparisons among state-of-the-art techniques are reported to confirm the potentiality of the benchmark. Costantino Grana, Federico Bolelli, Lorenzo Baraldi 0001, Roberto Vezzani |
ICPR | 4 |
| 2015 | Automatic configuration and calibration of modular sensing floorsabstractSensing floors are becoming an emerging solution for many privacy-compliant and large area surveillance systems. Many research and even commercial technologies have been proposed in the last years. Similarly to distributed camera networks, the problem of calibration is crucial, specially when installed in wide areas. This paper addresses the general problem of automatic calibration and configuration of modular and scalable sensing floors. Working on training data only, the system automatically finds the spatial placement of each sensor module and estimates threshold parameters needed for people detection. Tests on several training sequences captured with a commercial sensing floor are provided to validate the method. Roberto Vezzani, Martino Lombardi, Rita Cucchiara |
AVSS | 1 |
| 2015 | Mapping Appearance Descriptors on 3D Body Models for People Re-identification
Davide Baltieri, Roberto Vezzani, Rita Cucchiara |
Int. J. Comput. Vis. | 2 |
| 2015 | A General-Purpose Sensing Floor Architecture for Human-Environment InteractionabstractSmart environments are now designed as natural interfaces to capture and understand human behavior without a need for explicit human-computer interaction. In this article, we present a general-purpose architecture that acquires and understands human behaviors through a sensing floor. The pressure field generated by moving people is captured and analyzed. Specific actions and events are then detected by a low-level processing engine and sent to high-level interfaces providing different functions. The proposed architecture and sensors are modular, general-purpose, cheap, and suitable for both small- and large-area coverage. Some sample entertainment and virtual reality applications that we developed to test the platform are presented. Roberto Vezzani, Martino Lombardi, Augusto Pieracci, Paolo Santinelli, Rita Cucchiara |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2014 | 3D Hough transform for sphere recognition on point clouds - A systematic study and a new method proposal
Marco Camurri, Roberto Vezzani, Rita Cucchiara |
Mach. Vis. Appl. | 2 |
| 2014 | Detection of static groups and crowds gathered in open spaces by texture classification
Marco Manfredi, Roberto Vezzani, Simone Calderara, Rita Cucchiara |
Pattern Recognit. Lett. | 2 |
| 2013 | Sensing floors for privacy-compliant surveillance of wide areasabstractSurveillance systems can really benefit from the integration of multiple and heterogeneous sensors. In this paper we describe an innovative sensing floor. Thanks to its low cost and ease of installation, the floor is suitable for both private and public environments, from narrow zones to wide areas. The floor is made adding a sensing layer below commercial floating tiles. The sensor is scalable, reliable, and completely invisible to the users. The temporal and spatial resolutions of the data are high enough to identify the presence of people, to recognize their behavior and to detect events in a privacy compliant way. Experimental results on a real prototype implementation confirm the potentiality of the framework. Martino Lombardi, Augusto Pieracci, Paolo Santinelli, Roberto Vezzani, Rita Cucchiara |
AVSS | 4 |
| 2013 | Learning articulated body models for people re-identificationabstractPeople re-identification is a challenging problem in surveillance and forensics and it aims at associating multiple instances of the same person which have been acquired from different points of view and after a temporal gap. Image-based appearance features are usually adopted but, in addition to their intrinsically low discriminability, they are subject to perspective and view-point issues. We propose to completely change the approach by mapping local descriptors extracted from RGB-D sensors on a 3D body model for creating a view-independent signature. An original bone-wise color descriptor is generated and reduced with PCA to compute the person signature. The virtual bone set used to map appearance features is learned using a recursive splitting approach. Finally, people matching for re-identification is performed using the Relaxed Pairwise Metric Learning, which simultaneously provides feature reduction and weighting. Experiments on a specific dataset created with the Microsoft Kinect sensor and the OpenNi libraries prove the advantages of the proposed technique with respect to state of the art methods based on 2D or non-articulated 3D body models. Davide Baltieri, Roberto Vezzani, Rita Cucchiara |
ACM Multimedia | 2 |
| 2013 | Video surveillance online repository (ViSOR): www.openvisor.orgabstractThis paper describe the ViSOR (Video Surveillance Online Repository) repository, designed with the aim of establishing an open platform for collecting, annotating, retrieving, and sharing surveillance videos, as well as evaluating the performance of automatic surveillance systems. The repository is free and researchers can collaborate sharing their own videos or datasets. Most of the included videos are annotated. Annotations are based on a reference ontology which has been defined integrating hundreds of concepts, some of them coming from the LSCOM and MediaMill ontologies. A new annotation classification schema is also provided, which is aimed at identifying the spatial, temporal and domain detail level used. The web interface allows video browsing, querying by annotated concepts or by keywords, compressed video previewing, media downloading and uploading. Finally, ViSOR includes a performance evaluation desk which can be used to compare different annotations. Roberto Vezzani, Rita Cucchiara |
MMSys | 1 |
| 2013 | Editorial to the 'pattern recognition and artificial intelligence for human behaviour analysis' special sectionabstractThe Pattern Recognition (PR) and Artificial Intelligence (AI) scientific communities have shared knowledge and effort in order to obtain more effective solutions for many different research areas. However, although the techniques and approaches are somewhat similar, the two communities often tackle problems from rather different perspectives. In the first paper ‘Social Interactions by Visual Focus of Attention in a Three-Dimensional Environment', by Bazzani, Tosato, Cristani, Farenzena, Paggetti, Menegaz and Murino, a novel approach to social interaction discovery is presented; instead of using global or local appearance features, the authors exploit the Subjective View Frustum, which approximates the visual field of a person in a three-dimensional representation of the scene. The main contribution of the second paper 'Human action recognition using an ensemble of body-part detectors', by Chakraborty, Bagdanov, Gonzalez and Roca, is to transform the problem of action recognition into that of recognising the distinctive motion of specific body parts, for instance, the legs for walking, the hands for boxing, etc. The intuition behind the approach is that several human actions can be described more compactly and effectively by considering only the relevant motions of the body parts actually performing the actions. We hope you enjoy the special section. Luca Iocchi is Associate Professor at Sapienza University of Rome, Italy. His main research interests are in the areas of cognitive robotics, action planning, multi-robot coordination, robot perception, robot learning, sensor data fusion. He is being involved in several projects aiming at developing intelligent robotic systems and intelligent surveillance systems. He is active in many conferences and journals related to artificial intelligence and robotics, as well as in the organisation of scientific competitions, such as RoboCup@Home. Andrea Prati is Associate Professor at the University IUAV of Venice. He collaborated in several research projects at regional, national and international level. His research interests belong to different themes, from embedded devices for sensor networks in computer vision applications, to robotic vision, to multimedia, to performance analysis for multimedia computers. However, his main research activity is on video-surveillance topics: object tracking in distributed, multi-camera environments; analysis and removal of the shadows; behaviour analysis through trajectory classification. Andrea Prati is author of more than 130 papers in international journals and conference proceedings; he has been invited speaker and reviewer for many international journals. He is also a member of the Editorial Board of Journal of Optical Engineering (SPIE) and Journal on Ambient Intelligence and Smart Environments (IOS Press). He has also been the Program Chair of ICIAP 2007. He has been the PC of ACM/IEEE Intl Conf on Distributed Smart Cameras (ICDSC) in 2011 and 2012, and will be for 2013 edition in Palm Springs, CA (USA). He is also organising as General Chair the 2014 ICDSC edition in Venice. He is a senior member of IEEE, and a member of ACM and GIRPR. Roberto Vezzani is an Assistant Professor at University of Modena and Reggio Emilia and he works in the Engineering Department 'Enzo Ferrari'. His research interests mainly belong to video surveillance systems, with particular focus on behaviour analysis, people tracking and re-identification. He is the author of the ViSOR web repository, an online platform for sharing research videos and annotations developed within the European project VidiVideo. He was the technical coordinator of the European Project THIS, for transport hub intelligent video surveillance. He is author of more than 50 papers on international journals and conferences. Luca Iocchi, Andrea Prati 0001, Roberto Vezzani |
Expert Syst. J. Knowl. Eng. | 3 |
| 2012 | People Orientation Recognition by Mixtures of Wrapped Distributions on Random Trees
Davide Baltieri, Roberto Vezzani, Rita Cucchiara |
ECCV (5) | 2 |
| 2011 | Probabilistic people tracking with appearance models and occlusion classification: The AD-HOC system
Roberto Vezzani, Costantino Grana, Rita Cucchiara |
Pattern Recognit. Lett. | 1 |
| 2010 | Fast Background Initialization with Recursive Hadamard TransformabstractIn this paper, we present a new and fast technique for background estimation from cluttered image sequences. Most of the background initialization approaches developed so far collect a number of initial frames and then require a slow estimation step which introduces a delay whenever it is applied. Conversely, the proposed technique redistributes the computational load among all the frames by means of a patch by patch preprocessing, which makes the overall algorithm more suitable for real-time applications. For each patch location a prototype set is created and maintained. The background is then iteratively estimated by choosing from each set the most appropriate candidate patch, which should verify a sort of frequency coherence with its neighbors. To this aim, the Hadamard transform has been adopted which requires less computation time than the commonly used DCT Finally, a refinement step exploits spatial continuity constraints along the patch borders to prevent erroneous patch selections. The approach has been compared with the state of the art on videos from available datasets (ViSOR and CAVIAR), showing a speed up of about 10 times and an improved accuracy. Davide Baltieri, Roberto Vezzani, Rita Cucchiara |
AVSS | 2 |
| 2010 | Video Surveillance Online Repository (ViSOR): an integrated framework
Roberto Vezzani, Rita Cucchiara |
Multim. Tools Appl. | 1 |
| 2009 | An efficient Bayesian framework for on-line action recognitionabstractOn-line action recognition from a continuous stream of actions is still an open problem with fewer solutions proposed compared to time-segmented action recognition. The most challenging task is to classify the current action while finding its time boundaries at the same time. In this paper we propose an approach capable of performing on-line action segmentation and recognition by means of batteries of HMM taking into account all the possible time boundaries and action classes. A suitable Bayesian normalization is applied to make observation sequences of different length comparable and computational optimizations are introduce to achieve real-time performances. Results on a well known action dataset prove the efficacy of the proposed method. Roberto Vezzani, Massimo Piccardi, Rita Cucchiara |
ICIP | 1 |
| 2008 | Annotation Collection and Online Performance Evaluation for Video Surveillance: The ViSOR ProjectabstractThis paper presents the Visor (video surveillance online repository) project designed with the aim of establishing an open platform for collecting, annotating, retrieving, sharing surveillance videos, and of evaluating the performance of automatic surveillance systems. The main idea is to exploit the collaborative paradigm spreading in the web community to join together the ontology based annotation and retrieval concepts and the requirements of the computer vision and video surveillance communities. The ViSOR open repository is based on a reference ontology which integrates many concepts, also coming from LSCOM and MediaMill ontologies.The web interface allows video browse, query by annotated concepts or by keywords, compressed video preview, media download and upload. The repository contains metadata annotations, which can be either manually created as ground truth or automatically generated by video surveillance systems. Their automatic annotations can be compared each other or with the reference ground-truth exploiting an integrated on-line performance evaluator. Roberto Vezzani, Rita Cucchiara |
AVSS | 1 |
| 2008 | ViSOR: VIdeo Surveillance On-line Repository for annotation retrievalabstractThe Imagelab Laboratory of the University of Modena and Reggio Emilia has designed a large video repository, aiming at containing annotated video surveillance footages. The Web interface, named ViSOR (video surveillance online repository), allows video browse, query by annotated concepts or by keywords, compressed preview, video download and upload. The repository contains metadata annotation, both manually annotated ground-truth data and automatically obtained outputs of a particular system. In such a manner, the users of the repository are able to perform validation tasks of their own algorithms as well as comparative activities. Roberto Vezzani, Rita Cucchiara |
ICME | 1 |
| 2007 | A multi-camera vision system for fall detection and alarm generationabstractAbstract: In‐house video surveillance can represent an excellent support for people with some difficulties (e.g. elderly or disabled people) living alone and with a limited autonomy. New hardware technologies and in particular digital cameras are now affordable and they have recently gained credit as tools for (semi‐)automatically assuring people's safety. In this paper a multi‐camera vision system for detecting and tracking people and recognizing dangerous behaviours and events such as a fall is presented. In such a situation a suitable alarm can be sent, e.g. by means of an SMS. A novel technique of warping people's silhouette is proposed to exchange visual information between partially overlapped cameras whenever a camera handover occurs. Finally, a multi‐client and multi‐threaded transcoding video server delivers live video streams to operators/remote users in order to check the validity of a received alarm. Semantic and event‐based transcoding algorithms are used to optimize the bandwidth usage. A two‐room setup has been created in our laboratory to test the performance of the overall system and some of the results obtained are reported. Rita Cucchiara, Andrea Prati 0001, Roberto Vezzani |
Expert Syst. J. Knowl. Eng. | 3 |
| 2006 | 3-D Virtual Environments on Mobile Devices for Remote SurveillanceabstractIn this paper we present a distributed videosurveillance framework. Our end is the remote monitoring of the behavior of people moving in a scene exploiting a virtual reconstruction on low capabilities devices, like PDAs and cell phones. The main novelty of this system is the effective integration of the computer vision and computer graphics modules. The first, using a probabilistic frameworks, can detect the position, the trajectory and the posture of peoples moving in the scene. The second exploits the new possibility of both standard 3D graphics libraries on mobile (namely JSR184 and M3G graphic format) and new PDAs processing capability in order to reconstruct the remote surveillance data in real-time. Roberto Vezzani, Rita Cucchiara, Alessio Malizia, Luigi Cinque |
AVSS | 1 |
| 2006 | A Semi-Automatic Video Annotation tool with MPEG-7 Content CollectionsabstractIn this work, we present a general purpose system for hierarchical structural segmentation and automatic annotation of video clips, by means of standardized low level features. We propose to automatically extract some prototypes for each class with a context based intra-class clustering. Clips are annotated following the MPEG-7 standard directives to provide easier portability. Results of automatic annotation and semiautomatic metadata creation are provided Roberto Vezzani, Costantino Grana, Daniele Bulgarelli, Rita Cucchiara |
ISM | 1 |
| 2006 | PEANO: pictorial enriched annotation of videoabstractIn this DEMO, we present a tool set for video digital library management that allows i) structural annotation of edited videos in MPEG-7 by automatically extracting shots and clips; ii) automatic semantic annotation based on perceptual similarity against a taxonomy enriched with pictorial concepts iii) video clip access and hierarchical summarization with stand-alone and web interface iv) access to clips from mobile platform in GPRS-UMTS video-streaming. The tools can be applied in different domain-specific Video Digital Libraries. The main novelty is the possibility to enrich the annotation with pictorial concepts that are added to a textual taxonomy in order to make the automatic annotation process more fast and often effective. The resulting multimedia ontology is described in the MPEG-7 framework. The PEANO (Perceptual Annotation of Video) tool has been tested over video art , sport (Soccer, Olimpic Games 2006, Formula 1) and news clips. Costantino Grana, Roberto Vezzani, Daniele Bulgarelli, Giovanni Gualdi, Rita Cucchiara, Marco Bertini 0001, Carlo Torniai, Alberto Del Bimbo |
ACM Multimedia | 2 |
| 2006 | A system for automatic face obscuration for privacy purposes
Rita Cucchiara, Andrea Prati 0001, Roberto Vezzani |
Pattern Recognit. Lett. | 3 |
| 2005 | Entry edge of field of view for multi-camera tracking in distributed video surveillanceabstractEfficient solution to people tracking in distributed video surveillance is requested to monitor crowded and large environments. This paper proposes a novel use of the entry edges of field of view (E/sup 2/oFoV) to solve the consistent labeling problem between partially overlapped views. An automatic and reliable procedure allows obtaining the homographic transformation between two overlapped views, without any manual calibration of the cameras. Through the homography, the consistent labeling is established each time a new track is detected in one of the cameras. A camera transition graph (CTG) is defined to speed up the establishment process by reducing the search space. Experimental results prove the effectiveness of the proposed solution also in challenging conditions. Simone Calderara, Roberto Vezzani, Andrea Prati 0001, Rita Cucchiara |
AVSS | 2 |
| 2005 | Posture classification in a multi-camera indoor environmentabstractPosture classification is a key process for analyzing the people's behaviour. Computer vision techniques can be helpful in automating this process, but cluttered environments and consequent occlusions make this task often difficult. Different views provided by multiple cameras can be exploited to solve occlusions by warping known object appearance into the occluded view. To this aim, this paper describes an approach to posture classification based on projection histograms, reinforced by HMM for assuring temporal coherence of the posture. The single camera posture classification is then exploited in the multi-camera system to solve the cases in which the occlusions make the classification impossible. Experimental results of the classification from both the single camera and the multi-camera system are provided. Rita Cucchiara, Andrea Prati 0001, Roberto Vezzani |
ICIP (1) | 3 |
| 2005 | Probabilistic posture classification for Human-behavior analysisabstractComputer vision and ubiquitous multimedia access nowadays make feasible the development of a mostly automated system for human-behavior analysis. In this context, our proposal is to analyze human behaviors by classifying the posture of the monitored person and, consequently, detecting corresponding events and alarm situations, like a fall. To this aim, our approach can be divided in two phases: for each frame, the projection histograms (Haritaoglu et al., 1998) of each person are computed and compared with the probabilistic projection maps stored for each posture during the training phase; then, the obtained posture is further validated exploiting the information extracted by a tracking module in order to take into account the reliability of the classification of the first phase. Moreover, the tracking algorithm is used to handle occlusions, making the system particularly robust even in indoors environments. Extensive experimental results demonstrate a promising average accuracy of more than 95% in correctly classifying human postures, even in the case of challenging conditions. Rita Cucchiara, Costantino Grana, Andrea Prati 0001, Roberto Vezzani |
IEEE Trans. Syst. Man Cybern. Part A | 4 |
| 2003 | Object Segmentation in Videos from Moving Camera with MRFs on Color and Motion FeaturesabstractIn this paper we address the problem of fast segmenting moving objects in video acquired by moving camera or more generally with a moving background. We present an approach based on a color segmentation followed by a region-merging on motion through Markov random fields (MRFs). The technique we propose is inspired by the work of Gelgon and Bouthemy (2000), that has been modified to reduce computational cost in order to achieve a fast segmentation (about ten frame per second). To this aim a modified region matching algorithm (namely partitioned region matching) and an innovative arc-based MRF optimization algorithm with a suitable definition of the motion reliability are proposed. Results on both synthetic and real sequences are reported to confirm validity of our solution. Rita Cucchiara, Andrea Prati 0001, Roberto Vezzani |
CVPR (1) | 3 |