Guido Borghi

dblp:183/3554 · DBLP profile ↗
← Back
34ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0003-2441-7524ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 8 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 6 first-author · 16 since 2021Security and privacy · 6 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 GazeD: Context-Aware Diffusion for Accurate 3D Gaze Estimation
abstract
We introduce GazeD, a new 3D gaze estimation method that jointly provides 3D gaze and human pose from a single RGB image. Leveraging the ability of diffusion models to deal with uncertainty, it generates multiple plausible 3D gaze and pose hypotheses based on the 2D context information extracted from the input image. Specifically, we condition the denoising process on the 2D pose, the surroundings of the subject, and the context of the scene. With GazeD we also introduce a novel way of representing the 3D gaze by positioning it as an additional body joint at a fixed distance from the eyes. The rationale is that the gaze is usually closely related to the pose, and thus it can benefit from being jointly denoised during the diffusion process. Evaluations across three benchmark datasets demonstrate that GazeD achieves state-of-the-art performance in 3D gaze estimation, even surpassing methods that rely on temporal information. Project details will be available at https://aimagelab.ing.unimore.it/go/gazed
Riccardo Catalini, Davide Di Nucci, Guido Borghi, Davide Davoli 0002, Lorenzo Garattoni, Giampiero Francesca, Yuki Kawana, Roberto Vezzani
3DV3
2026 Quality-driven Adaptive Morphing Attack Detection in Operational Scenarios via Online Learning
abstract
Morphing Attack Detection (MAD) systems often suffer from performance degradation when deployed in operational environments, such as airports, that differ from the training domain.We propose an adaptive differential MAD framework that continuously refines a pre-trained detector using live bona fide samples acquired at the gate.The system is memoryless, so no samples are stored in memory to mitigate privacy concerns about the collection of personal data.To prevent the loss of discriminative power caused by bona fideonly adaptation, the method generates synthetic morph samples on-the-fly by combining the current operational subject with identities from an external public or synthetic face dataset.The adaptation process further relies on a quality-aware bona fide selection strategy and a controlled balancing mechanism for synthetic morph generation.Experimental results show that the proposed method improves target-domain specialization while maintaining robustness against morph attacks.
Nicolò Di Domenico, Annalisa Franco, Guido Borghi, Davide Maltoni
FG3
2026 Fake3DGS: A Benchmark for 3D Manipulation Detection in Neural Rendering
Davide Di Nucci, Riccardo Catalini, Guido Borghi, Roberto Vezzani
ICPR (5)3
2026 SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses
Alessandro Simoni, Riccardo Catalini, Davide Di Nucci, Guido Borghi, Davide Davoli 0002, Lorenzo Garattoni, Gianpiero Francesca, Yuki Kawana, Roberto Vezzani
ICPR (3)4
2026 Evaluating Age Estimation Robustness Under Realistic Facial Occlusions
Waqar Tanveer, Annalisa Franco, Guido Borghi, Laura Fernández-Robles, Eduardo Fidalgo
ICPR (13)3
2025 BioGaze: a Framework for Evaluating the Photographic Requirements of the ISO/IEC 39794-5 Standard
abstract
Facial recognition is a key biometric technology, especially for using electronic documents in real-world applications. The accuracy of this recognition technology strictly depends on the image quality, i.e. the face appearance in the image included in the document. Then, adherence to ISO/ICAO standards, which contain guidelines to standardize the image quality in official documents, is of paramount importance. However, ensuring compliance is challenging due to high subject variability. Furthermore, controls are often executed manually, making them subjective and time-consuming. Therefore, in this work, we introduce BioGaze, an automated framework for ISO/ICAO compliance verification that combines classical computer vision and deep learning algorithms to perform the checks contained in the latest standard version. The framework is tested on a synthetic dataset, achieving state-of-the-art performance across multiple ISO/ICAO requirements, surpassing public algorithms and commercial SDKs. BioGaze is publicly available to advance automated compliance verification and support standardization efforts1.1https://github.com/MI-BioLab/BioGaze
Osama Elatfi, Nicolò Di Domenico, Guido Borghi, Annalisa Franco, Davide Maltoni
FG3
2025 Towards Zero-Shot ISO/ICAO Face Compliance Verification via CLIP-IQA and Natural Language Prompting
abstract
Ensuring compliance of face images with ISO/ICAO quality standards is essential for boosting the document enrollment process. Indeed, traditional manual checks are slow, subjective, and difficult to scale. Therefore, we propose a system that aims to fully automate compliance verification by directly analyzing the official requirements without relying on predefined hand-crafted features or manual thresholds. Our method combines a Large Language Model, a novel prompt learning procedure, and a contrastive learning framework to evaluate the adherence of a face image to quality requirements. Tested on a recent dataset, our proposed system achieves high accuracy, surpassing existing academic and commercial solutions. By streamlining the implementation and updates to the compliance rules, our approach represents a significant step toward simple, scalable, and regulation-driven image verification. Code and models are publicly available1.
Nicolò Di Domenico, Guido Borghi, Annalisa Franco, Davide Maltoni
IJCB2
2025 Adversarial Attack Challenge for Secure Face Recognition 2025
abstract
Adversarial attacks pose a significant threat to the reliability of biometric systems, particularly in security-critical applications such as identity verification and access control. Ensuring robustness against such attacks is essential for the safe deployment of face recognition technologies in real-world scenarios. To advance this goal, the 2025 Adversarial Attack Challenge for Secure Face Recognition was organized as part of the International Joint Conference on Biometrics (IJCB) 2025.The competition focused on two main tracks: Detection, where the objective was to determine whether a given face image is clean or adversarial, and Resilience, which aimed to evaluate recognition systems under adversarial perturbations. Participants were provided with a standardized dataset derived from CelebA and LFW, encompassing both clean samples and adversarial images crafted using ten diverse attack methods targeting evasion and impersonation scenarios. To ensure fairness and reproducibility, all models were trained solely on the data provided, with support from a custom open source adversarial attack package tailored for face recognition.In addition to benchmarking adversarial robustness, the challenge contributes to the research community by releasing the data set and the extensible attack package, allowing further investigation of secure and reliable face recognition systems.
João Tremoço, Iurii Medvedev, Nuno R. Freitas, Andreia M. Costa, Diogo Nunes, Niklas Bunzel, Lukas Graner, Nicholas Göller, Lorenzo Pellegrini, Nicolò Di Domenico, Guido Borghi, Monson Verghese, Shruti Bhilare, Avik Hati, Miguel Lourenço, Nuno Gonçalves 0001
IJCB11
2025 BRUM: Robust 3D Vehicle Reconstruction from 360° Sparse Images
abstract
Accurate 3D reconstruction of vehicles is vital for applications such as vehicle inspection, predictive maintenance, and urban planning. Existing methods like Neural Radiance Fields and Gaussian Splatting have shown impressive results but remain limited by their reliance on dense input views, which hinders real-world applicability. This paper addresses the challenge of reconstructing vehicles from sparse-view inputs, leveraging depth maps and a robust pose estimation architecture to synthesize novel views and augment training data. Specifically, we enhance Gaussian Splatting by integrating a selective photometric loss, applied only to high-confidence pixels, and replacing standard Structure-from-Motion pipelines with the DUSt3R architecture to improve camera pose estimation. Furthermore, we present a novel dataset featuring both synthetic and real-world public transportation vehicles, enabling extensive evaluation of our approach. Experimental results demonstrate state-of-the-art performance across multiple benchmarks, showcasing the method's ability to achieve high-quality reconstructions even under constrained input conditions. Code and data are publicly available at https://aimagelab.ing.unimore.it/go/brum.
Davide Di Nucci, Matteo Tomei, Guido Borghi, Luca Ciuffreda, Roberto Vezzani, Rita Cucchiara
IV3
2025 3D Pose Nowcasting: Forecast the future to improve the present
abstract
Technologies to enable safe and effective collaboration and coexistence between humans and robots have gained significant importance in the last few years. A critical component useful for realizing this collaborative paradigm is the understanding of human and robot 3D poses using non-invasive systems. Therefore, in this paper, we propose a novel vision-based system leveraging depth data to accurately establish the 3D locations of skeleton joints. Specifically, we introduce the concept of Pose Nowcasting, denoting the capability of the proposed system to enhance its current pose estimation accuracy by jointly learning to forecast future poses. The experimental evaluation is conducted on two different datasets, providing accurate and real-time performance and confirming the validity of the proposed method on both the robotic and human scenarios. • We introduce the novel task of 3D Pose Nowcasting. • Our Pose Nowcasting system is based on both 3D Pose Estimation and Forecasting. • We show that knowledge about pose forecasting improves the accuracy of pose estimation. • We apply the proposed system both to human and robots. • Result on different dataset show state-of-the-art performance and robustness.
Alessandro Simoni, Francesco Marchetti, Guido Borghi, Federico Becattini, Lorenzo Seidenari, Roberto Vezzani, Alberto Del Bimbo
Comput. Vis. Image Underst.3
2025 Towards on-device continual learning with Binary Neural Networks in industrial scenarios
Lorenzo Vorabbi, Angelo Carraggi, Davide Maltoni, Guido Borghi, Stefano Santi
Image Vis. Comput.4
2024 ONOT: a High-Quality ICAO-compliant Synthetic Mugshot Dataset
abstract
Nowadays, state-of-the-art AI-based generative models represent a viable solution to overcome privacy issues and biases in the collection of datasets containing personal information, such as faces. Following this intuition, in this paper we introduce ONOT11One, No one and One hundred Thousand (L. Pirandello, 1926), a synthetic dataset specifically focused on the generation of high-quality faces in adherence to the requirements of the ISO/IEC 39794–5 standards that, following the guidelines of the International Civil Aviation Organization (ICAO), defines the interchange formats of face images in electronic Machine-Readable Travel Documents (eMRTD). The strictly controlled and varied mugshot images included in ONOT are useful in research fields related to the analysis of face images in eMRTD, such as Morphing Attack Detection and Face Quality Assessment. The dataset is publicly released22https://miatbiolab.csr.unibo.it/icao-synthetic-dataset, in combination with the generation procedure details in order to improve the reproducibility and enable future extensions.
Nicolò Di Domenico, Guido Borghi, Annalisa Franco, Davide Maltoni
FG2
2024 SDFR: Synthetic Data for Face Recognition Competition
abstract
Large-scale face recognition datasets are collected by crawling the Internet and without individuals' consent, raising legal, ethical, and privacy concerns. With the recent advances in generative models, recently several works proposed generating synthetic face recognition datasets to mitigate concerns in web-crawled face recognition datasets. This paper presents the summary of the Synthetic Data for Face Recognition (SDFR) Competition held in conjunction with the 18th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2024) and established to investigate the use of synthetic data for training face recognition models. The SDFR competition was split into two tasks, allowing participants to train face recognition systems using new synthetic datasets and/or existing ones. In the first task, the face recognition backbone was fixed and the dataset size was limited, while the second task provided almost complete freedom on the model backbone, the dataset, and the training pipeline. The submitted models were trained on existing and also new synthetic datasets and used clever methods to improve training with synthetic data. The submissions were evaluated and ranked on a diverse set of seven benchmarking datasets. The paper gives an overview of the submitted face recognition models and reports achieved performance compared to baseline models trained on real and synthetic datasets. Furthermore, the evaluation of submissions is extended to bias assessment across different demography groups. Lastly, an outlook on the current state of the research in training face recognition models using synthetic data is presented, and existing problems as well as potential future directions are also discussed.
Hatef Otroshi-Shahreza, Christophe Ecabert, Anjith George, Alexander Unnervik, Sébastien Marcel, Nicolò Di Domenico, Guido Borghi, Davide Maltoni, Fadi Boutros, Julia Vogel, Naser Damer, Ángela Sánchez-Pérez, Enrique Mas-Candela, Jorge Calvo-Zaragoza, Bernardo Biesseck, Pedro Vidal 0001, Roger Granada, David Menotti, Ivan DeAndres-Tame, Simone Maurizio La Cava, Sara Concas, Pietro Melzi, Ruben Tolosana, Rubén Vera-Rodríguez, Gianpaolo Perelli, Giulia Orrù, Gian Luca Marcialis, Julian Fierrez
FG7
2024 V-MAD: Video-based Morphing Attack Detection in Operational Scenarios
abstract
In response to the rising threat of the face morphing attack, this paper introduces and explores the potential of Video-based Morphing Attack Detection (V-MAD) systems in real-world operational scenarios. While current morphing attack detection methods primarily focus on a single or a pair of images, V-MAD is based on video sequences, exploiting the video streams acquired by face verification tools available, for instance, at airport gates. We show for the first time the advantages that the availability of multiple probe frames brings to the morphing attack detection task, especially in scenarios where the quality of probe images is varied. Experimental results on a real operational database demonstrate that video sequences represent valuable information for increasing the performance of morphing attack detection systems.
Guido Borghi, Annalisa Franco, Nicolò Di Domenico, Matteo Ferrara, Davide Maltoni
IJCB1
2024 Towards Federated Learning for Morphing Attack Detection
abstract
Through the Face Morphing attack is possible to use the same legal document by two different people, destroying the unique biometric link between the document and its owner. In other words, a morphed face image has the potential to bypass face verification-based security controls, then representing a severe security threat. Unfortunately, the lack of public, extensive and varied training datasets severely hampers the development of effective and robust Morphing Attack Detection (MAD) models, key tools in contrasting the Face Morphing attack since able to automatically detect the presence of morphing images. Indeed, privacy regulations limit the possibility of acquiring, storing, and transferring MAD-related data that contain personal information, such as faces. Therefore, in this paper, we investigate the use of Federated Learning to train a MAD model on local training samples across multiple sites, eliminating the need for a single centralized training dataset, as common in Machine Learning, and then overcoming privacy limitations. Experimental results suggest that FL is a viable solution that will need to be considered in future research works in MAD.
Marta Robledo-Moreno, Guido Borghi, Nicolò Di Domenico, Annalisa Franco, Kiran B. Raja, Davide Maltoni
IJCB2
2023 Detecting Morphing Attacks via Continual Incremental Training
abstract
Scenarios in which restrictions in data transfer and storage limit the possibility to compose a single dataset – also exploiting different data sources – to perform a batch-based training procedure, make the development of robust models particularly challenging. We hypothesize that the recent Continual Learning (CL) paradigm may represent an effective solution to enable incremental training, even through multiple sites. Indeed, a basic assumption of CL is that once a model has been trained, old data can no longer be used in successive training iterations and in principle can be deleted. Therefore, in this paper, we investigate the performance of different Continual Learning methods in this scenario, simulating a learning model that is updated every time a new chunk of data, even of variable size, is available. Experimental results reveal that a particular CL method, namely Learning without Forgetting (LwF), is one of the best-performing algorithms. Then, we investigate its usage and parametrization in Morphing Attack Detection and Object Classification tasks, specifically with respect to the amount of new training data that became available.
Lorenzo Pellegrini, Guido Borghi, Annalisa Franco, Davide Maltoni
IJCB2
2023 Depth-based 3D human pose refinement: Evaluating the refinet framework
abstract
In recent years, Human Pose Estimation has achieved impressive results on RGB images. The advent of deep learning architectures and large annotated datasets have contributed to these achievements. However, little has been done towards estimating the human pose using depth maps, and especially towards obtaining a precise 3D body joint localization. To fill this gap, this paper presents RefiNet, a depth-based 3D human pose refinement framework. Given a depth map and an initial coarse 2D human pose, RefiNet regresses a fine 3D pose. The framework is composed of three modules, based on different data representations, i.e. 2D depth patches, 3D human skeletons, and point clouds. An extensive experimental evaluation is carried out to investigate the impact of the model hyper-parameters and to compare RefiNet with off-the-shelf 2D methods and literature approaches. Results confirm the effectiveness of the proposed framework and its limited computational requirements.
Andrea D'Eusanio, Alessandro Simoni, Stefano Pini, Guido Borghi, Roberto Vezzani, Rita Cucchiara
Pattern Recognit. Lett.4
2022 Incremental Training of Face Morphing Detectors
abstract
Recently, the Face Morphing Attack has emerged as a serious and concrete security threat, and then several Morphing Attack Detectors (MADs) have been proposed in the literature. Unfortunately, most MAD approaches, especially if based on single input images, are not yet mature for real-world deployment, mainly due to low accuracy and limited generalization capabilities with data distribution different from those used for training. While better models and techniques will certainly contribute to advancing the state-of-the-art, one of the main obstacles is on the data side: indeed, training data available to research groups are often limited in size and variety, and privacy issues restrict the possibility of data release and exchange. The proposed approach envisages the incremental training of MADs on a collection of datasets even owned by different research groups, adopting a learning strategy that involves model transfer as an alternative to data sharing. Specifically, in this paper, we propose and publicly release a framework to support the adoption of Continual Learning strategies, enabling an efficient incremental training for MADs on new data that progressively become available at different places or times. Different learning strategies are analyzed and compared in terms of the effectiveness and stability of the results achieved.
Guido Borghi, Gabriele Graffieti, Annalisa Franco, Davide Maltoni
ICPR1
2022 SHREC 2022 track on online detection of heterogeneous gestures
Marco Emporio, Ariel Caputo, Andrea Giachetti 0001, Marco Cristani, Guido Borghi, Andrea D'Eusanio, Minh-Quan Le, Hai-Dang Nguyen, Minh-Triet Tran, Felix Ambellan, Martin Hanik, Esfandiar Nava-Yazdani, Christoph von Tycowicz
Comput. Graph.5
2021 SHREC 2021: Skeleton-based hand gesture recognition in the wild
Ariel Caputo, Andrea Giachetti 0001, Simone Soso, Deborah Pintani, Andrea D'Eusanio, Stefano Pini, Guido Borghi, Alessandro Simoni, Roberto Vezzani, Rita Cucchiara, Andrea Ranieri, Franca Giannini, Katia Lupinetti, Marina Monti, Mehran Maghoumi, Joseph J. LaViola Jr., Minh-Quan Le, Hai-Dang Nguyen, Minh-Triet Tran
Comput. Graph.7
2021 Video Frame Synthesis Combining Conventional and Event Cameras
abstract
Event cameras are biologically-inspired sensors that gather the temporal evolution of the scene. They capture pixel-wise brightness variations and output a corresponding stream of asynchronous events. Despite having multiple advantages with respect to conventional cameras, their use is limited due to the scarce compatibility of asynchronous event streams with traditional data processing and vision algorithms. In this regard, we present a framework that synthesizes RGB frames from the output stream of an event camera and an initial or a periodic set of color key-frames. The deep learning-based frame synthesis framework consists of an adversarial image-to-image architecture and a recurrent module. Two public event-based datasets, DDD17 and MVSEC, are used to obtain qualitative and quantitative per-pixel and perceptual results. In addition, we converted into event frames two additional well-known datasets, namely Kitti and Cityscapes, in order to present semantic results, in terms of object detection and semantic segmentation accuracy. Extensive experimental evaluation confirms the quality and the capability of the proposed approach of synthesizing frame sequences from color key-frames and sequences of intermediate events.
Stefano Pini, Guido Borghi, Roberto Vezzani
Int. J. Pattern Recognit. Artif. Intell.2
2020 A Transformer-Based Network for Dynamic Hand Gesture Recognition
abstract
Transformer-based neural networks represent a successful self-attention mechanism that achieves state-of-the-art results in language understanding and sequence modeling. However, their application to visual data and, in particular, to the dynamic hand gesture recognition task has not yet been deeply investigated. In this paper, we propose a transformer-based architecture for the dynamic hand gesture recognition task. We show that the employment of a single active depth sensor, specifically the usage of depth maps and the surface normals estimated from them, achieves state-of-the-art results, overcoming all the methods available in the literature on two automotive datasets, namely NVidia Dynamic Hand Gesture and Briareo. Moreover, we test the method with other data types available with common RGB-D devices, such as infrared and color data. We also assess the performance in terms of inference time and number of parameters, showing that the proposed framework is suitable for an online in-car infotainment system.
Andrea D'Eusanio, Alessandro Simoni, Stefano Pini, Guido Borghi, Roberto Vezzani, Rita Cucchiara
3DV4
2020 Baracca: a Multimodal Dataset for Anthropometric Measurements in Automotive
abstract
The recent spread of depth sensors has enabled new methods to automatically estimate anthropometric measurements, in place of manual procedures or expensive 3D scanners. Generally, the use of depth data is limited by the lack of depth-based public datasets containing accurate anthropometric annotations. Therefore, in this paper we propose a new dataset, called Baracca, specifically designed for the automotive context, including in-car and outside views. The dataset is multimodal: it has been acquired with synchronized depth, infrared, thermal and RGB cameras in order to deal with the requirements imposed by the automotive context. In addition, we propose several baselines to test the challenges of the presented dataset and provide considerations for future work.
Stefano Pini, Andrea D'Eusanio, Guido Borghi, Roberto Vezzani, Rita Cucchiara
IJCB3
2020 RefiNet: 3D Human Pose Refinement with Depth Maps
abstract
Human Pose Estimation is a fundamental task for many applications in the Computer Vision community and it has been widely investigated in the 2D domain, i.e. intensity images. Therefore, most of the available methods for this task are mainly based on 2D Convolutional Neural Networks and huge manually-annotated RGB datasets, achieving stunning results. In this paper, we propose RefiNet, a multi-stage framework that regresses an extremely-precise 3D human pose estimation from a given 2D pose and a depth map. The framework consists of three different modules, each one specialized in a particular refinement and data representation, i.e. depth patches, 3D skeleton and point clouds. Moreover, we present a new dataset, called Baracca, acquired with RGB, depth and thermal cameras and specifically created for the automotive context. Experimental results confirm the quality of the refinement procedure that largely improves the human pose estimations of off-the-shelf 2D methods.
Andrea D'Eusanio, Stefano Pini, Guido Borghi, Roberto Vezzani, Rita Cucchiara
ICPR3
2020 Anomaly Detection, Localization and Classification for Railway Inspection
abstract
The ability to detect, localize and classify objects that are anomalies is a challenging task in the computer vision community. In this paper, we tackle these tasks developing a framework to automatically inspect the railway during the night. Specifically, it is able to predict the presence, the image coordinates and the class of obstacles. To deal with the low-light environment, the framework is based on thermal images and consists of three different modules that address the problem of detecting anomalies, predicting their image coordinates and classifying them. Moreover, due to the absolute lack of publicly-released datasets collected in the railway context for anomaly detection, we introduce a new multi-modal dataset, acquired from a rail drone, used to evaluate the proposed framework. Experimental results confirm the accuracy of the framework and its suitability, in terms of computational load, performance, and inference time, to be implemented on a self-powered inspection system.
Riccardo Gasparini, Andrea D'Eusanio, Guido Borghi, Stefano Pini, Giuseppe Scaglione, Simone Calderara, Eugenio Fedeli, Rita Cucchiara
ICPR3
2020 Face-from-Depth for Head Pose Estimation on Depth Images
abstract
Depth cameras allow to set up reliable solutions for people monitoring and behavior understanding, especially when unstable or poor illumination conditions make unusable common RGB sensors. Therefore, we propose a complete framework for the estimation of the head and shoulder pose based on depth images only. A head detection and localization module is also included, in order to develop a complete end-to-end system. The core element of the framework is a Convolutional Neural Network, called POSEidon+, that receives as input three types of images and provides the 3D angles of the pose as output. Moreover, a Face-from-Depth component based on a Deterministic Conditional GAN model is able to hallucinate a face from the corresponding depth image. We empirically demonstrate that this positively impacts the system performances. We test the proposed framework on two public datasets, namely Biwi Kinect Head Pose and ICT-3DHP, and on Pandora, a new challenging dataset mainly inspired by the automotive setup. Experimental results show that our method overcomes several recent state-of-art works based on both intensity and depth input data, running in real-time at more than 30 frames per second.
Guido Borghi, Matteo Fabbri, Roberto Vezzani, Simone Calderara, Rita Cucchiara
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Learning to Generate Facial Depth Maps
abstract
In this paper, an adversarial architecture for facial depth map estimation from monocular intensity images is presented. By following an image-to-image approach, we combine the advantages of supervised learning and adversarial training, proposing a conditional Generative Adversarial Network that effectively learns to translate intensity face images into the corresponding depth maps. Two public datasets, namely Biwi database and Pandora dataset, are exploited to demonstrate that the proposed model generates high-quality synthetic depth images, both in terms of visual appearance and informative content. Furthermore, we show that the model is capable of predicting distinctive facial details by testing the generated depth maps through a deep model trained on authentic depth maps for the face verification task.
Stefano Pini, Filippo Grazioli, Guido Borghi, Roberto Vezzani, Rita Cucchiara
3DV3
2018 Face Verification from Depth using Privileged Information
Guido Borghi, Stefano Pini, Filippo Grazioli, Roberto Vezzani, Rita Cucchiara
BMVC1
2018 Hands on the wheel: A Dataset for Driver Hand Detection and Tracking
abstract
The ability to detect, localize and track the hands is crucial in many applications requiring the understanding of the person behavior, attitude and interactions. In particular, this is true for the automotive context, in which hand analysis allows to predict preparatory movements for maneuvers or to investigate the driver's attention level. Moreover, due to the recent diffusion of cameras inside new car cockpits, it is feasible to use hand gestures to develop new Human-Car Interaction systems, more user-friendly and safe. In this paper, we propose a new dataset, called Turms, that consists of infrared images of driver's hands, collected from the back of the steering wheel, an innovative point of view. The Leap Motion device has been selected for the recordings, thanks to its stereo capabilities and the wide view-angle. Besides, we introduce a method to detect the presence and the location of driver's hands on the steering wheel, during driving activity tasks.
Guido Borghi, Elia Frigieri, Roberto Vezzani, Rita Cucchiara
FG1
2018 Fully Convolutional Network for Head Detection with Depth Images
abstract
Head detection and localization are one of the most investigated and demanding tasks of the Computer Vision community. These are also a key element for many disciplines, like Human Computer Interaction, Human Behavior Understanding, Face Analysis and Video Surveillance. In last decades, many efforts have been conducted to develop accurate and reliable head or face detectors on standard RGB images, but only few solutions concern other types of images, such as depth maps. In this paper, we propose a novel method for head detection on depth images, based on a deep learning approach. In particular, the presented system overcomes the classic sliding-window approach, that is often the main computational bottleneck of many object detectors, through a Fully Convolutional Network. Two public datasets, namely Pandora and Watch-n-Patch, are exploited to train and test the proposed network. Experimental results confirm the effectiveness of the method, that is able to exceed all the state-of-art works based on depth images and to run with real time performance.
Diego Ballotta, Guido Borghi, Roberto Vezzani, Rita Cucchiara
ICPR2
2018 Domain Translation with Conditional GANs: from Depth to RGB Face-to-Face
abstract
Can faces acquired by low-cost depth sensors be useful to catch some characteristic details of the face? Typically the answer is no. However, new deep architectures can generate RGB images from data acquired in a different modality, such as depth data. In this paper, we propose a new Deterministic Conditional GAN, trained on annotated RGB-D face datasets, effective for a face-to-face translation from depth to RGB. Although the network cannot reconstruct the exact somatic features for unknown individual faces, it is capable to reconstruct plausible faces; their appearance is accurate enough to be used in many pattern recognition tasks. In fact, we test the network capability to hallucinate with some Perceptual Probes, as for instance face aspect classification or landmark detection. Depth face can be used in spite of the correspondent RGB images, that often are not available due to difficult luminance conditions. Experimental results are very promising and are as far as better than previously proposed approaches: this domain translation can constitute a new way to exploit depth data in new future applications.
Matteo Fabbri, Guido Borghi, Fabio Lanzi, Roberto Vezzani, Simone Calderara, Rita Cucchiara
ICPR2
2017 POSEidon: Face-from-Depth for Driver Pose Estimation
abstract
Fast and accurate upper-body and head pose estimation is a key task for automatic monitoring of driver attention, a challenging context characterized by severe illumination changes, occlusions and extreme poses. In this work, we present a new deep learning framework for head localization and pose estimation on depth images. The core of the proposal is a regressive neural network, called POSEidon, which is composed of three independent convolutional nets followed by a fusion layer, specially conceived for understanding the pose by depth. In addition, to recover the intrinsic value of face appearance for understanding head position and orientation, we propose a new Face-from-Depth model for learning image faces from depth. Results in face reconstruction are qualitatively impressive. We test the proposed framework on two public datasets, namely Biwi Kinect Head Pose and ICT-3DHP, and on Pandora, a new challenging dataset mainly inspired by the automotive setup. Results show that our method overcomes all recent state-of-art works, running in real time at more than 30 frames per second.
Guido Borghi, Marco Venturelli, Roberto Vezzani, Rita Cucchiara
CVPR1
2017 Embedded recurrent network for head pose estimation in car
abstract
An accurate and fast driver's head pose estimation is a rich source of information, in particular in the automotive context. Head pose is a key element for driver's behavior investigation, pose analysis, attention monitoring and also a useful component to improve the efficacy of Human-Car Interaction systems. In this paper, a Recurrent Neural Network is exploited to tackle the problem of driver head pose estimation, directly and only working on depth images to be more reliable in presence of varying or insufficient illumination. Experimental results, obtained from two public dataset, namely Biwi Kinect Head Pose and ICT-3DHP Database, prove the efficacy of the proposed method that overcomes state-of-art works. Besides, the entire system is implemented and tested on two embedded boards with real time performance.
Guido Borghi, Riccardo Gasparini, Roberto Vezzani, Rita Cucchiara
Intelligent Vehicles Symposium1
2016 Fast gesture recognition with Multiple Stream Discrete HMMs on 3D skeletons
abstract
HMMs are widely used in action and gesture recognition due to their implementation simplicity, low computational requirement, scalability and high parallelism. They have worth performance even with a limited training set. All these characteristics are hard to find together in other even more accurate methods. In this paper, we propose a novel double-stage classification approach, based on Multiple Stream Discrete Hidden Markov Models (MSD-HMM) and 3D skeleton joint data, able to reach high performances maintaining all advantages listed above. The approach allows both to quickly classify pre-segmented gestures (offline classification), and to perform temporal segmentation on streams of gestures (online classification) faster than real time. We test our system on three public datasets, MSRAction3D, UTKinect-Action and MSRDailyAction, and on a new dataset, Kinteract Dataset, explicitly created for Human Computer Interaction (HCI). We obtain state of the art performances on all of them.
Guido Borghi, Roberto Vezzani, Rita Cucchiara
ICPR1