Daniele Bonatto

dblp:193/6412 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0001-6502-5354ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Structure-from-motion in micro-image domain for uncalibrated plenoptic 2.0 cameras
abstract
We introduce a structure-from-motion method specifically designed to process the raw micro-images captured by plenoptic 2.0 cameras. Unlike traditional monocular cameras, plenoptic cameras incorporate a micro-lens array between the sensor and the main lens, capturing depth information at the expense of a more complex set of parameters to evaluate. Instead of simply integrating their projection model into the classical structure-from-motion pipeline, our contribution identifies the pinhole cameras-driven constraints and takes advantage of the inherent disparity information present in plenoptic cameras. This facilitates a robust initialization of the reconstruction. Our method shortcuts two of the limitations of the classical structure-from-motion: the ambiguity found in scenes captured with low angular disparity and the scale ambiguity. It enables the reconstruction of scenes captured by multiple uncalibrated plenoptic cameras, without using any calibration pattern or subaperture view extraction step. Our method undergoes experimental validation on both natural and synthetic datasets, showing a 10% error accuracy for relative pose estimation, which is comparable to calibration-pattern based methods. The results are robust to coarse initialization. Contrary to classical structure-from-motion, it is able to reconstruct scenes with parallel facing cameras. It also shows greater accuracy than reconstruction methods based on pinhole camera conversion.
Sarah Dury, Daniele Bonatto, Jaime Sancho, Eduardo Juárez Martínez, Mehrdad Teratani, Gauthier Lafruit
Int. J. Comput. Vis.2
2025 Synchronization and Calibration of Video Sequences Acquired Using Multiple Plenoptic 2.0 Cameras
Daniele Bonatto, Sarah Fachada, Jaime Sancho, Eduardo Juárez Martínez, Gauthier Lafruit, Mehrdad Teratani
MMM (4)1
2025 PaintUnicorn: Multi-view video dataset
abstract
We present PaintUnicorn, a real-world multiview video dataset designed for advancing research in dynamic neural rendering, 4D reconstruction, and view synthesis. Captured with 28 synchronized cameras at 1080p and 30 fps over 300 frames, the dataset includes complex non-rigid motion, specular and transparent surfaces, and varying lighting conditions. We provide calibrated camera intrinsics and extrinsics, color-corrected images, and generated depth maps combined from multiview geometry and learning-based estimates. PaintUnicorn is a dataset to support dense temporal and spatial benchmarking for methods like 3D Gaussian Splatting and dynamic NeRFs. The dataset is accessible in https://doi.org/10.5281/zenodo.16312256
Eva Dubar, Daniele Bonatto, Sarah Dury, Gauthier Lafruit
VCIP2
2025 Micro-Image Domain View Synthesizer for Free Navigation With Focused Plenoptic Cameras
abstract
We present a novel, first-of-its-kind view synthesis method for plenoptic images, which enables the direct manipulation of images in the micro-images array format, thereby bypassing intermediate transformation steps. Current plenoptic imaging approaches typically rely on an initial conversion to dense multiview images, also known as subaperture images extraction. However, the use of subaperture images presents two main limitations that ultimately impact further processing. First, existing subaperture view extraction methods offer limited control over camera parameters, resolutions, and poses of the subaperture views, which are also constrained to a small area around the main lens, thus restricting free navigation. Second, subaperture images are susceptible to artifacts which can propagate to subsequent processes such as calibration, depth estimation and view synthesis. In this paper, we propose a camera model that enables depth image-based rendering with plenoptic cameras, in a way that allows for the direct synthesis of any target viewpoint. In our evaluation, we show that our method expands view synthesis extrapolation to a range that is two to three times greater than that of pipelines requiring a conversion to subaperture images, including generally accepted tools such as depth image-based rendering and learning-based rendering approaches.
Sarah Fachada, Daniele Bonatto, Gauthier Lafruit, Mehrdad Teratani
IEEE Trans. Multim.2
2024 A Practical Approach to Depth-Aware Augmentation for Neural Radiance Fields
abstract
Neural Radiance Fields (NeRF) have demonstrated exceptional performance in generating novel views of scenes by learning implicit volumetric representations from calibrated RGB images, without depth information. A major limitation is the need for large training datasets in neural network-based view synthesis frameworks. The challenge of effective data augmentation for view synthesis remains unresolved. NeRF models require extensive scene coverage from multiple views to accurately estimate radiance and density. Insufficient coverage reduces the model’s ability to interpolate or extrapolate unseen parts of the scene effectively. In this paper, we propose a novel pipeline to address this data augmentation issue using depth map information. We use depth image-based rendering (DIBR) to overcome the lack of enough views for training NeRF. Experimental results indicate that our approach enhances the quality of rendered images using the NeRF framework, achieving an average peak signal-to-noise ratio (PSNR) increase of 7.2 dB, with a maximum improvement of 12 dB.
Hamed Razavi Khosroshahi, Jaime Sancho, Daniele Bonatto, Sarah Fachada, Gun Bang, Gauthier Lafruit, Eduardo Juárez Martínez, Mehrdad Teratani
VCIP3
2022 Pattern-free Plenoptic 2.0 Camera Calibration
abstract
The plenoptic 2.0 camera is a light field acquisition system consisting of a main lens and a micro-lens array (MLA) at a non-focal distance of the main lens. While it allows to retrieve the geometry of the scene, the distances between the main lens, the MLA and the sensor are usually unknown. Therefore, the use cases for plenoptic cameras stay limited while they have more potential applications such as virtual reality, provided that their camera parameters are precisely known. In this paper, we present a pattern-free calibration method to retrieve the plenoptic camera's intrinsic parameters from the rendered subaperture images, i.e. the distance from main lens to the MLA and sensor, the focal length and the principal point. The proposed method utilises the relation between scene's distances, i.e. depth, and the subaperture images' micro-disparities, which contains information about the camera parameters. To the best of our knowledge, it is the first pattern-free calibration method for plenoptic 2.0 cameras. We compare the parameters obtained using our method applied to subaperture images rendered from three different software tools (Plenoptic Toolbox, Reference Lenslet content Convertor, and Lenslet to Multiview) with an open-source pattern-based method (Compote). The proposed pattern-free calibration method has consistent camera parameters with the traditional pattern-based calibration method. We therefore reliably obtain the intrinsic parameters of any plenoptic camera from its captured images, even in the absence of any calibration pattern.
Sarah Fachada, Daniele Bonatto, Armand Losfeld, Gauthier Lafruit, Mehrdad Teratani
MMSP2
2022 3D Tensor Display for Non-Lambertian Content
abstract
A tensor display is a type of 3D light field display, composed of multiple transparent screens and a back-light that can render a scene with correct depth, allowing to view a 3D scene without wearing glasses. The analysis of state-of-the-art tensor displays assumes that the content is Lambertian. In order to extend its capabilities, we analyze the limitations of displaying non-Lambertian scenes and propose a new method to factorize the non-Lambertian scenes using disparity analysis. Moreover, we demonstrate a new prototype of a tensor display with three layers of full HD content at 60 fps. Compared with state-of-the-art, the evaluation results verify that the proposed non-Lambertian rendering method can display a higher quality for non-Lambertian scenes on both simulation and a prototyped tensor display.
Armand Losfeld, Eline Soetens, Daniele Bonatto, Sarah Fachada, Laurie Van Bogaert, Gauthier Lafruit, Mehrdad Teratani
VCIP3
2021 MPEG Immersive Video tools for Light Field Head Mounted Displays
abstract
Light field displays project hundreds of micro-parallax views for users to perceive 3D without wearing glasses. It results in gigantic bandwidth requirements if all views would be transmitted, even using conventional video compression per view. MPEG Immersive Video (MIV) follows a smarter strategy by transmitting only key images and some metadata to synthesize all the missing views. We developed (and will demonstrate) a real-time Depth Image Based Rendering software that follows this approach for synthesizing all light field micro-parallax views from a couple of RGBD input views.
Daniele Bonatto, Grégoire Hirt, Alexander Kvasov, Sarah Fachada, Gauthier Lafruit
VCIP1
2021 Polynomial Image-Based Rendering for non-Lambertian Objects
abstract
Non-Lambertian objects present an aspect which depends on the viewer's position towards the surrounding scene. Contrary to diffuse objects, their features move non-linearly with the camera, preventing rendering them with existing Depth Image-Based Rendering (DIBR) approaches, or to triangulate their surface with Structure-from-Motion (SfM). In this paper, we propose an extension of the DIBR paradigm to describe these non-linearities, by replacing the depth maps by more complete multi-channel “non-Lambertian maps”, without attempting a 3D reconstruction of the scene. We provide a study of the importance of each coefficient of the proposed map, measuring the trade-off between visual quality and data volume to optimally render non-Lambertian objects. We compare our method to other state-of-the-art image-based rendering methods and outperform them with promising subjective and objective results on a challenging dataset.
Sarah Fachada, Daniele Bonatto, Mehrdad Teratani, Gauthier Lafruit
VCIP2