Mehrdad Teratani

dblp:312/0719 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
14since 2021 · last 2026
0000-0001-9332-1409ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 13 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Structure-from-motion in micro-image domain for uncalibrated plenoptic 2.0 cameras
abstract
We introduce a structure-from-motion method specifically designed to process the raw micro-images captured by plenoptic 2.0 cameras. Unlike traditional monocular cameras, plenoptic cameras incorporate a micro-lens array between the sensor and the main lens, capturing depth information at the expense of a more complex set of parameters to evaluate. Instead of simply integrating their projection model into the classical structure-from-motion pipeline, our contribution identifies the pinhole cameras-driven constraints and takes advantage of the inherent disparity information present in plenoptic cameras. This facilitates a robust initialization of the reconstruction. Our method shortcuts two of the limitations of the classical structure-from-motion: the ambiguity found in scenes captured with low angular disparity and the scale ambiguity. It enables the reconstruction of scenes captured by multiple uncalibrated plenoptic cameras, without using any calibration pattern or subaperture view extraction step. Our method undergoes experimental validation on both natural and synthetic datasets, showing a 10% error accuracy for relative pose estimation, which is comparable to calibration-pattern based methods. The results are robust to coarse initialization. Contrary to classical structure-from-motion, it is able to reconstruct scenes with parallel facing cameras. It also shows greater accuracy than reconstruction methods based on pinhole camera conversion.
Sarah Dury, Daniele Bonatto, Jaime Sancho, Eduardo Juárez Martínez, Mehrdad Teratani, Gauthier Lafruit
Int. J. Comput. Vis.5
2025 Synchronization and Calibration of Video Sequences Acquired Using Multiple Plenoptic 2.0 Cameras
Daniele Bonatto, Sarah Fachada, Jaime Sancho, Eduardo Juárez Martínez, Gauthier Lafruit, Mehrdad Teratani
MMM (4)6
2025 An Analytical Method for Rendering Plenoptic Cameras 2.0 on 3D Multi-layer Displays
Armand Losfeld, Nicolas Seznec, Laurie Van Bogaert, Gauthier Lafruit, Mehrdad Teratani
MMM (1)5
2025 DA4NeRF: Depth-aware Augmentation technique for Neural Radiance Fields
abstract
Neural Radiance Fields (NeRF) demonstrate impressive capabilities in rendering novel views of specific scenes by learning an implicit volumetric representation from posed RGB images without any depth information. View synthesis is the computational process of synthesizing novel images of a scene from different viewpoints, based on a set of existing images. One big problem is the need for a large number of images in the training datasets for neural network-based view synthesis frameworks. The challenge of data augmentation for view synthesis applications has not been addressed yet. NeRF models require comprehensive scene coverage in multiple views to accurately estimate radiance and density at any point. In cases without sufficient coverage of scenes with different viewing directions, cannot effectively interpolate or extrapolate unseen scene parts. In this paper, we introduce a new pipeline to tackle this data augmentation problem using depth data. We use MPEG's Depth Estimation Reference Software and Reference View Synthesizer to add novel non-existent views to the training sets needed for the NeRF framework. Experimental results show that our approach improves the quality of the rendered images using NeRF's model. The average quality increased by 6.4 dB in terms of Peak Signal-to-Noise Ratio (PSNR), with the highest increase being 11 dB. Our approach not only adds the ability to handle the sparsely captured multiview content to be used in the NeRF framework, but also makes NeRF more accurate and useful for creating high-quality virtual views.
Hamed Razavi Khosroshahi, Jaime Sancho, Gun Bang, Gauthier Lafruit, Eduardo Juárez Martínez, Mehrdad Teratani
J. Vis. Commun. Image Represent.6
2025 Micro-Image Domain View Synthesizer for Free Navigation With Focused Plenoptic Cameras
abstract
We present a novel, first-of-its-kind view synthesis method for plenoptic images, which enables the direct manipulation of images in the micro-images array format, thereby bypassing intermediate transformation steps. Current plenoptic imaging approaches typically rely on an initial conversion to dense multiview images, also known as subaperture images extraction. However, the use of subaperture images presents two main limitations that ultimately impact further processing. First, existing subaperture view extraction methods offer limited control over camera parameters, resolutions, and poses of the subaperture views, which are also constrained to a small area around the main lens, thus restricting free navigation. Second, subaperture images are susceptible to artifacts which can propagate to subsequent processes such as calibration, depth estimation and view synthesis. In this paper, we propose a camera model that enables depth image-based rendering with plenoptic cameras, in a way that allows for the direct synthesis of any target viewpoint. In our evaluation, we show that our method expands view synthesis extrapolation to a range that is two to three times greater than that of pipelines requiring a conversion to subaperture images, including generally accepted tools such as depth image-based rendering and learning-based rendering approaches.
Sarah Fachada, Daniele Bonatto, Gauthier Lafruit, Mehrdad Teratani
IEEE Trans. Multim.4
2024 Single RGBD to Multilayer 3D Display Pipeline
abstract
Tensor displays are multiview glasses-free 3D displays composed of a stack of LCD layers. The images displayed on each layer are usually generated from a dense set of viewpoints (light field) or a set of focused images (focal stack). This paper presents an alternative using a single color view with depth for the generation of the images’ layers. To do so various single-RGBD pipelines using intermediate data formats to generate the layers’ images, including a direct method, are compared. Experiments show that by a single-color view with depth as input, one can provide a quality close to the light field or focal stack pipelines, hence reducing memory storage costs.
Laurie Van Bogaert, Armand Losfeld, Gauthier Lafruit, Mehrdad Teratani
ICME4
2024 Advancements in Lenslet Video Coding: Insights from MPEG LVC
abstract
Being a general representation format of the dense light field, lenslet video, where each frame consists of a 2D grid of micro-images, shows high potential for applications in immersive media such as glasses-free 3D displays and virtual reality. However, its distinct spatial-temporal-angular distribution places a significant challenge on conventional video coding. In July 2021, the Moving Picture Experts Group (MPEG) established an Ad-Hoc group, Lenslet Video Coding (LVC), to explore use cases, efficient compression methods, testing sequences, conversion tools, and coding architectures towards a new compression standard. This paper provides an overview of recent progress in the LVC Ad Hoc group and presents the compression efficiency of state-of-the-art codec agnostic coding tools to encourage contributions in the future.
Mehrdad Teratani, Byeungwoo Jeon, Toshiaki Fujii, Ruibo Zhao, Eline Soetens
VCIP2
2024 A Practical Approach to Depth-Aware Augmentation for Neural Radiance Fields
abstract
Neural Radiance Fields (NeRF) have demonstrated exceptional performance in generating novel views of scenes by learning implicit volumetric representations from calibrated RGB images, without depth information. A major limitation is the need for large training datasets in neural network-based view synthesis frameworks. The challenge of effective data augmentation for view synthesis remains unresolved. NeRF models require extensive scene coverage from multiple views to accurately estimate radiance and density. Insufficient coverage reduces the model’s ability to interpolate or extrapolate unseen parts of the scene effectively. In this paper, we propose a novel pipeline to address this data augmentation issue using depth map information. We use depth image-based rendering (DIBR) to overcome the lack of enough views for training NeRF. Experimental results indicate that our approach enhances the quality of rendered images using the NeRF framework, achieving an average peak signal-to-noise ratio (PSNR) increase of 7.2 dB, with a maximum improvement of 12 dB.
Hamed Razavi Khosroshahi, Jaime Sancho, Daniele Bonatto, Sarah Fachada, Gun Bang, Gauthier Lafruit, Eduardo Juárez Martínez, Mehrdad Teratani
VCIP8
2024 Codec-agnostic Lenslet Video Coding with Smoothing Transform
abstract
Plenoptic cameras are light field capturing devices able to acquire large amounts of angular and spatial information. The lenslet video produced by such cameras presents on each frame a distinctive hexagonal pattern of micro-images. Due to the particular structure of lenslet images, traditional video codecs perform poorly on lenslet video. Previous works have proposed a preprocessing scheme that cuts and realigns the micro-images on each lenslet frame. While effective, this method introduces high frequency components into the processed image. In this paper, we propose an additional step to the aforementioned scheme by applying an invertible smoothing transform. We evaluate the enhanced scheme on lenslet video sequences captured with single-focused and multi-focused plenoptic cameras. On average, the enhanced scheme achieves 9.85% bitrate reduction compared to the existing scheme.
Eline Soetens, Gauthier Lafruit, Mehrdad Teratani
VCIP3
2023 Perspective Camera Model for Layer Optimization of 3D Layered Display
abstract
3D layered displays are composed of a backlight and multiple LCD panels. To reproduce a 3D scene with such displays, a light field is given as input to optimize the layers' images that are displayed by each LCD panel. Current works optimize the layers using parallel rays (orthographic camera model) or do not take into account the distance of the user during the layers' optimization. In this paper, we present the use of the perspective model during the layer's optimization and guidelines for camera parameters based on the functioning of these displays. By using a more realistic camera model during the layers' optimization, a higher quality rendering is expected for close and middle viewing distances while giving similar quality for far distances. In the experiments, we observe that cameras' positioning can be adjusted to enhance the quality of high-parallax light fields. As a final result, the proposed method achieves higher quality on closer ranges to the 3D display when compared to the conventional method using parallel rays.
Armand Losfeld, Laurie Van Bogaert, Eline Soetens, Gauthier Lafruit, Mehrdad Teratani
MMSP5
2022 Pattern-free Plenoptic 2.0 Camera Calibration
abstract
The plenoptic 2.0 camera is a light field acquisition system consisting of a main lens and a micro-lens array (MLA) at a non-focal distance of the main lens. While it allows to retrieve the geometry of the scene, the distances between the main lens, the MLA and the sensor are usually unknown. Therefore, the use cases for plenoptic cameras stay limited while they have more potential applications such as virtual reality, provided that their camera parameters are precisely known. In this paper, we present a pattern-free calibration method to retrieve the plenoptic camera's intrinsic parameters from the rendered subaperture images, i.e. the distance from main lens to the MLA and sensor, the focal length and the principal point. The proposed method utilises the relation between scene's distances, i.e. depth, and the subaperture images' micro-disparities, which contains information about the camera parameters. To the best of our knowledge, it is the first pattern-free calibration method for plenoptic 2.0 cameras. We compare the parameters obtained using our method applied to subaperture images rendered from three different software tools (Plenoptic Toolbox, Reference Lenslet content Convertor, and Lenslet to Multiview) with an open-source pattern-based method (Compote). The proposed pattern-free calibration method has consistent camera parameters with the traditional pattern-based calibration method. We therefore reliably obtain the intrinsic parameters of any plenoptic camera from its captured images, even in the absence of any calibration pattern.
Sarah Fachada, Daniele Bonatto, Armand Losfeld, Gauthier Lafruit, Mehrdad Teratani
MMSP5
2022 3D Tensor Display for Non-Lambertian Content
abstract
A tensor display is a type of 3D light field display, composed of multiple transparent screens and a back-light that can render a scene with correct depth, allowing to view a 3D scene without wearing glasses. The analysis of state-of-the-art tensor displays assumes that the content is Lambertian. In order to extend its capabilities, we analyze the limitations of displaying non-Lambertian scenes and propose a new method to factorize the non-Lambertian scenes using disparity analysis. Moreover, we demonstrate a new prototype of a tensor display with three layers of full HD content at 60 fps. Compared with state-of-the-art, the evaluation results verify that the proposed non-Lambertian rendering method can display a higher quality for non-Lambertian scenes on both simulation and a prototyped tensor display.
Armand Losfeld, Eline Soetens, Daniele Bonatto, Sarah Fachada, Laurie Van Bogaert, Gauthier Lafruit, Mehrdad Teratani
VCIP7
2021 A Calibration Method for Subaperture Views of Plenoptic 2.0 Camera Arrays
abstract
We present a novel methodology to precisely calibrate the subaperture views of an array of plenoptic 2.0 cameras. Such cameras consist of a micro lens array, and the image captured through them is a lenslet image that can be converted to a dense set of pinhole views, the so-called subaperture images. This cam-era array provides several dense multiview images at some sparse points of 3D space. To find the relative position of those views, simply using structure-from-motion creates misalignments due to the small disparities within each set. Additionally, a traditional calibration using calibration patterns will also fail due to the complicated objectives of plenoptic 2.0 cameras and artifacts when they are converted to subaperture views. In this paper, we propose two calibration steps (a) to register the sparse central subaperture views using Structure-from-Motion which makes it robust to artifacts in the subaperture views, and (b) to register all dense multiview sets per plenoptic camera using camera’s lenses specifications, disparity and distance to the scene. These two steps are followed by a novel merging process of the former registrations, to achieve precise calibration parameters for all the subaperture views of the multi-plenoptic array. Experimental results objectively and subjectively demonstrate high accuracy of the calibration. We show a 10% smaller reprojection error than using a naive structure-from-motion approach and verify that our method is suitable for high precision view synthesis applications such as virtual reality and holography.
Sarah Fachada, Armand Losfeld, Takanori Senoh, Gauthier Lafruit, Mehrdad Teratani
MMSP5
2021 Polynomial Image-Based Rendering for non-Lambertian Objects
abstract
Non-Lambertian objects present an aspect which depends on the viewer's position towards the surrounding scene. Contrary to diffuse objects, their features move non-linearly with the camera, preventing rendering them with existing Depth Image-Based Rendering (DIBR) approaches, or to triangulate their surface with Structure-from-Motion (SfM). In this paper, we propose an extension of the DIBR paradigm to describe these non-linearities, by replacing the depth maps by more complete multi-channel “non-Lambertian maps”, without attempting a 3D reconstruction of the scene. We provide a study of the importance of each coefficient of the proposed map, measuring the trade-off between visual quality and data volume to optimally render non-Lambertian objects. We compare our method to other state-of-the-art image-based rendering methods and outperform them with promising subjective and objective results on a challenging dataset.
Sarah Fachada, Daniele Bonatto, Mehrdad Teratani, Gauthier Lafruit
VCIP3