EDBT 2026 Demo / reviewers in the wild / expert
Gauthier Lafruit
dblp:79/5867
· DBLP profile ↗
73ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-2349-1936ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 65 · 3 first-author · 16 since 2021Systems, architecture and hardware · 6 · 2 first-authorArtificial intelligence and machine learning · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Structure-from-motion in micro-image domain for uncalibrated plenoptic 2.0 camerasabstractWe introduce a structure-from-motion method specifically designed to process the raw micro-images captured by plenoptic 2.0 cameras. Unlike traditional monocular cameras, plenoptic cameras incorporate a micro-lens array between the sensor and the main lens, capturing depth information at the expense of a more complex set of parameters to evaluate. Instead of simply integrating their projection model into the classical structure-from-motion pipeline, our contribution identifies the pinhole cameras-driven constraints and takes advantage of the inherent disparity information present in plenoptic cameras. This facilitates a robust initialization of the reconstruction. Our method shortcuts two of the limitations of the classical structure-from-motion: the ambiguity found in scenes captured with low angular disparity and the scale ambiguity. It enables the reconstruction of scenes captured by multiple uncalibrated plenoptic cameras, without using any calibration pattern or subaperture view extraction step. Our method undergoes experimental validation on both natural and synthetic datasets, showing a 10% error accuracy for relative pose estimation, which is comparable to calibration-pattern based methods. The results are robust to coarse initialization. Contrary to classical structure-from-motion, it is able to reconstruct scenes with parallel facing cameras. It also shows greater accuracy than reconstruction methods based on pinhole camera conversion. Sarah Dury, Daniele Bonatto, Jaime Sancho, Eduardo Juárez Martínez, Mehrdad Teratani, Gauthier Lafruit |
Int. J. Comput. Vis. | 6 |
| 2025 | Synchronization and Calibration of Video Sequences Acquired Using Multiple Plenoptic 2.0 Cameras
Daniele Bonatto, Sarah Fachada, Jaime Sancho, Eduardo Juárez Martínez, Gauthier Lafruit, Mehrdad Teratani |
MMM (4) | 5 |
| 2025 | An Analytical Method for Rendering Plenoptic Cameras 2.0 on 3D Multi-layer Displays
Armand Losfeld, Nicolas Seznec, Laurie Van Bogaert, Gauthier Lafruit, Mehrdad Teratani |
MMM (1) | 4 |
| 2025 | PaintUnicorn: Multi-view video datasetabstractWe present PaintUnicorn, a real-world multiview video dataset designed for advancing research in dynamic neural rendering, 4D reconstruction, and view synthesis. Captured with 28 synchronized cameras at 1080p and 30 fps over 300 frames, the dataset includes complex non-rigid motion, specular and transparent surfaces, and varying lighting conditions. We provide calibrated camera intrinsics and extrinsics, color-corrected images, and generated depth maps combined from multiview geometry and learning-based estimates. PaintUnicorn is a dataset to support dense temporal and spatial benchmarking for methods like 3D Gaussian Splatting and dynamic NeRFs. The dataset is accessible in https://doi.org/10.5281/zenodo.16312256 Eva Dubar, Daniele Bonatto, Sarah Dury, Gauthier Lafruit |
VCIP | 4 |
| 2025 | DA4NeRF: Depth-aware Augmentation technique for Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) demonstrate impressive capabilities in rendering novel views of specific scenes by learning an implicit volumetric representation from posed RGB images without any depth information. View synthesis is the computational process of synthesizing novel images of a scene from different viewpoints, based on a set of existing images. One big problem is the need for a large number of images in the training datasets for neural network-based view synthesis frameworks. The challenge of data augmentation for view synthesis applications has not been addressed yet. NeRF models require comprehensive scene coverage in multiple views to accurately estimate radiance and density at any point. In cases without sufficient coverage of scenes with different viewing directions, cannot effectively interpolate or extrapolate unseen scene parts. In this paper, we introduce a new pipeline to tackle this data augmentation problem using depth data. We use MPEG's Depth Estimation Reference Software and Reference View Synthesizer to add novel non-existent views to the training sets needed for the NeRF framework. Experimental results show that our approach improves the quality of the rendered images using NeRF's model. The average quality increased by 6.4 dB in terms of Peak Signal-to-Noise Ratio (PSNR), with the highest increase being 11 dB. Our approach not only adds the ability to handle the sparsely captured multiview content to be used in the NeRF framework, but also makes NeRF more accurate and useful for creating high-quality virtual views. Hamed Razavi Khosroshahi, Jaime Sancho, Gun Bang, Gauthier Lafruit, Eduardo Juárez Martínez, Mehrdad Teratani |
J. Vis. Commun. Image Represent. | 4 |
| 2025 | Micro-Image Domain View Synthesizer for Free Navigation With Focused Plenoptic CamerasabstractWe present a novel, first-of-its-kind view synthesis method for plenoptic images, which enables the direct manipulation of images in the micro-images array format, thereby bypassing intermediate transformation steps. Current plenoptic imaging approaches typically rely on an initial conversion to dense multiview images, also known as subaperture images extraction. However, the use of subaperture images presents two main limitations that ultimately impact further processing. First, existing subaperture view extraction methods offer limited control over camera parameters, resolutions, and poses of the subaperture views, which are also constrained to a small area around the main lens, thus restricting free navigation. Second, subaperture images are susceptible to artifacts which can propagate to subsequent processes such as calibration, depth estimation and view synthesis. In this paper, we propose a camera model that enables depth image-based rendering with plenoptic cameras, in a way that allows for the direct synthesis of any target viewpoint. In our evaluation, we show that our method expands view synthesis extrapolation to a range that is two to three times greater than that of pipelines requiring a conversion to subaperture images, including generally accepted tools such as depth image-based rendering and learning-based rendering approaches. Sarah Fachada, Daniele Bonatto, Gauthier Lafruit, Mehrdad Teratani |
IEEE Trans. Multim. | 3 |
| 2024 | Single RGBD to Multilayer 3D Display PipelineabstractTensor displays are multiview glasses-free 3D displays composed of a stack of LCD layers. The images displayed on each layer are usually generated from a dense set of viewpoints (light field) or a set of focused images (focal stack). This paper presents an alternative using a single color view with depth for the generation of the images’ layers. To do so various single-RGBD pipelines using intermediate data formats to generate the layers’ images, including a direct method, are compared. Experiments show that by a single-color view with depth as input, one can provide a quality close to the light field or focal stack pipelines, hence reducing memory storage costs. Laurie Van Bogaert, Armand Losfeld, Gauthier Lafruit, Mehrdad Teratani |
ICME | 3 |
| 2024 | A Practical Approach to Depth-Aware Augmentation for Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have demonstrated exceptional performance in generating novel views of scenes by learning implicit volumetric representations from calibrated RGB images, without depth information. A major limitation is the need for large training datasets in neural network-based view synthesis frameworks. The challenge of effective data augmentation for view synthesis remains unresolved. NeRF models require extensive scene coverage from multiple views to accurately estimate radiance and density. Insufficient coverage reduces the model’s ability to interpolate or extrapolate unseen parts of the scene effectively. In this paper, we propose a novel pipeline to address this data augmentation issue using depth map information. We use depth image-based rendering (DIBR) to overcome the lack of enough views for training NeRF. Experimental results indicate that our approach enhances the quality of rendered images using the NeRF framework, achieving an average peak signal-to-noise ratio (PSNR) increase of 7.2 dB, with a maximum improvement of 12 dB. Hamed Razavi Khosroshahi, Jaime Sancho, Daniele Bonatto, Sarah Fachada, Gun Bang, Gauthier Lafruit, Eduardo Juárez Martínez, Mehrdad Teratani |
VCIP | 6 |
| 2024 | Codec-agnostic Lenslet Video Coding with Smoothing TransformabstractPlenoptic cameras are light field capturing devices able to acquire large amounts of angular and spatial information. The lenslet video produced by such cameras presents on each frame a distinctive hexagonal pattern of micro-images. Due to the particular structure of lenslet images, traditional video codecs perform poorly on lenslet video. Previous works have proposed a preprocessing scheme that cuts and realigns the micro-images on each lenslet frame. While effective, this method introduces high frequency components into the processed image. In this paper, we propose an additional step to the aforementioned scheme by applying an invertible smoothing transform. We evaluate the enhanced scheme on lenslet video sequences captured with single-focused and multi-focused plenoptic cameras. On average, the enhanced scheme achieves 9.85% bitrate reduction compared to the existing scheme. Eline Soetens, Gauthier Lafruit, Mehrdad Teratani |
VCIP | 2 |
| 2023 | Perspective Camera Model for Layer Optimization of 3D Layered Displayabstract3D layered displays are composed of a backlight and multiple LCD panels. To reproduce a 3D scene with such displays, a light field is given as input to optimize the layers' images that are displayed by each LCD panel. Current works optimize the layers using parallel rays (orthographic camera model) or do not take into account the distance of the user during the layers' optimization. In this paper, we present the use of the perspective model during the layer's optimization and guidelines for camera parameters based on the functioning of these displays. By using a more realistic camera model during the layers' optimization, a higher quality rendering is expected for close and middle viewing distances while giving similar quality for far distances. In the experiments, we observe that cameras' positioning can be adjusted to enhance the quality of high-parallax light fields. As a final result, the proposed method achieves higher quality on closer ranges to the 3D display when compared to the conventional method using parallel rays. Armand Losfeld, Laurie Van Bogaert, Eline Soetens, Gauthier Lafruit, Mehrdad Teratani |
MMSP | 4 |
| 2022 | Pattern-free Plenoptic 2.0 Camera CalibrationabstractThe plenoptic 2.0 camera is a light field acquisition system consisting of a main lens and a micro-lens array (MLA) at a non-focal distance of the main lens. While it allows to retrieve the geometry of the scene, the distances between the main lens, the MLA and the sensor are usually unknown. Therefore, the use cases for plenoptic cameras stay limited while they have more potential applications such as virtual reality, provided that their camera parameters are precisely known. In this paper, we present a pattern-free calibration method to retrieve the plenoptic camera's intrinsic parameters from the rendered subaperture images, i.e. the distance from main lens to the MLA and sensor, the focal length and the principal point. The proposed method utilises the relation between scene's distances, i.e. depth, and the subaperture images' micro-disparities, which contains information about the camera parameters. To the best of our knowledge, it is the first pattern-free calibration method for plenoptic 2.0 cameras. We compare the parameters obtained using our method applied to subaperture images rendered from three different software tools (Plenoptic Toolbox, Reference Lenslet content Convertor, and Lenslet to Multiview) with an open-source pattern-based method (Compote). The proposed pattern-free calibration method has consistent camera parameters with the traditional pattern-based calibration method. We therefore reliably obtain the intrinsic parameters of any plenoptic camera from its captured images, even in the absence of any calibration pattern. Sarah Fachada, Daniele Bonatto, Armand Losfeld, Gauthier Lafruit, Mehrdad Teratani |
MMSP | 4 |
| 2022 | 3D Tensor Display for Non-Lambertian ContentabstractA tensor display is a type of 3D light field display, composed of multiple transparent screens and a back-light that can render a scene with correct depth, allowing to view a 3D scene without wearing glasses. The analysis of state-of-the-art tensor displays assumes that the content is Lambertian. In order to extend its capabilities, we analyze the limitations of displaying non-Lambertian scenes and propose a new method to factorize the non-Lambertian scenes using disparity analysis. Moreover, we demonstrate a new prototype of a tensor display with three layers of full HD content at 60 fps. Compared with state-of-the-art, the evaluation results verify that the proposed non-Lambertian rendering method can display a higher quality for non-Lambertian scenes on both simulation and a prototyped tensor display. Armand Losfeld, Eline Soetens, Daniele Bonatto, Sarah Fachada, Laurie Van Bogaert, Gauthier Lafruit, Mehrdad Teratani |
VCIP | 6 |
| 2021 | A Calibration Method for Subaperture Views of Plenoptic 2.0 Camera ArraysabstractWe present a novel methodology to precisely calibrate the subaperture views of an array of plenoptic 2.0 cameras. Such cameras consist of a micro lens array, and the image captured through them is a lenslet image that can be converted to a dense set of pinhole views, the so-called subaperture images. This cam-era array provides several dense multiview images at some sparse points of 3D space. To find the relative position of those views, simply using structure-from-motion creates misalignments due to the small disparities within each set. Additionally, a traditional calibration using calibration patterns will also fail due to the complicated objectives of plenoptic 2.0 cameras and artifacts when they are converted to subaperture views. In this paper, we propose two calibration steps (a) to register the sparse central subaperture views using Structure-from-Motion which makes it robust to artifacts in the subaperture views, and (b) to register all dense multiview sets per plenoptic camera using camera’s lenses specifications, disparity and distance to the scene. These two steps are followed by a novel merging process of the former registrations, to achieve precise calibration parameters for all the subaperture views of the multi-plenoptic array. Experimental results objectively and subjectively demonstrate high accuracy of the calibration. We show a 10% smaller reprojection error than using a naive structure-from-motion approach and verify that our method is suitable for high precision view synthesis applications such as virtual reality and holography. Sarah Fachada, Armand Losfeld, Takanori Senoh, Gauthier Lafruit, Mehrdad Teratani |
MMSP | 4 |
| 2021 | MPEG Immersive Video tools for Light Field Head Mounted DisplaysabstractLight field displays project hundreds of micro-parallax views for users to perceive 3D without wearing glasses. It results in gigantic bandwidth requirements if all views would be transmitted, even using conventional video compression per view. MPEG Immersive Video (MIV) follows a smarter strategy by transmitting only key images and some metadata to synthesize all the missing views. We developed (and will demonstrate) a real-time Depth Image Based Rendering software that follows this approach for synthesizing all light field micro-parallax views from a couple of RGBD input views. Daniele Bonatto, Grégoire Hirt, Alexander Kvasov, Sarah Fachada, Gauthier Lafruit |
VCIP | 5 |
| 2021 | Polynomial Image-Based Rendering for non-Lambertian ObjectsabstractNon-Lambertian objects present an aspect which depends on the viewer's position towards the surrounding scene. Contrary to diffuse objects, their features move non-linearly with the camera, preventing rendering them with existing Depth Image-Based Rendering (DIBR) approaches, or to triangulate their surface with Structure-from-Motion (SfM). In this paper, we propose an extension of the DIBR paradigm to describe these non-linearities, by replacing the depth maps by more complete multi-channel “non-Lambertian maps”, without attempting a 3D reconstruction of the scene. We provide a study of the importance of each coefficient of the proposed map, measuring the trade-off between visual quality and data volume to optimally render non-Lambertian objects. We compare our method to other state-of-the-art image-based rendering methods and outperform them with promising subjective and objective results on a challenging dataset. Sarah Fachada, Daniele Bonatto, Mehrdad Teratani, Gauthier Lafruit |
VCIP | 4 |
| 2021 | A Lightweight Depth Estimation Network for Wide-Baseline Light FieldsabstractExisting traditional and ConvNet-based methods for light field depth estimation mainly work on the narrow-baseline scenario. This paper explores the feasibility and capability of ConvNets to estimate depth in another promising scenario: wide-baseline light fields. Due to the deficiency of training samples, a large-scale and diverse synthetic wide-baseline dataset with labelled data is introduced for depth prediction tasks. Considering the practical goal for real-world applications, we design an end-to-end trained lightweight convolutional network to infer depths from light fields, called LLF-Net. The proposed LLF-Net is built by incorporating a cost volume which allows variable angular light field inputs and an attention module that enables to recover details at occlusion areas. Evaluations are made on the synthetic and real-world wide-baseline light fields, and experimental results show that the proposed network achieves the best performance when compared to recent state-of-the-art methods. We also evaluate our LLF-Net on narrow-baseline datasets, and it consequently improves the performance of previous methods. Yan Li 0083, Lu Zhang 0037, Gauthier Lafruit |
IEEE Trans. Image Process. | 4 |
| 2021 | Low-Rank Constrained Super-Resolution for Mixed-Resolution Multiview VideoabstractMultiview video allows for simultaneously presenting dynamic imaging from multiple viewpoints, enabling a broad range of immersive applications. This paper proposes a novel super-resolution (SR) approach to mixed-resolution (MR) multiview video, whereby the low-resolution (LR) videos produced by MR camera setups are up-sampled based on the neighboring HR videos. Our solution analyzes the statistical correlation of different resolutions between multiple views, and introduces a low-rank prior based SR optimization framework using local linear embedding and weighted nuclear norm minimization. The target HR patch is reconstructed by learning texture details from the neighboring HR camera views using local linear embedding. A low-rank constrained patch optimization solution is introduced to effectively restrain visual artifacts and the ADMM framework is used to solve the resulting optimization problem. Comprehensive experiments including objective and subjective test metrics demonstrate that the proposed method outperforms the state-of-the-art SR methods for MR multiview video. Shao-Ping Lu, Senmao Li, Gauthier Lafruit, Ming-Ming Cheng, Adrian Munteanu 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Manet: Multi-Scale Aggregated Network For Light Field Depth EstimationabstractWe present a novel end-to-end network, MANet, for light field depth estimation. MANet is a parameter-effective and effi-cient multi-scale aggregated network, which is about 3 times smaller and 3 times faster than the current top-performing method Epinet. The MANet architecture is performed for estimating depth from light field plenoptic cameras, and experimental results show that the proposed MANet outperforms state-of-the-art methods on HCI, CVIA-HCI and EPFL Lytro light field datasets. Yan Li 0083, Lu Zhang 0037, Gauthier Lafruit |
ICASSP | 4 |
| 2018 | Robust Multiview Synthesis for Wide-Baseline Camera ArraysabstractIn many advanced multimedia systems, multiview content can offer more immersion compared to classical stereoscopy. The feeling of immersiveness is increased substantially by offering motion-parallax, as well as stereopsis. This drives both the so-called free-navigation and super-multiview technologies. However, it is currently still challenging to acquire, store, process, and transmit this type of content. This paper presents a novel multiview-interpolation framework for wide-baseline camera arrays. The proposed method comprises several novel components, including point cloud-based filtering, improved de-ghosting, multireference color blending, and depth-aware MRF-based disocclusion in painting. The method offers robustness against depth errors caused by quantization and smoothing across object boundaries. Furthermore, the available input color and depth are maximally exploited while preventing propagation of unreliable information to virtual viewpoints. The experimental results show that the proposed method outperforms the state-of-the-art View Synthesis Reference Software (VSRS 4.1) both in objective terms as well as subjectively, based on a visual assessment on a high-end light-field three-dimensional display. Beerend Ceulemans, Shao-Ping Lu, Gauthier Lafruit, Adrian Munteanu 0001 |
IEEE Trans. Multim. | 3 |
| 2017 | Erratum to: A model for adapting 3D graphics based on scalable coding, real-time simplification and remote rendering
Marius Preda, Paulo Villegas, Francisco Morán, Gauthier Lafruit, Robert-Paul Berretty |
Vis. Comput. | 4 |
| 2016 | Efficient MRF-based disocclusion inpainting in multiview videoabstractView synthesis using depth image-based rendering generates virtual viewpoints of a 3D scene based on texture and depth information from a set of available cameras. One of the core components in view synthesis is image inpainting which performs the reconstruction of areas that were occluded in the available cameras but are visible from the virtual viewpoint. Inpainting methods based on Markov random fields (MRFs) have been shown to be very effective in inpainting large areas in images. In this paper, we propose a novel MRF-based in-painting method for multiview video. The proposed method steers the MRF optimization towards completion from background to foreground and exploits the available depth information in order to avoid bleeding artifacts. The proposed approach allows for efficiently filling-in large disocclusion areas and greatly accelerates execution compared to traditional MRF-based inpainting techniques. The experimental results show that view synthesis based on the proposed inpainting method systematically improves performance over the state-of-the-art in multiview view synthesis. Average PSNR gains up to 1.88 dB compared to the MPEG View Synthesis Reference software were observed. Beerend Ceulemans, Shao-Ping Lu, Gauthier Lafruit, Peter Schelkens, Adrian Munteanu 0001 |
ICME | 3 |
| 2016 | Modular Parallelization Framework for Multi-Stream Video ProcessingabstractA flow-based software framework specialized for 3D and video is presented, which in particular handles automatic parallelization of multi-stream video processing. The workflow is decomposed into a set of filter units which become nodes in a graph. Execution with support for time windows on node inputs is automatically handled by the framework, as well as low-level manipulation of multi-dimensional data. Tim Lenertz, Gauthier Lafruit |
ACM Multimedia | 2 |
| 2015 | Generalized inpainting method for hyperspectral image acquisitionabstractA recently designed hyperspectral imaging device enables multiplexed acquisition of an entire data volume in a single snapshot thanks to monolithically-integrated spectral filters. Such an agile imaging technique comes at the cost of a reduced spatial resolution and the need for a demosaicing procedure on its interleaved data. In this work, we address both issues and propose an approach inspired by recent developments in compressed sensing and analysis sparse models. We formulate our superresolution and demosaicing task as a 3-D generalized inpainting problem. Interestingly, the target spatial resolution can be adjusted for mitigating the compression level of our sensing. The reconstruction procedure uses a fast greedy method called Pseudo-inverse IHT. We also show on simulations that a random arrangement of the spectral filters on the sensor is preferable to regular mosaic layout as it improves the quality of the reconstruction. The efficiency of our technique is demonstrated through numerical experiments on both synthetic and real data as acquired by the snapshot imager. Kévin Degraux, Valerio Cambareri, Laurent Jacques, Bert Geelen, Carolina Blanch, Gauthier Lafruit |
ICIP | 6 |
| 2015 | Color retargeting: Interactive time-varying color image composition from time-lapse sequencesabstractIn this paper, we present an interactive static image composition approach, namely color retargeting , to flexibly represent time-varying color editing effect based on time-lapse video sequences. Instead of performing precise image matting or blending techniques, our approach treats the color composition as a pixel-level resampling problem. In order to both satisfy the user’s editing requirements and avoid visual artifacts, we construct a globally optimized interpolation field. This field defines from which input video frames the output pixels should be resampled. Our proposed resampling solution ensures that (i) the global color transition in the output image is as smooth as possible, (ii) the desired colors/objects specified by the user from different video frames are well preserved, and (iii) additional local color transition directions in the image space assigned by the user are also satisfied. Various examples have been shown to demonstrate that our efficient solution enables the user to easily create time-varying color image composition results. Shao-Ping Lu, Guillaume Dauphin, Gauthier Lafruit, Adrian Munteanu 0001 |
Comput. Vis. Media | 3 |
| 2014 | Derivative-Based Scale Invariant Image Feature Detector With Error ResilienceabstractWe present a novel scale-invariant image feature detection algorithm (D-SIFER) using a newly proposed scale-space optimal 10th-order Gaussian derivative (GDO-10) filter, which reaches the jointly optimal Heisenberg's uncertainty of its impulse response in scale and space simultaneously (i.e., we minimize the maximum of the two moments). The D-SIFER algorithm using this filter leads to an outstanding quality of image feature detection, with a factor of three quality improvement over state-of-the-art scale-invariant feature transform (SIFT) and speeded up robust features (SURF) methods that use the second-order Gaussian derivative filters. To reach low computational complexity, we also present a technique approximating the GDO-10 filters with a fixed-length implementation, which is independent of the scale. The final approximation error remains far below the noise margin, providing constant time, low cost, but nevertheless high-quality feature detection and registration capabilities. D-SIFER is validated on a real-life hyperspectral image registration application, precisely aligning up to hundreds of successive narrowband color images, despite their strong artifacts (blurring, low-light noise) typically occurring in such delicate optical system setups. Pradip Mainali, Gauthier Lafruit, Klaas Tack, Luc Van Gool, Rudy Lauwereins |
IEEE Trans. Image Process. | 2 |
| 2013 | SIFER: Scale-Invariant Feature Detector with Error Resilience
Pradip Mainali, Gauthier Lafruit, Qiong Yang, Bert Geelen, Luc Van Gool, Rudy Lauwereins |
Int. J. Comput. Vis. | 2 |
| 2012 | Constant Time Joint Bilateral Filtering Using Joint Integral HistogramsabstractIn this brief, we present a constant time method for the joint bilateral filtering. First, we propose an image data structure, coined as joint integral histograms (JIHs). Extending the classic integral images and the integral histograms, it represents the global information of two correlated images. In a JIH, the value at each bin indicates an integral determined by the two images. Then, the joint bilateral filtering is transformed to computation and manipulation of histograms. Utilizing the JIHs, we are capable of joint bilateral filtering in constant time. Its performance is validated in a digital photography approach using Flash-noFlash image pairs. Compared with the brute-force method, the proposed method achieves a speedup factor of 2-3 orders of magnitude while producing similar filtering results. Ke Zhang 0012, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool |
IEEE Trans. Image Process. | 2 |
| 2011 | Real-time high-definition stereo matching on FPGAabstractAlthough many fast stereo matching designs have been proposed in the past decades, it is still very challenging to achieve real-time speed at high definition resolution while maintaining high matching accuracy. In this paper, we propose a real-time high definition stereo matching design on FPGA. By using the Mini-Census transform and the Cross-based cost aggregation, the proposed algorithm is robust to radiometric differences and produces accurate disparity maps. The algorithm modules have been optimized for efficient hardware implementations and instantiated in an SoC environment. Implemented on a single EP3SL150 FPGA, our design achieves 60 frames per second for 1024 × 768 stereo images. Evaluated with the Middlebury stereo benchmark, the proposed design also delivers leading stereo matching accuracy among prior related work. Lu Zhang 0018, Ke Zhang 0012, Tian-Sheuan Chang, Gauthier Lafruit, Georgi Kuzmanov, Diederik Verkest |
FPGA | 4 |
| 2011 | Robust Low Complexity Corner DetectorabstractCorner feature point detection with both the high-speed and high-quality is still very demanding for many real-time computer vision applications. The Harris and Kanade-Lucas-Tomasi (KLT) are widely adopted good quality corner feature point detection algorithms due to their invariance to rotation, noise, illumination, and limited view point change. Although they are widely adopted corner feature point detectors, their applications are rather limited because of their inability to achieve real-time performance due to their high complexity. In this paper, we redesigned Harris and KLT algorithms to reduce their complexity in each stage of the algorithm: Gaussian derivative, cornerness response, and non-maximum suppression (NMS). The complexity of the Gaussian derivative and cornerness stage is reduced by using an integral image. In NMS stage, we replaced a highly complex sorting and NMS by the efficient NMS followed by sorting the result. The detected feature points are further interpolated for sub-pixel accuracy of the feature point location. Our experimental results on publicly available evaluation data-sets for the feature point detectors show that our low complexity corner detector is both very fast and similar in feature point detection quality compared to the original algorithm. We achieve a complexity reduction by a factor of 9.8 and attain 50 f/s processing speed for images of size 640$\,\times\,$480 on a commodity central processing unit with 2.53 GHz and 3 GB random access memory. Pradip Mainali, Qiong Yang, Gauthier Lafruit, Luc Van Gool, Rudy Lauwereins |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | Real-Time and Accurate Stereo: A Scalable Approach With Bitwise Fast Voting on CUDAabstractThis paper proposes a real-time design for accurate stereo matching on compute unified device architecture (CUDA). We present a leading local algorithm and then accelerate it by parallel computing. High matching accuracy is achieved by cost aggregation over shape-adaptive support regions and disparity refinement using reliable initial estimates. A novel sample-and-restore scheme is proposed to make the algorithm scalable, capable of attaining several times speedup at the expense of minor accuracy degradation. The refinement and the restoration are jointly realized by a local voting method. To accelerate the voting on CUDA, a graphics processing unit (GPU)-oriented bitwise fast voting method is proposed, faster than the traditional histogram-based approach with two orders of magnitude. The whole algorithm is parallelized on CUDA at a fine granularity, efficiently exploiting the computing resources of GPUs. Our design is among the fastest stereo matching methods on GPUs. Evaluated in the Middlebury stereo benchmark, the proposed design produces the most accurate results among the real-time methods. The advantages of speed, accuracy, and desirable scalability advocate our design for practical applications such as robotics systems and multiview teleconferencing. Ke Zhang 0012, Jiangbo Lu, Qiong Yang, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2010 | Lococo: low complexity corner detectorabstractThe high-speed feature detection is still very demanding for many computer vision applications. In this paper, the Harris and KLT corner detectors are redesigned to reduce the complexity of the algorithm. The complexity of Harris and KLT corner detectors are reduced by using the box kernel, the integral image and efficient non-maximum suppression, achieving complexity reduction by a factor of 8. For the image of size 1000×700, our method costs only 74ms on a commodity CPU with 2GHz and 1GB RAM. Pradip Mainali, Qiong Yang, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool |
ICASSP | 3 |
| 2010 | Robust low complexity feature trackingabstractIn this paper, we present the Kanade-Lucas-Tomasi (KLT) tracking algorithm coupled with a varying integration window, which tracks a small subset of feature points to initialize the approximate motion model between images. For the remaining larger subset of the feature points initial tracking location is predicted by using this motion model, thus improving the tracking result. For an image of size 1000×700, the computational cost is reduced by a factor of 9.5 and tracking 500 features with our method runs in only 64 ms on a commodity 2GHz CPU with 1GB RAM. Pradip Mainali, Qiong Yang, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool |
ICIP | 3 |
| 2010 | Joint integral histograms and its application in stereo matchingabstractIn this paper, we first propose a technique, referred as joint integral histograms, for weighted filtering with O(1) computational complexity. The technique is built on the classic integral images and the recent integral histograms. In a joint integral histogram, instead of remembering bin occurrences, the value at each bin indicates an integral defined by two signals. Beyond the integral histograms, our method supports weighted filtering with a more general form, where the weight could be a function of a signal different from the signal to be filtered. Then, we present a local stereo matching approach as an instantiation of the technique. Using the joint integral histograms, we achieve a speedup factor of about two orders of magnitude. Thanks to the huge speedup, the stereo method is among the best local approaches in terms of the trade-off between matching accuracy and execution speed. Experimental results demonstrate the advantages of both the joint integral histograms technique and the stereo matching approach. Ke Zhang 0012, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool |
ICIP | 2 |
| 2010 | Modeling and exploiting spatial locality trade-offs in wavelet-based applications under varying resource requirementsabstractFuture dynamic applications will require new mapping strategies to deliver power-efficient performance. Fully static design-time mappings will not be able to optimally address the unpredictably varying application characteristics and system resource requirements. Instead, the platforms will not only need to be programmable in terms of instruction set processors, but also at least partial reconfigurability will be required, while the applications themselves will need to exploit this increased freedom at runtime to adapt to the dynamism. In this context, it is important for applications to optimally exploit the memory hierarchy under varying memory availability. This article presents an analysis of spatial locality trade-offs in wavelet-based applications, to be used in dynamic execution environments: Depending on the encountered runtime conditions, the execution switches to different memory optimized instantiations or localizations, optimally exploiting temporal and spatial locality under these conditions. This is enabled by systematic mapping guidelines, indicating how the miss-rate behavior of a localization is influenced by a specific execution condition, under which conditions a certain localization is optimal and which miss-rate gains may be obtained by switching to that localization. Bert Geelen, Vissarion Ferentinos, Francky Catthoor, Gauthier Lafruit, Diederik Verkest, Rudy Lauwereins, Thanos Stouraitis |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2009 | Interpolation error as a quality metric for stereo: Robust, or not?abstractTo properly benchmark and stimulate current stereo algorithms specifically in the application context of view interpolation, a robust quantitative evaluation approach is important. As a prevailing quality assessment method, interpolation error has been widely used. It measures the distortions between an interpolated image and a real camera image for a desired virtual viewpoint. However, is it a robust quality metric, especially when state-of-the-art stereo technology is developing so fast? This paper hence focuses on revealing several rarely attended weaknesses that make the interpolation error evaluation paradigm vulnerable. In addition, we propose an alternative evaluation method as an early attempt at addressing these challenges, from a perspective of communication system. Evaluation of representative stereo methods from the Middlebury Web site shows that the new approach yields consistent quality assessment outcomes. Jiangbo Lu, Qiong Yang, Gauthier Lafruit |
ICASSP | 3 |
| 2009 | Real-time stereo matching: A cross-based local approachabstractWe propose an area-based local stereo matching algorithm that yields accurate disparity estimates, while achieving the real-time speed completely on the graphics processing unit (GPU). For a local stereo method, the key challenge is to decide an appropriate support window for the pixel under consideration. Our stereo method starts with computing an upright local cross adaptively for each anchor pixel, which defines a per-pixel support skeleton. Next, based on this compact local cross representation, we aggregate the matching costs in a shape adaptive full support region using two orthogonal integration steps. Approximating scene structures accurately, the proposed method is among the best-performing real-time stereo methods according to the benchmark Middlebury stereo evaluation. Additionally, our method is very easy to implement, memory efficient, and hence it is promising for many practical applications. Jiangbo Lu, Ke Zhang 0012, Gauthier Lafruit, Francky Catthoor |
ICASSP | 3 |
| 2009 | Robust stereo matching with fast Normalized Cross-Correlation over shape-adaptive regionsabstractNormalized cross-correlation (NCC) is a common matching technique to tolerate radiometric differences between stereo images. However, traditional rectangle-based NCC tends to blur the depth discontinuities. This paper proposes an efficient stereo algorithm with NCC over shape-adaptive matching regions, producing depth-discontinuity preserving disparity maps while remaining the advantage of robustness to radiometric differences. To alleviate the computational intensity, we propose an acceleration algorithm using an orthogonal integral image technique, achieving a speedup factor of 10~27. In addition, a voting scheme on reliable estimates is applied to refine the initial estimates. Experiments show that, besides the robustness, the proposed method obtains accurate disparity maps at fast speed. Our method highly ranks among the local approaches in the Middlebury stereo benchmark. Ke Zhang 0012, Jiangbo Lu, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool |
ICIP | 3 |
| 2009 | Accurate and efficient stereo matching with robust piecewise votingabstractIn this paper, we propose an efficient local stereo algorithm for accurate disparity estimation. First, we attain initial disparity estimates by iterating a cross-based cost aggregation process. Then, we propose a robust voting scheme to refine the initial estimates based on a piecewise smoothness prior, improving the quality in occluded regions and low-textured regions effectively. The refinement is guided by the segmentation result of input images. Unreliable initial estimates, which are detected using an efficient left-right consistency check, are rejected to increase the reliability of the voting results. Evaluated with the Middlebury stereo benchmark, our method is among the top performing local methods in accuracy. Compared to other local methods with similar accuracy, our method is faster by a factor of about two orders. Ke Zhang 0012, Jiangbo Lu, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool |
ICME | 3 |
| 2009 | Real-time stereo-based view synthesis algorithms: A unified framework and evaluation on commodity GPUs
Sammy Rogmans, Jiangbo Lu, Philippe Bekaert, Gauthier Lafruit |
Signal Process. Image Commun. | 4 |
| 2009 | Scheduling and Resource Allocation for SVC Streaming Over OFDM Downlink SystemsabstractWe consider the problem of scheduling and resource allocation for multiuser video streaming over downlink orthogonal frequency division multiplexing (OFDM) channels. The video streams are precoded using the scalable video coding (SVC) scheme that offers both quality and temporal scalabilities. The OFDM technology provides the flexibility of resource allocation in terms of time, frequency, and power. We propose a gradient-based scheduling and resource allocation algorithm, which prioritizes the transmissions of different users by considering video contents, deadline requirements, and transmission history. Simulation results show that the proposed algorithm outperforms the content-blind and deadline-blind algorithms with a gain of as much as 6 dB in terms of average PSNR when the network is congested. Xin Ji, Jianwei Huang 0001, Mung Chiang, Gauthier Lafruit, Francky Catthoor |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2009 | Stream-Centric Stereo Matching and View Synthesis: A High-Speed Approach on GPUsabstractIn this paper, we propose a real-time image-based rendering (IBR) system. It is specifically designed for photorealistic view synthesis at high-speed on the graphics processing unit (GPU). We steer the proposed IBR system design with two high-level ideas. First, for cost-effective IBR, as long as the synthesized views look visually plausible, the estimated disparity and occlusion need not be correct. Hence, we jointly optimize stereo matching and view synthesis for a favorable end-to-end performance. Second, for great real-time acceleration on GPUs, all functional modules need be shaped at an early design stage, fitting the massively parallel streaming architecture of GPUs. Based on these two guidelines, we first propose a stream-centric local stereo matching algorithm. The key idea is to construct a versatile set of variable support patterns in a highly efficient manner, and then an optimal local support pattern is selected to approximate varying image structures adaptively. Next, a low-complexity adaptive view synthesis technique is proposed. It efficiently tackles visual artifacts in synthesized images, using a novel photometric outlier detection and handling scheme. We evaluated both the disparity estimation accuracy and novel view synthesis quality of the proposed approach, based on the benchmark Middlebury stereo datasets. The experiments show that our local stereo method produces consistently reliable disparity estimates for both homogeneous regions and depth discontinuities, outperforming several previous GPU-based local methods. More importantly, visually plausible intermediate views are generated by our IBR approach at high-speed on the GPU. With stereo matching and view synthesis completely running on an NVIDIA GeForce 8800 GT graphics card, the proposed IBR system reaches about 100 f/s for 450times375 stereo images with 60 disparity levels. Jiangbo Lu, Sammy Rogmans, Gauthier Lafruit, Francky Catthoor |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | Cross-Based Local Stereo Matching Using Orthogonal Integral ImagesabstractWe propose an area-based local stereo matching algorithm for accurate disparity estimation across all image regions. A well-known challenge to local stereo methods is to decide an appropriate support window for the pixel under consideration, adapting the window shape or the pixelwise support weight to the underlying scene structures. Our stereo method tackles this problem with two key contributions. First, for each anchor pixel an upright cross local support skeleton is adaptively constructed, with four varying arm lengths decided on color similarity and connectivity constraints. Second, given the local cross-decision results, we dynamically construct a shape-adaptive full support region on the fly, merging horizontal segments of the crosses in the vertical neighborhood. Approximating image structures accurately, the proposed method is among the best performing local stereo methods according to the benchmark Middlebury stereo evaluation. Additionally, it reduces memory consumption significantly thanks to our compact local cross representation. To accelerate matching cost aggregation performed in an arbitrarily shaped 2-D region, we also propose an orthogonal integral image technique, yielding a speedup factor of 5-15 over the straightforward integration. Ke Zhang 0012, Jiangbo Lu, Gauthier Lafruit |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | Spatial locality exploitation for runtime reordering of JPEG2000 wavelet data layoutsabstractExploitation of spatial locality is essential for memories to increase the access bandwidth and to reduce the access-related latency and energy per word. Spatial locality exploitation of a kernel can be improved by modifying placement of data in memory, but this may be felt not only by the kernel itself, but also in other application components accessing the same data. Thus care is needed to avoid global miss-rate improvements are thwarted by miss-rate increases in other application components. This article examines application-level miss-rate increases due to handling modified Wavelet Transform data layouts by explicitly reordering at runtime, exploiting the execution order freedom within a reordering buffer when the layout of surrounding components is known. For the JPEG2000 application, taking into account the reordering costs still results in 80% net WT miss-rate gains. Bert Geelen, Vissarion Ferentinos, Francky Catthoor, Gauthier Lafruit, Diederik Verkest, Rudy Lauwereins, Thanos Stouraitis |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2008 | Spatial locality trade-offs of wavelet-based applications in dynamic execution environmentsabstractFuture dynamic applications will require new mapping strategies to deliver power-efficient performance. Fully static design-time mappings will not address the unpredictably varying application characteristics and resource requirements. Instead, the platforms will not only need to be programmable in terms of instruction set processors, but also at least partial reconfigurability will be required, while the applications themselves will need to exploit this freedom at run-time to adapt to the dynamism. In this context, it is important for applications to exploit the memory hierarchy under varying memory availability. This paper presents a mapping strategy for wavelet-based applications: depending on the run-time conditions, it switches to different memory optimized instantiations, optimally exploiting temporal and spatial locality under these conditions. A comparison is performed between the gains of fully in-placed lifting-based wavelet transforms, and non in-placed versions with higher spatial locality. Bert Geelen, Aris Ferentinos, Francky Catthoor, Gauthier Lafruit, Diederik Verkest |
ICASSP | 4 |
| 2008 | Scalable stereo matching with Locally Adaptive Polygon ApproximationabstractWe present a scalable stereo matching algorithm based on a Locally Adaptive Polygon Approximation (LAPA) technique. For accurate local stereo matching, pixel-wise adaptive polygon-based support windows are constructed to approximate spatially varying image structures. Central to building these pixel-wise polygons is a fast algorithm that adaptively decides a set of directional scales, utilizing intensity and spatial information. Thanks to the locally adaptive support window, the proposed method achieves high stereo reconstruction quality both in depth-discontinuity regions and homogenous regions. Moreover, our LAPA-based method offers flexible scalability in terms of quality-complexity trade-off. As a specific instantiation favoring high-quality stereo estimation, our 8-direction stereo method outperforms most of the other local stereo methods and even some global optimization techniques. Another low-complexity alternative is also presented, achieving a significant speedup of up to a factor 20 with graceful accuracy degradation. Within a unified LAPA framework, our stereo method hence facilitates more flexibility in conciliating different algorithm design needs with processing performance issues. Ke Zhang 0012, Jiangbo Lu, Gauthier Lafruit |
ICIP | 3 |
| 2008 | Cross-layer optimization for multi-user video streaming over IEEE 802.11E HCCA wireless networksabstractIn this paper, we propose a cross-layer optimization scheme for multi-user video streaming over the uplink of IEEE 802.11e HCCA wireless network. The objective is to minimize the energy consumption of all users, including both video encoder energy and wireless transmission energy, while delivering desired video quality for each user. In our proposed scheme, both cross-layer optimization of individual mobile terminal and inter-user resource allocation are carried out, so that the video encoder and wireless transmitter configurations can be jointly steered. Experimental results show that when compared to the transmission system without cross-layer optimization, our proposed scheme is able to reduce total energy consumption from 35% up to 80%, while satisfying the same video quality target. Tong Gan, Antoine Dejonghe 0001, Gregory Lenoir, Kristof Denolf, Gauthier Lafruit, Iole Moccagatta |
ICME | 5 |
| 2008 | A model for adapting 3D graphics based on scalable coding, real-time simplification and remote rendering
Marius Preda, Paulo Villegas, Francisco Morán, Gauthier Lafruit, Robert-Paul Berretty |
Vis. Comput. | 4 |
| 2007 | Adaptive 3D Content for Multi-Platform On-Line GamesabstractMost current multi-player 3D games can only be played on dedicated platforms, requiring specifically designed content and communication over a prede-fined network. To overcome these limitations, the OLGA (On-Line GAming) consortium has devised a framework to develop distributive, multi-player 3D games. Scalability at the level of content, platforms and networks is exploited to achieve the best trade-offs between complexity and quality. Besides, stan-dardized content compression formats (MPEG-4, JPEG 2000) are used in OLGA's framework, enabling easy deployment over existing infrastructure, while keeping hooks to well-established practices in the game industry. Francisco Morán, Marius Preda, Gauthier Lafruit, Paulo Villegas, Robert-Paul Berretty |
CW | 3 |
| 2007 | Channel-Aware Rate Adaptation for Energy Optimization and Congestion AvoidanceabstractAchieving low energy consumption is one of the main challenges for wireless video transmission on battery limited devices. Moreover, the bandwidth is scarce and needs to be properly shared amongst different users. Congestion in the network can result in packet losses, with a significant impact on video quality. In this paper we propose the use of a channel-adaptive rate control mechanism in a multi-user WLAN up-link scenario. The benefit is twofold: the communication energy is reduced and congestion is strongly alleviated allowing an increase of the video quality or a network capacity increase for a similar quality. Carolina Blanch, Sofie Pollin, Gauthier Lafruit, Antoine Dejonghe 0001, Gregory Lenoir |
ICASSP (1) | 3 |
| 2007 | Energy-Efficient Bandwidth Allocation for Multi-User Video Streaming Over WlanabstractWe consider the problem of packet scheduling for the transmission of multiple video streams over a wireless local area network (WLAN). A cross-layer optimization framework is proposed to minimize the wireless transceiver energy consumption while reaching the user required visual quality. The framework relies on the IEEE 802.11 standard and on a wavelet-based scalable video coding scheme. It extends our previous work on energy-efficient scheduling by introducing an application-level video quality metric as QoS constraint (instead of a quality metric at the level of the communication layers) and by reformulating the energy minimization problem subject to the QoS constraint in order to also consider the fairness among users. Simulation results demonstrate significant additional energy gains by means of these extensions. Xin Ji, Sofie Pollin, Gauthier Lafruit, Iole Moccagatta, Antoine Dejonghe 0001, Francky Catthoor |
ICASSP (2) | 3 |
| 2007 | Fast Reliable Multi-Scale Motion Region Detection in Video ProcessingabstractMotion region detection is an important vision topic usually tackled by a background subtraction principle, which has some practical restrictions. We hence propose a multi-scale motion region detection technique that can fast and reliably segment foreground motion regions from two successive video frames. The key idea is to leverage multi-scale structural aggregation to effectively accentuate real motion changes while suppressing trivial noisy changes. Consequently, this technique can be effectively applied to motion region-of-interest (ROI) based video coding. Our experiments show that the proposed algorithm can reliably extract motion regions and is less sensitive to thresholds than single-scale methods. Compared with a H.264/AVC encoder, the proposed semantic video encoder achieves a bitrate saving ratio of up to 34% at the similar video quality, besides an overall speedup factor of 2.6 to 3.6. The motion-ROI detection can process a 352 × 288 size video at 20 fps on an Intel Pentium 4 processor. Jiangbo Lu, Gauthier Lafruit, Francky Catthoor |
ICASSP (1) | 2 |
| 2007 | Fast Variable Center-Biased Windowing for High-Speed Stereo on Programmable Graphics HardwareabstractWe present a high-speed dense stereo algorithm that achieves both good quality results and very high disparity estimation throughput on the graphics processing unit (GPU). The key idea is a variable center-biased windowing approach, enabling an adaptive selection of the most suitable support patterns with varying sizes and shapes. As the fundamental construct for variable windows, a truncated separable Laplacian kernel approximation is proposed for the efficient pixel-wise weighted cost aggregation. We also present a number of critical optimization schemes to boost the real-time speed on GPUs. Our method outperforms previous GPU-based local stereo methods and even some methods using global optimization on the Middlebury stereo database. Our optimized implementation completely running on an Nvidia GeForce 7900 graphics card achieves over 605 million disparity estimations per second (Mde/s) including all the overhead, about 2.1 to 12.1 times faster than the existing GPU-based solutions. Jiangbo Lu, Gauthier Lafruit, Francky Catthoor |
ICIP (6) | 2 |
| 2007 | Modelling Energy Consumption of an ASIC MPEG-4 Simple Profile EncoderabstractIn wireless video streaming, it is desirable to minimize the energy consumption of mobile devices while achieving target video quality. For this purpose, it is essential to model (i.e., estimate) the energy cost of the streaming system, including both video encoder and wireless transmission energy, under different system settings. In this paper, an energy consumption model is proposed for a MPEG-4 simple profile encoder implemented in ASIC. Our proposed model consists of three major steps: 1) compute the encoder energy consumption through power simulations; 2) empirically model the impact of configuration parameters through curve fitting; and 3) online updating of model coefficients. Experimental results show that our proposed model works well for different types of video, and the average estimation error is below 6%. Tong Gan, Kristof Denolf, Gauthier Lafruit, Iole Moccagatta, Antoine Dejonghe 0001, Gregory Lenoir |
ICME | 3 |
| 2007 | Real-Time Stereo Correspondence using a Truncated Separable Laplacian Kernel Approximation on Graphics HardwareabstractWe present a novel real-time stereo algorithm that achieves both good quality results and very high disparity estimation throughput on the graphics processing unit (GPU). As the key idea of this paper, a truncated separable approximation to an isotropic Laplacian kernel is proposed. This truncated 2D Laplacian kernel variant combines the advantages of large support windows and shiftable windows, while support-weights on geometric proximity can still be appropriately applied to each pixel in truncated support windows. Our method outperforms previous GPU-based local stereo methods and even some methods using global optimization on the benchmark Middlebury stereo database. Because of its separable and regular property, the proposed kernel can be very efficiently implemented on CPUs. Our optimized implementation completely running on an Nvidia GeForce 7900 graphics card achieves over 668 million disparity estimations per second (Mde/s) including all the overhead, about 2.3 to 13.4 times faster than the existing GPU-based solutions. Jiangbo Lu, Sammy Rogmans, Gauthier Lafruit, Francky Catthoor |
ICME | 3 |
| 2007 | High-Speed Stream-Centric Dense Stereo and View Synthesis on Graphics HardwareabstractThis paper presents an efficient image-based rendering system capable of performing online stereo matching and view synthesis at high speed, completely on the graphics processing unit (GPU). Given two rectified stereo images, our algorithm first extracts the disparity map with a stream-centric dense depth estimation approach. For high-quality view synthesis, multi-label masks are then automatically generated to postprocess occlusions and ambiguously estimated regions adaptively. To allow even faster interactive view generation, an alternative forward warping method is also integrated. The experiments show that photorealistic intermediate views of high image quality are yielded by our algorithm. The optimized implementation also provides the state-of-the-art stereo analysis and view synthesis speed, achieving over 47 fps with 450x375 stereo images and 60 disparity levels on an Nvidia GeForce 7900 graphics card. Jiangbo Lu, Sammy Rogmans, Gauthier Lafruit, Francky Catthoor |
MMSP | 3 |
| 2006 | Platform independent optimisation of multi-resolution 3D content to enable universal media access
Klaas Tack, Gauthier Lafruit, Francky Catthoor, Rudy Lauwereins |
Vis. Comput. | 2 |
| 2005 | Memory Centric Design of an MPEG-4 Video EncoderabstractThe cost-efficient implementation of video codecs requires a set of methodologies and decision taking at different levels in the design flow. We combine upfront algorithmic tuning with memory centric optimizations to transform the video application into a system consisting of functional blocks with localized data processing and a tailored memory hierarchy. This memory optimized functional description is the leverage for the cost-efficient mapping of the system on integrated multimedia platforms. It closely reflects the real implementation constraints and consequently allows for steering the architecture selection in a correct way. The proposed approach is demonstrated on a MPEG-4 video encoder and leads to its implementation as a pipelined system. Hardware development of the motion estimation validates that the high-level memory centric concepts are applicable and realizable at the lowest level. The motion estimation kernel supports up to 30 CIF f/s with minimized processing element requirements and data input rates. Kristof Denolf, Christophe De Vleeschouwer, Robert D. Turney, Gauthier Lafruit, Jan Bormans |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2004 | Introduction to the Special Issue on MPEG-4's Animation Framework eXtension (AFX)
Mikaël Bourges-Sévenier, Euee S. Jang, Gauthier Lafruit, Francisco Morán |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2004 | View-dependent, scalable texture streaming in 3-D QoS with MPEG-4 visual texture codingabstractMultimedia applications are characterized by high resource demands (computing power, memory, network bandwidth, and power consumption). Efficient implementations aim at reducing these resources to a minimum, which is of the utmost importance for small, low-cost terminals in low-bandwidth networks. Resource savings can also be obtained by content adaptation without impeding the quality of the decoded audio-visual media. In this context, the paper analyzes texture adaptation and streaming for three-dimensional applications, using MPEG-4's Visual Texture Coding tool, in conjunction with eXtensible Markup Language (XML)-based description techniques. Augmented features for content adaptation are supported, such as region selection, accompanied by resolution and SNR settings. As a result, quality is optimized for the terminal's computing capabilities and display resolution, taking the user's viewing conditions into account. Moreover, the instantaneous bandwidth utilization is highly reduced in streaming scenarios. Gauthier Lafruit, Eric Delfosse, Roberto R. Osorio, Wolfgang van Raemdonck, Vissarion Ferentinos |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2004 | MESHGRID-a compact, multiscalable and animation-friendly surface representationabstractMESHGRID is a novel, compact, multiscalable and animation-friendly surface representation method, which has been introduced in MPEG-4 . The MESHGRID representation attaches a description of the "global connectivity" between the vertices on the object's surface (i.e., the 3-D connectivity wireframe) to a regular 3-D grid of points (i.e., the reference grid). MESHGRID efficiently encodes the 3-D connectivity wireframe by using a new type of 3-D extension of Freeman chain-code. MESHGRID does not explicitly store the polygons of the surface, since the 3-D connectivity wireframe has particular connectivity properties allowing for the unambiguous derivation of the triangulation. The reference grid is a smooth vector field defined on a regular discrete 3-D space. This grid is efficiently compressed by using an embedded 3-D wavelet-based multiresolution intra-band coding algorithm. MESHGRID can be efficiently exploited for QoS since it allows for three types of scalability in both view-dependent and view-independent scenarios, including: 1) resolution scalability, i.e., the adaptation of the number of transmitted vertices; 2) shape precision, i.e., the adaptive reconstruction of the reference grid positions; and 3) vertex position scalability, i.e., the change of the precision of known vertex positions with respect to the reference grid. Furthermore, in addition to the classical vertex-based animation, MESHGRID also supports specific animation capabilities, such as: 1) rippling effects by changing the position of the vertices relative to corresponding reference grid points and 2) reshaping on a hierarchical basis of the regular reference grid and its attached vertices. Ioan Alexandru Salomie, Adrian Munteanu 0001, Augustin Gavrilescu, Gauthier Lafruit, Peter Schelkens, Rudi Deklerck, Jan Cornelis 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2003 | A framework for mapping scalable networked applications on run-time reconfigurable platformsabstractDue to the heterogeneous nature of networks and end-systems in distributed multimedia systems, multimedia applications should ideally be designed to counteract fluctuations in network bandwidth and end-system processing capacities for providing end users a certain degree of quality of service (QoS). This requirement can be satisfied with scalable applications. In addition, with the current evolution in run-time reconfigurable computing, run-time reconfigurable multimedia platforms are becoming increasingly viable. In this paper, an end-to-end delivery chain framework for mapping scalable networked multimedia applications on reconfigurable platforms is presented. The framework is demonstrated by a case study of a 3D game running on a prototype run-time reconfigurable platform. Nam Pham Ngoc 0001, Gauthier Lafruit, Jean-Yves Mignolet, Serge Vernalde, Geert Deconinck, Rudy Lauwereins |
ICME | 2 |
| 2003 | Terminal QoS for real-time 3-D visualization using scalable MPEG-4 codingabstractTerminal quality of service is the process of optimally scaling down the decoding and rendering computations to the available processing power, while maximizing the overall perceived quality. This process is applied in a real-time three-dimensional (3-D) decoding and rendering engine, exploiting scalable MPEG-4 coding algorithms. We derive a relation between the quality of the 3-D rendered objects and their processing requirements, expressed by simple CPU time models. The performance dependency parameters are in direct relation to the high-level 3-D content characteristics (number of triangles, number of rendered screen pixels per object) and are calibrated for the platform under test. Examples show that, for any pre-established frame rate, the platform workload can reliably be anticipated for adjusting the content and process parameters accordingly. Gauthier Lafruit, Nam Pham Ngoc 0001, Wolfgang van Raemdonck, Nicolaas Tack, Jan Bormans |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2003 | A scalable MPEG-4 wavelet-based visual texture compression system with optimized memory organizationabstractThe realization of new MPEG-4 functionality, applicable to three-dimensional graphics texture compression and image database access over the Internet, is demonstrated on a heterogeneous platform with several unique features. First, applying our system-level design methodologies effectively removes all data transfer and storage overhead that comprises the main bottleneck in the original system description. Second, a first-of-a-kind application specific solution, called Ozone, accelerates the embedded-zero-tree based encoding and is capable of compressing 30 color CIF images per second. The entire application is running on the Ozone coupled to a PC. Bart Vanhoof, Lode Nachtergaele, Gauthier Lafruit, Mercedes Peón, Bart Masschelein, Francky Catthoor, Jan Bormans, Ivo Bolsens |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2002 | Streaming MPEG-4 textures: A 3D view-dependent approachabstractOne of the key functionalities in today's multimedia applications is streaming. However unlike video, streaming 3D content remains rare, especially for textures. In this paper, we present an intelligent, view-dependent approach for streaming and decoding MPEG-4 textures in order to save network and platform resources, while minimizing the perceived quality loss. Different regions of the 3D object's texture are independently decoded up to a desired quality. Whenever viewing conditions change, newly visible texture regions are transmitted and/or already transmitted regions are refined. This 3D view-dependent decoding of MPEG-4 textures is possible by exploiting the error resilience feature in the MPEG-4 texture-coding tool. It allows a high flexibility (i.e. number of possible quality levels) but at the cost of an increased bitstream size. However, since streaming is possible, the transmission of the bitstream is spread over time, reducing the instantaneous bandwidth and processing cost. Eric Delfosse, Gauthier Lafruit, Jan Bormans |
ICASSP | 2 |
| 2002 | Bitstream Syntax Description Language for 3D MPEG-4 view-dependent texture streamingabstractIn modern multimedia applications, scalability is a key functionality that allows transmission and representation of content in a wide variety of networks and terminals. In order to obtain full advantage of scalability features, techniques for detailed description and transformation of multimedia contents are needed. In this paper, the Bitstream Syntax Description Language is used for describing the structure of an MPEG-4 wavelet-coded texture. An XML-based bitstream description transformation is then applied for selecting some texture regions at an appropriate quality, effectively scaling down the processing and bandwidth requirements for view-dependent texture transmission. Appropriately applying this technique to 3D streaming guarantees quality-of-service, i.e. it certifies the best quality at limited network/processing resources. Roberto R. Osorio, Sylvain Devillers, Eric Delfosse, Myriam Amielh, Gauthier Lafruit |
ICIP (3) | 5 |
| 2002 | MESHGRID - a compact, multi-scalable and animation-friendly surface representationabstractMESHGRID is a novel, compact, multi-scalable and animation-friendly surface representation method, which has been introduced in MPEG-4. The MESHGRID representation attaches a description of the "global connectivity" between the vertices on the object's surface (i.e. the 3D connectivity wireframe) to a regular 3D grid of points (i.e. the reference-grid). The 3D connectivity wireframe is efficiently encoded by using a new type of 3D extension of Freeman chain-code. MESHGRID does not explicitly store the polygons of the surface, since the 3D connectivity wireframe has particular connectivity properties allowing for the unambiguous derivation of the triangulation. The reference-grid is a smooth vector field defined on a regular discrete 3D space. This grid is efficiently compressed by using an embedded 3D wavelet-based multi-resolution intra-band coding algorithm. MESHGRID allows for three types of scalability in both view-dependent and view-independent scenarios: resolution scalability, shape precision, and vertex position scalability. Furthermore, in addition to the classical vertex-based animation, MESHGRID supports specific animation capabilities, such as rippling effects and reshaping on a hierarchical basis of the regular reference-grid and its attached vertices. Ioan Alexandru Salomie, Adrian Munteanu 0001, Augustin Gavrilescu, Gauthier Lafruit, Peter Schelkens, Rudi Deklerck, Jan Cornelis 0001 |
ICIP (3) | 4 |
| 2002 | Scalable 3D graphics processing in consumer terminalsabstractIn this paper, we present how scalable 3D graphics fits into a quality-of-service resource management approach for high volume electronics consumer terminals. Resource management is based on guaranteed and enforced resource budgets, which are allocated by a global quality manager. A specific 3D QoS manager makes sure that a 3D graphics service performs acceptably within the limitations of its budget. Based on models that estimate the quality and workload of functional tasks, such as rendering and decoding, for 3D objects, the scaling parameters of these tasks are adapted to guarantee real-time operation, possibly at a slightly degraded perceived quality. The validity of the approach is demonstrated with a 3D graphics application running in conjunction with video decoding on a Philips set-top box containing a TriMedia/spl trade/ processor. Wolfgang van Raemdonck, Gauthier Lafruit, Elisabeth F. M. Steffens, Clara Otero Pérez, Reinder J. Bril |
ICME (1) | 2 |
| 2001 | A local wavelet transform implementation versus an optimal row-column algorithm for the 2D multilevel decompositionabstractA new method for the implementation of the binary-tree decomposition of the convolution-based wavelet transform, called the local wavelet transform (LWT) has been recently proposed in the literature. While it produces exactly the same results as the classical row-column implementation of the transform, it has many implementation benefits. This fact is shown experimentally for the first time for a general-purpose processor-based architecture, by comparing our C implementation of the LWT with an optimal C implementation of the lifting-scheme row-column algorithm. The comparisons are made for the forward multilevel binary-tree decomposition using the 9/7 filter pair, in the typical Intel Pentium processor family. Yiannis Andreopoulos, Nikolaos D. Zervas, Gauthier Lafruit, Peter Schelkens, Thanos Stouraitis, Constantinos E. Goutis, Jan Cornelis 0001 |
ICIP (3) | 3 |
| 2000 | 3D computational graceful degradationabstractinfo:eu-repo/semantics/published Gauthier Lafruit, Lode Nachtergaele, Kristof Denolf, Jan Bormans |
ISCAS | 1 |
| 1999 | Implementation of a Scalable MPEG-4 Wavelet-Based Visual Texture Compression SystemabstractThe realization of new MPEG-4 functionality, applicable to 3D graphics texture compression and image database access over the Internet, is demonstrated in a PC-based compression system.Applying our system-level design methodologies effectively removes all implementation bottlenecks.A first-of-a-kind ASIC, called Ozone, accelerates the Embedded Zero Tree based encoding and is capable of compressing 30 color CIF images per second. Lode Nachtergaele, Bart Vanhoof, Mercedes Peón, Gauthier Lafruit, Jan Bormans, Ivo Bolsens |
DAC | 4 |
| 1999 | Optimal memory organization for scalable texture codecs in MPEG-4abstractThis paper addresses the problem of minimizing memory size and memory accesses in multiresolution texture coding architectures for discrete cosine transform (DCT) and wavelet-based schemes used, for example, in virtual-world walk-throughs or facial animation scenes of an MPEG-4 system. The problem of minimizing the memory cost is important since memory accesses, memory bandwidth limitations, and in general the correct handling of the data flows have become the true critical issues in designing high-speed and low-power video-processing architectures and in efficiently using multimedia processors. For instance, the straightforward implementation of a multiresolution texture codec typically needs an extra memory buffer of the same size as the image to be encoded/decoded. We propose a new calculation schedule that reduces this buffer memory size with up to two orders of magnitude, while still ensuring a number of external (off-chip) memory accesses that is very close to the theoretical minimum. The analysis is generic and is therefore useful for both wavelet and multiresolution DCT codecs. Gauthier Lafruit, Lode Nachtergaele, Jan Bormans, Marc Engels, Ivo Bolsens |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1999 | An efficient VLSI architecture for 2-D wavelet image coding with novel image scanabstractA folded very large scale integration (VLSI) architecture is presented for the implementation of the two-dimensional discrete wavelet transform, without constraints on the choice of the wavelet-filter bank. The proposed architecture is dedicated to flexible block-oriented image processing, such as adaptive vector quantization used in wavelet image coding. We show that reading the image along a two-dimensional (2-D) pseudo-fractal scan creates a very modular and regular data flow and, therefore, considerably reduces the folding complexity and memory requirements for VLSI implementation. This leads to significant area savings for on-chip storage (up to a factor of two) and reduces the power consumption. Furthermore, data scheduling and memory management remain very simple. The end result is an efficient VLSI implementation with a reduced area cost compared to the conventional approaches, reading the input data line by line. Gauthier Lafruit, Francky Catthoor, Jan Cornelis 0001, Hugo De Man |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1998 | Video coding based on motion estimation in the wavelet detail imagesabstractThis work proposes a new block based motion estimation and compensation technique applied on the detail images of the wavelet pyramidal decomposition. The algorithm uses two matching criteria, namely the absolute difference and the absolute sum. For a wavelet decomposed one-dimensional step function, it is shown that for odd translations of the step, the absolute sum reaches a smaller minimum than the absolute difference. We also derive in this case a constraint on the highpass filter coefficients so that a zero prediction error can be reached by using the absolute sum. Although this cannot be easily generalized for an arbitrary signal profile, experimental results obtained with photorealistic image sequences indicate that the prediction error can be reduced with respect to techniques that only use the absolute difference as matching criterion. Geert Van der Auwera, Adrian Munteanu 0001, Gauthier Lafruit, Jan Cornelis 0001 |
ICASSP | 3 |