Roger Olsson

dblp:44/521 · DBLP profile ↗
← Back
25ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Deep Semantic Inference Over the Air: An Efficient Task-Oriented Communication System
abstract
Empowered by deep learning, semantic communication marks a paradigm shift from transmitting raw data to conveying task-relevant meaning, enabling more efficient and intelligent wireless systems. In this study, we explore a deep learning-based task-oriented communication framework that jointly considers classification performance, computational latency, and communication cost. We evaluate ResNets-based models on the CIFAR-10 and CIFAR-100 datasets to simulate real-world classification tasks in wireless environments. We partition the model at various points to simulate split inference across a wireless channel. By varying the split location and the size of the transmitted semantic feature vector, we systematically analyze the trade-offs between task accuracy and resource efficiency. Experimental results show that, with appropriate model partitioning and semantic feature compression, the system can retain over 85\% of baseline accuracy while significantly reducing both computational load and communication overhead.
Chenyang Wang 0004, Roger Olsson, Stefan Forsström
WCNC2
2026 DUALF-D: Disentangled dual-hyperprior approach for light field image compression
abstract
Light field (LF) imaging captures spatial and angular information, offering a 4D scene representation enabling enhanced visual understanding. However, high dimensionality and redundancy across spatial and angular domains present major challenges for compression, particularly where storage, transmission bandwidth, or processing latency are constrained. We present a novel Variational Autoencoder (VAE)-based framework that explicitly disentangles spatial and angular features using two parallel latent branches. Each branch is coupled with an independent hyperprior model, allowing more precise distribution estimation for entropy coding and finer rate–distortion control. This dual-hyperprior structure enables the network to adaptively compress spatial and angular information based on their unique statistical characteristics, improving coding efficiency. To further enhance latent feature specialization and promote disentanglement, we introduce a mutual information-based regularization term that minimizes redundancy between the two branches while preserving feature diversity. Unlike prior methods relying on covariance-based penalties prone to collapse, our information-theoretic regularizer provides more stable and interpretable latent separation. Experimental results on publicly available LF datasets demonstrate our method achieves strong compression performance, yielding an average BD-PSNR gain of 2.91 dB over HEVC and high compression ratios (e.g., 200:1). Additionally, our design enables fast inference, with a total end-to-end time over 19x faster than the JPEG Pleno standard, making it well-suited for real-time and bandwidth-sensitive applications. By jointly leveraging disentangled representation learning, dual-hyperprior modeling, and information-theoretic regularization, our approach offers a scalable, effective solution for practical light field image compression.
Soheib Takhtardeshir, Roger Olsson, Christine Guillemot, Mårten Sjöström
Signal Process. Image Commun.2
2025 Subjective Visual Quality Assessment of Compressed Light Field Images: Learning-based vs. Conventional Methods
abstract
Light fields (LF) technology enables the capture and reproduction of a 3D scene accurately, which enhances visual experience in various applications. The sheer volume of multiplicity of the captured views creates logistical problems in both storage and data transmission, which makes LF compression crucial. Even though many LF compression techniques have been proposed and evaluated in recent years, the leading-edge learning-based approaches have not been subject to the same level of scrutiny. This paper presents a subjective quality assessment study on four different LF compression methods, including two learning-based LF compression methods, which have not been studied before from a subjective quality point of view. For this purpose, subjective opinion scores were collected from viewers in two different universities for a cross-lab study. The results indicate that the learning-based compression methods have different behavior in their rate-distortion curves, and that there is room for improvement for learning-based methods. A qualitative analysis also shows that their artifact structures are different from conventional ones. The results highlight the need for a perceptual objective quality metric that takes different types of artifacts into account. The obtained subjective quality database (MiX-LFQDB) is made public to support further research in this area: https://doi.org/10.5281/zenodo.16778670
Emin Zerman, Soheib Takhtardeshir, Anthony Trioux, Jianlong Qin, Roger Olsson, Mårten Sjöström
MMSP6
2025 Exploring Peripheral Visualization for Enhanced Manual Welding Quality
abstract
Despite the rise of robotic automation, manual welding remains essential in modern manufacturing where flexibility and precision are required, yet maintaining consistent quality remains a challenge. This paper presents a prototype system that uses peripheral visualization to deliver real-time, non-foveal feedback on travel speed via dynamic LED cues embedded in a welding visor. The system translates welding movement into peripheral visual cues designed to support focus and reduce visual distraction. In tests with both novice and expert welders, the feedback improved speed consistency by over 30%, indicating strong potential for enhancing task performance and training through subtle, low-cognitive-load feedback. While tailored to welding, the findings highlight peripheral visualization as a promising strategy for improving quality of experience in attention-critical AR tasks.
Roger Olsson, Abdulrahman Kanaa, Joakim Wahlsten
QoMEX1
2025 DUALF-C: Disentangled Light Field Compression with Entropy-Aware Bitstream Generation
abstract
Recent advancements in learned light field (LF) image compression highlight the advantages of modeling spatial and angular redundancies using deep generative models. Among these, Variational Autoencoders (VAEs) have shown strong potential in learning compact latent representations of LF data. However, many existing approaches rely on simple uniform quantization and basic entropy coding, which limits compression efficiency in practical applications. This paper introduces DUALF-C, a lightweight and retraining-free compression pipeline that augments pretrained VAE-based LF compression models using structured post-encoding transformations. The proposed framework integrates bitplane slicing, latent channel reordering, non-uniform quantization, and patch-based vector quantization to improve bitrate efficiency while preserving reconstruction quality. Experimental evaluations demonstrate that DUALF-C significantly reduces bit-per-pixel (BPP) without degrading image quality, making it a practical solution for bandwidth-constrained immersive imaging systems.
Soheib Takhtardeshir, Roger Olsson, Christine Guillemot, Mårten Sjöström
VCIP2
2025 3D-Gaussian Splatting Representation of Rendered Views from Plenoptic 2.0 Lenslet Images
abstract
Thanks to plenoptic cameras, rich information on the radiance of a scene can be conveniently captured without heavy devices like camera arrays. However, rendering techniques are needed to generate views for human visual perception. Existing patch extraction-based rendering techniques can generate views from the lenslet images captured by plenoptic cameras, but they suffer from the inherent problem of artifacts and the limited views to be rendered. In this paper, we present a new view rendering technique from plenoptic 2.0 camera-captured lenslet image by using 3-dimensional gaussian splatting (3DGS). At its first step, the reference lenslet converter (RLC) provided by MPEG LVC AhG, one of the existing patch extraction methods, generates initial views with the help of estimated disparity between adjacent micro images in the lenslet image. At its second step, the 3DGS generates the final views after being trained by the initial views. The rendering results obtained by the proposed 2-step approach show significantly fewer artifacts in the rendered views than the patch stitching process of the existing method.
Jonghoon Yim, Byeungwoo Jeon, Roger Olsson, Mårten Sjöström
VCIP3
2024 A Spherical Light Field Database for Immersive Telecommunication and Telepresence Applications
abstract
Immersive imaging technologies provide an enhanced user experience for visual applications and are getting ready for commonplace use by the industry and the general populace. In particular, light field is a promising technology that enables the capture and reproduction of real light rays from the scene, which can provide a backbone for immersive telecommunication and telepresence applications. Nevertheless, there are still many challenges in transmitting and reproducing light field data. This paper proposes a spherical light field dataset that can be used as a foundation for developing telepresence applications. The Spherical Light Field Database (SLFDB) consists of a light field of 60 views captured with an omnidirectional camera in 20 scenes. To show the usefulness of the proposed database, we provide two use cases: compression and viewpoint estimation. The initial results validate that the publicly available SLFDB will benefit the scientific community.
Emin Zerman, Manu Gond, Soheib Takhtardeshir, Roger Olsson, Mårten Sjöström
QoMEX4
2020 Shearlet Transform-Based Light Field Compression Under Low Bitrates
abstract
Light field (LF) acquisition devices capture spatial and angular information of a scene. In contrast with traditional cameras, the additional angular information enables novel postprocessing applications, such as 3D scene reconstruction, the ability to refocus at different depth planes, and synthetic aperture. In this paper, we present a novel compression scheme for LF data captured using multiple traditional cameras. The input LF views were divided into two groups: key views and decimated views. The key views were compressed using the multi-view extension of high-efficiency video coding (MV-HEVC) scheme, and decimated views were predicted using the shearlet-transform-based prediction (STBP) scheme. Additionally, the residual information of predicted views was also encoded and sent along with the coded stream of key views. The proposed scheme was evaluated over a benchmark multi-camera based LF datasets, demonstrating that incorporating the residual information into the compression scheme increased the overall peak signal to noise ratio (PSNR) by 2 dB. The proposed compression scheme performed significantly better at low bit rates compared to anchor schemes, which have a better level of compression efficiency in high bit-rate scenarios. The sensitivity of the human vision system towards compression artifacts, specifically at low bit rates, favors the proposed compression scheme over anchor schemes.
Waqas Ahmad 0002, Suren Vagharshakyan, Mårten Sjöström, Atanas P. Gotchev, Robert Bregovic, Roger Olsson
IEEE Trans. Image Process.6
2018 Shearlet Transform Based Prediction Scheme for Light Field Compression
abstract
Light field acquisition technologies capture angular and spatial information of the scene. The spatial and angular information enables various post processing applications, e.g. 3D scene reconstruction, refocusing, synthetic aperture etc at the expense of an increased data size. In this paper, we present a novel prediction tool for compression of light field data acquired with multiple camera system. The captured light field (LF) can be described using two plane parametrization as, L(u, v, s, t), where (u, v) represents each view image plane coordinates and (s, t) represents the coordinates of the capturing plane. In the proposed scheme, the captured LF is uniformly decimated by a factor d in both directions (in s and t coordinates), resulting in a sparse set of views also referred to as key views. The key views are converted into a pseudo video sequence and compressed using high efficiency video coding (HEVC). The shearlet transform based reconstruction approach, presented in [1], is used at the decoder side to predict the decimated views with the help of the key views. Four LF images (Truck, Bunny from Stanford dataset, Set2 and Set9 from High Density Camera Array dataset) are used in the experiments. Input LF views are converted into a pseudo video sequence and compressed with HEVC to serve as anchor. Rate distortion analysis shows the average PSNR gain of 0.98 dB over the anchor scheme. Moreover, in low bit-rates, the compression efficiency of the proposed scheme is higher compared to the anchor and on the other hand the performance of the anchor is better in high bit-rates. Different compression response of the proposed and anchor scheme is a consequence of their utilization of input information. In the high bit-rate scenario, high quality residual information enables the anchor to achieve efficient compression. On the contrary, the shearlet transform relies on key views to predict the decimated views without incorporating residual information. Hence, it has inherit reconstruction error. In the low bit-rate scenario, the bit budget of the proposed compression scheme allows the encoder to achieve high quality for the key views. The HEVC anchor scheme distributes the same bit budget among all the input LF views that results in degradation of the overall visual quality. The sensitivity of human vision system toward compression artifacts in low-bit-rate cases favours the proposed compression scheme over the anchor scheme.
Waqas Ahmad 0002, Suren Vagharshakyan, Mårten Sjöström, Atanas P. Gotchev, Robert Bregovic, Roger Olsson
DCC6
2018 Towards a Generic Compression Solution for Densely and Sparsely Sampled Light Field Data
abstract
Light field (LF) acquisition technologies capture the spatial and angular information present in scenes. The angular information paves the way for various post-processing applications such as scene reconstruction, refocusing, and synthetic aperture. The light field is usually captured by a single plenop-tic camera or by multiple traditional cameras. The former captures a dense LF, while the latter captures a sparse LF. This paper presents a generic compression scheme that efficiently compresses both densely and sparsely sampled LFs. A plenoptic image is converted into sub-aperture images, and each sub-aperture image is interpreted as a frame of a multiview sequence. In comparison, each view of the multi-camera system is treated as a frame of a multi-view sequence. The multi-view extension of high efficiency video coding (MV-HEVC) is used to encode the pseudo multi-view sequence. This paper proposes an adaptive prediction and rate allocation scheme that efficiently compresses LF data irrespective of the acquisition technology used.
Waqas Ahmad 0002, Roger Olsson, Mårten Sjöström
ICIP2
2017 Interpreting plenoptic images as multi-view sequences for improved compression
abstract
Over the last decade, advancements in optical devices have made it possible for new novel image acquisition technologies to appear. Angular information for each spatial point is acquired in addition to the spatial information of the scene that enables 3D scene reconstruction and various post-processing effects. Current generation of plenoptic cameras spatially multiplex the angular information, which implies an increase in image resolution to retain the level of spatial information gathered by conventional cameras. In this work, the resulting plenoptic image is interpreted as a multi-view sequence that is efficiently compressed using the multi-view extension of high efficiency video coding (MV-HEVC). A novel two-dimensional weighted prediction and rate allocation scheme is proposed to adopt the HEVC compression structure to the plenoptic image properties. The proposed coding approach is a response to ICIP 2017 Grand Challenge: Light field Image Coding. The proposed scheme outperforms all ICME-contestants, and improves on the JPEG-anchor of ICME with an average PSNR gain of 7.5 dB and the HEVC-anchor of ICIP 2017 Grand Challenge with an average PSNR gain of 2.4 dB.
Waqas Ahmad 0002, Roger Olsson, Mårten Sjöström
ICIP2
2016 SMART: a light field image quality dataset
abstract
In this contribution, the design of a Light Field image dataset is presented. It can be useful for design, testing, and benchmarking Light Field image processing algorithms. As first step, image content selection criteria have been defined based on selected image quality key-attributes, i.e. spatial information, colorfulness, texture key features, depth of field, etc. Next, image scenes have been selected and captured by using the Lytro Illum Light Field camera. Performed analysis shows that the proposed set of images is sufficient for addressing a wide range of attributes relevant for assessing Light Field image quality.
Pradip Paudyal, Roger Olsson, Mårten Sjöström, Federica Battisti, Marco Carli
MMSys2
2016 Virtual view synthesis using layered depth image generation and depth-based inpainting for filling disocclusions and translucent disocclusions
Suryanarayana Murthy Muddala, Mårten Sjöström, Roger Olsson
J. Vis. Commun. Image Represent.3
2016 Coding of Focused Plenoptic Contents by Displacement Intra Prediction
abstract
A light field is commonly described by a two-plane representation with four dimensions. Refocused 3D contents can be rendered from light field images. A method for capturing these images is using cameras with microlens arrays. A dense sampling of the light field results in large amounts of redundant data. Therefore, an efficient compression is vital for a practical use of these data. In this paper, we propose a displacement intra prediction scheme with a maximum of two hypotheses for the compression of plenoptic contents from focused plenoptic cameras. The proposed scheme is further implemented into High Efficiency Video Coding (HEVC). The work is aiming at efficiently coding plenoptic captured contents without knowing underlying camera geometries. In addition, the theoretical analysis of the displacement intra prediction for plenoptic images is explained; the relationship between the compressed captured images and their rendered quality is also analyzed. Evaluation results show that plenoptic contents can be efficiently compressed by the proposed scheme. Bit rate reduction up to 60% over HEVC is obtained for plenoptic images, and more than 30% is achieved for the tested video sequences.
Yun Li 0004, Mårten Sjöström, Roger Olsson, Ulf Jennehag
IEEE Trans. Circuits Syst. Video Technol.3
2016 Scalable Coding of Plenoptic Images by Using a Sparse Set and Disparities
abstract
One of the light field capturing techniques is the focused plenoptic capturing. By placing a microlens array in front of the photosensor, the focused plenoptic cameras capture both spatial and angular information of a scene in each microlens image and across microlens images. The capturing results in a significant amount of redundant information, and the captured image is usually of a large resolution. A coding scheme that removes the redundancy before coding can be of advantage for efficient compression, transmission, and rendering. In this paper, we propose a lossy coding scheme to efficiently represent plenoptic images. The format contains a sparse image set and its associated disparities. The reconstruction is performed by disparity-based interpolation and inpainting, and the reconstructed image is later employed as a prediction reference for the coding of the full plenoptic image. As an outcome of the representation, the proposed scheme inherits a scalable structure with three layers. The results show that plenoptic images are compressed efficiently with over 60 percent bit rate reduction compared with High Efficiency Video Coding intra coding, and with over 20 percent compared with an High Efficiency Video Coding block copying mode.
Yun Li 0004, Mårten Sjöström, Roger Olsson, Ulf Jennehag
IEEE Trans. Image Process.3
2015 Depth and angular resolution in plenoptic cameras
abstract
We present a model-based approach to extract the depth and angular resolution in a plenoptic camera. Obtained results for the depth and angular resolution are validated against Ze-max ray tracing results. The provided model-based approach gives the location and number of the resolvable depth planes in a plenoptic camera as well as the angular resolution with regards to disparity in pixels. The provided model-based approach is straightforward compared to practical measurements and can reflect on the plenoptic camera parameters such as the microlens f-number in contrast with the principal-ray-model approach. Easy and accurate quantification of different resolution terms forms the basis for designing the capturing setup and choosing a reasonable system configuration for plenoptic cameras. Results from this work will accelerate customization of the plenoptic cameras for particular applications without the need for expensive measurements.
Mitra Damghanian, Roger Olsson, Mårten Sjöström
ICIP2
2015 Coding of plenoptic images by using a sparse set and disparities
abstract
A focused plenoptic camera not only captures the spatial information of a scene but also the angular information. The capturing results in a plenoptic image consisting of multiple microlens images and with a large resolution. In addition, the microlens images are similar to their neighbors. Therefore, an efficient compression method that utilizes this pattern of similarity can reduce coding bit rate and further facilitate the usage of the images. In this paper, we propose an approach for coding of focused plenoptic images by using a representation, which consists of a sparse plenoptic image set and disparities. Based on this representation, a reconstruction method by using interpolation and inpainting is devised to reconstruct the original plenoptic image. As a consequence, instead of coding the original image directly, we encode the sparse image set plus the disparity maps and use the reconstructed image as a prediction reference to encode the original image. The results show that the proposed scheme performs better than HEVC intra with more than 5 dB PSNR or over 60 percent bit rate reduction.
Yun Li 0004, Mårten Sjöström, Roger Olsson
ICME3
2014 Performance analysis in Lytro camera: Empirical and model based approaches to assess refocusing quality
abstract
In this paper we investigate the performance of Lytro camera in terms of its refocusing quality. The refocusing quality of the camera is related to the spatial resolution and the depth of field as the contributing parameters. We quantify the spatial resolution profile as a function of depth using empirical and model based approaches. The depth of field is then determined by thresholding the spatial resolution profile. In the model based approach, the previously proposed sampling pattern cube (SPC) model for representation and evaluation of the plenoptic capturing systems is utilized. For the experimental resolution measurements, camera evaluation results are extracted from images rendered by the Lytro full reconstruction rendering method. Results from both the empirical and model based approaches assess the refocusing quality of the Lytro camera consistently, highlighting the usability of the model based approaches for performance analysis of complex capturing systems.
Mitra Damghanian, Roger Olsson, Mårten Sjöström
ICASSP2
2014 Efficient intra prediction scheme for light field image compression
abstract
Interactive photo-realistic graphics can be rendered by using light field datasets. One way of capturing the dataset is by using light field cameras with microlens arrays. The captured images contain repetitive patterns resulted from adjacent mi-crolenses. These images don't resemble the appearance of a natural scene. This dissimilarity leads to problems in light field image compression by using traditional image and video encoders, which are optimized for natural images and video sequences. In this paper, we introduce the full inter-prediction scheme in HEVC into intra-prediction for the compression of light field images. The proposed scheme is capable of performing both unidirectional and bi-directional prediction within an image. The evaluation results show that above 3 dB quality improvements or above 50 percent bit-rate saving can be achieved in terms of BD-PSNR for the proposed scheme compared to the original HEVC intra-prediction for light field images.
Yun Li 0004, Mårten Sjöström, Roger Olsson, Ulf Jennehag
ICASSP3
2014 Spatial resolution in a multi-focus plenoptic camera
abstract
Evaluation of the state of the art plenoptic cameras is necessary for design and application purposes. In this work, spatial resolution is investigated in a multi-focus plenoptic camera using two approaches: empirical and model-based. The Raytrix R29 plenoptic camera is studied which utilizes three types of micro lenses with different focal lengths in a hexagonal array structure to increase the depth of field. The modelbased approach utilizes the previously proposed sampling pattern cube (SPC) model for representation and evaluation of the plenoptic capturing systems. For the experimental resolution measurements, spatial resolution values are extracted from images reconstructed by the provided Raytrix reconstruction method. Both the measurement and the SPC model based approaches demonstrate a gradual variation of the resolution values in a wide depth range for the multi focus R29 camera. Moreover, the good agreement between the results from the model-based approach and those from the empirical approach confirms suitability of the SPC model in evaluating high-level camera parameters such as the spatial resolution in a complex capturing system as R29 multi-focus plenoptic camera.
Mitra Damghanian, Roger Olsson, Mårten Sjöström, A. Erdmann, Christian Perwass
ICIP2
2014 A Weighted Optimization Approach to Time-of-Flight Sensor Fusion
abstract
Acquiring scenery depth is a fundamental task in computer vision, with many applications in manufacturing, surveillance, or robotics relying on accurate scenery information. Time-of-flight cameras can provide depth information in real-time and overcome short-comings of traditional stereo analysis. However, they provide limited spatial resolution and sophisticated upscaling algorithms are sought after. In this paper, we present a sensor fusion approach to time-of-flight super resolution, based on the combination of depth and texture sources. Unlike other texture guided approaches, we interpret the depth upscaling process as a weighted energy optimization problem. Three different weights are introduced, employing different available sensor data. The individual weights address object boundaries in depth, depth sensor noise, and temporal consistency. Applied in consecutive order, they form three weighting strategies for time-of-flight super resolution. Objective evaluations show advantages in depth accuracy and for depth image based rendering compared with state-of-the-art depth upscaling. Subjective view synthesis evaluation shows a significant increase in viewer preference by a factor of four in stereoscopic viewing conditions. To the best of our knowledge, this is the first extensive subjective test performed on time-of-flight depth upscaling. Objective and subjective results proof the suitability of our approach to time-of-flight super resolution approach for depth scenery capture.
Sebastian Schwarz, Mårten Sjöström, Roger Olsson
IEEE Trans. Image Process.3
2012 The Sampling Pattern Cube - A Representation and Evaluation Tool for Optical Capturing Systems
Mitra Damghanian, Roger Olsson, Mårten Sjöström
ACIVS2
2012 Adaptive depth filtering for HEVC 3D video coding
abstract
Consumer interest in 3D television (3DTV) is growing steadily, but current available 3D displays still need additional eye-wear and suffer from the limitation of a single stereo view pair. So it can be assumed that autostereoscopic multiview displays are the next step in 3D-at-home entertainment, since these displays can utilize the Multiview Video plus Depth (MVD) format to synthesize numerous viewing angles from only a small set of given input views. This motivates efficient MVD compression as an important keystone for commercial success of 3DTV. In this paper we concentrate on the compression of depth information in an MVD scenario. There have been several publications suggesting depth down- and upsampling to increase coding efficiency. We follow this path, using our recently introduced Edge Weighted Optimization Concept (EWOC) for depth upscaling. EWOC uses edge information from the video frame in the upscaling process and allows the use of sparse, non-uniformly distributed depth values. We exploit this fact to expand the depth down-/upsampling idea with an adaptive low-pass filter, reducing high energy parts in the original depth map prior to subsampling and compression. Objective results show the viability of our approach for depth map compression with up-to-date High-Efficiency Video Coding (HEVC). For the same Y-PSNR in synthesized views we achieve up to 18.5% bit rate decrease compared to full-scale depth and around 10% compared to competing depth down-/upsampling solutions. These results were confirmed by a subjective quality assessment, showing a statistical significant preference for 87.5% of the test cases.
Sebastian Schwarz, Roger Olsson, Mårten Sjöström, Sylvain Tourancheau
PCS2
2007 Evaluation of a combined pre-processing and H.264-compression scheme for 3D integral images
abstract
To provide sufficient 3D-depth fidelity, integral imaging (II) requires an increase in spatial resolution of several orders of magnitude from today's 2D images. We have recently proposed a pre-processing and compression scheme for still II-frames based on forming a pseudo video sequence (PVS) from sub images (SI), which is later coded using the H.264/MPEG-4 AVC video coding standard. The scheme has shown good performance on a set of reference images. In this paper we first investigate and present how five different ways to select the SIs when forming the PVS affect the schemes compression efficiency. We also study how the II-frame structure relates to the performance of a PVS coding scheme. Finally we examine the nature of the coding artifacts which are specific to the evaluated PVS-schemes. We can conclude that for all except the most complex reference image, all evaluated SI selection orders significantly outperforms JPEG 2000 where compression ratios of up to 342:1, while still keeping PSNR > 30 dB, is achieved. We can also confirm that when selecting PVS-scheme, the scheme which results in a higher PVS-picture resolution should be preferred to maximize compression efficiency. Our study of the coded II-frames also indicates that the SI-based PVS, contrary to other PVS schemes, tends to distribute its coding artifacts more homogenously over all 3D-scene depths.
Roger Olsson, Mårten Sjöström, Youzhi Xu
VCIP1
2006 A Combined Pre-Processing and H.264-Compression Scheme for 3D Integral Images
abstract
The next evolutionary step in enhancing video communication fidelity is taken by adding scene depth. 3D video using integral imaging (II) is widely considered as the technique able to take this step. However, an increase in spatial resolution of several orders of magnitude from todays 2D video is required to provide a sufficient depth fidelity, which includes motion parallax. In this paper we propose a pre-processing and compression scheme that aims to enhance the compression efficiency of integral images. We first transform a still integral image into a pseudo video sequence consisting of sub-images, which is then compressed using an H.264 video encoder. The improvement in compression efficiency of using this scheme is evaluated and presented. An average PSNR increase of 5.7 dB or more, compared to JPEG 2000, is observed on a set of reference images.
Roger Olsson, Mårten Sjöström, Youzhi Xu
ICIP1