Mårten Sjöström

dblp:99/11531 · DBLP profile ↗
← Back
45ranked-venue papers
1as first author
21since 2021 · last 2026
0000-0003-3751-6089ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 43 · 20 since 2021Human-computer interaction and ubiquitous computing · 7 · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 FireSegUNet: Exploring computationally efficient fire segmentation network for Unmanned Aerial Vehicles
abstract
Semantic segmentation on resource-constrained hardware remains a key challenge in deep learning, particularly for deployment on edge devices and embedded systems. In this study, we propose FireSegUNet, a lightweight and computationally efficient deep-learning architecture tailored for such environments. The model integrates an optimized inverted bottleneck layer for feature extraction within an encoder–decoder framework, reducing computational complexity by up to 51%. It also improves segmentation accuracy through an efficient squeeze-and-excitation block, while reducing inference time by up to 7.3 × and energy consumption by up to 4.6 × compared to conventional attention mechanisms. Extensive evaluation on diverse fire segmentation datasets demonstrates that FireSegUNet achieves competitive segmentation accuracy while reducing the number of parameters and storage requirements by up to 81%. Additionally, we provide a detailed analysis of the relationship between model complexity metrics and actual inference time, memory usage, and energy consumption. This comprehensive evaluation confirms that FireSegUNet delivers better performance on edge devices and generalizes well to unseen datasets. These findings position FireSegUNet as a practical solution for efficient image segmentation in resource-constrained environments. Although primarily validated on fire segmentation, the modular design of FireSegUNet makes it adaptable to other computer vision tasks. The source code of FireSegUNet will be publicly available at https://github.com/Realistic3D-MIUN/FireSegUNet .
Ali Hassan 0007, Johan Johansson, Karen Egiazarian, Mårten Sjöström
Knowl. Based Syst.6
2026 DUALF-D: Disentangled dual-hyperprior approach for light field image compression
abstract
Light field (LF) imaging captures spatial and angular information, offering a 4D scene representation enabling enhanced visual understanding. However, high dimensionality and redundancy across spatial and angular domains present major challenges for compression, particularly where storage, transmission bandwidth, or processing latency are constrained. We present a novel Variational Autoencoder (VAE)-based framework that explicitly disentangles spatial and angular features using two parallel latent branches. Each branch is coupled with an independent hyperprior model, allowing more precise distribution estimation for entropy coding and finer rate–distortion control. This dual-hyperprior structure enables the network to adaptively compress spatial and angular information based on their unique statistical characteristics, improving coding efficiency. To further enhance latent feature specialization and promote disentanglement, we introduce a mutual information-based regularization term that minimizes redundancy between the two branches while preserving feature diversity. Unlike prior methods relying on covariance-based penalties prone to collapse, our information-theoretic regularizer provides more stable and interpretable latent separation. Experimental results on publicly available LF datasets demonstrate our method achieves strong compression performance, yielding an average BD-PSNR gain of 2.91 dB over HEVC and high compression ratios (e.g., 200:1). Additionally, our design enables fast inference, with a total end-to-end time over 19x faster than the JPEG Pleno standard, making it well-suited for real-time and bandwidth-sensitive applications. By jointly leveraging disentangled representation learning, dual-hyperprior modeling, and information-theoretic regularization, our approach offers a scalable, effective solution for practical light field image compression.
Soheib Takhtardeshir, Roger Olsson, Christine Guillemot, Mårten Sjöström
Signal Process. Image Commun.4
2026 Scalable Hole-Filling for Real-Time Multi-GPU Light Field Path Tracing
abstract
Light field (LF) displays address the mismatch in focus cues present in traditional displays by triggering natural defocus blur and enabling motion parallax. They rely on geometrical optics, displaying rays from multiple angles of view. LF path tracing is computationally expensive for real-time applications, since it requires rendering multiple views. To reduce this computational complexity, spatially reprojecting pixels between views is commonly performed. Reusing pixels that are already rendered is cheaper than path tracing additional ones. However, when occluded areas are uncovered in some views, reprojection is not possible, creating holes in these views. Filling-in the holes requires extra path tracing computation. This paper investigates scalable hole-filling strategies for LF path tracing, using multiple GPUs to reach real-time performance. We propose an algorithm search optimization procedure to determine whether a specific assignment algorithm can be generalized across scenes, using the hole-filling time as a minimization function. In addition, we introduce DaSH (Discarded and Subsampled path traced Hole-filling), a novel method that reduces computation and divergence overhead in fixed-size hardware thread-blocks. Based on local pixel sparsity within pixel patches, DaSH adaptively subsamples and discards hole-filling rays. Our evaluation demonstrates that DaSH achieves significant performance gains while preserving the visual and structural quality of refocused light field images at the retina plane. The experiments demonstrate an average speedup factor of $1.8\times$1.8× for DaSH, compared to prior work, in a multi-GPU rendering system.
Erwan Leria, Markku Mäkitalo, Pekka Jääskeläinen, Mårten Sjöström
IEEE Trans. Vis. Comput. Graph.4
2025 Real-Time View Synthesis with Multiplane Image Network using Multimodal Supervision
abstract
Recent advances in view synthesis from a single image have increased visual quality of the newly synthesized viewpoints significantly. However, the high computational cost of state-of-the-art methods remains a critical bottleneck, limiting their adoption in real-time applications such as immersive telepresence. To address this limitation, we present a multiplane image (MPI) network that achieves real-time view synthesis. Unlike existing approaches that often rely on a separate depth estimation network to guide the network for estimating MPI parameters, our framework directly predicts the MPI parameters from a single RGB image. To guide the network for estimating correct parameters, we introduce a training strategy that leverages joint supervision from both view synthesis and depth estimation losses to ensure visual fidelity. During inference, our method exclusively utilizes the optimized view synthesis branch, while the depth decoder is only used for training. Our end-to-end approach renders views from a single input image in real-time. Extensive experiments validate that our method provides a compelling rendering speed with visual quality in par with state-of-the-art methods, highlighting its suitability for live, interactive applications.
Manu Gond, Mohammadreza Shamshirgarha, Emin Zerman, Sebastian Knorr, Mårten Sjöström
MMSP5
2025 EPINET-Lite: Rethinking Mixed Convolutions for Efficient Light Field Disparity Estimation Network
abstract
Convolutional neural networks are widely used for light field disparity estimation. However, many state-of-the-art deep learning models are computationally expensive due to their reliance on standard convolutions with varying kernel sizes. In this paper, we analyze the effect of various advanced convolution operations with different kernel sizes for feature extraction in a state-of-the-art light field disparity estimation network. Based on this investigation, we propose an optimized mixed convolution layer to extract relevant features using multiple kernel sizes in parallel, while maintaining significantly lower computational cost. Experimental results demonstrate that our approach reduces model complexity by up to 4.2× while also improving disparity estimation accuracy. These findings make the proposed convolutional operation more practical for light field applications, where efficient spatial and angular feature extraction is essential for improved model performance.
Ali Hassan 0007, Karen Egiazarian, Mårten Sjöström
MMSP4
2025 Subjective Visual Quality Assessment of Compressed Light Field Images: Learning-based vs. Conventional Methods
abstract
Light fields (LF) technology enables the capture and reproduction of a 3D scene accurately, which enhances visual experience in various applications. The sheer volume of multiplicity of the captured views creates logistical problems in both storage and data transmission, which makes LF compression crucial. Even though many LF compression techniques have been proposed and evaluated in recent years, the leading-edge learning-based approaches have not been subject to the same level of scrutiny. This paper presents a subjective quality assessment study on four different LF compression methods, including two learning-based LF compression methods, which have not been studied before from a subjective quality point of view. For this purpose, subjective opinion scores were collected from viewers in two different universities for a cross-lab study. The results indicate that the learning-based compression methods have different behavior in their rate-distortion curves, and that there is room for improvement for learning-based methods. A qualitative analysis also shows that their artifact structures are different from conventional ones. The results highlight the need for a perceptual objective quality metric that takes different types of artifacts into account. The obtained subjective quality database (MiX-LFQDB) is made public to support further research in this area: https://doi.org/10.5281/zenodo.16778670
Emin Zerman, Soheib Takhtardeshir, Anthony Trioux, Jianlong Qin, Roger Olsson, Mårten Sjöström
MMSP7
2025 A Visual Quality of Experience Toolkit for Realistic Immersive Telepresence Applications
abstract
Immersive imaging applications gained a lot of traction in the last decade with the advances in capture, processing, compression, transmission, and display technologies. Recent works on spherical light fields and new view synthesis bring new challenges with respect to fast rendering of new views and a smooth quality of experience (QoE). The real-time rendering capabilities enable realistic immersive telepresence applications, which extend the traditional telepresence systems with improved visual realism and possible addition of depth cues. The recent efforts on the spherical light fields and new view synthesis approaches show that a combined light field and spherical data visualization will be beneficial for the scientific community. In this paper, we provide a WebGL- and WebXR-based visual QoE toolkit which can be used both on traditional displays and extended reality headsets, presenting various visual modalities. To validate the usefulness of the proposed toolkit, we conducted a pilot test on our publicly available spherical light field database. This toolkit can be used for easier and faster quality assessment and can support the scientific community for subjective visual QoE studies focusing on telecommunication, telepresence, and augmented telepresence applications.
Manu Gond, Emin Zerman, Mohammadreza Shamshirgarha, Sebastian Knorr, Mårten Sjöström
QoMEX5
2025 User Study on Visual Interface Helpfulness in Remote Inspection Tasks Using Teleoperation
abstract
Teleoperation systems are increasingly used in remote inspection and maintenance, especially in safety-critical settings. This paper presents findings from a lab study examining user interaction with a teleoperation for airport safety inspection. The study investigates five visual interface configurations: First Person View (FPV), Third Person View (TPV), Augmented FPV, FPV+TPV, and Augmented FPV+TPV. Eighteen participants completed inspection tasks under each condition and rated interface helpfulness, workload (NASA-TLX), and simulator sickness. The result was confirmed by bootstrapping with 1000 iterations. Results showed that all view configurations were rated more helpful than TPV alone, with FPV-based combinations improving depth perception and spatial understanding. NASA-TLX revealed cognitive effort, while SSQ scores decreased post-task, particularly in oculomotor symptoms, suggesting adaptation. These findings support continued research into optimising visual interfaces for teleoperation in safety-critical settings.
Shirin Rafiei, Kjell Brunnström, Gabriele Pifferi, Bo N. Schenkman, Anders Djupsjobacka, Börje Andérn, Mårten Sjöström
QoMEX7
2025 3D SMoE Splatting for Edge-aware Realtime Radiance Field Rendering
abstract
Steered Mixtures-of-Experts (SMoE) is an existing regression framework that has previously been applied for modeling and compression of 2D images and higher-dimensional imagery, including compression of light fields and light-field video. SMoE models are sparse, edge-aware representations that allow rendering of imagery with few Gaussians with excellent quality. In this paper a novel, edge-aware "3D SMoE Splatting" (3DSMoES) framework for 3D rendering is introduced, adopted to fit into the existing "3D Gaussian Splatting" (3DGS) CUDA optimization pipeline. Here, SMoE regression serves as a "plug-and-play" solution that replaces the established 3DGS regression as a novel workhorse. 3DSMoES achieves significant visual quality gains with drastically fewer Gaussian kernels compared to 3DGS. We observe up to approximately 4dB improvement in PSNR on individual scenes with kernel reductions between 20 to 50 percent. The sparse models are significantly faster to train and allow up to 30-50 percent improved rendering speeds.
Yi-Hsin Li, Thomas Sikora, Sebastian Knorr, Mårten Sjöström
SIGGRAPH Asia4
2025 DUALF-C: Disentangled Light Field Compression with Entropy-Aware Bitstream Generation
abstract
Recent advancements in learned light field (LF) image compression highlight the advantages of modeling spatial and angular redundancies using deep generative models. Among these, Variational Autoencoders (VAEs) have shown strong potential in learning compact latent representations of LF data. However, many existing approaches rely on simple uniform quantization and basic entropy coding, which limits compression efficiency in practical applications. This paper introduces DUALF-C, a lightweight and retraining-free compression pipeline that augments pretrained VAE-based LF compression models using structured post-encoding transformations. The proposed framework integrates bitplane slicing, latent channel reordering, non-uniform quantization, and patch-based vector quantization to improve bitrate efficiency while preserving reconstruction quality. Experimental evaluations demonstrate that DUALF-C significantly reduces bit-per-pixel (BPP) without degrading image quality, making it a practical solution for bandwidth-constrained immersive imaging systems.
Soheib Takhtardeshir, Roger Olsson, Christine Guillemot, Mårten Sjöström
VCIP4
2025 3D-Gaussian Splatting Representation of Rendered Views from Plenoptic 2.0 Lenslet Images
abstract
Thanks to plenoptic cameras, rich information on the radiance of a scene can be conveniently captured without heavy devices like camera arrays. However, rendering techniques are needed to generate views for human visual perception. Existing patch extraction-based rendering techniques can generate views from the lenslet images captured by plenoptic cameras, but they suffer from the inherent problem of artifacts and the limited views to be rendered. In this paper, we present a new view rendering technique from plenoptic 2.0 camera-captured lenslet image by using 3-dimensional gaussian splatting (3DGS). At its first step, the reference lenslet converter (RLC) provided by MPEG LVC AhG, one of the existing patch extraction methods, generates initial views with the help of estimated disparity between adjacent micro images in the lenslet image. At its second step, the 3DGS generates the final views after being trained by the initial views. The rendering results obtained by the proposed 2-step approach show significantly fewer artifacts in the rendered views than the patch stitching process of the existing method.
Jonghoon Yim, Byeungwoo Jeon, Roger Olsson, Mårten Sjöström
VCIP4
2025 Adaptive Segmentation-Based Initialization for Steered Mixture of Experts Image Regression
abstract
Kernel image regression methods have demonstrated excellent efficiency in various image processing tasks, including image and light-field compression, Gaussian Splatting, denoising and super-resolution. The estimation of parameters for these methods commonly employs gradient descent iterative optimization, which poses a significant computational burden for many applications. In this paper, we introduce a novel adaptive segmentation-based initialization method targeted for optimizing Steered-Mixture-of Experts (SMoE) gating networks and RadialBasis-Function (RBF) networks with steering kernels. The novel initialization method allocates kernels into pre-calculated image segments. The optimal number of kernels, kernel positions, and steering parameters are derived per segment in an iterative optimization and kernel sparsification procedure. The kernel information from local segments is then transferred into a global initialization, ready for use in iterative optimization of SMoE, RBF, and related kernel image regression methods. Results demonstrate significant improvements in both objective and subjective quality compared to regular grid, K-Means, deeplearning-based, and previous segmentation-based initialization methods. The proposed initialization method reduces kernel usage by 70% compared to other initialization methods while maintaining the same reconstruction quality. Furthermore, by generating initial parameters closer to optimized results, convergence time is reduced, achieving overall runtime savings of up to 50% compared to prior methods. Additionally, the method supports parallel computation, with initialization time halved when using four GPUs compared to one
Yi-Hsin Li, Sebastian Knorr, Mårten Sjöström, Thomas Sikora
IEEE Trans. Multim.3
2024 Laboratory study: Human Interaction using Remote Control System for Airport Safety Management
abstract
Remote control technology for machines and robots has experienced significant advancement in many domains where visual information delivery is essential. Safety management at airports is one field that benefits from remote control systems, enabling operators to scan the airstrip for obstacles and debris. Depth perception and proper view position are critical in remote control systems, where operators need to accurately perceive 3D coordinates and maintain an appropriate perspective for performing tasks from a distance. This Demo paper presents a laboratory platform for an experimental study on human interaction with remote control systems using visual interfaces for inspection tasks to enhance airport safety management. The platform can evaluate the user experience of interfaces provided by First Person View, Third Person View, and Augmentation technologies. This platform enables exploration through controlled experiments and by user tests. These provide an avenue for assessing how these interfaces may affect human performance, depth perception, and user experience when conducting inspection tasks remotely. The findings will shed light on the strengths and limitations of each interface type, offering insights into their potential applications in various domains such as industrial inspection, surveillance, and remote exploration.
Shirin Rafiei, Kjell Brunnström, Bo N. Schenkman, Jonas Andersson 0005, Mårten Sjöström
QoMEX5
2024 A Spherical Light Field Database for Immersive Telecommunication and Telepresence Applications
abstract
Immersive imaging technologies provide an enhanced user experience for visual applications and are getting ready for commonplace use by the industry and the general populace. In particular, light field is a promising technology that enables the capture and reproduction of real light rays from the scene, which can provide a backbone for immersive telecommunication and telepresence applications. Nevertheless, there are still many challenges in transmitting and reproducing light field data. This paper proposes a spherical light field dataset that can be used as a foundation for developing telepresence applications. The Spherical Light Field Database (SLFDB) consists of a light field of 60 views captured with an omnidirectional camera in 20 scenes. To show the usefulness of the proposed database, we provide two use cases: compression and viewpoint estimation. The initial results validate that the publicly available SLFDB will benefit the scientific community.
Emin Zerman, Manu Gond, Soheib Takhtardeshir, Roger Olsson, Mårten Sjöström
QoMEX5
2024 Interactive Multi-GPU Light Field Path Tracing Using Multi-Source Spatial Reprojection
abstract
Path tracing combined with multiview displays enables progress towards achieving ultrarealistic virtual reality. However, multiview displays based on light field technology impose a heavy workload for real-time graphics due to the large number of views to be rendered. In order to achieve low latency performance, computational effort can be reduced by path tracing only some views (source views), and synthesizing the remaining views (target views) through spatial reprojection, which reuses path traced pixels from source views to target views. Deciding the number of source views with respect to the computational resources is not trivial, since spatial reprojection introduces dependencies in the otherwise trivially parallel rendering pipeline and path tracing multiple source views increases the computation time.
Erwan Leria, Markku Mäkitalo, Pekka Jääskeläinen, Mårten Sjöström
VRST4
2023 Human Interaction in Industrial Tele-Operated Driving: Laboratory Investigation
abstract
Tele-operated driving enables industrial operators to control heavy machinery remotely. By doing so, they could work in improved and safe workplaces. However, some challenges need to be investigated while presenting visual information from on-site scenes for operators sitting at a distance in a remote site. This paper discusses the impact of video quality (spatial resolution), field of view, and latency on users' depth perception, experience, and performance in a lab-based tele-operated application. We performed user experience evaluation experiments to study these impacts. Overall, the user experience and comfort decrease while the users' performance error increases with an increase in the glass-to-glass latency. The user comfort reduces, and the user performance error increases with reduced video quality (spatial resolution).
Shirin Rafiei, Chetna Singhal 0001, Kjell Brunnström, Mårten Sjöström
QoMEX4
2023 Segmentation-based Initialization for Steered Mixture of Experts
abstract
The Steered-Mixture-of-Experts (SMoE) model is an edge-aware kernel representation that has successfully been explored for the compression of images, video, and higher-dimensional data such as light fields. The present work aims to leverage the potential for enhanced compression gains through efficient kernel reduction. We propose a fast segmentation-based strategy to identify a sufficient number of kernels for representing an image and giving initial kernel parametrization. The strategy implies both reduced memory footprint and reduced computational complexity for the subsequent parameter optimization, resulting in an overall faster processing time. Fewer kernels, when combined with the inherent sparsity of the SMoEs, further enhance the overall compression performance. Empirical evaluations demonstrate a gain of 0.3-1.0 dB in PSNR for a constant number of kernels, and the use of 23 % less kernels and 25 % less time for constant PSNR. The results highlight the feasibility and practicality of the approach, positioning it as a valuable solution for various image-related applications, including image compression.
Yi-Hsin Li, Mårten Sjöström, Sebastian Knorr, Thomas Sikora
VCIP2
2022 Light-Weight EPINET Architecture for Fast Light Field Disparity Estimation
abstract
Recent deep learning-based light field disparity estimation algorithms require millions of parameters, which demand high computational cost and limit the model deployment. In this paper, an investigation is carried out to analyze the effect of depthwise separable convolution and ghost modules on state-of-the-art EPINET architecture for disparity estimation. Based on this investigation, four convolutional blocks are proposed to make the EPINET architecture a fast and light-weight network for disparity estimation. The experimental results exhibit that the proposed convolutional blocks have significantly reduced the computational cost of EPINET architecture by up to a factor of 3.89, while achieving comparable disparity maps on HCI Benchmark dataset.
Ali Hassan 0007, Mårten Sjöström, Karen Egiazarian
MMSP2
2022 Event detection in surveillance videos: a review
abstract
Abstract Since 2008, a variety of systems have been designed to detect events in security cameras. There are also more than a hundred journal articles and conference papers published in this field. However, no survey has focused on recognizing events in the surveillance system. Thus, motivated us to provide a comprehensive review of the different developed event detection systems. We start our discussion with the pioneering methods that used the TRECVid-SED dataset and then developed methods using VIRAT dataset in TRECVid evaluation. To better understand the designed systems, we describe the components of each method and the modifications of the existing method separately. We have outlined the significant challenges related to untrimmed security video action detection. Suitable metrics are also presented for assessing the performance of the proposed models. Our study indicated that the majority of researchers classified events into two groups on the basis of the number of participants and the duration of the event for the TRECVid-SED Dataset. Depending on the group of events, one or more models to identify all the events were used. For the VIRAT dataset, object detection models to localize the first stage activities were used throughout the work. Except one study, a 3D convolutional neural network (3D-CNN) to extract Spatio-temporal features or classifying different activities were used. From the review that has been carried, it is possible to conclude that developing an automatic surveillance event detection system requires three factors: accurate and fast object detection in the first stage to localize the activities, and classification model to draw some conclusion from the input values.
Abdolamir Karbalaie, Farhad Abtahi, Mårten Sjöström
Multim. Tools Appl.3
2022 A TV regularisation sparse light field reconstruction model based on guided-filtering
Gangrong Qu, Mårten Sjöström
Signal Process. Image Commun.3
2021 Analysis of Top-Down Connections in Multi-Layered Convolutional Sparse Coding
abstract
Convolutional Neural Networks (CNNs) have been instrumental in the recent advances in machine learning, with applications to media applications. Multi-Layered Convolutional Sparse Coding (ML-CSC) based on a cascade of convolutional layers in which each layer can be approximately explained by the following layer can be seen as a biologically inspired framework. However, both CNNs and ML-CSC networks lack top-down information flows that are studied in neuroscience for understanding the mechanisms of the mammal cortex. A successful implementation of such top-down connections could lead to another leap in machine learning and media applications. This study analyses the effects of a feedback connection on an ML-CSC network, considering trade-off between sparsity and reconstruction error, support recovery rate, and mutual coherence in trained dictionaries. We find that using the feedback connection during training impacts the mutual coherence of the dictionary in a way that the equivalence between the l0-and l1-norm is verified for a smaller range of sparsity values. Experimental results show that the use of feedback during training does not favour inference with feedback, in terms of sparse support recovery rates. However, when the sparsity constraints are given a lower weight, the use of feedback at inference time is beneficial, in terms of support recovery rates.
Joakim Edlund, Christine Guillemot, Mårten Sjöström
MMSP3
2020 Latency impact on Quality of Experience in a virtual reality simulator for remote control of machines
abstract
In this article, we have investigated a VR simulator of a forestry crane used for loading logs onto a truck. We have mainly studied the Quality of Experience (QoE) aspects that may be relevant for task completion, and whether there are any discomfort related symptoms experienced during the task execution. QoE experiments were designed to capture the general subjective experience of using the simulator, and to study task performance. The focus was to study the effects of latency on the subjective experience, with regards to delays in the crane control interface. Subjective studies were performed with controlled delays added to the display update and hand controller (joystick) signals. The added delays ranged from 0 to 30 ms for the display update, and from 0 to 800 ms for the hand controller. We found a strong effect on latency in the display update and a significant negative effect for 800 ms added delay on latency in the hand controller (in total approx. 880 ms latency including the system delay). The Simulator Sickness Questionnaire (SSQ) gave significantly higher scores after the experiment compared to before the experiment, but a majority of the participants reported experiencing only minor symptoms. Some test subjects ceased the test before finishing due to their symptoms, particularly due to the added latency in the display update.
Kjell Brunnström, Elijs Dima, Tahir Qureshi, Mathias Johanson, Mattias Andersson 0004, Mårten Sjöström
Signal Process. Image Commun.6
2020 Shearlet Transform-Based Light Field Compression Under Low Bitrates
abstract
Light field (LF) acquisition devices capture spatial and angular information of a scene. In contrast with traditional cameras, the additional angular information enables novel postprocessing applications, such as 3D scene reconstruction, the ability to refocus at different depth planes, and synthetic aperture. In this paper, we present a novel compression scheme for LF data captured using multiple traditional cameras. The input LF views were divided into two groups: key views and decimated views. The key views were compressed using the multi-view extension of high-efficiency video coding (MV-HEVC) scheme, and decimated views were predicted using the shearlet-transform-based prediction (STBP) scheme. Additionally, the residual information of predicted views was also encoded and sent along with the coded stream of key views. The proposed scheme was evaluated over a benchmark multi-camera based LF datasets, demonstrating that incorporating the residual information into the compression scheme increased the overall peak signal to noise ratio (PSNR) by 2 dB. The proposed compression scheme performed significantly better at low bit rates compared to anchor schemes, which have a better level of compression efficiency in high bit-rate scenarios. The sensitivity of the human vision system towards compression artifacts, specifically at low bit rates, favors the proposed compression scheme over anchor schemes.
Waqas Ahmad 0002, Suren Vagharshakyan, Mårten Sjöström, Atanas P. Gotchev, Robert Bregovic, Roger Olsson
IEEE Trans. Image Process.3
2019 Depth-Assisted Demosaicing for Light Field Data in Layered Object Space
abstract
Light field technology, which emerged as a solution to the increasing demands of visually immersive experience, has shown its extraordinary potential for scene content representation and reconstruction. Unlike conventional photography that maps the 3D scenery onto a 2D plane by a projective transformation, light field preserves both the spatial and angular information, enabling further processing steps such as computational refocusing and image-based rendering. However, there are still gaps that have been barely studied, such as the light field demosaicing process. In this paper, we propose a depth-assisted demosaicing method for light field data. First, we exploit the sampling geometry of the light field data with respect to the scene content using the ray-tracing technique and develop a sampling model of light field capture. Then we carry out the demosaicing process in a layered object space with object-space sampling adjacencies rather than pixel placement. Finally, we compare our results with state-of-art approaches and discuss about the potential research directions of the proposed sampling model to show the significance of our approach.
Mårten Sjöström
ICIP2
2019 View Position Impact on QoE in an Immersive Telepresence System for Remote Operation
abstract
In this paper, we investigate how different viewing positions affect a user's Quality of Experience (QoE) and performance in an immersive telepresence system. A QoE experiment has been conducted with 27 participants to assess the general subjective experience and the performance of remotely operating a toy excavator. Two view positions have been tested, an overhead and a ground-level view, respectively, which encourage reliance on stereoscopic depth cues to different extents for accurate operation. Results demonstrate a significant difference between ground and overhead views: the ground view increased the perceived difficulty of the task, whereas the overhead view increased the perceived accomplishment as well as the objective performance of the task. The perceived helpfulness of the overhead view was also significant according to the participants.
Elijs Dima, Kjell Brunnström, Mårten Sjöström, Mattias Andersson 0004, Joakim Edlund, Mathias Johanson, Tahir Qureshi
QoMEX3
2018 Shearlet Transform Based Prediction Scheme for Light Field Compression
abstract
Light field acquisition technologies capture angular and spatial information of the scene. The spatial and angular information enables various post processing applications, e.g. 3D scene reconstruction, refocusing, synthetic aperture etc at the expense of an increased data size. In this paper, we present a novel prediction tool for compression of light field data acquired with multiple camera system. The captured light field (LF) can be described using two plane parametrization as, L(u, v, s, t), where (u, v) represents each view image plane coordinates and (s, t) represents the coordinates of the capturing plane. In the proposed scheme, the captured LF is uniformly decimated by a factor d in both directions (in s and t coordinates), resulting in a sparse set of views also referred to as key views. The key views are converted into a pseudo video sequence and compressed using high efficiency video coding (HEVC). The shearlet transform based reconstruction approach, presented in [1], is used at the decoder side to predict the decimated views with the help of the key views. Four LF images (Truck, Bunny from Stanford dataset, Set2 and Set9 from High Density Camera Array dataset) are used in the experiments. Input LF views are converted into a pseudo video sequence and compressed with HEVC to serve as anchor. Rate distortion analysis shows the average PSNR gain of 0.98 dB over the anchor scheme. Moreover, in low bit-rates, the compression efficiency of the proposed scheme is higher compared to the anchor and on the other hand the performance of the anchor is better in high bit-rates. Different compression response of the proposed and anchor scheme is a consequence of their utilization of input information. In the high bit-rate scenario, high quality residual information enables the anchor to achieve efficient compression. On the contrary, the shearlet transform relies on key views to predict the decimated views without incorporating residual information. Hence, it has inherit reconstruction error. In the low bit-rate scenario, the bit budget of the proposed compression scheme allows the encoder to achieve high quality for the key views. The HEVC anchor scheme distributes the same bit budget among all the input LF views that results in degradation of the overall visual quality. The sensitivity of human vision system toward compression artifacts in low-bit-rate cases favours the proposed compression scheme over the anchor scheme.
Waqas Ahmad 0002, Suren Vagharshakyan, Mårten Sjöström, Atanas P. Gotchev, Robert Bregovic, Roger Olsson
DCC3
2018 Towards a Generic Compression Solution for Densely and Sparsely Sampled Light Field Data
abstract
Light field (LF) acquisition technologies capture the spatial and angular information present in scenes. The angular information paves the way for various post-processing applications such as scene reconstruction, refocusing, and synthetic aperture. The light field is usually captured by a single plenop-tic camera or by multiple traditional cameras. The former captures a dense LF, while the latter captures a sparse LF. This paper presents a generic compression scheme that efficiently compresses both densely and sparsely sampled LFs. A plenoptic image is converted into sub-aperture images, and each sub-aperture image is interpreted as a frame of a multiview sequence. In comparison, each view of the multi-camera system is treated as a frame of a multi-view sequence. The multi-view extension of high efficiency video coding (MV-HEVC) is used to encode the pseudo multi-view sequence. This paper proposes an adaptive prediction and rate allocation scheme that efficiently compresses LF data irrespective of the acquisition technology used.
Waqas Ahmad 0002, Roger Olsson, Mårten Sjöström
ICIP3
2017 Interpreting plenoptic images as multi-view sequences for improved compression
abstract
Over the last decade, advancements in optical devices have made it possible for new novel image acquisition technologies to appear. Angular information for each spatial point is acquired in addition to the spatial information of the scene that enables 3D scene reconstruction and various post-processing effects. Current generation of plenoptic cameras spatially multiplex the angular information, which implies an increase in image resolution to retain the level of spatial information gathered by conventional cameras. In this work, the resulting plenoptic image is interpreted as a multi-view sequence that is efficiently compressed using the multi-view extension of high efficiency video coding (MV-HEVC). A novel two-dimensional weighted prediction and rate allocation scheme is proposed to adopt the HEVC compression structure to the plenoptic image properties. The proposed coding approach is a response to ICIP 2017 Grand Challenge: Light field Image Coding. The proposed scheme outperforms all ICME-contestants, and improves on the JPEG-anchor of ICME with an average PSNR gain of 7.5 dB and the HEVC-anchor of ICIP 2017 Grand Challenge with an average PSNR gain of 2.4 dB.
Waqas Ahmad 0002, Roger Olsson, Mårten Sjöström
ICIP3
2016 SMART: a light field image quality dataset
abstract
In this contribution, the design of a Light Field image dataset is presented. It can be useful for design, testing, and benchmarking Light Field image processing algorithms. As first step, image content selection criteria have been defined based on selected image quality key-attributes, i.e. spatial information, colorfulness, texture key features, depth of field, etc. Next, image scenes have been selected and captured by using the Lytro Illum Light Field camera. Performed analysis shows that the proposed set of images is sufficient for addressing a wide range of attributes relevant for assessing Light Field image quality.
Pradip Paudyal, Roger Olsson, Mårten Sjöström, Federica Battisti, Marco Carli
MMSys3
2016 Virtual view synthesis using layered depth image generation and depth-based inpainting for filling disocclusions and translucent disocclusions
Suryanarayana Murthy Muddala, Mårten Sjöström, Roger Olsson
J. Vis. Commun. Image Represent.2
2016 Coding of Focused Plenoptic Contents by Displacement Intra Prediction
abstract
A light field is commonly described by a two-plane representation with four dimensions. Refocused 3D contents can be rendered from light field images. A method for capturing these images is using cameras with microlens arrays. A dense sampling of the light field results in large amounts of redundant data. Therefore, an efficient compression is vital for a practical use of these data. In this paper, we propose a displacement intra prediction scheme with a maximum of two hypotheses for the compression of plenoptic contents from focused plenoptic cameras. The proposed scheme is further implemented into High Efficiency Video Coding (HEVC). The work is aiming at efficiently coding plenoptic captured contents without knowing underlying camera geometries. In addition, the theoretical analysis of the displacement intra prediction for plenoptic images is explained; the relationship between the compressed captured images and their rendered quality is also analyzed. Evaluation results show that plenoptic contents can be efficiently compressed by the proposed scheme. Bit rate reduction up to 60% over HEVC is obtained for plenoptic images, and more than 30% is achieved for the tested video sequences.
Yun Li 0004, Mårten Sjöström, Roger Olsson, Ulf Jennehag
IEEE Trans. Circuits Syst. Video Technol.2
2016 Scalable Coding of Plenoptic Images by Using a Sparse Set and Disparities
abstract
One of the light field capturing techniques is the focused plenoptic capturing. By placing a microlens array in front of the photosensor, the focused plenoptic cameras capture both spatial and angular information of a scene in each microlens image and across microlens images. The capturing results in a significant amount of redundant information, and the captured image is usually of a large resolution. A coding scheme that removes the redundancy before coding can be of advantage for efficient compression, transmission, and rendering. In this paper, we propose a lossy coding scheme to efficiently represent plenoptic images. The format contains a sparse image set and its associated disparities. The reconstruction is performed by disparity-based interpolation and inpainting, and the reconstructed image is later employed as a prediction reference for the coding of the full plenoptic image. As an outcome of the representation, the proposed scheme inherits a scalable structure with three layers. The results show that plenoptic images are compressed efficiently with over 60 percent bit rate reduction compared with High Efficiency Video Coding intra coding, and with over 20 percent compared with an High Efficiency Video Coding block copying mode.
Yun Li 0004, Mårten Sjöström, Roger Olsson, Ulf Jennehag
IEEE Trans. Image Process.2
2015 Depth and angular resolution in plenoptic cameras
abstract
We present a model-based approach to extract the depth and angular resolution in a plenoptic camera. Obtained results for the depth and angular resolution are validated against Ze-max ray tracing results. The provided model-based approach gives the location and number of the resolvable depth planes in a plenoptic camera as well as the angular resolution with regards to disparity in pixels. The provided model-based approach is straightforward compared to practical measurements and can reflect on the plenoptic camera parameters such as the microlens f-number in contrast with the principal-ray-model approach. Easy and accurate quantification of different resolution terms forms the basis for designing the capturing setup and choosing a reasonable system configuration for plenoptic cameras. Results from this work will accelerate customization of the plenoptic cameras for particular applications without the need for expensive measurements.
Mitra Damghanian, Roger Olsson, Mårten Sjöström
ICIP3
2015 Coding of plenoptic images by using a sparse set and disparities
abstract
A focused plenoptic camera not only captures the spatial information of a scene but also the angular information. The capturing results in a plenoptic image consisting of multiple microlens images and with a large resolution. In addition, the microlens images are similar to their neighbors. Therefore, an efficient compression method that utilizes this pattern of similarity can reduce coding bit rate and further facilitate the usage of the images. In this paper, we propose an approach for coding of focused plenoptic images by using a representation, which consists of a sparse plenoptic image set and disparities. Based on this representation, a reconstruction method by using interpolation and inpainting is devised to reconstruct the original plenoptic image. As a consequence, instead of coding the original image directly, we encode the sparse image set plus the disparity maps and use the reconstructed image as a prediction reference to encode the original image. The results show that the proposed scheme performs better than HEVC intra with more than 5 dB PSNR or over 60 percent bit rate reduction.
Yun Li 0004, Mårten Sjöström, Roger Olsson
ICME2
2014 Performance analysis in Lytro camera: Empirical and model based approaches to assess refocusing quality
abstract
In this paper we investigate the performance of Lytro camera in terms of its refocusing quality. The refocusing quality of the camera is related to the spatial resolution and the depth of field as the contributing parameters. We quantify the spatial resolution profile as a function of depth using empirical and model based approaches. The depth of field is then determined by thresholding the spatial resolution profile. In the model based approach, the previously proposed sampling pattern cube (SPC) model for representation and evaluation of the plenoptic capturing systems is utilized. For the experimental resolution measurements, camera evaluation results are extracted from images rendered by the Lytro full reconstruction rendering method. Results from both the empirical and model based approaches assess the refocusing quality of the Lytro camera consistently, highlighting the usability of the model based approaches for performance analysis of complex capturing systems.
Mitra Damghanian, Roger Olsson, Mårten Sjöström
ICASSP3
2014 Efficient intra prediction scheme for light field image compression
abstract
Interactive photo-realistic graphics can be rendered by using light field datasets. One way of capturing the dataset is by using light field cameras with microlens arrays. The captured images contain repetitive patterns resulted from adjacent mi-crolenses. These images don't resemble the appearance of a natural scene. This dissimilarity leads to problems in light field image compression by using traditional image and video encoders, which are optimized for natural images and video sequences. In this paper, we introduce the full inter-prediction scheme in HEVC into intra-prediction for the compression of light field images. The proposed scheme is capable of performing both unidirectional and bi-directional prediction within an image. The evaluation results show that above 3 dB quality improvements or above 50 percent bit-rate saving can be achieved in terms of BD-PSNR for the proposed scheme compared to the original HEVC intra-prediction for light field images.
Yun Li 0004, Mårten Sjöström, Roger Olsson, Ulf Jennehag
ICASSP2
2014 Spatial resolution in a multi-focus plenoptic camera
abstract
Evaluation of the state of the art plenoptic cameras is necessary for design and application purposes. In this work, spatial resolution is investigated in a multi-focus plenoptic camera using two approaches: empirical and model-based. The Raytrix R29 plenoptic camera is studied which utilizes three types of micro lenses with different focal lengths in a hexagonal array structure to increase the depth of field. The modelbased approach utilizes the previously proposed sampling pattern cube (SPC) model for representation and evaluation of the plenoptic capturing systems. For the experimental resolution measurements, spatial resolution values are extracted from images reconstructed by the provided Raytrix reconstruction method. Both the measurement and the SPC model based approaches demonstrate a gradual variation of the resolution values in a wide depth range for the multi focus R29 camera. Moreover, the good agreement between the results from the model-based approach and those from the empirical approach confirms suitability of the SPC model in evaluating high-level camera parameters such as the spatial resolution in a complex capturing system as R29 multi-focus plenoptic camera.
Mitra Damghanian, Roger Olsson, Mårten Sjöström, A. Erdmann, Christian Perwass
ICIP3
2014 A Weighted Optimization Approach to Time-of-Flight Sensor Fusion
abstract
Acquiring scenery depth is a fundamental task in computer vision, with many applications in manufacturing, surveillance, or robotics relying on accurate scenery information. Time-of-flight cameras can provide depth information in real-time and overcome short-comings of traditional stereo analysis. However, they provide limited spatial resolution and sophisticated upscaling algorithms are sought after. In this paper, we present a sensor fusion approach to time-of-flight super resolution, based on the combination of depth and texture sources. Unlike other texture guided approaches, we interpret the depth upscaling process as a weighted energy optimization problem. Three different weights are introduced, employing different available sensor data. The individual weights address object boundaries in depth, depth sensor noise, and temporal consistency. Applied in consecutive order, they form three weighting strategies for time-of-flight super resolution. Objective evaluations show advantages in depth accuracy and for depth image based rendering compared with state-of-the-art depth upscaling. Subjective view synthesis evaluation shows a significant increase in viewer preference by a factor of four in stereoscopic viewing conditions. To the best of our knowledge, this is the first extensive subjective test performed on time-of-flight depth upscaling. Objective and subjective results proof the suitability of our approach to time-of-flight super resolution approach for depth scenery capture.
Sebastian Schwarz, Mårten Sjöström, Roger Olsson
IEEE Trans. Image Process.2
2012 The Sampling Pattern Cube - A Representation and Evaluation Tool for Optical Capturing Systems
Mitra Damghanian, Roger Olsson, Mårten Sjöström
ACIVS3
2012 Adaptive depth filtering for HEVC 3D video coding
abstract
Consumer interest in 3D television (3DTV) is growing steadily, but current available 3D displays still need additional eye-wear and suffer from the limitation of a single stereo view pair. So it can be assumed that autostereoscopic multiview displays are the next step in 3D-at-home entertainment, since these displays can utilize the Multiview Video plus Depth (MVD) format to synthesize numerous viewing angles from only a small set of given input views. This motivates efficient MVD compression as an important keystone for commercial success of 3DTV. In this paper we concentrate on the compression of depth information in an MVD scenario. There have been several publications suggesting depth down- and upsampling to increase coding efficiency. We follow this path, using our recently introduced Edge Weighted Optimization Concept (EWOC) for depth upscaling. EWOC uses edge information from the video frame in the upscaling process and allows the use of sparse, non-uniformly distributed depth values. We exploit this fact to expand the depth down-/upsampling idea with an adaptive low-pass filter, reducing high energy parts in the original depth map prior to subsampling and compression. Objective results show the viability of our approach for depth map compression with up-to-date High-Efficiency Video Coding (HEVC). For the same Y-PSNR in synthesized views we achieve up to 18.5% bit rate decrease compared to full-scale depth and around 10% compared to competing depth down-/upsampling solutions. These results were confirmed by a subjective quality assessment, showing a statistical significant preference for 87.5% of the test cases.
Sebastian Schwarz, Roger Olsson, Mårten Sjöström, Sylvain Tourancheau
PCS3
2011 Layer Assignment Based on Depth Data Distribution for Multiview-Plus-Depth Scalable Video Coding
abstract
Three dimensional (3-D) video is experiencing a rapid growth in a number of areas, including 3-D cinema, 3-D TV, and mobile phones. Several problems must be addressed to display captured 3-D video at another location. One problem is how to represent the data. The multiview plus depth representation of a scene requires a lower bit rate than transmitting all views required by an application and provides more information than a 2-D-plus-depth sequence. Another problem is how to handle transmission in a heterogeneous network. Scalable video coding enables adaption of a 3-D video sequence to the conditions at the receiver. In this paper, we present a scheme that combines scalability based on the position in depth of the data and the distance to the center view. The general scheme preserves the center view data, whereas the data of the remaining views are extracted in enhancement layers depending on distance to the viewer and to the center camera. The data is assigned into enhancement layers within a view based on depth data distribution. Strategies concerning the layer assignment between adjacent views are proposed. In general, each extracted enhancement layer increases the visual quality and peak signal-to-noise ratio compared to only using center view data. The bit-rate per layer can be further decreased if depth data is distributed over the enhancement layers. The choice of strategy to assign layers between adjacent views depends on whether quality of the fore-most objects in the scene or the quality of the views close to the center is important.
Linda S. Karlsson, Mårten Sjöström
IEEE Trans. Circuits Syst. Video Technol.2
2007 Evaluation of a combined pre-processing and H.264-compression scheme for 3D integral images
abstract
To provide sufficient 3D-depth fidelity, integral imaging (II) requires an increase in spatial resolution of several orders of magnitude from today's 2D images. We have recently proposed a pre-processing and compression scheme for still II-frames based on forming a pseudo video sequence (PVS) from sub images (SI), which is later coded using the H.264/MPEG-4 AVC video coding standard. The scheme has shown good performance on a set of reference images. In this paper we first investigate and present how five different ways to select the SIs when forming the PVS affect the schemes compression efficiency. We also study how the II-frame structure relates to the performance of a PVS coding scheme. Finally we examine the nature of the coding artifacts which are specific to the evaluated PVS-schemes. We can conclude that for all except the most complex reference image, all evaluated SI selection orders significantly outperforms JPEG 2000 where compression ratios of up to 342:1, while still keeping PSNR > 30 dB, is achieved. We can also confirm that when selecting PVS-scheme, the scheme which results in a higher PVS-picture resolution should be preferred to maximize compression efficiency. Our study of the coded II-frames also indicates that the SI-based PVS, contrary to other PVS schemes, tends to distribute its coding artifacts more homogenously over all 3D-scene depths.
Roger Olsson, Mårten Sjöström, Youzhi Xu
VCIP2
2006 A Combined Pre-Processing and H.264-Compression Scheme for 3D Integral Images
abstract
The next evolutionary step in enhancing video communication fidelity is taken by adding scene depth. 3D video using integral imaging (II) is widely considered as the technique able to take this step. However, an increase in spatial resolution of several orders of magnitude from todays 2D video is required to provide a sufficient depth fidelity, which includes motion parallax. In this paper we propose a pre-processing and compression scheme that aims to enhance the compression efficiency of integral images. We first transform a still integral image into a pseudo video sequence consisting of sub-images, which is then compressed using an H.264 video encoder. The improvement in compression efficiency of using this scheme is evaluated and presented. An average PSNR increase of 5.7 dB or more, compared to JPEG 2000, is observed on a set of reference images.
Roger Olsson, Mårten Sjöström, Youzhi Xu
ICIP2
2005 Improved ROI video coding using variable Gaussian pre-filters and variance in intensity
abstract
In applications involving video over mobile phones or Internet, the limited quality depending on the transmission rate can be further improved by region-of-interest (ROI) coding. In this paper we present a preprocessing method using variable Gaussian filters controlled by a quality map indicating the distance to the ROI border. The border effects are reduced introducing a small improvement of the PSNR of the intensity component within the ROI after compression, compared to using only one low pass filter. With the compressed original sequence as a reference, the average PSNR was increased by 1.25 dB and 2.3 dB for 100 kbit/s and 150 kbit/s, respectively. A modified quality map is introduced using variance to exclude pixels, which are not visibly affected by the Gaussian filters, reducing computational complexity. Using less than 76% of the pixels gives no noticeable change in quality.
Linda S. Karlsson, Mårten Sjöström
ICIP (2)2
2000 Properties of smoothing with time gating
abstract
Estimates of transfer functions from measurements contain undesired information-noise-whose effect can partly be taken away by applying different smoothing techniques. One such smoothing technique suggests the removal of the noise by gating the impulse response in the time domain, hence the name time-gating. This paper shows that time-gating is a special case of a more general smoothing technique. The paper further contains a discussion concerning a preferable choice of time-gating window. An example applied to longitudinal beam transfer functions reflects the necessary trade-off between bias and variance.
Mårten Sjöström
ISCAS1