Markku Mäkitalo

dblp:86/9200 · also Markku J. Mäkitalo · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0001-8164-0031ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 6 since 2021Computer networks · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Scalable Hole-Filling for Real-Time Multi-GPU Light Field Path Tracing
abstract
Light field (LF) displays address the mismatch in focus cues present in traditional displays by triggering natural defocus blur and enabling motion parallax. They rely on geometrical optics, displaying rays from multiple angles of view. LF path tracing is computationally expensive for real-time applications, since it requires rendering multiple views. To reduce this computational complexity, spatially reprojecting pixels between views is commonly performed. Reusing pixels that are already rendered is cheaper than path tracing additional ones. However, when occluded areas are uncovered in some views, reprojection is not possible, creating holes in these views. Filling-in the holes requires extra path tracing computation. This paper investigates scalable hole-filling strategies for LF path tracing, using multiple GPUs to reach real-time performance. We propose an algorithm search optimization procedure to determine whether a specific assignment algorithm can be generalized across scenes, using the hole-filling time as a minimization function. In addition, we introduce DaSH (Discarded and Subsampled path traced Hole-filling), a novel method that reduces computation and divergence overhead in fixed-size hardware thread-blocks. Based on local pixel sparsity within pixel patches, DaSH adaptively subsamples and discards hole-filling rays. Our evaluation demonstrates that DaSH achieves significant performance gains while preserving the visual and structural quality of refocused light field images at the retina plane. The experiments demonstrate an average speedup factor of $1.8\times$1.8× for DaSH, compared to prior work, in a multi-GPU rendering system.
Erwan Leria, Markku Mäkitalo, Pekka Jääskeläinen, Mårten Sjöström
IEEE Trans. Vis. Comput. Graph.2
2025 Open Software Stack for Compression-Aware Adaptive Edge Offloading
abstract
Offloading a computationally complex task from an edge device can improve latency and its battery life. The additional network transfers increase power consumption and latency, but can be mitigated with compression at the cost of additional computation and distortion. Thus, a balance between compression efficiency and complexity must be found and maintained as the network conditions change. In this paper, we propose an open -source edge offloading software stack that decides between local and remote task execution and chooses the optimal compression method based on continuously monitored system metrics under user-defined constraints. We evaluate the system offloading a semantic segmentation task from a smartphone over WiFi-6 and 5G networks, using latency, intersection over union (IoU), and power consumption metrics. Portability, multi-tenancy, and granular profiling are achieved by leveraging the PoCL-R OpenCL implementation. In simulated network impairments, dynamically selecting the compression strategy achieves 2.1-10.7% average latency improvement but maintains the highest possible quality when network conditions allow meeting the latency budget. If the computational overhead of compression surpasses the network transfer overhead, the system can transmit images uncompressed. Field measurements under network impairments confirm the usability of the system and its ability to fall back to local execution.
Jakub Zádník, Robin Bijl, Jan Solanti, Erno Joensuu, Markku Mäkitalo, Pekka Jääskeläinen
WCNC5
2025 CV-Cast: Computer Vision-Oriented Linear Coding and Transmission
abstract
Remote inference allows lightweight edge devices, such as autonomous drones, to perform vision tasks exceeding their computational, energy, or processing delay budget. In such applications, reliable transmission of information is challenging due to high variations of channel quality. Traditional approaches involving spatio-temporal transforms, quantization, and entropy coding followed by digital transmission may be affected by a sudden decrease in quality (thedigital cliff) when the channel quality is less than expected during design. This problem can be addressed by using Linear Coding and Transmission (LCT), a joint source and channel coding scheme relying on linear operators only, allowing to achieve reconstructed per-pixel error commensurate with the wireless channel quality. In this paper, we propose CV-Cast: the first LCT scheme optimized for computer vision task accuracy instead of per-pixel distortion. Using this approach, for instance at 10 dB channel signal-to-noise ratio, CV-Cast requires transmitting 28% less symbols than a baseline LCT scheme in semantic segmentation and 15% in object detection tasks. Simulations involving a realistic 5G channel model confirm the smooth decrease in accuracy achieved with CV-Cast, while images encoded by JPEG or learned image coding (LIC) and transmitted using classical schemes at low Eb/N0 are subject to digital cliff.
Jakub Zádník, Michel Kieffer, Anthony Trioux, Markku Mäkitalo, Pekka Jääskeläinen
IEEE Trans. Mob. Comput.4
2025 Correction to "CV-Cast: Computer Vision-Oriented Linear Coding and Transmission"
abstract
In the above article [1], on page 1151, eq. (6), there is an error in the equation. The correct equation is: \begin{equation*} \min.\,\,D,\,\,\text{s.t.} \sum\limits_{k = 1}^K {{{\lambda }_k}\beta _k^2 \leqslant P.} \tag{6} \end{equation*} min.D,s.t.∑k=1Kλkβk2⩽P.(6)
Jakub Zádník, Michel Kieffer, Anthony Trioux, Markku Mäkitalo, Pekka Jääskeläinen
IEEE Trans. Mob. Comput.4
2025 Dynamic load balancing for real-time multiview path tracing on multi-GPU architectures
abstract
Stereoscopic and multiview rendering are used for virtual reality and the synthetic generation of light fields from three-dimensional scenes. Because rendering multiple views using ray tracing techniques is computationally expensive, the utilization of multiprocessor machines is necessary to achieve real-time frame rates. In this study, we propose a dynamic load-balancing algorithm for real-time multiview path tracing on multi-compute device platforms. The proposed algorithm was adapted to heterogeneous hardware combinations and dynamic scenes in real time. We show that on a heterogeneous dual-GPU platform, our implementation reduces the rendering time by an average of approximately 30%–50% compared with that of a uniform workload distribution, depending on the scene and number of views.
Erwan Leria, Markku Mäkitalo, Julius Ikkala, Pekka Jääskeläinen
Virtual Real. Intell. Hardw.2
2024 Interactive Multi-GPU Light Field Path Tracing Using Multi-Source Spatial Reprojection
abstract
Path tracing combined with multiview displays enables progress towards achieving ultrarealistic virtual reality. However, multiview displays based on light field technology impose a heavy workload for real-time graphics due to the large number of views to be rendered. In order to achieve low latency performance, computational effort can be reduced by path tracing only some views (source views), and synthesizing the remaining views (target views) through spatial reprojection, which reuses path traced pixels from source views to target views. Deciding the number of source views with respect to the computational resources is not trivial, since spatial reprojection introduces dependencies in the otherwise trivially parallel rendering pipeline and path tracing multiple source views increases the computation time.
Erwan Leria, Markku Mäkitalo, Pekka Jääskeläinen, Mårten Sjöström
VRST2
2024 Performance of Linear Coding and Transmission in Low-Latency Computer Vision Offloading
abstract
Image communication increasingly involves machine-to-machine delivery. For example, images acquired by an autonomous drone can be compressed and sent to an edge server over a wireless network for resource-intensive processing. Traditional compression techniques involving transform, quantization, and entropy coding reach high compression efficiency, but channel conditions worse than expected may lead to a sharp decrease in the decoded image quality. As an alternative, Linear Coding and Transmission (LCT) systems have been proposed to avoid this digital cliff problem: The reconstructed image quality decreases gradually as channel conditions degrade. This paper presents a comprehensive evaluation of computer vision tasks with input images processed and transmitted using LCT. It also analyses the benefits of network retraining, accounting for impairments due to LCT and noisy channel. Considering object detection and semantic segmentation over images transmitted and received by LCT systems, we show that the task accuracy degrades smoothly when the channel quality decreases, avoiding the cliff effect. Retraining with noisy images processed by LCT restores detection mAP degradation from 23.8% to 4.4% and segmentation mIoU degradation from 43.2% to 8.1 % when the channel signal-to-noise ratio is 10 dB.
Jakub Zádník, Anthony Trioux, Michel Kieffer, Markku Mäkitalo, François-Xavier Coudoux, Patrick Corlay, Pekka Jääskeläinen
WCNC4
2022 Real-Time Light Field Path Tracing
Markku Mäkitalo, Erwan Leria, Julius Ikkala, Pekka Jääskeläinen
CGI1
2022 Pruned Lightweight Encoders for Computer Vision
abstract
Latency-critical computer vision systems, such as autonomous driving or drone control, require fast image or video compression when offloading neural network inference to a remote computer. To ensure low latency on a near-sensor edge device, we propose the use of lightweight encoders with constant bitrate and pruned encoding configurations, namely, ASTC and JPEG XS. Pruning introduces significant distortion which we show can be recovered by retraining the neural network with compressed data after decompression. Such an approach does not modify the network architecture or require coding format modifications. By retraining with compressed datasets, we reduced the classification accuracy and segmentation mean intersection over union (mIoU) degradation due to ASTC compression to 4.9-5.0 percentage points (pp) and 4.4-4.0 pp, respectively. With the same method, the mIoU lost due to JPEG XS compression at the main profile was restored to 2.7-2.3 pp. In terms of encoding speed, our ASTC encoder implementation is 2.3x faster than JPEG. Even though the JPEG XS reference encoder requires optimizations to reach low latency, we showed that disabling significance flag coding saves 22–23% of encoding time at the cost of 0.4-0.3 mIoU after retraining.
Jakub Zádník, Markku Mäkitalo, Pekka Jääskeläinen
MMSP2
2021 DDISH-GI: Dynamic Distributed Spherical Harmonics Global Illumination
Julius Ikkala, Petrus E. J. Kivi, Joel Alanko, Markku Mäkitalo, Pekka Jääskeläinen
CGI4
2019 Blockwise Multi-Order Feature Regression for Real-Time Path-Tracing Reconstruction
abstract
Path tracing produces realistic results including global illumination using a unified simple rendering pipeline. Reducing the amount of noise to imperceptible levels without post-processing requires thousands of samples per pixel (spp), while currently it is only possible to render extremely noisy 1 spp frames in real time with desktop GPUs. However, post-processing can utilize feature buffers, which contain noise-free auxiliary data available in the rendering pipeline. Previously, regression-based noise filtering methods have only been used in offline rendering due to their high computational cost. In this article we propose a novel regression-based reconstruction pipeline, called Blockwise Multi-Order Feature Regression (BMFR), tailored for path-traced 1 spp inputs that runs in real time. The high speed is achieved with a fast implementation of augmented QR factorization and by using stochastic regularization to address rank-deficient feature data. The proposed algorithm is 1.8× faster than the previous state-of-the-art real-time path-tracing reconstruction method while producing better quality frame sequences.
Matias Koskela, Kalle Immonen, Markku Mäkitalo, Alessandro Foi, Timo Viitanen, Pekka Jääskeläinen, Heikki Kultala, Jarmo Takala
ACM Trans. Graph.3
2014 Noise Parameter Mismatch in Variance Stabilization, With an Application to Poisson-Gaussian Noise Estimation
abstract
In digital imaging, there is often a need to produce estimates of the parameters that define the chosen noise model. We investigate how the mismatch between the estimated and true parameter values affects the stabilization of variance of signal-dependent noise. As a practical application of the general theoretical considerations, we devise a novel approach for estimating Poisson–Gaussian noise parameters from a single image, combining variance-stabilization and noise estimation for additive Gaussian noise. Furthermore, we construct a simple algorithm implementing the devised approach. We observe that when combined with optimized rational variance-stabilizing transformations, the algorithm produces results that are competitive with those of a state-of-the-art Poisson–Gaussian estimator.
Markku Mäkitalo, Alessandro Foi
IEEE Trans. Image Process.1
2013 Optimal Inversion of the Generalized Anscombe Transformation for Poisson-Gaussian Noise
abstract
Many digital imaging devices operate by successive photon-to-electron, electron-to-voltage, and voltage-to-digit conversions. These processes are subject to various signal-dependent errors, which are typically modeled as Poisson-Gaussian noise. The removal of such noise can be effected indirectly by applying a variance-stabilizing transformation (VST) to the noisy data, denoising the stabilized data with a Gaussian denoising algorithm, and finally applying an inverse VST to the denoised data. The generalized Anscombe transformation (GAT) is often used for variance stabilization, but its unbiased inverse transformation has not been rigorously studied in the past. We introduce the exact unbiased inverse of the GAT and show that it plays an integral part in ensuring accurate denoising results. We demonstrate that this exact inverse leads to state-of-the-art results without any notable increase in the computational complexity compared to the other inverses. We also show that this inverse is optimal in the sense that it can be interpreted as a maximum likelihood inverse. Moreover, we thoroughly analyze the behavior of the proposed inverse, which also enables us to derive a closed-form approximation for it. This paper generalizes our work on the exact unbiased inverse of the Anscombe transformation, which we have presented earlier for the removal of pure Poisson noise.
Markku Mäkitalo, Alessandro Foi
IEEE Trans. Image Process.1
2012 Poisson-gaussian denoising using the exact unbiased inverse of the generalized anscombe transformation
abstract
The characteristic errors of many digital imaging devices can be modelled as Poisson-Gaussian noise, the removal of which can be approached indirectly through variance stabilization. The generalized Anscombe transformation (GAT) is commonly used for stabilization, but rigorous studies regarding its unbiased inverse transformation have been neglected. We introduce the exact unbiased inverse of the GAT, show that it is of essential importance for ensuring accurate denoising, and demonstrate that our approach leads to state-of-the-art results. This paper generalizes our earlier work, in which we presented an exact unbiased inverse of the Anscombe transformation for the case of pure Poisson noise removal.
Markku Mäkitalo, Alessandro Foi
ICASSP1
2011 Optimal Inversion of the Anscombe Transformation in Low-Count Poisson Image Denoising
abstract
The removal of Poisson noise is often performed through the following three-step procedure. First, the noise variance is stabilized by applying the Anscombe root transformation to the data, producing a signal in which the noise can be treated as additive Gaussian with unitary variance. Second, the noise is removed using a conventional denoising algorithm for additive white Gaussian noise. Third, an inverse transformation is applied to the denoised signal, obtaining the estimate of the signal of interest. The choice of the proper inverse transformation is crucial in order to minimize the bias error which arises when the nonlinear forward transformation is applied. We introduce optimal inverses for the Anscombe transformation, in particular the exact unbiased inverse, a maximum likelihood (ML) inverse, and a more sophisticated minimum mean square error (MMSE) inverse. We then present an experimental analysis using a few state-of-the-art denoising algorithms and show that the estimation can be consistently improved by applying the exact unbiased inverse, particularly at the low-count regime. This results in a very efficient filtering solution that is competitive with some of the best existing methods for Poisson image denoising.
Markku Mäkitalo, Alessandro Foi
IEEE Trans. Image Process.1
2011 A Closed-Form Approximation of the Exact Unbiased Inverse of the Anscombe Variance-Stabilizing Transformation
abstract
We presented an exact unbiased inverse of the Anscombe variance-stabilizing transformation in [M. Mäkitalo and A. Foi, "Optimal inversion of the Anscombe transformation in low-count Poisson image denoising," IEEE Trans. Image Process., vol. 20, no. 1, pp. 99-109, Jan. 2011.] and showed that when applied to Poisson image denoising, the combination of variance stabilization and state-of-the-art Gaussian denoising algorithms is competitive with some of the best Poisson denoising algorithms. We also provided a MATLAB implementation of our method, where the exact unbiased inverse transformation appears in nonanalytical form. Here, we propose a closed-form approximation of the exact unbiased inverse in order to facilitate the use of this inverse. The proposed approximation produces results equivalent to those obtained with the accurate (nonanalytical) exact unbiased inverse, and thus, notably better than one would get with the asymptotically unbiased inverse transformation that is commonly used in applications.
Markku Mäkitalo, Alessandro Foi
IEEE Trans. Image Process.1