VLDB 2026 Research / reviewers in the wild / expert
Tomas Akenine-Möller
dblp:a/TomasAkenineMoller · also Tomas Möller
· DBLP profile ↗
42ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0001-6226-3170ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
16 papers |
Rendering · 62% Image and video coding · 16% Image and video processing · 15% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
GPUs and heterogeneous computing · 80% Processor architecture and microarchitecture · 15% Memory systems · 5% |
Topics — the 30 heaviest of 36, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video coding › neural compression
neural texture compression |
0.7 | 1 | 2023 | Random-Access Neural Compression of Material Textures · ACM Trans. Graph. 2023 |
Image and video processing › texture analysis
texture representation |
0.7 | 1 | 2023 | Random-Access Neural Compression of Material Textures · ACM Trans. Graph. 2023 |
Rendering
visibility culling |
0.4 | 3 | 2015 | Masked depth culling for graphics hardware · ACM Trans. Graph. 2015 Automatic pre-tessellation culling · ACM Trans. Graph. 2009 PCU: the programmable culling unit · ACM Trans. Graph. 2007 |
Rendering
antialiasing |
0.3 | 2 | 2013 | A4: asynchronous adaptive anti-aliasing using shared memory · ACM Trans. Graph. 2013 High-quality spatio-temporal rendering using semi-analytical visibility · ACM Trans. Graph. 2011 |
Rendering
visibility computation |
0.3 | 2 | 2012 | High-quality curve rendering using line sampled visibility · ACM Trans. Graph. 2012 High-quality spatio-temporal rendering using semi-analytical visibility · ACM Trans. Graph. 2011 |
Rendering › ray tracing
monte carlo ray tracing |
0.2 | 1 | 2016 | Texture space caching and reconstruction for ray tracing · ACM Trans. Graph. 2016 |
Rendering
ray tracing |
0.2 | 2 | 2014 | Dynamic ray stream traversal · ACM Trans. Graph. 2014 Soft shadow volumes for ray tracing · ACM Trans. Graph. 2005 |
GPUs and heterogeneous computing
graphics hardware |
0.2 | 1 | 2015 | Masked depth culling for graphics hardware · ACM Trans. Graph. 2015 |
Geometric modeling and processing › spatial data structures
bounding volume hierarchy |
0.2 | 1 | 2014 | Dynamic ray stream traversal · ACM Trans. Graph. 2014 |
Rendering › shading
shading reuse |
0.2 | 1 | 2014 | AMFS: adaptive multi-frequency shading for future graphics processors · ACM Trans. Graph. 2014 |
Rendering › GPU rendering
variable rate shading |
0.2 | 1 | 2014 | AMFS: adaptive multi-frequency shading for future graphics processors · ACM Trans. Graph. 2014 |
Rendering
real-time rendering |
0.2 | 2 | 2013 | A4: asynchronous adaptive anti-aliasing using shared memory · ACM Trans. Graph. 2013 High dynamic range texture compression for graphics hardware · ACM Trans. Graph. 2006 |
Rendering › geometric rendering
curve rendering |
0.1 | 1 | 2012 | High-quality curve rendering using line sampled visibility · ACM Trans. Graph. 2012 |
Rendering › temporal rendering
motion blur |
0.1 | 1 | 2011 | High-quality spatio-temporal rendering using semi-analytical visibility · ACM Trans. Graph. 2011 |
Rendering
graphics hardware |
0.1 | 2 | 2007 | PCU: the programmable culling unit · ACM Trans. Graph. 2007 Graphics for the masses: a hardware rasterization architecture for mobile phones · ACM Trans. Graph. 2003 |
Rendering
global illumination |
0.1 | 2 | 2005 | Precomputed local radiance transfer for real-time lighting design · ACM Trans. Graph. 2005 Wavelet importance sampling: efficiently evaluating products of complex functions · ACM Trans. Graph. 2005 |
Geometric modeling and processing › mesh generation
surface meshing |
0.1 | 1 | 2009 | Automatic pre-tessellation culling · ACM Trans. Graph. 2009 |
GPUs and heterogeneous computing › GPU architecture
energy-efficient GPU design |
0.1 | 1 | 2008 | Graphics Processing Units for Handhelds · Proc. IEEE 2008 |
GPUs and heterogeneous computing › embedded GPU
mobile GPU |
0.1 | 1 | 2008 | Graphics Processing Units for Handhelds · Proc. IEEE 2008 |
Image and video processing › image restoration
denoising |
0.1 | 1 | 2016 | Texture space caching and reconstruction for ray tracing · ACM Trans. Graph. 2016 |
Image and video coding › texture compression
high dynamic range texture compression |
0.1 | 1 | 2006 | High dynamic range texture compression for graphics hardware · ACM Trans. Graph. 2006 |
Image and video coding
texture compression |
0.1 | 1 | 2006 | High dynamic range texture compression for graphics hardware · ACM Trans. Graph. 2006 |
Rendering › monte carlo rendering
importance sampling |
0.1 | 1 | 2005 | Wavelet importance sampling: efficiently evaluating products of complex functions · ACM Trans. Graph. 2005 |
Rendering › global illumination
precomputed radiance transfer |
0.1 | 1 | 2005 | Precomputed local radiance transfer for real-time lighting design · ACM Trans. Graph. 2005 |
Rendering › shadow rendering
shadow volumes |
0.1 | 1 | 2005 | Soft shadow volumes for ray tracing · ACM Trans. Graph. 2005 |
Rendering › shadow rendering
soft shadows |
0.1 | 1 | 2005 | Soft shadow volumes for ray tracing · ACM Trans. Graph. 2005 |
Rendering
shadow rendering |
0.0 | 1 | 2003 | A geometry-based soft shadow volume algorithm using graphics hardware · ACM Trans. Graph. 2003 |
Rendering
texture mapping |
0.0 | 1 | 2003 | Graphics for the masses: a hardware rasterization architecture for mobile phones · ACM Trans. Graph. 2003 |
Rendering › global illumination
ambient occlusion |
0.0 | 1 | 2011 | High-quality spatio-temporal rendering using semi-analytical visibility · ACM Trans. Graph. 2011 |
Memory systems › memory bandwidth management
memory bandwidth reduction |
0.0 | 1 | 2008 | Graphics Processing Units for Handhelds · Proc. IEEE 2008 |
Methods — techniques the papers use, named apart from their topics
neural network compression · 0.7custom training implementation · 0.7per-sample mask · 0.4layered depth representation · 0.4SIMD · 0.4texture space filtering · 0.2secondary ray queries · 0.2linear regression · 0.2multi-frequency shading · 0.2cache hierarchy optimization · 0.2bandwidth reduction algorithms · 0.1fragment program unit · 0.1discard instruction · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Estimates of Temporal Edge Detection Filters in Human VisionabstractEdge detection is an important process in human visual processing. However, as far as we know, few attempts have been made to map the temporal edge detection filters in human vision. To that end, we devised a user study and collected data from which we derived estimates of human temporal edge detection filters based on three different models, including the derivative of the infinite symmetric exponential function and temporal contrast sensitivity function. We analyze our findings using several different methods, including extending the filter to higher frequencies than were shown during the experiment. In addition, we show a proof of concept that our filter may be used in spatiotemporal image quality metrics by incorporating it into a flicker detection pipeline. Pontus Ebelin, Gyorgy Denes, Tomas Akenine-Möller, Kalle Åström, Magnus Oskarsson, William McIlhagga |
ACM Trans. Appl. Percept. | 3 |
| 2023 | Random-Access Neural Compression of Material TexturesabstractThe continuous advancement of photorealism in rendering is accompanied by a growth in texture data and, consequently, increasing storage and memory demands. To address this issue, we propose a novel neural compression technique specifically designed for material textures. We unlock two more levels of detail, i.e., 16× more texels, using low bitrate compression, with image quality that is better than advanced image compression techniques, such as AVIF and JPEG XL. At the same time, our method allows on-demand, real-time decompression with random access similar to block texture compression on GPUs, enabling compression on disk and memory. The key idea behind our approach is compressing multiple material textures and their mipmap chains together, and using a small neural network, that is optimized for each material, to decompress them. Finally, we use a custom training implementation to achieve practical compression speeds, whose performance surpasses that of general frameworks, like PyTorch, by an order of magnitude. Karthikeyan Vaidyanathan, Marco Salvi, Bartlomiej Wronski, Tomas Akenine-Möller, Pontus Ebelin, Aaron E. Lefohn |
ACM Trans. Graph. | 4 |
| 2022 | Spatiotemporal Blue Noise MasksabstractBlue noise error patterns are well suited to human perception, and when applied to stochastic rendering techniques, blue noise masks (blue noise textures) minimize unwanted low-frequency noise in the final image. Current methods of applying blue noise masks at each frame independently produce white noise frequency spectra temporally. This white noise results in slower integration convergence over time and unstable results when filtered temporally. Unfortunately, achieving temporally stable blue noise distributions is non-trivial since 3D blue noise does not exhibit the desired 2D blue noise properties, and alternative approaches degrade the spatial blue noise qualities. We propose novel blue noise patterns that, when animated, produce values at a pixel that are well distributed over time, converge rapidly for Monte Carlo integration, and are more stable under TAA, while still retaining spatial blue noise properties. To do so, we propose an extension to the well-known void and cluster algorithm that reformulates the underlying energy function to produce spatiotemporal blue noise masks. These masks exhibit blue noise frequency spectra in both the spatial and temporal domains, resulting in visually pleasing error patterns, rapid convergence speeds, and increased stability when filtered temporally. We demonstrate these improvements on a variety of applications, including dithering, stochastic transparency, ambient occlusion, and volumetric rendering. By extending spatial blue noise to spatiotemporal blue noise, we overcome the convergence limitations of prior blue noise works, enabling new applications for blue noise distributions. Alan Wolfe, Nathan Morrical, Tomas Akenine-Möller, Ravi Ramamoorthi |
EGSR (ST) | 3 |
| 2017 | Ray Accelerator: Efficient and Flexible Ray Tracing on a Heterogeneous ArchitectureabstractAbstract We present a hybrid ray tracing system, where the work is divided between the CPU cores and the GPU in an integrated chip, and communication occurs via shared memory. Rays are organized in large packets that can be distributed among the two units as needed. Testing visibility between rays and the scene is mostly performed using an optimized kernel on the GPU, but the CPU can help as necessary. The CPU cores typically handle most or all shading, which makes it easy to support complex appearances. For efficiency, the CPU cores shade whole batches of rays by sorting them on material and shading each material using a vectorized kernel. In addition, we introduce a method to support light paths with arbitrary recursion, such as multiple recursive Whitted‐style ray tracing and adaptive sampling where the result of a ray is examined before sending the next, while still batching up rays for the benefit of GPU‐accelerated traversal and vectorized shading. This allows our system to achieve high rendering performance while maintaining the flexibility to accommodate different rendering algorithms. Rasmus Barringer, Tomas Akenine-Möller |
Comput. Graph. Forum | 3 |
| 2017 | Time-Continuous Quasi-Monte Carlo Ray TracingabstractAbstract Domain‐continuous visibility determination algorithms have proved to be very efficient at reducing noise otherwise prevalent in stochastic sampling. Even though they come with an increased overhead in terms of geometrical tests and visibility information management, their analytical nature provides such a rich integral that the pay‐off is often worth it. This paper presents a time‐continuous, primary visibility algorithm for motion blur aimed at ray tracing. Two novel intersection tests are derived and implemented. The first is for ray versus moving triangle and the second for ray versus moving AABB intersection. A novel take on shading is presented as well, where the time continuum of visible geometry is adaptively point‐sampled. Static geometry is handled using supplemental stochastic rays in order to reduce spatial aliasing. Finally, a prototype ray tracer with a full time‐continuous traversal kernel is presented in detail. The results are based on a variety of test scenarios and show that even though our time‐continuous algorithm has limitations, it outperforms multi‐jittered quasi‐Monte Carlo ray tracing in terms of image quality at equal rendering time, within wide sampling rate ranges. Carl Johan Gribel, Tomas Akenine-Möller |
Comput. Graph. Forum | 2 |
| 2016 | Texture space caching and reconstruction for ray tracingabstractWe present a texture space caching and reconstruction system for Monte Carlo ray tracing. Our system gathers and filters shading on-demand, including querying secondary rays, directly within a filter footprint around the current shading point. We shade on local grids in texture space with primary visibility decoupled from shading. Unique filters can be applied per material, where any terms of the shader can be chosen to be included in each kernel. This is a departure from recent screen space image reconstruction techniques, which typically use a single, complex kernel with a set of large auxiliary guide images as input. We show a number of high-performance use cases for our system, including interactive denoising of Monte Carlo ray tracing with motion/defocus blur, spatial and temporal shading reuse, cached product importance sampling, and filters based on linear regression in texture space. Jacob Munkberg, Jon Hasselgren, Petrik Clarberg, Tomas Akenine-Möller |
ACM Trans. Graph. | 5 |
| 2015 | Filtered Stochastic Shadow Mapping Using a Layered ApproachabstractAbstract Given a stochastic shadow map rendered with motion blur, our goal is to render an image from the eye with motion‐blurred shadows with as little noise as possible. We use a layered approach in the shadow map and reproject samples along the average motion vector, and then perform lookups in this representation. Our results include substantially improved shadow quality compared to previous work and a fast graphics processing unit (GPU) implementation. In addition, we devise a set of scenes that are designed to bring out and show problematic cases for motion‐blurred shadows. These scenes have difficult occlusion characteristics, and may be used in future research on this topic. Jon Hasselgren, Jacob Munkberg, Tomas Akenine-Möller |
Comput. Graph. Forum | 4 |
| 2015 | Masked depth culling for graphics hardwareabstractHierarchical depth culling is an important optimization, which is present in all modern high performance graphics processors. We present a novel culling algorithm based on a layered depth representation, with a per-sample mask indicating which layer each sample belongs to. Our algorithm is feed forward in nature in contrast to previous work, which rely on a delayed feedback loop. It is simple to implement and has fewer constraints than competing algorithms, which makes it easier to load-balance a hardware architecture. Compared to previous work our algorithm performs very well, and it will often reach over 90% of the efficiency of an optimal culling oracle. Furthermore, we can reduce bandwidth by up to 16% by compressing the hierarchical depth buffer. Jon Hasselgren, Tomas Akenine-Möller |
ACM Trans. Graph. | 3 |
| 2015 | A performance and energy evaluation of many-light rendering algorithms
Björn Johnsson, Tomas Akenine-Möller |
Vis. Comput. | 2 |
| 2014 | Adaptive texture space shading for stochastic renderingabstractAbstract When rendering effects such as motion blur and defocus blur, shading can become very expensive if done in a naïve way, i.e. shading each visibility sample. To improve performance, previous work often decouple shading from visibility sampling using shader caching algorithms. We present a novel technique for reusing shading in a stochastic rasterizer. Shading is computed hierarchically and sparsely in an object‐space texture, and by selecting an appropriate mipmap level for each triangle, we ensure that the shading rate is sufficiently high so that no noticeable blurring is introduced in the rendered image. Furthermore, with a two‐pass algorithm, we separate shading from reuse and thus avoid GPU thread synchronization. Our method runs at real‐time frame rates and is up to 3 × faster than previous methods. This is an important step forward for stochastic rasterization in real time. Jon Hasselgren, Robert Toth, Tomas Akenine-Möller |
Comput. Graph. Forum | 4 |
| 2014 | Layered Reconstruction for Defocus and Motion BlurabstractAbstract Light field reconstruction algorithms can substantially decrease the noise in stochastically rendered images. Recent algorithms for defocus blur alone are both fast and accurate. However, motion blur is a considerably more complex type of camera effect, and as a consequence, current algorithms are either slow or too imprecise to use in high quality rendering. We extend previous work on real‐time light field reconstruction for defocus blur to handle the case of simultaneous defocus and motion blur. By carefully introducing a few approximations, we derive a very efficient sheared reconstruction filter, which produces high quality images even for a low number of input samples. Our algorithm is temporally robust, and is about two orders of magnitude faster than previous work, making it suitable for both real‐time rendering and as a post‐processing pass for offline rendering. Jacob Munkberg, Karthikeyan Vaidyanathan, Jon Hasselgren, Petrik Clarberg, Tomas Akenine-Möller |
Comput. Graph. Forum | 5 |
| 2014 | Dynamic ray stream traversalabstractWhile each new generation of processors gets larger caches and more compute power, external memory bandwidth capabilities increase at a much lower pace. Additionally, processors are equipped with wide vector units that require low instruction level divergence to be efficiently utilized. In order to exploit these trends for ray tracing, we present an alternative to traditional depth-first ray traversal that takes advantage of the available cache hierarchy, and provides high SIMD efficiency, while keeping memory bus traffic low. Our main contribution is an efficient algorithm for traversing large packets of rays against a bounding volume hierarchy in a way that groups coherent rays during traversal. In contrast to previous large packet traversal methods, our algorithm allows for individual traversal order for each ray, which is essential for efficient ray tracing. Ray tracing algorithms is a mature research field in computer graphics, and despite this, our new technique increases traversal performance by 36--53%, and is applicable to most ray tracers. Rasmus Barringer, Tomas Akenine-Möller |
ACM Trans. Graph. | 2 |
| 2014 | AMFS: adaptive multi-frequency shading for future graphics processorsabstractWe propose a powerful hardware architecture for pixel shading, which enables flexible control of shading rates and automatic shading reuse between triangles in tessellated primitives. The main goal is efficient pixel shading for moderately to finely tessellated geometry, which is not handled well by current GPUs. Our method effectively decouples the cost of pixel shading from the geometric complexity. It thereby enables a wider use of tessellation and fine geometry, even at very limited power budgets. The core idea is to shade over small local grids in parametric patch space, and reuse shading for nearby samples. We also support the decomposition of shaders into multiple parts, which are shaded at different frequencies. Shading rates can be locally and adaptively controlled, in order to direct the computations to visually important areas and to provide performance scaling with a graceful degradation of quality. Another important benefit of shading in patch space is that it allows efficient rendering of distribution effects, which further closes the gap between real-time and offline rendering. Petrik Clarberg, Robert Toth, Jon Hasselgren, Jim Nilsson, Tomas Akenine-Möller |
ACM Trans. Graph. | 5 |
| 2013 | Stochastic Depth Buffer Compression using Generalized Plane EncodingabstractAbstract In this paper, we derive compact representations of the depth function for a triangle undergoing motion or defocus blur. Unlike a static primitive, where the depth function is planar, the depth function is a rational function in time and the lens parameters. Furthermore, we show how these compact depth functions can be used to design an efficient depth buffer compressor/decompressor, which significantly lowers total depth buffer bandwidth usage for a range of test scenes. In addition, our compressor/decompressor is simpler in the number of operations needed to execute, which makes our algorithm more amenable for hardware implementation than previous methods. Jacob Munkberg, Tomas Akenine-Möller |
Comput. Graph. Forum | 3 |
| 2013 | A4: asynchronous adaptive anti-aliasing using shared memoryabstractEdge aliasing continues to be one of the most prominent problems in real-time graphics, e.g., in games. We present a novel algorithm that uses shared memory between the GPU and the CPU so that these two units can work in concert to solve the edge aliasing problem rapidly. Our system renders the scene as usual on the GPU with one sample per pixel. At the same time, our novel edge aliasing algorithm is executed asynchronously on the CPU. First, a sparse set of important pixels is created. This set may include pixels with geometric silhouette edges, discontinuities in the frame buffer, and pixels/polygons under user-guided artistic control. After that, the CPU runs our sparse rasterizer and fragment shader, which is parallel and SIMD:ified, and directly accesses shared resources (e.g., render targets created by the GPU). Our system can render a scene with shadow mapping with adaptive anti-aliasing with 16 samples per important pixel faster than the GPU with 8 samples per pixel using multi-sampling anti-aliasing. Since our system consists of an extensive code base, it will be released to the public for exploration and usage. Rasmus Barringer, Tomas Akenine-Möller |
ACM Trans. Graph. | 2 |
| 2012 | Efficient Depth of Field Rasterization Using a Tile Test Based on Half-Space CullingabstractAbstract For depth of field (DOF) rasterization, it is often desired to have an efficient tile versus triangle test, which can conservatively compute which samples on the lens that need to execute the sample‐in‐triangle test. We present a novel test for this, which is optimal in the sense that the region on the lens cannot be further reduced. Our test is based on removing half‐space regions of the (u, v) ‐space on the lens, from where the triangle definitely cannot be seen through a tile of pixels. We find the intersection of all such regions exactly, and the resulting region can be used to reduce the number of sample‐in‐triangle tests that need to be performed. Our main contribution is that the theory we develop provides a limit for how efficient a practical tile versus defocused triangle test ever can become. To verify our work, we also develop a conceptual implementation for DOF rasterization based on our new theory. We show that the number of arithmetic operations involved in the rasterization process can be reduced. More importantly, with a tile test, multi‐sampling anti‐aliasing can be used which may reduce shader executions and the related memory bandwidth usage substantially. In general, this can be translated to a performance increase and/or power savings. Tomas Akenine-Möller, Robert Toth, Jacob Munkberg, Jon Hasselgren |
Comput. Graph. Forum | 1 |
| 2012 | Per-Vertex Defocus Blur for Stochastic RasterizationabstractAbstract We present user‐controllable and plausible defocus blur for a stochastic rasterizer. We modify circle of confusion coefficients per vertex to express more general defocus blur, and show how the method can be applied to limit the foreground blur, extend the in‐focus range, simulate tilt‐shift photography, and specify per‐object defocus blur. Furthermore, with two simplifying assumptions, we show that existing triangle coverage tests and tile culling tests can be used with very modest modifications. Our solution is temporally stable and handles simultaneous motion blur and depth of field. Jacob Munkberg, Robert Toth, Tomas Akenine-Möller |
Comput. Graph. Forum | 3 |
| 2012 | High-quality curve rendering using line sampled visibilityabstractComputing accurate visibility for thin primitives, such as hair strands, fur, grass, at all scales remains difficult or expensive. To that end, we present an efficient visibility algorithm based on spatial line sampling, and a novel intersection algorithm between line sample planes and Bézier splines with varying thickness. Our algorithm produces accurate visibility both when the projected width of the curve is a tiny fraction of a pixel, and when the projected width is tens of pixels. In addition, we present a rapid resolve procedure that computes final visibility. Using an optimized implementation running on graphics processors, we can render tens of thousands long hair strands with noise-free visibility at near-interactive rates. Rasmus Barringer, Carl Johan Gribel, Tomas Akenine-Möller |
ACM Trans. Graph. | 3 |
| 2011 | High-quality spatio-temporal rendering using semi-analytical visibilityabstractWe present a novel visibility algorithm for rendering motion blur with per-pixel anti-aliasing. Our algorithm uses a number of line samples over a rectangular group of pixels, and together with the time dimension, a two-dimensional spatio-temporal visibility problem needs to be solved per line sample. In a coarse culling step, our algorithm first uses a bounding volume hierarchy to rapidly remove geometry that does not overlap with the current line sample. For the remaining triangles, we approximate each triangle's depth function, along the line and along the time dimension, with a number of patch triangles. We resolve for the final color using an analytical visibility algorithm with depth sorting, simple occlusion culling, and clipping. Shading is decoupled from visibility, and we use a shading cache for efficient reuse of shaded values. In our results, we show practically noise-free renderings of motion blur with high-quality spatial anti-aliasing and with competitive rendering times. We also demonstrate that our algorithm, with some adjustments, can be used to accurately compute motion blurred ambient occlusion. Carl Johan Gribel, Rasmus Barringer, Tomas Akenine-Möller |
ACM Trans. Graph. | 3 |
| 2011 | Efficient multi-view ray tracing using edge detection and shader reuse
Björn Johnsson, Jacob Munkberg, Petrik Clarberg, Jon Hasselgren, Tomas Akenine-Möller |
Vis. Comput. | 6 |
| 2010 | An Optimizing Compiler for Automatic Shader BoundingabstractAbstract Programmable shading provides artistic control over materials and geometry, but the black box nature of shaders makes some rendering optimizations difficult to apply. In many cases, it is desirable to compute bounds of shaders in order to speed up rendering. A bounding shader can be automatically derived from the original shader by a compiler using interval analysis, but creating optimized interval arithmetic code is non‐trivial. A key insight in this paper is that shaders contain metadata that can be automatically extracted by the compiler using data flow analysis. We present a number of domain‐specific optimizations that make the generated code faster, while computing the same bounds as before. This enables a wider use and opens up possibilities for more efficient rendering. Our results show that on average 42–44% of the shader instructions can be eliminated for a common use case: single‐sided bounding shaders used in lightcuts and importance sampling. Petrik Clarberg, Robert Toth, Jon Hasselgren, Tomas Akenine-Möller |
Comput. Graph. Forum | 4 |
| 2010 | Error-bounded lossy compression of floating-point color buffers using quadtree decomposition
Jim Rasmusson, Jacob Ström, Tomas Akenine-Möller |
Vis. Comput. | 3 |
| 2009 | Bounding Volume Hierarchies of Slab Cut BallsabstractAbstract We introduce a bounding volume hierarchy based on the Slab Cut Ball. This novel type of enclosing shape provides an attractive balance between tightness of fit, cost of overlap testing, and memory requirement. The hierarchy construction algorithm includes a new method for the construction of tight bounding volumes in worst case O(n) time, which means our tree data structure is constructed in O(n log n) time using traditional top‐down building methods. A fast overlap test method between two slab cut balls is also proposed, requiring as few as 28–99 arithmetic operations, including the transformation cost. Practical collision detection experiments confirm that our tree data structure is amenable for high performance collision queries. In all the tested benchmarks, our bounding volume hierarchy consistently gives performance improvements over the sphere tree, and it is also faster than the OBB tree in five out of six scenes. In particular, our method is asymptotically faster than the sphere tree, and it also outperforms the OBB tree, in close proximity situations. Thomas Larsson, Tomas Akenine-Möller |
Comput. Graph. Forum | 2 |
| 2009 | Automatic pre-tessellation cullingabstractGraphics processing units supporting tessellation of curved surfaces with displacement mapping exist today. Still, to our knowledge, culling only occurs after tessellation, that is, after the base primitives have been tessellated into triangles. We introduce an algorithm for automatically computing tight positional and normal bounds on the fly for a base primitive. These bounds are derived from an arbitrary vertex shader program, which may include a curved surface evaluation and different types of displacements, for example. The obtained bounds are used for backface, view frustum, and occlusion culling before tessellation. For highly tessellated scenes, we show that up to 80% of the vertex shader instructions can be avoided, which implies an “instruction speedup” of 5×. Our technique can also be used for offline software rendering. Jon Hasselgren, Jacob Munkberg, Tomas Akenine-Möller |
ACM Trans. Graph. | 3 |
| 2008 | Practical Product Importance Sampling for Direct IlluminationabstractAbstract We present a practical algorithm for sampling the product of environment map lighting and surface reflectance. Our method builds on wavelet‐based importance sampling, but has a number of important advantages over previous methods. Most importantly, we avoid using precomputed reflectance functions by sampling the BRDF on‐the‐fly. Hence, all types of materials can be handled, including anisotropic and spatially varying BRDFs, as well as procedural shaders. This also opens up for using very high resolution, uncompressed, environment maps. Our results show that this gives a significant reduction of variance compared to using lower resolution approximations. In addition, we study the wavelet product, and present a faster algorithm geared for sampling purposes. For our application, the computations are reduced to a simple quadtree‐based multiplication. We build the BRDF approximation and evaluate the product in a single tree traversal, which makes the algorithm both faster and more flexible than previous methods. Petrik Clarberg, Tomas Akenine-Möller |
Comput. Graph. Forum | 2 |
| 2008 | Exploiting Visibility Correlation in Direct IlluminationabstractAbstract The visibility function in direct illumination describes the binary visibility over a light source, e.g., an environment map. Intuitively, the visibility is often strongly correlated between nearby locations in time and space, but exploiting this correlation without introducing noticeable errors is a hard problem. In this paper, we first study the statistical characteristics of the visibility function. Then, we propose a robust and unbiased method for using estimated visibility information to improve the quality of Monte Carlo evaluation of direct illumination. Our method is based on the theory of control variates, and it can be used on top of existing state‐of‐the‐art schemes for importance sampling. The visibility estimation is obtained by sparsely sampling and caching the 4D visibility field in a compact bitwise representation. In addition to Monte Carlo rendering, the stored visibility information can be used in a number of other applications, for example, ambient occlusion and lighting design. Petrik Clarberg, Tomas Akenine-Möller |
Comput. Graph. Forum | 2 |
| 2008 | Practical HDR Texture CompressionabstractAbstract The use of high dynamic range (HDR) textures in real‐time graphics applications can increase realism and provide a more vivid experience. However, the increased bandwidth and storage requirements for uncompressed HDR data can become a major bottleneck. Hence, several recent algorithms for HDR texture compression have been proposed. In this paper, we discuss several practical issues one has to confront in order to develop and implement HDR texture compression schemes. These include improved texture filtering and efficient offline compression. For compression, we describe how Procrustes analysis can be used to quickly match a predefined template shape against chrominance data. To reduce the cost of HDR texture filtering, we perform filtering prior to the colour transformation, and use a simple trick to reduce the incurred errors. We also introduce a number of novel compression modes, which can be combined with existing compression schemes, or used on their own. Jacob Munkberg, Petrik Clarberg, Jon Hasselgren, Tomas Akenine-Möller |
Comput. Graph. Forum | 4 |
| 2008 | Graphics Processing Units for HandheldsabstractDuring the past few years, mobile phones and other handheld devices have gone from only handling dull text-based menu systems to, on an increasing number of models, being able to render high-quality three-dimensional graphics at high frame rates. This paper is a survey of the special considerations that must be taken when designing graphics processing units (GPUs) on such devices. Starting off by introducing desktop GPUs as a reference, the paper discusses how mobile GPUs are designed, often with power consumption rather than performance as the primary goal. Lowering the bus traffic between the GPU and the memory is an efficient way of reducing power consumption, and therefore some high-level algorithms for bandwidth reduction are presented. In addition, an overview of the different APIs that are used in the handheld market to handle both two-dimensional and three-dimensional graphics is provided. Finally, we present our outlook for the future and discuss directions of future research on handheld GPUs. Tomas Akenine-Möller, Jacob Ström |
Proc. IEEE | 1 |
| 2007 | PCU: the programmable culling unitabstractCulling techniques have always been a central part of computer graphics, but graphics hardware still lack efficient and flexible support for culling. To improve the situation, we introduce the programmable culling unit, which is as flexible as the fragment program unit and capable of quickly culling entire blocks of fragments. Furthermore, it is very easy for the developer to use the PCU as culling programs can be automatically derived from fragment programs containing a discard instruction. Our PCU can be integrated into an existing fragment program unit with a modest hardware overhead of only about 10%. Using the PCU, we have observed shader speedups between 1.4 and 2.1 for relevant scenes. Jon Hasselgren, Tomas Akenine-Möller |
ACM Trans. Graph. | 2 |
| 2006 | An Efficient Multi-View Rasterization Architecture
Jon Hasselgren, Tomas Akenine-Möller |
Rendering Techniques | 2 |
| 2006 | A dynamic bounding volume hierarchy for generalized collision detection
Thomas Larsson, Tomas Akenine-Möller |
Comput. Graph. | 2 |
| 2006 | High dynamic range texture compression for graphics hardwareabstractIn this paper, we break new ground by presenting algorithms for fixed-rate compression of high dynamic range textures at low bit rates. First, the S3TC low dynamic range texture compression scheme is extended in order to enable compression of HDR data. Second, we introduce a novel robust algorithm that offers superior image quality. Our algorithm can be efficiently implemented in hardware, and supports textures with a dynamic range of over 10 9 :1. At a fixed rate of 8 bits per pixel, we obtain results virtually indistinguishable from uncompressed HDR textures at 48 bits per pixel. Our research can have a big impact on graphics hardware and real-time rendering, since HDR texturing suddenly becomes affordable. Jacob Munkberg, Petrik Clarberg, Jon Hasselgren, Tomas Akenine-Möller |
ACM Trans. Graph. | 4 |
| 2005 | Interactive rendering of caustics using interpolated warped volumes
Manfred Ernst, Tomas Akenine-Möller, Henrik Wann Jensen |
Graphics Interface | 2 |
| 2005 | A Family of Inexpensive Sampling SchemesabstractAbstract To improve image quality in computer graphics, antialiazing techniques such as supersampling and multisampling are used. We explore a family of inexpensive sampling schemes that cost as little as 1.25 samples per pixel and up to 2.0 samples per pixel. By placing sample points in the corners or on the edges of the pixels, sharing can occur between pixels, and this makes it possible to create inexpensive sampling schemes. Using an evaluation and optimization framework, we present optimized sampling patterns costing 1.25, 1.5, 1.75 and 2.0 samples per pixel. Jon Hasselgren, Tomas Akenine-Möller, Samuli Laine |
Comput. Graph. Forum | 2 |
| 2005 | Wavelet importance sampling: efficiently evaluating products of complex functionsabstractWe present a new technique for importance sampling products of complex functions using wavelets. First, we generalize previous work on wavelet products to higher dimensional spaces and show how this product can be sampled on-the-fly without the need of evaluating the full product. This makes it possible to sample products of high-dimensional functions even if the product of the two functions in itself is too memory consuming. Then, we present a novel hierarchical sample warping algorithm that generates high-quality point distributions, which match the wavelet representation exactly. One application of the new sampling technique is rendering of objects with measured BRDFs illuminated by complex distant lighting --- our results demonstrate how the new sampling technique is more than an order of magnitude more efficient than the best previous techniques. Petrik Clarberg, Wojciech Jarosz, Tomas Akenine-Möller, Henrik Wann Jensen |
ACM Trans. Graph. | 3 |
| 2005 | Precomputed local radiance transfer for real-time lighting designabstractThis paper introduces a new method for real-time relighting of scenes illuminated by local light sources. We extend previous work on precomputed radiance transfer for distant lighting to local lighting by introducing the concept of unstructured light clouds. The unstructured light cloud enables a compact representation of local lights in the model and real-time rendering of complex models with full global illumination due to local light sources. We use simplification of lights, and clustered PCA to obtain a compressed representation. When storing only the indirect component of the illumination, we are able to get high quality with only 8-16 lighting coefficients per vertex. Our results demonstrate real-time rendering of scenes with moving lights, dynamic cameras, glossy materials and global illumination. Anders Wang Kristensen, Tomas Akenine-Möller, Henrik Wann Jensen |
ACM Trans. Graph. | 2 |
| 2005 | Soft shadow volumes for ray tracingabstractWe present a new, fast algorithm for rendering physically-based soft shadows in ray tracing-based renderers. Our method replaces the hundreds of shadow rays commonly used in stochastic ray tracers with a single shadow ray and a local reconstruction of the visibility function. Compared to tracing the shadow rays. our algorithm produces exactly the same image while executing one to two orders of magnitude faster in the test scenes used. Our first contribution is a two-stage method for quickly determining the silhouette edges that overlap an area light source, as seen from the point to be shaded. Secondly, we show that these partial silhouettes of occluders, along with a single shadow ray, are sufficient for reconstructing the visibility function between the point and the light source. Samuli Laine, Timo Aila, Ulf Assarsson, Jaakko Lehtinen, Tomas Akenine-Möller |
ACM Trans. Graph. | 5 |
| 2004 | Occlusion culling and z-fail for soft shadow volume algorithms
Ulf Assarsson, Tomas Akenine-Möller |
Vis. Comput. | 2 |
| 2003 | Graphics for the masses: a hardware rasterization architecture for mobile phonesabstractThe mobile phone is one of the most widespread devices with rendering capabilities. Those capabilities have been very limited because the resources on such devices are extremely scarce; small amounts of memory, little bandwidth, little chip area dedicated for special purposes, and limited power consumption. The small display resolutions present a further challenge; the angle subtended by a pixel is relatively large, and therefore reasonably high quality rendering is needed to generate high fidelity images.To increase the mobile rendering capabilities, we propose a new hardware architecture for rasterizing textured triangles. Our architecture focuses on saving memory bandwidth, since an external memory access typically is one of the most energy-consuming operations, and because mobile phones need to use as little power as possible. Therefore, our system includes three new key innovations: I) an inexpensive multisampling scheme that gives relatively high quality at the same cost of previous inexpensive schemes, II) a texture minification system, including texture compression, which gives quality relatively close to trilinear mipmapping at the cost of 1.33 32-bit memory accesses on average, III) a scanline-based culling scheme that avoids a significant amount of z-buffer reads, and that only requires one context. Software simulations show that these three innovations together significantly reduce the memory bandwidth, and thus also the power consumption. Tomas Akenine-Möller, Jacob Ström |
ACM Trans. Graph. | 1 |
| 2003 | A geometry-based soft shadow volume algorithm using graphics hardwareabstractMost previous soft shadow algorithms have either suffered from aliasing, been too slow, or could only use a limited set of shadow casters and/or receivers. Therefore, we present a strengthened soft shadow volume algorithm that deals with these problems. Our critical improvements include robust penumbra wedge construction, geometry-based visibility computation, and also simplified computation through a four-dimensional texture lookup. This enables us to implement the algorithm using programmable graphics hardware, and it results in images that most often are indistinguishable from images created as the average of 1024 hard shadow images. Furthermore, our algorithm can use both arbitrary shadow casters and receivers. Also, one version of our algorithm completely avoids sampling artifacts which is rare for soft shadow algorithms. As a bonus, the four-dimensional texture lookup allows for small textured light sources, and, even video textures can be used as light sources. Our algorithm has been implemented in pure software, and also using the GeForce FX emulator with pixel shaders. Our software implementation renders soft shadows at 0.5--5 frames per second for the images in this paper. With actual hardware, we expect that our algorithm will render soft shadows in real time. An important performance measure is bandwidth usage. For the same image quality, an algorithm using the accumulated hard shadow images uses almost two orders of magnitude more bandwidth than our algorithm. Ulf Assarsson, Tomas Akenine-Möller |
ACM Trans. Graph. | 2 |
| 2003 | Efficient collision detection for models deformed by morphing
Thomas Larsson, Tomas Akenine-Möller |
Vis. Comput. | 2 |
| 2001 | Occlusion horizons for driving through urban sceneryabstractWe present a rapid occlusion culling algorithm specifically designed for urban environments. For each frame, an occlusion horizon is being built on-the-fly during a hierarchical front-to-back traversal of the scene. All conservatively hidden objects are culled, while all the occluding impostors of all conservatively visible objects are added to the 2½D occlusion horizon. Our framework also supports levels-of-detail (LOD) rendering by estimating the visible area of the projection of an object in order to select the appropriate LOD for each object. This algorithm requires no substantial preprocessing and no excessive storage. In a test scene of 10,000 buildings, the cull phase took 11 ms on a PentiumII 333 MHz and 45 ms on an SGI Octane per frame on average. In typical views, the occlusion horizon culled away 80-90% of the objects that were within the view frustum, giving a 10 times speedup over view frustum culling alone. Combining the occlusion horizon with LOD rendering gave a 17 times speedup on an SGI Octane, and 23 times on a PII. Laura Downs, Tomas Akenine-Möller, Carlo H. Séquin |
SI3D | 2 |