Carsten Benthin

dblp:23/80 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0003-3337-1636ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 5 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Axis-Normalized Ray-Box Intersection
abstract
Abstract Ray‐axis aligned bounding box intersection tests play a crucial role in the runtime performance of many rendering applications, driven not by complexity but mainly by the volume of tests required. While existing solutions were believed to be pretty much optimal in terms of runtime on current hardware, our paper introduces a new intersection test requiring fewer arithmetic operations compared to all previous methods. By transforming the ray we eliminate the need for one third of the traditional bounding‐slab tests and achieve a speed enhancement of approximately 13.8% or 10.9%, depending on the compiler. We present detailed runtime analyses in various scenarios.
Fabian Friederichs, Carsten Benthin, Steve Grogorick, Elmar Eisemann, Marcus A. Magnor, Martin Eisemann
Comput. Graph. Forum2
2024 Ray Tracing Animated Displaced Micro-Meshes
abstract
Abstract We present a new method that allows efficient ray tracing of virtually artefact‐free animated displaced micro‐meshes (DMMs) [MMT23] and preserves their low memory footprint and low BVH build and update cost. DMMs allow for compact representation of micro‐triangle geometry through hierarchical encoding of displacements. Displacements are computed with respect to a coarse base mesh and are used to displace new vertices introduced during 1 : 4 subdivision of the base mesh. Applying non‐rigid transformation to the base mesh can result in silhouette and normal artefacts (see Figure 1) during animation. We propose an approach which prevents these artefacts by interpolating transformation matrices before applying them to the DMM representation. Our interpolation‐based algorithm does not change DMM data structures and it allows for efficient bounding of animated micro‐triangle geometry which is essential for fast tessellation‐free ray tracing of animated DMMs.
Holger Grün, Carsten Benthin, Andrew Kensler, Joshua Barczak, David McAllister
Comput. Graph. Forum2
2023 Real-Time Ray Tracing of Micro-Poly Geometry with Hierarchical Level of Detail
abstract
Abstract In recent work, Nanite has demonstrated how to rasterize virtualized micro‐poly geometry in real time, thus enabling immense geometric complexity. We present a system that employs similar methods for real‐time ray tracing of micro‐poly geometry. The geometry is preprocessed in almost the same fashion: Nearby triangles are clustered together and clusters get merged and simplified to obtain hierarchical level of detail (LOD). Then these clusters are compressed and stored in a GPU‐friendly data structure. At run time, Nanite selects relevant clusters, decompresses them and immediately rasterizes them. Instead of rasterization, we decompress each selected cluster into a small bounding volume hierarchy (BVH) in the format expected by the ray tracing hardware. Then we build a complete BVH on top of the bounding volumes of these clusters and use it for ray tracing. Our BVH build reaches more than 74% of the attainable peak memory bandwidth and thus it can be done per frame. Since LOD selection happens per frame at the granularity of clusters, all triangles cover a small area in screen space.
Carsten Benthin, Christoph Peters 0002
Comput. Graph. Forum1
2023 Stochastic Subsets for BVH Construction
abstract
Abstract BVH construction is a critical component of real‐time and interactive ray‐tracing systems. However, BVH construction can be both compute and bandwidth intensive, especially when a large degree of dynamic geometry is present. Different build algorithms vary substantially in the traversal performance that they produce, making high quality construction algorithms desirable. However, high quality algorithms, such as top‐down construction, are typically more expensive, limiting their benefit in real‐time and interactive contexts. One particular challenge of high quality top‐down construction algorithms is that the large working set at the top of the tree can make constructing these levels bandwidth‐intensive, due to O(nlog(n)) complexity, limited cache locality, and less dense compute at these levels. To address this limitation, we propose a novel stochastic approach to GPU BVH construction that selects a representative subset to build the upper levels of the tree. As a second pass, the remaining primitives are clustered around the BVH leaves and further processed into a complete BVH. We show that our novel approach significantly reduces the construction time of top‐down GPU BVH builders by a factor up to 1.8×, while achieving competitive rendering performance in most cases, and exceeding the performance in others.
Lorenzo Tessari, Addis Dittebrandt, Michael J. Doyle, Carsten Benthin
Comput. Graph. Forum4
2021 A Survey on Bounding Volume Hierarchies for Ray Tracing
abstract
Abstract Ray tracing is an inherent part of photorealistic image synthesis algorithms. The problem of ray tracing is to find the nearest intersection with a given ray and scene. Although this geometric operation is relatively simple, in practice, we have to evaluate billions of such operations as the scene consists of millions of primitives, and the image synthesis algorithms require a high number of samples to provide a plausible result. Thus, scene primitives are commonly arranged in spatial data structures to accelerate the search. In the last two decades, the bounding volume hierarchy (BVH) has become the de facto standard acceleration data structure for ray tracing‐based rendering algorithms in offline and recently also in real‐time applications. In this report, we review the basic principles of bounding volume hierarchies as well as advanced state of the art methods with a focus on the construction and traversal. Furthermore, we discuss industrial frameworks, specialized hardware architectures, other applications of bounding volume hierarchies, best practices, and related open problems.
Daniel Meister 0002, Shinji Ogaki, Carsten Benthin, Michael J. Doyle, Michael Guthe, Jirí Bittner
Comput. Graph. Forum3
2014 Embree: a kernel framework for efficient CPU ray tracing
abstract
We describe Embree, an open source ray tracing framework for x86 CPUs. Embree is explicitly designed to achieve high performance in professional rendering environments in which complex geometry and incoherent ray distributions are common. Embree consists of a set of low-level kernels that maximize utilization of modern CPU architectures, and an API which enables these kernels to be used in existing renderers with minimal programmer effort. In this paper, we describe the design goals and software architecture of Embree, and show that for secondary rays in particular, the performance of Embree is competitive with (and often higher than) existing state-of-the-art methods on CPUs and GPUs.
Ingo Wald, Sven Woop, Carsten Benthin, Gregory S. Johnson, Manfred Ernst
ACM Trans. Graph.3
2012 Combining Single and Packet-Ray Tracing for Arbitrary Ray Distributions on the Intel MIC Architecture
abstract
Wide-SIMD hardware is power and area efficient, but it is challenging to efficiently map ray tracing algorithms to such hardware especially when the rays are incoherent. The two most commonly used schemes are either packet tracing, or relying on a separate traversal stack for each SIMD lane. Both work great for coherent rays, but suffer when rays are incoherent: The former experiences a dramatic loss of SIMD utilization once rays diverge; the latter requires a large local storage, and generates multiple incoherent streams of memory accesses that present challenges for the memory system. In this paper, we introduce a single-ray tracing scheme for incoherent rays that uses just one traversal stack on 16-wide SIMD hardware. It uses a bounding-volume hierarchy with a branching factor of four as the acceleration structure, exploits four-wide SIMD in each box and primitive intersection test, and uses 16-wide SIMD by always performing four such node or primitive tests in parallel. We then extend this scheme to a hybrid tracing scheme that automatically adapts to varying ray coherence by starting out with a 16-wide packet scheme and switching to the new single-ray scheme as soon as rays diverge. We show that on the Intel Many Integrated Core architecture this hybrid scheme consistently, and over a wide range of scenes and ray distributions, outperforms both packet and single-ray tracing.
Carsten Benthin, Ingo Wald, Sven Woop, Manfred Ernst, William R. Mark
IEEE Trans. Vis. Comput. Graph.1
2003 Interactive Ray Tracing on Commodity PC Clusters
Ingo Wald, Carsten Benthin, Andreas Dietrich 0001, Philipp Slusallek
Euro-Par2
2003 A Scalable Approach to Interactive Global Illumination
abstract
Abstract The addition of global illumination can dramatically increase the realism achievable when rendering virtual environments.In particular with interactive applications we expect the environment to reflect changes in the scenedue to global lighting effects instead of it being just a static backdrop. However, a sufficiently fast and accuratecomputation of global illumination at interactive rates has been difficult even with recent approaches based onrealtime ray tracing. In this paper we present a highly scalable approach to interactive global illumination. It fully recomputes a high‐qualitysolution for each frame and thus offers immediate feedback even for dynamic scenes, achieving more than20 fps for simple scenes. Compared to previous systems we increased the raw performance by a factor of up toeight and removed the bottlenecks that were limiting scalability. The system now scales linearly in quality andavailable computing resources, tested with up to 48 CPUs in a commodity PC‐cluster. Due to its logarithmicscaling property with respect to scene complexity it even supports lighting simulation in complex scenes with morethan 50 million triangles. This scalability allows applications to perform flexible performance trade‐offs. We alsoargue that the realism achievable through interactive global illumination will make it a standard feature of future3D graphics systems once the required computing resources are readily available.
Carsten Benthin, Ingo Wald, Philipp Slusallek
Comput. Graph. Forum1
2001 Interactive Rendering with Coherent Ray Tracing
abstract
For almost two decades researchers have argued that ray tracing will eventually become faster than the rasterization technique that completely dominates todays graphics hardware. However, this has not happened yet. Ray tracing is still exclusively being used for off-line rendering of photorealistic images and it is commonly believed that ray tracing is simply too costly to ever challenge rasterization-based algorithms for interactive use. However, there is hardly any scientific analysis that supports either point of view. In particular there is no evidence of where the crossover point might be, at which ray tracing would eventually become faster, or if such a point does exist at all. This paper provides several contributions to this discussion: We first present a highly optimized implementation of a ray tracer that improves performance by more than an order of magnitude compared to currently available ray tracers. The new algorithm make better use of computational resources such as caches and SIMD instructions and better exploits image and object space coherence. Secondly, we show that this software implementation can challenge and even outperform high-end graphics hardware in interactive rendering performance for complex environments. We also provide an brief overview of the benefits of ray tracing over rasterization algorithms and point out the potential of interactive ray tracing both in hardware and software.
Ingo Wald, Philipp Slusallek, Carsten Benthin, Markus Wagner 0004
Comput. Graph. Forum3