Ingo Wald

dblp:73/3396 · DBLP profile ↗
← Back
40ranked-venue papers
15as first author
11since 2021 · last 2026
0000-0003-0046-713XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 38 · 14 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 3Systems, architecture and hardware · 2 · 1 first-author
YearPublicationVenuePosition
2026 Materializing Inter-Channel Relationships With Multi-Density Woodcock Tracking
abstract
Volume rendering techniques for scientific visualization have recently shifted toward Monte Carlo (MC) methods for their flexibility and robustness, but their use in multi-channel visualization remains underexplored. Traditional multi-channel volume rendering often relies on arbitrary, non-physically based color blending functions that hinder interpretation. We introduce multi-density Woodcock tracking, a simple extension of Woodcock tracking that leverages an MC method to produce high-fidelity, physically grounded multi-channel renderings without arbitrary blending. By generalizing Woodcock's distance tracking, we provide a unified blending modality that also integrates blending functions from prior works. We further implement effects that enhance boundary and feature recognition. By accumulating frames in real-time, our approach delivers high-quality visualizations with perceptual benefits, demonstrated on diverse datasets.
Alper Sahistan, Stefan Zellmann, Haichao Miao, Nathan Morrical, Ingo Wald, Valerio Pascucci
IEEE Trans. Vis. Comput. Graph.5
2025 Visualization of Large Non-Trivially Partitioned Unstructured Data With Native Distribution on High-Performance Computing Systems
abstract
Interactively visualizing large finite element simulation data on High-Performance Computing (HPC) systems poses several difficulties. Some of these relate to unstructured data, which, even on a single node, is much more expensive to render compared to structured volume data. Worse yet, in the data parallel rendering context, such data with highly non-convex spatial domain boundaries will cause rays along its silhouette to enter and leave a given rank's domains at different distances. This straddling, in turn, poses challenges for both ray marching, which usually assumes successive elements to share a face, and compositing, which usually assumes a single fragment per pixel per rank. We holistically address these issues using a combination of three inter-operating techniques: first, we use a highly optimized GPU ray marching technique that, given an entry point, can march a ray to its exit point with high-performance by exploiting an exclusive-or (XOR) based compaction scheme. Second, we use hardware-accelerated ray tracing to efficiently find the proper entry points for these marching operations. Third, we use a "deep" compositing scheme to properly handle cases where different ranks' ray segments interleave in depth. We use GPU-to-GPU remote direct memory access (RDMA) to achieve interactive frame rates of 10-15 frames per second and higher for our motivating use case, the Fun3D NASA Mars Lander.
Alper Sahistan, Serkan Demirci, Ingo Wald, Stefan Zellmann, João Barbosa, Nathan Morrical, Ugur Güdükbay
IEEE Trans. Vis. Comput. Graph.3
2024 Beyond ExaBricks: GPU Volume Path Tracing of AMR Data
abstract
Abstract Adaptive Mesh Refinement (AMR) is becoming a prevalent data representation for HPC, and thus also for scientific visualization. AMR data is usually cell centric (which imposes numerous challenges), complex, and generally hard to render. Recent work on GPU‐accelerated AMR rendering has made much progress towards real‐time volume and isosurface rendering of such data, but so far this work has focused exclusively on ray marching, with simple lighting models and without scattering events or global illumination. True high‐quality rendering requires a modified approach that is able to trace arbitrary incoherent paths; but this may not be a perfect fit for the types of data structures recently developed for ray marching. In this paper, we describe a novel approach to high‐quality path tracing of complex AMR data, with a specific focus on analyzing and comparing different data structures and algorithms to achieve this goal.
Stefan Zellmann, Qi Wu 0015, Alper Sahistan, Kwan-Liu Ma, Ingo Wald
Comput. Graph. Forum5
2023 State-of-the-art in Large-Scale Volume Visualization Beyond Structured Data
abstract
Abstract Volume data these days is usually massive in terms of its topology, multiple fields, or temporal component. With the gap between compute and memory performance widening, the memory subsystem becomes the primary bottleneck for scientific volume visualization. Simple, structured, regular representations are often infeasible because the buses and interconnects involved need to accommodate the data required for interactive rendering. In this state‐of‐the‐art report, we review works focusing on large‐scale volume rendering beyond those typical structured and regular grid representations. We focus primarily on hierarchical and adaptive mesh refinement representations, unstructured meshes, and compressed representations that gained recent popularity. We review works that approach this kind of data using strategies such as out‐of‐core rendering, massive parallelism, and other strategies to cope with the sheer size of the ever‐increasing volume of data produced by today's supercomputers and acquisition devices. We emphasize the data management side of large‐scale volume rendering systems and also include a review of tools that support the various volume data types discussed.
Jonathan Sarton, Stefan Zellmann, Serkan Demirci, Ugur Güdükbay, Welcome Alexandre-Barff, Laurent Lucas, Jean-Michel Dischler, Stefan Wesner, Ingo Wald
Comput. Graph. Forum9
2023 Data Parallel Multi-GPU Path Tracing using Ray Queue Cycling
abstract
Abstract We propose a novel approach to data‐parallel path tracing on single‐node/multi‐GPU hardware that builds on ray forwarding, but which aims—above all else—at generality and practicability. We do this by avoiding any attempts at reducing the number of traces or forward operations performed, and instead focus on always using all GPUs' aggregate compute and bandwidth to effectively trace each ray on every GPU. We show that—counter‐intuitively—this is both feasible and desirable; and that when run on typical data‐center/cloud hardware, the resulting framework not only achieves good performance and scalability, but also comes with significantly fewer limitations, assumptions, or preprocessing requirements than existing techniques.
Ingo Wald, Milan Jaros, Stefan Zellmann
Comput. Graph. Forum1
2023 Memory-Efficient GPU Volume Path Tracing of AMR Data Using the Dual Mesh
abstract
Abstract A common way to render cell‐centric adaptive mesh refinement (AMR) data is to compute the dual mesh and visualize that with a standard unstructured element renderer. While the dual mesh provides a high‐quality interpolator, the memory requirements of the dual meshdata structureare significantly higher than those of the original grid, which prevents rendering very large data sets. We introduce a GPU‐friendly data structure and a clustering algorithm that allow for efficient AMR dual mesh rendering with a competitive memory footprint. Fundamentally, any off‐the‐shelf unstructured element renderer running on GPUs could be extended to support our data structure just by adding agridletelement type in addition to the standard tetrahedra, pyramids, wedges, and hexahedra supported by default. We integrated the data structure into a volumetric path tracer to compare it to various state‐of‐the‐art unstructured element sampling methods. We show that our data structure easily competes with these methods in terms of rendering performance, but is much more memory‐efficient.
Stefan Zellmann, Qi Wu 0015, Kwan-Liu Ma, Ingo Wald
Comput. Graph. Forum4
2023 Quick Clusters: A GPU-Parallel Partitioning for Efficient Path Tracing of Unstructured Volumetric Grids
abstract
We propose a simple yet effective method for clustering finite elements to improve preprocessing times and rendering performance of unstructured volumetric grids without requiring auxiliary connectivity data. Rather than building bounding volume hierarchies (BVHs) over individual elements, we sort elements along with a Hilbert curve and aggregate neighboring elements together, improving BVH memory consumption by over an order of magnitude. Then to further reduce memory consumption, we cluster the mesh on the fly into sub-meshes with smaller indices using a series of efficient parallel mesh re-indexing operations. These clusters are then passed to a highly optimized ray tracing API for point containment queries and ray-cluster intersection testing. Each cluster is assigned a maximum extinction value for adaptive sampling, which we rasterize into non-overlapping view-aligned bins allocated along the ray. These maximum extinction bins are then used to guide the placement of samples along the ray during visualization, reducing the number of samples required by multiple orders of magnitude (depending on the dataset), thereby improving overall visualization interactivity. Using our approach, we improve rendering performance over a competitive baseline on the NASA Mars Lander dataset from 6× (1 frame per second (fps) and 1.0 M rays per second (rps) up to now 6 fps and 12.4 M rps, now including volumetric shadows) while simultaneously reducing memory consumption by 3×(33 GB down to 11 GB) and avoiding any offline preprocessing steps, enabling high-quality interactive visualization on consumer graphics cards. Then by utilizing the full 48 GB of an RTX 8000, we improve the performance of Lander by 17 × (1 fps up to 17 fps, 1.0 M rps up to 35.6 M rps).
Nathan Morrical, Alper Sahistan, Ugur Güdükbay, Ingo Wald, Valerio Pascucci
IEEE Trans. Vis. Comput. Graph.4
2022 Accelerating Unstructured Mesh Point Location With RT Cores
abstract
We present a technique that leverages ray tracing hardware available in recent Nvidia RTX GPUs to solve a problem other than classical ray tracing. Specifically, we demonstrate how to use these units to accelerate the point location of general unstructured elements consisting of both planar and bilinear faces. This unstructured mesh point location problem has previously been challenging to accelerate on GPU architectures; yet, the performance of these queries is crucial to many unstructured volume rendering and compute applications. Starting with a CUDA reference method, we describe and evaluate three approaches that reformulate these point queries to incrementally map algorithmic complexity to these new hardware ray tracing units. Each variant replaces the simpler problem of point queries with a more complex one of ray queries. Initial variants exploit ray tracing cores for accelerated BVH traversal, and subsequent variants use ray-triangle intersections and per-face metadata to detect point-in-element intersections. Although these later variants are more algorithmically complex, they are significantly faster than the reference method thanks to hardware acceleration. Using our approach, we improve the performance of an unstructured volume renderer by up to 4× for tetrahedral meshes and up to 15× for general bilinear element meshes, matching, or out-performing state-of-the-art solutions while simultaneously improving on robustness and ease-of-implementation.
Nathan Morrical, Ingo Wald, Will Usher 0001, Valerio Pascucci
IEEE Trans. Vis. Comput. Graph.2
2022 A Memory Efficient Encoding for Ray Tracing Large Unstructured Data
abstract
In theory, efficient and high-quality rendering of unstructured data should greatly benefit from modern GPUs, but in practice, GPUs are often limited by the large amount of memory that large meshes require for element representation and for sample reconstruction acceleration structures. We describe a memory-optimized encoding for large unstructured meshes that efficiently encodes both the unstructured mesh and corresponding sample reconstruction acceleration structure, while still allowing for fast random-access sampling as required for rendering. We demonstrate that for large data our encoding allows for rendering even the 2.9 billion element Mars Lander on a single off-the-shelf GPU-and the largest 6.3 billion version on a pair of such GPUs.
Ingo Wald, Nathan Morrical, Stefan Zellmann
IEEE Trans. Vis. Comput. Graph.1
2021 Multi-level tetrahedralization-based accelerator for ray-tracing animated scenes
abstract
Abstract We describe a hybrid acceleration structure for ray tracing. The hybrid structure is a Bounding Volume Hierarchy (BVH) where the leaf nodes are tetrahedralized for a decent ray‐surface intersection performance. We use the hybrid acceleration structure (BTH) in a two‐level acceleration structure for rendering animated scenes. There is a BVH at the top level in this two‐level structure and the proposed hybrid structure (BTH) at the bottom level. We test the proposed two‐level structure (BVH‐BTH) for various animated scenes and obtained promising results against other acceleration structures in terms of rendering times. The two‐level BVH‐BTH structure outperforms the two‐level BVH structure for the tested dynamic scenes.
Aytek Aman, Serkan Demirci, Ugur Güdükbay, Ingo Wald
Comput. Animat. Virtual Worlds4
2021 Ray Tracing Structured AMR Data Using ExaBricks
abstract
Structured Adaptive Mesh Refinement (Structured AMR) enables simulations to adapt the domain resolution to save computation and storage, and has become one of the dominant data representations used by scientific simulations; however, efficiently rendering such data remains a challenge. We present an efficient approach for volume- and iso-surface ray tracing of Structured AMR data on GPU-equipped workstations, using a combination of two different data structures. Together, these data structures allow a ray tracing based renderer to quickly determine which segments along the ray need to be integrated and at what frequency, while also providing quick access to all data values required for a smooth sample reconstruction kernel. Our method makes use of the RTX ray tracing hardware for surface rendering, ray marching, space skipping, and adaptive sampling; and allows for interactive changes to the transfer function and implicit iso-surfacing thresholds. We demonstrate that our method achieves high performance with little memory overhead, enabling interactive high quality rendering of complex AMR data sets on individual GPU workstations.
Ingo Wald, Stefan Zellmann, Will Usher 0001, Nathan Morrical, Ulrich Lang 0002, Valerio Pascucci
IEEE Trans. Vis. Comput. Graph.1
2019 Ray Tracing Generalized Tube Primitives: Method and Applications
abstract
We present a general high-performance technique for ray tracing generalized tube primitives. Our technique efficiently supports tube primitives with fixed and varying radii, general acyclic graph structures with bifurcations, and correct transparency with interior surface removal. Such tube primitives are widely used in scientific visualization to represent diffusion tensor imaging tractographies, neuron morphologies, and scalar or vector fields of 3D flow. We implement our approach within the OSPRay ray tracing framework, and evaluate it on a range of interactive visualization use cases of fixed- and varying-radius streamlines, pathlines, complex neuron morphologies, and brain tractographies. Our proposed approach provides interactive, high-quality rendering, with low memory overhead.
Mengjiao Han, Ingo Wald, Will Usher 0001, Qi Wu 0015, Feng Wang 0013, Valerio Pascucci, Charles D. Hansen, Chris R. Johnson 0001
Comput. Graph. Forum2
2019 Scalable Ray Tracing Using the Distributed FrameBuffer
abstract
Abstract Image‐ and data‐parallel rendering across multiple nodes on high‐performance computing systems is widely used in visualization to provide higher frame rates, support large data sets, and render data in situ. Specifically for in situ visualization, reducing bottlenecks incurred by the visualization and compositing is of key concern to reduce the overall simulation runtime. Moreover, prior algorithms have been designed to support either image‐ or data‐parallel rendering and impose restrictions on the data distribution, requiring different implementations for each configuration. In this paper, we introduce the Distributed FrameBuffer, an asynchronous image‐processing framework for multi‐node rendering. We demonstrate that our approach achieves performance superior to the state of the art for common use cases, while providing the flexibility to support a wide range of parallel rendering algorithms and data distributions. By building on this framework, we extend the open‐source ray tracing library OSPRay with a data‐distributed API, enabling its use in data‐distributed and in situ visualization applications.
Will Usher 0001, Ingo Wald, Jefferson Amstutz, Johannes Günther 0001, Carson Brownlee, Valerio Pascucci
Comput. Graph. Forum2
2019 CPU Isosurface Ray Tracing of Adaptive Mesh Refinement Data
abstract
Adaptive mesh refinement (AMR) is a key technology for large-scale simulations that allows for adaptively changing the simulation mesh resolution, resulting in significant computational and storage savings. However, visualizing such AMR data poses a significant challenge due to the difficulties introduced by the hierarchical representation when reconstructing continuous field values. In this paper, we detail a comprehensive solution for interactive isosurface rendering of block-structured AMR data. We contribute a novel reconstruction strategy-the octant method-which is continuous, adaptive and simple to implement. Furthermore, we present a generally applicable hybrid implicit isosurface ray-tracing method, which provides better rendering quality and performance than the built-in sampling-based approach in OSPRay. Finally, we integrate our octant method and hybrid isosurface geometry into OSPRay as a module, providing the ability to create high-quality interactive visualizations combining volume and isosurface representations of BS-AMR data. We evaluate the rendering performance, memory consumption and quality of our method on two gigascale block-structured AMR datasets.
Feng Wang 0013, Ingo Wald, Qi Wu 0015, Will Usher 0001, Chris R. Johnson 0001
IEEE Trans. Vis. Comput. Graph.2
2017 OSPRay - A CPU Ray Tracing Framework for Scientific Visualization
abstract
Scientific data is continually increasing in complexity, variety and size, making efficient visualization and specifically rendering an ongoing challenge. Traditional rasterization-based visualization approaches encounter performance and quality limitations, particularly in HPC environments without dedicated rendering hardware. In this paper, we present OSPRay, a turn-key CPU ray tracing framework oriented towards production-use scientific visualization which can utilize varying SIMD widths and multiple device backends found across diverse HPC resources. This framework provides a high-quality, efficient CPU-based solution for typical visualization workloads, which has already been integrated into several prevalent visualization packages. We show that this system delivers the performance, high-level API simplicity, and modular device support needed to provide a compelling new rendering framework for implementing efficient scientific visualization workflows.
Ingo Wald, Gregory P. Johnson, Jefferson Amstutz, Carson Brownlee, Aaron Knoll, Jim Jeffers, Johannes Günther 0001, Paul A. Navrátil
IEEE Trans. Vis. Comput. Graph.1
2014 RBF Volume Ray Casting on Multicore and Manycore CPUs
abstract
Abstract Modern supercomputers enable increasingly large N‐body simulations using unstructured point data. The structures implied by these points can be reconstructed implicitly. Direct volume rendering of radial basis function (RBF) kernels in domain‐space offers flexible classification and robust feature reconstruction, but achieving performant RBF volume rendering remains a challenge for existing methods on both CPUs and accelerators. In this paper, we present a fast CPU method for direct volume rendering of particle data with RBF kernels. We propose a novel two‐pass algorithm: first sampling the RBF field using coherent bounding hierarchy traversal, then subsequently integrating samples along ray segments. Our approach performs interactively for a range of data sets from molecular dynamics and astrophysics up to 82 million particles. It does not rely on level of detail or subsampling, and offers better reconstruction quality than structured volume rendering of the same data, exhibiting comparable performance and requiring no additional preprocessing or memory footprint other than the BVH. Lastly, our technique enables multi‐field, multi‐material classification of particle data, providing better insight and analysis.
Aaron Knoll, Ingo Wald, Paul A. Navrátil, Anne Bowen, Khairi Reda, Michael E. Papka, Kelly P. Gaither
Comput. Graph. Forum2
2014 Embree: a kernel framework for efficient CPU ray tracing
abstract
We describe Embree, an open source ray tracing framework for x86 CPUs. Embree is explicitly designed to achieve high performance in professional rendering environments in which complex geometry and incoherent ray distributions are common. Embree consists of a set of low-level kernels that maximize utilization of modern CPU architectures, and an API which enables these kernels to be used in existing renderers with minimal programmer effort. In this paper, we describe the design goals and software architecture of Embree, and show that for secondary rays in particular, the performance of Embree is competitive with (and often higher than) existing state-of-the-art methods on CPUs and GPUs.
Ingo Wald, Sven Woop, Carsten Benthin, Gregory S. Johnson, Manfred Ernst
ACM Trans. Graph.1
2012 Extending a C-like language for portable SIMD programming
abstract
SIMD instructions are common in CPUs for years now. Using these instructions effectively requires not only vectorization of code, but also modifications to the data layout. However, automatic vectorization techniques are often not powerful enough and suffer from restricted scope of applicability; hence, programmers often vectorize their programs manually by using intrinsics: compiler-known functions that directly expand to machine instructions. They significantly decrease programmer productivity by enforcing a very error-prone and hard-to-read assembly-like programming style. Furthermore, intrinsics are not portable because they are tied to a specific instruction set.
Roland Leißa, Sebastian Hack, Ingo Wald
PPoPP3
2012 Combining Single and Packet-Ray Tracing for Arbitrary Ray Distributions on the Intel MIC Architecture
abstract
Wide-SIMD hardware is power and area efficient, but it is challenging to efficiently map ray tracing algorithms to such hardware especially when the rays are incoherent. The two most commonly used schemes are either packet tracing, or relying on a separate traversal stack for each SIMD lane. Both work great for coherent rays, but suffer when rays are incoherent: The former experiences a dramatic loss of SIMD utilization once rays diverge; the latter requires a large local storage, and generates multiple incoherent streams of memory accesses that present challenges for the memory system. In this paper, we introduce a single-ray tracing scheme for incoherent rays that uses just one traversal stack on 16-wide SIMD hardware. It uses a bounding-volume hierarchy with a branching factor of four as the acceleration structure, exploits four-wide SIMD in each box and primitive intersection test, and uses 16-wide SIMD by always performing four such node or primitive tests in parallel. We then extend this scheme to a hybrid tracing scheme that automatically adapts to varying ray coherence by starting out with a 16-wide packet scheme and switching to the new single-ray scheme as soon as rays diverge. We show that on the Intel Many Integrated Core architecture this hybrid scheme consistently, and over a wide range of scenes and ray distributions, outperforms both packet and single-ray tracing.
Carsten Benthin, Ingo Wald, Sven Woop, Manfred Ernst, William R. Mark
IEEE Trans. Vis. Comput. Graph.2
2012 Fast Construction of SAH BVHs on the Intel Many Integrated Core (MIC) Architecture
abstract
We investigate how to efficiently build bounding volume hierarchies (BVHs) with surface area heuristic (SAH) on the Intel Many Integrated Core (MIC) Architecture. To achieve maximum performance, we use four key concepts: progressive 10-bit quantization to reduce cache footprint with negligible loss in BVH quality; an AoSoA data layout that allows efficient streaming and SIMD processing; high-performance SIMD kernels for binning and partitioning; and a parallelization framework with several build-specific optimizations. The resulting system is more than an order of magnitude faster than today's high-end GPU builders for comparable BVHs; it is usually faster even than spatial median builders; it can build SAH BVHs almost as fast as existing GPUs and CPUs- and CPU-based approaches can build regular grids; and in aggregate "build+render" performance is significantly faster than the best published numbers for either of these systems, be it CPU or GPU, BVH, kd-tree, or grid.
Ingo Wald
IEEE Trans. Vis. Comput. Graph.1
2011 Full-resolution interactive CPU volume rendering with coherent BVH traversal
abstract
We present an efficient method for volume rendering by ray casting on the CPU. We employ coherent packet traversal of an implicit bounding volume hierarchy, heuristically pruned using preintegrated transfer functions, to exploit empty or homogeneous space. We also detail SIMD optimizations for volumetric integration, trilinear interpolation, and gradient lighting. The resulting system performs well on low-end and laptop hardware, and can outperform out-of-core GPU methods by orders of magnitude when rendering large volumes without level-of-detail (LOD) on a workstation. We show that, while slower than GPU methods for low-resolution volumes, an optimized CPU renderer does not require LOD to achieve interactive performance on large data sets.
Aaron Knoll, Sebastian Thelen, Ingo Wald, Charles D. Hansen, Hans Hagen, Michael E. Papka
PacificVis3
2009 State of the Art in Ray Tracing Animated Scenes
abstract
Abstract Ray tracing has long been a method of choice for off‐line rendering, but traditionally was too slow for interactive use. With faster hardware and algorithmic improvements this has recently changed, and real‐time ray tracing is finally within reach. However, real‐time capability also opens up new problems that do not exist in an off‐line environment. In particular real‐time ray tracing offers the opportunity to interactively ray trace moving/animated scene content. This presents a challenge to the data structures that have been developed for ray tracing over the past few decades. Spatial data structures crucial for fast ray tracing must be rebuilt or updated as the scene changes, and this can become a bottleneck for the speed of ray tracing. This bottleneck has recently received much attention by researchers and that has resulted in a multitude of different algorithms, data structures and strategies for handling animated scenes. The effectiveness of techniques for ray tracing dynamic scenes vary dramatically depending on details such as scene complexity, model structure, type of motion and the coherency of the rays. Consequently, there is so far no approach that is best in all cases, and determining the best technique for a particular problem can be a challenge. In this State of the Art Report (STAR), we aim to survey the different approaches to ray tracing animated scenes, discussing their strengths and weaknesses, and their relationship to other approaches. The overall goal is to help the reader choose the best approach depending on the situation, and to expose promising areas where there is potential for algorithmic improvements.
Ingo Wald, William R. Mark, Johannes Günther 0001, Solomon Boulos, Thiago Ize, Warren A. Hunt, Steven G. Parker, Peter Shirley
Comput. Graph. Forum1
2009 Coherent multiresolution isosurface ray tracing
Aaron Knoll, Ingo Wald, Charles D. Hansen
Vis. Comput.2
2008 Fast, parallel, and asynchronous construction of BVHs for ray tracing animated scenes
Ingo Wald, Thiago Ize, Steven G. Parker
Comput. Graph.1
2008 Sequential Monte Carlo Adaptation in Low-Anisotropy Participating Media
abstract
Abstract This paper presents a novel method that effectively combines both control variates and importance sampling in a sequential Monte Carlo context. The radiance estimates computed during the rendering process are cached in a 5D adaptive hierarchical structure that defines dynamic predicate functions for both variance reduction techniques and guarantees well‐behaved PDFs, yielding continually increasing efficiencies thanks to a marginal computational overhead. While remaining unbiased, the technique is effective within a single pass as both estimation and caching are done online, exploiting the coherency in illumination while being independent of the actual scene representation. The method is relatively easy to implement and to tune via a single parameter, and we demonstrate its practical benefits with important gains in convergence rate and competitive results with state of the art techniques.
Vincent Pegoraro, Ingo Wald, Steven G. Parker
Comput. Graph. Forum2
2007 Interactive Iso-Surface Ray Tracing of Massive Volumetric Data Sets
Heiko Friedrich, Ingo Wald, Johannes Günther 0001, Gerd Marmitt, Philipp Slusallek
EGPGV2
2007 Asynchronous BVH Construction for Ray Tracing Dynamic Scenes on Parallel Multi-Core Architectures
Thiago Ize, Ingo Wald, Steven G. Parker
EGPGV2
2007 Packet-based whitted and distribution ray tracing
abstract
Much progress has been made toward interactive ray tracing, but most research has focused specifically on ray casting. A common approach is to use "packets" of rays to amortize cost across sets of rays. Whether "packets" can be used to speed up the cost of reflection and refraction rays is unclear. The issue is complicated since such rays do not share common origins and often have less directional coherence than viewing and shadow rays. Since the primary advantage of ray tracing over rasterization is the computation of global effects, such as accurate reflection and refraction, this lack of knowledge should be corrected. We are also interested in exploring whether distribution ray tracing, due to its stochastic properties, further erodes the effectiveness of techniques used to accelerate ray casting. This paper addresses the question of whether packet-based ray tracing algorithms can be effectively used for more than visibility computation. We show that by choosing an appropriate data structure and a suitable packet assembly algorithm we can extend the idea of "packets" from ray casting to Whitted-style and distribution ray tracing, while maintaining efficiency.
Solomon Boulos, David Edwards, J. Dylan Lacewell, Joe Michael Kniss, Jan Kautz, Peter Shirley, Ingo Wald
Graphics Interface7
2007 Ray tracing deformable scenes using dynamic bounding volume hierarchies
Ingo Wald, Solomon Boulos, Peter Shirley
ACM Trans. Graph.1
2007 A Coherent Grid Traversal Approach to Visualizing Particle-Based Simulation Data
abstract
We present an approach to visualizing particle-based simulation data using interactive ray tracing and describe an algorithmic enhancement that exploits the properties of these data sets to provide highly interactive performance and reduced storage requirements. This algorithm for fast packet-based ray tracing of multilevel grids enables the interactive visualization of large time-varying data sets with millions of particles and incorporates advanced features like soft shadows. We compare the performance of our approach with two recent particle visualization systems: one based on an optimized single ray grid traversal algorithm and the other on programmable graphics hardware. This comparison demonstrates that the new algorithm offers an attractive alternative for interactive particle visualization.
Christiaan P. Gribble, Thiago Ize, Andrew Kensler, Ingo Wald, Steven G. Parker
IEEE Trans. Vis. Comput. Graph.4
2007 Interactive Isosurface Ray Tracing of Time-Varying Tetrahedral Volumes
abstract
We describe a system for interactively rendering isosurfaces of tetrahedral finite-element scalar fields using coherent ray tracing techniques on the CPU. By employing state-of-the art methods in polygonal ray tracing, namely aggressive packet/frustum traversal of a bounding volume hierarchy, we can accomodate large and time-varying unstructured data. In conjunction with this efficiency structure, we introduce a novel technique for intersecting ray packets with tetrahedral primitives. Ray tracing is flexible, allowing for dynamic changes in isovalue and time step, visualization of multiple isosurfaces, shadows, and depth-peeling transparency effects. The resulting system offers the intuitive simplicity of isosurfacing, guaranteed-correct visual results, and ultimately a scalable, dynamic and consistently interactive solution for visualizing unstructured volumes.
Ingo Wald, Heiko Friedrich, Aaron Knoll, Charles D. Hansen
IEEE Trans. Vis. Comput. Graph.1
2006 Ray Tracing Animated Scenes using Motion Decomposition
abstract
Abstract Though ray tracing has recently become interactive, its high precomputation time for building spatial indices usually limits its applications to walkthroughs of static scenes. This is a major limitation, as most applications demand support for dynamically animated models. In this paper, we present a new approach to ray trace a special but important class of dynamic scenes, namely models whose connectivity does not change over time and for which all possible poses are known in advance. We support these kinds of models by introducing two new concepts: motion decomposition, and fuzzy kd‐trees. We analyze the animation and break the model down into submeshes with similar motion. For each of these submeshes and for every time step, we calculate a best affine transformation through a least square approach. Any residual motion is then captured in a single "fuzzy kd‐tree" for the entire animation. Together, these techniques allow for ray tracing animations without rebuilding the spatial index structures for the submeshes, resulting in interactive frame rates of 5 to 15 fps even on a single CPU. Categories and Subject Descriptors (according to ACM CCS): I.3.7 [Computer Graphics]: Ray tracing I.3.6 [Methodology and Techniques]: Graphics data structures and data types
Johannes Günther 0001, Heiko Friedrich, Ingo Wald, Hans-Peter Seidel, Philipp Slusallek
Comput. Graph. Forum3
2006 Ray tracing animated scenes using coherent grid traversal
abstract
We present a new approach to interactive ray tracing of moderate-sized animated scenes based on traversing frustum-bounded packets of coherent rays through uniform grids. By incrementally computing the overlap of the frustum with a slice of grid cells, we accelerate grid traversal by more than a factor of 10, and achieve ray tracing performance competitive with the fastest known packet-based kd-tree ray tracers. The ability to efficiently rebuild the grid on every frame enables this performance even for fully dynamic scenes that typically challenge interactive ray tracing systems.
Ingo Wald, Thiago Ize, Andrew Kensler, Aaron Knoll, Steven G. Parker
ACM Trans. Graph.1
2005 Faster Isosurface Ray Tracing Using Implicit KD-Trees
abstract
The visualization of high-quality isosurfaces at interactive rates is an important tool in many simulation and visualization applications. Today, isosurfaces are most often visualized by extracting a polygonal approximation that is then rendered via graphics hardware or by using a special variant of preintegrated volume rendering. However, these approaches have a number of limitations in terms of the quality of the isosurface, lack of performance for complex data sets, or supported shading models. An alternative isosurface rendering method that does not suffer from these limitations is to directly ray trace the isosurface. However, this approach has been much too slow for interactive applications unless massively parallel shared-memory supercomputers have been used. In this paper, we implement interactive isosurface ray tracing on commodity desktop PCs by building on recent advances in real-time ray tracing of polygonal scenes and using those to improve isosurface ray tracing performance as well. The high performance and scalability of our approach will be demonstrated with several practical examples, including the visualization of highly complex isosurface data sets, the interactive rendering of hybrid polygonal/isosurface scenes, including high-quality ray traced shading effects, and even interactive global illumination on isosurfaces.
Ingo Wald, Heiko Friedrich, Gerd Marmitt, Philipp Slusallek, Hans-Peter Seidel
IEEE Trans. Vis. Comput. Graph.1
2004 VRML Scene Graphs on an Interactive Ray Tracing Engine
Andreas Dietrich 0001, Ingo Wald, Markus Wagner 0004, Philipp Slusallek
VR2
2004 Colorplate: VRML Scene Graphs on an Interactive Ray Tracing Engine
Andreas Dietrich 0001, Ingo Wald, Markus Wagner 0004, Philipp Slusallek
VR2
2004 Balancing Considered Harmful - Faster Photon Mapping using the Voxel Volume Heuristic
abstract
Abstract Photon mapping is one of the most important algorithms for computing global illumination. Especially for efficiently producing convincing caustics, there are no real alternatives to photon mapping. On the other hand, photon mapping is also quite costly: Each radiance lookup requires to find the k nearest neighbors in a kd‐tree, which can be more costly than shooting several rays. Therefore, the nearest‐neighbor queries often dominate the rendering time of a photon map based renderer. In this paper, we present a method that reorganizes — i.e. un balances — the kd‐tree for storing the photons in a way that allows for finding the k‐nearest neighbors much more efficiently, thereby accelerating the radiance estimates by a factor of 1.2–3.4. Most importantly, our method still finds exactly the same k‐nearest‐neighbors as the original method, without introducing any approximations or loss of accuracy. The impact of our method is demonstrated with several practical examples. Categories and Subject Descriptors (according to ACM CCS): I.3.3 [Computer Graphics]: Global Illumination I.3.7 [Computer Graphics]: Raytracing
Ingo Wald, Johannes Günther 0001, Philipp Slusallek
Comput. Graph. Forum1
2003 Interactive Ray Tracing on Commodity PC Clusters
Ingo Wald, Carsten Benthin, Andreas Dietrich 0001, Philipp Slusallek
Euro-Par1
2003 A Scalable Approach to Interactive Global Illumination
abstract
Abstract The addition of global illumination can dramatically increase the realism achievable when rendering virtual environments.In particular with interactive applications we expect the environment to reflect changes in the scenedue to global lighting effects instead of it being just a static backdrop. However, a sufficiently fast and accuratecomputation of global illumination at interactive rates has been difficult even with recent approaches based onrealtime ray tracing. In this paper we present a highly scalable approach to interactive global illumination. It fully recomputes a high‐qualitysolution for each frame and thus offers immediate feedback even for dynamic scenes, achieving more than20 fps for simple scenes. Compared to previous systems we increased the raw performance by a factor of up toeight and removed the bottlenecks that were limiting scalability. The system now scales linearly in quality andavailable computing resources, tested with up to 48 CPUs in a commodity PC‐cluster. Due to its logarithmicscaling property with respect to scene complexity it even supports lighting simulation in complex scenes with morethan 50 million triangles. This scalability allows applications to perform flexible performance trade‐offs. We alsoargue that the realism achievable through interactive global illumination will make it a standard feature of future3D graphics systems once the required computing resources are readily available.
Carsten Benthin, Ingo Wald, Philipp Slusallek
Comput. Graph. Forum2
2001 Interactive Rendering with Coherent Ray Tracing
abstract
For almost two decades researchers have argued that ray tracing will eventually become faster than the rasterization technique that completely dominates todays graphics hardware. However, this has not happened yet. Ray tracing is still exclusively being used for off-line rendering of photorealistic images and it is commonly believed that ray tracing is simply too costly to ever challenge rasterization-based algorithms for interactive use. However, there is hardly any scientific analysis that supports either point of view. In particular there is no evidence of where the crossover point might be, at which ray tracing would eventually become faster, or if such a point does exist at all. This paper provides several contributions to this discussion: We first present a highly optimized implementation of a ray tracer that improves performance by more than an order of magnitude compared to currently available ray tracers. The new algorithm make better use of computational resources such as caches and SIMD instructions and better exploits image and object space coherence. Secondly, we show that this software implementation can challenge and even outperform high-end graphics hardware in interactive rendering performance for complex environments. We also provide an brief overview of the benefits of ray tracing over rasterization algorithms and point out the potential of interactive ray tracing both in hardware and software.
Ingo Wald, Philipp Slusallek, Carsten Benthin, Markus Wagner 0004
Comput. Graph. Forum1