EDBT 2026 Demo / reviewers in the wild / expert
Ingo Wald
dblp:73/3396
· DBLP profile ↗
40ranked-venue papers
15as first author
11since 2021 · last 2026
0000-0003-0046-713XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 14 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 3Systems, architecture and hardware · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Materializing Inter-Channel Relationships With Multi-Density Woodcock TrackingabstractVolume rendering techniques for scientific visualization have recently shifted toward Monte Carlo (MC) methods for their flexibility and robustness, but their use in multi-channel visualization remains underexplored. Traditional multi-channel volume rendering often relies on arbitrary, non-physically based color blending functions that hinder interpretation. We introduce multi-density Woodcock tracking, a simple extension of Woodcock tracking that leverages an MC method to produce high-fidelity, physically grounded multi-channel renderings without arbitrary blending. By generalizing Woodcock's distance tracking, we provide a unified blending modality that also integrates blending functions from prior works. We further implement effects that enhance boundary and feature recognition. By accumulating frames in real-time, our approach delivers high-quality visualizations with perceptual benefits, demonstrated on diverse datasets. Alper Sahistan, Stefan Zellmann, Haichao Miao, Nathan Morrical, Ingo Wald, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Visualization of Large Non-Trivially Partitioned Unstructured Data With Native Distribution on High-Performance Computing SystemsabstractInteractively visualizing large finite element simulation data on High-Performance Computing (HPC) systems poses several difficulties. Some of these relate to unstructured data, which, even on a single node, is much more expensive to render compared to structured volume data. Worse yet, in the data parallel rendering context, such data with highly non-convex spatial domain boundaries will cause rays along its silhouette to enter and leave a given rank's domains at different distances. This straddling, in turn, poses challenges for both ray marching, which usually assumes successive elements to share a face, and compositing, which usually assumes a single fragment per pixel per rank. We holistically address these issues using a combination of three inter-operating techniques: first, we use a highly optimized GPU ray marching technique that, given an entry point, can march a ray to its exit point with high-performance by exploiting an exclusive-or (XOR) based compaction scheme. Second, we use hardware-accelerated ray tracing to efficiently find the proper entry points for these marching operations. Third, we use a "deep" compositing scheme to properly handle cases where different ranks' ray segments interleave in depth. We use GPU-to-GPU remote direct memory access (RDMA) to achieve interactive frame rates of 10-15 frames per second and higher for our motivating use case, the Fun3D NASA Mars Lander. Alper Sahistan, Serkan Demirci, Ingo Wald, Stefan Zellmann, João Barbosa, Nathan Morrical, Ugur Güdükbay |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Beyond ExaBricks: GPU Volume Path Tracing of AMR DataabstractAbstract Adaptive Mesh Refinement (AMR) is becoming a prevalent data representation for HPC, and thus also for scientific visualization. AMR data is usually cell centric (which imposes numerous challenges), complex, and generally hard to render. Recent work on GPU‐accelerated AMR rendering has made much progress towards real‐time volume and isosurface rendering of such data, but so far this work has focused exclusively on ray marching, with simple lighting models and without scattering events or global illumination. True high‐quality rendering requires a modified approach that is able to trace arbitrary incoherent paths; but this may not be a perfect fit for the types of data structures recently developed for ray marching. In this paper, we describe a novel approach to high‐quality path tracing of complex AMR data, with a specific focus on analyzing and comparing different data structures and algorithms to achieve this goal. Stefan Zellmann, Qi Wu 0015, Alper Sahistan, Kwan-Liu Ma, Ingo Wald |
Comput. Graph. Forum | 5 |
| 2023 | State-of-the-art in Large-Scale Volume Visualization Beyond Structured DataabstractAbstract Volume data these days is usually massive in terms of its topology, multiple fields, or temporal component. With the gap between compute and memory performance widening, the memory subsystem becomes the primary bottleneck for scientific volume visualization. Simple, structured, regular representations are often infeasible because the buses and interconnects involved need to accommodate the data required for interactive rendering. In this state‐of‐the‐art report, we review works focusing on large‐scale volume rendering beyond those typical structured and regular grid representations. We focus primarily on hierarchical and adaptive mesh refinement representations, unstructured meshes, and compressed representations that gained recent popularity. We review works that approach this kind of data using strategies such as out‐of‐core rendering, massive parallelism, and other strategies to cope with the sheer size of the ever‐increasing volume of data produced by today's supercomputers and acquisition devices. We emphasize the data management side of large‐scale volume rendering systems and also include a review of tools that support the various volume data types discussed. Jonathan Sarton, Stefan Zellmann, Serkan Demirci, Ugur Güdükbay, Welcome Alexandre-Barff, Laurent Lucas, Jean-Michel Dischler, Stefan Wesner, Ingo Wald |
Comput. Graph. Forum | 9 |
| 2023 | Data Parallel Multi-GPU Path Tracing using Ray Queue CyclingabstractAbstract We propose a novel approach to data‐parallel path tracing on single‐node/multi‐GPU hardware that builds on ray forwarding, but which aims—above all else—at generality and practicability. We do this by avoiding any attempts at reducing the number of traces or forward operations performed, and instead focus on always using all GPUs' aggregate compute and bandwidth to effectively trace each ray on every GPU. We show that—counter‐intuitively—this is both feasible and desirable; and that when run on typical data‐center/cloud hardware, the resulting framework not only achieves good performance and scalability, but also comes with significantly fewer limitations, assumptions, or preprocessing requirements than existing techniques. Ingo Wald, Milan Jaros, Stefan Zellmann |
Comput. Graph. Forum | 1 |
| 2023 | Memory-Efficient GPU Volume Path Tracing of AMR Data Using the Dual MeshabstractAbstract A common way to render cell‐centric adaptive mesh refinement (AMR) data is to compute the dual mesh and visualize that with a standard unstructured element renderer. While the dual mesh provides a high‐quality interpolator, the memory requirements of the dual meshdata structureare significantly higher than those of the original grid, which prevents rendering very large data sets. We introduce a GPU‐friendly data structure and a clustering algorithm that allow for efficient AMR dual mesh rendering with a competitive memory footprint. Fundamentally, any off‐the‐shelf unstructured element renderer running on GPUs could be extended to support our data structure just by adding agridletelement type in addition to the standard tetrahedra, pyramids, wedges, and hexahedra supported by default. We integrated the data structure into a volumetric path tracer to compare it to various state‐of‐the‐art unstructured element sampling methods. We show that our data structure easily competes with these methods in terms of rendering performance, but is much more memory‐efficient. Stefan Zellmann, Qi Wu 0015, Kwan-Liu Ma, Ingo Wald |
Comput. Graph. Forum | 4 |
| 2023 | Quick Clusters: A GPU-Parallel Partitioning for Efficient Path Tracing of Unstructured Volumetric GridsabstractWe propose a simple yet effective method for clustering finite elements to improve preprocessing times and rendering performance of unstructured volumetric grids without requiring auxiliary connectivity data. Rather than building bounding volume hierarchies (BVHs) over individual elements, we sort elements along with a Hilbert curve and aggregate neighboring elements together, improving BVH memory consumption by over an order of magnitude. Then to further reduce memory consumption, we cluster the mesh on the fly into sub-meshes with smaller indices using a series of efficient parallel mesh re-indexing operations. These clusters are then passed to a highly optimized ray tracing API for point containment queries and ray-cluster intersection testing. Each cluster is assigned a maximum extinction value for adaptive sampling, which we rasterize into non-overlapping view-aligned bins allocated along the ray. These maximum extinction bins are then used to guide the placement of samples along the ray during visualization, reducing the number of samples required by multiple orders of magnitude (depending on the dataset), thereby improving overall visualization interactivity. Using our approach, we improve rendering performance over a competitive baseline on the NASA Mars Lander dataset from 6× (1 frame per second (fps) and 1.0 M rays per second (rps) up to now 6 fps and 12.4 M rps, now including volumetric shadows) while simultaneously reducing memory consumption by 3×(33 GB down to 11 GB) and avoiding any offline preprocessing steps, enabling high-quality interactive visualization on consumer graphics cards. Then by utilizing the full 48 GB of an RTX 8000, we improve the performance of Lander by 17 × (1 fps up to 17 fps, 1.0 M rps up to 35.6 M rps). Nathan Morrical, Alper Sahistan, Ugur Güdükbay, Ingo Wald, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Accelerating Unstructured Mesh Point Location With RT CoresabstractWe present a technique that leverages ray tracing hardware available in recent Nvidia RTX GPUs to solve a problem other than classical ray tracing. Specifically, we demonstrate how to use these units to accelerate the point location of general unstructured elements consisting of both planar and bilinear faces. This unstructured mesh point location problem has previously been challenging to accelerate on GPU architectures; yet, the performance of these queries is crucial to many unstructured volume rendering and compute applications. Starting with a CUDA reference method, we describe and evaluate three approaches that reformulate these point queries to incrementally map algorithmic complexity to these new hardware ray tracing units. Each variant replaces the simpler problem of point queries with a more complex one of ray queries. Initial variants exploit ray tracing cores for accelerated BVH traversal, and subsequent variants use ray-triangle intersections and per-face metadata to detect point-in-element intersections. Although these later variants are more algorithmically complex, they are significantly faster than the reference method thanks to hardware acceleration. Using our approach, we improve the performance of an unstructured volume renderer by up to 4× for tetrahedral meshes and up to 15× for general bilinear element meshes, matching, or out-performing state-of-the-art solutions while simultaneously improving on robustness and ease-of-implementation. Nathan Morrical, Ingo Wald, Will Usher 0001, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | A Memory Efficient Encoding for Ray Tracing Large Unstructured DataabstractIn theory, efficient and high-quality rendering of unstructured data should greatly benefit from modern GPUs, but in practice, GPUs are often limited by the large amount of memory that large meshes require for element representation and for sample reconstruction acceleration structures. We describe a memory-optimized encoding for large unstructured meshes that efficiently encodes both the unstructured mesh and corresponding sample reconstruction acceleration structure, while still allowing for fast random-access sampling as required for rendering. We demonstrate that for large data our encoding allows for rendering even the 2.9 billion element Mars Lander on a single off-the-shelf GPU-and the largest 6.3 billion version on a pair of such GPUs. Ingo Wald, Nathan Morrical, Stefan Zellmann |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2021 | Multi-level tetrahedralization-based accelerator for ray-tracing animated scenesabstractAbstract We describe a hybrid acceleration structure for ray tracing. The hybrid structure is a Bounding Volume Hierarchy (BVH) where the leaf nodes are tetrahedralized for a decent ray‐surface intersection performance. We use the hybrid acceleration structure (BTH) in a two‐level acceleration structure for rendering animated scenes. There is a BVH at the top level in this two‐level structure and the proposed hybrid structure (BTH) at the bottom level. We test the proposed two‐level structure (BVH‐BTH) for various animated scenes and obtained promising results against other acceleration structures in terms of rendering times. The two‐level BVH‐BTH structure outperforms the two‐level BVH structure for the tested dynamic scenes. Aytek Aman, Serkan Demirci, Ugur Güdükbay, Ingo Wald |
Comput. Animat. Virtual Worlds | 4 |
| 2021 | Ray Tracing Structured AMR Data Using ExaBricksabstractStructured Adaptive Mesh Refinement (Structured AMR) enables simulations to adapt the domain resolution to save computation and storage, and has become one of the dominant data representations used by scientific simulations; however, efficiently rendering such data remains a challenge. We present an efficient approach for volume- and iso-surface ray tracing of Structured AMR data on GPU-equipped workstations, using a combination of two different data structures. Together, these data structures allow a ray tracing based renderer to quickly determine which segments along the ray need to be integrated and at what frequency, while also providing quick access to all data values required for a smooth sample reconstruction kernel. Our method makes use of the RTX ray tracing hardware for surface rendering, ray marching, space skipping, and adaptive sampling; and allows for interactive changes to the transfer function and implicit iso-surfacing thresholds. We demonstrate that our method achieves high performance with little memory overhead, enabling interactive high quality rendering of complex AMR data sets on individual GPU workstations. Ingo Wald, Stefan Zellmann, Will Usher 0001, Nathan Morrical, Ulrich Lang 0002, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2019 | Ray Tracing Generalized Tube Primitives: Method and ApplicationsabstractWe present a general high-performance technique for ray tracing generalized tube primitives. Our technique efficiently supports tube primitives with fixed and varying radii, general acyclic graph structures with bifurcations, and correct transparency with interior surface removal. Such tube primitives are widely used in scientific visualization to represent diffusion tensor imaging tractographies, neuron morphologies, and scalar or vector fields of 3D flow. We implement our approach within the OSPRay ray tracing framework, and evaluate it on a range of interactive visualization use cases of fixed- and varying-radius streamlines, pathlines, complex neuron morphologies, and brain tractographies. Our proposed approach provides interactive, high-quality rendering, with low memory overhead. Mengjiao Han, Ingo Wald, Will Usher 0001, Qi Wu 0015, Feng Wang 0013, Valerio Pascucci, Charles D. Hansen, Chris R. Johnson 0001 |
Comput. Graph. Forum | 2 |
| 2019 | Scalable Ray Tracing Using the Distributed FrameBufferabstractAbstract Image‐ and data‐parallel rendering across multiple nodes on high‐performance computing systems is widely used in visualization to provide higher frame rates, support large data sets, and render data in situ. Specifically for in situ visualization, reducing bottlenecks incurred by the visualization and compositing is of key concern to reduce the overall simulation runtime. Moreover, prior algorithms have been designed to support either image‐ or data‐parallel rendering and impose restrictions on the data distribution, requiring different implementations for each configuration. In this paper, we introduce the Distributed FrameBuffer, an asynchronous image‐processing framework for multi‐node rendering. We demonstrate that our approach achieves performance superior to the state of the art for common use cases, while providing the flexibility to support a wide range of parallel rendering algorithms and data distributions. By building on this framework, we extend the open‐source ray tracing library OSPRay with a data‐distributed API, enabling its use in data‐distributed and in situ visualization applications. Will Usher 0001, Ingo Wald, Jefferson Amstutz, Johannes Günther 0001, Carson Brownlee, Valerio Pascucci |
Comput. Graph. Forum | 2 |
| 2019 | CPU Isosurface Ray Tracing of Adaptive Mesh Refinement DataabstractAdaptive mesh refinement (AMR) is a key technology for large-scale simulations that allows for adaptively changing the simulation mesh resolution, resulting in significant computational and storage savings. However, visualizing such AMR data poses a significant challenge due to the difficulties introduced by the hierarchical representation when reconstructing continuous field values. In this paper, we detail a comprehensive solution for interactive isosurface rendering of block-structured AMR data. We contribute a novel reconstruction strategy-the octant method-which is continuous, adaptive and simple to implement. Furthermore, we present a generally applicable hybrid implicit isosurface ray-tracing method, which provides better rendering quality and performance than the built-in sampling-based approach in OSPRay. Finally, we integrate our octant method and hybrid isosurface geometry into OSPRay as a module, providing the ability to create high-quality interactive visualizations combining volume and isosurface representations of BS-AMR data. We evaluate the rendering performance, memory consumption and quality of our method on two gigascale block-structured AMR datasets. Feng Wang 0013, Ingo Wald, Qi Wu 0015, Will Usher 0001, Chris R. Johnson 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | OSPRay - A CPU Ray Tracing Framework for Scientific VisualizationabstractScientific data is continually increasing in complexity, variety and size, making efficient visualization and specifically rendering an ongoing challenge. Traditional rasterization-based visualization approaches encounter performance and quality limitations, particularly in HPC environments without dedicated rendering hardware. In this paper, we present OSPRay, a turn-key CPU ray tracing framework oriented towards production-use scientific visualization which can utilize varying SIMD widths and multiple device backends found across diverse HPC resources. This framework provides a high-quality, efficient CPU-based solution for typical visualization workloads, which has already been integrated into several prevalent visualization packages. We show that this system delivers the performance, high-level API simplicity, and modular device support needed to provide a compelling new rendering framework for implementing efficient scientific visualization workflows. Ingo Wald, Gregory P. Johnson, Jefferson Amstutz, Carson Brownlee, Aaron Knoll, Jim Jeffers, Johannes Günther 0001, Paul A. Navrátil |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2014 | RBF Volume Ray Casting on Multicore and Manycore CPUsabstractAbstract Modern supercomputers enable increasingly large N‐body simulations using unstructured point data. The structures implied by these points can be reconstructed implicitly. Direct volume rendering of radial basis function (RBF) kernels in domain‐space offers flexible classification and robust feature reconstruction, but achieving performant RBF volume rendering remains a challenge for existing methods on both CPUs and accelerators. In this paper, we present a fast CPU method for direct volume rendering of particle data with RBF kernels. We propose a novel two‐pass algorithm: first sampling the RBF field using coherent bounding hierarchy traversal, then subsequently integrating samples along ray segments. Our approach performs interactively for a range of data sets from molecular dynamics and astrophysics up to 82 million particles. It does not rely on level of detail or subsampling, and offers better reconstruction quality than structured volume rendering of the same data, exhibiting comparable performance and requiring no additional preprocessing or memory footprint other than the BVH. Lastly, our technique enables multi‐field, multi‐material classification of particle data, providing better insight and analysis. Aaron Knoll, Ingo Wald, Paul A. Navrátil, Anne Bowen, Khairi Reda, Michael E. Papka, Kelly P. Gaither |
Comput. Graph. Forum | 2 |
| 2014 | Embree: a kernel framework for efficient CPU ray tracingabstractWe describe Embree, an open source ray tracing framework for x86 CPUs. Embree is explicitly designed to achieve high performance in professional rendering environments in which complex geometry and incoherent ray distributions are common. Embree consists of a set of low-level kernels that maximize utilization of modern CPU architectures, and an API which enables these kernels to be used in existing renderers with minimal programmer effort. In this paper, we describe the design goals and software architecture of Embree, and show that for secondary rays in particular, the performance of Embree is competitive with (and often higher than) existing state-of-the-art methods on CPUs and GPUs. Ingo Wald, Sven Woop, Carsten Benthin, Gregory S. Johnson, Manfred Ernst |
ACM Trans. Graph. | 1 |
| 2012 | Extending a C-like language for portable SIMD programmingabstractSIMD instructions are common in CPUs for years now. Using these instructions effectively requires not only vectorization of code, but also modifications to the data layout. However, automatic vectorization techniques are often not powerful enough and suffer from restricted scope of applicability; hence, programmers often vectorize their programs manually by using intrinsics: compiler-known functions that directly expand to machine instructions. They significantly decrease programmer productivity by enforcing a very error-prone and hard-to-read assembly-like programming style. Furthermore, intrinsics are not portable because they are tied to a specific instruction set. Roland Leißa, Sebastian Hack, Ingo Wald |
PPoPP | 3 |
| 2012 | Combining Single and Packet-Ray Tracing for Arbitrary Ray Distributions on the Intel MIC ArchitectureabstractWide-SIMD hardware is power and area efficient, but it is challenging to efficiently map ray tracing algorithms to such hardware especially when the rays are incoherent. The two most commonly used schemes are either packet tracing, or relying on a separate traversal stack for each SIMD lane. Both work great for coherent rays, but suffer when rays are incoherent: The former experiences a dramatic loss of SIMD utilization once rays diverge; the latter requires a large local storage, and generates multiple incoherent streams of memory accesses that present challenges for the memory system. In this paper, we introduce a single-ray tracing scheme for incoherent rays that uses just one traversal stack on 16-wide SIMD hardware. It uses a bounding-volume hierarchy with a branching factor of four as the acceleration structure, exploits four-wide SIMD in each box and primitive intersection test, and uses 16-wide SIMD by always performing four such node or primitive tests in parallel. We then extend this scheme to a hybrid tracing scheme that automatically adapts to varying ray coherence by starting out with a 16-wide packet scheme and switching to the new single-ray scheme as soon as rays diverge. We show that on the Intel Many Integrated Core architecture this hybrid scheme consistently, and over a wide range of scenes and ray distributions, outperforms both packet and single-ray tracing. Carsten Benthin, Ingo Wald, Sven Woop, Manfred Ernst, William R. Mark |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2012 | Fast Construction of SAH BVHs on the Intel Many Integrated Core (MIC) ArchitectureabstractWe investigate how to efficiently build bounding volume hierarchies (BVHs) with surface area heuristic (SAH) on the Intel Many Integrated Core (MIC) Architecture. To achieve maximum performance, we use four key concepts: progressive 10-bit quantization to reduce cache footprint with negligible loss in BVH quality; an AoSoA data layout that allows efficient streaming and SIMD processing; high-performance SIMD kernels for binning and partitioning; and a parallelization framework with several build-specific optimizations. The resulting system is more than an order of magnitude faster than today's high-end GPU builders for comparable BVHs; it is usually faster even than spatial median builders; it can build SAH BVHs almost as fast as existing GPUs and CPUs- and CPU-based approaches can build regular grids; and in aggregate "build+render" performance is significantly faster than the best published numbers for either of these systems, be it CPU or GPU, BVH, kd-tree, or grid. Ingo Wald |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2011 | Full-resolution interactive CPU volume rendering with coherent BVH traversalabstractWe present an efficient method for volume rendering by ray casting on the CPU. We employ coherent packet traversal of an implicit bounding volume hierarchy, heuristically pruned using preintegrated transfer functions, to exploit empty or homogeneous space. We also detail SIMD optimizations for volumetric integration, trilinear interpolation, and gradient lighting. The resulting system performs well on low-end and laptop hardware, and can outperform out-of-core GPU methods by orders of magnitude when rendering large volumes without level-of-detail (LOD) on a workstation. We show that, while slower than GPU methods for low-resolution volumes, an optimized CPU renderer does not require LOD to achieve interactive performance on large data sets. Aaron Knoll, Sebastian Thelen, Ingo Wald, Charles D. Hansen, Hans Hagen, Michael E. Papka |
PacificVis | 3 |
| 2009 | State of the Art in Ray Tracing Animated ScenesabstractAbstract Ray tracing has long been a method of choice for off‐line rendering, but traditionally was too slow for interactive use. With faster hardware and algorithmic improvements this has recently changed, and real‐time ray tracing is finally within reach. However, real‐time capability also opens up new problems that do not exist in an off‐line environment. In particular real‐time ray tracing offers the opportunity to interactively ray trace moving/animated scene content. This presents a challenge to the data structures that have been developed for ray tracing over the past few decades. Spatial data structures crucial for fast ray tracing must be rebuilt or updated as the scene changes, and this can become a bottleneck for the speed of ray tracing. This bottleneck has recently received much attention by researchers and that has resulted in a multitude of different algorithms, data structures and strategies for handling animated scenes. The effectiveness of techniques for ray tracing dynamic scenes vary dramatically depending on details such as scene complexity, model structure, type of motion and the coherency of the rays. Consequently, there is so far no approach that is best in all cases, and determining the best technique for a particular problem can be a challenge. In this State of the Art Report (STAR), we aim to survey the different approaches to ray tracing animated scenes, discussing their strengths and weaknesses, and their relationship to other approaches. The overall goal is to help the reader choose the best approach depending on the situation, and to expose promising areas where there is potential for algorithmic improvements. Ingo Wald, William R. Mark, Johannes Günther 0001, Solomon Boulos, Thiago Ize, Warren A. Hunt, Steven G. Parker, Peter Shirley |
Comput. Graph. Forum | 1 |
| 2009 | Coherent multiresolution isosurface ray tracing
Aaron Knoll, Ingo Wald, Charles D. Hansen |
Vis. Comput. | 2 |
| 2008 | Fast, parallel, and asynchronous construction of BVHs for ray tracing animated scenes
Ingo Wald, Thiago Ize, Steven G. Parker |
Comput. Graph. | 1 |
| 2008 | Sequential Monte Carlo Adaptation in Low-Anisotropy Participating MediaabstractAbstract This paper presents a novel method that effectively combines both control variates and importance sampling in a sequential Monte Carlo context. The radiance estimates computed during the rendering process are cached in a 5D adaptive hierarchical structure that defines dynamic predicate functions for both variance reduction techniques and guarantees well‐behaved PDFs, yielding continually increasing efficiencies thanks to a marginal computational overhead. While remaining unbiased, the technique is effective within a single pass as both estimation and caching are done online, exploiting the coherency in illumination while being independent of the actual scene representation. The method is relatively easy to implement and to tune via a single parameter, and we demonstrate its practical benefits with important gains in convergence rate and competitive results with state of the art techniques. Vincent Pegoraro, Ingo Wald, Steven G. Parker |
Comput. Graph. Forum | 2 |
| 2007 | Interactive Iso-Surface Ray Tracing of Massive Volumetric Data Sets
Heiko Friedrich, Ingo Wald, Johannes Günther 0001, Gerd Marmitt, Philipp Slusallek |
EGPGV | 2 |
| 2007 | Asynchronous BVH Construction for Ray Tracing Dynamic Scenes on Parallel Multi-Core Architectures
Thiago Ize, Ingo Wald, Steven G. Parker |
EGPGV | 2 |
| 2007 | Packet-based whitted and distribution ray tracingabstractMuch progress has been made toward interactive ray tracing, but most research has focused specifically on ray casting. A common approach is to use "packets" of rays to amortize cost across sets of rays. Whether "packets" can be used to speed up the cost of reflection and refraction rays is unclear. The issue is complicated since such rays do not share common origins and often have less directional coherence than viewing and shadow rays. Since the primary advantage of ray tracing over rasterization is the computation of global effects, such as accurate reflection and refraction, this lack of knowledge should be corrected. We are also interested in exploring whether distribution ray tracing, due to its stochastic properties, further erodes the effectiveness of techniques used to accelerate ray casting. This paper addresses the question of whether packet-based ray tracing algorithms can be effectively used for more than visibility computation. We show that by choosing an appropriate data structure and a suitable packet assembly algorithm we can extend the idea of "packets" from ray casting to Whitted-style and distribution ray tracing, while maintaining efficiency. Solomon Boulos, David Edwards, J. Dylan Lacewell, Joe Michael Kniss, Jan Kautz, Peter Shirley, Ingo Wald |
Graphics Interface | 7 |
| 2007 | Ray tracing deformable scenes using dynamic bounding volume hierarchies
Ingo Wald, Solomon Boulos, Peter Shirley |
ACM Trans. Graph. | 1 |
| 2007 | A Coherent Grid Traversal Approach to Visualizing Particle-Based Simulation DataabstractWe present an approach to visualizing particle-based simulation data using interactive ray tracing and describe an algorithmic enhancement that exploits the properties of these data sets to provide highly interactive performance and reduced storage requirements. This algorithm for fast packet-based ray tracing of multilevel grids enables the interactive visualization of large time-varying data sets with millions of particles and incorporates advanced features like soft shadows. We compare the performance of our approach with two recent particle visualization systems: one based on an optimized single ray grid traversal algorithm and the other on programmable graphics hardware. This comparison demonstrates that the new algorithm offers an attractive alternative for interactive particle visualization. Christiaan P. Gribble, Thiago Ize, Andrew Kensler, Ingo Wald, Steven G. Parker |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2007 | Interactive Isosurface Ray Tracing of Time-Varying Tetrahedral VolumesabstractWe describe a system for interactively rendering isosurfaces of tetrahedral finite-element scalar fields using coherent ray tracing techniques on the CPU. By employing state-of-the art methods in polygonal ray tracing, namely aggressive packet/frustum traversal of a bounding volume hierarchy, we can accomodate large and time-varying unstructured data. In conjunction with this efficiency structure, we introduce a novel technique for intersecting ray packets with tetrahedral primitives. Ray tracing is flexible, allowing for dynamic changes in isovalue and time step, visualization of multiple isosurfaces, shadows, and depth-peeling transparency effects. The resulting system offers the intuitive simplicity of isosurfacing, guaranteed-correct visual results, and ultimately a scalable, dynamic and consistently interactive solution for visualizing unstructured volumes. Ingo Wald, Heiko Friedrich, Aaron Knoll, Charles D. Hansen |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2006 | Ray Tracing Animated Scenes using Motion DecompositionabstractAbstract Though ray tracing has recently become interactive, its high precomputation time for building spatial indices usually limits its applications to walkthroughs of static scenes. This is a major limitation, as most applications demand support for dynamically animated models. In this paper, we present a new approach to ray trace a special but important class of dynamic scenes, namely models whose connectivity does not change over time and for which all possible poses are known in advance. We support these kinds of models by introducing two new concepts: motion decomposition, and fuzzy kd‐trees. We analyze the animation and break the model down into submeshes with similar motion. For each of these submeshes and for every time step, we calculate a best affine transformation through a least square approach. Any residual motion is then captured in a single "fuzzy kd‐tree" for the entire animation. Together, these techniques allow for ray tracing animations without rebuilding the spatial index structures for the submeshes, resulting in interactive frame rates of 5 to 15 fps even on a single CPU. Categories and Subject Descriptors (according to ACM CCS): I.3.7 [Computer Graphics]: Ray tracing I.3.6 [Methodology and Techniques]: Graphics data structures and data types Johannes Günther 0001, Heiko Friedrich, Ingo Wald, Hans-Peter Seidel, Philipp Slusallek |
Comput. Graph. Forum | 3 |
| 2006 | Ray tracing animated scenes using coherent grid traversalabstractWe present a new approach to interactive ray tracing of moderate-sized animated scenes based on traversing frustum-bounded packets of coherent rays through uniform grids. By incrementally computing the overlap of the frustum with a slice of grid cells, we accelerate grid traversal by more than a factor of 10, and achieve ray tracing performance competitive with the fastest known packet-based kd-tree ray tracers. The ability to efficiently rebuild the grid on every frame enables this performance even for fully dynamic scenes that typically challenge interactive ray tracing systems. Ingo Wald, Thiago Ize, Andrew Kensler, Aaron Knoll, Steven G. Parker |
ACM Trans. Graph. | 1 |
| 2005 | Faster Isosurface Ray Tracing Using Implicit KD-TreesabstractThe visualization of high-quality isosurfaces at interactive rates is an important tool in many simulation and visualization applications. Today, isosurfaces are most often visualized by extracting a polygonal approximation that is then rendered via graphics hardware or by using a special variant of preintegrated volume rendering. However, these approaches have a number of limitations in terms of the quality of the isosurface, lack of performance for complex data sets, or supported shading models. An alternative isosurface rendering method that does not suffer from these limitations is to directly ray trace the isosurface. However, this approach has been much too slow for interactive applications unless massively parallel shared-memory supercomputers have been used. In this paper, we implement interactive isosurface ray tracing on commodity desktop PCs by building on recent advances in real-time ray tracing of polygonal scenes and using those to improve isosurface ray tracing performance as well. The high performance and scalability of our approach will be demonstrated with several practical examples, including the visualization of highly complex isosurface data sets, the interactive rendering of hybrid polygonal/isosurface scenes, including high-quality ray traced shading effects, and even interactive global illumination on isosurfaces. Ingo Wald, Heiko Friedrich, Gerd Marmitt, Philipp Slusallek, Hans-Peter Seidel |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2004 | VRML Scene Graphs on an Interactive Ray Tracing Engine
Andreas Dietrich 0001, Ingo Wald, Markus Wagner 0004, Philipp Slusallek |
VR | 2 |
| 2004 | Colorplate: VRML Scene Graphs on an Interactive Ray Tracing Engine
Andreas Dietrich 0001, Ingo Wald, Markus Wagner 0004, Philipp Slusallek |
VR | 2 |
| 2004 | Balancing Considered Harmful - Faster Photon Mapping using the Voxel Volume HeuristicabstractAbstract Photon mapping is one of the most important algorithms for computing global illumination. Especially for efficiently producing convincing caustics, there are no real alternatives to photon mapping. On the other hand, photon mapping is also quite costly: Each radiance lookup requires to find the k nearest neighbors in a kd‐tree, which can be more costly than shooting several rays. Therefore, the nearest‐neighbor queries often dominate the rendering time of a photon map based renderer. In this paper, we present a method that reorganizes — i.e. un balances — the kd‐tree for storing the photons in a way that allows for finding the k‐nearest neighbors much more efficiently, thereby accelerating the radiance estimates by a factor of 1.2–3.4. Most importantly, our method still finds exactly the same k‐nearest‐neighbors as the original method, without introducing any approximations or loss of accuracy. The impact of our method is demonstrated with several practical examples. Categories and Subject Descriptors (according to ACM CCS): I.3.3 [Computer Graphics]: Global Illumination I.3.7 [Computer Graphics]: Raytracing Ingo Wald, Johannes Günther 0001, Philipp Slusallek |
Comput. Graph. Forum | 1 |
| 2003 | Interactive Ray Tracing on Commodity PC Clusters
Ingo Wald, Carsten Benthin, Andreas Dietrich 0001, Philipp Slusallek |
Euro-Par | 1 |
| 2003 | A Scalable Approach to Interactive Global IlluminationabstractAbstract The addition of global illumination can dramatically increase the realism achievable when rendering virtual environments.In particular with interactive applications we expect the environment to reflect changes in the scenedue to global lighting effects instead of it being just a static backdrop. However, a sufficiently fast and accuratecomputation of global illumination at interactive rates has been difficult even with recent approaches based onrealtime ray tracing. In this paper we present a highly scalable approach to interactive global illumination. It fully recomputes a high‐qualitysolution for each frame and thus offers immediate feedback even for dynamic scenes, achieving more than20 fps for simple scenes. Compared to previous systems we increased the raw performance by a factor of up toeight and removed the bottlenecks that were limiting scalability. The system now scales linearly in quality andavailable computing resources, tested with up to 48 CPUs in a commodity PC‐cluster. Due to its logarithmicscaling property with respect to scene complexity it even supports lighting simulation in complex scenes with morethan 50 million triangles. This scalability allows applications to perform flexible performance trade‐offs. We alsoargue that the realism achievable through interactive global illumination will make it a standard feature of future3D graphics systems once the required computing resources are readily available. Carsten Benthin, Ingo Wald, Philipp Slusallek |
Comput. Graph. Forum | 2 |
| 2001 | Interactive Rendering with Coherent Ray TracingabstractFor almost two decades researchers have argued that ray tracing will eventually become faster than the rasterization technique that completely dominates todays graphics hardware. However, this has not happened yet. Ray tracing is still exclusively being used for off-line rendering of photorealistic images and it is commonly believed that ray tracing is simply too costly to ever challenge rasterization-based algorithms for interactive use. However, there is hardly any scientific analysis that supports either point of view. In particular there is no evidence of where the crossover point might be, at which ray tracing would eventually become faster, or if such a point does exist at all. This paper provides several contributions to this discussion: We first present a highly optimized implementation of a ray tracer that improves performance by more than an order of magnitude compared to currently available ray tracers. The new algorithm make better use of computational resources such as caches and SIMD instructions and better exploits image and object space coherence. Secondly, we show that this software implementation can challenge and even outperform high-end graphics hardware in interactive rendering performance for complex environments. We also provide an brief overview of the benefits of ray tracing over rasterization algorithms and point out the potential of interactive ray tracing both in hardware and software. Ingo Wald, Philipp Slusallek, Carsten Benthin, Markus Wagner 0004 |
Comput. Graph. Forum | 1 |