EDBT 2026 Demo / reviewers in the wild / expert
Yuan-Hsi Chou
dblp:243/5096
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0006-9554-0545ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Treelet Accelerated Ray Tracing on GPUsabstractDespite advances in hardware acceleration, ray tracing use in real-time rendering is limited and often lowers frame rates, leading users such as video game players to disable the feature entirely. Prior work has shown that dividing the BVH tree into smaller subtrees (treelets) and traversing all rays that visit a treelet before switching treelets can significantly reduce memory traffic on a specialized accelerator, but there are many challenges to applying treelets to GPUs. We find that a naive treelet implementation is ineffective and propose optimizations to improve performance. Virtualized Treelet Queues consist of two main components. Ray virtualization increases the number of concurrent rays in flight to create more cache reuse opportunities by terminating raygen shaders that have already issued their trace ray instruction, reclaiming CUDA cores and allowing more raygen shaders to be executed. To take advantage of the increased concurrent rays, we propose a dynamic treelet queue architecture that dynamically switches between traversal modes to increase efficiency. We also find that performing warp repacking boosts SIMT efficiency of warps in the RT unit which is crucial to achieving good traversal performance with treelet queues. Our simulations show virtualized treelet queues achieve on average 95% speedup compared to a baseline GPU with ray tracing acceleration across all scenes in LumiBench rendered with path tracing at one sample per pixel with three max bounces per ray. Yuan-Hsi Chou, Tor M. Aamodt |
ASPLOS (2) | 1 |
| 2024 | Zatel: Sample Complexity-Aware Scale-Model Simulation for Ray TracingabstractRay tracing is a computationally intensive rendering technique that simulates the behavior of light rays as they interact with objects in a scene. It is becoming increasingly popular in video games and is already the de facto standard for animated movies. However, current hardware still struggles to efficiently ray trace complex scenes and requires further research. To evaluate early-stage hardware proposals that accelerate ray tracing for G PU s, one either uses cycle-accurate simulators, which are highly accurate and flexible but slow, or other models that are an order of magnitude faster but provide limited output with high error margins. In this paper, we propose Zatel, a prediction methodology for evaluating GPU performance on ray tracing workloads. We observe that the desired metrics can be estimated with reasonable accuracy by only tracing a representative subset of pixels. Furthermore, the parallel nature of GPUs allows us to split the scene into chunks, which lets Zatel execute faster using downscaled GPU configurations. We incorporate these two optimization steps into Zatel and evaluate it on a benchmark suite for ray tracing using Vulkan-Sim, a cycle-accurate simulator. By relying on Vulkan-Sim, architectural changes are captured through the simulator, and Zatel does not need to be updated to support each change. Zatel records less than 1 % error with 10 x simulation time speedup for measuring simulation cycles on a mobile G PU. Davit Grigoryan, Yuan-Hsi Chou, Tor M. Aamodt |
ISPASS | 2 |
| 2024 | Generalizing Ray Tracing Accelerators for Tree Traversals on GPUsabstractTree traversal is a fundamental operation in many applications, such as database indexing and physics simulations. Although tree traversals feature high parallelism, they are inherently divergent and irregular, leading to inefficient performance on GPUs. Tree traversals are also prevalent in ray tracing, which is executed on dedicated Ray-Tracing Accelerators (RTAs) in modern GPUs to mitigate inefficiencies such as control flow divergence and underutilization of memory bandwidth by irregular memory accesses. In this paper, we propose the Tree Traversal Accelerator (TTA) to replicate the success of RTAs in ray tracing for general tree traversal applications. TTAs extend RTAs to support tree structures and operations beyond those in ray tracing, such as B- Tree search and radius search algorithms, by modifying existing computing units. Despite TTAs' effectiveness, they still rely on fixed-function computations, making it challenging to support other tree-based applications such as N-Body simulation fully. Thus, we introduce TTA + as an alternative design, which modularizes the RTA computing units and makes them programmable, trading some efficiency for flexibility. With less than 1 % increase in RTA area, our proposals can achieve up to S.4x speedup for B-Tree search, 1.7x for N-Body simulation, and 1.2x for select ray-tracing applications. Dongho Ha, Lufei Liu 0001, Yuan-Hsi Chou, Seokjin Go, Won Woo Ro, Hung-Wei Tseng 0001, Tor M. Aamodt |
MICRO | 3 |
| 2023 | Treelet Prefetching For Ray Tracing
Yuan-Hsi Chou, Tyler Nowicki, Tor M. Aamodt |
MICRO | 1 |
| 2022 | Vulkan-Sim: A GPU Architecture Simulator for Ray TracingabstractRay tracing can generate photorealistic images with more convincing visual effects compared to rasterization. Recent hardware advances have enabled ray tracing to be applied in real-time. Current GPUs feature a dedicated ray tracing acceleration unit, and game developers have started to make use of ray tracing APIs to bring more realistic graphics to their players. Industry cooperatively contributed to Vulkan, which recently introduced an open-standard API for ray tracing. However, little has been disclosed about the mapping of this API to hardware. In this paper, we introduce Vulkan-Sim, a detailed cycle-level simulator for enabling architecture research for ray tracing. We extend GPGPU-Sim, integrating it with Mesa, an open-source graphics library to support the Vulkan API, and add dedicated ray traversal and intersection units. We also demonstrate an explicit mapping of the Vulkan ray tracing pipeline to a modern GPU using a technique we call delayed intersection and any-hit execution. Additionally we evaluate several ray tracing workloads with Vulkan-Sim, identifying bottlenecks and inefficiencies of the ray tracing hardware we model. To demonstrate the utility of Vulkan-Sim we conduct two case studies evaluating techniques recently proposed or deployed by industry targeting enhanced ray tracing performance. Mohammadreza Saed, Yuan-Hsi Chou, Lufei Liu 0001, Tyler Nowicki, Tor M. Aamodt |
MICRO | 2 |
| 2021 | Intersection Prediction for Accelerated GPU Ray TracingabstractRay tracing has been used for years in motion picture to generate photorealistic images while faster raster-based shading techniques have been preferred for video games to meet real-time requirements. However, recent Graphics Processing Units (GPUs) incorporate hardware accelerator units designed for ray tracing. These accelerator units target the process of traversing hierarchical tree data structures used to test for ray-object intersections. Distinct rays following similar paths through these structures execute many redundant ray-box intersection tests. We propose a ray intersection predictor that speculatively elides redundant operations during this process and proceeds directly to test primitives that the ray is likely to intersect. A key aspect of our predictor strategy involves identifying hash functions that preserve enough spatial information to identify redundant traversals. We explore how to integrate our ray prediction strategy into existing GPU pipelines along with improving the predictor effectiveness by predicting nodes higher in the tree as well as regrouping and scheduling traversal operations in a low cost, judicious manner. On a mobile class GPU with a ray tracing accelerator unit, we find the addition of a 5.5KB predictor per streaming multiprocessor improves performance for ambient occlusion workloads by a geometric mean of 26%. Lufei Liu 0001, Wesley Chang, Francois Demoullin, Yuan-Hsi Chou, Mohammadreza Saed, David Pankratz, Tyler Nowicki, Tor M. Aamodt |
MICRO | 4 |
| 2020 | Deterministic Atomic BufferingabstractDeterministic execution for GPUs is a desirable property as it helps with debuggability and reproducibility. It is also important for safety regulations, as safety critical workloads are starting to be deployed onto GPUs. Prior deterministic architectures, such as GPUDet, attempt to provide strong determinism for all types of workloads, incurring significant performance overheads due to the many restrictions that are required to satisfy determinism. We observe that a class of reduction workloads, such as graph applications and neural architecture search for machine learning, do not require such severe restrictions to preserve determinism. This motivates the design of our system, Deterministic Atomic Buffering (DAB), which provides deterministic execution with low area and performance overheads by focusing solely on ordering atomic instructions instead of all memory instructions. By scheduling atomic instructions deterministically with atomic buffering, the results of atomic operations are isolated initially and made visible in the future in a deterministic order. This allows the GPU to execute deterministically in parallel without having to serialize its threads for atomic operations as opposed to GPUDet. Our simulation results show that, for atomic-intensive applications, DAB performs 4× better than GPUDet and incurs only a 23% slowdown on average compared to a non-deterministic GPU architecture. We also characterize the bottlenecks and provide insights for future optimizations. Yuan-Hsi Chou, Christopher Ng, Shaylin Cattell, Jeremy Intan, Matthew D. Sinclair, Joseph Devietti, Timothy G. Rogers, Tor M. Aamodt |
MICRO | 1 |