Adrian Zhao

dblp:366/7772 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0002-9907-6314ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 ARC: Warp-level Adaptive Atomic Reduction in GPUs to Accelerate Differentiable Rendering
abstract
Differentiable rendering is widely used in emerging applications that represent any 3D scene as a model trained using gradient descent from 2D images. Recent works (e.g., 3D Gaussian Splatting) use rasterization to enable rendering photo-realistic imagery at high speeds from these learned 3D models. These rasterization-based differentiable rendering methods have been demonstrated to be very promising, providing state-of-art quality for various important tasks. However, training a model to represent a scene is still time-consuming even on powerful GPUs. In this work, we observe that the gradient computation step during model training is a significant bottleneck due to the large number of atomic operations. These atomics overwhelm the atomic units in the L2 cache of GPUs, causing long stalls.
Sankeerth Durvasula, Adrian Zhao, Ruofan Liang, Pawan Kumar Sanjaya, Yushi Guan, Christina Giannoula, Nandita Vijaykumar
ASPLOS (1)2
2024 ACE: Efficient GPU Kernel Concurrency for Input-Dependent Irregular Computational Graphs
abstract
GPUs are widely used to accelerate many important classes of workloads today. However, in this work, we observe that several important emerging classes of workloads, including simulation engines for deep reinforcement learning and dynamic neural networks, are unable to fully utilize the massive parallelism that GPUs offer. These applications tend to have kernels that are small in size, i.e., have few threads and thread blocks that cannot saturate the GPU’s compute resources. Executing independent kernels concurrently is a promising approach to improve parallelism and utilization. However, this inter-kernel concurrency is difficult to leverage in such workloads with existing approaches: First, the inter-kernel dependencies and computational graph are input-dependent and vary each time the application is executed. Second, the computational graphs tend to be irregular, requiring fine-grain scheduling and synchronization; thus incurring significant synchronization overheads if kernel execution is parallelized. In this work, we propose ACE, a new framework that enables lightweight detection of inter-kernel dependencies and low overhead kernel scheduling at runtime. The key idea behind ACE is to perform inter-kernel dependency checks for a small window of kernels at runtime, similar to out-of-order instruction scheduling. This enables concurrent execution of kernels in applications whose computational graphs are input-dependent and require fine-grained scheduling. We propose ACE-SW, a software-only open-source implementation of ACE and ACE-HW, a hardware-software cooperative implementation. ACE-HW further reduces synchronization overheads by reducing communication between the CPU and GPU. We evaluate ACE for deep RL simulation engines and dynamic and static DNNs on both real hardware and a GPU simulator. We demonstrate speedups of up to 2.19 × (1.56 × on average) by improving GPU utilization with concurrent kernel execution.
Sankeerth Durvasula, Adrian Zhao, Raymond Kiguru, Yushi Guan, Zhonghan Chen, Nandita Vijaykumar
PACT2
2024 Distributed Training of Neural Radiance Fields: A Performance Characterization
abstract
Implicit neural representation is an emerging method that leverages deep neural networks and learned parameters to represent 3D scenes efficiently and accurately. Neural radiance field (NeRF) is a state-of-art implicit representation that achieves photorealistic 3D reconstruction with compact neural network models. However, as the complexity and scale of the scene increase, training NeRF models with a single GPU proves insufficient for achieving fast training and high-quality reconstruction. To address this challenge, prior works proposed distributed NeRF training methods. This is the first work to conduct a detailed evaluation of two major distributed NeRF training methods and their tradeoffs: distributed data parallel (DDP) and spatial segmentation (SS). We find that DDP training requires cross-device synchronization during training, while SS training incurs additional fusion overhead during inference. Our analysis also reveals that sampling input images is a common key bottleneck in distributed NeRF training. At the beginning of each training iteration, the CPU generates input batches for all GPUs in the cluster by sampling all images in the dataset, causing significant stalls that constitute up to 43.3% of the total training time. To alleviate this bottleneck, we propose a pipelined input sampling strategy that precomputes input samples on the CPU concurrently with model training on the GPUs. Our evaluation demonstrates an average speedup in training time by$1.95\times($up to$2.24\times)$.
Adrian Zhao, Louis Zhang, Sankeerth Durvasula, Nilesh Jain, Selvakumar Panneer, Nandita Vijaykumar
ISPASS1