Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Rishkul Kulkami

dblp:320/8147 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 60% Processor architecture and microarchitecture · 40%
Computer graphics and multimedia
1 paper
Rendering · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
control flow divergence
0.612022
GPU Subwarp Interleaving · HPCA 2022
GPUs and heterogeneous computing
GPU microarchitecture
0.612022
GPU Subwarp Interleaving · HPCA 2022
Processor architecture and microarchitecture
latency hiding
0.612022
GPU Subwarp Interleaving · HPCA 2022
Processor architecture and microarchitecture › pipelining
pipeline stall
0.612022
GPU Subwarp Interleaving · HPCA 2022
GPUs and heterogeneous computing › GPU scheduling
warp scheduling
0.612022
GPU Subwarp Interleaving · HPCA 2022
Rendering
ray tracing
0.212022
GPU Subwarp Interleaving · HPCA 2022

Methods — techniques the papers use, named apart from their topics

simulation · 1.1microbenchmarking · 1.1
YearPublicationVenuePosition
2022 GPU Subwarp Interleaving
abstract
Raytracing applications have naturally high thread divergence, low warp occupancy and are limited by memory latency. In this paper, we present an architectural enhancement called Subwarp Interleaving that exploits thread divergence to hide pipeline stalls in divergent sections of low warp occupancy workloads. Subwarp Interleaving allows for fine-grained interleaved execution of diverged paths within a warp with the goal of increasing hardware utilization and reducing warp latency. However, notwithstanding the promise shown by early microbenchmark studies and an average performance upside of 6.3% (up to 20%) on a simulator across a suite of raytradng application traces, the Subwarp Interleaving design feature has shortcomings that preclude its near-term implementation. This paper introduces the reader to the challenges of raytradng and discusses a novel micro-architectural approach that, on paper, addresses many of the challenges. A thorough analysis of the idea on a production simulator reveals that the high-level motivating statistics are optimistic, and second-order effects, along with other architectural sharp edges, limit the idea’s potential. We identify Subwarp Interleaving’s primary limiters for an NVIDIA Tbring-like architecture, and we outline the conditions under which the approach could be more effective.
Sana Damani, Mark Stephenson, Ram Rangan, Daniel R. Johnson, Rishkul Kulkami, Stephen W. Keckler
HPCA5