Sanya Srivastava

dblp:266/9182 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0003-8859-683XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Consistency and Coherence of the NVIDIA Grace-Hopper Superchip
abstract
Modern heterogeneous processors like the NVIDIA Grace-Hopper Superchip tightly integrate CPU and GPU cores across a cache-coherent interconnect, with an implicit assumption that independently compiled CPU and GPU code can safely interact via shared memory. Yet the memory consistency and coherence of such systems remain empirically unvalidated. This paper presents the first systematic study of consistency and coherence on the Grace-Hopper. We empirically validate that the system enforces the Compound Memory Consistency Model (CMCM)---a theoretical prerequisite for correct independent compilation---using a novel heterogeneous litmus testing methodology spanning 1,960 test variants. We further introduce Value Propagation tests to reverse-engineer the underlying coherence mechanisms, revealing that internal GPU coherence relies on write-throughs and self-invalidations rather than classical writer-initiated invalidations, while global CPU-GPU coherence is maintained via directory-based invalidations consistent with an AMBA CHI-like protocol. These results establish the CMCM as a concrete architectural target for heterogeneous systems and provide the first empirical characterization of GPU and CPU-GPU coherence mechanisms in a commercial heterogeneous processor.
Soham Bagchi, Sanya Srivastava, Reese Levine, Tyler Sorensen 0001, Ryan Stutsman, Vijay Nagarajan
ISMM2
2023 Degree-Aware Kernel Mapping for Graph Processing on GPUs
abstract
Social network graphs follow a power-law distribution, enabling us to exploit the GPU’s hierarchical execution model for efficient computation through a degree-aware kernel computation approach. In this approach, nodes are mapped to different levels of parallelism on the GPU, depending upon their in-degree. To take advantage of this execution model, the nodes must be arranged in decreasing order of in-degree. However, doing so distorts the community structure (a property responsible for providing memory locality during computation), impacting the degree-aware kernel’s performance gain. To balance ordering by in-degree and community structure, we propose DRBS, a graph reordering algorithm that sorts the graph nodes while preserving enough community structure to enhance cache efficiency.
Sanya Srivastava, Tyler Sorensen 0001
ISPASS1