John Danskin

dblp:178/3195 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 61% GPUs and heterogeneous computing · 30% Processor architecture and microarchitecture · 9%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache coherence
0.212016
Selective GPU caches to eliminate CPU-GPU HW cache coherence · HPCA 2016
GPUs and heterogeneous computing › GPU memory management
GPU cache management
0.212016
Selective GPU caches to eliminate CPU-GPU HW cache coherence · HPCA 2016
Memory systems › cache management › cache insertion policy
selective caching
0.212016
Selective GPU caches to eliminate CPU-GPU HW cache coherence · HPCA 2016

Methods — techniques the papers use, named apart from their topics

variable-size transfer · 0.2request coalescing · 0.2page-level protection · 0.2
YearPublicationVenuePosition
2016 Pascal GPU with NVLink
abstract
Presents a collection of slides covering the following topics: Pascal GPU; NVLink; P100 SXM2 module; GP100 die; HBM; GPU performance; high bandwidth memory; Pascal unified memory; and CPU.
John Danskin, Denis Foley
Hot Chips Symposium1
2016 Selective GPU caches to eliminate CPU-GPU HW cache coherence
abstract
Cache coherence is ubiquitous in shared memory multiprocessors because it provides a simple, high performance memory abstraction to programmers. Recent work suggests extending hardware cache coherence between CPUs and GPUs to help support programming models with tightly coordinated sharing between CPU and GPU threads. However, implementing hardware cache coherence is particularly challenging in systems with discrete CPUs and GPUs that may not be produced by a single vendor. Instead, we propose, selective caching, wherein we disallow GPU caching of any memory that would require coherence updates to propagate between the CPU and GPU, thereby decoupling the GPU from vendor-specific CPU coherence protocols. We propose several architectural improvements to offset the performance penalty of selective caching: aggressive request coalescing, CPU-side coherent caching for GPU-uncacheable requests, and a CPU-GPU interconnect optimization to support variable-size transfers. Moreover, current GPU workloads access many read-only memory pages; we exploit this property to allow promiscuous GPU caching of these pages, relying on page-level protection, rather than hardware cache coherence, to ensure correctness. These optimizations bring a selective caching GPU implementation to within 93% of a hardware cache-coherent implementation without the need to integrate CPUs and GPUs under a single hardware coherence protocol.
David W. Nellans, Eiman Ebrahimi, Thomas F. Wenisch, John Danskin, Stephen W. Keckler
HPCA5