Stephen W. Redder

dblp:160/0634 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 38% Memory systems · 38% Parallel and multicore computing · 23%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache
0.212015
Priority-based cache allocation in throughput processors · HPCA 2015
GPUs and heterogeneous computing › GPU memory management
GPU cache management
0.212015
Priority-based cache allocation in throughput processors · HPCA 2015
Parallel and multicore computing › parallel scheduling
thread scheduling
0.112015
Priority-based cache allocation in throughput processors · HPCA 2015
Parallel and multicore computing › parallel programming runtimes › thread management
thread throttling
0.112015
Priority-based cache allocation in throughput processors · HPCA 2015

Methods — techniques the papers use, named apart from their topics

thread scheduling · 0.2priority-based allocation · 0.2
YearPublicationVenuePosition
2015 Priority-based cache allocation in throughput processors
abstract
GPUs employ massive multithreading and fast context switching to provide high throughput and hide memory latency. Multithreading can Increase contention for various system resources, however, that may result In suboptimal utilization of shared resources. Previous research has proposed variants of throttling thread-level parallelism to reduce cache contention and improve performance. Throttling approaches can, however, lead to under-utilizing thread contexts, on-chip interconnect, and off-chip memory bandwidth. This paper proposes to tightly couple the thread scheduling mechanism with the cache management algorithms such that GPU cache pollution is minimized while off-chip memory throughput is enhanced. We propose priority-based cache allocation (PCAL) that provides preferential cache capacity to a subset of high-priority threads while simultaneously allowing lower priority threads to execute without contending for the cache. By tuning thread-level parallelism while both optimizing caching efficiency as well as other shared resource usage, PCAL builds upon previous thread throttling approaches, improving overall performance by an average 17% with maximum 51%.
Minsoo Rhu, Daniel R. Johnson, Mike O'Connor, Mattan Erez, Doug Burger, Donald S. Fussell, Stephen W. Redder
HPCA8