Mick Wollman

dblp:173/9813 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 62% GPUs and heterogeneous computing · 38%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy
0.212015
WarpPool: sharing requests with inter-warp coalescing for throughput processors · MICRO 2015
Memory systems › memory hierarchy › cache hierarchy
l1 cache
0.212015
WarpPool: sharing requests with inter-warp coalescing for throughput processors · MICRO 2015
Memory systems
cache
0.112015
WarpPool: sharing requests with inter-warp coalescing for throughput processors · MICRO 2015
Memory systems › cache › cache performance
cache thrashing
0.112015
WarpPool: sharing requests with inter-warp coalescing for throughput processors · MICRO 2015
YearPublicationVenuePosition
2015 WarpPool: sharing requests with inter-warp coalescing for throughput processors
abstract
Although graphics processing units (GPUs) are capable of high compute throughput, their memory systems need to supply the arithmetic pipelines with data at a sufficient rate to avoid stalls. For benchmarks that have divergent access patterns or cause the L1 cache to run out of resources, the link between the GPU's load/store unit and the L1 cache becomes a bottleneck in the memory system, leading to low utilization of compute resources. While current GPU memory systems are able to coalesce requests between threads in the same warp, we identify a form of spatial locality between threads in multiple warps. We use this locality, which is overlooked in current systems, to merge requests being sent to the L1 cache. This relieves the bottleneck between the load/store unit and the cache, and provides an opportunity to prioritize requests to minimize cache thrashing. Our implementation, WarpPool, yields a 38% speedup on memory throughput-limited kernels by increasing the throughput to the L1 by 8% and the reducing the number of L1 misses by 23%. We also demonstrate that WarpPool can improve GPU programmability by achieving high performance without the need to optimize workloads' memory access patterns. A Verilog implementation including place-and route shows WarpPool requires 1.0% added GPU area and 0.8% added power.
John Kloosterman, Jonathan Beaumont, Mick Wollman, Ankit Sethia, Ronald G. Dreslinski, Trevor N. Mudge, Scott A. Mahlke
MICRO3