EDBT 2026 Demo / reviewers in the wild / expert
Mick Wollman
dblp:173/9813
· DBLP profile ↗
1ranked-venue papers
0as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 62% GPUs and heterogeneous computing · 38% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy |
0.2 | 1 | 2015 | WarpPool: sharing requests with inter-warp coalescing for throughput processors · MICRO 2015 |
Memory systems › memory hierarchy › cache hierarchy
l1 cache |
0.2 | 1 | 2015 | WarpPool: sharing requests with inter-warp coalescing for throughput processors · MICRO 2015 |
Memory systems
cache |
0.1 | 1 | 2015 | WarpPool: sharing requests with inter-warp coalescing for throughput processors · MICRO 2015 |
Memory systems › cache › cache performance
cache thrashing |
0.1 | 1 | 2015 | WarpPool: sharing requests with inter-warp coalescing for throughput processors · MICRO 2015 |
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | WarpPool: sharing requests with inter-warp coalescing for throughput processorsabstractAlthough graphics processing units (GPUs) are capable of high compute throughput, their memory systems need to supply the arithmetic pipelines with data at a sufficient rate to avoid stalls. For benchmarks that have divergent access patterns or cause the L1 cache to run out of resources, the link between the GPU's load/store unit and the L1 cache becomes a bottleneck in the memory system, leading to low utilization of compute resources. While current GPU memory systems are able to coalesce requests between threads in the same warp, we identify a form of spatial locality between threads in multiple warps. We use this locality, which is overlooked in current systems, to merge requests being sent to the L1 cache. This relieves the bottleneck between the load/store unit and the cache, and provides an opportunity to prioritize requests to minimize cache thrashing. Our implementation, WarpPool, yields a 38% speedup on memory throughput-limited kernels by increasing the throughput to the L1 by 8% and the reducing the number of L1 misses by 23%. We also demonstrate that WarpPool can improve GPU programmability by achieving high performance without the need to optimize workloads' memory access patterns. A Verilog implementation including place-and route shows WarpPool requires 1.0% added GPU area and 0.8% added power. John Kloosterman, Jonathan Beaumont, Mick Wollman, Ankit Sethia, Ronald G. Dreslinski, Trevor N. Mudge, Scott A. Mahlke |
MICRO | 3 |