EDBT 2026 Demo / reviewers in the wild / expert
John Kloosterman
dblp:173/9796
· DBLP profile ↗
4ranked-venue papers
2as first author
0since 2021 · last 2020
0000-0001-8180-1237ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-authorSecurity and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
GPUs and heterogeneous computing · 42% Processor architecture and microarchitecture · 31% Memory systems · 18% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
register file |
0.6 | 2 | 2018 | Scratch That (But Cache This): A Hybrid Register Cache/Scratchpad for GPUs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 Regless: just-in-time operand staging for GPUs · MICRO 2017 |
GPUs and heterogeneous computing
GPU microarchitecture |
0.3 | 1 | 2018 | Scratch That (But Cache This): A Hybrid Register Cache/Scratchpad for GPUs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
GPUs and heterogeneous computing
GPU architecture |
0.3 | 1 | 2017 | Regless: just-in-time operand staging for GPUs · MICRO 2017 |
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy |
0.2 | 1 | 2015 | WarpPool: sharing requests with inter-warp coalescing for throughput processors · MICRO 2015 |
Memory systems › memory hierarchy › cache hierarchy
l1 cache |
0.2 | 1 | 2015 | WarpPool: sharing requests with inter-warp coalescing for throughput processors · MICRO 2015 |
Energy-efficient computing › energy-efficient architecture
GPU energy efficiency |
0.1 | 1 | 2017 | Regless: just-in-time operand staging for GPUs · MICRO 2017 |
Energy-efficient computing
power management |
0.1 | 1 | 2017 | Regless: just-in-time operand staging for GPUs · MICRO 2017 |
Memory systems
cache |
0.1 | 1 | 2015 | WarpPool: sharing requests with inter-warp coalescing for throughput processors · MICRO 2015 |
Memory systems › cache › cache performance
cache thrashing |
0.1 | 1 | 2015 | WarpPool: sharing requests with inter-warp coalescing for throughput processors · MICRO 2015 |
Methods — techniques the papers use, named apart from their topics
register caching · 0.3compiler-managed scratchpad · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | AVGuardian: Detecting and Mitigating Publish-Subscribe Overprivilege for Autonomous Vehicle SystemsabstractAutonomous vehicle (AV) software systems are emerging to enable rapidly developed self-driving functionalities. Since such systems are responsible for safety-critical decisions, it is necessary to secure them in face of cyber attacks. Through an empirical study of representative AV software systems Baidu Apollo and Autoware, we discover a common over privilege problem with the publish-subscribe communication model widely adopted by AV systems: due to the coarse-grained message design for the publish-subscribe communication, some message fields are over-granted with publish/subscribe permissions. To comply with the least-privilege principle and reduce the attack surface resulting from such problem, we argue that the publish/subscribe permissions should be defined and enforced at the granularity of message fields instead of messages. To systematically address such publish-subscribe over-privilege problems, we present AVGuardian, a system that includes (1) a static analysis tool that detects overprivilege instances in AV software and generates the corresponding access control policies at the message field granularity, and (2) a low-overhead, module-transparent, runtime pub-lish/subscribe permission policy enforcement mechanism to perform online policy violation detection and prevention. Using our detection tool, we are able to automatically detect 581 overprivilege instances in total in Baidu Apollo. To demonstrate the severity, we further constructed several concrete exploits that can lead to vehicle collision and identity theft for AV owners, which have been reported to Baidu Apollo and confirmed as valid. For defense, we prototype and evaluate the policy enforcement mechanism, and find that it has very low overhead, does not affect original AV decision logic, and also is resilient to message replay attacks. David Ke Hong, John Kloosterman, Yuqi Jin, Qi Alfred Chen, Scott A. Mahlke, Z. Morley Mao |
EuroS&P | 2 |
| 2018 | Scratch That (But Cache This): A Hybrid Register Cache/Scratchpad for GPUsabstractGraphics processing units (GPUs) are throughput-oriented architectures that implement massive multithreading. Large, power-hungry register files are required in GPUs to support the simultaneous execution of thousands of threads on the hardware. Prior work proposed reducing register access energy by adding a small register cache (RC) to the GPU. The cache stores recently referenced registers and services subsequent accesses to these registers, reducing accesses to the main register file. Later work obtained further energy savings by replacing this cache with a compiler-managed scratchpad. We note that registers are allocated to the cache dynamically and reactively whereas registers are allocated to the scratchpad statically and proactively. Our insight is that these allocation schemes are complimentary because the cache leverages runtime information unavailable to the compiler and the scratchpad leverages compile time information unavailable to the cache. Further, there exist register access patterns that are easily captured by one structure but for which the other structure is ineffective. Instead of implementing either an RC or scratchpad alone, we propose dividing temporary register storage capacity between a cache and a scratchpad in order to capture a broader range of register accesses. Given 12 KB of storage per streaming multiprocessor, our hybrid design reduces register energy to 38.7% of the baseline, compared to 47.9% for a RC and 47.1% for a register scratchpad. Jonathan Bailey, John Kloosterman, Scott A. Mahlke |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | Regless: just-in-time operand staging for GPUsabstractThe register file is one of the largest and most power-hungry structures in a Graphics Processing Unit (GPU), because massive multithreading requires all the register state for every active thread to be available. Previous approaches to making register accesses more efficient have optimized how registers are stored, but they must keep all values for active threads in a large, high-bandwidth structure. If operand storage is to be reduced further, there will not be enough capacity for every live value to be stored at the same time. Our insight is that computation graphs can be sliced into regions and operand storage can be allocated to these regions as they are encountered at run time, allowing a small operand staging unit to replace the register file. Most operand values have a short lifetime that is contained in one region, so their value does not need to persist in the staging unit past the end of that region. The small number of longer-lived operands can be stored in lower-bandwidth global memory, but the hardware must anticipate their use to fetch them early enough to avoid stalls. In RegLess, hardware uses compiler annotations to anticipate warps' operand usage at run time, allowing the register file to be replaced with an operand staging unit 25% of the size, saving 75% of register file energy and 11% of total GPU energy with no average performance loss. John Kloosterman, Jonathan Beaumont, Davoud Anoushe Jamshidi, Jonathan Bailey, Trevor N. Mudge, Scott A. Mahlke |
MICRO | 1 |
| 2015 | WarpPool: sharing requests with inter-warp coalescing for throughput processorsabstractAlthough graphics processing units (GPUs) are capable of high compute throughput, their memory systems need to supply the arithmetic pipelines with data at a sufficient rate to avoid stalls. For benchmarks that have divergent access patterns or cause the L1 cache to run out of resources, the link between the GPU's load/store unit and the L1 cache becomes a bottleneck in the memory system, leading to low utilization of compute resources. While current GPU memory systems are able to coalesce requests between threads in the same warp, we identify a form of spatial locality between threads in multiple warps. We use this locality, which is overlooked in current systems, to merge requests being sent to the L1 cache. This relieves the bottleneck between the load/store unit and the cache, and provides an opportunity to prioritize requests to minimize cache thrashing. Our implementation, WarpPool, yields a 38% speedup on memory throughput-limited kernels by increasing the throughput to the L1 by 8% and the reducing the number of L1 misses by 23%. We also demonstrate that WarpPool can improve GPU programmability by achieving high performance without the need to optimize workloads' memory access patterns. A Verilog implementation including place-and route shows WarpPool requires 1.0% added GPU area and 0.8% added power. John Kloosterman, Jonathan Beaumont, Mick Wollman, Ankit Sethia, Ronald G. Dreslinski, Trevor N. Mudge, Scott A. Mahlke |
MICRO | 1 |