VLDB 2026 Research / reviewers in the wild / expert
Erik Lindholm
dblp:33/4447
· DBLP profile ↗
3ranked-venue papers
1as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Processor architecture and microarchitecture · 53% GPUs and heterogeneous computing · 24% Energy-efficient computing · 13% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
register file |
0.3 | 2 | 2012 | A Hierarchical Thread Scheduler and Register File for Energy-Efficient Throughput Processors · ACM Trans. Comput. Syst. 2012 Energy-efficient mechanisms for managing thread context in throughput processors · ISCA 2011 |
GPUs and heterogeneous computing
GPU architecture |
0.2 | 2 | 2012 | A Hierarchical Thread Scheduler and Register File for Energy-Efficient Throughput Processors · ACM Trans. Comput. Syst. 2012 A user-programmable vertex engine · SIGGRAPH 2001 |
Processor architecture and microarchitecture › register file
hierarchical register file |
0.1 | 1 | 2012 | A Hierarchical Thread Scheduler and Register File for Energy-Efficient Throughput Processors · ACM Trans. Comput. Syst. 2012 |
Processor architecture and microarchitecture
throughput processor |
0.1 | 1 | 2012 | A Hierarchical Thread Scheduler and Register File for Energy-Efficient Throughput Processors · ACM Trans. Comput. Syst. 2012 |
GPUs and heterogeneous computing
GPU microarchitecture |
0.1 | 1 | 2011 | Energy-efficient mechanisms for managing thread context in throughput processors · ISCA 2011 |
Processor architecture and microarchitecture › register file
register caching |
0.1 | 1 | 2011 | Energy-efficient mechanisms for managing thread context in throughput processors · ISCA 2011 |
Parallel and multicore computing › parallel scheduling
thread scheduling |
0.1 | 1 | 2011 | Energy-efficient mechanisms for managing thread context in throughput processors · ISCA 2011 |
Energy-efficient computing › energy-efficient architecture
GPU energy efficiency |
0.0 | 1 | 2012 | A Hierarchical Thread Scheduler and Register File for Energy-Efficient Throughput Processors · ACM Trans. Comput. Syst. 2012 |
Energy-efficient computing › energy-efficient architecture
processor energy efficiency |
0.0 | 1 | 2012 | A Hierarchical Thread Scheduler and Register File for Energy-Efficient Throughput Processors · ACM Trans. Comput. Syst. 2012 |
Energy-efficient computing › low-power design
low-power processor design |
0.0 | 1 | 2011 | Energy-efficient mechanisms for managing thread context in throughput processors · ISCA 2011 |
Energy-efficient computing
power management |
0.0 | 1 | 2011 | Energy-efficient mechanisms for managing thread context in throughput processors · ISCA 2011 |
GPUs and heterogeneous computing › GPU rendering
programmable graphics pipeline |
0.0 | 1 | 2001 | A user-programmable vertex engine · SIGGRAPH 2001 |
Methods — techniques the papers use, named apart from their topics
execution-driven simulation · 0.1multithreading · 0.0fixed-function pipeline · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | A Hierarchical Thread Scheduler and Register File for Energy-Efficient Throughput ProcessorsabstractModern graphics processing units (GPUs) employ a large number of hardware threads to hide both function unit and memory access latency. Extreme multithreading requires a complex thread scheduler as well as a large register file, which is expensive to access both in terms of energy and latency. We present two complementary techniques for reducing energy on massively-threaded processors such as GPUs. First, we investigate a two-level thread scheduler that maintains a small set of active threads to hide ALU and local memory access latency and a larger set of pending threads to hide main memory latency. Reducing the number of threads that the scheduler must consider each cycle improves the scheduler’s energy efficiency. Second, we propose replacing the monolithic register file found on modern designs with a hierarchical register file. We explore various trade-offs for the hierarchy including the number of levels in the hierarchy and the number of entries at each level. We consider both a hardware-managed caching scheme and a software-managed scheme, where the compiler is responsible for orchestrating all data movement within the register file hierarchy. Combined with a hierarchical register file, our two-level thread scheduler provides a further reduction in energy by only allocating entries in the upper levels of the register file hierarchy for active threads. Averaging across a variety of real world graphics and compute workloads, the active thread count can be reduced by a factor of 4 with minimal impact on performance and our most efficient three-level software-managed register file hierarchy reduces register file energy by 54%. Mark Gebhart, Daniel R. Johnson, David Tarjan, Stephen W. Keckler, William J. Dally, Erik Lindholm, Kevin Skadron |
ACM Trans. Comput. Syst. | 6 |
| 2011 | Energy-efficient mechanisms for managing thread context in throughput processorsabstractModern graphics processing units (GPUs) use a large number of hardware threads to hide both function unit and memory access latency. Extreme multithreading requires a complicated thread scheduler as well as a large register file, which is expensive to access both in terms of energy and latency. We present two complementary techniques for reducing energy on massively-threaded processors such as GPUs. First, we examine register file caching to replace accesses to the large main register file with accesses to a smaller structure containing the immediate register working set of active threads. Second, we investigate a two-level thread scheduler that maintains a small set of active threads to hide ALU and local memory access latency and a larger set of pending threads to hide main memory latency. Combined with register file caching, a two-level thread scheduler provides a further reduction in energy by limiting the allocation of temporary register cache resources to only the currently active subset of threads. We show that on average, across a variety of real world graphics and compute workloads, a 6-entry per-thread register file cache reduces the number of reads and writes to the main register file by 50% and 59% respectively. We further show that the active thread count can be reduced by a factor of 4 with minimal impact on performance, resulting in a 36% reduction of register file energy. Mark Gebhart, Daniel R. Johnson, David Tarjan, Stephen W. Keckler, William J. Dally, Erik Lindholm, Kevin Skadron |
ISCA | 6 |
| 2001 | A user-programmable vertex engineabstractIn this paper we describe the design, programming interface, and implementation of a very efficient user-programmable vertex engine. The vertex engine of NVIDIA's GeForce3 GPU evolved from a highly tuned fixed-function pipeline requiring considerable knowledge to program. Programs operate only on a stream of independent vertices traversing the pipe. Embedded in the broader fixed function pipeline, our approach preserves parallelism sacrificed by previous approaches. The programmer is presented with a straightforward programming model, which is supported by transparent multi-threading and bypassing to preserve parallelism and performance. Erik Lindholm, Mark J. Kligard, Henry P. Moreton |
SIGGRAPH | 1 |