EDBT 2026 Demo / reviewers in the wild / expert
Derek Hower
dblp:62/5426 · also Derek R. Hower
· DBLP profile ↗
10ranked-venue papers
4as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Memory systems · 46% Processor architecture and microarchitecture · 26% Parallel and multicore computing · 13% | |
| Software engineering, system software, and programming languages
2 papers |
Concurrent programming · 100% |
Topics — the 20 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › memory consistency
memory consistency model |
0.5 | 3 | 2015 | HRF-Relaxed: Adapting HRF to the Complexities of Industrial Heterogeneous Memory Models · ACM Trans. Archit. Code Optim. 2015 Heterogeneous-race-free memory models · ASPLOS 2014 Calvin: Deterministic or not? Free will to choose · HPCA 2011 |
Processor architecture and microarchitecture › value prediction
load value prediction |
0.4 | 1 | 2019 | Efficient Load Value Prediction Using Multiple Predictors and Filters · HPCA 2019 |
Processor architecture and microarchitecture
value prediction |
0.4 | 1 | 2019 | Efficient Load Value Prediction Using Multiple Predictors and Filters · HPCA 2019 |
Memory systems › memory bandwidth management
bandwidth partitioning |
0.3 | 1 | 2017 | PABST: Proportionally Allocated Bandwidth at the Source and Target · HPCA 2017 |
Memory systems
memory bandwidth management |
0.3 | 1 | 2017 | PABST: Proportionally Allocated Bandwidth at the Source and Target · HPCA 2017 |
Memory systems
memory controller |
0.3 | 1 | 2017 | PABST: Proportionally Allocated Bandwidth at the Source and Target · HPCA 2017 |
Concurrent programming
memory models |
0.3 | 2 | 2015 | HRF-Relaxed: Adapting HRF to the Complexities of Industrial Heterogeneous Memory Models · ACM Trans. Archit. Code Optim. 2015 Heterogeneous-race-free memory models · ASPLOS 2014 |
Memory systems
cache coherence |
0.3 | 3 | 2014 | QuickRelease: A throughput-oriented approach to release consistency on GPUs · HPCA 2014 Calvin: Deterministic or not? Free will to choose · HPCA 2011 Rerun: Exploiting Episodes for Lightweight Memory Race Recording · ISCA 2008 |
Parallel and multicore computing › synchronization
fine-grain synchronization |
0.2 | 1 | 2014 | QuickRelease: A throughput-oriented approach to release consistency on GPUs · HPCA 2014 |
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy |
0.2 | 1 | 2014 | QuickRelease: A throughput-oriented approach to release consistency on GPUs · HPCA 2014 |
GPUs and heterogeneous computing › heterogeneous architecture
heterogeneous shared memory |
0.2 | 1 | 2014 | Heterogeneous-race-free memory models · ASPLOS 2014 |
Memory systems › memory consistency › memory consistency model
release consistency |
0.2 | 1 | 2014 | QuickRelease: A throughput-oriented approach to release consistency on GPUs · HPCA 2014 |
Parallel and multicore computing
deterministic execution |
0.1 | 1 | 2011 | Calvin: Deterministic or not? Free will to choose · HPCA 2011 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2011 | Calvin: Deterministic or not? Free will to choose · HPCA 2011 |
Processor architecture and microarchitecture
data dependence |
0.1 | 1 | 2019 | Efficient Load Value Prediction Using Multiple Predictors and Filters · HPCA 2019 |
Electronic design automation › hardware verification and test
debugging |
0.1 | 1 | 2008 | Rerun: Exploiting Episodes for Lightweight Memory Race Recording · ISCA 2008 |
Parallel and multicore computing › parallel computing › parallel program debugging
deterministic replay |
0.1 | 1 | 2008 | Rerun: Exploiting Episodes for Lightweight Memory Race Recording · ISCA 2008 |
Processor architecture and microarchitecture › multicore design
memory race recording |
0.1 | 1 | 2008 | Rerun: Exploiting Episodes for Lightweight Memory Race Recording · ISCA 2008 |
Processor architecture and microarchitecture
multiprocessor architecture |
0.1 | 1 | 2008 | Rerun: Exploiting Episodes for Lightweight Memory Race Recording · ISCA 2008 |
GPUs and heterogeneous computing
GPU architecture |
0.1 | 1 | 2014 | QuickRelease: A throughput-oriented approach to release consistency on GPUs · HPCA 2014 |
Methods — techniques the papers use, named apart from their topics
scope inclusion · 0.4formalization · 0.4predictor filtering · 0.4predictor combination · 0.4formal modeling · 0.4workload evaluation · 0.3hardware architecture design · 0.3cycle-level simulation · 0.2simulation · 0.1hardware implementation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Efficient Load Value Prediction Using Multiple Predictors and FiltersabstractValue prediction [1], [2] has the potential to break through the performance limitations imposed by true data dependencies. Aggressive value predictors can deliver significant performance improvements, but usually require large hardware budgets. While predicting values of all instruction types is possible, prior work has shown that predicting just load values is most effective with a modest hardware budget (e.g., 8KB of prediction state [3], [4]). However, with hardware budget constraints and high prediction accuracy requirements (99%), prior work has struggled to increase the fraction of predicted loads (coverage) beyond the low 30s. In this paper, we analyzed four state-of-the-art load value predictors, and found that they complement each other. Based on that finding, we evaluated a new composite predictor that combines all four component predictors. Our results show that the composite predictor, combined with several optimizations we proposed, improve the benefit of load value prediction by 54%-74% depending on the total predictor budget. Moreover, our composite predictor delivers more than twice the coverage of first championship value prediction winner predictor (EVES [4]), and substantially increases the delivered speedup by more than 50%. Rami Sheikh, Derek Hower |
HPCA | 2 |
| 2017 | PABST: Proportionally Allocated Bandwidth at the Source and TargetabstractHigher integration lowers total cost of ownership (TCO) in the data center by reducing equipment cost and lowering energy consumption. However, higher integration also makes it difficult to achieve guaranteed quality of service (QoS) for shared resources. Unlike many other resources, memory bandwidth cannot be finely controlled by software in existing systems. As a result, many systems running critical, bandwidth-sensitive applications remain underutilized to protect against bandwidth interference. In this paper, we propose a novel hardware architecture allowing practical, software-controlled partitioning of memory bandwidth. Proportionally Allocated Bandwidth at the Source and Target (PABST) precisely controls the bandwidth of applications by throttling request rates at the source and prioritizes requests at the target. We show that PABST is work conserving, such that excess bandwidth beyond the requested allocation will not go unused. For applications sensitive to memory latency, we pair PABST with a simple priority scheme at the memory controller. We show that when combined, the system is able to lower TCO by providing performance isolation across a wide range of workloads, even when co-located with memory-intensive background jobs. Derek Hower, Harold W. Cain, Carl A. Waldspurger |
HPCA | 1 |
| 2017 | Jenga: Efficient Fault Tolerance for Stacked DRAMabstractIn this paper, we introduce Jenga, a new scheme for protecting 3D DRAM, specifically high bandwidth memory (HBM), from failures in bits, rows, banks, channels, dies, and TSVs. By providing redundancy at the granularity of a cache block-rather than across blocks, as in the current state of the art-Jenga achieves greater error-free performance and lower error recovery latency. We show that Jenga's runtime is on average only 1.03x the runtime of our Baseline across a range of benchmarks. Additionally, for memory intensive benchmarks, Jenga is on average 1.11x faster than prior work. Georgios Mappouras, Alireza Vahid, A. Robert Calderbank, Derek Hower, Daniel J. Sorin |
ICCD | 4 |
| 2015 | HRF-Relaxed: Adapting HRF to the Complexities of Industrial Heterogeneous Memory ModelsabstractMemory consistency models, or memory models, allow both programmers and program language implementers to reason about concurrent accesses to one or more memory locations. Memory model specifications balance the often conflicting needs for precise semantics, implementation flexibility, and ease of understanding. Toward that end, popular programming languages like Java, C, and C++ have adopted memory models built on the conceptual foundation of Sequential Consistency for Data-Race-Free programs (SC for DRF). These SC for DRF languages were created with general-purpose homogeneous CPU systems in mind, and all assume a single, global memory address space. Such a uniform address space is usually power and performance prohibitive in heterogeneous Systems on Chips (SoCs), and for that reason most heterogeneous languages have adopted split address spaces and operations with nonglobal visibility. There have recently been two attempts to bridge the disconnect between the CPU-centric assumptions of the SC for DRF framework and the realities of heterogeneous SoC architectures. Hower et al. proposed a class of Heterogeneous-Race-Free (HRF) memory models that provide a foundation for understanding many of the issues in heterogeneous memory models. At the same time, the Khronos Group developed the OpenCL 2.0 memory model that builds on the C++ memory model. The OpenCL 2.0 model includes features not addressed by HRF: primarily support for relaxed atomics and a property referred to as scope inclusion. In this article, we generalize HRF to allow formalization of and reasoning about more complicated models using OpenCL 2.0 as a point of reference. With that generalization, we (1) make the OpenCL 2.0 memory model more accessible by introducing a platform for feature comparisons to other models, (2) consider a number of shortcomings in the current OpenCL 2.0 model, and (3) propose changes that could be adopted by future OpenCL 2.0 revisions or by other, related, models. Benedict R. Gaster, Derek Hower, Lee W. Howes |
ACM Trans. Archit. Code Optim. | 2 |
| 2014 | Heterogeneous-race-free memory modelsabstractCommodity heterogeneous systems (e.g., integrated CPUs and GPUs), now support a unified, shared memory address space for all components. Because the latency of global communication in a heterogeneous system can be prohibi-tively high, heterogeneous systems (unlike homogeneous CPU systems) provide synchronization mechanisms that only guarantee ordering among a subset of threads, which we call a scope. Unfortunately, the consequences and se-mantics of these scoped operations are not yet well under-stood. Without a formal and approachable model to reason about the behavior of these operations, we risk an array of portability and performance issues. Derek Hower, Blake A. Hechtman, Bradford M. Beckmann, Benedict R. Gaster, Mark D. Hill, Steven K. Reinhardt, David A. Wood 0001 |
ASPLOS | 1 |
| 2014 | QuickRelease: A throughput-oriented approach to release consistency on GPUsabstractGraphics processing units (GPUs) have specialized throughput-oriented memory systems optimized for streaming writes with scratchpad memories to capture locality explicitly. Expanding the utility of GPUs beyond graphics encourages designs that simplify programming (e.g., using caches instead of scratchpads) and better support irregular applications with finer-grain synchronization. Our hypothesis is that, like CPUs, GPUs will benefit from caches and coherence, but that CPU-style “read for ownership” (RFO) coherence is inappropriate to maintain support for regular streaming workloads. This paper proposes QuickRelease (QR), which improves on conventional GPU memory systems in two ways. First, QR uses a FIFO to enforce the partial order of writes so that synchronization operations can complete without frequent cache flushes. Thus, non-synchronizing threads in QR can re-use cached data even when other threads are performing synchronization. Second, QR partitions the resources required by reads and writes to reduce the penalty of writes on read performance. Simulation results across a wide variety of general-purpose GPU workloads show that QR achieves a 7% average performance improvement compared to a conventional GPU memory system. Furthermore, for emerging workloads with finer-grain synchronization, QR achieves up to 42% performance improvement compared to a conventional GPU memory system without the scalability challenges of RFO coherence. To this end, QR provides a throughput-oriented solution to provide fine-grain synchronization on GPUs. Blake A. Hechtman, Shuai Che, Derek Hower, Yingying Tian, Bradford M. Beckmann, Mark D. Hill, Steven K. Reinhardt, David A. Wood 0001 |
HPCA | 3 |
| 2013 | FreshCache: Statically and dynamically exploiting dataless waysabstractAbstract- Last level caches (LLCs) account for a substantial fraction of the area and power budget in many modern processors. Two recent trends — dwindling die yield that falls off sharply with larger chips and increasing static power — make a strong case for a fresh look at LLC design. Inclusive caches are particularly interesting because many commercially successful processors use inclusion to ease coherence at a cost of some data being stale or redundant. Prior works have demonstrated that LLC designs could be improved through static (at design time) or dynamic (at runtime) use of “dataless ways”. The static dataless ways removes the data—but not tags—from some cache ways to save energy and area without complicating inclusive-LLC coherence. A dynamic version (dynamic dataless ways) could dynamically turn off data, but not tags, effectively adapting the classic selective cache ways idea to save energy in LLC but not area. Our data show that (a) all our benchmarks benefit from dataless ways, but (b) the best number of dataless ways varies by workload. Thus, a pure static dataless design leaves energy-saving opportunity on the table, while a pure dynamic dataless design misses area-saving opportunity. To surpass both pure static and dynamic approaches, we develop the FreshCache LLC design that both statically and dynamically exploits dataless ways, including a predictor to adapt the number of dynamic dataless ways as well as detailed cache management policies. Results show that FreshCache saves more energy than static dataless ways alone (e.g., 72 % vs. 9 % of LLC) and more area by dynamic dataless ways only (e.g., 8 % vs. 0 % of LLC). I. Arkaprava Basu, Derek Hower, Mark D. Hill, Michael M. Swift |
ICCD | 2 |
| 2011 | Calvin: Deterministic or not? Free will to chooseabstractMost shared memory systems maximize performance by unpredictably resolving memory races. Unpredictable memory races can lead to nondeterminism in parallel programs, which can suffer from hard-to-reproduce hiesenbugs. We introduce Calvin, a shared memory model capable of executing in a conventional nondeterministic mode when performance is paramount and a deterministic mode when execution repeatability is important. Unlike prior hardware proposals for deterministic execution, Calvin exploits the flexibility of a memory consistency model weaker than sequential consistency. Specifically, Calvin logically orders memory operations into strata that are compatible with the Total Store Order (TSO). Calvin is also designed with the needs of future power-aware processors in mind, and does not require any speculation support. We develop a Calvin-MIST implementation that uses an unordered coalescing write cache, multiple-write coherence protocol, and delayed (timebomb) invalidations while maintaining TSO compatibility. Results show that Calvin-MIST can execute workloads in conventional mode at speeds comparable to a conventional system (providing compatibility) or execute deterministically for a modest average slowdown of less than 20% (when determinism is valued). Derek Hower, Polina Dudnik, Mark D. Hill, David A. Wood 0001 |
HPCA | 1 |
| 2008 | Rerun: Exploiting Episodes for Lightweight Memory Race RecordingabstractMultiprocessor deterministic replay has many potential uses in the era of multicore computing, including enhanced debugging, fault tolerance, and intrusion detection. While sources of nondeterminism in a uniprocessor can be recorded efficiently in software, it seems likely that hardware support will be needed in a multiprocessor environment where the outcome of memory races must also be recorded. We develop a memory race recording mechanism, called Rerun, that uses small hardware state (~166 bytes/core), writes a small race log (~4 bytes/kilo- instruction), and operates well as the number of cores per system scales (e.g., to 16 cores). Rerun exploits the dual of conventional wisdom in race recording: Rather than record information about individual memory accesses that conflict, we record how long a thread executes without conflicting with other threads. In particular, Rerun passively creates atomic episodes. Each episode is a dynamic instruction sequence that a thread happens to execute without interacting with other threads. Rerun uses Lamport Clocks to order episodes and enable replay of an equivalent execution. Derek Hower, Mark D. Hill |
ISCA | 1 |
| 2006 | Self-Checking and Self-Diagnosing 32-bit Microprocessor MultiplierabstractIn this paper, we propose a low-cost fault tolerance technique for microprocessor multipliers, both non-pipelined (NP) and pipelined (P). Our fault tolerant multiplier designs are capable of detecting and correcting errors, diagnosing hard faults, and reconfiguring to take the faulty sub-unit off-line. We utilize the branch misprediction recovery mechanism in the microprocessor core to take the error detection process off the critical path. Our analysis shows that our scheme provides 99% fault security and, compared to a baseline unprotected multiplier, achieves this fault tolerance with low performance overhead (5% for NP and 2.5% for P multiplier) and reasonably low area (38% NP and 26% P) and power consumption (36% NP and 28.5% P) overheads Mahmut Yilmaz, Derek Hower, Sule Ozev, Daniel J. Sorin |
ITC | 2 |