Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Derek Hower

dblp:62/5426 · also Derek R. Hower · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 4 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Memory systems · 46% Processor architecture and microarchitecture · 26% Parallel and multicore computing · 13%
Software engineering, system software, and programming languages
2 papers
Concurrent programming · 100%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory consistency
memory consistency model
0.532015
HRF-Relaxed: Adapting HRF to the Complexities of Industrial Heterogeneous Memory Models · ACM Trans. Archit. Code Optim. 2015
Heterogeneous-race-free memory models · ASPLOS 2014
Calvin: Deterministic or not? Free will to choose · HPCA 2011
Processor architecture and microarchitecture › value prediction
load value prediction
0.412019
Efficient Load Value Prediction Using Multiple Predictors and Filters · HPCA 2019
Processor architecture and microarchitecture
value prediction
0.412019
Efficient Load Value Prediction Using Multiple Predictors and Filters · HPCA 2019
Memory systems › memory bandwidth management
bandwidth partitioning
0.312017
PABST: Proportionally Allocated Bandwidth at the Source and Target · HPCA 2017
Memory systems
memory bandwidth management
0.312017
PABST: Proportionally Allocated Bandwidth at the Source and Target · HPCA 2017
Memory systems
memory controller
0.312017
PABST: Proportionally Allocated Bandwidth at the Source and Target · HPCA 2017
Concurrent programming
memory models
0.322015
HRF-Relaxed: Adapting HRF to the Complexities of Industrial Heterogeneous Memory Models · ACM Trans. Archit. Code Optim. 2015
Heterogeneous-race-free memory models · ASPLOS 2014
Memory systems
cache coherence
0.332014
QuickRelease: A throughput-oriented approach to release consistency on GPUs · HPCA 2014
Calvin: Deterministic or not? Free will to choose · HPCA 2011
Rerun: Exploiting Episodes for Lightweight Memory Race Recording · ISCA 2008
Parallel and multicore computing › synchronization
fine-grain synchronization
0.212014
QuickRelease: A throughput-oriented approach to release consistency on GPUs · HPCA 2014
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy
0.212014
QuickRelease: A throughput-oriented approach to release consistency on GPUs · HPCA 2014
GPUs and heterogeneous computing › heterogeneous architecture
heterogeneous shared memory
0.212014
Heterogeneous-race-free memory models · ASPLOS 2014
Memory systems › memory consistency › memory consistency model
release consistency
0.212014
QuickRelease: A throughput-oriented approach to release consistency on GPUs · HPCA 2014
Parallel and multicore computing
deterministic execution
0.112011
Calvin: Deterministic or not? Free will to choose · HPCA 2011
Parallel and multicore computing
parallel programming models
0.112011
Calvin: Deterministic or not? Free will to choose · HPCA 2011
Processor architecture and microarchitecture
data dependence
0.112019
Efficient Load Value Prediction Using Multiple Predictors and Filters · HPCA 2019
Electronic design automation › hardware verification and test
debugging
0.112008
Rerun: Exploiting Episodes for Lightweight Memory Race Recording · ISCA 2008
Parallel and multicore computing › parallel computing › parallel program debugging
deterministic replay
0.112008
Rerun: Exploiting Episodes for Lightweight Memory Race Recording · ISCA 2008
Processor architecture and microarchitecture › multicore design
memory race recording
0.112008
Rerun: Exploiting Episodes for Lightweight Memory Race Recording · ISCA 2008
Processor architecture and microarchitecture
multiprocessor architecture
0.112008
Rerun: Exploiting Episodes for Lightweight Memory Race Recording · ISCA 2008
GPUs and heterogeneous computing
GPU architecture
0.112014
QuickRelease: A throughput-oriented approach to release consistency on GPUs · HPCA 2014

Methods — techniques the papers use, named apart from their topics

scope inclusion · 0.4formalization · 0.4predictor filtering · 0.4predictor combination · 0.4formal modeling · 0.4workload evaluation · 0.3hardware architecture design · 0.3cycle-level simulation · 0.2simulation · 0.1hardware implementation · 0.1
YearPublicationVenuePosition
2019 Efficient Load Value Prediction Using Multiple Predictors and Filters
abstract
Value prediction [1], [2] has the potential to break through the performance limitations imposed by true data dependencies. Aggressive value predictors can deliver significant performance improvements, but usually require large hardware budgets. While predicting values of all instruction types is possible, prior work has shown that predicting just load values is most effective with a modest hardware budget (e.g., 8KB of prediction state [3], [4]). However, with hardware budget constraints and high prediction accuracy requirements (99%), prior work has struggled to increase the fraction of predicted loads (coverage) beyond the low 30s. In this paper, we analyzed four state-of-the-art load value predictors, and found that they complement each other. Based on that finding, we evaluated a new composite predictor that combines all four component predictors. Our results show that the composite predictor, combined with several optimizations we proposed, improve the benefit of load value prediction by 54%-74% depending on the total predictor budget. Moreover, our composite predictor delivers more than twice the coverage of first championship value prediction winner predictor (EVES [4]), and substantially increases the delivered speedup by more than 50%.
Rami Sheikh, Derek Hower
HPCA2
2017 PABST: Proportionally Allocated Bandwidth at the Source and Target
abstract
Higher integration lowers total cost of ownership (TCO) in the data center by reducing equipment cost and lowering energy consumption. However, higher integration also makes it difficult to achieve guaranteed quality of service (QoS) for shared resources. Unlike many other resources, memory bandwidth cannot be finely controlled by software in existing systems. As a result, many systems running critical, bandwidth-sensitive applications remain underutilized to protect against bandwidth interference. In this paper, we propose a novel hardware architecture allowing practical, software-controlled partitioning of memory bandwidth. Proportionally Allocated Bandwidth at the Source and Target (PABST) precisely controls the bandwidth of applications by throttling request rates at the source and prioritizes requests at the target. We show that PABST is work conserving, such that excess bandwidth beyond the requested allocation will not go unused. For applications sensitive to memory latency, we pair PABST with a simple priority scheme at the memory controller. We show that when combined, the system is able to lower TCO by providing performance isolation across a wide range of workloads, even when co-located with memory-intensive background jobs.
Derek Hower, Harold W. Cain, Carl A. Waldspurger
HPCA1
2017 Jenga: Efficient Fault Tolerance for Stacked DRAM
abstract
In this paper, we introduce Jenga, a new scheme for protecting 3D DRAM, specifically high bandwidth memory (HBM), from failures in bits, rows, banks, channels, dies, and TSVs. By providing redundancy at the granularity of a cache block-rather than across blocks, as in the current state of the art-Jenga achieves greater error-free performance and lower error recovery latency. We show that Jenga's runtime is on average only 1.03x the runtime of our Baseline across a range of benchmarks. Additionally, for memory intensive benchmarks, Jenga is on average 1.11x faster than prior work.
Georgios Mappouras, Alireza Vahid, A. Robert Calderbank, Derek Hower, Daniel J. Sorin
ICCD4
2015 HRF-Relaxed: Adapting HRF to the Complexities of Industrial Heterogeneous Memory Models
abstract
Memory consistency models, or memory models, allow both programmers and program language implementers to reason about concurrent accesses to one or more memory locations. Memory model specifications balance the often conflicting needs for precise semantics, implementation flexibility, and ease of understanding. Toward that end, popular programming languages like Java, C, and C++ have adopted memory models built on the conceptual foundation of Sequential Consistency for Data-Race-Free programs (SC for DRF). These SC for DRF languages were created with general-purpose homogeneous CPU systems in mind, and all assume a single, global memory address space. Such a uniform address space is usually power and performance prohibitive in heterogeneous Systems on Chips (SoCs), and for that reason most heterogeneous languages have adopted split address spaces and operations with nonglobal visibility. There have recently been two attempts to bridge the disconnect between the CPU-centric assumptions of the SC for DRF framework and the realities of heterogeneous SoC architectures. Hower et al. proposed a class of Heterogeneous-Race-Free (HRF) memory models that provide a foundation for understanding many of the issues in heterogeneous memory models. At the same time, the Khronos Group developed the OpenCL 2.0 memory model that builds on the C++ memory model. The OpenCL 2.0 model includes features not addressed by HRF: primarily support for relaxed atomics and a property referred to as scope inclusion. In this article, we generalize HRF to allow formalization of and reasoning about more complicated models using OpenCL 2.0 as a point of reference. With that generalization, we (1) make the OpenCL 2.0 memory model more accessible by introducing a platform for feature comparisons to other models, (2) consider a number of shortcomings in the current OpenCL 2.0 model, and (3) propose changes that could be adopted by future OpenCL 2.0 revisions or by other, related, models.
Benedict R. Gaster, Derek Hower, Lee W. Howes
ACM Trans. Archit. Code Optim.2
2014 Heterogeneous-race-free memory models
abstract
Commodity heterogeneous systems (e.g., integrated CPUs and GPUs), now support a unified, shared memory address space for all components. Because the latency of global communication in a heterogeneous system can be prohibi-tively high, heterogeneous systems (unlike homogeneous CPU systems) provide synchronization mechanisms that only guarantee ordering among a subset of threads, which we call a scope. Unfortunately, the consequences and se-mantics of these scoped operations are not yet well under-stood. Without a formal and approachable model to reason about the behavior of these operations, we risk an array of portability and performance issues.
Derek Hower, Blake A. Hechtman, Bradford M. Beckmann, Benedict R. Gaster, Mark D. Hill, Steven K. Reinhardt, David A. Wood 0001
ASPLOS1
2014 QuickRelease: A throughput-oriented approach to release consistency on GPUs
abstract
Graphics processing units (GPUs) have specialized throughput-oriented memory systems optimized for streaming writes with scratchpad memories to capture locality explicitly. Expanding the utility of GPUs beyond graphics encourages designs that simplify programming (e.g., using caches instead of scratchpads) and better support irregular applications with finer-grain synchronization. Our hypothesis is that, like CPUs, GPUs will benefit from caches and coherence, but that CPU-style “read for ownership” (RFO) coherence is inappropriate to maintain support for regular streaming workloads. This paper proposes QuickRelease (QR), which improves on conventional GPU memory systems in two ways. First, QR uses a FIFO to enforce the partial order of writes so that synchronization operations can complete without frequent cache flushes. Thus, non-synchronizing threads in QR can re-use cached data even when other threads are performing synchronization. Second, QR partitions the resources required by reads and writes to reduce the penalty of writes on read performance. Simulation results across a wide variety of general-purpose GPU workloads show that QR achieves a 7% average performance improvement compared to a conventional GPU memory system. Furthermore, for emerging workloads with finer-grain synchronization, QR achieves up to 42% performance improvement compared to a conventional GPU memory system without the scalability challenges of RFO coherence. To this end, QR provides a throughput-oriented solution to provide fine-grain synchronization on GPUs.
Blake A. Hechtman, Shuai Che, Derek Hower, Yingying Tian, Bradford M. Beckmann, Mark D. Hill, Steven K. Reinhardt, David A. Wood 0001
HPCA3
2013 FreshCache: Statically and dynamically exploiting dataless ways
abstract
Abstract- Last level caches (LLCs) account for a substantial fraction of the area and power budget in many modern processors. Two recent trends — dwindling die yield that falls off sharply with larger chips and increasing static power — make a strong case for a fresh look at LLC design. Inclusive caches are particularly interesting because many commercially successful processors use inclusion to ease coherence at a cost of some data being stale or redundant. Prior works have demonstrated that LLC designs could be improved through static (at design time) or dynamic (at runtime) use of “dataless ways”. The static dataless ways removes the data—but not tags—from some cache ways to save energy and area without complicating inclusive-LLC coherence. A dynamic version (dynamic dataless ways) could dynamically turn off data, but not tags, effectively adapting the classic selective cache ways idea to save energy in LLC but not area. Our data show that (a) all our benchmarks benefit from dataless ways, but (b) the best number of dataless ways varies by workload. Thus, a pure static dataless design leaves energy-saving opportunity on the table, while a pure dynamic dataless design misses area-saving opportunity. To surpass both pure static and dynamic approaches, we develop the FreshCache LLC design that both statically and dynamically exploits dataless ways, including a predictor to adapt the number of dynamic dataless ways as well as detailed cache management policies. Results show that FreshCache saves more energy than static dataless ways alone (e.g., 72 % vs. 9 % of LLC) and more area by dynamic dataless ways only (e.g., 8 % vs. 0 % of LLC). I.
Arkaprava Basu, Derek Hower, Mark D. Hill, Michael M. Swift
ICCD2
2011 Calvin: Deterministic or not? Free will to choose
abstract
Most shared memory systems maximize performance by unpredictably resolving memory races. Unpredictable memory races can lead to nondeterminism in parallel programs, which can suffer from hard-to-reproduce hiesenbugs. We introduce Calvin, a shared memory model capable of executing in a conventional nondeterministic mode when performance is paramount and a deterministic mode when execution repeatability is important. Unlike prior hardware proposals for deterministic execution, Calvin exploits the flexibility of a memory consistency model weaker than sequential consistency. Specifically, Calvin logically orders memory operations into strata that are compatible with the Total Store Order (TSO). Calvin is also designed with the needs of future power-aware processors in mind, and does not require any speculation support. We develop a Calvin-MIST implementation that uses an unordered coalescing write cache, multiple-write coherence protocol, and delayed (timebomb) invalidations while maintaining TSO compatibility. Results show that Calvin-MIST can execute workloads in conventional mode at speeds comparable to a conventional system (providing compatibility) or execute deterministically for a modest average slowdown of less than 20% (when determinism is valued).
Derek Hower, Polina Dudnik, Mark D. Hill, David A. Wood 0001
HPCA1
2008 Rerun: Exploiting Episodes for Lightweight Memory Race Recording
abstract
Multiprocessor deterministic replay has many potential uses in the era of multicore computing, including enhanced debugging, fault tolerance, and intrusion detection. While sources of nondeterminism in a uniprocessor can be recorded efficiently in software, it seems likely that hardware support will be needed in a multiprocessor environment where the outcome of memory races must also be recorded. We develop a memory race recording mechanism, called Rerun, that uses small hardware state (~166 bytes/core), writes a small race log (~4 bytes/kilo- instruction), and operates well as the number of cores per system scales (e.g., to 16 cores). Rerun exploits the dual of conventional wisdom in race recording: Rather than record information about individual memory accesses that conflict, we record how long a thread executes without conflicting with other threads. In particular, Rerun passively creates atomic episodes. Each episode is a dynamic instruction sequence that a thread happens to execute without interacting with other threads. Rerun uses Lamport Clocks to order episodes and enable replay of an equivalent execution.
Derek Hower, Mark D. Hill
ISCA1
2006 Self-Checking and Self-Diagnosing 32-bit Microprocessor Multiplier
abstract
In this paper, we propose a low-cost fault tolerance technique for microprocessor multipliers, both non-pipelined (NP) and pipelined (P). Our fault tolerant multiplier designs are capable of detecting and correcting errors, diagnosing hard faults, and reconfiguring to take the faulty sub-unit off-line. We utilize the branch misprediction recovery mechanism in the microprocessor core to take the error detection process off the critical path. Our analysis shows that our scheme provides 99% fault security and, compared to a baseline unprotected multiplier, achieves this fault tolerance with low performance overhead (5% for NP and 2.5% for P multiplier) and reasonably low area (38% NP and 26% P) and power consumption (36% NP and 28.5% P) overheads
Mahmut Yilmaz, Derek Hower, Sule Ozev, Daniel J. Sorin
ITC2