Richard E. Hank

dblp:92/6679 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 2 first-authorSoftware engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
10 papers
Processor architecture and microarchitecture · 90% Memory systems · 10%
Software engineering, system software, and programming languages
6 papers
Compilers and program optimization · 88% Runtime systems and virtual machines · 12%

Topics — the 23 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
instruction-level parallelism
0.171995
A Comparison of Full and Partial Predicated Execution Support for ILP Processors · ISCA 1995
Characterizing the impact of predicated execution on branch prediction · MICRO 1994
Sentinel Scheduling for VLIW and Superscalar Processors · ACM Trans. Comput. Syst. 1993
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution
0.021995
A Comparison of Full and Partial Predicated Execution Support for ILP Processors · ISCA 1995
Characterizing the impact of predicated execution on branch prediction · MICRO 1994
Compilers and program optimization
predicated execution
0.021995
A Comparison of Full and Partial Predicated Execution Support for ILP Processors · ISCA 1995
Effective compiler support for predicated execution using the hyperblock · MICRO 1992
Processor architecture and microarchitecture
speculative execution
0.021993
Sentinel Scheduling for VLIW and Superscalar Processors · ACM Trans. Comput. Syst. 1993
Speculative execution exception recovery using write-back suppression · MICRO 1993
Compilers and program optimization
code generation
0.011995
A Comparison of Full and Partial Predicated Execution Support for ILP Processors · ISCA 1995
Compilers and program optimization
dependence analysis
0.011995
Compiler technology for future microprocessors · Proc. IEEE 1995
Compilers and program optimization › instruction scheduling
instruction-level parallelism
0.011995
Compiler technology for future microprocessors · Proc. IEEE 1995
Compilers and program optimization
interprocedural optimization
0.011995
Region-based compilation: an introduction and motivation · MICRO 1995
Compilers and program optimization
predicated compilation
0.011995
Compiler technology for future microprocessors · Proc. IEEE 1995
Runtime systems and virtual machines › dynamic compilation › just-in-time compilation
region-based compilation
0.011995
Region-based compilation: an introduction and motivation · MICRO 1995
Processor architecture and microarchitecture
instruction set architecture
0.021993
Register Connection: A New Approach to Adding Registers into Instruction Set Architectures · ISCA 1993
Superblock formation using static program analysis · MICRO 1993
Processor architecture and microarchitecture
branch prediction
0.011994
Characterizing the impact of predicated execution on branch prediction · MICRO 1994
Compilers and program optimization › instruction scheduling
software pipelining
0.011993
Superblock formation using static program analysis · MICRO 1993
Processor architecture and microarchitecture › instruction-level parallelism
compiler-controlled speculative execution
0.011993
Sentinel Scheduling for VLIW and Superscalar Processors · ACM Trans. Comput. Syst. 1993
Processor architecture and microarchitecture › instruction-level parallelism
multiple instruction issue
0.011993
Speculative execution exception recovery using write-back suppression · MICRO 1993
Processor architecture and microarchitecture
register file
0.011993
Register Connection: A New Approach to Adding Registers into Instruction Set Architectures · ISCA 1993
Processor architecture and microarchitecture › instruction set architecture
instruction set design
0.011995
A Comparison of Full and Partial Predicated Execution Support for ILP Processors · ISCA 1995
Compilers and program optimization
register allocation
0.011993
Register Connection: A New Approach to Adding Registers into Instruction Set Architectures · ISCA 1993
Memory systems
cache
0.011993
Speculative execution exception recovery using write-back suppression · MICRO 1993
Memory systems › cache › CPU cache
data cache
0.011993
Speculative execution exception recovery using write-back suppression · MICRO 1993
Processor architecture and microarchitecture
exception handling
0.011993
Sentinel Scheduling for VLIW and Superscalar Processors · ACM Trans. Comput. Syst. 1993
Memory systems
cache design
0.011992
An efficient architecture for loop based data preloading · MICRO 1992
Processor architecture and microarchitecture › instruction-level parallelism
superscalar and VLIW processors
0.011992
Effective compiler support for predicated execution using the hyperblock · MICRO 1992

Methods — techniques the papers use, named apart from their topics

static program analysis · 0.0compiler optimization · 0.0execution-driven simulation · 0.0simulation · 0.0software pipelining · 0.0selective scheduling · 0.0memory disambiguation · 0.0compile-time scheduling · 0.0
YearPublicationVenuePosition
2013 Instant profiling: Instrumentation sampling for profiling datacenter applications
abstract
Profile-guided optimization possesses huge potential to save costs for datacenters. Hardware performance monitoring units enable profiling with negligible overhead and they have been proven to be effective to help programmers find code regions to optimize by monitoring datacenter applications continuously on live traffic. However, these hardware features are inflexible and often buggy, limiting the types of data that can be gathered. Instrumentation-based profiling can complement or replace hardware functionality by providing more flexible and targeted information gathering. Unfortunately, the overhead of existing instrumentation mechanisms prevents their use in production runs. In order to be used in datacenters, we need a profiling mechanism to impose overheads of less than a few percent, in terms of both throughput and latency, while still generating meaningful profile data. This paper presents instant profiling, an instrumentation sampling technique using dynamic binary translation. Instead of instrumenting the entire execution, instant profiling periodically interleaves native execution and instrumented execution according to configurable profiling duration and frequency parameters. It further reduces the latency degradation of initial profiling phases by pre-populating a software code cache. We evaluate the performance and effectiveness of this new profiling technique on the SPEC CINT2006 benchmark suite and two datacenter application benchmarks. We show that it is well-suited for deployment to datacenters by incurring less than 6% slowdown and 3% computational overhead on average.
Hyoun Kyu Cho, Tipp Moseley, Richard E. Hank, Derek Bruening, Scott A. Mahlke
CGO3
2006 Prematerialization: reducing register pressure for free
abstract
Modern compiler transformations that eliminate redundant computations or reorder instructions, such as partial redundancy elimination and instruction scheduling, are very effective in improving application performance but tend to create longer and potentially more complex live ranges. Typically the task of dealing with the increased register pressure is left to the register allocator. To avoid introduction of spill code which can reduce or completely eliminate the benefit of earlier optimizations, researchers have developed techniques such as live range splitting and rematerializatio.This paper describes prematerialization (PM), a novel method for reducing register pressure for VLIW architectures with nop instructions. PM and rematerialization both select "never killed" live ranges and break them up by introducing one or more definitions close to the uses. However, while rematerialization is applied to live ranges selected for spilling during register allocation, PM relies on the availability of nop instructions and occurs prior to register allocation. PM simplifies register allocation by creating live ranges that are easier to color and less likely to spill. We have implemented prematerialization in HP-UX production compilers for the Intel® Itanium® architecture. Performance evaluation indicates that the proposed technique is effective in reducing register pressure inherent in highly optimized code.
Ivan D. Baev, Richard E. Hank, David H. Gross
PACT2
1995 A Comparison of Full and Partial Predicated Execution Support for ILP Processors
abstract
One can effectively utilize predicated execution to improve branch handling in instruction-level parallel processors. Although the potential benefits of predicated execution are high, the tradeoffs involved in the design of an instruction set to support predicated execution can be difficult. On one end of the design spectrum, architectural support for full predicated execution requires increasing the number of source operands for all instructions. Full predicate support provides for the most flexibility and the largest potential performance improvements. On the other end, partial predicated execution support, such as conditional moves, requires very little change to existing architectures. This paper presents a preliminary study to qualitatively and quantitatively address the benefit of full and partial predicated execution support. With our current compiler technology, we show that the compiler can use both partial and full predication to achieve speedup in large control-intensive programs. Some details of the code generation techniques are shown to provide insight into the benefit of going from partial to full predication. Preliminary experimental results are very encouraging: partial predication provides an average of 33% performance improvement for an 8-issue processor with no predicate support while full predication provides an additional 30% improvement.
Scott A. Mahlke, Richard E. Hank, James E. McCormick, David I. August, Wen-Mei W. Hwu
ISCA2
1995 Region-based compilation: an introduction and motivation
abstract
As the amount of instruction-level parallelism required to fully utilize VLIW and superscalar processors increases, compilers must perform increasingly more aggressive analysis, optimization, parallelization and scheduling on the input programs. Traditionally, compilers have been built assuming functions as the unit of compilation. In this framework, function boundaries tend to hide valuable optimization opportunities from the compiler. Function inlining may be applied to assemble strongly coupled functions into the same compilation unit at the cost of very large function bodies. This paper introduces a new technique, called region-based compilation, where the compiler is allowed to repartition the program into more desirable compilation units. Region-based compilation allows the compiler to control problem size while exposing inter-procedural optimization and code motion opportunities.
Richard E. Hank, Wen-Mei W. Hwu, Bob Rau
MICRO1
1995 Compiler technology for future microprocessors
abstract
Advances in hardware technology have made it possible for microprocessors to execute a large number of instructions concurrently (i.e., in parallel). These microprocessors take advantage of the opportunity to execute instructions in parallel to increase the execution speed of a program. As in other forms of parallel processing, the performance of these microprocessors can vary greatly depending on the qualify of the software. In particular the quality of compilers can make an order of magnitude difference in performance. This paper presents a new generation of compiler technology that has emerged to deliver the large amount of instruction-level-parallelism that is already required by some current state-of-the-art microprocessors and will be required by more future microprocessors. We introduce critical components of the technology which deal with difficult problems that are encountered when compiling programs for a high degree of instruction-level-parallelism. We present examples to illustrate the functional requirements of these components. To provide more insight into the challenges involved, we present in-depth case studies on predicated compilation and maintenance of dependence information, two of the components that are largely missing from most current commercial compilers.
Wen-Mei W. Hwu, Richard E. Hank, David M. Gallagher, Scott A. Mahlke, Daniel M. Lavery, Grant E. Haab, John C. Gyllenhaal, David I. August
Proc. IEEE2
1994 Characterizing the impact of predicated execution on branch prediction
abstract
Branch instructions are recognized as a major impediment to exploiting instruction level parallelism. Even with sophisticated branch prediction techniques, many frequently executed branches remain difficult to predict. An architecture supporting predicated execution may allow the compiler to remove many of these hard-to-predict branches, reducing the number of branch mispredictions and thereby improving performance. We present an in-depth analysis of the characteristics of those branches which are frequently mispredicted and examine the effectiveness of an advanced compiler to eliminate these branches. Over the benchmarks studied, an average of 27% of the dynamic branches and 56% of the dynamic branch mispredictions are eliminated with predicated execution support.
Scott A. Mahlke, Richard E. Hank, Roger A. Bringmann, John C. Gyllenhaal, David M. Gallagher, Wen-Mei W. Hwu
MICRO2
1993 Register Connection: A New Approach to Adding Registers into Instruction Set Architectures
abstract
Code optimization and scheduling for superscalar and superpipelined processors often increase the register requirement of programs. For existing instruction sets with a small to moderate number of registers, this increased register requirement can be a factor that limits the effectivess of the compiler. In this paper, we introduce a new architectural method for adding a set of extended registers into an architecture. Using a novel concept of connection, this method allows the data stored in the extended registers to be accessed by instructions that apparently reference core registers. Furthermore, we address the technical issues involved in applying the new method to an architecture: instruction set extension, procedure call convention, context switching considerations, upward compatibility, efficient implementation, compiler support, and performance. Experimental results based on a prototype compiler and execution driven simulation show that the proposed method can significantly improve the performance of superscalar processors with a small or moderate number of registers.
Tokuzo Kiyohara, Scott A. Mahlke, William Y. Chen, Roger A. Bringmann, Richard E. Hank, Sadun Anik, Wen-Mei W. Hwu
ISCA5
1993 Speculative execution exception recovery using write-back suppression
abstract
One of the key design concerns of multiple instruction issue (MII) processors is deciding how many memory ports need to be provided, considering performance and efficiency of the target processor. For an MII processor that exploits instruction-level parallelism (ILP) in non-numerical code, this decision is difficult to make due to its irregularity. The authors perform an empirical study aimed at characterizing a suitable MII organization that best exploits irregular ILP. The study is based on the selective scheduling compiler that performs precise memory disambiguation for concurrent execution of multiple memory operations, along with renaming, speculation, and software pipelining. The result indicates that a small number of memory ports (i.e. less than half of the issue rate) is enough for exploiting most of irregular ILP. The authors also examine related issues such as the utilization of memory ports and additional data cache misses caused by speculative loads.>
Roger A. Bringmann, Scott A. Mahlke, Richard E. Hank, John C. Gyllenhaal, Wen-Mei W. Hwu
MICRO3
1993 Superblock formation using static program analysis
abstract
To achieve higher instruction-level parallelism, the constraint imposed by a single control flow must be relaxed. Control operations should execute in parallel just like data operations. We present a new software pipelining method called GPMB (Global Pipelining with Multiple Branches) which is based on architectures supporting multi-way branching and multiple control flows. Preliminary experimental results show that, for IFless loops, GPMB performs as well as modulo scheduling, and for branch-intensive loops, GPMB performs much better than software pipelining assuming the constraint of one two-way branch per cycle.>
Richard E. Hank, Scott A. Mahlke, Roger A. Bringmann, John C. Gyllenhaal, Wen-Mei W. Hwu
MICRO1
1993 The superblock: An effective technique for VLIW and superscalar compilation
Wen-Mei W. Hwu, Scott A. Mahlke, William Y. Chen, Pohua P. Chang, Nancy J. Warter, Roger A. Bringmann, Roland G. Ouellette, Richard E. Hank, Tokuzo Kiyohara, Grant E. Haab, John G. Holm, Daniel M. Lavery
J. Supercomput.8
1993 Sentinel Scheduling for VLIW and Superscalar Processors
abstract
Speculative execution is an important source of parallelism for VLIW and superscalar processors. A serious challenge with compiler-controlled speculative execution is to efficiently handle exceptions for speculative instructions. In this article, a set of architectural features and compile-time scheduling support collectively referred to assentinel schedulingis introduced. Sentinel scheduling provides an effective framework for both compiler-controlled speculative execution and exception handling. All program exceptions are accurately detected and reported in a timely manner with sentinel scheduling. Recovery from exceptions is also ensured with the model. Experimental results show the effectiveness of sentinel scheduling for exploiting instruction-level parallelism and overhead associated with exception handling.
Scott A. Mahlke, William Y. Chen, Roger A. Bringmann, Richard E. Hank, Wen-Mei W. Hwu, Bob Rau, Mike Schlansker
ACM Trans. Comput. Syst.4
1992 An efficient architecture for loop based data preloading
William Y. Chen, Roger A. Bringmann, Scott A. Mahlke, Richard E. Hank, James E. Sicolo
MICRO4
1992 Effective compiler support for predicated execution using the hyperblock
abstract
Predicated execution is an effective technique for dealing with conditional branches in application programs. However, there are several problems associated with conventional compiler support for predicated execution. First, all paths of control are combined into a single path regardless of their execution frequency and size with conventional if-conversion techniques. Second, speculative execution is difficult to combine with predicated execution. In this paper, we propose the use of a new structure, referred to as the hyperblock, to overcome these problems. The hyperblock is an efficient structure to utilize predicated execution for both compiletime optimization and scheduling. Preliminary experimental results show that the hyperblock is highly effective for a wide range of superscalar and VLIW processors. 1 Introduction Superscalar and VLIW processors can potentially provide large performance improvements over their scalar predecessors by providing multiple data paths and function u...
Scott A. Mahlke, David C. Lin, William Y. Chen, Richard E. Hank, Roger A. Bringmann
MICRO4