Bob Rau

dblp:83/5996 · also B. Ramakrishna Rau · DBLP profile ↗
← Back
29ranked-venue papers
15as first author
0since 2021 · last 2002
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 14 first-authorSoftware engineering, systems software and programming languages · 7 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
12 papers
Compilers and program optimization · 94% Runtime systems and virtual machines · 6%
Computer architecture, parallel and distributed computing, and storage systems
18 papers
Processor architecture and microarchitecture · 79% Parallel and multicore computing · 8% Memory systems · 7%

Topics — the 30 heaviest of 41, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
instruction-level parallelism
0.161996
Profile-driven Instruction Level Parallel Scheduling with Application to Super Blocks · MICRO 1996
Sentinel Scheduling for VLIW and Superscalar Processors · ACM Trans. Comput. Syst. 1993
Dynamically scheduled VLIW processors · MICRO 1993
Compilers and program optimization › instruction scheduling
software pipelining
0.041994
Iterative modulo scheduling: an algorithm for software pipelining loops · MICRO 1994
Reverse If-Conversion · PLDI 1993
Register Allocation for Software Pipelined Loops · PLDI 1992
Compilers and program optimization
instruction scheduling
0.041996
Profile-driven Instruction Level Parallel Scheduling with Application to Super Blocks · MICRO 1996
Reverse If-Conversion · PLDI 1993
Register Allocation for Software Pipelined Loops · PLDI 1992
Compilers and program optimization › instruction scheduling
instruction-level parallelism
0.012001
Compiling for EPIC architectures · Proc. IEEE 2001
Processor architecture and microarchitecture
speculative execution
0.031993
Sentinel Scheduling for VLIW and Superscalar Processors · ACM Trans. Comput. Syst. 1993
Dynamically scheduled VLIW processors · MICRO 1993
Sentinel Scheduling for VLIW and Superscalar Processors · ASPLOS 1992
Compilers and program optimization › instruction scheduling › software pipelining
modulo scheduling
0.021994
Iterative modulo scheduling: an algorithm for software pipelining loops · MICRO 1994
Code generation schema for modulo scheduled loops · MICRO 1992
Processor architecture and microarchitecture › instruction-level parallelism
compiler-controlled speculative execution
0.021993
Sentinel Scheduling for VLIW and Superscalar Processors · ACM Trans. Comput. Syst. 1993
Sentinel Scheduling for VLIW and Superscalar Processors · ASPLOS 1992
Processor architecture and microarchitecture
instruction set architecture
0.032001
Compiling for EPIC architectures · Proc. IEEE 2001
Optimization of Machine Descriptions for Efficient Use · MICRO 1996
Architectural Support for the Efficient Generation of Code for Horizontal Architectures · ASPLOS 1982
Compilers and program optimization
interprocedural optimization
0.011995
Region-based compilation: an introduction and motivation · MICRO 1995
Runtime systems and virtual machines › dynamic compilation › just-in-time compilation
region-based compilation
0.011995
Region-based compilation: an introduction and motivation · MICRO 1995
Processor architecture and microarchitecture › instruction-level parallelism
VLIW
0.031993
Dynamically scheduled VLIW processors · MICRO 1993
Architectural Support for the Efficient Generation of Code for Horizontal Architectures · ASPLOS 1982
Efficient code generation for horizontal architectures: Compiler techniques and architectural support · ISCA 1982
Compilers and program optimization › loop transformation
loop scheduling
0.011994
Iterative modulo scheduling: an algorithm for software pipelining loops · MICRO 1994
Memory systems › memory architecture
interleaved memory
0.031991
Pseudo-Randomly Interleaved Memory · ISCA 1991
Program Behavior and the Performance of Interleaved Memories · IEEE Trans. Computers 1979
Interleaved Memory Bandwidth in a Model of a Muyltiprocessor Computer System · IEEE Trans. Computers 1979
Compilers and program optimization › program transformation › control flow transformation
if-conversion
0.011993
Reverse If-Conversion · PLDI 1993
Parallel and multicore computing › task scheduling
dynamic scheduling
0.011993
Dynamically scheduled VLIW processors · MICRO 1993
Processor architecture and microarchitecture › instruction set architecture
EPIC architecture
0.012001
Compiling for EPIC architectures · Proc. IEEE 2001
Compilers and program optimization
code generation
0.011992
Code generation schema for modulo scheduled loops · MICRO 1992
Compilers and program optimization
register allocation
0.011992
Register Allocation for Software Pipelined Loops · PLDI 1992
Processor architecture and microarchitecture › instruction-level parallelism
superscalar and VLIW processors
0.011992
Sentinel Scheduling for VLIW and Superscalar Processors · ASPLOS 1992
Processor architecture and microarchitecture
instruction scheduling
0.021994
Iterative modulo scheduling: an algorithm for software pipelining loops · MICRO 1994
Code generation schema for modulo scheduled loops · MICRO 1992
Compilers and program optimization › instruction scheduling
compile-time scheduling
0.021993
Dynamically scheduled VLIW processors · MICRO 1993
Sentinel Scheduling for VLIW and Superscalar Processors · ASPLOS 1992
Parallel and multicore computing
dataflow computing
0.011993
Predictability of load/store instruction latencies · MICRO 1993
Processor architecture and microarchitecture
exception handling
0.011993
Sentinel Scheduling for VLIW and Superscalar Processors · ACM Trans. Comput. Syst. 1993
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution
0.011993
Reverse If-Conversion · PLDI 1993
Memory systems
memory bandwidth
0.011991
Pseudo-Randomly Interleaved Memory · ISCA 1991
Embedded and real-time systems › model-based design
code generation
0.011982
Architectural Support for the Efficient Generation of Code for Horizontal Architectures · ASPLOS 1982
Performance modeling and evaluation › performance model construction › memory system performance modeling
memory bandwidth analysis
0.011979
Interleaved Memory Bandwidth in a Model of a Muyltiprocessor Computer System · IEEE Trans. Computers 1979
Memory systems › memory bandwidth
memory bandwidth modeling
0.011979
Program Behavior and the Performance of Interleaved Memories · IEEE Trans. Computers 1979
Parallel and multicore computing
multiprocessor system
0.011979
Interleaved Memory Bandwidth in a Model of a Muyltiprocessor Computer System · IEEE Trans. Computers 1979
Performance modeling and evaluation
queueing models
0.011979
Interleaved Memory Bandwidth in a Model of a Muyltiprocessor Computer System · IEEE Trans. Computers 1979

Methods — techniques the papers use, named apart from their topics

profile-driven scoring · 0.0list scheduling · 0.0static program analysis · 0.0iterative modulo scheduling · 0.0top-down and bottom-up clustering · 0.0register renaming · 0.0predicate internal representation · 0.0memory alias disambiguation · 0.0flow prediction · 0.0compile-time scheduling · 0.0reverse if-conversion · 0.0
YearPublicationVenuePosition
2002 Constructing and exploiting linear schedules with prescribed parallelism
abstract
We present two new results of importance in code generation for and synthesis of synchronously scheduled parallel processor arrays and multicluster VLIWs. The first is a new practical method for constructing a linear schedule for the iterations of a loop nest that schedules precisely one iteration per cycle on each of a prescribed set of processors. While this problem goes back to the era in which systolic computation was in vogue, it has defied practical solution until now. We provide a closed form solution that enables the enumeration of all such schedules. The second result is a new technique that reduces the cost of code or hardware whose function is to control the flow of data and predicate operations, and to generate memory addresses. The key idea is that by using the mathematical structure of any of the conflict-free schedules we construct, a very shallow recurrence can be developed to inexpensively update these quantities.
Alain Darte, Robert Schreiber, Bob Rau, Frédéric Vivien
ACM Trans. Design Autom. Electr. Syst.3
2001 Compiling for EPIC architectures
abstract
Designing compilers for Explicitly Parallel Instruction Computing (EPIC.) architectures presents challenges substantially different from those encountered in designing compilers for traditional sequential architectures. These challenges are addressed not only by employing new optimizations that are specific to EPIC, but also by employing new ways to architect compilers. EPIC architectures provide features that allow compilers to take a proactive role in exploiting instruction level parallelism. Compiler technology is intimately intertwined with the target processor architecture, and compiler architects must solve new analysis and optimization problems to achieve the highest levels of performance. When complex optimizations are uniformly applied to large applications, the resulting slow compile speeds are unacceptable. Demanding requirements to produce high-quality code at high compile speed shapes the fundamental structure of EPIC compilers.
Vinod Kathail, Mike Schlansker, Bob Rau
Proc. IEEE3
2000 High-Level Synthesis of Nonprogrammable Hardware Accelerators
abstract
The PICO-N system automatically synthesizes embedded nonprogrammable accelerators to be used as co-processors for functions expressed as loop nests in C. The output is synthesizable VHDL that defines the accelerator at the register transfer level (RTL). The system generates a synchronous array of customized VLIW (very-long instruction word) processors, their controller local memory, and interfaces. The system also modifies the user's application software to make use of the generated accelerator. The user indicates the throughput to be achieved by specifying the number of processors and their initiation interval. In experimental comparisons, PICO-N designs are slightly more costly than hand-designed accelerators with the same performance.
Robert Schreiber, Shail Aditya, Bob Rau, Vinod Kathail, Scott A. Mahlke, Santosh G. Abraham, Greg Snider
ASAP3
2000 Efficient design space exploration in PICO
abstract
Automated design tools must understand and exploit the hierarchical structure of large design spaces.We have developed a general methodology for decomposing system design spaces into smaller component design spaces, followed by component-level evaluation, filtering, recomposition and system-level evaluation.This methodology greatly reduces the time and cost of design space exploration, since the typical number of system-level evaluations is greatly reduced.This paper describes the application of our decomposition methodology in the context of PICO.PICO is a design space exploration system that automatically generates embedded designs consisting of a stylized processor, hardware accelerator and a cache hierarchy, each customized to a benchmark.First, PICO splits the specified system design space into smaller design spaces, one for each of the components, viz.processor, accelerator and data/instruction/unified caches.PICO further partitions each component design space into predicated design spaces, so that all designs in a predicated design space satisfy a specified predicate.PICO uses component-level evaluations to identify the performance-cost optimal component-level Pareto designs in each predicated design space.PICO generates all compositions of Pareto designs from compatible predicated design spaces and uses a system-level evaluation to identify the Pareto designs at the system level.For reasonable design spaces, PICO reduces the design exploration time by over four orders of magnitude compared to an exhaustive approach.
Santosh G. Abraham, Bob Rau
CASES2
2000 The era of embedded computing
abstract
No abstract available.
Bob Rau
CASES1
2000 Embedded Computing: New Directions in Architecture and Automation
Bob Rau, Mike Schlansker
HiPC1
2000 A Constructive Solution to the Juggling Problem in Processor Array Synthesis
abstract
We describe a new, practical, constructive method for solving the well-known conflict-free scheduling problem for the locally sequential, globally parallel (LSGP) case of processor array synthesis. First, we provide a closed form solution that enables the enumeration of all conflict-free schedules. Then, we discuss the reduction of the cost of hardware whose function is to control the flow of data, enable or disable functional units, and generate memory addresses. We present a new technique for controlling the complexity of these housekeeping functions in a processor array. Both of these techniques have been incorporated into a software system for the automatic synthesis of hardware accelerators developed by HP Labs.
Alain Darte, Robert Schreiber, Bob Rau, Frédéric Vivien
IPDPS3
2000 Code size minimization and retargetable assembly for custom EPIC and VLIW instruction formats
abstract
PICO is a fully automated system for designing the architecture and the microarchitecture of VLIW and EPIC processors. A serious concern with this class of processors, due to their very long instructions, is their code size. One focus of this paper is to describe a series of code size minimization techniques used within PICO, some of which are applied during the automatic design of the instruction format, while others are applied during program assembly. The design of a retargetable assembler to support these techniques also poses certain novel challenges, which constitute the second focus of this paper. Contrary to widely held perceptions, we demonstrate that it is entirely possible to design VLIW and EPIC processors that are capable of issuing large numbers of operational per cycle, but whose code size is only moderately larger than that for a sequential CISC processor.
Shail Aditya, Scott A. Mahlke, Bob Rau
ACM Trans. Design Autom. Electr. Syst.3
1996 Profile-driven Instruction Level Parallel Scheduling with Application to Super Blocks
abstract
Code scheduling to exploit instruction level parallelism (ILP) is a critical problem in compiler optimization research in light of the increased use of long-instruction-word machines. Unfortunately optimum scheduling is computationally intractable, and one must resort to carefully crafted heuristics in practice. If the scope of application of a scheduling heuristic is limited to basic blocks, considerable performance loss may be incurred at block boundaries. To overcome this obstacle, basic blocks can be coalesced across branches to form larger regions such as super blocks. In the literature, these regions are typically scheduled using algorithms that are either oblivious to profile information (under the assumption that the process of forming the region has fully utilized the profile information), or use the profile information as an addendum to classical scheduling techniques. We believe that even for the simple case of linear code regions such as super blocks, additional performance improvement can be gained by utilizing the profile information in scheduling as well. We propose a general paradigm for converting any profile-insensitive list scheduler to a profile-sensitive scheduler. Our technique is developed via a theoretical analysis of a simplified abstract model of the general problem of profile-driven scheduling over any acyclic code region, yielding a scoring measure for ranking branch instructions.
Chandra Chekuri, Rajeev Motwani 0001, B. Natarajan, Bob Rau, Mike Schlansker
MICRO5
1996 Optimization of Machine Descriptions for Efficient Use
John C. Gyllenhaal, Wen-Mei W. Hwu, Bob Rau
MICRO3
1995 Region-based compilation: an introduction and motivation
abstract
As the amount of instruction-level parallelism required to fully utilize VLIW and superscalar processors increases, compilers must perform increasingly more aggressive analysis, optimization, parallelization and scheduling on the input programs. Traditionally, compilers have been built assuming functions as the unit of compilation. In this framework, function boundaries tend to hide valuable optimization opportunities from the compiler. Function inlining may be applied to assemble strongly coupled functions into the same compilation unit at the cost of very large function bodies. This paper introduces a new technique, called region-based compilation, where the compiler is allowed to repartition the program into more desirable compilation units. Region-based compilation allows the compiler to control problem size while exposing inter-procedural optimization and code motion opportunities.
Richard E. Hank, Wen-Mei W. Hwu, Bob Rau
MICRO3
1994 Iterative modulo scheduling: an algorithm for software pipelining loops
abstract
Module scheduling is a framework within which a wide variety of algorithms and heuristics may be defined for software pipelining innermost loops. This paper presents a practical algorithm, iterative module scheduling, that is capable of dealing with realistic machine models. This paper also characterizes the algorithm in terms of the quality of the generated schedules as well the computational expense incurred.
Bob Rau
MICRO1
1993 Predictability of load/store instruction latencies
abstract
Presents a model of coarse grain dataflow execution. The authors present one top down and two bottom up methods for generation of multithreaded code, and evaluate their effectiveness. The bottom up techniques start from a fine-grain dataflow graph and coalesce this into coarse-grain clusters. The top down technique generates clusters directly from the intermediate data dependence graph used for compiler optimizations. The authors discuss the relevant phases in the compilation process. They compare the effectiveness of the strategies by measuring the total number of clusters executed, the total number of instructions executed, cluster size, and number of matches per cluster. It turns out that the top down method generates more efficient code, and larger clusters. However the number of matches per cluster is larger for the top down method, which could incur higher cluster synchronization costs.>
Santosh G. Abraham, Rabin A. Sugumar, Daniel Windheiser, Bob Rau
MICRO4
1993 Dynamically scheduled VLIW processors
abstract
Instruction-level parallelism in a single stream of code for non-numerical applications has been the subject of many recent researches. This work extends the analysis to symbolic applications described with logic programming. In particular, the authors analyze the effects on performance of speculative execution, memory alias disambiguation, renaming and flow prediction. The obtained results indicate that one can reach a sustained parallelism of 4 (comparable with imperative languages), with the proper optimizations. The authors also show a comparison between static and dynamic scheduled approaches, outlining the conditions under which a dynamic solution can reach substantial improvements over a static one. In this way, they point out some important optimizations and parameters of a dynamic scheduling approach, indicating a guideline for future architectural implementations.>
Bob Rau
MICRO1
1993 Reverse If-Conversion
abstract
In this paper we present a set of isomorphic control transformations that allow the compiler to apply local scheduling techniques to acyclic subgraphs of the control flow graph. Thus, the code motion complexities of global scheduling are eliminated. This approach relies on a new technique, Reverse If-Conversion (RIC), that transforms scheduled If-Converted code back to the control flow graph representation. This paper presents the predicate internal representation, the algorithms for RIC, and the correctness of RIC. In addition, the scheduling issues are addressed and an application to software pipelining is presented.
Nancy J. Warter, Scott A. Mahlke, Wen-Mei W. Hwu, Bob Rau
PLDI4
1993 Guest editors' introduction
Joseph A. Fisher, Bob Rau
J. Supercomput.2
1993 Instruction-level parallel processing: History, overview, and perspective
Bob Rau, Joseph A. Fisher
J. Supercomput.1
1993 Sentinel Scheduling for VLIW and Superscalar Processors
abstract
Speculative execution is an important source of parallelism for VLIW and superscalar processors. A serious challenge with compiler-controlled speculative execution is to efficiently handle exceptions for speculative instructions. In this article, a set of architectural features and compile-time scheduling support collectively referred to assentinel schedulingis introduced. Sentinel scheduling provides an effective framework for both compiler-controlled speculative execution and exception handling. All program exceptions are accurately detected and reported in a timely manner with sentinel scheduling. Recovery from exceptions is also ensured with the model. Experimental results show the effectiveness of sentinel scheduling for exploiting instruction-level parallelism and overhead associated with exception handling.
Scott A. Mahlke, William Y. Chen, Roger A. Bringmann, Richard E. Hank, Wen-Mei W. Hwu, Bob Rau, Mike Schlansker
ACM Trans. Comput. Syst.6
1992 Sentinel Scheduling for VLIW and Superscalar Processors
abstract
Speculative execution is an important source of parallelism for VLIW and superscalar processors. A serious challenge with compiler-controlled speculative execution is to accurately detect and report all program execution errors at the time of occurrence. In this paper, a set of architectural features and compile-time scheduling support referred to as sentinel scheduling is introduced. Sentinel scheduling provides an effective framework for compiler-controlled speculative execution that accurately detects and reports all exceptions. Sentinel scheduling also supports speculative execution of store instructions by providing a store buffer which allows probationary entries. Experimental results show that sentinel scheduling is highly effective for a wide range of VLIW and superscalar processors.
Scott A. Mahlke, William Y. Chen, Wen-Mei W. Hwu, Bob Rau, Mike Schlansker
ASPLOS4
1992 Code generation schema for modulo scheduled loops
Bob Rau, Mike Schlansker, Parthasarathy P. Tirumalai
MICRO1
1992 Register Allocation for Software Pipelined Loops
abstract
Software pipelining is an important instruction scheduling technique for efficiently overlapping successive iterations of loops and executing them in parallel. This paper studies the task of register allocation for software pipelined loops, both with and without hardware features that are specifically aimed at supporting software pipelines. Register allocation for software pipelines presents certain novel problems leading to unconventional solutions, especially in the presence of hardware support. This paper formulates these novel problems and presents a number of alternative solution strategies. These alternatives are comprehensively tested against over one thousand loops to determine the best register allocation strategy, both with and without the hardware support for software pipelining.
Bob Rau, Meng Lee, Parthasarathy P. Tirumalai, Mike Schlansker
PLDI1
1991 Pseudo-Randomly Interleaved Memory
abstract
Interleaved memories are often used to provide the high bandwidth needed by multi- processors and high performance uniprocessors. The manner in which memory locations are distributed across the memory modules has a significant influence on whether, and for which types of reference patterns, the full bandwidth of the memory system is achieved. The most common interleaved memory architecture is the sequentially interleaved memory in which successive memory locations are assigned to successive memory modules. Although such an architecture is the simplest to implement and provides good performance with strides that are odd integers, it can degrade badly in the face of even strides, especially strides that are a power of two. This happens because all the memory references are concentrated on a subset of the memory modules. Pseudo-
Bob Rau
ISCA1
1989 The Cydram 5 Stride-Insensitive Memory System
Bob Rau, Mike Schlansker, David W. L. Yen
ICPP (1)1
1982 Architectural Support for the Efficient Generation of Code for Horizontal Architectures
abstract
Horizontal architectures, such as the CDC Advanced Flexible Processor [I] and the FPS APi20-B [2}, consist of a number of resources that can operate in parallel, each of which is controlled by a field in the wide instruction word. Such architectures have been developed to perform high speed scientific computations at a modest cost: Figure 1 displays those characteristics of horizontal architectures that are germane to the issues discussed in this paper. The simultaneous requirements of high performance and low cost lead to an architecture consisting of multiple pipelined processing elements (PEs) such as adders and multipliers, a memory (which for scheduling purposes may be viewed as yet another PE with two operations: a READ and a WRITE), and an interconnect which ties them all together. The interconnect allows the result of one operation to be directly routed to another PE as one of the inputs for an operation that is to be performed there. The required memory bandwidth is reduced since temporary values need not be written to and read from the memory. The final aspect of horizontal processors that is of interest is that their program memories emit wide instructions which synchronously specify the actions of the multiple and possibly dissimilar PEs. The program memory is sequenced by a conventional sequencer that assumes sequential flow of control unless a branch is explicitly specified.
Bob Rau, Christopher D. Glaeser, E. M. Greenawalt
ASPLOS1
1982 Efficient code generation for horizontal architectures: Compiler techniques and architectural support
abstract
A horizontal architecture consists of a number of resources that can operate in parallel, each of which is controlled by a field in the wide instruction word. Such architectures offer the potential for high performance scientific computing at a modest cost. If this potential performance is to be realized, the multiple resources of a horizontal processor must be scheduled effectively. The scheduling task for conventional horizontal processors is quite complex and the construction of highly optimizing compilers for them is a difficult and expensive project. The polycyclic architecture is a horizontal architecture with architectural support for the scheduling task. The complexity of scheduling conventional horizontal processors and the ease of scheduling polycyclic processors is demonstrated by means of an example.
Bob Rau, Christopher D. Glaeser, Raymond L. Picard
ISCA1
1979 Interleaved Memory Bandwidth in a Model of a Muyltiprocessor Computer System
abstract
An approximate analysis is performed of an often studied model of an interleaved memory, multiprocessor system consisting of M memory modules and N processors. A closed-form solution is obtained and the one approximation used is found to result in negligible error. This solution is about an order of magnitude more accurate than the best previous result.
Bob Rau
IEEE Trans. Computers1
1979 Program Behavior and the Performance of Interleaved Memories
abstract
One of the major factors influencing the performance of an interleaved memory system is the behavior of the request sequence, but this is normally ignored. This paper examines this issue. Using trace driven simulations it is shown that the commonly used assumption, that each request is independently and equally likely to be to any module, is not valid. The duality of memory interference with paging behavior is noted and this suggests the use of the least-recently used stack model to model program behavior. Simulations indicate that this model is reasonably accurate. An accurate, though approximate, expression for the bandwidth is derived based upon this model.
Bob Rau
IEEE Trans. Computers1
1977 The Effect of Instruction Fetch Strategies upon the Performance of Pipelined Instruction Units
Bob Rau, George E. Rossman
ISCA1
1976 A new philosophy for interconnection on multilayer boards
abstract
Alternative approaches to the interconnection problem on multilayer boards are studied and evaluated with respect to their desirability and feasability. A brief review of the current methodology is included to highlight those aspects of the current approach that might be improved upon. The conclusion is that a more integrated approach to interconnection on multilayer boards is essential to deriving the maximum benefit from the additional degrees of freedom afforded by a multilayer board. Various such integrated approaches are proposed.
Bob Rau
DAC1