Rodric M. Rabbah

dblp:87/394 · also Rodric Michel Rabbah, Rodric Rabbah · DBLP profile ↗
← Back
25ranked-venue papers
2as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 2 first-authorSoftware engineering, systems software and programming languages · 13 · 1 first-authorTheory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
10 papers
Parallel and multicore computing · 37% GPUs and heterogeneous computing · 22% Electronic design automation · 14%
Software engineering, system software, and programming languages
10 papers
Compilers and program optimization · 79% Programming languages and type systems · 15% Program synthesis and code generation · 5%

Topics — the 26 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization › accelerator compilation
heterogeneous compilation
0.322012
A compiler and runtime for heterogeneous computing · DAC 2012
Lime: a Java-compatible and synthesizable language for heterogeneous architectures · OOPSLA 2010
Parallel and multicore computing
parallel programming models
0.222014
Translating imperative code to MapReduce · OOPSLA 2014
Teleport messaging for distributed stream programs · PPoPP 2005
Compilers and program optimization
program transformation
0.212014
Translating imperative code to MapReduce · OOPSLA 2014
Parallel and multicore computing › data-parallel programming
mapreduce
0.212014
Translating imperative code to MapReduce · OOPSLA 2014
Compilers and program optimization › accelerator compilation
GPU compiler
0.112012
Compiling a high-level language for GPUs: (via language support for architectures and compilers) · PLDI 2012
GPUs and heterogeneous computing
GPU programming
0.112012
Compiling a high-level language for GPUs: (via language support for architectures and compilers) · PLDI 2012
GPUs and heterogeneous computing › GPU programming
high-level language compilation
0.112012
Compiling a high-level language for GPUs: (via language support for architectures and compilers) · PLDI 2012
Programming languages and type systems › domain-specific languages
hardware description languages
0.112011
Virtualization of heterogeneous machines hardware description in a synthesizable object-oriented language · DAC 2011
Electronic design automation
high-level synthesis
0.112011
Virtualization of heterogeneous machines hardware description in a synthesizable object-oriented language · DAC 2011
Compilers and program optimization › vectorization
SIMD vectorization
0.112010
MacroSS: macro-SIMDization of streaming applications · ASPLOS 2010
Parallel and multicore computing
data parallelism
0.112010
MacroSS: macro-SIMDization of streaming applications · ASPLOS 2010
Electronic design automation › high-level synthesis
hardware compilation
0.112010
Lime: a Java-compatible and synthesizable language for heterogeneous architectures · OOPSLA 2010
Reconfigurable computing and FPGAs › FPGA accelerator
FPGA-based stream processing
0.112009
A computing origami: folding streams in FPGAs · DAC 2009
GPUs and heterogeneous computing
heterogeneous architecture
0.122011
Virtualization of heterogeneous machines hardware description in a synthesizable object-oriented language · DAC 2011
Lime: a Java-compatible and synthesizable language for heterogeneous architectures · OOPSLA 2010
Program synthesis and code generation › syntax-guided synthesis
sketch-based synthesis
0.112005
Programming by sketching for bit-streaming programs · PLDI 2005
Compilers and program optimization
vectorization
0.112005
Exploiting Vector Parallelism in Software Pipelined Loops · MICRO 2005
Processor architecture and microarchitecture
instruction set architecture
0.112005
Exploiting Vector Parallelism in Software Pipelined Loops · MICRO 2005
Parallel and multicore computing › parallel programming models
stream programming
0.112005
Teleport messaging for distributed stream programs · PPoPP 2005
Embedded and real-time systems › model-based design › dataflow modeling
synchronous dataflow
0.112005
Teleport messaging for distributed stream programs · PPoPP 2005
Processor architecture and microarchitecture › SIMD
vector instructions
0.112005
Exploiting Vector Parallelism in Software Pipelined Loops · MICRO 2005
Compilers and program optimization › prefetching
compiler-inserted prefetching
0.012004
Compiler orchestrated prefetching via speculation and predication · ASPLOS 2004
Compilers and program optimization
prefetching
0.012004
Compiler orchestrated prefetching via speculation and predication · ASPLOS 2004
Memory systems
memory access latency
0.012004
Compiler orchestrated prefetching via speculation and predication · ASPLOS 2004
Memory systems › cache
prefetching
0.012004
Compiler orchestrated prefetching via speculation and predication · ASPLOS 2004
Distributed systems
stream processing
0.012009
A computing origami: folding streams in FPGAs · DAC 2009
Program analysis › program representation
dependence graphs
0.012004
Compiler orchestrated prefetching via speculation and predication · ASPLOS 2004

Methods — techniques the papers use, named apart from their topics

rewrite rules · 0.4group-by operations · 0.4fold operations · 0.4runtime orchestration · 0.3high-level language compilation · 0.3logic synthesis · 0.2behavioral synthesis · 0.2type system · 0.2synthesis · 0.2macro-SIMDization · 0.2
YearPublicationVenuePosition
2014 Stream Processing with a Spreadsheet
Mandana Vaziri, Olivier Tardieu, Rodric M. Rabbah, Philippe Suter, Martin Hirzel
ECOOP3
2014 Translating imperative code to MapReduce
abstract
We present an approach for automatic translation of sequential, imperative code into a parallel MapReduce framework. Automating such a translation is challenging: imperative updates must be translated into a functional MapReduce form in a manner that both preserves semantics and enables parallelism. Our approach works by first translating the input code into a functional representation, with loops succinctly represented by fold operations. Then, guided by rewrite rules, our system searches a space of equivalent programs for an effective MapReduce implementation. The rules include a novel technique for handling irregular loop-carried dependencies using group-by operations to enable greater parallelism. We have implemented our technique in a tool called Mold. It translates sequential Java code into code targeting the Apache Spark runtime. We evaluated Mold on several real-world kernels and found that in most cases Mold generated the desired MapReduce program, even for codes with complex indirect updates.
Cosmin Radoi, Stephen J. Fink, Rodric M. Rabbah, Manu Sridharan
OOPSLA3
2013 The Liquid Metal IP bridge
abstract
Programmers are increasingly turning to heterogeneous systems to achieve performance. Examples include FPGA-based systems that integrate reconfigurable architectures with conventional processors. However, the burden of managing the coding complexity that is intrinsic to these systems falls entirely on the programmer. This limits the proliferation of these systems as only highly-skilled programmers and FPGA developers can unlock their potential. The goal of the Liquid Metal project at IBM Research is to address the programming complexity attributed to heterogeneous FPGA-based systems. A feature of this work is a vertically integrated development lifecycle that appeals to skilled software developers. A primary enabler for this work is a canonical IP bridge, designed to offer a uniform communication methodology between software and hardware, and that is applicable across a wide range of platforms available off-the-shelf.
Perry Cheng, Stephen J. Fink, Rodric M. Rabbah, Sunil Shukla
ASP-DAC3
2013 The Shape of Things to Run - Compiling Complex Stream Graphs to Reconfigurable Hardware in Lime
Joshua S. Auerbach, David F. Bacon, Perry Cheng, Steve Fink, Rodric M. Rabbah
ECOOP5
2013 The Liquid Metal Blokus Duo Design
abstract
This paper describes the Liquid Metal entry in the 2013 ICFPT Design Competition. The Liquid Metal system provides a high-level language called Lime and a toolchain targeting FPGAs. Lime allowed us to use standard software development processes for programming, debugging, and performance tuning our FPGA design. We believe such iteration and refinement are far more challenging with low-level languages and design tools commonly used for FPGA development.
Erik R. Altman, Joshua S. Auerbach, David F. Bacon, Ioana Baldini, Perry Cheng, Stephen J. Fink, Rodric M. Rabbah
FPT7
2012 A compiler and runtime for heterogeneous computing
abstract
Heterogeneous systems show a lot of promise for extracting high-performance by combining the benefits of conventional architectures with specialized accelerators in the form of graphics processors (GPUs) and reconfigurable hardware (FPGAs). Extracting this performance often entails programming in disparate languages and models, making it hard for a programmer to work equally well on all aspects of an application. Further, relatively little attention is paid to co-execution---the problem of orchestrating program execution using multiple distinct computational elements that work seamlessly together.
Joshua S. Auerbach, David F. Bacon, Ioana Burcea, Perry Cheng, Stephen J. Fink, Rodric M. Rabbah, Sunil Shukla
DAC6
2012 Compiling a high-level language for GPUs: (via language support for architectures and compilers)
abstract
Languages such as OpenCL and CUDA offer a standard interface for general-purpose programming of GPUs. However, with these languages, programmers must explicitly manage numerous low-level details involving communication and synchronization. This burden makes programming GPUs difficult and error-prone, rendering these powerful devices inaccessible to most programmers.
Christophe Dubach, Perry Cheng, Rodric M. Rabbah, David F. Bacon, Stephen J. Fink
PLDI3
2011 Virtualization of heterogeneous machines hardware description in a synthesizable object-oriented language
abstract
Lime is a new Java-compatible and object-oriented language designed to make programming of reconflgurable hardware significantly more accessible to skilled software developers. Lime programs may run either in software (via Java bytecodes) or in hardware (via behavioral and logic synthesis). This paper illustrates the salient synthesis-oriented features of the language using a photo-mosaic algorithm with inherent bit, pipeline, and data parallelism. The result is a virtual machine abstraction that extends across a heterogeneous architecture comprising a CPU, FPGA, and other computational structures.
Joshua S. Auerbach, David F. Bacon, Perry Cheng, Rodric M. Rabbah, Sunil Shukla
DAC4
2010 MacroSS: macro-SIMDization of streaming applications
Amir Hormati, Yoonseo Choi, Mark Woh, Manjunath Kudlur, Rodric M. Rabbah, Trevor N. Mudge, Scott A. Mahlke
ASPLOS5
2010 FPGA-based combined architecture for stream categorization and intrusion detection
abstract
This paper presents a working solution for the MEMOCODE 2010 design contest. The design presented in this paper is implemented in the Xilinx V5LX330 FPGA as a custom circuit. The solution implements pattern matching logic for all the mandatory and optional patterns while maintaining the required line rate of 500 Mbps.
Sunil Shukla, Rodric M. Rabbah, Martin Vorbach
MEMOCODE2
2010 Lime: a Java-compatible and synthesizable language for heterogeneous architectures
abstract
The halt in clock frequency scaling has forced architects and language designers to look elsewhere for continued improvements in performance. We believe that extracting maximum performance will require compilation to highly heterogeneous architectures that include reconfigurable hardware.
Joshua S. Auerbach, David F. Bacon, Perry Cheng, Rodric M. Rabbah
OOPSLA4
2009 Flextream: Adaptive Compilation of Streaming Applications for Heterogeneous Architectures
abstract
Increasing demand for performance and efficiency has driven the computer industry toward multicore systems. These systems have become the industry standard in almost all segments of the computer market from high-end servers to handheld devices. In order to efficiently use these systems, an extensive amount of research and industry support has been devoted to developing explicitly parallel programming paradigms, such as streaming models, and new compiler techniques. One important challenge that arises in multicore systems is the ability to dynamically adapt a running application to a target architecture in the face of changes in resource availability (e.g., number of cores, available memory or bandwidth). In this paper, we focus on the increasingly important area of streaming computing and introduce Flextream as a flexible compilation framework that can dynamically adapt applications to the changing characteristics of the underlying architecture. We believe this is an important contribution as software developers grapple with the details of parallelism in a rapidly changing architecture landscape. Flextream achieves its goals through a combination of static compilation and dynamic adaptation techniques. Our results indicate that Flextreampsilas approach can achieve high-performance resource allocations that are within an average of 9% of the optimal solution with low overhead for a wide range of streaming applications.
Amir Hormati, Yoonseo Choi, Manjunath Kudlur, Rodric M. Rabbah, Trevor N. Mudge, Scott A. Mahlke
PACT4
2009 A computing origami: folding streams in FPGAs
abstract
Stream processing represents an important class of applications that spans telecommunications, multimedia and the Internet. The implementation of streaming programs in FPGAs has attracted significant attention because of their inherent parallelism and high performance requirements. Languages, tools, and even custom hardware for streaming have been proposed, some of which are commercially available.
Andrei Hagiescu, Weng-Fai Wong, David F. Bacon, Rodric M. Rabbah
DAC4
2008 Optimus: efficient realization of streaming applications on FPGAs
abstract
In this paper, we introduce Optimus: an optimizing synthesis compiler for streaming applications. Optimus compiles programs written in a high level streaming language to either software or hardware implementations. The compiler uses a hierarchical compilation strategy that separates concerns between macro- and micro-functional requirements. Macro-functional concerns address how components (modules) are assembled to implement larger more complex applications. Micro-functional issues deal with synthesis issues of the module internals. Optimus thus allows software developers who lack deep hardware design expertise to transparently leverage the advantages of hardware customization without crossing the semantic gap between high level languages and hardware description languages. Optimus generates streaming hardware that achieves on average 40x speedup over our baseline embedded processor for a fraction of the energy. Additionally, our results show that streaming-specific optimizations can further improve performance by 255% and reduce the area requirements by 16% in average. These designs are competitive with Handel-C implementations for some of the same benchmarks.
Amir Hormati, Manjunath Kudlur, Scott A. Mahlke, David F. Bacon, Rodric M. Rabbah
CASES5
2008 How to Do a Million Watchpoints: Efficient Debugging Using Dynamic Instrumentation
Rodric M. Rabbah, Saman P. Amarasinghe, Larry Rudolph, Weng-Fai Wong
CC2
2008 Liquid Metal: Object-Oriented Programming Across the Hardware/Software Boundary
Shan Shan Huang, Amir Hormati, David F. Bacon, Rodric M. Rabbah
ECOOP4
2007 Ubiquitous Memory Introspection
abstract
Modern memory systems play a critical role in the performance of applications, but a detailed understanding of the application behavior in the memory system is not trivial to attain. It requires time consuming simulations and detailed modeling of the memory hierarchy, often using long address traces. It is increasingly possible to access hardware performance counters to count relevant events in the memory system, but the measurements are coarse-grained and better suited for performance summaries than providing instruction level feedback. The availability of a low cost, online, and accurate methodology for deriving finegrained memory behavior profiles can prove extremely useful for runtime analysis and optimization of programs. This paper presents a new methodology for Ubiquitous Memory Introspection (UMI). It is an online and lightweight methodology that uses fast mini-simulations to analyze short memory access traces recorded from frequently executed code regions. The simulations provide profiling results at varying granularities, down to that of a single instruction or address. UMI naturally complements runtime optimizations and enables new opportunities for online memory specific optimizations. We present a prototype runtime system implementing UMI. The prototype has an average runtime overhead of 14%. This overhead is only 1% more than a state of the art binary instrumentation tool. We used 32 benchmarks, including the full suite of SPEC CPU2000 benchmarks, for evaluation. We show that the mini-simulations accurately reflect the cache performance of two existing memory systems, an Intel Pentium 4 and an AMD Athlon MP (K7). We also demonstrate that UMI predicts delinquent load instructions with an 88% rate of accuracy for applications with a relatively high number of cache misses, and 61% overall. The online profiling results are used at runtime to implement a simple software prefetching strategy that achieves an overall speedup of 64% in the best case.
Rodric M. Rabbah, Saman P. Amarasinghe, Larry Rudolph, Weng-Fai Wong
CGO2
2006 MPEG-2 decoding in a stream programming language
abstract
Image and video codecs are prevalent in multimedia devices, ranging from embedded systems, to desktop computers, to high-end servers such as HDTV editing consoles. It is not uncommon however that developers create and customize separate coder and decoder implementations for each of the architectures they target. This practice is time consuming and error prone, leading to code that is neither malleable nor portable. This paper describes an implementation of the MPEG-2 decoder using the StreamIt programming language. StreamIt is an architecture-independent stream language that aims to improve programmer productivity, while concomitantly exposing the inherent parallelism and communication topology of the application. The paper shows that MPEG is a good match for the streaming programming model and illustrates the malleability of the implementation using a simple modification to the decoder to support alternate color compression formats. StreamIt allows for modular application development, which increases code reuse, and reduces the complexity of the debugging process since stream components can be verified independently. This in turn leads to greater programmer productivity.
Matthew Drake, Henry Hoffmann, Rodric M. Rabbah, Saman P. Amarasinghe
IPDPS3
2005 Cache aware optimization of stream programs
abstract
Effective use of the memory hierarchy is critical for achieving high performance on embedded systems. We focus on the class of streaming applications, which is increasingly prevalent in the embedded domain. We exploit the widespread parallelism and regular communication patterns in stream programs to formulate a set of cache aware optimizations that automatically improve instruction and data locality. Our work is in the context of the Synchronous Dataflow model, in which a program is described as a graph of independent actors that communicate over channels. The communication rates between actors are known at compile time, allowing the compiler to statically model the caching behavior.We present three cache aware optimizations: 1) execution scaling, which judiciously repeats actor executions to improve instruction locality, 2) cache aware fusion, which combines adjacent actors while respecting instruction cache constraints, and 3) scalar replacement, which converts certain data buffers into a sequence of scalar variables that can be register allocated. The optimizations are founded upon a simple and intuitive model that quantifies the temporal locality for a sequence of actor executions. Our implementation of cache aware optimizations in the StreamIt compiler yields a 249% average speedup (over unoptimized code) for our streaming benchmark suite on a StrongARM 1110 processor. The optimizations also yield a 154% speedup on a Pentium 3 and a 152% speedup on an Itanium 2.
Janis Sermulins, William Thies, Rodric M. Rabbah, Saman P. Amarasinghe
LCTES3
2005 Exploiting Vector Parallelism in Software Pipelined Loops
abstract
An emerging trend in processor design is the addition of short vector instructions to general-purpose and embedded ISAs. Frequently, these extensions are employed using traditional vectorization technology first developed for supercomputers. In contrast, scalar hardware is typically targeted using ILP techniques such as software pipelining. This paper presents a novel approach for exploiting vector parallelism in software pipelined loops. The proposed methodology (i) lowers the burden on the scalar resources by offloading computation to the vector functional units, (ii) explicitly manages communication of operands between scalar and vector instructions, (in) naturally handles misaligned vector memory operations, and (iv) partially (or fully) inhibits the optimization when vectorization will decrease performance. Our approach results in better resource utilization and allows for software pipelining with shorter initiation intervals. The proposed optimization is applied in the compiler backend, where vectorization decisions are more amenable to cost analysis. This is unique in that traditional vectorization optimizations are usually carried out at the statement level. Although our technique most naturally complements statically scheduled machines, we believe it is applicable to any architecture that tightly integrates support for instruction and data level parallelism. We evaluate our methodology using nine SPEC FP benchmarks. In comparison to software pipelining, our approach achieves a maximum speedup of 1.38times, with an average of 1.11times
Samuel Larsen, Rodric M. Rabbah, Saman P. Amarasinghe
MICRO2
2005 Programming by sketching for bit-streaming programs
Armando Solar-Lezama, Rodric M. Rabbah, Rastislav Bodík, Kemal Ebcioglu
PLDI2
2005 Teleport messaging for distributed stream programs
abstract
In this paper, we develop a new language construct to address one of the pitfalls of parallel programming: precise handling of events across parallel components. The construct, termed teleport messaging, uses data dependences between components to provide a common notion of time in a parallel system. Our work is done in the context of the Synchronous Dataflow (SDF) model, in which computation is expressed as a graph of independent components (or actors) that communicate in regular patterns over data channels. We leverage the static properties of SDF to compute a stream dependence function, SDEP, that compactly describes the ordering constraints between actor executions.Teleport messaging utilizes SDEP to provide powerful and precise event handling. For example, an actor A can specify that an event should be processed by a downstream actor B as soon as B sees the "effects" of the current execution of A. We argue that teleport messaging improves readability and robustness over existing practices. We have implemented messaging as part of the StreamIt compiler, with a backend for a cluster of workstations. As teleport messaging exposes optimization opportunities to the compiler, it also results in a 49% performance improvement for a software radio benchmark.
William Thies, Michal Karczmarek, Janis Sermulins, Rodric M. Rabbah, Saman P. Amarasinghe
PPoPP4
2004 Compiler orchestrated prefetching via speculation and predication
abstract
This paper introduces a compiler orchestrated prefetching system as a unified framework geared toward ameliorating the gap between processing speeds and memory access latencies. We focus the scope of the optimization on specific subsets of the program dependence graph that succinctly characterize the memory access pattern of both regular array-based applications and irregular pointer-intensive programs. We illustrate how program embedded precomputation via speculative execution can accurately predict and effectively prefetch future memory references with negligible overhead. The proposed techniques reduce the total running time of seven SPEC benchmarks and two OLDEN benchmarks by 27% on an Itanium 2 processor. The improvements are in addition to several state-of-the-art optimizations including software pipelining and data prefetching. In addition, we use cycle-accurate simulations to identify important and lightweight architectural innovations that further mitigate the memory system bottleneck. In particular, we focus on the notoriously challenging class of pointer-chasing applications, and demonstrate how they may benefit from a novel scheme of it sentineled prefetching. Our results for twelve SPEC benchmarks demonstrate that 45% of the processor stalls that are caused by the memory system are avoidable. The techniques in this paper can effectively mask long memory latencies with little instruction overhead, and can readily contribute to the performance of processors today.
Rodric M. Rabbah, Hariharan Sandanagobalane, Mongkol Ekpanyapong, Weng-Fai Wong
ASPLOS1
2003 Data remapping for design space optimization of embedded memory systems
abstract
In this article, we present a novel linear time algorithm for data remapping , that is, (i) lightweight; (ii) fully automated; and (iii) applicable in the context of pointer-centric programming languages with dynamic memory allocation support. All previous work in this area lacks one or more of these features. We proceed to demonstrate a novel application of this algorithm as a key step in optimizing the design of an embedded memory system. Specifically, we show that by virtue of locality enhancements via data remapping, we may reduce the memory subsystem needs of an application by 50%, and hence concomitantly reduce the associated costs in terms of size, power, and dollar-investment (61%). Such a reduction overcomes key hurdles in designing high-performance embedded computing solutions. Namely, memory subsystems are very desirable from a performance standpoint, but their costs have often limited their use in embedded systems. Thus, our innovative approach offers the intriguing possibility of compilers playing a significant role in exploring and optimizing the design space of a memory subsystem for an embedded design. To this end and in order to properly leverage the improvements afforded by a compiler optimization, we identify a range of measures for quantifying the cost-impact of popular notions of locality, prefetching, regularity of memory access, and others . The proposed methodology will become increasingly important, especially as the needs for application specific embedded architectures become prevalent. In addition, we demonstrate the wide applicability of data remapping using several existing microprocessors, such as the Pentium and UltraSparc. Namely, we show that remapping can achieve a performance improvement of 20% on the average. Similarly, for a parametric research HPL-PD microprocessor, which characterizes the new Itanium machines, we achieve a performance improvement of 28% on average. All of our results are achieved using applications from the DIS, Olden and SPEC2000 suites of integer and floating point benchmarks.
Rodric M. Rabbah, Krishna V. Palem
ACM Trans. Embed. Comput. Syst.1
2002 PD-XML: extensible markup language for processor description
abstract
This paper introduces PD-XML, a meta-language for describing instruction processors in general and with an emphasis on embedded processors, with the specific aim of enabling their rapid prototyping, evaluation and eventual design and implementation. PD-XML is not specific to any one architecture, compiler or simulation environment and hence provides greater flexibility than related machine description methodologies. We demonstrate how PD-XML can be interfaced to existing description methodologies and tool-flows. In particular we show how PD-XML specifications can be translated into appropriate machine descriptions for the parametric HPL-PD VLIW processor, and for the Flexible Instruction Processor (FIP) approach targeting reconfigurable implementations.
Shay Ping Seng, Krishna V. Palem, Rodric M. Rabbah, Weng-Fai Wong, Wayne Luk, Peter Y. K. Cheung
FPT3