VLDB 2026 Research / reviewers in the wild / expert
Robert S. Cohn
dblp:c/RobertSCohn · also Robert Cohn
· DBLP profile ↗
20ranked-venue papers
4as first author
1since 2021 · last 2024
0009-0005-7234-4886ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 11 · 1 first-authorArtificial intelligence and machine learning · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Processor architecture and microarchitecture · 64% Performance modeling and evaluation · 32% Interconnection networks and networks-on-chip · 2% | |
| Artificial intelligence
1 paper |
Planning, search and constraint satisfaction · 67% Multi-agent systems · 33% | |
| Software engineering, system software, and programming languages
5 papers |
Program analysis · 65% Compilers and program optimization · 32% Software maintenance and evolution · 3% |
Topics — the 20 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
branch prediction |
0.2 | 2 | 2009 | Virtual Program Counter (VPC) Prediction: Very Low Cost Indirect Branch Prediction Using Conditional Branch Prediction Hardware · IEEE Trans. Computers 2009 VPC prediction: reducing the cost of indirect branches via hardware-based dynamic devirtualization · ISCA 2007 |
Processor architecture and microarchitecture › branch prediction
indirect branch prediction |
0.2 | 2 | 2009 | Virtual Program Counter (VPC) Prediction: Very Low Cost Indirect Branch Prediction Using Conditional Branch Prediction Hardware · IEEE Trans. Computers 2009 VPC prediction: reducing the cost of indirect branches via hardware-based dynamic devirtualization · ISCA 2007 |
Knowledge, reasoning and agents › Multi-agent systems
human-agent interaction |
0.1 | 1 | 2011 | Comparing Action-Query Strategies in Semi-Autonomous Agents · AAAI 2011 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
planning under uncertainty |
0.1 | 1 | 2011 | Comparing Action-Query Strategies in Semi-Autonomous Agents · AAAI 2011 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › decision making under uncertainty
value of information |
0.1 | 1 | 2011 | Comparing Action-Query Strategies in Semi-Autonomous Agents · AAAI 2011 |
Performance modeling and evaluation
workload characterization |
0.1 | 2 | 2004 | Pinpointing Representative Portions of Large Intel® Itanium® Programs with Dynamic Instrumentation · MICRO 2004 Code layout optimizations for transaction processing workloads · ISCA 2001 |
Program analysis › dynamic analysis › instrumentation
binary instrumentation |
0.1 | 1 | 2005 | Pin: building customized program analysis tools with dynamic instrumentation · PLDI 2005 |
Program analysis › dynamic analysis
dynamic instrumentation |
0.1 | 1 | 2005 | Pin: building customized program analysis tools with dynamic instrumentation · PLDI 2005 |
Performance modeling and evaluation › program instrumentation
dynamic instrumentation |
0.0 | 1 | 2004 | Pinpointing Representative Portions of Large Intel® Itanium® Programs with Dynamic Instrumentation · MICRO 2004 |
Compilers and program optimization
code layout optimization |
0.0 | 1 | 2001 | Code layout optimizations for transaction processing workloads · ISCA 2001 |
Performance modeling and evaluation
profiling |
0.0 | 1 | 2005 | Pin: building customized program analysis tools with dynamic instrumentation · PLDI 2005 |
Compilers and program optimization
compiler optimization |
0.0 | 1 | 1996 | Hot Cold Optimization of Large Windows/NT Applications · MICRO 1996 |
Transaction processing and concurrency control
OLTP |
0.0 | 1 | 2001 | Code layout optimizations for transaction processing workloads · ISCA 2001 |
Parallel and multicore computing
parallel architecture |
0.0 | 2 | 1990 | Warp: an integrated solution of high-speed parallel computing · SC 1988 Supporting Systolic and Memory Communciation in iWarp · ISCA 1990 |
Interconnection networks and networks-on-chip
interprocessor communication |
0.0 | 1 | 1990 | Supporting Systolic and Memory Communciation in iWarp · ISCA 1990 |
Compilers and program optimization
instruction scheduling |
0.0 | 1 | 1989 | Architecture and Compiler Tradeoffs for a Long Instruction Word Microprocessor · ASPLOS 1989 |
Processor architecture and microarchitecture › instruction-level parallelism
VLIW |
0.0 | 1 | 1989 | Architecture and Compiler Tradeoffs for a Long Instruction Word Microprocessor · ASPLOS 1989 |
High-performance computing
distributed memory systems |
0.0 | 1 | 1988 | Warp: an integrated solution of high-speed parallel computing · SC 1988 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 1 | 1989 | Architecture and Compiler Tradeoffs for a Long Instruction Word Microprocessor · ASPLOS 1989 |
Interconnection networks and networks-on-chip
low-latency communication |
0.0 | 1 | 1988 | Warp: an integrated solution of high-speed parallel computing · SC 1988 |
Methods — techniques the papers use, named apart from their topics
hardware-based dynamic devirtualization · 0.1uncertainty minimization · 0.1expected gain maximization · 0.1liveness analysis · 0.1instruction scheduling · 0.1inlining · 0.1dynamic compilation · 0.1virtual program counter prediction · 0.1performance and energy evaluation · 0.1register reallocation · 0.1register re-allocation · 0.1pin · 0.0hardware performance counters · 0.0profile-guided optimization · 0.0compiler optimization · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Distributed Ranges: A Model for Distributed Data Structures, Algorithms, and ViewsabstractData structures and algorithms are essential building blocks for programs, and distributed data structures, which automatically partition data across multiple memory locales, are essential to writing high-level parallel programs. While many projects have designed and implemented C++ distributed data structures and algorithms, there has not been widespread adoption of an interoperable model allowing algorithms and data structures from different libraries to work together. This paper introduces distributed ranges, which is a model for building generic data structures, views, and algorithms. A distributed range extends a C++ range, which is an iterable sequence of values, with a concept of segmentation, thus exposing how the distributed range is partitioned over multiple memory locales. Distributed data structures provide this distributed range interface, which allows them to be used with a collection of generic algorithms implemented using the distributed range interface. The modular nature of the model allows for the straightforward implementation of distributed views, which are lightweight objects that provide a lazily evaluated view of another range. Views can be composed together recursively and combined with algorithms to implement computational kernels using efficient, flexible, and high-level standard C++ primitives. We evaluate the distributed ranges model by implementing a set of standard concepts and views as well as two execution runtimes, a multi-node, MPI-based runtime and a single-process, multi-GPU runtime. We demonstrate that high-level algorithms implemented using generic, high-level distributed ranges can achieve performance competitive with highly-tuned, expert-written code. Benjamin Brock, Robert S. Cohn, Suyash Bakshi, Tuomas Kärnä, Jeongnim Kim, Mateusz Nowak 0001, Lukasz Slusarczyk, Kacper Stefanski, Timothy G. Mattson |
ICS | 2 |
| 2016 | Simulation and Analysis Engine for Scale-Out WorkloadsabstractWe introduce a system-level Simulation and Analysis Engine (SAE) framework based on dynamic binary instrumentation for fine-grained and customizable instruction-level introspection of everything that executes on the processor. SAE can instrument the BIOS, kernel, drivers, and user processes. It can also instrument multiple systems simultaneously using a single instrumentation interface, which is essential for studying scale-out applications. SAE is an x86 instruction set simulator designed specifically to enable rapid prototyping, evaluation, and validation of architectural extensions and program analysis tools using its flexible APIs. It is fast enough to execute full platform workloads---a modern operating system can boot in a few minutes---thus enabling research, evaluation, and validation of complex functionalities related to multicore configurations, virtualization, security, and more. To reach high speeds, SAE couples tightly with a virtual platform and employs both a just-in-time (JIT) compiler that helps simulate simple instructions efficiently and a fast interpreter for simulating new or complex instructions. We describe SAE's architecture and instrumentation engine design and show the framework's usefulness for single- and multi-system architectural and program analysis studies. Nadav Chachmon, Daniel Richins, Robert S. Cohn, Magnus Christensson, Wenzhi Cui, Vijay Janapa Reddi |
ICS | 3 |
| 2014 | Characterizing EVOI-Sufficient k-Response Query Sets in Decision ProblemsabstractIn finite decision problems where an agent can query its human user to obtain information about its environment before acting, a query’s usefulness is in terms of its Expected Value of Information (EVOI). The usefulness of a query set is similarly measured in terms of the EVOI of the queries it contains. When the only constraint on what queries can be asked is that they have exactly k possible responses (with k \ge 2), we show that the set of k-response decision queries (which ask the user to select his/her preferred decision given a choice of k decisions) is EVOI-Sufficient, meaning that no single k-response query can have higher EVOI than the best single k-response decision query for any decision problem. When multiple queries can be asked before acting, we provide a negative result that shows the set of depth-n query trees constructed from k-response decision queries is not EVOI-Sufficient. However, we also provide a positive result that the set of depth-n query trees constructed from k-response decision-set queries, which ask the user to select from among k sets of decisions as to which set contains the best decision, is EVOI-Sufficient. We conclude with a discussion and analysis of algorithms that draws on a connection to other recent work on decision-theoretic knowledge elicitation. Robert S. Cohn, Satinder Singh 0001, Edmund H. Durfee |
AISTATS | 1 |
| 2011 | Comparing Action-Query Strategies in Semi-Autonomous AgentsabstractWe consider settings in which a semi-autonomous agent has uncertain knowledge about its environment, but can ask what action the human operator would prefer taking in the current or in a potential future state. Asking queries can improve behavior, but if queries come at a cost (e.g., due to limited operator attention), the value of each query should be maximized. We compare two strategies for selecting action queries: 1) based on myopically maximizing expected gain in long-term value, and 2) based on myopically minimizing uncertainty in the agent's policy representation. We show empirically that the first strategy tends to select more valuable queries, and that a hybrid method can outperform either method alone in settings with limited computation. Robert S. Cohn, Edmund H. Durfee, Satinder Singh 0001 |
AAAI | 1 |
| 2011 | Portable trace compression through instruction interpretationabstractExecution traces are a useful tool in studying processor and program behavior. However, the amount of information that needs to be stored makes them impractical in uncompressed form. This is especially true for full-state traces that can capture up to kilobytes of processor state for every instruction. In this paper we present Zcompr-a compression scheme that allows practical usage of full-state traces that are billions of instructions long. It allows complete state reproducibility, sufficient even for validation purposes, that is fully portable between different operating systems and host platforms. The compression scheme exploits the general similarity between compression and prediction. A simplified functional simulator is used to predict instruction effects in a repeatable manner. Its predictions can be used to reproduce those effects at decompression time, limiting the amount of information that needs to be stored per instruction. Final trace densities achieved by our scheme are on the order of two bits per instruction, with typical decompression speeds of 300 KIPS. Svilen Kanev, Robert S. Cohn |
ISPASS | 2 |
| 2010 | Dynamic program analysis of Microsoft Windows applicationsabstractSoftware instrumentation is a powerful and flexible technique for analyzing the dynamic behavior of programs. By inserting extra code in an application, it is possible to study the performance and correctness of programs and systems. Pin is a software system that performs run-time binary instrumentation of unmodified applications. Pin provides an API for writing custom instrumentation, enabling its use in a wide variety of performance analysis tasks such as workload characterization, program tracing, cache modeling, and simulation. Most of the prior work on instrumentation systems has focused on executing Unix applications, despite the ubiquity and importance of Windows applications. This paper identifies the Windows-specific obstacles for implementing a process-level instrumentation system, describes a comprehensive, robust solution, and discusses some of the alternatives. The challenges lie in managing the kernel/application transitions, injecting the runtime agent into the process, and isolating the instrumentation from the application. We examine Pin's overhead on typical Windows applications being instrumented with simple tools up to commercial program analysis products. The biggest factor affecting performance is the type of analysis performed by the tool. While the proprietary nature of Windows makes measurement and analysis difficult, Pin opens the door to understanding program behavior. Alex Skaletsky, Tevi Devor, Nadav Chachmon, Robert S. Cohn, Kim M. Hazelwood, Vladimir Vladimirov, Moshe Bach |
ISPASS | 4 |
| 2009 | Scalable support for multithreaded applications on dynamic binary instrumentation systemsabstractDynamic binary instrumentation systems are used to inject or modify arbitrary instructions in existing binary applications; several such systems have been developed over the past decade. Much of the literature describing the internal architecture and performance of these systems has focused on executing single-threaded guest applications. In this paper, we discuss the specific design decisions necessary for supporting large, multithreaded applications on JIT-based dynamic instrumentation systems. While implementing a working solution for multithreading is straightforward, providing a system that scales in terms of memory and performance is much more intricate. We highlight the design decisions in the latest version of the Pin dynamic instrumentation system, including the just-in-time compiler, the emulator, and the code cache. The overall design strives to provide scalable performance and memory footprints on modern applications. Kim M. Hazelwood, Greg Lueck, Robert S. Cohn |
ISMM | 3 |
| 2009 | Virtual Program Counter (VPC) Prediction: Very Low Cost Indirect Branch Prediction Using Conditional Branch Prediction HardwareabstractIndirect branches have become increasingly common in modular programs written in modern object-oriented languages and virtual-machine-based runtime systems. Unfortunately, the prediction accuracy of indirect branches has not improved as much as that of conditional branches. Furthermore, previously proposed indirect branch predictors usually require a significant amount of extra hardware storage and complexity, which makes them less attractive to implement. This paper proposes a new technique for handling indirect branches, called Virtual Program Counter (VPC) prediction. The key idea of VPC prediction is to use the existing conditional branch prediction hardware to predict indirect branch targets, avoiding the need for a separate storage structure. Our comprehensive evaluation shows that VPC prediction improves average performance by 26.7 percent and reduces average energy consumption by 19 percent compared to a commonly used branch target buffer based predictor on 12 indirect branch intensive C/C++ applications. Moreover, VPC prediction improves the average performance of the full set of object-oriented Java DaCapo applications by 21.9 percent, while reducing their average energy consumption by 22 percent. We show that VPC prediction can be used with any existing conditional branch prediction mechanism and that the accuracy of VPC prediction improves when a more accurate conditional branch predictor is used. Hyesoon Kim, José A. Joao, Onur Mutlu, Chang Joo Lee, Yale N. Patt, Robert S. Cohn |
IEEE Trans. Computers | 6 |
| 2007 | Persistent Code Caching: Exploiting Code Reuse Across Executions and ApplicationsabstractRun-time compilation systems are challenged with the task of translating a program's instruction stream while maintaining low overhead. While software managed code caches are utilized to amortize translation costs, they are ineffective for programs with short run times or large amounts of cold code. Such program characteristics are prevalent in real-life computing environments, ranging from graphical user interface (GUI) programs to large-scale applications such as database management systems. Persistent code caching addresses these issues. It is described and evaluated in an industry-strength dynamic binary instrumentation system - Pin. The proposed approach improves the intra-execution model of code reuse by storing and reusing translations across executions, thereby achieving inter-execution persistence. Dynamically linked programs leverage inter-application persistence by using persistent translations of library code generated by other programs. New translations discovered across executions are automatically accumulated into the persistent code caches, thereby improving performance over time. Inter-execution persistence improves the performance of GUI applications by nearly 90%, while inter-application persistence achieves a 59% improvement. In more specialized uses, the SPEC2K INT benchmark suite experiences a 26% improvement under dynamic binary instrumentation. Finally, a 400% speedup is achieved in translating the Oracle database in a regression testing environment Vijay Janapa Reddi, Daniel A. Connors, Robert S. Cohn, Michael D. Smith 0001 |
CGO | 3 |
| 2007 | VPC prediction: reducing the cost of indirect branches via hardware-based dynamic devirtualizationabstractIndirect branches have become increasingly common in modular programs written in modern object-oriented languages and virtual machine based runtime systems. Unfortunately, the prediction accuracy of indirect branches has not improved as much as that of conditional branches. Furthermore, previously proposed indirect branch predictors usually require a significant amount of extra hardware storage and complexity, which makes them less attractive to implement. Hyesoon Kim, José A. Joao, Onur Mutlu, Chang Joo Lee, Yale N. Patt, Robert S. Cohn |
ISCA | 6 |
| 2006 | A Cross-Architectural Interface for Code Cache ManipulationabstractSoftware code caches help amortize the overhead of dynamic binary transformation by enabling reuse of transformed code. Since code caches contain a potentially-altered copy of every instruction that executes, run-time access to a code cache can be a very powerful opportunity. Unfortunately, current research infrastructures lack the ability to model and direct code caching, and as a result, past code cache investigations have required access to the source code of the binary transformation system. This paper presents a code cache-aware interface to the pin dynamic instrumentation system. While a program executes, our interface allows a user to inspect the code cache, receive callbacks when key events occur, and manipulate the code cache contents at will. We demonstrate the utility of this interface on four architectures (IA32, EM64T, IPF, XScale) and present several tools written using our API. These tools include a self-modifying code handler, a two-phase instrumentation analyzer, a code cache visualizer, and custom code cache replacement policies. We also show that tools written using our interface have comparable performance to direct, source-level implementations. Both our interface and sample open-source tools that utilize the interface have been incorporated into the standard distribution of the pin dynamic instrumentation engine, which has been downloaded over 5,000 times in 18 months. Kim M. Hazelwood, Robert S. Cohn |
CGO | 2 |
| 2005 | Pin: building customized program analysis tools with dynamic instrumentationabstractRobust and powerful software instrumentation tools are essential for program analysis tasks such as profiling, performance evaluation, and bug detection. To meet this need, we have developed a new instrumentation system called Pin. Our goals are to provide easy-to-use, portable, transparent, and efficient instrumentation. Instrumentation tools (called Pintools) are written in C/C++ using Pin's rich API. Pin follows the model of ATOM, allowing the tool writer to analyze an application at the instruction level without the need for detailed knowledge of the underlying instruction set. The API is designed to be architecture independent whenever possible, making Pintools source compatible across different architectures. However, a Pintool can access architecture-specific details when necessary. Instrumentation with Pin is mostly transparent as the application and Pintool observe the application's original, uninstrumented behavior. Pin uses dynamic compilation to instrument executables while they are running. For efficiency, Pin uses several techniques, including inlining, register re-allocation, liveness analysis, and instruction scheduling to optimize instrumentation. This fully automated approach delivers significantly better instrumentation performance than similar tools. For example, Pin is 3.3x faster than Valgrind and 2x faster than DynamoRIO for basic-block counting. To illustrate Pin's versatility, we describe two Pintools in daily use to analyze production software. Pin is publicly available for Linux platforms on four architectures: IA32 (32-bit x86), EM64T (64-bit x86), Itanium®, and ARM. In the ten months since Pin 2 was released in July 2004, there have been over 3000 downloads from its website. Chi-Keung Luk, Robert S. Cohn, Robert Muth, Harish Patil, Artur Klauser, P. Geoffrey Lowney, Steven Wallace, Vijay Janapa Reddi, Kim M. Hazelwood |
PLDI | 2 |
| 2004 | Ispike: A Post-link Optimizer for the Intel®Itanium®ArchitectureabstractIspike is a post-link optimizer developed for the Intel/spl reg/ Itanium Processor Family (IPF) processors. The IPF architecture poses both opportunities and challenges to post-link optimizations. IPF offers a rich set of performance counters to collect detailed profile information at a low cost, which is essential to post-link optimization being practical. At the same time, the predication and bundling features on IPF make post-link code transformation more challenging than on other architectures. In Ispike, we have implemented optimizations like code layout, instruction prefetching, data layout, and data prefetching that exploit the IPF advantages, and strategies that cope with the IPF-specific challenges. Using SPEC CINT2000 as benchmarks, we show that Ispike improves performance by as much as 40% on the ltanium/spl reg/2 processor, with average improvement of 8.5% and 9.9% over executables generated by the Intel/spl reg/ Electron compiler and by the Gcc compiler, respectively. We also demonstrate that statistical profiles collected via IPF performance counters and complete profiles collected via instrumentation produce equal performance benefit, but the profiling overhead is significantly lower for performance counters. Chi-Keung Luk, Robert Muth, Harish Patil, Robert S. Cohn, P. Geoffrey Lowney |
CGO | 4 |
| 2004 | Pinpointing Representative Portions of Large Intel® Itanium® Programs with Dynamic InstrumentationabstractDetailed modeling of the performance of commercial applications is difficult. The applications can take a very long time to run on real hardware and it is impractical to simulate them to completion on performance models. Furthermore, these applications have complex execution environments that cannot easily be reproduced on a simulator, making porting the applications to simulators difficult. We attack these problems using the well-known SimPoint methodology to find representative portions of an application to simulate, and a dynamic instrumentation framework called Pin to avoid porting altogether. Our system uses dynamic instrumentation instead of simulation to find representative portions - called Pin-Points - for simulation. We have developed a toolkit that automatically detects PinPoints, validates whether they are representative using hardware performance counters, and generates traces for large Itanium® programs. We compared SimPoint-based selection to random selection of simulation points. We found for 95% of the SPEC2000 programs we tested, the PinPoints prediction was within 8% of the actual whole-program CPI, as opposed to 18% for random selection. We measure the end-to-end error, comparing real hardware to a performance model, and have a simple and efficient methodology to determine the step that introduced the error. Finally, we evaluate the system in the context of multiple configurations of real hardware, commercial applications, and industrial-strength performance models to understand the behavior of a complete and practical workload collection system. We have successfully used our system with many commercial Itanium® programs, some running for trillions of instructions, and have used the resulting traces for predicting performance of those applications on future Itanium processors. Harish Patil, Robert S. Cohn, Mark Charney, Rajiv Kapoor, Andrew Sun, Anand Karunanidhi |
MICRO | 2 |
| 2002 | Profile-guided post-link stride prefetchingabstractData prefetching is an e ective approach to addressing the memory latency problem. While a few processors have implemented hardware-based data prefetching, the majority of modern processors support data-prefetch instructions and rely on compilers to automatically insert prefetches. However, most prefetching schemes in commercial compilers suffer from two limitations: (1) the source code must be available before prefetching can be applied, and (2) these prefetching schemes target only loops with statically-known strided accesses. In this study, we broaden the scope of softwarecontrolled prefetching by addressing the above two limitations. We use pro ling to discover strided accesses that frequently occur during program execution but are not determinable by the compiler. We then use the strides discovered to insert prefetches into the executable directly, without the need for re-compilation. Performance evaluation was done on an Alpha 21264-based system with a 64KB data cache and an 8MB secondary cache. We nd that even with such large caches, our technique o ers speedups ranging from 3% to 56 % in 11 out of the 26 SPEC2000 benchmarks. Our technique has been incorporated into Pixie and Spike, two products in Compaq's Tru64 Unix. Chi-Keung Luk, Robert Muth, Harish Patil, Richard Weiss 0001, P. Geoffrey Lowney, Robert S. Cohn |
ICS | 6 |
| 2001 | Code layout optimizations for transaction processing workloadsabstractCommercial applications such as databases and Web servers constitute the most important market segment for high-performance servers. Among these applications, on-line transaction processing (OLTP) workloads provide a challenging set of requirements for system designs since they often exhibit inefficient executions dominated by a large memory stall component. This behavior arises from large instruction and data footprints and high communication miss rates. A number of recent studies have characterized the behavior of commercial workloads and proposed architectural features to improve their performance. However, there has been little research on the impact of software and compiler-level optimizations for improving the behavior of such workloads. Alex Ramírez, Luiz André Barroso, Kourosh Gharachorloo, Robert S. Cohn, Josep Lluís Larriba-Pey, P. Geoffrey Lowney, Mateo Valero |
ISCA | 4 |
| 1996 | Hot Cold Optimization of Large Windows/NT ApplicationsabstractA dynamic instruction trace often contains many unnecessary instructions that are required only by the unexecuted portion of the program. Hot-cold optimization (HCO) is a technique that realizes this performance opportunity. HCO uses profile information to partition each routine into frequently executed (hot) and infrequently executed (cold) parts. Unnecessary operations in the hot portion are removed and compensation code is added on transitions from hot to cold as needed. We evaluate HCO on a collection of large Windows/NT applications. HCO is most effective on the programs that are call intensive and have flat profiles, providing a 3-8% reduction in path length beyond conventional optimization. Robert S. Cohn, P. Geoffrey Lowney |
MICRO | 1 |
| 1990 | Supporting Systolic and Memory Communciation in iWarpabstractiWarp is a parallel architecture developed jointly by Carnegie Mellon University and Intel Corporation. The iWarp communication system supports two widely used interprocessor communication styles: memory communication and systolic communication. This paper describes the rationale, architecture, and implementation for the iWarp communication system. Shekhar Borkar, Robert S. Cohn, George W. Cox, Thomas R. Gross, H. T. Kung 0001, Monica S. Lam, Margie Levine, Brian Moore 0004, Wire Moore, Craig Peterson, Jim Susman, Jim Sutton, John Urbanski, Jon A. Webb |
ISCA | 2 |
| 1989 | Architecture and Compiler Tradeoffs for a Long Instruction Word MicroprocessorabstractA very long instruction word (VLIW) processor exploits parallelism by controlling multiple operations in a single instruction word. This paper describes the architecture and compiler tradeoffs in the design of iWarp, a VLIW single-chip microprocessor developed in a joint project with Intel Corp. The iWarp processor is capable of specifying up to nine operations in an instruction word and has a peak performance of 20 million floating-point operations and 20 million integer operations per second. An optimizing compiler has been constructed and used as a tool to evaluate the different architectural proposals in the development of iWarp. We present here the analysis and compiler optimizations for those architectural features that address two key issues in the design of a VLIW microprocessor: code density and a streamlined execution cycle. We support the results of our analysis with performance data for the Livermore Loops and a selection of programs from the LINPACK library. Robert S. Cohn, Thomas R. Gross, Monica S. Lam, P. S. Tseng |
ASPLOS | 1 |
| 1988 | Warp: an integrated solution of high-speed parallel computingabstractA description is given of the iWarp architecture and how it supports various communication models and system configurations. The heart of an iWarp system is the iWarp component: a single-chip processor that requires only the addition of memory chips to form a complete system building block, called the iWarp cell. Each iWarp component contains both a powerful computation engine that runs at 20 MFLOPS (million floating-point operations per second) and a high-throughput (320 Mb/s), low-latency (100-150-ns) communication engine for interfacing with other iWarp cells. Because of their strong computation and communication capabilities, the iWarp components provide a versatile building block for high-performance parallel systems ranging from special-purpose systolic arrays to general-purpose distributed memory computers. They can support both fine-grain parallel and coarse-grain distributed computation models simultaneously in the same system. The initial iWarp demonstration system consists of an 8*8 torus of iWarp cells, delivering more than 1.2 GFLOP (billions of FLOPS). It can be expanded to include up to 1024 cells.> Shekhar Borkar, Robert S. Cohn, George W. Cox, Sha Gleason, Thomas R. Gross |
SC | 2 |