VLDB 2026 Research / reviewers in the wild / expert
Mike Schlansker
dblp:00/8857 · also Michael S. Schlansker
· DBLP profile ↗
26ranked-venue papers
8as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 4 first-authorComputer networks · 5 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
15 papers |
Processor architecture and microarchitecture · 70% Parallel and multicore computing · 15% Electronic design automation · 7% | |
| Computer networks
2 papers |
Internet architecture and protocols · 50% Software-defined and programmable networks · 50% | |
| Software engineering, system software, and programming languages
12 papers |
Compilers and program optimization · 94% Program analysis · 6% |
Topics — the 30 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Software-defined and programmable networks › SDN-based network management
SDN failure recovery |
0.1 | 1 | 2012 | CORONET: Fault tolerance for Software Defined Networks · ICNP 2012 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.1 | 7 | 1999 | Control CPR: A Branch Height Reduction Optimization for EPIC Architectures · PLDI 1999 Analysis Techniques for Predicated Code · MICRO 1996 Profile-driven Instruction Level Parallel Scheduling with Application to Super Blocks · MICRO 1996 |
Internet architecture and protocols › local area network
ethernet |
0.1 | 1 | 2007 | High-performance ethernet-based communications for future multi-core processors · SC 2007 |
Internet architecture and protocols
network interface card |
0.1 | 1 | 2007 | High-performance ethernet-based communications for future multi-core processors · SC 2007 |
Compilers and program optimization
instruction scheduling |
0.1 | 5 | 1996 | Profile-driven Instruction Level Parallel Scheduling with Application to Super Blocks · MICRO 1996 Critical path reduction for scalar programs · MICRO 1995 Spill-free parallel scheduling of basic blocks · MICRO 1995 |
Parallel and multicore computing › synchronization
barrier synchronization |
0.1 | 1 | 2006 | Exploiting Fine-Grained Data Parallelism with Chip Multiprocessors and Fast Barriers · MICRO 2006 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 1 | 2006 | Exploiting Fine-Grained Data Parallelism with Chip Multiprocessors and Fast Barriers · MICRO 2006 |
Compilers and program optimization
register allocation |
0.0 | 3 | 1996 | Global Predicate Analysis and Its Application to Register Allocation · MICRO 1996 Spill-free parallel scheduling of basic blocks · MICRO 1995 Register Allocation for Software Pipelined Loops · PLDI 1992 |
Compilers and program optimization › instruction scheduling
instruction-level parallelism |
0.0 | 1 | 2001 | Compiling for EPIC architectures · Proc. IEEE 2001 |
Electronic design automation
high-level synthesis |
0.0 | 1 | 2001 | Bitwidth cognizant architecture synthesis of custom hardwareaccelerators · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2001 |
Compilers and program optimization › compiler optimization
branch optimization |
0.0 | 1 | 1999 | Control CPR: A Branch Height Reduction Optimization for EPIC Architectures · PLDI 1999 |
Interconnection networks and networks-on-chip
cluster interconnect |
0.0 | 1 | 2007 | High-performance ethernet-based communications for future multi-core processors · SC 2007 |
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution |
0.0 | 2 | 1996 | Analysis Techniques for Predicated Code · MICRO 1996 Global Predicate Analysis and Its Application to Register Allocation · MICRO 1996 |
Processor architecture and microarchitecture › instruction-level parallelism
compiler-controlled speculative execution |
0.0 | 2 | 1993 | Sentinel Scheduling for VLIW and Superscalar Processors · ACM Trans. Comput. Syst. 1993 Sentinel Scheduling for VLIW and Superscalar Processors · ASPLOS 1992 |
Processor architecture and microarchitecture
speculative execution |
0.0 | 2 | 1993 | Sentinel Scheduling for VLIW and Superscalar Processors · ACM Trans. Comput. Syst. 1993 Sentinel Scheduling for VLIW and Superscalar Processors · ASPLOS 1992 |
Processor architecture and microarchitecture
multicore design |
0.0 | 1 | 2006 | Exploiting Fine-Grained Data Parallelism with Chip Multiprocessors and Fast Barriers · MICRO 2006 |
Compilers and program optimization › instruction scheduling
software pipelining |
0.0 | 2 | 1992 | Register Allocation for Software Pipelined Loops · PLDI 1992 Code generation schema for modulo scheduled loops · MICRO 1992 |
Processor architecture and microarchitecture › instruction set architecture
EPIC architecture |
0.0 | 2 | 2001 | Compiling for EPIC architectures · Proc. IEEE 2001 Control CPR: A Branch Height Reduction Optimization for EPIC Architectures · PLDI 1999 |
Compilers and program optimization › instruction scheduling › software pipelining
modulo scheduling |
0.0 | 2 | 1992 | Code generation schema for modulo scheduled loops · MICRO 1992 Parallelization of loops with exits on pipelined architectures · SC 1990 |
Compilers and program optimization
compiler analysis |
0.0 | 1 | 1996 | Global Predicate Analysis and Its Application to Register Allocation · MICRO 1996 |
Program analysis
data flow analysis |
0.0 | 1 | 1996 | Analysis Techniques for Predicated Code · MICRO 1996 |
Compilers and program optimization › compiler analysis
predicate analysis |
0.0 | 1 | 1996 | Global Predicate Analysis and Its Application to Register Allocation · MICRO 1996 |
Processor architecture and microarchitecture
instruction set architecture |
0.0 | 1 | 2001 | Compiling for EPIC architectures · Proc. IEEE 2001 |
Compilers and program optimization
code generation |
0.0 | 1 | 1992 | Code generation schema for modulo scheduled loops · MICRO 1992 |
Processor architecture and microarchitecture › instruction-level parallelism
superscalar and VLIW processors |
0.0 | 1 | 1992 | Sentinel Scheduling for VLIW and Superscalar Processors · ASPLOS 1992 |
Compilers and program optimization
loop optimization |
0.0 | 1 | 1990 | Parallelization of loops with exits on pipelined architectures · SC 1990 |
Processor architecture and microarchitecture › pipelining
pipelined processor |
0.0 | 1 | 1990 | Parallelization of loops with exits on pipelined architectures · SC 1990 |
Processor architecture and microarchitecture › superscalar processor
wide-issue |
0.0 | 1 | 1990 | Parallelization of loops with exits on pipelined architectures · SC 1990 |
Processor architecture and microarchitecture
multiple functional units |
0.0 | 1 | 1995 | Spill-free parallel scheduling of basic blocks · MICRO 1995 |
Processor architecture and microarchitecture
exception handling |
0.0 | 1 | 1993 | Sentinel Scheduling for VLIW and Superscalar Processors · ACM Trans. Comput. Syst. 1993 |
Methods — techniques the papers use, named apart from their topics
prototype implementation · 0.1hardware-software co-design · 0.1simulation · 0.1compiler transformation · 0.1profiling · 0.0branch sequence recognition · 0.0resource allocation · 0.0operation scheduling · 0.0clustering · 0.0bitwidth analysis · 0.0compile-time scheduling · 0.0profile-driven scoring · 0.0modulo scheduling · 0.0live range analysis · 0.0list scheduling · 0.0graph-based data structure · 0.0control flow analysis · 0.0boolean expression manipulation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | CORONET: Fault tolerance for Software Defined NetworksabstractSoftware Defined Networking, or SDN, based networks are being deployed not only in testbed networks, but also in production networks. Although fault-tolerance is one of the most desirable properties in production networks, there are not much study in providing fault-tolerance to SDN-based networks. The goal of this work is to develop a fault tolerant SDN architecture that can rapidly recover from faults and scale to large network sizes. This paper presents CORONET, a SDN fault-tolerant system that recovers from multiple link failures in the data plane. We describe a prototype implementation based on NOX that demonstrates fault recovery for emulated topologies using Mininet. We also discuss possible extensions to handle control plane and controller faults. Hyojoon Kim, Mike Schlansker, Jose Renato Santos, Jean Tourrilhes, Yoshio Turner, Nick Feamster |
ICNP | 2 |
| 2011 | Routing Optimization for Ensemble RoutingabstractThe Ensemble Routing architecture (presented at ANCS 2010) implements multipath routing for data center networks. Rather than managing individual flows, ensemble routing manages flows in groups or ensembles to provide scalable responsive management using simple hardware. Ensemble Routing combines: routing VLANs that define a set of diverse paths through complex networks, and load balancing algorithms to split traffic among those VLANS and optimize traffic flow. This extended abstract describes improved algorithms for the formation of routing VLANs and traffic load-balancing for Ensemble routing. The VLAN formation algorithms are improved by incorporating traffic flow estimates into the VLAN formation heuristics. The previous load balancing algorithm using a greedy heuristic is replaced by linear programming that determines optimal traffic splitting among VLANs. Simulations show that these mechanisms significantly enhance performance. Wenfei Wu, Yoshio Turner, Mike Schlansker |
ANCS | 3 |
| 2010 | Ensemble routing for datacenter networksabstractThis paper describes Hash-Based Routing (HBR), an architecture that enhances Ethernet to support dynamic management for multipath networks in scalable datacenters. This work enhances HBR to support flow ensemble management for large-scale networks of arbitrary topology. Ensemble routing eliminates measurement and control for individual flows and instead manages using summary data thus providing a unique capability for reactive datacenter-wide network management. HBR provides seamless interoperability with Ethernet and supports the attachment of unmodified L2 hosts and devices including FCoE devices within converged fabrics. Simulation experiments demonstrate efficient multipath routing for a variety of scalable topologies. Optimized routing maintains efficiency in the presence of network faults and implements spatial Quality of Service to dynamically provision physical network hardware among co-hosted tenants or applications. Mike Schlansker, Yoshio Turner, Jean Tourrilhes, Alan Karp |
ANCS | 1 |
| 2010 | Killer Fabrics for Scalable DatacentersabstractMost datacenter networks are based on specialized edge-core topologies, which are costly to build, difficult to maintain and consume too much power. We propose enhancements to layer-two (L2) Ethernet switches to enable multipath L2 routing in scalable datacenters. This replaces an expensive router with commodity switches. Our hash-based routing approach reuses and minimally extends hardware structures in high-volume switches, while exposing a powerful network management inter-face for multipath load balancing, QoS differentiation, and resilience to faults. Simulation results demonstrate near-optimal load balancing for uniform and non-uniform traffic patterns, and effective management of large datacenter networks independent of the number of traffic flows. Mike Schlansker, Jean Tourrilhes, Yoshio Turner, Jose Renato Santos |
ICC | 1 |
| 2009 | Hash-based routing for scalable datacentersabstractMost datacenter networks are based on specialized edge-core topologies, which are costly to build, difficult to maintain and consume too much power. We propose enhancements to layer-two (L2) Ethernet switches to enable multipath L2 routing in scalable datacenters. Our hash-based routing approach reuses and minimally extends hardware structures in high-volume switches, while exposing a powerful network management interface for multipath load balancing, QoS differentiation, and resilience to faults. Mike Schlansker, Jean Tourrilhes, Yoshio Turner, Jose Renato Santos |
ANCS | 1 |
| 2007 | High-performance ethernet-based communications for future multi-core processorsabstractData centers and HPC clusters often incorporate specialized networking fabrics to satisfy system requirements. However, Ethernet's low cost and high performance are causing a shift from specialized fabrics toward standard Ethernet. Although Ethernet's low-level performance approaches that of specialized fabrics, the features that these fabrics provide such as reliable in-order delivery and flow control are implemented, in the case of Ethernet, by endpoint hardware and software. Unfortunately, current Ethernet endpoints are either slow (commodity NICs with generic TCP/IP stacks) or costly (offload engines). To address these issues, the JNIC project developed a novel Ethernet endpoint. JNIC's hardware and software were specifically designed for the requirements of high-performance communications within future data-centers and compute clusters. The architecture combines capabilities already seen in advanced network architectures with new innovations to create a comprehensive solution for scalable and high-performance Ethernet. We envision a JNIC architecture that is suitable for most in-data-center communication needs. Mike Schlansker, Bhushan Chitlur, Erwin Oertli, Paul M. Stillwell Jr., Linda Rankin, Dennis Bradford, Richard J. Carter, Jayaram Mudigonda, Nathan L. Binkert, Norman P. Jouppi |
SC | 1 |
| 2006 | Exploiting Fine-Grained Data Parallelism with Chip Multiprocessors and Fast BarriersabstractWe examine the ability of CMPs, due to their lower on-chip communication latencies, to exploit data parallelism at inner-loop granularities similar to that commonly targeted by vector machines. Parallelizing code in this manner leads to a high frequency of barriers, and we explore the impact of different barrier mechanisms upon the efficiency of this approach. To further exploit the potential of CMPs for fine-grained data parallel tasks, we present barrier filters, a mechanism for fast barrier synchronization on-chip multi-processors to enable vector computations to be efficiently distributed across the cores of a CMP. We ensure that all threads arriving at a barrier require an unavailable cache line to proceed, and, by placing additional hardware in the shared portions of the memory subsystem, we starve their requests until they all have arrived. Specifically, our approach uses invalidation requests to both make cache lines unavailable and identify when a thread has reached the barrier. We examine two types of barrier filters, one synchronizing through instruction cache lines, and the other through data cache lines Jack Sampson, Jean-Francois Collard, Norman P. Jouppi, Mike Schlansker, Brad Calder |
MICRO | 5 |
| 2003 | In Memory of Bob Rau
Mike Schlansker |
MICRO | 1 |
| 2001 | ShiftQ: a bufferred interconnect for custom loop acceleratorsabstractShiftQs are hardware structures consisting of registers and switches which buffer and transport operands among function units within custom hardware loop accelerators. ShiftQs help minimize buffering and interconnect costs by customizing the hardware to the given schedule and by intelligent sharing of register and interconnect resources. This paper describes the ShiftQ schema and a method to automatically synthesize them from modulo-scheduled loops. We also evaluate the cost savings by comparing them against traditional storage and interconnect mechanisms. Shail Aditya, Mike Schlansker |
CASES | 2 |
| 2001 | Compiling for EPIC architecturesabstractDesigning compilers for Explicitly Parallel Instruction Computing (EPIC.) architectures presents challenges substantially different from those encountered in designing compilers for traditional sequential architectures. These challenges are addressed not only by employing new optimizations that are specific to EPIC, but also by employing new ways to architect compilers. EPIC architectures provide features that allow compilers to take a proactive role in exploiting instruction level parallelism. Compiler technology is intimately intertwined with the target processor architecture, and compiler architects must solve new analysis and optimization problems to achieve the highest levels of performance. When complex optimizations are uniformly applied to large applications, the resulting slow compile speeds are unacceptable. Demanding requirements to produce high-quality code at high compile speed shapes the fundamental structure of EPIC compilers. Vinod Kathail, Mike Schlansker, Bob Rau |
Proc. IEEE | 2 |
| 2001 | Bitwidth cognizant architecture synthesis of custom hardwareacceleratorsabstractProgram-in chip-out (PICO) is a system for automatically synthesizing embedded hardware accelerators from loop nests specified in the C programming language. A key issue confronted when designing such accelerators is the optimization of hardware by exploiting information that is known about the varying number of bits required to represent and process operands. In this paper, we describe the handling and exploitation of integer bitwidth in PICO. A bitwidth analysis procedure is used to determine bitwidth requirements for all integer variables and operations in a C application. Given known bitwidths for all variables, complex problems arise when determining a program schedule that specifies on which function unit (FU) and at what time each operation executes. If operations are assigned to FUs with no knowledge of bitwidth, bitwidth-related cost benefit is lost when each unit is built to accommodate the widest operation assigned. By carefully placing operations of similar width on the same unit, hardware costs are decreased. This problem is addressed using a preliminary clustering of operations that is based jointly on width and implementation cost. These clusters are then honored during resource allocation and operation scheduling to create an efficient width-conscious design. Experimental results show that exploiting integer bitwidth substantially reduces the gate count of PICO-synthesized hardware accelerators across a range of applications. Scott A. Mahlke, Rajiv A. Ravindran, Mike Schlansker, Robert Schreiber, Timothy Sherwood |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2000 | Embedded Computing: New Directions in Architecture and Automation
Bob Rau, Mike Schlansker |
HiPC | 2 |
| 1999 | Control CPR: A Branch Height Reduction Optimization for EPIC ArchitecturesabstractThe challenge of exploiting high degrees of instruction-level parallelism is often hampered by frequent branching. Both exposed branch latency and low branch throughput can restrict parallelism. Control critical path reduction (control CPR) is a compilation technique to address these problems. Control CPR can reduce the dependence height of critical paths through branch operations as well as decrease the number of executed branches. In this paper, we present an approach to control CPR that recognizes sequences of branches using profiling statistics. The control CPR transformation is applied to the predominant path through this sequence. Our approach, its implementation, and experimental results are presented. This work demonstrates that control CPR enhances instruction-level parallelism for a variety of application programs and improves their performance across a range of processors. Mike Schlansker, Scott A. Mahlke |
PLDI | 1 |
| 1996 | Profile-driven Instruction Level Parallel Scheduling with Application to Super BlocksabstractCode scheduling to exploit instruction level parallelism (ILP) is a critical problem in compiler optimization research in light of the increased use of long-instruction-word machines. Unfortunately optimum scheduling is computationally intractable, and one must resort to carefully crafted heuristics in practice. If the scope of application of a scheduling heuristic is limited to basic blocks, considerable performance loss may be incurred at block boundaries. To overcome this obstacle, basic blocks can be coalesced across branches to form larger regions such as super blocks. In the literature, these regions are typically scheduled using algorithms that are either oblivious to profile information (under the assumption that the process of forming the region has fully utilized the profile information), or use the profile information as an addendum to classical scheduling techniques. We believe that even for the simple case of linear code regions such as super blocks, additional performance improvement can be gained by utilizing the profile information in scheduling as well. We propose a general paradigm for converting any profile-insensitive list scheduler to a profile-sensitive scheduler. Our technique is developed via a theoretical analysis of a simplified abstract model of the general problem of profile-driven scheduling over any acyclic code region, yielding a scoring measure for ranking branch instructions. Chandra Chekuri, Rajeev Motwani 0001, B. Natarajan, Bob Rau, Mike Schlansker |
MICRO | 6 |
| 1996 | Global Predicate Analysis and Its Application to Register AllocationabstractTo fully utilize the wide machine resources in modern high-performance microprocessors it is necessary to exploit parallelism beyond individual basic blocks. Architectural support for predicated execution increases the degree of instruction level parallelism by allowing instructions from different basic blocks to be converted to straight-line code guarded by boolean predicates. However predicated execution also presents significant challenges to an optimizing compiler. For example, in live range analysis, a predicated definition does not necessarily end the live range of a virtual register. This paper describes techniques to analyze the relations among predicates in order to improve the precision and effectiveness of various compiler analysis and transformation phases in the presence of predicated code. Our predicate analysis operates globally to obtain relations among predicates. Moreover, we analyze control flow and predication in a single unified framework. The result can be queried by subsequent optimization and analysis phases. Based on this framework, we extend a traditional method to a predicate-aware register allocator which takes global predicate relations into account. We have implemented the proposed algorithms to effectively reduce register pressure. Our experimental results show 24.6% of a large test suite obtain, on average, 20.71% and better register allocation due to the algorithms presented in this paper. David M. Gillies, Roy Dz-Ching Ju, Mike Schlansker |
MICRO | 4 |
| 1996 | Analysis Techniques for Predicated CodeabstractPredicated execution offers new approaches to exploiting instruction-level parallelism (ILP), but it also presents new challenges for compiler analysis and optimization. In predicated code, each operation is guarded by a boolean operand whose run-time value determines whether the operation is executed or nullified. While research has shown the utility of predication in enhancing ILP, there has been little discussion of the difficulties surrounding compiler support for predicated execution. Conventional program analysis tools (e.g. data flow analysis) assume that operations execute unconditionally within each basic block and thus make incorrect assumptions about the run-rime behavior of predicated code. These tools can be modified to be correct without requiring predicate analysis, but this yields overly-conservative results in crucial areas such as scheduling and register allocation. To generate high-quality code for machines offering predicated execution, a compiler must incorporate information about relations between predicates into its analysis. We present new techniques for analyzing predicated code. Operations which compute predicates are analyzed to determine relations between predicate values. These relations are captured in a graph-based data structure, which supports efficient manipulation of boolean expression representing facts about predicated code. This approach forms the basis for predicate-sensitive data flow analysis. Conventional data flow algorithms can be systematically upgraded to be predicate sensitive by incorporating information about predicates. Predicate-sensitive data flow analysis yields significantly more accurate results than conventional data flow analysis when applied to predicated code. Mike Schlansker |
MICRO | 2 |
| 1995 | Spill-free parallel scheduling of basic blocksabstractThis paper concerns the problem of spill-free scheduling of basic blocks on a processor with multiple functional units and a limited number of registers. The problem of minimizing the schedule length is well known to be computationally intractable. We present a heuristic for the problem, a general divide-and-conquer paradigm that converts any insensitive scheduling algorithm-one that is insensitive to register constraints-to one that respects register constraints. We estimate the goodness of the heuristic by relating its performance to that of the insensitive algorithm. We also present experimental results obtained by applying the heuristic to basic blocks from the SPEC benchmark programs, for several machine models. B. Natarajan, Mike Schlansker |
MICRO | 2 |
| 1995 | Critical path reduction for scalar programsabstractScalar performance on processors with instruction level parallelism (ILP) is often limited by control and data dependences. This paper describes a family of compiler techniques, called critical path reduction (CPR) techniques, which reduce the length of critical paths through control and data dependences. Control CPR reduces the number of branches on the critical path and improves the performance of branch intensive codes on processors with inadequate branch throughput or excessive branch latency. Data CPR reduces the number of arithmetic operations on the critical path. Optimization and scheduling are adapted to support CPR. Mike Schlansker, Vinod Kathail |
MICRO | 1 |
| 1994 | Height reduction of control recurrences for ILP processorsabstractThe performance of applications executing on processors with instruction level parallelism is often limited by control and data dependences. Performance bottlenecks caused by dependences can frequently be eliminated through transformations which reduce the height of critical paths through the program. While height reduction techniques are not always helpful, their utility can be demonstrated in a broad range of important situations. Mike Schlansker, Vinod Kathail, Sadun Anik |
MICRO | 1 |
| 1993 | Sentinel Scheduling for VLIW and Superscalar ProcessorsabstractSpeculative execution is an important source of parallelism for VLIW and superscalar processors. A serious challenge with compiler-controlled speculative execution is to efficiently handle exceptions for speculative instructions. In this article, a set of architectural features and compile-time scheduling support collectively referred to assentinel schedulingis introduced. Sentinel scheduling provides an effective framework for both compiler-controlled speculative execution and exception handling. All program exceptions are accurately detected and reported in a timely manner with sentinel scheduling. Recovery from exceptions is also ensured with the model. Experimental results show the effectiveness of sentinel scheduling for exploiting instruction-level parallelism and overhead associated with exception handling. Scott A. Mahlke, William Y. Chen, Roger A. Bringmann, Richard E. Hank, Wen-Mei W. Hwu, Bob Rau, Mike Schlansker |
ACM Trans. Comput. Syst. | 7 |
| 1992 | Sentinel Scheduling for VLIW and Superscalar ProcessorsabstractSpeculative execution is an important source of parallelism for VLIW and superscalar processors. A serious challenge with compiler-controlled speculative execution is to accurately detect and report all program execution errors at the time of occurrence. In this paper, a set of architectural features and compile-time scheduling support referred to as sentinel scheduling is introduced. Sentinel scheduling provides an effective framework for compiler-controlled speculative execution that accurately detects and reports all exceptions. Sentinel scheduling also supports speculative execution of store instructions by providing a store buffer which allows probationary entries. Experimental results show that sentinel scheduling is highly effective for a wide range of VLIW and superscalar processors. Scott A. Mahlke, William Y. Chen, Wen-Mei W. Hwu, Bob Rau, Mike Schlansker |
ASPLOS | 5 |
| 1992 | Code generation schema for modulo scheduled loops
Bob Rau, Mike Schlansker, Parthasarathy P. Tirumalai |
MICRO | 2 |
| 1992 | Register Allocation for Software Pipelined LoopsabstractSoftware pipelining is an important instruction scheduling technique for efficiently overlapping successive iterations of loops and executing them in parallel. This paper studies the task of register allocation for software pipelined loops, both with and without hardware features that are specifically aimed at supporting software pipelines. Register allocation for software pipelines presents certain novel problems leading to unconventional solutions, especially in the presence of hardware support. This paper formulates these novel problems and presents a number of alternative solution strategies. These alternatives are comprehensively tested against over one thousand loops to determine the best register allocation strategy, both with and without the hardware support for software pipelining. Bob Rau, Meng Lee, Parthasarathy P. Tirumalai, Mike Schlansker |
PLDI | 4 |
| 1991 | Parallelization of WHILE loops on pipelined architectures
Parthasarathy P. Tirumalai, Meng Lee, Mike Schlansker |
J. Supercomput. | 3 |
| 1990 | Parallelization of loops with exits on pipelined architecturesabstractModulo scheduling theory can be applied successfully to overlap Fortran DO loops on pipelined computers issuing multiple operations per cycle both with and without special loop architectural support. It is shown that a broader class of loops-repeat-until, while, and loops with more than one exit-where the trip count is not known beforehand, can also be overlapped efficiently on multiple issue pipelined machines. Special features that are required in the architecture as well as compiler representations for accelerating these loop constructions are discussed. The approach uses hardware architectural support, program transformation techniques, performance bounds calculations, and scheduling heuristics. Performance results are presented for a few select examples. A prototype scheduler is currently under construction for the Cydra 5 directed dataflow computer.> Parthasarathy P. Tirumalai, Meng Lee, Mike Schlansker |
SC | 3 |
| 1989 | The Cydram 5 Stride-Insensitive Memory System
Bob Rau, Mike Schlansker, David W. L. Yen |
ICPP (1) | 2 |