Andrew Putnam

dblp:85/5822 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
4since 2021 · last 2026
0000-0001-5241-5695ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 8 first-author · 3 since 2021Software engineering, systems software and programming languages · 6 · 2 first-authorComputer networks · 2 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 EZCache: Easy Action-Enabled FPGA Caches for Non-Stalling Datapaths in SmartNICs and Beyond
abstract
Caches are widely used in FPGA accelerators such as SmartNICs to hide DRAM latency, but conventional designs treat caches as passive storage. When workloads require read–modify–write (RMW) updates – such as flow tables, counters, or per-connection state – existing cache IPs force designers to either stall the pipeline or duplicate hazard-handling logic outside the cache. Both approaches waste bandwidth, complicate datapath design, and require extensive verification.We propose EZCache, a new FPGA cache IP that integrates programmable action blocks directly into the cache pipeline. These blocks perform user-defined operations on cached data in place, allowing RMW updates to complete without stalling other operations or exposing hazards to surrounding logic. By embedding compute into the cache, EZCache transforms it from a passive buffer into an active architectural primitive suitable for a wide range of FPGA accelerators.EZCache is fully parametric in associativity, latency, throughput, and action complexity, enabling a single design to support diverse workloads and memory systems. This parametric abstraction allows the surrounding datapath to remain stable even as cache configurations, action semantics, and external memories evolve across hardware generations. EZCache has been deployed across five generations of SmartNICs and millions of devices worldwide, proving both its maturity and its impact in production. Results across multiple FPGA platforms show EZCache’s generality and efficiency, establishing it as a new foundation for cache-based acceleration in reconfigurable systems.
Ahmed Abdelsalam, Vishal Gondaliya, Ezz Hamed, Pragati Medleri Hire Math, Marc Gepigon, Joshua Landgraf, Nadeen Gebara, Bob Groza, Anshuman Verma, Andrew Putnam
FCCM11
2026 AI-Assisted Copilot Automation for Reliable FPGA Verification at Hyperscale
abstract
Cloud-scale, Artificial Intelligence (AI)-centric deployments are increasingly accelerated and made flexible through Field-Programmable Gate Arrays (FPGAs). Traditional simulation-only verification has often been shown to miss system-level risks that surface late during hardware bring-up. A hybrid verification methodology has therefore been developed in which AI-driven automation is orchestrated with formal connectivity checks to ensure reliable end-to-end signal connectivity and robust reset behavior across heterogeneous FPGA Stock Keeping Unit (SKU) variants. Structured prompts are employed to guide Copilot in generating Python scripts that traverse - Register Transfer Level (RTL) designs, extract signal mappings, and auto-generate timing-aware connectivity assertions in Comma-Separated Values (CSV) format. Precise, context-rich prompts specifying signal roles, hierarchy depth, and reset domains are observed to yield consistent results, whereas generic prompts fail to capture architectural nuances. Reset validation strategies covering Function Level Resets (FLRs), Memory-Mapped Input Output (MMIO) handlers, and firmware-driven resets have been applied in FPGA contexts, while the underlying connectivity assertion framework has been shown to be portable to Application-Specific Integrated Circuit (ASIC) flows with minimal adaptation. Over nine months, the methodology was deployed on multiple hyperscale FPGA platforms, where critical bugs missed by weeks of simulation regressions were surfaced. Integration with commercial formal tools was achieved, scalability to new FPGA variants was demonstrated, and lessons in prompt engineering and cross-domain verification were documented.
Nguyen Le, Tony-Dat Tran, Jagannath Panduranga Rao, Andrew Putnam
FPGA5
2026 Hyperscale FPGA Engineering Systems at Microsoft
abstract
Microsoft has deployed FPGAs at hyperscale for over a decade, powering diverse application domains and products. While the underlying EDA tool flow remains familiar (synthesis, place & route, and verification), the engineering system that supports FPGA development at Microsoft looks nothing like a traditional hardware flow. Instead, it borrows heavily from modern cloud-scale software practices: Git for version control, Azure DevOps for automated pipelines, extensive regression suites, and daily compiles, effectively adapting the software mantra of ''ship every day'' to the hardware world as ''tape-out every day.''
Rob Rydberg, Madison N. Emas, John Demme, Ana Ibarra, Kara Kagi, Brandon Klouchek, Abhijeet Lawande, Todd Massengill, David J. Powers, Andrew Putnam
FPGA10
2026 Rules Offload Engine (ROE): Accelerating Host SDN Policy Evaluation
Anshuman Verma, Tian Tan 0007, Ahmed Abdelsalam, Milan Dasgupta, Jonathan Hunter, Zach Libby, Narayanan Ravichandran, Harish Srinivasan, Matt Reat, Nadeen Gebara, Vishal Gondaliya, Ezz Hamed, Rahul Garlapati, Lok Chand Koppaka, Abdullah Mughrabi, Dev Desai, Alexander Malysh, Shwetha Bhat, Rohan Kandi, Megan Sng, Tushar Garg, Muluken Hailesellasie, Andrew Putnam, Derek Chiou, Osman Ertugay, Alireza Dabagh, Vivek Bhanu, Daniel Firestone
SIGCOMM23
2020 What To Do With Datacenter FPGAs Besides Deep Learning
abstract
FPGAs have been deployed in datacenters worldwide and are now available for use by in both public and private clouds. Enormous focus has been given to optimizing machine learning workloads for FPGAs, especially for deep neural networks (DNNs) in areas like web search, image classification, and translation. However, major cloud applications encompasses a variety of areas that aren't primarily machine learning workloads, including databases, video encoding, text processing, gaming, bioinformatics, productivity and collaboration, file hosting and storage, e-mail, and many more. While machine learning can certainly play a role in each of these areas, is there more that can be done to accelerate these more traditional workloads? Even more challenging than identifying promising workloads is figuring out how developers can practically create and deploy useful applications using FPGAs to the cloud. While FPGAs-as-a-Service allow access to FPGAs in the cloud, there is a huge gap between raw programmable hardware and a customer paying money to use an application powered by that hardware. A wide variety of FPGA IP exists for developers to use, but individual IP blocks are a long way from being a fully functional cloud application. Building block IPs like Memcached, regex matching, protocol parsing, and linear algebra are only a subset of the necessary functionality for full cloud applications. Developing or acquiring IP and integrating it into a full application that customers will pay for is a significant task. And even when a customer pays, how should the money be distributed between IP vendors. Should it be a onetime fee? By usage? By number of FPGAs deployed? Who should have the burden for support if something goes wrong? In traditional cloud applications, FPGA IP block functions are implemented in software libraries. However, few examples of optimized software libraries are commercially successful, so is selling FPGA IP even a viable commercial model for cloud applications? High-level synthesis (HLS) tools promise to provide one path to enable software developers to make effective use of FPGAs for computing tasks, but are any tools really capable of accelerating cloud-scale applications? Many HLS tools require substantial microarchitectural guidance in the form of pragmas or configuration files to come out with good results. Real cloud applications also rarely have a single dominant function and have significant data movement, so without proper partitioning and tuning, the acceleration gains from the FPGA are quickly wiped out by data movement and Amdahl's Law. This panel will gather experts in using FPGAs for cloud application areas beyond machine learning, and how those applications can be built and successfully deployed. We will cover topics such as: -What are the most important cloud workloads for FPGAs to target besides machine learning? -Are there specific changes to the FPGA architecture that would benefit these cloud applications? -What are the economic models that will work for IP developers, application developers, and cloud providers? -How can we make development of FPGA applications easier for the Cloud? -Will open source IP make it impossible for IP vendors to make commercially successful libraries? -What advances are necessary for HLS tools to be practical in the Cloud? The panel is comprised of experts in applications, IP development, and cloud deployment. Each will give a short presentation of what they find as the most important applications and how they see FPGA development for the cloud going forward, then we will open the floor to an interactive discussion with the audience.
Andrew Putnam
FPGA1
2018 Azure Accelerated Networking: SmartNICs in the Public Cloud
Daniel Firestone, Andrew Putnam, Sambrama Mundkur, Derek Chiou, Alireza Dabagh, Mike Andrewartha, Hari Angepat, Vivek Bhanu, Adrian M. Caulfield, Eric S. Chung, Harish Kumar Chandrappa, Somesh Chaturmohta, Matt Humphrey, Jack Lavier, Norman Lam, Fengfen Liu, Kalin Ovtcharov, Jitendra Padhye, Gautham Popuri, Shachar Raindel, Tejas Sapre, Mark Shaw 0001, Gabriel Silva, Madhan Sivakumar, Nisheeth Srivastava, Anshuman Verma, Qasim Zuhair, Deepak Bansal, Doug Burger, Kushagra Vaid, David A. Maltz, Albert G. Greenberg
NSDI2
2018 Introduction to the Special Section on Deep Learning in FPGAs
abstract
International audience
Deming Chen, Andrew Putnam, Steve Wilton
ACM Trans. Reconfigurable Technol. Syst.2
2017 FPGAs in the Datacenter: Combining the Worlds of Hardware and Software Development
abstract
The Catapult project has brought the power and performance of FPGA-based reconfigurable computing to Microsoft's hyperscale datacenters, accelerating major production cloud applications such as Bing web search and Microsoft Azure, and enabling a new generation of machine learning and artificial intelligence applications. Catapult is now deployed in nearly every new server across the more than a million machines that make up the Microsoft hyperscale cloud.
Andrew Putnam
ACM Great Lakes Symposium on VLSI1
2017 KV-Direct: High-Performance In-Memory Key-Value Store with Programmable NIC
abstract
Performance of in-memory key-value store (KVS) continues to be of great importance as modern KVS goes beyond the traditional object-caching workload and becomes a key infrastructure to support distributed main-memory computation in data centers. Recent years have witnessed a rapid increase of network bandwidth in data centers, shifting the bottleneck of most KVS from the network to the CPU. RDMA-capable NIC partly alleviates the problem, but the primitives provided by RDMA abstraction are rather limited. Meanwhile, programmable NICs become available in data centers, enabling in-network processing. In this paper, we present KV-Direct, a high performance KVS that leverages programmable NIC to extend RDMA primitives and enable remote direct key-value access to the main host memory.
Bojie Li, Zhenyuan Ruan, Wencong Xiao, Yuanwei Lu, Yongqiang Xiong, Andrew Putnam, Enhong Chen
SOSP6
2016 Agile Co-Design for a Reconfigurable Datacenter
abstract
In 2015, a team of software and hardware developers at Microsoft shipped the world?s first commercial search engine accelerated using FPGAs in the datacenter. During the sprint to production, new algorithms in the Bing ranking service were ported into FPGAs and deployed to a production bed within several weeks of conception, leading to significant gains in latency and throughput. The fast turnaround time of new features demanded by an agile software culture would not have been possible without a disciplined and effective approach to co-design in the datacenter. This talk will describe some of the learnings and best practices developed from this unique experience.
Shlomi Alkalay, Hari Angepat, Adrian M. Caulfield, Eric S. Chung, Oren Firestein, Michael Haselman, Stephen Heil, Kyle Holohan, Matt Humphrey, Tamás Juhász, Puneet Kaur, Sitaram Lanka, Daniel Lo, Todd Massengill, Kalin Ovtcharov, Michael Papamichael, Andrew Putnam, Raja Seera, Rimon Tadros, Jason Thong, Lisa Woods, Derek Chiou, Doug Burger
FPGA17
2016 The configurable cloud - accelerating hyperscale datacenter services with FPGAs
abstract
Summary form only given. The Catapult project has brought the power and performance of FPGA-based reconfigurable computing to Microsoft's hyperscale datacenters, accelerating major production cloud applications such as Bing web search and Microsoft Azure, and enabling a new generation of machine learning and artificial intelligence applications. In this talk, I will describe the next generation of the Catapult configurable cloud architecture, show some of the applications have been accelerated by Catapult, and discuss the experiences and lessons learned while bringing up the first applications in the configurable cloud. I will also share stories of the evolution of Catapult - from its origins as a single development kit board through to the hyperscale deployment of hundreds of thousands of FPGAs across 15 countries and 5 continents - highlighting insights and experiences that transformed Catapult from a caffeine-fueled PowerPoint presentation into the largest deployment of reconfigurable computing in the world.
Andrew Putnam
FPT1
2016 A cloud-scale acceleration architecture
abstract
Hyperscale datacenter providers have struggled to balance the growing need for specialized hardware (efficiency) with the economic benefits of homogeneity (manageability). In this paper we propose a new cloud architecture that uses reconfigurable logic to accelerate both network plane functions and applications. This Configurable Cloud architecture places a layer of reconfigurable logic (FPGAs) between the network switches and the servers, enabling network flows to be programmably transformed at line rate, enabling acceleration of local applications running on the server, and enabling the FPGAs to communicate directly, at datacenter scale, to harvest remote FPGAs unused by their local servers. We deployed this design over a production server bed, and show how it can be used for both service acceleration (Web search ranking) and network acceleration (encryption of data in transit at high-speeds). This architecture is much more scalable than prior work which used secondary rack-scale networks for inter-FPGA communication. By coupling to the network plane, direct FPGA-to-FPGA messages can be achieved at comparable latency to previous work, without the secondary network. Additionally, the scale of direct inter-FPGA messaging is much larger. The average round-trip latencies observed in our measurements among 24, 1000, and 250,000 machines are under 3, 9, and 20 microseconds, respectively. The Configurable Cloud architecture has been deployed at hyperscale in Microsoft's production datacenters worldwide.
Adrian M. Caulfield, Eric S. Chung, Andrew Putnam, Hari Angepat, Jeremy Fowers, Michael Haselman, Stephen Heil, Matt Humphrey, Puneet Kaur, Joo-Young Kim 0001, Daniel Lo, Todd Massengill, Kalin Ovtcharov, Michael Papamichael, Lisa Woods, Sitaram Lanka, Derek Chiou, Doug Burger
MICRO3
2015 Accelerating Homomorphic Evaluation on Reconfigurable Hardware
Thomas Pöppelmann, Michael Naehrig, Andrew Putnam, Adrián Macías
CHES3
2014 A reconfigurable fabric for accelerating large-scale datacenter services
abstract
Datacenter workloads demand high computational capabilities, flexibility, power efficiency, and low cost. It is challenging to improve all of these factors simultaneously. To advance datacenter capabilities beyond what commodity server designs can provide, we have designed and built a composable, reconfigurable fabric to accelerate portions of large-scale software services. Each instantiation of the fabric consists of a 6×8 2-D torus of high-end Stratix V FPGAs embedded into a half-rack of 48 machines. One FPGA is placed into each server, accessible through PCIe, and wired directly to other FPGAs with pairs of 10 Gb SAS cables. In this paper, we describe a medium-scale deployment of this fabric on a bed of 1,632 servers, and measure its efficacy in accelerating the Bing web search engine. We describe the requirements and architecture of the system, detail the critical engineering challenges and solutions needed to make the system robust in the presence of failures, and measure the performance, power, and resilience of the system when ranking candidate documents. Under high load, the largescale reconfigurable fabric improves the ranking throughput of each server by a factor of 95% for a fixed latency distribution—or, while maintaining equivalent throughput, reduces the tail latency by 29%.
Andrew Putnam, Adrian M. Caulfield, Eric S. Chung, Derek Chiou, Kypros Constantinides, John Demme, Hadi Esmaeilzadeh, Jeremy Fowers, Gopi Prashanth Gopal, Jan Gray, Michael Haselman, Scott Hauck, Stephen Heil, Amir Hormati, Joo-Young Kim 0001, Sitaram Lanka, James R. Larus, Eric Peterson, Simon Pope, Aaron Smith, Jason Thong, Phillip Yi Xiao, Doug Burger
ISCA1
2013 How to implement effective prediction and forwarding for fusable dynamic multicore architectures
abstract
Dynamic multicore architectures, that fuse and split cores at run time, potentially offer a level of performance/energy agility that static multicore designs cannot achieve. Conventional ISAs, however, have scalability limits to fusion. EDGE-based designs offer greater scalability but to date have been performance limited by significant microarchitectural bottlenecks. This paper addresses these issues and makes three major contributions. First, it proposes Iterative Path Prediction to address low next block prediction accuracy and low speculation rates. It achieves close to taken/not-taken prediction accuracy for multi-exit instruction blocks while also speculating the predicated execution path within the block. Second, the paper proposes Exposed Operand Broadcasts to address the overhead of operand delivery for high fanout instructions by exposing a small number of broadcast operands in the ISA. Third, we present a scalable composable architecture called T3 that uses these mechanisms and show it can operate across a wide range of power and performance spectrum by increasing energy efficiency and performance significantly. Compared to previous EDGE designs, T3 improves energy efficiency by about 2x and performance by up to 50%.
Behnam Robatmili, Hadi Esmaeilzadeh, Madhu Saravana Sibi Govindan, Aaron Smith, Andrew Putnam, Doug Burger, Stephen W. Keckler
HPCA6
2012 Inspection resistant memory: Architectural support for security from physical examination
abstract
The ability to safely keep a secret in memory is central to the vast majority of security schemes, but storing and erasing these secrets is a difficult problem in the face of an attacker who can obtain unrestricted physical access to the underlying hardware. Depending on the memory technology, the very act of storing a 1 instead of a 0 can have physical side effects measurable even after the power has been cut. These effects cannot be hidden easily, and if the secret stored on chip is of sufficient value, an attacker may go to extraordinary means to learn even a few bits of that information. Solving this problem requires a new class of architectures that measurably increase the difficulty of physical analysis. In this paper we take a first step towards this goal by focusing on one of the backbones of any hardware system: on-chip memory. We examine the relationship between security, area, and efficiency in these architectures, and quantitatively examine the resulting systems through cryptographic analysis and microarchitectural impact. In the end, we are able to find an efficient scheme in which, even if an adversary is able to inspect the value of a stored bit with a probabilistic error of only 5%, our system will be able to prevent that adversary from learning any information about the original un-coded bits with 99.9999999999% probability.
Jonathan Valamehr, Melissa Chase, Seny Kamara, Andrew Putnam, Daniel Shumow, Vinod Vaikuntanathan, Timothy Sherwood
ISCA4
2010 MPI as a Programming Model for High-Performance Reconfigurable Computers
abstract
High-Performance Reconfigurable Computers (HPRCs) consist of one or more standard microprocessors tightly-coupled with one or more reconfigurable FPGAs. HPRCs have been shown to provide good speedups and good cost/performance ratios, but not necessarily ease of use, leading to a slow acceptance of this technology. HPRCs introduce new design challenges, such as the lack of portability across platforms, incompatibilities with legacy code, users reluctant to change their code base, a prolonged learning curve, and the need for a system-level Hardware/Software co-design development flow. This article presents the evolution and current work on TMD-MPI, which started as an MPI-based programming model for Multiprocessor Systems-on-Chip implemented in FPGAs, and has now evolved to include multiple X86 processors. TMD-MPI is shown to address current design challenges in HPRC usage, suggesting that the MPI standard has enough syntax and semantics to program these new types of parallel architectures. Also presented is the TMD-MPI Ecosystem , which consists of research projects and tools that are developed around TMD-MPI to further improve HPRC usability. Finally, we present preliminary communication performance measurements.
Manuel Saldaña, Arun Patel, Christopher A. Madill, Daniel Nunes, Danyao Wang, Paul Chow, Ralph Wittig, Henry Styles, Andrew Putnam
ACM Trans. Reconfigurable Technol. Syst.9
2009 Performance and power of cache-based reconfigurable computing
abstract
CHiMPS is a C-based compiler for high-performance computing (HPC) on heterogeneous CPU-FPGA computing platforms. CHiMPS efficiently supports random accesses to main memory through the many-cache memory model, enabling a broader range of applications to take advantage of FPGA-based acceleration. Many-cache creates multiple caches on top of an FGPA's small, independent memories, each targeting a particular data structure or region of memory in an application and each customized for the memory operations that access it. This poster presents the analyses and optimizations of the CHiMPS compiler that construct many-cache caches, and presents the details of the cache parameters on a Xilinx Virtex-5 LX110T FPGA. Detailed simulation results on HPC kernels demonstrate a 7.8x (geometric mean) performance boost over CPU-only execution of the same source code, FPGA power usage that is on average 4.1x less, and consequently performance per watt that is also greater, by a geometric mean of 21.3x.
Andrew Putnam, Susan J. Eggers, Dave Bennett, Eric Dellinger, Jeff Mason, Henry Styles, Prasanna Sundararajan, Ralph Wittig
FPGA1
2009 Performance and power of cache-based reconfigurable computing
abstract
Many-cache is a memory architecture that efficiently supports caching in commercially available FPGAs. It facilitates FPGA programming for high-performance computing (HPC) developers by providing them with memory performance that is greater and power consumption that is less than their current CPU platforms, but without sacrificing their familiar, C-based programming environment.
Andrew Putnam, Susan J. Eggers, Dave Bennett, Eric Dellinger, Jeff Mason, Henry Styles, Prasanna Sundararajan, Ralph Wittig
ISCA1
2008 CHiMPS: a high-level compilation flow for hybrid CPU-FPGA architectures
abstract
This poster describes CHiMPS, a toolflow that aims to provide software developers with a way to program hybrid CPU-FPGA platforms using familiar tools, languages, and techniques. CHiMPS starts with C and produces a specialized spatial dataflow architecture that supports coherent caches and the shared-memory programming model. The toolflow is designed to abstract away the complex details of data movement and separate memories on the hybrid platforms, as well as take advantage of memory management and computation techniques unique to reconfigurable hardware. This poster focuses on the memory design for CHiMPS, particularly the use of numerous small caches customized for various phases of program execution. The poster also addresses area vs. performance tradeoffs for various configurations. Applications compiled using CHiMPS show performance improvements of more than 36x on simple compute-intensive kernels, and 4.3x on the difficult-to-parallelize STSWM application without any special optimizations compared to running only on the CPU. The toolflow supports full ANSI-C, and produces hardware that runs on platforms that are expected to be available within one year
Andrew Putnam, Dave Bennett, Eric Dellinger, Jeff Mason, Prasanna Sundararajan
FPGA1
2008 CHiMPS: A C-level compilation flow for hybrid CPU-FPGA architectures
abstract
This paper describes CHiMPS, a C-based accelerator compiler for hybrid CPU-FPGA computing platforms. CHiMPS’s goal is to facilitate FPGA programming for high-performance computing developers. It inputs generic ANSIC code and automatically generates VHDL blocks for an FPGA. The accelerator architecture is customized with multiple caches that are tuned to the application. Speedups of 2.8x to 36.9x (geometric mean 6.7x) are achieved on a variety of HPC benchmarks with minimal source code changes.
Andrew Putnam, Dave Bennett, Eric Dellinger, Jeff Mason, Prasanna Sundararajan, Susan J. Eggers
FPL1
2007 The WaveScalar architecture
abstract
Silicon technology will continue to provide an exponential increase in the availability of raw transistors. Effectively translating this resource into application performance, however, is an open challenge that conventional superscalar designs will not be able to meet. We present WaveScalar as a scalable alternative to conventional designs. WaveScalar is a dataflow instruction set and execution model designed for scalable, low-complexity/high-performance processors. Unlike previous dataflow machines, WaveScalar can efficiently provide the sequential memory semantics that imperative languages require. To allow programmers to easily express parallelism, WaveScalar supports pthread-style, coarse-grain multithreading and dataflow-style, fine-grain threading. In addition, it permits blending the two styles within an application, or even a single function. To execute WaveScalar programs, we have designed a scalable, tile-based processor architecture called the WaveCache. As a program executes, the WaveCache maps the program's instructions onto its array of processing elements (PEs). The instructions remain at their processing elements for many invocations, and as the working set of instructions changes, the WaveCache removes unused instructions and maps new ones in their place. The instructions communicate directly with one another over a scalable, hierarchical on-chip interconnect, obviating the need for long wires and broadcast communication. This article presents the WaveScalar instruction set and evaluates a simulated implementation based on current technology. For single-threaded applications, the WaveCache achieves performance on par with conventional processors, but in less area. For coarse-grain threaded applications the WaveCache achieves nearly linear speedup with up to 64 threads and can sustain 7--14 multiply-accumulates per cycle on fine-grain threaded versions of well-known kernels. Finally, we apply both styles of threading to equake from Spec2000 and speed it up by 9x compared to the serial version.
Steven Swanson, Andrew Schwerin, Martha Mercaldi Kim, Andrew Petersen 0001, Andrew Putnam, Ken Michelson, Mark Oskin, Susan J. Eggers
ACM Trans. Comput. Syst.5
2006 Reducing control overhead in dataflow architectures
abstract
In recent years, computer architects have proposed tiled architectures in response to several emerging problems in processor design, such as design complexity, wire delay, and fabrication reliability. One of these architectures, WaveScalar, uses a dynamic, tagged-token dataflow execution model to simplify the design of the processor tiles and their interconnection network and to achieve good parallel performance. However, using a dataflow execution model reawakens old problems, including the instruction overhead required for control flow. Previous work compiling the functional language Id to the Monsoon Dataflow System found this overhead to be 2–3× that of programs written in C and targeted to a MIPS R3000.In this paper, we present and analyze three compiler optimizations that significantly reduce control overhead with minimal additional hardware. We begin by describing how to translate imperative code into dataflow assembly and analyze the resulting control overhead. We report a similar 2–4× instruction overhead, which suggests that the execution model, rather than a specific source language or target architecture, is responsible. Then, we present the compiler optimizations, each of which is designed to eliminate a particular type of control overhead, and analyze the extent to which they were able to do so. Finally, we evaluate the effect using all optimizations together has on program performance. Together, the optimizations reduce control overhead by 80% on average, increasing application performance between 21–37%.
Andrew Petersen 0001, Andrew Putnam, Martha Mercaldi Kim, Andrew Schwerin, Susan J. Eggers, Steven Swanson, Mark Oskin
PACT2
2006 Instruction scheduling for a tiled dataflow architecture
abstract
This paper explores hierarchical instruction scheduling for a tiled processor. Our results show that at the top level of the hierarchy, a simple profile-driven algorithm effectively minimizes operand latency. After this schedule has been partitioned into large sections, the bottom-level algorithm must more carefully analyze program structure when producing the final schedule.Our analysis reveals that at this bottom level, good scheduling depends upon carefully balancing instruction contention for processing elements and operand latency between producer and consumer instructions. We develop a parameterizable instruction scheduler that more effectively optimizes this trade-off. We use this scheduler to determine the contention-latency sweet spot that generates the best instruction schedule for each application. To avoid this application-specific tuning, we also determine the parameters that produce the best performance across all applications. The result is a contention-latency setting that generates instruction schedules for all applications in our workload that come within 17% of the best schedule for each.
Martha Mercaldi Kim, Steven Swanson, Andrew Petersen 0001, Andrew Putnam, Andrew Schwerin, Mark Oskin, Susan J. Eggers
ASPLOS4
2006 Area-Performance Trade-offs in Tiled Dataflow Architectures
abstract
Tiled architectures, such as RAW, SmartMemories, TRIPS, and WaveScalar, promise to address several issues facing conventional processors, including complexity, wire-delay, and performance. The basic premise of these architectures is that larger, higher-performance implementations can be constructed by replicating the basic tile across the chip. This paper explores the area-performance trade-offs when designing one such tiled architecture, WaveScalar. We use a synthesizable RTL model and cycle-level simulator to perform an area/performance pareto analysis of over 200 WaveScalar processor designs ranging in size from 19mm2to 575mm2and having a 22 FO4 cycle time. We demonstrate that, for multi-threaded workloads, WaveScalar performance scales almost ideally from 19 to 101mm2when optimized for area efficiency and from 44 to 202mm2when optimized for peak performance. Our analysis reveals that WaveScalar's hierarchical interconnect plays an important role in overall scalability, and that WaveScalar achieves the same (or higher) performance in substantially less area than either an aggressive out-of-order superscalar or Sun's Niagara CMP processor
Steven Swanson, Andrew Putnam, Martha Mercaldi Kim, Ken Michelson, Andrew Petersen 0001, Andrew Schwerin, Mark Oskin, Susan J. Eggers
ISCA2
2006 Modeling instruction placement on a spatial architecture
abstract
In response to current technology scaling trends, architects are developing a new style of processor, known as spatial computers. A spatial computer is composed of hundreds or even thousands of simple, replicated processing elements (or PEs), frequently organized into a grid. Several current spatial computers, such as TRIPS, RAW, SmartMemories, nanoFabrics and WaveScalar, explicitly place a program's instructions onto the grid. Designing instruction placement algorithms is an enormous challenge, as there are an exponential (in the size of the application) number of different mappings of instructions to PEs, and the choice of mapping greatly affects program performance. In this paper we develop an instruction placement performance model which can inform instruction placement. The model comprises three components, each of which captures a different aspect of spatial computing performance: inter-instruction operand latency, data cache coherence overhead, and contention for processing element resources. We evaluate the model on one spatial computer, WaveScalar, and find that predicted and actual performance correlate with a coefficient of -0.90. We demonstrate the model's utility by using it to design a new placement algorithm, which outperforms our previous algorithms. Although developed in the context of WaveScalar, the model can serve as a foundation for tuning code, compiling software, and understanding the microarchitectural trade-offs of spatial computers in general.
Martha Mercaldi Kim, Steven Swanson, Andrew Petersen 0001, Andrew Putnam, Andrew Schwerin, Mark Oskin, Susan J. Eggers
SPAA4