EDBT 2026 Demo / reviewers in the wild / expert
Greg Stitt
dblp:15/1667
· DBLP profile ↗
58ranked-venue papers
18as first author
6since 2021 · last 2026
0000-0001-7159-7439ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 53 · 17 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging the Gap: A Module-Context Modeling Methodology for Hyperscale FPGA ApplicationsabstractMicrosoft operates at hyperscale, deploying FPGA accelerators across global datacenters under development realities that diverge from common practices. When full builds take hours or days, or when the complete system doesn't yet exist, developers turn to module isolation to make progress. But existing isolation techniques sacrifice the physical context needed for accurate optimization, so improvements developed in isolation may not survive integration. Madison N. Emas, Austin Baylis, Greg Stitt |
FPGA | 3 |
| 2026 | EdgeSort: A Sub-100 ns, Line-Rate FPGA Streaming SorterabstractWhile high-throughput sorting is an extensively studied problem, low-latency sorting—despite its potential benefits for many streaming applications—remains largely underexplored by comparison. In many streaming use cases, optimizing for low latency is becoming increasingly relevant as the throughput of graphics-processing units (GPUs) and field-programmable gate arrays (FPGAs) far exceeds inherent I/O bandwidth limitations (e.g., 10 Gbps Ethernet), where further throughput improvements yield no additional benefit. Greg Stitt, Wesley Piard, Christopher Crary |
FPGA | 1 |
| 2024 | Low-Latency, Line-Rate Variable-Length Field Parsing for 100+ Gb/s EthernetabstractField-programmable gate arrays (FPGAs) are widely employed in network-interface cards across applications including cloud services, machine learning, and high-frequency trading. These applications often share a common optimization goal: minimizing latency while meeting throughput constraints. In addition, these applications ideally aim to achieve "line-rate" operation, where the FPGA operates at full bandwidth without using back-pressure to stall incoming data. However, these goals are often conflicting. For example, to minimize latency, application protocols must effectively utilize network bandwidth by encoding variable-length data in variable-length fields. However, variable-length fields often have prohibitively complex processing requirements that prevent line-rate throughput or have excessive latency. In this paper, we present a novel variable-length field parser capable of scaling to accommodate the bus widths and clock frequencies necessary for 100+ Gb/s Ethernet, while still achieving low latency. Our experiments demonstrate parsing variable-length fields at line rate for anticipated bus widths and throughputs, achieving ultra-low latencies under 2 ns for some use cases. To the best of our knowledge, this latency surpasses existing work, including fixed-length field parsing. Greg Stitt, Wesley Piard, Christopher Crary |
FPGA | 1 |
| 2023 | Using FPGA Devices to Accelerate Tree-Based Genetic Programming: A Preliminary Exploration with Recent Technologies
Christopher Crary, Wesley Piard, Greg Stitt, Caleb Bean, Benjamin M. Hicks |
EuroGP | 3 |
| 2023 | An Exploration of ATPG Methods for Redacted IP and Reconfigurable HardwareabstractAutomated test-pattern generation (ATPG) is an important step of testing flows that is responsible for generating test values that expose faults in post-fabrication hardware. Previous work has introduced numerous ATPG methods that analyze application functionality to minimize the number of required tests. However, this existing work is misaligned with the emerging trend to use reconfigurable hardware, such as eFPGAs, to redact security-critical IP. When using reconfigurable hardware, application functionality is only known after serially loading a bitstream into a set of configuration flip-flops, which requires ATPG to do more general tests of the reconfigurable hardware as opposed to the targeted application. This more general testing results in prohibitively slow testing times that are on average 14.6× longer than the original design. In this paper, we explore novel ATPG and test methods for reconfigurable hardware to maximize stuck-at fault coverage, while minimizing testing time. We show significantly improved testing times that are on average 1.9× slower than the unredacted designs, without requiring any knowledge of the original application. Jackson Fugate, Greg Stitt, Naren Vikram Raj Masna, Aritra Dasgupta 0002, Swarup Bhunia, Nij Dorairaj, David Kehlet |
VTS | 2 |
| 2022 | Work-in-Progress: Toward a Robust, Reconfigurable Hardware Accelerator for Tree-Based Genetic ProgrammingabstractGenetic programming (GP) is a general, broadly effective procedure by which computable solutions are constructed from high-level objectives. As with other machine-learning endeavors, one continual trend for GP is to exploit ever-larger amounts of parallelism. In this paper, we explore the possibility of accelerating GP by way of modern field-programmable gate arrays (FPGAs), which is motivated by the fact that FPGAs can sometimes leverage larger amounts of both function and data parallelism—common characteristics of GP— when compared to CPUs and GPUs. As a first step towards more general acceleration, we present a preliminary accelerator for the evaluation phase of "tree-based GP"—the original, and still popular, flavor of GP—for which the FPGA dynamically compiles programs of varying shapes and sizes onto a reconfigurable function tree pipeline. Overall, when compared to a recent open-source GPU solution implemented on a modern 8nm process node, our accelerator implemented on an older 20nm FPGA achieves an average speedup of 9.7×. Although our accelerator is 7.9× slower than most examples of a state-of-the-art CPU solution implemented on a recent 7nm process node, we describe future extensions that can make FPGA acceleration provide attractive Pareto-optimal tradeoffs. Christopher Crary, Wesley Piard, Britton Chesley, Greg Stitt |
CASES | 4 |
| 2020 | PANDORA: An Architecture-Independent Parallelizing Approximation-Discovery FrameworkabstractIn this article, we introduce a p arallelizing a pproximatio n - d isc o very f ra mework, PANDORA, for automatically discovering application- and architecture-specialized approximations of provided code. PANDORA complements existing compilers and runtime optimizers by generating approximations with a range of Pareto-optimal tradeoffs between performance and error, which enables adaptation to different inputs, different user preferences, and different runtime conditions (e.g., battery life). We demonstrate that PANDORA can create parallel approximations of inherently sequential code by discovering alternative implementations that eliminate loop-carried dependencies. For a variety of functions with loop-carried dependencies, PANDORA generates approximations that achieve speedups ranging from 2.3x to 81x, with acceptable error for many usage scenarios. We also demonstrate PANDORA’s architecture-specialized approximations via FPGA experiments, and highlight PANDORA’s discovery capabilities by removing loop-carried dependencies from a recurrence relation with no known closed-form solution. Greg Stitt, David Campbell |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2019 | Dynamic Scheduling on Heterogeneous MulticoresabstractHeterogeneous multicore systems help meet design goals by using disparate hardware components that are suitable for different application requirements/design goals. The individual cores may also have different tunable hardware parameters for additional specialization. However, this complicates scheduling since to reap the benefits of specialization, applications should be scheduled to the core that offers the best configuration based on the application's requirements and design goals. This scheduling decision could be made by exploring the design space to evaluate different configurations to determine the best configuration, or by executing the application in a base configuration to gather execution statistics to predict the best configuration. However, given increasingly complex systems, these methods may be infeasible given extremely large design spaces or difficulty in choosing a representative base configuration. In this paper, we present a dynamic scheduling methodology that uses predictive methods to schedule applications to best configurations for reduced energy consumption for a system with configurable caches. We use an artificial neural network (ANN) to train our predictive model using hardware counters. The trained ANN can then be used to predict the best core and a tuning heuristic explores the design space to determine the best configuration on non-best cores. If the best core is busy, our scheduler considers alternative idle cores or the application is stalled depending on which decision is energy advantageous. Our experiments show that system energy can be reduced by 28% on average as compared to a fixed-core system where all cores offer the same configuration. Ayobami S. Edun, Ruben Vazquez, Ann Gordon-Ross, Greg Stitt |
DATE | 4 |
| 2019 | Energy Prediction for Cache Tuning in Embedded SystemsabstractModern embedded systems are longer tasked at operating a single application or function and are increasingly required to operate more like general purpose desktop computers. Conforming to modern usage demands is extremely challenging given an embedded system's stringent design constraints, such as power, energy, and performance. Adherence to these constraints can be achieved by specializing/tuning the underlying system to application-specific execution requirements and characteristics by tuning a system's configurable parameters to meet these requirements given design constraints. Configurable parameters include architectural voltage, frequency, cache size, line size, and associativity, etc. However, given the complexity of modern systems, exploring these large design spaces is infeasible when the number of configurable parameters and valid parameter values increases beyond a trivial amount. In this paper, we propose using machine learning in lieu of traditional design space exploration techniques. In this work, we evaluate the potential for using an artificial neural network (ANN)-based prediction module for energy prediction. Since the cache hierarchy has a large impact on total energy consumption, without loss of generality, we study a configurable cache hierarchy with configurable cache size, associativity, and line size. We design and train an energy prediction module to infer the best cache configuration for an application based on the application's execution characteristics. Our approach requires only a single profiling run of the application to collect these characteristics. Our energy prediction module then predicts the energy consumption for all the configurations in the cache design space based on these characteristics, and outputs the configuration with the lowest energy consumption, thus essentially performing exhaustive design space exploration with a single execution. Our results show that our prediction module predicts the best instruction and data cache configurations for the majority of the applications, yielding an average energy degradation of less than 2% for both the instruction and data caches as compared to the optimal configuration determined by exhaustive design space exploration. Ruben Vazquez, Ann Gordon-Ross, Greg Stitt |
ICCD | 3 |
| 2019 | PANDORA: a parallelizing approximation-discovery framework (WIP paper)abstractIn this paper, we introduce PANDORA---a framework that complements existing parallelizing compilers by automatically discovering application- and architecture-specialized approximations. We demonstrate that PANDORA creates approximations that extract massive amounts of parallelism from inherently sequential code by eliminating loop-carried dependencies---a long-time goal of the compiler research community. Compared to exact parallel baselines, preliminary results show speedups ranging from 2.3x to 81x with acceptable error for many usage scenarios. Greg Stitt, David Campbell |
LCTES | 1 |
| 2018 | High-Frequency Absorption-FIFO Pipelining for Stratix 10 HyperFlexabstractFPGAs often have significantly lower clock frequencies than microprocessors and GPUs, due largely to propagation delays incurred by the reconfigurable interconnect. The Stratix 10 HyperFlex architecture reduces this problem by embedding numerous registers throughout the routing resources. However, such Hyper-Registers do not support back-pressure (i.e., pipeline stalls) that is commonly used in FPGA pipelines. In this paper, we present and evaluate pipeline transformations using absorption FIFOs, which avoid back-pressure limitations to enable numerous pipelines to benefit from HyperFlex, while also eliminating potentially expensive stall penalties incurred by existing techniques. We demonstrate that these transformations not only enable significant clock improvements on Stratix 10, but also for devices without HyperFlex, potentially making absorption FIFOs a better high-frequency strategy for any FPGA. Madison N. Emas, Austin Baylis, Greg Stitt |
FCCM | 3 |
| 2018 | Scalable Window Generation for the Intel Broadwell+Arria 10 and High-Bandwidth FPGA SystemsabstractEmerging FPGA systems are providing higher external memory bandwidth to compete with GPU performance. However, because FPGAs often achieve parallelism through deep pipelines, traditional FPGA design strategies do not necessarily scale well to large amounts of replicated pipelines that can take advantage of higher bandwidth. We show that sliding-window applications, an important subset of digital signal processing, demonstrate this scalability problem. We introduce a window generator architecture that enables replication to over 330 GB/s, which is an 8.7x improvement over previous work. We evaluate the window generator on the Intel Broadwell+Arria10 system for 2D convolution and show that for traditional convolution (one filter per image), our approach outperforms a 12-core Xeon Broadwell E5 by 81x and a high-end Nvidia P6000 GPU by an order of magnitude for most input sizes, while improving energy by 15.7x. For convolutional neural nets (CNNs), we show that although the GPU and Xeon typically outperform existing FPGA systems, projected performances of the window generator running on FPGAs with sufficient bandwidth can outperform high-end GPUs for many common CNN parameters. Greg Stitt, Abhay Gupta, Madison N. Emas, David Wilson 0004, Austin Baylis |
FPGA | 1 |
| 2018 | Scalable Behavioral Emulation of Extreme-Scale Systems Using Structural Simulation ToolkitabstractWith extremely large design spaces for algorithm and architecture to be explored, there is a need for fast and scalable performance modeling tools for preparing HPC application codes. Behavioral Emulation (BE) is a recent coarse-grained modeling and simulation methodology that has been proposed to solve this co-design problem. In this paper, we introduce a distributed parallel simulation library for Behavioral Emulation called BE-SST, integrated into the Structural Simulation Toolkit (SST). BE-SST provides simple interfaces and framework for development of coarse-grained BE models which can be extended to model new notional architectures. BE-SST also supports Monte Carlo simulations to generate meaningful distributions and summary statistics rather than a single datum for performance. In this paper, we present BE-SST simulations of two existing large DOE machines (Vulcan and Titan), which have been validated against actual testbed measurements and showed 5-10% error. These validated system models (up to 128k cores) are used to make blind predictions of application performance on systems larger than the current machines (up to 512k cores) - a crucial simulator feature for design-space exploration of notional systems. We further studied BE-SST in terms of scalability and performance, simulating up to a million cores, with BE-SST running on more than 2k parallel processes. BE-SST shows good scalability with a linear increase in memory usage and simulation time with increase in simulated system size, and a peak speedup of 7x over single process simulation. With ease of use and good scaling, we assert that BE-SST can significantly speed up design-space exploration. Ajay Ramaswamy, Nalini Kumar, Aravind Neelakantan, Herman Lam, Greg Stitt |
ICPP | 5 |
| 2017 | Serial Arithmetic Strategies for Improving FPGA ThroughputabstractSerial arithmetic has been shown to offer attractive advantages in area for field-programmable gate array (FPGA) datapaths but suffers from a significant reduction in throughput compared to traditional bit-parallel designs. In this work, we perform a performance and trade-off analysis that counterintuitively shows that, despite the decreased throughput of individual serial operators, replication of serial arithmetic can provide a 2.1 × average increase in throughput compared to bit-parallel pipelines for common FPGA applications. We complement this analysis with a novel SerDes architecture that enables existing FPGA pipelines to be replaced with serial logic with potentially higher throughput. We also present a serialized sliding-window architecture that improves average throughput 2.4 × compared to existing bit-parallel work. Aaron Landy, Greg Stitt |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2016 | Doubling FPGA Throughput via a Soft SerDes Architecture for Full-Bandwidth Serial Pipelining (Abstract Only)abstractSerial arithmetic has been shown to offer attractive advantages in area, clock frequency, and functional density for FPGA datapaths but suffers from a significant reduction in throughput compared to traditional bit-parallel designs that is prohibitive for many applications. In this work, we present a full-bandwidth SerDes architecture specialized for Xilinx FPGAs that enables serial pipelines to accept inputs and generate outputs at the same rate as bit-parallel pipelines. When combined with the clock improvements from serial pipelines, we show that this approach offers more than 2.1x average increase in throughput compared to bit-parallel pipelines. Although previous work has shown that serial pipelines can achieve similar results for some limited situations, the key contribution of this work is the ability to replace potentially any existing FPGA pipeline with a higher throughput serialized alternative. We also present a serialized sliding-window architecture that improves throughput up to 4x. Aaron Landy, Greg Stitt |
FPGA | 2 |
| 2016 | A Parallel Sliding-Window Generator for High-Performance Digital-Signal Processing on FPGAsabstractSliding-window applications, an important class of the digital-signal processing domain, are highly amenable to pipeline parallelism on field-programmable gate arrays (FPGAs). Although memory bandwidth often restricts parallelism for many applications, sliding-window applications can leverage custom buffers, referred to as sliding-window generators, that provide massive input bandwidth that far exceeds the capabilities of external memory. Previous work has introduced a variety of sliding-window generators, but those approaches typically generate at most one window per cycle, which significantly restricts parallelism. In this article, we address this limitation with a parallel sliding-window generator that can generate a configurable number of windows every cycle. Although in practice the number of parallel windows is limited by memory bandwidth, we show that even with common bandwidth limitations, the presented generator enables near-linear speedups up to 16x faster than previous FPGA studies that generate a single window per cycle, which were already in some cases faster than graphics-processing units and microprocessors. Greg Stitt, Eric Schwartz, Patrick Cooke |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2016 | The Unified Accumulator Architecture: A Configurable, Portable, and Extensible Floating-Point AccumulatorabstractApplications accelerated by field-programmable gate arrays (FPGAs) often require pipelined floating-point accumulators with a variety of different trade-offs. Although previous work has introduced numerous floating-point accumulation architectures, few cores are available for public use, which forces designers to use fixed-point implementations or vendor-provided cores that are not portable and are often not optimized for the desired set of trade-offs. In this article, we combine and extend previous floating-point accumulator architectures into a configurable, open-source core, referred to as the unified accumulator architecture (UAA), which enables designers to choose between different trade-offs for different applications. UAA is portable across FPGAs and allows designers to specialize the underlying adder core to take advantage of device-specific optimizations. By providing an extensible, open-source implementation, we hope for the research community to extend the provided core with new architectures and optimizations. David Wilson 0004, Greg Stitt |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2015 | An interpolation-based approach to multi-parameter performance modeling for heterogeneous systemsabstractTo effectively optimize applications for emerging heterogeneous architectures, compilers and synthesis tools must perform the challenging task of estimating the performance of different implementations and optimizations for different numbers and types of computational resources. Many performance-prediction techniques exist, but those approaches are specific to particular resources or applications, and are often not capable of prediction for all combinations of inputs. In this paper, we introduce an approach to multi-parameter performance modeling based on sampling and interpolation. This approach can be used in conjunction with execution time data, simulated or observed, to quickly perform performance estimation for any function, on any resource, with any combination of inputs. By evaluating a Kriging-based interpolator on a variety of functions and computational resources, we determine bounds on the accuracy of this approach, and show that an interpolation-based approach utilizing Kriging can effectively model execution time for most applications. We also show that Kriging is a highly effective interpolation technique for execution time, and can be up to four orders of magnitude more accurate than nearest-neighbor interpolation or radial basis function interpolation. Dylan Rudolph, Greg Stitt |
ASAP | 2 |
| 2015 | A scheduling and binding heuristic for high-level synthesis of fault-tolerant FPGA applicationsabstractSpace computing systems commonly use field-programmable gate arrays to provide fault tolerance by applying triple modular redundancy (TMR) to existing register-transfer-level (RTL) code. Although effective, this approach has a 3× area overhead that can be prohibitive for many designs that often allocate resources before considering effects of redundancy. Although a designer could modify existing RTL code to reduce resource usage, such a process is time consuming and error prone. Integrating redundancy into high-level synthesis is a more attractive approach that enables synthesis to rapidly explore different tradeoffs at no cost to the designer. In this paper, we introduce a scheduling and binding heuristic for high-level synthesis that explores tradeoffs between resource usage, latency, and the amount of redundancy. In many cases, an application will not require 100% error correction, which enables significant flexibility for scheduling and binding to reduce resources. Even for applications that require 100% error correction, our heuristic is able to explore solutions that sacrifice latency for reduced resources, and typically save up to 47% when relaxing the latency up to 2×. When the error constraint is reduced to 70%, our heuristic achieves typical resource savings ranging from 18% to 49% when relaxing the latency up to 2×, with a maximum of 77%. Even when comparing with optimized RTL designs, our heuristic uses up to 61% fewer resources than TMR. Aniruddha Shastri, Greg Stitt, Eduardo Riccio |
ASAP | 2 |
| 2015 | Adjustable-Cost Overlays for Runtime CompilationabstractPrevious work has shown that virtual architectures, or overlays, can greatly reduce lengthy FPGA compile times by providing application-specialized resources along with a flexible interconnect to support application changes. However, retaining full configurability of interconnect has also required significant area overhead. In this paper, we introduce a family of overlay architectures called super nets and an associated design methodology that uses data path merging to provide minimal-overhead support for multiple source net lists, and optionally provides an adjustable amount of source flexibility through a secondary interconnect network. We demonstrate that super nets can enable runtime compilation up to 13,000× faster than direct register-transfer logic (RTL) implementation, with up to 70% lower area than selectively enabled RTL data paths. Finally, we explore the design space of this family of overlays and show that it affords significant freedom to trade additional area for increased flexibility to support deviations from the source set, as introduced during development or by optimizations performed at runtime. James Coole, Greg Stitt |
FCCM | 2 |
| 2015 | Revisiting Serial Arithmetic: A Performance and Tradeoff Analysis for Parallel Applications on Modern FPGAsabstractSerial arithmetic cores reduce area compared to bit-parallel alternatives, but are generally assumed to be inappropriate for high-performance FPGA applications due to a significant reduction in throughput. In this paper, we perform a performance and tradeoff analysis of Xilinx 7-series specialized architectures for a novel serial adder tree and multiplier. We show that these serial arithmetic architectures significantly improve functional density due to an average 2x clock speedup compared to bit-parallel alternatives, which provides attractive tradeoffs for different usage scenarios. We also show that serial arithmetic can surprisingly provide better performance than bit-parallel alternatives when replication is solely limited by an area constraint and not application parallelism or input bandwidth. We evaluate this performance improvement on several highly parallel sliding-window applications, showing average speedups of 4.8x and 4.4x compared to bit-parallel implementations over a variety of area constraints. Aaron Landy, Greg Stitt |
FCCM | 2 |
| 2015 | Finite-State-Machine Overlay Architectures for Fast FPGA Compilation and Application PortabilityabstractDespite significant advantages, wider usage of field-programmable gate arrays (FPGAs) has been limited by lengthy compilation and a lack of portability. Virtual-architecture overlays have partially addressed these problems, but previous work focuses mainly on heavily pipelined applications with minimal control requirements. We expand previous work by enabling more flexible control via overlay architectures for finite-state machines. Although not appropriate for control-intensive circuits, the presented architectures reduced compilation times of control changes in a convolution case study from 7 hours to less than 1 second, with no performance overhead and an area overhead of 0.2%. Patrick Cooke, Lu Hao, Greg Stitt |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2015 | A Tradeoff Analysis of FPGAs, GPUs, and Multicores for Sliding-Window ApplicationsabstractThe increasing usage of hardware accelerators such as Field-Programmable Gate Arrays (FPGAs) and Graphics Processing Units (GPUs) has significantly increased application design complexity. Such complexity results from a larger design space created by numerous combinations of accelerators, algorithms, and hw/sw partitions. Exploration of this increased design space is critical due to widely varying performance and energy consumption for each accelerator when used for different application domains and different use cases. To address this problem, numerous studies have evaluated specific applications across different architectures. In this article, we analyze an important domain of applications, referred to as sliding-window applications , implemented on FPGAs, GPUs, and multicore CPUs. For each device, we present optimization strategies and analyze use cases where each device is most effective. The results show that, for large input sizes, FPGAs can achieve speedups of up to 5.6× and 58× compared to GPUs and multicore CPUs, respectively, while also using up to an order of magnitude less energy. For small input sizes and applications with frequency-domain algorithms, GPUs generally provide the best performance and energy. Patrick Cooke, Jeremy Fowers, Greg Brown, Greg Stitt |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2015 | Low-Overhead FPGA Middleware for Application Portability and ProductivityabstractReconfigurable computing devices such as field-programmable gate arrays (FPGAs) offer advantages over fixed-logic CPU and GPU architectures, including improved performance, superior power efficiency, and reconfigurability. The challenge of FPGA application development, however, has limited their acceptance in high-performance computing and high-performance embedded computing applications. FPGA development carries similar difficulties to hardware design, requiring that developers iterate through register-transfer level designs with cycle-level accuracy. Furthermore, the lack of hardware and software standards between FPGA platforms limits productivity and application portability, and makes porting applications between heterogeneous platforms a time-consuming and often challenging process. Recent efforts to improve FPGA productivity using high-level synthesis tools and languages show promise, but platform support remains limited and typically is left as a challenge for developers. To address these issues, we present RC Middleware (RCMW), a novel middleware that improves productivity and enables application and tool portability by abstracting away platform-specific details. RCMW provides an application-centric development environment, exposing only the resources and standardized interfaces required by an application, independent of the underlying platform. We demonstrate the portability and productivity benefits of RCMW using four heterogeneous platforms from three vendors. Our results indicate that RCMW enables application productivity and improves developer productivity, and that these benefits are achieved with less than 7% performance and 3% area overhead on average. Robert Kirchgessner, Alan D. George, Greg Stitt |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2014 | A High Memory Bandwidth FPGA Accelerator for Sparse Matrix-Vector MultiplicationabstractSparse matrix-vector multiplication (SMVM) is a crucial primitive used in a variety of scientific and commercial applications. Despite having significant parallelism, SMVM is a challenging kernel to optimize due to its irregular memory access characteristics. Numerous studies have proposed the use of FPGAs to accelerate SMVM implementations. However, most prior approaches focus on parallelizing multiply-accumulate operations within a single row of the matrix (which limits parallelism if rows are small) and/or make inefficient uses of the memory system when fetching matrix and vector elements. In this paper, we introduce an FPGA-optimized SMVM architecture and a novel sparse matrix encoding that explicitly exposes parallelism across rows, while keeping the hardware complexity and on-chip memory usage low. This system compares favorably with prior FPGA SMVM implementations. For the over 700 University of Florida sparse matrices we evaluated, it also performs within about two thirds of CPU SMVM performance on average, even though it has 2.4x lower DRAM memory bandwidth, and within almost one third of GPU SVMV performance on average, even at 9x lower memory bandwidth. Additionally, it consumes only 25W, for power efficiencies 2.6x and 2.3x higher than CPU and GPU, respectively, based on maximum device power. Jeremy Fowers, Kalin Ovtcharov, Karin Strauss, Eric S. Chung, Greg Stitt |
FCCM | 5 |
| 2014 | A framework for dynamic parallelization of FPGA-accelerated applicationsabstractHigh-level synthesis and compiler studies have introduced many compile-time techniques for parallelizing applications. However, one fundamental limitation of compile-time optimization is the requirement for pessimistic dependence assumptions that can significantly restrict parallelism. To avoid this limitation, many compilers require a restrictive coding style that is not practical for many designers. We present a more transparent approach that aggressively parallelizes applications by dynamically analyzing actual runtime dependencies and scheduling functions onto multiple devices when dependencies allow. In addition, the approach applies FPGA-specific pipelining optimizations to exploit deep parallelism in chains of dependent functions. Experimental results show a speedup of 4.9x for a video-processing application compared to sequential software execution, a speedup of 5.6x compared to traditional FPGA execution, with a framework overhead of only 4%. Jeremy Fowers, Jianye Liu, Greg Stitt |
SCOPES | 3 |
| 2013 | A comparison of correntropy-based feature tracking on FPGAs and GPUsabstractEmbedded signal-processing applications often require feature tracking to identify and track the motion of different objects (features) across a sequence of images. Common measures of similarity for real-time usage are either based on correlation, mean-squared error, or sum of absolute differences, which are not robust enough for safety-critical applications. A recent feature-tracking algorithm called C-Flow uses correntropy to significantly improve signal-to-noise ratio. In this paper, we present an FPGA accelerator for C-Flow that is typically 2-7x faster than a GPU and show that the FPGA is the only device capable of real-time usage for large features. Furthermore, we show the FPGA accelerator is generally more appropriate for embedded usage, with energy consumption that is often 1.2-7.9x less than the GPU. Patrick Cooke, Jeremy Fowers, Greg Stitt, Lee Hunt |
ASAP | 3 |
| 2013 | Virtual finite-state-machine architectures for fast compilation and portabilityabstractFPGA productivity suffers from lengthy compilation times and limited portability. To address these issues, previous work introduced virtual architecture overlays that enable application portability across FPGAs and place-androute that is orders-of-magnitude faster than device vendor tools. However, those previous approaches have limited applicability due to a focus on pipelines with few control requirements. In this paper, we expand control capabilities of previous work by introducing virtual control architectures for finite state machines. The presented architectures reduce lookup-table requirements by multiple orders-of-magnitude and enable larger finite state machines used in common benchmarks. Lu Hao, Greg Stitt |
ASAP | 2 |
| 2013 | Pseudo-constant logic optimizationabstractConstant folding reduces area and enables greater parallelism, but requires circuits with constant inputs. In this work, we extend constant folding to support pseudo-constants, which are values that change with low frequency. We present a method of pseudo-constant logic optimization based on dynamically reconfigurable capabilities of FPGAs, which optimizes logic for different pseudo-constant values and then reconfigures the logic whenever the pseudo-constant changes. Although not beneficial for all logic, we show this optimization achieves up to a 1.25x increase in functional density on Xilinx Virtex 5 FPGAs. Aaron Landy, Greg Stitt |
ASAP | 2 |
| 2013 | A high-performance, low-energy FPGA accelerator for correntropy-based feature tracking (abstract only)abstractComputer-vision and signal-processing applications often require feature tracking to identify and track the motion of different objects (features) across a sequence of images. Numerous algorithms have been proposed, but common measures of similarity for real-time usage are either based on correlation, mean-squared error, or sum of absolute differences, which are not robust enough for safety-critical applications. To improve robustness, a recent feature-tracking algorithm called C-Flow uses correntropy from Information Theoretic Learning to significantly improve signal-to-noise ratio. In this paper, we present an FPGA accelerator for C-Flow that is typically 3.6-8.5x faster than a GPU and show that the FPGA is the only device capable of real-time usage for large features. Furthermore, we show the FPGA accelerator is more appropriate for embedded usage, with energy consumption that is 2.5-22x less than the GPU. Patrick Cooke, Jeremy Fowers, Lee Hunt, Greg Stitt |
FPGA | 4 |
| 2013 | Dynafuse: dynamic dependence analysis for FPGA pipeline fusion and locality optimizationsabstractAlthough high-level synthesis improves FPGA productivity by enabling designers to use high-level code, the resulting performance is often significantly worse than register-transfer-level designs. One cause of such limited optimization is that high-level synthesis tools are restricted by multiple possible dependencies due to the undecidability of alias analysis. In this paper, we introduce the Dynafuse optimization, which analyzes dependencies dynamically to resolve aliases and enable runtime circuit optimizations. To resolve aliases, Dynafuse provides a specialized software data structure that dynamically determines definition-use chains between FPGA functions. In addition, Dynafuse statically creates a reconfigurable overlay network that uses detected dependencies to dynamically adjust connections between functions and memories in order to fuse pipelines and exploit data locality. Experimental results show that Dynafuse sped up two existing FPGA applications by 1.6-1.8x when exploiting locality and by 3-5x when fusing pipelines. Furthermore, the speedup from pipeline fusion increases linearly with the number of fused functions, which suggests larger applications will experience larger improvements. Jeremy Fowers, Greg Stitt |
FPGA | 2 |
| 2013 | A performance and energy comparison of convolution on GPUs, FPGAs, and multicore processorsabstractRecent architectural trends have focused on increased parallelism via multicore processors and increased heterogeneity via accelerator devices (e.g., graphics-processing units, field-programmable gate arrays). Although these architectures have significant performance and energy potential, application designers face many device-specific challenges when choosing an appropriate accelerator or when customizing an algorithm for an accelerator. To help address this problem, in this article we thoroughly evaluate convolution, one of the most common operations in digital-signal processing, on multicores, graphics-processing units, and field-programmable gate arrays. Whereas many previous application studies evaluate a specific usage of an application, this article assists designers with design space exploration for numerous use cases by analyzing effects of different input sizes, different algorithms, and different devices, while also determining Pareto-optimal trade-offs between performance and energy. Jeremy Fowers, Greg Brown, John Robert Wernsing, Greg Stitt |
ACM Trans. Archit. Code Optim. | 4 |
| 2012 | A low-overhead interconnect architecture for virtual reconfigurable fabricsabstractField-programmable gate arrays (FPGAs) have been widely shown to have significant performance and power advantages compared to microprocessors and graphics-processing units (GPUs), but remain a niche technology due in part to productivity challenges. Although such challenges have numerous causes, previous work has shown two significant contributing factors: 1) prohibitive place-and-route times preventing mainstream design methodologies, and 2) limited application portability preventing design reuse. Virtual reconfigurable architectures, referred to as intermediate fabrics (IFs), were recently introduced as a potential solution to these problems, providing 100x-1000x place-and-route speedup, while also enabling application portability across potentially any physical FPGA. However, one significant limitation of existing intermediate fabrics is area overhead incurred from virtualized interconnect resources. In this paper, we perform design-space exploration of virtual interconnect architectures and introduce an optimized virtual interconnect that reduces area overhead by 48% to 54% compared to previous work, while also improving clock frequencies by 24% with a modest routability overhead of 16%. Aaron Landy, Greg Stitt |
CASES | 2 |
| 2012 | The RACECAR heuristic for automatic function specialization on multi-core heterogeneous systemsabstractEmbedded systems increasingly combine multi-core processors and heterogeneous resources such as graphics-processing units and field-programmable gate arrays. However, significant application design complexity for such systems caused by parallel programming and device-specific challenges has often led to untapped performance potential. Application developers targeting such systems currently must determine how to parallelize computation, create different device-specialized implementations for each heterogeneous resource, and then determine how to apportion work to each resource. In this paper, we present the RACECAR heuristic to automate the optimization of applications for multi-core heterogeneous systems by automatically exploring implementation alternatives that include different algorithms, parallelization strategies, and work distributions. Experimental results show RACECAR-specialized implementations can effectively incorporate provided implementations and parallelize computation across multiple cores, graphics-processing units, and field-programmable gate arrays, improving performance by an average of 47x compared to a CPU, while the fastest provided implementations are only able to average 33x. John Robert Wernsing, Greg Stitt, Jeremy Fowers |
CASES | 2 |
| 2012 | Communication visualization for bottleneck detection of high-level synthesis applicationsabstractHigh-level synthesis tools increase FPGA productivity but can decrease performance compared to register-transfer level designs. To help optimize high-level synthesis applications, we introduce a bottleneck detection tool that provides a developer with a visualization of communication bandwidth between all application processes, while identifying potential bottlenecks via color coding. We evaluated the tool using third-party applications to identify and optimize bottlenecks in just several minutes, which achieved speedups ranging from 1.25x to 2.18x compared to the original FPGA execution. Overhead was modest with less than 2% resource overhead and 3% frequency overhead. John Curreri, Greg Stitt, Alan D. George |
FPGA | 2 |
| 2012 | A performance and energy comparison of FPGAs, GPUs, and multicores for sliding-window applicationsabstractWith the emergence of accelerator devices such as multicores, graphics-processing units (GPUs), and field-programmable gate arrays (FPGAs), application designers are confronted with the problem of searching a huge design space that has been shown to have widely varying performance and energy metrics for different accelerators, different application domains, and different use cases. To address this problem, numerous studies have evaluated specific applications across different accelerators. In this paper, we analyze an important domain of applications, referred to as sliding-window applications, when executing on FPGAs, GPUs, and multicores. For each device, we present optimization strategies and analyze use cases where each device is most effective. The results show that FPGAs can achieve speedup of up to 11x and 57x compared to GPUs and multicores, respectively, while also using orders of magnitude less energy. Jeremy Fowers, Greg Brown, Patrick Cooke, Greg Stitt |
FPGA | 4 |
| 2012 | VirtualRC: a virtual FPGA platform for applications and tools portabilityabstractNumerous studies have shown significant performance and power benefits of field-programmable gate arrays (FPGAs). Despite these benefits, FPGA usage has been limited by application design complexity caused largely by the lack of code and tool portability across different FPGA platforms, which prevents design reuse. This paper addresses the portability challenge by introducing a framework of architecture and middleware for virtualization of FPGA platforms, collectively named VirtualRC. Experiments show modest overhead of 5-6% in performance and 1% in area, while enabling portability of 11 applications and two high-level synthesis tools across three physical platforms. Robert Kirchgessner, Greg Stitt, Alan D. George, Herman Lam |
FPGA | 2 |
| 2012 | RACECAR: a heuristic for automatic function specialization on multi-core heterogeneous systemsabstractHigh-performance computing systems increasingly combine multi-core processors and heterogeneous resources such as graphics-processing units and field-programmable gate arrays. However, significant application design complexity for such systems has often led to untapped performance potential. Application designers targeting such systems currently must determine how to parallelize computation, create device-specialized implementations for each heterogeneous resource, and determine how to partition work for each resource. In this paper, we present the RACECAR heuristic to automate the optimization of applications for multi-core heterogeneous systems by automatically exploring implementation alternatives that include different algorithms, parallelization strategies, and work distributions. Experimental results show RACECAR-specialized implementations achieve speedups up to 117x and average 11x compared to a single CPU thread when parallelizing computation across multiple cores, graphics-processing units, and field-programmable gate arrays. John Robert Wernsing, Greg Stitt |
PPoPP | 2 |
| 2012 | Elastic computing: A portable optimization framework for hybrid computers
John Robert Wernsing, Greg Stitt |
Parallel Comput. | 2 |
| 2012 | RCML: An Environment for Estimation Modeling of Reconfigurable Computing SystemsabstractReconfigurable computing (RC) is emerging as a promising area for embedded computing, in which complex systems must balance performance, flexibility, cost, and power. The difficulty associated with RC development suggests improved strategic planning and analysis techniques can save significant development time and effort. This article presents a new abstract modeling language and environment, the RC Modeling Language (RCML), to facilitate efficient design space exploration of RC systems at the estimation modeling level, that is, before building a functional implementation. Two integrated analysis tools and case studies, one analytical and one simulative, are presented illustrating relatively accurate automated analysis of systems modeled in RCML. Casey Reardon, Brian Holland, Alan D. George, Greg Stitt, Herman Lam |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2012 | SCF: A Framework for Task-Level Coordination in Reconfigurable, Heterogeneous SystemsabstractHeterogeneous computing systems comprised of accelerators such as FPGAs, GPUs, and manycore processors coupled with standard microprocessors are becoming an increasingly popular solution for future computing systems due to their higher performance and energy efficiency. Although programming languages and tools are evolving to simplify device-level design, programming such systems is still difficult and time-consuming largely due to system-wide challenges involving communication between heterogeneous devices, which currently require ad hoc solutions. Most communication frameworks and APIs which have dominated parallel application development for decades were developed for homogeneous systems, and hence cannot be directly employed for hybrid systems. To solve this problem, this article presents the System Coordination Framework (SCF), which employs message passing to transparently enable communication between tasks described using different programming tools (and languages), and running on heterogeneous processing devices of systems from domains ranging from embedded systems to High-Performance Computing (HPC) systems. By hiding low-level architectural details of the underlying communication from an application designer, SCF can improve application development productivity, provide higher levels of application portability, and offer rapid design-space exploration of different task/device mappings. In addition, SCF enables custom communication synthesis that exploits mechanisms specific to different devices and platforms, which can provide performance improvements over generic solutions employed previously. Our results indicate a performance improvement of 28× and 682× by employing FPGA devices for two applications presented in this article, while simultaneously improving the developer productivity by approximately 2.5 to 5 times by using SCF. Vikas Aggarwal, Greg Stitt, Alan D. George, Changil Yoon |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2011 | Thread Warping: Dynamic and Transparent Synthesis of Thread AcceleratorsabstractWe introduce thread warping, a dynamic optimization technique that customizes multicore architectures to a given application by dynamically synthesizing threads into custom accelerator circuits on FPGAs (Field-Programmable Gate Arrays). Thread warping builds upon previous dynamic synthesis techniques for single-threaded applications, enabling dynamic architectural adaptation to different amounts of thread-level parallelism, while also exploiting parallelism within each thread to further improve performance. Furthermore, thread warping maintains the important separation of function from architecture, enabling portability of applications to architectures with different quantities of microprocessors and FPGAs, an advantage not shared by static compilation/synthesis approaches. We introduce an approach consisting of CAD tools and operating system support that enables thread warping on potentially any microprocessor/FPGA architecture. We evaluate thread warping using a simulator for high-performance computing systems with different interconnections in addition to multicore embedded systems having between 4 and 64 ARM11 microprocessors. On average, thread warping achieved approximately 3x speedup compared to a high-performance quad-core Intel Xeon and 109x compared to an embedded system consisting of 4 ARM11 cores, with a size cost approximately equal to 36 ARM11 cores. Greg Stitt, Frank Vahid |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2011 | Platform-aware bottleneck detection for reconfigurable computing applicationsabstractReconfigurable Computing (RC) has the potential to provide substantial performance benefits and yet simultaneously consume less power than traditional microprocessors or GPUs. While experimental performance analysis of RC applications has previously been shown crucial for achieving this potential, existing methods still require application designers to manually locate bottlenecks and determine appropriate optimizations, typically requiring significant designer expertise and effort. Worse, the diversity of platforms employed by RC applications further complicates the process of detecting bottlenecks and formulating optimizations. To address these shortcomings, we first discuss our platform-template system, which enables a performance analysis tool to perform more accurate bottleneck detection and achieve a higher degree of portability across diverse FPGA systems. We then provide details for our implementation of these concepts and techniques in the Reconfigurable Computing Application Performance (ReCAP) tool. Next, we present a taxonomy of common RC bottlenecks, providing associated detection and optimization strategies for each bottleneck, which we use to populate ReCAP's knowledge base for bottleneck detection. Finally, we demonstrate the utility of our approach via two application case studies across a total of three platforms. Seth Koehler, Greg Stitt, Alan D. George |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2010 | Elastic computing: a framework for transparent, portable, and adaptive multi-core heterogeneous computingabstractOver the past decade, system architectures have started on a clear trend towards increased parallelism and heterogeneity, often resulting in speedups of 10x to 100x. Despite numerous compiler and high-level synthesis studies, usage of such systems has largely been limited to device experts, due to significantly increased application design complexity. To reduce application design complexity, we introduce elastic computing - a framework that separates functionality from implementation details by enabling designers to use specialized functions, called elastic functions, which enable an optimization framework to explore thousands of possible implementations, even ones using different algorithms. Elastic functions allow designers to execute the same application code efficiently on potentially any architecture and for different runtime parameters such as input size, battery life, etc. In this paper, we present an initial elastic computing framework that transparently optimizes application code onto diverse systems, achieving significant speedups ranging from 1.3x to 46x on a hyper-threaded Xeon system with an FPGA accelerator, a 16-CPU Opteron system, and a quad-core Xeon system. John Robert Wernsing, Greg Stitt |
LCTES | 2 |
| 2008 | C is for circuits: capturing FPGA circuits as sequential code for portabilityabstractSynthesizing common sequential algorithms, captured in a language like C, to FPGA circuits is now well-known to provide dramatic speedups for numerous applications, and to provide tremendous portability and adaptability advantages over circuit implementations of an application. However, many applications targeted to FPGAs are still designed and distributed at the circuit level, due in part to tremendous human ingenuity being exercised at that level to achieve exceptional performance and efficiency. A question then arises as to whether applications for FPGAs will have to be distributed as circuits to achieve desired performance and efficiency, or if instead a more portable language like C might be used. Given a set of common synthesis transformations, we studied the extent to which circuits published in FCCM in the past 6 years could be captured as sequential code and then synthesized back to the published circuit. The study showed that a surprising 82% of the 35 circuits chosen for the study could be re-derived from some form of standard C code, suggesting that standard C code, without extensions, may be an effective means for distributing FPGA applications Scott Sirowy, Greg Stitt, Frank Vahid |
FPGA | 2 |
| 2008 | Hardware/software partitioning with multi-version implementation explorationabstractHardware/software partitioning is an increasingly common technique that maps critical regions of a software application into custom hardware to achieve application speedup. Most previous partitioning approaches assume that each application region has only a single hardware implementation. However, code regions typically can be implemented as many different versions that tradeoff performance and area, as in the case of a loop that can be unrolled by different amounts. We introduce a new formulation of hardware/software partitioning that integrates multiple versions of region implementations, improving performance by more than 27% on average compared to partitioning with a single implementation. We present an optimal ILP solution, and introduce an efficient heuristic that achieves solutions within 0% to 8% of the optimal while running in less than one second for large problem sizes. Greg Stitt |
ACM Great Lakes Symposium on VLSI | 1 |
| 2008 | Recursion flatteningabstractHigh-level synthesis tools automatically generate custom hardware circuits from high-level languages, including popular programming languages like standard ANSI C, but are unable to handle recursive functions. The convenience of recursive algorithms has made recursion a widespread programming practice, therefore limiting the applicability of high-level synthesis tools. We introduce a new synthesis technique, recursion flattening, that reduces the limitations caused by recursion for high-level synthesis. Recursion flattening can eliminate many instances of recursion by determining recursion depth, and then inlining recursive calls. Recursion flattening cannot eliminate all recursion, but we show that the technique succeeds for many common recursive algorithms. We applied the technique to seven recursive benchmarks that previously would not have been synthesizable, resulting in FPGA hardware circuits that run 75x faster on average than if the benchmark were run as microprocessor software. Furthermore, we compared those hardware circuits to circuits synthesized from the same benchmarks coded using non-recursive algorithms, and show nearly identical performance and area for many examples, and significantly increased performance for several examples. Greg Stitt, Jason R. Villarreal |
ACM Great Lakes Symposium on VLSI | 1 |
| 2007 | Binary synthesisabstractRecent high-level synthesis approaches and C-based hardware description languages attempt to improve the hardware design process by allowing developers to capture desired hardware functionality in a well-known high-level source language. However, these approaches have yet to achieve wide commercial success due in part to the difficulty of incorporating such approaches into software tool flows. The requirement of using a specific language, compiler, or development environment may cause many software developers to resist such approaches due to the difficulty and possible instability of changing well-established robust tool flows. Thus, in the past several years, synthesis from binaries has been introduced, both in research and in commercial tools, as a means of better integrating with tool flows by supporting all high-level languages and software compilers. Binary synthesis can be more easily integrated into a software development tool-flow by only requiring an additional backend tool, and it even enables completely transparent dynamic translation of executing binaries to configurable hardware circuits. In this article, we survey the key technologies underlying the important emerging field of binary synthesis. We compare binary synthesis to several related areas of research, and we then describe the key technologies required for effective binary synthesis: decompilation techniques necessary for binary synthesis to achieve results competitive with source-level synthesis, hardware/software partitioning methods necessary to find critical binary regions suitable for synthesis, synthesis methods for converting regions to custom circuits, and binary update methods that enable replacement of critical binary regions by circuits. Greg Stitt, Frank Vahid |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2006 | A code refinement methodology for performance-improved synthesis from CabstractAlthough many recent advances have been made in hardware synthesis techniques from software programming languages such as C, the performance of synthesized hardware commonly suffers due to the use of C constructs and coding practices that are not appropriate for hardware. Most previous approaches to addressing this problem require drastic changes to coding practice. We present an approach that instead requires only minimal changes but yields significant speedups. In this approach, a software developer initially writes C code as they normally would, and then applies simple refinement guidelines to only the performance-critical code regions, which are the regions most likely to be synthesized to hardware. Alternatively, if a designer is aware of performance-critical parts of the application, the guidelines could be followed during development. In this study, we analyze dozens of embedded benchmarks to determine the most common C coding practices that limit hardware performance, and introduce coding guidelines to make the code more amenable to synthesis. Those guidelines typically require minimal coding effort, generally consisting of less than ten lines of code for each guideline. The guidelines typically represent modifications that require designer knowledge, making the guidelines difficult or impossible for synthesis tools to automate. We apply these guidelines to six benchmarks, resulting in average speedups of 3.5x compared to synthesis from the original code with a negligible software size and performance overhead. Greg Stitt, Frank Vahid, Walid A. Najjar |
ICCAD | 1 |
| 2006 | Warp ProcessorsabstractWe describe a new processing architecture, known as a warp processor, that utilizes a field-programmable gate array (FPGA) to improve the speed and energy consumption of a software binary executing on a microprocessor. Unlike previous approaches that also improve software using an FPGA but do so using a special compiler, a warp processor achieves these improvements completely transparently and operates from a standard binary. A warp processor dynamically detects the binary's critical regions, reimplements those regions as a custom hardware circuit in the FPGA, and replaces the software region by a call to the new hardware implementation of that region. While not all benchmarks can be improved using warp processing, many can, and the improvements are dramatically better than those achievable by more traditional architecture improvements. The hardest part of warp processing is that of dynamically reimplementing code regions on an FPGA, requiring partitioning, decompilation, synthesis, placement, and routing tools, all having to execute with minimal computation time and data memory so as to coexist on chip with the main processor. We describe the results of developing our warp processor. We developed a custom FPGA fabric specifically designed to enable lean place and route tools, and we developed extremely fast and efficient versions of partitioning, decompilation, synthesis, technology mapping, placement, and routing. Warp processors achieve overall application speedups of 6.3X with energy savings of 66% across a set of embedded benchmark applications. We further show that our tools utilize acceptably small amounts of computation and memory which are far less than traditional tools. Our work illustrates the feasibility and potential of warp processing, and we can foresee the possibility of warp processing becoming a feature in a variety of computing domains, including desktop, server, and embedded applications. Roman L. Lysecky, Greg Stitt, Frank Vahid |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2005 | A Decompilation Approach to Partitioning Software for Microprocessor/FPGA PlatformsabstractWe present a software compilation approach for microprocessor/FPGA platforms that partitions a software binary onto custom hardware implemented in the FPGA. Our approach imposes fewer restrictions on software tool flow than previous compiler approaches, allowing software designers to use any software language and compiler. Our approach uses a back-end partitioning tool that utilizes decompilation techniques to recover important high-level information, resulting in performance comparable to high-level compiler-based approaches. Greg Stitt, Frank Vahid |
DATE | 1 |
| 2005 | Techniques for synthesizing binaries to an advanced register/memory structureabstractRecent works demonstrate several benefits of synthesizing software binaries onto FPGA hardware, including incorporating hardware design into established software tool flows with minimal impact, porting existing binaries to FPGAs, and even dynamically synthesizing software kernels to faster FPGA coprocessors. Those works showed that standard binary decompilation methods can recover enough high-level control information to result in reasonably-efficient hardware. However, recent synthesis methods for FPGAs utilize advanced memory structures, such as a "smart buffer," that require recovery of additional high-level information, specifically information about loops and arrays. We incorporate decompilation techniques into an existing binary synthesis tool flow to recover loops and arrays in order to take advantage of advanced memory structures when performing synthesis from a binary. We demonstrate through experiments on six benchmarks that our methods improve binary synthesis performance by 53%, by making effective use of smart buffers. Furthermore, we compare the binary results using smart buffers with results of synthesis directly from the original C code for the benchmarks, and show that our methods achieved almost identical performance results with only 10% area overhead. Greg Stitt, Zhi Guo, Walid A. Najjar, Frank Vahid |
FPGA | 1 |
| 2004 | Energy savings and speedups from partitioning critical software loops to hardware in embedded systemsabstractWe present results of extensive hardware/software partitioning experiments on numerous benchmarks. We describe our loop-oriented partitioning methodology for moving critical code from hardware to software. Our benchmarks included programs from PowerStone, MediaBench, and NetBench. Our experiments included estimated results for partitioning using an 8051 8-bit microcontroller or a 32-bit MIPS microprocessor for the software, and using on-chip configurable logic or custom application-specific integrated circuit hardware for the hardware. Additional experiments involved actual measurements taken from several physical implementations of hardware/software partitionings on real single-chip microprocessor/configurable-logic devices. We also estimated results assuming voltage scalable processors. We provide performance, energy, and size data for all of the experiments. We found that the benchmarks spent an average of 80% of their execution time in only 3% of their code, amounting to only about 200 bytes of critical code. For various experiments, we found that moving critical code to hardware resulted in average speedups of 3 to 5 and average energy savings of 35% to 70%, with average hardware requirements of only 5000 to 10,000 gates. To our knowledge, these experiments represent the most comprehensive hardware/software partitioning study published to date. Greg Stitt, Frank Vahid, Shawn Nematbakhsh |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2003 | Dynamic hardware/software partitioning: a first approachabstractPartitioning an application among software running on a microprocessor and hardware co-processors in on-chip configurable logic has been shown to improve performance and energy consumption in embedded systems. Meanwhile, dynamic software optimization methods have shown the usefulness and feasibility of runtime program optimization, but those optimizations do not achieve as much as partitioning. We introduce a first approach to dynamic hardware/software partitioning. We describe our system architecture and initial on-chip tools, including profiler, decompiler, synthesis, and placement and routing tools for a simplified configurable logic fabric, able to perform dynamic partitioning of real benchmarks. We show speedups averaging 2.6 for five benchmarks taken from Powerstone, NetBench, and our own benchmarks. Greg Stitt, Roman L. Lysecky, Frank Vahid |
DAC | 1 |
| 2003 | Profiling tools for hardware/software partitioning of embedded applicationsabstractLoops constitute the most executed segments of programs and therefore are the best candidates for hardware software partitioning. We present a set of profiling tools that are specifically dedicated to loop profiling and do support combined function and loop profiling. One tool relies on an instruction set simulator and can therefore be augmented with architecture and micro-architecture features simulation while the other is based on compile-time instrumentation of gcc and therefore has very little slow down compared to the original program We use the results of the profiling to identify the compute core in each benchmark and study the effect of compile-time optimization on the distribution of cores in a program. We also study the potential speedup that can be achieved using a configurable system on a chip, consisting of a CPU embedded on an FPGA, as an example application of these tools in hardware/software partitioning. Dinesh C. Suresh, Walid A. Najjar, Frank Vahid, Jason R. Villarreal, Greg Stitt |
LCTES | 5 |
| 2002 | Using On-Chip Configurable Logic to Reduce Embedded System Software EnergyabstractWe examine the energy savings possible by re-mapping critical software loops from a microprocessor to configurable logic appearing on the same-chip in commodity chips now commercially available. That logic is typically intended to implement peripherals and coprocessors without increasing chip count-but we show that reduced software energy is an additional benefit, making such chips even more useful. We find critical software loops and re-implement them in the configurable logic such that a repeating software task completes sooner, allowing us to put the system in a low-power state for longer periods, thus reducing energy. We use simulations and estimations for a hypothetical device having a 32-bit MIPS processor plus configurable logic, yielding energy savings of 25%, increasing to 39% assuming voltage scaling. We physically measured several examples running on two commercial single-chip devices having an 8-bit 8051 microprocessor plus configurable logic and a 32-bit ARM microprocessor with configurable logic, with energy savings of 71% and 53% respectively, increasing to an estimated 89% and 75% assuming voltage scaling. Greg Stitt, Brian Grattan, Jason R. Villarreal, Frank Vahid |
FCCM | 1 |
| 2002 | Hardware/software partitioning of software binariesabstractPartitioning an embedded system application among a microprocessor and custom hardware has been shown to improve the performance, power or energy of numerous examples. The advent of single-chip microprocessor/FPGA platforms makes such partitioning even more attractive. Previous partitioning approaches have partitioned sequential program source code, such as C or C++. We introduce a new approach that partitions at the software binary level. Although source code partitioning is preferable from a purely technical viewpoint, binary-level partitioning provides several very practical benefits for commercial acceptance. We demonstrate that binary-level partitioning yields competitive speedup results compared to source-level partitioning, achieving an average speedup of 1.4 compared to 1.5 for eight benchmarks partitioned on a single-chip microprocessor/FPGA device. Greg Stitt, Frank Vahid |
ICCAD | 1 |
| 2000 | A first-step towards an architecture tuning methodology for low powerabstractWe describe an automated environment to assist a system-on-achip designer to tune a microprocessor core to a particular application program that will run on the microprocessor, and vice-versa, with the goal of reducing embedded system power consumption.We limit such tuning to modifications that do not change the microprocessor instruction set, thus avoiding the large costs that would come with such a change.Our tuning environment for the 8051 microcontroller is freely-available on the web. Greg Stitt, Frank Vahid, Tony Givargis, Roman L. Lysecky |
CASES | 1 |