VLDB 2026 Research / reviewers in the wild / expert
Kyle Rupnow
dblp:27/5122
· DBLP profile ↗
39ranked-venue papers
7as first author
0since 2021 · last 2019
0000-0003-2908-2225ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 7 first-authorSoftware engineering, systems software and programming languages · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
14 papers |
Electronic design automation · 51% Reconfigurable computing and FPGAs · 16% Hardware accelerators and domain-specific architectures · 14% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 82% Debugging and program repair · 18% |
Topics — the 30 heaviest of 38, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
high-level synthesis |
1.9 | 8 | 2019 | Hybrid Quick Error Detection: Validation and Debug of SoCs Through High-Level Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 FCUDA-HB: Hierarchical and Scalable Bus Architecture Generation on FPGAs With the FCUDA Flow · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 Automated Verification Code Generation in HLS Using Software Execution Traces (Abstract Only) · FPGA 2016 |
Electronic design automation
hardware verification and test |
0.6 | 2 | 2019 | Hybrid Quick Error Detection: Validation and Debug of SoCs Through High-Level Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 Debugging and verifying SoC designs through effective cross-layer hardware-software co-simulation · DAC 2016 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator |
0.4 | 1 | 2019 | FPGA/DNN Co-Design: An Efficient Design Methodology for IoT Intelligence on the Edge · DAC 2019 |
Hardware reliability and fault tolerance
error detection |
0.4 | 1 | 2019 | Hybrid Quick Error Detection: Validation and Debug of SoCs Through High-Level Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Reconfigurable computing and FPGAs › FPGA accelerator
FPGA accelerator design |
0.4 | 1 | 2019 | FPGA/DNN Co-Design: An Efficient Design Methodology for IoT Intelligence on the Edge · DAC 2019 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.4 | 1 | 2019 | FPGA/DNN Co-Design: An Efficient Design Methodology for IoT Intelligence on the Edge · DAC 2019 |
Electronic design automation › hardware verification and test › design validation
post-silicon validation |
0.4 | 1 | 2019 | Hybrid Quick Error Detection: Validation and Debug of SoCs Through High-Level Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Hardware accelerators and domain-specific architectures
bioinformatics accelerator |
0.3 | 1 | 2017 | Hardware Acceleration of the Pair-HMM Algorithm for DNA Variant Calling · FPGA 2017 |
Electronic design automation › hardware verification and test › debugging
bug localization |
0.2 | 1 | 2016 | Debugging and verifying SoC designs through effective cross-layer hardware-software co-simulation · DAC 2016 |
GPUs and heterogeneous computing
control flow divergence |
0.2 | 1 | 2016 | An Accurate GPU Performance Model for Effective Control Flow Divergence Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 |
Electronic design automation › hardware/software co-design
co-simulation |
0.2 | 1 | 2016 | Debugging and verifying SoC designs through effective cross-layer hardware-software co-simulation · DAC 2016 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.2 | 1 | 2016 | High Level Synthesis of Complex Applications: An H.264 Video Decoder · FPGA 2016 |
Reconfigurable computing and FPGAs › FPGA design flow
FPGA design space exploration |
0.2 | 1 | 2016 | FCUDA-HB: Hierarchical and Scalable Bus Architecture Generation on FPGAs With the FCUDA Flow · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 |
Reconfigurable computing and FPGAs
FPGA routing architecture |
0.2 | 1 | 2016 | FCUDA-HB: Hierarchical and Scalable Bus Architecture Generation on FPGAs With the FCUDA Flow · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 |
Performance modeling and evaluation › processor performance modeling › accelerator performance modeling
GPU performance modeling |
0.2 | 1 | 2016 | An Accurate GPU Performance Model for Effective Control Flow Divergence Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 |
Electronic design automation › hardware verification and test
hardware verification |
0.2 | 1 | 2016 | Automated Verification Code Generation in HLS Using Software Execution Traces (Abstract Only) · FPGA 2016 |
Electronic design automation › hardware verification and test › hardware verification
high-level synthesis verification |
0.2 | 1 | 2016 | Automated Verification Code Generation in HLS Using Software Execution Traces (Abstract Only) · FPGA 2016 |
Electronic design automation › hardware verification and test › hardware verification
soc design verification |
0.2 | 1 | 2016 | Debugging and verifying SoC designs through effective cross-layer hardware-software co-simulation · DAC 2016 |
Hardware accelerators and domain-specific architectures
video decoder |
0.2 | 1 | 2016 | High Level Synthesis of Complex Applications: An H.264 Video Decoder · FPGA 2016 |
GPUs and heterogeneous computing
GPU resource management |
0.2 | 1 | 2015 | Efficient GPU Spatial-Temporal Multitasking · IEEE Trans. Parallel Distributed Syst. 2015 |
GPUs and heterogeneous computing
GPU sharing |
0.2 | 1 | 2015 | Efficient GPU Spatial-Temporal Multitasking · IEEE Trans. Parallel Distributed Syst. 2015 |
Electronic design automation › timing analysis
critical path analysis |
0.2 | 1 | 2014 | High-Level Synthesis With Behavioral-Level Multicycle Path Analysis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014 |
Electronic design automation
false positive elimination |
0.2 | 1 | 2014 | High-Level Synthesis With Behavioral-Level Multicycle Path Analysis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014 |
Reconfigurable computing and FPGAs
FPGA design flow |
0.2 | 1 | 2014 | Fast and effective placement and routing directed high-level synthesis for FPGAs · FPGA 2014 |
Electronic design automation
timing analysis |
0.2 | 1 | 2014 | High-Level Synthesis With Behavioral-Level Multicycle Path Analysis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014 |
Compilers and program optimization
loop transformation |
0.2 | 1 | 2013 | Improving high level synthesis optimization opportunity through polyhedral transformations · FPGA 2013 |
Compilers and program optimization
polyhedral model |
0.2 | 1 | 2013 | Improving high level synthesis optimization opportunity through polyhedral transformations · FPGA 2013 |
Edge and fog computing
edge intelligence |
0.1 | 1 | 2019 | FPGA/DNN Co-Design: An Efficient Design Methodology for IoT Intelligence on the Edge · DAC 2019 |
Performance modeling and evaluation › simulation
simulation-based evaluation |
0.1 | 1 | 2010 | Accurately evaluating application performance in simulated hybrid multi-tasking systems · FPGA 2010 |
Bioinformatics and computational biology › genomics
variant calling |
0.1 | 1 | 2017 | Hardware Acceleration of the Pair-HMM Algorithm for DNA Variant Calling · FPGA 2017 |
Methods — techniques the papers use, named apart from their topics
high-level synthesis · 0.8hardware-oriented model search · 0.8FPGA/DNN co-design · 0.8Pair-HMM · 0.6code conversion for synthesizability · 0.5HLS optimization · 0.5quick error detection · 0.4watchdog timer · 0.2software execution traces · 0.2instrumentation · 0.2cycle-accurate systemc simulation · 0.2cross-layer co-simulation · 0.2CUDA · 0.2polyhedral model · 0.2loop transformation · 0.2code generation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | FPGA/DNN Co-Design: An Efficient Design Methodology for IoT Intelligence on the EdgeabstractWhile embedded FPGAs are attractive platforms for DNN acceleration on edge-devices due to their low latency and high energy efficiency, the scarcity of resources of edge-scale FPGA devices also makes it challenging for DNN deployment. In this paper, we propose a simultaneous FPGA/DNN co-design methodology with both bottom-up and top-down approaches: a bottom-up hardware-oriented DNN model search for high accuracy, and a top-down FPGA accelerator design considering DNN-specific characteristics. We also build an automatic co-design flow, including an Auto-DNN engine to perform hardware-oriented DNN model search, as well as an Auto-HLS engine to generate synthesizable C code of the FPGA accelerator for explored DNNs. We demonstrate our co-design approach on an object detection task using PYNQ-Z1 FPGA. Results show that our proposed DNN model and accelerator outperform the state-of-the-art FPGA designs in all aspects including Intersection-over-Union (IoU) (6.2% higher), frames per second (FPS) (2.48× higher), power consumption (40% lower), and energy efficiency (2.5× higher). Compared to GPU-based solutions, our designs deliver similar accuracy but consume far less energy. Cong Hao, Xiaofan Zhang 0001, Sitao Huang, Jinjun Xiong, Kyle Rupnow, Wen-Mei W. Hwu, Deming Chen |
DAC | 6 |
| 2019 | Hybrid Quick Error Detection: Validation and Debug of SoCs Through High-Level SynthesisabstractValidation and debug challenges of system-on-chips (SoCs) are getting increasingly difficult. As we reach the limits of Dennard scaling, efforts to improve system performance and energy efficiency have resulted in the integration of a wide variety of complex hardware accelerators in SoCs. Hence, it is essential to address the validation and debug of hardware accelerators. High-level synthesis (HLS) is a promising technique to rapidly create customized hardware accelerators. In this paper, we present the hybrid quick error detection (H-QED) approach that overcomes validation and debug challenges for hardware accelerators by leveraging HLS techniques in both the presilicon and post-silicon stages. H-QED improves error detection latencies (time elapsed from when a bug is activated to when it is detected) by 2-5 orders of magnitude with one cycle latencies in presilicon scenarios and bug coverage threefold higher compared to traditional validation techniques. H-QED also uncovered previously unknown bugs in the CHStone benchmark suite, which is widely used by the HLS community. H-QED incurs an 8% accelerator area overhead with negligible silicon performance impact for post-silicon stage, and we also introduce techniques to minimize any possible intrusiveness introduced by H-QED. Keith A. Campbell, Leon He, Swathi T. Gurumani, Kyle Rupnow, Subhasish Mitra, Deming Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2017 | Hardware Acceleration of the Pair-HMM Algorithm for DNA Variant Calling
Sitao Huang, Gowthami Jayashri Manikandan, Anand Ramachandran 0001, Kyle Rupnow, Wen-Mei W. Hwu, Deming Chen |
FPGA | 4 |
| 2017 | High-performance video content recognition with long-term recurrent convolutional network for FPGAabstractFPGA is a promising candidate for the acceleration of Deep Neural Networks (DNN) with improved latency and energy consumption compared to CPU and GPU-based implementations. DNNs use sequences of layers of regular computation that are well suited for HLS-based design for FPGA. However, optimizing large neural networks under resource constraints is still a key challenge. HLS must manage on-chip computation, buffering resources, and off-chip memory accesses to minimize the total latency. In this paper, we present a design framework for DNNs that uses highly configurable IPs for neural network layers together with a new design space exploration engine for Resource Allocation Management (REALM). We also carry out efficient memory subsystem design and fixed-point weight re-training to further improve our FPGA solution. We demonstrate our design framework on the Long-term Recurrent Convolution Network for video inputs. Our implementation on a Xilinx VC709 board achieves 3.1X speedup compared to an NVIDIA K80 and 4.75X speedup compared to an Intel Xeon with 17.5X lower energy per image. Xiaofan Zhang 0001, Xinheng Liu, Anand Ramachandran 0001, Chuanhao Zhuge, Shibin Tang, Zuofu Cheng, Kyle Rupnow, Deming Chen |
FPL | 8 |
| 2017 | Machine learning on FPGAs to face the IoT revolutionabstractFPGAs have been rapidly adopted for acceleration of Deep Neural Networks (DNNs) with improved latency and energy efficiency compared to CPU and GPU-based implementations. High-level synthesis (HLS) is an effective design flow for DNNs due to improved productivity, debugging, and design space exploration ability. However, optimizing large neural networks under resource constraints for FPGAs is still a key challenge. In this paper, we present a series of effective design techniques for implementing DNNs on FPGAs with high performance and energy efficiency. These include the use of configurable DNN IPs, performance and resource modeling, resource allocation across DNN layers, and DNN reduction and re-training. We showcase several design solutions including Long-term Recurrent Convolution Network (LRCN) for video captioning, Inception module for FaceNet face recognition, as well as Long Short-Term Memory (LSTM) for sound recognition. These and other similar DNN solutions are ideal implementations to be deployed in vision or sound based IoT applications. Xiaofan Zhang 0001, Anand Ramachandran 0001, Chuanhao Zhuge, Di He 0004, Wei Zuo, Zuofu Cheng, Kyle Rupnow, Deming Chen |
ICCAD | 7 |
| 2017 | Machine learning on FPGAs to face the IoT revolutionabstractFPGAs have been rapidly adopted for acceleration of Deep Neural Networks (DNNs) with improved latency and energy efficiency compared to CPU and GPU-based implementations. High-level synthesis (HLS) is an effective design flow for DNNs due to improved productivity, debugging, and design space exploration ability. However, optimizing large neural networks under resource constraints for FPGAs is still a key challenge. In this paper, we present a series of effective design techniques for implementing DNNs on FPGAs with high performance and energy efficiency. These include the use of configurable DNN IPs, performance and resource modeling, resource allocation across DNN layers, and DNN reduction and re-training. We showcase several design solutions including Long-term Recurrent Convolution Network (LRCN) for video captioning, Inception module for FaceNet face recognition, as well as Long Short-Term Memory (LSTM) for sound recognition. These and other similar DNN solutions are ideal implementations to be deployed in vision or sound based IoT applications. Xiaofan Zhang 0001, Anand Ramachandran 0001, Chuanhao Zhuge, Di He 0004, Wei Zuo, Zuofu Cheng, Kyle Rupnow, Deming Chen |
ICCAD | 7 |
| 2016 | Designing high-quality hardware on a development effort budget: A study of the current state of high-level synthesisabstractHigh-level synthesis (HLS) promises high-quality hardware with minimal development effort. In this paper, we evaluate the current state-of-the-art in HLS and design techniques based on software references and architecture references. We present a software reference study developing a JPEG encoder from pre-existing software, and an architecture reference study developing an AES block encryption module from scratch in SystemC and SystemVerilog based on a desired architecture. Additionally, we develop micro-benchmarks to demonstrate best-practices in C coding styles that produce high-quality hardware with minimal development effort. Finally, we suggest language, tool, and methodology improvements to improve upon the current state-of-the-art in HLS. Zelei Sun, Keith A. Campbell, Wei Zuo, Kyle Rupnow, Swathi T. Gurumani, Frederic Doucet, Deming Chen |
ASP-DAC | 4 |
| 2016 | Debugging and verifying SoC designs through effective cross-layer hardware-software co-simulationabstractVerification of modern day electronic circuits has become the bottleneck for the timely delivery of complex SoC designs. We develop a novel cross-layer hardware/software co-simulation framework that can effectively debug and verify an SoC design. We combine high-level C/C++ software with cycle-accurate SystemC hardware, uniquely identify various types of bugs, and help the hardware designer localize them. Experimental results show that we are able to detect and aid in localization of logic bugs from both C/C++ specifications as well as the high-level synthesis engine itself. Our framework is fully automated, representing an important step forward targeting fast and effective SoC design verification. Keith A. Campbell, Leon He, Swathi T. Gurumani, Kyle Rupnow, Deming Chen |
DAC | 5 |
| 2016 | Real-time system-level implementation of a telepresence robot using an embedded GPU platform
Muhammad Teguh Satria, Swathi T. Gurumani, Keng Peng Tee, Augustine Koh, Pan Yu, Kyle Rupnow, Deming Chen |
DATE | 7 |
| 2016 | Acceleration of the Pair-HMM Algorithm for DNA Variant CallingabstractIn this project, we propose an SoC solution to accelerate the Pair-HMM's forward algorithm which is the key performance bottleneck in the GATK's HaplotypeCaller tool for DNA variant calling. We develop two versions of the Pair-HMM accelerator: one using High Level Synthesis (HLS), and another ring-based manual RTL implementation. We investigate the performance of the manual RTL design and HLS design in terms of design flexibility and overall run-time. We achieve a significant speed-up of up to 19x through the HLS implementation and speed-up of up to 95x through the RTL implementation of the algorithm. Gowthami Jayashri Manikandan, Sitao Huang, Kyle Rupnow, Wen-Mei W. Hwu, Deming Chen |
FCCM | 3 |
| 2016 | AutoSLIDE: Automatic Source-Level Instrumentation and Debugging for HLSabstractImproved quality of results from high level synthesis (HLS) tools have led to their increased adoption in hardware design. However, functional verification of HLS-produced designs remains a major challenge. Once a bug is exposed, designers must backtrace thousands of signals and simulation cycles to determine the underlying cause. The challenge is further exacerbated with HLS-produced non-human-readable RTL. In this paper, we present AutoSLIDE, an automated cross-layer verification framework that instruments critical operations, detects discrepancies between software and hardware execution, and traces the suspect datapath tree to identify bug source for the detected discrepancy. AutoSLIDE also maintains mappings between RTL datapath operations, LLVM-IR operations, and C/C++ source code to precisely pinpoint the root-cause of bugs to the exact line/operation in source code, substantially reducing user effort to localize bugs. We demonstrate the effectiveness by detecting and localizing bugs from former versions of the CHStone benchmark suite. Furthermore, we demonstrate the efficiency of AutoSLIDE, with low overhead in HLS time (27%), software trace gathering (10%), and significantly reduced trace size and simulation time compared to exhaustive instrumentation. Swathi T. Gurumani, Deming Chen, Kyle Rupnow |
FCCM | 4 |
| 2016 | High Level Synthesis of Complex Applications: An H.264 Video DecoderabstractHigh level synthesis (HLS) is gaining wider acceptance for hardware design due to its higher productivity and better design space exploration features. In recent years, HLS techniques and design flows have also advanced significantly, and as a result, many new FPGA designs are developed with HLS. However, despite many studies using HLS, the size and complexity of such applications remain generally small, and it is not well understood how to design and optimize for HLS with large, complex reference code. Typical HLS benchmark applications contain somewhere between 100 to 1400 lines of code and about 20 sub-functions, but typical input applications may contain many times more code and functions. To study such complex applications, we present a case study using HLS for a full H.264 decoder: an application with over 6000 lines of code and over 100 functions. We share our experience on code conversion for synthesizability, various HLS optimizations, HLS limitations while dealing with complex input code, and general design insights. Through our optimization process, we achieve 34 frames/s at 640x480 resolution (480p). To enable future study and benefit the research community, we open-source our synthe- sizable H.264 implementation. Xinheng Liu, Yao Chen 0008, Swathi T. Gurumani, Kyle Rupnow, Deming Chen |
FPGA | 5 |
| 2016 | FCUDA-SoC: Platform Integration for Field-Programmable SoC with the CUDA-to-FPGA CompilerabstractThroughput oriented high level synthesis allows efficient design and optimization using parallel input languages. Parallel languages offer the benefit of parallelism extraction at multiple levels of granularity, offering effective design space exploration to select efficient single core implementations, and easy scaling of parallelism through multiple core instantiations. However, study of high level synthesis for parallel languages has concentrated on optimization of core and on-chip communications, while neglecting platform integration, which can have a significant impact on achieved performance. In this paper, we create an automated flow to perform efficient platform integration for an existing CUDA-to-RTL throughput oriented HLS, and we open source the FCUDA tool, platform integration, and benchmark applications. We demonstrate platform integration of 16 benchmarks on two Zynq-based systems in bare-metal and OS mode. We study implementation optimization for platform integration, compare to an embedded GPU (Tegra TK1) and verify designs on a Zedboard Zynq 7020 (bare-metal) and Omnitek Zynq 7045 (OS). Swathi T. Gurumani, Kyle Rupnow, Deming Chen |
FPGA | 3 |
| 2016 | Automated Verification Code Generation in HLS Using Software Execution Traces (Abstract Only)abstractImproved quality of results from high level synthesis (HLS) tools has led to their increased adoption. Despite the automated translation from high level descriptions to register-transfer level (RTL) implementations, functional verification remains a major challenge. Verification can take significantly more time than the design process; if there is a functional mismatch, developers must back-trace thousands of signals and cycles to determine underlying cause. The challenge is further exacerbated with HLS-produced RTL, which is often not human readable. To overcome these challenges, we present a verification technique that uses software-execution traces and automated insertion of verification code into the HLS-generated RTL to assist in debugging. The verification code helps pinpoint the earliest instance of RTL simulation mismatch, either caused by HLS engine bugs or design bugs, and related instructions. We also integrate a watchdog timer to examine the execution of control-flow and perform source-to-source transformation on benchmarks to take advantage of our proposed instrumentation. We also create a framework to insert various types of bugs, e.g. data-flow, control-flow and operational bugs, to evaluate our technique. We use the CHStone benchmark suite and demonstrate that our verification detects over 90% of the inserted bugs, with over 70% of them detected within 10 cycles. In addition, the proposed flow can detect real-life bugs existing in previously released versions of CHStone suite as well. Swathi T. Gurumani, Suhaib A. Fahmy, Deming Chen, Kyle Rupnow |
FPGA | 5 |
| 2016 | FCUDA-HB: Hierarchical and Scalable Bus Architecture Generation on FPGAs With the FCUDA FlowabstractRecent progress in high-level synthesis (HLS) has helped raise the abstraction level of hardware design. HLS flows reduce designer effort by allowing development in a high-level language, which improves debugging, code reuse and ability to explore different implementation options. However, although the HLS process is fast, implementation and performance analysis still require lengthy logic synthesis and physical design. For design optimization, HLS tools require design space exploration to obtain parallelism at multiple levels of granularity including parallelism within a single HLS-generated core and parallelism between multiple instances of cores. Core interconnect and external bandwidth limitations can significantly impact feasible options in the design space. With many dimensions in a design space exploration, it quickly becomes infeasible to perform full logic synthesis and physical design for each possible design point. However, generation and evaluation of communications infrastructure as part of the exploration is critical to determine the system performance. Thus, in this paper, we extend the prior multilevel granularity parallelism exploration in the FCUDA HLS flow, which takes CUDA code as design input and generates a corresponding field programmable gate array implementation. Our framework performs an initial characterization of the application design space, then analytically explores the design space considering parallelism, core interconnect, and external memory bandwidth, and selects a pare-to-optimal set of designs. Our flow is completely automated to perform the exploration to characterize the analytical model, perform the exploration, select a solution, and integrate multiple instantiations of FCUDA cores via an advanced extensible interface bus interconnect. Our results demonstrate that this new FCUDA flow efficiently identifies and generates implementations with up to 5× improved system performance compared to single-level granularity parallelism (core-level optimization). Yao Chen 0008, Swathi T. Gurumani, Yun Liang 0001, Kyle Rupnow, Jason Cong, Wen-Mei W. Hwu, Deming Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2016 | An Accurate GPU Performance Model for Effective Control Flow Divergence OptimizationabstractGraphic processing units (GPUs) are composed of a group of single-instruction multiple data (SIMD) streaming multiprocessors (SMs). GPUs are able to efficiently execute highly data parallel tasks through SIMD execution on the SMs. However, if those threads take diverging control paths, all divergent paths are executed serially. In the worst case, every thread takes a different control path and the highly parallel architecture is used serially by each thread. This control flow divergence problem is well known in GPU development; code transformation, memory access redirection, and data layout reorganization are commonly used to reduce the impact of divergence. These techniques attempt to eliminate divergence by grouping together threads or data to ensure identical behavior. However, prior efforts using these techniques do not model the performance impact of any particular divergence or consider that complete elimination of divergence may not be possible. Thus, we perform analysis of the performance impact of divergence and potential thread regrouping algorithms that eliminate divergence or minimize the impact of remaining divergence. Finally, we develop a divergence optimization framework that analyzes and transforms the kernel at compile-time and regroups the threads at runtime. For the compute-bound applications, our proposed metrics achieve performance estimation accuracy within 6.2% of measured performance. Using these metrics, we develop thread regrouping algorithms, which consider the impact of divergence, and speed up these applications by 2.2× on average on NVIDIA GTX480. Yun Liang 0001, Muhammad Teguh Satria, Kyle Rupnow, Deming Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | FCUDA-NoC: A Scalable and Efficient Network-on-Chip Implementation for the CUDA-to-FPGA FlowabstractHigh-level synthesis (HLS) of data-parallel input languages, such as the Compute Unified Device Architecture (CUDA), enables efficient description and implementation of independent computation cores. HLS tools can effectively translate the many threads of computation present in the parallel descriptions into independent, optimized cores. The generated hardware cores often heavily share input data and produce outputs independently. As the number of instantiated cores grows, the off-chip memory bandwidth may be insufficient to meet the demand. Hence, a scalable system architecture and a data-sharing mechanism become necessary for improving system performance. The network-on-chip (NoC) paradigm for intrachip communication has proved to be an efficient alternative to a hierarchical bus or crossbar interconnect, since it can reduce wire routing congestion, and has higher operating frequencies and better scalability for adding new nodes. In this paper, we present a customizable NoC architecture along with a directory-based data-sharing mechanism for an existing CUDA-to-FPGA (FCUDA) flow to enable scalability of our system and improve overall system performance. We build a fully automated FCUDA-NoC generator that takes in CUDA code and custom network parameters as inputs and produces synthesizable register transfer level (RTL) code for the entire NoC system. We implement the NoC system on a VC709 Xilinx evaluation board and evaluate our architecture with a set of benchmarks. The results demonstrate that our FCUDA-NoC design is scalable and efficient and we improve the system execution time by up to 63× and reduce external memory reads by up to 81% compared with a single hardware core implementation. Yao Chen 0008, Swathi T. Gurumani, Yun Liang 0001, Guofeng Li, Donghui Guo, Kyle Rupnow, Deming Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2015 | Behavioral-level IP integration in high-level synthesisabstractHigh level synthesis (HLS) quality improvements have led to its increased adoption in hardware design. In the design flow, IP reuse is critical for achieving quality of results, yet current HLS tools allow only a small set of tool-provided IPs integrated during HLS. General IP integration is then handled as an additional step either manually or using other system level tools. Performing post-HLS integration of IPs requires a clear separation of IPs from HLS-generated cores, requiring significant partitioning effort. In contrast, behavioral-level IP integration during HLS can simplify the design flow while still supporting HLS-based optimization and design space exploration. In this paper, we develop a general IP integration framework for HLS that supports fixed- and variable-latency IPs without requiring application partitioning. Using this framework that allows user-specified function/instruction-to-IP mapping, we demonstrate integration of both synthesizable and non-synthesizable IPs. Swathi T. Gurumani, Deming Chen, Kyle Rupnow |
FPT | 4 |
| 2015 | JIT trace-based verification for high-level synthesisabstractHigh level synthesis (HLS) tools are increasingly adopted for hardware design as the quality of tools consistently improves. Concerted development effort on HLS tools represents significant software development effort, and debugging and validation represents a significant portion of that effort. However, HLS tools are different from typical large-scale software systems; HLS tool output must be subsequently verified through functional verification of the generated RTL implementation. Debugging machine-generated functionally incorrect RTL is time-consuming and cumbersome requiring back-tracing through hundreds of signals and simulation cycles to determine the underlying error. This challenging process requires support framework in the HLS flow to enable fast and efficient pinpointing of the incorrectness in the tool. In this paper, we present a debug framework that uses just-in-time (JIT) traces and automated insertion of verification code into the generated RTL to assist in debugging an HLS tool. This framework aids the user by quickly pinpointing the earliest instance of execution mismatch, paired with detailed information on the faulty signal, and the corresponding instruction from the application source. Using CHStone benchmarks, we demonstrate that this technique can significantly reduce bug detection latency: often with zero cycle detection. Magzhan Ikram, Swathi T. Gurumani, Suhaib A. Fahmy, Deming Chen, Kyle Rupnow |
FPT | 6 |
| 2015 | Efficient GPU Spatial-Temporal MultitaskingabstractHeterogeneous computing nodes are now pervasive throughout computing, and GPUs have emerged as a leading computing device for application acceleration. GPUs have tremendous computing potential for data-parallel applications, and the emergence of GPUs has led to proliferation of GPU-accelerated applications. This proliferation has also led to systems in which many applications are competing for access to GPU resources, and efficient utilization of the GPU resources is critical to system performance. Prior techniques of temporal multitasking can be employed with GPU resources as well, but not all GPU kernels make full use of the GPU resources. There is, therefore, an unmet need for spatial multitasking in GPUs. Resources used inefficiently by one kernel can be instead assigned to another kernel that can more effectively use the resources. In this paper we propose a software-hardware solution for efficient spatial-temporal multitasking and a software based emulation framework for our system. We pair an efficient heuristic in software with hardware leaky-bucket based thread-block interleaving to implement spatial-temporal multitasking. We demonstrate our techniques on various GPU architecture using nine representative benchmarks from CUDA SDK. Our experiments on Fermi GTX480 demonstrate performance improvement by up to 46% (average 26%) over sequential GPU task execution and 37% (average 18%) over default concurrent multitasking. Compared with the state-of-the-art Kepler K20 using Hyper-Q technology, our technique achieves up to 40% (average 17%) performance improvement over default concurrent multitasking. Yun Liang 0001, Huynh Phung Huynh, Kyle Rupnow, Rick Siow Mong Goh, Deming Chen |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2014 | Integrated CUDA-to-FPGA Synthesis with Network-on-ChipabstractData parallel languages such as CUDA and Open CL efficiently describe many parallel threads of computation, and HLS tools can effectively translate these descriptions into independent optimized cores. As the number of instantiated cores grows, average external memory access latency can be a significant factor in system performance. However, although each core produces outputs independently, the cores often heavily share input data. Exploiting on-chip data sharing both reduces external bandwidth demand and improves the average memory access latency, allowing the system to improve performance at the same number of cores. In this paper, we develop a network-on-chip coupled with computation cores synthesized from CUDA for FPGAs that enables on-chip data sharing. We demonstrate reduced external bandwidth demand by up to 60% (average 56%) and total application latency in cycles by up to 43% (average 27%). Swathi T. Gurumani, Jacob Tolar, Yao Chen 0008, Yun Liang 0001, Kyle Rupnow, Deming Chen |
FCCM | 5 |
| 2014 | Fast and effective placement and routing directed high-level synthesis for FPGAsabstractAchievable frequency (fmax) is a widely used input constraint for designs targeting Field-Programmable Gate Arrays (FPGA), because of its impact on design latency and throughput. Fmax is limited by critical path delay, which is highly influenced by lower-level details of the circuit implementation such as technology mapping, placement and routing. However, for high-level synthesis~(HLS) design flows, it is challenging to evaluate the real critical delay at the behavioral level. Current HLS flows typically use module pre-characterization for delay estimates. However, we will demonstrate that such delay estimates are not sufficient to obtain high fmax and also minimize total execution latency. Hongbin Zheng, Swathi T. Gurumani, Kyle Rupnow, Deming Chen |
FPGA | 3 |
| 2014 | High-Level Synthesis With Behavioral-Level Multicycle Path AnalysisabstractHigh-level synthesis (HLS) tools generate register-transfer level (RTL) hardware descriptions from behavioral-level specifications through resource allocation, scheduling and binding. Traditionally, HLS tools build datapath pipelines by inserting pipeline registers to break combinational logic into single-cycle segments; accurately analyzing that the number of available cycles for signal propagation is proven to be infeasible at the RT-level. Thus, RT-level timing analyses must pessimistically assume each path has at most one cycle for signal propagation. This leads to false positives in critical-path analyses, prevents RTL synthesis tools from optimizing real critical paths, and forces HLS flows to insert pipeline registers without improving hardware quality. In this paper, we present an efficient behavioral-level multicycle path analysis (BL-MCPA) algorithm that leverages control-data flow information to reduce time complexity of multicycle path analysis from exponential to polynomial. BL-MCPA helps eliminate false positives in timing analysis, and improves the reported fmaxby 15% on average. With BL-MCPA, we avoid unnecessary pipeline register insertion, and reduce execution latency by 25% and register usage by 29% under a user fmaxconstraint of 300 MHz. Using BL-MCPA, we replace large multiplexers (MUXs) by pipelined MUX-trees and reduce execution latency of hardware by up to 67% on designs whose performance is limited by the large MUXs. Hongbin Zheng, Swathi T. Gurumani, Deming Chen, Kyle Rupnow |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2013 | High-level synthesis of multiple dependent CUDA kernels on FPGAabstractHigh-level synthesis (HLS) tools provide automatic generation of hardware at the register transfer level (RTL) from algorithm descriptions written in high-level languages, enabling faster creation of custom accelerators for FPGA architectures. Existing HLS tools support a wide variety of input languages, and assist users in design space exploration through automation and feedback on designs' performance bottlenecks. This design space exploration applies techniques such as pipelining, partitioning and resource sharing in order to improve performance, and resource utilization. However, although automated exploration can find some inherent parallelism, data-parallel input source code is still superior for exposing a greater variety of parallelism. In prior work, we demonstrated automated design space exploration of GPU multi-threaded (CUDA) language source code for efficient RTL generation. In this paper, we examine the challenges in extending this automated design space exploration to multiple dependent CUDA kernels, demonstrate a step-by-step procedure for efficiently performing multi-kernel synthesis, and demonstrate the potential of this approach through a case study of a stereo matching algorithm. This study demonstrates that HLS of multiple dependent CUDA kernels can maintain performance parity with the GPU implementation, while consuming over 16X less energy than the GPU. Based on our manual procedure, we identify the key challenges in fully automating the synthesis of multi-kernel CUDA programs. Swathi T. Gurumani, Hisham Cholakkal, Yun Liang 0001, Kyle Rupnow, Deming Chen |
ASP-DAC | 4 |
| 2013 | Register and thread structure optimization for GPUsabstractGPUs are an increasingly popular implementation platform for a variety of general purpose applications from mobile and embedded devices to high performance computing. The CUDA and OpenCL parallel programming models enable easy utilization of the GPU's resources. However, tuning GPU applications' performance is a complex and labor intensive task. Software programmers employ a variety of optimization techniques to explore tradeoffs between the thread parallelism and performance of a single thread. However, prior techniques ignore register allocation, a significant factor in single thread performance and, indirectly affects the number of simultaneously active threads. In this paper, we show that joint optimization of register allocation and thread structure has great potential to significantly improve performance. However, the design space for this joint optimization can be large; therefore, we develop performance metrics appropriate for evaluation within a compiler's inner loop and efficient design space exploration techniques that use the metrics to narrow the search space. Across a range of GPU applications, we achieve average performance speedup of 1.33X (up to 1.73X) with design space exploration 355X faster than the exhaustive search. Yun Liang 0001, Kyle Rupnow, Deming Chen |
ASP-DAC | 3 |
| 2013 | Improving high level synthesis optimization opportunity through polyhedral transformationsabstractHigh level synthesis (HLS) is an important enabling technology for the adoption of hardware accelerator technologies. It promises the performance and energy efficiency of hardware designs with a lower barrier to entry in design expertise, and shorter design time. State-of-the-art high level synthesis now includes a wide variety of powerful optimizations that implement efficient hardware. These optimizations can implement some of the most important features generally performed in manual designs including parallel hardware units, pipelining of execution both within a hardware unit and between units, and fine-grained data communication. We may generally classify the optimizations as those that optimize hardware implementation within a code block (intra-block) and those that optimize communication and pipelining between code blocks (inter-block). However, both optimizations are in practice difficult to apply. Real-world applications contain data-dependent blocks of code and communicate through complex data access patterns. Existing high level synthesis tools cannot apply these powerful optimizations unless the code is inherently compatible, severely limiting the optimization opportunity. In this paper we present an integrated framework to model and enable both intra- and inter-block optimizations. This integrated technique substantially improves the opportunity to use the powerful HLS optimizations that implement parallelism, pipelining, and fine-grained communication. Our polyhedral model-based technique systematically defines a set of data access patterns, identifies effective data access patterns, and performs the loop transformations to enable the intra- and inter-block optimizations. Our framework automatically explores transformation options, performs code transformations, and inserts the appropriate HLS directives to implement the HLS optimizations. Furthermore, our framework can automatically generate the optimized communication blocks for fine-grained communication between hardware blocks. Experimental evaluation demonstrates that we can achieve an average of 6.04X speedup over the high level synthesis solution without our transformations to enable intra- and inter-block optimizations. Wei Zuo, Yun Liang 0001, Peng Li 0031, Kyle Rupnow, Deming Chen, Jason Cong |
FPGA | 4 |
| 2013 | High-level synthesis with behavioral level multi-cycle path analysisabstractHigh-level synthesis (HLS) tools generate register transfer level (RTL) hardware descriptions through a process of resource allocation, scheduling and binding. Intuitively, RTL quality influences the logic synthesis quality. Specifically, the achievable clock rate, area, and latency in clock cycles will be determined by the RTL description. However, not all paths should receive equal logic synthesis effort - multi-cycle paths represent an opportunity to spend logic synthesis effort elsewhere to achieve better design quality. In this paper, we perform multi-cycle optimisation on chained functional operations. We couple HLS and logic synthesis synergistically so multi-cycle paths can be identified and optimised coherently across both behavioral and logic levels. In addition, we perform multi-cycle path analysis at the behavioral level efficiently. We prove that our technique examines all reachable circuit state and finds multi-cycle paths including control flow and guarding conditions that improve the flexibility and power of the technique. Compared to LegUp, we achieve average 55% execution time improvement, 29% area improvement, and 68% time-area product improvement targeting FPGA architecture. Hongbin Zheng, Swathi T. Gurumani, Deming Chen, Kyle Rupnow |
FPL | 5 |
| 2012 | Real-time implementation and performance optimization of 3D sound localization on GPUsabstractReal-time 3D sound localization is an important technology for various applications such as camera steering systems, robotics audition, and gunshot direction. 3D sound localization adds a new dimension, but also significantly increases the computational requirements. Real-time 3D sound localization continuously processes large volumes of data for each possible 3D direction and acoustic frequency range. Such highly demanding compute requirements outpace current CPU compute abilities. This paper develops a real-time implementation of 3D sound localization on Graphical Processing Units (GPUs). Massively parallel GPU architectures are shown to be well suited for 3D sound localization. We optimize various aspects of GPU implementation, such as number of threads per thread block, register allocation per thread, and memory data layout for performance improvement. Experiments indicate that our GPU implementation achieves 501X and 130X speedup compared to a single-thread and a multi-thread CPU implementation respectively, thus enabling real-time operation of 3D sound localization. Yun Liang 0001, Shengkui Zhao, Kyle Rupnow, Douglas L. Jones, Deming Chen |
DATE | 4 |
| 2012 | An Accurate GPU Performance Model for Effective Control Flow Divergence OptimizationabstractGraphics processing units (GPUs) are increasingly critical for general-purpose parallel processing performance. GPU hardware is composed of many streaming multiprocessors, each of which employs the single-instruction multiple-data (SIMD) execution style. This massively parallel architecture allows GPUs to execute tens of thousands of threads in parallel. Thus, GPU architectures efficiently execute heavily data-parallel applications. However, due to this SIMD execution style, resource utilization and thus overall performance can be significantly affected if computation threads must take diverging control paths. Control flow divergence in GPUs is a well-known problem: prior approaches have attempted to reduce control flow divergence through code transformations, memory access indirection, and input data reorganization. However, as we will demonstrate, the utility of these transformations is seriously affected by the lack of a guiding metric that properly estimates how control flow divergence affects application performance. In this paper, we introduce a metric that simply and accurately estimates performance of computation-bound GPU kernels with control flow divergence, and use the metric as a value function for thread re-grouping algorithms. We measure the performance on NVIDIA GTS250 GPU. For the tested set of applications, our experiments demonstrate that the proposed metric correlates well with actual GPU application performance. Through thread re-grouping guided by our metric, control flow divergence optimization can improve application performance by up to 3.19X. Yun Liang 0001, Kyle Rupnow, Deming Chen |
IPDPS | 3 |
| 2011 | High level synthesis of stereo matching: Productivity, performance, and software constraintsabstractFPGAs are an attractive platform for applications with high computation demand and low energy consumption requirements. However, design effort for FPGA implementations remains high - often an order of magnitude larger than design effort using high level languages. Instead of this time-consuming process, high level synthesis (HLS) tools generate hardware implementations from high level languages (HLL) such as C/C++/SystemC. Such tools reduce design effort: high level descriptions are more compact and less error prone. HLS tools promise hardware development abstracted from software designer knowledge of the implementation platform. In this paper, we examine several implementations of stereo matching, an active area of computer vision research that uses techniques also common for image de-noising, image retrieval, feature matching and face recognition. We present an unbiased evaluation of the suitability of using HLS for typical stereo matching software, usability and productivity of AutoPilot (a state of the art HLS tool), and the performance of designs produced by AutoPilot. Based on our study, we provide guidelines for software design, limitations of mapping general purpose software to hardware using HLS, and future directions for HLS tool development. For the stereo matching algorithms, we demonstrate between 3.5X and 67.9X speedup over software (but less than achievable by manual RTL design) with a five-fold reduction in design effort vs. manual hardware design. Kyle Rupnow, Yun Liang 0001, Dongbo Min, Minh N. Do, Deming Chen |
FPT | 1 |
| 2011 | Scientific Application Demands on a Reconfigurable Functional Unit InterfaceabstractModern scientific applications are large, complex, and highly parallel they are commonly executed on supercomputers with tens of thousands of processors. Yet these applications still commonly require weeks or even months to execute. Thus, single-thread performance remains a concern for highly parallel scientific applications. Adding a reconfigurable accelerator to each CPU could improve system performance; however, scientific applications have design constraints that differ from most application domains commonly accelerated by reconfigurable logic. In this article, we discuss the constraints imposed by scientific applications on the computation model, the accelerator architecture, and the accelerator’s communication interface with the CPU. Based on these constraints and application analysis, we have previously proposed adding a Reconfigurable Functional Unit (RFU) to accelerate integer graphs that calculate complex memory addresses. In this work, we now propose a flexible multi-instruction interface technique that allows dataflow graphs implemented on the RFU to access a large number of inputs and outputs with minor CPU datapath modifications. We present an in-depth examination of the performance effects of different communication interfaces that use this technique, and select one that best matches the needs of Sandia’s scientific applications. Although RFU execution overall improves performance, we also isolate two key negative performance effects introduced by aggregating CPU instructions into dataflow graphs: delayed issue and graph serialization. Finally, to demonstrate the marketability of an RFU beyond scientific applications, we reanalyze the proposed interfaces using the SPEC-fp benchmark suite. We show that although choosing an interface based on SPEC-fp needs is detrimental to Sandia application performance, choosing an interface based on Sandia demands works well for more general-purpose applications. Kyle Rupnow, Keith D. Underwood, Katherine Compton |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2010 | Accurately evaluating application performance in simulated hybrid multi-tasking systemsabstractEvaluating the performance of reconfigurable computing applications in multi-tasking systems using simulation (as can be needed in early design-space exploration) faces several challenges. The complexity of full-system, cycle-accurate simulation prevents executing applications of any appreciable size to completion. One must sample only a portion of execution; yet unless care is taken, the measured performance for the sampled interval will not be indicative of the complete execution. Although this is generally a problem for simulation-based evaluation, the problem is exacerbated for multi-tasking systems. This paper therefore presents work to develop a performance evaluation methodology that accurately measures hybrid (both hardware and software) application performance, accounts for additional overhead introduced by hybrid resource management (such as run-time allocation of reconfigurable hardware), and correctly compensates for momentary imbalances in processor time allocation that are only artifacts of the (necessarily) short simulated execution timespan and would balance out over time. Kyle Rupnow, Jacob Adriaens, Wenyin Fu, Katherine Compton |
FPGA | 1 |
| 2010 | A Reconfigurable Computing Scheduler Optimized for Multicore SystemsabstractThe operating system plays an important role in managing the reconfigurable hardware (RH) resources in a reconfigurable computing system to maximize system performance. These systems are becoming increasingly more complex, with multi-core CPUs and multi-tasking applications, each containing multiple kernels of computation that could be accelerated in RH. In such a system, the dynamic allocation of RH resources also becomes increasingly difficult. In this paper we examine one of the important issues a scheduler may encounter: the performance benefits of accelerating a particular kernel in an application may be enhanced by also accelerating other kernels in the application. We show that a scheduler that ignores these interdependencies will not always find the best allocation. We also propose a new RH scheduler designed specifically to handle these interdependencies, improving system performance. Philip Garcia, Kyle Rupnow, Katherine Compton |
FPL | 2 |
| 2010 | Dynamic Binding and Scheduling of Firm-Deadline Tasks on Heterogeneous Compute ResourcesabstractEmbedded systems increasingly include heterogeneous compute resources. Yet the vast majority of real-time scheduling methods are designed for single-resource or homogeneous multi-resource systems. Heterogeneity complicates scheduling; task execution time is resource-dependent. Furthermore, the best resource for one task may not necessarily be the best resource for all tasks, so one resource may not be universally more valuable than another. This paper presents new algorithms designed specifically for heterogeneous real-time scheduling. We evaluate the algorithms' deadline miss rates for heterogeneous task sets that represent a variety of execution scenarios, and show that two of our algorithms have lower deadline miss rates than the Earliest Deadline First or Least Laxity First approaches. We also discuss how task set and system characteristics affect the schedulers' abilities to achieve a quality schedule. Hsiang-Kuo Tang, Kyle Rupnow, Parameswaran Ramanathan, Katherine Compton |
RTCSA | 2 |
| 2009 | Block, Drop or Roll(back): Alternative Preemption Methods for RH Multi-TaskingabstractSave and restore of context data is traditionally used in process preemption in multi-tasking operating systems. Multi-tasking, and by consequence, preemption, is key to effective CPU sharing. However, it is much more expensive to save and restore context data in reconfigurable hardware than it is in traditional software. The configuration and current state comprises a large amount of data, making the transfer a long and expensive operation. In this paper, we explore alternatives to the save and restore operation for hardware multi-tasking. We compare the system performance of three alternate policies for reconfigurable hardware kernel preemption in a multi-process system: block, drop and roll. The best-performing policy is able to achieve on average within 4% of the performance of an idealized, zero-overhead save and restore method on a mixed application workload. Kyle Rupnow, Wenyin Fu, Katherine Compton |
FCCM | 1 |
| 2009 | Performance metrics for hybrid multi-tasking systemsabstractPerformance evaluation of hybrid (heterogeneous ISA) computing systems faces three major challenges: hybrid execution, multi-tasking, and system-level simulation variation. To evaluate system-level design decisions, a metric must encompass all forms of execution in a system, and incorporate any overheads introduced by hybrid execution. Differences in relative application speedups in a multi-tasking system complicate overall system performance evaluation. In full-system simulation, the relatively limited time-span for feasible tests compounds the evaluation problem. This paper discusses these challenges and presents metrics that address them. Kyle Rupnow, Jacob Adriaens, Wenyin Fu, Katherine Compton |
FPL | 1 |
| 2007 | Scientific Application Acceleration with Reconfigurable Functional UnitsabstractWhile scientific applications in the past were limited by floating point computations, modern scientific applications use more unstructured formulations. These applications have a significant percentage of integer computation - increasingly a limiting factor in scientific application performance. In real scientific applications employed at Sandia National Labs, integer computations constitute on average 37% of the application operations, forming large and complex dataflow graphs. Reconfigurable functional units (RFUs) are a particularly attractive accelerator for these graphs because they can potentially accelerate many unique graphs with a small amount of additional hardware. In this study, we analyze application traces of Sandia's scientific applications and the SPEC-FP benchmark suite. First we select a set of dataflow graphs to accelerate using the RFU, then we use execution-based simulation to determine the acceleration potential of the applications when using an RFU. On average, a set of 32 or fewer graphs is sufficient to capture the dataflow behavior of 30% of the integer computation, and more than half of Sandia applications show an improvement of 5% or more. Kyle Rupnow, Keith D. Underwood, Katherine Compton |
FCCM | 1 |
| 2007 | Reconfigurable Functional Units for Scientific Superscalar ProcessorsabstractAs it becomes more difficult to increase single-threaded performance, focus on multi-core processor designs increases. However, individual core performance is still important, especially for long-executing applications such as in scientific computing. Based on scientific application needs as modeled by SPEC-FP and a set of applications from Sandia National Labs, we have created several different reconfigurable functional unit (RFU) designs for superscalar multi-processor supercomputers. This paper discusses the design process and evaluates the RFUs' ability to implement instruction dataflow graphs from scientific workloads. Our best-performing RFU design is able to implement 89% of the dataflow graphs in the benchmarks. Jonathan Evans, Kyle Rupnow, Katherine Compton |
FPT | 2 |
| 2006 | Scientific applications vs. SPEC-FP: a comparison of program behaviorabstractMany modern scientific applications execute on massively parallel collections of microprocessors. Supercomputers such as the Cray XT3 (Red Storm) and Blue Gene/L support thousands to tens of thousands of processors per parallel job. However, individual microprocessor performance remains a critical component of overall performance. Traditional approaches to improve scientific application performance concentrate on floating-point (FP) instructions; however, our studies show that in the scientific applications used at Sandia National Labs, integer instructions constitute a large and critical part of the instruction mix. Although the SPEC-FP benchmark suite is considered representative of FP workloads, it has a much smaller proportion of integer computation instructions than the Sandia scientific applications, with 22.9% as compared to 36.9%. Integer instructions in Sandia applications also behave differently than in SPEC-FP. Integer instruction outputs are reused 8.8x to 13.1x more often in SPEC-FP benchmarks, and integer dataflow in Sandia applications is more complex than in the SPEC-FP suite. In this work, we examine common dataflow and usage patterns of integer instructions---information essential to develop hardware techniques to accelerate critical scientific applications. We present statistics for SPEC-FP and Sandia applications, summarizing integer computation usage and the size, shape and interface (number of inputs/outputs) of dataflow graphs. Kyle Rupnow, Arun Rodrigues, Keith D. Underwood, Katherine Compton |
ICS | 1 |