VLDB 2026 Research / reviewers in the wild / expert
Paul Chow
dblp:c/PaulChow
· DBLP profile ↗
122ranked-venue papers
8as first author
12since 2021 · last 2023
0000-0002-0523-7117ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 111 · 8 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5Software engineering, systems software and programming languages · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 3Computer networks · 2Artificial intelligence and machine learning · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A Lightweight Routing Layer Using a Reliable Link-Layer ProtocolabstractIn today’s data centers, the performance of interconnects plays a pivotal role. However, many of the underlying technologies for these interconnects have a history of several decades and existed long before data centers came into being. To better cater to the requirements of data center networks, particularly in the context of intra-rack communication, we have developed a new interconnect. This interconnect is based on a lossless link layer protocol, named RIFL. In this work, we designed and implemented RIFL Layer 2, a scalable network that supports up to multi-hundred Gbps communication. RIFL Layer 2 includes the RIFL switch and RIFL NIC. By utilizing a simple Batcher Banyan and iSLIP RIFL switch, we effectively keep the typical intra-rack latency under 400 nanoseconds. Moreover, for a 32-port 100Gbps network, under both Bernoulli arrival and bursty arrival traffic patterns, we ensure that the 99% tail latency does not exceed 12 microseconds. Qiangfeng Shen, Paul Chow |
CloudCom | 2 |
| 2023 | Partitioning Large-Scale, Multi-FPGA Applications for the Data CenterabstractWith the deployment of FPGAs in a data center, there is the opportunity to build large multi-FPGA applications. In this paper, we design a partitioner to address the problem of efficiently assigning the various tasks of a large multi-FPGA application to individual network-connected FPGAs according to constraints that consider resource usage, communication bandwidth and communication latency. By using simulated annealing, we can modify the cost function as new objectives and constraints are determined. We build on the Galapagos multi-FPGA platform by introducing a multi-die shell to extend Galapagos to more recent FPGA boards and design the partitioner to work on any collection of single- and multi-die FPGAs. Finally, We evaluate the new shell and partitioner using micro-benchmarks and analyze the partitioning of a real-world multi-FPGA application, a Transformer model. Mohammadmahdi Mazraeli, Paul Chow |
FPL | 3 |
| 2022 | Parallel CRC On An FPGA At Terabit SpeedsabstractThe Cyclic Redundancy Check Algorithm (CRC) is critical for ensuring high data reliability in serial communication such as Ethernet networks, allowing for the detection of corrupted packets with a programmable and arbitrarily small probability of failure. The baseline algorithm, however, is highly serialized due to read after write (RAW) dependencies, preventing efficient parallelization of the algorithm for use in hardware. We built a fully parameterizable open-source IP core that has no such dependencies to produce the equivalent result as the baseline CRC algorithm but in a form that can be fully parallelized, with fully automated pipelining, which works for any CRC polynomial, and with a low-resource end-of-packet alignment. This allows for up to 64-bit CRC to be computed in an FPGA at 4 Tbps. Qianfeng Shen, Juan Camilo Vega, Paul Chow |
FPT | 3 |
| 2022 | The Future of FPGA Acceleration in Datacenters and the CloudabstractIn this article, we survey existing academic and commercial efforts to provide Field-Programmable Gate Array (FPGA) acceleration in datacenters and the cloud. The goal is a critical review of existing systems and a discussion of their evolution from single workstations with PCI-attached FPGAs in the early days of reconfigurable computing to the integration of FPGA farms in large-scale computing infrastructures. From the lessons learned, we discuss the future of FPGAs in datacenters and the cloud and assess the challenges likely to be encountered along the way. The article explores current architectures and discusses scalability and abstractions supported by operating systems, middleware, and virtualization. Hardware and software security becomes critical when infrastructure is shared among tenants with disparate backgrounds. We review the vulnerabilities of current systems and possible attack scenarios and discuss mitigation strategies, some of which impact FPGA architecture and technology. The viability of these architectures for popular applications is reviewed, with a particular focus on deep learning and scientific computing. This work draws from workshop discussions, panel sessions including the participation of experts in the reconfigurable computing field, and private discussions among these experts. These interactions have harmonized the terminology, taxonomy, and the important topics covered in this manuscript. Christophe Bobda, Joel Mandebi, Paul Chow, Mohammad Ewais, Naif Tarafdar, Juan Camilo Vega, Kenneth Eguro, Dirk Koch, Suranga Handagala, Miriam Leeser, Martin C. Herbordt, Hafsah Shahzad, H. Peter Hofstee, Burkhard Ringlein, Jakub Szefer, Ahmed Sanaullah, Russell Tessier |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2022 | AIgean: An Open Framework for Deploying Machine Learning on Heterogeneous ClustersabstractAIgean , pronounced like the sea, is an open framework to build and deploy machine learning (ML) algorithms on a heterogeneous cluster of devices (CPUs and FPGAs). We leverage two open source projects: Galapagos , for multi-FPGA deployment, and hls4ml , for generating ML kernels synthesizable using Vivado HLS. AIgean provides a full end-to-end multi-FPGA/CPU implementation of a neural network. The user supplies a high-level neural network description, and our tool flow is responsible for the synthesizing of the individual layers, partitioning layers across different nodes, as well as the bridging and routing required for these layers to communicate. If the user is an expert in a particular domain and would like to tinker with the implementation details of the neural network, we define a flexible implementation stack for ML that includes the layers of Algorithms, Cluster Deployment & Communication, and Hardware. This allows the user to modify specific layers of abstraction without having to worry about components outside of their area of expertise, highlighting the modularity of AIgean . We demonstrate the effectiveness of AIgean with two use cases: an autoencoder, and ResNet-50 running across 10 and 12 FPGAs. AIgean leverages the FPGA’s strength in low-latency computing, as our implementations target batch-1 implementations. Naif Tarafdar, Giuseppe Di Guglielmo, Philip C. Harris, Jeffrey D. Krupa, Vladimir Loncar, Dylan S. Rankin, Zhenbin Wu, Qianfeng Shen, Paul Chow |
ACM Trans. Reconfigurable Technol. Syst. | 10 |
| 2021 | Pharos: a Performance Monitor for Multi-FPGA SystemsabstractField-Programmable Gate Arrays have been increasingly deployed in datacenters and cloud computing infrastructure over the recent years. The most famous example, is perhaps the Microsoft Catapult project [1] , which describes a reconfigurable fabric used to accelerate the Bing web search engine. One important point highlighted by the Catapult paper is the need for performance monitoring tools that provide visibility into the state of multi-FPGA systems. Such tools are not only useful in functional debugging, but can also prove to be valuable in identifying the performance bottlenecks of the hardware. Arzhang Rafii, Paul Chow, Welson Sun |
FCCM | 2 |
| 2021 | FFIVE: An FPGA Framework for Interactive VNF EnvironmentsabstractSummary form only given. In the world of telecommunications, there is greater focus on using Virtual Network Functions (VNFs) managed by Software Defined Networking (SDN). VNFs are tradition-ally implemented as software functions, but as technology evolves and application demands dramatically increase, the high performance and low latency of FPGAs make them more suited for use in VNF implementations. We propose FFIVE, a framework for the creation of FPGA-based VNF containers that can be deployed and man-aged in the same way as software-based VNF containers, but with improved bandwidth, efficiency, and latency. Our framework offers an approach for the virtualization of FPGA devices, the deployment of FPGA-based Virtual Network Functions (VNFs), and configuring the VNFs. Juan Camilo Vega, Mohammad Ewais, Alberto Leon-Garcia, Paul Chow |
FCCM | 4 |
| 2021 | Interactive Debugging at IP Block Interfaces in FPGAsabstractRecent developments have shown FPGAs to be effective for data centre applications, but debugging support in that environment has not evolved correspondingly. This presents an additional barrier to widespread adoption. This work proposes Debug Governors, a new open-source debugger designed for controllability and interactive debugging that can help to locate issues across multiple FPGAs. Marco Antonio Merlini, Isamu Poy, Paul Chow |
FPGA | 3 |
| 2021 | Exploring PGAS Communication for Heterogeneous Clusters with FPGAsabstractThis work presents a heterogeneous communication library for generic clusters of processors and FPGAs. This library, Shoal, supports the Partitioned Global Address Space (PGAS) memory model for applications. PGAS is a shared memory model for clusters that creates a distinction between local and remote memory access. Through Shoal and its common application programming interface for hardware and software, applications can be more freely migrated to the optimal platform and deployed onto dynamic cluster topologies. Paul Chow |
FPGA | 2 |
| 2021 | RIFL: A Reliable Link Layer Network Protocol for FPGA-to-FPGA CommunicationabstractMore and more latency-sensitive applications are being introduced into the data center. Performance of such applications can be limited by the high latency of the network interconnect. Because the conventional network stack is designed not only for LAN, but also for WAN, it carries a great amount of redundancy that is not required in a data center network. This paper introduces the concept of a three-layer protocol stack that can replace the conventional network stack and fulfill the exact demands of data center network communications. The detailed design and implementation of the first layer of the stack, which we call RIFL, is presented. A novel low latency in-band hop-by-hop re-transmission protocol is proposed and adopted in RIFL, which guarantees lossless transmission for links whose longest wire segment is no more than 150 meters. Experimental results show that RIFL achieves 218 nanoseconds round-trip latency on 3 meter zero-hop links, at a throughput of 104.7 Gbps. RIFL is a multi-lane protocol with scalable throughput from 500 Mbps to above 200 Gbps. It is portable to most of the recent FPGAs. It can be the enabler of low latency, high throughput, flexible, scalable, and lossless data center networks. Qianfeng Shen, Paul Chow |
FPGA | 3 |
| 2021 | Pharos: a Multi-FPGA Performance MonitorabstractIn recent years, there has been a lot of focus on tools that help the development of FPGA applications. However, unlike the software world, there are not many FPGA tools that help analyze the performance of FPGA applications. Also, as the application platforms scale from one FPGA to many FPGAs, such as the well-known Microsoft Catapult platform, it is no longer feasible to rely on low-level FPGA debugging tools such as the embedded logic analyzers. We present Pharos, a lightweight, generic performance monitor for multi-FPGA systems. Pharos is capable of measuring unidirectional latency between multiple network-connected FPGAs, as well as measuring throughput and tracing events across the datacenter. Arzhang Rafii, Welson Sun, Paul Chow |
FPL | 3 |
| 2021 | FPGA Implementation of an Improved OMP for Compressive Sensing ReconstructionabstractThis article proposes an improved orthogonal matching pursuit (OMP) algorithm and its implementation with Xilinx Vivado high-level synthesis (HLS). We use the Gram-Schmidt orthogonalization to improve the update process of signal residuals so that the signal recovery only needs to perform the least-squares solution once, which greatly reduces the number of matrix operations in a hardware implementation. Simulation results show that our OMP algorithm has the same signal reconstruction accuracy as the original OMP algorithm. Our approach provides a fast and reconfigurable implementation for different signal sizes, different measurement matrix sizes, and different sparsity levels. The proposed design can recover a 128-length signal with measurement number M = 32 and sparsity K = 5 and K = 8 in 13.2 and 21 μs, which is at least a 21.9% and 22.2% improvement compared with the existing HLS-based works; a 256-length signal with M = 64 and K = 8 in 20.6 μs, which is a 24% improvement compared with the existing work; and a 1024-length signal with measurement number M = 256 and sparsity K = 12 and K = 36 in 150.3 and 423 μs, respectively, which are close to the results of traditional hardware description language (HDL) implementations. Our results show that our improved OMP algorithm not only offers a superior reconstruction time compared with other recent HLS-based works but also can compete with existing works that are implemented using the traditional field-programmable gate array (FPGA) design route. Jun Li 0094, Paul Chow, Yuanxi Peng |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Hierarchical Modelling of Generators in Design-Space ExplorationabstractModern FPGA designs are often composed of multiple interacting components. In a design reuse methodology, selecting the best IP for each component as well as values for its parameters is a challenging task due to the time needed to evaluate each possible variant. Model-based approaches for searching the design space have proven to be effective for helping select the best design. However, they do not take advantage of hierarchical information available in the FPGA tool flow. In this work, we describe how to decompose a probabilistic model of a system into hierarchical components and use the model to improve the optimization speed for this design-space exploration problem. Charles Lo, Paul Chow |
FCCM | 2 |
| 2020 | AIgean: An Open Framework for Machine Learning on Heterogeneous ClustersabstractMachine learning (ML) in the past decade has been one of the most popular topics of research within the computing community. Interest within the computing field ranges across all levels of the computation stack. We show this stack in Figure 1. This work introduces an open framework, called AIgean, to build and deploy machine learning (ML) algorithms on a heterogeneous cluster of devices (CPUs and FPGAs). Users can flexibly modify any layer of the machine learning stack in Figure 1 to suit their need. This allows both machine learning domain experts to focus on higher algorithmic layers, and distributed systems experts to create the communication layers below. Naif Tarafdar, Giuseppe Di Guglielmo, Philip C. Harris, Jeffrey D. Krupa, Vladimir Loncar, Dylan S. Rankin, Zhenbin Wu, Qianfeng Shen, Paul Chow |
FCCM | 10 |
| 2020 | FFShark: A 100G FPGA Implementation of BPF Filtering for WiresharkabstractWireshark-based debugging can be performed on ordinary desktop computers at 1G speeds, but only powerful computers can keep up with 10G. At 100G, this debugging becomes virtually impossible to perform on a single machine.This work presents FFShark, a Fast FPGA implementation of Wireshark. The result is a compact, relatively inexpensive passthrough device that can be inserted into any running 100G network. Packets will travel through FFShark with no interruption and minimal additional latency. A developer can send standard Wireshark filter programs to the FFShark device at any time; packets that satisfy the filter will be copied and sent back to the developer’s workstation over a separate connection.We show that our open source passthrough device has lower latency than commercial 100G switches, and that our design is already capable of handling 400G speeds. Juan Camilo Vega, Marco Antonio Merlini, Paul Chow |
FCCM | 3 |
| 2020 | SHIP: Storage for Hybrid Interconnected ProcessorsabstractDrivers for accessing storage are complex. In addition to the complexity involved in using the NVMe protocol, navigating filesystems requires multiple serialized storage accesses and data processing between each access. As a result, efforts to create storage drivers for FPGAs, so that FPGAs can directly access storage without help from a CPU, have either failed, require too many resources/time, or remove some of the functionality expected by storage users (such as removing the filesystem) [1], [2]. Juan Camilo Vega, Qianfeng Shen, Paul Chow |
FCCM | 3 |
| 2019 | Sonar: Writing Testbenches through PythonabstractDesign verification is an important though time-consumingaspect of hardware design. A good testbench should supportperforming functional coverage of a design by making it easy to implement tests and determine which tests are being performed. However, for complex designs, creating and main-taining effective testbenches can take increasing amounts of time away from actual design. A further complication is there may be two development flows: conventional hardware written in a hardware description language (HDL) such as Verilog orVHDL and high-level synthesis (HLS). In the HLS approach, the hardware is specified in a higher-level language (HLL) and then converted to an HDL through HLS tools. In this flow, testbenches for the design are written in the same HLLand cosimulation is used to verify the generated HDL. Due totool restrictions, cosimulation may not always work. In VivadoHLS [1] for example, the design must contain control signals to define when to start and stop the module or the initiation interval for new data must be one cycle. Without cosimulation, the user must write an HDL testbench manually in addition to a testbench in the HLL for preliminary verification. To simplify writing testbenches, we present Sonar: an open-source Python library to write cross-language testbenches. From a common source script, Sonar can generate testbenches written in SystemVerilog (SV) and C++. These files can then be imported into standard simulation tools such as ModelSim[2] or Vivado HLS and run. The use of Python makes it easy to extend Sonar with higher layers of abstraction for testbenches and integrate it with other software platforms.Sonar is available at https://github.com/UofT-HPRC/sonar. Naif Tarafdar, Paul Chow |
FCCM | 3 |
| 2019 | A Modular Heterogeneous Stack for Deploying FPGAs and CPUs in the Data CenterabstractIn this work we present a heterogeneous deployment stack, calledGalapagos, that includes the abstraction of individual nodes (FPGAsand CPUs), the communication protocols between nodes and theorchestration and connection of these nodes into clusters. The stackwe create is also highly modular, allowing users to explore a designspace in the implementation of their cluster such as different net-work protocols or communication layers. The communication layerwe have currently implemented within our hardware stack, calledHUMboldt, handles heterogeneous communication between multi-ple FPGAs and CPUs. We implementHUMboldtusing High-LevelSynthesis (HLS) to ensure functional portability of communicatingkernels, allowing us to prototype hardware kernels in software. Ourresults have shown that our modular approach to this heterogeneousdeployment stack has introduced very little area and latency over-head in the FPGAs and can still perform at line-rate, bottleneckedsolely by the network links connecting the nodes. Our results alsohighlight the scalability of our design as our performance remainslimited by the network links when the cluster size increases. Nariman Eskandari, Naif Tarafdar, Daniel Ly-Ma, Paul Chow |
FPGA | 4 |
| 2019 | The Network Management Unit (NMU): Securing Network Access for Direct-Connected FPGAsabstractReconfigurable compute devices, namely Field Programmable Gate Arrays (FPGA), have increasingly been deployed in datacenters and cloud infrastructures. Such devices have proven effective at performing certain types of compute tasks, often performing these tasks faster, with lower latency, at a higher throughput, and/or at lower power than traditional compute devices (e.g. CPUs). Some recent works have demonstrated the benefits of deploying such devices as direct-connected nodes, i.e., the FPGAs are connected directly to the datacenters' network infrastructure. In this work, we introduce the concept of the Network Management Unit (NMU), which secures the network from potentially unwarranted access from malicious or malfunctioning FPGA applications; specifically, the NMU targets network traffic originating from (or targeted towards) an application resident solely on the FPGA itself. We argue that this is a necessary feature of direct-connected reconfigurable compute devices. We present the design of several NMUs, introduce a taxonomy for describing these different designs, and analyze the trade-offs of each design. The NMUs discussed range from those that employ hairpin routing techniques to push management to the next level switch, to those that perform routing and access control directly. Daniel Rozhko, Paul Chow |
FPGA | 2 |
| 2019 | Introducing ReCPRI: A Field Re-configurable Protocol for Backhaul Communication in a Radio Access Network
Juan Camilo Vega, Qianfeng Shen, Alberto Leon-Garcia, Paul Chow |
IM | 4 |
| 2019 | Accelerating Apache Spark with FPGAsabstractSummary Apache Spark has become one of the most popular engines for big data processing. Spark provides a platform‐independent, high‐abstraction programming paradigm for large‐scale data processing by leveraging the Java framework. Though it provides software portability across various machines, Java also limits the performance of distributed environments, such as Spark. While it may be unrealistic to rewrite platforms like Spark in a faster language, a more viable approach to mitigate its poor performance is to accelerate the computations while still working within the Java‐based framework. This paper demonstrates the feasibility of incorporating Field‐Programmable Gate Array (FPGA) acceleration into Spark and presents the performance benefits and bottlenecks of our FPGA‐accelerated Spark environment using a MapReduce implementation of the k‐means clustering algorithm, to show that acceleration is possible even when using a hardware platform that is not well optimized for performance. An important feature of our approach is that the use of FPGAs is completely transparent to the user through the use of library functions, which is a common way by which users access functions provided by Spark. Power users can further develop other computations using high‐level synthesis. Ehsan Ghasemi, Paul Chow |
Concurr. Comput. Pract. Exp. | 2 |
| 2018 | A High-Level Synthesis Case Study on Light Propagation Simulation in Turbid MediaabstractIn this work, we look into the benefit of using High-Level Synthesis (HLS) in building and accelerating complex systems with floating-point operations. We present a highly-optimized Monte-Carlo (MC) simulator for light propagation in 3D voxel-based biological tissue representations using HLS. We show how to utilize HLS in creating efficient structures that help achieve the desired throughput. We use Vivado to implement the design on a Xilinx Kintex Ultrascale FPGA running at 150 MHz. With a design time of 1.5 months, experimental results show a 3x speedup against the fastest software simulator published to date. Abdul-Amir Yassine, Yasmin Afsharnejad, Omar Ragheb, Vaughn Betz, Paul Chow |
FCCM | 5 |
| 2018 | Multi-fidelity Optimization for High-Level Synthesis DirectivesabstractHigh-Level Synthesis (HLS) tools enable rapid hardware development, but design expertise and effort are necessary to tune the high-level descriptions into optimized circuits. To improve designer productivity, automated design-space exploration techniques have been proposed. However, the optimization processes sample expensive CAD flows. In this paper, we adapt multi-fidelity optimization methods to incorporate low-fidelity estimates available in the FPGA CAD flow and speed up tuning of HLS parameters. We find that multi-fidelity optimization techniques can significantly reduce optimization time compared to previous approaches. Charles Lo, Paul Chow |
FPL | 2 |
| 2017 | Enabling network function virtualization over heterogeneous resourcesabstractThe economies of scale afforded by cloud computing has been a driving force behind the rapid development and deployment of new cloud-based network applications and services. With the massive growth of IoT devices, we expect a sharp rise in the volume of traffic seen going to and coming from cloud datacenters, which will continue to grow over the next several years. Network Function Virtualization (NFV) is a recent concept which promises to grant network operators the required flexibility to quickly develop and provision new network functions and services in the cloud. As NFV is agnostic to the computing resource, we foresee scenarios where unconventional resources such as FPGAs and GPUs will be of benefit. To this end, we present an architecture based on Software-Defined Infrastructure (SDI) which offers an abstracted control and management interface over virtualized heterogeneous resources in the cloud. Through a unified set of APIs, this architecture enables both application developers and network operators to dynamically deploy and manage new services in the cloud alongside the underlying network that interconnects them, all in a fully software-defined manner. We demonstrate and evaluate an implementation of our NFV-enablement architecture using the SAVI testbed, a multi-tier and SDN-enabled cloud containing virtualized heterogeneous compute resources. Thomas Lin, Naif Tarafdar, Byungchul Park, Paul Chow, Alberto Leon-Garcia |
APNOMS | 4 |
| 2017 | Packet Matching on FPGAs Using HMC Memory: Towards One Million Rules
Daniel Rozhko, Geoffrey Elliott, Daniel Ly-Ma, Paul Chow, Hans-Arno Jacobsen |
FPGA | 4 |
| 2017 | Enabling Flexible Network FPGA Clusters in a Heterogeneous Cloud Data Center
Naif Tarafdar, Thomas Lin, Eric Fukuda, Hadi Bannazadeh, Alberto Leon-Garcia, Paul Chow |
FPGA | 6 |
| 2017 | Heterogeneous virtualized network function framework for the data centerabstractWe present a framework for creating heterogeneous virtualized network function (VNF) service chains from cloud data center resources. Traditionally, these functions are packaged in software images within a catalog of networking applications that can be loaded onto a virtual machine CPU, and can be offered to users as a service. Our framework combines the best of both software and hardware by allowing users to chain traditional software-based VNFs with hardware-based VNFs that the user provides as an IP to generate a bitstream or a pre-generated VNF as part of a library. To accomplish this, our framework first creates the hardware bitstreams and programs the FPGA VNFs, loads any software VNFs requested, and programs the network to daisy chain the VNFs together. Furthermore, this enables an incremental design flow where the user can start by implementing a chain of VNFs in software and incrementally substitute software VNFs for their hardware counterparts. Our paper investigates two case studies to show the ability to switch between hardware and software VNFs in our framework and to demonstrate the benefit of using hardware VNFs. The first study is signature matching at fixed offsets, similar to matching packet headers. In this case study, the CPU can keep up at line-rate using specialized networking drivers. The second case study involves string matching within a packet, which requires scanning through the entire frame. In this case, the CPU performance drops to approximately 20 percent of the input rate, whereas the FPGA can continue to keep up at line-rate. Naif Tarafdar, Thomas Lin, Nariman Eskandari, David Lion, Alberto Leon-Garcia, Paul Chow |
FPL | 6 |
| 2017 | FPGA-based training of convolutional neural networks with a reduced precision floating-point libraryabstractConvolutional Neural Networks (CNNs) have been shown to have high accuracy for classification tasks in numerous applications, which has resulted in their widespread adoption. However, the high accuracy of CNNs comes at the cost of high compute and bandwidth requirements for both classification and training. In this work we discuss an FPGA-based CNN training engine: FCTE, implemented using High-Level Synthesis (HLS), targeting the Xilinx Kintex Ultrascale XCKU115 device. Furthermore, we detail custom-precision floating-point (CPFP) cores for multiplication and addition implemented using HLS, which allows for reduced area utilization. We use these cores with our engine to train networks to demonstrate that an exponent width of 6 and mantissa width of 5 achieves accuracy comparable to single-precision floating-point for the MNIST and CIFAR-10 datasets. These results are achieved using round-to-zero for the CPFP multipliers and round-to-nearest for the CPFP adders, allowing for LUT savings of 32.6% for the multipliers and 21.7% for the adders when compared to half-precision floating-point, while using the same number of DSPs. Roberto DiCecco, Paul Chow |
FPT | 3 |
| 2017 | An FPGA-based processor for training convolutional neural networksabstractConvolutional neural networks (CNNs) have gained great success in various computer vision applications. However, training a CNN model is computation-intensive and time-consuming. Hence training is mainly processed on large clusters of high-performance processors like server CPUs and GPUs. In this paper, we propose an FPGA-based processor design to accelerate the training process of CNNs. We first analyze the operations in all types of CNN layers in the training process. A uniform computation engine design is proposed to efficiently carry out all kinds of operations based on the analysis. Then a scalable accelerator framework is presented that exploits the parallelism further by unrolling the loops in two levels. The proposed accelerator design is demonstrated by implementing a processor on the Xilinx ZU19EG FPGA working at 200 MHz. The evaluation results on a group of CNN models show that our processor is 5.7 to 10.7-fold faster than the software implementations on the Intel Core i5-4440 CPU(@3.10GHz). Yong Dou, Jingfei Jiang, Qiang Wang 0006, Paul Chow |
FPT | 5 |
| 2016 | Accelerating Apache Spark Big Data Analysis with FPGAsabstractSummary form only given. Apache Spark has become one of the most popular engines for big data processing. Spark provides a platform-independent, high-abstraction programming paradigm for large-scale data processing by leveraging the Java frame-work. Though it provides software portability across various machines, Java also limits the performance of distributed environments, such as Spark. While it may be unrealistic to rewrite platforms like Spark in a faster language, a more viable approach to mitigate its poor performance is to accelerate the computations while still working within the Java-based framework. This work demonstrates the feasibility of incorporating FPGA acceleration into Spark, and uses a MapReduce implementation of the k-means clustering algorithm to show that acceleration is possible even when using a hardware platform that is not well-optimized for performance. An important feature of our approach is that the use of FPGAs is completely transparent to the user through the use of library functions, which is a common way by which users access functions provided by Spark. Power users can further develop other computations using high-level synthesis. Ehsan Ghasemi, Paul Chow |
FCCM | 2 |
| 2016 | A Scalable Heterogeneous Dataflow Architecture For Big Data Analytics Using FPGAs (Abstract Only)abstractDue to rapidly expanding data size, there is increasing need for scalable, high-performance, and low-energy frameworks for large- scale data computation. We build a dataflow architecture that harnesses FPGA resources within a distributed analytics platform creating a heterogeneous data analytics framework. This approach leverages the scalability of existing distributed processing environments and provides easy access to custom hardware accelerators for large-scale data analysis. We prototype our framework within the Apache Spark analytics tool running on a CPU-FPGA heterogeneous cluster. As a specific application case study, we have chosen the MapReduce paradigm to implement a multi-purpose, scalable, and customizable RTL accelerator inside the FPGA, capable of incorporating custom High-Level Synthesis (HLS) MapReduce kernels. We demonstrate how a typical MapReduce application can be simply adapted to our distributed framework while retaining the scalability of the Spark platform. Ehsan Ghasemi, Paul Chow |
FPGA | 2 |
| 2016 | Model-based optimization of High Level Synthesis directivesabstractHigh Level Synthesis (HLS) tools improve the speed of FPGA hardware design entry compared to traditional hardware description languages by raising the level of design abstraction. Using compiler directives to guide the tool, a wide variety of hardware architectures can be obtained without modification of the original behavioural code. However, selecting an optimal application of directives from this large design space can be daunting and time-consuming for a designer since evaluating a particular setting of directives requires running the FPGA tool flow. This work considers the use of sequential model-based optimization (SMBO) methods for automatically selecting directive settings. These methods construct models of the design space to guide the optimization process and minimize the number of tool evaluations. In this paper, we evaluate the use of SMBO for selecting HLS directives and extend the method to relate multiple uses of the same directive within a design. We observe that SMBO can quickly find optimal directive settings in a space of tens of thousands of possible directive configurations and find that our proposed extension can further improve the convergence rate over the standard method. Charles Lo, Paul Chow |
FPL | 2 |
| 2016 | Caffeinated FPGAs: FPGA framework For Convolutional Neural NetworksabstractConvolutional Neural Networks (CNNs) have gained significant traction in the field of machine learning, particularly due to their high accuracy in visual recognition. Recent works have pushed the performance of GPU implementations of CNNs showing significant improvements in their classification and training times. With these improvements, many frameworks have become available for implementing CNNs on both CPUs and GPUs, with no support for FPGA implementations. In this work we present a modified version of the popular CNN framework Caffe, with FPGA support. This allows for classification using CNN models and specialized FPGA implementations with the flexibility of reprogramming the device when necessary, seamless memory transactions between host and device, simple-to-use test benches, and the ability to create pipelined layer implementations. To validate the framework, we use the Xilinx SDAccel environment to implement an FPGA-based Winograd convolution engine and show that it can be used alongside other layers running on a host processor to run several popular CNNs (AlexNet, GoogleNet, VGG A, Overfeat). The results show that our framework achieves 50 GFLOPS across 3×3 convolutions in the benchmarks. This is achieved within a practical framework, which will aid in future development of FPGA-based CNNs. Roberto DiCecco, Griffin Lacey, Jasmina Vasiljevic, Paul Chow, Graham W. Taylor, Shawki Areibi |
FPT | 4 |
| 2016 | Extracting Designs of Secure IPs Using FPGA CAD ToolsabstractIn today's competitive market, a company's success is strongly dependent on delivering sophisticated and state-of-the-art IPs prior to their competitors. To take a short cut, a company may resort to reverse engineering or pirating their competitor's IP. Vincent Mirian, Paul Chow |
ACM Great Lakes Symposium on VLSI | 2 |
| 2016 | CORDIC-Based Enhanced Systolic Array Architecture for QR DecompositionabstractMultiple input multiple output (MIMO) with orthogonal frequency division multiplexing (OFDM) systems typically use orthogonal-triangular (QR) decomposition. In this article, we present an enhanced systolic array architecture to realize QR decomposition based on the Givens rotation (GR) method for a 4 × 4 real matrix. The coordinate rotation digital computer (CORDIC) algorithm is adopted and modified to speed up and simplify the process of GR. To verify the function and evaluate the performance, the proposed architectures are validated on a Virtex 5 FPGA development platform. Compared to a commercial implementation of vectoring CORDIC, the enhanced vectoring CORDIC is presented that uses 37.7% less hardware resources, dissipates 71.6% less power, and provides a 1.8 times speedup while maintaining the same computation accuracy. The enhanced QR systolic array architecture based on the enhanced vectoring CORDIC saves 24.5% in power dissipation, provides a factor of 1.5-fold improvement in throughput, and the hardware efficiency is improved 1.45-fold with no accuracy penalty when compared to our previously proposed QR systolic array architecture. Paul Chow, Hengzhu Liu |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2015 | Expanding OpenFlow Capabilities with Virtualized Reconfigurable HardwareabstractWe present a novel method of using cloud-based virtualized reconfigurable hardware to enhance the functionality of OpenFlow Software-Defined Networks. OpenFlow is a capable and popular SDN implementation, but when users require new or unsupported packet-processing, software processing in the OpenFlow controller cannot provide multi-gigabit rates. Our method sees packet flows redirected through virtualized hardware with custom-designed packet-processing engines that can add new capabilities to an OpenFlow network, while retaining line-rate processing. A case study shows this can be achieved with virtually no loss in throughput and minimal latency overheads. Stuart Byma, Naif Tarafdar, Talia Xu, Hadi Bannazadeh, Alberto Leon-Garcia, Paul Chow |
FPGA | 6 |
| 2015 | Exploring pipe implementations using an OpenCL framework for FPGAsabstractIn the last decade, OpenCL has sparked the interest of the computing world as it is a language based on an open standard that can run on many different heterogeneous platforms. This standard is continuously evolving to adapt to various use cases of different platforms. For example, with requests from the FPGA community, the pipe construct was added to the standard to facilitate the implementation of streaming applications. The versatility of the pipe construct allows several different usage modes. In this paper, we explore various implementations of pipes to evaluate the pipe construct. Our results show that for the FIFO mode, an implementation using an off-the-shelf component is favourable due to its performance and resource usage. However for the remaining modes, our proposed pipe implementation performs significantly better than other implementations at a resource utilization cost that is insignificant when compared to the abundance of resources in modern FPGAs. Vincent Mirian, Paul Chow |
FPT | 2 |
| 2015 | OpenCL library of stream memory components targeting FPGAsabstractIn recent years, high-level languages and compilers, such as OpenCL have improved both productivity and FPGA adoption on a wider scale. One of the challenges in the design of high-performance stream FPGA applications is iterative manual optimization of the numerous application buffers (e.g., arrays, FIFOs and scratch-pads). First, to achieve the desired throughput, the programmer faces the burden of analyzing the memory accesses of each application buffer, and based on observed data locality determines the optimal on-chip buffering, and off-chip read/write data access strategy. Second, to minimize throughput bottlenecks, the programmer has to carefully partition the limited on-chip memory resources among many application buffers. In this work we present an FPGA OpenCL library of pre-optimized stream memory components (SMCs). The library contains three types of SMCs, which implement frequently applied data transformations: 1) stencil, 2) transpose and 3) tiling. The library generates SMCs that are optimized both for the specific data transformation they perform as well as the user specified data set size. Further, to ease the partitioning of on-chip memory resources among many application memories, the library automatically maps application buffers to on-chip and off-chip memory resources. This is achieved by enabling the programmer to specify an on-chip memory budget for each component. In terms of on-chip memory, the SMCs perform data buffering to exploit data locality and maximize reuse. In terms of off-chip memory accesses, the SMCs optimize read/write memory operations by performing data coalescing, bursting and prefetching. We show that using the SMC library, the programmer can quickly generate scalable, pre-optimized stream application memory components, thus reaching throughput targets without time consuming manual memory optimization. Jasmina Vasiljevic, Ralph Wittig, Paul Schumacher, Jeff Fifield, Fernando Martinez-Vallina, Henry Styles, Paul Chow |
FPT | 7 |
| 2015 | FPGA implementation of low-power and high-PSNR DCT/IDCT architecture based on adaptive recoding CORDICabstractThe discrete cosine transform (DCT) and its inverse (IDCT) are widely used in image and video compression standards. In this paper, we propose a novel unified architecture for DCT and IDCT based on adaptive recoding coordinate rotation digital computer (ARC). The proposed architecture requires two types of ARC rotators. In addition, an efficient adder and shifter-based scale factor approximation is used in the proposed architecture. To verify the function and evaluate the performance, the proposed architecture is validated on a Virtex 5 FPGA development platform. Under DCT-only mode, compared with the proposed architecture, a state-of-the-art DCT architecture uses 12% more hardware resources, increases the critical path delay by 7.12%, consumes 10.1% more power and decreases 4.8 dB in PSNR. Under DCT/IDCT mode, the latest unified DCT/IDCT architecture has a factor of 2.17-fold in latency, needs 74.9% more hardware resources and dissipates 52.5% more power when compared to the proposed architecture. In addition, PSNR of the proposed architecture is better by 2 dB. Paul Chow, Hengzhu Liu |
FPT | 2 |
| 2015 | An Enhanced Adaptive Recoding Rotation CORDICabstractThe Conventional Coordinate Rotation Digital Computer (CORDIC) algorithm has been widely used in many applications, particularly in Direct Digital Frequency Synthesizers (DDS) and Fast Fourier Transforms (FFT). However, CORDIC is constrained by the excessive number of iterations, angle data path, and scaling factor compensation. In this article, an enhanced adaptive recoding CORDIC (EARC) is proposed. It uses the enhanced adaptive recoding method to reduce the required iterations and adopts the trigonometric transformation scheme to scale up the rotation angles. Computing sine and cosine is used first to compare the core functionality of EARC with basic CORDIC; then a 16-bit DDS and a 1,024-point FFT based on EARC are evaluated to demonstrate the benefits of EARC in larger applications. All the proposed architectures are validated on a Virtex 5 FPGA development platform. Compared with a commercial implementation of CORDIC, EARC requires 33.3% less hardware resources, provides a twofold speedup, dissipates 70.4% less power, and improves accuracy in terms of the Bit Error Position (BEP). Compared to the state-of-the-art Hybrid CORDIC, EARC reduces latency by 11.1% and consumes 17% less power. Compared with a commercial implementation of DDS, the dissipated power of the proposed DDS is reduced by 27.2%. The proposed DDS improves Spurious-Free Dynamic Range (SFDR) by nearly 7 dBc and dissipates 21.8% less power when compared with a recently published DDS circuit. The FFT based on EARC dissipates a factor of 2.05 less power than the commercial FFT even when choosing the 100% toggle rate for the FFT based on EARC and the 12.5% toggle rate for the commercial FFT. Compared with a recently published FFT, the FFT based on EARC improves Signal-to-Noise Ratio (SNR) by 8.9 dB and consumes 7.78% less power. Paul Chow, Hengzhu Liu |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2014 | FPGAs in the Cloud: Booting Virtualized Hardware Accelerators with OpenStackabstractWe present a new approach for integrating virtualized FPGA-based hardware accelerators into commercial-scale cloud computing systems, with minimal virtualization overhead. Partially reconfigurable regions across multiple FPGAs are offered as generic cloud resources through OpenStack (opensource cloud software), thereby allowing users to “boot” custom designed or predefined network-connected hardware accelerators with the same commands they would use to boot a regular Virtual Machine. We propose a hardware and software framework to enable this virtualization. This is a first attempt at closely fitting FPGAs into existing cloud computing models, where resources are virtualized, flexible, and have the illusion of infinite scalability. Our system can set up and tear down virtual accelerators in approximately 2.6 seconds on average, much faster than regular virtual machines. The static virtualization hardware on the physical FPGAs causes only a three cycle latency increase and a one cycle pipeline stall per packet in accelerators when compared to a non-virtualized system. We present a case study analyzing the design and performance of an application-level load balancer using a fully implemented prototype of our system. Our study shows that FPGA cloud compute resources can easily outperform virtual machines, while the system's virtualization and abstraction significantly reduces design iteration time and design complexity. Stuart Byma, J. Gregory Steffan, Hadi Bannazadeh, Alberto Leon-Garcia, Paul Chow |
FCCM | 5 |
| 2014 | MPack: global memory optimization for stream applications in high-level synthesisabstractOne of the challenges in designing high-performance FPGA applications is fine-tuning the use of limited on-chip memory storage among many buffers in an application. To achieve desired performance the designer faces the burden of packaging such buffers into on-chip memories and manually optimizing the utilization of each memory and the throughput of each buffer. In addition, the application memories may not match the word width or depth of the physical on-chip memories available on the FPGA. This process is time consuming and non-trivial, particularly with a large number of buffers of various depths and bit widths. We propose a tool, MPack, which globally optimizes on-chip memory use across all buffers for stream applications. The goal is to speed up development time by providing rapid design space exploration and relieving the designer of lengthy low-level iterations. We introduce new high-level pragmas allowing the user to specify global memory requirements, such as an application's on-chip memory budget and data throughput. We allow the user to quickly generate a large number of memory solutions and explore the trade-off between memory usage and achievable throughput. To demonstrate the effectiveness of our tool, we apply the new high-level pragmas to an image processing benchmark. MPack effectively explores the design space and is able to produce a large number of memory solutions ranging from 10 to 100% in throughput, and from 12 to 100% in on-chip memory usage. Jasmina Vasiljevic, Paul Chow |
FPGA | 2 |
| 2014 | Using an OpenCL framework to evaluate interconnect implementations on FPGAsabstractField Programmable Gate Arrays (FPGAs) are an ideal platform for building systems with custom hardware accelerators, however managing these systems is still a major challenge. The OpenCL standard has become accepted as a good programming model for managing heterogeneous platforms due to its rich constructs. Although commercial OpenCL frameworks are now emerging, there is a need for an open-source OpenCL framework that facilitates the exploration of the overall system architecture and software, as well as the implementation and architectures of the custom hardware accelerators (devices). In this paper, we use an OpenCL framework to compare interconnect implementations for a simple multiprocessor accelerator. Vincent Mirian, Paul Chow |
FPL | 2 |
| 2014 | Using buffer-to-BRAM mapping approaches to trade-off throughput vs. memory useabstractOne of the challenges in designing high-performance FPGA applications is fine-tuning the use of limited on-chip memory storage among many buffers in an application. To achieve desired performance and meet the on-chip memory budget requirements, the designer faces the burden of manually assigning application buffers to physical on-chip memories. Mismatches between dimensions (bit-width and depth) of buffers and physical on-chip memories lead to underutilized memories. Memory utilization can be increased via buffer packing - grouping buffers together and implementing them as a single memory, at the expense of data throughput. However, identifying buffer groups that result in the least amount of physical memory is a combinatorial problem with a large search space. This process is time consuming and non-trivial, particularly with a large number of buffers of various depths and bit widths. Previous work [1] introduced a tool that provides high-level pragmas allowing the user to specify global memory requirements, such as an application's on-chip memory budget and data throughput. This paper extends the previous work by introducing two low-level pragmas that specify information about memory access patterns, resulting in an improved on-chip memory utilization up to 22%. Further, we develop a simulated annealing based buffer packing algorithm, which reduces the tool's run-time from over 30 mins down to 15 sec, with an improvement in performance in the generated memory solution. Finally, we demonstrate the effectiveness of our tool with four stream application benchmarks. Jasmina Vasiljevic, Paul Chow |
FPL | 2 |
| 2014 | An efficient FPGA implementation of QR decomposition using a novel systolic array architecture based on enhanced vectoring CORDICabstractMultiple input multiple output (MIMO) - Orthogonal frequency division multiplexing (OFDM) systems typically use Orthogonal-triangular (QR) decomposition. In this paper, we present a novel systolic array architecture to realize QR decomposition based on the Givens rotation method for a 4 × 4 real matrix. The coordinate rotation digital computer (CORDIC) algorithm is adopted and modified to speed up and simplify the Givens rotation. To verify the function and evaluate the performance, the proposed architectures are validated on a Virtex 5 FPGA development platform. Compared to a commercial implementation of vectoring CORDIC, an enhanced vectoring CORDIC is presented that uses 37.7% less hardware resources, dissipates 76.8% less power and provides a 1.8 times speed-up while maintaining the same computation accuracy. The novel QR systolic array architecture based on the enhanced vectoring CORDIC saves 5% in hardware and the throughput is improved by a factor of two with no accuracy penalty when compared with the best previous version of the QR systolic array. Paul Chow, Hengzhu Liu |
FPT | 2 |
| 2014 | Benefits of Adding Hardware Support for Broadcast and Reduce Operations in MPSoC ApplicationsabstractMPI has been used as a parallel programming model for supercomputers and clusters and recently in MultiProcessor Systems-on-Chip (MPSoC). One component of MPI is collective communication and its performance is key for certain parallel applications to achieve good speedups. Previous work showed that, with synthetic communication-only benchmarks, communication improvements of up to 11.4-fold and 22-fold for broadcast and reduce operations, respectively, can be achieved by providing hardware support at the network level in a Network-on-Chip (NoC). However, these numbers do not provide a good estimation of the advantage for actual applications, as there are other factors that affect performance besides communications, such as computation. To this end, we extend our previous work by evaluating the impact of hardware support over a set of five parallel application kernels of varying computation-to-communication ratios. By introducing some useful computation to the performance evaluation, we obtain more representative results of the benefits of adding hardware support for broadcast and reduce operations. The experiments show that applications with lower computation-to-communication ratios benefit the most from hardware support as they highly depend on efficient collective communications to achieve better scalability. We also extend our work by doing more analysis on clock frequency, resource usage, power, and energy. The results show reasonable scalability for resource utilization and power in the network interfaces as the number of channels increases and that, even though more power is dissipated in the network interfaces due to the added hardware, the total energy used can still be less if the actual speedup is sufficient. The application kernels are executed in a 24-embedded-processor system distributed across four FPGAs. Yuanxi Peng, Manuel Saldaña, Christopher A. Madill, Xiaofeng Zou, Paul Chow |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2014 | Software/Hardware Parallel Long-Period Random Number Generation Framework Based on the WELL MethodabstractThis paper presents a hardware architecture for efficient implementation of the well equidistributed long-period linear (WELL) algorithm. Our design achieves a throughput of one sample-per-cycle and runs as fast as 423 MHz on a Xilinx XC5VFX130T field-programmable gate array (FPGA) device. This performance is 7.1-fold faster than a dedicated software implementation. The proposed architecture is also implemented on targeting different devices for the comparison of other types of pseudorandom number generators. In addition, we design a software/hardware framework that is capable of dividing the WELL stream into an arbitrary number of independent parallel substreams. With support from software, this framework can obtain speedup roughly proportional to the number of parallel cores. The sequences produced by the single design are verified to be consistent with the standard software generator. In addition, the statistical tests of interleaved sequences are also performed to check for correlations between different substreams of the parallel framework. We apply our framework to two applications. Experimental results verify the correctness of our framework as well as the better characteristics of the WELL algorithm compared with the Mersenne Twister method. Paul Chow, Minxuan Zhang, Shaojun Wei |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | A remote memory access infrastructure for global address space programming models in FPGAsabstractWe are proposing a shared-memory communication infrastructure that provides a common parallel programming interface for FPGA and CPU components in a heterogeneous system. Our intent is to ease the integration of reconfigurable hardware into parallel programming models like Partitioned Global Address Space (PGAS). For this purpose, we introduce a remote memory access component based on Active Messages that implements the core API of the Berkeley GASNet communication library, and a simple controller that manages communication and synchronization for custom FPGA cores. We demonstrate how these components deliver a simple and easily configurable communication mechanism between distributed memories in a multi-FPGA system with processors as well as custom hardware nodes. Ruediger Willenberg, Paul Chow |
FPGA | 2 |
| 2013 | NetThreads-10G: Software packet processing on NetFPGA-10G in a virtualized networking environment demonstration abstractabstractFPGAs are often used in high speed networking and telecommunications environments, where they have been shown to be very capable of line rate forwarding and routing. However, complex processes are more easily described in high-level software. In addition, many researchers do not have backgrounds in complex hardware design. NetThreads 10G is a solution to both of these problems - a soft, multithreaded multicore network processor implemented on the NetFPGA-10G[1], and software programmable using C. NefThreads10G is a port and upgrade of the original NetThreads [2] system designed for the NetFPGA: the number of cores has been doubled, packet buffer capacity increased, and a new Ethernet packet based programming system has been implemented. NetThreads 10G has a bus-based architecture connecting four MIPS-like processors to a shared data cache and a shared packet I/O buffer (Figure 1). Each core has a private instruction cache and four independent threads executed in a round robin fashion. Sixteen hardware locks are included for protecting critical code sections. The NetFPGA-10G onboard RLDRAM provides up to 128MB of main memory. During the demonstration, a sample application is developed and compiled using the NetThreads cross compiler tool. NefThreads10G is configured on the NetFPGA10G, and the application is downloaded remotely via Ethernet packets. The application is a deep packet inspection program that can detect suspicious keywords in packet payloads and keeps a record in shared memory. The demo shows how NetThreads affords us complete programmable and stateful control over OSI Layer 2 and above. The demonstration also shows NetThreads in the context of the SAVI (Smart Applications on Virtual Infrastructure) testbed. SAVI [3] is a new approach to network and Internet infrastructure - completely virtualized and extremely flexible, it views infrastructure as "converged", where processing, compute, networking and reconfigurable resources are all part of a shared and managed pool. Having reconfigurable hardware in such a virtualized and programmable environment will open up new avenues of research in reconfigurable systems. Stuart Byma, J. Gregory Steffan, Paul Chow |
FPL | 3 |
| 2013 | Simulation-based HW/SW co-debugging for field-programmable systems-on-chipabstractWe are presenting SimXMD (Simulation-based eXperimental Microprocessor Debugger), a tool that allows developers to debug microcontroller code and custom hardware simultaneously. SimXMD connects a GNU debugger instance to a full-system simulation of an embedded FPGA system. This enables free-roaming investigation of hardware-software interactions inside the system, including reverting back to an earlier point in simulation time. A custom memory logging mechanism enables access to variables in on-chip, off-chip and cached memory. SimXMD is open source, and its modular architecture facilitates extension to other embedded processors as well as different simulators and debuggers. Ruediger Willenberg, Paul Chow |
FPL | 2 |
| 2013 | SimXMD: Simulation-based HW/SW co-debuggingabstractThe unique promise of embedded systems in FPGAs is that designers can develop and modify their own peripheral hardware with a high degree of flexibility. However, the task of verifying the hardware commonly involves writing software to interact with it. This software itself is prone to design errors. To debug a system with two untested interacting components, it is preferable if their interaction can be precisely traced. We are presenting SimXMD (Simulation-based eXperimental Microprocessor Debugger), a tool that allows developers to debug microcontroller code and custom hardware simultaneously. The concept of debugging hardware and software together is not a new one. However, we take two established tools already used by the respective developers and connect them in a transparent way. SimXMD connects a GNU debugger (GDB) instance to a full-system simulation of an embedded FPGA system in ModelSim. This enables free-roaming investigation of hardware-software interactions inside the system, including reverting back to an earlier moment in simulation time. Software can be debugged in the same way that it would be commonly done with a real implementation on an FPGA board. Ruediger Willenberg, Paul Chow |
FPL | 2 |
| 2013 | Why Put FPGAs in your CPU socket?abstractSummary form only given. Ever since FPGAs were invented, there has been great interest in using them as computing devices, and with the logic densities of today's devices, many interesting functions have been shown to have significant performance and energy benefits when implemented in FPGAs. However, when an application requires the combination of a high-performance CPU and an FPGA accelerator, the effectiveness of the FPGA is highly determined by the latency and bandwidth between the CPU, the CPU memory system and the FPGA and its memory system. Putting FPGAs into the CPU socket is one way to address this issue. This talk will present the history, the advantages and disadvantages, the challenges, architectures, programming models and applications of "insocket" accelerator systems. Paul Chow |
FPT | 1 |
| 2013 | ZCluster: A Zynq-based Hadoop clusterabstractARM-based servers are garnering increasing interest in big data processing for their low power consumption. However, they are ill-suited for compute-intensive tasks due to their poor processing capability compared to the CPUs used in a traditional server. This paper describes our early efforts to integrate the processing power of the FPGA with the ARM processor inside the Xilinx Zynq SoC. An eight-slave Zynq-based Hadoop cluster is built and a customized hardware accelerator for a standard FIR filter is implemented to demonstrate the effectiveness of hardware acceleration. The Xillybus is used for communication between the ARM processor and the FPGA fabric, achieving a bandwidth of 103MB/s. The Hadoop cluster is proved to be linearly scalable with different input sizes and numbers of slaves. Overall, the cluster achieves a 3.3-fold speedup compared to a native pure software implementation on a single ARM processor and about a 20% improvement compared to an ARM-based cluster without hardware accelerators. Zhongduo Lin, Paul Chow |
FPT | 2 |
| 2012 | OpenCL memory infrastructure for FPGAs (abstract only)abstractProgramming models assist developers in creating high performance computing systems by forming a higher level abstraction of the target platform. OpenCL has emerged as a standard programming model for heterogeneous systems and there has been recent activity combining OpenCL and FPGAs. This work introduces memory infrastructure for FPGAs and is designed for OpenCL style computation, complementing previous work. An Aggregating Memory Controller is implemented in hardware and aims to maximize bandwidth to external, large, high-latency, high-bandwidth memories by finding the minimal number of external memory burst requests from a vector of requests. A template processing array with soft-processor and hand-coded hardware elements was also designed to drive the memory controller. The Aggregating Memory Controller is described in terms of operation and future scalability and the created processing array is described as a flexible structure that can support many types of processing solutions. A hardware prototype of the memory controller and processing array was implemented on a Virtex-5 LX110T FPGA. Two micro-benchmarks were run on both the soft-processor elements and the hand-coded hardware cores to exercise the memory controller. Results for effective memory bandwidth within the system show that the high-latency can be hidden using the Aggregating Memory Controller by increasing the number of threads within the processing array. S. Alexander Chin, Paul Chow |
FPGA | 2 |
| 2012 | FCache: a system for cache coherent processing on FPGAsabstractMuch like other computing platforms in the world today, FPGAs are becoming increasingly larger and contain large amounts of reconfigurable logic. This makes FPGAs an acceptable platform for multiprocessor systems. However in today's world of FPGA computing, very limited infrastructure is available to facilitate the creation of cache coherent shared memory systems for FPGAs. This paper introduces FCache, a system for shared memory cache coherent processing on FPGAs. The paper also describes the mapping of the conventional shared bus to FPGAs using two distinct network implemented in FCache. FCache also provides flushing and multithreaded synchronization functionalities, such as locking and unlocking of a mutex variable, which is embedded in its cache component. Despite these additional functionalities, results show that FCache has little resource overhead compared to a previous more simplistic cache coherent system that was targeted for FPGAs. Vincent Mirian, Paul Chow |
FPGA | 2 |
| 2012 | K-means implementation on FPGA for high-dimensional data using triangle inequalityabstractOne of the challenges to data mining raised by technology development is that both data size and dimensionality is growing rapidly. K-means, one of the most popular clustering algorithms in data mining, suffers in computational time when used for large data sets and data with high dimensionality. In this paper, we propose a hardware architecture for K-means with triangle inequality optimization on FPGA. An optimal 8-bit square calculator for 6-LUT architectures is described to minimize the hardware cost and an approximation solution is proposed to avoid square root calculation in the original triangle inequality optimization. Our software and hardware experiments are tested with the MNIST benchmark and uniform random numbers of various size. This approximation results in 2% more distance calculations for MNIST and 5% for uniform random numbers than the original optimization. Compared to the baseline hardware system without optimization, our approach achieves up to 77% improvement in processing time with about 10% logic overhead. We demonstrate that the hardware can achieve 55-fold speed up compared to software for the 1024 MNIST. Zhongduo Lin, Charles Lo, Paul Chow |
FPL | 3 |
| 2012 | Software/hardware framework for generating parallel Gaussian random numbers based on the Monty Python methodabstractWe present a hardware architecture for efficient implementation of a Gaussian random number generator (GRNG), using the Monty Python method. To maximize the performance/complexity efficiency, an efficient word-length optimization model is proposed to find out both the optimal integer and fractional word-lengths for signals. Experimental results show that our optimized Fixed-Point design achieves a throughput of almost 1 sample-per-cycle and runs as fast as 375.9 MHz on a Xilinx XC6VLX240T FPGA device. This performance is 23.4-fold faster than a dedicated software version running on a 2.67-GHz Intel core i5 processor. It takes 1976 LUTs, 1785 Flip-Flops, 12 BRAMs and 35 DSPs, which is only about 1% of the device as well as a great reduction compared to its corresponding Floating-Point implementations. Furthermore, we develop a framework that is capable of partitioning the Gaussian distribution stream into an arbitrary number of parallel sub-streams. With support from software, this framework can obtain speedup roughly linearly with the number of parallel cores. The quality of the variables produced by our design are verified via the standard Gaussian statistical test suit, the chi-square (X2) test. Paul Chow, Minxuan Zhang, Shaojun Wei |
FPT | 2 |
| 2012 | A high-performance architecture for training Viola-Jones object detectorsabstractThe object detection algorithm developed by Viola and Jones has become very popular due to its high quality and detection speed. However, the complexity of the computation required to train a detector makes it difficult to develop and test potential improvements to this algorithm. Furthermore, improving or training new detectors in the field is problematic. In this paper, we present a flexible FPGA architecture to accelerate this training process. The proposed systolic architecture is constructed to provide high throughput and make efficient use of the available external memory bandwidth. The design is implemented on a Xilinx ML605 development platform running at 200 MHz and obtains a 14-fold speed-up over a multi-threaded OpenCV implementation running on a high-end processor. Charles Lo, Paul Chow |
FPT | 2 |
| 2012 | Managing mutex variables in a cache-coherent shared-memory system for FPGAsabstractModern FPGAs have the ability to place many processing elements on a single die that can access shared memory. In a multiprocessing system, mutex variables are often used to provide proper synchronization and access to memory locations shared by the processing elements. This paper introduces a novel technique to manage mutex variables in caches for FPGAs, and is compared to an off-the-shelf system built mostly by components from an FPGA vendor. Results show that a cache-coherent system using our proposed technique performs more barriers per second than the off-the-shelf system with comparable hardware resources while also providing coherent caches. Vincent Mirian, Paul Chow |
FPT | 2 |
| 2012 | SimXMD: Integrated debugging of C code and hardware componentsabstractIn our demonstration, we present SimXMD, a tool that enables developers to debug microcontroller code and custom hardware simultaneously. SimXMD (Simulated eXperimental Microprocessor Debugger). SimXMD connects a GNU Debugger instance to a ModelSim instance simulating an embedded FPGA system with a Xilinx Microblaze processor. We will demonstrate debugging a multiprocessor FPGA system where the processor cores are connected through custom-designed network hardware. SimXMD is Open Source, and its modular architecture facilitates extending it to other embedded processors as well as different simulators or debuggers. Ruediger Willenberg, Paul Chow |
FPT | 2 |
| 2011 | Building a multi-FPGA virtualized restricted boltzmann machine architecture using embedded MPIabstractSeveral FPGA architectures exist for accelerating Restricted Boltzmann Machines (RBMs). However, the network size for most is limited by the amount of available on-chip memory. Therefore, many FPGAs are required to implement very large networks for use in real-world applications. A virtualized design is able to time-multiplex the hardware resources and handle much larger networks but suffers a performance penalty due to the context switch. In this paper, we present a number of improvements to a virtualized FPGA architecture for RBMs. First, we take advantage of 16-bit arithmetic to pack larger networks onto a chip. Second, a custom DMA engine is designed to reduce the performance impact of the large amount of memory transactions. Finally, the architecture is scaled to multiple FPGAs to gain additional performance through coarse grain parallelism. The design effort required to implement these changes is minimized through the use of an embedded MPI framework. The architecture, tested on a Berkeley Emulation Engine 3 platform running at 100 Mhz, achieves a speed of 12.563 GCUPS on a 8192x8192 network. Charles Lo, Paul Chow |
FPGA | 2 |
| 2011 | Software/Hardware Framework for Generating Parallel Long-Period Random Numbers Using the WELL MethodabstractThe Well Equidistributed Long-period Linear (WELL) algorithm is proven to have better characteristics than the Mersenne Twister (MT), one of the most widely used long-period pseudo-random number generators (PRNGs). In this paper, we propose a hardware architecture for efficient implementation of WELL. Our design achieves a throughput of 1 sample-per-cycle and runs as fast as 449.4 MHz on a Xilinx XC6VLX240T FPGA. This performance is 7.6-fold faster than a dedicated software implementation, and is comparable to a MT hardware generator built on the same device. It takes up 633 LUTs, 537 Flip-Flops and 4 BRAMs, which is only 0.5% of the device. Furthermore, we design a software/hardware framework that is capable of dividing the WELL stream into an arbitrary number of independent parallel sub-streams. With support from software, this framework can obtain speedup roughly proportional to the number of parallel cores. The quality of the random numbers generated by our design is verified by the standard statistical test suites Diehard and TestU01. We also apply our framework to a Monte-Carlo simulation for estimating p. Experimental results verify the correctness of our framework as well as the better characteristics of the WELL algorithm. Paul Chow, Minxuan Zhang |
FPL | 2 |
| 2011 | Hardware Support for Broadcast and Reduce in MPSoCabstractMPI has been used as a parallel programming model for supercomputers and clusters but also in Multiprocessor System-on-Chip. One component of MPI is collective communication and its performance is key for parallel applications to achieve good speedups. Considerable research has been done to optimize such communication by improving the MPI library algorithms. However, these optimizations are focused on the processing nodes (end-points in a network) rather than on the network itself. In this paper, we target a Network-on-Chip (NoC) and modify it to provide hardware support for broadcast and reduce operations for the ArchES-MPI library. This library is a subset implementation of the MPI standard targeting embedded processors and hardware accelerators implemented in FPGAs. The experimental results show that for a system with 24 embedded processors, the broadcast and reduce operations improved up to 11.4-fold and 22-fold, respectively. Higher benefits are expected for larger systems at the expense of a modest increase resource utilization. Yuanxi Peng, Manuel Saldaña, Paul Chow |
FPL | 3 |
| 2011 | FPGA Acceleration of MultiFactor CDO PricingabstractThe last decade has seen a significant growth in the financial industry. The recent widespread use of Internet technology has increased the accessibility of the general population to financial data, thereby increasing the average portfolio size. This increase, compounded by the need for accurate real-time results, has led to a rising demand for faster risk simulations. Often, accurately pricing widespread instruments, such as Collateralized Debt Obligations (CDOs), can take excessively long due to their multifactor assets dependency. We present a hardware implementation for a MultiFactor Gaussian Copula (MFGC) CDO pricing algorithm. Through a detailed benchmark exploration we demonstrate how reconfigurable hardware could be used to exploit fine-grain parallelism. Our results show that our implementation mapped onto a Xilinx Virtex 5 (XC5VSX50T) FPGA is over 71 times faster than corresponding software running on a single core 3.4 GHz Intel Xeon processor. Alexander Kaganov, Asif Lakhany, Paul Chow |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2011 | Leveraging reconfigurability in the hardware/software codesign processabstractCurrent technology allows designers to implement complete embedded computing systems on a single FPGA. Using an FPGA as the implementation platform introduces greater flexibility into the design process and allows a new approach to embedded system design. Since there is no cost to reprogramming an FPGA, system performance can be measured on-chip in the runtime environment and the system's architecture can be altered based on an evaluation of the data to meet design requirements. In this article, we discuss a new hardware/software codesign methodology tailored to reconfigurable platforms and a design infrastructure created to incorporate on-chip design tools. This methodology utilizes the FPGA's reconfigurability during the design process to profile and verify system performance, thereby reducing system design time. Our current design infrastructure includes: a system specification tool, two on-chip profiling tools, and an on-chip system verification tool. Lesley Shannon, Paul Chow |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2010 | Integrating High-Level Synthesis into MPIabstractIn this paper, we investigate how easily we can port existing HPC applications that use MPI to run on HPRC systems, using three commercial high-level synthesis tools in conjunction with the ArchES-MPI software/hardware communication layer. Specifically, we examine how each tool interfaces with our existing message-passing hardware, and we present a sample application that illustrates how the interface can be used. Andrew W. H. House, Manuel Saldaña, Paul Chow |
FCCM | 3 |
| 2010 | Acceleration of an analytical approach to collateralized debt obligation pricingabstractThis paper proposes a hardware implementation for pricing Collateralized Debt Obligations using a recursive analytical method. A novel approach using FIFOs for storage is implemented for the recursive convolution resulting in significant memory savings and a 41-fold speedup over a software implementation for the hardware implementation. Dharmendra P. Gupta, Paul Chow |
FPGA | 2 |
| 2010 | High-Performance Reconfigurable Hardware Architecture for Restricted Boltzmann MachinesabstractDespite the popularity and success of neural networks in research, the number of resulting commercial or industrial applications has been limited. A primary cause for this lack of adoption is that neural networks are usually implemented as software running on general-purpose processors. Hence, a hardware implementation that can exploit the inherent parallelism in neural networks is desired. This paper investigates how the restricted Boltzmann machine (RBM), which is a popular type of neural network, can be mapped to a high-performance hardware architecture on field-programmable gate array (FPGA) platforms. The proposed modular framework is designed to reduce the time complexity of the computations through heavily customized hardware engines. A method to partition large RBMs into smaller congruent components is also presented, allowing the distribution of one RBM across multiple FPGA resources. The framework is tested on a platform of four Xilinx Virtex II-Pro XC2VP70 FPGAs running at 100 MHz through a variety of different configurations. The maximum performance was obtained by instantiating an RBM of 256 × 256 nodes distributed across four FPGAs, which resulted in a computational speed of 3.13 billion connection-updates-per-second and a speedup of 145-fold over an optimized C program running on a 2.8-GHz Intel processor. Daniel Le Ly, Paul Chow |
IEEE Trans. Neural Networks | 2 |
| 2010 | MPI as a Programming Model for High-Performance Reconfigurable ComputersabstractHigh-Performance Reconfigurable Computers (HPRCs) consist of one or more standard microprocessors tightly-coupled with one or more reconfigurable FPGAs. HPRCs have been shown to provide good speedups and good cost/performance ratios, but not necessarily ease of use, leading to a slow acceptance of this technology. HPRCs introduce new design challenges, such as the lack of portability across platforms, incompatibilities with legacy code, users reluctant to change their code base, a prolonged learning curve, and the need for a system-level Hardware/Software co-design development flow. This article presents the evolution and current work on TMD-MPI, which started as an MPI-based programming model for Multiprocessor Systems-on-Chip implemented in FPGAs, and has now evolved to include multiple X86 processors. TMD-MPI is shown to address current design challenges in HPRC usage, suggesting that the MPI standard has enough syntax and semantics to program these new types of parallel architectures. Also presented is the TMD-MPI Ecosystem , which consists of research projects and tools that are developed around TMD-MPI to further improve HPRC usability. Finally, we present preliminary communication performance measurements. Manuel Saldaña, Arun Patel, Christopher A. Madill, Daniel Nunes, Danyao Wang, Paul Chow, Ralph Wittig, Henry Styles, Andrew Putnam |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2009 | FPGA-based Monte Carlo Computation of Light Absorption for Photodynamic Cancer TherapyabstractPhotodynamic therapy (PDT) is a method of treating cancer that combines light and light-sensitive drugs to selectively destroy cancerous tumours without harming the healthy tissue. The success of PDT depends on the accurate computation of light dose distribution. Monte Carlo (MC) simulations can provide an accurate solution for light dose distribution, but have high computation time that prevents them from being used in treatment planning. To alleviate this problem, a hardware design of an MC simulation based on the gold standard software in biophotonics was implemented on a large modern FPGA. This implementation achieved a 28-fold speedup and 716-fold lower power-delay product compared to the gold standard software executed on a 3 GHz Intel Xeon 5160 processor. The accuracy of the hardware was compared to the gold standard using a realistic skin model. An experiment using 100 million photon packets yielded a light dose distribution that diverged by less than 0.1 mm. We also describe our development methodology, which employs an intermediate hardware description in SystemC prior to Verilog coding that led to significant design effort efficiency. Jason Luu, Keith Redmond, William Lo, Paul Chow, Lothar Lilge, Jonathan Rose |
FCCM | 4 |
| 2009 | A high-performance FPGA architecture for restricted boltzmann machinesabstractDespite the popularity and success of neural networks in research, the number of resulting commercial or industrial applications have been limited. A primary cause of this lack of adoption is due to the fact that neural networks are usually implemented as software running on general-purpose processors. Algorithms to implement a neural network in software are typically O(n2) problems -- as a result, neural networks are unable to provide the performance and scalability required in non-academic settings. Daniel Le Ly, Paul Chow |
FPGA | 2 |
| 2009 | A multi-FPGA architecture for stochastic Restricted Boltzmann MachinesabstractAlthough there are many neural network FPGA architectures, there is no framework for designing large, high-performance neural networks suitable for the real world. In this paper, we present two concepts to support a multi-FPGA architecture for stochastic restricted Boltzmann machines (RBM), a popular type of neural network. First, a hardware core, called the kth stage piecewise linear interpolator, is used to implement a high-precision, pipelined function generator. The interpolator increases the resolution of a look up table implementation, guaranteeing an additional bit of precision for every pipeline stage. This function generator is used to implement a sigmoid function required in stochastic node selection. Next, a partitioning algorithm is used to efficiently divide a RBM amongst multiple FPGAs. The partitioning algorithm optimizes performance by minimizing the inter-FPGA communication. The architecture is tested on the Berkeley Emulation Engine 2 running at 100 MHz. One board supports a RBM of 256 times 256 nodes, and results in a computational speed of 1.85 billion connection-updatesper- second and a speed-up of 85-fold over an optimized C program running on a 2.8 GHz Intel processor. Daniel Le Ly, Paul Chow |
FPL | 2 |
| 2009 | The challenges of using an embedded MPI for hardware-based processing nodesabstractThis paper presents several challenges and solutions in designing an efficient Message Passing Interface (MPI) implementation for embedded FPGA applications. Popular MPI implementations are designed for general-purpose computers which have significantly different properties and trade-offs than embedded platforms. Our work focuses on two types of interactions that are not present in typical MPI implementations. First, a number of improvements designed to accelerate software-hardware interactions are introduced, including a Direct Memory Access (DMA) engine with MPI functionality; the use of non-interrupting, non-blocking messages; and a proposed function, called MPI_Coalesce, to reduce the function call overhead from a series of sequential messages. These improvements resulted in a speed-up of 5-fold compared to an embedded software-only MPI implementation. Next, a novel dataflow message passing model is presented for hardware-hardware interactions to overcome the limitations of atomic messages, allowing hardware engines to communicate and compute simultaneously. This dataflow model provides a natural method for hardware designers to build high performance, MPI systems. Finally, two hardware cores, Tee cores and message watchdog timers, are introduced to provide a transparent method of debugging hardware MPI designs. Daniel Le Ly, Manuel Saldaña, Paul Chow |
FPT | 3 |
| 2009 | Programming the Nallatech Xeon + multi-FPGA heterogeneous platform
Paul Chow, Manuel Saldaña, Arun Patel, Christopher A. Madill |
Hot Chips Symposium | 1 |
| 2008 | Investigation of Programming Models for Emerging FPGA-Based High Performance Computing SystemsabstractThis work proposes a set of requirements for programming emerging FPGA-based high performance computing systems, and uses them to evaluate a number of existing parallel programming models. Andrew W. H. House, Paul Chow |
FCCM | 2 |
| 2008 | FPGA acceleration of Monte-Carlo based credit derivative pricingabstractIn recent years the financial world has seen an increasing demand for faster risk simulations, driven by growth in client portfolios. Traditionally many financial models employ Monte-Carlo simulation, which can take excessively long to compute in software. This paper describes a hardware implementation for Collateralized Debt Obligations (CDOs) pricing, using the One-Factor Gaussian Copula (OFGC) model. We explore the precision requirements and the resulting resource utilization for each number representation. Our results show that our hardware implementationmapped onto a Xilinx XC5VSX50T is over 63 times faster than a software implementation running on a 3.4 GHz Intel Xeon processor. Alexander Kaganov, Paul Chow, Asif Lakhany |
FPL | 2 |
| 2008 | A profiler for a heterogeneous multi-core multi-FPGA systemabstractUnderstanding the behavior of an application is rarely a trivial task, due to the complexity of the system in which the application is executed, and the complexity of the application itself. The task becomes even more troublesome, if the application is being run in a parallel environment where relationships between each application execution are needed to grasp the necessary understanding of the application behavior. FPGA flexibility increases the complexity of such tasks by allowing not only changes to the application, to adapt to the hardware, but also to tailor the hardware for a specific application. To take full advantage of these systems, a tool that will help the user to understand an application is paramount. In this paper, we present a profiler for the TMD, a heterogeneous multicore multiFPGA system designed at the University of Toronto. The profiler can be configured for a specific application running on a specific hardware configuration. It allows retrieval of all communication calls and any user state defined by instrumentation of the source code. We test the profiler with two simple case studies: MPI Barrier, where we compare a sequential with a binary tree algorithm, and a heat equation solver that uses the Jacobi iterations method, where we compare blocking with non-blocking MPI calls. Daniel Nunes, Manuel Saldaña, Paul Chow |
FPT | 3 |
| 2008 | Compile-time and instruction-set methods for improving floating- to fixed-point conversion accuracyabstractThis paper proposes and evaluates compile time and instruction-set techniques for improving the accuracy of signal-processing algorithms run on fixed-point embedded processors. These techniques are proposed in the context of a profile guided floating- to fixed-point compiler-based conversion process. A novel fixed-point scaling algorithm (IRP) is introduced that exploits correlations between values in a program by applying fixed-point scaling, retaining as much precision as possible without causing overflow. This approach is extended into a more aggressive scaling algorithm (IRP-SA) by leveraging the modulo nature of 2's complement addition and subtraction to discard most significant bits that may not be redundant sign-extension bits. A complementary scaling technique (IDS) is then proposed that enables the fixed-point scaling of a variable to be parameterized, depending upon the context of its definitions and uses. Finally, a novel instruction-set enhancement—fractional multiplication with internal left shift(FMLS)—is proposed to further leverage interoperand correlations uncovered by the IRP-SA scaling algorithm. FMLS preserves a different subset of the full product's bits than traditional fractional fixed-point or integer multiplication. On average, FMLS combined with IRP-SA improves accuracy on processors with uniform bitwidth register architectures by the equivalent of 0.61 bits of additional precision for a set of signal-processing benchmarks (up to 2 bits). Even without employing FMLS, the IRP-SA scaling algorithm achieves additional accuracy over two previous fixed-point scaling algorithms by averages of 1.71 and 0.49 bits. Furthermore, as FMLS combines multiplication with a scaling shift, it reduces execution time by an average of 9.8%. An implementation of IDS, specialized to single-nested loops, is found to improve accuracy of a lattice filter benchmark by the equivalent of more than 16-bits of precision. Tor M. Aamodt, Paul Chow |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2007 | Integrating FPGAs in high-performance computing: introductionabstractNo abstract available. Paul Chow, Mike Hutton |
FPGA | 1 |
| 2007 | Design of a versatile and cost-effective hybrid floating-point/LNS arithmetic processorabstractLNS (logarithmic number system) arithmetic has the advantages of high-precision and high performance in complex function computation. However, the large hardware problem in LNS addition/subtraction computation has made the large word-length LNS arithmetic implementation impractical. In this research, we proposed a hybrid floating-point (FLP)/LNS processor that can utilize the FLP multiplication-addition-fused (MAF) unit and the FLP division unit for implementing the computation of LNS addition/subtraction. With unified representation format in FLP and LNS numbers, this hybrid processor is versatile because it can execute the FLP-to-LNS and LNS-to-FLP conversions easily, without any extra hardware cost, in addition to the FLP multiplication-addition/subtraction, FLP division, and LNS addition/subtraction instructions. It is cost-effective because the FLP hardware is shared by the LNS unit. A 32-bit hybrid FLP/LNS processor is implemented on the Xilinx Virtex II multimedia FF896 development board. From the synthesis results, the hardware of the 32-bit hybrid processor is at most three times that of a 32-bit pure FLP processor. Our proposed hybrid FLP/LNS approach has made the design of very large word-length LNS arithmetic processors become practical. Chichyang Chen, Paul Chow |
ACM Great Lakes Symposium on VLSI | 2 |
| 2007 | Optimization of data prefetch helper threads with path-expression based statistical modelingabstractThis paper investigates helper threads that improve performance by prefetching data on behalf of an application's main thread. The focus is data prefetch helper threads that lack branch instructions and which generate prefetches for one dynamic instance of a delinquent load instruction per spawned helper thread. This form of helper thread, some-times called a simple p-thread, has been studied previously by Roth et al. [29, 26] who proposed a framework for optimizing their impact. A key step in that framework is predicting the performance impact of a helper thread. In this paper we propose and evaluate a novel performance prediction technique that achieves comparable results yet requires less detailed information about dynamic program behavior. This technique extends a path expression based statistical modeling framework [2] by incorporating information about branch correlation (which we show is important) and by considering data flow information in a statistical manner. Significantly, the profile information we use is similar to that provided within current optimizing compilers. This paper also provides the first comprehensive assessment of the sources of modeling error relevant to predicting the performance impact of simple p-threads. Tor M. Aamodt, Paul Chow |
ICS | 2 |
| 2007 | Routability of Network Topologies in FPGAsabstractA fundamental difference between application-specific integrated circuits (ASICs) and field-programmable gate arrays (FPGAs) is that the wires in ASICs are designed to match the requirements of a particular design. Conversely, in an FPGA, the area is fixed and the routing resources exist whether or not they are used. In this paper, we investigate how well several common network topologies map onto a modern FPGA routing fabric. Different multiprocessor network topologies with between 8 and 64 nodes are mapped to a single large FPGA. Except for the fully-connected networks, it is observed that the difference in logic resources used and routing overhead among these topologies is insignificant for the systems tested. Fully-connected networks up to about 22 nodes are also feasible on the same FPGA although the logic and routing utilization clearly grows much faster. The conclusion is that a modern FPGA fabric is very rich in resources and capable of supporting highly interconnected topologies. For systems with a modest number of nodes implemented on current large FPGAs, it is not necessary to use the connectivity-limited topologies typically used for networks-on-chip. Rather, direct point-to-point connections between all communicating nodes can be considered. Manuel Saldaña, Lesley Shannon, Jia Shuo Yue, Sikang Bian, John Craig, Paul Chow |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2007 | SIMPPL: An Adaptable SoC Framework Using a Programmable Controller IP Interface to Facilitate Design ReuseabstractAs the complexity of designing system-on-chips increases, so does the need to abstract low-level design issues to improve designer productivity. The reuse of previously designed Intellectual Property (IP) modules is a common form of abstraction used to reduce design time. However, different applications typically use a variety of physical interfaces, communication protocols, and global system-level control for IP modules, which complicates design reuse. In this paper, we describe the SIMPPL system model and an abstraction for IP modules, called the computing element (CE), that facilitate the SoC design for both field-programmable gate array (FPGA) and application-specific integrated circuit (ASIC) platforms. The CE abstraction decouples the datapath and system-level communication from the application-specific control to promote design reuse by localizing control redesign of IP for new applications. The SIMPPL model facilitates multi-clock domain SoC designs and expedites system integration by defining the intermodule links and communication protocols Lesley Shannon, Paul Chow |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | A Scalable FPGA-based MultiprocessorabstractIt has been shown that a small number of FPGAs can significantly accelerate certain computing tasks by up to two or three orders of magnitude. However, particularly intensive large-scale computing applications, such as molecular dynamics simulations of biological systems, underscore the need for even greater speedups to address relevant length and time scales. In this work, we propose an architecture for a scalable computing machine built entirely using FPGA computing nodes. The machine enables designers to implement large-scale computing applications using a heterogeneous combination of hardware accelerators and embedded microprocessors spread across many FPGAs, all interconnected by a flexible communication network. Parallelism at multiple levels of granularity within an application can be exploited to obtain the maximum computational throughput. By focusing on applications that exhibit a high computation-to-communication ratio, we narrow the extent of this investigation to the development of a suitable communication infrastructure for our machine, as well as an appropriate programming model and design flow for implementing applications. By providing a simple, abstracted communication interface with the objective of being able to scale to thousands of FPGA nodes, the proposed architecture appears to the programmer as a unified, extensible FPGA fabric. A programming model based on the MPI message-passing standard is also presented as a means for partitioning an application into independent computing tasks that can be implemented on our architecture. Finally, we demonstrate the first use of our design flow by developing a simple molecular dynamics simulation application for the proposed machine, which runs on a small platform of development boards Arun Patel, Christopher A. Madill, Manuel Saldaña, Chris Comis, Régis Pomès, Paul Chow |
FCCM | 6 |
| 2006 | The routability of multiprocessor network topologies in FPGAsabstractA fundamental difference between ASICs and FPGAs is that wires in ASICs are designed such that they match the requirements of a particular design. Wire parameters such as length, width, layout and the number of wires can be varied to implement a desired circuit. Conversely, in an FPGA, area is fixed and routing resources exist whether or not they are used, so the goal becomes implementing a circuit within the limits of available resources. The architecture for existing routing structures in FPGAs has evolved over time to suit the requirements of large, localized digital circuits. However, FPGAs now have the capacity to host networks of such circuits, and system-level interconnection becomes a key element of the design process.Following a standard design flow and using commercial tools, we investigate how this fundamental difference in resource usage affects the mapping of various network topologies to a modern FPGA routing structure. By exploring the routability of different multiprocessor network topologies with 8, 16 and 32 nodes on a single FPGA, we show that the difference between resource utilization of a ring, star, hypercube and mesh topologies is not significant up to 32 nodes. We also show that a fully-connected network can be implemented with at least 16 nodes, but with 32 nodes it exceeds the routing resources available on the FPGA. We also derive a cost metric that helps to estimate the impact of the topology selection based on the number of nodes. Manuel Saldaña, Lesley Shannon, Paul Chow |
FPGA | 3 |
| 2006 | TMD-MPI: An MPI Implementation for Multiple Processors Across Multiple FPGAsabstractWith current FPGAs, designers can now instantiate several embedded processors, memory units, and a wide variety of IP blocks to build a single-chip, high-performance multiprocessor embedded system. Furthermore, multi-FPGA systems can be built to provide massive parallelism given an efficient programming model. In this paper, we present a lightweight subset implementation of the standard message-passing interface, MPI, that is suitable for embedded processors. It does not require an operating system and uses a small memory footprint. With our MPI implementation (TMD-MPI), we provide a programming model capable of using multiple-FPGAs that hides hardware complexities from the programmer, facilitates the development of parallel code and promotes code portability. To enable intra-FPGA and inter-FPGA communications, a simple network-on-chip is also developed using a low overhead network packet protocol. Together, TMD-MPI and the network provide a homogeneous view of a cluster of embedded processors to the programmer. Performance parameters such as link latency, link bandwidth, and synchronization cost are measured by executing a set of microbenchmarks Manuel Saldaña, Paul Chow |
FPL | 2 |
| 2006 | A System Design Methodology for Reducing System Integration Time and Facilitating Modular Design VerificationabstractThis paper provides a realistic case study of using the previously introduced SIMPPL system architectural model, which fixes the physical interface and communication protocols between processing elements (PEs) using PE-specific SIMPPL controllers. The implementation of a real-time MPEG-1 video decoder using SIMPPL provides a practical demonstration of how the complexity of system-level design issues are reduced by enabling rapid system-level integration and on-chip verification. The adaptation of the MPEG-1 PEs into the SIMPPL framework combined with the system-level integration was accomplished in 72.5 hours, which is only 4.5% of the overall system design time, instead of the more typical system integration times that can be as much as 30% of the design time Lesley Shannon, Blair Fort, Samir Parikh, Arun Patel, Manuel Saldaña, Paul Chow |
FPL | 6 |
| 2005 | Simplifying the Integration of Processing Elements in Computing Systems Using a Programmable ControllerabstractAs technology sizes decrease and die area increases, designers are creating increasingly complex computing systems using FPGAs. To reduce design time for new products, the reuse of previously designed intellectual property (IP) cores is essential. However, since no universally accepted interface standards exist for IP cores, there is often a certain amount of redesign necessary before they are incorporated into the new system. Furthermore, the core's functionality may need updating to support the requirements of the new application. This paper demonstrates how the SIMPPL system model allows designers to rapidly implement on-chip systems comprising multiple computing elements (CEs). Furthermore, using a controller-based interface to manage inter-CE transfers enables users to easily adapt the control sequence of individual CEs to suit the needs of new applications without necessitating the redesign of other elements in the system. Two systems using three different hardware modules adapted to CEs are described to illustrate the power and simplicity of the SIMPPL model. It required a total of six hours to implement both designs on-chip once the individual CEs had been designed. Lesley Shannon, Paul Chow |
FCCM | 2 |
| 2005 | Leveraging Reconfigurability in the Design ProcessabstractWe are investigating an on-chip design methodology for embedded computing systems implemented on FPGAs. The objective is to exploit the platform's reconfigurability, such that system development can occur on the final implementation platform. This is comparable to the software development process where applications are typically designed on a workstation that is representative of the final product's platform, not a simulator that models the processor. Designing on the target technology is also appealing for FPGA designs. An on-chip design methodology better leverages the main advantage of reconfigurability - the user may redesign the system while avoiding non-recurring costs, such as mask redesign costs. It also allows the user to quickly obtain real-time information about system performance. Lesley Shannon, Paul Chow |
FPL | 2 |
| 2005 | Designing an FPGA SoC Using a Standardized IP Block Interface
Lesley Shannon, Blair Fort, Samir Parikh, Arun Patel, Manuel Saldaña, Paul Chow |
FPT | 6 |
| 2004 | Reconfigurable Molecular Dynamics SimulatorabstractCurrent high-performance applications are typically implemented on large-scale general-purpose distributed or multiprocessing systems often based on commodity microprocessors. Field-Programmable Gate Arrays (FPGAs) have now reached a level of sophistication that they too could be used for such applications. In this paper we explore the feasibility of using FPGAs to implement large-scale application-specific computations by way of a case study that implements a novel molecular dynamics system. The system has been designed such that it is scalable and parallelizable. On the Transmogrifier 3 (TM3), the system performs calculations on an 8,192 particle system in 37 seconds at 26 MHz. This implementation shows that by scaling to more modern parts running at 100 MHz, a speedup of over 20 x can be achieved compared to a state-of-the-art microprocessor. This can also be achieved at less cost, using less power and taking less space than a standard microprocessor-based system, while maintaining the computational precision required. Navid Azizi, Ian Kuon, Aaron Egier, Ahmad Darabiha, Paul Chow |
FCCM | 5 |
| 2004 | FPGA-based supercomputing: an implementation for molecular dynamicsabstractCurrent high-performance supercomputing applications are typically implemented on large-scale general-purpose distributed or multiprocessing systems often based on commodity microprocessors. FPGAs have now reached a level of sophistication that they too could be used for such applications. We explore the feasibility of using FPGAs to implement large-scale application-specific computations by way of a case study that implements a novel Molecular Dynamics system. The system has been designed such that it is scalable and parallelizable. On the Transmogrifier 3, the system performs calculations on an 8,192 particle system in 37 seconds at 26MHz. This implementation shows that by scaling to more modern parts running at 100MHz and using a better architecture, a speedup of over 20x can be achieved compared to a state-of-the-art microprocessor. This can also be achieved at less cost, using less power and taking less space than a standard microprocessor-based system, while maintaining the computational precision required. Ian Kuon, Navid Azizi, Ahmad Darabiha, Aaron Egier, Paul Chow |
FPGA | 5 |
| 2004 | Using reconfigurability to achieve real-time profiling for hardware/software codesignabstractEmbedded systems combine a processor with dedicated logic to meet design specifications at a reasonable cost. The attempt to amalgamate two distinct design environments introduces many problems, one being how to partition a single design for the two platforms to achieve the best performance with the least effort. Since the latest FPGA technology allows the integration of soft or hard CPU cores with dedicated logic on a single chip, this presents new opportunities for addressing hardware/software codesign issues in the FPGA design process by utilizing the reconfigurable environment.This paper introduces SnoopP, a non-intrusive, real time, profiling tool. The user is able to obtain a clock cycle accurate profile of the real time performance of a software program running on a soft-core processor instantiated on an FPGA. SnoopP is an essential tool for hardware/software codesign on a reconfigurable platform. It allows the user to quickly obtain accurate profiling information that may greatly influence the partitioning of the design. Lesley Shannon, Paul Chow |
FPGA | 2 |
| 2004 | Maximizing system performance: using reconfigurability to monitor system communicationsabstractCommercial FPGA companies now provide tools that allow users to implement designs comprising soft-core processors and modules of dedicated logic. If a designer chooses to partition a system into multiple processors and hardware modules, tools and techniques for design analysis are necessary to understand system performance. This work introduces WOoDSTOCK, a tool that profiles system performance by adding monitors to the circuit running in real time on the chip. The user is able to generate a system specific profiler tailored to monitor the communication links between the different computing elements. This provides a macroscopic picture of system performance, which highlights the computing elements that cause bottlenecks in the design. Lesley Shannon, Paul Chow |
FPT | 2 |
| 2004 | Hardware Support for Prescient Instruction PrefetchabstractThis paper proposes and evaluates hardware mechanisms for supporting prescient instruction prefetch — an approach to improving single-threaded application performance by using helper threads to perform instruction prefetch. We demonstrate the need for enabling store-to-load communication and selective instruction execution when directly pre-executing future regions of an application that suffer I-cache misses. Two novel hardware mechanisms, safe-store and YAT-bits, are introduced that help satisfy these requirements. This paper also proposes and evaluates .nite state machine recall, a technique for limiting pre-execution to branches that are hard to predict by leveraging a counted I-prefetch mechanism. On a research Itanium®SMT processor with next line and streaming I-prefetch mechanisms that incurs latencies representative of next generation processors, prescient instruction prefetch can improve performance by an average of 10.0% to 22% on a set of SPEC 2000 benchmarks that suffer significant I-cache misses. Prescient instruction prefetch is found to be competitive against even the most aggressive research hardware instruction prefetch technique: fetch directed instruction prefetch. Tor M. Aamodt, Paul Chow, Per Hammarlund, Hong Wang 0003, John Paul Shen |
HPCA | 2 |
| 2003 | Standardizing the Performance Assessment of Reconfigurable Processor ArchitecturesabstractThis paper presents the Reconfigurable Architecture TEsting Suite, or RATES, which defines a standard for describing and using benchmarks for reconfigurable architectures. RATES is a set of functional benchmarks, is totally independent from the architecture and language, and usable on any processing platform be it general purpose or reconfigurable. It requires standard algorithms to allow comparisons amongst architectures but allows custom algorithms to highlight specific features. Lesley Shannon, Paul Chow |
FCCM | 2 |
| 2003 | A framework for modeling and optimization of prescient instruction prefetchabstractThis paper describes a framework for modeling macroscopic program behavior and applies it to optimizing prescient instruction prefetch -- novel technique that uses helper threads to improve single-threaded application performance by performing judicious and timely instruction prefetch. A helper thread is initiated when the main thread encounters a spawn point, and prefetches instructions starting at a distant target point. The target identifies a code region tending to incur I-cache misses that the main thread is likely to execute soon, even though intervening control flow may be unpredictable. The optimization of spawn-target pair selections is formulated by modeling program behavior as a Markov chain based on profile statistics. Execution paths are considered stochastic outcomes, and aspects of program behavior are summarized via path expression mappings. Mappings for computing reaching, and posteriori probability; path length mean, and variance; and expected path footprint are presented. These are used with Tarjan's fast path algorithm to efficiently estimate the benefit of spawn-target pair selections. Using this framework we propose a spawn-target pair selection algorithm for prescient instruction prefetch. This algorithm has been implemented, and evaluated for the Itanium Processor Family architecture. A limit study finds 4.8%to 17% speedups on an in-order simultaneous multithreading processor with eight contexts, over nextline and streaming I-prefetch for a set of benchmarks with high I-cache miss rates. The framework in this paper is potentially applicable to other thread speculation techniques. Tor M. Aamodt, Pedro Marcuello, Paul Chow, Antonio González 0001, Per Hammarlund, Hong Wang 0003, John Paul Shen |
SIGMETRICS | 3 |
| 2001 | The effect of reconfigurable units in superscalar processorsabstractThis paper describes OneChip, a third generation reconfigurable processor architecture that integrates a Reconfigurable Functional Unit (RFU) into a superscalar Reduced Instruction Set Computer (RISC) processor's pipeline. The architecture allows dynamic scheduling and dynamic reconfiguration. It also provides support for pre-loading configurations and for Least Recently Used (LRU) configuration management. Jorge E. Carrillo, Paul Chow |
FPGA | 2 |
| 2000 | Embedded ISA support for enhanced floating-point to fixed-point ANSI-C compilationabstractRecently tools for automating the translation of floatingpoint signal-processing applications written in ANSI C into fixed-point have been presented [34, 17, 8]. This paper introduces a novel fixed-point instruction-set operation, Fractional Multiplication with internal Left Shift (FMLS), and an associated translation algorithm—Intermediate-Result-Profiling based Shift Absorption (IRP-SA), that enhance fixedpoint rounding-noise and runtime performance. A significant feature of FMLS is that it is well suited to the latest generation of embedded processors that maintain relatively homogeneous register architectures. FMLS may improve the rounding-noise performance of fractional multiplication operations in three ways depending upon the specific fixed-point scaling properties an application exhibits. The IRP-SA algorithm enhances this by exploiting the modular nature of 2’s-complement addition which allows the discarding of most-significant-bits that are redundant due to inter-operand correlations. Rounding-noise reductions equivalent to carrying as much as 2.0 additional bits of precision throughout the computation are presented. Furthermore, by encoding a very limited set of output shift values (two left, one left, none, and one right) into the FMLS operation, speedups of up to 13 percent are observed. 1. Tor M. Aamodt, Paul Chow |
CASES | 2 |
| 1999 | DES Cracking on the Transmogrifier 2a
Ivan Hamer, Paul Chow |
CHES | 2 |
| 1999 | Memory Interfacing and Instruction Specification for Reconfigurable ProcessorsabstractAs custom computing machines evolve, it is clear that a major bottleneck is the slow interconnection architecture between the logic and memory.This paper describes the architecture of a custom computing machine that overcomes the interconnection bottleneck by closely integrating a fixed-logic processor, a reconfigurable logic array, and memory into a single chip, called OneChip-98.The OneChip-system has a seamless programming model that enables the programmer to easily specify instructions without additional complex instruction decoding hardware.As well, there is a simple scheme for mapping instructions to the corresponding programming bits.To allow the processor and the reconfigurable array to execute concurrently, the programming model utilizes a novel memory-consistency scheme implemented in the hardware.To evaluate the feasibility of the OneChip-architecture, a 32-bit MIPS-like processor and several performance enhancement applications were mapped to the Transmogrifier-2 field programmable system.For two typical applications, the 2-dimensional discrete cosine transform and the 64-tap FIR filter, we were capable of achieving a performance speedup of over 30 times that of a stand-alone state-of-the-art processor.1. Jeffrey A. Jacob, Paul Chow |
FPGA | 2 |
| 1999 | The design of an SRAM-based field-programmable gate array. I. ArchitectureabstractField-programmable gate arrays (FPGAs) are now widely used for the implementation of digital systems, and many commercial architectures are available. Although the literature and data books contain detailed descriptions of these architectures, there is very little information on how the high-level architecture was chosen, and no information on the circuit-level or physical design of the devices. This paper describes the high-level architectural design of a static-random-access memory programmable FPGA. A forthcoming Part II will address the circuit design issues through to the physical layout. The logic block and routing architecture of the FPGA was determined through experimentation with benchmark circuits and custom-built computer-aided design tools. The resulting logic block is an asymmetric tree of four-input lookup tables that are hard-wired together and a segmented routing architecture with a carefully chosen segment length distribution. Paul Chow, Soon Ong Seo, Jonathan Rose, Kevin Chung, Gerard Páez-Monzón, Immanuel Rahardja |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1999 | The design of a SRAM-based field-programmable gate array-Part II: Circuit design and layoutabstractFor Pt.I see ibid., vol.7, pp.191-7 (1999). Field-programmable gate arrays (FPGA's) are now widely used for the implementation of digital systems, and many commercial architectures are available. Although the literature and data books contain detailed descriptions of these architectures, there is very little information on how the high-level architecture was chosen and no information on the circuit-level or physical design of the devices. In Part I of this paper, we described the high-level architectural design of a static random-access memory programmable FPGA. This paper will address the circuit-design issues through to the physical layout. We address area-speed tradeoffs in the design of the logic block circuits and in the connections between the logic and the routing structure. All commercial FPGA designs are done using full-custom hand layout to obtain absolute minimum die sizes. This is both labor and time intensive. We propose a design style with a minitile that contains a portion of all the components in the logic tile, resulting in less full-custom effort. The minitile is replicated in a 4/spl times/4 array to create a macro tile. The minitile is optimized for layout density and speed, and is customized in the array by adding appropriate vias. This technique also permits easy changing of the hard-wired connections in the logic block architecture and the segmentation length distribution in the routing architecture. Paul Chow, Soon Ong Seo, Jonathan Rose, Kevin Chung, Gerard Páez-Monzón, Immanuel Rahardja |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1998 | The Transmogrifier-2: a 1 million gate rapid-prototyping systemabstractThis paper describes the Transmogrifier-2 (TM-2), a second-generation multifield programmable gate array (FPGA) rapid-prototyping system. The largest version of the system will comprise 16 boards that each contain two Altera 10K50 FPGA's, four I-Cube interconnect chips, and up to 8 Mbytes of memory. The inter-FPGA routing architecture of the TM-2 uses a novel interconnect structure, a nonuniform partial crossbar, that provides a constant delay between any two FPGA's in the system. The TM-2 architecture is modular and scalable, meaning that systems of various sizes can be constructed from copies of the same board, while maintaining routability and the constant delay feature. Other features include a system-level programmable clock that allows single-cycle access to off-chip memory, and programmable clock waveforms with edge resolution of 10 ns. The first Transmogrifier-2 boards have been manufactured and are functional. They have recently been used successfully in some simple graphics acceleration applications. David M. Lewis, David R. Galloway, Marcus van Ierssel, Jonathan Rose, Paul Chow |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 1997 | The Transmogrifier-2: A 1 Million Gate Rapid Prototyping SystemabstractThis paper describes the Transmogrifier-2, a second generation multi-FPGA system. The largest version of the system will comprise 16 boards that each contain two Altera 10K50 FPGAs, four I-cube interconnect chips, and up to 8 Mbytes of memory. The inter-FPGA routing architecture of the TM-2 uses a novel interconnect structure, a non-uniform partial crossbar, that provides a constant delay between any two FPGAs in the system. The TM-2 architecture is modular and scalable, meaning that various sized systems can be constructed from the same board, while maintaining routability and the constant delay feature. Other features include a system-level programmable clock that allows single-cycle access to off-chip memory, and programmable clock waveforms with resolution to 10ns. The first Transmogrifier-2 boards have been manufactured and are functional. They have recently been used successfully in some simple graphics acceleration applications. David M. Lewis, David R. Galloway, Marcus van Ierssel, Jonathan Rose, Paul Chow |
FPGA | 5 |
| 1997 | Memory-System Design Considerations for Dynamically-Scheduled ProcessorsabstractIn this paper, we identify performance trends and design relationships between the following components of the data memory hierarchy in a dynamically-scheduled processor: the register file, the lockup-free data cache, the stream buffers, and the interface between these components and the lower levels of the memory hierarchy. Similar performance was obtained from all systems having support for fewer than four in-flight misses, irrespective of the register-file size, the issue width of the processor, and the memory bandwidth. While providing support for more than four in-flight misses did increase system performance, the improvement was less than that obtained by increasing the number of registers. The addition of stream buffers to the investigated systems led to a significant performance increase, with the larger increases for systems having less in-flight-miss support, greater memory bandwidth, or more instruction issue capability. The performance of these systems was not significantly affected by the inclusion of traffic filters, dynamic-stride calculators, or the inclusion of the per-load non-unity stride-predictor and the incremental-prefetching techniques, which we introduce. However, the incremental prefetching technique reduces the bandwidth consumed by stream buffers by 50% without a significant impact on performance. Keith I. Farkas, Paul Chow, Norman P. Jouppi, Zvonko G. Vranesic |
ISCA | 2 |
| 1997 | The Multicluster Architecture: Reducing Cycle Time Through PartitioningabstractThe multicluster architecture that we introduce offers a decentralized, dynamically scheduled architecture, in which the register files, dispatch queue, and functional units of the architecture are distributed across multiple clusters, and each cluster is assigned a subset of the architectural registers. The motivation for the multicluster architecture is to reduce the clock cycle time, relative to a single-cluster architecture with the same number of hardware resources, by reducing the size and complexity of components on critical timing paths. Resource partitioning, however, introduces instruction-execution overhead and may reduce the number of concurrently executing instructions. To counter these two negative by-products of partitioning, we developed a static instruction scheduling algorithm. We describe this algorithm, and using trace-driven simulations of SPEC92 benchmarks, evaluate its effectiveness. This evaluation indicates that for the configurations considered the multicluster architecture may have significant performance advantages at feature sizes below 0.35 /spl mu/m, and warrants further investigation. Keith I. Farkas, Paul Chow, Norman P. Jouppi, Zvonko G. Vranesic |
MICRO | 2 |
| 1996 | Exploiting Dual Data-Memory Banks in Digital Signal ProcessorsabstractOver the past decade, digital signal processors (DSPs) have emerged as the processors of choice for implementing embedded applications in high-volume consumer products. Through their use of specialized hardware features and small chip areas, DSPs provide the high performance necessary for embedded applications at the low costs demanded by the high-volume consumer market. One feature commonly found in DSPs is the use of dual data-memory banks to double the memory system's bandwidth. When coupled with high-order data interleaving, dual memory banks provide the same bandwidth as more costly memory organizations such as a dual-ported memory. However, making effective use of dual memory banks remains difficult, especially for high-level language (HLL) DSP compilers.In this paper, we describe two algorithms --- compaction-based (CB) data partitioning and partial data duplication --- that we developed as part of our research into the effective exploitation of dual data-memory banks in HLL DSP compilers. We show that CB partitioning is an effective technique for exploiting dual data-memory banks, and that partial data duplication can augment CB partitioning in improving execution performance. Our results show that CB partitioning improves the performance of our kernel benchmarks by 13%-40% and the performance of our application benchmarks by 3%-15%. For one of the application benchmarks, partial data duplication boosts performance from 3% to 34%. Mazen A. R. Saghir, Paul Chow, Corinna G. Lee |
ASPLOS | 2 |
| 1996 | OneChip: an FPGA processor with reconfigurable logicabstractThis paper describes a processor architecture called OneChip, which combines a fixed-logic processor core with reconfigurable logic resources. Using the programmable components of this the performance of speed-critical can be improved by customizing OneChip's execution units, or flexibility can be added to the glue logic interfaces of embedded controller applications. OneChip eliminates the shortcomings of other custom compute machines by tightly integrating its reconfigurable resources into a MIPS-like processor. Speedups of close to 50 over strict software implementations on a MIPS R4400 are achievable for computing the DCT. Ralph Wittig, Paul Chow |
FCCM | 2 |
| 1996 | RACER: a reconfigurable constraint-length 14 Viterbi decoderabstractThis paper describes the architecture and implementation of a constraint-length 14 Viterbi decoder that achieves a decoding rate of 41 Kbits/s. The system uses 36 Xilinx XC4010 FPGAs with seven processor cards and a custom backplane to implement a multi-ring general cascade Viterbi decoder architecture. The paper also shows how to achieve decoding rates of 1 Mbit/s using current FPGA technology. Comparisons are made to JPL's big Viterbi decoder, which uses custom ASICs. David Yeh, Gennady Feygin, Paul Chow |
FCCM | 3 |
| 1996 | Register File Design Considerations in Dynamically Scheduled ProcessorsabstractWe have investigated the register file requirements of dynamically scheduled processors using register renaming and dispatch queues running the SPEC92 benchmarks. We looked at processors capable of issuing either four or eight instructions per cycle and found that in most cases implementing precise exceptions requires a relatively small number of additional registers compared to imprecise exceptions. Systems with aggressive non-blacking load support were able to achieve performance similar to processors with perfect memory systems at the cost of some additional registers. Given our machine assumptions, we found that the performance of a four-issue machine with a 32-entry dispatch queue tends to saturate around 80 registers. For an eight-issue machine with a 64-entry dispatch queue performance does not saturate until about 128 registers. Assuming the machine cycle time is proportional to the register file cycle time, the 8-issue machine yields only 20% higher performance than the 4-issue machine due in part to the cycle time impact of additional hardware. Keith I. Farkas, Norman P. Jouppi, Paul Chow |
HPCA | 3 |
| 1995 | A Field-Programmable Mixed-Analog-Digital ArrayabstractA novel field-programmable mixed-analog-digital array (FPMA) is proposed, which contains a field-programmable analog array, a field-programmable digital array, and a mixed-signal interface. This device is intended to be used for the rapid implementation of mixed-signal circuits. The resource and architectural requirements for this array are determined by analyzing a set of sample circuits. The mixed-signal interface is constructed from converter blocks that contain configurable A/D and D/A converters, which gives some flexibility in the specification of the interface. A 1.2 μm CMOS prototype IC has been designed to demonstrate the feasibility of FPMA technology. Paul Chow, P. Glenn Gulak |
FPGA | 1 |
| 1995 | How Useful Are Non-Blocking Loads, Stream Buffers and Speculative Execution in Multiple Issue Processors?abstractWe investigate the relative performance impact of non-blocking loads, stream buffers, and speculative execution both used individually and in conjunction with each other. We have simulated the SPEC92 benchmarks on a statically scheduled quad-issue processor model, running code from the Multiflow compiler. Non-blocking loads and stream buffers both provide a significant performance advantage, and their combination performs significantly better than either alone. For example, with a 64-byte, 2-way set associative cache with 32 cycle fetch latency, non-blocking loads reduce the run-time by 21% while stream-buffers reduce it by 26%, and the combined use of the two yields a 47% reduction. The addition of speculative execution further improves the performance of the systems that we have simulated, with or without non-blocking loads and stream buffers, by an additional 20% to 4O%. We expect that the use of all three of these techniques will be important in future generations of microprocessors.> Keith I. Farkas, Norman P. Jouppi, Paul Chow |
HPCA | 3 |
| 1994 | Architectural Advances in the VLSI Implementation of Arithmetic Coding for Binary Image CompressionabstractThis paper presents some recent advances in the architecture for the data compression technique known as arithmetic coding. The new architecture employs loop unrolling and speculative execution of the inner loop of the algorithm to achieve a significant speed-up relative to the Q-coder architecture. This approach reduces the number of iterations required to compress a block of data by a factor that is on the order of the compression ratio. While the speed-up technique has been previously discovered independently by researchers at IBM, no systematic study of the architectural trade-offs has ever been published. For the CCITT facsimile documents, the new architecture achieves a speed-up of approximately seven compared to the IBM Q-coder when four lookahead units are employed in parallel. A structure for fast input/output processing based on run length pre-coding of the data stream to accompany the new architecture is also presented.> Gennady Feygin, P. Glenn Gulak, Paul Chow |
Data Compression Conference | 3 |
| 1994 | Application-driven design of DSP architectures and compilersabstractCurrent DSP architectures are designed to enhance the execution of computationally-intensive, kernel-like loops. Their peculiar architectural features are often difficult for high-level language compilers to exploit. Moreover, their tightly-encoded instruction sets usually restrict the exploitation of instruction-level parallelism beyond a few instances. The quality of compiler-generated code is therefore poor when compared to hand-coded assembly language. We argue for an application-driven approach to designing flexible DSP architectures and effective compilers. We show that the run-time behavior and architectural characteristics of DSP kernels are different from those of DSP applications. We also show that when given a sufficiently flexible target architecture, a compiler is capable of effectively exploiting instances of instruction-level parallelism and DSP-specific architectural features. Finally, we show that a suitable DSP architecture is one that provides the functionality to support digital signal processing requirements, and the flexibility that enables a compiler to generate efficient code.> Mazen A. R. Saghir, Paul Chow, Corinna G. Lee |
ICASSP (2) | 2 |
| 1994 | Minimizing Excess Code Length and VLSI Complexity in the Multiplication Free Approximation of Arithmetic Coding
Gennady Feygin, P. Glenn Gulak, Paul Chow |
Inf. Process. Manag. | 3 |
| 1993 | Minimizing Error and VLSI Complexity in the Multiplication-Free Approximation of Arithmetic CodingabstractTwo new algorithms for performing arithmetic coding without multiplication are presented. The first algorithm, suitable for an alphabet of arbitrary size, reduces the worst-case normalized excess length to under 0.8% versus 1.911% for the previously known best method of Chevion et al. The second algorithm, suitable only for alphabets of less than twelve symbols, allows even greater reduction in the excess code length. For the important binary alphabet the worst-case excess code length is reduced to less than 0.1% versus 1.1% for the method of Chevion et al. The implementation requirements of the proposed new algorithms are discussed and shown to be similar.> Gennady Feygin, P. Glenn Gulak, Paul Chow |
Data Compression Conference | 3 |
| 1993 | A VLSI Implementation of a Cascade Viterbi Decoder with Traceback
Gennady Feygin, Paul Chow, P. Glenn Gulak, John Chappel, Grant Goodes, Oswin Hall, Ahmad Sayes, Satwant Singh, Michael B. Smith, Steve Wilton |
ISCAS | 2 |
| 1991 | Generalized cascade Viterbi decoder-a locally connected multiprocessor with linear speed-upabstractA family of multiprocessor architectures implementing the Viterbi algorithm is presented. The family of architectures is shown to be capable of achieving an increase in throughput which is directly proportional to the number of processors when the number of processors is smaller than the constraint length of the code v. The hardware utilization of nearly 100% and availability of deep pipelining inside each processor are demonstrated. An implementation proposal for a constraint length v=14 Viterbi decoder based on the proposed family of architectures is presented. The proposed thirteen-processor decoder is intended to be compatible with specifications for the Galileo space probe being developed by the Jet Propulsion Laboratory.> Gennady Feygin, P. Glenn Gulak, Paul Chow |
ICASSP | 3 |
| 1991 | A streamlined DSP microprocessor architectureabstractA microprocessor architecture for digital signal processing is described. The architecture offers almost twice the performance of the Motorola DSP56000 microprocessor, while maintaining assembly-code compatibility. Improved performance is achieved by streamlining the instruction set, using a seven-stage pipeline, and adding a second memory write stage late in the pipeline.> Michael Takefman, Paul Chow |
ICASSP | 2 |
| 1987 | Architectural Tradeoffs in the Design of MIPS-XabstractThe design of a RISC processor requires a careful analysis of the tradeoffs that can be made between hardware complexity and software. As new generations of processors are built to take advantage of more advanced technologies, new and different tradeoffs must be considered. We examine the design of a second generation VLSI RISC processor, MIPS-X. Paul Chow, Mark Horowitz |
ISCA | 1 |
| 1983 | A Pipelined Distributed Arithmetic PFFT ProcessorabstractPrevious experience in implementing the prime factor Fourier transform (PFFT) showed that it was much more difficult to do than the FF because of its complicated structure. In most FFT implementations the "butterfly" structure is the basic arithmetic element implemented. It is much simpler than the equivalent PFFT unit. Paul Chow, Zvonko G. Vranesic, Jui Lin Yen |
IEEE Trans. Computers | 1 |