VLDB 2026 Research / reviewers in the wild / expert
John Wawrzynek
dblp:w/JohnWawrzynek
· DBLP profile ↗
76ranked-venue papers
5as first author
13since 2021 · last 2025
0009-0003-1466-4553ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 56 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 4 since 2021Computer networks · 4Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DEMOTIC: A Differentiable Sampler for Multi-Level Digital CircuitsabstractEfficient sampling of satisfying formulas for circuit satisfiability (CircuitSAT), a well-known NP-complete problem, is essential in modern front-end applications for thorough testing and verification of digital circuits. Generating such samples is a hard computational problem due to the inherent complexity of digital circuits, size of the search space, and resource constraints involved in the process. Addressing these challenges has prompted the development of specialized algorithms that heavily rely on heuristics. However, these heuristic-based approaches frequently encounter scalability issues when tasked with sampling from a larger number of solutions, primarily due to their sequential nature. Different from such heuristic algorithms, we propose a novel differentiable sampler for multi-level digital circuits, called Demotic, that utilizes gradient descent (GD) to solve the CircuitSAT problem and obtain a wide range of valid and distinct solutions. Demotic leverages the circuit structure of the problem instance to learn valid solutions using GD by re-framing the CircuitSAT problem as a supervised multi-output regression task. This differentiable approach allows bit-wise operations to be performed independently on each element of a tensor, enabling parallel execution of learning operations, and accordingly, GPU-accelerated sampling with significant runtime improvements compared to state-of-the-art heuristic samplers. We demonstrate the superior runtime performance of Demotic in the sampling task across various CircuitSAT instances from the ISCAS-85 benchmark suite. Specifically, Demotic outperforms the state-of-the-art sampler by more than two orders of magnitude in most cases. Arash Ardakani, Kevin He, Qijing Huang 0001, Vighnesh M. Iyer, Suhong Moon, John Wawrzynek |
ASP-DAC | 7 |
| 2025 | High-Effort Logic Synthesis Using Randomized TransductionabstractHigh-effort logic synthesis has become an important research direction due to the increase in silicon cost and the growth of design complexity. The emphasis on security leads to complex cryptographic circuits, while the acceleration of AI/ML results in custom arithmetic blocks---all of which need to be highly optimized by EDA tools. In such applications, high-effort logic synthesis allows for an efficient exploration of larger solution spaces, leading to area and power savings beyond the capacity of traditional methods. This paper presents a novel variation of high-effort logic synthesis called transduction, which performs transformation and reduction using don't-cares to restructure the circuit. Integrating the proposed method into a stochastic optimization flow with dynamic scheduling saved 6.8% AIG nodes on average, compared to the original flow using the same runtime. An additional experiment further demonstrated the strength of the proposed method, which derived smaller AIGs than the previously synthesized minimum AIGs for 46 out of 100 benchmarks. Yukio Miyasaka, Alan Mishchenko, John Wawrzynek, Dino Ruic |
ASP-DAC | 3 |
| 2025 | High-Throughput SAT SamplingabstractIn this work, we present a novel technique for GPU-accelerated Boolean satisfiability (SAT) sampling. Unlike conventional sampling algorithms that directly operate on conjunctive normal form (CNF), our method transforms the logical constraints of SAT problems by factoring their CNF representations into simplified multilevel, multi-output Boolean functions. It then leverages gradient-based optimization to guide the search for a diverse set of valid solutions. Our method operates directly on the circuit structure of refactored SAT instances, reinterpreting the SAT problem as a supervised multi-output regression task. This differentiable technique enables independent bit-wise operations on each tensor element, allowing parallel execution of learning processes. As a result, we achieve GPU-accelerated sampling with significant runtime improvements ranging from 33.6x to 523.6x over state-of-the-art heuristic samplers. We demonstrate the superior performance of our sampling method through an extensive evaluation on 60 instances from a public domain benchmark suite utilized in previous studies. Arash Ardakani, Kevin He, Qijing Huang 0001, John Wawrzynek |
DATE | 5 |
| 2025 | Lessons from 40 Years of Reconfigurable ComputingabstractSince the introduction of the first FPGAs in the mid 1980's, reconfigurable devices have offered a promising alternative to conventional computing devices. Computing structures are organized by spatially wiring a fabric of simple logic units, as opposed to serially executing instructions, as with conventional processors. Reconfigurable devices occupy the region between conventional processors, which offer highly flexible, yet relatively inefficient execution, and application specific integrated circuits (ASICs), which offer highly optimized yet inflexible behavior. Field Programmable Gate Arrays (FPGAs) and other reconfigurable devices provide the flexibility of software processors with their ability to be customized or specialized on a per application basis, along with the power, performance, and cost efficiency approaching that of fixed function application specific integrated circuits (ASICs). It has long been understood that reconfigurable devices can provide significant advantages over conventional processors; and because of their fine-grain parallelism and flexibility, they can be adapted over a wide range of applications and scale over a variety of problem sizes. John Wawrzynek |
FPGA | 1 |
| 2025 | Chip Placement with Diffusion ModelsabstractMacro placement is a vital step in digital circuit design that defines the physical location of large collections of components, known as macros, on a 2D chip. Because key performance metrics of the chip are determined by the placement, optimizing it is crucial. Existing learning-based methods typically fall short because of their reliance on reinforcement learning (RL), which is slow and struggles to generalize, requiring online training on each new circuit. Instead, we train a diffusion model capable of placing new circuits zero-shot, using guided sampling in lieu of RL to optimize placement quality. To enable such models to train at scale, we designed a capable yet efficient architecture for the denoising model, and propose a novel algorithm to generate large synthetic datasets for pre-training. To allow zero-shot transfer to real circuits, we empirically study the design decisions of our dataset generation algorithm, and identify several key factors enabling generalization. When trained on our synthetic data, our models generate high-quality placements on unseen, realistic circuits, achieving competitive performance on placement benchmarks compared to state-of-the-art methods. Vint Lee, Leena Elzeiny, Chun Deng, Pieter Abbeel, John Wawrzynek |
ICML | 6 |
| 2024 | Late Breaking Results: Differential and Massively Parallel Sampling of SAT FormulasabstractDiverse solutions to the Boolean satisfiability (SAT) problem are essential for thorough testing and verification of software and hardware designs, ensuring reliability and applicability to real-world scenarios. We introduce a novel differentiable sampling method, called DiffSampler, which employs gradient descent (GD) to learn diverse solutions to the SAT problem. By formulating SAT as a supervised multi-output regression task and minimizing its loss function using GD, our approach enables performing the learning operations in parallel, leading to GPU-accelerated sampling and comparable run time performance w.r.t. heuristic samplers. We demonstrate that DiffSampler can generate diverse uniform-like solutions similar to conventional samplers. Arash Ardakani, Kevin He, Vighnesh M. Iyer, Suhong Moon, John Wawrzynek |
DAC | 6 |
| 2024 | Synthesis of LUT Networks for Random-Looking Dense Functions with Don't Cares - Towards Efficient FPGA Implementation of DNNabstractMany EDA applications deal with logic functions representing complex mathematical computations. Although in many cases, these functions depend on a small number of inputs, they often resemble random functions, making it hard to synthesize them using the traditional methods based on SOP minimization. This paper describes efficient synthesis and LUT mapping for this class of functions using a novel method that implements BDD-based minimization based on truth tables. The paper also investigates optimization with don't cares, when the outputs of a function are unspecified for some inputs, which is particularly useful in machine learning applications that trade accuracy for area. Compared to optimization and mapping used in academic and industrial tools, our method works faster and results in 1.5x smaller networks, while extra 20% area reduction was possible with don't cares at almost no accuracy cost. Yukio Miyasaka, Alan Mishchenko, John Wawrzynek, Nicholas J. Fraser |
FCCM | 3 |
| 2023 | Narrowing the Synthesis Gap: Academic FPGA Synthesis is Catching Up With the IndustryabstractHistorically, open-source FPGA synthesis and technology mapping tools have been considered far inferior to industry-standard tools. We show that this is no longer true. Improvements in recent years to Yosys (Verilog elaborator) and ABC (technology mapper) have resulted in substantially better performance, evident in both the reduction of area utilization and the increase in the maximum achievable clock frequency. More specifically, we describe how ABC9 — a set of feature additions to ABC — was integrated into Yosys upstream and available in the latest version. Technology mapping now has a complete view of the circuit, including support for hard blocks (e.g., carry chains) and multiple clock domains for timing-aware mapping. We demonstrate how these improvements accumulate in dramatically better synthesis results, with Yosys-ABC9 reducing the delay gap from 30% to 0% on a commercial FPGA target for the commonly used VTR benchmark, thus matching Vivado's performance in terms of maximum clock frequency. We also measured the performance on a selection of circuits from OpenCores as well as literature, comparing the results produced by Vivado, Yosys-ABC1 (existing work), and the proposed Yosys-ABC9 integration. Benjamin Lukas Cajus Barzen, Arya Reais-Parsi, Eddie Hung, Alan Mishchenko, Jonathan W. Greene, John Wawrzynek |
DATE | 7 |
| 2023 | SPADES: A Productive Design Flow for Versal Programmable LogicabstractWith the increasing growth of complexity and heterogeneity of modern FPGA fabrics, the conventional “flat” design flow relying on the standard tools, from Synthesis, Implementation, to Bitstream Generation, has become more arduous than ever. This leads to an inordinate turn-around time which severely impacts the productivity of application developers in quest of design space exploration. We propose an open-source tool flow built around a customizable overlay of Spatially Distributed Socket Engines (SPADES) to address the FPGA productivity issue. SPADES organizes the computation and communication of an application in a parallel, distributed execution of socket engines. To tackle the compilation time issue, we exploit the hardened Network-on-Chip present in Versal, a novel commercial FPGA architecture from AMD, to alleviate the inter-socket routing task, as well as implement reusable socket netlist by utilizing regular programmable fabric regions. We use three data-parallel benchmarks to demonstrate that our tool flow achieves shorter compilation time than the standard, top-down AMD Vitis flow by 7x on average (from hours down to minutes) with comparable or better performance on the AMD Versal VCK5000 data center card. Zachary Blair, Stephen Neuendorffer, John Wawrzynek |
FPL | 4 |
| 2022 | Learning A Continuous and Reconstructible Latent Space for Hardware Accelerator DesignabstractThe hardware design space is high-dimensional and discrete. Systematic and efficient exploration of this space has been a significant challenge. Central to this problem is the intractable search complexity that grows exponentially with the design choices and the discrete nature of the search space. This work investigates the feasibility of learning a meaningful low-dimensional continuous representation for hardware designs to reduce such complexity and facilitate the search process. We devise a variational autoencoder (VAE)-based design space exploration framework called VAESA, to encode the hardware design space in a compact and continuous representation. We show that black-box and gradient-based design space exploration algorithms can be applied to the latent space, and design points optimized in the latent space can be reconstructed to high-performance realistic hardware designs. Our experiments show that performing the design space search on the latent space consistently leads to the optimal design point under a fixed number of samples. In addition, the latent space can improve the sample efficiency of the original algorithm by 6.8$\times$ and can discover hardware designs that are up to 5% more efficient than the optimal design searched directly in the high-dimensional input space. Qijing Huang 0001, Charles Hong, John Wawrzynek, Mahesh Subedar, Sophia Shao |
ISPASS | 3 |
| 2021 | HAO: Hardware-aware Neural Architecture Optimization for Efficient InferenceabstractAutomatic algorithm-hardware co-design for DNN has shown great success in improving the performance of DNNs on FPGAs. However, this process remains challenging due to the intractable search space of neural network architectures and hardware accelerator implementation. Differing from existing hardware-aware neural architecture search (NAS) algorithms that rely solely on the expensive learning-based approaches, our work incorporates integer programming into the search algorithm to prune the design space. Given a set of hardware resource constraints, our integer programming formulation directly outputs the optimal accelerator configuration for mapping a DNN subgraph that minimizes latency. We use an accuracy predictor for different DNN subgraphs with different quantization schemes and generate accuracy-latency pareto frontiers. With low computational cost, our algorithm can generate quantized networks that achieve state-of-the-art accuracy and hardware performance on Xilinx Zynq (ZU3EG) FPGA for image classification on ImageNet dataset. The solution searched by our algorithm achieves 72.5% top-1 accuracy on ImageNet at framerate 50, which is 60% faster than MnasNet [37] and 135% faster than FBNet [43] with comparable accuracy. Zhen Dong 0003, Yizhao Gao 0002, Qijing Huang 0001, John Wawrzynek, Hayden Kwok-Hay So, Kurt Keutzer |
FCCM | 4 |
| 2021 | CoDeNet: Efficient Deployment of Input-Adaptive Object Detection on Embedded FPGAsabstractDeploying deep learning models on embedded systems for computer vision tasks has been challenging due to limited compute resources and strict energy budgets. The majority of existing work focuses on accelerating image classification, while other fundamental vision problems, such as object detection, have not been adequately addressed. Compared with image classification, detection problems are more sensitive to the spatial variance of objects, and therefore, require specialized convolutions to aggregate spatial information. To address this need, recent work introduces dynamic deformable convolution to augment regular convolutions. Regular convolutions process a fixed grid of pixels across all the spatial locations in an image, while dynamic deformable convolution may access arbitrary pixels in the image with the access pattern being input-dependent and varying with spatial location. These properties lead to inefficient memory accesses of inputs with existing hardware. Qijing Huang 0001, Dequan Wang, Zhen Dong 0003, Yizhao Gao 0002, Yaohui Cai, Bichen Wu, Kurt Keutzer, John Wawrzynek |
FPGA | 9 |
| 2021 | CoSA: Scheduling by Constrained Optimization for Spatial AcceleratorsabstractRecent advances in Deep Neural Networks (DNNs) have led to active development of specialized DNN accelerators, many of which feature a large number of processing elements laid out spatially, together with a multi-level memory hierarchy and flexible interconnect. While DNN accelerators can take advantage of data reuse and achieve high peak throughput, they also expose a large number of runtime parameters to the programmers who need to explicitly manage how computation is scheduled both spatially and temporally. In fact, different scheduling choices can lead to wide variations in performance and efficiency, motivating the need for a fast and efficient search strategy to navigate the vast scheduling space.To address this challenge, we present CoSA, a constrained-optimization-based approach for scheduling DNN accelerators. As opposed to existing approaches that either rely on designers’ heuristics or iterative methods to navigate the search space, CoSA expresses scheduling decisions as a constrained-optimization problem that can be deterministically solved using mathematical optimization techniques. Specifically, CoSA leverages the regularities in DNN operators and hardware to formulate the DNN scheduling space into a mixed-integer programming (MIP) problem with algorithmic and architectural constraints, which can be solved to automatically generate a highly efficient schedule in one shot. We demonstrate that CoSA-generated schedules significantly outperform state-of-the-art approaches by a geometric mean of up to 2.5× across a wide range of DNN networks while improving the time-to-solution by 90×. Qijing Huang 0001, Aravind Kalaiah, James Demmel, Grace Dinh, John Wawrzynek, Thomas Norell, Sophia Shao |
ISCA | 6 |
| 2019 | GraphSAR: a sparsity-aware processing-in-memory architecture for large-scale graph processing on ReRAMsabstractLarge-scale graph processing has drawn great attention in recent years. The emerging metal-oxide resistive random access memory (ReRAM) and ReRAM crossbars have shown huge potential in accelerating graph processing. However, the sparse feature of natural graphs hinders the performance of graph processing on ReRAMs. Previous work of graph processing on ReRAMs stored and computed edges separately, leading to high energy consumption and long latency of transferring data. In this paper, we present GraphSAR, a sparsity-aware processing-in-memory large-scale graph processing accelerator on ReRAMs. Computations over edges are performed in the memory, eliminating overheads of transferring edges. Moreover, graphs are divided considering the sparsity. Subgraphs with low densities are further divided into smaller ones to minimize the waste of memory space. According to our extensive experimental results, GraphSAR achieves 4.43x energy reduction and 1.85x speedup (8.19x lower energy-delay product, EDP) against previous graph processing architecture on ReRAMs (GraphR [1]). Guohao Dai 0001, Yu Wang 0002, Huazhong Yang, John Wawrzynek |
ASP-DAC | 5 |
| 2019 | AutoPhase: Compiler Phase-Ordering for HLS with Deep Reinforcement LearningabstractThe performance of the code generated by a compiler depends on the order in which the optimization passes are applied. In high-level synthesis, the quality of the generated circuit relates directly to the code generated by the front-end compiler. Choosing a good order-often referred to as the phase-ordering problem-is an NP-hard problem. In this paper, we evaluate a new technique to address the phase-ordering problem: deep reinforcement learning. We implement a framework in the context of the LLVM compiler to optimize the ordering for HLS programs and compare the performance of deep reinforcement learning to state-of-the-art algorithms that address the phase-ordering problem. Overall, our framework runs one to two orders of magnitude faster than these algorithms, and achieves a 16% improvement in circuit performance over the -O3 compiler flag. Qijing Huang 0001, Ameer Haj-Ali, William S. Moses, John Xiang, Ion Stoica, Krste Asanovic, John Wawrzynek |
FCCM | 7 |
| 2019 | Synetgy: Algorithm-hardware Co-design for ConvNet Accelerators on Embedded FPGAsabstractUsing FPGAs to accelerate ConvNets has attracted significant attention in recent years. However, FPGA accelerator design has not leveraged the latest progress of ConvNets. As a result, the key application characteristics such as frames-per-second (FPS) are ignored in favor of simply counting GOPs, and results on accuracy, which is critical to application success, are often not even reported. In this work, we adopt an algorithm-hardware co-design approach to develop a ConvNet accelerator called Synetgy and a novel ConvNet model called DiracDeltaNet. Both the accelerator and ConvNet are tailored to FPGA requirements. DiracDeltaNet, as the name suggests, is a ConvNet with only $1\times 1$ convolutions while spatial convolutions are replaced by more efficient shift operations. DiracDeltaNet achieves competitive accuracy on ImageNet (89.0% top-5), but with 48× fewer parameters and 65× fewer OPs than VGG16. We further quantize DiracDeltaNet's weights to 1-bit and activations to 4-bits, with less than 1% accuracy loss. These quantizations exploit well the nature of FPGA hardware. In short, DiracDeltaNet's small model size, low computational OP count, ultra-low precision and simplified operators allow us to co-design a highly customized computing unit for an FPGA. We implement the computing units for DiracDeltaNet on an Ultra96 SoC system through high-level synthesis. Our accelerator's final top-5 accuracy of 88.2% on ImageNet, is higher than all the previously reported embedded FPGA accelerators. In addition, the accelerator reaches an inference speed of 96.5 FPS on the ImageNet classification task, surpassing prior works with similar accuracy by at least 16.9×. Qijing Huang 0001, Bichen Wu, Tianjun Zhang, Liang Ma 0003, Giulio Gambardella, Michaela Blott, Luciano Lavagno, Kees A. Vissers, John Wawrzynek, Kurt Keutzer |
FPGA | 10 |
| 2019 | Centrifuge: Evaluating full-system HLS-generated heterogenous-accelerator SoCs using FPGA-AccelerationabstractTo overcome the end of traditional scaling, modern SoC systems consist of general-purpose compute augmented with large numbers of specialized accelerators. However, building and evaluating these systems is extremely expensive and time-consuming, even in early stages of development. While high-level modeling and back-of-the-envelope calculations can provide early insights into a new system, there are key effects that only manifest at the full-system level. However, full-system design has traditionally required writing RTL or developing complex software models for the entire design. In this paper, we describe a methodology and implement an open-source flow (“Centrifuge”) that can rapidly generate and evaluate heterogeneous SoCs by combining an HLS toolchain with the open-source FireSim FPGA-accelerated simulation platform. Our system can quickly produce complete SoC systems with many integrated HLS-generated accelerators as specified by the user, simulate them quickly and cycle-accurately on FPGAs, and run complete software stacks on top, including booting Linux and running full application frameworks. Our system allows users to easily explore a variety of accelerator integration techniques, by automatically integrating accelerators in several ways-as tightly coupled RoCC accelerators, as accelerators that communicate over the standard on-chip network, and lastly as “disaggregated” accelerators that are directly attached to an Ethernet network between SoCs. By integrating these tools, our methodology allows users to rapidly generate an entire hardware/software stack for a customized SoC that can be fabricated as an ASIC and evaluate its end-to-end performance using cycle-exact FPGA simulation, allowing for agile design-space exploration of novel accelerator-based systems. Qijing Huang 0001, Christopher Yarp, Sagar Karandikar, Nathan Pemberton, Benjamin Brock, Liang Ma 0003, Guohao Dai 0001, Robert Quitt, Krste Asanovic, John Wawrzynek |
ICCAD | 10 |
| 2019 | Antenna Array Geometries for Directional Wireless NetworksabstractAs we move to higher carrier frequencies, directional wireless networks using planar arrays with many antenna elements will become common. Directional radios have been implemented using a variety of antenna array geometries. While it is clear that the optimal choice of array geometry is effected by the physical extent of the network space, there has been limited study of the interaction of array geometry and system performance under the realistic assumption of a finite operating space. In this study, we examine antenna array geometries in directional wireless networks and their effect on interference using probabilistic analysis. We treat the nodes as having a uniform distribution in a given physical space and calculate the expected interference. We observe that linear antenna arrays, independent of position, perform significantly better than other rectangular antenna array geometries given a fixed number of antennas. James C. Martin, Robert W. Brodersen, John Wawrzynek |
WCNC | 3 |
| 2019 | HyVE: Hybrid Vertex-Edge Memory Hierarchy for Energy-Efficient Graph ProcessingabstractHigh energy consumption of conventional memory modules (e.g., DRAMs) hinders the further improvement of large-scale graph processing's energy efficiency. The emerging resistive random-access memory (ReRAM) has shown great potential in providing an energy-efficient memory module. However, the performance of ReRAMs suffers from data access patterns with poor locality and large amounts of written data, which are common in graph processing. In this paper, we propose HyVE, a Hybrid Vertex-Edge memory hierarchy for energy-efficient graph processing. In HyVE, we avoid random access and data written to ReRAM modules. HyVE can reduce memory energy consumption by 86.17 percent compared with conventional memory systems. We have also proposed data sharing and bank-level power-gating schemes, which improve the energy efficiency by 1.60x and 1.53x. By analyzing the graph processing model on ReRAMs, we show that ReRAMs are good for read-intensive operations in graph processing (e.g., reading edges), while ReRAM crossbars are not suitable for processing edges because of heavy writing overheads. Our evaluations show that the optimized design achieves two orders of magnitude and 5.90x energy efficiency improvement compared with the CPU-based and conventional memory hierarchy based designs, respectively. Moreover, HyVE achieves 2.83x energy reduction compared with the previous ReRAM-based graph processing architecture. Guohao Dai 0001, Yu Wang 0002, Huazhong Yang, John Wawrzynek |
IEEE Trans. Computers | 5 |
| 2018 | NewGraph: Balanced Large-Scale Graph Processing on FPGAs with Low Preprocessing OverheadsabstractLarge-scale graph processing has been widely required in various domains, including social network analysis, neural network modeling, database computing, etc. Performance of large-scale graph suffers from random and unpredictable data access pattern, which leads to drastic bandwidth degradation on caches, DRAMs, and disks. The support for high bandwidth random access makes SRAMs the promising solution for graph processing. Many FPGA based large-scale graph processing systems have been proposed in previous works and taken advantage of the SRAM resources. Guohao Dai 0001, Yu Wang 0002, Huazhong Yang, John Wawrzynek |
FCCM | 5 |
| 2018 | Receiver Adaptive Beamforming and Interference of Indoor Environments in mmWaveabstractWe consider networks consisting of nodes equipped with large antenna count mmWave arrays enabling narrow beam patterns. For the first time, we investigate and quantify the adverse impact of interference on network capacity for these very narrow beams. Network capacity is studied in terms of both node density as well as antenna array size. Finally, we show a distributed adaptive receiver algorithm that can reduce the adverse interference impact up to 60%. James C. Martin, Robert W. Brodersen, John Wawrzynek |
PIMRC | 3 |
| 2018 | AWStream: adaptive wide-area streaming analyticsabstractThe emerging class of wide-area streaming analytics faces the challenge of scarce and variable WAN bandwidth. Non-adaptive applications built with TCP or UDP suffer from increased latency or degraded accuracy. State-of-the-art approaches that adapt to network changes require developer writing sub-optimal manual policies or are limited to application-specific optimizations. Ben Zhang 0003, Xin Jin 0008, Sylvia Ratnasamy, John Wawrzynek, Edward A. Lee |
SIGCOMM | 4 |
| 2017 | OLAF'17: Third International Workshop on Overlay Architectures for FPGAs
Hayden Kwok-Hay So, John Wawrzynek |
FPGA | 2 |
| 2017 | Synthesis of program binaries into FPGA accelerators with runtime dependence validationabstractWith the emergence of readily available FPGA cloud computing platforms, ease of use for application developers becomes increasingly crucial to widespread adoption. Synthesis directly from binaries has been proposed as an option to alleviate the design burden. However, in program binaries, loop bounds and loop invariants used for memory index calculation are often compiled into runtime data stored in registers or memories, making static loop dependence analysis infeasible. In this work, a two-phase approach is presented to address this issue with: 1) an offine phase to recover memory access patterns in the loop for data dependence analysis based on software profiling. and 2) an online phase to dynamically check for parallelization assertions. We use this method to discover and exploit coarse-grained parallelism for accelerating compute-intensive affine loops in binaries. With our target platform, the Zynq-7000 FPGA SoC, we ran and examined four benchmarks with our flow: GemsFDTD, Matrix Multiply, Sobel Edge Detection, and K-Nearest Neighbors. Results show up to 9.5x speedup with our flow compared to the pure software flow on the 667 MHz ARM Cortex A9 processor. Shaoyi Cheng, Qijing Huang 0001, John Wawrzynek |
FPT | 3 |
| 2017 | Selection and Aggregation of Location Information Provisioning ServicesabstractAggregation of location estimates from multiple services for provisioning of location information enhances the accuracy and robustness of the final location information. Recent Internet of Things (IoT)-based localization service architectures therefore envision a "manager" for selecting and invoking provisioning services and, in the later step, aggregating the received information and providing it to location-based applications. The selection of provisioning services should take into account the accuracy and latency requirements from the applications and accuracy, latency, and power consumption characteristics of provisioning services. However, it is yet unclear how such selection should be made. We propose two algorithms for the selection of provisioning services aiming at meeting latency and subsequently accuracy requirements from the applications, one subject to minimizing per-request power consumption, while the other subject to a per-time bucket power minimization. In the considered examples, we show that the per-time bucket optimization achieves around 25% better performance in terms of power consumption, while trading-off accuracy satisfaction. Filip Lemic, Vlado Handziski, Mladen Miksa, Jan M. Rabaey, John Wawrzynek, Adam Wolisz |
ICCCN | 5 |
| 2017 | SLSR: A flexible middleware localization service architectureabstractLocation information of mobile devices is a foundational input to location-based services and a valuable source of context information in wireless networks. To maximize the value, we need location information that is accurate, robust, and promptly and seamlessly available. Unfortunately, individual localization services seldom satisfy all these requirements. For achieving that vision, a set of challenges has to be addressed, pertaining to handover, fusion, and integration of different sources of location information. Current approaches for integration of individual localization services are either not specific enough or are limited in scope and lack flexibility. In the following, we provide a detailed design and a prototypical implementation of the Standardized Localization Service (SLSR), a middleware architecture for achieving those goals. We instantiate the service in an office environment and perform exhaustive performance benchmarking in a testbed specifically designed for supporting such experimentation. Our results characterize the effects of different functional components envisioned in the SLSR on its performance. Our results also quantify the accuracy benefits of fusion of representative sources of location information. Filip Lemic, Vlado Handziski, Ivan Azcarate, John Wawrzynek, Jan M. Rabaey, Adam Wolisz |
IPIN | 4 |
| 2016 | OLAF'16: Second International Workshop on Overlay Architectures for FPGAsabstractThe Second International Workshop on Overlay Architec- tures for FPGAs is held in Monterey, California, USA, on February 21, 2016 and co-located with FPGA 2016: The 24th ACM/SIGDA International Symposium on Field Pro- grammable Gate Arrays. The main objective of the work- shop is to address how overlay architectures can help address the challenges and opportunities provided by FPGA-based reconfigurable computing. The workshop provides a venue for researchers to present and discuss the latest develop- ments in FPGA overlay architecture and related areas. We have assembled a program of six refereed papers and a panel discussion with prominent experts in the field. Hayden Kwok-Hay So, John Wawrzynek |
FPGA | 2 |
| 2016 | Synthesis of statically analyzable accelerator networks from sequential programsabstractThis paper describes a general framework for transforming a sequential program into a network of processes, which are then converted to hardware accelerators through high level synthesis. Also proposed is a complementing technique for performing static deadlock analysis of the generated accelerator network. The interactions between the accelerators' schedules, the capacity of the communication channels in the network and the memory access mechanisms are all incorporated into our model, such that potential artificial deadlocks can be detected and resolved a priori. An algorithm optimized for FPGA implementation is developed and applied through our transformation framework. A set of irregular computation kernels are converted into networks of FPGA accelerators. Compared to hardware accelerators generated without our transformation, the accelerator networks achieve significantly better performance. Shaoyi Cheng, John Wawrzynek |
ICCAD | 2 |
| 2016 | Toward standardized localization serviceabstractLocalization services available on today's mobile devices are proprietary and leverage a limited set of sources of location information. Integration of new location estimation methods is therefore cumbersome, requiring adaptation to the specific interfaces of the proprietary location service. In addition, location-based applications are tightly interwoven with the location service that is typically provided by the operating system, hence these applications require significant restructuring to be able to run with another location service. To address these problems, we propose a modular localization service architecture that consists of location-based applications, an integrated location service enabling a fusion of so-called elementary location services, and resources for generating location information. A unified style of interaction among these components is enabled by a set of well-defined Application Programming Interfaces (APIs). The practicability and advantages of the proposed design is demonstrated by outlining how the APIs can be realized using modern types of component interactions. Filip Lemic, Vlado Handziski, Nitesh Mor, Jan M. Rabaey, John Wawrzynek, Adam Wolisz |
IPIN | 5 |
| 2016 | Localization as a feature of mmWave communicationabstractmmWave (millimeter-Wave) is a very promising technology for the future wireless communication. To mitigate its high attenuation characteristics, mmWave communication frequently employs directional beamforming for both transmission and reception. Localization commonly takes advantage of directionality in RF frequencies in urban and indoor environments. In this paper, we use lessons learned from classical RF-based localization for discussing a set of feasible localization approaches in the context of mmWave bands. We further map the requirements of each discussed localization approach to design requirements for future mmWave devices and assess the expected accuracy of such approaches for a set of realistic scenarios. Our results show that mmWave-based localization is promising in both its availability and accuracy, even in the presence of a limited number of localization anchor nodes. Filip Lemic, James C. Martin, Christopher Yarp, Douglas S. Chan, Vlado Handziski, Robert W. Brodersen, Gerhard P. Fettweis, Adam Wolisz, John Wawrzynek |
IWCMC | 9 |
| 2015 | The Cloud is Not Enough: Saving IoT from the Cloud
Ben Zhang 0003, Nitesh Mor, John Kolb, Douglas S. Chan, Ken Lutz, Eric Allman, John Wawrzynek, Edward A. Lee, John Kubiatowicz |
HotStorage | 7 |
| 2014 | Architectural synthesis of computational pipelines with decoupled memory accessabstractAs high level synthesis (HLS) moves towards mainstream adoption among FPGA designers, it has proven to be an effective method for rapid hardware generation. However, in the context of offloading compute intensive software kernels to FPGA accelerators, current HLS tools do not always take full advantage of the hardware platforms. In this paper, we present an automatic flow to refactor and restructure processor-centric software implementations, making them better suited for FPGA platforms. The methodology generates pipelines that decouple memory operations and data access from computation. The resulting pipelines have much better throughput due to their efficient use of the memory bandwidth and improved tolerance to data access latency. The methodology complements existing work in high-level synthesis, easing the creation of heterogeneous systems with high performance accelerators and general purpose processors. With this approach, for a set of non-regular algorithm kernels written in C, a performance improvement of 3.3 to 9.1x is observed over direct C-to-Hardware mapping using a state-of-the-art HLS tool. Shaoyi Cheng, John Wawrzynek |
FPT | 2 |
| 2013 | Reconfigurable computing in the era of post-silicon scaling [panel discussion]abstractSummary form only given, as follows. Although transistor densities continue to scale exponentially, the failure of Dennard Scaling prevents us from maximally utilizing die area in future power-constrained multicore processors-a phenomenon referred to as "Dark Silicon". Alternative energy-efficient architectures based on FPGAs, GPGPUs, ASICs, MPPAs, etc. are likely to continue on an exponential scaling trajectory while outperforming conventional architectures by an order-of-magnitude or more. With the impending threat of dark silicon, there is a critical window of opportunity for reconfigurable computing to become a mainstream ingredient and driver of future, scalable computer architectures. Before this can happen, major challenges and opportunities must be addressed: (1) how to gracefully integrate reconfigurable computing into existing software and hardware ecosystems, (2) how to build tools, languages, and compilers for agile application development and debugging, (3) how to identify and exploit emerging applications in datacenters and in energy-constrained form factors, (4) how to train and educate students and practitioners to use these systems in sustainable ways, and (5) how to define new and stable boundaries between software and hardware that make it easier to exploit reconfigurable computing. This panel brings together pioneers and experts in computer architecture and reconfigurable computing to discuss opportunities and challenges in the wake of dark silicon. Eric S. Chung, Doug Burger, Mike Butts, Jan Gray, Charles P. Thacker, Kees A. Vissers, John Wawrzynek |
FCCM | 7 |
| 2012 | Chisel: constructing hardware in a Scala embedded languageabstractIn this paper we introduce Chisel, a new hardware construction language that supports advanced hardware design using highly parameterized generators and layered domain-specific hardware languages. By embedding Chisel in the Scala programming language, we raise the level of hardware design abstraction by providing concepts including object orientation, functional programming, parameterized types, and type inference. Chisel can generate a high-speed C++-based cycle-accurate software simulator, or low-level Verilog designed to map to either FPGAs or to a standard ASIC flow for synthesis. This paper presents Chisel, its embedding in Scala, hardware examples, and results for C++ simulation, Verilog emulation and ASIC synthesis. Jonathan Bachrach, Huy Vo, Brian C. Richards, Yunsup Lee, Andrew Waterman, Rimas Avizienis, John Wawrzynek, Krste Asanovic |
DAC | 7 |
| 2012 | Exploiting Memory-Level Parallelism in Reconfigurable AcceleratorsabstractAs memory accesses increasingly limit the overall performance of reconfigurable accelerators, it is important for high level synthesis (HLS) flows to discover and exploit memory-level parallelism. This paper develops 1) a framework where parallelism between memory accesses can be revealed from runtime profile of applications and provided to a high level synthesis flow, and 2) a novel multi-accelerator/multi-cache architecture to support parallel memory accesses, taking advantage of the high aggregated memory bandwidth found in modern FPGA devices. Our experimental results have shown that for 10 accelerators generated from 9 benchmark applications, circuits using our proposed memory structure achieve on average 52% improved performance over accelerators using a traditional memory interface. We believe that our study represents a solid advance towards achieving memory-parallel embedded computing on hybrid CPU+FPGA platforms. Shaoyi Cheng, Mingjie Lin, Hao Jun Liu, Simon Scott, John Wawrzynek |
FCCM | 5 |
| 2011 | Bridging the GPGPU-FPGA efficiency gapabstractThis paper compares an implementation of a Bayesian inference algorithm across several FPGAs and GPGPUs, while embracing both the execution model and high-level architecture of a GPGPU. Our study is motivated by recent work in template-based programming and architectural models for FPGA computing. The comparison we present is meant to demonstrate the FPGA's potential, while constraining the design to follow the microarchitectural template of more programmable devices such as GPGPUs. Christopher W. Fletcher, Ilia A. Lebedev, Narges Bani Asadi, Daniel Burke, John Wawrzynek |
FPGA | 5 |
| 2011 | Using many-core architectural templates for FPGA-based computing (abstract only)abstractTruly unleashing the computing potential of FPGAs, as well as widening their applicability, demands alleviating cumbersome HDL programming and relieving laborious manual optimization. Towards this end, we propose a Many-core Approach to Reconfigurable Computing (MARC) that enables efficient high-performance computing for applications expressed with imperative programming languages such as C/C++ without constructing FPGA computing machines from scratch when targeting various applications within the same or similar problem domains. A MARC system achieves high computing performance by leveraging a many-core architectural template, sophisticated logic synthesizing techniques, and state-of-art compiler optimization technology. In addition, MARC exploits abundant special FPGA resources such as distributed block memories and DSP blocks to implement complete single-chip high efficiency many-core microarchitectures. The key benefits of MARC include (i) allowing programmers to easily express parallelism through a high-level programming language, (ii) supporting coarse-grain multithreading and dataflow-style fine-grain threading while permitting bit-level resource control, and (iii) greatly reducing the effort required to re-purpose the hardware system for different algorithms or different applications. Mingjie Lin, Shaoyi Cheng, John Wawrzynek |
FPGA | 3 |
| 2011 | Should the academic community launch an open-source FPGA device and tools effort?: evening panelabstractFor years, many academic researchers in reconfigurable computing have been frustrated by their reliance on commercial FPGAs and tools. Commercial FPGAs have highly complex micro-architectures, come with undocumented binary interfaces, have no compatibility between generations, and come with difficult to use proprietary place and route tools. The FPGA vendors are making the right moves for serving their commercial customer base, but it seems at times these moves are in conflict with the needs of the academic research community. John Wawrzynek |
FPGA | 1 |
| 2010 | High-throughput bayesian computing machine with reconfigurable hardwareabstractWe use reconfigurable hardware to construct a high throughput Bayesian computing machine (BCM) capable of evalu- ating probabilistic networks with arbitrary DAG (directed acyclic graph) topology. Our BCM achieves high throughput by exploiting the FPGA's distributed memories and abundant hardware structures (such as long carry-chains and registers), which enables us to 1) develop an innovative memory allocation scheme based on a maximal matching algorithm that completely avoids memory stalls, 2) optimize and deeply pipeline the logic design of each processing node, and 3) optimally schedule them. The BCM architecture we present not only can be applied to many important algorithms in artificial intelligence, signal processing, and digital communications, but also has high reusability, i.e., a new application needs not change a BCM's hardware design, only new task graph processing and code compilation are necessary. Moreover, the throughput of a BCM scales almost linearly with the size of the FPGA on which it is implemented. Mingjie Lin, Ilia A. Lebedev, John Wawrzynek |
FPGA | 3 |
| 2010 | OpenRCL: Low-Power High-Performance Computing with Reconfigurable DevicesabstractThis work presents the Open Reconfigurable Computing Language (OpenRCL) system designed to enable low-power high-performance reconfigurable computing with imperative programming language such as C/C++. The key idea is to expose the FPGA platform as a compiler target for applications expressed in the OpenCL paradigm. To this end, we present a combination of low-level virtual machine instruction set, execution model, many-core architecture, and associated compiler to achieve high performance and power efficiency by exploiting the FPGA's distributed memories and abundant hardware structures (such as DSP blocks, long carry-chains, and registers). Our resulting OpenRCL system not only allows programmers to easily express parallelism through the API defined in the OpenCL standard but also supports coarse-grain multithreading and dataflow-style fine-grain threading while permitting bit-level resource control. An OpenRCL prototype machine with 30 processing nodes was implemented using a Virtex-5 (XCV5LX155T-2) FPGA. For the well-known Parallel Prefix Sum (Scan) problem, comparing the runtime of the same problem on a GeForce 9400m using the OpenCL SDK from Apple Inc., the OpenRCL machine demonstrates comparable performance with a 5x reduction in core power consumption. Mingjie Lin, Ilia A. Lebedev, John Wawrzynek |
FPL | 3 |
| 2010 | ParaLearn: a massively parallel, scalable system for learning interaction networks on FPGAsabstractParaLearn is a scalable, parallel FPGA-based system for learning interaction networks using Bayesian statistics. ParaLearn includes problem specific parallel/scalable algorithms, system software and hardware architecture to address this complex problem. Narges Bani Asadi, Christopher W. Fletcher, Greg Gibeling, John Wawrzynek, Wing H. Wong, Garry P. Nolan |
ICS | 4 |
| 2010 | Improving FPGA Placement With Dynamically Adaptive Stochastic TunnelingabstractThis paper develops a dynamically adaptive stochastic tunneling (DAST) algorithm to avoid the “freezing” problem commonly found when using simulated annealing for circuit placement on field-programmable gate arrays (FPGAs). The main objective is to reduce the placement runtime and improve the quality of final placement. We achieve this by allowing the DAST placer to tunnel energetically inaccessible regions of the potential solution space, adjusting the stochastic tunneling schedule adaptively by performing detrended fluctuation analysis, and selecting move types dynamically by a multi-modal scheme based on Gibbs sampling. A prototype annealing-based placer, called DAST, was developed as part of this paper. It targets the same computer-aided design flow as the standard versatile placement and routing (VPR) but replaces its original annealer with the DAST algorithm. Our experimental results using the benchmark suite and FPGA architecture file which comes with the Toronto VPR5 software package have shown a 18.3% reduction in runtime and a 7.2% improvement in critical-path delay over that of conventional VPR. Mingjie Lin, John Wawrzynek |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2010 | Exploring FPGA Routing Architecture StochasticallyabstractThis paper proposes a systematic strategy to efficiently explore the design space of field-programmable gate array (FPGA) routing architectures. The key idea is to use stochastic methods to quickly locate near-optimal solutions in designing FPGA routing architectures without exhaustively enumerating all design points. The main objective of this paper is not as much about the specific numerical results obtained, as it is to show the applicability and effectiveness of the proposed optimization approach. To demonstrate the utility of the proposed stochastic approach, we developed the tool for optimizing routing architecture (TORCH) software based on the versatile place and route tool. Given FPGA architecture parameters and a set of benchmark designs, TORCH simultaneously optimizes the routing channel segmentation and switch box patterns using the performance metric of average interconnect power-delay product estimated from placed and routed benchmark designs. Special techniques - such as incremental routing, infrequent placement, multi-modal move selection, and parallelized metric evaluation - are developed to reduce the overall run time and improve the quality of results. Our experimental results have shown that the stochastic design strategy is quite effective in co-optimizing both routing channel segmentation and switch patterns. With the optimized routing architecture, relative to the performance of our chosen architecture baseline, TORCH can achieve average improvements of 24% and 15% in delay and power consumption for the 20 largest Microelectronics Center of North Carolina benchmark designs, and 27% and 21% for the eight benchmark designs synthesized with the Altera Quartus II University Interface Program tool. Additionally, we found that the average segment length in an FPGA routing channel should decrease with technology scaling. Finally, we demonstrate the versatility of TORCH by illustrating how TORCH can be used to optimize other aspects of the routing architecture in an FPGA. Mingjie Lin, John Wawrzynek, Abbas El Gamal |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2009 | Using adaptive routing to compensate for performance heterogeneityabstractScalable and power efficient multi-core architectures must be performance heterogeneous to accommodate semi-conductor parametric variations and non-uniform access to shared resources. Due to its rate matching, a NoC on a Voltage-Frequency Island architecture can connect cores without forcing each one to give up its own operating point for the chip-wide common worst case. With run-time adaptive routing and task-to-core mapping, a NoC can run at the average not the worst case network saturation bandwidth. These run-time processes compensate for variations because they match application resource requirements with heterogeneous cores and routers. We focus on adaptive routing that simultaneously combats communication load imbalance from on-die variations and application topology. We show that even with static, fixed task-to-core mapping on multi-core architectures affected by stochastic variations, our MATC router increases the expected saturation bandwidth by 7-25% vs Dimension Order router. With systematic variations, the improvements are 5-50%. These gains compensate for saturation bandwidth degradation due to manufacturing variations and help to reduce design guard-bands. Yury Markovsky, Yatish Patel, John Wawrzynek |
NOCS | 3 |
| 2009 | A design methodology for domain-optimized power-efficient supercomputingabstractAs power has become the pre-eminent design constraint for future HPC systems, computational efficiency is being emphasized over simply peak performance. Recently, static benchmark codes have been used to find a power efficient architecture. Unfortunately, because compilers generate sub-optimal code, benchmark performance can be a poor indicator of the performance potential of architecture design points. Therefore, we present hardware/software cotuning as a novel approach for system design, in which traditional architecture space exploration is tightly coupled with software auto-tuning for delivering substantial improvements in area and power efficiency. We demonstrate the proposed methodology by exploring the parameter space of a Tensilica-based multi-processor running three of the most heavily used kernels in scientific computing, each with widely varying micro-architectural requirements: sparse matrix vector multiplication, stencil-based computations, and general matrix-matrix multiplication. Results demonstrate that co-tuning significantly improves hardware area and energy efficiency -- a key driver for next generation of HPC system design. Marghoob Mohiyuddin, Mark Murphy, Leonid Oliker, John Shalf, John Wawrzynek, Samuel Williams 0001 |
SC | 5 |
| 2007 | RAMP Blue: A Message-Passing Manycore System in FPGAsabstractWe are developing a set of reusable design blocks and several prototype systems for emulation of multi-core architectures in FPGAs. RAMP Blue is the first of these prototypes and was designed to emulate a distributed-memory message-passing architecture. The system consists of 768-1008 MicroBlaze cores in 64-84 Virtex-II Pro 70 FPGAs on 16-21 BEE2 boards, surpassing the milestone of 1000 cores in a standard 42U rack. An architecture based on point-to-point channels and switches using a combination of custom and generic hardware provides the functionality. Virtual-cut-through dimensional routing on one of two hybrid topologies with virtual channels provides the connectivity. A control network with a tree topology provides management and debugging capabilities. A software infrastructure consisting of GCC, uClinux and UPC allows running off-the-shelf applications and scientific benchmarks. Initial performance is encouraging for emulation purposes. In this paper we report on the design and implementation of RAMP Blue and discuss our experiences and lessons learned. Alex Krasnov, John Wawrzynek, Greg Gibeling, Pierre-Yves Droz |
FPL | 3 |
| 2007 | Adventures with a Reconfigurable Research PlatformabstractSummary form only given. The computer industry is at a cross-roads. The problems associated with scaling uniprocessor performance has forced all major computer manufactures to turn to multi-and many-core architectures. This sea change in processor design has created many opportunities for field programmable logic. In the RAMP project, we are developing an affordable and versatile multiprocessor emulation platform being built as a large collaborative effort. RAMP hardware, from processors to caches to networks, is implemented in FPGAs for flexibility, accuracy, visibility, cost and performance. It is designed to be composable, where different components can be quickly written, assembled and run. By using hardware rather than simulation, RAMP will be fast enough to run real codes and be useful to software. John Wawrzynek |
FPL | 1 |
| 2006 | Research accelerator for multiple processors
David A. Patterson 0001, Arvind 0001, Krste Asanovic, Derek Chiou, James C. Hoe, Christoforos E. Kozyrakis, Shih-Lien Lu, Mark Oskin, Jan M. Rabaey, John Wawrzynek |
Hot Chips Symposium | 10 |
| 2005 | Defect Tolerance in Multiple-FPGA SystemsabstractSRAM-based FPGAs have an inherent capacity for defect tolerance. We propose a simple scheme that exploits this potential in multiple-FPGA systems. The symmetry of the system is exploited to yield a large number of possible mappings of bitstreams on FPGAs, which results in a high probability that at least one functional mapping exists. We show that the behavior of a system built using a large number of defective FPGAs approaches that of the ideal defect-free system. Various interconnection topologies such as the tree, the crossbar, and a hybrid form are compared. Zohair Hyder, John Wawrzynek |
FPL | 2 |
| 2004 | The SFRA: a corner-turn FPGA architectureabstractFPGAs normally operate at whatever clock rate is appropriate for the loaded configuration. When FPGAs are used as computational devices in a larger system, however, it is better to employ fixed-frequency FPGAs operating at a high clock frequency. Such fixed-frequency arrays require pipelined interconnect structures, which are difficult to support in a traditional FPGA architecture. We have developed a novel approach, called a interconnect, based on a Manhattan array of logically depopulated S-boxes with full connectivity but limited routability. This interconnect supports new polynomial-time routing techniques while maintaining conventional placement and other upstream toolflow. We have used the corner-turn interconnect to define a fixed-frequency FPGA architecture, the SFRA, that is largely compatible with the Xilinx Virtex while providing higher speed, pipelined operation. Our tools automatically repipeline designs to operate at the SFRA's intrinsic clock frequency. Since the arrays are largely compatible, we directly compare the SFRA with the Virtex on four benchmark designs. On these benchmarks, the SFRA offers higher throughput and competitive throughput per area. The SFRA routing and retiming tools also run one to two orders of magnitude faster than their Xilinx counterparts. Nicholas Weaver, John R. Hauser, John Wawrzynek |
FPGA | 3 |
| 2003 | Stochastic, spatial routing for hypergraphs, trees, and meshesabstractFPGA place and route is time consuming, often serving as the major obstacle inhibiting a fast edit-compile-test loop in prototyping and development and the major obstacle preventing late-bound hardware and design mapping for reconfigurable systems. Previous work showed that hardware-assisted routing can accelerate fanout-free routing on Fat-Trees by three orders of magnitude with modest modifications to the network itself. In this paper, we show how these techniques can be applied to any FPGA and how they can be implemented on top of LUT networks in cases where modification of the FPGA itself is not justified. We further show how to accommodate fanout and how to achieve comparable route quality to software-based methods. For a tree network, we estimate an FPGA implementation of our routing logic could route the Toronto Place and Route Benchmarks at least two orders of magnitude faster than a software Pathfinder while achieving within 3% of the aggregate quality. Preliminary results on small mesh benchmarks achieve within one track of vpr-fast. Randy Huang, John Wawrzynek, André DeHon |
FPGA | 2 |
| 2003 | Post-placement C-slow retiming for the xilinx virtex FPGAabstractC-slow retiming is a process of automatically increasing the throughput of a design by enabling fine grained pipelining of problems with feedback loops. This transformation is especially appropriate when applied to FPGA designs because of the large number of available registers. To demonstrate and evaluate the benefits of C-slow retiming, we constructed an automatic tool which modifies designs targeting the Xilinx Virtex family of FPGAs. Applying our tool to three benchmarks: AES encryption, Smith/Waterman sequence matching, and the LEON 1 synthesized microprocessor core, we were able to substantially increase the total throughput. For some parameters, throughput is effectively doubled. Nicholas Weaver, Yury Markovsky, Yatish Patel, John Wawrzynek |
FPGA | 4 |
| 2003 | Quality based compute-resource allocation in real-time signal processingabstractWe present a novel method for controlling the complexity of real-time signal processing computational tasks, in order to make sure that a total quality metric for all the signal processing tasks is maximized. The method makes decisions about how much compute power is allocated to each task through past observations of the input and output data of each task. We present preliminary results from filtering applications that demonstrate the ability of the system to maximize the total quality of a large number of tasks under a real-time computational constraint. Joseph Yeh, John Wawrzynek |
ICASSP (2) | 2 |
| 2002 | Hardware-Assisted Fast RoutingabstractTo fully realize the benefits of partial and rapid reconfiguration of field-programmable devices, we often need to dynamically schedule computing tasks and generate instance-specific configurations-new graphs which must be routed during program execution. Consequently, route time can be a significant overhead cost reducing the achievable net benefits of dynamic configuration generation. BY adding hardware to accelerate routing, we show that it is possible to compute routes in one thousandth the time of a traditional, software router and achieve routes that are within 5% of the state-of-the-art offline routing algorithms for a sample set of application netlists and within 25% for a set of difficult synthetic benchmarks. We further outline how strategic use of parallelism can allow the total route time to scale substantially less than linearly in graph size. We detail the source of the benefits in our approach and survey a range of options for hardware assistance that van, from a speedup of over 10/spl times/ with modest hardware overhead to speedups in excess of 1000/spl times/. André DeHon, Randy Huang, John Wawrzynek |
FCCM | 3 |
| 2002 | The Effects of Datapath Placement and C-Slow Retiming on Three Computational BenchmarksabstractSummary form only given. Two important optimizations within the FPGA design process, C-slow retiming and datapath placement, offer significant benefits for designers. Many have advocated and implemented tools to use these techniques in both automatic and semiautomatic manner but they have not made their way into conventional FPGA toolflows. C-slow retiming is a method of accelerating computations that include feedback loops. Instead of having a single instance of the computation, the feedback loop is pipelined so that C separate instances are all calculated simultaneously. This allows fine grained pipelining to occur even in designs that include feedback loops, such as single round cryptographic implementations or microprocessors. Done properly, it imposes a significant but not imposing latency penalty for single computations while offering huge increases in throughput. Datapath placement is simply constructing the design in a manner that accounts for the higher level data flows. This offers several benefits, including improved performance, more physically compact designs, shorter wires, and faster place and route times when the FPGA is heavily utilized. Even for designs with less structure which are amenable to simulated annealing, datapath placement may still offer a significant benefit. To clearly demonstrate the importance of these optimizations we have hand-modified three computational benchmarks which represent significant themes within FPGA computation: Rijndael/AES encryption, Smith/Waterman, and a simplified 32-bit microprocessor datapath. All three represent significantly different modes of computation within FPGAs, but all gain significantly from the use of these techniques. Nicholas Weaver, John Wawrzynek |
FCCM | 2 |
| 2002 | Analysis of quasi-static scheduling techniques in a virtualized reconfigurable machineabstractThe SCORE compute model uses fixed-size, virtual compute and memory pages connected by stream links to capture the definition ofa computation abstracted from the detailed size ofthe physical hardware. When the number of physical compute pages is smaller than the number of virtual compute pages in the abstract computation graph, the design is time-multiplexed onto the available physical hardware. A key component ofthis strategy is an automatic scheduler that selects the temporal sequencing ofvirtual resources onto the physical device. We describe a quasistatic scheduling strategy that retains the full semantic power of the dynamic SCORE flow graph while taking advantage ofstatic scheduling techniques at program load time to hoist most ofthe computational work out of the inner scheduling loops. This strategy reduces online scheduling work per reconfiguration epoch by an order of magnitude. In addition, a more global perspective available from offline-scheduling improves schedule quality, resulting in a net reduction oftotal execution time by 46–81%. 1. Yury Markovsky, Eylon Caspi, Randy Huang, Joseph Yeh, Michael Chu, John Wawrzynek, André DeHon |
FPGA | 6 |
| 2001 | A case for network musical performanceabstractA Network Musical Performance (NMP) occurs when a group of musicians, located at different physical locations, interact over a network to perform as they would if located in the same room. In this paper, we present a case for NMP as a practical Internet application, and describe a method to ameliorate the effect of late and lost packets on NMP. We describe an NMP system that embodies this concept, that combines several existing standards (MIDI, MPEG 4 Structured Audio, RTP/AVP, and SIP) with a new RTP packetization for MIDI performance. We analyze NMP experiments performed on CalREN2 hosts on the UC Berkeley, Stanford, and Caltech campuses. John Lazzaro, John Wawrzynek |
NOSSDAV | 2 |
| 2000 | Adapting software pipelining for reconfigurable computingabstractThe Garp compiler and architecture have been developed in parallel, in part to help investigate whether features of the architecture help facilitate rapid, automatic compilation utilizing the Garp's rapidly reconfigurable coprocessor. Previously reported work for compiling to Garp has drawn heavily on techniques from software compilation rather than high-level synthesis. That trend continues in this paper, which describes the extension of those techniques to support pipelined execution of loops on the coprocessor. Even though it targets hardware, our approach resembles VLIW software pipelining much more than it resembles hardware synthesis retiming algorithms. Timothy J. Callahan, John Wawrzynek |
CASES | 2 |
| 1999 | Reconfigurable Computing: What, Why, and Implications for Design AutomationabstractReconfigurable Computing is emerging as an important new organizational structure for implementing computations.It combines the post-fabrication programmability of processors with the spatial computational style most commonly employed in hardware designs.The result changes traditional "hardware" and "software" boundaries, providing an opportunity for greater computational capacity and density within a programmable media.Reconfigurable Computing must leverage traditional CAD technology for building spatial designs.Beyond that, however, reprogrammablility introduces new challenges and opportunities for automation, including binding-time and specialization optimizations, regularity extraction and exploitation, and temporal partitioning and scheduling. André DeHon, John Wawrzynek |
DAC | 2 |
| 1999 | HSRA: High-Speed, Hierarchical Synchroous Reconfigurable ArrayabstractThere is no inherent characteristic forcing Field Programmable Gate Array (FPGA) or Reconfigurable Computing (RC) Array cycle times to be greater than processors in the same process. Modern FPGAs seldom achieve application clock rates close to their processor cousins because (1) resources in the FPGAs are not balanced appropriately for high-speed operation, (2) FPGA CAD does not automatically provide the requisite transforms to support this operation, and (3) interconnect delays can be large and vary almost continuously, complicating high frequency mapping. We introduce a novel reconfigurable computing array, the High-Speed, Hierarchical Synchronous Reconfigurable Array (HSRA), and its supporting tools. This packagedemonstrates that computing arrays can achieve efficient, high-speedoperation. We have designedand implemented a prototype component in a 0.4 m logic design on a DRAM process which will support 250MHz operation for CAD mapped designs. William Tsu, Kip Macy, Atul Joshi, Randy Huang, Norman Walker, Tony Tung, Omid Rowhani, George Varghese, John Wawrzynek, André DeHon |
FPGA | 9 |
| 1999 | A fixed-point recursive digital oscillator for additive synthesis of audioabstractThis paper summarizes our work adapting a recursive digital resonator for use on sixteen-bit fixed-point hardware. Our modified oscillator is a two-pole filter that maintains frequency precision at a cost of two additional operations per filter sample. The new filter's error properties are expressly matched to use in the range of frequencies relevant to additive synthesis of digital audio and sinusoidal modelling of speech in order to minimize the additional computational overhead. We present the algorithm, an error analysis, a performance analysis, and measurements of an implementation on a fixed-point vector microprocessor system. Todd D. Hodes, John R. Hauser, John Wawrzynek, Adrian Freed, David Wessel |
ICASSP | 3 |
| 1999 | JPEG Quality Transcoding Using Neural Networks Trained With a Perceptual Error MeasureabstractA JPEG Quality Transcoder (JQT) converts a JPEG image file that was encoded with low image quality to a larger JPEG image file with reduced visual artifacts, without access to the original uncompressed image. In this article, we describe technology for JQT design that takes a pattern recognition approach to the problem, using a database of images to train statistical models of the artifacts introduced through JPEG compression. In the training procedure for these models, we use a model of human visual perception as an error measure. Our current prototype system removes 32.2% of the artifacts introduced by moderate compression, as measured on an independent test database of linearly coded images using a perceptual error metric. This improvement results in an average PSNR reduction of 0.634 dB. John Lazzaro, John Wawrzynek |
Neural Comput. | 2 |
| 1998 | Object Oriented Circuit-Generators in JavaabstractGenerators, parameterized code which produces a digital design, have long been a staple of the VLSI community. In recent years, several field programmable gate array (FPGA) design tools have adopted generators, as it is a convenient way to specify reusable designs in a familiar programming environment. We have built a generator framework in Java as a basis for programming reconfigurable devices and as a tool to be embedded in larger development systems. In addition to the conventional benefits of generators, this powerful framework allows for partial evaluation, simulation, specialization, and easy inclusion of other automatic services. In order to verify the utility of this system, we have implemented several applications using this framework and compared them with implementations using schematic capture and HDL synthesis. Our system runs significantly faster and produces comparable or superior results when mapped to a target FPGA. Michael Chu, Nicholas Weaver, Kolja Sulimma, André DeHon, John Wawrzynek |
FCCM | 5 |
| 1998 | Fast Module Mapping and Placement for Datapaths in FPGAsabstractBy tailoring a compiler tree-parsing tool for datapath module mapping, we produce good quality results for datapath synthesis in very fast run time. Rather than flattening the design to gates, we preserve the datapath structure; this allows exploitation of specialized datapath features in FPGAs, retains regularity, and also results in a smaller problem size. To further achive high mapping speed, we formulate the problem as tree covering and solve it efficiently with a linear-time dynamic programming algorithm. In a novel extension to the tree-covering algorithm, we perform module placement simultaneously with the mapping, still in linear time. Integrating placement has the potential to increase the quality of the result since we can optimize total delay including routing delays. Timothy J. Callahan, Philip Chong, André DeHon, John Wawrzynek |
FPGA | 4 |
| 1997 | Datapath-oriented FPGA mapping and placement for configurable computingabstractWidespread acceptance of FPGA-based reconfigurable coprocessors will be expedited if compilation time for FPGA configurations can be reduced to be comparable to software compilation. This research achieves this goal, generating complete datapath layouts in fractions of a second rather than hours. Our algorithm, adapted from instruction selection in compilers, packs multiple operations into single rows of CLBs when possible, while preserving a regular bit-slice layout. Furthermore, placement and thus routing delays are considered simultaneously with packing, so that the total delay, not just the CLB delay, is optimized. Timothy J. Callahan, John Wawrzynek |
FCCM | 2 |
| 1997 | Garp: a MIPS processor with a reconfigurable coprocessorabstractTypical reconfigurable machines exhibit shortcomings that make them less than ideal for general-purpose computing. The Garp Architecture combines reconfigurable hardware with a standard MIPS processor on the same die to retain the better features of both. Novel aspects of the architecture are presented, as well as a prototype software environment and preliminary performance results. Compared to an UltraSPARC, a Garp of similar technology could achieve speedups ranging from a factor of 2 to as high as a factor of 24 for some useful applications. John R. Hauser, John Wawrzynek |
FCCM | 2 |
| 1996 | A Micropower Analog VLSI HMM State Decoder for Wordspotting
John Lazzaro, John Wawrzynek, Richard Lippmann |
NIPS | 2 |
| 1995 | Silicon Models for Auditory Scene Analysis
John Lazzaro, John Wawrzynek |
NIPS | 2 |
| 1995 | SPERT-II: A Vector Microprocessor System and its Application to Large Problems in Backpropagation Training
John Wawrzynek, Krste Asanovic, Brian Kingsbury, James Beck, David Johnson 0001, Nelson Morgan |
NIPS | 1 |
| 1993 | Designing A Connectionist Network SupercomputerabstractThis paper describes an effort at UC Berkeley and the International Computer Science Institute to develop a supercomputer for artificial neural network applications. Our perspective has been strongly influenced by earlier experiences with the construction and use of a simpler machine. In particular, we have observed Amdahl's Law in action in our designs and those of others. These observations inspire attention to many factors beyond fast multiply-accumulate arithmetic. We describe a number of these factors along with rough expressions for their influence and then give the applications targets, machine goals and the system architecture for the machine we are currently designing. Krste Asanovic, James Beck, Jerome A. Feldman, Nelson Morgan, John Wawrzynek |
Int. J. Neural Syst. | 5 |
| 1993 | Silicon auditory processors as computer peripheralsabstractSeveral research groups are implementing analog integrated circuit models of biological auditory processing. The outputs of these circuit models have taken several forms, including video format for monitor display, simple scanned output for oscilloscope display, and parallel analog outputs suitable for data-acquisition systems. Here, an alternative output method for silicon auditory models, suitable for direct interface to digital computers, is described. As a prototype of this method, an integrated circuit model of temporal adaptation in the auditory nerve that functions as a peripheral to a workstation running Unix is described. Data from a working hybrid system that includes the auditory model, a digital interface, and asynchronous software are given. This system produces a real-time X-window display of the response of the auditory nerve model. John Lazzaro, John Wawrzynek, Misha Mahowald, Massimo Sivilotti, Dave Gillespie |
IEEE Trans. Neural Networks | 2 |
| 1993 | The design of a neuro-microprocessorabstractThe architecture of a neuro-microprocessor is presented. This processor was designed using the results of careful analysis of a set of applications and extensive simulation of moderate-precision arithmetic for back-propagation networks. Simulated performance results and test-chip results for the processor are presented. This work is an important intermediate step in the development of a connectionist network supercomputer. John Wawrzynek, Krste Asanovic, Nelson Morgan |
IEEE Trans. Neural Networks | 1 |
| 1992 | SPERT: a VLIW/SIMD microprocessor for artificial neural network computationsabstractSPERT (synthetic perceptron testbed) is a fully programmable single chip microprocessor designed for efficient execution of artificial neural network algorithms. The first implementation is in a 1.2 mu m CMOS technology with a 50 MHz clock rate, and a prototype system is being designed to occupy a double SBus slot within a Sun Sparcstation. SPERT sustains over 300*10/sup 6/ connections per second during pattern classification, and around 100*10/sup 6/ connection updates per second while running the popular error backpropagation training algorithm. This represents a speedup of around two orders of magnitude over a Sparcstation-2 for algorithms of interest. An earlier system produced by the group, the Ring Array Processor (RAP), used commercial DSP chips. Compared with a RAP multiprocessor of similar performance, SPERT represents over an order of magnitude reduction in cost for problems where fixed-point arithmetic is satisfactory.> Krste Asanovic, James Beck, Brian Kingsbury, Phil Kohn, Nelson Morgan, John Wawrzynek |
ASAP | 6 |
| 1992 | Silicon Auditory Processors as Computer Peripherals
John Lazzaro, John Wawrzynek, Misha Mahowald, Massimo Sivilotti, Dave Gillespie |
NIPS | 2 |
| 1991 | Fine-Grain Parallelism with Minimal Hardware Support: A Compiler-Controlled Threaded Abstract Machineabstractarticle Free Access Share on Fine-grain parallelism with minimal hardware support: a compiler-controlled threaded abstract machine Authors: David E. Culler View Profile , Anurag Sah View Profile , Klaus E. Schauser View Profile , Thorsten von Eicken View Profile , John Wawrzynek View Profile Authors Info & Claims ACM SIGOPS Operating Systems ReviewVolume 25Issue Special IssueApr. 1991pp 164–175https://doi.org/10.1145/106974.106990Published:01 April 1991Publication History 229citation1,379DownloadsMetricsTotal Citations229Total Downloads1,379Last 12 Months108Last 6 weeks19 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF David E. Culler, Anurag Sah, Klaus E. Schauser, Thorsten von Eicken, John Wawrzynek |
ASPLOS | 5 |
| 1991 | A Two-Dimensional Topological Compactor With Octagonal GeometryabstractWe present a two-dimensional layout compactor with octagonal geometry.We discuss layout optimization algorithms used within a topological framework.We present the results of comparing the compactor against a standard set of benchmarks.This version of the compactor uses an iterative greedy layout optimization algorithm with wire minimization and produces leafcells more compact than others published. Paul de Dood, John Wawrzynek, Erwin Liu, Roberto Suaya |
DAC | 2 |