EDBT 2026 Demo / reviewers in the wild / expert
Miriam Leeser
dblp:l/MiriamLeeser · also Miriam E. Leeser
· DBLP profile ↗
88ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-5624-056XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 77 · 6 first-author · 10 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Computer networks · 2 · 1 since 2021Theory of computation · 2Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Artifact Evaluation in the FPGA Community and in ACM TRETSabstractArtifact evaluation (AE) is gaining traction across the computer science community as a means of advancing reproducible research and strengthening readers’ confidence in published results. Applying reproducibility to computer systems and architecture research has proven particularly challenging, and the FPGA community faces its own distinct hurdles—namely, the use of non-standard hardware platforms and dependence on specific software tools and versions. In this editorial, we review the history of AE in the FPGA community, compare it to practices in related fields, and discuss challenges and future directions. To date, AE has meaningfully improved the availability and accessibility of artifacts and their documentation, while also increasing readers’ confidence in published findings. Looking ahead, AE is poised to continue growing across the reconfigurable hardware community. Notably, ACM Transactions on Reconfigurable Technology and Systems will now offer AE with the opportunity to earn artifact badges for all accepted papers. Miriam Leeser |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2025 | Miniature: Fast AI Supercomputer Networks Simulation on FPGAs
Yicheng Qian, Ran Shu 0001, Rui Ma 0021, Yang Wang 0053, Derek Chiou, Nadeen Gebara, Luca Piccolboni, Miriam Leeser, Yongqiang Xiong |
APNet | 8 |
| 2025 | LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network InferenceabstractFor FPGA-based neural network accelerators, digital signal processing (DSP) blocks have traditionally been the cornerstone for handling multiplications. This paper introduces LUTMUL, which harnesses the potential of look-up tables (LUTs) for performing multiplications. The availability of LUTs typically outnumbers that of DSPs by a factor of 100, offering a significant computational advantage. By exploiting this advantage of LUTs, our method demonstrates a potential boost in the performance of FPGA-based neural network accelerators with a reconfigurable dataflow architecture. Our approach challenges the conventional peak performance on DSP-based accelerators and sets a new benchmark for efficient neural network inference on FPGAs. Experimental results demonstrate that our design achieves the best inference speed among all FPGA-based accelerators, achieving a throughput of 1627 images per second and maintaining a top-1 accuracy of 70.95% on the ImageNet dataset. Yanyue Xie, Zhengang Li 0001, Dana Diaconu, Suranga Handagala, Miriam Leeser, Xue Lin 0001 |
ASP-DAC | 5 |
| 2025 | Transfer Learning on the Edge for a Wireless Application Using an SoC PlatformabstractEdge devices with limited resources are critical components of modern wireless communication systems. As communication environments become increasingly complex, neural networks are playing a larger role in processing large amounts of data to enable Machine Learning (ML) within these systems. While most FPGA-based accelerators focus on neural network inference, deploying the training phase on resource-constrained edge devices remains a significant challenge. Training on a System on Chip (SoC) that combines ARM processors with FPGA fabric provides unique benefits, including the ability to quickly adapt models to dynamic environments. This work leverages the Tiny Transfer Learning (TinyTL) framework for on-device training, which allows edge devices to continuously adapt neural network models to new data with minimal memory requirements. To the best of our knowledge, this is the first use of a heterogeneous platform to accelerate training using TinyTL. We present the Accelerating TinyTL-based Digital PreDistortion (ATDPD) system, designed to adapt to varying behaviors of power amplifiers in wireless communication systems and implement it on an AMD RFSoC. Our heterogeneous approach achieves comparable training accuracy to ARM-based systems, while accelerating the training phase by more than 20%. Yiyue Jiang, John Dooley, Aidan Edward Colgan, Jonathan Guimaraes Ribeiro, Zhilin Ren, Miriam Leeser |
FCCM | 6 |
| 2024 | Efficient Neural Networks on the Edge with FPGAs by Optimizing an Adaptive Activation FunctionabstractThe implementation of neural networks (NN) on edge devices enables local processing of wireless data but faces challenges such as high computational complexity and memory requirements when deep neural networks (DNN) are used. Shallow neural networks customized for specific problems are more efficient, requiring fewer resources, and resulting in a lower latency solution. An additional benefit of the smaller network size is that it is suitable for real-time processing on edge devices. The main concern with shallow neural networks is their accuracy performance compared to DNNs. In this paper, we demonstrate that a customized adaptive activation function (AAF) can meet the accuracy of a DNN. We designed an efficient FPGA implementation for a customized segmented spline curve neural network (SSCNN) structure to replace the traditional fixed activation function with an AAF. We compared our SSCNN with different neural network structures such as real-valued time delay neural network (RVTDNN), augmented real-valued time delay neural network (ARVTDNN), and deep neural networks with different parameters. Our proposed SSCNN implementation uses 40% fewer hardware resources and no Block RAMS compared to the DNN with similar accuracy. We experimentally validate this computationally efficient and memory-saving FPGA implementation of SSCNN for digital predistortion of RF power amplifiers using the AMD/Xilinx RFSoC ZCU111 while using less than 3% of the available resources, leaving space for additional real-time processing, while achieving the speed of 221 MHz. Yiyue Jiang, Andrius Vaicaitis, John Dooley, Miriam Leeser |
FPGA | 4 |
| 2023 | Neural Network on the Edge: Efficient and Low Cost FPGA Implementation of Digital Predistortion in MIMO SystemsabstractBase stations in cellular networks must operate linearly, power efficiently, and with ever increasing flexibility. Recent FPGA hardware advances have demonstrated linearization using neural networks, however the latency introduced by these solutions is a concern. We present a novel hardware implementation for a low digital cost, high throughput pipelined Real Valued Time Delay Neural Network (RVTDNN) structure with a hardware-efficient activation function. Network training times are reduced by minimizing the training signal samples used, based on a biased probability density function (pdf). The design has been experimen-tally validated using an AMD/Xilinx RFSoC ZCU216 board and surpasses the data throughput of conventional RVTDNN-based DPD while using a fraction of their hardware utilization. Yiyue Jiang, Andrius Vaicaitis, Miriam Leeser, John Dooley |
DATE | 3 |
| 2023 | Artifact Evaluation for ACM TRETS Papers Submitted from the FPT Journal TrackabstractAuthors of papers that were accepted to ACM TRETS via the FPT 2022 journal track had the option of participating in Artifact Evaluation (AE). Four papers from this track volunteered to participate in the AE process. All of these papers have been awarded badges from ACM as described below. Miriam Leeser |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2022 | FPGA-aware automatic acceleration framework for vision transformer with mixed-scheme quantization: late breaking resultsabstractVision transformers (ViTs) are emerging with significantly improved accuracy in computer vision tasks. However, their complex architecture and enormous computation/storage demand impose urgent needs for new hardware accelerator design methodology. This work proposes an FPGA-aware automatic ViT acceleration framework based on the proposed mixed-scheme quantization. To the best of our knowledge, this is the first FPGA-based ViT acceleration framework exploring model quantization. Compared with state-of-the-art ViT quantization work (algorithmic approach only without hardware acceleration), our quantization achieves 0.31% to 1.25% higher Top-1 accuracy under the same bit-width. Compared with the 32-bit floating-point baseline FPGA accelerator, our accelerator achieves around 5.6× improvement on the frame rate (i.e., 56.4 FPS vs. 10.0 FPS) with 0.83% accuracy drop for DeiT-base. Mengshu Sun, Zhengang Li 0001, Alec Lu, Geng Yuan, Yanyue Xie, Hao Tang 0005, Yanyu Li, Miriam Leeser, Zhangyang Wang, Xue Lin 0001, Zhenman Fang |
DAC | 9 |
| 2022 | Auto-ViT-Acc: An FPGA-Aware Automatic Acceleration Framework for Vision Transformer with Mixed-Scheme QuantizationabstractVision transformers (ViTs) are emerging with significantly improved accuracy in computer vision tasks. However, their complex architecture and enormous computation/storage demand impose urgent needs for new hardware accelerator design methodology. This work proposes an FPGA-aware automatic ViT acceleration framework based on the proposed mixed-scheme quantization. To the best of our knowledge, this is the first FPGA-based ViT acceleration framework exploring model quantization. Compared with state-of-the-art ViT quantization work (algorithmic approach only without hardware acceleration), our quantization achieves 0.47% to 1.36% higher Top-l accuracy under the same bit-width. Compared with the 32-bit floating-point baseline FPGA accelerator, our accelerator achieves around 5.6x improvement on the frame rate (i.e., 56.8 FPS vs. 10.0 FPS) with 0.71% accuracy drop on ImageNet dataset for DeiT-base. Zhengang Li 0001, Mengshu Sun, Alec Lu, Geng Yuan, Yanyue Xie, Hao Tang 0005, Yanyu Li, Miriam Leeser, Zhangyang Wang, Xue Lin 0001, Zhenman Fang |
FPL | 9 |
| 2022 | The Future of FPGA Acceleration in Datacenters and the CloudabstractIn this article, we survey existing academic and commercial efforts to provide Field-Programmable Gate Array (FPGA) acceleration in datacenters and the cloud. The goal is a critical review of existing systems and a discussion of their evolution from single workstations with PCI-attached FPGAs in the early days of reconfigurable computing to the integration of FPGA farms in large-scale computing infrastructures. From the lessons learned, we discuss the future of FPGAs in datacenters and the cloud and assess the challenges likely to be encountered along the way. The article explores current architectures and discusses scalability and abstractions supported by operating systems, middleware, and virtualization. Hardware and software security becomes critical when infrastructure is shared among tenants with disparate backgrounds. We review the vulnerabilities of current systems and possible attack scenarios and discuss mitigation strategies, some of which impact FPGA architecture and technology. The viability of these architectures for popular applications is reviewed, with a particular focus on deep learning and scientific computing. This work draws from workshop discussions, panel sessions including the participation of experts in the reconfigurable computing field, and private discussions among these experts. These interactions have harmonized the terminology, taxonomy, and the important topics covered in this manuscript. Christophe Bobda, Joel Mandebi, Paul Chow, Mohammad Ewais, Naif Tarafdar, Juan Camilo Vega, Kenneth Eguro, Dirk Koch, Suranga Handagala, Miriam Leeser, Martin C. Herbordt, Hafsah Shahzad, H. Peter Hofstee, Burkhard Ringlein, Jakub Szefer, Ahmed Sanaullah, Russell Tessier |
ACM Trans. Reconfigurable Technol. Syst. | 10 |
| 2021 | Evaluation of Optimized CNNs on Heterogeneous Accelerators Using a Novel Benchmarking ApproachabstractNumerous algorithmic optimization techniques have been proposed to alleviate the computational complexity of convolutional neural networks. Given the broad selection of AI accelerators, it is not obvious which approach benefits from which optimization most. The design space includes a large number of deployment settings (batch sizes, power modes, etc.) and unclear measurement methods. This research provides clarity into this design space, leveraging a novel benchmarking approach. We provide a theoretical evaluation of different CNNs and hardware platforms, focusing on understanding the impact of pruning and quantization as primary optimization techniques. We benchmark across a spectrum of FPGA, GPU, TPU, and VLIW processors for systematically pruned and quantized neural networks (ResNet50, GoogLeNetv1, MobileNetv1, a VGG derivative, a multilayer perceptron) over many deployment options, considering power, latency, and throughput at a specific accuracy. Our findings show that channel pruning is most effective and works across most hardware platforms, with speedups directly correlated to the reduction in compute load, while FPGAs benefit the most from quantization. Pruning and quantization are orthogonal, and yield optimal design points when combined. Further in-depth results can be found at our web portal, where we share all experimental data, provide data analytics, and invite the community to contribute. Michaela Blott, Nicholas J. Fraser, Giulio Gambardella, Lisa Halder, Johannes Kath, Zachary Neveu, Yaman Umuroglu, Alina Vasilciuc, Miriam Leeser, Linda Doyle |
IEEE Trans. Computers | 9 |
| 2020 | 3D CNN Acceleration on FPGA using Hardware-Aware PruningabstractThere have been many recent attempts to extend the successes of convolutional neural networks (CNNs) from 2-dimensional (2D) image classification to 3-dimensional (3D) video recognition by exploring 3D CNNs. Considering the emerging growth of mobile or Internet of Things (IoT) market, it is essential to investigate the deployment of 3D CNNs on edge devices. Previous works have implemented standard 3D CNNs (C3D) on hardware platforms, however, they have not exploited model compression for acceleration of inference. This work proposes a hardware-aware pruning approach that can fully adapt to the loop tiling technique of FPGA design and is applied onto a novel 3D network called R(2+1)D. Leveraging the powerful ADMM, the proposed pruning method achieves simultaneous high accuracy and significant acceleration of computation on FPGA. With layer-wise pruning rates up to 10× and negligible accuracy loss, the pruned model is implemented on a Xilinx ZCU102 FPGA board, where the pruned model achieves 2.6× speedup compared with the unpruned version, and 2.3× speedup and 2.3× power efficiency improvement compared with state-of-the-art FPGA implementation of C3D. Mengshu Sun, Pu Zhao 0001, Mehmet Güngör, Massoud Pedram, Miriam Leeser, Xue Lin 0001 |
DAC | 5 |
| 2020 | Evaluation of Optimized CNNs on FPGA and non-FPGA based Accelerators using a Novel Benchmarking ApproachabstractNumerous algorithmic optimization techniques have been proposed to alleviate the computational complexity of convolutional neural networks (CNNs). However, given the broad selection of inference accelerators, it is not obvious which approach benefits from which optimization and to what degree. In addition, the design space is further obscured by many deployment settings such as power and operating modes, batch sizes, as well as ill-defined measurement methodologies. In this paper, we systematically benchmark different types of CNNs leveraging both pruning and quantization as the most promising optimization techniques leveraging a novel benchmarking approach. We evaluate a spectrum of FPGA implementations, GPU, TPU and VLIW processor, for a selection of systematically pruned and quantized neural networks (including ResNet50, GoogleNetv1, MobileNetv1, a VGG derivative, and a multilayer perceptron) taking the full design space into account including batch sizes, thread counts, stream sizes and operating modes, and considering power, latency, and throughput at a specific accuracy as figure of merit. Our findings show that channel pruning is effective across most hardware platforms, with resulting speedups directly correlated to the reduction in compute load, while FPGAs benefit the most from quantization. FPGAs outperform regarding latency and latency variation for the majority of CNNs, in particular with feed-forward dataflow implementations. Finally, pruning and quantization are orthogonal techniques and yield the majority of all optimal design points when combined. With this benchmarking approach, both in terms of methodology and measured results, we aim to drive more clarity in the choice of CNN implementations and optimizations. Michaela Blott, Johannes Kath, Lisa Halder, Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Miriam Leeser, Linda Doyle |
FPGA | 7 |
| 2019 | QuTiBench: Benchmarking Neural Networks on Heterogeneous HardwareabstractNeural Networks have become one of the most successful universal machine-learning algorithms. They play a key role in enabling machine vision and speech recognition and are increasingly adopted in other application domains. Their computational complexity is enormous and comes along with equally challenging memory requirements in regards to capacity and access bandwidth, which limits deployment in particular within energy constrained, embedded environments. To address these implementation challenges, a broad spectrum of new customized and heterogeneous hardware architectures have emerged, often accompanied with co-designed algorithms to extract maximum benefit out of the hardware. Furthermore, numerous optimization techniques are being explored for neural networks to reduce compute and memory requirements while maintaining accuracy. This results in an abundance of algorithmic and architectural choices, some of which fit specific use cases better than others. For system-level designers, there is currently no good way to compare the variety of hardware, algorithm, and optimization options. While there are many benchmarking efforts in this field, they cover only subsections of the embedded design space. None of the existing benchmarks support essential algorithmic optimizations such as quantization, an important technique to stay on chip, or specialized heterogeneous hardware architectures. We propose a novel benchmark suite, QuTiBench , that addresses this need. QuTiBench is a novel multi-tiered benchmarking methodology ( Ti ) that supports algorithmic optimizations such as quantization ( Qu ) and helps system developers understand the benefits and limitations of these novel compute architectures in regard to specific neural networks and will help drive future innovation. We invite the community to contribute to QuTiBench to support the full spectrum of choices in implementing machine-learning systems. Michaela Blott, Lisa Halder, Miriam Leeser, Linda Doyle |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2018 | Cross Component Optimization for Modern LTE Downlink Shared Channel ImplementationabstractField-programmable Gate Arrays (FPGAs) are the preferred technology for implementing Software Defined Radios (SDRs). This is increasingly true for LTE as the pace of LTE releases is very fast, and flexibility as well as speed of implementation is required. We have modeled and implemented the whole LTE Downlink Shared Channel (DL-SCH) processing chain in FPGA hardware. We show that it is important to consider the whole DL-SCH processing chain's design when allocating data buffers. Prior work looked at each element in the processing chain in isolation, while we make use of cross-component optimization, which can save 20% or more of BRAMs. This savings helps us to keep up with the requirements of newer LTE releases. Jieming Xu, Miriam Leeser |
FCCM | 2 |
| 2018 | Digital Pre-distortion Implemented Using FPGAabstractMassive-MIMO and beamforming techniques have long been proposed as a means of increasing cellular network capacity and improving signal to interference ratio performance. The implementation of such systems requires a large number of signal transmission paths. To realize this, a distributed array of power amplifiers (PAs) is likely to be needed. These PAs will possess similar, but unique, characteristics which will alter over time independently due to temperature drift and component ageing. In order to operate all PAs in both a linear and efficient fashion a linearisation technique, such as Digital Pre-Distortion (DPD), must be used. DPD algorithms benefit from reconfigurability, low latency and power efficiency, all traits associated with Field Programmable Gate Arrays (FPGAs). This demonstration shows how an FPGA, specifically a ZYNQ System on a Chip (SoC), can be used in tandem with a transceiver board, the FMCOMMS2, to implement a DPD system. Declan Byrne, Ronan Farrell, Sidath Madhuwantha, Miriam Leeser, John Dooley |
FPL | 4 |
| 2018 | Local and Global Shared Memory for Task Based HPC Applications on Heterogeneous PlatformsabstractWith the prevalence of multicore and manycore processors, developing parallel applications to benefit from massively parallel resources is important. In this work, we introduce a hybrid shared memory mechanism based on a high-level task design. We implemented task scoped global shared data based on the one-sided communication feature of MPI-3 and enable users to implement and create multi-threaded tasks that can execute either on a single node or on multiple nodes. Task threads of distributed nodes can share data sets through global shared data objects using one-sided remote memory access. We ported and developed a set of benchmark applications and tested on a cluster platform. The high-level task design and hybrid shared memory help users develop and maintain parallel programs easily, and the results show that the global shared data can deliver good RMA performance; the multi-threaded task implementations perform up to 20% faster than ordinary OpenMP programs and have better scaling performance than MPI programs on multiple nodes. Chao Liu 0061, Miriam Leeser |
PDP | 2 |
| 2018 | FINN-R: An End-to-End Deep-Learning Framework for Fast Exploration of Quantized Neural NetworksabstractConvolutional Neural Networks have rapidly become the most successful machine-learning algorithm, enabling ubiquitous machine vision and intelligent decisions on even embedded computing systems. While the underlying arithmetic is structurally simple, compute and memory requirements are challenging. One of the promising opportunities is leveraging reduced-precision representations for inputs, activations, and model parameters. The resulting scalability in performance, power efficiency, and storage footprint provides interesting design compromises in exchange for a small reduction in accuracy. FPGAs are ideal for exploiting low-precision inference engines leveraging custom precisions to achieve the required numerical accuracy for a given application. In this article, we describe the second generation of the FINN framework, an end-to-end tool that enables design-space exploration and automates the creation of fully customized inference engines on FPGAs. Given a neural network description, the tool optimizes for given platforms, design targets, and a specific precision. We introduce formalizations of resource cost functions and performance predictions and elaborate on the optimization algorithms. Finally, we evaluate a selection of reduced precision neural networks ranging from CIFAR-10 classifiers to YOLO-based object detection on a range of platforms including PYNQ and AWS F1, demonstrating new unprecedented measured throughput at 50 TOp/s on AWS F1 and 5 TOp/s on embedded devices. Michaela Blott, Thomas B. Preußer, Nicholas J. Fraser, Giulio Gambardella, Kenneth O'Brien, Yaman Umuroglu, Miriam Leeser, Kees A. Vissers |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2017 | FIM: Performance Prediction for Parallel Computation in Iterative Data Processing ApplicationsabstractPredicting performance of an application running on high performance computing (HPC) platforms in a cloud environment is increasingly becoming important because of its influence on development time and resource management. However, predicting the performance with respect to parallel processes is complex for iterative, multi-stage applications. This research proposes a performance approximation approach FiM to model the computing performance of iterative, multi-stage applications running on a master-compute framework. FiM consists of two key components that are coupled with each other: 1) Stochastic Markov Model to capture non-deterministic runtime that often depends on parallel resources, e.g., number of processes. 2) Machine Learning Model that extrapolates the parameters for calibrating our Markov model when we have changes in application parameters such as dataset. Our new modeling approach considers different design choices along multiple dimensions, namely (i) process level parallelism, (ii) distribution of cores on multi-core processors in cloud computing, (iii) application related parameters, and (iv) characteristics of datasets. The major contribution of our prediction approach is that FiM is able to provide an accurate prediction of parallel computation time for the datasets which have much larger size than that of the training datasets. Such calculation prediction provides data analysts a useful insight of optimal configuration of parallel resources (e.g., number of processes and number of cores) and also helps system designers to investigate the impact of changes in application parameters on system performance. Janki Bhimani, Ningfang Mi, Miriam Leeser, Zhengyu Yang 0001 |
CLOUD | 3 |
| 2017 | Secure Function Evaluation Using an FPGA Overlay Architecture
Xin Fang 0001, Stratis Ioannidis, Miriam Leeser |
FPGA | 3 |
| 2017 | FPGA modeling techniques for detecting and demodulating multiple wireless protocolsabstractIn an increasingly interconnected world, the rising number of wireless devices in the Internet of Things has caused heavy congestion on particular bandwidths (BWs). Due to spectrum scarcity, the need has arisen for these devices to operate on the same BWs. However, existing wireless devices are inflexible and have no capabilities to coexist with devices using other protocols. In this work, we propose new FPGA-based design techniques to receive multiple protocols on the same computing platform. Our methods incorporate tunable parameters, such as FIR filter length and number of bits per fixed-point word, to explore design tradeoffs regarding clock cycle, resource utilization, power consumption, and detection accuracy. We separate the physical (PHY) layer receive chains into a set of building blocks, including rate transition, pattern detection, and OFDM demodulation. We investigate implementation of the LTE physical downlink shared channel (PDSCH) and 802.11a protocols to test our techniques. LTE is the standard for high-speed wireless communication for mobile phones and data terminals; Wi-Fi uses variants of IEEE 802.11a. To ease the system development process, we develop our models using MathWorks Simulink, which supports auto-generation of HDL code for the non-critical sections and incorporation of hand-tuned HDL code as part of its black box interface. Our building blocks can be used by the wireless system modeling community to meet the needs of modern evolving wireless standards. In the future, our framework will allow researchers to achieve high-performance transceiver implementations on FPGA fabric for multiple cutting edge protocols. Benjamin Drozdenko, Suranga Handagala, Kaushik R. Chowdhury, Miriam Leeser |
FPL | 4 |
| 2017 | Scaling Neural Network Performance through Customized Hardware Architectures on Reconfigurable LogicabstractConvolutional Neural Networks have dramatically improved in recent years, surpassing human accuracy on certain problems and performance exceeding that of traditional computer vision algorithms. While the compute pattern in itself is relatively simple, significant compute and memory challenges remain as CNNs may contain millions of floating-point parameters and require billions of floating-point operations to process a single image. These computational requirements, combined with storage footprints that exceed typical cache sizes, pose a significant performance and power challenge for modern compute architectures. One of the promising opportunities to scale performance and power efficiency is leveraging reduced precision representations for all activations and weights as this allows to scale compute capabilities, reduce weight and feature map buffering requirements as well as energy consumption. While a small reduction in accuracy is encountered, these Quantized Neural Networks have been shown to achieve state-of-the-art accuracy on standard benchmark datasets, such as MNIST, CIFAR-10, SVHN and even ImageNet, and thus provide highly attractive design trade-offs. Current research has focused mainly on the implementation of extreme variants with full binarization of weights and or activations, as well typically smaller input images. Within this paper, we investigate the scalability of dataflow architectures with respect to supporting various precisions for both weights and activations, larger image dimensions, and increasing numbers of feature map channels. Key contributions are a formalized approach to understanding the scalability of the existing hardware architecture with cost models and a performance prediction as a function of the target device size. We provide validating experimental results for an ImageNet classification on a server class platform, namely the AWS F1 node. Michaela Blott, Thomas B. Preußer, Nicholas J. Fraser, Giulio Gambardella, Kenneth O'Brien, Yaman Umuroglu, Miriam Leeser |
ICCD | 7 |
| 2016 | State-Action Based Link Layer Design for IEEE 802.11b Compliant MATLAB-Based SDRabstractSoftware defined radio (SDR) allows unprecedented levels of flexibility by transitioning the radio communication system from a rigid hardware platform to a more user-controlled software paradigm. However, it can still be time consuming to design and implement such SDRs as they typically require thorough knowledge of the operating environment and a careful tuning of the program. In this work, we describe a systems contribution and outline strategies on how to create a state-action based design in implementing the CSMA/CA/ACK MAC layer in MATLAB®that runs on the USRP®platform, a commonly used SDR. Our design allows optimal selection of the parameters so that all operations remain functionally compliant with the IEEE 802.11b standard (1Mbps specification). The code base of the system is enabled through the Communications System ToolboxTMand incorporates channel sensing and exponential random back-off for contention resolution. The current work provides a testbed to experiment with and enables creation of new MAC protocols starting from the fundamental IEEE 802.11b compliant standard. Our system design approach guarantees the consistent performance of the bi-directional link and we include the experimental results for the three node system to demonstrate the robustness of the MAC layer in mitigating packet collisions and enforcing fairness among nodes. Subramanian Ramanathan, Eric Doyle, Benjamin Drozdenko, Miriam Leeser, Kaushik R. Chowdhury |
DCOSS | 4 |
| 2016 | Modeling considerations for the hardware-software co-design of flexible modern wireless transceiversabstractSoftware-defined radios have introduced new platforms for dynamically modifying wireless system designs, and heterogeneous computing has opened up implementing such designs on different computing elements. Our goal is to develop a hardware-software modeling environment that captures reusability of various processing blocks at the physical layer for several modern protocols, and makes decisions regarding whether each processing block should be part of reconfigurable hardware or embedded processor software based on timing constraints and power budgets for the overlying applications. Our approach creates several different MathWorks Simulink model variants for both the transmitter and the receiver, each with a different boundary between hardware and software components. Using the 802.11a standard as an example, we use these models to generate a bitstream for the FPGA and executable code for the ARM processor on a Xilinx Zynq system-on-chip. Our results collect such metrics as data path delay, resource utilization, and power usage and demonstrate how to enhance the SDR design. Benjamin Drozdenko, Matthew Zimmermann, Tuan Dao, Kaushik R. Chowdhury, Miriam Leeser |
FPL | 5 |
| 2016 | Performance prediction techniques for scalable large data processing in distributed MPI systemsabstractPredicting performance of an application running on parallel computing platforms is increasingly becoming important due to the long development time of an application and the high resource management cost of parallel computing platforms. However, predicting overall performance is complex and must take into account both parallel calculation time and communication time. Difficulty in accurate performance modeling is compounded by myriad design choices along multiple dimensions, namely (i) process level parallelism, (ii) distribution of cores on multi-processor platforms, (iii) application related parameters, and (iv) characteristics of datasets. This research proposes a fast and accurate performance prediction approach to predict the calculation and communication time of an application running on a distributed computing platform. The major contribution of our prediction approach is that it can provide an accurate prediction of execution times for new datasets which have much larger sizes than the training datasets. Our approach consists of two models, i.e., a probabilistic self-learning model to predict calculation time and a simulation queuing model to predict network communication time. The combination of these two models provides data analysts a useful insight of optimal configuration of parallel resources (e.g., number of processes and number of cores) and application parameters setting. Janki Bhimani, Ningfang Mi, Miriam Leeser |
IPCCC | 3 |
| 2016 | Open-Source Variable-Precision Floating-Point Library for Major Commercial FPGAsabstractThere is increased interest in implementing floating-point designs for different precisions that take advantage of the flexibility offered by Field-Programmable Gate Arrays (FPGAs). In this article, we present updates to the Variable-precision FLOATing Point Library (VFLOAT) developed at Northeastern University and highlight recent improvements in implementations for implementing reciprocal, division, and square root components that scale to double precision for FPGAs from the two major vendors: Altera and Xilinx. Our library is open source and flexible and provides the user with many options. A designer has many tradeoffs to consider including clock frequency, total latency, and resource usage as well as target architecture. We compare the generated cores to those produced by each vendor and to another popular open-source tool: FloPoCo. VFLOAT has the advantage of not tying the user’s design to a specific target architecture and of providing the maximum flexibility for all options including clock frequency and latency compared to other alternatives. Our results show that variable-precision as well as double-precision designs can easily be accommodated and the resulting components are competitive and in many cases superior to the alternatives. Xin Fang 0001, Miriam Leeser |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2015 | Balance power leakage to fight against side-channel analysis at gate level in FPGAsabstractSide-channel attacks have been a serious threat to the security of embedded cryptographic systems, and various countermeasures have been devised to mitigate the leakages. Power balance technologies such as wave dynamic differential logic (WDDL) aim to balance the power by introducing differential logic. However, different routing length leads to different capacitance of wire, and this hampers the strength of the power balance countermeasure. In this paper, we further balance the power of differential signals by manipulating the lower level primitives and placement constraints on a Field Programmable Gate Array (FPGA). We choose Advanced Encryption Standard (AES) as the encryption algorithm and apply Hamming weight model to demonstrate the amount of leakage for different implementations. Results show that our method not only efficiently mitigates the side-channel leakage but also saves FPGA logic block resources and dynamic power consumption. Xin Fang 0001, Pei Luo, Yunsi Fei, Miriam Leeser |
ASAP | 4 |
| 2015 | Behavioral Non-portability in Scientific Numeric Computing
Yijia Gu, Thomas Wahl, Mahsa Bayati, Miriam Leeser |
Euro-Par | 4 |
| 2015 | Kernel Specialization Provides Adaptable GPU Code for Particle Image VelocimetryabstractGraphics Processing Units (GPUs) are increasingly used to accelerate scientific applications. The state-of-the-art limits the adaptability of GPU kernels to both problem parameters and hardware characteristics. This makes writing high performance libraries for GPUs challenging. We address these challenges through Kernel Specialization (KS) which supports both user and hardware parameters and produces highly optimized GPU code. We apply KS to Particle Image Velocimetry (PIV), a technique used to obtain instantaneous velocity measurements in fluids for such diverse applications as aircraft design and artificial heart design. KS helps the user search PIV’s highly non-linear design space, supports a wide range of PIV parameters, and results in improved acceleration times over existing kernels. Nicholas Moore, Miriam Leeser, Laurie A. Smith King |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | Make it real: Effective floating-point reasoning via exact arithmeticabstractFloating-point arithmetic is widely used in scientific computing. While many programmers are subliminally aware that floating-point numbers only approximate the reals, few are cognizant of the dangers this entails for programming. Such dangers range from tolerable rounding errors in sequential programs, to unexpected, divergent control flow in parallel code. To address these problems, we present a decision procedure for floating-point arithmetic (FPA) that exploits the proximity to real arithmetic (RA), via a loss-less reduction from FPA to RA. Our procedure does not involve any form of bit-blasting or bit-vectorization, and can thus generate much smaller back-end decision problems, albeit in a more complex logic. This tradeoff is beneficial for the exact and reliable analysis of parallel scientific software, which tends to give rise to large but benignly structured formulas. We have implemented a prototype decision engine and present encouraging results analyzing such software for numerical accuracy. Miriam Leeser, Saoni Mukherjee, Jaideep Ramachandran, Thomas Wahl |
DATE | 1 |
| 2014 | Reducing Processing Latency with a Heterogeneous FPGA-Processor FrameworkabstractBoth Xilinx and Altera have released SoCs that tightly couple programmable logic with a dual core Cortex A9 ARM processor. These SoCs show promise in accelerating applications that exploit both the FPGA's parallel processing architecture and the CPU's sequential processing. For example, before accessing a wireless channel, a cognitive radio does spectrum sensing to detect channel occupancy and then makes a decision based on spectrum policies. Spectrum sensing maps well to FPGA fabric, while spectrum decision can be implemented with a CPU. Both algorithms are highly sensitive to latency as a faster decision improves spectrum utilization. This paper introduces CRASH: Cognitive Radio Accelerated with Software and Hardware - a new software and programmable logic framework for Xilinx's Zynq SoC targeting cognitive radio. We implement spectrum sensing and the spectrum decision in three configurations: both algorithms in the FPGA, both in software only, and spectrum sensing on the FPGA and spectrum decision on the CPU. We measure the end-to-end latency to detect and acquire unoccupied spectrum for these configurations. Results show that CRASH can successfully partition algorithms between FPGA and CPU and reduce processing latency. Jonathon Pendlum, Miriam Leeser, Kaushik R. Chowdhury |
FCCM | 2 |
| 2013 | Minimum energy operation for clustered island-style FPGAsabstractDespite the advantages offered by field-programmable gate arrays (FPGAs) for low-power systems requiring flexible computing resources, applications with the lowest power budgets still favor microprocessors and application-specific integrated circuits (ASICs). In order for such systems to exploit FPGAs, an FPGA achieving minimum energy operation is needed. Minimum energy points have been found for ASICs and microprocessors to occur at operating voltages that are typically below the transistor threshold voltage. This paper presents two clustered island-style test chips capable of operating with a single supply voltage as low as 260 mV. This supply voltage represents the lowest voltage at which an FPGA has been successfully programmed. Test chip measurements show that the minimum energy point of both circuits is at or below this minimum operating voltage. Operation at 260 mV leads to a 40X power-delay product reduction vs. 1.5V operation. The results demonstrate a clear path forward for fabricating low voltage FPGAs that are fully compatible with existing tool flows. Peter Grossmann, Miriam Leeser, Marvin Onabajo |
FPGA | 2 |
| 2013 | Kernel Specialization for Improved Adaptability and Performance on Graphics Processing Units (GPUs)abstractGraphics processing units (GPUs) offer significant speedups over CPUs for certain classes of applications. However, programming for GPUs is challenging. There are many parameters that affect performance and their values may change depending on both problem instance and GPU hardware specifics. In addition, most GPU kernels are compiled once; performance optimizations are applied at application compile time. As a result, many GPU libraries and programs have limited adaptability to variations among problem instances and hardware configurations. These factors limit code reuse and the applicability of GPU computing to a wider variety of problems. This paper introduces GPGPU kernel specialization, a technique used to describe highly adaptable kernels that exhibit high performance across a wide range of programmer variables as well as different generations of GPUs. We also introduce our GPU Prototyping Framework (GPU-PF) for dynamic runtime generation of customized GPU kernels incorporating both problem and implementation-specific parameters. GPU-PF fully separates the GPU and CPU code so the GPU code can be compiled during program execution once all the parameters are known. This work explores the implementation and parameterization of two real world applications targeting two generations of NVIDIA CUDA-enabled GPUs using kernel specialization and GPU-PF: large template matching and cone-beam image reconstruction via backprojection. Starting with high performance GPU kernels that compare favorably to multi-threaded reference implementations, kernel specialization is shown to increase adaptability while providing performance improvements including improved run time and reduction in resource usage. Kernel specialization offers productivity benefits, improved library code, and a means to increase the parameterizability of GPGPU implementations. Nicholas Moore, Miriam Leeser, Laurie A. Smith King |
IPDPS | 2 |
| 2012 | Cognitive Radio Universal Software HardwareabstractCognitive radio (CR) allows a pair of wireless communicating devices to intelligently adapt their transmission parameters based on the characteristics of the radio frequency (RF) environment, possibly using vacant portions of licensed frequency bands. A key step in this process is spectrum sensing, where the presence of primary users of the spectrum is detected by sending raw samples from the front-end to the host computer, a process that incurs significant processing delay. In this paper we introduce the Cognitive Radio Universal Software Hardware (CRUSH) platform that revisits the classical software defined radio (SDR) architecture by including a Xilinx ML605 FPGA board, thereby flexibly sharing the processing load with the host computer. CRUSH shifts the spectrum sensing overhead closer to the front-end. Our design results in orders of magnitude speedup and can flexibly connect up to three SDRs together, resulting in a powerful platform suitable for time-sensitive CR functions. George Eichinger, Kaushik R. Chowdhury, Miriam Leeser |
FCCM | 3 |
| 2012 | Implementing Murf: Accelerating Large State Space Exploration on FPGAsabstractPHAST, a Pipelined Hardware Accelerated State Checker, achieves a 30x end-to-end speedup of a large state space exploration application in the form of an explicit state model checker. PHAST is a re-implementation, to accommodate FPGA hardware, of the Murphi verifier developed at Stanford University. Explicit state model checking explores a large state space and checks properties defined by the user. The FPGA infrastructure for PHAST can be reused for many different models and properties. Our model of the DASH protocol is similar in size and complexity to models Intel uses to validate proposed features of future processors: state sizes between 1200 and 1800 bits and a transition relation with more than 100 rules. Mary Ellen Tie, Miriam Leeser |
FCCM | 2 |
| 2012 | Incremental clustering applied to radar deinterleaving: a parameterized FPGA implementationabstractICED (Incremental Clustering of Evolving Data) is a novel incremental clustering algorithm designed for data whose characteristics change over time. ICED is an unsupervised clustering technique that assumes no prior knowledge of the incoming data, and supports removing clusters that contain stale data. The user controls the FPGA implementation through a combination of compile time parameters (number of clusters) and run time parameters (distance threshold, fade cycle length). ICED has been applied to a radar application: pulse deinterleaving. ICED is the first implementation of incremental clustering on an FPGA of which we are aware. The implementation runs 39 times faster than an equivalent C implementation on a 3GHz Intel Xeon processor, and is capable of processing radar data in real time. Scott Bailie, Miriam Leeser |
FPGA | 2 |
| 2012 | CRUSH: Cognitive Radio Universal Software HardwareabstractThe FPGA is an integral component of a software defined radio (SDR) that provides the needed reconfigurability for dynamically adapting its transceiver and data processing functions. Because of the desire to process data faster and with less latency, researchers are looking at FPGA-based SDR. Our architecture, called CRUSH, is composed of a Xilinx ML605 connected to an Ettus USRP through a a custom interface board allowing flexible data transfer between them. In addition, we provide a framework that supports ease of use, independent programming on both devices, and integration with software running on the host. To demonstrate our platform we implemented spectrum sensing, a key step in determining channel availability before transmission in dynamic spectrum access networks. Spectrum sensing is implemented on CRUSH using FFTs for a 100× speedup; the complete sensing cycle is 10× faster than the same design without CRUSH. By reducing the load of transferring raw samples to the host and allowing a powerful FPGA extension for off-the-shelf devices, CRUSH enables advances in both protocol design and reconfigurable hardware targeting radio applications. George Eichinger, Kaushik R. Chowdhury, Miriam Leeser |
FPL | 3 |
| 2012 | VForce: An environment for portable applications on high performance systems with accelerators
Nicholas Moore, Miriam Leeser, Laurie A. Smith King |
J. Parallel Distributed Comput. | 2 |
| 2011 | An Autonomous Vector/Scalar Floating Point Coprocessor for FPGAsabstractWe present a Floating Point Vector Coprocessor that works with the Xilinx embedded processors. The FPVC is completely autonomous from the embedded processor, exploiting parallelism and exhibiting greater speedup than alternative vector processors. The FPVC supports scalar computation so that loops can be executed independently of the main embedded processor. Floating point addition, multiplication, division and square root are implemented with the Northeastern University VFLOAT library. The FPVC is parameterized so that the number of vector lanes and maximum vector length can be easily modified. We have implemented the FPVC on a Xilinx Virtex 5 connected via the Processor Local Bus (PLB) to the embedded PowerPC. Our results show more than five times improved performance over the PowerPC augmented with the Xilinx Floating Point Unit on applications from linear algebra: QR and Cholesky decomposition. Jainik Kathiara, Miriam Leeser |
FCCM | 2 |
| 2011 | A prototype FPGA for subthreshold-optimized CMOS (abstract only)abstractField-programmable gate arrays (FPGAs) are frequently used in low power systems because they can implement the same functionality as a microprocessor in a more energy-efficient manner while still offering the benefits of low development time and cost relative to an ASIC. Similar advantages for FPGAs may be found in ultra-low power applications operating at subthreshold supply voltages, where performance is sacrificed in favor of increased energy efficiency. Process technology research has demonstrated the benefits of tailoring device design to subthreshold operation. Subthreshold FPGA research is only beginning, and has yet to consider use of subthreshold-optimized devices. A simplified FPGA implemented in subthreshold-optimized CMOS is presented. The results obtained show that this technology provides the capability to implement an FPGA suitable for ultra-low power applications consuming tens to hundreds of microwatts of average power. Peter Grossmann, Miriam Leeser |
FPGA | 2 |
| 2010 | VFloat: A Variable Precision Fixed- and Floating-Point Library for Reconfigurable HardwareabstractOptimal reconfigurable hardware implementations may require the use of arbitrary floating-point formats that do not necessarily conform to IEEE specified sizes. We present a variable precision floating-point library (VFloat) that supports general floating-point formats including IEEE standard formats. Most previously published floating-point formats for use with reconfigurable hardware are subsets of our format. Custom datapaths with optimal bitwidths for each operation can be built using the variable precision hardware modules in the VFloat library, enabling a higher level of parallelism. The VFloat library includes three types of hardware modules for format control, arithmetic operations, and conversions between fixed-point and floating-point formats. The format conversions allow for hybrid fixed- and floating-point operations in a single design. This gives the designer control over a large number of design possibilities including format as well as number range within the same application. In this article, we give an overview of the components in the VFloat library and demonstrate their use in an implementation of the K-means clustering algorithm applied to multispectral satellite images. Miriam Leeser |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2009 | Implementing a Highly Parameterized Digital PIV System on Reconfigurable HardwareabstractThis paper presents PARPIV the design and prototyping of a highly parameterized digital Particle Image Velocimetry (PIV) system implemented on reconfigurable hardware. Despite many improvements to PIV methods over the last twenty years, PIV post-processing remains a computationally intensive task. It becomes a serious bottleneck as camera acquisition rates reach 1000 frames per second. In this research, we aim to substantially speed up PIV processing by implementing it in reconfigurable hardware. Furthermore, this implementation is highly parameterized, supporting adaptation to a variety of setups and application domains. The circuit is parameterized by the dimensions of the captured images as well as the dimensions of the interrogation windows and sub-areas, pixel representation, board memory width, displacement and overlap. Through this work a parameterized library of different VHDL components was built. To the best of the authorspsila knowledge, this is the first highly parameterized PIV system implemented on reconfigurable hardware reported in the literature. For a typical PIV configuration with images of 512times512 pixels, 40times40 pixel interrogation windows and 32times32 pixel sub-areas, we achieved about 65\ times speedup in hardware over a standard software implementation. Abderrahmane Bennis, Miriam Leeser, Gilead Tadmor |
ASAP | 2 |
| 2009 | A truly two-dimensional systolic array FPGA implementation of QR decompositionabstractWe have implemented a two-dimensional systolic array QR decomposition on a Xilinx Virtex5 FPGA using the Givens rotation algorithm. QR decomposition is a key step in many DSP applications including sonar beamforming, channel equalization, and 3G wireless communication. Compared to previous work that implements Givens rotations using a one-dimensional systolic array, our implementation uses a truly two-dimensional systolic array architecture. As a result, latency scales well for larger matrices. In addition, prior work avoids divide and square root operations in the Givens rotation algorithm by using special operations such as CORDIC or special number systems such as the logarithmic number system (LNS). In contrast, our design uses straightforward floating-point divide and square root implementations, which makes it easier to be used within a larger system. In our design, the input matrix size can be configured at compile time to many different sizes, making it easily scalable to future large FPGAs or over multiple FPGAs. The QR module is fully pipelined with a throughput of over 130MHz for the IEEE single-precision floating-point format. The peak performance for a 12 × 12 input matrix is approximately 35 GFLOPs. Miriam Leeser |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2008 | An efficient implementation of a phase unwrapping kernel on reconfigurable hardwareabstractThe optical quadrature method of microscopy (OQM) uses phase data to capture information about the sample being studied. This phase data need to be unwrapped before it can be of use. Phase unwrapping is the process by which an integer multiple of 2pi is added to a measured, wrapped phase value in order to generate a continuous function. The algorithm used is the minimum LPnorm method which uses a two dimensional discrete cosine transform (2-D DCT) which forms the most computationally expensive part of the minimum LPnorm method. This paper presents an implementation on reconfigurable hardware that performs the 2-D DCT over the entire 1024 times 512 image, solves the intermediate equation and then performs the two dimensional Inverse discrete cosine transform (2-D IDCT) using a novel FPGA implementation of the DCT with a semi-floating point data representation. This represents the largest 2-D DCT FPGA implementation in the literature, with most previous work focusing on the 8 times 8 transform. Sherman Braganza, Miriam Leeser |
ASAP | 2 |
| 2008 | An Efficient Implementation of a Phase Unwrapping Kernel on Reconfigurable HardwareabstractThe optical quadrature method of microscopy (OQM) was developed at Northeastern University for the purpose of non-invasively capturing phase data to image the sample being studied. This phase data need to be unwrapped before it can be of use. Phase unwrapping is the process by which an integer multiple of 2¿ is added to a measured, wrapped phase value in order to generate a continuous function. The algorithm used is the minimum LPnorm method which uses a two dimensional discrete cosine transform (2-D DCT) to solve the discrete Poisson equation. This calculation forms the most computationally expensive part of the minimum LPnorm method. This paper presents an implementation on reconfigurable hardware that performs the 2-D DCT over the entire image, solves the Poisson equation and then performs the two dimensional inverse discrete cosine transform (2-D IDCT) using a novel FPGA implementation of the DCT with a semi-floating point data representation. Sherman Braganza, Miriam Leeser |
FCCM | 2 |
| 2008 | An FPGA Implementation of Explicit-State Model CheckingabstractWe present PHAST, a pipelined hardware accelerated explicit-state model checker. The algorithms and methodologies used to perform the state checking in PHAST are based on the Mur¿ verifier, developed at Stanford University. Mur¿ has been used to verify hardware and protocols, cache coherency protocols in particular. Mur¿ is used in industry due to its success in finding errors in real designs. Until now, Mur¿ and other model checkers have been solely performed in software. PHAST takes advantage of the flexible memory architecture on FPGA chips and the inherent concurrency available in hardware designs to accelerate model checking. Using PHAST in simulation, we achieved over 200× application speedup over Mur¿ on an example of a counter. Mary Ellen Fuess, Miriam Leeser, Tim Leonard |
FCCM | 2 |
| 2008 | Efficient FPGA implementation of qr decomposition using a systolic array architectureabstractQR decomposition is used in many signal processing applications. We have implemented a systolic array QR decomposition on a Xilinx Virtex5 FPGA using the Givens rotation algorithm. It uses a truly two dimensional systolic array architecture so latency scales well for large matrices. To accommodate the dynamic range of input data, floating-point arithmetic is chosen, using the Northeastern University Variable Precision Floating-Point (VFloat) library. We support any general floating-point format including IEEE single precision. Our design uses straightforward floating-point divide and square root implementations, compared to prior work which uses special operations or formats such as CORDIC or the logarithmic number system (LNS). This makes our design more standard and portable to different systems, thus easier to fit into a larger design. We support square, tall and short matrices. The input matrix size can be configured at compile-time to virtually any size. Therefore, it can be easily scaled to future larger FPGA devices, or over multiple FPGAs. The QR module is fully pipelined with a throughput of over 130 MHz for IEEE single precision floating-point format. 35 GFlops throughput peak performance is achieved for a 12 by 12 matrix with this implementation Miriam Leeser |
FPGA | 2 |
| 2008 | Special issue: General-purpose processing using graphics processing units
David R. Kaeli, Miriam Leeser |
J. Parallel Distributed Comput. | 2 |
| 2008 | Acknowledgment to special issue reviewers
David R. Kaeli, Miriam Leeser |
J. Parallel Distributed Comput. | 2 |
| 2008 | Efficient Communication Between the Embedded Processor and the Reconfigurable Logic on an FPGAabstractIncreasing device densities have prompted FPGA manufacturers, such as Xilinx and Altera, to incorporate larger embedded components, including multipliers, DSP blocks and even embedded processors. One of the recent architectural enhancements in the Xilinx Virtex family architecture is the introduction of the PowerPC405 hard-core embedded processor. In this paper we present a software defined radio application that serves as a vehicle for investigating effective communication between the PowerPC405 processor and the surrounding FPGA fabric. Joshua Noseworthy, Miriam Leeser |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2007 | The 1D Discrete Cosine Transform For Large Point Sizes Implemented On Reconfigurable HardwareabstractThe discrete cosine transform (DCT) is used in place of the discrete Fourier transform (DFT) in a wide variety of audio and image processing applications due to its energy compaction properties which approach those of the optimal Karhunen-Love transform. Previous work in reconfigurable hardware has focused on implementations of 2D 8times8 transforms of the type commonly used in the JPEG and MPEG standards. Several applications for larger DCTs exist, including those involving the extraction of features from image data, solving partial differential equations (PDEs) and those utilizing the preconditioned conjugate gradient (PCG) method such as phase-unwrapping. This paper presents an indirect algorithm and implementation on a Xilinx FPGA that performs 1D DCTs on large block sizes using a block floating point format. The DCT was designed to use fewer resources than other popular approaches due to the larger point sizes supported which would otherwise consume all available chip area, but at the cost of higher latency. This latency is similar to that required for an identically sized FFT A 512-point DCT has been shown to take 1771 cycles or 13.3 us at 133 MHz as compared to a similarly sized FFT that takes 1757 cycles or 13.2 us (including all component transfer times). Sherman Braganza, Miriam Leeser |
ASAP | 2 |
| 2007 | Writing Portable Applications that Dynamically Bind at Run Time to Reconfigurable HardwareabstractPowerful multicomputer platforms that combine FPGAs and programmable processors promise tremendous performance benefits for applications that take advantage of these rapidly emerging architectures. Portable applications are desirable because they can be easily adapted to take advantage of different reconfigurable computing platforms. raditional practices, however, intertwine application code with hardware specific code such that porting entails a significant rewrite of the application and reuse is difficult. Vforce, based on the VSIPL++ standard, is an exten sible framework we created that allows the same application code to run on different reconfigurable computing platforms. Vforce offers application-level portability, framework-level extensibility to new hardware, and system-level run time resource management. In particular, Vforce supports very late binding of the application to a specific hardware platform such that binding does not occur until run time. This paper describes Vforce with a focus on late run time binding to a specific hardware platform. Results using Vforce to implement an FFT and a time domain adaptive beam- former are presented. Nicholas Moore, Albert Conti, Miriam Leeser, Laurie A. Smith King |
FCCM | 3 |
| 2007 | K-means Clustering for Multispectral Images Using Floating-Point DivideabstractMany signal processing algorithms can be accelerated using reconfigurable hardware. To achieve a good speedup compared to running software on a general purpose processor, fine-grained control over the bitwidth of each component in the datapath is desired. This goal can be achieved by using NU's variable precision floating-point library. To analyze the usefulness of the floating-point divide unit, we incorporate it into our previous implementation of the Kmeans clustering algorithm applied to multispectral satellite images. With the lack of a floating-point divide hardware implementation, the mean updating step in each iteration of the K-means algorithm had to be moved to the host computer for calculation. The new means calculated on the host then had to be moved back to the FPGA board for the next iteration of the algorithm. This added data transfer overhead between the host and the FPGA board. In this work, we use the new fp div module to implement the mean updating step in FPGA hardware. This greatly reduces the communication overhead between host and FPGA board and further accelerates run time. The Kmeans clustering example illustrates the use of the fp div, fix2float and float2fix modules seamlessly assembled together in a real application. It is the first implementation that has the complete K-means computation done in FPGA hardware. Our results show that the hardware implementation achieves a speedup of over 2150x for core computation time and about 11x for total run time including data transfer time. They also show that the divide in FPGA hardware is 100 times faster than in software. Moreover, implementing divide in the FPGA frees the host to work on other tasks concurrently with K-means clustering, thus providing further speedup by allowing the image analyst to exploit this coarse grained parallelism. Miriam Leeser |
FCCM | 2 |
| 2006 | Advanced Components in the Variable Precision Floating-Point LibraryabstractOptimal reconfigurable hardware implementations may require the use of arbitrary floating-point formats that do not necessarily conform to IEEE specified sizes. The authors have previously presented a variable precision floating-point library for use with reconfigurable hardware. The authors recently added three advanced components: floating-point division, floating-point square root and floating-point accumulation to our library. These advanced components use algorithms that are well suited to FPGA implementations and exhibit a good tradeoff between area, latency and throughput. The floating-point format of our library is both general and flexible. All IEEE formats, including 64-bit double-precision format, are a subset of our format. All previously published floating-point formats for reconfigurable hardware are a subset of our format as well. The generic floating-point format supported by all of our library components makes it easy and convenient to create a pipelined, custom data path with optimal bitwidth for each operation. Our library can be used to achieve more parallelism and less power dissipation than adhering to a standard format. To further increase parallelism and reduce power dissipation, our library also supports hybrid fixed and floating point operations in the same design. The division and square root designs are based on table lookup and Taylor series expansion, and make use of memories and multipliers embedded on the FPGA chip. The iterative accumulator utilizes the library addition module as well as buffering and control logic to achieve performance similar to that of the addition by itself. They are all fully pipelined designs with clock speed comparable to that of other library components to aid the designer in implementing fast, complex, pipelined designs Sherman Braganza, Miriam Leeser |
FCCM | 3 |
| 2006 | Automatic Sliding Window Operation Optimization for FPGA-BasedabstractFPGA-based computing boards are frequently used as hardware accelerators for image processing algorithms based on sliding window operations (SWOs). SWOs are both computationally intensive and data intensive and benefit from hardware acceleration with FPGAs, especially for delay sensitive applications. The current design process requires that, for each specific application using SWOs with different size of window, image, etc.; a detail design must be completed before a realistic estimate of the achievable speedup can be obtained. We present an automated tool, sliding window operation optimization (SWOOP), that generates the estimate of speedup for a high performance design before detailed implementation is complete. The achievable speedup is determined by the area of the FPGA, or, more often, the memory bandwidth to the processing elements. The memory bandwidth to each processing element is a combination of bandwidth to the FPGA and the efficient use of on-chip RAM as a data cache. SWOOP uses analytic techniques to automatically determine the number of parallel processing elements to implement on the FPGA, the assignment of input and output data to on-board memory, and the organization of data in on-chip memory to most effectively keep the processing elements busy. The result is a block layout of the final design, its memory architecture, and a measure of the achievable speedup. The results, compared to manual designs, show that the estimates obtained usinq SWOOP are very accurate Haiqian Yu, Miriam Leeser |
FCCM | 2 |
| 2006 | Efficient use of communications between an FPGA's embedded processor and its reconfigurable logicabstractAbstract — Increasing device densities allow chip manufacturers to integrate more functionality onto a single piece of silicon. FPGA manufacturers, such as Xilinx and Altera, use these additional resources to further diversify and improve the processing capabilities of their architectures. One of the more recent architectual enhancements that has been made to Xilinx’s Virtex family architecture is the introduction of the PowerPC405 hard-core embedded processor. In this paper we present a Software Defined Radio Application that serves as a vehicle for investigating effective communications between the PowerPC405 Processor and the surrounding FPGA fabric.. A challenging aspect of developing applications that target the PowerPC is the interfacing of the processor with the surrounding reconfigurable logic. We have implemented a dozen different versions of a Software Defined Radio (SDR) application to exercise the various interfaces that enable communication between the processor and the surrounding FPGA fabric. The implementations differ only in the interfaces used. Our results indicate that the performance of the SDR application application can be increased by as much as 60 percent just by choosing the interfaces that are most appropriate for the different types of data in the implementation. This demonstrates that the performance of FPGA applications that use the embedded processor are dramatically effected by the mechanisms chosen to enable communication between the processor and its surrounding resources. I. Joshua Noseworthy, Miriam Leeser |
FPGA | 2 |
| 2006 | Poster reception - Improving the performance of parallel backprojection on a reconfigurable supercomputerabstractReconfigurable supercomputing is a new direction of research in the quest for novel high-performance architectures. A handful of products are available in this field, but no architecture has emerged as the clear winner in performance. Using the US Air Force Research Labs' Heterogeneous High Performance Cluster (HHPC), we have developed a successful reconfigurable supercomputing implementation of synthetic aperture radar (SAR) image formation using backprojection, a highly parallel algorithm. By studying the architectural features of the HHPC and hand-tuning the application to those features, we estimate that better than 1000x speedup over a single-node software-only solution is feasable. In particular, we consider the movement of data between nodes, the movement of data between host PC and FPGA board, and the use of on-board memories available to the FPGA for data storage. These results indicate that reconfigurable supercomputers are an excellent choice to tackle highly-parallel large-dataset problems. Ben Cordes, Miriam Leeser, Eric L. Miller 0001, Richard W. Linderman |
SC | 2 |
| 2006 | Enabling MPEG-2 video playback in embedded systems through improved data cache efficiencyabstractDigital video decoding, enabled by the MPEG-2 video standard, is an important future application for embedded systems, particularly personal digital assistants and other information appliances. Many such systems require portability and wireless communication capabilities, and thus face severe limitations in size and power consumption. This places a premium on integration and efficiency, and favors software solutions for video functionality over specialized hardware. Apart from computation, an equally important problem in video decoding is the data bandwidth and the need to insure adequate data supply. MPEG data sets are very large, and generate significant amounts of excess memory traffic for standard data caches, up to 100 times the amount required for decoding. Yet MPEG data has locality which caches can exploit if properly optimized, providing fast, flexible, and automatic data supply. We propose a set of enhancements which target the specific needs of the heterogeneous types within the MPEG decoder working set. These optimizations significantly improve the efficiency of small caches, reducing cache-memory traffic by almost 70%, and can make an enhanced 4-kB cache perform better than a standard 1 MB cache. This performance improvement can enable high-resolution, full frame rate video playback in cheaper, smaller systems than would otherwise be possible. Peter Soderquist, Miriam Leeser, Juan Carlos Rojas |
IEEE Trans. Multim. | 2 |
| 2005 | Enabling a RealTime Solution for Neuron Detection with Reconfigurable Hardware (abstract only)abstractFPGAs provide a speed advantage in processing for embedded systems, especially when processing is moved close to the sensors. Perhaps the ultimate embedded system is a neural prosthetic, where probes are inserted into the brain and recorded electrical activity is analyzed to determine which neurons have fired. In turn, this information can be used to manipulate an external device such as a robot arm or a computer mouse. To make the detection of these signals possible, some baseline data must be processed to correlate impulses to particular neurons. One method for processing this data uses a statistical clustering algorithm called Expectation Maximization, or EM. In this paper, we examine the EM clustering algorithm, determine the most computationally intensive portion, map it onto a reconfigurable device, and show several areas of performance gain. Ben Cordes, Jennifer G. Dy, Miriam Leeser, James Goebel |
FPGA | 3 |
| 2004 | Smart Camera Based on Reconfigurable Hardware Enables Diverse Real-Time ApplicationsabstractWe demonstrate the use of a "smart camera " to accelerate two very different image processing applications. The smart camera consists of a high quality video camera and frame grabber connected directly to an FPGA processing board. The advantages of this setup include minimizing the movement of large datasets and minimizing the latency by starting to process data before a complete frame has been acquired. The two applications, one from the area of medical image processing and the other from computational fluid dynamics both exhibit speedups of more than 20 times over software implementations on a 1.5 GHz PC. This smart camera setup is enabling image processing implementations that have not, before now, been achievable in real time. Miriam Leeser, Shawn Miller, Haiqian Yu |
FCCM | 1 |
| 2004 | An FPGA implementation of the two-dimensional finite-difference time-domain (FDTD) algorithmabstractUnderstanding and predicting electromagnetic behavior is needed more and more in modern technology. The Finite-Difference Time-Domain (FDTD) method is a powerful computational electromagnetic technique for modelling the electromagnetic space. The 3D FDTD buried object detection forward model is emerging as a useful application in mine detection and other subsurface sensing areas. However, the computation of this model is complex and time consuming. Implementing this algorithm in hardware will greatly increase its computational speed and widen its use in many other areas. We present an FPGA implementation to speedup the pseudo-2D FDTD algorithm which is a simplified version of the 3D FDTD model. The pseudo-2D model can be upgraded to 3D with limited modification of structure. We implement the pseudo-2D FDTD model for layered media and complete boundary conditions on an FPGA. The computational speed on the reconfigurable hardware design is about 24 times faster than a software implementation on a 3.0GHz PC. The speedup is due to pipelining, parallelism, use of fixed point arithmetic, and careful memory architecture design. Panos Kosmas, Miriam Leeser, Carey M. Rappaport |
FPGA | 3 |
| 2003 | Runtime Assignment of Reconfigurable Hardware Components for Image Processing PipelinesabstractThe combination of hardware acceleration and flexibility make FPGAs (field programmable gate arrays) important to image processing applications. There is also a need for efficient, flexible hardware/software codesign environments that can balance the benefits and costs of using FPGAs. Image processing applications often consist of pipeline of components where each component applies a different processing algorithm. Components can be implemented for FPGAs or software. Such systems enable an image analyst to work with either FPGA or software implementations of image processing algorithms for a given problem. The pipeline assignment problem chooses from alternative implementations of pipeline components to yield the fastest pipeline. Our codesign system solves the pipeline assignment problem to provide the most effective implementation automatically, so the image analyst can focus solely on choosing components, which make up the pipeline. However, the pipeline assignment problem is NP complete. An efficient, dynamic solution to the pipeline assignment problem is a desirable enabler of codesign systems which use both FPGA and software implementations. This paper is concerned with solving pipeline assignment in this context. Consequently, we focus on optimal and heuristic methods for fast (fixed time limit) runtime pipeline assignment are investigated. We present experimental finding for pipelines of twenty or fewer components, which show that in our environment, optimal runtime solutions are possible for smaller pipelines and nearly optimal heuristic solutions are possible for larger pipelines. Heather M. Quinn, Laurie A. Smith King, Miriam Leeser, Waleed Meleis |
FCCM | 3 |
| 2003 | Programming portable optimized multimedia applicationsabstractMultimedia computer architectures can speed-up applications significantly when programmed manually. Optimized programs have been non-portable up to now, because of differences in instruction sets, register lengths, alignment requirements and programming styles. We solve all these problems by using a library of C pre-processor macros called MMM. We implemented three examples from video compression in MMM, and automatically translated them into optimized code for four distinct multimedia processors. Their performance is comparable, and in several cases better, than equivalent examples optimized by the processor vendors. Juan Carlos Rojas, Miriam Leeser |
ACM Multimedia | 2 |
| 2002 | Parallel-beam backprojection: an FPGA implementation optimized for medical imagingabstractMedical image processing in general and computerized tomography (CT) in particular can benefit greatly from hardware acceleration. This application domain is marked by computationally intensive algorithms requiring the rapid processing of large amounts of data. To date, reconfigurable hardware has not been applied to this important area. For efficient implementation and maximum speedup, fixed-point implementations are required. The associated quantization errors must be carefully balanced against the requirements of the medical community. Specifically, care must be taken so that very little error is introduced compared to floating-point implementations and the visual quality of the images is not compromised. In this paper, we present an FPGA implementation of the parallel-beam backprojection algorithm used in CT for which all of these requirements are met. We explore a number of quantization issues arising in backprojection and concentrate on minimizing error while maximizing efficiency. Our implementation shows significant speedup over software versions of the same algorithm, and is more flexible than an ASIC implementation. Our FPGA implementation can easily be adapted to both medical sensors with different dynamic ranges as well as tomographic scanners employed in a wider range of application areas including nondestructive evaluation and baggage inspection in airport terminals. Srdjan Coric, Miriam Leeser, Eric L. Miller 0001, Marc Trepanier |
FPGA | 2 |
| 2002 | A Library of Parameterized Floating-Point Modules and Their Use
Pavle Belanovic, Miriam Leeser |
FPL | 2 |
| 2001 | Algorithmic transformations in the implementation of K- means clustering on reconfigurable hardwareabstractIn mapping the k-means algorithm to FPGA hardware, we examined algorithm level transforms that dramatically increased the achievable parallelism. We apply the k-means algorithm to multi-spectral and hyper-spectral images, which have tens to hundreds of channels per pixel of data. K-means is an iterative algorithm that assigns assigns to each pixel a label indicating which of K clusters the pixel belongs to. Mike Estlick, Miriam Leeser, James Theiler, John J. Szymanski |
FPGA | 2 |
| 2001 | Run-Time Execution of Reconfigurable Hardware in a Java EnvironmentabstractWe present tools that support the run-time execution of applications that mix software running on networks of workstations and reconfigurable hardware. We use JHDL to describe the reconfigurable hardware, and JavaPorts to handle the communications between nodes in the network. The heterogeneous resources are handled by interposing a communication layer between the application and the hardware. The communication layer provides (i) the ability to modify the target hardware without modifying the application, (ii) co-design of the application and hardware, (iii) simulation of the entire system before the hardware design is complete, and (iv) remote execution so the application can reside on a different host from the hardware. We demonstrate the feasibility of this approach with a Java-based system which has a communication layer called the packet exchange platform (PEP). We present the system, describe the PEP and its implementation, and show how this approach has been applied to an image processing application. Laurie A. Smith King, Heather M. Quinn, Miriam Leeser, Demetris G. Galatopoullos, Elias S. Manolakos |
ICCD | 3 |
| 2001 | Design and analysis of a dynamically reconfigurable three-dimensional FPGAabstractThis paper presents the design and analysis of a dynamically reconfigurable field programmable gate array (FPGA) that consists of three physical layers: routing and logic block layer, routing layer, and memory layer. The architecture was developed using a methodology that examines different architectural parameters and how they affect different performance criteria such as speed, area, and reconfiguration time. The resulting architecture has high performance while the requirement of balancing the areas of its constituent layers is satisfied. Silviu M. S. A. Chiricescu, Miriam Leeser, Mankuan Michael Vai |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | Implementing a RAKE receiver for wireless communications on an FPGA-based computer systemabstractRAKE receivers are widely used in the wireless communications industry. Currently, custom VLSI is the most popular implementation. Programmable and reconfigurable logic implementations are becoming more attractive because of their flexibility and due to technology advancements. We have implemented a RAKE receiver on an Annapolis Wildforce board with four Xilinx 4000 family chips for a total of 100,000 gate equivalents. Our system is able to implement a RAKE receiver for underwater data communication systems that works in real time. We also investigate mapping a RAKE receiver to a Virtex chip for real-time atmospheric wireless communication. Ali M. Shankiti, Miriam Leeser |
FPGA | 2 |
| 2000 | A data-centric approach to high-level synthesisabstractMoving data between various components of a system is fast becoming the performance bottleneck in digital design today. This is especially true in high-throughput, memory-intensive applications, like those in multimedia and video processing. With improved fabrication technology, the capacities of application-specific integrated circuits (ASICs) that implement these digital systems are increasing as well. Higher levels of design abstraction are used to prevent the design process from becoming untenable. High-level synthesis (HLS) is one such level of abstraction. We present Midas, an HLS system for ASIC design that treats the data produced and used by a system very differently from any previous HLS system. Midas uses a novel model for HLS centered around data-transfers (DTs), instead of operations as is more traditional. Midas also incorporates floorplanning information within the main HLS flow. The consideration of data-transfers and floorplanning during synthesis allows Midas to design architectures whose storage units and execution units show a close temporal and spatial integration. Data is stored near where it is produced and used. DTs happen over short distances instead of long ones. The total effect is better utilization of internal DT bandwidth on the ASIC. Shantanu Tarafdar, Miriam Leeser |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2000 | HML, a novel hardware description language and its translation to VHDLabstractWe present hardware ML (HML), an innovative hardware description language (HDL) based on the functional programming language SML. Features of HML not found in other HDL's include polymorphic types and advanced type checking and type inference techniques. We have implemented an HML type checker and a translator for automatically generating VHDL from HML descriptions. We generate a synthesizable subset of VHDL and automatically infer types and interfaces. This paper gives an overview of HML and discusses the translation from HML to VHDL and the type inference process. Yanbing Li, Miriam Leeser |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1998 | The DT-Model: High-Level Synthesis Using Data TransfersabstractWe presen t a new model for formulating the classic HLS sub-problems: scheduling, allocation, and binding. The model is unique in its use of data-transfers as the basic entity in syn thesis. A data transfer represents the movement of one instance of data and con tains the operation sourcing the data and all the operations using it. Our model compels the storage architecture of the design to be optimized concurren tly with the execution unit. We ha ve built a high-level syn thesis system, Midas, based on our data transfer model. Midas generates designs with smaller storage and data transfer requirements than other HLS systems. Shantanu Tarafdar, Miriam Leeser |
DAC | 2 |
| 1998 | High Level Synthesis for Designing Custom Computing HardwareabstractWe apply High Level Synthesis (HLS) to the design of FPGA based computing systems. HLS allows for a level of design space exploration unrealizable with Register Transfer Level (RTL) techniques. The use of HLS tools allow designers to prototype their designs with high quality results and fast turn around times. Our design flow makes use of Synopsys Behavioral Compiler (BC) followed by logic synthesis to map designs onto the Altera RIPP10 board. We illustrate our approach with a case study: the design of a DTMF receiver from a high-level behavioral description down to implementation on the RIPP10 board. We were able to design working hardware, meet our delay constraints and achieve 90% utilization of the available FPGAs. The final design had approximately 90000 gate equivalents. Goran Doncev, Miriam Leeser, Shantanu Tarafdar |
FCCM | 2 |
| 1998 | Integrating floorplanning in data-transfer based high-level synthesisabstractModern digital systems move and process vast amounts of data. Designing good ASIC architectures for these systems requires efficient data routing and storage. A high-level synthesis (HLS) system must consider spatial aspects of the architecture it synthesizes to achieve this. In this paper, we discuss using floorplanning information in the main HLS flow. Our HLS system, Midas, uses two floorplanners and formulates HLS using a data-transfer model. Midas synthesizes architectures whose data storage and transfer subsystems are spatially integrated with its execution unit. Midas also generates a high-level floorplan for the architecture, which contains the shapes and coordinates of its components and global routing channel specifications for its buses. Our experiments comparing Midas's architectures to those generated by a HLS system that does not use the data-transfer model or floorplanning show that Midas's architectures are smaller and yet allow for large amounts of simultaneous data mo... Shantanu Tarafdar, Miriam Leeser, Zixin Yin |
ICCAD | 2 |
| 1997 | Memory Traffic and Data Cache Behavior of an MPEG-2 Software DecoderabstractThe authors investigate the impact of multimedia applications on the cache behavior of desktop systems. Specifically they consider the memory bandwidth and data cache challenges associated with MPEG-2 software decoding. Recent extensions to instruction set architectures, including Intel's MMX, address the computational aspects of MPEG decoding. The large amount of data traffic generated, however has received little attention. Standard data caches consistently generate an excess of cache-memory traffic. Varying basic cache parameters only reduces traffic to double the minimum required at best. Incremental changes in cache size have a negligible effect for most feasible values. Increasing set associativity yields rapidly diminishing returns, and manipulating line size is similarly unproductive. Achieving higher efficiency requires understanding the composition and behavior of the decoder data set. They present a model of MPEG-2 decoder memory behavior and describe how to exploit this knowledge to minimize required memory bandwidth. Their results show that simply eliminating one component, video output data, from the cache can reduce traffic by as much as 50 percent. Peter Soderquist, Miriam Leeser |
ICCD | 2 |
| 1997 | Optimizing the Data Cache Performance of a Software MPEG-2 Video DecoderabstractMultimedia functionality has become an established component of core computer worHoads.MPEG-2 video decoding represents a particularly important and computationally demanding application example.Instruction set extensions like Intel's MMX significantly reduce the computational challenges of this and other multimedia algorithms.However, memory subsystem deficiencies have now become the major barrier to increased performance, partly as a consequence of this improved CPU performance.Decoding MPEG-2 video data in software makes significant bandwidth demands on memory subsystems, which is seriously aggravated by cache ineficiencies.Conventional data caches generate many times more cache-memory trafic than required, at best double the minimum necessary to support decoding.Improving eficiency requires understanding the behavior of the decoder and composition of its data set.We provide an analysis of the memory and cache behavior of software MPEG-2 video decoding, and lay out a set of cache-oriented architectural enhancements which offer relief for the problem of excess cache-memory bandwidth.Our results show that cache-sensitive handling of different data types can reduce trafic by 50 percent or more. Peter Soderquist, Miriam Leeser |
ACM Multimedia | 2 |
| 1995 | An Area/Performance Comparison of Subtractive and Multiplicative Divide/Square Root ImplementationsabstractThe implementations of division and square root in the FPU's of current microprocessors are based on one of two categories of algorithms. Multiplicative techniques, exemplified by the Newton-Raphson method and Goldschmidt's algorithm, share functionality with the floating-point multiplier. Subtractive methods, such as the many variations of radix-4 SRT, generally use dedicated, parallel hardware. These different approaches give rise to the distinct area and performance characteristics which are explored in this paper. Area comparisons are derived from measurements of commercial and academic hardware implementations. Representative divide/square root implementations are paired with typical add-multiply structures and simulated, using data from current microprocessor and arithmetic coprocessor designs, to obtain performance estimates. The results suggest that subtractive implementations offer a superior balance of area and performance, and stand to benefit most decisively from improvements in technology and growing transistor budgets due to their parallel operation. Multiplicative methods lend themselves best to situations where hardware re-use is mandated due to area or architectural constraints.> Peter Soderquist, Miriam Leeser |
IEEE Symposium on Computer Arithmetic | 2 |
| 1995 | Verification of a subtractive radix-2 square root algorithm and implementationabstractMany modern microprocessors implement floating point square root hardware using subtractive algorithms. Such processors include the HP PA7200, the MIPS R4400, and the Intel Pentium. The Intel Pentium division bug highlights the importance of verifying such implementations. In this paper we discuss the verification of a radix-2 square root unit similar to that used in the MIPS R4400. The verification is done by theorem proving to bridge the gap between the algorithm and the implementation. At the top level, we verify that a subtractive, non-restoring algorithm correctly calculates the square root function. We then show a series of optimizing transformations that refine the top level algorithm into the hardware implementation. Each transformation can be verified. We show the transformation of the top level proof to a level that is closer to the hardware implementation. The implementation is at the RTL level, and consists of a structural description of the hardware including an adder/subtracter, simple combinational hardware and some registers. Miriam Leeser, John W. O'Leary |
ICCD | 1 |
| 1995 | An Automaton Model for Scheduling Constraints in Synchronous MachinesabstractWe present a finite-state model for scheduling constraints in digital system design. We define a two-level hierarchy of finite-state machines: a behavior FSM's input and output events are partially ordered in time; a register-transfer FSM is a traditional FSM whose inputs and outputs are totally ordered in time. Explicit modeling of scheduling constraints is useful for both high-level synthesis and verification-we can explicitly search the space of register-transfer FSM's which implement a desired schedule. State-based models for scheduling are particularly important in the design of control-dominated systems. This paper describes the BFSM I model, describes several important operations and algorithms on BFSM's and networks of communicating BFSM's, and illustrates the use of BFSM's in high-level synthesis. Andrés Takach, Marilyn Wolf, Miriam Leeser |
IEEE Trans. Computers | 3 |
| 1995 | Verifying a Logic-Synthesis Algorithm and Implementation: A Case Study in Software VerificationabstractWe describe the verification of a logic synthesis tool with the Nuprl proof development system. The logic synthesis tool, Pbs, implements the weak division algorithm. Pbs consists of approximately 1000 lines of code implemented in a functional subset of Standard ML. It is a proven and usable implementation and is an integral part of the Bedroc high level synthesis system. The program was verified by embedding the subset of Standard ML in Nuprl and then verifying the correctness of the implementation of Pbs in the Nuprl logic. The proof required approximately 500 theorems. In the process of verifying Pbs we developed a consistent approach for using a proof development system to reason about functional programs. The approach hides implementation details and uses higher order theorems to structure proofs and aid in abstract reasoning. Our approach is quite general, should be applicable to any higher order proof system, and can aid in the future verification of large software implementations.> Mark D. Aagaard, Miriam Leeser |
IEEE Trans. Software Eng. | 2 |
| 1994 | Simulation of digital circuits in the presence of uncertainty
Mark H. Linderman, Miriam Leeser |
ICCAD | 2 |
| 1994 | A Methodology for Efficient Hardware Verification
Mark D. Aagaard, Miriam Leeser |
Formal Methods Syst. Des. | 2 |
| 1994 | PBS: proven Boolean simplificationabstractWe describe PBS, a formally proven implementation of multi-level logic synthesis based on the weak division algorithm. We have proved that for all legal input circuits, PBS generates an output circuit that is functionally correct and has minimal size. PBS runs on large examples in reasonable time for a prototype system. PBS, was verified using the Nuprl proof development system. The proof of PBS, which required several months, was well worth the effort since the benefits are realized every time the program is run. When engineers use a verified synthesis tool, they get the increased confidence of applying formal methods to their designs without the cost of varying each design produced.> Mark D. Aagaard, Miriam Leeser |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1993 | A Framework for Specifying and Designing PipelinesabstractWe present a framework for designing control circuitry for pipelined circuits. Our framework supports any mixture of data-dependent, data-stationary and time-stationary control for pipelines with out-of-order execution. There are four parameters to the framework: (1) Protocol schemes describe how transactions are transferred between stages in the pipeline. (2) Arbitration schemes specify how to prevent and/or handle structural hazards in the pipeline. (3) Control schemes determine how transactions are routed through the pipeline and how stages know what operation to perform. (4) Ordering schemes describe a method for matching up transactions as they leave the pipeline with transactions that entered the pipeline. Based on our framework we have developed an algorithm for compiling the control circuitry for a wide class of pipelines from a high-level description of the schedule for the pipeline.> Mark D. Aagaard, Miriam Leeser |
ICCD | 2 |
| 1991 | A Formally Verified System for Logic SynthesisabstractThe correctness of a logic synthesis system is implemented and proved. The algorithm is based on the weak division algorithm for Boolean simplification previously presented. The implementation is in the programming language ML; and the proof is in the Nuprl proof development system. This study begins with a proof of the algorithm previously presented and extends it to a level of detail sufficient for proving the implementation of the system. In the process of developing the proof many definitions presented in previous accounts of the algorithms were clarified, and several errors in the implementation were discovered. The result is that the designs generated by the implementation can be claimed to be correct by construction, since the correctness of the system was proven.> Mark D. Aagaard, Miriam Leeser |
ICCD | 2 |
| 1991 | Formally verified synthesis of combinational CMOS circuits
David A. Basin, Geoffrey M. Brown, Miriam Leeser |
Integr. | 3 |
| 1989 | Reasoning about the function and timing of integrated circuits with interval temporal logicabstractImportant aspects of behavior at the transistor level are discussed, including timing and capacitance. In the approach described here, the structures of circuits and their functional behavior are described with interval temporal logic (ITL). These specifications are expressed in Prolog, and the logical manipulations of the proof process are achieved with the Prolog system. To demonstrate the flexibility of this approach, the behavior of several CMOS circuits designed with different design styles is described. These examples include a dynamic latch and a 1-b adder, both of which use a two-phase clocking scheme and exploit charge storage. The 1-b adder is a sophisticated full adder implemented with a dynamic CMOS design style. Timing as well as functional aspects of behavior are derived, and constraints on the way a circuit interacts with its environment are reasoned about formally.> Miriam Leeser |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1986 | Automatic determination of signal flow through MOS transistor networks
W. F. Clocksin, Miriam Leeser |
Integr. | 2 |