VLDB 2026 Research / reviewers in the wild / expert
David Andrews 0001
dblp:25/4212-1 · also David L. Andrews 0001
· DBLP profile ↗
36ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0003-1464-7107ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 1 first-author · 13 since 2021Software engineering, systems software and programming languages · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimized Coding and Parameter Selection for Efficient FPGA Design of Attention MechanismsabstractEfficient utilization of on-chip computational and memory resources, along with optimized high-level synthesis (HLS) coding, is vital to maximize parallelism and minimize latency. This paper demonstrates the HLS algorithms to achieve high utilization of processing elements to enhance parallelism. It also analyzes how various parameters of an attention layer impact latency, employs an efficient tiling technique, and explains the process of selecting an optimized tile size (TS). Ehsan Kabir, Austin R. J. Downey, Jason D. Bakos, David Andrews 0001, Miaoqing Huang |
FCCM | 4 |
| 2025 | N-TORC: Native Tensor Optimizer for Real-Time ConstraintsabstractCompared to overlay-based tensor architectures like VTA or Gemmini, compilers that directly translate machine learning models into a dataflow architecture as HLS code, such as HLS4ML and FINN, generally can achieve lower latency by generating customized matrix-vector multipliers and memory structures tailored to the specific fundamental tensor operations required by each layer. However, this approach has significant drawbacks: the compilation process is highly time-consuming and the resulting deployments have unpredictable area and latency, making it impractical to constrain the latency while simultaneously minimizing area. Currently, no existing methods address this type of optimization. In this paper, we present N-TORC (Native Tensor Optimizer for Real-Time Constraints), a novel approach that utilizes data-driven performance and resource models to optimize individual layers of a dataflow architecture. When combined with model hyperparameter optimization, N-TORC can quickly generate architectures that satisfy latency constraints while simultaneously optimizing for both accuracy and resource cost (i.e. offering a set of optimal trade-offs between cost and accuracy). To demonstrate its effectiveness, we applied this framework to a cyber-physical application, DROPBEAR (Dynamic Reproduction of Projectiles in Ballistic Environments for Advanced Research). N-TORC's HLS4ML performance and resource models achieve higher accuracy than prior efforts, and its Mixed Integer Program (MIP)-based solver generates equivalent solutions to a stochastic search in 1000X less time. Suyash Vardhan Singh, Iftakhar Ahmad, David Andrews 0001, Miaoqing Huang, Austin R. J. Downey, Jason D. Bakos |
FCCM | 3 |
| 2025 | Resource Scheduling for Real-Time Machine Learning
Suyash Vardhan Singh, Iftakhar Ahmad, David Andrews 0001, Miaoqing Huang, Austin R. J. Downey, Jason D. Bakos |
FPGA | 3 |
| 2025 | DA-VinCi: A Deep-Learning Accelerator Overlay Using In-Memory ComputingabstractThe matrix operations that underpin today’s deep learning models are routinely implemented in Single Instruction Multiple Data (SIMD) domain specific accelerators. SIMD accelerators including GPUs and array processors can effectively leverage parallelism in models that are compute-bound, but their effectiveness can be diminished for models that are memory-bound. Processing-in-Memory (PIM) architectures are being explored to provide better energy efficiency and scalable performance for these memory-bound models. Modern Field Programmable Gate Arrays (FPGAs) feature hundreds of megabits of Static Random Access Memory (SRAM) distributed across the device as disaggregated memory resources. This makes FPGAs ideal programmable platforms for developing custom Processor In/Near Memory accelerators. Several PIM array-based accelerator designs have been proposed to leverage this substantial internal bandwidth. However, results reported to date show the FPGA based PIM architectures operating at system clock frequencies well below a chips Block-RAM (BRAM) Fmax clock frequency. Results also show that the compute densities of the designs do not scale linearly with BRAM densities. These results indicate that FPGA PIM architectures will never be competitive with their custom Application-Specific Integrated Circuit (ASIC) counterparts. In this article, we introduce DA-VinCi, a D eep-Learning A ccelerator O v erlay using In -Memory C omput i ng. DA-VinCi is the first scalable FPGA based PIM deep-learning accelerator overlay capable of clocking at the maximum frequency of a device’s BRAM. Further, the architecture of DA-VinCi allows the number of compute units to scale linearly up to the maximum capacity of a devices BRAM, and at the maximum clock frequency of the BRAM. The DA-VinCi overlay has a programmable Instruction Set Architecture (ISA) that allows the same synthesized design to provide low-latency inferencing of a range of memory-bound deep-learning models, including Multilayer Perceptrons, Recurrent Neural Network, Long Short-Term Memory, and Gated Recurrent Unit networks. The scalability and high clocking frequency of DA-VinCi is achieved through a new Processor In Memory (PIM) tile architecture and a highly scalable system-level framework. We present results showing DA-VinCi linearly scaling the number of Processing Elements (PEs) to 100% of the BRAM capacity (over 60K PEs) on an Alveo U55 clocking at 737 MHz, the chips BRAM Fmax. We provide comparative studies on inference latency across multiple deep-learning applications that show DA-VinCi achieves up to a 201 \(\times\) improvement over a state-of-the-art PIM overlay accelerator, up to 87 \(\times\) improvement over existing PIM-based FPGA accelerators, and up to 57 \(\times\) improvement over custom deep-learning accelerators on FPGAs. M. D. Arafat Kabir, Nathaniel Fredricks, Tendayi Kamucheka, Joel Mandebi, Miaoqing Huang, Jason D. Bakos, David Andrews 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2024 | The BRAM is the Limit: Shattering Myths, Shaping Standards, and Building Scalable PIM AcceleratorsabstractMany recent FPGA-based Processor-in-Memory (PIM) architectures have appeared with promises of impressive levels of parallelism but with performance that falls short of expectations due to reduced maximum clock frequencies, an inability to scale processing elements up to the maximum BRAM capacity, and minimal hardware support for large reduction operations. In this paper, we propose a “Standard” set of design objectives for PIM array-based FPGA designs. We then propose a PIM array-based GEMV accelerator architecture as a case study to show the proposed Standard can be realized in practice. The GEMV accelerator serves as existence proof that dispels several myths surrounding what is normally accepted as clocking and scaling FPGA performance limitations. Specifically, the proposed accelerator clocks at the maximum frequency of the BRAM and scales to 100% of the available BRAMs. Comparative analyses show execution speeds over existing PIM-based GEMV engines on FPGAs and achieving a 2.65Χ – 3.2Χ faster clock. An AMD Alveo U55 implementation achieves a system clock speed of 737 MHz, providing 64K bit serial multiply-accumulate (MAC) units for GEMV operation. M. D. Arafat Kabir, Tendayi Kamucheka, Nathaniel Fredricks, Joel Mandebi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001 |
FCCM | 7 |
| 2024 | Ph.D. Project: A Compiler-Driven Approach to HW/SW Co-Design of Deep-Learning AcceleratorsabstractThis work introduces SPAR, a compiler framework based on MLIR, tailored for deep learning inference applications. Alongside SPAR, we present SPAR HL, an extensible Instruction Set Architecture (ISA) designed to seamlessly interface custom FPGA-based accelerators with the SPAR compiler. Historically, custom accelerators on FPGA have posed a challenge as a compiler target. Prior compiler initiatives have addressed this issue by generating application-specific hardware during compilation, necessitating expertise across both application and hardware domains. In response, we offer an alternative approach by proposing an ISA as a unified compiler target and hardware interface for custom accelerators. Furthermore, we introduce a compiler capable of translating high-level machine learning models encoded in ONNX into code compatible with our proposed ISA, thus enabling efficient deployment on FPGA-based custom accelerators. Tendayi Kamucheka, David Andrews 0001 |
FCCM | 2 |
| 2024 | IMAGine: An In-Memory Accelerated GEMV Engine OverlayabstractProcessor-in-Memory (PIM) overlays and alternative reconfigurable tile fabrics have been proposed to eliminate the von Neumann bottleneck and enable processing performance to scale with BRAM capacity. The performance of these FPGA-based PIM architectures has been limited due to a reduction of the BRAMs maximum clock frequencies and less than ideal scaling of processing elements with increased BRAM capacity. This paper presents IMAGine, an In-Memory Accelerated GEMV engine, a PIM-array accelerator that clocks at the maximum frequency of the BRAM and scales to 100% of the available BRAMs. Comparative analyses are presented showing execution speeds over existing PIM-based GEMV engines on FPGAs and achieving a $2.65 \times-3.2 \times$ faster clock. An AMD Alveo U55 implementation is presented that achieves a system clock speed of 737 MHz, providing 64 K bit-serial multiply-accumulate (MAC) units for GEMV operation. This establishes IMAGine as the fastest PIM-based GEMV overlay, outperforming even the custom PIM-based FPGA accelerators reported to date. Additionally, it surpasses TPU v1-v2 and Alibaba Hanguang 800 in clock speed while offering an equal or greater number of multiply-accumulate (MAC) units. M. D. Arafat Kabir, Tendayi Kamucheka, Nathaniel Fredricks, Joel Mandebi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001 |
FPL | 7 |
| 2023 | Making BRAMs Compute: Creating Scalable Computational Memory Fabric OverlaysabstractThe increasing density of distributed BRAMs diffused throughout modern Field Programmable Gate Arrays (FP-GAs) is ideal for forming processor in/near memory architectures. This breaks the traditional von Neumann memory bottleneck limiting concurrency and degrading energy efficiency. Ideally, processing density should scale linearly with BRAM capacity, and clock frequencies should be set by the read/write access times of the BRAM. In this paper, we present a PIM overlay that achieves these goals. We observe an improvement of performance by 2.25 x, logic resource utilization by 2 x, and accumulation delay by 17 x compared to prior published work. M. D. Arafat Kabir, Joshua Hollis, Atiyehsadat Panahi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001 |
FCCM | 6 |
| 2023 | Accelerating LSTM-Based High-Rate Dynamic System ModelsabstractIn this paper, we evaluate the use of a trained Long Short-Term Memory (LSTM) network as a surrogate for a Euler-Bernoulli beam model, and then we describe and characterize an FPGA-based deployment of the model for use in real-time structural health monitoring applications. The focus of our efforts is the DROPBEAR (Dynamic Reproduction of Projectiles in Ballistic Environments for Advanced Research) dataset, which was generated as a benchmark for the study of real-time structural modeling applications. The purpose of DROPBEAR is to evaluate models that take vibration data as input and give the initial conditions of the cantilever beam on which the measurements were taken as output. DROPBEAR is meant to serve an exemplar for emerging high-rate “active structures” that can be actively controlled with feedback latencies of less than one microsecond. Although the Euler-Bernoulli beam model is a well-known solution to this modeling problem, its computational cost is prohibitive for the time scales of interest. It has been previously shown that a properly structured LSTM network can achieve comparable accuracy with less workload, but achieving sub-microsecond model latency remains a challenge. Our approach is to deploy the LSTM optimized specifically for latency on FPGA. We designed the model using both high-level synthesis (HLS) and hardware description language (HDL). The lowest latency of$1.42\ \mu\mathrm{S}$and the highest throughput of 7.87 Gops/s were achieved on Alveo U55C platform for HDL design. Ehsan Kabir, Daniel Coble, Joud N. Satme, Austin R. J. Downey, Jason D. Bakos, David Andrews 0001, Miaoqing Huang |
FPL | 6 |
| 2023 | FPGA Processor In Memory Architectures (PIMs): Overlay or Overhaul ?abstractThe dominance of machine learning and the ending of Moore's law have renewed interests in Processor in Memory (PIM) architectures. This interest has produced several recent proposals to modify an FPGA's BRAM architecture to form a next-generation PIM reconfigurable fabric [1], [2]. PIM architectures can also be realized within today's FPGAs as overlays without the need to modify the underlying FPGA architecture. To date, there has been no study to understand the comparative advantages of the two approaches. In this paper, we present a study that explores the comparative advantages between two proposed custom architectures and a PIM overlay running on a commodity FPGA. We created PiCaSO, a Processor in/near Memory Scalable and Fast Overlay architecture as a representative PIM overlay. The results of this study show that the PiCaSO overlay achieves up to 80% of the peak throughput of the custom designs with 2.56 x shorter latency and 25% - 43% better BRAM memory utilization efficiency. We then show how several key features of the PiCaSO overlay can be integrated into the custom PIM designs to further improve their throughput by 18%, latency by 19.5%, and memory efficiency by 6.2%. M. D. Arafat Kabir, Ehsan Kabir, Joshua Hollis, Eli Levy-Mackay, Atiyehsadat Panahi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001 |
FPL | 8 |
| 2022 | High-Rate Machine Learning for Forecasting Time-Series Signalsabstract"Active structures" are physical structures that incorporate real-time monitoring and control. Examples include active vibration damping or blast mitigation systems. Evaluating physics-based models in real-time is generally not feasible for such systems having high-rate dynamics which require microsecond response times, but data-driven machine-learning-based models can potentially offer a solution. This paper compares the cost and performance of two FPGA-based implementations of real-time, continuously-trained models for forecasting time-series signals with non-stationarities, with one using High-Level Synthesis (HLS) and the other a programmable overlay architecture. The proposed model accepts a uni-variate vibration signal and seeks to forecast future samples to inform high-rate controllers. The proposed forecasting method performs two concurrent neural inference operations. One inference forecasts the state of the signal f samples into the future as a function of the most recent h samples, while the other forecasts the current sample given h samples starting from h+f−1 samples into the past. The first forecast produces the forecast while the second forecast allows the system to calculate the model’s loss and perform an immediate model update before the next sample period. Atiyehsadat Panahi, Ehsan Kabir, Austin R. J. Downey, David Andrews 0001, Miaoqing Huang, Jason D. Bakos |
FCCM | 4 |
| 2022 | A Masked Pure-Hardware Implementation of Kyber Cryptographic AlgorithmabstractQuantum computing-specifically Shor's algorithm [1]-presents an existential threat to some standard cryptographic algorithms. In preparation, post-quantum cryptography (PQC) algorithms have been in development and are nearing mathematical and cryptanalytic maturity. Standardization efforts through the National Institute of Standards and Technology (NIST) PQC standardization process have chosen one PKE/KEM algorithm (i.e., CRYSTALS-Kyber) and three digital signature algorithms (i.e., CRYSTALS-Dilithium, Falcon, and SPHINCS+). CRYSTALS-Kyber is a lattice-based, IND-CCA2-secure, key-encapsulation mechanism (KEM) based on the learning-with-errors problem over module lattices. This paper presents a masked hardware implementation of Kyber that is demonstrably secure against side-channel power analysis methods. Tendayi Kamucheka, Alexander Nelson 0001, David Andrews 0001, Miaoqing Huang |
FPT | 3 |
| 2021 | A Customizable Domain-Specific Memory-Centric FPGA Overlay for Machine Learning ApplicationsabstractThis paper presents an overview and performance analysis of a software-programmable domain-customizable System-on-Chip (SoC) overlay for low-latency inferencing of variable and low-precision Machine Learning (ML) networks targeting Internet-of-Things (IoT) edge devices. The SoC includes a 2-D processor array that can be customized at design time for FPGA logic families. The overlay resolves historic issues of poor designer productivity associated with traditional Field Programmable Gate Array (FPGA) design flows without the performance losses normally incurred by overlays. A standard Instruction Set Architecture (ISA) allows different ML networks to be quickly compiled and run on the overlay without the need to resynthesize. Performance results are presented that show the overlay achieves $1.3\times-8.0\times$ speedup over custom designs while still allowing rapid changes to ML algorithms on the FPGA through standard compilation. Atiyehsadat Panahi, Suhail Basalama, Ange-Thierry Ishimwe, Joel Mandebi, David Andrews 0001 |
FPL | 5 |
| 2020 | FPGA-Based Gesture Recognition with Capacitive Sensor Array using Recurrent Neural NetworksabstractThis work presents a prototype of an FPGA-based hand motion recognition system using a capacitive sensor array (CSA). The prototype system is being developed as a tool to evaluate upper-limb motor skills for assistive or rehabilitative applications. A light-weight gesture segmentation algorithm was developed that uses summation and moving average filtering of quantized capacitive sensing data to segment motions. The time-series hand motions are then recognized through a recurrent classifier based on long short-term memory (LSTM) neural networks. The classifier model is trained on uni-stroke hand written digit ('0'–'9') samples obtained from four volunteers. A total of 12,000 hand motion samples are collected. The accuracy of 10-fold and leave-one-user-out cross-validation accuracy is respectively 97.5% and 91.3% using a two-layer LSTM network. The LSTM classifier is implemented on a Zynq FPGA device. The experiment demonstrated that the FPGA implementation of the LSTM-based classifier can achieve real-time gesture classification with capacitive sensor data. Haoyan Liu 0002, Atiyehsadat Panahi, David Andrews 0001, Alexander Nelson 0001 |
FCCM | 3 |
| 2018 | FPGAVirt: A Novel Virtualization Framework for FPGAs in the CloudabstractField-Programmable Gate Arrays (FPGAs) are becoming important components within commercially available cloud computing systems. However, the FPGAs are not yet sufficiently abstracted within existing software ecosystems. Contrary to how applications are transparently scheduled across general purpose processors, software processes need to explicitly provision and control communications with hardware circuits within the FPGAs. In this paper, we introduce a novel virtualization framework called FPGAVirt that leverages Virtio to implement an efficient communication scheme between virtual machines and the FPGAs. FPGAVirt avoids the overhead of context switches between virtual machine and host address spaces by using the in-kernel network stack for transferring packets to FPGAs. Experimental results show FPGAVirt can deliver an additional 2x to 35x performance increase compared to current state of the art virtualization approaches. Joel Mandebi, Festus Hategekimana, Danielle Tchuinkou, David Andrews 0001, Christophe Bobda |
IEEE CLOUD | 4 |
| 2018 | Enabling Transparent Acceleration of OpenCV Library Kernels on a Hybrid Memory Cube ComputerabstractThis paper presents a CPU-FPGA heterogeneous system that provides hardware support for computer vision libraries to attain acceleration over image processing applications. The architecture achieves this improvement by staying completely transparent to the developer while providing necessary acceleration on conventional software designs. Md Jubaer Hossain Pantho, Joel Mandebi, Christophe Bobda, David Andrews 0001, Marjan Asadinia |
FCCM | 4 |
| 2018 | Transparent Acceleration of Image Processing Kernels on FPGA-Attached Hybrid Memory Cube ComputersabstractThe Hybrid Memory Cube (HMC) is representative of emerging architectures that integrate FPGAs with multichannel interconnected 3-D stacked memory, offering great potential for high bandwidth streaming applications. However, creating new hardware components that tap the full potential of the concurrent communications channels requires the structural understanding of the memory layout and interconnect configurations. In this paper, we present a new development framework aimed at removing the need for software programmers to understand the underlying physical architecture. The proposed framework automates the creation of hardware/software co-designs for computer vision applications in a transparent way to the developer. The development system dynamically detects function calls in software kernels and replaces those calls by a hardware wrapper function that exploits the HMCs memory hierarchy and multichannel interconnect with the FPGA. Results show our flow can exploit the 3-D stacked memory and concurrent communications channels to achieve speed-up with no need to tune the original software application to the memory hierarchy. Md Jubaer Hossain Pantho, Joel Mandebi, Christophe Bobda, David Andrews 0001 |
FPT | 4 |
| 2016 | Run time interpretation for creating custom accelerators
Sen Ma, Zeyad Aklah, David Andrews 0001 |
DATE | 3 |
| 2016 | Just In Time Assembly of AcceleratorsabstractDespite the significant advancements that have been made in High Level Synthesis, the reconfigurable computing community has failed at getting programmers to use Field Programmable Gate Arrays (FPGAs). Existing barriers that prevent programmers from using FPGAs include the need to work within vendor specific CAD tools, knowledge of hardware programming models, and the requirement to pass each design through synthesis, place and route. In this paper we present a new approach that takes these barriers out of the design flows for programmers. Synthesis is eliminated from the application programmers path by becoming part of the initial coding process when creating the programming patterns that define a Domain Specific Language. Programmers see no difference between creating software or hardware functionality when using the DSL. A run time interpreter is introduced that assembles hardware accelerators within a configurable tile array of partially reconfigurable slots at run time. Initial results show the approach allows hardware accelerators to be compiled 100x faster compared to the time required to synthesize the same functionality. Initial performance results further show a compilation/interpretation approach can achieve approximately equivalent performance for matrix operations and filtering compared to synthesizing a custom accelerator. Sen Ma, Zeyad Aklah, David Andrews 0001 |
FPGA | 3 |
| 2015 | Automatic support for multi-module parallelism from computational patternsabstractField Programmable Gate Arrays (FPGAs) can be customized into application-specific architectures to achieve high performance and energy-efficiency. Unfortunately, they are yet to gain significant adoption by application developers due to their low-level programming model. Moreover, to obtain good performance in an FPGA design, one often needs to correctly parallelize computation and balance the computational throughput with the available data access bandwidth. To address the programming model problem, recent efforts have focused on composing applications out of parallel computational patterns, such as map, reduce, zipWith and foreach, and leveraging the properties of these patterns to generate highly parallel hardware modules capable of high performance. In this work, we focus on the problem of further improving the performance and show that we can utilize the knowledge of how data is consumed and produced by these computational patterns in conjunction with the information of the system architecture to automatically parallelize computations across multiple hardware modules. To achieve this, we automatically infer synchronization needs arising due to parallelization and generate a complete system that can obtain high performance for a given application. We evaluate our approach using seven applications from different domains and show that our automatically generated designs achieve performance improvements ranging from 1.8 to 9.4 times. Nithin George, HyoukJoong Lee, David Novo, Muhsen Owaida, David Andrews 0001, Kunle Olukotun, Paolo Ienne |
FPL | 5 |
| 2015 | A run time interpretation approach for creating custom acceleratorsabstractThe world of software development has the notion of just-in-time compilation, run time binary translation, and language interpretation. These dynamic run time techniques support increased code portability and designer productivity. There are no such equivalences to increase the productivity or portability of creating new hardware components within Field Programmable Gate Arrays (FPGAs). Instead, creating a new hardware component requires hardware design skills and the overhead of running through synthesis, place and route. If a change is made to even a single line of code, the synthesis, place and route steps must be repeated. In this paper we present a new approach that allows hardware accelerators to be built and run using compilation and run time interpretation. Our results show the approach can enable software programmers without any hardware skills to create hardware accelerators at productivity levels consistent with software development and compilation. The same accelerator can be compiled 100× faster than synthesis. Even though the approach is focused on productivity, our observed performance results are promising. Our initial application test cases show the same accelerator written by a software programmer and synthesized through Vivado HLS or written using our DSL and compiled within our approach achieves equivalent performance. Sen Ma, Zeyad Aklah, David Andrews 0001 |
FPL | 3 |
| 2014 | On energy efficiency and amdahl's law in FPGA based chip heterogeneous multiprocessor systems (abstract only)abstractThis poster presents our preliminary findings on the relationship between speedup and energy efficiency on FPGA based Chip Heterogeneous Multiprocessor Systems (CHMPs). While researchers have investigated how to tailor combinations of heterogeneous compute engines within a CHMP system to best meet the performance needs of specific applications, exploring how these optimized architectures also effect energy efficiency is not as well studied. We show that a simple relationship exists between the speedup these systems gain and their associated energy efficiency. We show that the simple relationship between Amdahl's law and energy efficiency. All the experiments result achieved through actual run time measurements on homogeneous and heterogeneous multiprocessor systems implemented within a Xilinx Virtex6 FPGA. We further show how a systems with 6 MicroBlaze soft processors' dynamic power and hence the overall energy efficiency of the system can be effected through transparent operating system control of the compute resources. We also present how to use clock gating to control the dynamic power consumption for each processor and with this careful power-aware management unit, the system's dynamic power consumption can follow the requirements of each application. Sen Ma, David Andrews 0001 |
FPGA | 2 |
| 2014 | Achieving portability and efficiency over chip heterogeneous multiprocessor systemsabstractEmerging programming models for chip heterogeneous multiprocessor (CHMP) systems elevate architecture details up into the source code. This eliminates portability and requires designers to navigate a multidimensional search space when trying to optimize designs. In this paper, we present an approach that reinstates portability through a combination of polymorphic functions and an adaptive runtime system. Together they enable runtime profiling and dynamic scheduling of unaltered source code across systems with different combinations of heterogeneous resources. Our results verify the ability of our programming model and runtime system to re-enable the notion of writing code once and run anywhere. Runtime results show how runtime tuning can increase resource utilization and provide performance increases as the number and heterogeneity of computing resources increases. Eugene Cartwright, Alborz Sadeghian, Sen Ma, David Andrews 0001 |
FPL | 4 |
| 2013 | Modular Design of Fully Pipelined Reduction Circuits on FPGAsabstractFast and efficient reduction circuits are critical for a broad range of scientific and embedded system applications. High throughput reduction circuits are typically hand designed for specific vector lengths. These circuits need to be modified when the set lengths are changed. In this paper, we present a new design approach that can handle any set length or combination of different consecutive set lengths without stalling and generates in-order results. The flexibility of the design allows it to be used for any reduction operations, such as floating-point addition and multiplication. By providing a simple and efficient interface to the user and a modular architecture for the designer, the proposed technique has a broad impact across a wide range of custom hardware designs. Miaoqing Huang, David Andrews 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2012 | Automating the design of mLUT MPSoPC FPGAs in the cloudabstractModern platform FPGAs are over the million-LUT level, large enough to support complete heterogeneous Multiprocessor System-On-Chips (MPSoCs). Constructing systems with 10's of processors is currently feasible using existing manual methods within vendor-specific CAD tools. However these manual, by-hand, approaches will not be feasible for constructing future systems with 100's to 1,000's of processors. Instead, new automated system assembly approaches will be required to handle these levels of system complexity and diversity. In this paper we present a new automated design flow for creating such next generation heterogeneous MPSoCs. An integral part of the MPSoPC system created is the inclusion of a general purpose PThreads-compliant HW/SW co-designed operating system and heterogeneous compiler. Our design flow has been placed in the cloud and is freely accessible across the Internet. Eugene Cartwright, Azad Fahkari, Sen Ma, Christina Smith, Miaoqing Huang, David Andrews 0001, Jason Agron |
FPL | 6 |
| 2010 | Distributed Hardware-Based Microkernels: Making Heterogeneous OS Functionality a System PrimitiveabstractAs chips have moved from homogeneous single core systems to much more complex, heterogeneous multi-core systems, the ability to create both uniform and efficient operating system services has begun to diminish. The importance of these services suggests that these primitives should no longer be virtual, but rather physical services built into modern computing devices. In this paper we outline some of the challenges involved in building traditional OS services in heterogeneous computing systems. We present a hardware-based solution that provides basic OS primitives to heterogeneous systems that are both efficient and uniformly accessible to heterogeneous compute elements. A prototype system utilizing a hardware-based microkernel is demonstrated that allows programmers to target systems with ISA-level heterogeneity using a familiar, uniform multithreaded programming model. Jason Agron, David Andrews 0001 |
FCCM | 2 |
| 2010 | Modular design of fully pipelined accumulatorsabstractFast and efficient accumulation arithmetic circuits are critical for a broad range of scientific and embedded system applications. High throughput accumulation circuits are typically hand designed for specific vector lengths requiring the circuit to be modified when the lengths are changed. In this work we present a new design approach that can achieve low latency and near optimal throughput for input data vectors of arbitrary length. The flexibility of the design allows it to be used for both integer and floating-point operations. By providing a simple and efficient interface to the user and a modular architecture for the designer, the proposed technique has broad impact across a wide range of custom hardware designs. Miaoqing Huang, David Andrews 0001 |
FPT | 2 |
| 2009 | Building heterogeneous reconfigurable systems using threadsabstractField Programmable Gate Arrays (FPGAs) have long held the promise of allowing designers to create systems with performance levels close to custom circuits but with a software-like productivity for reconfiguring the gates. Unfortunately achieving this promise has been elusive. Modern FPGAs can now support a complete Multi-processor System on Chip (MPSoC) architecture that raises the design abstraction level from gates to processors. In this paper we present a new design flow and run-time system that enables developers to create a complete heterogeneous MPSoC from high-level programming model abstractions. This approach allows designers to eliminate synthesis times by using soft core processors thus enabling the creation of custom heterogeneous MPSoC architectures at software productivity levels. Jason Agron, David Andrews 0001 |
FPL | 2 |
| 2008 | An Infrastructure for Hardware-Software Co-Design of Embedded Real-Time Java ApplicationsabstractThe partitioning of applications into hardware and software is an important issue in embedded systems, opening room for high level specifications as well as the exploration of different implementation strategies. This paper presents a software architecture to specify threads in hardware in the context of the real time specification for Java (RTSJ) standard. There is a Java class that encapsulates hardware components, providing an abstraction layer to the application developer. Below this Java class, a wrapper hardware component provides a standard interface between RTSJ-based software components and the hardware that implements the thread behavior. This approach provides a high flexibility in choosing either a hardware or software implementation, allowing to postpone hardware/software partitioning to the very end of system development. The paper includes some quantitative data from an example containing hardware and software threads. While both implementations are compatible with the rest of the application from an interface point- of-view, they lead to very different timing and area results. Elias Teodoro Silva Jr., David Andrews 0001, Carlos Eduardo Pereira, Flávio Rech Wagner |
ISORC | 2 |
| 2008 | Achieving Programming Model Abstractions for Reconfigurable ComputingabstractThis paper introduces hthreads, a unifying programming model for specifying application threads running within a hybrid computer processing unit (CPU)/field-programmable gate-array (FPGA) system. Presently accepted hybrid CPU/FPGA computational models-and access to these computational models via high level languages-focus on programming language extensions to increase accessibility and portability. However, this paper argues that new high-level programming models built on common software abstractions better address these goals. The hthreads system, in general, is unique within the reconfigurable computing community as it includes operating system and middleware layer abstractions that extend across the CPU/FPGA boundary. This enables all platform components to be abstracted into a unified multiprocessor architecture platform. Application programmers can then express their computations using threads specified from a single POSIX threads (pthreads) multithreaded application program and can then compile the threads to either run on the CPU or synthesize them to run within an FPGA. To enable this seamless framework, we have created the hardware thread interface (HWTI) component to provide an abstract, platform-independent compilation target for hardware-resident computations. The HWTI enables the use of standard thread communication and synchronization operations across the software/hardware boundary. Key operating system primitives have been mapped into hardware to provide threads running in both hardware and software uniform access to a set of sub-microsecond, minimal-jitter services. Migrating the operating system into hardware removes the potential bottleneck of routing all system service requests through a central CPU. David Andrews 0001, Ron Sass, Erik K. Anderson, Jason Agron, Wesley Peck, Jim Stevens, Fabrice Baijot, Ed Komp |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2007 | Supporting High Level Language Semantics Within Hardware Resident ThreadsabstractThe paper presents the new Hardware Thread Interface (HWTI), a meaningful and semantic rich target for a high level language to hardware descriptive language translator. The HWTI provides a hardware thread with the same hthread system calls available to software threads, a fast global distributed memory, support for pointers, a generalized function call model including recursion, local variable declaration, dynamic memory allocation, and a remote procedural call model that enables hardware threads access to any library function. Erik K. Anderson, Wesley Peck, Jim Stevens, Jason Agron, Fabrice Baijot, Seth Warn, David Andrews 0001 |
FPL | 7 |
| 2006 | Enabling a Uniform Programming Model Across the Software/Hardware BoundaryabstractIn this paper, we present hthreads, a unifying programming model for specifying application threads running within a hybrid CPU/FPGA system. Threads are specified from a single pthreads multithreaded application program and compiled to run on the CPU or synthesized to run on the FPGA. The hthreads system, in general, is unique within the reconfigurable computing community as it abstracts the CPU/FPGA components into a unified custom threaded multiprocessor architecture platform. To support the abstraction of the CPU/FPGA component boundary, we have created the hardware thread interface (HWTI) component that frees the designer from having to specify and embed platform specific instructions to form customized hardware/software interactions. Instead, the hardware thread interface supports the generalized pthreads API semantics, and allows passing of abstract data types between hardware and software threads. Thus the hardware thread interface provides an abstract, platform independent compilation target that enables thread and instruction-level parallelism across the software/hardware boundary Erik K. Anderson, Jason Agron, Wesley Peck, Jim Stevens, Fabrice Baijot, Ed Komp, Ron Sass, David Andrews 0001 |
FCCM | 8 |
| 2006 | Hthreads: A Computational Model for Reconfigurable DevicesabstractRecent architectural advancements in reconfigurable devices have exposed the ability to support massive parallelism inside of small, low-cost, embedded devices. The massive parallelism inside of these reconfigurable devices has promised to bring an unprecedented level of performance to the embedded systems domain. However, the complexity of programming these reconfigurable devices is daunting - too daunting for the average programmer. This paper presents Hthreads. Hthreads is a computational architecture which aims to bridge the gap between regular programmers and powerful but complex reconfigurable devices. Hthreads accomplishes this goal using three layers of abstraction built upon standard reconfigurable devices: an operating system capable of supporting a diverse collection of computational models within a reconfigurable device, an intermediate form representation which eases the development of applications on reconfigurable devices, and support for high level languages which are familiar to most programmers Wesley Peck, Erik K. Anderson, Jason Agron, Jim Stevens, Fabrice Baijot, David Andrews 0001 |
FPL | 6 |
| 2006 | Run-Time Services for Hybrid CPU/FPGA Systems on ChipabstractModern FPGA devices, which include (multiple) processor core(s) as diffused IP on the silicon die, provide an excellent platform for developing custom multiprocessor systems-on-programmable chip (MPSoPC) architectures. As researchers are investigating new methods for migrating portions of applications into custom hardware circuits, it is also critical to develop new run-time service frameworks to support these capabilities. Hthreads (HybridThreads) is a multithreaded RTOS kernel for hybrid FPGA/CPU systems designed to meet this new growing need. A key capability of hthreads is the migration of thread management, synchronization primitives, and run-time scheduling services for both hardware and software threads into hardware. This paper describes the hthreads scheduler, a key component for controlling both software-resident threads (SW threads) and threads implemented in programmable logic (HW threads). Run-time analysis shows that the hthreads scheduler module helps in reducing unwanted system overhead and jitter when compared to historical software schedulers, while fielding scheduling requests from both hardware and software threads in parallel with application execution. Run time analysis shows the scheduler achieves constant time scheduling for up to 256 active threads with a total of 128 different priority levels, while using uniform APIs for threads requesting OS services from either side of the hardware/software boundary Jason Agron, Wesley Peck, Erik K. Anderson, David Andrews 0001, Ed Komp, Ron Sass, Fabrice Baijot, Jim Stevens |
RTSS | 4 |
| 2002 | Interprocess communications in the AN/BSY-2 distributed computer system: a case study
David Andrews 0001, Paul Austin, Peter Costello, David O. LeVan |
J. Syst. Softw. | 1 |
| 1994 | Rapid prototype of an SIMD processor array (using FPGA's)abstractA custom chip set implementing a single instruction multiple data (SIMD) architecture has been designed bringing the benefits of massively parallel processing to the embedded systems domain. A scaled, rapid prototype was first implemented providing an exact duplicate of the functionality and interfaces of the custom chips, but using off the shelf technology. This scaled version was specified to allow development and debugging of software, and provide early feedback for verification of the interfaces and instruction operations. The rapid prototype provides full functionality, allowing any design errors or beneficial modifications to the design to be identified.> David Andrews 0001, Andrew Wheeler, Barry Wealand, Cliff Kancler |
RSP | 1 |