VLDB 2026 Research / reviewers in the wild / expert
Martin C. Herbordt
dblp:87/324
· DBLP profile ↗
91ranked-venue papers
17as first author
19since 2021 · last 2025
0000-0002-3443-9113ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 85 · 13 first-author · 19 since 2021Artificial intelligence and machine learning · 4 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SmartNIC-GPU-CPU Heterogeneous System for Large Machine Learning Model with Software-Hardware CodesignabstractThe rapid growth of large machine learning models, from billions to trillions of parameters, has led to powerful AI capabilities that increasingly impact everyday life.However, this expansion in model size has surpassed the capacity of GPU memory.As a result, GPU clusters-built by aggregating multiple GPUs-have scaled up significantly to accommodate these models.To address this scalability challenge, and to make large-model training more widely accessible, researchers have proposed heterogeneous systems.These systems leverage CPUs and secondary memory to offload storage and computation onto these devices, thereby reducing the total number of GPUs required for training.Despite their promise, such heterogeneous systems have so far faced challenges in achieving high efficiency and performance. Anqi Guo, Yuchen Hao, Xiteng Yao, Shining Yang, Tong Geng, Martin C. Herbordt |
ICS | 7 |
| 2024 | SmartFuse: Reconfigurable Smart Switches to Accelerate Fused Collectives in HPC ApplicationsabstractCommunication switches have sometimes been augmented to process collectives, e.g., in the IBM BlueGene and Mellanox SHArP switches. In this work, we find that there is a great acceleration opportunity through the further augmentation of switches to accelerate more complex functions that combine communication with computation. We consider three types of such functions. The first is fully-fused collectives built by fusing multiple existing collectives like Allreduce with Alltoall. The second is semi-fused collectives built by combining a collective with another computation. The third are higher-order collectives built by combining multiple computations and communications, such as to perform matrix-matrix multiply (PGEMM). Pouya Haghi, Cheng Tan 0002, Anqi Guo, Chunshu Wu, Dongfang Liu, Ang Li 0006, Anthony Skjellum, Tong Geng, Martin C. Herbordt |
ICS | 9 |
| 2024 | FPGA-Accelerated Range-Limited Molecular DynamicsabstractLong timescale Molecular Dynamics (MD) simulation of small molecules is crucial in drug design and basic science. To accelerate a small data set that is executed for a large number of iterations, high-efficiency is required. Recent work in this domain has demonstrated that among COTS devices only FPGA-centric clusters can scale beyond a few processors. The problem addressed here is that, as the number of on-chip processors has increased from fewer than 10 into the hundreds, previous intra-chip routing solutions are no longer viable. We find, however, that through various design innovations, high efficiency can be maintained. These include replacing the previous broadcast networks with ring-routing and then augmenting the rings with out-of-order and caching mechanisms. Others are adding a level of hierarchical filtering and memory recycling. Two novel optimized architectures emerge, together with a number of variations. These are validated, analyzed, and evaluated. We find that in the domain of interest speed-ups over GPUs are achieved. The potential impact is that this system promises to be the basis for scalable long timescale MD with commodity clusters. Chunshu Wu, Chen Yang 0010, Sahan Bandara, Tong Geng, Anqi Guo, Pouya Haghi, Ang Li 0006, Martin C. Herbordt |
IEEE Trans. Computers | 8 |
| 2024 | Long-Range MD Electrostatics Force Computation on FPGAsabstractStrong scaling of long-range electrostatic force computation, which is a central concern of long timescale molecular dynamics simulations, is challenging for CPUs and GPUs due to its complex communication structure and global communication requirements. The scalability challenge is seen especially in small simulations of tens to hundreds of thousands of atoms that are of interest to many important applications such as physics-driven drug discovery. FPGA clusters, with their direct, tightly coupled, low-latency interconnects, are able to address these requirements. For FPGA MD clusters to be effective, however, single device performance must also be competitive. In this work, we leverage the inherent benefits of FPGAs to implement a long-range electrostatic force computation architecture. We present an overall framework with numerous algorithmic, mapping, and architecture innovations, including a unified interleaved memory, a spatial scheduling algorithm, and a design for seamless integration with the larger MD system. We examine a number of alternative configurations based on different resource allocation strategies and user parameters. We show that the best configuration of this architecture, implemented on an Intel Agilex FPGA, can achieve$2124 ns$and$287 ns$of simulated time per day of wall-clock time for the two molecular dynamics benchmarks DHFR and ApoA1; simulating 23K and 92K particles, respectively. Sahan Bandara, Anthony Ducimo, Chunshu Wu, Martin C. Herbordt |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | Software-Hardware Co-design of Heterogeneous SmartNIC System for Recommendation Models Inference and TrainingabstractDeep Learning Recommendation Models (DLRMs) are important applications in various domains and have evolved into one of the largest and most important machine learning applications. With their trillions of parameters necessarily exceeding the high bandwidth memory (HBM) capacity of GPUs, ever more massive DLRMs require large-scale multi-node systems for distributed training and inference. However, these all suffer from the all-to-all communication bottleneck, which limits scalability. Anqi Guo, Yuchen Hao, Chunshu Wu, Pouya Haghi, Zhenyu Pan, Min Si, Dingwen Tao, Ang Li 0006, Martin C. Herbordt, Tong Geng |
ICS | 9 |
| 2023 | FLASH: FPGA-Accelerated Smart Switches with GCN Case StudyabstractSome communication switches, e.g., the Mellanox SHArP and those in the IBM BlueGene clusters, are augmented to process packets at the application level with fixed-function collectives. This approach, however, lacks flexibility, which limits their applicability in diverse and dynamic workloads. Recently, a new type of programmable packet processor, which uses high-level languages, e.g., P4, has emerged as a possible candidate. P4-based switches, however, fall short in certain applications, including machine learning, where capabilities not currently supported by P4 are needed. These include more complex calculation, such as sparse computation and fused multiply-accumulate, data-intensive floating point operations, data reuse, and significant memory. The problem addressed here is that such a switch augmentation needs to support: a large amount of state, significant flexible compute capability, and ease of programming, all while maintaining full functionality, including ensuring high throughput, and demonstrating utility. Pouya Haghi, William Krska, Cheng Tan 0002, Tong Geng, Po-Hao Chen 0002, Connor Greenwood, Anqi Guo, Thomas M. Hines, Chunshu Wu, Ang Li 0006, Anthony Skjellum, Martin C. Herbordt |
ICS | 12 |
| 2023 | FASDA: An FPGA-Aided, Scalable and Distributed Accelerator for Range-Limited Molecular DynamicsabstractConducting long-timescale simulations of small molecules using Molecular Dynamics (MD) is crucial in drug design. However, traditional methods to accelerate the process, including ASICs or GPUs, have limitations. ASIC solutions are not always generally available, while GPU solutions may not scale when processing small molecules. FPGAs are both communication processors and accelerators, with tight coupling between these capabilities, and so could be used to address strong scaling in this domain. Chunshu Wu, Tong Geng, Anqi Guo, Sahan Bandara, Pouya Haghi, Chuan Liu 0001, Ang Li 0006, Martin C. Herbordt |
SC | 8 |
| 2023 | Introduction to the Special Section on FCCM 2022abstractNo abstract available. Jing Jane Li, Martin C. Herbordt |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2022 | FCsN: A FPGA-Centric SmartNIC Framework for Neural NetworksabstractNetwork communication is increasingly becoming the performance bottleneck for scaled-out HPC and warehouse applications, as enormous CPU processing is devoted to packet processing, contributing to long latencies. To reduce this latency, advanced network interface cards known as SmartNICs have been introduced to handle networking functions. Dozens of commercial FPGA-based SmartNICs have been released (e.g., [1] – [3] and see surveys [4] , [5] ). Other commercial SmartNICs have been developed also with the aim of near-network processing [6] – [9] . There is also prior art that uses SmartNICs as compute resources [10] , [11] . For instance, COPA [12] , INCA [13] , sPIN [14] provide a portable programming model to offload simple packet processing. Other work (e.g., [15] – [21] ) supports collectives in FPGA-based hardware. Anqi Guo, Tong Geng, Yongan Zhang, Pouya Haghi, Chunshu Wu, Cheng Tan 0002, Yingyan (Celine) Lin, Ang Li 0006, Martin C. Herbordt |
FCCM | 9 |
| 2022 | COPA Use Case: Distributed Secure Joint ComputationabstractData centers provide good environments for distributed computing as they are easily accessible and may have low-latency communication between nodes [1] ; often, however, performance is limited by network bandwidth. These network bottlenecks drive the need for alternative communication resources to improve performance of large-scale applications. SmartNICs [2] – [4] have been introduced to perform the same tasks of standard NICs, but contain additional resources to allow for network function optimization with additional hardware. Adoption of SmartNICs continues to increase as a means to accelerate network functions and offload packet processing tasks away from CPU resources [5] – [13] . Rushi Patel, Pouya Haghi, Shweta Jain 0005, Andriy Kot, Venkata Krishnan, Mayank Varia, Martin C. Herbordt |
FCCM | 7 |
| 2022 | H-GCN: A Graph Convolutional Network Accelerator on Versal ACAP ArchitectureabstractGraph Neural Networks (GNNs) have drawn tremendous attention due to their unique capability to extend Machine Learning (ML) approaches to applications broadly-defined as having unstructured data, especially graphs. Compared with other Machine Learning (ML) modalities, the acceleration of Graph Neural Networks (GNNs) is more challenging due to the irregularity and heterogeneity derived from graph typologies. Existing efforts, however, have focused mainly on handling graphs' irregularity and have not studied their heterogeneity. To this end we propose H-GCN, a PL (Programmable Logic) and AIE (AI Engine) based hybrid accelerator that leverages the emerging heterogeneity of Xilinx Versal Adaptive Compute Acceleration Platforms (ACAPs) to achieve high-performance GNN inference. In particular, H-GCN partitions each graph into three subgraphs based on its inherent heterogeneity, and processes them using PL and AIE, respectively. To further improve performance, we explore the sparsity support of AIE and develop an efficient density-aware method to automatically map tiles of sparse matrix-matrix multiplication (SpMM) onto the systolic tensor array. Compared with state-of-the-art GCN accelerators, H-GCN achieves, on average, speedups of 1.1~2.3x. Chengming Zhang 0006, Tong Geng, Anqi Guo, Jiannan Tian, Martin C. Herbordt, Ang Li 0006, Dingwen Tao |
FPL | 5 |
| 2022 | A Framework for Neural Network Inference on FPGA-Centric SmartNICsabstractFPGA-based SmartNICs offer great potential to significantly improve the performance of high-performance computing and warehouse data processing by tightly coupling support for reconfigurable data-intensive computation with cross-node communication thereby mitigating the von Neumann bottleneck. Existing work however has generally been limited in that it assumes an accelerator model where kernels are offloaded to SmartNICs with most control tasks left to the CPUs. This leads to frequent waiting reduced performance and scaling challenges. In this work we propose a new distributive data-centric computing framework named FCsN for reconfigurable SmartNIC-based systems. Through a lightweight task circulation execution model and its implementation architecture FCsN allows the complete detaching of NN kernel execution control logic system scheduling and network communication to the SmartNICs. This boosts performance by (i) avoiding control dependency with CPUs and (ii) supporting streaming NN kernel execution and network communication at line rate and in a very fine-grained manner. We demonstrate the efficiency and flexibility of FCsN using various types of neural network kernels and applications including deep neural networks (DNN) and graph neural networks (GNN) as these last are both irregular and data intensive they offer an especially robust demonstration. Evaluations using commonly-used neural network models and graph datasets show that a system with FCsN can achieve 10 × speedups over the MPI-based standard CPU baselines Anqi Guo, Tong Geng, Yongan Zhang, Pouya Haghi, Chunshu Wu, Cheng Tan 0002, Yingyan (Celine) Lin, Ang Li 0006, Martin C. Herbordt |
FPL | 9 |
| 2022 | Optimized Mappings for Symmetric Range-Limited Molecular Force Calculations on FPGAsabstractIn N-body applications, the efficient evaluation of range-limited forces depends on applying certain constraints, including a cut-off radius and force symmetry (Newton's Third Law). When computing the pair-wise forces in parallel, finding the optimal mapping of particles and computations to memories and processors is surprisingly challenging, but can result in greatly reduced data movement and computation. Despite FPGAs having a distinct compute model (BRAMs/network/pipelines) from CPUs and ASICs, mappings on FPGAs have not previously been studied in depth: it was thought that the half-shell method was preferred. In this work, we find that the Manhattan method is sur-prisingly compatible with FPGA hardware. With the cache overlapping technique proposed in this paper, the ultra-fine-grained data access demanded by the Manhattan method can be satisfied, despite the fact that the memory blocks on FPGAs appear to be insufficiently fine-grained. We further demonstrate that, compared to the traditional baseline half-shell method, approximately a half of the filters (preprocessors) can be removed without performance degradation. For communication, the amount of data transferred can be reduced by 40% - 75% in the most common multi-FPGA scenarios. Moreover, data transfers are almost perfectly balanced along all directions, and the optimization requires only minimal hardware resources. The practical consequence is that nearly 2 x to 4 x the workload can be handled without upgrading the network connections between FPGAs. This is a critical finding given the relatively limited bandwidth available in many common accelerator boards and the strong-scaling applications to which FPGA clusters are being applied. Chunshu Wu, Sahan Bandara, Tong Geng, Anqi Guo, Pouya Haghi, Vipin Sachdeva, Woody Sherman, Martin C. Herbordt |
FPL | 8 |
| 2022 | Reconfigurable switches for high performance and flexible MPI collectivesabstractAbstract There has been much effort in offloading MPI collective operations into hardware. But while NIC‐based collective acceleration is well‐studied, offloading their processing into the switching fabric, despite numerous advantages, has been much more limited. A major problem with fixed logic implementations is that either only a fraction of the possible collective communication is accelerated or that logic is wasted in the applications that do not need a particular capability. Using reconfigurable logic has numerous advantages: exactly the required operations can be implemented; the level of desired performance can be specified; and new, possibly complex, operations can be defined and implemented. We have designed an in‐switch collective accelerator,MPI‐FPGA, and demonstrated its use with seven MPI collectives and over a set of benchmarks and proxy applications (MiniApps). The accelerator uses a novel two‐level switch design containing fully pipelined vectorized aggregation logic units. Essential to this work is providing support for sub‐communicator collectives that enables communicators of arbitrary shape, and that is scalable to large systems. A streaming interface improves the performance for long messages. While this reconfigurable design is generally applicable, we prototype it with an FPGA‐centric cluster. A sampleMPI‐FPGAdesign in a direct network achieves considerable speedups over conventional clusters in the most likely scenarios. We also present results for indirect networks with reconfigurable high‐radix switches and show that this approach is competitive withSHArPtechnology for the subset of operations thatSHArPsupports.MPI‐FPGAis fully integrated into MPICH and is transparent to MPI applications. Pouya Haghi, Anqi Guo, Qingqing Xiong, Chen Yang 0010, Tong Geng, Justin T. Broaddus, Ryan J. Marshall, Derek Schafer, Anthony Skjellum, Martin C. Herbordt |
Concurr. Comput. Pract. Exp. | 10 |
| 2022 | The Future of FPGA Acceleration in Datacenters and the CloudabstractIn this article, we survey existing academic and commercial efforts to provide Field-Programmable Gate Array (FPGA) acceleration in datacenters and the cloud. The goal is a critical review of existing systems and a discussion of their evolution from single workstations with PCI-attached FPGAs in the early days of reconfigurable computing to the integration of FPGA farms in large-scale computing infrastructures. From the lessons learned, we discuss the future of FPGAs in datacenters and the cloud and assess the challenges likely to be encountered along the way. The article explores current architectures and discusses scalability and abstractions supported by operating systems, middleware, and virtualization. Hardware and software security becomes critical when infrastructure is shared among tenants with disparate backgrounds. We review the vulnerabilities of current systems and possible attack scenarios and discuss mitigation strategies, some of which impact FPGA architecture and technology. The viability of these architectures for popular applications is reviewed, with a particular focus on deep learning and scientific computing. This work draws from workshop discussions, panel sessions including the participation of experts in the reconfigurable computing field, and private discussions among these experts. These interactions have harmonized the terminology, taxonomy, and the important topics covered in this manuscript. Christophe Bobda, Joel Mandebi, Paul Chow, Mohammad Ewais, Naif Tarafdar, Juan Camilo Vega, Kenneth Eguro, Dirk Koch, Suranga Handagala, Miriam Leeser, Martin C. Herbordt, Hafsah Shahzad, H. Peter Hofstee, Burkhard Ringlein, Jakub Szefer, Ahmed Sanaullah, Russell Tessier |
ACM Trans. Reconfigurable Technol. Syst. | 11 |
| 2021 | Particle Mesh Ewald for Molecular Dynamics in OpenCL on an FPGA ClusterabstractMolecular Dynamics (MD) simulations play a central role in physics-driven drug discovery. MD applications often use the Particle Mesh Ewald (PME) algorithm to accelerate electrostatic force computations, but efficient parallelization has proven difficult due to the high communication requirements of distributed 3D FFTs. In this paper, we present the design and implementation of a scalable PME algorithm that runs on a cluster of Intel Stratix 10 FPGAs and can handle FFT sizes appropriate to address real-world drug discovery projects (grids up to 1283). To our knowledge, this is the first work to fully integrate all aspects of the PME algorithm (charge spreading, 3D FFT/IFFT, and force interpolation) within a distributed FPGA framework. The design is fully implemented with OpenCL for flexibility and ease of development and uses 100 Gbps links for direct FPGA-to-FPGA communications without the need for host interaction. We present experimental data up to 4 FPGAs (e.g., 206 microseconds per timestep for a 65536 atom simulation and 643 3D FFT), outperforming GPUs. Additionally, we discuss design scalability on clusters with differing topologies up to 64 FPGAs (with expected performance greater than all known GPU implementations) and integration with other hardware components to form a complete molecular dynamics application. We predict best-case performance of 6.6 microseconds per timestep on 64 FPGAs. Lawrence C. Stewart, Carlo Pascoe, Emery Davis, Brian W. Sherman, Martin C. Herbordt, Vipin Sachdeva |
FCCM | 5 |
| 2021 | Upgrade of FPGA Range-Limited Molecular Dynamics to Handle Hundreds of ProcessorsabstractWith the current pandemic, the central role that Molecular Dynamics simulation (MD) plays in drug discovery makes advances in MD performance urgent. Recent work has demonstrated that among COTS devices only FPGA-centric clusters can scale beyond a few processors for relevant targets; other work has shown that single FPGA performance compares favorably to that of a GPU. In this study we demonstrate that an additional factor of 4× performance can be achieved which results in a factor of 5× speed up over a GPU. The problem addressed is that the designs of the last decade no longer scale when the number of processing pipelines grows from around ten to the hundreds. We begin by systematically evaluating existing work, exposing its flaws, and proposing a series of new design solutions. There are four major contributions. First, we address the massive routing problem by augmenting the design with three minimal networks in logic and latency. Second, we have developed a novel asynchronous out-of-order communication mechanism that removes nearly all bubbles from the routing networks. Third, we find that inverting the standard particle access algorithm results in improved locality and performance. Finally, we have created a custom numerical format that increases precision while saving space and logic. Chunshu Wu, Tong Geng, Sahan Bandara, Chen Yang 0010, Vipin Sachdeva, Woody Sherman, Martin C. Herbordt |
FCCM | 7 |
| 2021 | I-GCN: A Graph Convolutional Network Accelerator with Runtime Locality Enhancement through IslandizationabstractGraph Convolutional Networks (GCNs) have drawn tremendous attention in the past three years. Compared with other deep learning modalities, high-performance hardware acceleration of GCNs is as critical but even more challenging. The hurdles arise from the poor data locality and redundant computation due to the large size, high sparsity, and irregular non-zero distribution of real-world graphs. Tong Geng, Chunshu Wu, Yongan Zhang, Cheng Tan 0002, Chenhao Xie 0001, Haoran You, Martin C. Herbordt, Yingyan (Celine) Lin, Ang Li 0006 |
MICRO | 7 |
| 2021 | O3BNN-R: An Out-of-Order Architecture for High-Performance and Regularized BNN InferenceabstractBinarized Neural Networks (BNN), which significantly reduce computational complexity and memory demand, have shown potential in cost- and power-restricted domains, such as IoT and smart edge-devices, where reaching certain accuracy bars is sufficient and real-time is highly desired. In this article, we demonstrate that the highly-condensed BNN model can be shrunk significantly by dynamically pruning irregular redundant edges. Based on two new observations on BNN-specific properties, an out-of-order (OoO) architecture, O3BNN-R, which can curtail edge evaluation in cases where the binary output of a neuron can be determined early at runtime during inference, is proposed. Similar to instruction level parallelism (ILP), fine-grained, irregular, and runtime pruning opportunities are traditionally presumed to be difficult to exploit. To further enhance the pruning opportunities, we conduct an algorithm/architecture co-design approach where we augment the loss function during the training stage with specialized regularization terms favoring edge pruning. We evaluate our design on an embedded FPGA using networks that include VGG-16, AlexNet for ImageNet, and a VGG-like network for Cifar-10. Results show that O3BNN-R without regularization can prune, on average, 30 percent of the operations, without any accuracy loss, bringing 2.2× inference-speedup, and on average 34× energy-efficiency improvement over state-of-the-art BNN implementations on FPGA/GPU/CPU. With regularization at training, the performance is further improved, on average, by 15 percent. Tong Geng, Ang Li 0006, Chunshu Wu, Runbin Shi, Wei Wu 0016, Martin C. Herbordt |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2020 | FP-AMG: FPGA-Based Acceleration Framework for Algebraic Multigrid SolversabstractPartial Differential Equations (PDEs) are fundamental to many real-world scientific computing applications and so their optimization has undergone decades of study. Algebraic multigrid (AMG) is one of the most well-known solvers, being widely adopted in High Performance Computing (HPC) due to its good scalability. Acceleration of AMG is known to be very challenging, due to the following reasons: (1) irregular computation patterns, (2) random memory access, and (3) a large number of kernels with various computation types. To the best of our knowledge, there is no prior work on FPGA-based acceleration of AMG. To tackle these challenges, we propose an efficient FPGA-based reconfigurable framework, called FP-AMG, for high-performance AMG calculation. In order to obtain full pipeline utilization, we propose a novel and scalable architecture that can be reused for all kernels in AMG. Given that AMG is strictly memory-bound, we propose algorithmic and architectural optimizations to ensure nearly ideal use of memory bandwidth. The efficiency of FP-AMG is evaluated with six well-known benchmarks on two FPGA devices: one with and one without high bandwidth memory (HBM). The experimental results are compared with a highly optimized Intel Xeon E5-2680-V4 implementation of the state-of-the-art HYPRE library. Our experiments show that FP-AMG can achieve average speedups of $ 2.5\times$ and $ 6.6\times$, for FPGAs without and with HBM, respectively. Pouya Haghi, Tong Geng, Anqi Guo, Martin C. Herbordt |
FCCM | 5 |
| 2020 | Accelerating MPI Collectives with FPGAs in the Network and Novel Communicator SupportabstractMPI collective operations can often be performance killers in HPC applications; we seek to solve this bottleneck by offloading them to reconfigurable hardware within the switch itself, rather than, e.g., the NIC. We have designed a hardware accelerator MPI-FPGA to implement six MPI collectives in the network. Preliminary results show that MPI-FPGA achieves $10 \times $ speedup in the most likely scenarios over conventional clusters. We introduce a novel mechanism that enables the hardware to support a large number of communicators of arbitrary shape, and that is scalable to very large systems. MPI-FPGA is fully integrated into MPICH and so transparent to MPI applications. Qingqing Xiong, Chen Yang 0010, Pouya Haghi, Anthony Skjellum, Martin C. Herbordt |
FCCM | 5 |
| 2020 | Secret Sharing MPC on FPGAs in the DatacenterabstractMulti-Party Computation (MPC) is a technique enabling data from several sources to be used in a secure computation revealing only the result while protecting the original data, facilitating shared utilization of data sets gathered by different entities. The presence of Field Programmable Gate Array (FPGA) hardware in datacenters can provide accelerated computing as well as low latency, high bandwidth communication that bolsters the performance of MPC and lowers the barrier to using MPC for many applications. In this work, we propose a Secret Sharing FPGA design based on the protocol described by Araki et al. We compare our hardware design to the original authors' software implementations of Secret Sharing and to work accelerating MPC protocols based on Garbled Circuits with FPGAs. Our conclusion is that Secret Sharing in the datacenter is competitive and when implemented on FPGA hardware was able to use at least 10× fewer computer resources than the original work using CPUs. Pierre-François W. Wolfe, Rushi Patel, Robert Munafo, Mayank Varia, Martin C. Herbordt |
FPL | 5 |
| 2020 | CSB-RNN: a faster-than-realtime RNN acceleration framework with compressed structured blocksabstractRecurrent neural networks (RNNs) have been widely adopted in temporal sequence analysis, where realtime performance is often in demand. However, RNNs suffer from heavy computational workload as the model often comes with large weight matrices. Pruning (a model compression method) schemes have been proposed for RNNs to eliminate the redundant (close-to-zero) weight values. On one hand, the non-structured pruning methods achieve a high pruning rate but introducing computation irregularity (random sparsity), which is unfriendly to parallel hardware. On the other hand, hardware-oriented structured pruning suffers from low pruning rate due to restricted constraints on allowable pruning structure. Runbin Shi, Peiyan Dong, Tong Geng, Yuhao Ding, Hayden Kwok-Hay So, Martin C. Herbordt, Ang Li 0006, Yanzhi Wang 0001 |
ICS | 7 |
| 2020 | AWB-GCN: A Graph Convolutional Network Accelerator with Runtime Workload RebalancingabstractDeep learning systems have been successfully applied to Euclidean data such as images, video, and audio. In many applications, however, information and their relationships are better expressed with graphs. Graph Convolutional Networks (GCNs) appear to be a promising approach to efficiently learn from graph data structures, having shown advantages in many critical applications. As with other deep learning modalities, hardware acceleration is critical. The challenge is that real-world graphs are often extremely large and unbalanced; this poses significant performance demands and design challenges. In this paper, we propose Autotuning-Workload-Balancing GCN (AWB-GCN) to accelerate GCN inference. To address the issue of workload imbalance in processing real-world graphs, three hardware-based autotuning techniques are proposed: dynamic distribution smoothing, remote switching, and row remapping. In particular, AWB-GCN continuously monitors the sparse graph pattern, dynamically adjusts the workload distribution among a large number of processing elements (up to 4K PEs), and, after converging, reuses the ideal configuration. Evaluation is performed using an Intel D5005 FPGA with five commonly-used datasets. Results show that 4K-PE AWB-GCN can significantly elevate PE utilization by 7.7× on average and demonstrate considerable performance speedups over CPUs (3255×), GPUs (80.3×), and a prior GCN accelerator (5.1×). Tong Geng, Ang Li 0006, Runbin Shi, Chunshu Wu, Pouya Haghi, Antonino Tumeo, Shuai Che, Steven K. Reinhardt, Martin C. Herbordt |
MICRO | 11 |
| 2020 | FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA ClustersabstractDeep convolutional Neural Networks (CNNs) have revolutionized numerous applications, but the demand for ever more performance remains unabated. Scaling CNN computations to larger clusters is generally done by distributing tasks in batch mode using methods such as distributed synchronous SGD. Among the issues with this approach is that, to make the distributed cluster work with high utilization, the workload distributed to each node must be large; this implies nontrivial growth in the SGD mini-batch size. In this article we propose a framework, called FPDeep, which uses a hybrid of model and layer parallelism to configure distributed reconfigurable clusters to train CNNs. This approach has numerous benefits. First, the design does not suffer from performance loss due to batch size growth. Second, work and storage are balanced among nodes through novel workload and weight partitioning schemes. Part of the mechanism is the surprising finding that it is preferable to store excess weights in neighboring devices rather than in local off-chip memory. Third, the entire system is a fine-grained pipeline. This leads to high parallelism and utilization and also minimizes the time that features need to be cached while waiting for back-propagation. As a result, storage demand is reduced to the point where only on-chip memory is used for the convolution layers. And fourth, we find that the simplest topology, a 1D array, is preferred for interconnecting the FPGAs thus enabling widespread applicability. We evaluate FPDeep with the Alexnet, VGG-16, and VGG-19 benchmarks. Results show that FPDeep has good scalability to a large number of FPGAs, with the limiting factor being the FPGA-to-FPGA bandwidth. But with 250 Gb/s bidirectional bandwidth per FPGA, which is easily supported by current generation FPGAs, FPDeep performance shows linearity up to 100 FPGAs. Energy efficiency is evaluated with respect to GOPs/J. FPDeep provides, on average, 6.4× higher energy efficiency than comparable GPU servers. Tong Geng, Ang Li 0006, Xi Jin 0002, Martin C. Herbordt |
IEEE Trans. Computers | 5 |
| 2019 | LP-BNN: Ultra-low-Latency BNN Inference with Layer ParallelismabstractHigh inference latency seriously limits the deployment of DNNs in real-time domains such as autonomous driving, robotic control, and many others. To address this emerging challenge, researchers have proposed approximate DNNs with reduced precision, e.g., Binarized Neural Networks (BNNs). While BNNs can be built to have little loss in accuracy, latency reduction still has much room for improvement. In this paper, we propose a single-FPGA-based BNN accelerator that achieves microsecond-level ultra-low-latency inference of ImageNet, LP-BNN. We obtain this performance via several design optimizations. First, we optimize the network structure by removing Batch Normalization (BN) functions which leads to significant latency in BNNs without any loss on accuracy. Second, we propose a parameterized architecture which is based on layer parallelism and supports nearly perfect load balancing for various types of BNNs. Third, we fuse all the convolution layers and the first fully connected layer. We process them in parallel through fine-grained inter-layer pipelining. With our proposed accelerator, the inference of binarized AlexNet, VGGNet, and ResNet are completed within 21.5us, 335us, and 67.8us respectively, with no loss in accuracy as compared with other BNN implementations. Tong Geng, Chunshu Wu, Chen Yang 0010, Shuaiwen Song, Ang Li 0006, Martin C. Herbordt |
ASAP | 7 |
| 2019 | Accelerating AP3M-Based Computational Astrophysics Simulations with Reconfigurable ClustersabstractIn this paper, we present a case study of using a reconfigurable computing cluster to accelerate AP3M-based computational astrophysics simulations. AP3M is an adaptive particle-particle, particle-mesh method. Many computational astrophysics simulations are based on this method. AP3M can dynamically and adaptively apply computational resources non-uniformly to emphasize regions of interest. Therefore, AP3M can be faster and more energy-efficient than the traditional P3M (particle-particle, particle-mesh) approach. However, the dynamic and pointer-based data structure used by AP3M makes it extremely difficult to accelerate with FPGAs. In this work, we use a custom data structure and hardware kernel to overcome these challenges. All CPU-based dynamic and pointer-based tasks are mapped to FPGAs. Our experiments show that a single FPGA outperforms a Xeon E5-2660 CPU server (8 cores) by from 21x to 23x depending on problem size and data distribution. Tong Geng, Xi Jin 0002, Martin C. Herbordt |
ASAP | 4 |
| 2019 | Molecular Dynamics Range-Limited Force Evaluation Optimized for FPGAsabstractFPGA Molecular Dynamics was much studied from 2004-2010. Due to limited chip resources of that era, and the inherent variety and complexity of tasks comprising Molecular Dynamics simulations (MD), those FPGA accelerators relied on host or embedded processors to organize and pre-process input and output data. This introduced long latency for data movement between simulation iterations and, as technology advanced, drastically limited performance. Current generation FPGAs are equipped not only with abundant on-chip resources, but also have hardware support for floating point operations; these advances provide an opportunity for creating self-contained MD simulation systems on a single device. In this paper, we demonstrate such a system based on the range-limited force, which comprises 90% of the flops in a typical MD simulation. It features online particle-pair generation, hundreds of force evaluation pipelines, motion update, and particle data migration. We integrate into OpenMM and find that, for a representative dataset (liquid argon with 20K atoms), we can achieve a simulation throughput of 1.4us/day with a single FPGA, more than twice the performance of a comparable generation GPU. The bulk of the work presented here explores the design of an independent MD range-limited force evaluation system tailored for modern FPGAs without data exchange with any off-chip devices. The primary contributions are the designs of the new features, the methods for coupling those features into an integrated system, and, especially, the analysis of the most likely mappings among particles/cells, on-chip memories (BRAMs), and on-chip compute units (pipelines). Chen Yang 0010, Tong Geng, Charles Lin, Jiayi Sheng, Vipin Sachdeva, Woody Sherman, Martin C. Herbordt |
ASAP | 8 |
| 2019 | FP-AMR: A Reconfigurable Fabric Framework for Adaptive Mesh Refinement ApplicationsabstractAdaptive mesh refinement (AMR) is one of the most widely used methods in High Performance Computing accounting a large fraction of all supercomputing cycles. AMR operates by dynamically and adaptively applying computational resources non-uniformly to emphasize regions of the model as a function of their complexity. Because AMR generally uses dynamic and pointer-based data structures, acceleration is challenging, especially in hardware. As far as we are aware there has been no previous work published on accelerating AMR with FPGAs. In this paper, we introduce a reconfigurable fabric framework called FP-AMR. The work is in two parts. In the first FP-AMR offloads the bulk per-timestep computations to the FPGA; analogous systems have previously done this with GPUs. In the second part we show that the rest of the CPU-based tasks-including particle mesh mapping, mesh refinement, and coarsening-can also be mapped efficiently to the FPGA. We have evaluated FP-AMR using the widely used program AMReX and found that a single FPGA outperforms a Xeon E5-2660 CPU server (8 cores) by from 21x -23x depending on problem size and data distribution. Tong Geng, Xi Jin 0002, Martin C. Herbordt |
FCCM | 4 |
| 2019 | GhostSZ: A Transparent FPGA-Accelerated Lossy Compression FrameworkabstractHigh-performance computing (HPC) applications often generate enormous amounts of data that must be transferred for check-pointing, in situ processing, or post-execution analysis. To reduce the related network traffic and storage consumption, lossy compression schemes that target scientific data are often used. SZ compression emerged three years ago and has gained much attention because of its high compression ratio. However, performing SZ compression can take half a day per Terabyte of data; this could be a drawback to adoption. We propose GhostSZ an FPGA framework for accelerating tasks in SZ at line rate, and so transparently. The critical problem to be overcome is the tight data dependence central to SZ. GhostSZ solves this with a data transfer path having novel staged hardware. We test our implementation with both synthetic and real HPC application data and show 9.5×-80× core versus pipeline speedup over the optimized production version running on a state-of-the-art CPU and 8.2× per chip. Much of the variance in performance is due to the FPGA already running at line rate and so benefiting less from optimizations applicable to the CPU only on the most favorable data sets. The significance of this work is the possibility of a major reduction in required networking and storage in HPC installations. For example, using GhostSZ, fewer than 10 FPGAs would be sufficient to handle the entire I/O bandwidth of the top entry on the latest IO-500 list. Qingqing Xiong, Rushi Patel, Chen Yang 0010, Tong Geng, Anthony Skjellum, Martin C. Herbordt |
FCCM | 6 |
| 2019 | O3BNN: an out-of-order architecture for high-performance binarized neural network inference with fine-grained pruningabstractBinarized Neural Networks (BNN) have drawn tremendous attention due to significantly reduced computational complexity and memory demand. They have especially shown great potential in cost- and power-restricted domains, such as IoT and smart edge-devices, where reaching a certain accuracy bar is often sufficient, and real-time is highly desired. Tong Geng, Chunshu Wu, Chen Yang 0010, Wei Wu 0016, Ang Li 0006, Martin C. Herbordt |
ICS | 7 |
| 2019 | BSTC: a novel binarized-soft-tensor-core design for accelerating bit-based approximated neural netsabstractBinarized neural networks (or BNNs) promise tremendous performance improvement over traditional DNNs through simplified bit-level computation and significantly reduced memory access/storage cost. In addition, it has advantages of low-cost, low-energy, and high-robustness, showing great potential in resources-constrained, volatile, and latency-critical applications, which are critical for future HPC, cloud, and edge applications. However, the promised significant performance gain of BNN inference has never been fully demonstrated on general-purpose processors, particularly on GPUs, due to: (i) the challenge of extracting and leveraging sufficient finegrained bit-level-parallelism to saturate GPU cores when the batch size is small; (ii) the fundamental design conflict between bit-based BNN algorithm and word-based architecture; and (iii) architecture & performance unfriendly to BNN network design. To address (i) and (ii), we propose a binarized-soft-tensor-core as a software-hardware codesign approach to construct bit-manipulation capability for modern GPUs and thereby effectively harvest bit-level-parallelism (BLP). To tackle (iii), we propose intra- and inter-layer fusion techniques so that the entire BNN inference execution can be packed into a single GPU kernel, and so avoid the high-cost of frequent launching and releasing. Experiments show that our Singular-Binarized-Neural-Network (SBNN) design can achieve over 1000X speedup for raw inference latency over the state-of-the-art full-precision BNN inference for AlexNet on GPUs. Comparisons with CPU, GPU, FPGA and Xeon-Phi demonstrate the effectiveness of our design. SBNN is opensourced and available at https://github.com/uuudown/SBNN. Ang Li 0006, Tong Geng, Martin C. Herbordt, Shuaiwen Song, Kevin J. Barker |
SC | 4 |
| 2019 | Fully integrated FPGA molecular dynamics simulationsabstractThe implementation of Molecular Dynamics (MD) on FPGAs has received substantial attention. Previous work, however, has consisted of either proof-of-concept implementations of components, usually the range-limited force; full systems, but with much of the work shared by the host CPU; or prototype demonstrations, e.g., using OpenCL, that neither implement a whole system nor have competitive performance. In this paper, we present what we believe to be the first full-scale FPGA-based simulation engine, and show that its performance is competitive with a GPU (running Amber in an industrial production environment). The system features on-chip particle data storage and management, short- and long-range force evaluation, as well as bonded forces, motion update, and particle migration. Other contributions of this work include exploring numerous architectural trade-offs and analysis of various mappings schemes among particles/cells and the various on-chip compute units. The potential impact is that this system promises to be the basis for long timescale Molecular Dynamics with a commodity cluster. Chen Yang 0010, Tong Geng, Rushi Patel, Qingqing Xiong, Ahmed Sanaullah, Chunshu Wu, Jiayi Sheng, Charles Lin, Vipin Sachdeva, Woody Sherman, Martin C. Herbordt |
SC | 12 |
| 2018 | FPDeep: Acceleration and Load Balancing of CNN Training on FPGA ClustersabstractFPGA-based CNN accelerators have advantages in flexibility and power efficiency and so are being deployed by a number of cloud computing service providers, including Microsoft, Amazon, Tencent, and Alibaba. Given the increasing complexity of neural networks, however, it is becoming challenging to efficiently map CNNs to multi-FPGA platforms. In this work, we present a scalable framework, FPDeep, which helps engineers map a specific CNN's training logic to a multi-FPGA cluster or cloud and to build RTL implementations for the target network. With FPDeep, multi-FPGA accelerators work in a deeply-pipelined manner using a simple 1-D topology; this enables the accelerators to map directly onto many existing platforms, including Catapult, Catapult2, and almost any tightly-coupled FPGA cluster. FPDeep uses two mechanisms to facilitate high-performance and energy-efficiency. First, FPDeep provides a strategy to balance workload among FPGAs, leading to improved utilization. Second, training of CNNs is executed in a fine-grained inter- and intra-layer pipelined manner, minimizing the time that features need to remain available while waiting for back-propagation. This reduces the storage demand to where only on-chip memory is required for convolution layers. Experiments show that FPDeep has good scalability to a large number of FPGAs, with the limiting factor being the FPGA-to-FPGA bandwidth. Using six transceivers per FPGA, FPDeep shows linearity up to 60 FPGAs. We evaluate energy efficiency in GOPs/J and find that FPDeep provides up to 3.4 times higher energy efficiency than the Tesla K80 GPU. Tong Geng, Ahmed Sanaullah, Chen Yang 0010, Rushi Patel, Martin C. Herbordt |
FCCM | 7 |
| 2018 | High Performance Dynamic Communication on Reconfigurable ClustersabstractFPGA clusters with the FPGAs directly linked through their Multi-Gigabit Transceiver (MGT) ports have a proven advantage over other commodity architectures for communication-bound applications. We find that the standard wormhole routers need some modification to be appropriate for clusters with tightly coupled FPGAs, and create such a router. We generalize this router so that it is parameterized by several parameters including routing algorithm, arbitration policy, virtual channels, and buffers. We have evaluated these designs with respect to a number standard communication patterns and packet sizes. These results enable selection of the appropriate router for any resource budget. Finally, We find that the optimality of the router design varies significantly with workload. We observe that for a 512 FPGA cluster, connected in an 8^3 torus, compared with the router configuration with the best average performance, application-aware router configurations reduce average batch latency by 3%, improve the throughput by 6% on average, and improve area consumption by 50%. Jiayi Sheng, Chen Yang 0010, Martin C. Herbordt |
FCCM | 4 |
| 2018 | A Framework for Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters with Work and Weight Load BalancingabstractTo improve flexibility and energy efficiency of Convolutional Neural Networks, a number of cloud computing service providers-including Microsoft, Amazon, and Alibaba-are using FPGA-based CNN accelerators. However, the growing size and complexity of neural networks, coupled with communication and off-chip memory bottlenecks, make it increasingly difficult for multi-FPGA designs to achieve high resource utilization and performance, especially when training. In this work, we present new results for a scalable framework, FPDeep, which helps users efficiently map CNN training logic to multiple FPGAs and automatically generates the resulting RTL implementation. FPDeep is equipped with two mechanisms to facilitate high-performance and energy-efficient training. First, FPDeep improves DSP slice utilization across FPGAs by balancing workload using dedicated partition and mapping strategies. Second, only on-chip memory is used in the CONV layers: a) FPDeep balances CNN weight allocation among FPGAs to improve BRAM utilization; b) training of CNNs is executed in a fine-grained pipelined manner, minimizing the time features need to be cached while waiting for back-propagation leading to a reduced storage demand. We evaluate our framework by training AlexNet, VGG-16, and VGG-19. Experimental results show FPDeep has good scalability to a large number of FPGAs, with the limiting factor being the inter-FPGA bandwidth. With 6 transceivers per FPGA, FPDeep shows linearity up to 83 FPGAs. FPDeep provides, on average, 6.36x higher energy efficiency than GPU servers. Tong Geng, Ahmed Sanaullah, Chen Yang 0010, Rushi Patel, Martin C. Herbordt |
FPL | 6 |
| 2018 | High Performance Communication on Reconfigurable ClustersabstractFPGA clusters with the FPGAs directly linked through their Multi-Gigabit Transceivers (MGT) have a proven advantage over other commodity architectures for communication-bound applications. To date, however, communication infrastructure for such clusters has generally taken one of two approaches: nearest neighbor only, which is fast but has limited utility, and processor-based, which is general, but relatively slow. What is needed is for communication microarchitecture of these systems to be systematically explored, as has been done for HPC clusters and for Networks on Chip (NoC) on both FPGAs and ASICs. Our first contribution is finding that the properties of clusters of tightly coupled FPGAs substantially influence the router design space. We create a candidate router and generalize it so that it is parameterized by routing algorithm, arbitration policy, and virtual channels (VC). We have created a cycle-accurate simulator validated on a four-FPGA system. We evaluate the design space with respect to a number of standard communication patterns and packet sizes. These results enable selection of the appropriate router for any resource budget. We find that the optimality of the router design varies significantly with workloads. We present a framework that helps to determine appropriate parameters based on different applications and generate the HDL design. We observe that for a 512 FPGA cluster, compared with the router configuration with the best average performance, application-aware router selection can lead to substantial improvement in performance or reduction in area. Jiayi Sheng, Chen Yang 0010, Martin C. Herbordt |
FPL | 3 |
| 2018 | Accelerating MPI Message Matching through FPGA OffloadabstractThe Message Passing Interface (MPI) is the de facto communication standard for distributed-memory High-Performance Computing (HPC) systems. Ultra-low latency communication in HPC is difficult to achieve because of MPI processing requirements, in particular matching requests and messages done by traversing the corresponding queues. Many researchers have addressed this issue by redesigning queues or by offloading them to hardware accelerators. However, state-of-art software approaches cannot free CPUs “from the misery” and hardware approaches either lack scalability or still leave substantial room for further improvement. With the emergence of numerous tightly coupled CPU-FPGA computing architectures, offload of MPI functionality to user-controlled hardware is now becoming viable; we find it productive to revisit hardware approaches. To maintain the generality necessary to support MPI while preventing high resource utilization, we design our MPI queue processing offload based on a recent analysis of performance characteristics in HPC applications. We propose a novel, two-level message queue design: a content addressable memory (CAM) coupled with a resource-saving hardware linked-list. We also propose an optimization that maintains high speed in the cases when the queue is long. To test our design, we create an SOC-based testbed consisting of softcore processors and hardware implementations of the MPI communication stacks. Even while using only a small fraction of the Stratix-V logic, our design can be one to two orders of magnitude faster than two well-known hardware designs. Qingqing Xiong, Anthony Skjellum, Martin C. Herbordt |
FPL | 3 |
| 2018 | An Empirically Guided Optimization Framework for FPGA OpenCLabstractFPGAs have been demonstrated to be capable of very high performance, especially power-performance, but generally at the cost of hand-tuned HDL code by FPGA experts. OpenCL is the leading industry effort in improving performance-programmability. But while it is recognized that optimizing OpenCL code using published best practices is critical to achieving good performance, even optimized code has so far rarely matched that of HDL code, or that available with competing technologies such as GPUs. In this paper we propose a series of systematic and empirically guided code optimizations that augment current best practices and substantially improve achieved performance. Our work characterizes and measures the impact of all of these optimizations. This enables programmers to not only follow a script when optimizing their own kernels, but also opens the way for the development of autotuners to perform optimizations automatically. We also demonstrate that, by applying these proposed code design practices to a number of parallel computing dwarfs, our optimized kernels outperform CPU and previous FPGA OpenCL implementations by 1.2× and 5× respectively. Moreover, our optimizations enable OpenCL FPGA codes to consistently achieve performance within striking distance of approximately 2× best current equivalent code for GPUs and HDL. To the best of our knowledge, this is at least 2× better than previous characterizations of OpenCL FPGA optimizations. Ahmed Sanaullah, Rushi Patel, Martin C. Herbordt |
FPT | 3 |
| 2018 | MPI Derived Datatypes: Performance and Portability IssuesabstractThis paper addresses performance-portability and overall performance issues when derived datatypes are used with four MPI implementations: Open MPI, MPICH, MVAPICH2, and Intel MPI. These comparisons are particularly relevant today since most vendor implementations are now based on Open MPI or MPICH rather than on vendor proprietary code as was more prevalent in the past. Our findings are that, within a single MPI implementation, there are significant differences in performance as a function of it reasonable encodings of derived datatypes as supported by the MPI standard. While this finding may not be surprising, it is important to understand how fundamental vs. arbitrary choices made in early implementation impact the use of derived datatypes to date. Qingqing Xiong, Purushotham V. Bangalore, Anthony Skjellum, Martin C. Herbordt |
EuroMPI | 4 |
| 2018 | Real-time data analysis for medical diagnosis using FPGA-accelerated neural networksabstractBACKGROUND: Real-time analysis of patient data during medical procedures can provide vital diagnostic feedback that significantly improves chances of success. With sensors becoming increasingly fast, frameworks such as Deep Neural Networks are required to perform calculations within the strict timing constraints for real-time operation. However, traditional computing platforms responsible for running these algorithms incur a large overhead due to communication protocols, memory accesses, and static (often generic) architectures. In this work, we implement a low-latency Multi-Layer Perceptron (MLP) processor using Field Programmable Gate Arrays (FPGAs). Unlike CPUs and Graphics Processing Units (GPUs), our FPGA-based design can directly interface sensors, storage devices, display devices and even actuators, thus reducing the delays of data movement between ports and compute pipelines. Moreover, the compute pipelines themselves are tailored specifically to the application, improving resource utilization and reducing idle cycles. We demonstrate the effectiveness of our approach using mass-spectrometry data sets for real-time cancer detection. RESULTS: We demonstrate that correct parameter sizing, based on the application, can reduce latency by 20% on average. Furthermore, we show that in an application with tightly coupled data-path and latency constraints, having a large amount of computing resources can actually reduce performance. Using mass-spectrometry benchmarks, we show that our proposed FPGA design outperforms both CPU and GPU implementations, with an average speedup of 144x and 21x, respectively. CONCLUSION: In our work, we demonstrate the importance of application-specific optimizations in order to minimize latency and maximize resource utilization for MLP inference. By directly interfacing and processing sensor data with ultra-low latency, FPGAs can perform real-time analysis during procedures and provide diagnostic feedback that can be critical to achieving higher percentages of successful patient outcomes. Ahmed Sanaullah, Chen Yang 0010, Yuri Alexeev, Kazutomo Yoshii, Martin C. Herbordt |
BMC Bioinform. | 5 |
| 2017 | Bonded Force Computations on FPGAsabstractWhile acceleration of Molecular Dynamics has received much attention, a significant part of that application, the bonded force calculation, has not. We present what we believe to be the first description and analysis of bonded force calculations outside of ASICs. We characterize the computational requirements. We find that a naive direct implementation requires FPGA resources out of proportion with its proportion of the workload. We investigate other options including various softcores and speed/area tradeoffs. These result in an assortment of solutions optimal for various combinations of problem and cluster size. Qingqing Xiong, Martin C. Herbordt |
FCCM | 2 |
| 2017 | HPC on FPGA clouds: 3D FFTs and implications for molecular dynamicsabstractThe architecture of the Microsoft Catapult II cloud places the accelerator (FPGA) as a bump-in-the-wire on the way to the network and thus promises a dramatic reduction in latency as layers of hardware and software are avoided. We demonstrate this capability with an implementation of the 3D FFT. Next we examine phased application elasticity, i.e., the use of a reduced set of nodes for some phases of an HPC application. We find that, for the FFT phase within Molecular Dynamics, such contraction is beneficial with a 13%–14% performance improvement. Turning to MD, we show how this elasticity can be integrated into the existing data transformation to hide its communication overhead and increase the performance benefit to 16%–29%. Jiayi Sheng, Chen Yang 0010, Ahmed Sanaullah, Michael Papamichael, Adrian M. Caulfield, Martin C. Herbordt |
FPL | 6 |
| 2016 | FPGA-Accelerated Particle-Grid MappingabstractComputing the forces derived from long-range electrostatics is a critical application and also a central part of Molecular Dynamics. Part of that computation, the transformation of a charge grid to a potential grid via a 3D FFT, has received some attention recently and has been found to work extremely well on FPGAs. Here we report on the rest of the computation, which consists of two mappings: charges onto a grid and a potential grid onto the particles. These mappings are interesting in their own right as they are far more compute intensive than the FFTs; each is typically done using tricubic interpolation. We believe that these mappings have been studied only once previously for FPGAs and then found to be exorbitantly expensive; i.e., only bicubic would lit on the chip. In the current work we lind that, when using the Altera Arria 10, not only do both mappings lit, but also an appropriately sized 3D FFT. This enables the building of a balanced accelerator for the entire long-range electrostatics computation on a single FPGA. This design scales directly to FPGA clusters. Other contributions include a new mapping scheme based on table lookup and a measure of the utility of the floating point support of the Arria-10. Ahmed Sanaullah, Arash Khoshparvar, Martin C. Herbordt |
FCCM | 3 |
| 2016 | Application-Aware Collective Communication (Extended Abstract)abstractPreliminary results are presented of hardware support for collective communication that takes advantage of a priori routing information. Jiayi Sheng, Qingqing Xiong, Chen Yang 0010, Martin C. Herbordt |
FCCM | 4 |
| 2016 | Communication and cooling aware job allocation in data centers for communication-intensive workloads
Eduard Llamosí, Fulya Kaplan, Chulian Zhang, Jiayi Sheng, Martin C. Herbordt, Gunar Schirner, Ayse K. Coskun |
J. Parallel Distributed Comput. | 6 |
| 2015 | NCBI BLASTP on High-Performance Reconfigurable Computing SystemsabstractThe BLAST sequence alignment program is a central application in bioinformatics. The de facto standard version, NCBI BLAST, uses complex heuristics that make it challenging to simultaneously achieve both high performance and exact agreement. We propose a system that uses novel FPGA-based filters that reduce the input database by over 99.97% without loss of sensitivity. There are several contributions. First is design of the filters themselves, which perform two-hit seeding, exhaustive ungapped alignment, and exhaustive gapped alignments, respectively. Second is the coupling of the filters, especially the two-hit seeding and the ungapped alignment. Third is pipelining the filters in a single design, including maintaining load balancing as data are reduced by orders of magnitude at each stage. Fourth is the optimization required to maintain operating frequency for the resulting complex design. And finally, there is system integration both in hardware (the Convey HC1-EX) and software (NCBI BLASTP). We present results for various usage scenarios and find complete agreement and a factor of nearly 5x speedup over a fully parallel implementation of the reference code on a contemporaneous CPU. We believe that the resulting system is the leading per-socket-accelerated NCBI BLAST. Atabak Mahram, Martin C. Herbordt |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2014 | 3D FFTs on a Single FPGAabstractThe 3D FFT is critical in many physical simulations and image processing applications. On FPGAs, however, the 3D FFT was thought to be inefficient relative to other methods such as convolution-based implementations of multigrid. We find the opposite: a simple design, operating at a conservative frequency, takes 4μs for 163, 21μs for 323, and 215μs for 643single precision data points. The first two of these compare favorably with the 25μs and 29μs obtained running on a current Nvidia GPU. Some broader significance is that this is a critical piece in implementing a large scale FPGA-based MD engine: even a single FPGA is capable of keeping the FFT off of the critical path for a large fraction of possible MD simulations. Benjamin Humphries, Hansen Zhang, Jiayi Sheng, Raphael Landaverde, Martin C. Herbordt |
FCCM | 5 |
| 2012 | FMSA: FPGA-Accelerated ClustalW-Based Multiple Sequence Alignment through Pipelined PrefilteringabstractMultiple Sequence Alignment (MSA) is perhaps second only to sequence alignment in overall importance in Bioinformatics, being critical, e.g., in determining the structure and function of molecules from putative families of sequences. But while pair wise sequence alignment has been the subject of scores of FPGA acceleration studies, MSA only a few. The most important of these accelerate Clustal-W, the most commonly used MSA code, by either implementing the first of three phases (over 90% of the run time) with Dynamic Programming (DP) methods, or by accelerating the third phase which consumes most of the remaining time. We use a new approach: we apply prefiltering of the kind commonly used in BLAST to perform the initial all-pairs alignments. This results in a speedup of from 80× to 190× over the CPU code (8 cores) and speedup of from 2.5× to 8× over DP/FPGA- and GPU-based methods. When combined with a recently published method for phase 3, and using the original software for phase 2, the end-to-end speedup is at least 50× over an 8-core implementation of the original code. The quality is comparable to the original according to a commonly used benchmark suite evaluated with respect to multiple distance metrics. Atabak Mahram, Martin C. Herbordt |
FCCM | 2 |
| 2012 | CAAD BLASTP 2.0: NCBI BLASTP accelerated with pipelined filtersabstractBLAST is a central application in bioinformatics and so has been the subject of numerous acceleration studies. The de facto standard version of this code, NCBI BLAST, uses complex heuristics which make it challenging to simultaneously achieve both high performance and exact agreement with the original output. In previous work, we have used novel FPGA-based filters that reduce the input database by over 99.99% without loss of sensitivity. In the present work there are two primary contributions. The first is a new mechanism to couple two of the filters in such a way that promising alignments can be found in a fraction of the previous time. The second is the pipelining of the three filters. This is a challenging load balancing problem since the work per filter drops by 5× - 10× at both of the interfaces. Pipelining the filters has two benefits: it removes the need to reconfigure between passes and it reduces the off-chip bandwidth requirement. Together, these two enhancements more than double the performance over the previous best implementation. We currently have CAAD BLASTP working on Virtex-6 and Stratix-IV FPGAs with speed-ups of 9× and 15×, respectively, over the multithreaded original code running on an 8-core PC. We discuss FPGA features that cause this performance disparity. CAAD BLASTP scales easily and is appropriate for use in large FPGA-based servers. Atabak Mahram, Martin C. Herbordt |
FPL | 2 |
| 2011 | Efficient Calculation of Pairwise Nonbonded ForcesabstractA major bottleneck in molecular dynamics (MD) simulations is the calculation of the pair wise nonbonded interactions. Previous work on FPGAs has shown that these calculations can be implemented with a number of force computation pipelines operating in parallel (4 and 8 for the Stratix-III and Stratix-V, respectively). Optimization has received some attention previously in CPU, GPU, FPGA, and ASIC implementations, with direct computation of the equations of interaction being replaced with table lookup with interpolation, and the order and granularity of those interpolations being optimized. FPGAs lend themselves to a particularly rich design space both of opportunities and constraints. We explore and evaluate this space with respect to both resource requirements and simulation quality. We find that FPGAs' BRAM architecture makes them well suited to support unusually fine-grained intervals. This leads to a reduction in other logic and a proportional increase in performance. We demonstrate these designs with prototype implementations supporting full electrostatics and integrated into NAMD-lite. Throughput is improved by 50% over the previous best FPGA implementation while simulation quality is maintained. Matt Chiu, Md. Ashfaquzzaman Khan, Martin C. Herbordt |
FCCM | 3 |
| 2010 | Fast and accurate NCBI BLASTP: acceleration with multiphase FPGA-based prefilteringabstractNCBI BLAST has become the de facto standard in bioinformatic approximate string matching and so its acceleration is of fundamental importance. The problem is that it uses complex heuristics which make it difficult to simultaneously achieve both substantial speed-up and exact agreement with the original output. We have previously described how a novel FPGA-based prefilter that performs exhaustive ungapped alignment (EUA) could be used to reduce the computation by over 99.9% without loss of sensitivity. The primary contribution here is to show how the EUA filter can be combined with another filter, this one based on standard 2-hit seeding. The result is a doubling of performance over the previous best implementation, which itself is an order of magnitude faster than the unaccelerated original. Other contributions include new algorithms for both the original EUA and the 2-hit filters and experimental results demonstrating their utility. This new multiphase FPGA-accelerated NCBI BLASTP scales easily and is appropriate for use in large FPGA-based servers such as the Novo-G. Atabak Mahram, Martin C. Herbordt |
ICS | 2 |
| 2010 | CAAD BLASTn: Accelerated NCBI BLASTn with FPGA prefilteringabstractThe canonical bioinformatics application is determining the biological similarity of a new sequence (protein or DNA) with respect to databases of known sequences. The BLAST algorithm is used for the vast majority of these searches. Of the various BLAST implementations, the one published by NCBI is a recognized standard. In previous work we described FPGA acceleration of the protein version of NCBI BLAST (BLASTp) using our TreeBLAST-based filter. Here we apply this filter to NCBI BLASTn, the DNA version. We show the modifications to the structures of the filtering components needed to handle DNA, as opposed to protein, sequences. The design has been implemented on an Altera Stratix III family chip. Our experimental results show that the speedup is greater than 12x and the accuracy is 100%. Jin H. Park, Yunfei Qiu, Martin C. Herbordt |
ISCAS | 3 |
| 2010 | Molecular Dynamics Simulations on High-Performance Reconfigurable Computing SystemsabstractThe acceleration of molecular dynamics (MD) simulations using high-performance reconfigurable computing (HPRC) has been much studied. Given the intense competition from multicore and GPUs, there is now a question whether MD on HPRC can be competitive. We concentrate here on the MD kernel computation: determining the short-range force between particle pairs. In one part of the study, we systematically explore the design space of the force pipeline with respect to arithmetic algorithm, arithmetic mode, precision, and various other optimizations. We examine simplifications and find that some have little effect on simulation quality. In the other part, we present the first FPGA study of the filtering of particle pairs with nearly zero mutual force, a standard optimization in MD codes. There are several innovations, including a novel partitioning of the particle space, and new methods for filtering and mapping work onto the pipelines. As a consequence, highly efficient filtering can be implemented with only a small fraction of the FPGA's resources. Overall, we find that, for an Altera Stratix-III EP3ES260, 8 force pipelines running at nearly 200 MHz can fit on the FPGA, and that they can perform at 95% efficiency. This results in an 80-fold per core speed-up for the short-range force, which is likely to make FPGAs highly competitive for MD. Matt Chiu, Martin C. Herbordt |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2009 | Parallel Discrete Event Simulation of Molecular Dynamics Through Event-Based DecompositionabstractMolecular dynamics simulation based on discrete event simulation (DMD) is emerging as an alternative to time-step driven molecular dynamics (MD). DMD uses simplified discretized models, enabling simulations to be advanced by event, with a resulting performance increase of several orders of magnitude. Even so, DMD is compute bound. Moreover, unlike MD, causality issues make DMD difficult to scale. Here we present a microarchitecture-inspired parallel algorithm for DMD: speculative execution enables multithreading, while in-order commitment ensures correctness. Our initial not-yet optimized implementation obtains scalability for a multicore processor when running realistic simulation models. Martin C. Herbordt, Md. Ashfaquzzaman Khan, Tony Dean |
ASAP | 1 |
| 2009 | CAAD BLASTP: NCBI BLASTP Accelerated with FPGA-Based Accelerated Pre-FilteringabstractNCBI BLAST has become the de facto standard in bioinformatic approximate string matching and so its acceleration is of fundamental importance. The problem is that it uses complex heuristics which make it difficult to simultaneously achieve both substantial speed-up and exact agreement with the original output. Our approach is to prefilter the database. To make this work we have developed a novel heuristic which we append to a previously described structure for ungapped alignment. This enables us to quickly reduce the database by factors of 300 and 1100, for the ungapped and gapped options, respectively, while rejecting no significant sequences. On current hardware we anticipate a speed-up of at least a factor of 10 for NCBI BLASTP, independent of sensitivity settings. This filter is portable to other BLAST codes, and other filters can be similarly integrated into NCBI BLAST. Jin H. Park, Yunfei Qiu, Martin C. Herbordt |
FCCM | 3 |
| 2009 | Efficient particle-pair filtering for acceleration of molecular dynamics simulationabstractThe acceleration of molecular dynamics (MD) simulations using high performance reconfigurable computing (HPRC) has been much studied. Given the intense competition from multicore and GPUs, there has been a question whether MD on HPRC can be competitive. We concentrate here on the MD kernel computation: determining the short-range force between particle pairs. In particular, we present the first FPGA study on the filtering of particle pairs with nearly zero mutual force, a standard optimization in MD codes. There are several innovations, including a novel partitioning of the particle space, and new methods for filtering and mapping work onto the pipelines. As a consequence, highly efficient filtering can be implemented with only a small fraction of the FPGA's resources. Overall, we find that, for an Altera Stratix-III EP3ES260, 8 force pipelines running at 200MHz can fit on the FPGA, and that they can perform at 95% efficiency. This results in a 80-fold per core speed-up for the short-range force, which is likely to make FPGAs highly competitive for MD. Matt Chiu, Martin C. Herbordt |
FPL | 2 |
| 2008 | An Efficient O(1) Priority Queue for Large FPGA-Based Discrete Event Simulations of Molecular DynamicsabstractMolecular dynamics simulation based on discrete event simulation (DMD) is emerging as an alternative to time-step driven molecular dynamics (MD). Although DMD improves performance by several orders of magnitude, it is still compute bound. In previous work, we found that FPGAs are extremely well suited to accelerating DMD, with speed-ups of 200× to 400× being achieved. Large models, however, are problematic because they require that most predicted events be stored in off-chip memory, rather than on the FPGA. Here we present a solution that allows the priority queue to be extended seamlessly into off-chip memory, resulting in a throughput equal to the hardware-only priority queue, or about 30× faster than the best software-only algorithm. The solution is based on the observation that--when an event is predicted to occur far in the future--not only can its processing be imprecise, but the time when the processing itself occurs can also be substantially delayed. This allows numerous optimizations and restructurings. We demonstrate the resulting design on standard hardware and present the experimental results used to tune the data structures. Martin C. Herbordt, Francois Kosie, Josh Model |
FCCM | 1 |
| 2008 | Acceleration of a production rigid molecule docking codeabstractModeling the interactions of biological molecules, or docking is critical to both understanding basic life processes and to designing new drugs. Here we describe the FPGA-based acceleration of a recently developed, complex, production docking code. We find that it is necessary to extend our previous 3D correlation structure in several ways, most significantly to support simultaneous computation of several correlation functions. The result is a hundred-fold speed-up of a section of the code that represents over 92% of the original run-time. An additional 4% is accelerated through a previously described method, yielding a total acceleration of almost 25times for typical protein-ligand combinations. Bharat Sukhwani, Martin C. Herbordt |
FPL | 2 |
| 2008 | Explicit design of FPGA-based coprocessors for short-range force computations in molecular dynamics simulations
Yongfeng Gu, Tom Van Court, Martin C. Herbordt |
Parallel Comput. | 3 |
| 2007 | FPGA-Based Multigrid Computation for Molecular Dynamics SimulationsabstractFPGA-based acceleration of molecular dynamics (MD) has been the subject of several recent studies. Implementing long-range forces, however, has only recently been addressed. Here we describe a solution based on the multigrid method. We show that multigrid is, in general, an excellent match to FPGAs: the primary operations take advantage of the large number of independently addressable RAMs and the efficiency with which complex systolic structures can be implemented. The multigrid accelerator has been integrated into our existing MD system, and an overall performance gain of 5x to 7x has been obtained, depending on hardware configuration and reference code. The simulation accuracy is comparable to the original double precision serial code. Yongfeng Gu, Martin C. Herbordt |
FCCM | 2 |
| 2007 | Discrete Event Simulation of Molecular Dynamics with Configurable LogicabstractMolecular dynamics simulation based on discrete event simulation (DMD) is emerging as an alternative to time-step driven molecular dynamics (MD). DMD uses simplified discretized models, enabling simulations to be advanced by event, with a resulting performance increase of several orders of magnitude. Even so, DMD is compute bound. Moreover, unlike MD, causality issues make DMD difficult to scale, with O(√p) being the best so far achieved. We find that FPGAs are extremely well suited to accelerating DMD. The chaotic execution, which results in there being virtually no prediction window, is overcome with a long processing pipeline augmented with associative structures analogous to those used in CPU reorder buffers. Our primary result is a microarchitecture for DMD that processes events with a throughput equal to a small multiple of the FPGA's clock, resulting in a hundred-fold speed-up over serial implementations. Josh Model, Martin C. Herbordt |
FPL | 2 |
| 2007 | Single pass streaming BLAST on FPGAs
Martin C. Herbordt, Josh Model, Bharat Sukhwani, Yongfeng Gu, Tom Van Court |
Parallel Comput. | 1 |
| 2006 | Application-Specific Memory Interleaving Enables High Performance in FPGA-based Grid ComputationsabstractCurrent generations of FPGAs create possibilities for innovative, application-specific computation pipelines. In many cases, the pipeline can fully exploit the FPGA's parallelism only when multiple operands are available concurrently, requiring clusters of values to be fetched from memory. These clusters of values often have fixed organization, as in the eight grid points around an off-grid position that are needed for 3D interpolation of a value at that position. This paper presents a technique for creating custom interleaving of the FPGA's on-chip memories, giving access to the entire cluster of values in one memory cycle. This technique works on grids of 2, 3, or more dimensions, on many non-rectangular grids, and on cluster organization specific to each application. The authors report the initial version of a design tool that inputs the relative positions of grid points in the access cluster, and produces synthesizable HDL code for the custom-interleaved memory Tom Van Court, Martin C. Herbordt |
FCCM | 2 |
| 2006 | Integrating FPGA Acceleration into the Protomol Molecular Dynamics Code: Preliminary ReportabstractThe authors describe a new pipeline for computing non-bonded forces and its integration into the ProtoMol molecular dynamics (MD) code. There are several innovations: a novel interpolation strategy, including use of higher order terms; coefficient generation with orthonormal functions; the introduction of "semi-floating point" numbering; and various issues related to system integration. As a result, we are able to model far more particle types, without relying on complex buffering, and obtain higher accuracy than previously. A two pipeline accelerator has been implemented on a 2004-era Xilinx VirtexII Pro VP70, integrated into ProtoMol, and tested with an enzyme inhibitor model having 8000 particles and 26 particle types. Despite performing all O(n) work on the host PC, as well as the data conversion and communication overhead, this implementation yields 5.5x to 15.7x speed-ups over a 2.8GHz PC (depending on whether cell lists are used), and with accuracy comparable to the serial code Yongfeng Gu, Tom Van Court, Martin C. Herbordt |
FCCM | 3 |
| 2006 | Single Pass, BLAST-Like, Approximate String Matching on FPGAsabstractApproximate string matching is fundamental to bioinformatics, and has been the subject of numerous FPGA acceleration studies. We address issues with respect to FPGA implementations of both BLAST- and dynamic programming- (DP) based methods. Our primary contributions are two new algorithms for emulating the seeding and extension phases of BLAST. These operate in a single pass through a database at streaming rate (110 Maa/sec on a VP70 for query sizes up to 600 and 170 Maa/sec on a Virtex4 for query sizes up to 1024), and with no preprocessing other than loading the query string. Further, they use very high sensitivity with no slowdown. While current DP-based methods also operate at streaming rate, generating results can be cumbersome. We address this with a new structure for data extraction. We present results from several implementations Martin C. Herbordt, Josh Model, Yongfeng Gu, Bharat Sukhwani, Tom Van Court |
FCCM | 1 |
| 2006 | Application-Specific Memory Interleaving for FPGA-Based Grid Computations: A General Design TechniqueabstractMany compute-intensive applications generate single result values by accessing clusters of nearby points in grids of one, two, or more dimensions. Often, the performance of FGPA implementations of such algorithms would improve if there were concurrent, non-interfering access to all points in each cluster. When clusters contain dozens of points and access patterns are irregular, multiported memories are infeasible and vector-oriented approaches are inapplicable. Instead, the grid points can be distributed across multiple interleaved memory banks so that, when accessing any cluster, each point comes from a different memory bank. We present a general technique based on the knowledge of the application's multidimensional indexing. This technique maps access clusters into a custom-interleaved memory using the FPGA's multiple on-chip RAMs and configurable data paths. Case studies examine rectangular and non-rectangular grids of different dimensionality, including performance vs. resource tradeoffs when cluster sizes are not powers of two. We also present a prototype tool for generating interleaved memories automatically from concise, application-specific definitions. Tom Van Court, Martin C. Herbordt |
FPL | 2 |
| 2006 | Sizing of Processing Arrays for FPGA-Based ComputationabstractComputing applications in FPGAs are commonly built from repetitive structures of computing and/or memory elements. In many cases, application performance depends on the degree of parallelism - ideally, the most that will fit into the fabric of the FPGA being used. Several factors complicate determination of the largest structure that will fit the FPGA: arrays that grow nonlinearly and in uneven step sizes, coupled structures that grow in different polynomial order, multiple design parameters controlling different aspects of the computing structure, and interlocked usage of different hardware resources. Combined with resource usage that depends on application-specific data elements and arithmetic details, these factors defeat any simple approach for scaling the computing structures up to the FPGA's capacity. We present a formal analysis of maximizing FPGA utilization, with adaptations that simplify the optimization problem. We also report on design tools containing extensions that support automated sizing of FPGA-based computation arrays Tom Van Court, Martin C. Herbordt |
FPL | 2 |
| 2006 | Improved Interpolation and System Integration for FPGA-Based Molecular Dynamics SimulationsabstractFPGA-based acceleration of molecular dynamics (MD) has been the subject of several recent studies. The paper describes a new non-bonded force computation pipeline implemented on a 2004-era COTS FPGA board and its integration into the ProtoMol MD code. There are several innovations: a novel interpolation strategy; the introduction of a "semi-floating point" format; and various issues related to system integration. As a result, the authors are able to model far more particle types, without relying on complex buffering, and obtain higher accuracy than previously. A two pipeline accelerator has been implemented on a Xilinx VirtexII Pro VP70, integrated into ProtoMol, and tested with an enzyme inhibitor model having 8000 particles and 26 particle types. Despite performing all O(n) work on the host PC, as well as the data conversion and communication overhead, this implementation yields a 5.5times speed-up over a 2.8GHz PC, and with accuracy comparable to the serial code Yongfeng Gu, Tom Van Court, Martin C. Herbordt |
FPL | 3 |
| 2005 | Preliminary Report: FPGA Acceleration of Molecular Dynamics ComputationsabstractMolecular dynamics (MD) is of central importance to computational chemistry and its myriad applications. In this paper we show that, at even a preliminary stage of development, MD can be implemented efficiently on a COTS FPGA board, and that a 57x speed-up over a PC implementation can be obtained. We sketch our FPGA implementation and describe how performance tuning and precision management could double this factor. Yongfeng Gu, Tom Van Court, Douglas DiSabello, Martin C. Herbordt |
FCCM | 4 |
| 2005 | LAMP: A Tool Suite for Families of FPGA-Based Application AcceleratorsabstractField-programmable gate arrays (FPGAs) are becoming increasingly attractive as computation engines: they are currently being integrated into supercomputers as application accelerators. In order for widespread use of FPGA-based accelerators to be practical, however, design tools must resolve a number of conflicting needs: application-specific tuning versus wide applicability, stability of software investment versus use of the most recent and powerful acceleration hardware, and customization of complex computing structures by end users who lack logic design skills. We describe how the LAMP tools address these conflicts and report performance results from experiments in creating families of application accelerators. Tom Van Court, Martin C. Herbordt |
FPL | 2 |
| 2005 | Accelerating Molecular Dynamics Simulations With Configurable CircuitsabstractMolecular dynamics (MD) is of central importance to computational chemistry. Here we show that MD can be implemented efficiently on a COTS FPGA board, and that speed-ups from 31/spl times/ to 88/spl times/ over a PC implementation can be obtained. Although the amount of speed-up depends on the stability required, 46/spl times/ can be obtained with virtually no detriment, and the upper end of the range is apparently viable in many cases. We sketch our FPGA implementations and describe the effects of precision on the trade-off between performance and quality of the MD simulation. Yongfeng Gu, Tom Van Court, Martin C. Herbordt |
FPL | 3 |
| 2004 | Families of FPGA-Based Algorithms for Approximate String Matching
Tom Van Court, Martin C. Herbordt |
ASAP | 2 |
| 2004 | FPGA Acceleration of Rigid Molecule InteractionsabstractWe show how an application of central importance to computational biochemistry can be implemented efficiently on an FPGA. This requires reformulating the algorithm and taking advantage of application characteristics to maximize device utilization. Tom Van Court, Yongfeng Gu, Martin C. Herbordt |
FCCM | 3 |
| 2004 | Processing Repetitive Sequence Structures with Mismatches at Streaming Rate
Albert A. Conti, Tom Van Court, Martin C. Herbordt |
FPL | 3 |
| 2004 | FPGA Acceleration of Rigid Molecule Interactions
Tom Van Court, Yongfeng Gu, Martin C. Herbordt |
FPL | 3 |
| 2004 | Array control for high-performance SIMD systems
Martin C. Herbordt, Jade Cravy, Honghai Zhang |
J. Parallel Distributed Comput. | 1 |
| 2003 | Case Study of a Functional Genomics Application
Tom Van Court, Martin C. Herbordt, Richard J. Barton |
FPL | 2 |
| 2000 | Control for High-Speed PE ArraysabstractAlthough arrays of SIMD PEs can be built with very high operating frequencies, problems exist in keeping the array busy. The inherent mismatch between host and array makes it difficult to maintain high array utilization: either the rate of instruction issue is very low or PE data locality is compromised, having the same effect. Our solution is based on an array control unit (ACU) design that expands macro instructions in two stages, first by data tile and then into microinstructions. The expansion itself solves the issue problem; decoupling the expansion modalities maintains data locality. Several issues involving host/ACU interaction need to be resolved to effect this solution. Martin C. Herbordt, Honghai Zhang, Calvin Lin, Hong Rao, Jade Cravy |
ASAP | 1 |
| 2000 | A System for Evaluating Performance and Cost of SIMD Array Designs
Martin C. Herbordt, Jade Cravy, Renoy Sam, Owais Kidwai, Calvin Lin |
J. Parallel Distributed Comput. | 1 |
| 1999 | Using Emulations to Enhance the Performance of Parallel ArchitecturesabstractWe illustrate the potential of techniques and results from the theory of network emulations to enhance the performance of a parallel architecture. The vehicle for this demonstration is a suite of algorithms that endow an N-processor bit-serial processor array A with a "meta-instruction" GAUGE k, which (logically) reconfigures A into an N/k-processor virtual machine B/sub k/ that has: 1) a datapath and memory bus whose emulated width is k bits, as opposed to A's 1-bit width and 2) an instruction set that operates on k-bit words, in contrast to A's instruction set, which operates on 1-bit words. In order to stress the strength of the approach, we show (via pseudocode) how our emulation techniques can be implemented efficiently even if A operates in strict SIMD mode, with only single-bit masking capabilities and with no indexed memory accesses. We describe at an algorithmic level how to implement our technique-including datapath conversion ("corner-turning") and the creation of the word-parallel instruction sets-on arrays of any regular network topology. We instantiate our technique in detail for arrays based on topologies with quite disparate characteristics: the hypercube, the de Bruijn network, and a genre of mesh with reconfigurable buses. Importantly, the emulations that underlie our technique do not alter the native machine's instruction set, hence allowing an invariant programming model across gauges. Bojana Obrenic, Martin C. Herbordt, Arnold L. Rosenberg, Charles C. Weems |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1997 | Preprototyping SIMD Coprocessors Using Virtual Machine Emulation and Trace CompilationabstractThe use of massively parallel SIMD array architectures is proliferating in the area of domain specific coprocessors. Even so, they have undergone few systematic empirical studies. The underlying problems include the size of the architecture space, the lack of portability of the test programs, and the inherent complexity of simulating up to hundreds of thousands of processing elements. We address the computational cost problem with a novel approach to trace-based simulation. Code is run on an abstract virtual machine to generate a coarse-grained trace, which is then refined through a series of transformations (a process we call trace compilation) wherein greater resolution is obtained with respect to the details of the target machine. We have found this technique to be one to two orders of magnitude faster than instruction-level simulation while still retaining much of the accuracy of the model. Furthermore, abstract machine traces must be regenerated for only a small fraction of the possible parameter combinations. Using virtual machine emulation and trace compilation also addresses program portability by allowing the user to code in a single data parallel language with a single compiler, regardless of the target architecture. This technique has already been used to generate significant results with respect to SIMD array architectures, a sample of which are presented here. Martin C. Herbordt, Owais Kidwai, Charles C. Weems |
SIGMETRICS | 1 |
| 1995 | An empirical study of datapath, memory hierarchy, and network in SIMD array architecturesabstractAlthough SIMD arrays have been built for 30 years, they have as a class been the subject of few empirical design studies. Using ENPASSANT, a simulation environment developed for that purpose, we analyze several aspects of SIMD array architecture with respect to a test suite of spatially mapped applications. Several surprising results are obtained. With respect to memory hierarchy, we find that adding a level of cache to current PE designs is likely to be advantageous, but that such a cache will look quite different than expected. In particular, we find that associativity has unusual significance and that performance varies inversely with block size. Router network results indicate the importance of support for local transfers, broadcast, and reduction even at the expense of arbitrary permutations. Other communication results point to the appropriate dimensionality of k-ary n-cube networks (2 or 3), and the criticality of supporting bidirectional transfers, even if the overall bandwidth remains unchanged. Martin C. Herbordt, Charles C. Weems |
ICCD | 1 |
| 1995 | Experimental Analysis of Some SIMD Array Memory Hierarchies
Martin C. Herbordt, Charles C. Weems |
ICPP (1) | 1 |
| 1995 | Enpassant: An Environment for Evaluating Massively Parallel Array Architectures for Spatially Mapped ApplicationsabstractAlthough massively parallel arrays for spatially mapped applications have been proposed since the 1950s42 and built since the 1960s,12 there have been very few systematic empirical studies that cover more than a small fraction of the design space. The problems have included the lack of a test suite of non-trivial application codes; inadequate language support; the difficulties of balancing evaluation performance with flexibility; and balancing test suite portability with accuracy of evaluation. We describe an environment that addresses these problems. A realistic workload including a series of applications currently being used as building blocks in vision research has been constructed. Both flexibility in architectural parameter selection and simulation efficiency are maintained with a novel new technique that combines virtual machine emulation with trace-driven simulation. The trade-off between fairness to diverse target architectures and programmability of the test suite is addressed through the use of operator and application libraries for a small set of critical functions. We also present examples of the type of results we are obtaining, including the effects of changing ALU designs and datapath widths, finding critical points in register set and cache sizes, the benefits of various types of router networks, and the performance cost of processor virtualization. Martin C. Herbordt, Charles C. Weems |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1994 | Practical Algorithms for Online Routing on Fixed and Reconfigurable Meshes
Martin C. Herbordt, James C. Corbett, Charles C. Weems, John Spalding |
J. Parallel Distributed Comput. | 1 |
| 1992 | Nonuniform region processing on SIMD arrays using the coterie network
Martin C. Herbordt, Charles C. Weems, Michael J. Scudder |
Mach. Vis. Appl. | 1 |
| 1991 | A computational framework and SIMD algorithms for low-level support of intermediate level vision processingabstractThe authors propose an additional level of parallelism, called multi-associativity, as a framework for simultaneously performing associative computation on data sets mapped to irregular, non-uniform, aggregates of processing elements (PEs). They introduce algorithms developed for the CAAPP to simulate efficiently within aggregates of PEs simultaneously the associative algorithms typically supported in hardware at the array level. Some of the results are: the efficient application of existing associative algorithms to arbitrary aggregates of PEs in parallel and the development of multi-associative algorithms, among them parallel prefix and convex hull. The multi-associative framework also extends the associative paradigm by allowing operation on and among aggregates themselves, operations not defined when the entity in question is always an entire array.> Martin C. Herbordt, Charles C. Weems, Michael J. Scudder |
CVPR | 1 |
| 1991 | Multi-associativity: A Framework for Solving Multiple Non-uniform Problem Instances Simultaneously on SIMD Arrays
Martin C. Herbordt, Charles C. Weems |
ICPP (3) | 1 |
| 1990 | Routing on the CAAPPabstractA collection of routing algorithms for the content-addressable array parallel processor (CAAPP) which allow many more classes of interprocessor communication to be executed efficiently than otherwise on machines with conventional mesh-connected topologies is presented. It is shown that routing on the CAAPP can be executed with simplicity and performance similar to that of a dedicated routing network. Experimental results are presented from random permutations as well as from several common machine vision applications.> Martin C. Herbordt, Charles C. Weems, David B. Shu |
ICPR (2) | 1 |
| 1990 | Message-Passing Algorithms for a SIMD Torus with CoteriesabstractThis paper describes the results of an investigation into routing algorithms to be used when programming the CAAPP (Content Addressable Array Parallel Processor) [19], a SIMD mesh-connected array processor enhanced with the coterie network, a mechanism similar to reconfigurable buses. We will show that the coterie network gives the CAAPP a capability far beyond solely meshconnected processors; in fact, the performance of routing on many classes of permutations is more comparable to the Connection Machine which has a dedicated hypercube routing network. Most of the current routing algorithms for meshconnected array processors (with N PEs in an n \\Theta n Martin C. Herbordt, Charles C. Weems, James C. Corbett |
SPAA | 1 |