VLDB 2026 Research / reviewers in the wild / expert
Alan D. George
dblp:g/AlanDGeorge
· DBLP profile ↗
73ranked-venue papers
3as first author
11since 2021 · last 2024
0000-0001-9665-2879ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 54 · 2 first-author · 9 since 2021Computer networks · 16 · 1 since 2021Artificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 2Security and privacy · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Ph.D. Project: Investigating oneAPI Performance & Scalability using Feedforward FFT ArchitecturesabstractImproving development time while maintaining high-performance and resource-efficient designs is an important area of advancement for FPGAs. While hardware description languages provide fine-grain control over architectures created, they are often tedious to use and debug. High-level synthesis (HLS) tools, however, allow designers to create systems using languages such as C and C++, greatly improving productivity. The oneAPI toolkit for FPGAs provides a high-level interface to describe FPGA architectures, enabling rapid design iteration through such productivity gains. This research conducts a design space exploration of feedforward FFT architectures using FPGAs. FFT architectures of varying frequency resolution are created, observing performance as designs are scaled up in size. Additionally, resource utilization is compared between designs to gain an understanding of the impact of design choices, as well as the efficacy of the oneAPI compiler. Through this combined analysis of performance and resource metrics, we conclude that the oneAPI toolkit is capable of producing high-throughput designs. However, preliminary results fail to provide improvements at larger scales when compared to highly parallelized libraries. James Bickerstaff, Alan D. George |
FCCM | 2 |
| 2024 | Ph.D. Project: Accelerated Machine Learning for On-Orbit OPIR Target DetectionabstractDetecting dim point-source targets in overhead persistent infrared (OPIR) imagery with high probability requires emerging machine-learning (ML) methods. These complex algorithms and increasing sensor resolution have made high-throughput, on-orbit processing challenging. This research ex-plores end-to-end acceleration of the Multiscale Local Contrast Learning Network (MLCLNet) for scalable on-orbit target detection with arbitrarily large input frames. An end-to-end pipeline for inference of batched subframes is developed and various inference accelerators are explored including GPU, Xilinx Deep Learning Processor Unit (DPU), and a custom FPGA design. The target detection performance for quantized MLCLNet is evaluated at various model sizes. The end-to-end pipeline is evaluated in terms of throughput, latency, resources, power, and efficiency for each inference accelerator considered. Initial results for FPGA-based acceleration using the Xilinx DPU are presented. Daniel C. Stumpp, Alan D. George |
FCCM | 2 |
| 2024 | Leveraging HBM2 for Accelerating k-mer Counting with oneAPI on FPGAs
Owen P. Lucas, Alan D. George |
FPL | 2 |
| 2024 | Flow-Based Visual Stream Compression for Event CamerasabstractAs the use of neuromorphic, event-based vision sensors expands, the need for compression of their output streams has increased. While their operational principle ensures event streams are spatially sparse, the high-temporal resolution of the sensors can result in high-data rates from the sensor depending on scene dynamics. For systems operating in communication-bandwidth-constrained and power-constrained environments, it is essential to compress these streams before transmitting them to a remote receiver. Therefore, we introduce a flow-based method for the real-time asynchronous compression of event streams as they are generated. This method leverages real-time optical flow estimates to predict future events without needing to transmit them, therefore, drastically reducing the amount of data transmitted. The flow-based compression introduced is evaluated using a variety of methods, including spatiotemporal distance between event streams. The introduced method itself is shown to achieve an average compression ratio (CR) of 2.70 on a variety of event-camera data sets with the evaluation configuration used. That compression is achieved with a median temporal error of 0.31 ms and an average spatiotemporal event-stream distance of 4.72. When combined with Lempel-Ziv–Markov chain algorithm compression for non-real-time applications, our method can achieve state-of-the-art average CRs ranging from 9.29 to 13.01. Additionally, we demonstrate that the proposed prediction algorithm is capable of performing real time, low-latency event prediction. Daniel C. Stumpp, Himanshu Akolkar, Alan D. George, Ryad Benosman |
IEEE Internet Things J. | 3 |
| 2024 | Characterizing Parameter Scaling with Quantization for Deployment of CNNs on Real-Time SystemsabstractModern deep-learning models tend to include billions of parameters, reducing real-time performance. Embedded systems are compute-constrained while frequently used to deploy these models for real-time systems given size, weight, and power requirements. Tools like parameter-scaling methods help to shrink models to ease deployment. This research compares two scaling methods for convolutional neural networks, uniform scaling and NeuralScale, and analyzes their impact on inference latency, memory utilization, and power. Uniform scaling scales the number of filters evenly across a network. NeuralScale adaptively scales the model to theoretically achieve the highest accuracy for a target parameter count. In this study, VGG-11, MobileNetV2, and ResNet-50 models were scaled to four ratios: 0.25×, 0.50×, 0.75×, 1.00×. These models were benchmarked on an ARM Cortex-A72 CPU, an NVIDIA Jetson AGX Xavier GPU, and a Xilinx ZCU104 FPGA. Additionally, quantization was applied to meet real-time objectives. The CIFAR-10 and tinyImageNet datasets were studied. On CIFAR-10, NeuralScale creates more computationally intensive models than uniform scaling for the same parameter count, with relative speeds of 41% on the CPU, 72% on the GPU, and 96% on the FPGA. The additional computational complexity is a tradeoff for accuracy improvements in VGG-11 and MobileNetV2 NeuralScale models but reduced ResNet-50 NeuralScale accuracy. Furthermore, quantization alone achieves similar or better performance on the CPU and GPU devices when compared with models scaled to 0.50×, despite slight reductions in accuracy. On the GPU, quantization reduces latency by 2.7× and memory consumption by 4.3×. Uniform-scaling models are 1.8× faster and use 2.8× less memory. NeuralScale reduces latency by 1.3× and dropped memory by 1.1×. We find quantization to be a practical first tool for improved performance. Uniform scaling can easily be applied for additional improvements. NeuralScale may improve accuracy but tends to negatively impact performance, so more care must be taken with it. Calvin B. Gealy, Alan D. George |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2023 | Accelerating Graph Analytics with oneAPI and Intel FPGAsabstractGraphs are a popular and effective way to represent relationships between data points in a network due to their simple data abstraction and low storage cost per data point. However, as data continues to grow, the processing complexity of graph operations also significantly increases. Thus, there is a need to accelerate graph processing through specialized techniques and hardware in order to improve analytical throughput. This research uses the oneAPI toolkit with SYCL for FPGAs and leverages the increased productivity when creating designs, as compared to traditional hardware description language methodologies. The oneAPI toolkit is used in the creation of minimum-spanning-tree (MST) and breadth-first search (BFS) accelerators, evaluating the impact of different high-level design choices on the overall performance observed. James Bickerstaff, Luke Kljucaric, Alan D. George |
FCCM | 3 |
| 2023 | Clustering Classification on FPGAs for Neuromorphic Feature ExtractionabstractState-of-the-art machine-learning (ML) apps are often expensive in terms of memory and computation required for high-accuracy feature extraction and object classification using frame-based sensors. By contrast, neuromorphic feature-extraction algorithms offer lower memory and compute complexity than traditional ML algorithms, which can improve the scalability of ML apps. These neuromorphic algorithms operate on neuromorphic, event-based sensor data as opposed to traditional, frame-based camera images. These neuromorphic sensors capture events at microsecond resolution with variable bandwidth. Therefore, neuromorphic algorithms require low-latency architectures to enable real-time processing. Field-programmable gate arrays (FPGAs) are reconfigurable-logic devices that can realize low-latency data paths required for real-time neuromorphic computation. Luke Kljucaric, Alan D. George |
FCCM | 2 |
| 2023 | Deep Learning Inferencing with High-performance Hardware AcceleratorsabstractAs computer architectures continue to integrate application-specific hardware, it is critical to understand the relative performance of devices for maximum app acceleration. The goal of benchmarking suites, such as MLPerf for analyzing machine learning (ML) hardware performance, is to standardize a fair comparison of different hardware architectures. However, there are many apps that are not well represented by these standards that require different workloads, such as ML models and datasets, to achieve similar goals. Additionally, many apps, like real-time video processing, are focused on latency of computations rather than strictly on throughput. This research analyzes multiple compute architectures that feature ML-specific hardware on a case study of handwritten Chinese character recognition. Specifically, AlexNet and a custom version of GoogLeNet are benchmarked in terms of their streaming latency and maximum throughput for optical character recognition. Considering that these models are composed of fundamental neural network operations yet architecturally different from each other, these models can stress devices in different yet insightful ways that generalizations of the performance of other models can be drawn from. Many devices featuring ML-specific hardware and optimizations are analyzed including Intel and AMD CPUs, Xilinx and Intel FPGAs, NVIDIA GPUs, and Google TPUs. Overall, ML-oriented hardware added to the Intel Xeon CPUs helps to boost throughput by 3.7× and to reduce latency by up to 34.7×, which makes the latency of Intel Xeon CPUs competitive on more parallel models. The TPU devices were limited in terms of throughput due to large data transfer times and not competitive in terms of latency. The FPGA frameworks showcase the lowest latency on the Xilinx Alveo U200 FPGA achieving 0.48 ms on AlexNet using Mipsology Zebra and 0.39 ms on GoogLeNet using Vitis-AI. Through their custom acceleration datapaths coupled with high-performance SRAM, the FPGAs are able to keep critical model data closer to processing elements for lower latency. The massively parallel and high-memory GPU devices with Tensor Core accelerators achieve the best throughput. The NVIDIA Tesla A100 GPU showcases the highest throughput at 42,513 and 52,484 images/second for AlexNet and GoogLeNet, respectively. 1 Luke Kljucaric, Alan D. George |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2021 | A streaming hardware architecture for real-time SIFT feature extractionabstractThe Scale-Invariant Feature Transform (SIFT) is a feature extractor that serves as a key step in many computer-vision pipelines. Real-time operation based on a software-only approach is often infeasible, but FPGAs can be employed to parallelize execution and accelerate the application to meet latency requirements. In this study, we present a stream-based hardware acceleration architecture for SIFT feature extraction. Using a novel strategy to store pixels required for descriptor computation, the execution time needed to generate SIFT descriptors is greatly improved relative to previous designs. This strategy also enables further reduction of the execution time by introducing multiple processing elements (PEs) for computation of several SIFT descriptors in parallel. Additionally, the proposed architecture supports keypoint detection at an arbitrary number of octaves and allows for runtime configuration of various parameters. An FPGA implementation targeting the Xilinx Zynq-7045 system-on-chip (SoC) device is deployed to demonstrate the efficiency of the proposed architecture. In the target hardware, the resulting system is capable of processing images with a resolution of 1280 × 720 pixels at up to 150 FPS while maintaining modest resource utilization. Hector A. Li Sanchez, Alan D. George |
FPT | 2 |
| 2021 | Real-time, High-resolution Depth Upsampling on Embedded AcceleratorsabstractHigh-resolution, low-latency apps in computer vision are ubiquitous in today’s world of mixed-reality devices. These innovations provide a platform that can leverage the improving technology of depth sensors and embedded accelerators to enable higher-resolution, lower-latency processing for 3D scenes using depth-upsampling algorithms. This research demonstrates that filter-based upsampling algorithms are feasible for mixed-reality apps using low-power hardware accelerators. The authors parallelized and evaluated a depth-upsampling algorithm on two different devices: a reconfigurable-logic FPGA embedded within a low-power SoC; and a fixed-logic embedded graphics processing unit. We demonstrate that both accelerators can meet the real-time requirements of 11 ms latency for mixed-reality apps. 1 David Langerman, Alan D. George |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2021 | Reconfigurable Framework for Resilient Semantic Segmentation for Space ApplicationsabstractDeep learning (DL) presents new opportunities for enabling spacecraft autonomy, onboard analysis, and intelligent applications for space missions. However, DL applications are computationally intensive and often infeasible to deploy on radiation-hardened (rad-hard) processors, which traditionally harness a fraction of the computational capability of their commercial-off-the-shelf counterparts. Commercial FPGAs and system-on-chips present numerous architectural advantages and provide the computation capabilities to enable onboard DL applications; however, these devices are highly susceptible to radiation-induced single-event effects (SEEs) that can degrade the dependability of DL applications. In this article, we propose Reconfigurable ConvNet (RECON), a reconfigurable acceleration framework for dependable, high-performance semantic segmentation for space applications. In RECON, we propose both selective and adaptive approaches to enable efficient SEE mitigation. In our selective approach, control-flow parts are selectively protected by triple-modular redundancy to minimize SEE-induced hangs, and in our adaptive approach, partial reconfiguration is used to adapt the mitigation of dataflow parts in response to a dynamic radiation environment. Combined, both approaches enable RECON to maximize system performability subject to mission availability constraints. We perform fault injection and neutron irradiation to observe the susceptibility of RECON and use dependability modeling to evaluate RECON in various orbital case studies to demonstrate a 1.5–3.0× performability improvement in both performance and energy efficiency compared to static approaches. Sebastian Sabogal, Alan D. George, Gary Crum |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2020 | Reconfigurable Framework for Environmentally Adaptive Resilience in Hybrid Space SystemsabstractDue to ongoing innovations in both sensor technology and spacecraft autonomy, onboard space processing continues to be outpaced by the escalating computational demands required for next-generation missions. Commercial-off-the-shelf, hybrid system-on-chips, combining fixed-logic CPUs with reconfigurable-logic FPGAs, present numerous architectural advantages that address onboard computing challenges. However, commercial devices are highly susceptible to space radiation and require dependable computing strategies to mitigate radiation-induced single-event effects. Depending upon the mission, the dynamics of the near-Earth space-radiation environment expose spacecraft to radiation fluxes that can vary by several orders of magnitude. By adopting an adaptive approach to dependable computing, spacecraft computers can reconfigure system resources to efficiently accommodate changing environmental conditions to maximize system performance while satisfying availability constraints throughout the mission. In this article, we propose Hybrid, Adaptive, Reconfigurable Fault Tolerance (HARFT), a reconfigurable framework for environmentally adaptive resilience in hybrid space systems. Furthermore, we describe a methodology to model adaptive systems, represented as phased-mission systems using Markov chains, subject to the near-Earth space-radiation environment, using a combination of orbital perturbation, geomagnetic field, and single-event effect rate prediction tools. We apply this methodology to evaluate the HARFT architecture using various static and adaptive strategies for several orbital case studies and demonstrate the achievable performability gains. Sebastian Sabogal, Alan D. George |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2018 | Deep Learning for Hyperspectral Image Classification on Embedded PlatformsabstractHyperspectral image (HSI) analysis refers to the processes used to identify and classify objects photographed using equipment that can image photons from a broad range of the electromagnetic spectrum. Downlinking such large images from space on radiation-resistant platforms with limited on-board computing power takes a large amount of time, memory, and other mission-critical resources. Performing such analysis in space before downlinking all images will save these resources by enabling a subset of images of interest to be downloaded rather than the entire set. The goal of this study is to benchmark and evaluate HSI-classification methods which incorporate deep learning on embedded platforms with limited computing resources. Support Vector Machine (SVM), Multi-Layer Perceptron (MLP), and Convolutional Neural Network (CNN) are the classification methods used in this study. These algorithms were executed on a desktop PC and two embedded platforms: the ODROID-C2 and the Raspberry Pi 3B. Accuracy, run-time, and memory benchmarks determined the optimal model for each platform. Based on results gathered in this research, CNN classification is recommended for the desktop PC due to its high accuracy of 97%. MLP classification is recommended for the embedded platforms under study, as it showcased the shortest run-time and second-highest accuracy. Siddharth Balakrishnan, David Langerman, Evan Gretok, Alan D. George |
IPAS | 4 |
| 2018 | Accelerating Real-Time, High-Resolution Depth Upsampling on FPGAsabstractWhile the popularity of high-resolution, computer-vision applications (e.g. mixed reality, autonomous vehicles) is increasing, there have been complementary advances in time-of-flight depth sensor resolution and quality. These advances in time-of-flight sensors provide a platform for new research into real-time, depth-upsampling algorithms targeted at high-resolution video systems with low-latency requirements. This paper describes a case study in which a previously developed bilateral-filter-style upsampling algorithm is profiled, parallelized, and accelerated on an FPGA using high-level synthesis tools from Xilinx. We show that our accelerated algorithm can effectively upsample the resolution and reduce the noise of time-of-flight sensors. We also demonstrate that this algorithm exceeds the real-time requirements of 90 frames per second necessitated by mixed-reality hardware, achieving a lower-bound speedup of 40 times over the fastest CPU-only version. David Langerman, Sebastian Sabogal, Barath Ramesh, Alan D. George |
IPAS | 4 |
| 2018 | Onboard Processing With Hybrid and Reconfigurable Computing on Small SatellitesabstractDue to the increasing demands of onboard sensor and autonomous processing, one of the principal needs and challenges for future spacecraft is onboard computing. Space computers must provide high performance and reliability (which are often at odds), using limited resources (power, size, weight, and cost), in an extremely harsh environment (due to radiation, temperature, vacuum, and vibration). As spacecraft shrink in size, while assuming a growing role for science and defense missions, the challenges for space computing become particularly acute. For example, processing capabilities on CubeSats (smaller class of SmallSats) have been extremely limited to date, often featuring microcontrollers with performance and reliability barely sufficient to operate the vehicle let alone support various sensor and autonomous applications. This article surveys the challenges and opportunities of onboard computers for small satellites (SmallSats) and focuses upon new concepts, methods, and technologies that are revolutionizing their capabilities, in terms of two guiding themes: hybrid computing and reconfigurable computing. These innovations are of particular need and value to CubeSats and other Smallsats. With new technologies, such as CHREC Space Processor (CSP), we demonstrate how system designers can exploit hybrid and reconfigurable computing on SmallSats to harness these advantages for a variety of purposes, and we highlight several recent missions by NASA and industry that feature these principles and technologies. Alan D. George |
Proc. IEEE | 1 |
| 2017 | Optimizing FPGA Performance, Power, and Dependability with Linear ProgrammingabstractField-programmable gate arrays (FPGA) are an increasingly attractive alternative to traditional microprocessor-based computing architectures in extreme-computing domains, such as aerospace and supercomputing. FPGAs offer several resource types that offer different tradeoffs between speed, power, and area, which make FPGAs highly flexible for varying application computational requirements. However, since an application’s computational operations can map to different resource types, a major challenge in leveraging resource-diverse FPGAs is determining the optimal distribution of these operations across the device’s available resources for varying FPGA devices, resulting in an extremely large design space. In order to facilitate fast design-space exploration, this article presents a method based on linear programming (LP) that determines the optimal operation distribution for a particular device and application with respect to performance, power, or dependability metrics. Our LP method is an effective tool for exploring early designs by quickly analyzing thousands of FPGAs to determine the best FPGA devices and operation distributions, which significantly reduces design time. We demonstrate our LP method’s effectiveness with two case studies involving dot-product and distance-calculation kernels on a range of Virtex-5 FPGAs. Results show that our LP method selects optimal distributions of operations to within an average of 4% of actual values. Nicholas Wulf, Alan D. George, Ann Gordon-Ross |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2016 | Novo-G#: a multidimensional torus-based reconfigurable cluster for molecular dynamicsabstractSummary Molecular dynamics (MD) is a large‐scale, communication‐intensive problem that has been the subject of high‐performance computing research and acceleration for years. Not surprisingly, the most success in accelerating MD comes from specialized systems such as the Anton machine. In this paper, we describe Novo‐G# (novo‐jee‐sharp), a multi‐node reconfigurable system designed for the acceleration of communication‐intensive scientific problems in general, and MD in particular. This system provides a high‐bandwidth, low‐latency 3D torus network to allow direct communication between kernels running on multiple field‐programmable gate arrays. We also present a performance model for Novo‐G# running the 3D Fast Fourier Transform (FFT) kernel that forms the core of MD simulations. We validate the model against published Anton performance data and through initial hardware experiments on Novo‐G#. Finally, through simulation studies, we show that this system at scale performs better than specialized systems like Anton and outperforms established CPU‐based clusters like Blue Gene/Q by an order of magnitude for the 3D FFT kernel, with greater flexibility and lower costs. Copyright © 2015 John Wiley & Sons, Ltd. Abhijeet Lawande, Alan D. George, Herman Lam |
Concurr. Comput. Pract. Exp. | 2 |
| 2016 | Analysis of Fixed, Reconfigurable, and Hybrid Devices with Computational, Memory, I/O, & Realizable-Utilization MetricsabstractThe modern processor landscape is a varied and diverse community. As such, developers need a way to quickly and fairly compare various devices for use with particular applications. This article expands the authors’ previously published computational-density metrics and presents an analysis of a new generation of various device architectures, including CPU, DSP, FPGA, GPU, and hybrid architectures. Also, new memory metrics are added to expand the existing suite of metrics to characterize the memory resources on various processing devices. Finally, a new relational metric, realizable utilization (RU) , is introduced, which quantifies the fraction of the computational density metric that an application achieves within an individual implementation. The RU metric can be used to provide valuable feedback to application developers and architecture designers by highlighting the upper bound on specific application optimization and providing a quantifiable measure of theoretical and realizable performance. Overall, the analysis in this article quantifies the performance tradeoffs among the architectures studied, the memory characteristics of different device types, and the efficiency of device architectures. Justin Richardson, Alan D. George, Herman Lam |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2016 | A Framework for Evaluating and Optimizing FPGA-Based SoCs for Aerospace ComputingabstractOn-board processing systems are often deployed in harsh aerospace environments and must therefore adhere to stringent constraints such as low power, small size, and high dependability in the presence of faults. Field-programmable gate arrays (FPGAs) are often an attractive option for designers seeking low-power, high-performance devices. However, unlike nonreconfigurable devices, radiation effects can alter an FPGA’s functionality instead of just the device’s data, requiring designers to consider fault-tolerant strategies to mitigate these effects. In this article, we present a framework to ease these system design challenges and aid designers in considering a broad range of devices and fault-tolerant strategies for on-board processing, highlighting the most promising options and tradeoffs early in the design process. This article focuses on the power, dependability, and lifetime evaluation metrics, which our framework calculates and leverages to evaluate the effectiveness of varying system-on-chip (SoC) designs. Finally, we use our framework to evaluate SoC designs for a case study on a hyperspectral-imaging (HSI) mission to demonstrate our framework’s ability to identify efficient and effective SoC designs. Nicholas Wulf, Alan D. George, Ann Gordon-Ross |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2015 | Comparative analysis of OpenCL vs. HDL with image-processing kernels on Stratix-V FPGAabstractApplication development with hardware description languages (HDLs) such as VHDL or Verilog involves numerous productivity challenges, limiting the potential impact of reconfigurable computing (RC) with FPGAs in high-performance computing. Major challenges with HDL design include steep learning curves, large and complex codes, long compilation times, and lack of development standards across platforms. A relative newcomer to RC, the Open Computing Language (OpenCL) reduces productivity hurdles by providing a platform-independent, C-based programming language. In this study, we conduct a performance and productivity comparison between three image-processing kernels (Canny edge detector, Sobel filter, and SURF feature-extractor) developed using Altera's SDK for OpenCL and traditional VHDL. Our results show that VHDL designs achieved a more efficient use of resources (59% to 70% less logic), however, both OpenCL and VHDL designs resulted in similar timing constraints (255MHzmax<; 325MHz). Furthermore, we observed a 6× increase in productivity when using OpenCL development tools, as well as the ability to efficiently port the same OpenCL designs without change to three different RC platforms, with similar performance in terms of frequency and resource utilization. Kenneth Hill, Stefan Craciun, Alan D. George, Herman Lam |
ASAP | 3 |
| 2015 | CMT-bone: A Mini-App for Compressible Multiphase Turbulence Simulation SoftwareabstractDesigned with the goal of mimicking key features of real HPC workloads, mini-apps have become an important tool for co-design. An investigation of mini-app behavior can provide system designers with insight into the impact of architectures, programming models, and tools on application performance. Mini-apps can also serve as a platform for fast algorithm design space exploration, allowing the application developers to evaluate their design choices before significantly redesigning the application codes. Consequently, it is prudent to develop a mini-app alongside the full blown application it is intended to represent. In this paper, we present CMT-bone a mini-app for the compressible multiphase turbulence (CMT) application, CMT-nek, being developed to extend the physics of the CESAR Nek5000 application code. CMT-bone consists of the most computationally intensive kernels of CMT-nek and the communication operations involved in nearest-neighbor updates and vector reductions. The mini-app represents CMT-nek in its most mature state and going forward it will be developed in parallel with the CMT-nek application to keep pace with key new performance impacting changes. We describe these kernels and discuss the role that CMT-bone has played in enabling interdisciplinary collaboration by allowing application developers to work with computer scientists on performance optimization on current architectures and performance analysis on notional future systems. Nalini Kumar, Mrugesh Sringarpure, Tania Banerjee, Jason Hackl, S. Balachandar 0001, Herman Lam, Alan D. George, Sanjay Ranka |
CLUSTER | 7 |
| 2015 | Low-level PGAS computing on many-core processors with TSHMEMabstractSummary Diminishing returns from increased clock frequencies and instruction‐level parallelism have forced computer architects to adopt architectures that exploit wider parallelism through multiple processor cores. While emerging many‐core architectures have progressed at a remarkable rate, concerns arise regarding the performance and productivity of numerous parallel‐programming tools for application development. Development of parallel applications on many‐core processors often requires developers to familiarize themselves with unique characteristics of a target platform while attempting to maximize performance and maintain correctness of their applications. The family of partitioned global address space (PGAS) programming models comprises the current state of the art in balancing performance and programmability. One such PGAS approach is SHMEM, a lightweight, shared‐memory programming library that has demonstrated high performance and productivity potential for parallel‐computing systems with distributed‐memory architectures. In the paper, we present research, design, and analysis of a new SHMEM infrastructure specifically crafted for low‐level PGAS on modern and emerging many‐core processors featuring dozens of cores and more. Our approach (with a new library known as TSHMEM) is investigated and evaluated atop two generations of Tilera architectures, which are among the most sophisticated and scalable many‐core processors to date, and is intended to enable similar libraries atop other architectures now emerging. In developing TSHMEM, we explore design decisions and their impact on parallel performance for the Tilera TILE‐Gx and TILEPro many‐core architectures, and then evaluate the designs and algorithms within TSHMEM through microbenchmarking and applications studies with other communication libraries. Our results with barrier primitives provided by the Tilera libraries show dissimilar performance between the TILE‐Gx and TILEPro; therefore, TSHMEM's barrier design takes an alternative approach and leverages the on‐chip mesh network to provide consistent low‐latency performance. In addition, our experiments with TSHMEM show that naive collective algorithms consistently outperformed linear distributed collective algorithms when executed in an SMP‐centric environment. In leveraging these insights for the design of TSHMEM, our approach outperforms the OpenSHMEM reference implementation, achieves similar to positive performance over OpenMP and OSHMPI atop MPICH, and supports similar libraries in delivering high‐performance parallel computing to emerging many‐core systems. Copyright © 2015 John Wiley & Sons, Ltd. Bryant C. Lam, Alan D. George, Herman Lam, Vikas Aggarwal |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | Low-Overhead FPGA Middleware for Application Portability and ProductivityabstractReconfigurable computing devices such as field-programmable gate arrays (FPGAs) offer advantages over fixed-logic CPU and GPU architectures, including improved performance, superior power efficiency, and reconfigurability. The challenge of FPGA application development, however, has limited their acceptance in high-performance computing and high-performance embedded computing applications. FPGA development carries similar difficulties to hardware design, requiring that developers iterate through register-transfer level designs with cycle-level accuracy. Furthermore, the lack of hardware and software standards between FPGA platforms limits productivity and application portability, and makes porting applications between heterogeneous platforms a time-consuming and often challenging process. Recent efforts to improve FPGA productivity using high-level synthesis tools and languages show promise, but platform support remains limited and typically is left as a challenge for developers. To address these issues, we present RC Middleware (RCMW), a novel middleware that improves productivity and enables application and tool portability by abstracting away platform-specific details. RCMW provides an application-centric development environment, exposing only the resources and standardized interfaces required by an application, independent of the underlying platform. We demonstrate the portability and productivity benefits of RCMW using four heterogeneous platforms from three vendors. Our results indicate that RCMW enables application productivity and improves developer productivity, and that these benefits are achieved with less than 7% performance and 3% area overhead on average. Robert Kirchgessner, Alan D. George, Greg Stitt |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2013 | A scalable RC architecture for mean-shift clusteringabstractThe mean-shift algorithm provides a unique non-parametric and unsupervised clustering solution to image segmentation and has a proven record of very good performance for a wide variety of input images. It is essential to image processing because it provides the initial and vital steps to numerous object recognition and tracking applications. However, image segmentation using mean-shift clustering is widely recognized as one of the most compute-intensive tasks in image processing, and suffers from poor scalability with respect to the image size (N pixels) and number of iterations (k): O(kN2). Our novel approach focuses on creating a scalable hardware architecture fine-tuned to the computational requirements of the mean-shift clustering algorithm. By efficiently parallelizing and mapping the algorithm to reconfigurable hardware, we can effectively cluster hundreds of pixels independently. Each pixel can benefit from its own dedicated pipeline and can move independently of all other pixels towards its respective cluster. By using our mean-shift FPGA architecture, we achieve a speedup of three orders of magnitude with respect to our software baseline. Stefan Craciun, Gongyu Wang, Alan D. George, Herman Lam, José C. Príncipe |
ASAP | 3 |
| 2013 | Reconfigurable computing middleware for application portability and productivityabstractReconfigurable computing (RC) devices such as field-programmable gate arrays (FPGAs) offer significant advantages over fixed-logic, many-core CPU and GPU architectures, including increased performance for many computationally challenging applications, superior power efficiency, and reconfigurability. Difficulties of using FPGAs, however, has limited their acceptance in high-performance computing (HPC) and high-performance embedded computing (HPEC) applications. These difficulties stem from a lack of standards between FPGA platforms and the complexities of hardware design, and lead to higher costs and time to market over competing technologies. Differences in FPGA platform resources such as the type and number of FPGAs, memories and interconnects, as well as vendor-specific procedural APIs and hardware interfaces, inhibits application portability and code reusability. Despite efforts to reduce FPGA application design complexity through technologies such as high-level synthesis (HLS) tools, platform support and portability remains limited, and is typically left as a challenge for application developers. In this paper, we present a novel RC Middleware (RCMW), an extensible framework which enables FPGA application portability and enhances developer productivity by providing an application-centric development environment. Developers focus specifically on the optimal resources and interfaces required by their application, and RCMW handles the mapping and translation of those resources onto a target platform. We demonstrate that RCMW enables application portability over three heterogeneous platforms from two vendors, using both Xilinx and Altera FPGAs, with less than 10% performance and area overhead for several application kernels, and microbenchmarks for the common case. We present the productivity benefits of RCMW, showing that RCMW reduces required number of hardware and software driver lines of code and total development time with respect to native platform deployment methods for several application kernels. Robert Kirchgessner, Alan D. George, Herman Lam |
ASAP | 2 |
| 2012 | Communication visualization for bottleneck detection of high-level synthesis applicationsabstractHigh-level synthesis tools increase FPGA productivity but can decrease performance compared to register-transfer level designs. To help optimize high-level synthesis applications, we introduce a bottleneck detection tool that provides a developer with a visualization of communication bandwidth between all application processes, while identifying potential bottlenecks via color coding. We evaluated the tool using third-party applications to identify and optimize bottlenecks in just several minutes, which achieved speedups ranging from 1.25x to 2.18x compared to the original FPGA execution. Overhead was modest with less than 2% resource overhead and 3% frequency overhead. John Curreri, Greg Stitt, Alan D. George |
FPGA | 3 |
| 2012 | VirtualRC: a virtual FPGA platform for applications and tools portabilityabstractNumerous studies have shown significant performance and power benefits of field-programmable gate arrays (FPGAs). Despite these benefits, FPGA usage has been limited by application design complexity caused largely by the lack of code and tool portability across different FPGA platforms, which prevents design reuse. This paper addresses the portability challenge by introducing a framework of architecture and middleware for virtualization of FPGA platforms, collectively named VirtualRC. Experiments show modest overhead of 5-6% in performance and 1% in area, while enabling portability of 11 applications and two high-level synthesis tools across three physical platforms. Robert Kirchgessner, Greg Stitt, Alan D. George, Herman Lam |
FPGA | 3 |
| 2012 | Overhead and reliability analysis of algorithm-based fault tolerance in FPGA systemsabstractCommercial SRAM-based, field-programmable gate arrays (FPGAs) have the capability to provide space applications with the necessary performance, energy-efficiency, and adaptability to meet next-generation mission requirements. However, mitigating an FPGA's susceptibility to radiation-induced faults is challenging. Triple-modular redundancy (TMR) techniques are traditionally used to mitigate radiation effects, but TMR incurs substantial overheads such as increased area and power requirements. In order to reduce these overheads while still providing sufficient radiation mitigation, we propose the use of algorithm-based fault tolerance (ABFT). We investigate the effectiveness of hardware-based ABFT logic in COTS FPGAs by developing multiple ABFT-enabled matrix multiplication designs, carefully analyzing resource usage and reliability tradeoffs, and proposing design modifications for higher reliability. We perform fault-injection testing on a Xilinx Virtex-5 platform to validate these ABFT designs, measure design vulnerability, and compare ABFT effectiveness to other fault-tolerance methods. Our hybrid ABFT design reduces total design vulnerability by 99% while only incurring 25% overhead over a baseline, non-protected design. Adam Jacobs, Grzegorz Cieslewski, Alan D. George |
FPL | 3 |
| 2012 | RCML: An Environment for Estimation Modeling of Reconfigurable Computing SystemsabstractReconfigurable computing (RC) is emerging as a promising area for embedded computing, in which complex systems must balance performance, flexibility, cost, and power. The difficulty associated with RC development suggests improved strategic planning and analysis techniques can save significant development time and effort. This article presents a new abstract modeling language and environment, the RC Modeling Language (RCML), to facilitate efficient design space exploration of RC systems at the estimation modeling level, that is, before building a functional implementation. Two integrated analysis tools and case studies, one analytical and one simulative, are presented illustrating relatively accurate automated analysis of systems modeled in RCML. Casey Reardon, Brian Holland, Alan D. George, Greg Stitt, Herman Lam |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2012 | SCF: A Framework for Task-Level Coordination in Reconfigurable, Heterogeneous SystemsabstractHeterogeneous computing systems comprised of accelerators such as FPGAs, GPUs, and manycore processors coupled with standard microprocessors are becoming an increasingly popular solution for future computing systems due to their higher performance and energy efficiency. Although programming languages and tools are evolving to simplify device-level design, programming such systems is still difficult and time-consuming largely due to system-wide challenges involving communication between heterogeneous devices, which currently require ad hoc solutions. Most communication frameworks and APIs which have dominated parallel application development for decades were developed for homogeneous systems, and hence cannot be directly employed for hybrid systems. To solve this problem, this article presents the System Coordination Framework (SCF), which employs message passing to transparently enable communication between tasks described using different programming tools (and languages), and running on heterogeneous processing devices of systems from domains ranging from embedded systems to High-Performance Computing (HPC) systems. By hiding low-level architectural details of the underlying communication from an application designer, SCF can improve application development productivity, provide higher levels of application portability, and offer rapid design-space exploration of different task/device mappings. In addition, SCF enables custom communication synthesis that exploits mechanisms specific to different devices and platforms, which can provide performance improvements over generic solutions employed previously. Our results indicate a performance improvement of 28× and 682× by employing FPGA devices for two applications presented in this article, while simultaneously improving the developer productivity by approximately 2.5 to 5 times by using SCF. Vikas Aggarwal, Greg Stitt, Alan D. George, Changil Yoon |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2012 | Reconfigurable Fault Tolerance: A Comprehensive Framework for Reliable and Adaptive FPGA-Based Space ComputingabstractCommercial SRAM-based, field-programmable gate arrays (FPGAs) have the potential to provide space applications with the necessary performance to meet next-generation mission requirements. However, mitigating an FPGA’s susceptibility to single-event upset (SEU) radiation is challenging. Triple-modular redundancy (TMR) techniques are traditionally used to mitigate radiation effects, but TMR incurs substantial overheads such as increased area and power requirements. In order to reduce these overheads while still providing sufficient radiation mitigation, we propose a reconfigurable fault tolerance (RFT) framework that enables system designers to dynamically adjust a system’s level of redundancy and fault mitigation based on the varying radiation incurred at different orbital positions. This framework includes an adaptive hardware architecture that leverages FPGA reconfigurable techniques to enable significant processing to be performed efficiently and reliably when environmental factors permit. To accurately estimate upset rates, we propose an upset rate modeling tool that captures time-varying radiation effects for arbitrary satellite orbits using a collection of existing, publically available tools and models. We perform fault-injection testing on a prototype RFT platform to validate the RFT architecture and RFT performability models. We combine our RFT hardware architecture and the modeled upset rates using phased-mission Markov modeling to estimate performability gains achievable using our framework for two case-study orbits. Adam Jacobs, Grzegorz Cieslewski, Alan D. George, Ann Gordon-Ross, Herman Lam |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2011 | SHMEM+: A multilevel-PGAS programming model for reconfigurable supercomputingabstractReconfigurable Computing (RC) systems based on FPGAs are becoming an increasingly attractive solution to building parallel systems of the future. Applications targeting such systems have demonstrated superior performance and reduced energy consumption versus their traditional counterparts based on microprocessors. However, most of such work has been limited to small system sizes. Unlike traditional HPC systems, lack of integrated, system-wide, parallel-programming models and languages presents a significant design challenge for creating applications targeting scalable, reconfigurable HPC systems. In this article, we extend the traditional Partitioned Global Address Space (PGAS) model to provide a multilevel integration of memory, which simplifies development of parallel applications for such systems and improves developer productivity. The new multilevel-PGAS programming model captures the unique characteristics of reconfigurable HPC systems, such as the existence of multiple levels of memory hierarchy and heterogeneous computation resources. Based on this model, we extend and adapt the SHMEM communication library to become what we call SHMEM+, the first known SHMEM library enabling coordination between FPGAs and CPUs in a reconfigurable, heterogeneous HPC system. Applications designed with SHMEM+ yield improved developer productivity compared to current methods of multidevice RC design and exhibit a high degree of portability. In addition, our design of SHMEM+ library itself is portable and provides peak communication bandwidth comparable to vendor-proprietary versions of SHMEM. Application case studies are presented to illustrate the advantages of SHMEM+. Vikas Aggarwal, Alan D. George, Changil Yoon, Kishore Yalamanchili, Herman Lam |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2011 | An analytical model for multilevel performance prediction of Multi-FPGA systemsabstractPower limitations in semiconductors have made explicitly parallel device architectures such as Field-Programmable Gate Arrays (FPGAs) increasingly attractive for use in scalable systems. However, mitigating the significant cost of FPGA development requires efficient design-space exploration to plan and evaluate a range of potential algorithm and platform choices prior to implementation. The authors propose the RC Amenability Test for Scalable Systems (RATSS), an analytical model which enables straightforward, fast, and reasonably accurate performance prediction prior to implementation by extending current modeling concepts to multi-FPGA designs. RATSS provides a comprehensive strategic model to evaluate applications based on the computation and communication requirements of the algorithm and capabilities of the FPGA platform. The RATSS model targets data-parallel applications on current scalable FPGA systems. Three case studies with RATSS demonstrate nearly 90% prediction accuracy as compared to corresponding implementations. Brian Holland, Alan D. George, Herman Lam, Melissa C. Smith |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2011 | Platform-aware bottleneck detection for reconfigurable computing applicationsabstractReconfigurable Computing (RC) has the potential to provide substantial performance benefits and yet simultaneously consume less power than traditional microprocessors or GPUs. While experimental performance analysis of RC applications has previously been shown crucial for achieving this potential, existing methods still require application designers to manually locate bottlenecks and determine appropriate optimizations, typically requiring significant designer expertise and effort. Worse, the diversity of platforms employed by RC applications further complicates the process of detecting bottlenecks and formulating optimizations. To address these shortcomings, we first discuss our platform-template system, which enables a performance analysis tool to perform more accurate bottleneck detection and achieve a higher degree of portability across diverse FPGA systems. We then provide details for our implementation of these concepts and techniques in the Reconfigurable Computing Application Performance (ReCAP) tool. Next, we present a taxonomy of common RC bottlenecks, providing associated detection and optimization strategies for each bottleneck, which we use to populate ReCAP's knowledge base for bottleneck detection. Finally, we demonstrate the utility of our approach via two application case studies across a total of three platforms. Seth Koehler, Greg Stitt, Alan D. George |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2010 | Optimizing rapidIO architectures for onboard processingabstractIn this article, we study optimization of a RapidIO network and FPGA-based computation engines to address the taxing requirements of a set of real-time Ground-Moving Target Indicators (GMTI) and Synthetic Aperture Radar (SAR) kernels for Space-Based Radar (SBR). By employing a RapidIO hardware testbed and validated simulation, we determine key trade-offs in design of reconfigurable systems for GMTI and SAR in terms of processing, memory, and network throughput. In addition, we study considerations for timely delivery of latency-sensitive control traffic present in many satellite applications. Based on our results, we propose architectural modifications to further improve performance of SBR systems. David Bueno, Chris Conger, Alan D. George |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2010 | Performance Analysis Framework for High-Level Language Applications in Reconfigurable ComputingabstractHigh-Level Languages (HLLs) for Field-Programmable Gate Arrays (FPGAs) facilitate the use of reconfigurable computing resources for application developers by using familiar, higher-level syntax, semantics, and abstractions, typically enabling faster development times than with traditional Hardware Description Languages (HDLs). However, programming at a higher level of abstraction is typically accompanied by some loss of performance as well as reduced transparency of application behavior, making it difficult to understand and improve application performance. While runtime tools for performance analysis are often featured in development with traditional HLLs for sequential and parallel programming, HLL-based development for FPGAs has an equal or greater need yet lacks these tools. This article presents a novel and portable framework for runtime performance analysis of HLL applications for FPGAs, including an automated tool for performance analysis of designs created with Impulse C, a commercial HLL for FPGAs. As a case study, this tool is used to successfully locate performance bottlenecks in a molecular dynamics kernel in order to gain speedup. John Curreri, Seth Koehler, Alan D. George, Brian Holland, Rafael García |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2010 | A Simulation Framework for Rapid Analysis of Reconfigurable Computing SystemsabstractReconfigurable computing (RC) is rapidly emerging as a promising technology for the future of high-performance and embedded computing, enabling systems with the computational density and power of custom-logic hardware and the versatility of software-driven hardware in an optimal mix. Novel methods for rapid virtual prototyping, performance prediction, and evaluation are of critical importance in the engineering of complex reconfigurable systems and applications. These techniques can yield insightful tradeoff analyses while saving valuable time and resources for researchers and engineers alike. The research described herein provides a methodology for mapping arbitrary applications to targeted reconfigurable platforms in a simulation environment called RCSE. By splitting the process into two domains, the application and simulation domains, characterization of each element can occur independently and in parallel, leading to fast and accurate performance prediction results for large and complex systems. This article presents the design of a novel framework for system-level simulative performance prediction of RC systems and applications. The article also presents a set of case studies analyzing two applications, Hyperspectral Imaging (HSI) and Molecular Dynamics (MD), across three disparate RC platforms within the simulation framework. The validation results using each of these applications and systems show that our framework can quickly obtain performance prediction results with reasonable accuracy on a variety of platforms. Finally, a set of simulative case studies are presented to illustrate the various capabilities of the framework to quickly obtain a wide range of performance prediction results and power consumption estimates. Casey Reardon, Eric Grobelny, Alan D. George, Gongyu Wang |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2010 | Characterization of Fixed and Reconfigurable Multi-Core Devices for Application AccelerationabstractAs on-chip transistor counts increase, the computing landscape has shifted to multi- and many-core devices. Computational accelerators have adopted this trend by incorporating both fixed and reconfigurable many-core and multi-core devices. As more, disparate devices enter the market, there is an increasing need for concepts, terminology, and classification techniques to understand the device tradeoffs. Additionally, computational performance, memory performance, and power metrics are needed to objectively compare devices. These metrics will assist application scientists in selecting the appropriate device early in the development cycle. This article presents a hierarchical taxonomy of computing devices, concepts and terminology describing reconfigurability, and computational density and internal memory bandwidth metrics to compare devices. Chris Massie, Alan D. George, Justin Richardson, Kunal Gosrani, Herman Lam |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2009 | Bitstream relocation with local clock domains for partially reconfigurable FPGAsabstractPartial Reconfiguration (PR) of FPGAs presents many opportunities for application design flexibility, enabling tasks to dynamically swap in and out of the FPGA without entire system interruption. However, mapping a task to any available PR region (PRR) requires a unique partial bitstream for each PRR. This replication can introduce significant overheads in terms of bitstream storage and communication requirements. Previous research in partial bitstream relocation can alleviate these overheads by transforming a single partial bitstream to map to any available PRR. However, careful steps are necessary to ensure proper functionality of relocated partial bitstreams and may result in clock routing inefficiencies. These routing inefficiencies can be alleviated by using regional clock resources introduced in the Virtex-4 FPGAs to implement local clock domains. PRRs can internally drive local clock domains, enabling each PRR to vary its clock frequency with respect to a single global clock signal, as opposed to sending multiple global clock signals (one for each desired clock frequency) to each PRR. We introduce this novel local clock domain (LCD) concept, which provides enhanced PR design flexibility. However, integration of LCDs and partial bitstream relocation introduces new challenges. In this paper, we identify motivating application domains for this integration, analyze integration benefits, and provide a detailed integration methodology. Adam Flynn, Ann Gordon-Ross, Alan D. George |
DATE | 3 |
| 2009 | Exploiting Partially Reconfigurable FPGAs for Situation-Based Reconfiguration in Wireless Sensor NetworksabstractWireless sensor networks (WSNs) are typically composed of very small, battery-operated devices (sensor nodes) containing simple microprocessors with few computational resources. However, the rapidly increasing popularity of WSNs has placed increased computational demands upon these systems, due to increasingly complex operating environments and enhanced data-sensing technology. Whereas introducing more powerful microprocessors into sensor nodes addresses these demands, sensor nodes do not contain sufficient energy reserves to support these microprocessors. In this paper, we present a partially reconfigurable FPGA-based architecture and methodology to provide increased WSN flexibility and computational resources, resulting in superior power consumption and performance compared to a microprocessor capable of satisfying similar demands. Rafael García, Ann Gordon-Ross, Alan D. George |
FCCM | 3 |
| 2009 | Reconfigurable fault tolerance: A framework for environmentally adaptive fault mitigation in spaceabstractCommercial SRAM-based FPGAs have the potential to provide aerospace applications with the necessary performance to meet next generation mission requirements. However, the susceptibility of these devices to radiation in the form of single-event upsets is a significant drawback. TMR techniques are traditionally used to mitigate these effects, but with an overwhelming amount of extra area and power. We propose a framework for reconfigurable fault tolerance which enables systems engineers to dynamically change the amount of redundancy and fault mitigation that is used in an FPGA design. This approach leverages the reconfigurable nature of the FPGA to allow significant processing to be performed safely and reliably when environmental factors permit. Phased-mission Markov modeling is used to estimate performability gains that can be achieved using the framework for two case-study orbits. Adam Jacobs, Alan D. George, Grzegorz Cieslewski |
FPL | 2 |
| 2009 | A generalized, distributed analysis system for optimization of Parallel ApplicationsabstractDeveloping a high performance parallel application is difficult. An application must often be analyzed and optimized by the programmer before reaching an acceptable level of performance. Performance tools that collect and visualize performance data can reduce the effort needed by the user in the nontrivial optimization process. However, as the size of the performance dataset grows, it becomes nearly impossible for the user to manually examine the data and find performance issues. To address this problem, we have developed a new analysis system to automatically detect, diagnose, and possibly resolve bottlenecks. In this paper, we present the architecture and the distributed, peer-to-peer processing mechanism of a programming model-independent analysis system, which includes a range of useful analyses such as scalability analysis and common-bottleneck detection. We then describe the details of an initial sequential implementation of the system that has been integrated into our parallel performance wizard (PPW) tool. Finally, we provide correctness and performance results for this initial version and demonstrate the effectiveness of the system through two case studies. Hung-Hsun Su, Max Billingsley, Alan D. George |
IPDPS | 3 |
| 2009 | RAT: RC Amenability Test for Rapid Performance PredictionabstractWhile the promise of achieving speedup and additional benefits such as high performance per watt with FPGAs continues to expand, chief among the challenges with the emerging paradigm of reconfigurable computing is the complexity in application design and implementation. Before a lengthy development effort is undertaken to map a given application to hardware, it is important that a high-level parallel algorithm crafted for that application first be analyzed relative to the target platform, so as to ascertain the likelihood of success in terms of potential speedup. This article presents the RC Amenability Test, or RAT, a methodology and model developed for this purpose, supporting rapid exploration and prediction of strategic design tradeoffs during the formulation stage of application development. Brian Holland, Karthik Nagarajan, Alan D. George |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2008 | Performance Analysis with High-Level Languages for High-Performance Reconfigurable ComputingabstractHigh-Level Languages (HLLs) for FPGAs (Field-Programmable Gate Arrays) facilitate the use of reconfigurable computing resources for application developers by using familiar, higher-level syntax, semantics, and abstractions, typically enabling faster development times than with traditional Hardware Description Languages (HDLs). However, this abstraction is typically accompanied by some loss of performance as well as reduced transparency of application behavior, making it difficult to understand and improve application performance. While runtime tools for performance analysis are often featured in development with traditional HLLs for serial and parallel programming, HLL-based applications for FPGAs have an equal or greater need yet lack these tools. This paper presents a novel and portable framework for runtime performance analysis of HLL applications for FPGAs, including a prototype tool for performance analysis with Impulse C, a commercial HLL for FPGAs. As a case study, this tool is used to locate performance bottlenecks in a molecular dynamics application. John Curreri, Seth Koehler, Brian Holland, Alan D. George |
FCCM | 4 |
| 2008 | Scalable and Portable Architecture for Probability Density Function Estimation on FPGAsabstractThis paper describes the design, development, and analysis of a scalable and portable architecture for multi-dimensional, non-parametric PDF estimation using Gaussian kernels on FPGAs. Karthik Nagarajan, Brian Holland, K. Clint Slatton, Alan D. George |
FCCM | 4 |
| 2008 | Parallel performance wizard: A performance analysis tool for partitioned global-address-space programmingabstractGiven the complexity of parallel programs, developers often must rely on performance analysis tools to help them improve the performance of their code. While many tools support the analysis of message-passing programs, no tool exists that fully supports programs written in programming models that present a partitioned global address space (PGAS) to the programmer, such as UPC and SHMEM. Existing tools with support for message-passing models cannot be easily extended to support PGAS programming models, due to the differences between these paradigms. Furthermore, the inclusion of implicit and one-sided communication in PGAS models renders many of the analyses performed by existing tools irrelevant. For these reasons, there exists a need for a new performance tool capable of handling the challenges associated with PGAS models. In this paper, we first present background research and the framework for Parallel Performance Wizard (PPW), a modularized, event-based performance analysis tool for PGAS programming models. We then discuss features of PPW and how they are used in the analysis of PGAS applications. Finally, we illustrate how one would use PPW in the analysis and optimization of PGAS applications by presenting a small case study using the PPW version 1.0 implementation. Hung-Hsun Su, Max Billingsley, Alan D. George |
IPDPS | 3 |
| 2008 | Real-time performance analysis of Adaptive Link RateabstractHigh speed links are widely deployed in modern day computer networks to meet the ever growing needs for increasing data bandwidth. However, with the increase in the link rate, the power consumption of the network interfaces increases exponentially, compounding growing concerns about network power consumption. Fortunately, network traffic characteristics show that rapid link rates are not always required. During times of reduced network traffic, the Adaptive Link Rate (ALR) mechanism allows link rates to be reduced with little impact on network performance. Current research has focused on policies to control when and how to change link rates, and have shown promising energy savings. However, these works have been largely simulative, and have not addressed many of the challenges involved in implementation. In this paper, we develop a hardware prototype ALR system and address real-time challenges involved in realizing such an implementation. We also identify new considerations for control policy development given current technology capabilities as well as future projections. Baoke Zhang, Karthik Sabhanatarajan, Ann Gordon-Ross, Alan D. George |
LCN | 4 |
| 2008 | Performance analysis challenges and framework for high-performance reconfigurable computing
Seth Koehler, John Curreri, Alan D. George |
Parallel Comput. | 3 |
| 2008 | Optimization of checkpointing-related I/O for high-performance parallel and distributed computing
Rajagopal Subramaniyan, Eric Grobelny, R. Scott Studham, Alan D. George |
J. Supercomput. | 4 |
| 2007 | RapidIO for radar processing in advanced space systemsabstractSpace-based radar is a suite of applications that presents many unique system design challenges. In this paper, we investigate use of RapidIO, a new high-performance embedded systems interconnect, in addressing issues associated with the high network bandwidth requirements of real-time ground moving target indicator (GMTI), and synthetic aperture Radar (SAR) applications in satellite systems. Using validated simulation, we study several critical issues related to the RapidIO network and algorithms under study. The results show that RapidIO is a promising platform for space-based radar using emerging technology, providing network bandwidth to enable parallel computation previously unattainable in an embedded satellite system. David Bueno, Chris Conger, Alan D. George, Ian A. Troxel, Adam Leko |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2006 | Adaptable and Autonomic Mission Manager for Dependable Aerospace ComputingabstractAs NASA and other agencies continue to undertake ever challenging remote sensing missions, the ability of satellites and space probes to diagnose and autonomously recover from faults will be paramount. In addition, a more pronounced use of radiation-susceptible components in order to reduce cost makes the challenge of ensuring system dependability even more difficult. To meet these and other needs, a processing platform for space is currently under development at Honeywell Inc. and the University of Florida for an upcoming NASA New Millennium Program mission. Among other features, the platform deploys an autonomic software management system to increase system dependability. In addition, a mission manager has been investigated and developed to provide an autonomous means to adapt to environmental conditions and system failures. This paper provides a detailed analysis of the management system philosophy with a focus on the adaptable mission manager. A case study is presented that highlights the dependability and performance improvement provided by the mission manager and autonomic health monitoring scheme Ian A. Troxel, Alan D. George |
DASC | 2 |
| 2006 | Reconfigurable computing with multiscale data fusion for remote sensingabstractRecent advances in sensor technologies have resulted in tremendous increases in the amount of data collected for imaging applications such as airborne and space-based remote sensing of the Earth. Data acquisition and dissemination systems need to perform more processing than ever before to support real-time applications and reduce bandwidth demands on the downlink. FPGA-based reconfigurable computing systems are emerging as cost-effective solutions that offer enormous computation potential in the embedded systems arena. Research in this paper explores the potential capability offered by deploying reconfigurable computing systems in a remote sensing system by means of a commonly employed application. Multiple designs for a multiscale data-fusion algorithm were developed for an FPGA-based platform. These designs are used to demonstrate speedup over processor-based solutions and study the demands posed by such applications upon the system. Due to the vast number of sensors inputs, such applications pose high demands on the memory capacity and bandwidth which becomes a critical factor in determining the overall system performance. Results of our experiments depict that over an order of magnitude improvement can be obtained with efficient designs and appropriate hardware resources. Projections of enhanced performance with emerging system architectures are also presented. Vikas Aggarwal, Alan D. George, K. Clint Slatton |
FPGA | 2 |
| 2006 | Simulative analysis of dynamic scheduling heuristics for reconfigurable computing of parallel applicationsabstractThe demand for processing power has been ever increasing with the growth of high-performance computing (HPC) applications and so have the constraints restricting the solutions to such requirements. High-performance distributed and parallel computing and custom-built, hardware-based computing have attempted to address this problem with some success but at a substantial cost. Recently, systems augmented with Field-Programmable Gate Arrays (FPGAs) offering a fusion of traditional parallel and distributed machines with customizable and dynamically reconfigurable hardware have emerged as a cost-effective alternative to traditional systems. However, providing a robust runtime environment for such systems to which HPC users have become accustomed has been fraught with numerous challenges. Dynamic scheduling of large-scale HPC applications in such parallel reconfigurable computing (RC) environments is one such challenge and has not been sufficiently studied to our knowledge. In this paper, we simulatively analyze the performance of several common HPC scheduling heuristics that can be used by an automated job management service to schedule application tasks on a parallel RC system. We also present a performance prediction model which the scheduling heuristics employ to schedule several common HPC applications on a collection of typical FPGA processing platforms. Rajagopal Subramaniyan, Ian A. Troxel, Alan D. George, Melissa C. Smith |
FPGA | 3 |
| 2006 | Ethernet Adaptive Link Rate (ALR): Analysis of a MAC Handshake ProtocolabstractIn this paper, a handshake protocol at the medium access control (MAC) layer is proposed and analyzed for dynamically changing the link rate in the network interface card (NIC), adapting to network utilization, and thus decreasing average power consumption. Simulation results show that this protocol can be used to change link rate in Ethernet network devices without causing user-perceivable delays Himanshu Anand, Casey Reardon, Rajagopal Subramaniyan, Alan D. George |
LCN | 4 |
| 2006 | Power-Proxying on the NIC: A Case Study with the Gnutella File-Sharing ProtocolabstractEdge devices such as desktop and laptop computers constitute a majority of the devices connected to the Internet today. Peer-to-peer (P2P) file-sharing applications generally require edge devices to maintain network presence whenever possible to enhance the robustness of the file-sharing network, which in turn can lead to considerable wastage of energy. We show that energy can be saved by permitting edge devices to enter into standby state and still maintain network connectivity by proxying protocols in the network interface card (NIC) Pradeep Purushothaman, Mukund Navada, Rajagopal Subramaniyan, Casey Reardon, Alan D. George |
LCN | 5 |
| 2006 | Reconfigurable supercomputing - Is high-performance reconfigurable computing the next supercomputing paradigm?abstractHigh-Performance Reconfigurable Computers (HPRCs) based on the combination of conventional processors and FPGAs have been gaining attention in the past few years. Their benefits were particularly harnessed in compute-intensive integer applications. However, there has been doubt that the same benefits can be attained for general scientific applications. Fortunately, the trend in reconfigurable chip sizes and diversity of resources may be relieving some of those concerns. Yet, with the hardware reconfigurability, it is feared that domain scientists have to learn how to design hardware if they were to use such machines effectively. In order to address the overarching question, this panel will address the following questions: Can FPGAs deliver order-of-magnitude performance gains in scientific floating-point applications in the foreseeable future? Can programming HPRCs programmability become similar to that of HPCs in its level of difficulty? What are the major developments in the industry or the community that make all this possible? Tarek A. El-Ghazawi, Dave Bennett, Daniel S. Poznanovic, Allan Cantle, Keith D. Underwood, Rob Pennington, Duncan A. Buell, Alan D. George, Volodymyr V. Kindratenko |
SC | 8 |
| 2006 | Poster reception - Parallel performance wizard: a performance analysis tool for partitioned global-address-space programming modelsabstractScientific programmers must optimize the total time-to-solution, the combination of software development and refinement time and actual execution time. The increasing complexity at all levels of supercomputing architectures, coupled with advancements in sequential performance and a growing degree of hardware parallelism, has increasingly placed the bulk of the time-to-solution cost into the software development and tuning phase. Performance analysis tools have been useful for reducing the time-to-solution for message-passing applications; however, there is insufficient tool support for programs developed using Global-Address-Space (GAS) programming models. With the aim of maximizing user productivity, the Parallel Performance Wizard (PPW) fills this void by providing a full range of visualizations and analyses specifically designed for GAS models. To facilitate accurate instrumentation and measurement of GAS programs in PPW, a portable, model-independent performance tool interface (GASP) has been developed and successfully used with Berkeley UPC. Adam Leko, Hung-Hsun Su, Dan Bonachea, Bryan Golden, Max Billingsley, Alan D. George |
SC | 6 |
| 2004 | Simulative Analysis of the RapidIO Embedded Interconnect Architecture for Real-Time, Network-Intensive ApplicationsabstractRapidIO is an emerging standard for switched interconnection of processors and boards in embedded systems. We use discrete-event simulation to evaluate and prototype RapidIO-based systems with respect to their performance in an environment targeted towards space-based radar applications. This application class makes an ideal test case for a RapidIO feasibility study due to high system throughput requirements and real-time processing constraints. Our results show that a baseline RapidIO system is well suited to space-based radar, providing significant improvements over typical bus-based architectures. Our results also show that extensions to the RapidIO protocol such as cut-through routing and transmitter-controlled flow-control would provide minimal performance improvements for the applications under study. David Bueno, Adam Leko, Chris Conger, Ian A. Troxel, Alan D. George |
LCN | 5 |
| 2004 | The next frontier for communications networks: power management
Kenneth J. Christensen, Chamara Gunaratne, Bruce Nordman, Alan D. George |
Comput. Commun. | 4 |
| 2003 | Performance analysis of HP AlphaServer ES80 vs. SAN-based clustersabstractThe last decade has introduced various affordable computing platforms to the parallel computing community. Distributed shared-memory systems and clusters built with commercial-off-the-shelf (COTS) parts and interconnected with high-performance networks have proven to be serious alternatives to expensive supercomputers in terms of both performance and cost. HP's new AlphaServer ES80 is an example of distributed shared-memory systems, while SCI and Myrinet are the two most widely used high-performance interconnects in building parallel-computing clusters. In this study, we experimentally compare the performance of these parallel computer systems. The emphasis is pointing out the strengths and weakness of the HPs AlphaServer ES80 in comparison with high-performance SCI and Myrinet clusters. We evaluated the systems in terms of sustainable memory bandwidth, interprocess communication and overall parallel computation performance using various widely-accepted benchmarks such as STREAM, PALLAS PMB-MP1, and NAS2.3 parallel suite. It was observed that the HPs AlphaServer ES80, executing Linux, provides remarkable computing power while its communication subsystem cannot handle heavy loads of small messages as effectively. Burt Gordon, Sarp Oral, Hung-Hsun Su, Alan D. George |
IPCCC | 5 |
| 2003 | High-performance embedded computing for Conventional Matched-Field ProcessingabstractAdvanced sonar algorithms with complex-wave propagation processing are of critical importance in the field of acoustic signal processing, as they possess the ability to localize signal sources more precisely in a cluttered environment. Conventional Matched-Field Processing (CMFP) is such an algorithm that provides range and depth results by employing environmental parameters. However, the enhancement of features is limited by the extensive computational and memory requirements of the algorithm even for a small problem size. High-performance computing, when applied to matched-field processing on a microprocessor-based distributed system, can provide solutions for this challenging problem in performance, scalability, and cost. Based on domain decomposition techniques, two parallel algorithms for matched-field processing are introduced in this paper for in-situ signal processing. Performance results collected on low-power distributed embedded systems are presented in terms of execution times and parallel efficiencies. The results of these analyses demonstrate that parallel in-situ processing holds the potential to meet the needs of matched-field processing in a scalable fashion. Keonwook Kim, Alan D. George |
IPCCC | 2 |
| 2003 | A User-level Multicast Performance Comparison of Scalable Coherent Interface and Myrinet InterconnectsabstractThis paper compares and evaluates the multicast performance of two of the most widely deployed system-area networks (SANs), Dolphin's scalable coherent interface (SCI) and Myricom's Myrinet. Both networks deliver low latency and high bandwidth to applications, but do not support multicast in hardware. We compared SCI and Myrinet in terms of their user-level performance using various software-based multicast algorithms under various networking and multicasting scenarios. The strengths and weaknesses of each network are comparatively presented in terms of numerous metrics, such as multicast completion latency, CPU utilization, link concentration and concurrency. Sarp Oral, Alan D. George |
LCN | 2 |
| 2003 | Multiple-path execution for chip multiprocessors
Matthew C. Chidester, Alan D. George, Matthew A. Radlinski |
J. Syst. Archit. | 2 |
| 2003 | A high-performance communication service for parallel computing on distributed DSP systems
James Kohout, Alan D. George |
Parallel Comput. | 2 |
| 2002 | Gigabit COTS Ethernet Switch Evaluation for AvionicsabstractThe evolving network needs of both commercial and military aircraft are expanding the requirements from slow speeds (1 Mbps) to over 1 Gbps for many new aircraft designs, driven primarily by advanced video technology. Low-cost Ethernet switch technology is a key enabler for advanced aircraft system interconnection. Our research focuses on the evaluation of several gigabit Ethernet (GE) COTS switches to meet avionic requirements. The results indicate that GE COTS switches are beginning to approach the required performance for future avionic applications. Jack L. Meier, Alan D. George, Sarp Oral |
LCN | 3 |
| 2002 | A Comparative Throughput Analysis of Scalable Coherent Interface and MyrinetabstractIt has become increasingly popular to construct large parallel computers by connecting many inexpensive nodes built with commercial-off-the-shelf (COTS) parts. These clusters can be built at a much lower cost than traditional supercomputers of comparable performance. A key decision that will greatly affect the overall performance of the cluster is the method used to connect the nodes together. Choosing the best interconnect and topology is not at all trivial since performance and cost will change as the system size is scaled. This paper presents throughput models used for the analysis and comparison of performance in two leading system area networks (SAN), Myrinet and Scalable Coherent Interface (SCI). First, analytical models for throughput are developed by determining the theoretical bandwidth of all internal buses and links that are part of the interconnect architecture. Then, experiments are conducted to measure the actual bandwidth available at each of these components, and the models are calibrated so they accurately represent the experimental results. Finally, the models are used to compare the maximum throughput of Myrinet and SCI systems with respect to system size and overall dollar cost. Sarp Millich, Alan D. George, Sarp Oral |
LCN | 2 |
| 2002 | Multicast Performance Analysis for High-Speed Torus NetworksabstractOverall efficiency of high-performance computing clusters not only relies on the computing power of the individual nodes, but also on the performance that the underlying network can provide to the computational application. Although modern high-performance networks, especially system area networks (SAN), have high unicast performance, they do not support multicast communication in hardware. This research experimentally evaluates the performance of various protocols for unicast-based and path-based multicast communication on high-speed torus networks. Software-based multicast performance results of selected algorithms on a 16-node Scalable Coherent Interface (SCI) torus are given. The strengths and weaknesses of the various protocols are illustrated in terms of startup and completion latency, CPU utilization, and link utilization and concurrency. Sarp Oral, Alan D. George |
LCN | 2 |
| 2002 | Design and Analysis of a Dynamically Reconfigurable Network ProcessorabstractThe combination of high-performance processing power and flexibility found in network processors (NPs) has made them a good solution for today's packet processing needs. Similarly, the emerging technology of reconfigurable computing (RC) has made advances in packet processing as well as other point-solution markets. Current NP designs offer configurable elements but generally do not use dynamic RC techniques for run-time reconfiguration. Incorporating RC into NP designs to enhance packet processing is a natural progression for both of these emerging technologies. This paper presents the simulation results of a novel design for a RC-enhanced NP based on the Intel IXP1200 NIC design philosophy. The enhanced NP's performance is compared to that of the baseline NP in terms of three normalized traffic patterns and a case-study traffic pattern based on a military application. The results demonstrate that the enhanced NP significantly outperforms the baseline NP design in terms of latency for prioritized traffic that is non-uniform. Ian A. Troxel, Alan D. George, Sarp Oral |
LCN | 2 |
| 2001 | Performance Analysis of Flat and Layered Gossip Services for Failure Detection and Consensus in Scalable Heterogeneous ClustersabstractGossip protocols and services provide a means bywhich failures can be detected in large, distributed systemsin an asynchronous manner without the limits associatedwith reliable multicasting for group communications.Gossiping with consensus can take place throughout thesystem via a flat structure, or it can be hierarchicallydistributed across cooperating layers of nodes. In thispaper, the performance of flat and layered protocols isanalyzed on an experimental testbed in terms of consensustime and scalability. Performance associated with layeredgossip is analyzed with varying group sizes and is shownto scale well in a heterogeneous environment. Alan D. George, Krishnakanth Sistla, Robert W. Todd, Raghukul Tilak |
IPDPS | 1 |
| 2001 | Achieving Scalable Cluster System Analysis and Management with a Gossip-Based Network ServiceabstractClusters of workstations are increasingly used for applications requiring high levels of both performance and reliability. Certain fundamental services are highly desirable to achieve these twin goals of network-based cluster system analysis and management. Among these services is the ability to detect network and node failures and the capability to efficiently determine computer and network load levels. Furthermore, the ability to allow for the distribution of administrative directives is also integral to the goal of cluster management. This paper presents a scalable approach to providing these vital support capabilities for distributed computing integrated into a cluster management system. Previous approaches to cluster management have suffered from problems of scalability and the inability to properly support heterogeneous systems in a non-proprietary fashion. This cluster management system employs gossip techniques to address the problem of scalability in network-based system management. The results of two case studies show that the cluster management system is scalable and has little adverse impact on the performance of sequential and parallel applications running on the managed system. David E. Collins, Alan D. George, R. A. Quander |
LCN | 2 |
| 2001 | Comparative performance analysis of directed flow control for real-time SCI
Robert W. Todd, Matthew C. Chidester, Alan D. George |
Comput. Networks | 3 |
| 2000 | Reliability Modeling of SCI Ring-Based TopologiesabstractReliability prediction is an important factor in the study of multiprocessor and cluster interconnects. One such interconnect is the scalable coherent interface (SCI), a point-to-point, ring-based interconnect that can be connected into switched ring topologies. However, the reliability of SCI-based topologies cannot be deduced from earlier network reliability work as link failures within an SCI-based interconnect are not independent of one another. This paper presents the results of a reliability study on 1D and 2D k-ary n-cube switching fabrics for SCI based on the elimination of rings rather than of links. The reliability models were created in UltraSAN and verified using analytical modeling. The results show that the reliability of a single-ring system can be greatly enhanced by the addition of a redundant ring yet the reliability of a torus does not increase significantly with additional redundant rings. Hence, the benefit of adding redundant rings is dependent both upon the topology and the degree of reliability sought. M. A. Sarwar, Alan D. George, David E. Collins |
LCN | 2 |
| 2000 | Real-time sonar beamforming on high-performance distributed computers
Alan D. George, Jeff Markwell, Ryan Fogarty |
Parallel Comput. | 1 |