EDBT 2026 Demo / reviewers in the wild / expert
Bingyi Zhang
dblp:18/267
· DBLP profile ↗
23ranked-venue papers
13as first author
20since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 11 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Model-Architecture Codesign for High-Performance and Energy-Efficient SAR ATR on FPGAabstractSynthetic Aperture Radar (SAR) Automatic Target Recognition (ATR) is a crucial technology in remote sensing. SAR devices, such as those on satellites like Sentinel-1A, collect SAR data for various applications. However, state-of-the-art CNN-based approaches for SAR ATR have high computational complexity. This makes them unsuitable for deployment on resource-limited platforms. In this paper, we present a novel model-architecture co-design for SAR ATR on Field Programmable Gate Arrays (FPGAs). Our proposed co-design consists of: (1) A novel multi-layer Graph Neural Network (GNN) model for SAR ATR with low computational complexity and a small number of parameters. (2) An optimized hardware architecture on FPGAs, ensuring low-latency, high-throughput, and energy-efficient execution of the GNN model. For model design, we leverage attention mechanisms to enhance classification accuracy. We then use knowledge distillation to train a simplified GNN model. This reduces computational cost while maintaining nearly the same accuracy. Additionally, we apply model pruning to further reduce computational complexity without compromising performance. To maximize computational parallelism, we introduce a customized hardware accelerator on FPGA. This accelerator uses the Scatter-Gather paradigm to efficiently manage the irregular computation and memory access patterns inherent to GNNs. For productivity, we develop parameterized hardware templates using High-level Synthesis (HLS) and create user-friendly Application Programming Interfaces (APIs). We deploy our accelerator design on both datacenter and embedded FPGAs. Specifically, we test it on Alveo U280, AMD/Xilinx PYNQ-Z1, and ZCU104, which are comparable to state-of-the-art space-grade FPGAs. We evaluate the proposed model on multiple datasets, including MSTAR, SynthWakeSAR, and GBSAR. Compared with the state-of-the-art models, the proposed GNN model achieves higher or comparable accuracy with substantially less computational complexity. Furthermore, compared with CPU and GPU implementations, our FPGA-based accelerators consistently deliver lower latency, improved throughput, and higher energy efficiency. Bingyi Zhang, Rajgopal Kannan, Carl E. Busart, Viktor Prasanna 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | ViTeGNN: Towards Versatile Inference of Temporal Graph Neural Networks on FPGAabstractTemporal Graph Neural Networks (TGNNs) are powerful models to capture temporal, structural, and contextual information on temporal graphs, outperforming other methods in many high-impact downstream tasks. However, achieving high-performance TGNN inference in production environments is challenging because TGNN models suffer from high computation complexity and intrinsic temporal data dependency that hinders data parallelism. In addition, real-world TGNN applications have different latency and throughput requirements. This work presents ViTeGNN, a versatile TGNN inference solution for memory-based TGNNs on FPGAs. ViTeGNN performs algorithm-model-architecture co-design to meet the latency and throughput requirements of real-world TGNN applications. Besides the vanilla inference mode ViTeGNN-bal that updates embeddings for nodes interacting with others, we propose ViTeGNN-lat and ViTeGNN-thpt, optimized for latency and throughput. Our model optimizations include a lightweight method to compute attention scores and a related temporal neighbor pruning strategy to reduce computation and memory accesses. These are holistically coupled with key hardware optimizations that leverage the FPGA hardware. We propose a novel hardware module to execute the complex neighbor update process efficiently. To ensure similar accuracy vis-á-vis the original model, the simplified models are trained using the knowledge distillation technique. We propose a unified hardware design that supports all of these three inference modes without FPGA reconfiguration. Enabled by our flexible hardware architecture, we further propose ViTeGNN-auto, which automatically selects the best inference mode at runtime based on latency and throughput requirements, guided by our accurate performance model. We evaluate the performance of the proposed hardware accelerator on five real-world datasets. ViTeGNN-bal reduces the computation complexity by an average of 62% and memory accesses by an average of 36% with only 0.0042 accuracy loss. Compared with state-of-the-art implementations on CPU and GPU, our FPGA implementation achieves$53.9/26.0/16.1\times$speedup and$8.2/4.0/2.5\times$speedup for ViTeGNN-lat/-bal/-thpt, respectively. Bingyi Zhang, Rajgopal Kannan, Carl E. Busart, Viktor Prasanna 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | Accelerating ViT Inference on FPGA through Static and Dynamic PruningabstractVision Transformers (ViTs) have achieved state-of-the-art accuracy on various computer vision tasks. However, their high computational complexity prevents them from being applied to many real-world applications. Weight and token pruning methods are well-known in reducing ViT model complexity. However, naively combining and integrating both the methods results in irregular computation patterns leading to accuracy drops and difficulties in hardware acceleration. This limits the net complexity reduction offered by integrating such pruning methods. To address the above challenges, we propose a comprehensive algorithm-hardware codesign for accelerating ViT on FPGA through simultaneous pruning - combining static weight pruning and dynamic token pruning. For algorithm design, we systematically combine a hardware-aware structured block-pruning method for pruning model parameters and a dynamic token pruning method for removing unimportant token vectors. Moreover, we design a novel training algorithm to reduce the accuracy drop due to such simultaneous pruning. For hardware design, we develop a novel hardware accelerator for executing the pruned model. The proposed hardware design employs multi-level parallelism with a load-balancing strategy to efficiently deal with the irregular computation pattern presented by the two pruning approaches. Moreover, we develop an efficient hardware mechanism for executing the on-the-fly token pruning. We apply our codesign approach to the widely used DeiT-Small model. We implement the proposed accelerator on a state-of-the-art FPGA. The evaluation results show that the proposed algorithm reduces computation complexity by up to 3.4× with ≈ 3% accuracy drop and a model compression ratio of up to 1.6×. Compared with state-of-the-art implementation on CPU, GPU, and FPGA, our codesign on FPGA achieves an average latency reduction of 12.8×, 3.2×, and 0.7 – 2.1×, respectively. Dhruv Parikh, Shouyi Li, Bingyi Zhang, Rajgopal Kannan, Carl E. Busart, Viktor Prasanna 0001 |
FCCM | 3 |
| 2024 | GCV-Turbo: End-to-end Acceleration of GNN-based Computer Vision Tasks on FPGAabstractGraph neural networks (GNNs) have recently em-powered various novel computer vision (CV) tasks. In GNN-based CV tasks, a combination of CNN layers and GNN layers or only GNN layers are employed. This paper introduces GCV-Turbo, a domain-specific accelerator on FPGA for end-to-end acceleration of GNN-based CV tasks. GCV-Turbo consists of two key components: (1) a novel hardware architecture optimized for the computation kernels in both CNNs and GNNs using the same set of computation resources. (2) a compiler that takes a user-defined model as input, performs end-to-end optimization for the computation graph of a given GNN-based CV task, and produces optimized code for hardware execution. The hardware architecture and the compiler work synergistically to support a variety of GNN-based CV tasks. We implement GCV-Turbo on a state-of-the-art FPGA and evaluate its performance across six representative GNN-based CV tasks with diverse input data modalities (e.g., image, human skeleton, point cloud). Compared with state-of-the-art CPU (GPU) implementations, GCV-Turbo achieves an average latency reduction of 68.4× (4.1x) on these six GNN-based CV tasks. Moreover, GCV-Turbo supports the execution of the standalone CNNs or GNNs, achieving performance comparable to that of state-of-the-art CNN (GNN) accelerators for widely used CNN-only (GNN-only) models. Bingyi Zhang, Rajgopal Kannan, Carl E. Busart, Viktor Prasanna 0001 |
FCCM | 1 |
| 2024 | A Single Graph Convolution is All You Need: Efficient Grayscale Image ClassificationabstractImage classifiers for domain-specific tasks like Synthetic Aperture Radar Automatic Target Recognition (SAR ATR) and chest X-ray classification often rely on convolutional neural networks (CNNs). These networks, while powerful, experience high latency due to the number of operations they perform, which can be problematic in real-time applications. Many image classification models are designed to work with both RGB and grayscale datasets, but classifiers that operate solely on grayscale images are less common. Grayscale image classification has critical applications in fields such as medical imaging and SAR ATR. In response, we present a novel grayscale image classification approach using a vectorized view of images. By leveraging the lightweight nature of Multi-Layer Perceptrons (MLPs), we treat images as vectors, simplifying the problem to grayscale image classification. Our approach incorporates a single graph convolutional layer in a batch-wise manner, enhancing accuracy and reducing performance variance. Additionally, we develop a customized accelerator on FPGA for our model, incorporating several optimizations to improve performance. Experimental results on benchmark grayscale image datasets demonstrate the effectiveness of our approach, achieving significantly lower latency (up to $16 \times$ less on MSTAR) and competitive or superior performance compared to state-of-the-art models for SAR ATR and medical image classification. Jacob Fein-Ashley, Sachini Wickramasinghe, Bingyi Zhang, Rajgopal Kannan, Viktor Prasanna 0001 |
ICIP | 3 |
| 2024 | HitGNN: High-Throughput GNN Training Framework on CPU+Multi-FPGA Heterogeneous PlatformabstractAs the size of real-world graphs increases, training Graph Neural Networks (GNNs) has become time-consuming and requires acceleration. While previous works have demonstrated the potential of utilizing FPGA for accelerating GNN training, few works have been carried out to accelerate GNN training with multiple FPGAs due to the necessity of hardware expertise and substantial development effort. To this end, we propose HitGNN, a framework that enables users to effortlessly map GNN training workloads onto a CPU+Multi-FPGA platform for acceleration. In particular, HitGNN takes the user-defined synchronous GNN training algorithm, GNN model, and platform metadata as input, determines the design parameters based on the platform metadata, and performs hardware mapping onto the CPU+Multi-FPGA platform, automatically. HitGNN consists of the following building blocks: (1) high-level application programming interfaces (APIs) that allow users to specify various synchronous GNN training algorithms and GNN models with only a handful of lines of code; (2) a software generator that generates a host program that performs mini-batch sampling, manages CPU-FPGA communication, and handles workload balancing among the FPGAs; (3) an accelerator generator that generates GNN kernels with optimized datapath and memory organization. We show that existing synchronous GNN training algorithms such as DistDGL and PaGraph can be easily deployed on a CPU+Multi-FPGA platform using our framework, while achieving high training throughput. Compared with the state-of-the-art frameworks that accelerate synchronous GNN training on a multi-GPU platform, HitGNN achieves up to 27.21× bandwidth efficiency, and up to 4.26× speedup using much less compute power and memory bandwidth than GPUs. In addition, HitGNN demonstrates good scalability to 16 FPGAs on a CPU+Multi-FPGA platform. Yi-Chien Lin, Bingyi Zhang, Viktor Prasanna 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | VisionAGILE: A Versatile Domain-Specific Accelerator for Computer Vision TasksabstractThe emergence of diverse machine learning (ML) models has led to groundbreaking revolutions in computer vision (CV). These ML models include convolutional neural networks (CNNs), graph neural networks (GNNs), and vision transformers (ViTs). However, existing hardware accelerators designed for CV lack the versatility to support various ML models, potentially limiting their applicability to real-world scenarios. To address this limitation, we introduce VisionAGILE, a domain-specific accelerator designed to be versatile and capable of accommodating a range of ML models, including CNNs, GNNs, and ViTs. VisionAGILE comprises a compiler, a runtime system, and a hardware accelerator. For the hardware accelerator, we develop a novel unified architecture with a flexible data path and memory organization to support the computation primitives in various ML models. Regarding the compiler design, we develop a unified compilation workflow that maps various ML models to the proposed hardware accelerator. The runtime system executes dynamic sparsity exploitation to reduce inference latency and dynamic task scheduling for workload balance. The compiler, the runtime system, and the hardware accelerator work synergistically to support a variety of ML models in CV, enabling low-latency inference. We deploy the hardware accelerator on a state-of-the-art data center FPGA (Xilinx Alveo U250). We evaluate VisionAGILE on diverse ML models for CV, including CNNs, GNNs, hybrid models (comprising both CNN and GNN), and ViTs. The experimental results indicate that, compared with state-of-the-art CPU (GPU) implementations, VisionAGILE achieves a speedup of$81.7\times$($4.8\times$) in terms of latency. Evaluated on standalone CNNs, GNNs, and ViTs, VisionAGILE demonstrates comparable or higher performance with state-of-the-art CNN accelerators, GNN accelerators, and ViT accelerators, respectively. Bingyi Zhang, Rajgopal Kannan, Carl E. Busart, Viktor Prasanna 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | Exploiting On-Chip Heterogeneity of Versal Architecture for GNN Inference AccelerationabstractGraph Neural Networks (GNNs) have revolutionized many Machine Learning (ML) applications, such as social network analysis, bioinformatics, etc. GNN inference can be accelerated by exploiting data sparsity in the input graph, vertex features, and intermediate data in GNN computations. For dynamic sparsity exploitation, we leverage the heterogeneous computing capabilities of AMD Versal ACAP architecture to accelerate GNN inference. We develop a custom hardware module that executes the sparse primitives of the computation kernel on the Programmable Logic (PL) and efficiently computes the dense primitives using the AI Engine (AIE). To exploit data sparsity during inference, we devise a runtime kernel mapping strategy that dynamically assigns computation tasks to the PL and AIE based on data sparsity. Our implementation on the VCK5000 ACAP platform leads to superior performance compared with the state-of-the-art implementations on CPU, GPU, ACAP, and other custom GNN accelerators. Compared with these implementations, we achieve significant average runtime speedup across various models and datasets of 162.42x, 17.01×, 9.90×, and 27.23×, respectively. Furthermore, for Graph Convolutional Network (GCN) inference, our approach leads to a speedup of 3.9-96.7× compared to designs using PL only on the same ACAP device. Paul Chen, Pavan Manjunath, Sasindu Wijeratne, Bingyi Zhang, Viktor Prasanna 0001 |
FPL | 4 |
| 2023 | Dynasparse: Accelerating GNN Inference through Dynamic Sparsity ExploitationabstractGraph Neural Network (GNN) inference is used in many real-world applications. Data sparsity in GNN inference, including sparsity in the input graph and the GNN model, offer opportunities to further speed up inference. Also, many pruning techniques have been proposed for model compression that increase the data sparsity of GNNs.We propose Dynasparse, a comprehensive hardware-software codesign on FPGA to accelerate GNN inference through dynamic sparsity exploitation. For this, we decouple the GNN computation kernels from the basic computation primitives, and explore hardware-software codesign as follows: 1) Hardware design: We propose a novel unified accelerator design on FPGA to efficiently execute various computation primitives. We develop a customized soft processor that is tightly coupled with the accelerator to execute a runtime system. Moreover, we develop efficient hardware mechanisms to profile the data sparsity and perform on-the-fly data format transformation to prepare the input data for various computation primitives; 2) Software design: We develop a runtime system that works synergistically with the accelerator to perform dynamic kernel-to-primitive mapping based on data sparsity. We implement Dynasparse on a state-of-the-art FPGA platform, Xilinx Alveo U250, and evaluate the design using widely used GNN models (GCN, GraphSAGE, GIN and SGC). For the above GNN models and various input graphs, the proposed accelerator and dynamic kernel-to-primitive mapping reduces the inference latency by 3.73× on the average compared with the static mapping strategies employed in the state-of-the-art GNN accelerators. Compared with state-of-the-art CPU (GPU) implementations, Dynasparse achieves up to 56.9× (2.37×) speedup in end-to-end latency. Compared with state-of-the-art FPGA implementations, Dynasparse achieves 2.7× speedup in accelerator execution latency. Bingyi Zhang, Viktor Prasanna 0001 |
IPDPS | 1 |
| 2023 | GraphAGILE: An FPGA-Based Overlay Accelerator for Low-Latency GNN InferenceabstractThis article presents GraphAGILE, a domain-specific FPGA-based overlay accelerator for graph neural network (GNN) inference. GraphAGILE consists of (1)a novel unified architecture designwith aninstruction set, and (2)a compilerbuilt upon the instruction set that can quickly generate optimized code. Due to the proposed instruction set architecture (ISA) and the compiler, GraphAGILE does not require any FPGA reconfiguration when performing inference on various GNN models and input graphs. For the architecture design, we propose a novel hardware module named Adaptive Computation Kernel (ACK), that can execute various computation kernels of GNNs, including general matrix multiplication (GEMM), sparse-dense matrix multiplication (SpDMM), and sampled dense-dense matrix multiplication (SDDMM). The compiler takes the specifications of a GNN model and the graph meta data (e.g., the number of vertices and edges) as input, and generates a sequence of instructions for inference execution. We develop the following compiler optimizations to reduce inference latency: 1) computation order optimization that automatically reorders the computation graph to reduce the total computation complexity, 2) layer fusion that merges adjacent layers to reduce data communication volume, 3) data partitioning with a partition-centric execution scheme that partitions the input graph to fit the available on-chip memory of FPGA, 4) kernel mapping that automatically selects execution mode for ACK, and performs task scheduling to overlap computation with data communication and achieves dynamic load balance. We implement GraphAGILE on a state-of-the-art FPGA platform, Xilinx Alveo U250. GraphAGILE can execute widely used GNN models, including GCN, GAT, GIN, GraphSAGE, SGC and other GNN models supported by GraphGym. Experimental results show that GraphAGILE achieves up to$47.1\times$($3.9\times$) reduction in end-to-end latency, including the latency of compilation and hardware execution, compared with the state-of-the-art implementations on CPU (GPU), and achieves up to$2.9\times$reduction in hardware execution latency compared with the state-of-the-art FPGA accelerators. Bingyi Zhang, Hanqing Zeng, Viktor Prasanna 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | HP-GNN: Generating High Throughput GNN Training Implementation on CPU-FPGA Heterogeneous PlatformabstractGraph Neural Networks (GNNs) have shown great success in many applications such as recommendation systems, molecular property prediction, traffic prediction, etc. Recently, CPU-FPGA heterogeneous platforms have been used to accelerate many applications by exploiting customizable data path and abundant user-controllable on-chip memory resources of FPGAs. Yet, accelerating and deploying GNN training on such platforms requires not only expertise in hardware design but also substantial development efforts. Yi-Chien Lin, Bingyi Zhang, Viktor Prasanna 0001 |
FPGA | 2 |
| 2022 | DecGNN: A Framework for Mapping Decoupled GNN Models onto CPU-FPGA Heterogeneous PlatformabstractA well-known issue in mini-batch GNN inference is neighborhood explosion. This results in two challenges for its hardware acceleration: (1) high computation and communication costs resulting in high latency, and (2) low computation-to-communication ratio leading to low hardware utilization. To address these challenges, we propose a hardware mapping framework following the recently proposed GNN design principle of model depth-receptive field decoupling. We show that Decoupled GNNs enjoy significantly higher computation-to-communication ratio, therefore, are more suitable for hardware acceleration. To efficiently map Decoupled GNNs onto CPU-FPGA heterogeneous platforms, we propose the following model-architecture co-optimizations: (1) Model instantiation: according to the bandwidth and computation resources, we determine the number of GNN layers to achieve high hardware utilization; (2) Neighbor selection: to meet the application constraint, we select a small number of important neighbors surrounding the target vertices to improve the throughput without sacrificing accuracy; (3) Hardware mapping: given the model and neighborhood defined above, we determine the accelerator parameters based on our novel hardware templates, enabling fast computation of GNN inference workloads. We evaluate our framework on two state-of-the-art FPGA platforms, using two models (GCN, GraphSAGE). Experiments show that the resulting designs achieve high hardware utilization 88%-94% and significant speedup (1.1x-2.5x) compared with the implementations on state-of-the-art CPU-GPU platform. Bingyi Zhang, Hanqing Zeng, Viktor Prasanna 0001 |
FPGA | 1 |
| 2022 | Accurate, Low-latency, Efficient SAR Automatic Target Recognition on FPGAabstractSynthetic aperture radar (SAR) automatic target recognition (ATR) is the key technique for remote-sensing image recognition. The state-of-the-art convolutional neural networks (CNNs) for SAR ATR suffer from high computation cost and large memory footprint, making them unsuitable to be deployed on resource-limited platforms, such as small/micro satellites. In this paper, we propose a comprehensive GNN-based model-architecture co-design on FPGA to address the above issues. Model design: we design a novel graph neural network (GNN) for SAR ATR. The proposed GNN model incorporates GraphSAGE layer operators and attention mechanism, achieving comparable accuracy as the state-of-the-art work with near 1/100 computation cost. Then, we propose a pruning approach including weight pruning and input pruning. While weight pruning through lasso regression reduces most parameters without accuracy drop, input pruning eliminates most input pixels with negligible accuracy drop. Architecture design: to fully unleash the computation parallelism within the proposed model, we develop a novel unified hardware architecture that can execute various computation kernels (feature aggregation, feature transformation, graph pooling). The proposed hardware design adopts the Scatter-Gather paradigm to efficiently handle the irregular computation patterns of various computation kernels. We deploy the proposed design on an embedded FPGA (AMD Xilinx ZCU104) and evaluate the performance using MSTAR dataset. Compared with the state-of-the-art CNNs, the proposed GNN achieves comparable accuracy with 1/3258 computation cost and 1/83 model size. Compared with the state-of-the-art CPU/GPU, our FPGA accelerator achieves 14.8×/2.5× speedup (latency) and is 62×/39× more energy efficient. Bingyi Zhang, Rajgopal Kannan, Viktor Prasanna 0001, Carl E. Busart |
FPL | 1 |
| 2022 | Hypersort: High-performance Parallel Sorting on HBM-enabled FPGAabstractAccelerating sorting on FPGA has been extensively studied by leveraging the fine-grained data parallelism of FPGAs. However, with the optimized hardware pipelines, the performance of sorting algorithms is bounded by the off-chip memory band-width. The integration of high-bandwidth memory (HBM) on FPGAs offers significantly more off-chip memory bandwidth compared with traditional DDR memory, which enables new opportunities for accelerating sorting. In this paper, we develop Hypersort, a hardware accelerator to accelerate sorting on HBM-enabled FPGA. We use columnsort to merge HBM channels. To support the data communication pat-terns of Columnsort, we propose several optimizations to reduce external memory (HBM) traffic and hide data communication latency to further improve the overall throughput. We implement our accelerator on a state-of-the-art HBM-enabled FPGA. Ex-perimental results show that our implementation achieves overall sorting throughput of 34 GB/s, which is up to 14.8×, 4.73× and 2.18 ×faster than the state-of-the-art implementations on CPU, FPGA with external DDR and HBM-enabled FPGA, respectively. The proposed approach demonstrates higher efficiency for merging sorted arrays in HBM channels compared with the state-of-the-art implementation on HBM-enabled FPGA. Soundarya Jayaraman, Bingyi Zhang, Viktor Prasanna 0001 |
FPT | 2 |
| 2022 | Low-latency Mini-batch GNN Inference on CPU-FPGA Heterogeneous PlatformabstractMini-batch inference of Graph Neural Networks (GNNs) is a key problem in many real-world applications. In this paper, we develop a computationally efficient mapping of GNNs onto CPU-FPGA heterogeneous platforms to achieve low-latency mini-batch inference. While the lightweight preprocessing algorithm of GNNs can be efficiently mapped onto the CPU platform, on the FPGA platform, we design a novel GNN hardware accelerator with an adaptive datapath denoted as Adaptive Computation Kernel (ACK) that can execute various computation kernels of GNNs with low-latency: (1) for dense computation kernels expressed as matrix multiplication, ACK works as a systolic array with fully localized connections, (2) for sparse computation kernels, ACK follows the scatter-gather paradigm and works as multiple parallel pipelines to support the irregular connectivity of graphs. The proposed task scheduling hides the CPU-FPGA data communication overhead to reduce the inference latency. We develop a fast design space exploration algorithm to generate a single accelerator for multiple target GNN models. We implement our accelerator on a state-of-the-art CPU-FPGA platform and evaluate the performance using three representative models (GCN, GraphSAGE, GAT). Results show that our CPU-FPGA implementation achieves 21.4−50.8×, 2.9 − 21.6×, 4.7× latency reduction compared with state-of-the-art implementations on CPU-only, CPU-GPU and CPU-FPGA platforms. Bingyi Zhang, Hanqing Zeng, Viktor Prasanna 0001 |
HIPC | 1 |
| 2022 | Model-Architecture Co-Design for High Performance Temporal GNN Inference on FPGAabstractTemporal Graph Neural Networks (TGNNs) are powerful models to capture temporal, structural, and contextual information on temporal graphs. The generated temporal node embeddings outperform other methods in many downstream tasks. Real-world applications require high performance inference on real-time streaming dynamic graphs. However, these models usually rely on complex attention mechanisms to capture relationships between temporal neighbors. In addition, maintaining vertex memory suffers from intrinsic temporal data dependency that hinders task-level parallelism, making it inefficient on general-purpose processors. In this work, we present a novel model-architecture co-design for inference in memory-based TGNNs on FPGAs. The key modeling optimizations we propose include a light-weight method to compute attention scores and a related temporal neighbor pruning strategy to further reduce computation and memory accesses. These are holistically coupled with key hardware optimizations that leverage FPGA hardware. We replace the temporal sampler with an on-chip FIFO based hardware sampler and the time encoder with a look-up-table. We train our simplified models using knowledge distillation to ensure similar accuracy vis-á-vis the original model. Taking advantage of the model optimizations, we propose a principled hardware architecture using batching, pipelining, and prefetching techniques to further improve the performance. We also propose a hardware mechanism to ensure the chronological vertex updating without sacrificing the computation parallelism. We evaluate the performance of the proposed hardware accelerator on three real-world datasets. The proposed model reduces the computation complexity by 84% and memory accesses by 67% with less than 0.33% accuracy loss. Compared with CPU/GPU, our FPGA accelerator achieves 16.4/2.3× speedup in latency and 0.27% improvement in accuracy compared with the state-of-the-art inference algorithm. To the best of our knowledge, this is the first work that performs model-architecture co-design on memory-based Temporal Graph Neural Networks. Bingyi Zhang, Rajgopal Kannan, Viktor Prasanna 0001, Carl E. Busart |
IPDPS | 2 |
| 2021 | BoostGCN: A Framework for Optimizing GCN Inference on FPGAabstractGraph convolutional networks (GCNs) have revolutionized many big data applications, such as recommendation systems, traffic prediction, etc. However, accelerating GCN inference is challenging due to (1) massive external memory traffic and irregular memory access, (2) workload imbalance due to skewed degree distribution, and (3) intra-stage load imbalance caused by two heterogeneous computation phases of the algorithm. To address the above challenges, we propose a framework named BoostGCN to optimize GCN inference on FPGA. First, we develop a novel hardware-aware Partition-Centric Feature Aggregation (PCFA) scheme that leverages 3-D partitioning with the vertex-centric computing paradigm. This increases on-chip data reuse and reduces the total data communication volume with external memory. Second, we design a novel hardware architecture to enable pipelined execution of the two heterogeneous computation phases. We develop a low-overhead task scheduling strategy to reduce the pipeline stalls caused by the two computation phases. Third, we provide a complete GCN acceleration framework on FPGA with optimized RTL templates. It can generate hardware designs based on the customized configuration and is adaptable to various GCN models. Using our framework, we generate accelerators for various GCN models on a state-of-the-art FPGA platform and evaluate our designs using widely used datasets. Experimental results show that the accelerators produced by our framework achieve significant speedup compared with state-of-the-art implementations on CPU (≈ 100×), GPU (≈ 30×), prior FPGA accelerator (3-45)×. Bingyi Zhang, Rajgopal Kannan, Viktor Prasanna 0001 |
FCCM | 1 |
| 2021 | A Framework for Optimizing GCN Inference on FPGAabstractGraph convolutional networks (GCNs) have revolutionized many big data applications. However, accelerating GCN inference is still challenging due to (1) massive external memory traffic and irregular memory access, (2) workload imbalance because of the skewed degree distribution, and (3) intra-stage load imbalance between feature aggregation and feature transformation steps. To address the above challenges, we propose a framework to optimize GCN inference on FPGA. First, we propose a novel Partition-Centric Feature Aggregation (PCFA) scheme to increase the data locality and reduce the number of random memory accesses in feature aggregation step. Second, we propose a novel hardware architecture to enable pipelined execution of the two heterogeneous computation steps. Then, a low-overhead task scheduling strategy is proposed to achieve stall-free execution of the two computation steps. Third, we provide a complete GCN acceleration framework on FPGA, and define key parameters for users to fine-tune the throughput. The model-specific operators can be customized to support a wide-range of GCN models. Using our framework, we design accelerators on a state-of-the-art FPGA. We evaluate our work using widely used datasets and. Experimental results show the accelerators produced by our framework achieve significant speedup compared with state-of-the-art implementations on CPU (≈100x), GPU (≈30x), and FPGA (4.5-32x). Bingyi Zhang, Rajgopal Kannan, Viktor Prasanna 0001 |
FPGA | 1 |
| 2021 | Performance of Local Push Algorithms for Personalized PageRank on Multi-core PlatformsabstractPersonalized PageRank (PPR) is used to measure the importance of vertices with respect to a source vertex. PPR is a key kernel used in many real-world applications, such as information retrieval, recommendations, knowledge discovery, etc. Local push algorithms have been widely used for developing state-of-the-art fast PPR algorithms. In this paper, we analyze the computational characteristics of local push algorithms at the algorithm (error tolerance, damping factor) and hardware architecture (available memory, cache features) levels. First, we profile the algorithm to understand the effect of various algorithm parameters. We study the trade-offs between latency and accuracy of local push algorithms using various performance metrics including Top-K accuracy and scalability. Then, we perform our analysis on two state-of-the-art multi-core platforms to understand the latency of the algorithm for PPR computation on a single source vertex and its scalability for multiple vertices using thread-level parallelism. We analyze the impact of error tolerance and damping factor on the overall performance of the algorithms. Madhav Aggarwal, Bingyi Zhang, Viktor Prasanna 0001 |
HiPC | 2 |
| 2021 | TEANS: A Target Enhancement and Attenuated Nonmaximum Suppression Object Detector for Remote Sensing ImagesabstractIn this letter, we propose an effective approach to learn a convolutional neural network (CNN) model with target enhancement and attenuated nonmaximum suppression (NMS) technique (TEANS) for object detection in optical remote sensing images. TEANS mainly consists of two steps. First, the target enhancement architecture, including target upsampling and reconvolution, is designed into a given deep ResNet-101 model for accurate object detection, especially for small ones. Second, the attenuated NMS technique is used for overcoming wrong eliminations of serried object proposals. For verifying the effectiveness of the TEANS method, evaluations are implemented on a publicly available 15-class optical remote sensing object detection data set. Experimental results show that TEANS can achieve 5.55%, 18.77%, 26.81%, 55.07%, 28.48%, 6.01%, and 5.51% improvements in mean Average Precision (mAP), respectively, compared with standard Faster R-CNN, R-FCN, YOLOv2, SSD, USB-BBR, YOLOv3, and MS-VANs frameworks. Haibao Chen, Guanghui He 0002, Bingyi Zhang, Hao Yu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2020 | Hardware Acceleration of Large Scale GCN InferenceabstractGraph Convolutional Networks (GCNs) have become state-of-the-art deep learning models for representation learning on graphs. Hardware acceleration of GCN inference is challenging due to: 1) massive size of the input graph, 2) heterogeneous workload of the GCN inference that consists of sparse and dense matrix operations, and 3) irregular information propagation along the edges during the computation. To address the above challenges, we propose the algorithm-architecture co-optimization to accelerate large-scale GCN inference on FPGA. We first perform data partitioning to fit each partition in the limited on-chip memory. Then, we use a two-phase preprocessing algorithm consisting of sparsification and node reordering. The first phase (sparsification) eliminates edge connections of high-degree nodes by merging common neighbor nodes. The second phase (re-ordering) effectively groups adjacent nodes to improve on-chip data reuse. Incorporating the above algorithmic optimizations, we propose a generic FPGA architecture to pipeline the two major computational kernels in GCN: aggregation and transformation. The flexible data path and task scheduling strategy of our design support various GCN models and lead to high throughput inference. We evaluate our design on state-of-the-art FPGA platform using three large scale datasets: Flickr, Reddit, Yelp. Compared with the state-of-the-art multi-core and GPU baselines, our design improves the throughput by up to $30 \times$ and $2 \times$ respectively. Bingyi Zhang, Hanqing Zeng, Viktor Prasanna 0001 |
ASAP | 1 |
| 2020 | Accelerating Large Scale GCN Inference on FPGAabstractWe propose an algorithm-architecture co-optimization framework to accelerate large-scale GCN inference on FPGA. We first perform data partitioning to fit each partition in the limited on-chip memory of FPGA. Then, we use the two-phase pre-processing algorithm consisting of sparsification and node reordering. The first phase (sparsification) eliminates edge connections of high-degree nodes by merging common neighbor nodes. The second phase (re-ordering) effectively groups densely connected neighborhoods to improve on-chip data reuse. Incorporating the above algorithmic optimizations, we propose an FPGA architecture to efficiently execute the two key computational kernels of GCN - feature aggregation and weight transformation. We evaluate our design on a state-of-the-art FPGA device. Compared with multi-core and GPU baselines, our design reduces the inference latency by up to $30 \times $ and $2 \times $ respectively. Bingyi Zhang, Hanqing Zeng, Viktor Prasanna 0001 |
FCCM | 1 |
| 2004 | Simulation of network traffic and its applicationabstractNetwork traffic model based on Alpha stable process is a hot point. Many models without proof have been proposed. A fractal Alpha model is proposed in this paper and two proofs based on flow and session level respectively are given. Proof and simulation show the model in this paper can describe the self similarity, spike and the long range dependent between increments. Using the mode in this paper, a formula for the residual of the queueing distribution function (RDF) is deduced. Comparing the RDF based on the model in this paper with that based on other models, it can be found that the formula in this paper meet real RDF better. Bingyi Zhang, Yamin Sun |
ICARCV | 1 |