EDBT 2026 Demo / reviewers in the wild / expert
Yang Yang 0111
dblp:48/450-111
· DBLP profile ↗
10ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-7718-1925ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 7 first-author · 8 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | High Throughput Matrix Transposition on HBM-Enabled FPGAsabstractMatrix transposition is a classic operation in machine learning and scientific applications. HBM-enabled FPGAs, with their high-bandwidth capabilities, are increasingly deployed in the data centers. However, achieving high bandwidth utilization on HBM is challenging due to the large strided access patterns in matrix transposition, which significantly degrade bandwidth utilization. Additionally, saturating HBM bandwidth requires a large number of parallel accesses, further complicating the design. In this paper, we present a high throughput matrix transposition design for HBM-enabled FPGAs. Our design performs small strided accesses to HBM, ensuring optimized bandwidth utilization. We use on-chip SRAMs to reorganize data from HBM into the access pattern needed for matrix transposition. Inspired by Latin Squares, we propose a novel data layout for storing the matrix tiles in SRAM. This data layout is paired with a customized scheduling strategy to eliminate SRAM bank conflicts. We develop a fully pipelined architecture with multiple Processing Elements (PEs) to enable parallel HBM accesses. Our design is highly scalable, supporting various configurations of HBM channels and arbitrary matrix dimensions. We implement the proposed design on the AMD Alveo U280 FPGA. Experimental results show that our design achieves a matrix transposition throughput of up to 415 GB/s, more than 90% of the peak HBM bandwidth of the target FPGA platform. Our design outperforms state-of-the-art GPU implementations, delivering up to 1.44 × higher HBM memory bandwidth utilization. Yang Yang 0111, Rajgopal Kannan, Viktor Prasanna 0001 |
FCCM | 1 |
| 2025 | OLA: An FPGA-based Overlay Accelerator for Privacy Preserving Machine Learning with Homomorphic EncryptionabstractHomomorphic Encryption (HE) is a promising technique for protecting user privacy in cloud-based Machine Learning (ML) inference. However, homomorphically encrypted operations are orders of magnitude slower than the corresponding unencrypted operations due to high computation and memory bandwidth requirements. FPGAs are attractive platforms for designing domain-specific accelerators. Manually programming FPGAs for HE ML inference is challenging, as different HE parameters and operations require different mapping strategies. We propose OLA, the first FPGA-based overlay accelerator for low latency computation on homomorphically encrypted data. OLA eliminates the need for manual and time consuming FPGA programming by providing a Python-based interface and a compiler that efficiently maps OLA programs to hardware instructions for execution. The hardware architecture and instruction set are co-designed to accelerate common HE primitives, while the compiler maps all the required HE operations to these primitives for efficient execution. OLA features two optimizations to address memory bandwidth challenges in HE computation. First, we propose an asynchronous dataflow execution model where the compiler manages data processing order and the architecture enforces it at run time. This approach enables guaranteed data reuse via on-chip SRAMs. Second, we design latency-aware instruction scheduling in the compiler to reduce data reuse distance. We implement the overlay accelerator on the AMD Alveo U280 FPGA. We evaluate the effectiveness of the proposed overlay accelerator by executing various HE operations, HE linear algebra benchmarks, and end-to-end ML inference. Experimental results show that OLA reduces the latency of HE operations by up to 683× (4.4×) compared with State-Of-The-Art (SOTA) CPU (GPU) implementations. It achieves up to 3212× speedup in latency compared with the SOTA CPU implementations for HE ML inference. Yang Yang 0111, Rajgopal Kannan, Viktor Prasanna 0001 |
FPGA | 1 |
| 2024 | A Framework for Generating Accelerators for Homomorphic Encryption Operations on FPGAsabstractHomomorphic Encryption (HE) is a promising technique for preserving user data privacy in cloud computing. Nevertheless, HE operations are magnitudes slower than un-encrypted computations due to their high computational complexity. FPGAs are attractive platforms for designing domain-specific accelerators. However, manually programming FPGAs for HE applications is nontrivial because of the vastly different parameter settings and latency requirements. To close the gap, we propose a framework to generate low latency FPGA accelerators for all the operations supported by HE, enabling users to utilize FPGA-accelerated HE processing without requiring knowledge of FPGA implementation details. The framework takes HE parameters and hardware resource constraints as input, uses design space exploration to automatically determine the design parameters that minimize HE computation latency, and produce synthesizable Verilog code. We propose a layered approach that decomposes HE operations into basic HE primitives, coupled with a parameterized HE domain-specific architecture that can efficiently execute the HE primitives. This approach avoids allocating dedicated FPGA resources to different subroutines within HE operations and improves compute utilization. Our evaluation shows that the generated accelerators significantly reduce latency in various HE operations, achieving up to$215\times$improvement over state-of-the-art CPU implementations. We demonstrate our framework's capability to compose end-to-end HE applications using HE CNN inference. Our designs outperform state-of-the-art CPU designs in latency by up to$60\times$. Yang Yang 0111, Rajgopal Kannan, Viktor Prasanna 0001 |
ASAP | 1 |
| 2024 | Bandwidth Efficient Homomorphic Encrypted Discrete Fourier Transform Acceleration on FPGAabstractFully Homomorphic Encryption (FHE) plays an important role in privacy-preserving computation on the cloud. It allows computations on encrypted data without decryption. Bootstrapping is a fundamental operation in FHE, enabling an unlimited number of homomorphic encrypted computations, but at a significant time cost. A major bootstrapping component, the Homomorphic Encrypted Discrete Fourier Transform (HE DFT), is particularly time-consuming and requires the transfer of a large amount of data from external memory. In this paper, we propose a bandwidth-efficient FPGA implementation of HE DFT. We design a cost model to evaluate the on-chip memory requirement and the off-chip data transfer overhead for HE DFT. Our analysis shows that prior approaches can lead to significant off-chip data transfers, which process the entire ciphertext between subroutines. To address DRAM transfer overhead, we propose LimbFlow, an optimized dataflow approach for HE DFT that enhances fine-grained data reuse by rearranging the processing order of ciphertext and merging several subroutines. Leveraging the LimbFlow, we develop an FPGA-based accelerator tailored for HE DFT. We evaluate the accelerator on AMD U280 FPGA across various sets of security parameters. Our accelerator achieves up to 4.90 × and 1.98 × speedup compared with the State-Of-The-Art (SOTA) GPU and FPGA implementations. Zhihan Xu, Yang Yang 0111, Rajgopal Kannan, Viktor Prasanna 0001 |
FCCM | 2 |
| 2023 | FPGA Acceleration of Rotation in Homomorphic Encryption Using Dynamic Data LayoutabstractHomomorphic Encryption (HE) is a promising technique to guarantee the security and privacy of Machine Learning (ML) applications in the cloud. Rotation is a key operation in HE ML; however, the high computational complexity and memory bandwidth requirements severely limit its performance. This work proposes a low-latency HE rotation accelerator targeting HBM-enabled FPGAs. First, we identify memory inefficiencies due to the access patterns of various sub-routines in rotation. We propose a dynamic data layout technique that converts large stride memory accesses to unit stride accesses to improve the bandwidth utilization. We leverage this technique to develop an FPGA accelerator that supports rotation for various HE parameter settings. The accelerator utilizes an optimized dataflow and an architecture specially designed to perform the dynamic data layout. We evaluate the accelerator using AMD U280 FPGA. Our design achieves up to 2.1 x speedup compared with two commonly used static layout approaches and up to 1.47x speedup compared with state-of-the-art GPU implementation across various rotation benchmarks. Yang Yang 0111, Weihang Long, Rajgopal Kannan, Viktor Prasanna 0001 |
FPL | 1 |
| 2022 | NTTGen: a framework for generating low latency NTT implementations on FPGAabstractHomomorphic encryption (HE) is a promising technique to ensure the security and privacy of applications in the cloud. Number Theoretic Transform (NTT) is a key operation in HE-based applications. HE requires vastly different NTT parameters to meet the performance and security requirements of applications. The increasing compute capabilities and flexibility of FPGAs make them attractive to accelerate NTT. However, programming FPGA still involves hardware design expertise and significant development effort. To close the gap, we propose NTTGen, a framework to automatically generate low latency NTT designs targeting HE-based applications. NTTGen takes application parameters, latency and hardware resource constraints as input, determines the design parameters, and produces synthesizable Verilog code as output. Low latency NTT implementations are obtained by varying the data, pipeline and batch parallelism. NTTGen utilizes streaming permutation network to reduce the interconnect complexity between stages in the NTT computation. The framework supports two types of NTT cores to perform modular arithmetic, the key computation in NTT: a low latency and resource efficient NTT core for a specific class of prime moduli and a general purpose NTT core for other primes. We further develop a design space exploration flow to identify the hardware design parameters of an optimal design. We evaluate NTTGen by generating designs for various NTT parameters. The designs result in up to 2.9X improvement in latency over the state-of-the-art FPGA implementations. Yang Yang 0111, Sanmukh R. Kuppannagari, Rajgopal Kannan, Viktor Prasanna 0001 |
CF | 1 |
| 2022 | FPGA Accelerator for Homomorphic Encrypted Sparse Convolutional Neural Network InferenceabstractHomomorphic Encryption (HE) is a promising solution to the increasing concerns of privacy in machine learning. But HE-based CNN inference remains impractically slow. Pruning can significantly reduce the compute and memory footprint of CNNs. However, homomorphic encrypted Sparse Convolutional Neural Networks (SCNN) have vastly different compute and memory characteristics compared with unencrypted SCNN. Simply extending the design principles of existing SCNN accelerators may offset the potential acceleration offered by sparsity. To realize fast execution, we propose an FPGA accelerator to speedup the computation of linear layers, the main computational bottleneck in HE SCNN batch inference. First, we analyze the memory requirements of various linear layers in HE SCNN and discuss the unique challenges. Motivated by the analysis, we present a novel dataflow specially designed to optimize HE SCNN data reuse coupled with an efficient scheduling policy that minimizes on-chip SRAM access conflicts. Leveraging the proposed dataflow and scheduling algorithm, we demonstrate the first end-to-end acceleration of HE SCNN batch inference targeting CPU-FPGA heterogeneous platforms. For a batch of 8K images, our design achieves up to 5.6× speedup in inference latency compared with the CPU-only solution for widely studied 6-layer and 11-layer HE CNNs. Yang Yang 0111, Sanmukh R. Kuppannagari, Rajgopal Kannan, Viktor Prasanna 0001 |
FCCM | 1 |
| 2022 | Bandwidth Efficient Homomorphic Encrypted Matrix Vector Multiplication Accelerator on FPGAabstractHomomorphic Encryption (HE) is a promising solution to the increasing concerns of privacy in Machine Learning (ML) as it enables computations directly on encrypted data. However, it imposes significant overhead on the compute system and remains impractically slow. Prior works have proposed efficient FPGA implementations of basic HE primitives such as number theoretic transform (NTT), key switching, etc. Composing the primitives together to realize higher level ML computation is still a challenge due to the large data transfer overhead. In this work, we propose an efficient FPGA implementation of HE Matrix Vector Multiplication$(\mathbf{M}\times \mathbf{V})$, a key kernel in HE-based Machine Learning applications. By analyzing the data reuse characteristics and the encryption overhead of HE$\mathbf{M}\times \mathbf{V}$, we show that simply using the principles of unencrypted$\mathbf{M}\times \mathbf{V}$to design accelerators for HE$\mathbf{M}\times \mathbf{V}$can lead to a significant amount of DRAM data transfers. We tackle the computation and data transfer challenges by proposing a bandwidth efficient dataflow that is specially optimized for HE$\mathbf{M}\times \mathbf{V}$. We identify highly reused data entities in HE$\mathbf{M}\times \mathbf{V}$and efficiently utilize the on-chip SRAM to reduce the DRAM data transfers. To speed up the computation of HE$\mathbf{M}\times \mathbf{V}$, we exploit three types of parallelism: partial sum parallelism, residual polynomial parallelism and coefficient parallelism. Leveraging these innovations, we demonstrate the first FPGA accelerator for HE matrix vector multiplication. Evaluation on 7 HE$\mathbf{M}\times \mathbf{V}$benchmarks shows that our FPGA accelerator is up to$3.8\times$(GeoMean$2.8\times$) faster compared to the 64-thread CPU implementation. Yang Yang 0111, Sanmukh R. Kuppannagari, Rajgopal Kannan, Viktor Prasanna 0001 |
FPT | 1 |
| 2013 | Shared memory heterogeneous computation on PCIe-supported platformsabstractDomain-disparity between CPU and Hardware Accelerators(HA) leads to CPU under-utilization and inter-domain data copy overheads. By exposing HA memory to OS and host MMU, these overheads can be eliminated. In this paper, we present a shared virtual memory real system design for PCIe-based HAs to enable parallel heterogeneous execution in CPU and HAs without driver overheads. We extend Linux with a custom memory manager and scheduler to manage HA memory and application-cores respectively. Our FPGA-based multi-application logic design supports simultaneous execution of multiple heterogeneous applications. We show the advantages of heterogeneous execution and analyze how our design reduces OS overhead. Sambit Kumar Shukla, Yang Yang 0111, Laxmi N. Bhuyan, Philip Brisk |
FPL | 2 |
| 2012 | An efficient dynamic multiple-candidate motion vector approach for GPU-based hierarchical motion estimationabstractHierarchical or pyramid search is a widely used approach in motion estimation, a most expensive function in video encoding, for its low computational complexity and high efficiency. In this approach, multiple down-sampled resolutions from video frames are created. An initial motion estimation is quickly made at a lowest resolution. The final motion estimation result is achieved by propagating the initial estimation towards the original resolution. GPU or General purpose GPU embedded hundreds of number of SIMD-based cores is best suitable for motion estimation, especially with full-search-based approaches as the process can be efficiently parallelized. However, a common fundamental drawback of the hierarchical search is the erroneous estimation from the reduced resolutions may cause the final motion estimation inaccurate. Multiple-candidate motion vector approaches are proposed, however, they lack a mechanism to select the best multiple-candidate schemes considering diverse video encoding characteristics. In this paper we analyse and verify the computational complexity of the hierarchical search using NVIDIA's GPU with realistic workloads. Based on this analysis, we propose an efficient dynamic multiple-candidate motion vector approach to dynamically select best multiple-candidate motion vector schemes at runtime. This approach can achieve highest possible speedups and satisfy a desire motion estimation efficiency. Experiments on realistic workloads show the dynamic scheme selection outperforms the fixed scheme selection based on profiling. Dung Vu, Yang Yang 0111, Laxmi N. Bhuyan |
IPCCC | 2 |