Sergey Gribok

dblp:236/6718 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
11since 2021 · last 2024
0000-0003-3339-7705ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 4 first-author · 11 since 2021
YearPublicationVenuePosition
2024 if-ZKP: Intel FPGA-Based Acceleration of Zero Knowledge Proofs
abstract
Zero-Knowledge Proofs (ZKPs) allow proving a statement's correctness without revealing anything else, enabling privacy in applications like blockchains and digital voting. Recently, Zero-knowledge Succinct Non-interactive Arguments of Knowledge (zk-SNARKs) have helped address ZKPs' scalability challenges and gained significant attention. This paper presents a novel scalable FPGA architecture for accelerating the zk-SNARK prover's compute-intensive multi-scalar multiplication (MSM) operation. The architecture exploits MSM's inherent parallelism, using optimized IP for modular arithmetic. Implemented with Intel OneAPI for FPGAs, it achieves 110x-150x speedup over software for the BLS12-381 and BN128 elliptic curves, utilizing a generic Jacobian coordinate system.
Shahzad Ahmad Butt, Benjamin Reynolds, Veeraraghavan Ramamurthy, Pohrong Chu, Setareh Sharifian, Sergey Gribok, Bogdan Pasca 0001
FCCM7
2024 Efficient 8-bit Matrix Multiplication on Intel Agilex-5 FPGAs
abstract
Matrix multiplication is a fundamental operation in many fields including artificial intelligence and machine learning, and it often requires significant computational resources. FPGAs have always been a great platform for accelerating such calculations thanks to their inherent parallelism and flexibility regarding data movement. The newly released Agilex-5 FPGA devices introduce several AI-specific hardware features that complement the traditional DSP Block functionality. The fixed-point Tensor Mode of the DSP Block exposes twenty 8-bit signed multipliers organized into two 10-element dot products, having one set of inputs fed from internal DSP Block registers. In this paper, based on the philosophy of efficiently utilizing the low-level features of the new DSP Block, we make use of these structures and construct a flexible 8-bit matrix multiplication engine. The generic matrix-multiplication architecture presented here achieves at steady-state 100% compute resource utilization (no idle states). In the 810-DSP configuration, the engine achieves over 750MHz on an Agilex-5 (fastest speedgrade) device with a throughput of 24.75TOPs and an energy efficiency of 1.65TOPs/W.
Sergey Gribok, Bogdan Pasca 0001
FCCM1
2024 Stay Flexible: A High-Performance FPGA NPU Overlay for Graph Neural Networks
abstract
Graph neural networks (GNNs) are a class of deep learning (DL) models widely-used for learning latent representations of graph-structured data for a variety of node/graph-level prediction tasks. Real-time applications of GNNs are evolving in various domains such as 3D object detection from LiDAR point clouds in autonomous vehicles [1] and classifying collected data in particle physics colliders [2]. Typically, these use cases have stringent latency constraints but can still benefit from batch processing of multiple graphs from different input sources. Existing accelerators either rely on preprocessing input graphs [3], [4] or are extremely specialized streaming pipelines which are unable to support dynamically changing workloads for these applications [5]. In this work, we take a different approach by enhancing the neural processing unit (NPU) [6] to accelerate a wide variety of GNN models without sacrificing its flexibility, performance or ability to run any of its originally supported DL workloads (e.g. MLPs, RNNs, GRUs, LSTMs).
Taikun Zhang, Andrew Boutros, Sergey Gribok, Kwadwo Boateng, Vaughn Betz
FCCM3
2024 FPGA Modular Multipliers using Hybrid Reduction Techniques
abstract
Modular multiplication is a key kernel in many computing fields. What makes this function so challenging are the very large word sizes – sometimes in the thousands of bits – that are typically required for the target applications. In this paper we propose a modular multiplication implementation based on a multi-stage hybrid reduction technique. Our proposed approach uses a parameterized number of multiplier-based reduction stages followed by a memory-based reduction. This construction allows for the multiplier-based stages to take advantage of Karatsuba multiplication, resulting in a reduced number of DSP Blocks. Our method also allows specifying the number of multiplier-based stages which adjusts the ratio of multipliers to memory blocks. The resource utilization of the proposed architecture outperforms the existing state-of-the-art modular multiplication designs while offering a user-defined way of distributing resources between memory and DSP Blocks.
Sergey Gribok, Martin Langhammer, Bogdan Pasca 0001
FPL1
2024 A Software-Programmable Neural Processing Unit for Graph Neural Network Inference on FPGAs
abstract
Graph neural networks (GNNs) are a widely-used class of deep learning (DL) models for learning latent representations of graph-structured data for a variety of node/graph-level prediction tasks, some of which require real-time low latency inference. Most existing GNN accelerators rely on preprocessing input graphs on a host/embedded CPU to parallelize computations on different sub-graphs, making them unsuitable for real-time use cases. Others are extremely specialized streaming pipelines for only a specific type of GNN and therefore suffer from long FPGA bitstream compile times when the model is updated and cannot be used in applications that combine GNNs with other classes of DL models. In this work, we enhance the neural processing unit (NPU) FPGA overlay architecture, instruction set, and software stack to support a variety of GNN models. We achieve this without sacrificing the NPU flexibility; our enhanced NPU can be programmed purely through software to accelerate different GNNs or any of its originally supported DL workloads (e.g. MLPs, RNNs, GRUs, LSTMs). In addition, this flexibility enables our NPU software compiler to generate GNN kernels with different performance targets (throughput-optimized vs. latency-optimized) by exploiting different dimensions of compute parallelism on the same overlay architecture. Besides the flexibility benefits, our NPU implemented on an Intel Stratix 10 NX (14 nm) FPGA can process $7.8 \times$ more graphs per second at a similar latency on average compared to a state-of-the-art model-specific FPGA accelerator targeting real-time applications on an AMD Ultrascale+ same-generation FPGA. It also achieves 5.8 $\times$ higher throughput compared to an Nvidia RTX A6000 GPU (8 nm) and $2.6 \times$ lower latency than a state-of-the-art accelerator that combines CPU-based graph preprocessing with AMD Versal (7 nm) fabric and AI engine compute. Finally, we present a case study for using our enhanced NPU in real-time GNNbased multi-input multi-output (MIMO) antenna scheduling, highlighting that it meets the latency requirements of this task in 5G communication networks.
Taikun Zhang, Andrew Boutros, Sergey Gribok, Kwadwo Boateng, Vaughn Betz
FPL3
2024 CSAIL2019 Crypto-Puzzle Solver Architecture
abstract
tThe CSAIL2019 time-lock puzzle is an unsolved cryptographic challenge introduced by Ron Rivest in 2019, replacing the solved LCS35 puzzle. Solving these types of puzzles requires large amounts of intrinsically sequential computations, with each iteration performing a very large (3,072-bit for CSAIL2019) modular multiplication operation. The complexity of each iteration is several times greater than known field-programmable gate array (FPGA) implementations, and the number of iterations has been increased by about 1,000x compared with LCS35. Because of the high complexity of this new puzzle, a number of intermediate, or milestone, versions of the puzzle have been specified. In this article, we present several FPGA architectures for the CSAIL2019 solver, which we implement on a medium-sized Intel Agilex device. We develop a new multi-cycle modular multiplication method, which is flexible and can fit on a wide variety of sizes of current FPGAs. We introduce a class of multi-cycle squarer-based architectures that allow for better resource and area trade-offs. We also demonstrate a new approach for improving the fitting and timing closure of large, chip-filling arithmetic designs. We used the solver to compute the first 23 out of 28 milestone solutions of the puzzle, which are the first reported results for this problem.
Sergey Gribok, Bogdan Pasca 0001, Martin Langhammer
ACM Trans. Reconfigurable Technol. Syst.1
2023 CSAIL2019 Crypto-Puzzle Solver Architecture
abstract
The CSAIL2019 time-lock puzzle is an unsolved cryptographic challenge introduced by Ron Rivest in 2019, replacing the solved LCS35 puzzle. Solving these types of puzzles requires large amounts of intrinsically sequential computations (i.e. computations which cannot be parallelized), with each iteration performing a very large (3072-bit in the case of CSAIL2019) modular multiplication operation. The complexity of each iteration is several times greater than known FPGA implementations, and the number of iterations has been increased by about 1000x compared to LCS35. Because of the high complexity of this new puzzle, a number of intermediate, or milestone versions of the puzzle have been specified.
Sergey Gribok, Bogdan Pasca 0001, Martin Langhammer
FPGA1
2022 Low-Latency Modular Exponentiation for FPGAs
abstract
Modular exponentiation, especially for very large integers of hundreds or thousands of bits, is a commonly used function in popular cryptosystems such as RSA. The complexity of this algorithm is partly driven by the very large word sizes, which require many - often millions - of primitive operations in a CPU implementation, or a large amount of logic when accelerated by an ASIC. FPGAs, with their many embedded DSP resources have started to be used as well. In almost all cases, the calculations have required multiple - occasionally many - clock cycles to complete. Recently, blockchain algorithms have required very low-latency implementations of modular multiplications, motivating new implementations and approaches.In this paper we show nine different high performance modular exponentiation for 1024-bit operands, using a 1024-bit modular multiplication as it’s core. Rather than just showing a number of completed designs, our paper shows the evolution of architectures which lead to different resource mix options. This will allow the reader to apply the examples to different FPGA targets which may have differing ratios of logic, memory, and embedded DSP blocks. In one design, we show a 1024b modular multiplier requiring 83K ALMs and 2372 DSPs, with a delay of 21.21ns.
Martin Langhammer, Sergey Gribok, Bogdan Pasca 0001
FCCM2
2022 Stratix 10 NX Architecture
abstract
The advent of AI has driven the exploration of high-density low-precision arithmetic on FPGAs. This has resulted in new methods in mapping both arithmetic functions as well as dataflows onto the fabric, as well as some changes to the embedded DSP Blocks. Technologies outside of the FPGA realm have also evolved, such as the addition of tensor structures for GPUs, as well as the introduction of numerous AI ASSPs, all of which have a higher claimed performance and efficiency than current FPGAs. In this article, we will introduce the Stratix 10 NX device, which is a variant of FPGA specifically optimized for the AI application space. In addition to the computational capabilities of the standard programmable soft-logic fabric, a new type of DSP Block provides the dense arrays of low-precision multipliers typically used in AI implementations. The architecture of the block is tuned for the common matrix-matrix or vector-matrix multiplications in AI, with capabilities designed to work efficiently for both small and large matrix sizes. The base precisions are INT8 and INT4, along with shared exponent support to support block FP16 and block FP12 numerics. All additions/accumulations can be done in INT32 or IEEE-754 single precision floating point (FP32), and multiple blocks can be cascaded together to support larger matrices. We will also describe methods by which the smaller precision multipliers can be aggregated to create larger multipliers that are more applicable to standard signal processing requirements. In the AI market, the FPGA must compete directly with other types of devices, rather than occupy a unique niche. Deterministic system performance is as important as the performance of individual FPGA elements, such as logic, memory, and DSP. We will show that the feed forward datapath structures that are needed to support the typical AI matrix-vector and matrix-matrix multiplication operations can consistently close timing at over 500 MHz on a mid-speed grade device, even if all of the Tensor Blocks on the device are used. We will also show a full-chip NPU processor implementation that out performs GPUs at the same process node for a variety of AI inferencing workloads, even though it has a lower operating frequency of 365 MHz. In terms of overall compute throughput, Stratix 10 NX is specified at 143 INT8/FP16 TOPs/FLOPs or 286 INT4/FP12 TOPS/FLOPs. Depending on the configuration, power efficiency is in the range of 1–4 TOPs or TFLOPs/W.
Martin Langhammer, Eriko Nurvitadhi, Sergey Gribok, Bogdan Pasca 0001
ACM Trans. Reconfigurable Technol. Syst.3
2021 Stratix 10 NX Architecture and Applications
abstract
The advent of AI has driven the adoption of high density low precision arithmetic on FPGAs. This has resulted in new methods in mapping both arithmetic functions as well as dataflows onto the fabric, as well as some changes to the embedded DSP Blocks. Technologies outside of the FPGA realm have also evolved, such as the addition of tensor structures for GPUs, and also the introduction of numerous AI ASSPs, all of which have a higher claimed performance and efficiency than current FPGAs. In this paper we will introduce the Stratix 10 NX device (NX), which is a variant of FPGA specifically optimized for the AI application space. In addition to the computational capabilities of the standard programmable soft logic fabric, a new type of DSP Block provides the dense arrays of low precision multipliers typically used in AI implementations. The architecture of the block is tuned for the common matrix-matrix or vector-matrix multiplications in AI, with capabilities designed to work efficiently for both small and large matrix sizes. The base precisions are INT8 and INT4, along with shared exponent support for support block floating point FP16 and FP12 numerics. All additions/accumulations can be done in INT32 or IEEE754 single precision floating point (FP32), and multiple blocks can be cascaded together to support larger matrices. We will also describe methods by which the smaller precision multipliers can be aggregated to create larger multiplier that are more applicable to standard signal processing requirements. In terms of overall compute throughput, Stratix 10 NX achieves 143 INT8/FP16 TOPs/FLOPs, or 286 INT4/FP12 TOPS/FLOPs at 600MHz. Depending on the configuration, power efficiency is in the range of 1-4 TOPs or TFLOPs/W.
Martin Langhammer, Eriko Nurvitadhi, Bogdan Pasca 0001, Sergey Gribok
FPGA4
2021 Dense FPGA Compute Using Signed Byte Tuples
abstract
The importance of AI to FPGA has resulted in ever increasing low precision hard arithmetic features in newer devices. Many FPGAs, including those from Achronix, Intel, and Xilinx, have significantly increased the density of INT8 and INT9 embedded multipliers. Mainstream devices with these enhanced densities still support the traditional intermediate integer (typically 18-bit) multipliers, with IEEE-754 floating-point now becoming more prevalent as well.Recently, Intel introduced the Stratix 10 NX FPGA, which is targeted specifically at AI acceleration. This device contains a new type of AI-specific DSP Block with approximately an order of magnitude higher INT8 density than previous FPGA industry DSP Blocks. Larger standard FPGA integer precisions, however, are not directly supported. Intel has described some methods of aggregating larger multipliers from the NX Blocks, but these are somewhat smaller than typically used by DSP applications. Larger multiplications can also be useful for other AI applications, such as found in training. In this paper, we introduce the concept of signed tuples, which can be used to assemble signed multipliers into more useful larger precision multipliers by leveraging FPGA soft-logic inexpensively. We demonstrate several constructions of INT16 multipliers, with some modes requiring less than 3 ALMs per INT16 multiplier when implemented in a tensor format. We also describe the application of these methods to even larger multipliers and alternate constructs such as complex multiplication. We show that there is essentially no performance degradation or system fitting impact from our method. The mid-size NX device can support up 33 TOPs INT16 (from 29,700 constructed INT16 multipliers on a mid-speed grade device) with this approach, which is higher than any other current or announced monolithic die FPGA. Our methods are not limited to FPGA, or any particular starting precision, and so may be used for other aggregations as well.
Martin Langhammer, Simon Finn, Sergey Gribok, Bogdan Pasca 0001
FPL3
2020 High Density 8-Bit Multiplier Systolic Arrays For Fpga
abstract
Artificial Intelligence (AI) has become the fastest growing application area for FPGAs. Two types of numerics are needed. Training typically uses floating point arithmetic (which is now widely available as embedded functions in current FPGAs). Inference is typically calculated with lower precision integer numbers, which can be implemented with embedded functions, soft logic, or a combination of the two. INT8 performance is therefore used as a typical benchmarking metric for current FPGAs. Recent publications based on Xilinx devices show the extraction of two INT8 multipliers from a 24×18 multiplier. A paper from Intel describes how to obtain two INT8 multipliers from a 18×18 multiplier, with the help of a small amount of soft logic. In this paper we introduce a number of new INT8 multiplier techniques, starting with the Intel 18×18 multiplier approach. Using both memory and logic resources - for a more balanced use of the FPGA features - we improve the INT8 density, and also show a signed-magnitude (SM) 1.7 construct that is even smaller. To demonstrate the usability of these new multipliers, we develop a scalable systolic array, that contains up to 32,768 SM1.7 multipliers, or 28,800 INT8 multipliers, fit in an Intel Stratix 10 2800 device. Finally, we implement a system architecture that includes input and output flow buffering and control, which can be instantiated directly into a larger AI design, or can enable the FPGA to be used as a standalone accelerator. This system exceeds 400 MHz for the largest array on a mid-speed device (26 TOPS INT8), and can operate up to 600 MHz for smaller array sizes.
Martin Langhammer, Sergey Gribok, Gregg Baeckler
FCCM2
2020 High Density Pipelined 8bit Multiplier Systolic Arrays for FPGA
abstract
With the advent of AI and machine learning as the highest profile FPGA applications, INT8 performance is currently one of the key benchmarking metrics. In current devices, INT8 multipliers must be extracted from higher precision multipliers. Recently, we reported the implementation of a mixed DSP Block and soft logic design, with 22,400 INT8 multipliers, and a system clock rate of 416MHz, on the Intel Stratix 10 2800 chip.
Martin Langhammer, Sergey Gribok, Gregg Baeckler
FPGA2
2019 Why Compete When You Can Work Together: FPGA-ASIC Integration for Persistent RNNs
abstract
Interactive intelligent services, such as smart web search, are important datacenter workloads. They rely on dataintensive deep learning (DL) algorithms with strict latency constraints and thus require balancing both data movement and compute capabilities. As such, a persistent approach that keeps the entire DL model on-chip is becoming the new norm for realtime services to avoid the expensive off-chip memory accesses. This approach is adopted in Microsoft's Brainwave and is also provided by Nvidia's cuDNN libraries. This paper presents a comparative study of FPGA, GPU, and FPGA+ASIC in-package solutions for persistent DL. Unlike prior work, we offer a fair and direct comparison targeting common numerical precisions (FP32, INT8) and modern high-end FPGA (Intel® Stratix®10), GPU (Nvidia Volta), and ASIC (10 nm process), all using the persistent approach. We show that Stratix 10 FPGAs offer 2.7× (FP32) to 8.6× (INT8) lower latency than Volta GPUs across RNN, GRU, and LSTM workloads from DeepBench. The GPU can only utilize ~6% of its peak TOPS, while the FPGA with a more balanced on-chip memory and compute can achieve much higher utilization (~57%). We also study integrating an ASIC chiplet, TensorRAM, with an FPGA as system-in-package to enhance on-chip memory capacity and bandwidth, and provide compute throughput matching the required bandwidth. We show that a small 32 mm2 TensorRAM 10nm chiplet can offer 64 MB memory, 32 TB/s on-chiplet bandwidth, and 64 TOPS (INT8). A small Stratix 10 FPGA with a TensorRAM (INT8) offers 15.9× better latency than GPU (FP32) and 34× higher energy efficiency. It has 2× aggregate on-chip memory capacity compared to a large FPGA or GPU. Overall, our study shows that the FPGA is better than the GPU for persistent DL, and when integrated with an ASIC chiplet, it can offer a more compelling solution.
Eriko Nurvitadhi, Dongup Kwon, Andrew Boutros, Jaewoong Sim, Phillip Tomson, Huseyin Ekin Sumbul, Gregory K. Chen, Phil C. Knag, Raghavan Kumar, Ram Krishnamurthy 0001, Sergey Gribok, Bogdan Pasca 0001, Martin Langhammer, Debbie Marr, Aravind Dasu
FCCM12
2019 Fractal Synthesis: Invited Tutorial
abstract
This paper will describe Fractal Synthesis, which is a new set of synthesis, clustering, and packing algorithms for FPGA devices, which dramatically increases the utilization and effective performance for arithmetic rich designs. The emergence of AI inferencing as a significant new FPGA application has brought some of the shortcomings of the FPGA and current design flows into focus. We describe new results where near 100% logic utilization of the FPGA is not only possible, but deterministic, with consistent high clock rates. Alternately, smaller datapaths can be synthesized, and combined to make chip filling designs. In one benchmark consisting of purely arithmetic datapath for a large Stratix®10 FPGA (E-2 speedgrade), we will show 92% logic utilization at 460 MHz for an automatically placed arithmetic datapath, and 410MHz with 97% logic utilization. Furthermore, we describe new results, where these performance and density level can be applied to non-arithmetic designs, by extending these techniques to placement.
Martin Langhammer, Gregg Baeckler, Sergey Gribok
FPGA3
2019 Evaluating and Enhancing Intel® Stratix® 10 FPGAs for Persistent Real-Time AI
abstract
Interactive intelligent services (e.g., smart web search) are becoming essential datacenter workloads. They rely on data-intensive artificial intelligence (AI) algorithms that do not use batch computation due to their tight latency constraints. Since off-chip data accesses have higher latency and energy consumption than on-chip accesses, a persistent AI approach with the entire model stored in on-chip memory is becoming the new norm for real-time AI. This approach is the cornerstone of Microsoft's Brainwave FPGA-based AI cloud and was recently added to Nvidia's cuDNN library. In this work, we implement, optimize and evaluate a Brainwave-like neural processing unit (NPU) on a large Stratix-10 FPGA. We benchmark it against a large Nvidia Volta GPU running cuDNN persistent AI kernels. Across real-time persistent RNN, GRU, and LSTM workloads, we show that Stratix-10 offers ~3× (FP32) and ~10× (INT8) better latency than GPU (FP32), which uses only ~6% of its peak throughput. Then, we propose TensorRAM, an ASIC chiplet for persistent AI that is 2.5D integrated with an FPGA in the same package. TensorRAM enhances the on-chip memory capacity and bandwidth, with enough multi-precision INT8/4/2/1 throughput to match that bandwidth. Multiple TensorRAMs can be integrated with Stratix-10. Our evaluation shows that a small 32-mm2 TensorRAM on 10nm offers 64MB of SRAMs with 32TB/s on-chiplet bandwidth and 64 TOP/s (INT8). A small Stratix-10 with a TensorRAM (INT8) offers 16× better latency and 34× energy efficiency compared to GPU (FP32). Overall, Stratix-10 with TensorRAM offers compelling and scalable persistent AI solutions.
Eriko Nurvitadhi, Dongup Kwon, Andrew Boutros, Jaewoong Sim, Phillip Tomson, Huseyin Ekin Sumbul, Gregory K. Chen, Phil C. Knag, Raghavan Kumar, Ram Krishnamurthy 0001, Debbie Marr, Sergey Gribok, Bogdan Pasca 0001, Martin Langhammer, Aravind Dasu
FPGA13
2019 Extracting INT8 Multipliers from INT18 Multipliers
abstract
With the advent of machine learning as perhaps the most high-profile application area for FPGAs, there is a compelling reason to improve the provision of smaller precision arithmetic on these devices. INT8 is commonly used for AI inferencing, and along with some additional soft logic for exponent handling, can be an effective solution for training as well. This paper describes techniques for efficiently extracting INT8 multipliers from commonly available INT18 multipliers found in many modern FPGAs. A small amount of soft logic - as little as 7 ALMs per INT8 multiplier - is required to provide pre or post multiplier correction to calculate two INT8 multiplies from a single 18x18 multiplier. We present two configurations for both signed and unsigned representations where two multiplications share one input operand. In addition to the individual INT8 variants, we present full device cases of 22,400 INT8 multipliers organized as DOT32 product arrays, with the soft logic tightly bound to the INT18 based DSP Blocks. A majority of the soft logic and routing in the device is left untouched, and available for application development.
Martin Langhammer, Bogdan Pasca 0001, Gregg Baeckler, Sergey Gribok
FPL4