EDBT 2026 Demo / reviewers in the wild / expert
Rishov Sarkar
dblp:312/4942
· DBLP profile ↗
11ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0002-9168-0392ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FIFOAdvisor: A DSE Framework for Automated FIFO Sizing of High-Level Synthesis DesignsabstractDataflow hardware designs are important for efficient algorithm implementations on FPGAs across various domains using high-level synthesis (HLS). However, these designs pose a challenge: correctly and optimally sizing first-in-first-out (FIFO) channel buffers. FIFO sizes are user-defined parameters, introducing a trade-off between latency and area—undersized FIFOs cause stalls and increase latency, while oversized FIFOs waste on-chip memory. In many cases, insufficient FIFO sizes can also lead to deadlocks. Deciding the best FIFO sizes is non-trivial. Existing methods make limiting assumptions about FIFO access patterns, overallocate FIFOs conservatively, or use time-consuming RTL simulations to evaluate different FIFO sizes. Furthermore, we highlight that runtime-based analyses (i.e., simulation) are the only way to solve the FIFO optimization problem while ensuring a deadlock-free solution for designs with data-dependent control flow. To tackle this challenge, we propose FIFOAdvisor, a framework to automatically decide FIFO sizes in HLS designs. Our approach is powered by LightningSim, a fast simulator that is 99.9% cycle-accurate and supports millisecond-scale incremental simulations with new FIFO configurations. We formulate FIFO sizing as a dual-objective black-box optimization problem and explore various heuristic and search-based methods to analyze the latency–resource trade-off. We also integrate FIFOAdvisor with Stream-HLS, a recent framework for optimizing affine dataflow designs lowered from C++, MLIR, or PyTorch, enabling deeper optimization of the heavily-used FIFOs in these workloads. We evaluate FIFOAdvisor on a suite of Stream-HLS benchmarks, including linear algebra and deep learning workloads, to demonstrate our approach’s ability to optimize large and dynamic dataflow patterns. Our results show Pareto-optimal latency–memory usage frontiers for FIFO configurations generated via different optimization strategies. Compared to baseline designs with naïvely-sized FIFOs, FIFOAdvisor identifies configurations with much lower memory usage and minimal delay overhead. Additionally, we measure the runtime of our optimization process and demonstrate significant speedups compared to traditional HLS/RTL co-simulation-based approaches, making FIFOAdvisor practical for rapid design space exploration. Finally, we present a case study using FIFOAdvisor to optimize a complex hardware accelerator with non-trivial data-dependent control flow. Code and results open-sourced at https://github.com/sharc-lab/fifo-advisor. Stefan Abi-Karam, Rishov Sarkar, Suhail Basalama, Jason Cong, Cong Hao |
ASP-DAC | 2 |
| 2026 | From Acceleration to Accelerating Acceleration: Applications and Tools for Democratizing Hardware DesignabstractHigh-level synthesis allows users to express hardware designs in terms of software languages, raising the level of abstraction. However, most designers still prefer writing low-level RTL code. We present two classes of research works aiming to improve the appeal of HLS: (1) accelerators to showcase the power of HLS for ML applications, and (2) tools to improve the HLS developer experience. Rishov Sarkar, Cong Hao |
FCCM | 1 |
| 2025 | OmniSim: Simulating Hardware with C Speed and RTL Accuracy for High-Level Synthesis DesignsabstractHigh-Level Synthesis (HLS) is increasingly popular for hardware design using C/C++ instead of Register-Transfer Level (RTL).To express concurrent hardware behavior in a sequential language like C/C++, HLS tools introduce constructs such as infinite loops and dataflow modules connected by FIFOs (first-in first-out).While these constructs can represent concurrency, efficiently and accurately simulating them at C level remains challenging.First, without hardware timing information, functional verification typically requires slow RTL synthesis and simulation, as the current approaches in commercial HLS tools.Second, cycle-accurate performance metrics, such as end-to-end latency or throughput, also rely on RTL simulation.No existing HLS tool fully overcomes the first limitation.For the second, prior work such as LightningSim partially improves simulation speed but lacks support for advanced dataflow features like cyclic dependencies and non-blocking FIFO accesses.To overcome both limitations, we propose OmniSim, a framework that significantly extends the simulation capabilities of both academic and commercial HLS tools.First, OmniSim enables fast and accurate simulation of complex dataflow designs, especially those explicitly declared unsupported by commercial tools.It does so through sophisticated software multi-threading, where threads are orchestrated by querying and updating a set of FIFO tables that explicitly record exact hardware timing of each FIFO access.Second, OmniSim achieves near-C simulation speed with near-RTL accuracy for both functionality and performance, via flexibly coupled and overlapped functionality and performance simulations.We demonstrate that OmniSim successfully simulates eleven designs previously unsupported by any HLS tool, achieving up to 35.9× speedup over traditional C/RTL co-simulation, and up to 6.61× speedup over the state-of-the-art yet less capable simulator, LightningSim, on its own benchmark suite. Rishov Sarkar, Cong Hao |
MICRO | 1 |
| 2024 | Ph.D. Project: Modernizing High-Level Hardware Design WorkflowsabstractThe field of machine learning (ML) is growing at a break-neck pace. As ML models increasingly demand real-time performance, energy efficiency, and high throughput, researchers look to customized computing architectures, such as FPGAs. Rishov Sarkar, Cong Hao |
FCCM | 1 |
| 2024 | LightningSimV2: Faster and Scalable Simulation for High-Level Synthesis via Graph Compilation and OptimizationabstractHigh-Level Synthesis (HLS) enables rapid prototyping of complex hardware designs by translating C or C++ code to low-level RTL code. However, the testing and evaluation of HLS designs still typically rely on slow RTL-level simulators that can take hours to provide feedback, especially for complex designs. A recent work, LightningSim, helps to solve this problem by providing a simulation workflow one to two orders of magnitude faster than RTL simulation. However, it still exhibits inefficiencies due to several types of redundant computation, making it slow for large design simulation and design space exploration. Addressing these inefficiencies, we introduce LightningSimV2, a much faster and scalable simulation tool. LightningSimV2 features three main innovations. First, we perform compile-time static analysis, exploiting the repetitive structures in HLS designs, e.g., loops, to reduce the simulation workload. Second, we propose a novel graph-based simulation approach, with decoupled simulation graph construction step and graph traversal step, significantly reducing repeated computation. Third, benefiting from the decoupled approach, LightningSimV2 can perform incremental stall analysis extremely fast, enabling highly efficient design space exploration of large numbers of complex hardware parameters, e.g., optimal FIFO depths. Moreover, the DSE is well-suited for parallel computing, further improving the DSE efficiency. Compared with LightningSim, LightningSimV2 achieves up to 3.5× speedup in full simulation and up to 577× speed up for incremental DSE. Our code is open-source on GitHub at https://github.com/sharc-lab/LightningSim/tree/v0.2.0. Rishov Sarkar, Rachel Paul, Cong Hao |
FCCM | 1 |
| 2023 | LightningSim: Fast and Accurate Trace-Based Simulation for High-Level SynthesisabstractHigh-Level Synthesis allows hardware designers to create complex RTL designs using C/C++. The traditional HLS workflow involves iterations of C/C++ simulation for partial functional verification and HLS synthesis for coarse timing estimates. However, neither C/C++ simulation nor HLS synthesis estimates can account for complex behaviors like FIFO interactions and pipeline stalls, thereby obscuring problems like deadlocks and latency overheads. Such problems are revealed only through C/RTL co-simulation, which is typically orders of magnitude slower than either C/C++ simulation or HLS synthesis, far too slow to integrate into the edit-run development cycle. Addressing this, we propose LightningSim, a fast simulation tool for HLS that combines the speed of native C/C++ with the accuracy of C/RTL co-simulation. LightningSim directly operates on the LLVM intermediate representation (IR) code and accurately simulates a hardware design's dynamic behavior. First, it traces LLVM IR execution to capture the run-time information; second, it maps the static HLS scheduling information to the trace to simulate the dynamic behavior; third, it calculates stalls and deadlocks from inter-function interactions to get precise cycle counts. Evaluated on 33 benchmarks, LightningSim produces 99.9%-accurate timing estimates up to 95x faster than RTL simulation. Our code is publicly available on GitHub.11https://github.com/sharc-lab/LightningSim Rishov Sarkar, Cong Hao |
FCCM | 1 |
| 2023 | From Acceleration to Accelerating Acceleration: Modernizing the Accelerator Landscape using High-Level SynthesisabstractThe field of machine learning continues to grow at an exponential rate. As conventional CPU and GPU architectures struggle to keep up with demands of real-time performance, energy efficiency, and high throughput, researchers are increasingly turning towards FPGA acceleration to solve these challenges through customized computing architectures. Recently, High-Level Synthesis (HLS) tools have emerged, aiming to bring the ease of C/C++ software development to FPGA and ASIC hardware acceleration through an automatic conversion process (“synthesis”) to RTL code. However, HLS still remains far from the level of accessibility achieved by software design, retaining most of the challenges of traditional RTL-based hardware design. For instance: Rishov Sarkar, Cong Hao |
FCCM | 1 |
| 2023 | FlowGNN: A Dataflow Architecture for Real-Time Workload-Agnostic Graph Neural Network InferenceabstractGraph neural networks (GNNs) have recently exploded in popularity thanks to their broad applicability to graph-related problems such as quantum chemistry, drug discovery, and high energy physics. However, meeting demand for novel GNN models and fast inference simultaneously is challenging due to the gap between developing efficient accelerators and the rapid creation of new GNN models. Prior art focuses on accelerating specific classes of GNNs, such as Graph Convolutional Networks (GCN), but lacks generality to support a wide range of existing or new GNN models. Furthermore, most works rely on graph pre-processing to exploit data locality, making them unsuitable for real-time applications. To address these limitations, in this work, we propose a generic dataflow architecture for GNN acceleration, named FlowGNN, which is generalizable to the majority of message-passing GNNs. The contributions are three-fold. First, we propose a novel and scalable dataflow architecture, which generally supports a wide range of GNN models with message-passing mechanism. The architecture features a configurable dataflow optimized for simultaneous computation of node embedding, edge embedding, and message passing, which is generally applicable to all models. We also propose a rich library of model-specific components. Second, we deliver ultra-fast real-time GNN inference without any graph pre-processing, making it agnostic to dynamically changing graph structures. Third, we verify our architecture on the Xilinx Alveo U50 FPGA board and measure the on-board end-to-end performance. We achieve a speed-up of up to 24–254× against CPU (6226R) and 1.3–477× against GPU (A6000) (with batch sizes 1 through 1024); we also outperform the SOTA GNN accelerator I-GCN by 1.26× speedup and 1.55× energy efficiency over four datasets. Our implementation code and on-board measurement are publicly available on GitHub.1 Rishov Sarkar, Stefan Abi-Karam, Lakshmi Sathidevi, Cong Hao |
HPCA | 1 |
| 2023 | INR-Arch: A Dataflow Architecture and Compiler for Arbitrary-Order Gradient Computations in Implicit Neural Representation ProcessingabstractAn increasing number of researchers are finding use for nth-order gradient computations for a wide variety of applications, including graphics, meta-learning (MAML), scientific computing, and most recently, implicit neural representations (INRs). Recent work shows that the gradient of an INR can be used to edit the data it represents directly without needing to convert it back to a discrete representation. However, given a function represented as a computation graph, traditional architectures face challenges in efficiently computing its nth-order gradient due to the higher demand for computing power and higher complexity in data movement. This makes it a promising target for FPGA acceleration. In this work, we introduce INR-Arch, a framework that transforms the computation graph of an nth-order gradient into a hardware-optimized dataflow architecture. We address this problem in two phases. First, we design a dataflow architecture that uses FIFO streams and an optimized computation kernel library, ensuring high memory efficiency and parallel computation. Second, we propose a compiler that extracts and optimizes computation graphs, automatically configures hardware parameters such as latency and stream depths to optimize throughput, while ensuring deadlock-free operation, and outputs High-Level Synthesis (HLS) code for FPGA implementation. We utilize INR editing as our benchmark, presenting results that demonstrate 1.8-4.8x and 1.5-3.6x speedup compared to CPU and GPU baselines respectively. Furthermore, we obtain 3.1-8.9x and 1.7-4.3x lower memory usage, and 1.7-11.3x and 5.5-32.8x lower energy-delay product. Our framework will be made open-source and available on GitHub.****https://github.com/sharc-lab/inr-arch Stefan Abi-Karam, Rishov Sarkar, Dejia Xu, Zhiwen Fan, Zhangyang Wang, Cong Hao |
ICCAD | 2 |
| 2023 | Edge-MoE: Memory-Efficient Multi-Task Vision Transformer Architecture with Task-Level Sparsity via Mixture-of-ExpertsabstractThe computer vision community is embracing two promising learning paradigms: the Vision Transformer (ViT) and Multi-task Learning (MTL). ViT models show extraordinary performance over traditional convolution networks but are commonly recognized as computation-intensive, especially the self-attention with quadratic complexity. MTL uses one model to infer multiple tasks with better performance by enforcing shared representation among tasks, but a huge drawback is that, most MTL regimes require activation of the entire model even when only one or a few tasks are needed, causing significant computing waste. M3ViT is the latest multi-task Vi$T$model that introduces mixture-of-experts (MoE), where only a small portion of subnetworks (“experts”) are sparsely and dynamically activated based on the current task. M3Vi$T$achieves better accuracy and over 80% computation reduction and paves the way for efficient real-time MTL using ViT. Despite the algorithmic advantages of MTL, ViT, and even M3ViT, there are still many challenges for efficient deployment on FPGA. For instance, in general Transformer/ViT models, the self-attention is known as computational intensive and requires high bandwidth. In addition, softmax operations and the activation function GELU are extensively used, which unfortunately can consume more than half of the entire FPGA resource (LUTs). In the M3ViT model, the promising MoE mechanism for multi-task exposes new challenges for memory access overhead and also increases resource usage because of more layer types. To address these challenges in both general Transformer/ViT models and the state-of-the-art multi-task M3ViT with MoE, we propose Edge-MoE, the first end-to-end FPGA accelerator for multi-task ViT with a rich collection of architectural innovations. First, for general Transformer/ViT models, we propose (1) a novel reordering mechanism for self-attention, which reduces the bandwidth requirement from proportional to constant regardless of the target parallelism; (2) a fast single-pass softmax approximation; (3) an accurate and low-cost GELU approximation, which can significantly reduce the computation latency and resource usage; and (4) a unified and flexible computing unit that can be shared by almost all computational layers to maximally reduce resource usage. Second, for the advanced multi-task M3ViT with MoE, we propose a novel patch reordering method to completely eliminate any memory access overhead. Third, we deliver on-board implementation and measurement on Xilinx ZCU102 FPGA, with verified functionality and open-sourced hardware design, which achieves 2.24× and 4.90× better energy efficiency comparing with GPU (A6000) and CPU (Xeon 6226R), respectively. A real-time video demonstration of our accelerated multi-task ViT on an autonomous driving dataset is available in GitHub,11https://github.com/sharc-lab/Edge-MoE/raw/main/demo.mp4 together with our FPGA design using High-Level Synthesis, host code, FPGA bitstream, and on-board performance results. Rishov Sarkar, Hanxue Liang, Zhiwen Fan, Zhangyang Wang, Cong Hao |
ICCAD | 1 |
| 2022 | M³ViT: Mixture-of-Experts Vision Transformer for Efficient Multi-task Learning with Model-Accelerator Co-designabstractMulti-task learning (MTL) encapsulates multiple learned tasks in a single model and often lets those tasks learn better jointly. Multi-tasking models have become successful and often essential for many sophisticated systems such as autonomous driving and indoor robots. However, when deploying MTL onto those real-world systems that are often resource-constrained or latency-sensitive, two prominent challenges arise: (i) during training, simultaneously optimizing all tasks is often difficult due to gradient conflicts across tasks, and the challenge is amplified when a growing number of tasks have to be squeezed into one compact model; (ii) at inference, current MTL regimes have to activate nearly the entire model even to just execute a single task. Yet most real systems demand only one or two tasks at each moment, while flexibly switching between tasks per need: therefore such “all tasks activated” inference is also highly inefficient and non-scalable in practice. In this paper, we present a model-accelerator co-design framework to enable efficient on-device MTL, that tackles both training and inference bottlenecks. Our framework, dubbed M³ViT, customizes mixture-of-experts (MoE) layers into a vision transformer (ViT) backbone for MTL, and sparsely activates task-specific experts during training, which effectively disentangles the parameter spaces to avoid different tasks’ training conflicts. Then at inference with any task of interest, the same design allows for activating only the task-corresponding sparse “expert” pathway, instead of the full model. Our new model design is further enhanced by hardware-level innovations, in particular, a novel computation reordering scheme tailored for memory-constrained MTL that achieves zero-overhead switching between tasks and can scale to any number of experts. Extensive experiments on PASCAL-Context and NYUD-v2 datasets at both software and hardware levels are conducted to demonstrate the effectiveness of the proposed design. When executing the practical scenario of single-task inference, M³ViT achieves higher accuracies than encoder-focused MTL methods, while significantly reducing 88% inference FLOPs. When implemented on a hardware platform of one Xilinx ZCU104 FPGA, our co-design framework reduces the memory requirement by 2.40×, while achieving energy efficiency (as the product of latency and power) up to 9.23× times higher than a comparable FPGA baseline. Hanxue Liang, Zhiwen Fan, Rishov Sarkar, Ziyu Jiang, Tianlong Chen 0001, Yu Cheng 0001, Cong Hao, Zhangyang Wang |
NeurIPS | 3 |