EDBT 2026 Demo / reviewers in the wild / expert
Stefan Abi-Karam
dblp:312/4591
· DBLP profile ↗
7ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0002-6697-8517ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FIFOAdvisor: A DSE Framework for Automated FIFO Sizing of High-Level Synthesis DesignsabstractDataflow hardware designs are important for efficient algorithm implementations on FPGAs across various domains using high-level synthesis (HLS). However, these designs pose a challenge: correctly and optimally sizing first-in-first-out (FIFO) channel buffers. FIFO sizes are user-defined parameters, introducing a trade-off between latency and area—undersized FIFOs cause stalls and increase latency, while oversized FIFOs waste on-chip memory. In many cases, insufficient FIFO sizes can also lead to deadlocks. Deciding the best FIFO sizes is non-trivial. Existing methods make limiting assumptions about FIFO access patterns, overallocate FIFOs conservatively, or use time-consuming RTL simulations to evaluate different FIFO sizes. Furthermore, we highlight that runtime-based analyses (i.e., simulation) are the only way to solve the FIFO optimization problem while ensuring a deadlock-free solution for designs with data-dependent control flow. To tackle this challenge, we propose FIFOAdvisor, a framework to automatically decide FIFO sizes in HLS designs. Our approach is powered by LightningSim, a fast simulator that is 99.9% cycle-accurate and supports millisecond-scale incremental simulations with new FIFO configurations. We formulate FIFO sizing as a dual-objective black-box optimization problem and explore various heuristic and search-based methods to analyze the latency–resource trade-off. We also integrate FIFOAdvisor with Stream-HLS, a recent framework for optimizing affine dataflow designs lowered from C++, MLIR, or PyTorch, enabling deeper optimization of the heavily-used FIFOs in these workloads. We evaluate FIFOAdvisor on a suite of Stream-HLS benchmarks, including linear algebra and deep learning workloads, to demonstrate our approach’s ability to optimize large and dynamic dataflow patterns. Our results show Pareto-optimal latency–memory usage frontiers for FIFO configurations generated via different optimization strategies. Compared to baseline designs with naïvely-sized FIFOs, FIFOAdvisor identifies configurations with much lower memory usage and minimal delay overhead. Additionally, we measure the runtime of our optimization process and demonstrate significant speedups compared to traditional HLS/RTL co-simulation-based approaches, making FIFOAdvisor practical for rapid design space exploration. Finally, we present a case study using FIFOAdvisor to optimize a complex hardware accelerator with non-trivial data-dependent control flow. Code and results open-sourced at https://github.com/sharc-lab/fifo-advisor. Stefan Abi-Karam, Rishov Sarkar, Suhail Basalama, Jason Cong, Cong Hao |
ASP-DAC | 1 |
| 2026 | Focus Session: Stepping Stones Towards Domain Acceleration Using AI-Driven High-Level Synthesis DesignabstractField-programmable gate arrays (FPGAs) combined with high-level synthesis (HLS) enable domain-specific accelerators from high-level algorithmic descriptions, but HLS has not yet democratized hardware design in practice. High-performance HLS design for accelerators still requires expert knowledge in HLS optimization, design space exploration (DSE), performance analysis, and toolflows with limited debuggability and analysis even for expert human designers.To reduce this expertise burden, the community has explored ML-guided DSE, quality-of-results (QoR) proxy models, and LLMs for HLS code generation, editing, and optimization. However, progress remains constrained by scarce, accessible, high-quality HLS datasets and the difficulty of building diverse design collections. Many existing benchmarks sample design variations from a small pool of kernels rather than expanding the base pool of diverse designs. At the same time, most LLM hardware-design benchmarking emphasizes HDLs (e.g., Verilog), leaving comprehensive benchmarking infrastructure for HLS tasks comparatively limited. Furthermore, LLM-driven agents have shown strong automation and performance for software development tasks while agentic HLS design workflows still commonly lack robust interfaces to HLS tools, structured access to HLS reports, and feedback loops spanning HLS and downstream implementation.To address these gaps, we present progress toward end-to-end rapid design of domain-specific accelerators using AI-driven approaches for HLS by targeting three problems: (1) HLS design datasets, (2) LLM benchmarking for HLS design, and (3) end-to-end agentic HLS design.For HLS design datasets, we present HLSFactory, an end-to-end framework for building diverse HLS datasets that expands individual HLS designs into diverse, cross-vendor design spaces, runs HLS/FPGA toolflows at scale, and aggregates standardized outputs into packaged datasets for benchmarking and ML, deep learning, and LLM research. HLSFactory is API-driven and extensible, enabling users to contribute designs or results and plug in custom toolflows at any stage.For LLM benchmarking, we present HLS-Eval, a benchmark and evaluation framework for LLM-driven HLS design that targets HLS code generation from natural language and HLS-specific optimization edits to existing code. It includes an "LLM-ready" benchmark for zero-shot evaluation, primarily drawing designs from HLSFactory for benchmarking.Finally, we also present our ongoing effort toward an end-to-end agentic HLS design workflow that unifies accelerator design stages in a multi-agent framework, with integrated HLS design tools from our prior works including LightningSim (fast cycle-accurate simulation), FIFO-Advisor (automated HLS FIFO sizing), AutoDSE/OptDSL/HLSFactory (structured DSE for HLS), deep-learning proxy models (QoR prediction), and libvhls (programmatic HLS tool interaction and hierarchical report analysis). Through this effort, we also plan to extend HLS-Eval to enable large-scale, parallel evaluations on realistic open-source HLS accelerators from academic projects, collecting agentic metrics on correctness, performance, agent runtime, and compute cost. Stefan Abi-Karam, Miaoyan Zhou, Cong Hao |
DATE | 1 |
| 2023 | M5: Multi-modal Multi-task Model Mapping on Multi-FPGA with Accelerator Configuration SearchabstractRecent machine learning (ML) models have advanced from single-modality single-task to multi-modality multi-task (MMMT). MMMT models typically have multiple backbones of different sizes along with complicated connections, exposing great challenges for hardware deployment. For scalable and energy-efficient implementations, multi-FPGA systems are emerging as the ideal design choices. However, finding the optimal solutions for mapping MMMT models onto multiple FPGAs is non-trivial. Existing mapping algorithms focus on either streamlined linear deep neural network architectures or only the critical path of simple heterogeneous models. Direct extensions of these algorithms for MMMT models lead to sub-optimal solutions. To address these shortcomings, we propose M5, a novel MMMT Model Mapping framework for Multi-FPGA platforms. In addition to handling multiple modalities present in the models, M5 can flexibly explore accelerator configurations and possible resource sharing opportunities to significantly improve the system performance. For various computation-heavy MMMT models, experiment results demonstrate that M5 can remarkably outperform existing mapping methods and lead to an average reduction of 35%, 62%, and 70% in the number of low-end, mid-end, and high-end FPGAs required to achieve the same throughput, respectively. Code is publicly available1, Akshay Karkal Kamath, Stefan Abi-Karam, Ashwin Bhat, Cong Hao |
DATE | 2 |
| 2023 | GNNBuilder: An Automated Framework for Generic Graph Neural Network Accelerator Generation, Simulation, and OptimizationabstractThere are plenty of graph neural network (GNN) accelerators being proposed. However, they highly rely on users' hardware expertise and are usually optimized for one specific GNN model, making them challenging for practical use. Therefore, in this work, we propose GNNBuilder, the first automated, generic, end-to-end GNN accelerator generation framework. It features four advantages: (1) GNNBuilder can automatically generate GNN accelerators for a wide range of GNN models arbitrarily defined by users; (2) GNNBuilder takes standard PyTorch programming interface, introducing zero overhead for algorithm developers; (3) GNNBuilder supports end-to-end code generation, simulation, accelerator optimization, and hardware deployment, realizing push-button GNN accelerator design; (4) GNNBuilder is equipped with accurate performance models for its generated accelerators, enabling fast and flexible design space exploration (DSE). In the experiments, we show that our accelerator performance model has errors within 36% for latency prediction and 18% for BRAM count prediction. Additionally, we show that our generated accelerators can outperform CPU by 6.33× and GPU by 6.87×. This framework is open-source, and the code is available at https://github.com/sharc-lab/gnn-builder. Stefan Abi-Karam, Cong Hao |
FPL | 1 |
| 2023 | FlowGNN: A Dataflow Architecture for Real-Time Workload-Agnostic Graph Neural Network InferenceabstractGraph neural networks (GNNs) have recently exploded in popularity thanks to their broad applicability to graph-related problems such as quantum chemistry, drug discovery, and high energy physics. However, meeting demand for novel GNN models and fast inference simultaneously is challenging due to the gap between developing efficient accelerators and the rapid creation of new GNN models. Prior art focuses on accelerating specific classes of GNNs, such as Graph Convolutional Networks (GCN), but lacks generality to support a wide range of existing or new GNN models. Furthermore, most works rely on graph pre-processing to exploit data locality, making them unsuitable for real-time applications. To address these limitations, in this work, we propose a generic dataflow architecture for GNN acceleration, named FlowGNN, which is generalizable to the majority of message-passing GNNs. The contributions are three-fold. First, we propose a novel and scalable dataflow architecture, which generally supports a wide range of GNN models with message-passing mechanism. The architecture features a configurable dataflow optimized for simultaneous computation of node embedding, edge embedding, and message passing, which is generally applicable to all models. We also propose a rich library of model-specific components. Second, we deliver ultra-fast real-time GNN inference without any graph pre-processing, making it agnostic to dynamically changing graph structures. Third, we verify our architecture on the Xilinx Alveo U50 FPGA board and measure the on-board end-to-end performance. We achieve a speed-up of up to 24–254× against CPU (6226R) and 1.3–477× against GPU (A6000) (with batch sizes 1 through 1024); we also outperform the SOTA GNN accelerator I-GCN by 1.26× speedup and 1.55× energy efficiency over four datasets. Our implementation code and on-board measurement are publicly available on GitHub.1 Rishov Sarkar, Stefan Abi-Karam, Lakshmi Sathidevi, Cong Hao |
HPCA | 2 |
| 2023 | INR-Arch: A Dataflow Architecture and Compiler for Arbitrary-Order Gradient Computations in Implicit Neural Representation ProcessingabstractAn increasing number of researchers are finding use for nth-order gradient computations for a wide variety of applications, including graphics, meta-learning (MAML), scientific computing, and most recently, implicit neural representations (INRs). Recent work shows that the gradient of an INR can be used to edit the data it represents directly without needing to convert it back to a discrete representation. However, given a function represented as a computation graph, traditional architectures face challenges in efficiently computing its nth-order gradient due to the higher demand for computing power and higher complexity in data movement. This makes it a promising target for FPGA acceleration. In this work, we introduce INR-Arch, a framework that transforms the computation graph of an nth-order gradient into a hardware-optimized dataflow architecture. We address this problem in two phases. First, we design a dataflow architecture that uses FIFO streams and an optimized computation kernel library, ensuring high memory efficiency and parallel computation. Second, we propose a compiler that extracts and optimizes computation graphs, automatically configures hardware parameters such as latency and stream depths to optimize throughput, while ensuring deadlock-free operation, and outputs High-Level Synthesis (HLS) code for FPGA implementation. We utilize INR editing as our benchmark, presenting results that demonstrate 1.8-4.8x and 1.5-3.6x speedup compared to CPU and GPU baselines respectively. Furthermore, we obtain 3.1-8.9x and 1.7-4.3x lower memory usage, and 1.7-11.3x and 5.5-32.8x lower energy-delay product. Our framework will be made open-source and available on GitHub.****https://github.com/sharc-lab/inr-arch Stefan Abi-Karam, Rishov Sarkar, Dejia Xu, Zhiwen Fan, Zhangyang Wang, Cong Hao |
ICCAD | 1 |
| 2023 | GNNBuilder: An Automated Framework for Generic Graph Neural Network Accelerator Generation, Simulation, and Optimization
Stefan Abi-Karam, Cong Hao |
Softw. Pract. Exp. | 1 |