Shashank Obla

dblp:266/8593 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0003-0467-9439ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Lightweight Queueing Abstraction for Rapid Simulation and Automated Tuning of Input-Dependent Streaming Pipelines on FPGAs
abstract
The programmability of FPGAs enables designs to be tuned to specific deployment and use-cases. This capability is critical for input-dependent streaming pipelines, whose optimal configuration varies not only across resource limits and performance targets, but also with the data being processed. However, existing methods fail to scale for large, real-world designs with long workloads. To address this challenge, we propose RapidQ, a performance model that allows designers to capture the dataflow of their streaming pipeline as a queueing system, enabling fast case-by-case re-tuning for resource efficiency. Unburdened by the functional details of the design, a trace model drives a fast queueing simulation that can predict the performance of the pipeline across various buffer sizes and module throughput configurations without repeating full functional simulations. Our simulator is over 7x faster than the state-of-the-art and yields up to 42% resource savings for real-world workloads.
Shashank Obla, James C. Hoe
FCCM1
2026 Reconfigurable Computing Challenge: RapidScan High-Throughput Parameterized HLS-based Streaming String Matching Library for FPGAs
Shashank Obla, Tommy Tracy II, Matthew Beck, James C. Hoe, Kevin Skadron, Wajih Ul Hassan
FCCM1
2026 Analysis and Optimization of Input-Dependent Stream Processing Pipelines on FPGAs
abstract
The optimal tuning for input-dependent stream processing pipelines depends on the input characteristics, which can vary by deployment site and even time of day. FPGAs can provide these deployment-specific customized solutions, but the ability to tune is hindered by long RTL-level design turnaround times. A higher-level abstraction is essential for expressing the input-dependent performance-resource tradeoffs to support repeated case-by-case retuning. We present the FPGA-as-a-System (FaaS) queueing-based performance model that allows an FPGA designer to decouple the input-dependent performance behavior of their stream processing pipeline from the low-level functional details. A fast queueing-based performance simulator of the FaaS model incorporates the input characteristics, guiding the tuning and optimization of input-dependent streaming FPGA pipelines to meet performance targets with fewer resources. We present a case study on applying the FaaS workflow to tune a multi-string matching pipeline for network security and log monitoring. The results show that, starting from an implementation already manually tuned for the throughput target, the FaaS approach uncovers previously overlooked bottlenecks and underutilization. By addressing these deficiencies, a revised implementation achieves a reduction of more than 40% in resource footprint while maintaining the original implementation's throughput.
Shashank Obla, James C. Hoe
FPGA1
2023 Exploiting the Common Case When Accelerating Input-Dependent Stream Processing by FPGA
abstract
FPGAs have traditionally been successful in accelerating stream processing applications where the amount and type of work performed on each record—e.g., image, packet—do not depend on the record's contents. On the other hand, accelerating ‘input-dependent’ stream processing on FPGAs presents a much more challenging problem where different records in the stream can require widely different operations. It is inefficient and unnecessary to support all operations at the same throughput when the distributions of the operations are skewed. In this paper, we examine the application of the ”make the common case fast” strategy to efficiently accelerate ‘input-dependent’ stream processing on FPGAs. In particular, we study the use offast-slow pathandearly-exittechniques in the design of an FPGA-accelerated network intrusion prevention system (IPS). To avoid overfitting when common-case behavior is varied, we further examinecompile-time re-tuningandruntime adaptationtechniques. A quantitative analysis shows that fast-slow path and early-exit techniques can save an order of magnitude of resources compared to a common-case unaware IPS design. Compile-time re-tuning to specific conditions further achieves 30% – 94% BRAM savings relative to a generalized design. Adding runtime adaptation improves the zero-loss throughput by 1.43 – 2.75 × compared to a fixed design.
Joseph Melber, Siddharth Sahay, Shashank Obla, Eriko Nurvitadhi, James C. Hoe
IEEE Trans. Computers4