David Ojika

dblp:202/5014 · also Dave Ojika · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0003-0911-1893ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 RISCBench: Benchmarking RISC-V Orchestration Efficiency in FPGA and FPGA-Like Computing Engines: An Industry Evaluation of Control-Plane Bottlenecks and Sustained Throughput Metrics
abstract
Heterogeneous systems increasingly rely on RISC-V cores as orchestration engines to manage data movement, synchronization, and scheduling across accelerators and reconfigurable fabrics. Conventional performance metrics, such as FLOPs, TOPS/W, or energy per operation, do not capture orchestration efficiency, even though it often dictates sustained system behavior. This gap is increasingly relevant as systems evolve toward tightly coupled heterogeneous fabrics and co-packaged accelerators, where control-plane behavior determines whether these platforms achieve their promised performance. We present RISCBench, a kernel benchmark suite and open methodology for quantifying orchestration efficiency. RISCBench introduces the Sustained Instantaneous Throughput (SIT) metric, which accumulates instantaneous throughput over near-aggregate execution intervals, capturing sustained efficiency beyond peak rates. The methodology is evaluated across representative platforms spanning soft and hard RISC-V orchestration engines, including FPGA-based prototyping and accelerator-class implementations. Results highlight synchronization and data residency driven tradeoffs that limit realized throughput beyond peak performance, motivating SIT as a practical, platform-independent descriptor for evaluating orchestration efficiency in heterogeneous systems and AI inference applications.
David Ojika, Projjal Gupta, Preethi Budi, Herman Lam, Shreya Mehrotra
FPGA1
2022 FSHMEM: Supporting Partitioned Global Address Space on FPGAs for Large-Scale Hardware Acceleration Infrastructure
abstract
By providing highly efficient one-sided communication with globally shared memory space, Partitioned Global Address Space (PGAS) has become one of the most promising parallel computing models in high-performance computing (HPC). Meanwhile, FPGA is getting attention as an alternative compute platform for HPC systems with the benefit of custom computing and design flexibility. However, the exploration of PGAS has not been conducted on FPGAs, unlike the traditional message passing interface. This paper proposes FSHMEM, a software/hardware framework that enables the PGAS programming model on FPGAs. We implement the core functions of GASNet specification on FPGA for native PGAS integration in hardware, while its programming interface is designed to be highly compatible with legacy software. Our experiments show that FSHMEM achieves the peak bandwidth of 3813 MB/s, which is more than 95% of the theoretical maximum, outperforming the prior works by 9.5×. It records 0.35us and 0.59us latency for remote write and read operations, respectively. Finally, we conduct a case study on the two Intel D5005 FPGA nodes integrating Intel's deep learning accelerator. The two-node system programmed by FSHMEM achieves 1.94× and 1.98× speedup for matrix multiplication and convolution operation, respectively, showing its scalability notential for HPC infrastructure.
Yashael Faith Arthanto, David Ojika, Joo-Young Kim 0001
FPL2
2021 Optimized FPGA-based Deep Learning Accelerator for Sparse CNN using High Bandwidth Memory
abstract
Large Convolutional Neural Networks (CNNs) are often pruned and compressed to reduce the amount of parameters and memory requirement. However, the resulting irregularity in the sparse data makes it difficult for FPGA accelerators that contains systolic arrays of Multiply-and-Accumulate (MAC) units, such as Intel's FPGA-based Deep Learning Accelerator (DLA), to achieve their maximum potential. Moreover, FPGAs with low-bandwidth off-chip memory could not satisfy the memory bandwidth requirement for sparse matrix computation. In this paper, we present 1) a sparse matrix packing technique that condenses sparse inputs and filters before feeding them into the systolic array of MAC units in the Intel DLA, and 2) a customization of the Intel DLA which allows the FPGA to efficiently utilize a high bandwidth memory (HBM2) integrated in the same package. For end-to-end inference with randomly pruned ResNet-50/MobileNet CNN models, our experiments demonstrate 2.7x/3x performance improvement compared to an FPGA with DDR4, 2.2x/2.1x speedup against a server-class Intel SkyLake CPU, and comparable performance with 1.7x/2x power efficiency gain as compared to an NVidia V100 GPU.
David Ojika, Bhavesh Patel, Herman Lam
FCCM2
2019 Accelerating Scientific Discovery with SCAIGATE Science Gateway
abstract
The demand for computational accelerators (GPUs, FPGAs, ASICs, etc.) is growing due to the widening variety of datacenter applications fueled by recent scientific breakthroughs that leverage artificial intelligence (AI). As much as these applications (e.g., cosmology, physics, etc.) have continued to witness record-breaking accuracy in predictive capabilities due to AI widespread influence, the infrastructure and workflow to take these applications out of research labs into production and business use-cases continues to lag. To address these important infrastructural challenges, we present SCAIGATE, a prototype science gateway with a simplified workflow aimed at facilitating model building/validation workflows in large-scale scientific applications.
David Ojika, Bhavesh Patel, Ann Gordon-Ross, Herman Lam
eScience2
2017 SWiF: A Simplified Workload-Centric Framework for FPGA-Based Computing
abstract
In this paper, we introduce SWiF - Simplified Workload-intuitive Framework - a workload-centric, application programming framework designed to simplify the large-scale deployment of FPGAs in end-to-end applications. SWiF can intelligently mediate access to shared resources by orchestrating the distribution and scheduling of tasks across a heterogeneous mix of FPGA and CPU resources in order to improve utilization and maintain system requirements. We implemented SWiF atop Intel Accelerator Abstraction Layer (AAL) and deployed the resulting software stack in a datacenter with an Intel-based Xeon+FPGA server running Apache Spark. We demonstrate that by using SWiF's API, developers can flexibly and easily deploy FPGA-enabled applications and frameworks with almost no change to existing software stack. In particular, we demonstrate that by offloading through SWiF the compression workload of Spark unto FPGA, we gain a speedup of 3.2X in total job execution, and up to 5X when Spark's Resilient Distributed Datasets (RDDs) are persisted in memory.
David Ojika, Piotr Majcher, Wojciech Neubauer, Suchit Subhaschandra, Darin Acosta
FCCM1