Joel Mandebi

dblp:217/9172 · also Joel Mandebi Mbongue · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0002-9277-5043ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 6 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Multi-Tenant Cloud FPGA: A Survey on Security, Trust, and Privacy
abstract
With the growing demand for enhanced performance and scalability in cloud applications and systems, data center architectures are evolving to incorporate heterogeneous computing fabrics that leverage CPUs, GPUs, and FPGAs. Unlike traditional processing platforms like CPUs and GPUs, FPGAs offer the unique ability for hardware reconfiguration at runtime, enabling improved and tailored performance, flexibility, and acceleration. FPGAs excel at executing large-scale search optimization, acceleration, and signal processing tasks while consuming low power and minimizing latency. Major public cloud providers, such as Amazon, Huawei, Microsoft, Alibaba, and others, have already begun integrating FPGA-based cloud acceleration services into their offerings. Although FPGAs in cloud applications facilitate customized hardware acceleration, they also introduce new security challenges that demand attention. Granting cloud users the capability to reconfigure hardware designs after deployment may create potential vulnerabilities for malicious users, thereby jeopardizing entire cloud platforms. In particular, multi-tenant FPGA services, where a single FPGA is divided spatially among multiple users, are highly vulnerable to such attacks. This article examines the security concerns associated with multi-tenant cloud FPGAs, provides a comprehensive overview of the related security, privacy and trust issues, and discusses forthcoming challenges in this evolving field of study.
Muhammed Kawser Ahmed, Max Panoff, Joel Mandebi, Sujan Kumar Saha, Erman Nghonda, Peter Mbua, Christophe Bobda
ACM Trans. Reconfigurable Technol. Syst.3
2025 DA-VinCi: A Deep-Learning Accelerator Overlay Using In-Memory Computing
abstract
The matrix operations that underpin today’s deep learning models are routinely implemented in Single Instruction Multiple Data (SIMD) domain specific accelerators. SIMD accelerators including GPUs and array processors can effectively leverage parallelism in models that are compute-bound, but their effectiveness can be diminished for models that are memory-bound. Processing-in-Memory (PIM) architectures are being explored to provide better energy efficiency and scalable performance for these memory-bound models. Modern Field Programmable Gate Arrays (FPGAs) feature hundreds of megabits of Static Random Access Memory (SRAM) distributed across the device as disaggregated memory resources. This makes FPGAs ideal programmable platforms for developing custom Processor In/Near Memory accelerators. Several PIM array-based accelerator designs have been proposed to leverage this substantial internal bandwidth. However, results reported to date show the FPGA based PIM architectures operating at system clock frequencies well below a chips Block-RAM (BRAM) Fmax clock frequency. Results also show that the compute densities of the designs do not scale linearly with BRAM densities. These results indicate that FPGA PIM architectures will never be competitive with their custom Application-Specific Integrated Circuit (ASIC) counterparts. In this article, we introduce DA-VinCi, a D eep-Learning A ccelerator O v erlay using In -Memory C omput i ng. DA-VinCi is the first scalable FPGA based PIM deep-learning accelerator overlay capable of clocking at the maximum frequency of a device’s BRAM. Further, the architecture of DA-VinCi allows the number of compute units to scale linearly up to the maximum capacity of a devices BRAM, and at the maximum clock frequency of the BRAM. The DA-VinCi overlay has a programmable Instruction Set Architecture (ISA) that allows the same synthesized design to provide low-latency inferencing of a range of memory-bound deep-learning models, including Multilayer Perceptrons, Recurrent Neural Network, Long Short-Term Memory, and Gated Recurrent Unit networks. The scalability and high clocking frequency of DA-VinCi is achieved through a new Processor In Memory (PIM) tile architecture and a highly scalable system-level framework. We present results showing DA-VinCi linearly scaling the number of Processing Elements (PEs) to 100% of the BRAM capacity (over 60K PEs) on an Alveo U55 clocking at 737 MHz, the chips BRAM Fmax. We provide comparative studies on inference latency across multiple deep-learning applications that show DA-VinCi achieves up to a 201 \(\times\) improvement over a state-of-the-art PIM overlay accelerator, up to 87 \(\times\) improvement over existing PIM-based FPGA accelerators, and up to 57 \(\times\) improvement over custom deep-learning accelerators on FPGAs.
M. D. Arafat Kabir, Nathaniel Fredricks, Tendayi Kamucheka, Joel Mandebi, Miaoqing Huang, Jason D. Bakos, David Andrews 0001
ACM Trans. Reconfigurable Technol. Syst.4
2024 The BRAM is the Limit: Shattering Myths, Shaping Standards, and Building Scalable PIM Accelerators
abstract
Many recent FPGA-based Processor-in-Memory (PIM) architectures have appeared with promises of impressive levels of parallelism but with performance that falls short of expectations due to reduced maximum clock frequencies, an inability to scale processing elements up to the maximum BRAM capacity, and minimal hardware support for large reduction operations. In this paper, we propose a “Standard” set of design objectives for PIM array-based FPGA designs. We then propose a PIM array-based GEMV accelerator architecture as a case study to show the proposed Standard can be realized in practice. The GEMV accelerator serves as existence proof that dispels several myths surrounding what is normally accepted as clocking and scaling FPGA performance limitations. Specifically, the proposed accelerator clocks at the maximum frequency of the BRAM and scales to 100% of the available BRAMs. Comparative analyses show execution speeds over existing PIM-based GEMV engines on FPGAs and achieving a 2.65Χ – 3.2Χ faster clock. An AMD Alveo U55 implementation achieves a system clock speed of 737 MHz, providing 64K bit serial multiply-accumulate (MAC) units for GEMV operation.
M. D. Arafat Kabir, Tendayi Kamucheka, Nathaniel Fredricks, Joel Mandebi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001
FCCM4
2024 IMAGine: An In-Memory Accelerated GEMV Engine Overlay
abstract
Processor-in-Memory (PIM) overlays and alternative reconfigurable tile fabrics have been proposed to eliminate the von Neumann bottleneck and enable processing performance to scale with BRAM capacity. The performance of these FPGA-based PIM architectures has been limited due to a reduction of the BRAMs maximum clock frequencies and less than ideal scaling of processing elements with increased BRAM capacity. This paper presents IMAGine, an In-Memory Accelerated GEMV engine, a PIM-array accelerator that clocks at the maximum frequency of the BRAM and scales to 100% of the available BRAMs. Comparative analyses are presented showing execution speeds over existing PIM-based GEMV engines on FPGAs and achieving a $2.65 \times-3.2 \times$ faster clock. An AMD Alveo U55 implementation is presented that achieves a system clock speed of 737 MHz, providing 64 K bit-serial multiply-accumulate (MAC) units for GEMV operation. This establishes IMAGine as the fastest PIM-based GEMV overlay, outperforming even the custom PIM-based FPGA accelerators reported to date. Additionally, it surpasses TPU v1-v2 and Alibaba Hanguang 800 in clock speed while offering an equal or greater number of multiply-accumulate (MAC) units.
M. D. Arafat Kabir, Tendayi Kamucheka, Nathaniel Fredricks, Joel Mandebi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001
FPL4
2022 Accelerating Hybrid Quantized Neural Networks on Multi-tenant Cloud FPGA
abstract
The increasing adoption of Field-Programmable Gate Arrays (FPGA) into cloud and data center systems opens the way to the unprecedented acceleration of Machine Learning applications. Convolutional Neural Networks (CNN) have largely been adopted as algorithms for image classification and object detection. As we head towards FPGA multi-tenancy in the cloud, it becomes necessary to investigate architectures and mechanisms for the efficient deployment of CNN into multitenant FPGAs cloud Infrastructure. In this work, we propose an FPGA architecture and a design flow that support efficient integration of CNN applications into a cloud infrastructure that exposes multi-tenancy to cloud developers. We prototype the proposed approach on randomly allocated virtual regions to tenants. We study how space-sharing of a single device between multiple cloud tenants influence the design flow, the allocation of resources, and the performance in term of resource utilization and overall latency compared to single-tenant deployments. Prototyping results show a latency at most 8% lower than that of single-tenant deployment while achieving higher resource utilization. We also record a maximum frequency of up to 12% higher in multi-tenant implementations.
Danielle Tchuinkou, Erman Nghonda, Joel Mandebi, Christophe Bobda
ICCD3
2022 Towards a component-based acceleration of convolutional neural networks on FPGAs
Danielle Tchuinkou, Erman Nghonda, Joel Mandebi, Christophe Bobda
J. Parallel Distributed Comput.3
2022 The Future of FPGA Acceleration in Datacenters and the Cloud
abstract
In this article, we survey existing academic and commercial efforts to provide Field-Programmable Gate Array (FPGA) acceleration in datacenters and the cloud. The goal is a critical review of existing systems and a discussion of their evolution from single workstations with PCI-attached FPGAs in the early days of reconfigurable computing to the integration of FPGA farms in large-scale computing infrastructures. From the lessons learned, we discuss the future of FPGAs in datacenters and the cloud and assess the challenges likely to be encountered along the way. The article explores current architectures and discusses scalability and abstractions supported by operating systems, middleware, and virtualization. Hardware and software security becomes critical when infrastructure is shared among tenants with disparate backgrounds. We review the vulnerabilities of current systems and possible attack scenarios and discuss mitigation strategies, some of which impact FPGA architecture and technology. The viability of these architectures for popular applications is reviewed, with a particular focus on deep learning and scientific computing. This work draws from workshop discussions, panel sessions including the participation of experts in the reconfigurable computing field, and private discussions among these experts. These interactions have harmonized the terminology, taxonomy, and the important topics covered in this manuscript.
Christophe Bobda, Joel Mandebi, Paul Chow, Mohammad Ewais, Naif Tarafdar, Juan Camilo Vega, Kenneth Eguro, Dirk Koch, Suranga Handagala, Miriam Leeser, Martin C. Herbordt, Hafsah Shahzad, H. Peter Hofstee, Burkhard Ringlein, Jakub Szefer, Ahmed Sanaullah, Russell Tessier
ACM Trans. Reconfigurable Technol. Syst.2
2022 Deploying Multi-tenant FPGAs within Linux-based Cloud Infrastructure
abstract
Cloud deployments now increasingly exploit Field-Programmable Gate Array (FPGA) accelerators as part of virtual instances. While cloud FPGAs are still essentially single-tenant, the growing demand for efficient hardware acceleration paves the way to FPGA multi-tenancy. It then becomes necessary to explore architectures, design flows, and resource management features that aim at exposing multi-tenant FPGAs to the cloud users. In this article, we discuss a hardware/software architecture that supports provisioning space-shared FPGAs in Kernel-based Virtual Machine (KVM) clouds. The proposed hardware/software architecture introduces an FPGA organization that improves hardware consolidation and support hardware elasticity with minimal data movement overhead. It also relies on VirtIO to decrease communication latency between hardware and software domains. Prototyping the proposed architecture with a Virtex UltraScale+ FPGA demonstrated near specification maximum frequency for on-chip data movement and high throughput in virtual instance access to hardware accelerators. We demonstrate similar performance compared to single-tenant deployment while increasing FPGA utilization, which is one of the goals of virtualization. Overall, our FPGA design achieved about 2× higher maximum frequency than the state of the art and a bandwidth reaching up to 28 Gbps on 32-bit data width.
Joel Mandebi, Danielle Tchuinkou, Alex Shuping, Christophe Bobda
ACM Trans. Reconfigurable Technol. Syst.1
2021 ESCA: Event-Based Split-CNN Architecture with Data-Level Parallelism on UltraScale+ FPGA
abstract
This paper presents an event-based split-CNN architecture (ESCA) for running time-critical vision applications with comparatively less memory footprint while consuming low power. ESCA has a dedicated hardware architecture and scheduling of on-chip memory buffering using a split-CNN that reduces memory requirements by splitting the feature maps into small patches and independently executes them. The model emulates the concepts of the biological vision system to obtain possible events from each patch. We save energy and time by processing only the patches with possible events. The data-level parallelism is employed with a deep pipeline strategy that accelerates the system. We implement the design in the Virtex UltraScale+ FPGA at 320 MHz. Simulation results show that the architecture obtains significant speedup while power-saving depends on each image patch's features.
Pankaj Bhowmik, Md Jubaer Hossain Pantho, Joel Mandebi, Christophe Bobda
FCCM3
2021 A Customizable Domain-Specific Memory-Centric FPGA Overlay for Machine Learning Applications
abstract
This paper presents an overview and performance analysis of a software-programmable domain-customizable System-on-Chip (SoC) overlay for low-latency inferencing of variable and low-precision Machine Learning (ML) networks targeting Internet-of-Things (IoT) edge devices. The SoC includes a 2-D processor array that can be customized at design time for FPGA logic families. The overlay resolves historic issues of poor designer productivity associated with traditional Field Programmable Gate Array (FPGA) design flows without the performance losses normally incurred by overlays. A standard Instruction Set Architecture (ISA) allows different ML networks to be quickly compiled and run on the overlay without the need to resynthesize. Performance results are presented that show the overlay achieves $1.3\times-8.0\times$ speedup over custom designs while still allowing rapid changes to ML algorithms on the FPGA through standard compilation.
Atiyehsadat Panahi, Suhail Basalama, Ange-Thierry Ishimwe, Joel Mandebi, David Andrews 0001
FPL4
2021 Domain Isolation in FPGA-Accelerated Cloud and Data Center Applications
abstract
Cloud and data center applications increasingly leverage FPGAs because of their performance/watt benefits and flexibility advantages over traditional processing cores such as CPUs and GPUs. As the rising demand for hardware acceleration gradually leads to FPGA multi-tenancy in the cloud, there are rising concerns about the security challenges posed by FPGA virtualization. Exposing space-shared FPGAs to multiple cloud tenants may compromise the confidentiality, integrity, and availability of FPGA-accelerated applications. In this work, we present a hardware/software architecture for domain isolation in FPGA-accelerated clouds and data centers with a focus on software-based attacks aiming at unauthorized access and information leakage. Our proposed architecture implements Mandatory Access Control security policies from software down to the hardware accelerators on FPGA. Our experiments demonstrate that the proposed architecture protects against such attacks with minimal area and communication overhead.
Joel Mandebi, Sujan Kumar Saha, Christophe Bobda
ACM Great Lakes Symposium on VLSI1
2020 Architecture Support for FPGA Multi-tenancy in the Cloud
abstract
Cloud deployments now increasingly provision FPGA accelerators as part of virtual instances. While FPGAs are still essentially single-tenant, the growing demand for hardware acceleration will inevitably lead to the need for methods and architectures supporting FPGA multi-tenancy. In this paper, we propose an architecture supporting space-sharing of FPGA devices among multiple tenants in the cloud. The proposed architecture implements a network-on-chip (NoC) designed for fast data movement and low hardware footprint. Prototyping the proposed architecture on a Xilinx Virtex Ultrascale + demonstrated near specification maximum frequency for on-chip data movement and high throughput in virtual instance access to hardware accelerators. We demonstrate similar performance compared to single-tenant deployment while increasing FPGA utilization (we achieved $6 \times$ higher FPGA utilization with our case study), which is one of the major goals of virtualization. Overall, our NoC interconnect achieved about $2 \times$ higher maximum frequency than the state-of-the-art and a bandwidth of 25.6 Gbps.
Joel Mandebi, Alex Shuping, Pankaj Bhowmik, Christophe Bobda
ASAP1
2020 Accommodating Multi-Tenant FPGAs in the Cloud
abstract
This work presents a network-on-chip based architecture enabling multi-tenant access to FPGAs in cloud infrastructures. Prototyping the proposed architecture on Xilinx Virtex Ultrascale + and Intel Stratix IV demonstrated performance similar to single tenant mode while enabling hardware consolidation.
Joel Mandebi, Christophe Bobda
FCCM1
2018 FPGAVirt: A Novel Virtualization Framework for FPGAs in the Cloud
abstract
Field-Programmable Gate Arrays (FPGAs) are becoming important components within commercially available cloud computing systems. However, the FPGAs are not yet sufficiently abstracted within existing software ecosystems. Contrary to how applications are transparently scheduled across general purpose processors, software processes need to explicitly provision and control communications with hardware circuits within the FPGAs. In this paper, we introduce a novel virtualization framework called FPGAVirt that leverages Virtio to implement an efficient communication scheme between virtual machines and the FPGAs. FPGAVirt avoids the overhead of context switches between virtual machine and host address spaces by using the in-kernel network stack for transferring packets to FPGAs. Experimental results show FPGAVirt can deliver an additional 2x to 35x performance increase compared to current state of the art virtualization approaches.
Joel Mandebi, Festus Hategekimana, Danielle Tchuinkou, David Andrews 0001, Christophe Bobda
IEEE CLOUD1
2018 Inheriting Software Security Policies within Hardware IP Components
abstract
Domain isolation enforcement is one of the challenging issues in software environments. To address this problem, NSA, in conjunction with the Secure Computing Corporation and the University of Utah, developed the open-source Flux Advanced Security Kernel (Flask), the mandatory access control (MAC) security architecture underlying major Operating Systems/Hypervisors widely deployed in cloud/desktop environments. In this work, we extend this security architecture to FPGA-based heterogeneous systems. Specifically, we explore the design and implementation of a security framework for controlled sharing of FPGA hardware modules in MAC-based OS/Hypervisor environments. The proposed design guarantees that hardware modules execute in the same security context as of the processes calling them by propagating the latter security policies expressed at the software level, down to the hardware. We prototype the proposed framework with SELinux and demonstrate its utility by evaluating trade-offs between security performance and execution overhead incurred by example applications. The preliminary results show our proposed framework provides isolation with an average of 0.6% worst case performance overhead.
Festus Hategekimana, Joel Mandebi, Md Jubaer Hossain Pantho, Christophe Bobda
FCCM2
2018 Enabling Transparent Acceleration of OpenCV Library Kernels on a Hybrid Memory Cube Computer
abstract
This paper presents a CPU-FPGA heterogeneous system that provides hardware support for computer vision libraries to attain acceleration over image processing applications. The architecture achieves this improvement by staying completely transparent to the developer while providing necessary acceleration on conventional software designs.
Md Jubaer Hossain Pantho, Joel Mandebi, Christophe Bobda, David Andrews 0001, Marjan Asadinia
FCCM2
2018 Secure Hardware Kernels Execution in CPU+FPGA Heterogeneous Cloud
abstract
In this paper, we present a new security framework which allows controlled sharing and isolated execution of mutually distrusted FPGA-accelerators in heterogeneous cloud systems. The proposed framework enables the accelerators running in FPGAs in cloud computers to transparently inherit at run-time, software security policies of the virtual machines processes calling them. This capability allows system security policies enforcement mechanism to propagate access control privilege boundaries expressed at the hypervisor level, down to individual FPGA-accelerators. Furthermore, we present a software/hardware prototype implementation of the proposed security framework, showing that it can easily be transparently integrated within the virtual machine software stacks that run in today's cloud-based systems. Experimentation results show our proposed framework provides secure hardware execution with negligible execution overhead on guest VMs applications.
Festus Hategekimana, Joel Mandebi, Md Jubaer Hossain Pantho, Christophe Bobda
FPT2
2018 Transparent Acceleration of Image Processing Kernels on FPGA-Attached Hybrid Memory Cube Computers
abstract
The Hybrid Memory Cube (HMC) is representative of emerging architectures that integrate FPGAs with multichannel interconnected 3-D stacked memory, offering great potential for high bandwidth streaming applications. However, creating new hardware components that tap the full potential of the concurrent communications channels requires the structural understanding of the memory layout and interconnect configurations. In this paper, we present a new development framework aimed at removing the need for software programmers to understand the underlying physical architecture. The proposed framework automates the creation of hardware/software co-designs for computer vision applications in a transparent way to the developer. The development system dynamically detects function calls in software kernels and replaces those calls by a hardware wrapper function that exploits the HMCs memory hierarchy and multichannel interconnect with the FPGA. Results show our flow can exploit the 3-D stacked memory and concurrent communications channels to achieve speed-up with no need to tune the original software application to the memory hierarchy.
Md Jubaer Hossain Pantho, Joel Mandebi, Christophe Bobda, David Andrews 0001
FPT2
2018 FLexiTASK: A Flexible FPGA Overlay for Efficient Multitasking
abstract
One of the major obstacles to the adoption of FPGAs in high-performance computing is their programmability. It requires hardware design skills and long compilation times. Overlays have been proposed as a way to abstract FPGA resources. Unfortunately, most of the time, the topologies they use to connect computing cores impose restrictions on where tasks are placed and how they communicate. In this paper, we propose an overlay architecture designed for efficiency and flexibility. It features a novel Network-on-Chip (NoC) infrastructure making flexible, with no limitation, the placement of hardware tasks. The presented architecture allows tasks to communicate with a low latency and eases the reconfiguration of desired areas on the fabric at runtime. After prototyping the proposed architecture on an Altera Cyclone V FPGA, a maximum frequency of 282 MHz has been reached and a speedup ranging from 4x to 195x has been observed in some applications compared to the native execution.
Joel Mandebi, Danielle Tchuinkou, Christophe Bobda
ACM Great Lakes Symposium on VLSI1
2018 FPGA Virtualization in Cloud-Based Infrastructures Over Virtio
abstract
In this paper, we introduce a novel framework for virtualizing FPGA resources in the cloud. The proposed framework targets hardware/software architectures that leverage the Virtio paradigm for efficient communication between virtual machines (VMs) and the FPGAs. Furthermore, we present an FPGA overlay that uses reconfigurable hardware tiles and a flexible network-on-chip (NoC) architecture for transparent and optimized allocation of FPGA resources to VMs. The proposed overlay makes it possible to merge several FPGA regions allocated to a VM into a larger area, thus allowing resizing of FPGA's resources on demand. Hardware sandboxes are then provided as a means to enforce domain separation between hardware tasks belonging to different VMs. The framework introduced prevents the overhead of context switches between the virtual machine and host address spaces by using the in-kernel network stack for transferring packets to FPGAs. Experimental results show a 2x to 35x performance increase compared to current state of the art virtualization approaches.
Joel Mandebi, Festus Hategekimana, Danielle Tchuinkou, Christophe Bobda
ICCD1