Danielle Tchuinkou

dblp:217/9394 · also Danielle Tchuinkou Kwadjo · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
4since 2021 · last 2022
0000-0001-5487-6049ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2022 Accelerating Hybrid Quantized Neural Networks on Multi-tenant Cloud FPGA
abstract
The increasing adoption of Field-Programmable Gate Arrays (FPGA) into cloud and data center systems opens the way to the unprecedented acceleration of Machine Learning applications. Convolutional Neural Networks (CNN) have largely been adopted as algorithms for image classification and object detection. As we head towards FPGA multi-tenancy in the cloud, it becomes necessary to investigate architectures and mechanisms for the efficient deployment of CNN into multitenant FPGAs cloud Infrastructure. In this work, we propose an FPGA architecture and a design flow that support efficient integration of CNN applications into a cloud infrastructure that exposes multi-tenancy to cloud developers. We prototype the proposed approach on randomly allocated virtual regions to tenants. We study how space-sharing of a single device between multiple cloud tenants influence the design flow, the allocation of resources, and the performance in term of resource utilization and overall latency compared to single-tenant deployments. Prototyping results show a latency at most 8% lower than that of single-tenant deployment while achieving higher resource utilization. We also record a maximum frequency of up to 12% higher in multi-tenant implementations.
Danielle Tchuinkou, Erman Nghonda, Joel Mandebi, Christophe Bobda
ICCD1
2022 Coarse-Grained Floorplanning for streaming CNN applications on Multi-Die FPGAs
abstract
With the vast adoption of FPGAs in the cloud, it becomes necessary to investigate architectures and mechanisms for the efficient deployment of CNN into multi-FPGAs cloud Infrastructure. However, neural networks’ growing size and complexity, coupled with communication and off-chip memory bottlenecks, make it increasingly difficult for multi-FPGA designs to achieve high resource utilization. In this work, we introduce a scalable framework that supports the efficient integration of CNN applications into a cloud infrastructure that exposes multi-Die FPGAs to cloud developers. Our framework is equipped is with two mechanisms to facilitate the deployment of CNN inference on FPGA. First, we propose a model to find the parameters that maximize the parallelism within the resource budget while maintaining a balanced rate between the layers. Then, we propose an efficient Coarse-Grained graph partitioning algorithm for high-quality and scalable routability-drive placement of CNN’s components on the FPGAs. Prototyping results achieve an overall 37% higher frequency, with lower resource usage compared to a baseline implementation on the same number of FPGAs.
Danielle Tchuinkou, Erman Nghonda, Christophe Bobda
ISPDC1
2022 Towards a component-based acceleration of convolutional neural networks on FPGAs
Danielle Tchuinkou, Erman Nghonda, Joel Mandebi, Christophe Bobda
J. Parallel Distributed Comput.1
2022 Deploying Multi-tenant FPGAs within Linux-based Cloud Infrastructure
abstract
Cloud deployments now increasingly exploit Field-Programmable Gate Array (FPGA) accelerators as part of virtual instances. While cloud FPGAs are still essentially single-tenant, the growing demand for efficient hardware acceleration paves the way to FPGA multi-tenancy. It then becomes necessary to explore architectures, design flows, and resource management features that aim at exposing multi-tenant FPGAs to the cloud users. In this article, we discuss a hardware/software architecture that supports provisioning space-shared FPGAs in Kernel-based Virtual Machine (KVM) clouds. The proposed hardware/software architecture introduces an FPGA organization that improves hardware consolidation and support hardware elasticity with minimal data movement overhead. It also relies on VirtIO to decrease communication latency between hardware and software domains. Prototyping the proposed architecture with a Virtex UltraScale+ FPGA demonstrated near specification maximum frequency for on-chip data movement and high throughput in virtual instance access to hardware accelerators. We demonstrate similar performance compared to single-tenant deployment while increasing FPGA utilization, which is one of the goals of virtualization. Overall, our FPGA design achieved about 2× higher maximum frequency than the state of the art and a bandwidth reaching up to 28 Gbps on 32-bit data width.
Joel Mandebi, Danielle Tchuinkou, Alex Shuping, Christophe Bobda
ACM Trans. Reconfigurable Technol. Syst.2
2020 Late Breaking Results: Automated Hardware Generation of CNN Models on FPGAs
abstract
In this paper, we propose an automated framework that takes as input a TensorFlow inference graph and generates high-performance accelerators on FPGA by assembling CNN pre-implemented components as a puzzle, based on the graph topology. Using pre-implemented components allows us the only use the minimum of resources necessary, predict the performance and a gain in productivity We adopt a unified representation based on systolic array to perform the computational-hungry operations of the model and provide novel analysis of design trade-offs for FPGA CNN accelerators. Experimental results show the great performance, low latency and flexibility provided by the proposed framework.
Danielle Tchuinkou, Christophe Bobda
DAC1
2018 FPGAVirt: A Novel Virtualization Framework for FPGAs in the Cloud
abstract
Field-Programmable Gate Arrays (FPGAs) are becoming important components within commercially available cloud computing systems. However, the FPGAs are not yet sufficiently abstracted within existing software ecosystems. Contrary to how applications are transparently scheduled across general purpose processors, software processes need to explicitly provision and control communications with hardware circuits within the FPGAs. In this paper, we introduce a novel virtualization framework called FPGAVirt that leverages Virtio to implement an efficient communication scheme between virtual machines and the FPGAs. FPGAVirt avoids the overhead of context switches between virtual machine and host address spaces by using the in-kernel network stack for transferring packets to FPGAs. Experimental results show FPGAVirt can deliver an additional 2x to 35x performance increase compared to current state of the art virtualization approaches.
Joel Mandebi, Festus Hategekimana, Danielle Tchuinkou, David Andrews 0001, Christophe Bobda
IEEE CLOUD3
2018 FLexiTASK: A Flexible FPGA Overlay for Efficient Multitasking
abstract
One of the major obstacles to the adoption of FPGAs in high-performance computing is their programmability. It requires hardware design skills and long compilation times. Overlays have been proposed as a way to abstract FPGA resources. Unfortunately, most of the time, the topologies they use to connect computing cores impose restrictions on where tasks are placed and how they communicate. In this paper, we propose an overlay architecture designed for efficiency and flexibility. It features a novel Network-on-Chip (NoC) infrastructure making flexible, with no limitation, the placement of hardware tasks. The presented architecture allows tasks to communicate with a low latency and eases the reconfiguration of desired areas on the fabric at runtime. After prototyping the proposed architecture on an Altera Cyclone V FPGA, a maximum frequency of 282 MHz has been reached and a speedup ranging from 4x to 195x has been observed in some applications compared to the native execution.
Joel Mandebi, Danielle Tchuinkou, Christophe Bobda
ACM Great Lakes Symposium on VLSI2
2018 FPGA Virtualization in Cloud-Based Infrastructures Over Virtio
abstract
In this paper, we introduce a novel framework for virtualizing FPGA resources in the cloud. The proposed framework targets hardware/software architectures that leverage the Virtio paradigm for efficient communication between virtual machines (VMs) and the FPGAs. Furthermore, we present an FPGA overlay that uses reconfigurable hardware tiles and a flexible network-on-chip (NoC) architecture for transparent and optimized allocation of FPGA resources to VMs. The proposed overlay makes it possible to merge several FPGA regions allocated to a VM into a larger area, thus allowing resizing of FPGA's resources on demand. Hardware sandboxes are then provided as a means to enforce domain separation between hardware tasks belonging to different VMs. The framework introduced prevents the overhead of context switches between the virtual machine and host address spaces by using the in-kernel network stack for transferring packets to FPGAs. Experimental results show a 2x to 35x performance increase compared to current state of the art virtualization approaches.
Joel Mandebi, Festus Hategekimana, Danielle Tchuinkou, Christophe Bobda
ICCD3
2018 R-Covnet: Recurrent Neural Convolution Network for 3D Object Recognition
abstract
Pointcloud is a very precise digital format for recording objects in space. Pointclouds have received increasing attention lately, due to the higher amount of information it provides compared to images. In this paper, we propose a new deep learning architecture called R-CovNet, designed for 3D object recognition. Unlike previous architectures that usually sample or convert pointcloud into three-dimensional grids before processing, R-CovNet does not require any preprocessing. Our main goal is to provide a permutation invariant architecture specially designed for pointclouds data of any size. Experiments with well-known benchmarks show that R-CovNet can achieve an accuracy of 92.7%, thus outperforming all the volumetric methods.
Danielle Tchuinkou, Christophe Bobda
ICIP1