Najdet Charaf

dblp:295/8163 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2024
0000-0001-5005-0928ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-author · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2024 DA-CGRA: Domain-Aware Heterogeneous Coarse-Grained Reconfigurable Architecture for the Edge
abstract
Coarse-Grained Reconfigurable Architectures (CGRAs) are one of the promising solutions to be employed in power-hungry edge devices owing to providing a good balance between reconfigurability, performance and energy-efficiency. Most of the proposed CGRAs feature a homogeneous set of processing elements (PEs) which all support the same set of operations. Homogeneous PEs can lead to high unwanted power consumption. As application benchmarks utilize different operations irregularly, heterogeneous PE design is a powerful approach to reduce power consumption of CGRA. In this paper, we propose DA-CGRA, a domain-aware CGRA tailored to signal processing applications. To extract heterogeneous architecture, first, a set of signal processing applications has been profiled to derive the requirements of the applications in terms of type of operations, number of operations and memory usage. Then, domain-specific PEs are designed using Verilog RTL based on the profiling results. We have selected spatio-temporal or spatial execution model based on the application features to increase the overall performance and efficiency. Experimental results demonstrate DA-CGRA outperforms FLEX and RipTide state-of-the-art CGRAs in terms of energy-efficiency by 23% and 38%, respectively. Moreover, DA-CGRA can achieve 3.2x performance improvement over HM-HvCUBE.
Ensieh Aliagha, Najdet Charaf, Nitin Krishna Venkatesan, Diana Göhringer
DSD2
2024 NC-Library: Expanding SystemC Capabilities for Nested reConfigurable Hardware Modelling
abstract
As runtime reconfiguration is used in an increasing number of hardware architectures, new simulation and modeling tools are needed to support the developer during the design phases. In this article, a language extension for SystemC is presented, together with a design methodology for the description and simulation of dynamically reconfigurable hardware at different levels of abstraction. The library presented offers a high degree of flexibility in the description of reconfiguration features and their management, while allowing runtime reconfiguration simulation, removal, and replacement of custom modules as well as third-party components throughout the architecture development process. In addition, our approach supports the emerging concept of nested reconfiguration and split regions with a minimal simulation overhead of a maximum of three delta cycles for signal and transaction forwarding, and four delta cycles for the reconfiguration process.
Julian Haase, Najdet Charaf, Alexander Groß 0002, Diana Göhringer
ACM Trans. Reconfigurable Technol. Syst.2
2023 EuFRATE: European FPGA Radiation-hardened Architecture for Telecommunications
abstract
The EuFRATE project aims to research, develop and test radiation-hardening methods for telecommunication payloads deployed for Geostationary-Earth Orbit (GEO) using Commercial-Off- The-Shelf Field Programmable Gate Arrays (FPGAs). This project is conducted by Argotec Group (Italy) with the collaboration of two partners: Politecnico di Torino (Italy) and Technische Universität Dresden (Germany). The idea of the project focuses on high-performance telecommunication algorithms and the design and implementation strategies for connecting an FPGA device into a robust and efficient cluster of multi-FPGA systems. The radiation-hardening techniques currently under development are addressing both device and cluster levels, with redundant datapaths on multiple devices, comparing the results and isolating fatal errors. This paper introduces the current state of the project's hardware design description, the composition of the FPGA cluster node, the proposed cluster topology, and the radiation hardening techniques. Intermediate stage experimental results of the FPGA communication layer performance and fault detection techniques are presented. Finally, a wide summary of the project's impact on the scientific community is provided.1
Ludovica Bozzoli, Antonino Catanese, Emilio Fazzoletto, Eugenio Scarpa, Diana Göhringer, Sergio A. Pertuz 0001, Lester Kalms, Cornelia Wulf, Najdet Charaf, Luca Sterpone, Sarah Azimi, Daniele Rizzieri, Salvatore Gabriele La Greca, David Merodio Codinachs
DATE9
2023 RTASS: a RunTime Adaptable and Scalable System for Network-on-Chip-Based Architectures
abstract
In an ever-evolving digital world with complex algorithms like machine learning, we need new strategies for more flexibility to cope with the ever-changing environment. For this, runtime scalability and runtime adaptability for low-power and highly efficient hardware is a promising solution. By combining the runtime reconfiguration of FPGAs with the efficient communication of Networks-on-Chip (NoC), we are able to implement a highly scalable, high-performance, and energy-efficient computing architecture that fixed-function units and specialized static accelerators lack. In this work, we introduce a RunTime Adaptable and Scalable System for NoC-based architectures called RTASS. The hardware architecture includes a master subsystem, a network adapter, and an NoC subsystem with parametrizable routers and several various routing algorithms. Furthermore, RTASS provides a software architecture that includes advanced drivers for runtime management. The key benefit of RTASS is the ability to dynamically adjust the number of routers within the NoC at runtime based on the current application's requirements. That allows the system to support both homogeneous and inhomogeneous types of processing elements as well as regular and irregular shapes. The development of this runtime scalable and flexible architecture will establish the foundation for future highly adaptable applications such as machine learning and computer vision in the embedded computing field. We implemented and evaluated the proposed work with the Xilinx Zynq-7000 FPGA, with the possibility of porting it to other FPGAs that support runtime reconfiguration.
Najdet Charaf, Julian Haase, Adrian Kulisch, Christian von Elm, Diana Göhringer
DSD1
2022 MaNaBIT: A Versatile Tool for Manipulating and Analyzing FPGA Bitstreams
abstract
The ability of the reconfigurable systems to provide flexible and high-performance hardware has contributed to the fact that their popularity and multifaceted usage increased enormously in recent years [1] . They are occupying a central position in our modern complex systems. The development of Field Programmable Gate Arrays (FPGA) has taken hardware flexibility, in general, one step further. In recent years, many approaches have been developed that exploit dynamic reconfigurability of FPGAs, especially Xilinx FPGAs. Dynamic partial reconfiguration (DPR) and especially relocation are well-established and promising techniques in this area. In this context, the use of partial reconfiguration to add the adaptability feature to the design makes system design even more complex [2] . Therefore, solutions that help to reduce time and design efforts are needed. One of these solutions is bitstream manipulation. This approach leads to the user being able to perform modifications at runtime, thus reducing design time significantly.
Najdet Charaf, Christoph Tietz, Diana Göhringer
FCCM1
2022 Scheduling of Hardware Tasks in Reconfigurable Mixed-Criticality Systems
abstract
FPGA virtualization allows the shared usage of an FPGA by several operating systems with different criticality levels. To avoid mutual interference, most state-of-the-art systems strictly isolate subsystems in spatial respect at the expense of lower resource utilization. We present an allocation and scheduling strategy for hardware tasks that improves resource utilization while respecting different real-time levels (hard, soft, and no real-time) of guest operating systems. To not jeopardize deadlines, Dynamic Partial Reconfiguration (DPR) latencies are reduced by reusing, prefetching and reserving of hardware accelerators. Compared with an existing scheduler for hardware tasks, we could increase the resource usage by 156% while deadline misses were reduced by 6%.
Cornelia Wulf, Najdet Charaf, Diana Göhringer
FCCM2
2022 A Framework for Intrinsic Evolvable Systems
abstract
Systems with hardware that can dynamically and autonomously change their architecture and behavior by interacting with their environment are becoming very valuable in modern applications. Therefore, research and development in intrinsically evolvable embedded systems are becoming increasingly attractive. Runtime reconfiguration and relocation are a promising approach for designing self-adaptive and self-optimizing autonomous embedded systems. The vision behind this PhD work is to provide an all-encompassing framework to automate all the challenging tasks required for designing self-adaptive systems. This paper presents our framework and preliminary results and highlights our next steps and future work.
Najdet Charaf, Diana Göhringer
FPL1
2022 Virtualization of Reconfigurable Mixed-Criticality Systems
abstract
The increasing complexity of reconfigurable embedded systems often requires the integration of multiple applications with potentially different levels of criticality on the same hardware platform. As the deployment scales, there is a need for resource management, isolation, and performance that makes FPGA virtualization techniques a key consideration. FPGA virtualization enables multiple guest operating systems to run with different requirements, such as real-time, safety, or security. Most state-of-the-art systems incorporate mechanisms to strictly isolate subsystems in spatial respect at the expense of lower resource utilization. In this work, we present L4ReC, a microkernel-based virtualization layer that enables the sharing of reconfigurable resources among multiple virtual machines. The mapping and scheduling strategy for hardware threads considers not only deadlines, but also the real-time levels of guest operating systems. A POSIX thread-based interface facilitates the access to hardware accelerators. Compared with an existing scheduler for hardware threads, the average utilization factor - indicating the FPGA resource usage - is 1,9 times higher when threads are mapped and scheduled with L4ReC. Deadline misses are reduced by 3%.
Cornelia Wulf, Najdet Charaf, Diana Göhringer
FPL2
2021 AMAH-Flex: A Modular and Highly Flexible Tool for Generating Relocatable Systems on FPGAs
abstract
In this work, we present a solution to a common problem encountered when using FPGAs in dynamic, ever-changing environments. Even when using dynamic function exchange to accommodate changing workloads, partial bitstreams are typically not relocatable. So the runtime environment needs to store all reconfigurable partition/reconfigurable module combinations as separate bitstreams. We present a modular and highly flexible tool (AMAH-Flex) that converts any static and reconfigurable system into a 2 dimensional dynamically relocatable system. It also features a fully automated floorplanning phase, closing the automation gap between synthesis and bitstream relocation. It integrates with the Xilinx Vivado toolchain and supports both FPGA architectures, the 7-Series and the UltraScale+. In addition, AMAH-Flex can be ported to any Xilinx FPGA family, starting with the 7-Series. We demonstrate the functionality of our tool in several reconfiguration scenarios on four different FPGA families and show that AMAH-Flex saves up to 80% of partial bitstreams.
Najdet Charaf, Christoph Tietz, Michael Raitza, Akash Kumar 0001, Diana Göhringer
FPT1