Mariam Rakka

dblp:282/9705 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-2514-7960ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 SoftmAP: Software-Hardware Co-Design for Integer-Only Softmax on Associative Processors
abstract
Recent research efforts focus on reducing the computational and memory overheads of Large Language Models (LLMs) to make them feasible on resource-constrained devices. Despite advancements in compression techniques, nonlinear operators like Softmax and Layernorm remain bottlenecks due to their sensitivity to quantization. We propose SoftmAP, a software-hardware co-design methodology that implements an integer-only low-precision Softmax using In-Memory Compute (IMC) hardware. Our method achieves up to three orders of magnitude improvement in the energy-delay product compared to A100 and RTX3090 GPUs, making LLMs more deployable without compromising performance.
Mariam Rakka, Jinhao Li 0006, Guohao Dai 0001, Ahmed M. Eltawil, Mohamed E. Fouda, Fadi J. Kurdahi
DATE1
2024 HDRLPIM: A Simulator for Hyper-Dimensional Reinforcement Learning Based on Processing In-Memory
abstract
Processing In-Memory (PIM) is a data-centric computation paradigm that performs computations inside the memory, hence eliminating the memory wall problem in traditional computational paradigms used in Von-Neumann architectures. The associative processor, a type of PIM architecture, allows performing parallel and energy-efficient operations on vectors. This architecture is found useful in vector-based applications such as Hyper-Dimensional (HDC) Reinforcement Learning (RL). HDC is rising as a new powerful and lightweight alternative to costly traditional RL models such as Deep Q-Learning. The HDC implementation of Q-Learning relies on encoding the states in a high-dimensional representation where calculating Q-values and finding the maximum one can be done entirely in parallel. In this article, we propose to implement the main operations of a HDC RL framework on the associative processor. This acceleration achieves up to \(152.3\times\) and \(6.4\times\) energy and time savings compared to an FPGA implementation. Moreover, HDRLPIM shows that an SRAM-based AP implementation promises up to \(968.2\times\) energy-delay product gains compared to the FPGA implementation.
Mariam Rakka, Walaa Amer, Hanning Chen, Mohsen Imani, Fadi J. Kurdahi
ACM J. Emerg. Technol. Comput. Syst.1
2024 A Review of State-of-the-art Mixed-Precision Neural Network Frameworks
abstract
Mixed-precision Deep Neural Networks (DNNs) provide an efficient solution for hardware deployment, especially under resource constraints, while maintaining model accuracy. Identifying the ideal bit precision for each layer, however, remains a challenge given the vast array of models, datasets, and quantization schemes, leading to an expansive search space. Recent literature has addressed this challenge, resulting in several promising frameworks. This paper offers a comprehensive overview of the standard quantization classifications prevalent in existing studies. A detailed survey of current mixed-precision frameworks is provided, with an in-depth comparative analysis highlighting their respective merits and limitations. The paper concludes with insights into potential avenues for future research in this domain.
Mariam Rakka, Mohamed E. Fouda, Pramod P. Khargonekar, Fadi J. Kurdahi
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Information Processing Factory 2.0 - Self-awareness for Autonomous Collaborative Systems
abstract
This paper summarizes the talks of a special session on the IPF 2.0 project, a collaborative German-US research project that leverages self-awareness principles for the self-management of distributed systems of autonomous multiprocessor systems-on-chip (MPSoCs).
Nora Sperling, Alex Bendrick, Dominik Stöhrmann, Rolf Ernst, Bryan Donyanavard, Florian Maurer 0003, Oliver Lenke, Anmol Surhonne, Andreas Herkersdorf, Walaa Amer, Caio Batista de Melo, Ping-Xiang Chen, Quang Anh Hoang, Rachid Karami, Biswadip Maity, Paul Nikolian, Mariam Rakka, Dongjoo Seo, Saehanseul Yi, Minjun Seo, Nikil Dutt, Fadi J. Kurdahi
DATE17
2023 Hardware Implementation and Evaluation of an Information Processing Factory
abstract
The Information Processing Factory (IPF) utilizes factory management principles to tackle the complexities of integrated embedded systems, ensuring continuous safe operation and optimization at runtime. This paper presents a hardware implementation of IPF that enables dynamic task migration across system resources, ensuring reliability in the face of internal or external failures. We demonstrate the effectiveness of IPF through the efficient migration of tasks in multiprocessor SoCs using a safety-critical pacemaker application as a case study. Despite the additional software and hardware requirements, implementing IPF in a pacemaker results in comparable reliability to dual modular redundancy (DMR) with faster service resumption and improved resource utilization.
Walaa Amer, Mariam Rakka, Rachid Karami, Minjun Seo, Mazen A. R. Saghir, Rouwaida Kanj, Fadi J. Kurdahi
VLSI-SoC2
2021 Importance Splitting Sample Point Reuse for Efficient Memory Yield Estimation
abstract
In this paper, we propose and evaluate efficient sample point reuse methodologies for Importance Splitting- based yield estimation of memory designs to assess the impact of manufacturing variability induced design center shifts. The proposed methodologies enable yield estimation at almost no additional simulation cost to that incurred due to studying the yield at the nominal design center. In order to unbias the Importance Splitting conditional probabilities with respect to the manufacturing variability centers, we evaluate three techniques that rely on Importance Sampling, center projections and geometric ratios. The geometric-based approach is shown to be most accurate for center shifts up-to three sigma away from the nominal. The proposed reuse methodologies achieve up-to 4 orders of magnitude speedup for the analysis of a 132-D 16nm SRAM design compared to traditional Importance Splitting.
Mariam Rakka, Rouwaida Kanj
ISCAS1
2020 Hybrid Importance Splitting Importance Sampling Methodology for Fast Yield Analysis of Memory Designs
abstract
Rare fail event estimation methodologies suffer from inefficiency when dealing with high dimensional design space problems. Importance splitting overcomes this complexity by recursively computing the rare fail probability as a product of larger conditional probabilities in the 1-D performance metric space. Its efficiency, however, drops as the events become rarer. In this work, we propose a novel hybrid Importance Sampling Importance Splitting methodology for purposes of rare fail event estimation of high-dimensional memory designs. In this context, we propose and evaluate two methods for unbiasing the estimate, a geometric ratio-based and an Importance Sampling-based methodology. We demonstrate 3-5X reduction in runtime for both theoretical and 16nm SRAM design applications compared to traditional Importance Splitting approaches.
Mariam Rakka, Rouwaida Kanj, Ragheb Raad
ISCAS1