Albi Mema

dblp:343/4969 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0001-7841-1975ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Ferroelectric Digital In-Memory Computing for Scalable, Reliable, and Efficient Similarity Computation
abstract
Classification-based learning in deep neural networks, particularly few-shot learning, demands efficient similarity metrics such as Hamming distance. Conventional architectures suffer from high energy overheads due to frequent data movement between memory and processing units, hindering scalability. In-memory computing addresses this by integrating computation within memory, yet analog-based systems rely on power-hungry analog-to-digital converters (ADCs) and face scalability challenges due to device variability, especially in emerging memories. This work presents a fully digital Ferroelectric FET (FeFET)-based Logic-in-Memory (LiM) XOR cell, designed using GlobalFoundries’ 28 nm technology, eliminating ADCs and ensuring robust, energy-efficient, and scalable operation. Our 2T FeFET XOR cell, applied to 4096-bit Hamming distance calculations, achieves$23\times $lower energy,$3\times $faster latency, and$14\times $area reduction over state-of-the-art designs. Delivering 2337 Gsamples/(s$\cdot $W$\cdot $mm2) — a$300\times $improvement — this architecture offers a compelling solution for energy-efficient, reliable, and scalable AI hardware, driving sustainable computing.
Anirban Kar, Albi Mema, Thorgund Nemec, Stefan Dünkel, Halid Mulaosmanovic, Sven Beyer, Yogesh Singh Chauhan, Hussam Amrouch
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 Transistor-to-GDS Reliability Analysis in Sub-3nm: Impact of Self-Heating and Aging on Timing
abstract
As transistor scaling advances into the sub-3 nm regime, self-heating effects (SHE) and aging-induced degradation emerge as profound challenges that threaten timing closure, signal integrity, guardbands, and long-term reliability. This work presents a comprehensive transistor-to-GDS reliability analysis that captures the impact of SHE and aging in nanosheet field-effect transistors (NSFETs) and propagates it through the entire design stack to full-chip signoff. We evaluate a 64-bit RISC-V processor core and an AI accelerator containing 4096 Multiply-and-Accumulate (MAC) units, both implemented using gate-all-around (GAA) NSFET technology. TCAD simulations, carefully calibrated against measurement data, reveal local temperature rises up to 124 K in multi-stack sheet structures, which exacerbate aging and result in a threshold voltage shift of up to 42.3 mV. Incorporating these effects into standard cell characterization and commercial signoff timing analysis uncovers substantial End-of-Life (EOL) timing degradation—37.7 % for the RISC-V core and 61.7 % for the AI accelerator—highlighting the urgent need for SHE- and aging-aware methodologies, as well as reliability-optimized standard cell libraries for advanced nodes.
Swati Deshwal, Hadi Nour Eddine, Mahdi Benkhelifa, Albi Mema, Yogesh Singh Chauhan, Hussam Amrouch
ISLPED4
2023 Reliable Hyperdimensional Reasoning on Unreliable Emerging Technologies
abstract
While Graph Neural Networks (GNNs) have demonstrated remarkable achievements in knowledge graph reasoning, their computational efficiency on conventional computing platforms is impeded by the memory wall problem. To overcome these challenges, we introduce an innovative algorithm-hardware solution that harnesses the potential of hyperdimensional computing (HDC) for robust and memory-centric computation on computing in-memory (CiM) platforms. Departing from traditional graph neural networks, the proposed HDC reasoning model employs a symbolic approach to effectively encode graph entities and their relationships as high-dimensional neural activity. Complementing this approach is a customized Computing-in-Memory (CiM) architecture based on advanced Ferroelectric Field-Effect Transistor (FeFET) technology, which incorporates a precise characterization of non-idealities. This modeling enables the generation of an HDC-tailored model that faithfully represents the hardware architecture. Despite the non-idealities inherent in emerging CiM technologies, our platform demonstrates performance on par with traditional von Neumann architectures for substantial combinations of FeFET device parameters. Our solution overcomes FeFET CiM the increased non-idealities from down-scaled 3nm, operating effectively under all possible configurations when 50 graph edges are considered. Scenarios with less than 4-bit precision per FeFET device cannot handle graphs with more than 200 edges, whereas the 4-bit case can achieve a 90.3% graph reconstruction rate on the worst-case scenario of 80% of noise.
Hamza Errahmouni Barkam, Sanggeon Yun, Hanning Chen, Paul Gensler, Albi Mema, Andrew Ding, George Michelogiannakis, Hussam Amrouch, Mohsen Imani
ICCAD5
2023 FDSOI-Based Analog Computing for Ultra-Efficient Hamming Distance Similarity Calculation
abstract
Computing the similarity between two binary strings is a frequently used operation in cryptography, machine learning, and other areas. The Hamming distance is a simple yet costly to compute similarity metric. A common way is to XOR both binary input strings and then count the number of 1s. Especially the latter popcount part is inefficient with purely digital circuits. In this paper, a novel analog circuit is proposed to compute the Hamming distance in an ultra-efficient way. Contrary to the major trend in the state of the art, no emerging technology is required. Instead, the unique feature of the mature FDSOI transistor technology is exploited for the first time to perform analog-based similarity calculation. Thanks to the additional back gate available in this technology, the transistor’s threshold voltage can be modulated by more than 1 V. Through this key feature, an ultra-efficient analog computing is realized, replacing the inefficient digital popcount traditionally built from expensive adder tree structures. The design is evaluated with an FDSOI transistor model calibrated with industrial measurements. The energy-delay product is at least 24$\times $smaller than purely digital implementations and the transistor count is reduced by over 2.6$\times $.
Albi Mema, Simon Thomann, Paul R. Genssler, Hussam Amrouch
IEEE Trans. Circuits Syst. I Regul. Pap.1