EDBT 2026 Demo / reviewers in the wild / expert
Omar Al Kailani
dblp:389/6997
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0002-6696-6089ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BCIM: A Bit-Serial Approach for Block-Cipher-in-MemoryabstractThis abstract presents Block-Cipher-In-Memory (BCIM), a constant-time, high-throughput, bit-serial in-memory cryptography scheme to support versatile block ciphers. Exploiting the nature of constant-time bit-serial execution and column-wise single instruction multiple data (SIMD) processing, BCIM achieves inherent resiliency to timing-based side-channel attacks and high-throughput encryption/decryption. Architectural and circuit level innovations ensure minimal memory footprint for storing repetitive keys, and ultimately improve the peak operating frequency and energy efficiency. Experimental results suggest that BCIM shows substantial performance and energy improvements over state-of-the-art bit-parallel IMC ciphers, as well as competitive performance and orders of magnitude energy advantages over the bitsliced software implementations on both CPU/MCU platforms. Andrew Dervay, Omar Al Kailani |
FCCM | 2 |
| 2025 | JSA-CIM: A Joint-Sparse AdderNet Compute-In-Memory Accelerator for Energy-Efficient Edge AI ApplicationsabstractAdderNet has emerged as a compelling alternative to the prevailing Convolutional Neural Networks, offering comparable model accuracy while significantly reducing the computational complexity of Multiply-Accumulate operations with efficient Sum-Of-Absolute-Differences alternatives. Despite being promising, current Compute-In-Memory (CIM) designs for AdderNet fail to leverage the inherent model sparsity, thereby overlooking substantial opportunities for enhancing energy efficiency in CIM-based implementations. In this paper, we propose JSA-CIM, an 8T-SRAM-based Digital CIM (DCIM) for efficient AdderNet inference, exploiting joint-sparsity of both activations and weights in a bit-serial fashion. Specifically, JSA-CIM features (a) at architecture level, a sparsity- and kernel-size-aware CIM macro that facilitates word-level activation sparsity exploitation for energy reduction, and accommodates the memory organization to fit popular kernel sizes for maximizing the application-level throughput, (b) at circuit-level, a novel joint-sparse bitline architecture based on 8T-SRAM that can exploit both word-level activation sparsity and bit-level weight sparsity to minimize the array energy consumption, and (c) a robust and sparsity-gated minimum selection circuit to achieve reliable and energy-efficient bit-serial minimum operation. When implemented in 65nm CMOS technology, the proposed JSA-CIM demonstrates a peak throughput of 170.7 GOPS, 1.8 TOPS/mm2area efficiency, and 75.98 TOPS/W energy efficiency. Omar Al Kailani |
ICCAD | 1 |
| 2025 | AdderNet 2.0: Optimal AdderNet Accelerator Designs With Activation-Oriented Quantization and Fused Bias Removal-Based Memory OptimizationabstractConvolutional neural networks (CNNs) are computationally demanding due to expensive Multiply-ACcumulate (MAC) operations. Emerging neural network models, such as AdderNet, exploit efficient arithmetic alternatives like sum-of-absolute-difference (SAD) operations to replace the costly MAC operations, while still achieving competitive model accuracy as compared with the CNN counterparts. Nevertheless, existing AdderNet accelerators still face critical implementation challenges to achieve maximal hardware and energy efficiency at the cost of model inference accuracy loss. This paper presents AdderNet 2.0, an algorithm-hardware co-design framework featuring a novel Activation-Oriented Quantization (AOQ) strategy, a Fused Bias Removal (FBR) scheme for on-chip feature map memory bitwidth reduction, and optimal PE designs to improve the overall resource utilization towards optimal AdderNet accelerator designs. Multiple AdderNet 2.0 accelerator design variants were implemented on Xilinx KV-260 FPGA. Experimental results show that the INT6 AdderNet 2.0 accelerators achieve significant hardware resource and energy savings when compared to prior CNN and AdderNet designs. Omar Al Kailani |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | AdderNet 2.0: Optimal FPGA Acceleration of AdderNet with Activation-Oriented Quantization and Fused Bias Removal based Memory OptimizationabstractConvolutional neural networks (CNNs) are computationally demanding due to expensive Multiply-ACcumulate (MAC) operations. Emerging neural network models, such as AdderNet, exploit efficient arithmetic alternatives like sum-of-absolute-difference (SAD) operations to replace the costly MAC operations, while still achieving competitive model accuracy as the CNN counterparts. Nevertheless, existing AdderNet accelerators still face critical implementation challenges to achieve maximal hardware and energy efficiency. This paper presents AdderNet 2.0, an algorithm-hardware co-design framework featuring a novel Activation-Oriented Quantization (AOQ) strategy, a Fused Bias Removal (FBR) scheme for on-chip feature memory bitwidth reduction, and optimal PE designs to improve the overall resource utilization towards optimal AdderNet accelerator designs. Multiple AdderNet 2.0 accelerator design variants were implemented on Xilinx KV-260 FPGA. Experimental results show that the INT6 AdderNet 2.0 accelerators achieve significant hardware resource and energy savings when compared to prior CNN and AdderNet designs. Omar Al Kailani |
DAC | 2 |