Baker Mohammad

dblp:54/3227 · also Baker S. Mohammad · DBLP profile ↗
← Back
48ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0002-6063-473XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 40 · 4 first-author · 11 since 2021Security and privacy · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient On-the-Fly Twiddle Factor Generation for Falcon PQC NTT/INTT
Ghada Alsuhli, Hani Saleh, Mahmoud Al-Qutayri, Baker Mohammad, Thanos Stouraitis
ISCAS4
2026 Reconfigurable SRAM-CAM for Similarity Index Computation in 22-nm FDSOI
Eman Hassan, K. M. Shadid Hassan, Shaymaa Elshahaby, Mahmoud Al-Qutayri, Baker Mohammad
ISCAS5
2025 Enhanced CNN Performance without Retraining Via Weight Approximation and Data Reuse
abstract
This paper introduces an efficient CNN algorithm to address key limitations in Deep Neural Networks (DNNs) used for image recognition, focusing particularly on model size and retraining time. Traditional methods often require significant training durations; however, applying approximation techniques during retraining can exacerbate these time demands. We present an approach that enhances approximation techniques while eliminating the need for model retraining, thus enabling DNN compression with minimal accuracy loss. The proposed method integrates three core strategies: weight arrangement, approximation, and data reuse. The DNN weights are initially arranged in ascending order to optimize subsequent operations. During inference, the approximation is applied to reduce the model size and minimize computational complexity by reducing the number of operations required for each multiply-accumulate (MAC) unit. Then, the original weights are replaced with the approximated values, enabling the reuse of computations and data across different sets of weights. As a result, the method significantly reduces memory access, computational demands, and energy consumption. Experimental results on the CIFAR-10 and TinyImageNet datasets demonstrate that our method achieves a model reduction rate of approximately 198.6× while maintaining a minimal loss in accuracy. The proposed technique bypasses the need for retraining, offering a practical solution to the growing complexity of DNN models in modern applications.
Mohamed F. Tolba 0002, Hani Saleh, Baker Mohammad, Mahmoud Al-Qutayri, Thanos Stouraitis
ISCAS3
2025 RTSSNN: Efficient Image Classification for Latency-Critical and Energy-Constrained SNNs Through a Time-Step Reduction Technique
abstract
Spiking Neural Networks (SNNs) are emerging as potent alternatives to Convolutional Neural Networks (CNNs), especially for energy-constrained and latency-sensitive applications, due to their spiked activations and inherent sparsity. Rate-coded SNNs, trained with backpropagation through time (BPTT) and using static data over artificial time-steps, have achieved state-of-the-art results on benchmarks like MNIST. For resource-constrained devices, smaller models help meet memory and power limitations. However, shallow rate-coded SNNs often need numerous time-steps for accurate inference, increasing latency and computational cost. To address this, a technique named RTS (Reduced Time-Step) is proposed to reduce time-steps, optimizing the balance between model size and convergence latency. RTSSNN leverages the periodic dynamics of Leaky-Integrate-and-Fire (LIF) neurons’ membrane potentials when stimulated with constant input by adding a small Fully Connected (FC) layer at the end of the network. At this depth, spikes are sparse and stable, allowing reduced time-steps without losing accuracy. Demonstrated on four-bit quantized SNNs on Raspberry Pi, the method achieves 9x, 4x, and 4.9x operations reduction, 3x, 2x, and 2.2x time-step reduction, as well as 2.3x, 1.5x, and 1.9x runtime reduction during inference on MNIST, FashionMNIST, and GTSRB datasets, respectively, with maintained accuracy. It also shows a 4x, 2x, and 1.9x operations, time-step, and runtime reduction in one-bit quantized SNNs. Additional experiments conducted on the higher complexity CIFAR-10 dataset as well as the dynamic neuromorphic N-MNIST dataset confirmed that RTSSNN is effective primarily on shallow networks with static datasets, where stable activations support periodic membrane behavior in LIF neurons.
Nada Abu Hamra, Baker Mohammad, Mahmoud Al-Qutayri
IEEE Internet Things J.2
2025 Efficient NTT/INTT processor for FALCON post-quantum cryptography
abstract
FALCON is a lattice-based post-quantum cryptographic (PQC) digital signature standard known for its compact signatures and resistance to quantum attacks. Since its recent standardization, its hardware implementation remains an open challenge, particularly for key generation, which is significantly more complex than the simple and well-studied signature verification process. In this paper, targeting edge devices with constrained resources, we present an energy-efficient and area-optimized NTT/INTT architecture tailored to the specific requirements of FALCON key generation. By leveraging NTT-friendly primes and reducing the size of the multipliers in the Montgomery reduction algorithm — optimized for ASIC implementation — our design minimizes hardware complexity, achieving the lowest power and area consumption compared to state-of-the-art Montgomery reduction implementations. The proposed hardware architecture features a processing element array, distributed SRAMs, and ROMs, with three levels of reconfigurability, supporting both NTT and INTT operations. Designed using the Global Foundries’ 22 nm FD-SOI process, an Application-Specific Integrated Circuit (ASIC) is estimated to occupy 0.04 mm 2 and consume 18.2 mW at 1 GHz. The proposed processor achieves 700 times greater energy efficiency and performs computations 200 times faster than software implementations on the ARM Cortex-M4. It also achieves the lowest area–time product and highest energy efficiency among state-of-the-art NTT/INTT hardware accelerators. By carefully balancing power consumption and computational speed, this design offers an efficient solution for deploying FALCON key generation on devices with limited resources.
Ghada Alsuhli, Hani Saleh, Mahmoud Al-Qutayri, Baker Mohammad, Thanos Stouraitis
J. Inf. Secur. Appl.4
2025 A Survey and Comparative Analysis of Number Systems for Deep Neural Networks
abstract
Deep neural networks (DNNs) are indispensable in various artificial intelligence (AI) applications. However, their inherent complexity presents significant challenges, particularly when deploying them on resource-constrained devices. To overcome these hurdles, academia and industry are actively seeking ways to accelerate and optimize DNN implementations. A significant area of research revolves around discovering more effective methods to represent the enormous data volumes processed by DNNs. Traditional number systems (NSs) have proven nonoptimal for this task, prompting extensive exploration into alternative and bespoke systems for DNNs. This survey aims to comprehensively discuss various NSs utilized to efficiently represent DNN data. These systems are categorized mainly based on their impact on DNN performance and hardware implementation. This survey offers an overview of these categorized NSs and delves into different subsystems within each, outlining their effect on DNN performance and hardware design. Furthermore, these systems are compared quantitatively and qualitatively concerning their expected quantization error, memory utilization, and computational requirements. This survey also emphasizes the challenges linked with each system and the diverse proposed solutions to address them. Insights into the utilization of these NSs for sophisticated DNNs are also presented in this survey. Readers will acquire a deeper understanding of the importance of efficient NSs for DNNs, explore commonly used systems, comprehend the tradeoffs between these systems, delve into design considerations influencing their impact on DNN performance, and discover recent trends and potential research avenues in this field.
Ghada Alsuhli, Vasilis Sakellariou, Hani Saleh, Mahmoud Al-Qutayri, Baker Mohammad, Thanos Stouraitis
Proc. IEEE5
2024 Efficient and lightweight in-memory computing architecture for hardware security
Hala Ajmi, Fakhreddine Zayer, Amira Hadj Fredj, Belgacem Hamdi, Baker Mohammad, Naoufel Werghi, Jorge Dias 0001
J. Parallel Distributed Comput.5
2024 A Heterogeneous RISC-V Based SoC for Secure Nano-UAV Navigation
abstract
The rapid advancement of energy-efficient parallel ultra-low-power (ULP)$\mu$controllers units (MCUs) is enabling the development of autonomous nano-sized unmanned aerial vehicles (nano-UAVs). These sub-10cm drones represent the next generation of unobtrusive robotic helpers and ubiquitous smart sensors. However, nano-UAVs face significant power and payload constraints while requiring advanced computing capabilities akin to standard drones, including real-time Machine Learning (ML) performance and the safe co-existence of general-purpose and real-time OSs. Although some advanced parallel ULP MCUs offer the necessary ML computing capabilities within the prescribed power limits, they rely on small main memories ($<$1MB) and$\mu$controller-class CPUs with no virtualization or security features, and hence only support simple bare-metal runtimes. In this work, we present Shaheen, a 9mm$^{\textbf{2}}$200mW SoC implemented in 22nm FDX technology. Differently from state-of-the-art MCUs, Shaheen integrates a Linux-capable RV64 core, compliant with the v1.0 ratified Hypervisor extension and equipped with timing channel protection, along with a low-cost and low-power memory controller exposing up to 512MB of off-chip low-cost low-power HyperRAM directly to the CPU. At the same time, it integrates a fully programmable energy-and area-efficient multi-core cluster of RV32 cores optimized for general-purpose DSP as well as reduced-and mixed-precision ML. To the best of the authors’ knowledge, it is the first silicon prototype of a ULP SoC coupling the RV64 and RV32 cores in a heterogeneous host+accelerator architecture fully based on the RISC-V ISA. We demonstrate the capabilities of the proposed SoC on a wide range of benchmarks relevant to nano-UAV applications including general-purpose DSP as well as inference and online learning of quantized DNNs. The cluster can deliver up to 90GOp/s and up to 1.8TOp/s/W on 2-bit integer kernels and up to 7.9GFLOp/s and up to 150GFLOp/s/W on 16-bit FP kernels.
Luca Valente, Alessandro Nadalini, Asif Veeran, Mattia Sinigaglia, Bruno Sá, Nils Wistoff, Yvan Tortorella, Simone Benatti, Rafail Psiakis, Ari Kulmala, Baker Mohammad, Sandro Pinto 0001, Daniele Palossi, Luca Benini, Davide Rossi 0001
IEEE Trans. Circuits Syst. I Regul. Pap.11
2024 DRAM-Based PUF Utilizing the Variation of Adjacent Cells
abstract
The Physical Unclonable Function (PUF) is a security mechanism that takes advantage of the physical variations in a device to create a unique response that can be used as a device signature or secure key. However, many DRAM-based PUFs violate the operating rules of commodity DRAM to exploit a source of entropy in the DRAM read path. This work proposes a fast and reliable DRAM-based PUF that evaluates the variation of adjacent cells and produces the response through the normal read operation. The proposed design is implemented using 65-nm technology, and a detailed statistical SPICE simulation verifies its validity. The statistical analysis shows that the proposed PUF achieves 54.19% uniformity and 49.43% uniqueness, with 98% of the investigated responses achieving a Shannon entropy of 0.95. Additionally, the proposed design generates the response by 45, which is at least 66.7 times faster than existing systems. Furthermore, the proposed design uses the relative behavior of cells, which allows for stable responses against temperature and voltage variations, eliminating the need for error correction codes. The proposed PUF also shows resiliency against machine learning-based modeling attacks, as the prediction accuracy does not exceed 55% over 5K Challenge-Response Pairs (CRPs). The area overhead is negligible as the proposed design uses standard circuits, with the addition of only one 2x1 multiplexer at the inputs of the row buffer.
Enas E. Abulibdeh, Leen Younes, Baker Mohammad, Khaled Humood, Hani Saleh, Mahmoud Al-Qutayri
IEEE Trans. Inf. Forensics Secur.3
2023 Shaheen: An Open, Secure, and Scalable RV64 SoC for Autonomous Nano-UAVs
abstract
Open Source Hardware, the way it should be!
Luca Valente, Asif Veeran, Mattia Sinigaglia, Yvan Tortorella, Alessandro Nadalini, Nils Wistoff, Bruno Sá, Angelo Garofalo, Rafail Psiakis, M. Tolba, Ari Kulmala, Nimisha Limaye, Ozgur Sinanoglu, Sandro Pinto 0001, Daniele Palossi, Luca Benini, Baker Mohammad, Davide Rossi 0001
HCS17
2023 EACNN: Efficient CNN Accelerator Utilizing Linear Approximation and Computation Reuse
abstract
This paper proposes an efficient hardware accelerator named EACNN for use in Convolution Neural Networks. EACNN is an efficient CNN architecture that is based on co-optimization of algorithms and hardware. The proposed approach is based on linear approximation of the weights for pre-trained networks with low loss of accuracy. Furthermore, a weight substitution and remapping technique adopts linear approximation coefficients to replace CNN weights. That leads to a repetition of the weight values across different kernels and enables the reuse of CNN computations for various output feature maps. The input activations corresponding to the same linear co-efficient can be multiplied and accumulated first and then reused to generate multiple output feature maps. This computational reuse method reduces the number of multiplication and addition operations and memory accesses, which is efficiently supported by a dedicated element in the proposed EACNN. Experimental results on CIFAR 10 and CIFAR 100 datasets show that the proposed method eliminates around 61% of the multiplications in the network without significant loss of accuracy$(< 3\%)$. As a demonstration, a hardware accelerator based on EACNN was implemented on Xilinx FPGA Artix 7 and achieved a 50% reduction in the FPGA hardware resources.
Mohamed F. Tolba 0002, Hani Saleh, Baker Mohammad, Mahmoud Al-Qutayri, Thanos Stouraitis
ISCAS3
2023 On the Performance of NOMA-OFDM Systems with Time-Domain Interleaving
abstract
Non-orthogonal multiple access (NOMA) based on orthogonal frequency division multiplexing (OFDM) multicarrier modulation technique is a promising multiple access scheme for next-generation wireless communication systems. This paper analyzes the bit error rate (BER) performance of a downlink power domain NOMA-OFDM system with time-domain interleaving (TDI) over frequency selective Rayleigh channel. Theoretical BER expressions for a downlink power domain NOMA-OFDM system with TDI and using minimum mean squared error (MMSE) equalizer are developed for an arbitrary number of users. The Monte Carlo simulation results show that the proposed downlink power domain NOMA-OFDM system with TDI has better BER performance over frequency-selective fading multipath channels compared to the conventional NOMA-OFDM system without TDI.
Welelaw Yenieneh Lakew, Arafat Al-Dweik, Mahmoud Aldababsa, Mohamed Abou-Khousa, Baker Mohammad
VTC2023-Spring5
2022 Reduce Computing Complexity of Deep Neural Networks Through Weight Scaling
abstract
Large deep neural network (DNN) models are computation and memory intensive, which limits their deployment especially on edge devices. Therefore, pruning, quantization, data sparsity and data reuse have been applied to DNNs to reduce memory and computation complexity at the expense of some accuracy loss. The reduction in the bit-precision results in loss of information, and the aggressive bit-width reduction could result in noticeable accuracy loss. This paper introduces Scaling-Weight-based Convolution (SWC) technique to reduce the DNN model size and the complexity and number of arithmetic operations. This is achieved by, using a small set of high-precision weights (maximum absolute weight “MAW”) and a large set of low-precision weights (Scaling weights “SWs”). This results in decreasing the model size with minimum loss in accuracy compared to simply reducing the precision. Moreover, a scaling and quantized network-acceleration processor (SQNAP) is proposed based on the SWC method to achieve high-speed and low-power with reduced memory accesses. The proposed SWC eliminate >90% of the multiplications in the network. Moreover, the less important SWs are pruned, which has a small portion of the MAW. Retraining is applied in order to maintain accuracy. Full analysis for MNIST, Fashion MNIST, Cifar 10 and Cifar 100 datasets is presented for image recognition, where different DNN models are used including LeNet, ResNet, AlexNet and VGG 16.
Mohamed F. Tolba 0002, Hani Saleh, Mahmoud Al-Qutayri, Baker Mohammad
ISCAS4
2022 GNN-RE: Graph Neural Networks for Reverse Engineering of Gate-Level Netlists
abstract
This work introduces a generic, machine learning (ML)-based platform for functional reverse engineering (RE) of circuits. Our proposed platformGNN-REleverages the notion of graph neural networks (GNNs) to: 1) represent and analyze flattened/unstructured gate-level netlists; 2) automatically identify the boundaries between the modules or subcircuits implemented in such netlists; and 3) classify the subcircuits based on their functionalities. For GNNs in general, each graph node is tailored to learn about its own features and its neighboring nodes, which is a powerful approach for the detection of any kind of subgraphs of interest. ForGNN-RE, in particular, each node represents a gate and is initialized with a feature vector that reflects on the functional and structural properties of its neighboring gates.GNN-REalso learns the global structure of the circuit, which facilitates identifying the boundaries between subcircuits in a flattened netlist. Initially, to provide high-quality data for training ofGNN-RE, we deploy a comprehensive dataset of foundational designs/components with differing functionalities, implementation styles, bit widths, and interconnections.GNN-REis then tested on the unseen shares of this custom dataset, as well as the EPFL benchmarks, the ISCAS-85 benchmarks, and the 74X series benchmarks.GNN-REachieves an average accuracy of 98.82% in terms of mapping individual gates to modules, all without any manual intervention or postprocessing. We also release our code and source data.
Lilas Alrahis, Abhrajit Sengupta, Johann Knechtel, Satwik Patnaik, Hani Saleh, Baker Mohammad, Mahmoud Al-Qutayri, Ozgur Sinanoglu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2021 A 1: 4 Active Power Divider for 5G Phased-Array Transmitters in 22nm CMOS FDSOI
abstract
A CMOS broadband 1:4 active power divider is proposed in this paper. The power splitter can be used in Phased Array Transceivers at the Transmitter side. It is based on a cascode 1:4 current splitter and transmission lines. Compared to other passive power dividers and active power dividers, the proposed design exhibits 1-2 dB power gain and smaller area. The measured input 1dB compression point is 6 dB whereas the IIP3 is 7.2 dBm. The Noise Figure of the Power Divider is 10 dB at lower frequencies and 15 dB at 28 GHz. The measured results are performed across several chips. Realized in 22nm CMOS FDSOI from GF, the total power consumption is 29 mW from a 1 V power supply and the area occupied by the divider is 700μm × 600μm. A thorough analysis of the gain and noise of the divider is presented as well.
Nourhan Elsayed, Hani Saleh, Ademola Mustapha, Baker Mohammad, Mihai Sanduleanu
ISCAS4
2021 DTRNG: Low Cost and Robust True Random Number Generator Using DRAM Weak Write Scheme
abstract
True Random Number Generators (TRNGs) are used in a variety of applications including cryptography, computer simulation, gambling games, machine learning and hyperdimensional computing. Dynamic Random-Access Memory (DRAM)-based TRNGs have been widely used as they are low cost and available in most modern electronic devices. However, most existing DRAM-based TRNGs produce random numbers with either hardware overhead, or low throughput. In this work, we report a novel high dense, high throughput and high entropy DRAM-based TRNG, named DTRNG, that exploits the process variations inherited by the DRAM cell's access transistor and sense amplifiers. DTRNG can be employed in commodity DRAM chips with no hardware overhead and using the standard memory controller commands. The proposed method is based on controlling the word line voltage supply part of the DRAM chip during random number mode. The novel approach provided in this work has been validated using cadence circuit tools on a constructed 8x8 DRAM chip in 65nm technology. In addition, Monte Carlo simulations are performed to verify the entropy of the generated numbers. The random sequence generated by DTRNG has passed all the NIST test without any post-processing demonstrating randomness and suitability for security scheme. DTRNG is considered a milestone towards low cost and high efficient secure hardware for AI and IoT applications.
Khaled Humood, Baker Mohammad, Heba Abunahla
ISCAS2
2021 UNSAIL: Thwarting Oracle-Less Machine Learning Attacks on Logic Locking
abstract
Logic locking aims to protect the intellectual property (IP) of integrated circuit (IC) designs throughout the globalized supply chain. The SAIL attack, based on tailored machine learning (ML) models, circumvents combinational logic locking with high accuracy and is amongst the most potent attacks as it does not require a functional IC acting as an oracle. In this work, we propose UNSAIL, a logic locking technique that inserts key-gate structures with the specific aim to confuse ML models like those used in SAIL. More specifically, UNSAIL serves to prevent attacks seeking to resolve the structural transformations of synthesis-induced obfuscation, which is an essential step for logic locking. Our approach is generic; it can protect any local structure of key-gates against such ML-based attacks in an oracle-less setting. We develop a reference implementation for the SAIL attack and launch it on both traditionally locked and UNSAIL-locked designs. For SAIL, two ML models have been proposed (which we implement accordingly), namely a change-prediction model and a reconstruction model; the change-prediction model is used to determine which key-gate structures to restore using the reconstruction model. Our study on benchmarks ranging from the ISCAS-85 and ITC-99 suites to the OpenRISC Reference Platform System-on-Chip (ORPSoC) confirms that UNSAIL degrades the accuracy of the change-prediction model and the reconstruction model by an average of 20.13 and 17 percentage points (pp), respectively. When the aforementioned models are combined, which is the most powerful scenario for SAIL, UNSAIL reduces the attack accuracy of SAIL by an average of 11pp. We further demonstrate that UNSAIL thwarts other oracle-less attacks, i.e., SWEEP and the redundancy attack, indicating the generic nature and strength of our approach. Detailed layout-level evaluations illustrate that UNSAIL incurs minimal area and power overheads of 0.26% and 0.61%, respectively, on the million-gate ORPSoC design.
Lilas Alrahis, Satwik Patnaik, Johann Knechtel, Hani Saleh, Baker Mohammad, Mahmoud Al-Qutayri, Ozgur Sinanoglu
IEEE Trans. Inf. Forensics Secur.5
2020 A 28GHz, Asymmetrical, Modified Doherty Power Amplifier, in 22nm FDSOI CMOS
abstract
A 28GHz, Modified Doherty Power Amplifier (MDPA) was implemented in 22nm FDSOI CMOS technology from GF. The MDPA adopts an asymmetrical topology utilizing two cascode CMOS amplifiers as the main (Class-A) and auxiliary (Class-C). This allows a supply voltage of 2.5V and consequently higher output power. The use of a main Class-A amplifier is conducive to a higher linearity (IIP3). The integrated design implements the main and auxiliary amplifier, along with the matching and transmission line networks on chip. The fabricated amplifier occupies an area of 1.2mm2, exhibits 12 dBm saturated output power, a peak power gain of 10dB, 16% peak power-added efficiency (PAE) and 12.5% at 6-dB back-off. The measured IIP3 is 20dBm.
Nourhan Elsayed, Hani Saleh, Baker Mohammad, Mihai Sanduleanu
ISCAS3
2020 Ratioed Logic Comparator Based Digital LDO Regulator in 22nm FDSOI
abstract
This paper presents a fast and an efficient digital LDO (DLDO) regulator utilizing a clock-less ratioed logic comparator (RLC). In addition to eliminating the clock, the proposed RLC-DLDO removes the shift registers used in the conventional DLDO. It achieves a transient speed improvement in the ns range and a quiescent current reduction by 9X over the conventional DLDO design that targets μA load current. The RLC-DLDO consists of RLC, PMOS power switches and control unit. The RLC compares between the reference and the load voltage and generates a single bit that turns on/off the PMOS switches. Unlike the clocked comparator, the RLC is an event-driven design that continuously responds to the voltage difference. The control unit provides digital bits to control the power switches and the RLC circuit in order to support different output voltage levels. The RLC-DLDO has an input voltage range between 0.8V and 0.6V and generates an output voltage range between 0.7V to 0.5V for load current between 10μA and 500μA. The design is implemented in 22nm FDSOI and occupies an active area of 0.0171mm2. The simulation results show that the peak efficiency is 99.9% and the load transient response time is 5ns at VL=0.5V.
Dima Kilani, Baker Mohammad, Mihai Sanduleanu
ISCAS2
2020 ASIC Implementation of a Pre-Trained Neural Network for ECG Feature Extraction
abstract
The electrocardiogram signal (ECG), a record of electrical activity of the cardiac muscle, has been used in diagnosing many cardiopathies. Wearable devices equipped with readout sensors and circuits can be used to record and process weak ECG signals. In this paper, a pre-trained neural network was implemented for detecting the QRS feature of an ECG signal, which is crucial for auto-diagnostic of various cardiopathies. To take advantage of the fast evolution of artificial intelligence and its ability to find non-linear relationships, neural network based feature extraction of ECG signals for wearable devices was explored and tested using ASIC implementation flow. Firstly, a high-level simulation was carried out in MATLAB and verified with test data obtained from PhysioNET database. Recurrent neural network (RNN) MLP was created and trained using the data obtained from PhysioNET database. A high-level performance evaluation was carried out using the same network for P and T wave extraction. The weight and bias matrices obtained from the high-level trained network in MATLAB were used in the design of the hardware. An accuracy of 96.55% was achieved in the hardware implementation of the network.
Huruy Tekle Tefai, Hani Saleh, Temesghen Tekeste, Mahmoud Al-Qutayri, Baker Mohammad
ISCAS5
2019 ScanSAT: unlocking obfuscated scan chains
abstract
While financially advantageous, outsourcing key steps such as testing to potentially untrusted Outsourced Semiconductor Assembly and Test (OSAT) companies may pose a risk of compromising on-chip assets. Obfuscation of scan chains is a technique that hides the actual scan data from the untrusted testers; logic inserted between the scan cells, driven by a secret key, hide the transformation functions between the scan-in stimulus (scan-out response) and the delivered scan pattern (captured response). In this paper, we propose ScanSAT: an attack that transforms a scan obfuscated circuit to its logic-locked version and applies a variant of the Boolean satisfiability (SAT) based attack, thereby extracting the secret key. Our empirical results demonstrate that ScanSAT can easily break naive scan obfuscation techniques using only three or fewer attack iterations even for large key sizes and in the presence of scan compression.
Lilas Alrahis, Muhammad Yasin, Hani Saleh, Baker Mohammad, Mahmoud Al-Qutayri, Ozgur Sinanoglu
ASP-DAC4
2019 Functional Reverse Engineering on SAT-Attack Resilient Logic Locking
abstract
Logic locking is a solution that mitigates hardware security threats, such as Trojan insertion, piracy and counterfeiting. Research in this area has led to, in an iterative fashion, a series of logic locking defenses as well as attacks that circumvent these defenses by extracting the logic locking key. The most powerful attacks rely on a full access to a working chip/oracle that can be used to produce the input-output pairs utilized in recovering the secret key. A recently proposed technique Stripped Functionality Logic Locking (SFLL) provides resilience to all known attacks on combinational logic locking. In this paper, we propose a functional reverse engineering attack on SFLL: an attack that can detect the protection logic of SFLL which results in obtaining the original unlocked design with a high success rate. The restore and perturb blocks utilized by SFLL were detected with average coverage percentages of 93.95% and 85.42% respectively, proving that our attack is capable of breaking the state of the art logic locking technique.
Lilas Alrahis, Muhammad Yasin, Hani Saleh, Baker Mohammad, Mahmoud Al-Qutayri
ISCAS4
2019 Editorial TVLSI Positioning - Continuing and Accelerating an Upward Trajectory
abstract
I. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5].
Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber
IEEE Trans. Very Large Scale Integr. Syst.33
2019 A Gain-Controlled, Low-Leakage Dickson Charge Pump for Energy-Harvesting Applications
abstract
This paper presents a single-stage power management unit to boost and regulate a low supply voltage for CMOS system-on-chip (SoC) applications. It consists of low-leakage, enhanced Dickson charge pump (DCP) that utilizes both stage and frequency modulation (FM) techniques to achieve high efficiency and lower area. In addition, the proposed design uses an enhanced stage-switch structure for the charge pump, which significantly reduces the cross-stage leakage. A stage number controller is used to control the gain of the charge pump by changing the number of stages based on the desired output voltage. FM is utilized to further fine-tune the output voltage through a closed-loop control based on a predetermined reference voltage. Silicon measurement results for the four-stage charge pump in 65-nm CMOS technology show a maximum end-to-end efficiency of 66% at an input voltage of 0.7 V and an output power of$27~\mu \text{W}$. The proposed design achieved more than a$100\times $reduction in leakage compared to traditional DCP. The system supports a range of load currents between 0.1 and$34~\mu \text{A}$with a maximum operating frequency of 1.8 MHz. The proposed system supports an input voltage range of 0.55–0.7 V which makes it an excellent candidate for solar and thermal energy-harvesting applications targeting low-power internet-of-things SOC.
Abdulqader Nael Mahmoud, Mohammad Alhawari, Baker Mohammad, Hani Saleh, Mohammed Ismail 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2018 An Efficient and Small Area Multioutput Switched Capacitor Buck Converter for IoTs
abstract
This paper presents an area and power efficient multioutput switched capacitor (MOSC) DC-DC buck converter targeting ultra-low power and IoT devices. The MOSC converter has a variable input voltage range between 1.05V to 1.4V and generates regulated simultaneous multiple output voltage levels of 1V and 0.55V. Adaptive digital time multiplexing controller is employed to enable multiple power domains. In addition, a pulse frequency modulation is utilized to regulate the output voltages over a wide range of load current. The proposed converter supports a load current of 10μA to 350μA and to 10μA at a load voltage of 1V and 0.55V, respectively. Adaptive time multiplexing and pulse width modulation are implemented through a finite state machine to eliminate the reverse current issue. This problem arises during the switching from a low load voltage of 0.55V to a high load voltage of 1V. The MOSC circuit is fabricated in 65nm CMOS technology and it occupies an active area of 0.27mm2. Moreover, both MIM and MOS capacitors are utilized to further reduce the area of MOSC converter. Measured results shows that the peak efficiency of 78% is achieved at a load power of 300μW.
Dima Kilani, Mohammad Alhawari, Baker Mohammad, Hani Saleh, Mohammed Ismail 0001
ISCAS3
2018 A Charge Pump Based Power Management Unit With 66%-Efficiency in 65 nm CMOS
abstract
This paper presents a single stage power management unit that includes an enhanced stage-switch Dickson charge pump (DCP) to boost and regulate a low input voltage. A new switching mechanism is presented to significantly reduce the losses encountered in conventional DCP switches. Frequency and stage modulation are utilized in the proposed design. The stage modulation provides different gain levels (coarse) and the frequency modulation tunes the voltage level and regulates the output voltage based on a pre-determined reference voltage. Using four stages charge pump, silicon measurement results in 65 nm CMOS technology show a maximum efficiency of 66% at input voltage of 0.7 V and output power of 27 μW. The system supports a range of load current between 0.1 μA − 34 μA with a maximum operating frequency of 1.8MHz. The proposed system supports an input voltage range from 0.55 to 0.7 V which can be used in energy harvesting applications such as solar and thermal harvesting.
Abdulqader Nael Mahmoud, Mohammad Alhawari, Baker Mohammad, Hani Saleh, Mohammed Ismail 0001
ISCAS3
2018 Stateful Memristor-Based Search Architecture
abstract
Computer vision and recognition is emerging as one of the important pillars in artificial intelligence systems. It is a vital way to interpret the collected data and find matching patterns that will help in real-time decision making. CMOS-based search engines suffer from density and power limitations. Memristor is a feasible candidate that is capable of performing search within a stored structure (in-memory computing). This paper proposes the first memristor-based stateful search engine architecture based on a novel stateful heterogeneous memristive XOR gate. The design is suitable for 2-D media applications, such as image matching and pattern inspection. It performs bitwise comparison using the proposed XOR gate. The output states of all XOR gates are transferred into a single analog memristor value that is read via a digital comparator. The design assumes a single memristor device for each of the incoming data, template, and result bits. Each 2-D array of input, template, and output is reordered into a single 1-D array with 3 × (N × M) structure, where N represents the number of entry data and M is the number of bits per entry. This allows for a significantly higher storage density than conventional CMOSbased or other memristor-based search engines. Simulations of the proposed architecture demonstrate functionalities in search and compare modes using an LTSpice circuit simulator. The proposed architecture achieves a 3-ns search cycle time at 0.34 nJ/database at 1.5 V/1 GHz using 2N + 1 memristors.
Yasmin Halawani, Muath Abu Lebdeh, Baker Mohammad, Mahmoud Al-Qutayri, Said F. Al-Sarawi
IEEE Trans. Very Large Scale Integr. Syst.3
2018 Memristor-Based Hardware Accelerator for Image Compression
abstract
Memristor-based hardware accelerators are gaining an increased attention as a potential candidate to speed-up the vector-matrix operations commonly needed in many digital image processing tasks due to their area, speed, and energy efficiency. In this paper, a memristor-based image compression (MR-IC) architecture that exploits a lossy 2-D discrete wavelet transform is proposed. The architecture is composed of a computational memristor crossbar, an intermediate memory array that stores the row-transformed coefficients and a final memory that holds the compressed version of the original image. The computational memristor array performs in-memory computation on the initially stored transformation coefficients. Using the quantitative analysis approach, we demonstrate a 10× reduction in a number of operations compared with a conventional application-specific integrated circuit implementation. This translates to five orders of magnitude reduction in area, around 11× improvement in energy efficiency, and 1.28× speedup in computation time. Image quality metrics, such as peak signal-to-noise ratio (PSNR), structural similarity (SSIM) index, and complex wavelet-SSIM (CW-SSIM), are used to quantify the reduction in image quality due to lossy compression. The achieved metrics for conventional versus MR-IC are: PSNR 57.24 versus 33.29 dB, SSIM 0.9994 versus 0.8853, and CW-SSIM 1 versus 0.9983. Simulation results show that the 32 quantization levels proposed architecture provides significant improvements in energy, area, and performance compared to the 32 levels CMOS implementation with comparable CW-SSIM.
Yasmin Halawani, Baker Mohammad, Mahmoud Al-Qutayri, Said F. Al-Sarawi
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Novel hafnium oxide memristor device: Switching behaviour and size effect
abstract
Unipolar RRAM devices are of high interest due to their high resistance ratio and simple selector circuit. In this paper, we report on a measurements from nano-thick memristor featuring a novel Pd/Hf/HfO2/Pd stack. The fabricated device exhibits a unipolar switching behavior, due to the asymmetric device structure and the existence of the Pd metal as a bottom electrode. The forming voltage of the proposed memristive stack (Hf-10nm/HfO2-10nm) is found to be size dependent at the microscale and its average forming voltage decreases by 20% when the active area increases from 30×30 to 1000×1000 μm2. The findings presented in this work highlight the impact of device geometry on its electrical performance and power, which provide guidance to the design tradeoffs (size, power, resistance ratio) and fabrication process of memristor devices.
Heba Abunahla, Baker Mohammad, Maguy Abi Jaoude, Mahmoud Al-Qutayri
ISCAS2
2017 A sub-μW bio-potential front end in 65nm CMOS
abstract
A bio-potential amplifier intended for continuous monitoring of vitals characterized by its long operational lifetime is required to operate at the lowest power budget possible. Moreover, a compact active area directly related to portability is essential. This paper presents a 0.55pW auto gain controlled biopotential amplifier implemented in 65nm 1P7M CMOS for ECG signal classifier SoC. A chopper-stabilized amplifier is designed at 0.6V supply voltage to mitigate the DC offset and near DC flicker noise. The input ECG signal level is further set by the four gain levels of the variable gain amplifier (VGA) to provide maximum swing to the ADC. The whole system is integrated into a core are of 0.10mm2and can operate at a wide range of 0.6-1.2V supply voltage.
Yonatan Kifle, Hani Saleh, Baker Mohammad, Mohammed Ismail 0001
VLSI-SoC3
2016 Physics model of memristor devices with varying active materials
abstract
This paper presents a physics-based model for memristors with different active layer materials. The model predicts the effect of changing the active material on the electrical characteristics of the devices. It captures the essential characteristics of the memristor such as coupling between ion mobility and electron current in addition to the nonlinear effects of electric fields. The parameters in the model depend on material (metal-oxide) properties that have impact on the device behavior. These properties are activation energy, escape attempt frequency, hopping parameter and relative permittivity. In this work, the effect of each parameter is highlighted and explained. In addition, the physics-based Matlab model is used to analyze the electrical characteristics of simulated memristor device using the following oxide materials; ZnO, TiO2 and Ta2O5. The simulation results of the model are validated with experimental data reported in the literature. The value of this contribution is to enable the selection of suitable oxide materials for the target memristor using correlated mathematical models.
Heba Abunahla, Nadeen El Nachar, Dirar Homouz, Baker Mohammad, Maguy Abi Jaoude
ISCAS4
2016 An efficient thermal energy harvesting and power management for μWatt wearable BioChips
abstract
This paper presents an efficient thermal energy harvesting IC (EHIC) that supports a battery-less μWatt system-on-chips. The EHIC consists of an inductor-based DC-DC converter that boosts a low input voltage to a suitable output voltage level. Further, a switched capacitor buck converter is utilized to regulate the boost converter output voltage and to support multiple output voltage levels, namely 0.6V, 0.8V and 1V. In low energy mode and to enhance the efficiency, the EHIC is capable of bypassing the switched capacitor so that the load is driven directly from the boost converter. The prototype chip is fabricated in 65nm CMOS and occupies an area of less than 0.46mm2. Measured results confirm an efficiency of 65% at 0.6V output voltage and 42μW. In addition, the end-to-end peak efficiency is 71% at 0.8V output voltage and 182μW.
Mohammad Alhawari, Dima Kilani, Baker Mohammad, Hani Saleh, Mohammed Ismail 0001
ISCAS3
2016 A biomedical SoC architecture for predicting ventricular arrhythmia
abstract
Electrocardiography (ECG) represents the hearts electrical activity and has features such as QRS complex, P-wave and T-wave that provide critical clinical information for detection and prediction of cardiac diseases. This paper presents a novel ECG processing architecture for the prediction of ventricular arrhythmia (VA). The architecture implements a novel ECG feature extraction which is optimized for ultra-low power applications. The architecture is based on Curve Length Transform (CLT) for the detection of QRS complex and Discrete Wavelet Transform (DWT) for the delineation of TP waves. Features extracted from two consecutive ECG cycles are used to set innovative parameters for VA prediction up to 3 hours before VA onset. Two databases of the heart signal recordings from the American Heart Association (AHA) and the MIT PhysioNet were used as training, test and validation sets to evaluate the performance of the proposed system.
Temesghen Tekeste, Hani Saleh, Baker Mohammad, Ahsan H. Khandoker, Mohammed Ismail 0001
ISCAS3
2016 Low-Power ECG-Based Processor for Predicting Ventricular Arrhythmia
abstract
This paper presents the design of a fully integrated electrocardiogram (ECG) signal processor (ESP) for the prediction of ventricular arrhythmia using a unique set of ECG features and a naive Bayes classifier. Real-time and adaptive techniques for the detection and the delineation of the P-QRS-T waves were investigated to extract the fiducial points. Those techniques are robust to any variations in the ECG signal with high sensitivity and precision. Two databases of the heart signal recordings from the MIT PhysioNet and the American Heart Association were used as a validation set to evaluate the performance of the processor. Based on application-specified integrated circuit (ASIC) simulation results, the overall classification accuracy was found to be 86% on the out-of-sample validation data with 3-s window size. The architecture of the proposed ESP was implemented using 65-nm CMOS process. It occupied 0.112- ${\rm mm}^{2}$ area and consumed 2.78- $\mu \text{W}$ power at an operating frequency of 10 kHz and from an operating voltage of 1 V. It is worth mentioning that the proposed ESP is the first ASIC implementation of an ECG-based processor that is used for the prediction of ventricular arrhythmia up to 3 h before the onset.
Nourhan Bayasi, Temesghen Tekeste, Hani Saleh, Baker Mohammad, Ahsan H. Khandoker, Mohammed Ismail 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2016 Modeling and Optimization of Memristor and STT-RAM-Based Memory for Low-Power Applications
abstract
Conventional charge-based memory usage in low-power applications is facing major challenges. Some of these challenges are leakage current for static random access memory (SRAM) and dynamic random access memory (DRAM), additional refresh operation for DRAM, and high programming voltage for Flash. In this paper, two emerging resistive random access memory (ReRAM) technologies are investigated, memristor and spin-transfer torque (STT)-RAM, as potential universal memory candidates to replace traditional ones. Both of these nonvolatile memories support zero leakage and low-voltage operation during read access, which makes them ideal for devices with long sleep time. To date, high write energy for both memristor and STT-RAM is one of the major inhibitors for adopting the technologies. The primary contribution of this paper is centered on addressing the high write energy issue by trading off retention time with noise margin. In doing so, the memristor and STT-RAM power has been compared with the traditional six-transistor-SRAM-based memory power and potential application in wireless sensor nodes is explored. This paper uses 45-nm foundry process technology data for SRAM and physics-based mathematical models derived from real devices for memristor and STT-RAM. The simulations are conducted using MATLAB and the results show a potential power savings of 87% and 77% when using memristor and STT-RAM, respectively, at 1% duty cycle.
Yasmin Halawani, Baker Mohammad, Dirar Homouz, Mahmoud Al-Qutayri, Hani Saleh
IEEE Trans. Very Large Scale Integr. Syst.2
2015 A maximally stable extremal regions system-on-chip for real-time visual surveillance
abstract
This paper presents a novel implementation of the Maximally Stable Extremal Regions (MSER) detector on system-on-chip (SoC) using 65 nm CMOS technology. The novel SoC was developed following the Application Specific Integrated Circuit (ASIC) design flow which significantly enhanced its realization and fabrication, and overall performances. The SoC has very low area requirement (around 0.05 mm2) and is capable of detecting both bright and dark MSERs in a single run, while computing simultaneously their associated regions' moments, simplifying its interfacing with other image algorithms (e.g. SIFT and SURF). The novel MSER SoC is power-efficient (requires 2.25 mW) and memory-efficient as it saves more than 31% of the memory space reported in the state-of-the-art MSER implementation on FPGA, making it suitable for mobile devices. With 256×256 resolution and its operating frequency of 133 MHz, the SoC is expected to have a 200 frames/second processing rate, making it suitable (when integrated with other algorithms in the system) for time-critical real-time applications such as visual surveillance.
Ehab Salahat, Hani Saleh, Andrzej Stefan Sluzek, Mahmoud Al-Qutayri, Baker Mohammad, Mohammed Ismail 0001
IECON5
2015 Novel fast and scalable parallel union-find ASIC implementation for real-time digital image segmentation
abstract
This paper presents a new fast and scalable Parallel Union-Find algorithm for image segmentation and its System-on-Chip (SoC) implementation using 65nm CMOS technology following the Application-Specific Integrated Circuit (ASIC) design flow. The algorithm is capable of labeling all foreground and background pixels, using the least possible pixels scanning. This contrasts the classical labeling algorithms that label only foreground (or background) pixels in a single run. The new algorithm utilizes only two memory blocks. In one memory block, it labels image segments using their seeds as the label and, simultaneously, the segments sizes are used as the other label in second memory block. By this parallel labeling, monitoring the image segments is very fast and efficient. With 350 MHz operating frequency, the processing rate estimated to be 2100 frames/sec, the total chip area of 15950.5 μm2 (off-chip memory) and very low-power of 0.3 mW, the SoC tends to be an excellent candidate for mobile devices and real-time applications.
Ehab Salahat, Hani Saleh, Andrzej Stefan Sluzek, Mahmoud Al-Qutayri, Baker Mohammad, Mohammed Ismail 0001
IECON5
2015 Novel MSER-guided street extraction from satellite images
abstract
The paper presents a novel technique to segment and extract streets from satellite images. This technique utilizes, for the first time in the known literature, the Maximally Stable Extremal Regions (MSER) algorithm to robustly identify and segment streets from satellite images. The technique extracts dark MSERs and then classifies them based on multiple metrics such as the intensity of the pixels, the region stability, and the major-to-minor axes ratio. Testing results under multiple scenarios corroborate the accuracy of the proposed technique. The technique will allow fast and accurate implementation of a wide spectrum of applications such as in Global Positioning System (GPS) driving guidance.
Ehab Salahat, Hani Saleh, Andrzej Stefan Sluzek, Baker Mohammad, Mahmoud Al-Qutayri, Mohammed Ismail 0001
IGARSS4
2015 A 65-nm low power ECG feature extraction system
abstract
This paper presents a real-time adaptive ECG detection and delineation algorithm alongside an architecture based on time-domain signal processing of the ECG signal. The algorithm is enhanced to detect large number of different P-QRS-T waveform morphologies using adaptive search windows and adaptive threshold levels. The proposed architecture has been implemented in the state-of-the-art 65-nm CMOS technology. It occupied 0.03416 mm2 area and consumed 0.614 mW power. Furthermore, the non-complex nature of the architecture resulted with a realization using smaller number of computation and higher performance. The design of the QRS detector was tested on ECG records obtained from the Physionet QT database and achieved a sensitivity of Se =99.83% and a positive predictivity of P+= 98.65%. Similarly, the mean error values of the T peak, T offset, P peak and P offset were found to be -1.367, 6.36, 5.5 and -2.59 milliseconds, respectively, using the same database. The small area, low power, and high performance of our architecture makes it suitable for inclusion in System On Chips (SOCs) targeting wearable mobile medical devices.
Nourhan Bayasi, Temesghen Tekeste, Hani Saleh, Baker Mohammad, Mohammed Ismail 0001
ISCAS4
2015 Memory impact on the lifetime of a Wireless Sensor Node using a Semi-Markov model
abstract
The increase in demand for higher functionality, smaller size, lower cost and near perpetual operation of Wireless Sensor Nodes (WSNs) are posing big challenges for system designers. A major aspect is the operational lifetime of the system which is determined by the finite energy source supplied by the battery. In WSNs, the higher power incurred due to added system functionality and the increase leakage as a result of technology scaling have high impact on the battery lifetime. In this work, memory was used to further exploit the energy efficiency at the sensor node system-level. A detailed analysis using Semi-Markov model to investigate different operational modes of WSN with the present SRAM shows an improvement of 4x at 90% duty cycle in the node's lifetime. In addition, an emerging non-volatile memory (NVM) technology, Memristor, is also explored to further improve the WSN energy efficiency. Its non-volatility nature will suppress the power wasted as leakage in SRAM during idle periods which is typical for low duty-cycle WSNs. The results show that utilizing an on-chip NVM can further improve WSN lifetime by 1x for low activity μW range sensor nodes.
Yasmin Halawani, Baker Mohammad, Mahmoud Al-Qutayri, Hani Saleh
ISCAS2
2015 Adaptive ECG interval extraction
abstract
ECG intervals such as QRS, QT and PR provide significant information and are widely used as clinical parameters for diagnosing cardiac diseases. This paper presents a novel QRS detection technique based on Curve Length Transform (CLT) and a refined delineation of P-wave and T-wave using Discrete Wavelet Transform (DWT). The proposed technique was verified using the PhysioNet database. The QRS detection achieved a sensitivity of 98.59% and a positive predictivity of 97.86%. The QRS duration, QT interval and PR interval had a mean error of -1.56± 28.8ms, -5.39± 42.4ms and 0.86± 40.3ms respectively. The proposed algorithm is computationally efficient and is simpler to implement in hardware, hence, will lead to a faster execution time, smaller design area and consequently low power consumption.
Temesghen Tekeste, Nourhan Bayasi, Hani Saleh, Ahsan H. Khandoker, Baker Mohammad, Mahmoud Al-Qutayri, Mohammed Ismail 0001
ISCAS5
2015 Evolutionary QR-Based Traffic Sign Recognition System for Next-Generation Intelligent Vehicles
abstract
This paper introduces a dramatically novel traffic signs recognition (TSR) system that can perform traffic sign detection and tracking simultaneously. The proposed approach utilizes intensity images and the depth images, in parallel, to robustly detect and track traffic signs in real-time. Additionally, we suggest to supplement the ordinary traffic signs with the corresponding quick-response (QR) code plates that inherent the many advantages of the QR-codes, introducing the concept of QR-TSR systems.
Ehab Salahat, Hani Saleh, Andrzej Stefan Sluzek, Mahmoud Al-Qutayri, Baker Mohammad, Mohammed Ismail 0001
VTC Fall5
2015 Embedded Memory Interface Logic and Interconnect Testing
abstract
The increased size of embedded memory for system-on-chip (SoC) and multicore processors has a positive impact on performance yet poses a big challenge for chip yield, power consumption, and overall cost. Big percentage (>60%) of today's processors and SoC area in both 2-D (planar) and 3-D technologies, such as through silicon via (TSV), are dedicated to memory. Most of today's embedded memories are not as simple as a storage area with single interface of data, address, and control, but rather they compromise complex logic on their interface due to timing constrains and interconnect technologies (NoC and TSV). Memory core testing strategy is well understood and has mature tools and methodologies to screen for defects such as built-in-self-test (BIST). In addition, core and logic-based testing using scan and automatic test pattern generation (ATPG) tools and methodologies are intended for flop-based design. However, interface logic and complex interconnect like the one in 3-D chips are not thoroughly tested using BIST or ATPG as they are not designed for such logic. This becomes even more important for 3-D chips where a stack memory could have different testing strategies other than the base layer core which is interfacing with it. This brief presents a design for test methodology to achieve good coverage on interface logic for embedded and stack memory. The proposed approach uses modified ATPG and scan methodology to test the memory logic interface with minimum impact to existing design.
Baker Mohammad
IEEE Trans. Very Large Scale Integr. Syst.1
2015 Design Methodologies for Yield Enhancement and Power Efficiency in SRAM-Based SoCs
abstract
This paper comprises two new methodologies to improve yield and reduce system-on-a-chip power. The first methodology is based on faulty static random-access memory (SRAM) cells detections and cache resizing. The key advantage of this approach is that it enables the end user to control the system's parameters to be error tolerant. Furthermore, this technique enables aggressive voltage scaling which causes parametric (soft) failures in SRAM-based memory. As such, the proposed methodology can be utilized to exchange cache size for lower power or better yield. In the second methodology, data from faulty cells are treated as imposed noise. Depending on the application, this error percentage (imposed noise) can be mitigated through three options. First, ignore error if the percentage of the error is tolerable. Second, simple hardware filtration is needed. Finally, software-based filtration is required. The viability of this approach is that it allows aggressive voltage scaling below the traditional to be a 100% correct approach for SRAM supply, which results in substantial reduction of power, trading off quality for power. For both approaches, BIST is used as part of the powerup sequence to identify the faulty memory addresses per voltage level and compute the faulty cells percentage. Furthermore, the proposed methodologies help in improving reliability and counteracting long-term effects on memory cell stability and lifetime degradation caused by negative bias temperature instability.
Baker Mohammad, Hani Saleh, Mohammed Ismail 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2013 Robust Hybrid Memristor-CMOS Memory: Modeling and Design
abstract
In this paper, we explore various aspects of memristor modeling and use them to propose improved access operations and design of a memristor-based memory. We study the current mathematical and SPICE modeling of memristors and compare them with known device specifications. Based on this survey of existing models, we adopt an improved mathematical model of the memristor that captures the well-established features of memristive devices. This modeling is used to analyze the time and voltage characteristics of stable read and write operations. The tradeoffs between the various design parameters such as voltage, frequency, noise margin, and area are also analyzed. Based on the device modeling, we propose a hybrid CMOS-memristor memory cell and architecture that addresses the limitations of memristor such as state drift, cell-cell interference, and refresh requirements. Memristor is used as a state element, and CMOS-based transistors are used to isolate, control, decode, and inter operate the logic. We verify our design using SPICE simulation using a 28-nm model for CMOS and a modified memristor model.
Baker Mohammad, Dirar Homouz, Hazem Elgabra
IEEE Trans. Very Large Scale Integr. Syst.1
2012 Write-through method for embedded memory with compression Scan-based testing
abstract
Demands for low defects per million (DPM) rates are increasing as process technology scaling is able to increase transistor density and add more functionality to the integrated circuits. For stuck at fault and delay testing, Scan-based testing in conjunction with ATPG is the preferred approach to reduces DPM compared to functional testing. However embedded memories have been a challenge to ATPG gate level simulation due to limitation of gate level generation method and the additional logic needed to prevent unknowns (X's) to be propagated from memory during ATPG testing, this X-propagation becomes more of an issue when the design has a test compressor. This paper examines the challenges of ATPG memory write through method on the design with chip test compression logic and proposes new design strategy and ATPG pattern generation method. The proposed design will make the memory look like a one dimensional set of registers and ATPG pattern generation method will support write through mode without Xs propagation.
Geewhun Seok, Hong Kim, Baker Mohammad
VTS3
2008 Adaptive SRAM memory for low power and high yield
abstract
SRAMs typically represent half of the area and more than half of the transistors on a chip today. Variability increases as feature size decreases, and the impact of variability is especially pronounced on SRAMs since they make extensive use of minimum sized devices. Variability leads to a large amount of guard banding in the design phase in order to meet frequency and yield targets. We develop an SRAM architecture that eliminates guard banding. Specifically, our SRAM uses multiple supply voltages that are assigned post-manufacturing. We compensate for variation by powering up manufactured devices that are slower than designed. Specifically, we assign supply voltages to 6T cells on a per-column basis; this gives us sufficiently fine-grained control over devices without excessive area overhead. We show that post-manufacturing voltage assignment results in a 28% reduction in bitline energy compared to a fixed voltage design for the same yield using data from a real-world 45 nm process.
Baker Mohammad, Stephen Bijansky, Adnan Aziz, Jacob A. Abraham
ICCD1
2007 A 65-nm pulsed latch with a single clocked transistor
abstract
This paper discusses the technology limits placed on the clock switching energy in sequential elements. It proposes a novel pulsed latch that uses a single clocked transistor and consumes close to ten times less clock power than a conventional latch using six clocked transistors. It describes how the new circuit enables additional power savings when virtual grounds, instead of a regular clock, are locally distributed to a group of latches. Finally, the paper discusses how to further reduce the dynamic clock power consumption of the new latch without degrading its timing by feeding it a low-swing clock.
Martin Saint-Laurent, Baker Mohammad, Paul Bassett
ISLPED2