EDBT 2026 Demo / reviewers in the wild / expert
Hani Saleh
dblp:139/9023 · also Hani H. Saleh
· DBLP profile ↗
46ranked-venue papers
3as first author
19since 2021 · last 2027
0000-0002-7185-0278ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 2 first-author · 9 since 2021Security and privacy · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Chin electromyography-based explainable machine learning framework for obstructive sleep apnea severity assessment
Adil Rehman, Hani Saleh, Ahsan H. Khandoker, Mahmoud Al-Qutayri |
Expert Syst. Appl. | 2 |
| 2026 | Efficient On-the-Fly Twiddle Factor Generation for Falcon PQC NTT/INTT
Ghada Alsuhli, Hani Saleh, Mahmoud Al-Qutayri, Baker Mohammad, Thanos Stouraitis |
ISCAS | 2 |
| 2026 | Casting Ventricular Arrhythmia Detection as Anomaly Detection via One-Class Meta-LearningabstractVentricular arrhythmia detection is a critical yet challenging task in cardiac healthcare due to the rarity of abnormal episodes and the high inter-patient variability in cardiac signals. These challenges are further exacerbated in implantable cardioverter-defibrillators, which operate under stringent memory and computational constraints. In this paper, we analyze inter- and intra-patient variability using dimensionality reduction and divergence metrics, and leverage these observations to formulate ventricular arrhythmia detection as a deployment-aligned one-class meta-learning problem. Accordingly, we adopt a one-class formulation of model-agnostic meta-learning (OC-MAML) with a clinically grounded task design that reflects real-world deployment conditions. Specifically, patient-disjoint support and query sets are used to simulate realistic distribution shifts and inter-patient variability. By training primarily on normal intracardiac electrogram segments, the OC-MAML-based framework learns a task-agnostic initialization that rapidly adapts to new patients using only a few normal samples, thereby substantially reducing dependence on labeled arrhythmic data. Compared to conventionalvanillamodel-agnostic meta-learning (MAML), our proposed OC-MAML-based framework achieves a relative improvement of +14.1% in sensitivity, +2.5% in balanced accuracy, and +4.1% inF1-score, while reducing adaptation time by 5× and maintaining comparable memory efficiency. These results underscore the framework’s potential for scalable, nearly label-free deployment in edge-based cardiac monitoring systems. The code for the proposed framework is publicly available at https://github.com/jaradat/VAD-OC-MAML. Abeer A. Jaradat, Hani Saleh, Omar Alhussein, Ghada Alsuhli, Thanos Stouraitis |
IEEE Internet Things J. | 2 |
| 2025 | Detecting Sleep Stages and Transitions Using Chin Electromyography: A Two-Stage Machine Learning Approach
Adil Rehman, Mostafa Moussa, Hani Saleh, Ahsan H. Khandoker, Ali Khraibi, Mahmoud Al-Qutayri |
HealthCom | 3 |
| 2025 | Enhanced CNN Performance without Retraining Via Weight Approximation and Data ReuseabstractThis paper introduces an efficient CNN algorithm to address key limitations in Deep Neural Networks (DNNs) used for image recognition, focusing particularly on model size and retraining time. Traditional methods often require significant training durations; however, applying approximation techniques during retraining can exacerbate these time demands. We present an approach that enhances approximation techniques while eliminating the need for model retraining, thus enabling DNN compression with minimal accuracy loss. The proposed method integrates three core strategies: weight arrangement, approximation, and data reuse. The DNN weights are initially arranged in ascending order to optimize subsequent operations. During inference, the approximation is applied to reduce the model size and minimize computational complexity by reducing the number of operations required for each multiply-accumulate (MAC) unit. Then, the original weights are replaced with the approximated values, enabling the reuse of computations and data across different sets of weights. As a result, the method significantly reduces memory access, computational demands, and energy consumption. Experimental results on the CIFAR-10 and TinyImageNet datasets demonstrate that our method achieves a model reduction rate of approximately 198.6× while maintaining a minimal loss in accuracy. The proposed technique bypasses the need for retraining, offering a practical solution to the growing complexity of DNN models in modern applications. Mohamed F. Tolba 0002, Hani Saleh, Baker Mohammad, Mahmoud Al-Qutayri, Thanos Stouraitis |
ISCAS | 2 |
| 2025 | A Mixed-Precision RNS DNN AcceleratorabstractThe Residue Number System (RNS) has been used for the design of Deep Neural Network (DNN) processing architectures due to its efficient implementation of the multiply-accumulate (MAC) operation. Prior-art RNS DNN accelerators have demonstrated notable benefits compared to conventional fixed-point (FXP) representations for arithmetic precisions of at least 8 bits. However, advanced quantization techniques have recently enabled accurate ultra-low-precision FXP DNN inference. Thus, it remains an open research question whether RNS can still outperform FXP representations for smaller precisions and especially in mixed-precision (MXP) quantization settings, where optimal bit-width configurations with respect to overall accuracy drop constraints are sought. This work addresses this gap by presenting an RNS-based MXP DNN accelerator that supports 3–8-bit quantization and consistently achieves superior model performance vs. hardware cost tradeoffs for various DNN models, resulting in up to 1.2× energy efficiency improvements compared to the FXP counterpart. Synthesized on a 22-nm technology, the RNS MXP accelerator achieves 6.93–14.58 TOPS/W, outperforming the state-of-the-art uniform-precision RNS accelerator by 1.4× while maintaining the original model accuracy, as well as mixed-precision FXP accelerators. Vasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis |
ISCAS | 4 |
| 2025 | Encoder decoder-based Virtual Physically Unclonable Function for Internet of Things device authentication using split-learningabstractInternet of Things (IoT) networks have been deployed widely making device authentication a crucial requirement that poses challenges related to security vulnerabilities, power consumption, and maintenance overheads. While current cryptographic techniques secure device communication; storing keys in Non-Volatile Memory (NVM) poses challenges for edge devices. Physically Unclonable Functions (PUFs) offer robust hardware-based authentication but introduce complexities such as hardware production and conservation expenses and susceptibility to aging effects. This paper’s main contribution is a novel scheme based on split learning, utilizing an encoder–decoder architecture at the device and server nodes, to first create a Virtual PUF (VPUF) that addresses the shortcomings of the hardware PUF and secondly perform device authentication. The proposed VPUF reduces maintenance and power demands compared to the hardware PUF while enhancing security by transmitting latent space representations of responses between the node and the server. Also, since the encoder is placed on the node, while the decoder is on the server, this approach further reduces the computational load and processing time on the resource-constrained node. The obtained results demonstrate the effectiveness of the proposed VPUF scheme in modeling the behavior of the hardware-based PUF. Additionally, we investigate the impact of Gaussian noise in the communication channel between the server and the node on the system performance. The obtained results further reveal that the achieved authentication accuracy of the proposed scheme is 100%, as measured by the validation rate of the legitimate nodes. This highlights the superior performance of the proposed scheme in emulating the capabilities of a hardware-based PUF while providing secure and efficient authentication in IoT networks. Raviha Khan, Hossien B. Eldeeb, Brahim Mefgouda, Omar Alhussein, Hani Saleh, Sami Muhaidat |
Comput. Secur. | 5 |
| 2025 | Chin electromyography-based motor unit decomposition for alternative screening of obstructive sleep apnea events: A comprehensive analysisabstractObstructive Sleep Apnea-Hypopnea Syndrome (OSAHS) is a prevalent sleep disorder characterized by recurrent episodes of obstructed breathing due to the relaxation of muscles in the upper airway during sleep, often linked with neuromuscular and cardiovascular disorders. This study introduces a novel method using traditional machine learning classifiers and surface electromyography (SEMG) features extracted from motor units (MUs) decomposed from chin electromyography (EMG) signals to screen for OSA events in OSAHS subjects. SEMG features were extracted from individual MUs decomposed from chin EMG segments using a novel dataset. An apnea detection algorithm was designed to label these events for OSAHS subjects across sleep stages. Analysis of motor neuron firing patterns in OSAHS subjects revealed lower activation during OSA events and higher activation during non-OSA segments. Additionally, we evaluated the proposed system on a publicly available dataset, achieving a maximum accuracy of 72% for OSAHS subjects in the midlife phase age group (40–59 years) and 72.5% for subjects in the severe phase of OSAHS using Support Vector Machines (SVM). The random forest (RF) classifier demonstrated robust performance, achieving 97% accuracy, 93.2% sensitivity, 100% specificity, 100% precision, a 96.48% F1-score, and an area under the curve (AUC) of 0.996. This system facilitates early differentiation between OSA and non-OSA events, enabling timely intervention in the mild apnea phase to prevent progression to severe OSAHS. Moreover, it offers a convenient alternative to conventional polysomnography (PSG), enhancing diagnostic accessibility and clinical management. Adil Rehman, Mostafa Moussa, Hani Saleh, Ali Khraibi, Ahsan H. Khandoker |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Efficient NTT/INTT processor for FALCON post-quantum cryptographyabstractFALCON is a lattice-based post-quantum cryptographic (PQC) digital signature standard known for its compact signatures and resistance to quantum attacks. Since its recent standardization, its hardware implementation remains an open challenge, particularly for key generation, which is significantly more complex than the simple and well-studied signature verification process. In this paper, targeting edge devices with constrained resources, we present an energy-efficient and area-optimized NTT/INTT architecture tailored to the specific requirements of FALCON key generation. By leveraging NTT-friendly primes and reducing the size of the multipliers in the Montgomery reduction algorithm — optimized for ASIC implementation — our design minimizes hardware complexity, achieving the lowest power and area consumption compared to state-of-the-art Montgomery reduction implementations. The proposed hardware architecture features a processing element array, distributed SRAMs, and ROMs, with three levels of reconfigurability, supporting both NTT and INTT operations. Designed using the Global Foundries’ 22 nm FD-SOI process, an Application-Specific Integrated Circuit (ASIC) is estimated to occupy 0.04 mm 2 and consume 18.2 mW at 1 GHz. The proposed processor achieves 700 times greater energy efficiency and performs computations 200 times faster than software implementations on the ARM Cortex-M4. It also achieves the lowest area–time product and highest energy efficiency among state-of-the-art NTT/INTT hardware accelerators. By carefully balancing power consumption and computational speed, this design offers an efficient solution for deploying FALCON key generation on devices with limited resources. Ghada Alsuhli, Hani Saleh, Mahmoud Al-Qutayri, Baker Mohammad, Thanos Stouraitis |
J. Inf. Secur. Appl. | 2 |
| 2025 | A Survey and Comparative Analysis of Number Systems for Deep Neural NetworksabstractDeep neural networks (DNNs) are indispensable in various artificial intelligence (AI) applications. However, their inherent complexity presents significant challenges, particularly when deploying them on resource-constrained devices. To overcome these hurdles, academia and industry are actively seeking ways to accelerate and optimize DNN implementations. A significant area of research revolves around discovering more effective methods to represent the enormous data volumes processed by DNNs. Traditional number systems (NSs) have proven nonoptimal for this task, prompting extensive exploration into alternative and bespoke systems for DNNs. This survey aims to comprehensively discuss various NSs utilized to efficiently represent DNN data. These systems are categorized mainly based on their impact on DNN performance and hardware implementation. This survey offers an overview of these categorized NSs and delves into different subsystems within each, outlining their effect on DNN performance and hardware design. Furthermore, these systems are compared quantitatively and qualitatively concerning their expected quantization error, memory utilization, and computational requirements. This survey also emphasizes the challenges linked with each system and the diverse proposed solutions to address them. Insights into the utilization of these NSs for sophisticated DNNs are also presented in this survey. Readers will acquire a deeper understanding of the importance of efficient NSs for DNNs, explore commonly used systems, comprehend the tradeoffs between these systems, delve into design considerations influencing their impact on DNN performance, and discover recent trends and potential research avenues in this field. Ghada Alsuhli, Vasilis Sakellariou, Hani Saleh, Mahmoud Al-Qutayri, Baker Mohammad, Thanos Stouraitis |
Proc. IEEE | 3 |
| 2024 | DRAM-Based PUF Utilizing the Variation of Adjacent CellsabstractThe Physical Unclonable Function (PUF) is a security mechanism that takes advantage of the physical variations in a device to create a unique response that can be used as a device signature or secure key. However, many DRAM-based PUFs violate the operating rules of commodity DRAM to exploit a source of entropy in the DRAM read path. This work proposes a fast and reliable DRAM-based PUF that evaluates the variation of adjacent cells and produces the response through the normal read operation. The proposed design is implemented using 65-nm technology, and a detailed statistical SPICE simulation verifies its validity. The statistical analysis shows that the proposed PUF achieves 54.19% uniformity and 49.43% uniqueness, with 98% of the investigated responses achieving a Shannon entropy of 0.95. Additionally, the proposed design generates the response by 45, which is at least 66.7 times faster than existing systems. Furthermore, the proposed design uses the relative behavior of cells, which allows for stable responses against temperature and voltage variations, eliminating the need for error correction codes. The proposed PUF also shows resiliency against machine learning-based modeling attacks, as the prediction accuracy does not exceed 55% over 5K Challenge-Response Pairs (CRPs). The area overhead is negligible as the proposed design uses standard circuits, with the addition of only one 2x1 multiplexer at the inputs of the row buffer. Enas E. Abulibdeh, Leen Younes, Baker Mohammad, Khaled Humood, Hani Saleh, Mahmoud Al-Qutayri |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | A multiplier-Free RNS-Based CNN accelerator exploiting bit-Level sparsityabstractIn this work, a Residue Numbering System (RNS)-based Convolutional Neural Network (CNN) accelerator utilizing a multiplier-free distributed-arithmetic Processing Element (PE) is proposed. A method for maximizing the utilization of the arithmetic hardware resources is presented. It leads to an increase of the system's throughput, by exploiting bit-level sparsity within the weight vectors. The proposed PE design takes advantage of the properties of RNS and Canonical Signed Digit (CSD) encoding to achieve higher energy efficiency and effective processing rate, without requiring any compression mechanism or introducing any approximation. An extensive design space exploration for various parameters (RNS base, PE micro-architecture, encoding) using analytical models as well as experimental results from CNN benchmarks is conducted and the various trade-offs are analyzed. A complete end-to-end RNS accelerator is developed based on the proposed PE. The introduced accelerator is compared to traditional binary and RNS counterparts as well as to other state-of-the-art systems. Implementation results in a 22-nm process show that the proposed PE can lead to 1.85× and 1.54× more energy-efficient processing compared to binary and conventional RNS, respectively, with a 1.88× maximum increase of effective throughput for the employed benchmarks. Compared to a state-of-the-art, all-digital, RNS-based system, the proposed accelerator is 8.87× and 1.11× more energy- and area-efficient, respectively. Vasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis |
ARITH | 4 |
| 2023 | EACNN: Efficient CNN Accelerator Utilizing Linear Approximation and Computation ReuseabstractThis paper proposes an efficient hardware accelerator named EACNN for use in Convolution Neural Networks. EACNN is an efficient CNN architecture that is based on co-optimization of algorithms and hardware. The proposed approach is based on linear approximation of the weights for pre-trained networks with low loss of accuracy. Furthermore, a weight substitution and remapping technique adopts linear approximation coefficients to replace CNN weights. That leads to a repetition of the weight values across different kernels and enables the reuse of CNN computations for various output feature maps. The input activations corresponding to the same linear co-efficient can be multiplied and accumulated first and then reused to generate multiple output feature maps. This computational reuse method reduces the number of multiplication and addition operations and memory accesses, which is efficiently supported by a dedicated element in the proposed EACNN. Experimental results on CIFAR 10 and CIFAR 100 datasets show that the proposed method eliminates around 61% of the multiplications in the network without significant loss of accuracy$(< 3\%)$. As a demonstration, a hardware accelerator based on EACNN was implemented on Xilinx FPGA Artix 7 and achieved a 50% reduction in the FPGA hardware resources. Mohamed F. Tolba 0002, Hani Saleh, Baker Mohammad, Mahmoud Al-Qutayri, Thanos Stouraitis |
ISCAS | 2 |
| 2022 | Reduce Computing Complexity of Deep Neural Networks Through Weight ScalingabstractLarge deep neural network (DNN) models are computation and memory intensive, which limits their deployment especially on edge devices. Therefore, pruning, quantization, data sparsity and data reuse have been applied to DNNs to reduce memory and computation complexity at the expense of some accuracy loss. The reduction in the bit-precision results in loss of information, and the aggressive bit-width reduction could result in noticeable accuracy loss. This paper introduces Scaling-Weight-based Convolution (SWC) technique to reduce the DNN model size and the complexity and number of arithmetic operations. This is achieved by, using a small set of high-precision weights (maximum absolute weight “MAW”) and a large set of low-precision weights (Scaling weights “SWs”). This results in decreasing the model size with minimum loss in accuracy compared to simply reducing the precision. Moreover, a scaling and quantized network-acceleration processor (SQNAP) is proposed based on the SWC method to achieve high-speed and low-power with reduced memory accesses. The proposed SWC eliminate >90% of the multiplications in the network. Moreover, the less important SWs are pruned, which has a small portion of the MAW. Retraining is applied in order to maintain accuracy. Full analysis for MNIST, Fashion MNIST, Cifar 10 and Cifar 100 datasets is presented for image recognition, where different DNN models are used including LeNet, ResNet, AlexNet and VGG 16. Mohamed F. Tolba 0002, Hani Saleh, Mahmoud Al-Qutayri, Baker Mohammad |
ISCAS | 2 |
| 2022 | A High-performance RNS LSTM blockabstractThe Residue Number System (RNS) has been proposed as an alternative to conventional binary representations for use in AI hardware accelerators. While it has been successfully utilized in applications targeting Convolutional Neural Networks (CNNs), its usage in other network models such as Recurrent Neural Networks (RNNs) has been set back due to the difficulty of implementing more complex activations functions like tanh and sigmoid ($\sigma$) in the RNS domain. In this paper, we seek to extend its usage in such models, and in particular LSTM networks, by providing efficient RNS implementations of the activation functions. To this aim, we derive improved accuracy piecewise linear approximations of the tanh and $\sigma$ functions using the minimax approach and propose a fully RNS-based hardware realization. We show that our approximations can effectively mitigate accuracy degradation in LSTM networks compared to naive approximations, while the RNS LSTM block can be up to 40% more efficient in terms of performance per area unit compared to a binary counterpart, when used in high performance-targeted accelerators. Vasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis |
ISCAS | 4 |
| 2022 | GNN-RE: Graph Neural Networks for Reverse Engineering of Gate-Level NetlistsabstractThis work introduces a generic, machine learning (ML)-based platform for functional reverse engineering (RE) of circuits. Our proposed platformGNN-REleverages the notion of graph neural networks (GNNs) to: 1) represent and analyze flattened/unstructured gate-level netlists; 2) automatically identify the boundaries between the modules or subcircuits implemented in such netlists; and 3) classify the subcircuits based on their functionalities. For GNNs in general, each graph node is tailored to learn about its own features and its neighboring nodes, which is a powerful approach for the detection of any kind of subgraphs of interest. ForGNN-RE, in particular, each node represents a gate and is initialized with a feature vector that reflects on the functional and structural properties of its neighboring gates.GNN-REalso learns the global structure of the circuit, which facilitates identifying the boundaries between subcircuits in a flattened netlist. Initially, to provide high-quality data for training ofGNN-RE, we deploy a comprehensive dataset of foundational designs/components with differing functionalities, implementation styles, bit widths, and interconnections.GNN-REis then tested on the unseen shares of this custom dataset, as well as the EPFL benchmarks, the ISCAS-85 benchmarks, and the 74X series benchmarks.GNN-REachieves an average accuracy of 98.82% in terms of mapping individual gates to modules, all without any manual intervention or postprocessing. We also release our code and source data. Lilas Alrahis, Abhrajit Sengupta, Johann Knechtel, Satwik Patnaik, Hani Saleh, Baker Mohammad, Mahmoud Al-Qutayri, Ozgur Sinanoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | GNNUnlock: Graph Neural Networks-based Oracle-less Unlocking Scheme for Provably Secure Logic LockingabstractLogic locking is a holistic design-for-trust technique that aims to protect the design intellectual property (IP) from untrustworthy entities throughout the supply chain. Functional and structural analysis-based attacks successfully circumvent state-of-the-art, provably secure logic locking (PSLL) techniques. However, such attacks are not holistic and target specific implementations of PSLL. Automating the detection and subsequent removal of protection logic added by PSLL while accounting for all possible variations is an open research problem. In this paper, we propose GNNUnlock, the first-of-its-kind oracle-less machine learning-based attack on PSLL that can identify any desired protection logic without focusing on a specific syntactic topology. The key is to leverage a well-trained graph neural network (GNN) to identify all the gates in a given locked netlist that belong to the targeted protection logic, without requiring an oracle. This approach fits perfectly with the targeted problem since a circuit is a graph with an inherent structure and the protection logic is a sub-graph of nodes (gates) with specific and common characteristics. GNNs are powerful in capturing the nodes' neighborhood properties, facilitating the detection of the protection logic. To rectify any misclassifications induced by the GNN, we additionally propose a connectivity analysis-based post-processing algorithm to successfully remove the predicted protection logic, thereby retrieving the original design. Our extensive experimental evaluation demonstrates that GNNUnlock is 99.24% - 100% successful in breaking various benchmarks locked using stripped-functionality logic locking [1], tenacious and traceless logic locking [2], and Anti-SAT [3]. Our proposed post-processing enhances the detection accuracy, reaching 100% for all of our tested locked benchmarks. Analysis of the results corroborates that GNNUnlock is powerful enough to break the considered schemes under different parameters, synthesis settings, and technology nodes. The evaluation further shows that GNNUnlock successfully breaks corner cases where even the most advanced state-of-the-art attacks [4], [5] fail. We also open source our attack framework [6]. Lilas Alrahis, Satwik Patnaik, Faiq Khalid, Muhammad Abdullah Hanif, Hani Saleh, Muhammad Shafique 0001, Ozgur Sinanoglu |
DATE | 5 |
| 2021 | A 1: 4 Active Power Divider for 5G Phased-Array Transmitters in 22nm CMOS FDSOIabstractA CMOS broadband 1:4 active power divider is proposed in this paper. The power splitter can be used in Phased Array Transceivers at the Transmitter side. It is based on a cascode 1:4 current splitter and transmission lines. Compared to other passive power dividers and active power dividers, the proposed design exhibits 1-2 dB power gain and smaller area. The measured input 1dB compression point is 6 dB whereas the IIP3 is 7.2 dBm. The Noise Figure of the Power Divider is 10 dB at lower frequencies and 15 dB at 28 GHz. The measured results are performed across several chips. Realized in 22nm CMOS FDSOI from GF, the total power consumption is 29 mW from a 1 V power supply and the area occupied by the divider is 700μm × 600μm. A thorough analysis of the gain and noise of the divider is presented as well. Nourhan Elsayed, Hani Saleh, Ademola Mustapha, Baker Mohammad, Mihai Sanduleanu |
ISCAS | 2 |
| 2021 | UNSAIL: Thwarting Oracle-Less Machine Learning Attacks on Logic LockingabstractLogic locking aims to protect the intellectual property (IP) of integrated circuit (IC) designs throughout the globalized supply chain. The SAIL attack, based on tailored machine learning (ML) models, circumvents combinational logic locking with high accuracy and is amongst the most potent attacks as it does not require a functional IC acting as an oracle. In this work, we propose UNSAIL, a logic locking technique that inserts key-gate structures with the specific aim to confuse ML models like those used in SAIL. More specifically, UNSAIL serves to prevent attacks seeking to resolve the structural transformations of synthesis-induced obfuscation, which is an essential step for logic locking. Our approach is generic; it can protect any local structure of key-gates against such ML-based attacks in an oracle-less setting. We develop a reference implementation for the SAIL attack and launch it on both traditionally locked and UNSAIL-locked designs. For SAIL, two ML models have been proposed (which we implement accordingly), namely a change-prediction model and a reconstruction model; the change-prediction model is used to determine which key-gate structures to restore using the reconstruction model. Our study on benchmarks ranging from the ISCAS-85 and ITC-99 suites to the OpenRISC Reference Platform System-on-Chip (ORPSoC) confirms that UNSAIL degrades the accuracy of the change-prediction model and the reconstruction model by an average of 20.13 and 17 percentage points (pp), respectively. When the aforementioned models are combined, which is the most powerful scenario for SAIL, UNSAIL reduces the attack accuracy of SAIL by an average of 11pp. We further demonstrate that UNSAIL thwarts other oracle-less attacks, i.e., SWEEP and the redundancy attack, indicating the generic nature and strength of our approach. Detailed layout-level evaluations illustrate that UNSAIL incurs minimal area and power overheads of 0.26% and 0.61%, respectively, on the million-gate ORPSoC design. Lilas Alrahis, Satwik Patnaik, Johann Knechtel, Hani Saleh, Baker Mohammad, Mahmoud Al-Qutayri, Ozgur Sinanoglu |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | A 28GHz, Asymmetrical, Modified Doherty Power Amplifier, in 22nm FDSOI CMOSabstractA 28GHz, Modified Doherty Power Amplifier (MDPA) was implemented in 22nm FDSOI CMOS technology from GF. The MDPA adopts an asymmetrical topology utilizing two cascode CMOS amplifiers as the main (Class-A) and auxiliary (Class-C). This allows a supply voltage of 2.5V and consequently higher output power. The use of a main Class-A amplifier is conducive to a higher linearity (IIP3). The integrated design implements the main and auxiliary amplifier, along with the matching and transmission line networks on chip. The fabricated amplifier occupies an area of 1.2mm2, exhibits 12 dBm saturated output power, a peak power gain of 10dB, 16% peak power-added efficiency (PAE) and 12.5% at 6-dB back-off. The measured IIP3 is 20dBm. Nourhan Elsayed, Hani Saleh, Baker Mohammad, Mihai Sanduleanu |
ISCAS | 2 |
| 2020 | ASIC Implementation of a Pre-Trained Neural Network for ECG Feature ExtractionabstractThe electrocardiogram signal (ECG), a record of electrical activity of the cardiac muscle, has been used in diagnosing many cardiopathies. Wearable devices equipped with readout sensors and circuits can be used to record and process weak ECG signals. In this paper, a pre-trained neural network was implemented for detecting the QRS feature of an ECG signal, which is crucial for auto-diagnostic of various cardiopathies. To take advantage of the fast evolution of artificial intelligence and its ability to find non-linear relationships, neural network based feature extraction of ECG signals for wearable devices was explored and tested using ASIC implementation flow. Firstly, a high-level simulation was carried out in MATLAB and verified with test data obtained from PhysioNET database. Recurrent neural network (RNN) MLP was created and trained using the data obtained from PhysioNET database. A high-level performance evaluation was carried out using the same network for P and T wave extraction. The weight and bias matrices obtained from the high-level trained network in MATLAB were used in the design of the hardware. An accuracy of 96.55% was achieved in the hardware implementation of the network. Huruy Tekle Tefai, Hani Saleh, Temesghen Tekeste, Mahmoud Al-Qutayri, Baker Mohammad |
ISCAS | 2 |
| 2019 | ScanSAT: unlocking obfuscated scan chainsabstractWhile financially advantageous, outsourcing key steps such as testing to potentially untrusted Outsourced Semiconductor Assembly and Test (OSAT) companies may pose a risk of compromising on-chip assets. Obfuscation of scan chains is a technique that hides the actual scan data from the untrusted testers; logic inserted between the scan cells, driven by a secret key, hide the transformation functions between the scan-in stimulus (scan-out response) and the delivered scan pattern (captured response). In this paper, we propose ScanSAT: an attack that transforms a scan obfuscated circuit to its logic-locked version and applies a variant of the Boolean satisfiability (SAT) based attack, thereby extracting the secret key. Our empirical results demonstrate that ScanSAT can easily break naive scan obfuscation techniques using only three or fewer attack iterations even for large key sizes and in the presence of scan compression. Lilas Alrahis, Muhammad Yasin, Hani Saleh, Baker Mohammad, Mahmoud Al-Qutayri, Ozgur Sinanoglu |
ASP-DAC | 3 |
| 2019 | Functional Reverse Engineering on SAT-Attack Resilient Logic LockingabstractLogic locking is a solution that mitigates hardware security threats, such as Trojan insertion, piracy and counterfeiting. Research in this area has led to, in an iterative fashion, a series of logic locking defenses as well as attacks that circumvent these defenses by extracting the logic locking key. The most powerful attacks rely on a full access to a working chip/oracle that can be used to produce the input-output pairs utilized in recovering the secret key. A recently proposed technique Stripped Functionality Logic Locking (SFLL) provides resilience to all known attacks on combinational logic locking. In this paper, we propose a functional reverse engineering attack on SFLL: an attack that can detect the protection logic of SFLL which results in obtaining the original unlocked design with a high success rate. The restore and perturb blocks utilized by SFLL were detected with average coverage percentages of 93.95% and 85.42% respectively, proving that our attack is capable of breaking the state of the art logic locking technique. Lilas Alrahis, Muhammad Yasin, Hani Saleh, Baker Mohammad, Mahmoud Al-Qutayri |
ISCAS | 3 |
| 2019 | A Gain-Controlled, Low-Leakage Dickson Charge Pump for Energy-Harvesting ApplicationsabstractThis paper presents a single-stage power management unit to boost and regulate a low supply voltage for CMOS system-on-chip (SoC) applications. It consists of low-leakage, enhanced Dickson charge pump (DCP) that utilizes both stage and frequency modulation (FM) techniques to achieve high efficiency and lower area. In addition, the proposed design uses an enhanced stage-switch structure for the charge pump, which significantly reduces the cross-stage leakage. A stage number controller is used to control the gain of the charge pump by changing the number of stages based on the desired output voltage. FM is utilized to further fine-tune the output voltage through a closed-loop control based on a predetermined reference voltage. Silicon measurement results for the four-stage charge pump in 65-nm CMOS technology show a maximum end-to-end efficiency of 66% at an input voltage of 0.7 V and an output power of$27~\mu \text{W}$. The proposed design achieved more than a$100\times $reduction in leakage compared to traditional DCP. The system supports a range of load currents between 0.1 and$34~\mu \text{A}$with a maximum operating frequency of 1.8 MHz. The proposed system supports an input voltage range of 0.55–0.7 V which makes it an excellent candidate for solar and thermal energy-harvesting applications targeting low-power internet-of-things SOC. Abdulqader Nael Mahmoud, Mohammad Alhawari, Baker Mohammad, Hani Saleh, Mohammed Ismail 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | An Efficient and Small Area Multioutput Switched Capacitor Buck Converter for IoTsabstractThis paper presents an area and power efficient multioutput switched capacitor (MOSC) DC-DC buck converter targeting ultra-low power and IoT devices. The MOSC converter has a variable input voltage range between 1.05V to 1.4V and generates regulated simultaneous multiple output voltage levels of 1V and 0.55V. Adaptive digital time multiplexing controller is employed to enable multiple power domains. In addition, a pulse frequency modulation is utilized to regulate the output voltages over a wide range of load current. The proposed converter supports a load current of 10μA to 350μA and to 10μA at a load voltage of 1V and 0.55V, respectively. Adaptive time multiplexing and pulse width modulation are implemented through a finite state machine to eliminate the reverse current issue. This problem arises during the switching from a low load voltage of 0.55V to a high load voltage of 1V. The MOSC circuit is fabricated in 65nm CMOS technology and it occupies an active area of 0.27mm2. Moreover, both MIM and MOS capacitors are utilized to further reduce the area of MOSC converter. Measured results shows that the peak efficiency of 78% is achieved at a load power of 300μW. Dima Kilani, Mohammad Alhawari, Baker Mohammad, Hani Saleh, Mohammed Ismail 0001 |
ISCAS | 4 |
| 2018 | A Charge Pump Based Power Management Unit With 66%-Efficiency in 65 nm CMOSabstractThis paper presents a single stage power management unit that includes an enhanced stage-switch Dickson charge pump (DCP) to boost and regulate a low input voltage. A new switching mechanism is presented to significantly reduce the losses encountered in conventional DCP switches. Frequency and stage modulation are utilized in the proposed design. The stage modulation provides different gain levels (coarse) and the frequency modulation tunes the voltage level and regulates the output voltage based on a pre-determined reference voltage. Using four stages charge pump, silicon measurement results in 65 nm CMOS technology show a maximum efficiency of 66% at input voltage of 0.7 V and output power of 27 μW. The system supports a range of load current between 0.1 μA − 34 μA with a maximum operating frequency of 1.8MHz. The proposed system supports an input voltage range from 0.55 to 0.7 V which can be used in energy harvesting applications such as solar and thermal harvesting. Abdulqader Nael Mahmoud, Mohammad Alhawari, Baker Mohammad, Hani Saleh, Mohammed Ismail 0001 |
ISCAS | 4 |
| 2017 | A sub-μW bio-potential front end in 65nm CMOSabstractA bio-potential amplifier intended for continuous monitoring of vitals characterized by its long operational lifetime is required to operate at the lowest power budget possible. Moreover, a compact active area directly related to portability is essential. This paper presents a 0.55pW auto gain controlled biopotential amplifier implemented in 65nm 1P7M CMOS for ECG signal classifier SoC. A chopper-stabilized amplifier is designed at 0.6V supply voltage to mitigate the DC offset and near DC flicker noise. The input ECG signal level is further set by the four gain levels of the variable gain amplifier (VGA) to provide maximum swing to the ADC. The whole system is integrated into a core are of 0.10mm2and can operate at a wide range of 0.6-1.2V supply voltage. Yonatan Kifle, Hani Saleh, Baker Mohammad, Mohammed Ismail 0001 |
VLSI-SoC | 2 |
| 2016 | An efficient thermal energy harvesting and power management for μWatt wearable BioChipsabstractThis paper presents an efficient thermal energy harvesting IC (EHIC) that supports a battery-less μWatt system-on-chips. The EHIC consists of an inductor-based DC-DC converter that boosts a low input voltage to a suitable output voltage level. Further, a switched capacitor buck converter is utilized to regulate the boost converter output voltage and to support multiple output voltage levels, namely 0.6V, 0.8V and 1V. In low energy mode and to enhance the efficiency, the EHIC is capable of bypassing the switched capacitor so that the load is driven directly from the boost converter. The prototype chip is fabricated in 65nm CMOS and occupies an area of less than 0.46mm2. Measured results confirm an efficiency of 65% at 0.6V output voltage and 42μW. In addition, the end-to-end peak efficiency is 71% at 0.8V output voltage and 182μW. Mohammad Alhawari, Dima Kilani, Baker Mohammad, Hani Saleh, Mohammed Ismail 0001 |
ISCAS | 4 |
| 2016 | A biomedical SoC architecture for predicting ventricular arrhythmiaabstractElectrocardiography (ECG) represents the hearts electrical activity and has features such as QRS complex, P-wave and T-wave that provide critical clinical information for detection and prediction of cardiac diseases. This paper presents a novel ECG processing architecture for the prediction of ventricular arrhythmia (VA). The architecture implements a novel ECG feature extraction which is optimized for ultra-low power applications. The architecture is based on Curve Length Transform (CLT) for the detection of QRS complex and Discrete Wavelet Transform (DWT) for the delineation of TP waves. Features extracted from two consecutive ECG cycles are used to set innovative parameters for VA prediction up to 3 hours before VA onset. Two databases of the heart signal recordings from the American Heart Association (AHA) and the MIT PhysioNet were used as training, test and validation sets to evaluate the performance of the proposed system. Temesghen Tekeste, Hani Saleh, Baker Mohammad, Ahsan H. Khandoker, Mohammed Ismail 0001 |
ISCAS | 2 |
| 2016 | Low-Power ECG-Based Processor for Predicting Ventricular ArrhythmiaabstractThis paper presents the design of a fully integrated electrocardiogram (ECG) signal processor (ESP) for the prediction of ventricular arrhythmia using a unique set of ECG features and a naive Bayes classifier. Real-time and adaptive techniques for the detection and the delineation of the P-QRS-T waves were investigated to extract the fiducial points. Those techniques are robust to any variations in the ECG signal with high sensitivity and precision. Two databases of the heart signal recordings from the MIT PhysioNet and the American Heart Association were used as a validation set to evaluate the performance of the processor. Based on application-specified integrated circuit (ASIC) simulation results, the overall classification accuracy was found to be 86% on the out-of-sample validation data with 3-s window size. The architecture of the proposed ESP was implemented using 65-nm CMOS process. It occupied 0.112- ${\rm mm}^{2}$ area and consumed 2.78- $\mu \text{W}$ power at an operating frequency of 10 kHz and from an operating voltage of 1 V. It is worth mentioning that the proposed ESP is the first ASIC implementation of an ECG-based processor that is used for the prediction of ventricular arrhythmia up to 3 h before the onset. Nourhan Bayasi, Temesghen Tekeste, Hani Saleh, Baker Mohammad, Ahsan H. Khandoker, Mohammed Ismail 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Modeling and Optimization of Memristor and STT-RAM-Based Memory for Low-Power ApplicationsabstractConventional charge-based memory usage in low-power applications is facing major challenges. Some of these challenges are leakage current for static random access memory (SRAM) and dynamic random access memory (DRAM), additional refresh operation for DRAM, and high programming voltage for Flash. In this paper, two emerging resistive random access memory (ReRAM) technologies are investigated, memristor and spin-transfer torque (STT)-RAM, as potential universal memory candidates to replace traditional ones. Both of these nonvolatile memories support zero leakage and low-voltage operation during read access, which makes them ideal for devices with long sleep time. To date, high write energy for both memristor and STT-RAM is one of the major inhibitors for adopting the technologies. The primary contribution of this paper is centered on addressing the high write energy issue by trading off retention time with noise margin. In doing so, the memristor and STT-RAM power has been compared with the traditional six-transistor-SRAM-based memory power and potential application in wireless sensor nodes is explored. This paper uses 45-nm foundry process technology data for SRAM and physics-based mathematical models derived from real devices for memristor and STT-RAM. The simulations are conducted using MATLAB and the results show a potential power savings of 87% and 77% when using memristor and STT-RAM, respectively, at 1% duty cycle. Yasmin Halawani, Baker Mohammad, Dirar Homouz, Mahmoud Al-Qutayri, Hani Saleh |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2015 | A hardware accelerator for real-time extraction of the linear-time MSER algorithmabstractThis paper presents a novel hardware accelerator architecture for the linear-time Maximally Stable Extremal Regions (MSER) detector algorithm. In contrast to the standard MSER algorithm, the linear-time MSER implementation is more suitable for real-time applications of image retrieval in large-scale and high resolution datasets (e.g. satellite images). The linear-time MSER accelerator design is optimized by enhancing its flooding process (which is one of the major drawbacks of the standard linear-time MSER) using a structure that we called stack of pointers, which makes it memory-efficient as it reduces the memory requirement by nearly 90%. The accelerator is configurable and can be integrated with many image processing algorithms, allowing a wide spectrum of potential real-time applications to be realized even on small and power-limited devices. Sohailah Alyammahi, Ehab Salahat, Hani Saleh, Andrzej Stefan Sluzek |
IECON | 3 |
| 2015 | A maximally stable extremal regions system-on-chip for real-time visual surveillanceabstractThis paper presents a novel implementation of the Maximally Stable Extremal Regions (MSER) detector on system-on-chip (SoC) using 65 nm CMOS technology. The novel SoC was developed following the Application Specific Integrated Circuit (ASIC) design flow which significantly enhanced its realization and fabrication, and overall performances. The SoC has very low area requirement (around 0.05 mm2) and is capable of detecting both bright and dark MSERs in a single run, while computing simultaneously their associated regions' moments, simplifying its interfacing with other image algorithms (e.g. SIFT and SURF). The novel MSER SoC is power-efficient (requires 2.25 mW) and memory-efficient as it saves more than 31% of the memory space reported in the state-of-the-art MSER implementation on FPGA, making it suitable for mobile devices. With 256×256 resolution and its operating frequency of 133 MHz, the SoC is expected to have a 200 frames/second processing rate, making it suitable (when integrated with other algorithms in the system) for time-critical real-time applications such as visual surveillance. Ehab Salahat, Hani Saleh, Andrzej Stefan Sluzek, Mahmoud Al-Qutayri, Baker Mohammad, Mohammed Ismail 0001 |
IECON | 2 |
| 2015 | Novel fast and scalable parallel union-find ASIC implementation for real-time digital image segmentationabstractThis paper presents a new fast and scalable Parallel Union-Find algorithm for image segmentation and its System-on-Chip (SoC) implementation using 65nm CMOS technology following the Application-Specific Integrated Circuit (ASIC) design flow. The algorithm is capable of labeling all foreground and background pixels, using the least possible pixels scanning. This contrasts the classical labeling algorithms that label only foreground (or background) pixels in a single run. The new algorithm utilizes only two memory blocks. In one memory block, it labels image segments using their seeds as the label and, simultaneously, the segments sizes are used as the other label in second memory block. By this parallel labeling, monitoring the image segments is very fast and efficient. With 350 MHz operating frequency, the processing rate estimated to be 2100 frames/sec, the total chip area of 15950.5 μm2 (off-chip memory) and very low-power of 0.3 mW, the SoC tends to be an excellent candidate for mobile devices and real-time applications. Ehab Salahat, Hani Saleh, Andrzej Stefan Sluzek, Mahmoud Al-Qutayri, Baker Mohammad, Mohammed Ismail 0001 |
IECON | 2 |
| 2015 | Novel MSER-guided street extraction from satellite imagesabstractThe paper presents a novel technique to segment and extract streets from satellite images. This technique utilizes, for the first time in the known literature, the Maximally Stable Extremal Regions (MSER) algorithm to robustly identify and segment streets from satellite images. The technique extracts dark MSERs and then classifies them based on multiple metrics such as the intensity of the pixels, the region stability, and the major-to-minor axes ratio. Testing results under multiple scenarios corroborate the accuracy of the proposed technique. The technique will allow fast and accurate implementation of a wide spectrum of applications such as in Global Positioning System (GPS) driving guidance. Ehab Salahat, Hani Saleh, Andrzej Stefan Sluzek, Baker Mohammad, Mahmoud Al-Qutayri, Mohammed Ismail 0001 |
IGARSS | 2 |
| 2015 | A 65-nm low power ECG feature extraction systemabstractThis paper presents a real-time adaptive ECG detection and delineation algorithm alongside an architecture based on time-domain signal processing of the ECG signal. The algorithm is enhanced to detect large number of different P-QRS-T waveform morphologies using adaptive search windows and adaptive threshold levels. The proposed architecture has been implemented in the state-of-the-art 65-nm CMOS technology. It occupied 0.03416 mm2 area and consumed 0.614 mW power. Furthermore, the non-complex nature of the architecture resulted with a realization using smaller number of computation and higher performance. The design of the QRS detector was tested on ECG records obtained from the Physionet QT database and achieved a sensitivity of Se =99.83% and a positive predictivity of P+= 98.65%. Similarly, the mean error values of the T peak, T offset, P peak and P offset were found to be -1.367, 6.36, 5.5 and -2.59 milliseconds, respectively, using the same database. The small area, low power, and high performance of our architecture makes it suitable for inclusion in System On Chips (SOCs) targeting wearable mobile medical devices. Nourhan Bayasi, Temesghen Tekeste, Hani Saleh, Baker Mohammad, Mohammed Ismail 0001 |
ISCAS | 3 |
| 2015 | Memory impact on the lifetime of a Wireless Sensor Node using a Semi-Markov modelabstractThe increase in demand for higher functionality, smaller size, lower cost and near perpetual operation of Wireless Sensor Nodes (WSNs) are posing big challenges for system designers. A major aspect is the operational lifetime of the system which is determined by the finite energy source supplied by the battery. In WSNs, the higher power incurred due to added system functionality and the increase leakage as a result of technology scaling have high impact on the battery lifetime. In this work, memory was used to further exploit the energy efficiency at the sensor node system-level. A detailed analysis using Semi-Markov model to investigate different operational modes of WSN with the present SRAM shows an improvement of 4x at 90% duty cycle in the node's lifetime. In addition, an emerging non-volatile memory (NVM) technology, Memristor, is also explored to further improve the WSN energy efficiency. Its non-volatility nature will suppress the power wasted as leakage in SRAM during idle periods which is typical for low duty-cycle WSNs. The results show that utilizing an on-chip NVM can further improve WSN lifetime by 1x for low activity μW range sensor nodes. Yasmin Halawani, Baker Mohammad, Mahmoud Al-Qutayri, Hani Saleh |
ISCAS | 4 |
| 2015 | Adaptive ECG interval extractionabstractECG intervals such as QRS, QT and PR provide significant information and are widely used as clinical parameters for diagnosing cardiac diseases. This paper presents a novel QRS detection technique based on Curve Length Transform (CLT) and a refined delineation of P-wave and T-wave using Discrete Wavelet Transform (DWT). The proposed technique was verified using the PhysioNet database. The QRS detection achieved a sensitivity of 98.59% and a positive predictivity of 97.86%. The QRS duration, QT interval and PR interval had a mean error of -1.56± 28.8ms, -5.39± 42.4ms and 0.86± 40.3ms respectively. The proposed algorithm is computationally efficient and is simpler to implement in hardware, hence, will lead to a faster execution time, smaller design area and consequently low power consumption. Temesghen Tekeste, Nourhan Bayasi, Hani Saleh, Ahsan H. Khandoker, Baker Mohammad, Mahmoud Al-Qutayri, Mohammed Ismail 0001 |
ISCAS | 3 |
| 2015 | Novel Unified Analysis of Orthogonal Space-Time Block Codes over Generalized-K and AWGGN MIMO NetworksabstractThis paper presents a novel unified performance analysis of Space-Time Block Codes (STBCs) operating in the Multiple Input Multiple Output (MIMO) network. Specifically, we derive a novel unified expression for the average bit error rate for all coherent modulation schemes assuming independent and identically distributed generalized-K fading channels, which accounts for both multipath fading and shadowing, and modeled via the Nakagami-m distribution and the gamma distributions, respectively. The noise model in the network is assumed to be an Additive White Generalized Gaussian Noise (AWGGN), which encompasses the Laplacian and the Gaussian noise environments as special cases. The derived expression obviates the need to recalculate the diversity of various receivers for many fading and noise models in a piecemeal fashion. Published results from the literature as well as numerical integration methods corroborate the accuracy of the unified expression. Ehab Salahat, Hani Saleh |
VTC Spring | 2 |
| 2015 | Evolutionary QR-Based Traffic Sign Recognition System for Next-Generation Intelligent VehiclesabstractThis paper introduces a dramatically novel traffic signs recognition (TSR) system that can perform traffic sign detection and tracking simultaneously. The proposed approach utilizes intensity images and the depth images, in parallel, to robustly detect and track traffic signs in real-time. Additionally, we suggest to supplement the ordinary traffic signs with the corresponding quick-response (QR) code plates that inherent the many advantages of the QR-codes, introducing the concept of QR-TSR systems. Ehab Salahat, Hani Saleh, Andrzej Stefan Sluzek, Mahmoud Al-Qutayri, Baker Mohammad, Mohammed Ismail 0001 |
VTC Fall | 2 |
| 2015 | Design Methodologies for Yield Enhancement and Power Efficiency in SRAM-Based SoCsabstractThis paper comprises two new methodologies to improve yield and reduce system-on-a-chip power. The first methodology is based on faulty static random-access memory (SRAM) cells detections and cache resizing. The key advantage of this approach is that it enables the end user to control the system's parameters to be error tolerant. Furthermore, this technique enables aggressive voltage scaling which causes parametric (soft) failures in SRAM-based memory. As such, the proposed methodology can be utilized to exchange cache size for lower power or better yield. In the second methodology, data from faulty cells are treated as imposed noise. Depending on the application, this error percentage (imposed noise) can be mitigated through three options. First, ignore error if the percentage of the error is tolerable. Second, simple hardware filtration is needed. Finally, software-based filtration is required. The viability of this approach is that it allows aggressive voltage scaling below the traditional to be a 100% correct approach for SRAM supply, which results in substantial reduction of power, trading off quality for power. For both approaches, BIST is used as part of the powerup sequence to identify the faulty memory addresses per voltage level and compute the faulty cells percentage. Furthermore, the proposed methodologies help in improving reliability and counteracting long-term effects on memory cell stability and lifetime degradation caused by negative bias temperature instability. Baker Mohammad, Hani Saleh, Mohammed Ismail 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | FFT Implementation with Fused Floating-Point OperationsabstractThis paper describes two fused floating-point operations and applies them to the implementation of fast Fourier transform (FFT) processors. The fused operations are a two-term dot product and an add-subtract unit. The FFT processors use "butterfly” operations that consist of multiplications, additions, and subtractions of complex valued data. Both radix-2 and radix-4 butterflies are implemented efficiently with the two fused floating-point operations. When placed and routed using a high performance standard cell technology, the fused FFT butterflies are about 15 percent faster and 30 percent smaller than a conventional implementation. Also the numerical results of the fused implementations are slightly more accurate, since they use fewer rounding operations. Earl E. Swartzlander Jr., Hani Saleh |
IEEE Trans. Computers | 2 |
| 2009 | Digital Forensics: Electronic Evidence Collection, Examination and Analysis by Using Combine Moments in Spatial and Transform DomainabstractA novel digital forensics tool is developed by combining wavelet invariant with spatial moments. A forensic printed circuit board image matching system is presented that is capable of probing a large database of digital images of circuit boards and compare them for similarity to provide investigation leads for electronic crimes digital forensic science investigations. The developed system has been implemented, and proved to be very efficient in detection similarities between a target image and a large image database even when the target image is noisy, scaled or mirrored. Hani Saleh, Sos S. Agaian, Khader Mohammad |
SMC | 1 |
| 2008 | A floating-point fused dot-product unitabstractA floating-point fused dot-product unit is presented that performs single-precision floating-point multiplication and addition operations on two pairs of data in a time that is only 150% the time required for a conventional floating-point multiplication. When placed and routed in a 45 nm process, the fused dot-product unit occupied about 70% of the area needed to implement a parallel dot-product unit using conventional floating-point adders and multipliers. The speed of the fused dot-product is 27% faster than the speed of the conventional parallel approach. The numerical result of the fused unit is more accurate because one rounding operation is needed versus at least three for other approaches. Hani Saleh, Earl E. Swartzlander Jr. |
ICCD | 1 |
| 2008 | A family of scalable FFT architectures and an implementation of 1024-point radix-2 FFT for real-time communicationsabstractThe paper presents a family of architectures for FFT implementation based on the decomposition of the perfect shuffle permutation, which can be designed with variable number of processing elements. This provides designers with a trade-off choice of speed vs. complexity (cost and area.). A detailed case study is provided on the implementation of 1024-point FFT with 2 processing elements using 45 nm process technology, including area, timing, power and place-and-route results. Adnan Suleiman, Hani Saleh, Adel Hussein, David Akopian |
ICCD | 2 |
| 2007 | Contention-free switch-based implementation of 1024-point Radix-2 Fourier Transform EngineabstractThis paper examines the use of a switch based architecture to implement a Radix-2 decimation in frequency fast Fourier transform engine. The architecture interconnects M processing elements with 2*M memories. An algorithm to detect and resolve memory access contention is presented. The implementation of 1024-point FFTs with 2 processing elements is discussed in detail, including timing and place-and-route results. The switch based architecture provides a factor of M speedup over a single processing element realization. Hani Saleh, Bassam Jamil Mohd, Adnan Aziz, Earl E. Swartzlander Jr. |
ICCD | 1 |