VLDB 2026 Research / reviewers in the wild / expert
Khaled N. Salama
dblp:07/3956 · also Khaled Nabil Salama
· DBLP profile ↗
24ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0001-7742-1282ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 1 first-author · 7 since 2021Computer networks · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Efficient IMC Accelerator Design Through Joint Hardware-Workload Co-optimizationabstractDesigning generalized in-memory computing (IMC) hardware that efficiently supports a variety of workloads requires extensive design space exploration, which is infeasible to perform manually. Optimizing hardware individually for each workload or solely for the largest workload often fails to yield the most efficient generalized solutions. To address this, we propose a joint hardware-workload optimization framework that identifies optimised IMC chip architecture parameters, enabling more efficient, workload-flexible hardware. We show that joint optimization achieves 36%, 36%, 20%, and 69% better energy-latency-area scores for VGG16, ResNet18, AlexNet, and MobileNetV3, respectively, compared to the separate architecture parameters search optimizing for a single largest workload. Additionally, we quantify the performance trade-offs and losses of the resulting generalized IMC hardware compared to workload-specific IMC designs. Olga Krestinskaya, Mohamed E. Fouda, Ahmed M. Eltawil, Khaled N. Salama |
ISCAS | 4 |
| 2024 | A 19 fJ/op, Low-Offset StrongARM Latch Comparator for Low-Power High-Speed ApplicationsabstractIn this paper, a new StrongARM latch comparator design has been proposed for low-power high-speed applications. The proposed design improves the energy consumption and propagation delay when compared to the previous designs in the literature. The proposed design is post-layout simulated in TSMC 65 nm technology node and it achieves a low energy consumption of 19.31 fJ per operation and a low propagation delay of 211 ps. Moreover, the proposed design shows a highly favorable input offset voltage of 0.56 mV and achieves a maximum frequency of 8 GHz. Furthermore, the proposed design reduced the transistor stack that allows it to be used in the low-voltage supply application. Abdullah Alshehri 0003, Khaled N. Salama, Hossein Fariborzi |
ISCAS | 2 |
| 2023 | Resistive Neural Hardware AcceleratorsabstractDeep neural networks (DNNs), as a subset of machine learning (ML) techniques, entail that real-world data can be learned, and decisions can be made in real time. However, their wide adoption is hindered by a number of software and hardware limitations. The existing general-purpose hardware platforms used to accelerate DNNs are facing new challenges associated with the growing amount of data and are exponentially increasing the complexity of computations. Emerging nonvolatile memory (NVM) devices and the compute-in-memory (CIM) paradigm are creating a new hardware architecture generation with increased computing and storage capabilities. In particular, the shift toward resistive random access memory (ReRAM)-based in-memory computing has great potential in the implementation of area- and power-efficient inference and in training large-scale neural network architectures. These can accelerate the process of IoT-enabled AI technologies entering our daily lives. In this survey, we review the state-of-the-art ReRAM-based DNN many-core accelerators, and their superiority compared to CMOS counterparts was shown. The review covers different aspects of hardware and software realization of DNN accelerators, their present limitations, and prospects. In particular, a comparison of the accelerators shows the need for the introduction of new performance metrics and benchmarking standards. In addition, the major concerns regarding the efficient design of accelerators include a lack of accuracy in simulation tools for software and hardware codesign. Kamilya Smagulova, Mohamed E. Fouda, Fadi J. Kurdahi, Khaled N. Salama, Ahmed M. Eltawil |
Proc. IEEE | 4 |
| 2023 | Architectural Trade-Off Analysis for Accelerating LSTM Network Using Radix-r OBC SchemeabstractThis paper presents architectural trade-off analysis for accelerating two (Type I, II) fixed-point long short-term memory (LSTM) network based on circulant matrix-vector multiplications (MVMs) using radix-$r$offset binary coding (OBC) scheme. Type I MVM architecture rotates the weights with the proposed modulo-cum interleaver and uses partial product generators (PPGs) with a single generation unit across a column. It is hardware-optimized using a single adder tree through time-multiplexing. Meanwhile, Type II MVM architecture rotates the inputs with the proposed store-cum interleaver and uses single PPGs with a single generation unit across a row. It is time-optimized by unfolding shift-accumulate unit to a shift-add tree followed by pipelining. A new design for element-wise multiplication using radix-$r$PPG is also presented. Both the designs are extended to their block-circulant variants for certain accuracy requirements. Post-synthesis of Type I and II architectures for a different model, kernel, radix sizes and clock frequencies result in several efficient designs. Compared with the prior scheme, Type I architecture for$128 \times 128$with$r=2$on 28 nm FDSOI technology at 800 MHz occupies 32.27% lesser area, consumes 67.89% lesser power at the same throughput, while Type II architecture at the expense of area and power provides$40\times $higher throughput. Mohd. Tasleem Khan, Hasan Erdem Yantir, Khaled N. Salama, Ahmed M. Eltawil |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | Analog Image Denoising with an Adaptive Memristive Crossbar NetworkabstractNoise in image sensors led to the development of a whole range of denoising filters. A noisy image can become hard to recognize and often require several types of post-processing compensation circuits. This paper proposes an adaptive denoising system implemented using analog in-memory neural computing network. The proposed method can learn new noises and can be integrated into or alone with CMOS image sensors. Three denoising network configurations are implemented, namely, (1) single layer network, (2) convolution network, and (3) fusion network. The single layer network shows the processing time, energy consumption and on-chip area of 3.2$\mu$s, 21n J per image and 0.3mm2respectively, meanwhile, convolution denoising network correspondingly shows 72m s, 236$\mu$J and 0.48mm2. Among all the implemented networks, it is observed that performance metrics SSIM, MSE and PSNR show a maximum improvement of 3.61, 21.7 and 7.7 times respectively. Olga Krestinskaya, Khaled N. Salama, Alex James 0001 |
ISCAS | 2 |
| 2022 | The Internet of Bodies: A Systematic Survey on Propagation Characterization and Channel ModelingabstractThe Internet of Bodies (IoBs) is an imminent extension to the vast Internet of Things domain, where interconnected devices (e.g., worn, implanted, embedded, swallowed, etc.) are located in-on-and-around the human body form a network. Thus, the IoB can enable a myriad of services and applications for a wide range of sectors, including medicine, safety, security, wellness, entertainment, to name but a few. Especially, considering the recent health and economic crisis caused by the novel coronavirus pandemic, also known as COVID-19, the IoB can revolutionize today’s public health and safety infrastructure. Nonetheless, reaping the full benefit of IoB is still subject to addressing related risks, concerns, and challenges. Hence, this survey first outlines the IoB requirements and related communication and networking standards. Considering the lossy and heterogeneous dielectric properties of the human body, one of the major technical challenges is characterizing the behavior of the communication links in-on-and-around the human body. Therefore, this article presents a systematic survey of channel modeling issues for various link types of human body communication (HBC) channels below 100 MHz, the narrowband (NB) channels between 400 and 2.5 GHz, and ultrawideband (UWB) channels from 3 to 10 GHz. After explaining bio-electromagnetics attributes of the human body, physical, and numerical body phantoms are presented along with electromagnetic propagation tool models. Then, the first-order and the second-order channel statistics for NB and UWB channels are covered with a special emphasis on body posture, mobility, and antenna effects. For capacitively, galvanically, and magnetically coupled HBC channels, four different channel modeling methods (i.e., analytical, numerical, circuit, and empirical) are investigated, and electrode effects are discussed. Finally, interested readers are provided with open research challenges and potential future research directions. Abdulkadir Celik, Khaled N. Salama, Ahmed M. Eltawil |
IEEE Internet Things J. | 2 |
| 2022 | A hardware/software co-design methodology for in-memory processors
Hasan Erdem Yantir, Ahmed M. Eltawil, Khaled N. Salama |
J. Parallel Distributed Comput. | 3 |
| 2022 | Toward the Optimal Design and FPGA Implementation of Spiking Neural NetworksabstractThe performance of a biologically plausible spiking neural network (SNN) largely depends on the model parameters and neural dynamics. This article proposes a parameter optimization scheme for improving the performance of a biologically plausible SNN and a parallel on-field-programmable gate array (FPGA) online learning neuromorphic platform for the digital implementation based on two numerical methods, namely, the Euler and third-order Runge-Kutta (RK3) methods. The optimization scheme explores the impact of biological time constants on information transmission in the SNN and improves the convergence rate of the SNN on digit recognition with a suitable choice of the time constants. The parallel digital implementation leads to a significant speedup over software simulation on a general-purpose CPU. The parallel implementation with the Euler method enables around 180× ( 20× ) training (inference) speedup over a Pytorch-based SNN simulation on CPU. Moreover, compared with previous work, our parallel implementation shows more than 300× ( 240× ) improvement on speed and 180× ( 250× ) reduction in energy consumption for training (inference). In addition, due to the high-order accuracy, the RK3 method is demonstrated to gain 2× training speedup over the Euler method, which makes it suitable for online training in real-time applications. Wenzhe Guo, Hasan Erdem Yantir, Mohamed E. Fouda, Ahmed M. Eltawil, Khaled N. Salama |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Efficient Neuromorphic Hardware Through Spiking Temporal Online Local LearningabstractLocal learning schemes have shown promising performance in spiking neural networks (SNNs) training and are considered a step toward more biologically plausible learning. Despite many efforts to design high-performance neuromorphic systems, a fast and efficient on-chip training algorithm is still missing, which limits the deployment of neuromorphic systems in many real-time applications. This work proposes a scalable, fast, and efficient spiking neuromorphic hardware system with on-chip local learning capability. We introduce an effective hardware-friendly local training algorithm compatible with sparse temporal input coding and binary random classification weights. The algorithm is demonstrated to deliver competitive accuracy in different tasks. The proposed digital system explores spike sparsity in communication, parallelism in vector–matrix operations and process-level dataflow, and locality of training errors, which leads to low cost and fast training speed. The system is optimized under various performance metrics. Taking into consideration energy, speed, resources, and accuracy, the proposed method shows around$10\times $efficiency over a recent work with a direct feedback alignment (DFA) method and$4.5\times $efficiency over the spike-timing-dependent plasticity (STDP) method. Moreover, our hardware architecture can easily scale up with the network size at a linear rate. Thus, our method has demonstrated great potential for use in various applications, especially those demanding low latency. Wenzhe Guo, Mohamed E. Fouda, Ahmed M. Eltawil, Khaled N. Salama |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2021 | IMCA: An Efficient In-Memory Convolution AcceleratorabstractTraditional convolutional neural network (CNN) architectures suffer from two bottlenecks: computational complexity and memory access cost. In this study, an efficient in-memory convolution accelerator (IMCA) is proposed based on associative in-memory processing to alleviate these two problems directly. In the IMCA, the convolution operations are directly performed inside the memory as in-place operations. The proposed memory computational structure allows for a significant improvement in computational metrics, namely, TOPS/W. Furthermore, due to its unconventional computation style, the IMCA can take advantage of many potential opportunities, such as constant multiplication, bit-level sparsity, and dynamic approximate computing, which, while supported by traditional architectures, require extra overhead to exploit, thus reducing any potential gains. The proposed accelerator architecture exhibits a significant efficiency in terms of area and performance, achieving around 0.65 GOPS and 1.64 TOPS/W at 16-bit fixed-point precision with an area less than 0.25 mm2. Hasan Erdem Yantir, Ahmed M. Eltawil, Khaled N. Salama |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | Towards Hardware Optimal Neural Network Selection with Multi-Objective Genetic SearchabstractThe selection of hyperparameters and circuit components for optimum hardware implementation of a neural network is a challenging task, which has not been automated yet. This work proposes the method for the selection of optimum neural network architecture and hyperparameters using genetic algorithm based on the hardware-related performance metrics, such an on-chip area, power consumption, processing time and robustness to hardware non-idealities, and focus on memristor-based analog network architecture. The experimental results show that the proposed approach allows to select the optimum architecture based on the designers' preferences. Olga Krestinskaya, Khaled N. Salama, Alex James 0001 |
ISCAS | 2 |
| 2019 | FPGA implementation of dynamically reconfigurable IoT security module using algorithm hopping
Shady Mohamed Soliman, Mohammed A. Jaela, Abdelrhman Mohamed Abotaleb, Youssef Hassan, Mohamed Abdelghany, Amr Talaat Abdel-Hamid, Khaled N. Salama, Hassan Mostafa |
Integr. | 7 |
| 2018 | Analog Backpropagation Learning Circuits for Memristive Crossbar Neural NetworksabstractThe implementation of backpropagation algorithm using gradient descent operation with analog circuits is an open problem. In this paper, we present the analog learning circuits for realizing backpropagation algorithm for use with neural networks in memristive crossbar arrays. The circuits are simulated in SPICE using TSMC 180nm CMOS process models, and HP memristor models. The gradient descent operations are validated comprehensively using the relevant transfer characteristics and transient response of individual circuit modules. Olga Krestinskaya, Khaled N. Salama, Alex James 0001 |
ISCAS | 2 |
| 2016 | Stochastic synaptic plasticity with memristor crossbar arraysabstractMemristive devices have been shown to exhibit slow and stochastic resistive switching behavior under low-voltage, low-current operating conditions. Here we explore such mechanisms to emulate stochastic plasticity in memristor crossbar synapse arrays. Interfaced with integrate-and-fire spiking neurons, the memristive synapse arrays are capable of implementing stochastic forms of spike-timing dependent plasticity which parallel mean-rate models of stochastic learning with binary synapses. We present theory and experiments with spike-based stochastic learning in memristor crossbar arrays, including simplified modeling as well as detailed physical simulation of memristor stochastic resistive switching characteristics due to voltage and current induced filament formation and collapse. Rawan Naous, Maruan Al-Shedivat, Emre Neftci, Gert Cauwenberghs, Khaled N. Salama |
ISCAS | 5 |
| 2014 | An interference cancellation strategy for broadcast in hierarchical cell structureabstractIn this paper, a hierarchical cell structure is considered, where public safety broadcasting is fulfilled in a femtocell located within a macrocell. In the femtocell, also known as local cell, an access point broadcasts to each local node (LN) over an orthogonal frequency sub-band independently. Since the local cell shares the spectrum licensed to the macrocell, a given LN is interfered by transmissions of the macrocell user (MU) in the same sub-band. To improve the broadcast performance in the local cell, a novel scheme is proposed to mitigate the interference from the MU to the LN while achieving diversity gain. For the sake of performance evaluation, ergodic capacity of the proposed scheme is quantified and a corresponding closed-form expression is obtained. By comparing with the traditional scheme that suffers from the MU's interference, numerical results substantiate the advantage of the proposed scheme and provide a useful tool for the broadcast design in hierarchical cell systems. Yuli Yang 0003, Sonia Aïssa, Ahmed M. Eltawil, Khaled N. Salama |
GLOBECOM | 4 |
| 2014 | Hardware stream cipher with controllable chaos generator for colour image encryptionabstractThis study presents hardware realisation of chaos‐based stream cipher utilised for image encryption applications. A third‐order chaotic system with signum non‐linearity is implemented and a new post processing technique is proposed to eliminate the bias from the original chaotic sequence. The proposed stream cipher utilises the processed chaotic output to mask and diffuse input pixels through several stages of XORing and bit permutations. The performance of the cipher is tested with several input images and compared with previously reported systems showing superior security and higher hardware efficiency. The system is experimentally verified on XilinxVirtex 4 field programmable gate array (FPGA) achieving small area utilisation and a throughput of 3.62 Gb/s. M. L. Barakat, Abhinav S. Mansingka, Ahmed Gomaa Radwan, Khaled N. Salama |
IET Image Process. | 4 |
| 2013 | Modeling and fabrication of an RF MEMS variable capacitor with a fractal geometryabstractIn this paper, we model, fabricate, and measure an electrostatically actuated MEMS variable capacitor that utilizes a fractal geometry and serpentine-like suspension arms. Explicitly, a variable capacitor that possesses a top suspended plate with a specific fractal geometry and also possesses a bottom fixed plate complementary in shape to the top plate has been fabricated in the PolyMUMPS process. An important benefit that was achieved from using the fractal geometry in designing the MEMS variable capacitor is increasing the tuning range of the variable capacitor beyond the typical ratio of 1.5. The modeling was carried out using the commercially available finite element software COMSOL to predict both the tuning range and pull-in voltage. Measurement results show that the tuning range is 2.5 at a maximum actuation voltage of 10V. Amro M. Elshurafa, Khaled N. Salama, P. H. Ho |
ISCAS | 2 |
| 2012 | Memristor: the illusive deviceabstractThe memristor (M) is considered to be the fourth two-terminal passive element in electronics, alongside the resistor (R), the capacitor (C), and the inductor (L). Its existence was postulated in 1971 but its first implementation was reported in 2008. Where was it hiding all that time and what can we do with it? Come and learn how the memristor completes the roster of electronic devices much like a missing particle that physicists seek to complete their tableaus. Khaled N. Salama |
ACM Great Lakes Symposium on VLSI | 1 |
| 2012 | A Best-First Soft/Hard Decision Tree Searching MIMO Decoder for a 4 × 4 64-QAM SystemabstractThis paper presents the algorithm and VLSI architecture of a configurable tree-searching approach that combines the features of classical depth-first and breadth-first methods. Based on this approach, techniques to reduce complexity while providing both hard and soft outputs decoding are presented. Furthermore, a single programmable parameter allows the user to tradeoff throughput versus BER performance. The proposed multiple-input-multiple-output decoder supports a 4 × 4 64-QAM system and was synthesized with 65-nm CMOS technology at 333 MHz clock frequency. For the hard output scheme the design can achieve an average throughput of 257.8 Mbps at 24 dB signal-to-noise ratio (SNR) with area equivalent to 54.2 Kgates and a power consumption of 7.26 mW. For the soft output scheme it achieves an average throughput of 83.3 Mbps across the SNR range of interest with an area equivalent to 64 Kgates and a power consumption of 11.5 mW. Chung-An Shen, Ahmed M. Eltawil, Khaled N. Salama, Sudip Mondal |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | High performance technique for database applicationsusing a hybrid GPU/CPU platformabstractMany database applications, such as sequence comparing, sequence searching, and sequence matching, etc, process large database sequences. we introduce a novel and efficient technique to improve the performance of database applications by using a Hybrid GPU/CPU platform. In particular, our technique solves the problem of the low efficiency resulting from running short-length sequences in a database on a GPU. To verify our technique, we applied it to the widely used Smith-Waterman algorithm. The experimental results show that our Hybrid GPU/CPU technique improves the average performance by a factor of 2.2, and improves the peak performance by a factor of 2.8 when compared to earlier implementations. Mohammed Affan Zidan, Talal Bonny, Khaled N. Salama |
ACM Great Lakes Symposium on VLSI | 3 |
| 2010 | A best-first tree-searching approach for ML decoding in MIMO systemabstractIn MIMO communication systems maximum-likelihood (ML) decoding can be formulated as a tree-searching problem. This paper presents a tree-searching approach that combines the features of classical depth-first and breadth-first approaches to achieve close to ML performance while minimizing the number of visited nodes. A detailed outline of the algorithm is given, including the required storage. The effects of storage size on BER performance and complexity in terms of search space are also studied. Our result demonstrates that with a proper choice of storage size the proposed method visits 40% fewer nodes than a sphere decoding algorithm at signal to noise ratio (SNR) = 20dB and by an order of magnitude at 0 dB SNR. Chung-An Shen, Ahmed M. Eltawil, Sudip Mondal, Khaled N. Salama |
ISCAS | 4 |
| 2010 | Design and Implementation of a Sort-Free K-Best Sphere DecoderabstractThis paper describes the design and very-large-scale integration (VLSI) architecture for a 4 × 4 breadth-first K-best multiple-input-multiple-output (MIMO) decoder using a 64 quadrature-amplitude modulation (QAM) scheme. A novel sort-free approach to path extension, as well as quantized metrics result in a high-throughput VLSI architecture with lower power and area consumption compared to state-of-the-art published systems. Functionality is confirmed via a field-programmable gate array (FPGA) implementation on a Xilinx Virtex II Pro FPGA. Comparison of simulation and measurements are given, and FPGA utilization figures are provided. Finally, VLSI architectural tradeoffs are explored for a synthesized application-specific IC (ASIC) implementation in a 65-nm CMOS technology. Sudip Mondal, Ahmed M. Eltawil, Chung-An Shen, Khaled N. Salama |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2008 | A novel approach for K-best MIMO detection and its VLSI implementationabstractSince the complexity of MIMO detection algorithms is exponential, the K-best algorithm is often chosen for efficient VLSI implementation. This detection problem is often viewed as a tree search problem where the breadth first search (BFS) method is adopted and only the K-best branches are kept at each level of the tree. An earlier VLSI implementation of the K-best BFS has been reported, however it has an inherent speed bottleneck due to the calculation of many path metrics and then sorting among them to select the K-best. In this paper an alternative implementation of the BFS is presented, which is suitable for VLSI implementation. To test the performance of this approach it has been applied to a 4X4 MIMO detector with a 64 QAM constellation. The results show less than 1 dB degradation from the sphere decoding algorithm. The implementation of a single spiral cell, the basic block behind the system, occupies a 764 mum2of area and consumes a 52.58 muw of power a 0.13 mum CMOS technology. Sudip Mondal, Khaled N. Salama, Wersame H. Ali |
ISCAS | 2 |
| 2000 | A system for chaos generation and its implementation in monolithic formabstractWe propose a novel autonomous system for chaos generation based on a third-order abstract canonical mathematical model. Nonlinearity in this system is introduced by a bipolar switching constant which reflects the behaviour of a simple inverter circuit. Two implementations of the system are given. The first uses commercially available components while the second was designed on a CMOS chip. Numerical simulations and experimental results are provided. Ahmed S. Elwakil, Khaled N. Salama, Michael Peter Kennedy |
ISCAS | 2 |