EDBT 2026 Demo / reviewers in the wild / expert
Ashok Kumar 0001
dblp:55/3227-1
· DBLP profile ↗
17ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-9740-1219ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 2 first-author · 6 since 2021Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hardware Acceleration of CoAP Protocol for High-Speed and Low-Power Internet of Things CommunicationabstractThe Internet of Things (IoT) is a transformative technology facilitating seamless communication between diverse devices and systems, including resource-constrained devices. Speed efficiency and energy efficiency in communication protocols for IoT devices are crucial. The constrained application protocol (CoAP) is a promising, lightweight, and efficient protocol for IoT, offering robust messaging capabilities while conserving resources. An emerging research focus and challenge is designing hardware accelerators for CoAP that are fast, energy-efficient, and reliable. This article addresses that research challenge by proposing a CoAP hardware accelerator for optimizing message processing in resource-constrained IoT environments. The proposed accelerator’s architecture uses virtual channels (VCs) to manage incoming message traffic efficiently, enabling concurrent processing and enhancing throughput capacity. The accelerator minimizes processing delays and improves the system responsiveness by leveraging dynamic resource allocation and streamlined routing mechanisms. The proposed method is implemented using VHDL on Altera 10 GX FPGA. It reduces power consumption by consuming only 112.4 mW. Additionally, the accelerator demonstrates an impressive average latency of$58~\mu $s and energy consumption of$6.62~\mu $J, showcasing its superior performance metrics. The efficacy of the proposed CoAP hardware accelerator is tested through detailed evaluation and comparative analysis, affirming its superior performance over previously reported results in the literature. Kasem Khalil, Ashok Kumar 0001, Magdy A. Bayoumi |
IEEE Internet Things J. | 2 |
| 2025 | Accurate Hardware Predictor for Epileptic SeizureabstractEpilepsy triggers seizures, which develop before clinical onset in patients, and a timely and accurate prediction can save lives. A research challenge is to design accurate, fast, and energy-efficient hardware predictors. This work advances hardware-based seizure prediction research by proposing a new machine-learning-based predictor. It proposes a novel reconfigurable electroencephalogram (EEG) signal segmentation for increased learning. The proposed reconfigurable segmentation adaptively adjusts the overlap extent between consecutive segments and prepares new segments. Such prepared segments are fed into a Convolutional Auto-Encoder (CAE) using a proposed convolution module. The proposed convolution module uses optimized hyperparameters, including the number of layers, filters, filter size, pooling method, stride value, and padding for high learning and feature extraction. The learned CAE feeds into an Economic Long Short-Term Memory (ELSTM) to attain the final prediction result. The proposed predictor achieves high accuracy by exploiting the temporal dynamics of epileptic activity. It predicts seizures with an accuracy of 99.32%, a sensitivity of 99.29%, and a false alarm rate of 0.003 per hour, yielding high performance across classification thresholds, incurring low costs, and outperforming related hardware solutions. It is implemented in stand-alone VHDL, Altera Arria 10 GX FPGA, and synthesized into 45-nm technology. Kasem Khalil, Ashok Kumar 0001, Magdy A. Bayoumi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | An Economic Uniqueness-Improved Reliable Reconfigurable RO PUF for IoT SecurityabstractPhysical Unclonable Function (PUF) emerges as a promising solution for Internet of Things (IoT) and hardware security systems. Nevertheless, the advancement in the design of IoT constrained devices is limited due to area overhead and power consumption incurred from applying complex and expensive methods to strengthen hardware security systems in particular the PUF. This paper leverages a reconfigurable PUF design as an efficient solution for hardware security of IoT constrained devices. This work proposes ERRO PUF as a novel economic improved-uniqueness reliable reconfigurable RO PUF for IoT and hardware security. ERRO PUF reduces area overhead and power consumption of the design, improves system design reusability, and reduces hardware resources utilization. It generates more CRPs with considerably less hardware resources, improves hardware efficiency of the PUF design, and achieves better performance with significantly less hardware resources requirement when compared to other conventional ring oscillator (RO), configurable RO (CRO), and reconfigurable RO (RRO) PUF designs. ERRO PUF requires only 6.25% of hardware resources to generate 1-bit PUF response when compared to other state-of-the-art CRO and RRO PUFs. ERRO PUF’s results show that the proposed model is suitable for enhancing a lightweight hardware security system and eliminating the need for trading security with area and power reduction. Dominick Rizk, Rodrigue Rizk, Frederic Rizk, Ashok Kumar 0001 |
ISCAS | 4 |
| 2022 | A Cost-Efficient Reversible-Based Reconfigurable Ring Oscillator Physical Unclonable FunctionabstractPhysical Unclonable Function (PUF) has been advocated as a promising solution for the hardware security of the Internet of Things (IoT). Nevertheless, the development of the IoT constrained devices is limited due to an important factor which is power. By leveraging the reversible logic paradigm, PUF can mitigate the power limitation. Integrating reversible logic into PUF alleviates the security and power challenges that limit the advancement in IoT and hardware security. To the best of our knowledge, this work is the first to design a cost-efficient reversible-based reconfigurable ring oscillator PUF (denoted as R3OPUF) for IoT and hardware security. The proposed design leverages the reversible logic paradigm and achieves better performance with significantly less hardware resources requirement compared to other conventional ring oscillator (RO)PUF, configurable RO (CRO), and reconfigurable (RRO) PUF designs. Quantum and comparative analysis for the R3O PUF is discussed in this paper. The proposed R3O PUF is able to generate more Challenge Response Pairs (CRPs) compared with state-of the-art RO PUFs with using an equal number of configurable logic blocks (CLBs) of an FPGA. R3O PUF requires only 50% and 25% of hardware resources to generate 1-bit PUF response compared to other CRO PUFs and RRO PUFs, respectively. R3O PUF’s results show that the proposed model is suitable for enhancing a lightweight hardware security system especially for root of trust (RoT), authentication and key generation applications and eliminating the need for trading the security with power reduction. Frederic Rizk, Dominick Rizk, Rodrigue Rizk, Ashok Kumar 0001 |
ISCAS | 4 |
| 2022 | A Resource-Saving Energy-Efficient Reconfigurable Hardware Accelerator for BERT-based Deep Neural Network Language Models using FFT MultiplicationabstractBidirectional Encoder Representations from Transformers (BERT) based language models are a new class of deep neural networks with an attention mechanism. They emerge as a better alternative to the traditional recurrent neural networks for better sequence representation. They have achieved state-of-the-art performance in various natural language processing (NLP) tasks. Nevertheless, they demand intensive computation, energy, and memory requirements which pose a major challenge for their deployment on resource-constrained platforms and edge devices. To mitigate these limitations, this paper proposes a novel hardware accelerator design dedicated for BERT-based architectures with a reconfigurable functionality that improves circuit reusability and reduces hardware resources utilization. To the best of our knowledge, it is the first to present a holistic design and implementation of a reconfigurable hardware accelerator for BERT-based deep neural network language models. The proposed design leverages Fast Fourier Transform-based multiplication on block-circulant matrices for accelerating BERT weights matrices' multiplication. It is evaluated for different BERT-based model configurations on mainstream popular benchmarks while achieving a state-of-the-art performance. It is also evaluated for distinct batch sizes to study the impact of the batch size on the energy efficiency. A cross-platform comparative analysis shows that the proposed hardware accelerator achieves $6 \times, 27 \times, 3.18 \times$, and $8 \times$ improvement compared to $C P U$, and up to $1.17 \times, 1.77 \times$, $5 \times$, and $86 \times$ improvement compared to GPU in latency, throughput, power consumption, and energy efficiency, respectively. This design is suitable for efficient NLP on resource-constrained platforms where low latency and high throughput are critical. Rodrigue Rizk, Dominick Rizk, Frederic Rizk, Ashok Kumar 0001, Magdy A. Bayoumi |
ISCAS | 4 |
| 2022 | Designing Novel AAD Pooling in Hardware for a Convolutional Neural Network AcceleratorabstractConvolutional neural network (CNN) hardware accelerators for specialized Internet of Things (IoT) requiring high accuracy is an emerging research topic. The pooling module in a CNN pipeline impacts both the speed and accuracy of a classification task. This work proposes the design and hardware implementation of a novel pooling method absolute average deviation (AAD) for CNN accelerator. AAD utilizes the spatial locality of pixels using vertical and horizontal deviations to achieve higher accuracy, lower area, and lower power consumption than mixed pooling without increasing the computational complexity. AAD is tested on four different datasets: EEG, ImageNet, Common Objects in Context (COCO), United States Postal Service (USPS), and multiple CNN structures: CNN, VGG16, VGG19, ResNet, and DenseNet. In hardware, AAD is implemented using Very High Speed Integrated Circuit (VHSIC) Hardware Description Language (VHDL) on Altera Arria10 GX field-programmable gate array (FPGA) and 45-nm technology using Synopsys Design Compiler. The area and power consumption are found to be 244.46 nm2and 0.31 mW, respectively. AAD achieves 98% accuracy with lower computational and hardware costs compared to mixed pooling, making it an ideal pooling mechanism for an IoT CNN accelerator. Kasem Khalil, Omar Eldash, Ashok Kumar 0001, Magdy A. Bayoumi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2021 | A Reversible-Logic Based Architecture for Long Short-Term Memory (LSTM) NetworkabstractAny sequential learning task relies on the idea of connecting previous time-stamp information to the immediate present time-stamp task to predict the future. The underlying challenge is to understand the hidden patterns in the sequence by means of analyzing short- and long-term dependencies and temporal differences. Recurrent Neural Networks (RNNs) and their variants, such as Long Short-Term Memory (LSTM) are widely used in problem domains like speech recognition, Natural Language Processing (NLP), fault prediction, and language translation modeling over the past few years. Higher accuracy demands complex LSTM network models which lead to high computational cost, area overhead, and excessive power consumption. Reversible logic circuit synthesis, in the context of ideally Zero heat dissipation, has emerged as a new research paradigm for low power circuit designs. In this paper, we have proposed a novel design of LSTM architecture using reversible logic gates. To the best of our knowledge, the proposed approach is the first attempt to implement a complete feedforward LSTM circuit using only reversible logic gates. The hardware implementation of the proposed method is presented using VHDL and Altera Arria10 GX FPGA. The comparative analysis demonstrates that the proposed approach has achieved an approximately 17% reduction in overall power dissipation compared to traditional networks. The proposed approach also has better scalability than the classical design approach. Kasem Khalil, Bappaditya Dey, Ashok Kumar 0001, Magdy A. Bayoumi |
ISCAS | 3 |
| 2020 | A Novel Design Reversible Logic Based Configurable Fault-Tolerant Embryonic HardwareabstractWith the advancement of advanced node technology beyond sub-10 nm nodes, high-performance computing is facing a great challenge in the form of excessive levels of heat. Against this limitation, we can re-synthesis any complex digital circuits using reversible logic only, known for ideally Zero-heat dissipation. This paper proposes a novel reversible logic based on Configurable Fault-Tolerant Embryonic Hardware. We have reinvestigated the concept of Self-healing for hardware systems in the context of reversible logic and circuits. This paper presents a comparative analysis between conventional and proposed quantum approach on various parameters such as area-overhead, power dissipation and quantum cost along with the limitations of conventional computing. The reliability of the proposed approach is analyzed against other existing classical approaches with different failure rates. The overall power dissipation is almost 19% lower for the proposed approach compared to other conventional approaches using digital gates with cell number 32. The proposed approach is implemented for the ALU array using VHDL on Altera 10 GX FPGA. Kasem Khalil, Bappaditya Dey, Yasser Sherazi, Ashok Kumar 0001, Magdy A. Bayoumi |
ISCAS | 4 |
| 2019 | Demystifying Emerging Nonvolatile Memory Technologies: Understanding Advantages, Challenges, Trends, and Novel ApplicationsabstractWith CMOS scaling moving toward an end, some “out-of-the-box” non-volatile memory (NVM) technologies come to life and promise to break the “memory wall”, fill the gap between memory and processor computation, and to cope with the limitations of conventional memory technologies that become limited in fulfilling the new requirements of the changing market trends. With the emergence of distinct NVM technologies, such as STT-RAM, PCM, and ReRAM, the concept of “universal memory” seems to be now achievable and about to have a far-reaching impact on the computing market. This paper studies each of the distinctive emerging non-volatile memories (STT-RAM, PCM, ReRAM) and presents their benefits, current limitations and trends, as well as their potential novel applications in a broad range of fields. Rodrigue Rizk, Dominick Rizk, Ashok Kumar 0001, Magdy A. Bayoumi |
ISCAS | 3 |
| 2017 | Real-time streaming challenges in Internet of Video Things (IoVT)abstractThe limited capabilities of Internet of Things (IoT) devices make real-time video streaming a major challenge. Video encoding and transmission are computationally intensive processes. Applications; like urban surveillance and health care monitoring, require real-time high definition video streams. Current video encoders are not designed to meet these requirements, thus alternative architectures, algorithms and compression techniques are needed to meet these goals. In this paper, real-time video streaming challenges for IoT applications are presented. Ahmed Sammoud, Ashok Kumar 0001, Magdy A. Bayoumi, Tarek A. Elarabi |
ISCAS | 2 |
| 2007 | Pixel-Level Image Fusion Scheme based on Linear AlgebraabstractImage fusion refers to the process of integrating complementary image sources from multiple imaging sensor such that the resulting fused image improves the performance of computational analysis tasks such as segmentation, feature extraction and object recognition. The paper introduces a pixel-level image fusion scheme based on linear algebra. The image fusion process begins by computing the discrete wavelet transform of the source images. Then, the wavelet transform of the images are fused using a feature-based rule. A salient feature may extend to several pixels; therefore, a rule that can include a region of pixels containing it results in a more efficient integration. The fusion rule is based on a measurement of the linear dependency of a small window centered on the pixel under consideration. The linear dependency measurement is the Wronskian determinant that is a simple and rigorous test. The performance assessment of the proposed method is established by using mutual information measurement as well as root mean square error and peak signal to noise ratio. The simulation results show that the proposed method is an efficient approach to image fusion. Ruth Aguilar-Ponce, Jose Luis Tecpanecatl-Xihuitl, Ashok Kumar 0001, Magdy A. Bayoumi |
ISCAS | 3 |
| 2007 | Design and Realization of Analog Phi-Function for LDPC DecoderabstractOne of the ambitious design goals of future generations of wireless systems, including 4G, IEEE 802.11n/802.16 standards, is to reliably provide very high data rate transmission in real-time. This poses a challenge to find an optimal coding scheme that has good performance and can be efficiently implemented in hardware. The most well-known LDPC decoding algorithm is log sum product (log-SP) in which a set of calculations on a non-linear function called Phi-function is approximated by a minimum function. Until now this function has been implemented through look up tables (LUT). But this direct implementation is costly for hardware. Also LUTs are very sensitive to the number of quantization bits and number of LUT values. Therefore, we have proposed analog Phi-function. The design is easily scalable and reconfigurable for larger block sizes. Simulation results show that our proposed design dissipates only 18 nW. Abu Baker, Soumik Ghosh, Ashok Kumar 0001, Magdy A. Bayoumi, Rafic Ayoubi |
ISCAS | 3 |
| 2007 | A network of sensor-based framework for automated visual surveillance
Ruth Aguilar-Ponce, Ashok Kumar 0001, Jose Luis Tecpanecatl-Xihuitl, Magdy A. Bayoumi |
J. Netw. Comput. Appl. | 2 |
| 2006 | Design of Robust, Energy-Efficient Full Adders for Deep-Submicrometer Design Using Hybrid-CMOS Logic StyleabstractWe present a new design for a 1-b full adder featuring hybrid-CMOS design style. The quest to achieve a good-drivability, noise-robustness, and low-energy operations for deep submicrometer guided our research to explore hybrid-CMOS style design. Hybrid-CMOS design style utilizes various CMOS logic style circuits to build new full adders with desired performance. This provides the designer a higher degree of design freedom to target a wide range of applications, thus significantly reducing design efforts. We also classify hybrid-CMOS full adders into three broad categories based upon their structure. Using this categorization, many full-adder designs can be conceived. We will present a new full-adder design belonging to one of the proposed categories. The new full adder is based on a novel xor-xnor circuit that generates xor and xnor full-swing outputs simultaneously. This circuit outperforms its counterparts showing 5%-37% improvement in the power-delay product (PDP). A novel hybrid-CMOS output stage that exploits the simultaneous xor-xnor signals is also proposed. This output stage provides good driving capability enabling cascading of adders without the need of buffer insertion between cascaded stages. There is approximately a 40% reduction in PDP when compared to its best counterpart. During our experimentations, we found out that many of the previously reported adders suffered from the problems of low swing and high noise when operated at low supply voltages. The proposed full adder is energy efficient and outperforms several standard full adders without trading off driving capability and reliability. The new full-adder circuit successfully operates at low voltages with excellent signal integrity and driving capability. To evaluate the performance of the new full adder in a real circuit, we embedded it in a 4- and 8-b, 4-operand carry-save array adder with final carry-propagate adder. The new adder displayed better performance as compared to the standard full adders Sumeer Goel, Ashok Kumar 0001, Magdy A. Bayoumi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Efficient shield insertion for inductive noise reduction in nanometer technologiesabstractWith high clock frequencies, faster transistor rise/fall time, wider wires, and the use of Cu material interconnects, interconnect inductive noise is becoming an important design metric in digital circuits. An efficient technique to reduce the inductive noise of on-chip interconnects is to insert shields among signal wires. An efficient solution for the min-area shield insertion problem to satisfy given explicit noise bounds in multiple coupled nets is provided. The proposed algorithm determines the locations and number of shields needed to satisfy certain noise constraints. Experimental results show that the proposed approach minimizes the number of shields required to satisfy the noise constraints and uses less runtime than the best alternative reported approach. Mohamed A. Elgamel, Ashok Kumar 0001, Magdy A. Bayoumi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2004 | A methodology for low power scheduling with resources operating at multiple voltages
Ashok Kumar 0001, Magdy A. Bayoumi, Mohamed A. Elgamel |
Integr. | 1 |
| 1999 | Novel Formulations for Low-Power Binding of Function Units in High-Level SynthesisabstractThis paper considers minimizing the switchings of the function units through new binding formulations. Switching activities on the function units are gathered through profiling the data-flow graph of the design at hand. Several thousand random input streams are generated and used for such profiling. The switching activities are obtained and stored in a matrix. The problem of binding the function units for low power is then formulated and solved by using the proposed methods. Ashok Kumar 0001, Magdy A. Bayoumi |
ICCD | 1 |