EDBT 2026 Demo / reviewers in the wild / expert
Magdy A. Bayoumi
dblp:b/MagdyABayoumi
· DBLP profile ↗
188ranked-venue papers
10as first author
22since 2021 · last 2026
0000-0002-0630-5273ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 100 · 3 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 53 · 6 first-author · 2 since 2021Computer networks · 17 · 3 since 2021Artificial intelligence and machine learning · 8 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Databases, data management, data science and information retrieval · 3Theory of computation · 2 · 1 first-authorSecurity and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Novel TSV Model With Fault Characterization for High-Frequency Transmission in 3D ICsabstractThrough-silicon vias (TSVs) are essential for 3D integrated circuits (ICs) and advanced chiplet packaging. The semiconductor industry is transitioning toward 3D ICs, chiplets, and system-in-package (SiP) solutions due to the slowdown of Moore’s Law and limitations in conventional silicon scaling. In this paper, we propose an optimized TSV architecture for high-frequency transmission to enhance its suitability for 6G communication chips, and develop a comprehensive equivalent circuit model for fault-free and faulty TSVs. This model accounts for open-circuit and short-circuit fault conditions while considering the effects of higher frequencies, substrate type, doping concentration, and adjacent layers. At the physical level, the TSVs are simulated using the Ansys High-Frequency Structure Simulator (HFSS), and the equivalent circuits are designed using the Cadence Virtuoso tool. An experimental evaluation is also conducted to validate the physical design. We position the TSVs in a pattern of ground-signal-ground (G-S-G) to reduce the effective inductance of the signal TSV, thereby minimizing inductive reactance at ultra-high frequencies. Consequently, the reflection coefficient remains below -10 dB across the frequency range of 0.1 to 146.3 GHz. Furthermore, we compare simulation outcomes from HFSS and Cadence for both fault-free and faulty TSVs under varying operating conditions. Additionally, the parasitic circuit components are characterized through extensive theoretical derivations for in-depth circuit verification. Collectively, the rigorous analysis, experimental validation, and thorough investigation of the proposed design and its equivalent circuit demonstrate their potential for use in creating datasets for a fault prediction machine learning model. Prosen Kirtonia, Shelby Williams, Sonia Akter, Magdy A. Bayoumi, Kasem Khalil |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | Optimization of Chirality Variation in Carbon Nanotube Field Effect Transistor Spiking NeuronsabstractFor over half a century, silicon-based Complementary Metal Oxide Semiconductors (CMOS) have been the dominant technology used in the manufacture of nearly all integrated circuits. To fulfill Moore’s Law prediction, CMOS device dimensions were meticulously and carefully reduced to the single-digit nanometer regime. This gradual reduction has led to remarkable exponential performance increases over several decades, ushering in an unparalleled era of computation in human history. However, Short-Channel Effects (SCEs) present many impediments to further improvements in CMOS devices. SCEs occur when the channel length has been scaled down to the same order of magnitude as the depletion-layer widths of the source and drain junctions. To overcome the numerous SCEs caused by the miniaturization of CMOS devices, Carbon Nanotube Field Effect Transistors (CNFETs) aim to serve as their superior successors. CNFETs exhibit exceptional electrical properties, far surpassing those of CMOS, primarily due to their ballistic transport properties and excellent electrostatic scaling. To the best of our knowledge, this paper is the first to investigate chirality variation in CNFETs for spiking neurons. More specifically, chirality variation in CNFETs is used to determine two optimizations: (1) highest-frequency and (2) lowest-energy spiking neurons, using the Penta-Transistor Integrate & Fire (PTIF) architecture to demonstrate these dual mandates. These optimizations separately provide a 6.89x increase in spiking frequency or an 87.43% energy saving. Shelby Williams, Prosen Kirtonia, Kasem Khalil, Magdy A. Bayoumi |
ICASSP | 4 |
| 2025 | A High Performance and Efficient Method for Enhancing Randomness in Linear Feedback Shift Registers(LFSR)abstractThis paper presents an efficient Pseudo Random Number Generator (PRNG) design based on a 16-bit Linear Feedback Shift Register (LFSR) with 16 polynomials dynamically controlled by four distinct ring oscillators (ROs). RNGs are the fundamental components of modern digital communications systems. The deterministic properties of PRNGs are used for symmetric key encryption for bulk data transmission and other diverse applications. This work focuses on enhancing the randomness of LFSR-based RNGs to improve security and resource utilization. The proposed method achieves high randomness through dynamic polynomial (taps) selection, providing unpredictability in both sample size and selection. The whole design is synthesized and validated in the Xilinx Vivado 2023.2 tool using a Spartan-7 FPGA board. The randomness quality of the generated bitstream is evaluated using the NIST SP800-22 statistical test suite, with the proposed RNG passing all 16 tests and producing significant P-values. Then, P-values from the NIST tests are compared with the state-of-the-art PRNG methods. In addition, the design performs considerably better with respect to randomness and resource usage than traditional 64-bit Fibonacci LFSR, which did not pass all the NIST tests. The design is also synthesized in the Synopsis design compiler for 14nm, 32nm, and 45nm technology nodes to better understand resource usage. In addition, autocorrelation results are also presented to further validate the quality of generated random numbers. This work provides an efficient, lightweight LFSR-based PRNG architecture for IoT, automation, embedded systems, and symmetric key encryption applications, where high-quality random bits are critical. Sonia Akter, Kasem Khalil, Magdy A. Bayoumi |
ISCAS | 3 |
| 2025 | TRUE-BSG: A True Random Bit-Stream Generator for Fast and Efficient Stochastic ComputingabstractStochastic computing (SC) leverages random bitstreams to perform arithmetic operations, offering ultra-low-cost, fault-tolerant, and highly parallelizable computations. The quality of these bit-streams is crucial for the accuracy and reliability of SC. This paper introduces TRUE-BSG, a novel true random bit-stream generator designed for fast and energy-efficient SC. Unlike state-of-the-art (SoTA) pseudo-random and quasi-random bit-stream generators, TRUE-BSG utilizes a high-quality true random number generator (TRNG), capable of producing random bits at a rate of 1 Gigabit per second. Our TRNG ensures high entropy and minimal correlation. TRUE-BSG shows comparable accuracy to software-based generators and better energy efficiency than SoTA bit-stream generators, making it an ideal solution for resource-constrained devices. Mehran Shoushtari Moghadam, Shelby Williams, Abu Kaisar Mohammad Masum, M. Hassan Najafi, Sercan Aygün, Magdy A. Bayoumi |
ISCAS | 6 |
| 2025 | Lightweight FPGA Implementation of the Shadow PUF Module for Generating Reconfigurable Proxy PUFsabstractPhysically unclonable functions (PUFs) are hardware security primitives with notable success in secure key generation applications, but have been faced with challenges in device authentication applications. This paper provides an FPGA implementation of the Shadow PUF design for generating reconfigurable PUFs to mitigate the dangers posed by modeling attacks. The proposed design is scalable with sampled randomness and uniqueness values averaging 48.88% and 48.05%, respectively. The selected design for the proof-of-concept implementation utilized 239 logic cells (LCs) and 256 bits of memory on an iCE40 HX1K FPGA—highlighting the proposed design as a viable security core for even the most resource-constrained devices. Pablo Rojas, Sara Alahmadi, Magdy A. Bayoumi |
ISCAS | 3 |
| 2025 | S²RNN: Self-Supervised Reconfigurable Neural Network Hardware Accelerator for Machine Learning ApplicationsabstractHardware implementation of neural networks (NNs) is challenging due to varying application requirements. This often necessitates creating specific field programmable gate arrays (FPGAs) configurations from scratch for each application. This article proposes a flexible, self-supervised reconfigurable method to fit several application requirements by providing only the maximum available computational nodes a priori. The proposed method dynamically reconfigures the required number of hidden layers and nodes based on the application. The goal is to automatically determine the optimal NN configuration through reconfigurability to achieve maximum accuracy. Optimality is demonstrated through minimum average power, average delay, and area overhead, as well as maximum throughput and accuracy. Experimental results show that the proposed approach significantly reduces the optimized architecture search cost (the number of online training iterations) and associated average power consumption for successive datasets/applications. The method’s effectiveness is shown both quantitatively and qualitatively, verified against the MNIST and CIFAR-10 classification problems. Our reconfigurable method demonstrates stable accuracy of 98.97% and 98.95% compared to state-of-the-art NNs with fixed configurations (98.85% and 73.0% for MNIST and 93.47% and 70.21% for CIFAR-10, respectively). Additionally, the proposed method shows a 20.9% reduction in average power dissipation compared to state-of-the-art methods. Implemented and tested using VHDL and Altera FPGA, the results indicate resource utilization comparable with the state-of-the-art method. This reconfigurability is especially advantageous for Internet of Things applications where power efficiency and adaptability to different tasks are critical. Kasem Khalil, Bappaditya Dey, Magdy A. Bayoumi |
IEEE Internet Things J. | 3 |
| 2025 | Hardware Acceleration of CoAP Protocol for High-Speed and Low-Power Internet of Things CommunicationabstractThe Internet of Things (IoT) is a transformative technology facilitating seamless communication between diverse devices and systems, including resource-constrained devices. Speed efficiency and energy efficiency in communication protocols for IoT devices are crucial. The constrained application protocol (CoAP) is a promising, lightweight, and efficient protocol for IoT, offering robust messaging capabilities while conserving resources. An emerging research focus and challenge is designing hardware accelerators for CoAP that are fast, energy-efficient, and reliable. This article addresses that research challenge by proposing a CoAP hardware accelerator for optimizing message processing in resource-constrained IoT environments. The proposed accelerator’s architecture uses virtual channels (VCs) to manage incoming message traffic efficiently, enabling concurrent processing and enhancing throughput capacity. The accelerator minimizes processing delays and improves the system responsiveness by leveraging dynamic resource allocation and streamlined routing mechanisms. The proposed method is implemented using VHDL on Altera 10 GX FPGA. It reduces power consumption by consuming only 112.4 mW. Additionally, the accelerator demonstrates an impressive average latency of$58~\mu $s and energy consumption of$6.62~\mu $J, showcasing its superior performance metrics. The efficacy of the proposed CoAP hardware accelerator is tested through detailed evaluation and comparative analysis, affirming its superior performance over previously reported results in the literature. Kasem Khalil, Ashok Kumar 0001, Magdy A. Bayoumi |
IEEE Internet Things J. | 3 |
| 2025 | Accurate Hardware Predictor for Epileptic SeizureabstractEpilepsy triggers seizures, which develop before clinical onset in patients, and a timely and accurate prediction can save lives. A research challenge is to design accurate, fast, and energy-efficient hardware predictors. This work advances hardware-based seizure prediction research by proposing a new machine-learning-based predictor. It proposes a novel reconfigurable electroencephalogram (EEG) signal segmentation for increased learning. The proposed reconfigurable segmentation adaptively adjusts the overlap extent between consecutive segments and prepares new segments. Such prepared segments are fed into a Convolutional Auto-Encoder (CAE) using a proposed convolution module. The proposed convolution module uses optimized hyperparameters, including the number of layers, filters, filter size, pooling method, stride value, and padding for high learning and feature extraction. The learned CAE feeds into an Economic Long Short-Term Memory (ELSTM) to attain the final prediction result. The proposed predictor achieves high accuracy by exploiting the temporal dynamics of epileptic activity. It predicts seizures with an accuracy of 99.32%, a sensitivity of 99.29%, and a false alarm rate of 0.003 per hour, yielding high performance across classification thresholds, incurring low costs, and outperforming related hardware solutions. It is implemented in stand-alone VHDL, Altera Arria 10 GX FPGA, and synthesized into 45-nm technology. Kasem Khalil, Ashok Kumar 0001, Magdy A. Bayoumi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | ECO: Enhanced In-Stream Correlation Manipulation for Low-Discrepancy Stochastic ComputingabstractStochastic computing (SC) is a reemerging computing paradigm that offers low-cost and noise-resilient hardware designs for a variety of arithmetic functions. In SC, circuits operate on uniform bit-streams, where the value is encoded by the probability of observing ‘1’s in the stream. The accuracy of SC operations highly depends on the correlation between input bit-streams. Some operations, such as minimum and maximum, require highly correlated inputs, whereas others like multiplication demand uncorrelated or statistically independent inputs for accurate results. Developing low-cost and accurate correlation manipulation circuits is critical, as they allow correlation management without incurring the high cost of bit-stream regeneration. This work introduces novel in-streamcorrelatoranddecorrelatorcircuits capable of: 1) adjusting correlation between stochastic bit-streams and 2) controlling the distribution of ‘1’s in the output bit-streams. Compared to state-of-the-art (SoA) approaches, our designs offer improved accuracy and reduced hardware overhead. The output bit-streams enjoylow-discrepancy (LD)distribution, leading to higher quality of results. To further increase the accuracy when dealing with pseudo-random inputs, we propose an enhancement module that balances the number of ‘1’s across adjacent input segments. We show the effectiveness of the proposed techniques through two application case studies: SC design of sorting and median filtering. Sina Asadi, Amir Hossein Jalilvand, M. Hassan Najafi, Magdy A. Bayoumi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | Digital-Twin Architecture of a Spiking Neuron using Carbon Nanotube Field Effect TransistorsabstractThe rapid advancements in nanotechnology and neuromorphic engineering have paved the way for developing high-performance computational models that mimic biological neural networks. This paper presents a novel digital-twin architecture for a spiking neuron, leveraging the exceptional properties of Carbon Nanotube Field Effect Transistors (CNFETs). The proposed architecture aims to emulate the dynamic behavior of biological neurons with high fidelity, providing a robust platform for simulation and analysis. By integrating CNFETs, we achieve significant improvements in lowest-energy usage and highest-spiking frequency, compared to traditional silicon-based technologies. Furthermore, the digital twin not only replicates the electrical characteristics of a spiking neuron but also facilitates advanced functionalities such as adaptive learning, fault tolerance, and self-optimization. Experimental results demonstrate the potential of CNFET-based neurons in enhancing the performance of neuromorphic systems, offering promising applications in artificial intelligence, cognitive computing, and complex system modeling. Integrating digital twin technology with CNFETs allows for precise control and manipulation of spiking behavior, enabling the development of more sophisticated and efficient neural models. This research underscores the transformative impact of combining digital twin technology with cutting-edge nanomaterials, setting a new benchmark for future explorations in the field. The findings highlight the importance of continued research and development in this area, as it holds the promise of revolutionizing the way we design and implement neuromorphic systems, pushing the boundaries of what is possible in artificial intelligence and beyond. Shelby Williams, Prosen Kirtonia, Kasem Khalil, Magdy A. Bayoumi |
IPCCC | 4 |
| 2024 | Fortifying Strong PUFs: A Modeling Attack-Resilient Approach Using Weak PUF for IoT Device SecurityabstractStrong Physical Unclonable Functions (PUFs) have gained traction as lightweight authentication solutions for IoT devices. However, their vulnerability to machine learning attacks poses a security risk. Various strategies have been introduced in the literature to enhance its resilience against modeling attacks, introducing additional complexity and making them unsuitable for resource-constrained devices. In contrast, weak PUFs exhibit inherent resistance to modeling attacks, but they suffer from a restricted number of Challenge-Response Pairs (CRPs), thus unsuitable for authentication. This paper proposes a PUF design that incorporates weak PUFs to obscure the responses of Strong PUFs, effectively safeguarding them from modeling attacks. Our design shows resilience against a modeling attack, revealing a maximum accuracy of 57% despite using 107CRPs. The proposed method is implemented on the Artix-7 FPGA with Verilog HDL. The results demonstrate that the proposed method has a small footprint in terms of resource utilization. This innovative approach offers a lightweight solution for IoT device authentication, combining the strengths of strong and weak PUFs while mitigating the vulnerabilities associated with modeling attacks. Sara Alahmadi, Kasem Khalil, Haytham Idriss, Magdy A. Bayoumi |
ISCAS | 4 |
| 2023 | MMST-ViT: Climate Change-aware Crop Yield Prediction via Multi-Modal Spatial-Temporal Vision TransformerabstractPrecise crop yield prediction provides valuable information for agricultural planning and decision-making processes. However, timely predicting crop yields remains challenging as crop growth is sensitive to growing season weather variation and climate change. In this work, we develop a deep learning-based solution, namely Multi-Modal Spatial-Temporal Vision Transformer (MMST-ViT), for predicting crop yields at the county level across the United States, by considering the effects of short-term meteorological variations during the growing season and the long-term climate change on crops. Specifically, our MMST-ViT consists of a Multi-Modal Transformer, a Spatial Transformer, and a Temporal Transformer. The Multi-Modal Transformer leverages both visual remote sensing data and short-term meteorological data for modeling the effect of growing season weather variations on crop growth. The Spatial Transformer learns the high-resolution spatial dependency among counties for accurate agricultural tracking. The Temporal Transformer captures the long-range temporal dependency for learning the impact of long-term climate change on crops. Meanwhile, we also devise a novel multi-modal contrastive learning technique to pre-train our model without extensive human supervision. Hence, our MMST-ViT captures the impacts of both short-term weather variations and long-term climate change on crops by leveraging both satellite images and meteorological data. We have conducted extensive experiments on over 200 counties in the United States, with the experimental results exhibiting that our MMST-ViT outperforms its counterparts under three performance metrics of interest. Our dataset and code are available at https://github.com/fudong03/MMST-ViT. Fudong Lin, Summer Crawford, Kaleb Guillot, Yihe Zhang 0001, Xu Yuan 0001, Li Chen 0019, Shelby Williams, Robert Minvielle, Xiangming Xiao, Drew Gholson, Nicolas Ashwell, Tri Setiyono, Brenda Tubana, Lu Peng 0001, Magdy A. Bayoumi, Nian-Feng Tzeng |
ICCV | 16 |
| 2023 | Security Scalability of Arbiter PUF DesignsabstractPhysically Unclonable Functions (PUFs) are hardware security primitives that can offer an alternative lightweight security solution for authenticating constrained Internet of Things (IoT) devices. However, PUFs are susceptible to modeling attacks, requiring the adoption of various design approaches to increase their resiliency. Many research efforts propose design approaches that offer better security against modeling attacks. This work investigates state-of-the-art modeling attacks performed on well-known Arbiter-based PUF architectures highlighting the best-fit modeling algorithm for different design approaches. Furthermore, the area efficiency of studied PUF designs is examined, and the optimal PUF design approaches for various area constraints are suggested. Such an assessment is required to evaluate PUF security accurately and guide the PUF community toward better practices. The findings revealed that some machine-learning algorithms performed better on a particular design. Additionally, when considering area overhead, we found that some PUF designs offer less security per area unit than their simpler counterparts. Accordingly, certain design elements are more efficient and add more security. Sara Alahmadi, Haytham Idriss, Pablo Rojas, Magdy A. Bayoumi |
ISCAS | 4 |
| 2023 | Low-Cost Hardware Design Approach for Long Short-Term Memory (LSTM)abstractLong Short-Term Memory (LSTM) has become commonly used for problems with a sequence of data. Hardware implementation of LSTM is a challenge for lightweight applications. This paper proposes an optimized LSTM method with few hardware components. The proposed method utilizes one sigmoid function to perform both input and output gates. The proposed method also utilizes one shared adder instead of two adders to perform the accumulation function for both the input and output gates. This unit performs the two functions in a sequence based on a selection signal which is used as a guide. The proposed method is tested using two datasets: MNIST and IMDB. The simulation results show the proposed method achieves the desired performance in classification compared similarly to the traditional method with a few hardware units. The proposed method is implemented using VHDL on Altera Arria 10 GX FPGA. The simulation results show that the proposed method utilizes fewer resources than the traditional method. The proposed method has a 16% area reduction compared to the traditional method. The proposed method has a power consumption of al.546 W while the traditional method consumes 1.847 W. Thus, the proposed method is suitable for lightweight applications with low hardware costs and desired performance. Kasem Khalil, Tamador Mohaidat, Magdy A. Bayoumi |
ISCAS | 3 |
| 2022 | XFeed PUF: A Secure and Efficient Delay-based Strong PUF Using Cross-Feed ConnectionsabstractPhysical unclonable functions (PUFs) are hardware security primitives that offer a lightweight security solution for constrained devices in the Internet of Things. The challenges facing PUFs security scaling have so far hindered their wide-scale deployment beyond simple key generation primitives. Although physically unclonable, PUFs are vulnerable to soft modeling attacks. Many PUF security enhancements impose significant implementation overhead, which could be problematic for devices operating in a constrained environment. This work introduces the Cross-Feed (XFeed) PUF, a highly efficient PUF circuit resilient against machine learning attacks while requiring a small circuit implementation area. In the XFeed PUF, arbiters feed intermediate race conditions to adjacent PUF rows to increase the non-linearity of the PUF system. A systematic categorization and benchmarking of the possible interconnection strategies are performed to determine the near-optimal connection schemes for the introduced XFeed PUF. The results showed that the XFeed PUF has superior security efficiency and scalability compared to other arbiter PUF-based enhancements. Tarek A. Idriss, Alex Gavin, Adrian Gabales, Haytham Idriss, Magdy A. Bayoumi |
ISCAS | 5 |
| 2022 | Shadow PUFs: Generating Temporal PUFs with Properties Isomorphic to Delay-Based APUFsabstractPhysical Unclonable Functions (PUFs) are popular hardware security primitives that offer lightweight authentication for constrained devices. However, lightweight PUF-based authentication often limits the number of authentications provided or even throttle the device to prevent adversaries from collecting enough information that could compromise the device’s security. This work introduces a Shadow PUF design, a controlled Strong PUF, to secure challenge-response exchanges against attackers and allow for an unlimited generation of unique responses. The proposed Shadow PUF design ensures security by periodically reconfiguring its behavior. The reconfiguration bits are generated by a static PUF primitive and are never exposed, while all authentication exchanges are done using the reconfigurable Shadow PUF. An ASIC Synthesis of the Shadow PUF demonstrates its small implementation area requirements and low power consumption. The security of the proposed PUF design against modeling attacks has also been analyzed. Haytham Idriss, Pablo Rojas, Sara Alahmadi, Tarek A. Idriss, Albert H. Carlson, Magdy A. Bayoumi |
ISCAS | 6 |
| 2022 | A Resource-Saving Energy-Efficient Reconfigurable Hardware Accelerator for BERT-based Deep Neural Network Language Models using FFT MultiplicationabstractBidirectional Encoder Representations from Transformers (BERT) based language models are a new class of deep neural networks with an attention mechanism. They emerge as a better alternative to the traditional recurrent neural networks for better sequence representation. They have achieved state-of-the-art performance in various natural language processing (NLP) tasks. Nevertheless, they demand intensive computation, energy, and memory requirements which pose a major challenge for their deployment on resource-constrained platforms and edge devices. To mitigate these limitations, this paper proposes a novel hardware accelerator design dedicated for BERT-based architectures with a reconfigurable functionality that improves circuit reusability and reduces hardware resources utilization. To the best of our knowledge, it is the first to present a holistic design and implementation of a reconfigurable hardware accelerator for BERT-based deep neural network language models. The proposed design leverages Fast Fourier Transform-based multiplication on block-circulant matrices for accelerating BERT weights matrices' multiplication. It is evaluated for different BERT-based model configurations on mainstream popular benchmarks while achieving a state-of-the-art performance. It is also evaluated for distinct batch sizes to study the impact of the batch size on the energy efficiency. A cross-platform comparative analysis shows that the proposed hardware accelerator achieves $6 \times, 27 \times, 3.18 \times$, and $8 \times$ improvement compared to $C P U$, and up to $1.17 \times, 1.77 \times$, $5 \times$, and $86 \times$ improvement compared to GPU in latency, throughput, power consumption, and energy efficiency, respectively. This design is suitable for efficient NLP on resource-constrained platforms where low latency and high throughput are critical. Rodrigue Rizk, Dominick Rizk, Frederic Rizk, Ashok Kumar 0001, Magdy A. Bayoumi |
ISCAS | 5 |
| 2022 | Stochastic Selection of Responses for Physically Unclonable FunctionsabstractChallenges in securing the Internet of Things (IoT) has led to the development of novel technologies such as physically unclonable functions (PUFs). Having applications in both lightweight authentication and key generation protocols for IoT devices, PUFs have received a great deal of research. Despite their promise, delay-based PUFs such as Arbiter PUFs and 4-XOR PUFs are easily modeled with 600 and 50, 000 challenge-response pairs (CRPs), respectively. While it has been shown that delay-based PUFs can be further improved by XORing together an increasing number of PUF instances, it also tends to become area-inefficient. In this paper the authors propose a novel method that combats the effectiveness of machine learning algorithms for modeling PUF behaviors by randomly selecting responses from a pool of PUFs. Six variants to our Random Bit Selection (RBS) PUF are proposed and investigated. The yielded results show that specific variants of RBS PUF are machine learning resistant despite using a 5, 760, 000 CRP dataset for training. Furthermore, the results indicate no significant improvement in the modeling algorithm despite a 100 times increase in the number of CRPs used. Finally, the security of the proposed design is also evaluated through a brute-force analysis to show its resistance to brute-force attacks. Pablo Rojas, Haytham Idriss, Sara Alahmadi, Magdy A. Bayoumi |
ISCAS | 4 |
| 2022 | Designing Novel AAD Pooling in Hardware for a Convolutional Neural Network AcceleratorabstractConvolutional neural network (CNN) hardware accelerators for specialized Internet of Things (IoT) requiring high accuracy is an emerging research topic. The pooling module in a CNN pipeline impacts both the speed and accuracy of a classification task. This work proposes the design and hardware implementation of a novel pooling method absolute average deviation (AAD) for CNN accelerator. AAD utilizes the spatial locality of pixels using vertical and horizontal deviations to achieve higher accuracy, lower area, and lower power consumption than mixed pooling without increasing the computational complexity. AAD is tested on four different datasets: EEG, ImageNet, Common Objects in Context (COCO), United States Postal Service (USPS), and multiple CNN structures: CNN, VGG16, VGG19, ResNet, and DenseNet. In hardware, AAD is implemented using Very High Speed Integrated Circuit (VHSIC) Hardware Description Language (VHDL) on Altera Arria10 GX field-programmable gate array (FPGA) and 45-nm technology using Synopsys Design Compiler. The area and power consumption are found to be 244.46 nm2and 0.31 mW, respectively. AAD achieves 98% accuracy with lower computational and hardware costs compared to mixed pooling, making it an ideal pooling mechanism for an IoT CNN accelerator. Kasem Khalil, Omar Eldash, Ashok Kumar 0001, Magdy A. Bayoumi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2021 | Generative Adversarial Network Based Semi-supervised Learning for Epileptic Focus LocalizationabstractAccurate localization of the epileptic focus plays an important role in the success of epileptic surgery. Epileptic focus localization process helps the physicians to determine the epileptogenic source in the brain by classifying the intracranial Electroencephalogram (iEEG) recordings into focal and non-focal. Using machine learning to automatically localize the epileptic focus requires a large number of labeled EEG recordings to effectively tackle such challenging classification problem. However, obtaining that much of labeled EEG data is expensive and time consuming. In order to accurately automate the focus localization task while having limited number of labeled EEG recordings, we introduce a semi-supervised learning method based on convolutional generative adversarial network. The proposed method is able to accurately classify the EEG signals into focal and non-focal using small amount of labeled training data which helps in reducing the burden of EEG data annotations. Discriminative spatial features are automatically extracted and classified by convolutional layers from raw EEG data without any preprocessing rather than manual feature extraction in the previous work. High classification accuracy of 95.5% makes the proposed method the most efficient among the state of the art. Hisham G. Daoud, Magdy A. Bayoumi |
BIBM | 2 |
| 2021 | An Efficient Deep Learning System for Epileptic Seizure PredictionabstractPredicting epilepsy ahead of its occurrence has been an arduous job for scientists for a long time. Epileptic patients are still endeavoring to find a prosperous way to evade seizures to improve the quality of their lives. In this paper, we propose a novel deep learning system for epileptic seizure prediction using multi-channel electroencephalogram (EEG) recordings from the scalp of human brains. The proposed system is patient-specific and is predicated on the classification between the interictal and preictal brain states for the epileptic patient. The system uses a two-dimensional convolutional variational autoencoder and trains it once in a supervised way for automatic feature learning and classification. Within a prediction window of up to one hour, our proposed system achieved an average sensitivity of 94.45% and 0.06FP/h average false prediction rate which makes it one of the most efficient among state-of-the-art methods. Ahmed M. Abdelhameed, Magdy A. Bayoumi |
ISCAS | 2 |
| 2021 | A Reversible-Logic Based Architecture for Long Short-Term Memory (LSTM) NetworkabstractAny sequential learning task relies on the idea of connecting previous time-stamp information to the immediate present time-stamp task to predict the future. The underlying challenge is to understand the hidden patterns in the sequence by means of analyzing short- and long-term dependencies and temporal differences. Recurrent Neural Networks (RNNs) and their variants, such as Long Short-Term Memory (LSTM) are widely used in problem domains like speech recognition, Natural Language Processing (NLP), fault prediction, and language translation modeling over the past few years. Higher accuracy demands complex LSTM network models which lead to high computational cost, area overhead, and excessive power consumption. Reversible logic circuit synthesis, in the context of ideally Zero heat dissipation, has emerged as a new research paradigm for low power circuit designs. In this paper, we have proposed a novel design of LSTM architecture using reversible logic gates. To the best of our knowledge, the proposed approach is the first attempt to implement a complete feedforward LSTM circuit using only reversible logic gates. The hardware implementation of the proposed method is presented using VHDL and Altera Arria10 GX FPGA. The comparative analysis demonstrates that the proposed approach has achieved an approximately 17% reduction in overall power dissipation compared to traditional networks. The proposed approach also has better scalability than the classical design approach. Kasem Khalil, Bappaditya Dey, Ashok Kumar 0001, Magdy A. Bayoumi |
ISCAS | 4 |
| 2020 | A Convolutional Gated Recurrent Neural Network for Seizure Onset LocalizationabstractThe success of epileptic surgery highly depends on the accurate localization of the epileptic seizure. Seizure onset localization process is done using the intracranial Electroencephalogram (iEEG) recording which helps the physicians to determine the epileptogenic source in the brain. In this paper, we propose a supervised learning method based on a convolutional gated recurrent neural network to accurately analyze the non-stationary and nonlinear EEG signals. We study the effect of different hyperparameters like the number of convolutional layers and the number of kernels on the accuracy of such a difficult classification task. Discriminative spatio-temporal features are automatically extracted from the EEG signals by the convolutional neural network and the recurrent neural network. EEG feature extraction and classification applied to raw data are performed in a single automated system rather than extracting handcrafted features as in the previous work. High classification accuracy of 95.1% using ten-fold cross-validation testing strategy, makes the proposed method the most efficient among the state of the art. Hisham G. Daoud, Magdy A. Bayoumi |
BIBM | 2 |
| 2020 | A Novel Design Reversible Logic Based Configurable Fault-Tolerant Embryonic HardwareabstractWith the advancement of advanced node technology beyond sub-10 nm nodes, high-performance computing is facing a great challenge in the form of excessive levels of heat. Against this limitation, we can re-synthesis any complex digital circuits using reversible logic only, known for ideally Zero-heat dissipation. This paper proposes a novel reversible logic based on Configurable Fault-Tolerant Embryonic Hardware. We have reinvestigated the concept of Self-healing for hardware systems in the context of reversible logic and circuits. This paper presents a comparative analysis between conventional and proposed quantum approach on various parameters such as area-overhead, power dissipation and quantum cost along with the limitations of conventional computing. The reliability of the proposed approach is analyzed against other existing classical approaches with different failure rates. The overall power dissipation is almost 19% lower for the proposed approach compared to other conventional approaches using digital gates with cell number 32. The proposed approach is implemented for the ALU array using VHDL on Altera 10 GX FPGA. Kasem Khalil, Bappaditya Dey, Yasser Sherazi, Ashok Kumar 0001, Magdy A. Bayoumi |
ISCAS | 5 |
| 2020 | Reduced-gate convolutional long short-term memory using predictive coding for spatiotemporal predictionabstractAbstract Spatiotemporal sequence prediction is an important problem in deep learning. We study next‐frame(s) video prediction using a deep‐learning‐based predictive coding framework that uses convolutional LSTM (convLSTM) modules. We introduce a novel rgcLSTM architecture that requires a significantly lower parameter budget than a comparable convLSTM. By using a single multifunction gate, our reduced‐gate model achieves equal or better next‐frame(s) prediction accuracy than the original convolutional LSTM while using a smaller parameter budget, thereby reducing training time and memory requirements. We tested our reduced gate modules within a predictive coding architecture on the moving MNIST and KITTI datasets. We found that our reduced‐gate model has a significant reduction of approximately 40% of the total number of training parameters and a 25% reduction in elapsed training time in comparison with the standard convolutional LSTM model. The performance accuracy of the new model was also improved. This makes our model more attractive for hardware implementation, especially on small devices. We also explored a space of 20 different gated architectures to get insight into how our rgcLSTM fits into that space. Nelly Elsayed, Anthony S. Maida, Magdy A. Bayoumi |
Comput. Intell. | 3 |
| 2019 | An Analysis of Univariate and Multivariate Electrocardiography Signal ClassificationabstractHeart diseases are mainly diagnosed by the electrocardiogram (ECG) or (EKG). The correct classification of ECG signals helps in diagnosing heart diseases. In this paper, we study and analyze the univariate and multivariate ECG signal classification problems to find the optimal classifier for ECG signals from existing state-of-the-art time series classification models. Nelly Elsayed, Anthony S. Maida, Magdy A. Bayoumi |
ICMLA | 3 |
| 2019 | Reduced-Gate Convolutional LSTM Architecture for Next-Frame Video Prediction Using Predictive CodingabstractSpatiotemporal sequence prediction is an important problem in deep learning. We study next-frame video prediction using a deep-learning-based predictive coding framework that uses convolutional, long short-term memory (convLSTM) modules. We introduce a novel reduced-gate convolutional LSTM architecture that achieves better next-frame prediction accuracy than the original convolutional LSTM while using a smaller parameter budget, thereby reducing training time and memory requirements. We tested our reduced gate modules within a predictive coding architecture on the gray-scale and RGB video datasets. We found that our reduced-gate model has a significant reduction of approximately 40 percent of the total number of training parameters and training time in comparison with the standard LSTM model which makes it attractive for hardware implementation especially on small devices. Nelly Elsayed, Anthony S. Maida, Magdy A. Bayoumi |
IJCNN | 3 |
| 2019 | Demystifying Emerging Nonvolatile Memory Technologies: Understanding Advantages, Challenges, Trends, and Novel ApplicationsabstractWith CMOS scaling moving toward an end, some “out-of-the-box” non-volatile memory (NVM) technologies come to life and promise to break the “memory wall”, fill the gap between memory and processor computation, and to cope with the limitations of conventional memory technologies that become limited in fulfilling the new requirements of the changing market trends. With the emergence of distinct NVM technologies, such as STT-RAM, PCM, and ReRAM, the concept of “universal memory” seems to be now achievable and about to have a far-reaching impact on the computing market. This paper studies each of the distinctive emerging non-volatile memories (STT-RAM, PCM, ReRAM) and presents their benefits, current limitations and trends, as well as their potential novel applications in a broad range of fields. Rodrigue Rizk, Dominick Rizk, Ashok Kumar 0001, Magdy A. Bayoumi |
ISCAS | 4 |
| 2019 | Cost-Efficient Cloud-Based Video Streaming Through Measuring HotnessabstractVideo streaming providers generally have to store several formats of the same video and stream the appropriate format based on the characteristics of the viewer’s device. This approach, called pre-transcoding, incurs a significant cost to the stream providers that rely on cloud services. Furthermore, pre-transcoding proven to be inefficient due to the long-tail access pattern to video streams. To reduce the incurred cost, we propose to pre-transcode only frequently accessed videos (called hot videos) and partially pre-transcode others, depending on their hotness degree. Therefore, we need to measure video stream hotness. Accordingly, we first provide a model to measure the hotness of video streams. Then, we develop methods that operate based on the hotness measure and determine how to pre-transcode videos to minimize the cost of stream providers. The partial pre-transcoding methods operate at different granularity levels to capture different patterns in accessing videos. Particularly, one of the methods operates faster but cannot partially pre-transcode videos with the non-long-tail access pattern. Experimental results show the efficacy of our proposed methods, specifically, when a video stream repository includes a high percentage of the Frequently Accessed Video Streams and a high percentage of videos with the non-long-tail accesses pattern. Mahmoud Darwich, Mohsen Amini Salehi, Ege Beyazit, Magdy A. Bayoumi |
Comput. J. | 4 |
| 2019 | Semi-Supervised EEG Signals Classification System for Epileptic Seizure DetectionabstractIn the past few decades, measuring and recording the brain electrical activities using Electroencephalogram (EEG) has become a standout amongst the tools utilized for neurological disorders' diagnosis, especially seizure detection. In this letter, a novel epileptic seizure detection system based on classifying raw EEG signals' recordings, eliminating the overhead of engineered feature extraction, is proposed. The system employs a mixing of unsupervised and supervised deep learning utilizing a one-dimensional convolutional variational autoencoder. To ascertain the robustness of the system against classifying unseen data, the evaluation of the proposed system is done using k-fold cross-validation. The classification results between normal and ictal cases have achieved a 100% accuracy while the classification results between the normal, inter-ictal and ictal cases accomplished a 99% overall accuracy which makes our system one of the most efficient among other state-of-the-art systems. Ahmed M. Abdelhameed, Magdy A. Bayoumi |
IEEE Signal Process. Lett. | 2 |
| 2019 | Performance Analysis and Modeling of Video Transcoding Using Heterogeneous Cloud ServicesabstractHigh-quality video streaming, either in form of Video-On-Demand (VOD) or live streaming, usually requires converting (i.e., transcoding) video streams to match the characteristics of viewers' devices (e.g., in terms of spatial resolution or supported formats). Considering the computational cost of the transcoding operation and the surge in video streaming demands, Streaming Service Providers (SSPs) are becoming reliant on cloud services to guarantee Quality of Service (QoS) of streaming for their viewers. Cloud providers offer heterogeneous computational services in form of different types of Virtual Machines (VMs) with diverse prices. Effective utilization of cloud services for video transcoding requires detailed performance analysis of different video transcoding operations on the heterogeneous cloud VMs. In this research, for the first time, we provide a thorough analysis of the performance of the video stream transcoding on heterogeneous cloud VMs. Providing such analysis is crucial for efficient prediction of transcoding time on heterogeneous VMs and for the functionality of any scheduling methods tailored for video transcoding. Based upon the findings of this analysis and by considering the cost difference of heterogeneous cloud VMs, in this research, we also provide a model to quantify the degree of suitability of each cloud VM type for various transcoding tasks. The provided model can supply resource (VM) provisioning methods with accurate performance and cost trade-offs to efficiently utilize cloud services for video streaming. Xiangbo Li, Mohsen Amini Salehi, Yamini Joshi, Mahmoud Darwich, Brad Landreneau, Magdy A. Bayoumi |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2018 | A Comparative Analysis on Resource Discovery Protocols for The Internet of ThingsabstractResources discovery is a fundamental requirement to the full realization of the vision of Internet of Things. Discovery includes resource properties, capabilities, and metadata. It enables consumers to build IoT applications and services that utilize “smart things” with no prior knowledge about these things. This paper studies three of the most commonly used discovery protocols, CoAP, MQTT and UPnP. We compare between their performance and behavior in IoT deployments. Each protocol is implemented and deployed on a mobile phone as a client and a Raspberry Pi as a broker/server. The implementation includes a WeMo switch and a TI SensorTag as resources. The paper provides insights on the different features, behavioural attributes, and a comparative analysis between these protocols. We also present performance indexes of each protocol in terms of memory and CPU usage, latency, and traffic exchange between a publisher and a subscriber. Our analysis shows that despite the three protocols are fundamentally different and each one has pros and cons, they are all highly useful in IoT environments. However, CoAP is generally more flexible and scalable, but requires a higher memory footprint. Kasem Khalil, Khalid Elgazzar, Magdy A. Bayoumi |
GLOBECOM | 3 |
| 2018 | A Low Power Hardware Implementation of Multi-Object DPM Detector for Autonomous DrivingabstractObject detection is a fundamental process in traffic management systems and self-driving cars. Deformable part model (DPM) is a popular and competitive detector for its high precision. This paper presents a programmable, low power hardware implementation of DPM based object detection for real-time applications. Our approach employs a very fast object detection pipeline with complementary techniques such as fast feature pyramid, Fast Fourier Transform (FFT) and early classification to accelerate DPM with a reasonable accuracy loss and achieves a speed-up of 50x and 6x over original DPM and cascade DPM respectively on single core CPU. The hardware circuit uses 65nm CMOS technology and consumes only 36.5mW (0.81 nJ/pixel) based on the post-layout simulation. The ASIC has an area of 3362 kgates and 295.5 KB on-chip memory and the design utilizes two simultaneous engines to process two independent object categories with 8 deformable parts per category. Alaa Ali, Oladiran G. Olaleye, Bappaditya Dey, Kasem Khalil, Magdy A. Bayoumi |
ICASSP | 5 |
| 2018 | Semi-Supervised Deep Learning System for Epileptic Seizures Onset PredictionabstractThe advance prediction of seizures before its onset has been a challenging task for scientists for a long time. It is still the epileptic patients' hope to find an effective way of preventing seizures to improve the quality of their lives. In this paper, using an innovative mixing of unsupervised and supervised deep learning techniques, we propose a novel epileptic seizure prediction system using electroencephalogram (EEG) recordings from the human brains. The proposed system is built upon classifying between the interictal and the preictal brain states. The proposed system uses two-dimensional deep convolutional autoencoder for learning the best discriminative spatial features from the multichannel unlabeled raw EEG recordings. A Bidirectional Long Short-Term Memory recurrent neural network is used for classification based on the temporal information. To help achieve faster learning and reliable convergence for our system, the transfer learning technique is used for initializing the weights for the patient-specific networks. Within, up to one hour of prediction window, our system achieved an average sensitivity of 94.6% and average low false prediction alarm rate of 0.04FP/h which makes it one of the most efficient among state-of-the-art methods. Ahmed M. Abdelhameed, Magdy A. Bayoumi |
ICMLA | 2 |
| 2018 | Empirical Activation Function Effects on Unsupervised Convolutional LSTM LearningabstractThis paper empirically evaluates and analyzes the effect of the choice of recurrent activation and unit activation functions on the unsupervised convolutional LSTM learning process. The goal of this work is to provide guidance for selecting the optimal non-linear activation function for the convolutional LSTM models which target the video prediction problem. This paper shows an empirical analysis of different non-linear activation functions that are commonly implemented in different deep learning APIs. We used the moving MNIST dataset as the most common benchmark for video prediction problems. Nelly Elsayed, Anthony S. Maida, Magdy A. Bayoumi |
ICTAI | 3 |
| 2018 | A Cost-Effective Self-Healing Approach for Reliable Hardware SystemsabstractIn this paper, self-healing concept for hardware systems is investigated and a new approach is proposed. Hardware systems have been proposing imitations to biological organisms in the way they offer healing and recovery abilities. Digital systems with inspired homogeneous architecture have improved capabilities to compensate for any faults. Self-healing is defined by the ability of a system to detect faults or failures and fix them. One of the main problems in current self-healing approaches is area overhead and scalability for complex structures considering they are based on redundancy and spare blocks. This paper proposes a different approach for self-healing based on embryonic structures without a need for spare cells. The area overhead is lower compared to other approaches relying on spare cells. The proposed approach relies on time multiplexing two functions in one cell within one clock cycle. The reliability of the proposed technique is studied and compared to conventional system with different failure rates. This approach is capable of healing up to 50% of the cells where each cell can cover another neighbor failed cell at most. The area overhead is 9% for the proposed approach which is much lower compared to other approaches using spare cell. The proposed approach is applied to investigate two case studies; ALU array, and neural network. Kasem Khalil, Omar Eldash, Magdy A. Bayoumi |
ISCAS | 3 |
| 2018 | Neuro-NoC: Energy Optimization in Heterogeneous Many-Core NoC using Neural Networks in Dark Silicon EraabstractDue to the end of Dennard Scaling and the rise of dark silicon, it is essential to design energy-efficient heterogeneous NoC under critical power and thermal constraints. The challenge is to determine and configure NoC resources while meeting the application(s) requirements. Because of the large and complex many-core NoC design space (voltage/frequency scaling, link bandwidth, power-gating, etc.), design space becomes difficult to explore within a reasonable time for optimal decision at run-time. Furthermore, reactive resource management is not effective in preventing problems, such as creating thermal hotspots and exceeding power budget, from happening. Therefore, we propose a Neuro-NoC model, which utilizes neural networks learning algorithm to dynamically monitor, predict, and configure NoC resources based on online learning of the system status. Distributed cluster-wise neural network and a global neural network model for resource monitoring and configuration in many-core NoC has been proposed. Simulations demonstrate that Neuro-NoC can predict the global optimal NoC configuration with high accuracy (88%), sensitivity (97% true positive), and specificity (88% true negative). Md Farhadur Reza, Tung Thanh Le, Bappaditya Dey, Magdy A. Bayoumi, Danella Zhao |
ISCAS | 4 |
| 2018 | Cost-Efficient and Robust On-Demand Video Transcoding Using Heterogeneous Cloud ServicesabstractVideo streams, either in the form of Video On-Demand (VOD) or live streaming, usually have to be converted (i.e., transcoded) to match the characteristics of viewers' devices (e.g., in terms of spatial resolution or supported formats). Transcoding is a computationally expensive and time-consuming operation. Therefore, streaming service providers have to store numerous transcoded versions of a given video to serve various display devices. With the sharp increase in video streaming, however, this approach is becoming cost-prohibitive. Given the fact that viewers' access pattern to video streams follows a long tail distribution, for the video streams with low access rate, we propose to transcode them in an on-demand (i.e., lazy) manner using cloud computing services. The challenge in utilizing cloud services for on-demand video transcoding, however, is to maintain a robust QoS for viewers and cost-efficiency for streaming service providers. To address this challenge, in this paper, we present the Cloud-based Video Streaming Services (CVS2) architecture. It includes a QoS-aware scheduling component that maps transcoding tasks to the Virtual Machines (VMs) by considering the affinity of the transcoding tasks with the allocated heterogeneous VMs. To maintain robustness in the presence of varying streaming requests, the architecture includes a cost-efficient VM Provisioner component. The component provides a self-configurable cluster of heterogeneous VMs. The cluster is reconfigured dynamically to maintain the maximum affinity with the arriving workload. Simulation results obtained under diverse workload conditions demonstrate that CVS2 architecture can maintain a robust QoS for viewers while reducing the incurred cost of the streaming service provider by up to 85 percent. Xiangbo Li, Mohsen Amini Salehi, Magdy A. Bayoumi, Nian-Feng Tzeng, Rajkumar Buyya |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | Correlation-based detection of TCM signals for cognitive radiosabstractIn this work, the inter-dependency of TCM signals is studied. Using this inter-dependency, correlation-based detectors are proposed for spectrum sensing of TCM signals in white Gaussian noise. In particular, a constant false alarm rate (CFAR) detector is presented and its performance is evaluated using simulations. We also describe an application of our detector for the classification of uncoded modulation systems vs. their TCM counterparts and present numerical results on their performance. Reza Soosahabi, Mort Naraghi-Pour, Nasim Nasirian, Magdy A. Bayoumi |
ICASSP | 4 |
| 2017 | Fault tolerant techniques for TSV-based interconnects in 3-D ICsabstractTSV-based 3-D IC is one of the most promising approach in modern integrated circuit design. As TSV enables the integration of different dies with different technologies, it is considered as one of the key elements in IC design. Reliability of TSV is an important issue which needs to be addressed very carefully since a faulty TSV can result in the failure of the whole stack. TSV open defect alongside TSV short are the most prevalent faults in 3-D ICs. They can happen during TSV filling, bonding process or as a result of aging due to Electromigration. In this paper, we present two methods for reducing the amount of delay caused by a partially open TSV and TSV short defect in 3D-ICs. Simulation results demonstrate the effectiveness of the proposed method in improving the performance of 3-D ICs in presence of partially open and short TSV defects. Siroos Madani, Magdy A. Bayoumi |
ISCAS | 2 |
| 2017 | Dark silicon-power-thermal aware runtime mapping and configuration in heterogeneous many-core NoCabstractTo address power-thermal-dark silicon issues in many-core chip, run time task-resource and voltage co-allocation with reconfigurable network-on-chip (NoC) framework for energy and hotspots minimization is proposed in this work. At runtime, the global manager with the help of proposed MinEnergy mapping algorithm reconfigures the NoC links bandwidth and nodes voltage-level and power-gated the resources depending on the traffic demand and resource statistics collected from the distributed cluster-managers. MinEnergy mapping algorithm minimizes overall chip power and thermal hotspots in heterogeneous large-scale NoC. We have formulated the mapping and configuration problem into a linear optimization model and implemented a traditional minimum-path contiguous mapping for comparisons. Simulations show that MinEnergy dynamic mapping solution is 80-90% close to the optimal solution, and significantly better than the minimum-path mapping solution. Md Farhadur Reza, Danella Zhao, Magdy A. Bayoumi |
ISCAS | 3 |
| 2017 | Real-time streaming challenges in Internet of Video Things (IoVT)abstractThe limited capabilities of Internet of Things (IoT) devices make real-time video streaming a major challenge. Video encoding and transmission are computationally intensive processes. Applications; like urban surveillance and health care monitoring, require real-time high definition video streams. Current video encoders are not designed to meet these requirements, thus alternative architectures, algorithms and compression techniques are needed to meet these goals. In this paper, real-time video streaming challenges for IoT applications are presented. Ahmed Sammoud, Ashok Kumar 0001, Magdy A. Bayoumi, Tarek A. Elarabi |
ISCAS | 3 |
| 2016 | High Performance On-demand Video Transcoding Using Cloud ServicesabstractVideo streams, either in form of on-demand streaming or live streaming, usually have to be converted (i.e., transcoded) based on the characteristics (e.g., spatial resolution) of clients' devices. Transcoding is a computationally expensive operation, therefore, streaming service providers currently store numerous transcoded versions of the same video to serve different types of client devices. However, recent studies show that accessing video streams have a long tail distribution. That is, there are few popular videos that are frequently accessed while the majority of them are accessed infrequently. The idea we propose in this research is to transcode the infrequently accessed videos in a on-demand (i.e., lazy) manner. Due to the cost of maintaining infrastructure, streaming service providers (e.g., Netflix) are commonly using cloud services. However, the challenge in utilizing cloud services for video transcoding is how to deploy cloud resources in a cost-efficient manner without any major impact on the quality of video streams. To address the challenge, in this research, we present an architecture for on-demand transcoding of video streams. The architecture provides a platform for streaming service providers to utilize cloud resources in a cost-efficient manner and with respect to the Quality of Service (QoS) requirements of video streams. In particular, the architecture includes a QoS-aware scheduling component to efficiently map video streams to cloud resources, and a cost-efficient dynamic (i.e., elastic) resource provisioning policy that adapts the resource acquisition with respect to the video streaming QoS requirements. Xiangbo Li, Mohsen Amini Salehi, Magdy A. Bayoumi |
CCGrid | 3 |
| 2016 | CVSS: A Cost-Efficient and QoS-Aware Video Streaming Using Cloud ServicesabstractVideo streams, either in form of on-demand streaming or live streaming, usually have to be converted (i.e., transcoded) based on the characteristics of clients' devices (e.g., spatial resolution, network bandwidth, and supported formats). Transcoding is a computationally expensive and time-consuming operation, therefore, streaming service providers currently store numerous transcoded versions of the same video to serve different types of client devices. Due to the expense of maintaining and upgrading storage and computing infrastructures, many streaming service providers (e.g., Netflix) recently are becoming reliant on cloud services. However, the challenge in utilizing cloud services for video transcoding is how to deploy cloud resources in a cost-efficient manner without any major impact on the quality of video streams. To address this challenge, in this paper, we present the Cloud-based Video Streaming Service (CVSS) architecture to transcode video streams in an on-demand manner. The architecture provides a platform for streaming service providers to utilize cloud resources in a cost-efficient manner and with respect to the Quality of Service (QoS) demands of video streams. In particular, the architecture includes a QoS-aware scheduling method to efficiently map video streams to cloud resources, and a cost-aware dynamic (i.e., elastic) resource provisioning policy that adapts the resource acquisition with respect to the video streaming QoS demands. Simulation results based on realistic cloud traces and with various workload conditions, demonstrate that the CVSS architecture can satisfy video streaming QoS demands and reduces the incurred cost of stream providers up to 70%. Xiangbo Li, Mohsen Amini Salehi, Magdy A. Bayoumi, Rajkumar Buyya |
CCGrid | 3 |
| 2016 | Towards real-time DPM object detector for driver assistanceabstractAutomatic object detection is a rapidly evolving area in surveillance and autonomous vehicles. Deformable part model (DPM) is a well-known object detector for its high precision and speed bottleneck. This paper proposes a very fast object detection pipeline based on complementary techniques to accelerate DPM. A recent fast feature pyramid technique is employed with look-up table HOG features, Fast Fourier Transform and early classification technique to speed up DPM and maintain its accuracy. We exploit SIMD optimization and multiple cores to achieve a real time detector. Our results shows that we achieve a speed-up of 50x and 6x on a single core over DPM and cascade DPM respectively. Our optimized version processes a 640×480 pixel image at 38 fps. Alaa Ali, Magdy A. Bayoumi |
ICIP | 2 |
| 2016 | An Energy-Detection-Based Cooperative Spectrum Sensing Scheme for Minimizing the Effects of NPEE and RSPFabstractFor improved spectrum utilization, the key technique for acquiring spectrum situational awareness (SSA) -- spectrum sensing -- is greatly improved by cooperation among the active spectrum users, as network size increases. However, the many cooperative spectrum sensing (CSS) schemes that have been proposed are based on the assumptions of accurate noise power estimates, characterizable variation in noise level and absence of false or malicious users. As part of a series of SSA research projects, in this research work, we propose a novel scheme for minimizing the effects of noise power estimation error (NPEE) and received signal power falsification (RSPF) by energy-based reliability evaluation. The scheme adopts the Voting rule for fusing multiple spectrum sensing data. Based on simulation results, the proposed scheme yields significant improvement, 68.2 - 88.8%, over the conventional CSS schemes, when compared on the basis of the schemes' stability to uncertainties in noise and signal power. Oladiran G. Olaleye, Ahmed Aly, Dmitri D. Perkins, Magdy A. Bayoumi |
MSWiM | 5 |
| 2016 | Introducing a Novel Smart Design Framework for a Reconfigurable Multi-Processor Systems-on-Chip (MPSoC) ArchitectureabstractIn this work, a smart designing framework for a Systems- on-Chip has been proposed. We examined a reconfigurable MPSoC architecture which incorporates one Processor-FPGA core to ensure flexibility and better design parameters. The objective of this work is to propose and develop a smart framework which initiates this philosophy: the system would build a better system by itself. The system would learn about usage statistics by using an android application. Then the system would form a decision function and classify the user with the help of support vector machine-based machine learning algorithm. According to user's preference, it would re-design the reconfigurable MPSoC to ensure customized and superior user experience. The machine learning algorithm runs on cloud for saving computing power and resources. An image processing task has been performed as a case-study on a FPGA-SoC platform and on GPU as a proof of concept to ensure current standards. As of our knowledge, this is a novel and competent approach of designing system for hand-held devices which enables customization of the device after the manufacturer's end. Anandi Dutta, Magdy A. Bayoumi |
SMARTCOMP | 2 |
| 2015 | ASIC implementation of a computationally efficient compressive sensing detection method using least squares optimization in 45 nm CMOS technologyabstractThis paper presents a high speed architecture of a recently proposed compressive sensing detection method for wideband cognitive radios using least squares. Using least squares instead of the orthogonal matching pursuit for signal recovery reduces the computational complexity where, the index search and matrix inverse stages are avoided. The proposed architecture is fully pipelined where, 14 clock cycles are required to detect 1024-length signal occupying 8 channels from 16 measurements. The design is implemented in 45 nm CMOS operating at 165 MHz. Since, the sensing time for 1024-length signal is roughly 84.8 ns, the proposed design offers high speed signal detection compared with state of art orthogonal matching pursuit architecture. Mohamed Shaban, Tarek A. Idriss, Haytham Idriss, Magdy A. Bayoumi |
ICASSP | 4 |
| 2015 | Adaptive neural matching online spike sorting VLSI chip design for wireless BCI implantsabstractControlling the surrounding world by just the power of our thoughts has always seemed to be just a fictional dream. With recent advancements in technology and research, this dream has become a reality for some through the use of a Brain Computer/Machine Interface (BCI/BMI). One of the most important goals of BCI is to enable handicap people to control artificial limbs. Some research proposed wireless implants that do not require chronic wound in the skull. However, the communications consume a high bandwidth and power that exceeds the allowed limits, 8–10mW. This study proposes and implements a modified version of real-time spike sorting for wireless BCI [4] that simplifies and uses less computation via an adaptive neural-structure; which makes it simpler, faster and power and area efficient. The system was implemented, and simulated using Modalism and Cadence, with ideal case and worst case accuracy of 100% and 91.7%, respectively. Also, the chip layout of 0.704mm2, with power consumption of 4.7mW and was synthesized on 45nm technology using Synopsys. Zaghloul Saad Zaghloul, Magdy A. Bayoumi |
ICASSP | 2 |
| 2015 | Building and Evaluating COTS Based Optical Interlinks for NanosatellitesabstractNanosatellites in low earth orbit are becoming more popular due their proven performance using accessible COTS components. Optical interlinks for such satellites have been suggested before as an alternative to RF links. In this paper we present the steps involved in selecting suitable COTS components for such links and evaluating them as well. Finally we compare the achievable performance of these links to what current low power RF technologies offer. Tarek A. Idriss, Magdy A. Bayoumi |
ICCCN | 2 |
| 2015 | An overview of IEEE standardization efforts for cognitive radio networksabstractThe Cognitive Radio (CR) technology is envisioned as a primary solution for the spectrum scarcity problem in wireless communication networks. Cognitive radio networking allows the exploitation of the unutilized spectrum bands in an opportunistic manner, and thereby, improves the efficiency of the spectrum utilization. Cognitive radio Networks (CRNs) provide wireless connectivity via heterogeneous wireless architectures and dynamic spectrum access techniques. In this paper, we present an overview of the state-of-the-art of the IEEE standardization efforts for CR and CRNs. Ahmed K. F. Khattab, Magdy A. Bayoumi |
ISCAS | 2 |
| 2014 | Optimal Probabilistic Encryption for Secure Detection in Wireless Sensor NetworksabstractWe consider the problem of secure detection in wireless sensor networks operating over insecure links. It is assumed that an eavesdropping fusion center (EFC) attempts to intercept the transmissions of the sensors and to detect the state of nature. The sensor nodes quantize their observations using a multilevel quantizer. Before transmission to the ally fusion center (AFC), the senor nodes encrypt their data using a probabilistic encryption scheme, which randomly maps the sensor's data to another quantizer output level using a stochastic cipher matrix (key). The communication between the sensors and each fusion center is assumed to be over a parallel access channel with identical and independent branches, and with each branch being a discrete memoryless channel. We employ J-divergence as the performance criterion for both the AFC and EFC. The optimal solution for the cipher matrices is obtained in order to maximize J-divergence for AFC, whereas ensuring that it is zero for the EFC. With the proposed method, as long as the EFC is not aware of the specific cipher matrix employed by each sensor, its detection performance will be very poor. The cost of this method is a small degradation in the detection performance of the AFC. The proposed scheme has no communication overhead and minimal processing requirements making it suitable for sensors with limited resources. Numerical results showing the detection performance of the AFC and EFC verify the efficacy of the proposed method. Reza Soosahabi, Mort Naraghi-Pour, Dmitri D. Perkins, Magdy A. Bayoumi |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2013 | A novel clustering paradigm for key pre-distribution: Toward a better security in homogenous WSNsabstractHomogeneous wireless sensor networks (HWSNs) are a class of WSNs in which nodes have identical hardware configurations. A HWSN inherits the same constraints from WSNs such as limited processing capacity and resources. However, the unique characteristic of HWSNs gets more attention for recent security applications. The environment of these applications mandates lightweight security provisions, especially a key management protocol that endures security threats such as node capture attack. On the other hand, the limited computation and communication capabilities of sensor nodes make it hard to deal with such security threats. Thus, in this paper, we introduce a novel clustering paradigm to be used in a key pre-distribution key management scheme designed for HWSNs. Then, we develop a cluster-based key pre-distribution based on this paradigm. Our developed scheme improves the security dramatically by partitioning sensor nodes to the disjointed clusters and letting clusters to be inter-connected via randomly selected nodes as consular nodes from all the clusters to form a universal cluster. As opposed to the adopted schemes, we provide a comprehensive analysis for both sensor node connectivity, and resiliency against node capture attack, and also computational results are presented. Mohammad Rezaeirad, Mahdi Orooji, Sahar Mazloom, Dmitri D. Perkins, Magdy A. Bayoumi |
CCNC | 5 |
| 2013 | A Novel Authenticated Encryption Algorithm for RFID SystemsabstractIn this paper a light symmetric encryption algorithm is presented for resource constrained applications like RFID systems. In this algorithm some extra bits are distributed among plaintext bits where the location of these bits inside the cipher text is the secret key. The algorithm provides confidentiality, authentication, and integrity services. Experimental results confirm that its area overhead and power overhead is less than other known symmetric algorithms proposed for RFID systems. Zahra Jeddi, Esmaeil Amini, Magdy A. Bayoumi |
DSD | 3 |
| 2013 | Design, Implementation and Characterization of Practical Distributed Cognitive Radio NetworksabstractOpportunistic Spectrum Access (OSA) in distributed cognitive radio networks (CRNs) has been well studied in the literature from a theoretical perspective. However, such theoretically-optimized distributed OSA approaches are challenged by several practical implementation issues. In this paper, we design a custom cross-layer framework that enables the: (i) clean-slate implementation of a wide variety of OSA mechanisms; (ii) experimental evaluation of the individual practical OSA components of the Rate-Adaptive Probabilistic (RAP) framework; and (iii) detailed comparison of the performance of such a practical OSA approach against theoretical OSA approaches developed for fully-capable CRNs. Our evaluation reveals the multi-fold goodput improvement and remarkable fairness characteristics of the practical RAP OSA approach compared to the OSA approaches that overlook the OSA and CR practical limitations. However, the superior performance of practical OSA comes at the expense of more outages to the primary licensed networks but within the permissible bounds. Another key finding is that the wide family of existing theoretically-optimized OSA protocols can benefit from the gains available to the individual components of the practical RAP approach, namely, the random spectrum sensing and the probabilistic non-greedy access. Ahmed K. F. Khattab, Dmitri D. Perkins, Magdy A. Bayoumi |
IEEE Trans. Commun. | 3 |
| 2012 | RBS: Redundant Bit Security Algorithm for RFID SystemsabstractIn this paper, we propose a symmetric encryption algorithm for RFID systems called RBS (Redundant Bit Security). RBS is based on inserting redundant bits into the original data bits. The location of redundant bits inside the transmitted data represents the secret key. The redundant bits are generated by Light Weight Mac algorithm in order to provide integrity and authentication as well. The implemented hardware architecture meets the resource constraints of RFID systems, and encryption and decryption time is shorter compared to existing algorithms. Zahra Jeddi, Esmaeil Amini, Magdy A. Bayoumi |
ICCCN | 3 |
| 2012 | Structure generation and design of tracking ADCsabstractIn this paper, a new model for a low power CMOS current-mode Tracking Analog-to-Digital Converter (ADC) is proposed. The model generates several designs that differ in power consumption and signal bandwidth specifications. As an example, A 6-bit ADC, with a maximum acquisition speed of 130 MHz, is implemented in a 1 V analog supply voltage. Spectre simulation results for the proposed tracking ADC verifying the analytical results are also given. It shows that the circuit consumes a static power of 0.63 mW and a dynamic power of 0.49 mW in a commercial 90 nm CMOS process. Mohamed O. Shaker, Magdy A. Bayoumi |
ISCAS | 2 |
| 2012 | Experimental evaluation of Opportunistic Spectrum Access in distributed cognitive radio networksabstractOpportunistic Spectrum Access (OSA) is foreseen as the future of wireless communications. However, today's cognitive radio technologies lag far behind the OSA goals and do not allow for the applicability of exiting theoretical distributed OSA techniques. In this paper, we experimentally demonstrate the ability of realizing OSA despite the practical limitations of existing transceiver technologies. We use the general purpose Wireless open-Access Research Platform (WARP) to instrument the implementation of fundamental OSA functionalities. Then we implement a suite of OSA schemes using this implementation framework. Our experiments show that suboptimal but practical OSA approaches such as random spectrum sensing and non-greedy access achieve superior performance given cognitive radios with limited capabilities compared to OSA approaches that are optimized for fully-capable cognitive radio networks (e.g., with wide-band sensing capability and adopt winner-takes-all access relying on network-wide coordination mechanisms). Ahmed K. F. Khattab, Dmitri D. Perkins, Magdy A. Bayoumi |
IWCMC | 3 |
| 2012 | Secure localization for wireless sensor networks using decentralized dynamic key generationabstractSecurity frameworks for wireless sensor network localization application can no longer be ignored. Wireless sensor network are being deployed in sensitive environment that require high levels of confidentiality, integrity and authenticity. Employing already existing security algorithms dedicated for wireless networks are infeasible for sensor network environment due to limited resources. In this context, a security framework tailored for wireless sensor network localization applications is proposed. The algorithm uses symmetric key encryption where the keys are generated in a decentralized fashion. These keys are computed on the fly, never transmitted and, changes autonomously with each transmission. Furthermore, the algorithm utilizes the bitwise XOR function for all its encryption needs which presents low overhead. Simulations show the robustness of the proposed algorithm in case of malicious node attack on the localization process and on key deduction and computation. Zaher Merhi, Amin Haj-Ali, Samih Abdul-Nabi, Magdy A. Bayoumi |
IWCMC | 4 |
| 2012 | Hardware architecture for fast Intra mode and direction prediction in real-time MPEG-2 to H.264/AVC transcoderabstractIn this article, we propose a hardware architecture for our fast Intra mode and direction prediction algorithm to accelerate the MPEG-2 to H.264/AVC transcoding devices. In order to eliminate the redundant operations in the transcoder, our implemented algorithm uses the DCT coefficients from the MPEG-2 decoder to predict the Intra mode and reconstruction direction for the H.264/AVC encoder. In addition, the Intra prediction process in the H.264/AVC part of the transcoder has been dramatically accelerated by using our full-search elimination technique. The empirical results show 92% reduction in the transcoding time while reducing the PSNR for less than 3.5%. The proposed architecture has achieved an operating frequency of 323MHz at a power consumption of 112 mW when implemented on Virtex-5 FPGA Development Board. Tarek A. Elarabi, Randa Ayoubi, Hanan A. Mahmoud, Magdy A. Bayoumi |
WOWMOM | 4 |
| 2012 | High-speed Motion Estimation Architecture for Real-time Video TransmissionabstractMotion estimation (ME) process consumes up to 70% of the total encoding time of video transmission. Because it has a high coding efficiency and it is very scalable, the full search (FS) algorithm is considered to be the most popular ME algorithm. However, the main drawback of the FS algorithm is that it is computationally intensive. For this reason, FS is rarely used for real-time video coding. This paper proposes a simple and fast adaptive search window size (ASWS) algorithm that eliminates a significant amount of computations from the conventional FS algorithm. This is achieved by dynamically reducing the required search area for each reference block. Simulation results show that more than 94% of the candidate blocks are eliminated by our algorithm without significant loss in visual quality. This paper also presents an efficient and high-speed ME engine (MEE) architecture for the proposed ASWS algorithm. The MEE efficiently reuses the search area data to minimize the memory I/O while fully utilizing the available hardware resources. A smart processing element design along with an innovative data scheduling scheme allows the search area data to flow both horizontally and vertically, whereas the current block data remain stationary. This allows the proposed architecture a simple and highly regular dataflow through the core. Simulation results show that for a search range of [−16,+15] and a block size of 16×16, the proposed architecture performs the ME for 60 fps of 4CIF video at 100 MHz and easily outperforms many FS architectures. Sumeer Goel, Yasser Ismail, Magdy A. Bayoumi |
Comput. J. | 3 |
| 2012 | Fast Motion Estimation System Using Dynamic Models for H.264/AVC Video CodingabstractH.264/AVC offers many coding tools for achieving high compression gains of up to 50% more than other standards. These tools dramatically increase the computational complexity of the block based motion estimation (BB-ME) which consumes up to 80% of the entire encoder's computations. In this paper, computationally efficient accurate skipping models are proposed to speed up any BB-ME algorithm. First, an accurate initial search center (ISC) is decided using a smart prediction technique. Thereafter, a dynamic early stop search termination (DESST) is used to decide if the block at the ISC position can be considered as a best match candidate block or not. If the DESST algorithm fails, a less complex style of the motion estimation algorithm which incorporates dynamic padding window size technique will be used. Further reductions in computations are achieved by combining the following two techniques. First, a dynamic partial internal stop search technique which utilizes an accurate adaptive threshold model is exploited to skip the internal sum of absolute difference operations between the current and the candidate blocks. Second, a dynamic external stop search technique greatly reduces the unnecessary operations by skipping all the irrelevant blocks in the search area. The proposed techniques can be incorporated in any block matching motion estimation algorithm. Computational complexity reduction is reflected in the amount of savings in the motion estimation encoding time. The novelty of the proposed techniques comes from their superior saving in computations with an acceptable degradation in both peak signal-to-noise ratio (PSNR) and bit-rate compared to the state of the art and the recent motion estimation techniques. Simulation results using H.264/AVC reference software (JM 12.4) show up to 98% saving in motion estimation time using the proposed techniques compared to the conventional full search algorithm with a negligible degradation in the PSNR by approximately 0.05 dB and a small increase in the required bits per frame by only 2%. Experimental results also prove the effectiveness of the proposed techniques if they are incorporated with any fast BB-ME technique such as fast extended diamond enhanced predictive zonal search and predictive motion vector field adaptive search technique. Yasser Ismail, Jason McNeely, Mohsen Shaaban, Hanan A. Mahmoud, Magdy A. Bayoumi |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2011 | High-performance asic architecture for hysteresis thresholding and component feature extraction in limited-resource applicationsabstractHysteresis thresholding offers enhanced object detection but is time-consuming. Hence, it is often avoided in real-time, constrained streaming processors. We propose a high-performance solution for limited-resource applications. This features two memory-efficient and fast architectures for hysteresis thresholding followed by object feature extraction: a lower-area uniform implementation and a faster pipelined one with slight area increase. Both designs couple thresholding with candidate pixel handling, component analysis, and feature extraction on the fly. This avoids multiple scans of the image, additional label resolving, and large image buffers; thus offering an average speedup up to 37× with about 99% memory reduction compared to state of the art schemes for VGA images. Moreover, the ASIC implementation allows this process to be done at 460fps, making the integration of this part into real-time detection systems very feasible. Mayssaa Al Najjar, Swetha Karlapudi, Magdy A. Bayoumi |
ICIP | 3 |
| 2011 | Distributed Kalman Filter using fast polynomial filterabstractDistributed estimation algorithms have received a lot of attention in the past few years, particularly in the fusion framework of Wireless Sensor Network (WSN). Distributed Kalman Filter (DKF) for WSN is one of the most fundamental distributed estimation algorithms for scalable wireless sensor fusion. In the literature, most of DKF methods rely on consensus filter algorithms. The convergence rate of such distributed consensus algorithms is slow and typically depends on the network topology and the weights given to the edges between neighboring sensors. In this paper, we propose a DKF based on polynomial filter to accelerate the distributed average consensus in the static network topologies. The main contribution of the proposed methodology is to apply a polynomial filter on the network matrix that will shape its spectrum in order to increase the convergence rate by minimizing its second largest eigenvalue. The simulation results show that the proposed algorithm increases the convergence rate of DKF by 4 times compared to the standard iteration. The proposed methodology can contribute in the real time WSN's applications. Ahmed Abdelgawad 0001, Magdy A. Bayoumi |
ISCAS | 2 |
| 2011 | ETSSI: Energy-based Task Scheduling Simulator for wireless sensor networksabstractDistributed processing has been a viable solution for enabling the next generation of real-time wireless sensor networks (WSN). Efficient task scheduling and allocation (TSA) policies guarantee that efficiency of the distribution. However, TSA policies in WSN face the challenges imposed by the wireless communication medium. This makes the accurate evaluation and verification of TSA policies difficult in live systems. Hence, developing a TSA simulator becomes essential to decrease the time for successful development and testing of relevant algorithms. This work addresses the need for a TSA simulator for WSN and develops ETSSI an Energy-based Task Scheduling Simulator. ETSSI is an event-driven, scalable, simulator which provides a user-friendly graphical interface. Its accuracy is more than 80% compared to in-door live implementations on a test bed of Telosb nodes. Most importantly, the TSA policy designer using ETSSI is only concerned about the application model, not the actual application implementation, which is mandatory in today's WSN simulators. Sherine Abdelhak, Chandra Sekhar Gurram, Jared Tessier, Soumik Ghosh, Magdy A. Bayoumi |
ISCAS | 5 |
| 2011 | P2E-DWT: A parallel and pipelined efficient VLSI architecture of 2-D Discrete Wavelet TransformabstractDiscrete Wavelet Transforms has surpassed its counterparts due to its attractive properties, and hence been adopted by image processing algorithms. However, with the emergence of real-time resource constrained embedded imaging platforms, DWT manifests as a bottleneck. This article presents a hardware implementation for 2-D DWT. An area-efficient, parallel and pipelined architecture is proposed with a modified image scan coupled with "multiplier-free" multiplications. Through simulations and implementation, the proposed scheme proves to be a fast, area and power efficient solution for DWT. Milad Ghantous, Magdy A. Bayoumi |
ISCAS | 2 |
| 2011 | A clock gated flip-flop for low power applications in 90 nm CMOSabstractA new clock gated flip-flop is presented. The circuit is based on a new clock gating approach to reduce the consumption of clock signal's switching power. It operates with no redundant clock cycles and has reduced number of transistors to minimize the overhead and to make it suitable for data signals with higher switching activity. The proposed flip-flop is used to design 10 bits binary counter and 14 bits successive approximation register. These applications have been designed up to the layout level with 1 V power supply in 90 nm CMOS technology and have been simulated using Spectre. Simulations with the inclusion of parasitics have shown the effectiveness of the new approach on power consumption and transistor count. Mohamed O. Shaker, Magdy A. Bayoumi |
ISCAS | 2 |
| 2011 | TALS: Trigonometry-based Ad-hoc Localization System for wireless sensor networksabstractNode localizations techniques are a critical part in many wireless sensor network (WSN) applications. There are several challenges for building an efficient localization system that is suitable for wireless sensor networks, mainly, reducing computational complexity and communication overhead. For example, solving an over constraint set of linear equations via least square approximations and utilizing square root operations are computationally intensive. Moreover, flooding techniques used for transmitting the locations of the anchors wastes bandwidth and energy. In this context, the Trigonometric based Ad-hoc Localization System (TALS) is an anchor-based range-based localization system that utilizes trigonometric identities and properties to compute the position of the node. TALS is designed to address the above challenges without deteriorating the quality of the estimates by taking advantage of redundancy and data fusion techniques. TALS is simulated and compared against popular localization techniques where it presented superiority against those techniques. Zaher Merhi, Mohamed A. Elgamel, Rafic Ayoubi, Magdy A. Bayoumi |
IWCMC | 4 |
| 2011 | Rate-adaptive probabilistic spectrum management for cognitive radio networksabstractExisting distributed opportunistic spectrum management schemes do not consider the inability of today's cognitive transceivers to measure interference at primary receivers. Consequently, optimizing the constrained cognitive radio network performance based only on local interference measurements at cognitive senders does not lead to truly optimal performance due to hidden and exposed primary senders. In this paper, we present the rate-adaptive probabilistic medium access protocol (RAP-MAC) for opportunistic spectrum management. In contrast to existing techniques, our approach addresses the unavoidable inaccuracy in spectrum sensing via a probabilistic spectrum access strategy. Furthermore, RAP-MAC allows multiple cognitive flows to fairly share the available bandwidth without explicit coordination by adopting a non-greedy transmission approach that prevents a single cognitive transmitter from monopolizing an available spectral opportunity. We analytically formulate the cognitive user performance optimization problem as a mixed integer non-linear programming to derive the optimal parameter values. Packet-level simulations show that RAP-MAC achieves between 65% to 119.5% higher goodput with significantly better fairness characteristics compared to existing greedy approaches that overlook the spectrum sensing limitations. Ahmed K. F. Khattab, Dmitri D. Perkins, Magdy A. Bayoumi |
WOWMOM | 3 |
| 2011 | Energy-Aware Distributed QR Decomposition on Wireless Sensor NodesabstractWireless sensor networks (WSNs) are starting to mature into the next generation where they can be used for adaptive filtering and signal processing, breaking away from the current generation of microcontroller applications. The tasks involved, however, are computationally intensive and strain the energy resources of any single computational sensor node. Moreover, most sensor nodes do not have the computational resources to complete many of these tasks repeatedly. Hence, exploring distributed processing on WSNs becomes a necessity to enable such computational load to be processed in real-time. In this work, a new distributed QR decomposition algorithm, on WSNs, is developed and implemented. QR decomposition has prominent applications in adaptive filtering which is essential for many WSN applications, such as target tracking and beamforming. The contributions of this work can be summarized as follows: (i) developing a new scalable tile-based distributed QR decomposition algorithm, (ii) distributing the least-squares problem based on the proposed distribution of the QR decomposition, (iii) developing resource-aware task allocation and mapping and (iv) developing a simple decentralized transmission scheduling scheme to guarantee efficient operation. This work demonstrates that distributed processing on WSNs paves the way for larger computations beyond the capabilities of a single node. This is accomplished while decreasing the energy per node and increasing the speed of the computation versus the implementation on a single node. The experiments, on a test bed of Telosb sensor nodes, prove that the proposed distributed algorithm enables higher computational capabilities while reducing the energy per node by up to 91.93% and speeding up the computation by up to 79.29% compared with running the QR decomposition on a single node, thus laying the foundation for energy-feasible real-time in-network processing. Sherine Abdelhak, Rabi S. Chaudhuri, Chandra Sekhar Gurram, Soumik Ghosh, Magdy A. Bayoumi |
Comput. J. | 5 |
| 2011 | Memory-Efficient Architecture for Hysteresis Thresholding and Object Feature ExtractionabstractHysteresis thresholding is a method that offers enhanced object detection. Due to its recursive nature, it is time consuming and requires a lot of memory resources. This makes it avoided in streaming processors with limited memory. We propose two versions of a memory-efficient and fast architecture for hysteresis thresholding: a high-accuracy pixel-based architecture and a faster block-based one at the expense of some loss in the accuracy. Both designs couple thresholding with connected component analysis and feature extraction in a single pass over the image. Unlike queue-based techniques, the proposed scheme treats candidate pixels almost as foreground until objects complete; a decision is then made to keep or discard these pixels. This allows processing on the fly, thus avoiding additional passes for handling candidate pixels and extracting object features. Moreover, labels are reused so only one row of compact labels is buffered. Both architectures are implemented in MATLAB and VHDL. Simulation results on a set of real and synthetic images show that the execution speed can attain an average increase up to 24× for the pixel-based and 52× for the block-based when compared to state-of-the-art techniques. The memory requirements are also drastically reduced by about 99%. Mayssaa Al Najjar, Swetha Karlapudi, Magdy A. Bayoumi |
IEEE Trans. Image Process. | 3 |
| 2011 | An Efficient Adaptive High Speed Manipulation Architecture for Fast Variable Padding Frequency Domain Motion EstimationabstractMotion estimation (ME) consumes up to 70% of the entire video encoder's computations and is, therefore, the main encoding-time consuming process. Discrete cosine transform (DCT)-based phase correlation along with dynamic padding (DP) are the recently evolved frequency domain ME (FDME) techniques that promise to efficiently reduce the computational complexity of the ME process. DP uses dynamic padding thresholds to select the proper search area size according to a pre-estimated set of motion vectors (MVs). The main drawbacks of using conventional DP in the frequency domain are two-fold. First, the dynamic thresholds need to be estimated in the pixel (IDCT) domain which increases complexity. Second, the mismatched transformed search area is formed from different successive transformed blocks, which would lead to an inaccurate ME if the search area is not manipulated. In this paper, an efficient low complexity algorithm and high speed architecture are proposed to implement an adaptive manipulation unit engine (MUE). The MUE, the main module of the FDME system, adaptively decides the padding size and forges a matched transformed search area from the successive transformed blocks. Additionally, the proposed utilized dynamic thresholds are efficiently estimated in the frequency domain (FD). The MUE architecture is presented with two different design implementations trading off the VLSI design parameters. Implementation and simulation results project that the proposed MUE, when integrated in a whole FDME system, can perform ME for 60 fps of 4CIF video at 172 MHz. Yasser Ismail, Mohsen Shaaban, Jason McNeely, Magdy A. Bayoumi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2010 | A fast discrete transform architecture for Frequency Domain Motion EstimationabstractFrequency Domain Motion Estimation (FDME) is a recent technique that promises to efficiently reduce the computational complexity of ME process. Related Transformed-Discrete Cosine Transform (RT-DCT) is one of the main modules that build the FDME encoder. The RT-DCT module is responsible for generating four transforms that are required for the ME process in the frequency domain. The main problem of generating such transforms is the low speed of the FDME encoder that prevents its use in real time applications. In this paper an efficient fast RT-DCT architecture is proposed to accelerate the encoding process in the frequency domain. The proposed architecture achieves approximately 58%, 39%, and 50% reductions in gate count, power consumption, and area compared to the conventional state of the art pipelined RT-DCT generators. Implementation and Simulation results project that the proposed RT-DCT architecture, when integrated in a whole FDME system, can perform ME for 60 fps of 4CIF video at 118 MHz. Yasser Ismail, Jason McNeely, Mohsen Shaaban, Mayssaa Al Najjar, Magdy A. Bayoumi |
ICIP | 5 |
| 2010 | A compact single-pass architecture for hysteresis thresholding and component labelingabstractHysteresis thresholding offers enhanced edge/object detection in the presence of noise. However, due to its recursive nature, it requires a lot of memory and execution time. Thus, it is restricted and sometimes totally avoided in streaming processors with limited memory. We propose an efficient architecture coupling hysteresis thresholding with component labeling and feature extraction in a single pass over the image pixels. The operations are performed on the fly while recycling labels to avoid additional passes for handling candidate pixels and extracting object features. Moreover, only one row of compact labels is buffered. Hence, the execution speed of the algorithm is increased and the memory requirements are drastically reduced when compared to state of the art techniques. When implemented on FPGA, this technique promises to offer even more speed up and efficient resource utilization. Mayssaa Al Najjar, Swetha Karlapudi, Magdy A. Bayoumi |
ICIP | 3 |
| 2010 | An efficient area manipulation architecture for frequency domain encoding processabstractFrequency-Domain Motion Estimation (FD-ME) evolved as a technique that would greatly reduce ME computations and the whole encoding time. In Dynamic Padding FD-ME (DP-FD-ME), a dynamic padding threshold adaptively selects the proper search area size according to a pre-estimated set of motion vectors. The main drawback of DP is the mismatched transformed search area formed from different consecutive transformed blocks which would lead to inaccurate ME. In this paper, efficient area and high speed Manipulation Unit Engine (MUE) is proposed, a main module of DP-FD-ME system, to forge a matched transformed search area from successive transformed blocks. Implementation results nominate the proposed MUE architecture to those multimedia applications that favors reducing both area and power consumption. Simulation results project that MUE, when integrated in a whole FD-ME system, can perform ME for 60 fps of 4CIF video at 172 MHZ. For the best of our knowledge, it is the first attempt to address such architecture. Yasser Ismail, Mohsen Shaaban, Jason McNeely, Mohamed A. Elgamel, Magdy A. Bayoumi |
ISCAS | 5 |
| 2009 | A multi-modal automatic image registration technique based on complex waveletsabstractImage registration is considered one of the most fundamental and crucial pre-processing tasks in digital imaging. This paper describes a fast multimodal automatic image registration algorithm that handles the alignment of IR and visible images. A multiresolution approach based on dual tree-complex wavelet transform is employed to speed up the process. At the coarsest level, an accurate registration estimate for higher levels is achieved, using edge detection and cross correlation. Mutual information, on the other hand, is applied at higher levels as a matching criterion applied to the six orientation bands of the complex wavelet. The process is completely automatic, and was tested on several sets of synthetic and real data. Experimental results show that the proposed technique exhibits better accuracy than DWT-based algorithms for uni and multi-modal cases. Milad Ghantous, Soumik Ghosh, Magdy A. Bayoumi |
ICIP | 3 |
| 2009 | An efficient adaptive manipulation architecture for real time video coding in Frequency DomainabstractMotion estimation (ME) consumes approximately up to 70% of the entire video encoder's computations and is it's main exhaustive time consuming process. Frequency-Domain Motion Estimation (FDME) evolved as a technique that would greatly reduce ME computations and the whole encoding time. In Dynamic Padding FDME (DP-FDME), a dynamic padding threshold adaptively selects the proper search area size according to a pre-estimated set of motion vectors. The main drawback of DP is the mismatched transformed search area formed from different consecutive transformed blocks which would lead to inaccurate ME. In this paper, efficient high speed architecture is proposed to implement an adaptive Manipulation Unit Engine (MUE), a main module of DP-FDME system, to forge a matched transformed search area from successive transformed blocks. Implementation results nominate the proposed MUE architecture to those multimedia applications that favors high speed processing as a trade-off to an acceptable increase in area and power. Simulation results project that MUE, when integrated in a whole FDME system, can perform ME for 60 fps of 4CIF video at 172 MHZ. Yasser Ismail, Mohsen Shaaban, Jason McNeely, Magdy A. Bayoumi |
ICIP | 4 |
| 2009 | "Voodoo" error prediction for bit-depth scalable video codingabstractA new prediction method for bit-depth scalable video coding is proposed in this paper. Correlation between the dropped bits during linear scaling tone mapping is used as the basis for using the error as an additional aid in predicting the high bit-depth enhancement layer. The residual to be transmitted in the enhancement layer is reduced by this technique. The effectiveness of this technique is increased when using a small quantization parameter. We show up to a 20% reduction in bit rate of the CAVLC output in our simulation test frames. Bit-depth scalability may be needed in future systems that will support legacy display applications as well as future high dynamic range displays. Jason McNeely, Magdy A. Bayoumi |
ICIP | 2 |
| 2009 | Robust object tracking using correspondence voting for smart surveillance visual sensing nodesabstractThis paper presents a bottom-up tracking algorithm for surveillance applications where speed and reliability in the case of multiple matches and occlusions are major concerns. The algorithm is divided into four steps. First, moving objects are detected using an accurate hybrid scheme with selective Gaussian modeling. Simple object features balancing speed, reliability, and complexity are then extracted. Objects are matched based on their spatial proximity and feature similarity. Finally, correspondence voting solves multiple match conflicts, segmentation errors, and occlusion cases. This approach is very simple, which makes it suitable for implementation at smart surveillance visual sensing nodes. Moreover, the simulation results demonstrate its robustness in detecting occlusions and correcting segmentation errors without any prior knowledge about the objects models or constraints on the direction of their motion. Mayssaa Al Najjar, Soumik Ghosh, Magdy A. Bayoumi |
ICIP | 3 |
| 2009 | Data Fusion Framework for Sand Detection in PipelinesabstractReliable sand detection is an important component of oil production system. In practice, produced sand in oil pipelines poses a serious problem in many production situations, since a small amount of sand in the produced fluid can result in significant erosion in a very short time stage. A new data fusion framework for sand detection in pipeline is presented. The framework is collecting data from oil pipeline using acoustic sensors (SENACO AS100) and flow analyzer (MC-II) in real time. The framework combines two modules: a wireless receiving and transmission (ReT) module and a data fusion module (DaF). The ReT module implementation is based on TinyOS and Crossbow MICAz motes. In order to optimize between the complexity and accuracy needs, DaF module is implemented using two methods; fuzzy art (FA) and maximum likelihood estimator (MLE). The results show the efficient number of sensors needed and compare between FA and MLE redundant. Ahmed Abdelgawad 0001, Zaher Merhi, Mohamed A. Elgamel, Magdy A. Bayoumi, Amal Zaki |
ISCAS | 4 |
| 2009 | Enhanced Efficient Diamond Search Algorithm for Fast Block Motion EstimationabstractIn this paper, a modified diamond search (MDS) algorithm is proposed for fast motion estimation based on the well known diamond search (DS) algorithm. A set of computationally efficient algorithms that can be applied to any block matching algorithm and is applied to the DS as a study case achieves higher complexity reduction than DS algorithm without further relative PSNR (peak signal to noise ratio) degradation compared to full search (FS). First, dynamic internal stop search (DISS) algorithm is used to reduce the internal redundant SAD (sum of absolute difference) operations between the current and the candidate blocks using an accurate dynamic threshold. Second, a dynamic external stop search (DESS) greatly reduces the unnecessary operations by skipping all the irrelevant blocks in the search area. In addition, early search termination and adaptive pattern selections techniques are applied to the proposed MDS as initialization steps to achieve even higher complexity reduction. The accuracy of the proposed model threshold equations guarantee not to fall into a local minima. Experiments show that the proposed MDS algorithm reduces the computations greatly up to 99% and 20% compared with the conventional FS algorithm and DS respectively with no significant degradation in both the PSNR and the bit-rate. Yasser Ismail, Jason McNeely, Mohsen Shaaban, Magdy A. Bayoumi |
ISCAS | 4 |
| 2009 | A Hybrid Adaptive Scheme based on Selective Gaussian Modeling for Real-time Object DetectionabstractObject detection is receiving a growing attention with the emergence of surveillance systems. This paper presents a hybrid adaptive scheme based on selective Gaussian modeling for detecting objects in complex outdoor scenes with gradual illumination changes and dense, moving background objects like swinging tree branches. The proposed technique combines simple frame difference (FD), simple adaptive background subtraction (BS), and accurate Gaussian modeling to benefit from the high detection accuracy of Mixture of Gaussian solution (MoG) in outdoor scenes while reducing the computations required, thus, making it faster and more suitable for real time surveillance applications. Moreover, by applying selective component matching and updating and hysteresis thresholding, the probability of detecting a background pixel as foreground decreases leading to better detection accuracy than MoG as demonstrated in the quantitative and qualitative comparison. Mayssaa Al Najjar, Soumik Ghosh, Magdy A. Bayoumi |
ISCAS | 3 |
| 2009 | A Low Complexity Inter Mode Decision for MPEG-2 to H.264/AVC Video Transcoding in Mobile EnvironmentsabstractThis paper presents a fast variable block size inter mode decision algorithm suitable for low complexity MPEG-2 to H.264/AVC heterogeneous video transcoding in mobile environments. Macroblock inter coding mode prediction in H.264/AVC represents almost 70% of its computational complexity. An efficient transcoder would take advantage of the information stored in MPEG-2 bitstream (ex. motion vectors, DCT coefficients, etc.) to simplify macroblock mode decision in H.264/AVC. The proposed fast variable block size inter mode decision (FVBSMD) algorithm conditionally reuse MPEG-2 motion vectors along with only few DCT coefficients to eliminate unnecessarily complexity from H.264/AVC variable block size motion re-estimation process. Simulations results; using video sequences with different motion complexities, resolutions and frame rates; show about 80% computational complexity elimination with only a 0.2 dB and 2.5% degradation in PSNR and bit-rate respectively. Mohsen Shaaban, Magdy A. Bayoumi |
ISM | 2 |
| 2009 | Fast Variable Padding Motion Estimation Using Smart Zero Motion Prejudgment Technique for Pixel and Frequency DomainsabstractMotion estimation (ME) plays an important role in modern video coders since it consumes approximately 60-80% of the entire encoder's computations. In this paper, three novel techniques are proposed to effectively speed up the ME process. First, a smart prediction technique for effectively deciding an initial search center is proposed. Second, a zero motion prejudgment technique is proposed to accurately decide whether the pre-estimated ISC can be considered as a best match motion vector (MV) and consequently save the required computations for the MV refinement process. Finally, a variable padding pixels ME technique is proposed to adaptively determine the number of padding pixels required for the search window for more computational cost savings. The three techniques are combined and applied to the block-based ME for a superior computational complexity savings in the ME process. The performance of the proposed techniques is tested in both the pixel domain ME and the frequency domain ME in terms of their quantitative visual quality (peak signal-to-noise ratio, PSNR), their computational complexity, and their bit rate. Experimental results demonstrate that the proposed fast ME technique is able to achieve approximately a 99.4% reduction in ME time compared to the conventional full search block-based ME (FSSBB-ME) with negligible degradation in both the PSNR and the bit rate. Additionally, the experimental results prove the effectiveness of the proposed techniques if they are combined with any block-based ME technique such as the fast extended diamond enhanced predictive zonal search. Experimental results also demonstrate that there is at least an additional savings of 72% in ME time using the conventional discrete cosine transform phase correlation ME (DCT-PC-ME) in the frequency domain compared to the conventional FSBB-ME technique in pixel domain. Compared to the conventional DCT-PC-ME, applying the proposed novel techniques to the DCT-PC-ME saves up to 89% in ME time. Yasser Ismail, Mohamed A. Elgamel, Magdy A. Bayoumi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | A Lightweight Collaborative Fault Tolerant Target Localization System for Wireless Sensor NetworksabstractEfficient target localization in wireless sensor networks is a complex and challenging task. Many past assumptions for target localization are not valid for wireless sensor networks. Limited hardware resources, energy conservation, and noise disruption due to wireless channel contention and instrumentation noise pose new constraints on designers nowadays. In this work, a lightweight acoustic target localization system for wireless sensor networks based on time difference of arrival (TDOA) is presented. When an event is detected, each sensor belonging to a group calculates an estimate of the target's location. A fuzzyART data fusion center detects errors and fuses estimates according to a decision tree based on spatial correlation and consensus vote. Moreover, a MAC protocol for wireless sensor networks (EB-MAC) is developed which is tailored for event-based systems that characterizes acoustic target localization systems. The system was implemented on MicaZ motes with TinyOS and a PIC 18F8720 microcontroller board as a coprocessor. Errors were detected and eliminated hence acquiring a fault tolerant operation. Furthermore, EB-MAC provided a reliable communication platform where high channel contention was lowered while maintaining high throughput. Zaher Merhi, Mohamed A. Elgamel, Magdy A. Bayoumi |
IEEE Trans. Mob. Comput. | 3 |
| 2009 | Low-Power Clocked-Pseudo-NMOS Flip-Flop for Level Conversion in Dual Supply SystemsabstractClustered voltage scaling (CVS) is an effective way to decrease power dissipation. One of the design challenges is the design of an efficient level converter with fewer power and delay overheads. In this paper, level-shifting flip-flop topologies are investigated. Different level-shifting schemes are analyzed and classified into groups: differential style, n-type metal-oxide-semiconductor (NMOS) pass-transistor style, and precharged style. An efficient level-shifting scheme, the clocked-pseudo-NMOS (CPN) level conversion scheme, is presented. One novel level conversion flip-flop (CPN-LCFF) is proposed, which combines the conditional discharge technique and pseudo-NMOS technique. In view of power and delay, the new CPN-LCFF outperforms previous LCFF by over 8% and 15.6%, respectively. Peiyi Zhao, Jason McNeely, Pradeep Golconda, Soujanya Venigalla, Magdy A. Bayoumi, Weidong Kuang, Luke Downey |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2008 | Assumers for high-speed single and multi-cycle on-chip interconnect with low repeater countabstractAchieving high-speed signaling across narrow deep-submicron wires with reduced repeater count is a major design challenge. A clocked repeater circuit, called assumer, that allows high-speed point-to-point signaling with single repeater per single-cycle wirelength is presented in this paper. Simulations at the 90-nm node on 3-mm to 10-mm range of wirelength considering minimum pitch intermediate and global metal layers show up to 31% delay reduction and up to 80% less repeater count compared to conventional repeated wires. However, assumer interconnect suffers from switching power overhead. Therefore, the proposed method is only suitable for designs where speed and area are the primary concerns. Charbel J. Akl, Magdy A. Bayoumi |
ACM Great Lakes Symposium on VLSI | 2 |
| 2008 | A gradient-based hybrid image fusion scheme using object extractionabstractThis paper presents a new hybrid image fusion scheme that combines features of pixel and region based fusion, to be integrated in a surveillance system. In such systems, objects can be extracted from the different set of images due to background availability, and transferred to the new composite image with no additional processing usually imposed by other fusion approaches. The background information is then fused in a multi-resolution pixel-based fashion using gradient-based rules to yield a more reliable feature selection. According to Piella and Petrovic quantitative evaluation metrics, the proposed scheme exhibits a superior performance compared to existing fusion algorithms. Milad Ghantous, Soumik Ghosh, Magdy A. Bayoumi |
ICIP | 3 |
| 2008 | Power analysis of the Huffman decoding treeabstractA power analysis of the Huffman decoding tree for variable length decoding is presented. Leakage power is shown to be the primary component for large table sizes, while dynamic power has a role in smaller table sizes. Area increases nearly linearly with increasing table size, regardless of the probabilities of the symbols for a given table size, while the delay depends more on the probabilities. Although delay of the tree can be larger than other table lookup types, it is acceptable for mobile devices in order to reduce power. Jason McNeely, Yasser Ismail, Magdy A. Bayoumi, Peiyi Zhao |
ICIP | 3 |
| 2008 | An efficient frequency domain intra prediction for H.264/AVCabstractThis paper presents an efficient intra prediction implementation for H.264/AVC in the frequency domain. Intra prediction in the frequency domain is vital for transform domain heterogeneous video transcoding in wireless and mobile networks. A limited computation capabilities constraint constitutes a burden over such networks. New distributed intra prediction arithmetic is proposed with only addition and shift operations for transform domain intra prediction modes computations. Compared to previous attempts the proposed method reduces the computations extensively by eliminating the expensive matrix multiplications and omits the need for excessive memory storage. Mohsen Shaaban, Magdy A. Bayoumi |
ICIP | 2 |
| 2008 | Cost-effective and low-power memory address bus encodingsabstractThis paper presents encoding methods that build on TO-encoding to achieve considerable reduction in memory address bus wires with small performance overhead, while reducing address bus switching activity. The known TO-encoding is combined with a variable cycle transmission technique that uses the TO's increment signal (INC) as an indicator of the number of cycles. Two of the proposed methods maintain the energy efficiency that is significantly reduced due to time-multiplexing. Benchmark experiments show that wire count of the instruction memory address bus can be reduced by around half at the price of 12.89% average performance penalty, while maintaining the low switching activity of the original TO-encoded bus. Charbel J. Akl, Magdy A. Bayoumi |
ISCAS | 2 |
| 2008 | A low-area, low-power programmable frequency multiplier for DLL based clock synthesizersabstractA simple low-area and low-power clock frequency multiplier is proposed for Delay Locked Loop (DLL) based clock synthesizers. In this circuit, 2n voltage controlled delay lines (VCDL) are used to multiply the input frequency of a clock signal by n. This frequency multiplier is less susceptible to jitter accumulation as it is a DLL-based design. The proposed circuit can operate at a substantially low supply voltage. Simulation results show that the proposed frequency multiplier dissipates about 10% to 50% less power than similar clock multiplier circuits. In addition, an architecture for programmable frequency multiplication has been proposed in this paper. Md. Ibrahim Faisal, Magdy A. Bayoumi |
ISCAS | 2 |
| 2008 | A generalized fast motion estimation algorithm using external and internal stop search techniques for H.264 video coding standardabstractIn this paper, a set of computationally efficient accurate skipping techniques are proposed for motion estimation. First, a partial internal stop search (ISS) technique which utilizes an accurate adaptive threshold model is exploited to skip the internal SAD (sum of absolute difference) operations between the current and reference blocks. Second, an external stop search (ESS) technique greatly reduces the unnecessary operations by skipping all the irrelevant blocks in the search area. The proposed techniques can be incorporated in any block matching motion estimation algorithm. Computational complexity reduction is reflected on the amount of saving in motion estimation encoding time. Simulation results using H.264 reference software (JM 12.4) show up to 71.26% saving in motion estimation time using the proposed techniques compared to the fast full search algorithm adopted in JM 12.4 with a negligible degradation in the PSNR by approximately 0.03 dB and a small increase in the required bits per frame by only 2%. Yasser Ismail, Jason McNeely, Mohsen Shaaban, Magdy A. Bayoumi |
ISCAS | 4 |
| 2008 | High speed single-ended pseudo differential current sense amplifier for SRAM cellabstractWith reducing feature sizes, SRAM stability has become a major concern for future technologies. This critical issue can be solved by using highly stable separate bit-line read SRAM cell, but access time improvement becomes critical, since differential sense amplifier cannot be used for single bit-line read operation. In this paper, a novel pseudo differential single ended current mode sense amplifier is proposed. We demonstrate that this design can deliver a performance similar to that of conventional current mode differential amplifier without using dual bit-line for read operation. The overall read operation delay of the proposed single-ended design is almost 60% less than conventional single-ended design in 90nm CMOS technology. The proposed design consumes 51.6% less energy than conventional design counterpart. Abhijit Sil, Eswar Prasad Kolli, Soumik Ghosh, Magdy A. Bayoumi |
ISCAS | 4 |
| 2008 | Reducing wakeup latency and energy of MTCMOS circuits via keeper insertionabstractA simple yet effective technique that aims at reducing the energy and latency overheads incurred during the wakeup period of MTCMOS circuits is presented in this paper. One or more high-Vth keepers are inserted in MTCMOS combinational logic to reduce the metastability time that causes excessive short circuit current during mode transition and to minimize spurious glitches at internal circuit nodes. Employing the proposed keeper insertion technique in a 16-bit MTCMOS adder, up to 17.5% average wakeup energy and 54.6% wakeup latency reductions are achieved with negligible runtime power and latency overheads, while maintaining the standby energy efficiency of the original MTCMOS design. Charbel J. Akl, Magdy A. Bayoumi |
ISLPED | 2 |
| 2008 | Transition Skew Coding for Global On-Chip InterconnectabstractThis paper presents new simulation results of the previously proposed transition skew coding (TSC) for global on-chip interconnects. Considering 2-GHz global clock frequency at the 90-nm node, we show that TSC can be applied to broad range of wire length on both semiglobal and global metal layers, while maintaining its energy efficiency and its advantages in terms of crosstalk reduction and signal integrity, and wiring and repeater area minimization. Miriam J. Akl, Magdy A. Bayoumi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | Reducing Interconnect Delay Uncertainty via Hybrid Polarity Repeater InsertionabstractCapacitive crosstalk between adjacent signal wires has significant effect on performance and delay uncertainty of point-to-point on-chip buses in deep submicrometer (DSM) VLSI technologies. We propose a hybrid polarity repeater insertion technique that combines inverting and non-inverting repeater insertion to achieve constant average effective coupling capacitance per wire transition for all possible switching patterns. Theoretical analysis shows the superiority of the proposed method in terms of performance and delay uncertainty compared to conventional and staggered repeater insertion methods. Simulations at the 90-nm node on semi-global METAL5 layer show around 25% reduction in worst case delay and around 86% delay uncertainty minimization compared to standard bus with optimal repeater configuration. The reduction in worst case capacitive coupling reduces peak energy which is a critical factor for thermal regulation and packaging. Isodelay comparisons with standard bus show that the proposed technique achieves considerable reduction in total buffers area, which in turn reduces average energy and peak current. Comparisons with staggered repeater which is one of the simplest and most effective crosstalk reduction techniques in the literature show that hybrid polarity repeater offers higher performance, less delay uncertainty, and reduced sensitivity to repeater placement variation. Charbel J. Akl, Magdy A. Bayoumi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2007 | Transition Skew Coding: A Power and Area Efficient Encoding Technique for Global On-Chip InterconnectsabstractGlobal signaling is becoming more and more challenging as technology scales down toward the deep submicron. We propose a new bus encoding technique, transition skew coding, that targets many of the global interconnects challenges such as crosstalk, peak energy and current, switching and leakage power, repeaters area, wiring area, signal integrity and noise. Simulations are done on different bus lengths using a 90nm library. Repeaters sizing and spacing are optimized, and the proposed encoded bus is compared against a standard bus and a bus with shields inserted between every two wires. The encoding and decoding latencies are also analyzed. Simulations show that transition skew coding is efficient in terms of energy and area with low encoding and decoding latency overhead. Charbel J. Akl, Magdy A. Bayoumi |
ASP-DAC | 2 |
| 2007 | Low Power Lookup Tables for Huffman DecodingabstractMobile video devices are energy constrained and therefore need to contain circuits that consume a minimum amount of energy. In this paper, different architectures of lookup tables for Huffman decoding are studied. The PLA type structure is common in these Huffman lookup tables because of their speed and simplicity advantage. However, we determine that the tree structure for table lookup can be a lower power alternative than a PLA structure in certain situations. These situations accounted for 56% of our total simulations runs, and of these runs, the average power savings of the tree in those situations was 78%. Another goal of this paper is to show the effects of varying the table size and varying the probability distributions of a table on power, area, and delay. Jason McNeely, Magdy A. Bayoumi |
ICIP (6) | 2 |
| 2007 | High Speed and Area-Efficient Multiply Accumulate (MAC) Unit for Digital Signal Prossing ApplicationsabstractA high speed and area-efficient merged multiply accumulate (MAC) units is proposed in this work. To realize the area-efficient and high speed MAC unit proposed in this work, first we examine the critical delays and hardware complexities of conventional MAC architectures to derive at a unit with low critical delay and low hardware complexity. The new architecture is based on binary trees constructed using a modified 4:2 compressor circuits. Reducing the overall area is achieved by the full utilization of the compressors instead of putting zeros in free inputs. Increasing the speed of operation is achieved by avoid using the modified compressor in the critical path. Feeding the bits of the accumulated operand into the summation tree before the final adder helps to increase the speed too. The proposed MAC unit and the previous merged MAC unit are mapped on a field programmable gate array (FPGA) chip, in order to compare between them. The simulation result shows that the proposed system for 8-bit, 16-bit, and 32-bit MAC unit reduces area by 6.25%, 3.2 %, and 2.5% and increases the speed by 14%, 16%, and 19% respectively. The experimental test for the proposed 8-bit MAC is done using XESS demo board (XSA-100, Spartan-X2S100tq144). Ahmed Abdelgawad 0001, Magdy A. Bayoumi |
ISCAS | 2 |
| 2007 | Pixel-Level Image Fusion Scheme based on Linear AlgebraabstractImage fusion refers to the process of integrating complementary image sources from multiple imaging sensor such that the resulting fused image improves the performance of computational analysis tasks such as segmentation, feature extraction and object recognition. The paper introduces a pixel-level image fusion scheme based on linear algebra. The image fusion process begins by computing the discrete wavelet transform of the source images. Then, the wavelet transform of the images are fused using a feature-based rule. A salient feature may extend to several pixels; therefore, a rule that can include a region of pixels containing it results in a more efficient integration. The fusion rule is based on a measurement of the linear dependency of a small window centered on the pixel under consideration. The linear dependency measurement is the Wronskian determinant that is a simple and rigorous test. The performance assessment of the proposed method is established by using mutual information measurement as well as root mean square error and peak signal to noise ratio. The simulation results show that the proposed method is an efficient approach to image fusion. Ruth Aguilar-Ponce, Jose Luis Tecpanecatl-Xihuitl, Ashok Kumar 0001, Magdy A. Bayoumi |
ISCAS | 4 |
| 2007 | Design and Realization of Analog Phi-Function for LDPC DecoderabstractOne of the ambitious design goals of future generations of wireless systems, including 4G, IEEE 802.11n/802.16 standards, is to reliably provide very high data rate transmission in real-time. This poses a challenge to find an optimal coding scheme that has good performance and can be efficiently implemented in hardware. The most well-known LDPC decoding algorithm is log sum product (log-SP) in which a set of calculations on a non-linear function called Phi-function is approximated by a minimum function. Until now this function has been implemented through look up tables (LUT). But this direct implementation is costly for hardware. Also LUTs are very sensitive to the number of quantization bits and number of LUT values. Therefore, we have proposed analog Phi-function. The design is easily scalable and reconfigurable for larger block sizes. Simulation results show that our proposed design dissipates only 18 nW. Abu Baker, Soumik Ghosh, Ashok Kumar 0001, Magdy A. Bayoumi, Rafic Ayoubi |
ISCAS | 4 |
| 2007 | An Adaptive Block Size Phase Correlation Motion Estimation Using Adaptive Early Search Termination TechniqueabstractAn adaptive block size phase correlation motion estimation (ABSPC-ME) with a smart adaptive early termination technique is proposed and implemented in this paper. With its performance, efficiency and complexity ABSPC-ME is compared to that of the original phase correlation (PC) and full search block matching (FSBM) techniques. Since the phase correlation method measures the motion directly from the phase correlation map, it gives a more accurate and robust estimate of the motion vector. Besides increasing the encoding quality, the complexity of the encoder and computational cost are also decreased. Results show that there is approximately 78% reduction in computations compared with the original PC technique and 96% compared with FSBM without a significant loss in the visual quality. Yasser Ismail, Mohsen Shaaban, Magdy A. Bayoumi |
ISCAS | 3 |
| 2007 | A Low Power 4-bit Interleaved Burst Sampling ADC for Sub-GHz Impulse UWB RadioabstractThis paper presents a low power 4-bit ADC for sub-GHz Ultra Wideband (UWB) receivers. The power efficiency is achieved by taking advantage of the low duty cycle feature of UWB impulse. After the synchronization is achieved, the burst-mode sampling approach is employed to avoid unnecessary operations. So, the ADC only samples at the time when a pulse is expected and stays in standby during the rest of the time. The proposed burst sampling ADC employs five interleaved pipeline flash ADCs controlled by a low duty cycle 25 MHz sampling clock with five different phases. The resistor ladder reference circuit is eliminated by using a modified Quantum Voltage comparator, which can generate the reference voltages internally. The proposed ADC has been designed and simulated by using TSMC 0.18μm CMOS process. Simulation results show that the proposed 4-bit ADC can operate at 1G sample/s for UWB impulse with power consumption of 7.6 mW. Xiaodong Zhang 0008, Magdy A. Bayoumi |
ISCAS | 2 |
| 2007 | A Low Power Domino with Differential-Controlled-KeeperabstractDomino circuits are used to achieve higher system performance than static CMOS techniques. This work briefly surveys domino keeper designs for high fan-in domino circuits. A new domino circuit structure is shown in this paper that reduces the power-delay-product over 16% as compared to previous domino techniques with keepers Peiyi Zhao, Jason McNeely, Magdy A. Bayoumi, Pradeep Golconda, Weidong Kuang |
ISCAS | 3 |
| 2007 | Fully Decentralized Weighted Kalman Filter for Wireless Sensor Networks with FuzzyART Neural NetworksabstractA framework for developing a decentralized weighted Kalman filter with FuzzyART neural networks is presented. Active sensors communication is restricted to a node-to-node basis in the vicinity of each cluster where readings are only shared between neighbors. FuzzyART neural network takes these readings and classifies them into categories based on estimated measurement accuracy where appropriate weights are assigned to each measurement according to a decision tree. Few modifications have been made to the traditional Information Kalman Filter to adapt with the algorithm developed herein. The FuzzyART model is layered below the Kalman filter in a way it detects faulty measurements and spatial alteration of the target and accommodates these changes to better estimate the target. Sensors with faulty measurements are inhibited form participating in the estimation process and its readings are neither processed nor further communicated thus achieving higher energy efficiency. Furthermore, introducing a FuzzyART model lies within the complexity constraints of a sensor node and have acceptable overhead. Results and simulations show the superiority of the proposed algorithm when compared with the traditional Information Filter. Zaher Merhi, Mohamed A. Elgamel, Magdy A. Bayoumi |
ISCC | 3 |
| 2007 | Hybrid multiplierless FIR filter architecture based on NEDAabstractThis paper presents new hybrid multiplierless finite impulse response (FIR) architecture based on New Distributed Arithmetic (NEDA). The hybrid structure is a trade off between direct form and transposed direct form that results in a reduction of the critical path and the size of the delays elements as well as the fan-out. While the multiplications involved in the hybrid structure are replaced by a butterfly adder tree. Compared with previous methods, our proposed architecture achieves an average of 20% less additions. Moreover, the design method is simple and achieves better results than previous methods. Jose Luis Tecpanecatl-Xihuitl, Ruth Aguilar-Ponce, Magdy A. Bayoumi |
VLSI-SoC | 3 |
| 2007 | A network of sensor-based framework for automated visual surveillance
Ruth Aguilar-Ponce, Ashok Kumar 0001, Jose Luis Tecpanecatl-Xihuitl, Magdy A. Bayoumi |
J. Netw. Comput. Appl. | 4 |
| 2007 | Low-Power Clock Branch Sharing Double-Edge Triggered Flip-FlopabstractIn this paper, a new technique for implementing low-energy double-edge triggered flip-flops is introduced. The new technique employs a clock branch-sharing scheme to reduce the number of clocked transistors in the design. The newly proposed design also employs conditional discharge and split-path techniques to further reduce switching activity and short-circuit currents, respectively. As compared to the other state of the art double-edge triggered flip-flop designs, the newly proposed CBS_ip design has an improvement of up to 20% and 12.4% in view of power consumption and PDP, respectively Peiyi Zhao, Jason McNeely, Pradeep Golconda, Magdy A. Bayoumi, Robert A. Barcenas, Weidong Kuang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2006 | Area-Efficient NEDA Architecture for The 1-D DCT/IDCTabstractA new distributed arithmetic has been applied to the 1-D DCT to produce a low power, high throughput architecture. In this paper, we apply NEDA to the even-odd decomposition matrices of the 8times8 forward and inverse DCT. We show that, with the proposed approach, the number of adders required for the adder array for the forward DCT and the inverse DCT is fewer than required if NEDA is applied directly to the 8times8 DCT and IDCT matrices. This reduction results in power savings, without decreasing the throughput. Also, for the inverse DCT, the number of adder stages is reduced, resulting in faster decoding Archana Chidanandan, Magdy A. Bayoumi |
ICASSP (3) | 2 |
| 2006 | Multi-Path Search Algorithm for Block-Based Motion EstimationabstractBlock-based motion estimation algorithms are widely adopted by video coding standards due to their simplicity and good distortion performance. Amongst these, pattern based search algorithms are very popular. The basic assumption made by these algorithms is that there exists a unimodal error surface within the search window. As seen in real world video sequences, this is far from reality and the algorithms often get stuck in local minimums yielding sub-optimal results. We propose a multi-path search (MPS) algorithm that utilizes more than one path to find the absolute minima. Wrong search paths are avoided early resulting in faster search. Better motion vectors are estimated increasing the video quality, consequently reducing the bitrate requirement. MPS algorithm offers flexibility and robustness with better speed and quality. Results show that a speedup of upto 34% is achieved with slight increase in PSNR as compared to its best competitor. When compared to full search, the proposed algorithm looses only 0.1%~2% in PSNR while saving 92%~95% computations. Sumeer Goel, Magdy A. Bayoumi |
ICIP | 2 |
| 2006 | A low-power clock frequency multiplierabstractA low-power output feedback controlled frequency multiplier is proposed for delay locked loop (DLL) based clock synthesizers. It uses N voltage controlled delay lines (VCDL) to multiply the input clock frequency by a factor of N/2. This frequency multiplier is less susceptible to jitter-accumulation. The proposed circuit can operate at a substantially low supply voltage. Simulation results show that the proposed frequency multiplier dissipates about 27% to 36% less power than other similar circuits. In addition, the proposed circuit can be easily programmed for generating various output clock frequencies. Md. Ibrahim Faisal, Magdy A. Bayoumi, Peiyi Zhao |
ISCAS | 2 |
| 2006 | A power-efficient architecture for EBCOT tier-1 in JPEG 2000abstractPower reduction has become a serious issue in recent years. In this paper, a power-efficient architecture for EBCOT tier-1 is proposed. EBCOT tier-1 architecture is divided into BC (bit-plane coding), AE (arithmetic encoding), and FIFO that connects BC with AE and balances the different throughput between them. In BC, simple control logics are added to reduce computation in bit-plane coding; in FIFO, memory access is reduced since AE is fed with fixed values instead of reading from FIFO; in AE, simple control logics are added to reduce computation in AE and forwarding technique combined with clock gating is adopted to reduce switching activities in the last two pipeline stages. Experimental results, with standard test image benchmarks, show that the proposed power reduction techniques keep the same system throughput and achieve about 48%, 16%, and 20% improvement for BC, FIFO, and AE, respectively, in the power consumption by comparison with the original architecture. Magdy A. Bayoumi |
ISCAS | 2 |
| 2006 | A low power adaptive transmitter architecture for low band UWB applicationsabstractAn adaptive CMOS UWB transmitter architecture based on programmable pulse position modulation (PPPM) is proposed. It can operate in high-speed mode and power-saving mode depending on various QoS constrains. When the data rate is the primary concern, the transmitter operates at high-speed mode. When the power consumption becomes the major constrain, the transmitter can switch to power-saving mode. The high frequency 80 MHz clock is generated from a 10 MHz low frequency reference by a DLL-based frequency multiplier. The proposed adaptive UWB transmitter has been designed and simulated by using TSMC 0.18mum CMOS process. Simulation results shows proposed transmitter can provide a flexible knob for cross-layer QoS optimization with low power consumption Xiaodong Zhang 0008, Magdy A. Bayoumi |
ISCAS | 2 |
| 2006 | A Three-Level Parallel High-Speed Low-Power Architecture for EBCOT of JPEG 2000abstractFor JPEG 2000-based multimedia systems, embedded block coding with optimized truncation (EBCOT) tier-1 has become a bottleneck for the entire system. EBCOT tier-1 is full with bit operation, so hardware implementation is more efficient in both system throughput and power consumption. In this paper, a three-level parallel high-speed power-efficient architecture for EBCOT tier-1 is proposed. This architecture is divided into bit-plane coding (BC), arithmetic encoding (AE), and first-in first-out (FIFO) that connects BC with AE and balances the different throughput between them. To improve the system throughput, three levels of parallelism in BC are adopted: 1) the parallelism among bit planes; 2) the parallelism among three pass scans; and 3) the parallelism among coding bits. AE is implemented in four pipeline stages. To achieve power efficiency, several techniques are applied: in BC, simple control logics are added to reduce computation in BC; in FIFO, memory access is reduced since AE is fed with fixed values instead of reading from FIFO; in AE, simple control logics are added to reduce computation in AE and forwarding technique combined with clock gating is adopted to reduce switching activities in the last two pipeline stages. The proposed architecture can encode one code block with size NtimesN in only around (0.35~0.46)timesNtimesN clock cycles. Experimental results, with standard test image benchmarks, show that the proposed power reduction techniques keep the same system throughput and achieve about 27% improvement in the power consumption by comparison with the architecture without these techniques Magdy A. Bayoumi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Design of Robust, Energy-Efficient Full Adders for Deep-Submicrometer Design Using Hybrid-CMOS Logic StyleabstractWe present a new design for a 1-b full adder featuring hybrid-CMOS design style. The quest to achieve a good-drivability, noise-robustness, and low-energy operations for deep submicrometer guided our research to explore hybrid-CMOS style design. Hybrid-CMOS design style utilizes various CMOS logic style circuits to build new full adders with desired performance. This provides the designer a higher degree of design freedom to target a wide range of applications, thus significantly reducing design efforts. We also classify hybrid-CMOS full adders into three broad categories based upon their structure. Using this categorization, many full-adder designs can be conceived. We will present a new full-adder design belonging to one of the proposed categories. The new full adder is based on a novel xor-xnor circuit that generates xor and xnor full-swing outputs simultaneously. This circuit outperforms its counterparts showing 5%-37% improvement in the power-delay product (PDP). A novel hybrid-CMOS output stage that exploits the simultaneous xor-xnor signals is also proposed. This output stage provides good driving capability enabling cascading of adders without the need of buffer insertion between cascaded stages. There is approximately a 40% reduction in PDP when compared to its best counterpart. During our experimentations, we found out that many of the previously reported adders suffered from the problems of low swing and high noise when operated at low supply voltages. The proposed full adder is energy efficient and outperforms several standard full adders without trading off driving capability and reliability. The new full-adder circuit successfully operates at low voltages with excellent signal integrity and driving capability. To evaluate the performance of the new full adder in a real circuit, we embedded it in a 4- and 8-b, 4-operand carry-save array adder with final carry-propagate adder. The new adder displayed better performance as compared to the standard full adders Sumeer Goel, Ashok Kumar 0001, Magdy A. Bayoumi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2006 | MAC-SCC: a medium access control protocol with separate control channel for reconfigurable multi-hop wireless networksabstractIn this paper, we propose a novel medium access control protocol with a separate control channel (MAC-SCC) to increase the channel efficiency and address the unfairness and instability problems of IEEE 802.11 MAC protocol. In MAC-SCC, the available bandwidth is partitioned into two channels: a data channel and a control channel, each associated with a network allocation vector (NAV). To reduce hardware complexity, the station transmits or receives on one channel only at any given time. In the network employing MAC-SCC, the next data frame can be pre-scheduled during the current data transmission via the separate control channel, and thus reducing the frame collision probability and the bandwidth wasted during backoff. Moreover the use of the separate control channel helps to achieve fair medium access and solve the instability problem resulted from frequent link failures. The optimal bandwidth partitioning between the two channels is analyzed via a statistical model, which shows 10% bandwidth for the control channel and 90% bandwidth for the data channel. The performance of MAC-SCC is quantified via extensive simulations in both a stand-alone simulator developed by using PARSEC and a comprehensive network simulator called QualNet with whole protocol stack. Our results show that MAC-SCC can effectively reduce the link failure probability, achieve fair medium access when running multiple TCP sessions, and yield a throughput gain up to 60% under high traffic load Hongyi Wu, Nian-Feng Tzeng, Dmitri D. Perkins, Magdy A. Bayoumi |
IEEE Trans. Wirel. Commun. | 5 |
| 2005 | Noise-tolerant high fan-in dynamic CMOS circuit designabstractScaling CMOS technology to next generation improves performance, increases transistor density, and reduces power consumption per device. However, scaling also increases the subthreshold leakage current which greatly degrades the circuit's noise immunity. In this paper we propose a new circuit technique that makes domino dynamic CMOS more robust and more noise-tolerant with minimal performance degradation and energy overhead. Simulations for high fan-in gates show a noise immunity improvement of 2.13X using Berkeley Predictive Technology Models (BPTM) of 70nm with minimal performance and power degradations over standard domino circuits. Walid Elgharbawy, Pradeep Golconda, Magdy A. Bayoumi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2005 | Novel high-throughput EBCOT architecture for JPEG2000abstractEmbedded block coding with optimized truncation (EBCOT) consumes more than 50% of the processing time in JPEG200 encoding system. Hardware implementation with careful handling for the control nature of tier-1 is essential. Although, some architectures have been developed to speed up the coding operations, they still require a tedious checking mechanism to decide if each sample is eligible or not for coding. We propose a novel checking scheme for the three coding passes for EBCOT that works in parallel with the encoding process to achieve the required high throughput. The simulation results show that the proposed architecture increases the throughput by 19% on average compared, to other well known architectures. Ramy E. Aly, Beth Wilson, Magdy A. Bayoumi |
ICASSP (5) | 3 |
| 2005 | Three-level parallel high speed architecture for EBCOT in JPEG2000abstractA multi-level parallel high speed architecture for embedded block coding with optimized truncation (EBCOT) tier-1 in JPEG2000 is proposed. To increase the system throughput, this architecture adopts three levels of parallelism: 1) parallelism among bit-planes - all the bit-planes can be processed simultaneously; 2) parallelism among three pass scannings - three passes scan one bit-plane in parallel; 3) parallelism among coding bits - bits that are coded in different passes can be coded simultaneously without any conflict. Experimental results show that the proposed architecture can encode one code block with size N/spl times/N in only 0.6/spl times/N/spl times/N clock cycles, and is twice as fast as the fastest architecture in the literature so far. Magdy A. Bayoumi |
ICASSP (5) | 2 |
| 2005 | Efficient shield insertion for inductive noise reduction in nanometer technologiesabstractWith high clock frequencies, faster transistor rise/fall time, wider wires, and the use of Cu material interconnects, interconnect inductive noise is becoming an important design metric in digital circuits. An efficient technique to reduce the inductive noise of on-chip interconnects is to insert shields among signal wires. An efficient solution for the min-area shield insertion problem to satisfy given explicit noise bounds in multiple coupled nets is provided. The proposed algorithm determines the locations and number of shields needed to satisfy certain noise constraints. Experimental results show that the proposed approach minimizes the number of shields required to satisfy the noise constraints and uses less runtime than the best alternative reported approach. Mohamed A. Elgamel, Ashok Kumar 0001, Magdy A. Bayoumi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2004 | Dynamic profiling algorithms for low bit rate video applicationsabstractWe propose three algorithms for low bit-rate (LBR) video transmission, namely simple dynamic profiling (SimpleDP), minimum dynamic profiling (MinimumDP) and mean dynamic profiling (MeanDP). These techniques can be used in conjunction with other low bit rate techniques. Many of the techniques available for low bit rate video applications either require a lot of hardware resources or take advantage of the powerful computer platform that is used for their software implementation. On the other hand, the proposed algorithms are devised to target hardware implementation for systems with limited hardware resources. When compared to the coarse quantization technique, the new techniques not only achieve better compression ratio (twice that of coarse quantization), but also result in better PSNR results (27 dB for coarse quantization versus 33 to 36 dB for the new techniques). Tarek Darwish, Magdy A. Bayoumi |
ICASSP (5) | 2 |
| 2004 | A data merging technique for high-speed low-power multiply accumulate unitsabstractIn an attempt to meet the low-power requirements of high performance portable signal processing VLSI systems, while maintaining a high operating speed, a new data merging architecture for high-speed multiply accumulate units is proposed. The architecture can be applied on binary trees constructed using 4:2 compressor circuits. Increasing the speed of operation is achieved by taking advantage of the available free input lines of the compressor circuits, which result from the natural parallelogram shape of the generated partial products and using the bits of the accumulated value to fill in these gaps. This results in merging the accumulation operation within the multiplication process. Ayman A. Fayed, Walid Elgharbawy, Magdy A. Bayoumi |
ICASSP (5) | 3 |
| 2004 | A methodology for low power scheduling with resources operating at multiple voltages
Ashok Kumar 0001, Magdy A. Bayoumi, Mohamed A. Elgamel |
Integr. | 2 |
| 2004 | High-performance and low-power conditional discharge flip-flopabstractIn this paper, high-performance flip-flops are analyzed and classified into two categories: the conditional precharge and the conditional capture technologies. This classification is based on how to prevent or reduce the redundant internal switching activities. A new flip-flop is introduced: the conditional discharge flip-flop (CDFF). It is based on a new technology, known as the conditional discharge technology. This CDFF not only reduces the internal switching activities, but also generates less glitches at the output, while maintaining the negative setup time and small D-to-Q delay characteristics. With a data-switching activity of 37.5%, the proposed flip-flop can save up to 39% of the energy with the same speed as that for the fastest pulsed flip-flops. Peiyi Zhao, Tarek Darwish, Magdy A. Bayoumi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2003 | Noise tolerant low voltage XOR-XNOR for fast arithmeticabstractWith scaling down to deep submicron and nanometer technologies, noise immunity is becoming a metric of the same importance as power, speed, and area. Smaller feature sizes, low voltage, and high frequency are the characteristics for deep submicron circuits. This paper proposes a low voltage noise tolerant XOR-XNOR gate with 8 transistors. The proposed gate has been implanted in an already existing (5-2) compressor cell to test its driving capability. The proposed gate is characterized and compared with those published ones for reliability and energy efficiency. The average noise threshold energy (ANTE) and the energy normalized ANTE metrics are used for quantifying the noise immunity and energy efficiency respectively. Results using 0.18m CMOS technology and HSPICE for simulation show that the proposed XOR-XNOR circuit is more noise-immune and displays better power-delay product characteristics than the compared circuit. Also, the circuit proves to be faster in operation and works at all ranges of supply voltage starting from 0.6V to 3.3V. Mohamed A. Elgamel, Sumeer Goel, Magdy A. Bayoumi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2003 | Parallel high-speed architecture for EBCOT in JPEG2000abstractIn this paper, a parallel high-speed architecture for EBCOT is proposed, based on the parallel mode in JPEG2000. It discovers the parallelism among three passes in bit-plane coding: two passes can work on the coding operation without the addition of coding processing elements (PE) by using parallel context modeling. So it manages to make two bits encoded in one clock cycle. In order to keep high throughput and reduce the memory requirement, the pipelined pass-switching arithmetic encoder is adopted. The experimental results show that the proposed architecture reduces the processing time by more than 16.6% compared with the pass-parallel mode architecture in Chiang et al. (2002) and by more than 38% compared with the serial mode architecture in Chen et al. (2001). Ramy E. Aly, Magdy A. Bayoumi, Samia A. Mashali |
ICASSP (2) | 3 |
| 2003 | Memory accesses reduction for MIME algorithmabstractPower consumption of digital systems has become a critical design parameter. An important class of digital systems includes applications such as video image processing and speech recognition, which are extremely memory dominant. In such systems, a significant amount of power is consumed during memory accesses. Reducing the number of memory accesses can considerably impact the power dissipation in the rest of the design. Therefore, optimizing an application for reduced memory access can greatly effect the overall power consumption in the entire system. This paper presents an architectural enhancement multi-stage interval-based motion estimation (MIME) algorithm that not only saves power by reducing the number of memory accesses but also significantly increases the speedup. Sumeer Goel, Mohsen Shaaban, Tarek Darwish, Hanan A. Mahmoud, Magdy A. Bayoumi |
ICME | 5 |
| 2003 | Efficient Mapping Algorithm of Multilayer Neural Network on Torus ArchitectureabstractThis paper presents a new efficient parallel implementation of neural networks on mesh-connected SIMD machines. A new algorithm to implement the recall and training phases of the multilayer perceptron network with back-error propagation is devised. The developed algorithm is much faster than other known algorithms of its class and comparable in speed to more complex architecture such as hypercube, without the added cost; it requires O(1) multiplications and O(log N) additions, whereas most others require O(N) multiplications and O(N) additions. The proposed algorithm maximizes parallelism by unfolding the ANN computation to its smallest computational primitives and processes these primitives in parallel. Rafic Ayoubi, Magdy A. Bayoumi |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2002 | A merged multiplier-accumulator for high speed signal processing applicationsabstractIn an attempt to improve the speed of signal processing VLSI systems, a new architecture for high speed Multiply Accumulate Units is proposed. The architecture is based on Binary trees constructed using 4-2 compressor circuits. Increasing the speed of operation is achieved by taking advantage of the available free input lines of the 4-2 compressors, which result from the parallelogram shape of the generated partial products, and using the bits of the accumulated value to fill in these gaps. This results in merging the accumulation operation within the multiplication process. An 8-bit Multiplier Accumulator prototype circuit using the proposed architecture is prototyped in 0.35 micron double metal CMOS technology and simulated using hspice. Simulation results at 3.3 V show that the proposed architecture has a delay of 4.26 ns with a 16.8 delay savings. At 150 MHz operating frequency, the power consumption is 324 mWatts with a 23.04% power saving compared to other architectures not using the merging technique. Ayman A. Fayed, Magdy A. Bayoumi |
ICASSP | 2 |
| 2002 | Algorithm-based low-power VLSI architecture for 2D mesh video-object motion trackingabstractThe new VLSI architecture for video object (VO) motion tracking uses a novel hierarchical adaptive structured mesh topology. The structured mesh offers a significant reduction in the number of bits that describe the mesh topology. The motion of the mesh nodes represents the deformation of the VO. Motion compensation is performed using a multiplication-free algorithm for affine transformation, significantly reducing the decoder architecture complexity. Pipelining the affine unit contributes a considerable power saving. The VO motion-tracking architecture is based on a new algorithm. It consists of two main parts: a video object motion-estimation unit (VOME) and a video object motion-compensation unit (VOMC). The VOME processes two consequent frames to generate a hierarchical adaptive structured mesh and the motion vectors of the mesh nodes. It implements parallel block matching motion-estimation units to optimize the latency. The VOMC processes a reference frame, mesh nodes and motion vectors to predict a video frame. It implements parallel threads in which each thread implements a pipelined chain of scalable affine units. This motion-compensation algorithm allows the use of one simple warping unit to map a hierarchical structure. The affine unit warps the texture of a patch at any level of hierarchical mesh independently. The processor uses a memory serialization unit, which interfaces the memory to the parallel units. The architecture has been prototyped using top-down low-power design methodology. Performance analysis shows that this processor can be used in online object-based video applications such as MPEG-4 and VRML. Wael Badawy, Magdy A. Bayoumi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2002 | Performance analysis of low-power 1-bit CMOS full adder cellsabstractA performance analysis of 1-bit full-adder cell is presented. The adder cell is anatomized into smaller modules. The modules are studied and evaluated extensively. Several designs of each of them are developed, prototyped, simulated and analyzed. Twenty different 1-bit full-adder cells are constructed (most of them are novel circuits) by connecting combinations of different designs of these modules. Each of these cells exhibits different power consumption, speed, area, and driving capability figures. Two realistic circuit structures that include adder cells are used for simulation. A library of full-adder cells is developed and presented to the circuit designers to pick the full-adder cell that satisfies their specific applications. Ahmed M. Shams, Tarek Darwish, Magdy A. Bayoumi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2000 | A Multiplication-Free Parallel Architecture for Affine TransformationabstractThis paper presents a novel low power parallel architecture for computing affine transformation (AT). It is based on a new multiplication-free algorithm that employs the inherent algebraic properties of the AT. Low power has been achieved at the algorithmic level by replacing the multiplication with shifting operation, at the architecture level by using parallel computational units, and at the circuit level by using low power cells. The proposed architecture can be used as a computational kernel in object-based video processing. It is compatible with MPEG-4 and VRML standards. The architecture has been prototyped in 0.6 /spl mu/m CMOS technology with three layers of metal. Wael Badawy, Magdy A. Bayoumi |
ASAP | 2 |
| 2000 | A 108 Gbps, 1.5 GHz 1D-DCT ArchitectureabstractA high-performance ID-DCT architecture is proposed. It is based on the New Distributed Arithmetic Architecture algorithm (NEDA). Enhancements to NEDA are proposed to reduce the number of computations. Only addition operations are used, with 42 additions to compute the outputs for a 8/spl times/1 DCT. No subtractions, multiplications, or ROM are needed. High-throughput is achieved by pipelining the architecture. In every clock cycle, it receives eight pixels (each is 9-bits) as inputs, and produces eight DCT coefficients (each is 14-bits). The delay of one pipeline stage is the delay of a 3-level 4:2 compressor tree. The architecture is implemented in 0.35 /spl mu/m technology; it runs at 1.5 GHz, and processes 108 Gbps of image/video sequence data. Ahmed M. Shams, Magdy A. Bayoumi |
ASAP | 2 |
| 2000 | An Efficient Low-Bit Rate Motion Compensation Technique Based on QuadtreeabstractSummary form only given. A quadtree structured motion compensation technique effectively utilizes the motion content of a frame as oppose to a fixed size block motion compensation technique. In this paper, we propose a novel quadtree-structured region-wise motion compensation technique that divides a frame into equilateral triangle blocks using the quadtree structure. Arbitrary partition shapes are achieved by allowing 4-to-1, 3-to-1 and 2-to-1 merge/combine of sibling blocks having the same motion vector. We propose an optimal code scheme and a temporal predictive coding for the quadtree. Simulation results show that our techniques reduce the bit rate by 40% compared to other methods. Hanan A. Mahmoud, Magdy A. Bayoumi |
Data Compression Conference | 2 |
| 2000 | An Efficient Successive Elimination Algorithm for Block-Matching Motion EstimationabstractSummary form only given. Low power VLSI video compression processors are in high demand for the emerging wireless video applications. Video compression processors include VLSI implementation of a motion estimation algorithm. Many motion estimation algorithms are found in the literature. Some of them are fast but cannot guarantee an optimal solution; they can be stuck in local optima. Such algorithms are fast, consumes less power when implemented in VLSI, but they can result in high levels of distortion that cannot be accepted in many applications. On the other hand the full search block matching algorithm (FSBM) is computationally intensive and a VLSI implementation of such an algorithm has high power consumption. This paper presents an exhaustive search algorithm for block matching motion estimation. The proposed algorithm reduces the computational load with successive elimination of non-candidate blocks in the search window. Our proposed algorithm assigns each pixel to a category depending on its value. The number of categories is predetermined. The algorithm consists of number of stages, the first of which has the fewer categories, eliminating those search points that are the farthest from the match. The last stage is the FSBM but with fewer search points. This computational reduction leads to low-power VLSI implementation of the algorithm. Also, it leads to faster efficient motion estimation procedure. The correctness of this algorithm and its complexity are proved. Hanan A. Mahmoud, Magdy A. Bayoumi |
Data Compression Conference | 2 |
| 2000 | Compressed Domain Texture Classification from a Modified EZW Symbol StreamabstractSummary form only given. Researchers have demonstrated the effectiveness of wavelet subband energy values as a feature for representing and classifying texture images. However, the extraction of these texture features from compressed data can be cumbersome using traditional decompress-process approaches. A method has been developed for calculating wavelet energy features directly from a compressed embedded zerotree wavelet (EZW) symbol stream. The resulting technique is efficient and requires less memory than traditional approaches. In order to simplify the detection of subbands within the compressed data stream, end-of-subband markers have been inserted during the dominant pass of the EZW coding process. After compressing the test image, the reconstruction values described in Shapiro (1993) are used to calculate the energy of each subband. Following this technique, the memory requirements are reduced since the image is no longer reconstructed prior to the energy calculations. Additionally, the potentially large reconstruction matrix is no longer traversed which also reduces the time complexity. Beth Wilson, Magdy A. Bayoumi |
Data Compression Conference | 2 |
| 2000 | On minimizing hierarchical mesh coding overhead: (HASM) hierarchical adaptive structured mesh approachabstractThis paper presents an efficient mesh coding technique suitable for MPEG-4 video applications. The proposed technique significantly reduces the number of bits that are used to describe the mesh topology. It uses an adaptive structured mesh from coarse to fine, which can be coded as a count of splitting instead of nodes' locations. In the case of the quadtree, less than one bit per node can be achieved. This reduction induces an improvement of either the image quality or the global bit rate. Wael Badawy, Magdy A. Bayoumi |
ICASSP | 2 |
| 2000 | Low Power Video Object Motion-Tracking Architecture for Very Low Bit Rate Online Video ApplicationsabstractThis paper presents a low power VLSI architecture for video object motion-tracking that can be used in very low bit rate online video applications. Power has been reduced at both algorithmic and arithmetic levels. The video object motion-tracking architecture consists of two main parts, a mesh-based motion estimation unit and a mesh-based motion compensation unit. The mesh-based motion estimation unit implements parallel block matching motion estimation units to optimize the latency. The mesh-based motion compensation unit uses parallel multiplication-free affine core. The architecture has been prototyped and its performance measures have been evaluated. This processor can be used in online object-based video applications. Wael Badawy, Magdy A. Bayoumi |
ICCD | 2 |
| 2000 | A New Block-Matching Motion Estimation Algorithm Based on Successive EliminationabstractThis paper presents an exhaustive search algorithm for block matching motion estimation. The proposed algorithm reduces the computational load with successive elimination of non-candidate blocks in the search window. The proposed algorithm locates the global optima as located by the full search block-matching algorithm. This computational reduction leads to low-power VLSI implementation of the algorithm. Also, it leads to a faster efficient motion estimation procedure. The correctness of this algorithm and its complexity are presented. Simulation results on benchmark video clips are also presented. Hanan A. Mahmoud, Magdy A. Bayoumi |
ICIP | 2 |
| 2000 | A VLSI architecture for hierarchical mesh based motion compensation using scalable affine transformation coreabstractThis paper presents a VLSI architecture for hierarchical mesh based motion compensation. It uses a hierarchical adaptive structured mesh, which minimizes the number of bits describing the mesh. The mesh is constructed as triangular patches that describe the motion at different resolutions. Image warping is used to reconstruct the frame, whereas an affine transformation is used to texture map the triangular patches. The architecture implements a scalable affine transformation core, which can be used with any level of hierarchical mesh. The architecture has been prototyped and performance measures have been conducted. The prototype can be used as building block for MPEG-4 codec. Wael Badawy, Michael Talley, Magdy A. Bayoumi |
ISCAS | 4 |
| 2000 | A high-performance 1D-DCT architectureabstractA high performance 1D-DCT is proposed. It is based on a new distributed arithmetic architecture technique (NEDA). Only addition operations are used, with 35 additions to complete the first phase of the computations. The final phase is the primitive add-and-shift operation. No subtraction, multiplication, or ROM are needed. High-throughput is achieved by pipelining the architecture. The delay of one stage is the delay of one 12-bit addition. Low power consumption is achieved by reducing the computation requirements compared to other implementations of the 1D-DCT. Ahmed M. Shams, W. David Pan, Archana Chandanandan, Magdy A. Bayoumi |
ISCAS | 4 |
| 2000 | Low-bit-rate generalized quad-tree motion compensation algorithm and its optimal encoding schemes
Hanan A. Mahmoud, Magdy A. Bayoumi |
VCIP | 2 |
| 2000 | Video codec incorporating block-based multihypothesis motion-compensated prediction
Hanan A. Mahmoud, Magdy A. Bayoumi |
VCIP | 2 |
| 1999 | Novel Formulations for Low-Power Binding of Function Units in High-Level SynthesisabstractThis paper considers minimizing the switchings of the function units through new binding formulations. Switching activities on the function units are gathered through profiling the data-flow graph of the design at hand. Several thousand random input streams are generated and used for such profiling. The switching activities are obtained and stored in a matrix. The problem of binding the function units for low power is then formulated and solved by using the proposed methods. Ashok Kumar 0001, Magdy A. Bayoumi |
ICCD | 2 |
| 1999 | Performance Analysis for a New Automatic Error Control System for Overcoming Turbulent Fading Errors for Wireless ATM NetworksabstractThis paper suggests and analyzes the performance of a new wireless ATM adaptive error control system, which can be used in wireless ATM networks to improve quality of service in keeping the cell loss under a certain predefined limit and reducing the average delay of cell encoding. The new system is based on multi-level forward error correction (FEC), and specifically the Bose-Chaudhuri-Hocqenghem (BCH) code, which has multiple error correcting capability and a wide class of block codes, to overcome random errors. Realistic error patterns caused by multipath turbulent fading are analytically approximated and applied to the suggested system. Performance analysis for the new system showed that it improves cell loss rate to a value of less than 10/sup -9/ in case of non-fading and it keeps the average cell loss close to a reasonable value (approximately 10/sup -7/) in case of fading, while maintaining an average encoding delay considerably less than the delay of a system using a fixed BCH correction capability. M. Mazen Al-Khatib, Magdy A. Bayoumi |
ISCC | 2 |
| 1999 | Performance analysis for a new automatic error control system for overcoming atmospheric multipath fading with random bit errors for wireless ATM networksabstractThis paper suggests and analyzes the performance of a new wireless ATM adaptive error control system, which can be used in wireless ATM networks to improve the quality of service in keeping the cell loss under a certain predefined limit and reducing the average delay of cell encoding. The new system is based on multi-level forward error correction (FEC) and specifically the Bose-Chaudhuri-Hocqenghem (BCH) code, which has multiple error correcting capability and a wide class of block codes, to overcome random errors. Realistic error patterns caused by multipath fading are analytically approximated and applied to the suggested system. Performance analysis for the new system showed that it improves the cell loss rate to a value of less than 10/sup -9/ in case of non-fading and it keeps the average cell loss close to the limit achieved by using a fixed BCH error correction system, while maintaining an average encoding delay considerably less than the delay of a system using a fixed BCH correction capability. M. Mazen Al-Khatib, Magdy A. Bayoumi |
WCNC | 2 |
| 1998 | A New Full Adder Cell for Low-Power ApplicationsabstractA new low power CMOS 1-bit full adder cell is presented. It is based on recent design of XOR and XNOR gates, and pass-transistors, it has 17 transistors. This cell has been compared to two widely used efficient adder cells; the transmission function full adder cell (16 transistors) and the low power adder cell (14 transistors). The new cell has no short circuit power and lower dynamic power (than the other adder cells), because of less number and magnitude of circuit capacitances. It consumes 10% to 15% less power than the other two cells. A comparative analysis (using Magic and Hspice) for 8-bit ripple carry and carry select adders shows that the adders based on the new cell can save up to 25% of power consumption. Ahmed M. Shams, Magdy A. Bayoumi |
Great Lakes Symposium on VLSI | 2 |
| 1998 | Three-dimensional defect sensitivity modeling for open circuits in ULSI structuresabstractIn the last three decades, device feature size has been reduced from ten's of microns to sub-0.5 /spl mu/m, shaping what is commonly called the ultra large scale integration (ULSI) era. At the same time, there has also been an appreciable decrease in the defect density over these years. While some of the defects have been completely eliminated, many still remain in modem fab lines. In view of the continuing trend of feature size reduction, the three-dimensional (3-D) nature of these defects is likely to have an impact on the functionality of integrated circuit (IC) structures. An open circuit is conventionally modeled as a break in a conductor due to an insulating defect. For 3-D defects and conductors, this physical interpretation of open circuits may not be the most appropriate one. In this paper, we present a parametrized modeling of open circuits for 3-D defects in ULSI processing. In the proposed approach, the maximum allowable increase in a conductor's resistance due to a partially or completely embedded insulating defect is suggested as the criterion for an open circuit. Analytical expressions for ULSI defect sensitivity are then obtained. M. K. Kidambi, Akhilesh Tyagi, Mohammed R. Madani, Magdy A. Bayoumi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 1997 | A low power based system partitioning and binding technique for multi-chip module architecturesabstractIn this paper, we present a low power targeted high-level synthesis framework for the synthesis of Multi-Chip Modules (MCM). This new framework is based on minimizing the switching activity on the functional units as well as the inter-chip buses. The main focus of the developed method is minimizing the power during partitioning and binding phases of high-level synthesis. A Stochastic Evolution based technique has been used for system partitioning. Experimental results were highly encouraging with power reduction of up to 60% on certain benchmark designs. Raghava V. Cherabuddi, Magdy A. Bayoumi, H. Krishnamurthy |
Great Lakes Symposium on VLSI | 2 |
| 1997 | A prototype chipset for a large scaleable ATM switching nodeabstractThis paper presents a chipset for a 16/spl times/16 switching node for the distributed banyan network. This chipset enables the use of a larger and much more efficient switching node than was previously available. Very high performance is required of the chips and thus a number of special circuits have been designed to achieve this performance. The chipset resulting from this design consumes low power. The chips have been designed in 1.0 micron CMOS using a mixture of static and dynamic logic. To achieve the speed needed for a larger node, a register file has been employed to store the packet headers on the control chip. It has an area of 3,150/spl times/3,750 micron, and uses 130,000 transistors. The SRAM blocks on the switch chip, which store a bit-slice of the packets, uses 228,600 transistors. M. B. Maaz, H. Krishnamurthy, Paul Shipley, Magdy A. Bayoumi |
Great Lakes Symposium on VLSI | 5 |
| 1996 | A High Speed VLSI Architecture for Scaleable ATM SwitchesabstractThis paper presents a prototype of a VLSI chip to be used as a building block for an efficiently scaleable ATM switch with a link speed of 622.2 Mb/s. The chip is a 4/spl times/4 shared multibuffer ATM switch based on the distributing banyan architecture. It is efficient in storage space like a shared memory switch and scaleable in size like a space division switch. Since the architecture is self-routing, the chip contains all necessary routing control. Special high speed and low power circuitry is used. The chip is implemented in 1.0 micron static CMOS and measures only 25 mm/sup 2/ in area. Paul Shipley, Sherif Sayed, Magdy A. Bayoumi |
Great Lakes Symposium on VLSI | 3 |
| 1996 | The Extended Cube Connected Cycles: An Efficient Interconnection for Massively Parallel SystemsabstractThe hypercube structure is a very widely used interconnection topology because of its appealing topological properties. For massively parallel systems with thousands of processors, the hypercube suffers from a high node fanout which makes such systems impractical and infeasible. In this paper, we introduce an interconnection network called The Extended Cube Connected Cycles (ECCC) which is suitable for massively parallel systems. In this topology the processor fanout is fixed to four. Other attractive properties of the ECCC include a diameter of logarithmic order and a small average interprocessor communication distance which imply fast data transfer. The paper presents two algorithms for data communication in the ECCC. The first algorithm is for node-to-node communication and the second is for node-to-all broadcasting. Both algorithms take O(log N) time units, where N is the total number of processors in the system. In addition, the paper shows that a wide class of problems, the divide and conquer class, is easily and efficiently solvable on the ECCC topology. The solution of a divide and conquer problem of size N requires O(log N) time units. Rafic Ayoubi, Qutaibah M. Malluhi, Magdy A. Bayoumi |
IEEE Trans. Computers | 3 |
| 1996 | Correction to "Efficient Mapping of ANNs on Hypercube Massively Parallel Machines"
Qutaibah M. Malluhi, Magdy A. Bayoumi, T. R. N. Rao |
IEEE Trans. Computers | 2 |
| 1995 | A scalable shared buffer ATM switch architectureabstractA scalable shared buffer switch architecture for asynchronous transfer mode (ATM) with O(/spl radic/N) complexity for memory bandwidth requirement and maximum crosspoint switch size, and O(N) scalability for buffer memory size is proposed. Access time to buffer memories has been reduced by virtue of parallel access. The switch architecture features multiple buffer memories between the input and output side crosspoint switches. The new switch architecture is better than the standard shared buffer approach as it eliminates the use of input and output time division multiplexing and makes it possible to meet buffer memory access time limitations for larger switches. At the same time, the proposed switch architecture is able to keep the crosspoint switches from growing as O(N/sup 2/) as is the case in the pure multibuffer architecture. The proposed architecture offers a good compromise between the simple shared buffer and shared multibuffer architectures Architectural and implementation details are discussed and a quantitative comparison between the buffer architectures given. Implementation of an 8/spl times/8 switch in 1.0 /spl mu/m CMOS technology is described. Aditya Agrawal, Anand Raju, Sachidanand Varadarajan, Magdy A. Bayoumi |
Great Lakes Symposium on VLSI | 4 |
| 1995 | A scalable analog architecture for neural networks with on-chip learning and refreshingabstractThis paper discusses various techniques for analog storage and handling and proposes a new class of architecture suitable for modular and scalable analog neural networks with on-chip learning and refreshing. The new architecture is based on analog functional blocks and analog pass switches which enhance the system versatility. Supporting algorithms are also developed. A novel characteristic is the full-analog on-chip learning methodology which substantially increases the learning speed. The speedup is evidenced by the utilization of local analog synaptic updating scheme which utterly eliminates time-sharing components. Moreover, this localization scheme conceives unbounded scalability in this neural architecture. Bassem A. Alhalabi, Magdy A. Bayoumi |
Great Lakes Symposium on VLSI | 2 |
| 1995 | A new ATM congestion control scheme for shared buffer switch architecturesabstractIn this paper, we propose a new pre-emptive congestion control (PECC) scheme for deployment in shared buffer-based ATM networks. The proposed PECC uses a rate-based feedback scheme for dynamic bandwidth allocation to best-effort traffic coexisting with guaranteed service traffic. Policing of the guaranteed service continuous bit-rate and variable bit-rate traffic is done using the leaky bucket mechanism. The PECC is a predictive scheme rather than a reactive one. The internal state of the ATM switch is used to predict the onset of congestion. The backward explicit congestion notification feedback scheme developed in earlier works for ATM local area networks has been effectively used in our scheme for delays equivalent to those of MANs and WANs. A simulation model for an illustrative ATM network exhibiting congestion has been constructed. The effect of varying key parameters including mean inter-arrival time, buffer threshold throttle factor, cell reject throttle factor, buffer threshold, filter time and rate reduction delta has been studied. Transient behaviour of an example ATM network employing shared buffer switches under the event of congestion, and its control by backward explicit congestion notification has also been simulated. Aditya Agrawal, Magdy A. Bayoumi, Amr Elchouemi |
ICCCN | 2 |
| 1995 | A New Thinning Algorithm for Arabic Character Using Self-Organizing Neural NetworkabstractIn this paper, we propose a new thinning algorithm based on clustering the data image. We employ the ART2 network which is a self-organizing neural network for the clustering of Arabic characters. The skeleton is generated by plotting the cluster centers and connecting adjacent clusters by straight lines. This algorithm produces skeletons which are superior to the outputs of the conventional algorithms. It achieves a higher data reduction efficiency and much simpler skeletons with less noise spurs. Moreover, to make the algorithm appropriate for real-time applications, an optimization technique is developed to reduce the time complexity of the algorithm. Nevertheless, the algorithm is not limited to Arabic characters, it can, also, be used to skeletonize characters of other languages. Majid M. Altowairjri, Magdy A. Bayoumi |
ISCAS | 2 |
| 1995 | Efficient Mapping of ANNs on Hypercube Massively Parallel MachinesabstractThis paper presents a technique for mapping artificial neural networks (ANNs) on hypercube massively parallel machines. The paper starts by synthesizing a parallel structure, the mesh-of-appendixed-trees (MAT), for fast ANN implementation. Then, it presents a recursive procedure to embed the MAT structure into the hypercube topology. This procedure is used as the basis for an efficient mapping of ANN computations on hypercube systems. Both the multilayer feedforward with backpropagation (FFBP) and the Hopfield ANN models are considered. Algorithms to implement the recall and the training phases of the FFBP model as well as the recall phase of the Hopfield model are provided. The major advantage of our technique is high performance. Unlike the other techniques presented in the literature which require O(n) time, where N is the size of the largest layer, our implementation requires only O(log N) time. Moreover, it allows pipelining of more than one input pattern and thus further improves the performance.> Qutaibah M. Malluhi, Magdy A. Bayoumi, T. R. N. Rao |
IEEE Trans. Computers | 2 |
| 1994 | Automated system partitioning for synthesis of multi-chip modulesabstractWe present a system-level partitioning technique for the synthesis of multi-chip modules. It is based on the stochastic evolution heuristic, which is an effective heuristic for solving several combinatorial optimization problems. We perform the partitioning at the behavioral level. The advantage of partitioning at the behavioral level is that both area and time constraints can be taken care of at the system level and also that scheduling/allocation can be applied concurrently to system-level partitioning. We formulate the partitioning problem as an extension to the network-bisectioning problem for which the stochastic evolution heuristic has been shown to provide better results than the simulated annealing technique. Preliminary scheduling/allocation and pin sharing are also performed simultaneously to estimate the area and pincount of each of the partitions. Efficient partitions are obtained for some of the digital signal processing applications in reasonable CPU time.> Raghava V. Cherabuddi, Magdy A. Bayoumi |
Great Lakes Symposium on VLSI | 2 |
| 1994 | Arabic Text Recognition Using Neural NetworksabstractRecognizing multi-font Arabic texts is a difficult task in the area of optical character recognition (OCR) because Arabic is a cursive type language. This paper proposes a hybrid Arabic character recognition system based on Moment Invariants employing an Artificial Neural Network classifier. The feature extraction stage uses a set of moment invariants descriptors which are invariants under shift, scaling, and rotation. The actual classification is done using a multilayer perceptron network with back-propagation learning. As a preprocessing step, a new approach to segmentation of Arabic words is proposed in this paper. The system has been tested and has shown a very high accuracy.> Majid M. Altuwaijri, Magdy A. Bayoumi |
ISCAS | 2 |
| 1994 | Storage Allocation Strategies for Data Path Synthesis of ACICsabstractIn this paper, we concentrate on the storage allocation problem in datapath synthesis. Datapath allocation techniques can be classified into two main categories: iterative/constructive; and global. Storage allocation deals with determining the number of registers that are needed. Registers are used in datapaths to store values that are generated in one control step and used in an other. When lifetimes of such values do not overlap, they can then be mapped on to the same registers. Just as in the case of functional unit allocation, storage allocation also affects the steering and interconnection logic. Therefore storage allocation must take into consideration minimization of such logic. Storage allocation does not necessarily deal with register allocation. In lieu of registers, register files or memory units can be used to decrease overall area.> N. A. Ramakrishna, Magdy A. Bayoumi |
ISCAS | 2 |
| 1994 | The Hierarchical Hypercube: A New Interconnection Topology for Massively Parallel SystemsabstractInterconnection networks play a crucial role in the performance of parallel systems. This paper introduces a new interconnection topology that is called the hierarchical hypercube (HHC). This topology is suitable for massively parallel systems with thousands of processors. An appealing property of this network is the low number of connections per processor, which enhances the VLSI design and fabrication of the system. Other alluring features include symmetry and logarithmic diameter, which imply easy and fast algorithms for communication. Moreover, the HHC is scalable; that is it can embed HHC's of lower dimensions. The paper presents two algorithms for data communication in the HHC. The first algorithm is for one-to-one transfer, and the second is for one-to-all broadcasting. Both algorithms take O(log/sub 2/ k), where k is the total number of processors in the system. A wide class of problems, the divide & conquer class (D&Q), is shown to be easily and efficiently solvable on the HHC topology. Parallel algorithms are provided to describe how a D&Q problem can be solved efficiently on an HHC structure. The solution of a D&Q problem instance having up to k inputs requires a time complexity of O(log/sub 2/ k).> Qutaibah M. Malluhi, Magdy A. Bayoumi |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1993 | An application-specific array architecture for feedforward with backpropagation ANNsabstractAn application-specific array architecture for Artificial Neural Networks (ANNs) computation is proposed. This array is configured as a mesh-of-appendixed-trees (MAT). Algorithms to implement both the recall and the training phases of the multilayer feedforward with backpropagation ANN model are developed on MAT. The proposed MAT architecture requires only O(log N) time, while other reported techniques offer O(N) time, where N is the size of the largest layer. Beside the high speed performance, pipelining of more than one input pattern can be achieved which further improves the performance.> Qutaibah M. Malluhi, Magdy A. Bayoumi, T. R. N. Rao |
ASAP | 2 |
| 1993 | Parallel Implementation of a Cut and Paste Maze Routing Algorithm
Magdy A. Bayoumi, Akhilesh Tyagi, Nam Ling, R. Kalyan |
ISCAS | 2 |
| 1992 | SDS: a framework for the design of DSP ASICsabstractA brief overview is given of the Sphinx Design System (SDS), a high-level synthesis system consisting of an integrated and interacting set of tools for the synthesis of digital circuits. The system is specifically tuned to the synthesis of digital signal processor (DSP) application specific integrated circuits (ASICs) from behavioral specifications written in Verilog. SDS consists of tools for behavioral, structural and logical synthesis, technology mapping and for simulation. C and SKILL have been used as the implementation and extension languages and the system will be integrated into the Cadence Edge Framework. SPS, the scheduling and allocation subsystem of SDS, is discussed in some detail.> Magdy A. Bayoumi, N. A. Ramakrishna, V. Israni, R. K. Jayam |
ICASSP | 1 |
| 1991 | A high speed pipelined FFT processorabstractThe authors describe a novel modular design and a VLSI implementation of a bit-serial pipelined fast Fourier transform (FFT) coprocessor. The proposed architecture is based on a distributed hardwired control mechanism. The control of various subunits in the processor is done by local controllers and the synchronization of operations is provided by a global controller. This FFT processor is a custom-built chip which has a built-in self-test (BIST) capability. BIST is provided using a coprocessor to a microprocessor, and the data transfer is controlled by asynchronous signals. A prototype of the proposed processor was implemented in 3- mu m SCMOS technology; it can operate at a maximum frequency of 50 MHz. The chip is a 32-pin square package, and it has a total area of 2.2 cm/sup 2/.> Srinivasa R. Malladi, Seshagiri R. Myneni, Pardha V. Pothana, Magdy A. Bayoumi |
ICASSP | 4 |
| 1991 | VLSI implementation of a systolic database machine for relational algebra and hashing
Khaled M. Elleithy, Magdy A. Bayoumi, Lois M. L. Delcambre |
Integr. | 2 |
| 1990 | A formal design methodology for parallel architecturesabstractThe authors introduce a formal approach for synthesis of array architectures. The methodology provides two main features: completeness and correctness. Completeness means the ability to use the approach for any general algorithm. Correctness is achieved by using a set of transformations that are proved to be correct. Four different forms are used to express the input algorithm: simultaneous recursion, recursion with respect to different variables, fixed nesting, and variable nesting. Four different architectures for the same algorithm are obtained. As an example, a matrix-matrix multiplication algorithm is used to obtain four different optimal architectures. The different architectures of this example are compared in terms of area, time, broadcasting, and required hardware.> Khaled M. Elleithy, Magdy A. Bayoumi |
ASAP | 2 |
| 1990 | A formal high level synthesis approach for DSP architecturesabstractAn approach is presented for high-level synthesis of digital signal processing (DSP) algorithms. Two features are provided by the approach: completeness and correctness. A given algorithm is represented in a newly developed language termed the algorithm specification language (ASL). ASL had the ability to describe any general algorithm. An automatic procedure is used to transform an ASL representation into a specific realization specification using a correctness preserving set of transformations. The realization format is based on representing the digital architectures by another language called the realization specification language (RSL). Logic programming is used as a user interface for the synthesis procedure.> Khaled M. Elleithy, Magdy A. Bayoumi |
ICASSP | 2 |
| 1990 | Formal Synthesis of a Parallel Architectures from Recursive Equations
Khaled M. Elleithy, Magdy A. Bayoumi |
ICPP (1) | 2 |
| 1990 | Systolic array implementation of image segmentation by a directed split and merge procedureabstractA systolic array implementation of image segmentation by a split and merge procedure is proposed. The implementation enhances the input/output and memory bandwidth requirements, leading to a decrease in the computation time for the segmentation of an image. Image segmentation through the proposed approach can be achieved in linear time.> Aakash Tyagi, Magdy A. Bayoumi |
ICPR (2) | 2 |
| 1990 | Systolic temporal arithmetic: a new formalism for specification and verification of systolic arraysabstractA systolic temporal arithmetic (STA) formalism suitable for describing arithmetic operations in dynamic environments is introduced. It can be used for formal specification and verification of systolic arrays at the array architecture level. Besides providing value and operation abstraction from the lower level, it also exploits several features of systolic arrays, such as synchrony, regularity, repeatability, modularity, pipelinability, parallel processing ability, and spatial and temporal locality. STA provides constructs and verification techniques for simple, efficient, and effective systolic-array specification and verification. Verification techniques such as mathematical induction are suggested to exploit these systolic array features so as to speed up the process. STA overcomes several limitations of other existing specification and verification techniques and can be used with lower-level formalism for multilevel reasoning of systolic arrays. A brief description of this formalism and examples of how the formalism can be applied for formal specification and verification of two systolic arrays are discussed. Other applications of STA, such as simulations, fault diagnoses, and test generation for systolic arrays, are also suggested.> Nam Ling, Magdy A. Bayoumi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1989 | θ(logN) architectures for RNS arithmetic decodingabstractDecoding in residue-number-system (RNS)-based architectures can be a bottleneck. A high-speed, flexible modulo decoder is an essential computational element to maintain the advantages of RNS. A fast and flexible modulo decoder, based on the Chinese remainder theorem (CRT), is presented. It decodes a set of residues into its equivalent representation in either unsigned magnitude or two's-complement binary number system. Two different architectures are analyzed: the first one uses carry-save adders, and the other uses modified structure carry-save adders. Both architectures are modular and are based on simple cells, which leads to efficient VLSI implementation. The decoder has a time complexity of theta (log N).> Khaled M. Elleithy, Magdy A. Bayoumi |
IEEE Symposium on Computer Arithmetic | 2 |
| 1989 | The design and implementation of multidimensional systolic arrays for DSP applicationsabstractThe authors present a technique for transforming DSP (digital signal processing) algorithms to a form suitable for multidimensional systolic array implementation. The aim of the transformation is to speed up computation without much increase in area requirement. The price to be paid is the small amount of additional circuitry (usually in the form of adders and interconnection wires) required for interrow or interplane communications. The application of the technique to some DSP algorithms is presented. The systolic networks produced are implemented using NORA CMOS logic structure and laid out using 3- mu m CMOS p-well technology. Areas and times for the resulting architectures are evaluated and discussed.> Nam Ling, Magdy A. Bayoumi |
ICASSP | 2 |
| 1989 | Systematic Algorithm Mapping for Multidimensional Systolic Arrays
Nam Ling, Magdy A. Bayoumi |
J. Parallel Distributed Comput. | 2 |
| 1988 | Reliable modulo systolic arrays for DSP algorithmsabstractThe algebraic and structural properties of a residue number system (RNS) to optimize the design of DSP (digital signal processing) systolic arrays is discussed. RNS establishes parallelism on the algorithmic level by decomposing the original computational field into a set of finite fields, in which arithmetic operations are performed independently for each field. The impact of the moduli size on the structural choices is analyzed. Three different structures, namely, word-parallel, bit-parallel, and bit-serial, are presented. Redundant RNS has fault tolerance capabilities. An efficient fault-detection technique is developed as an alternative to the standard procedure based on a mixed radix algorithm. A finite-impulse-response (FIR) filter algorithm is used as a case study.> Magdy A. Bayoumi |
ICASSP | 1 |
| 1988 | Multi-dimensional systolic networks for DSP algorithmsabstractThe authors present a novel technique for transforming a class of digital signal processing (DSP) algorithms, and some arithmetic algorithms, to specific forms that can be directly mapped onto higher-dimensional systolic networks. The latency, as well as the order of complexity of computation time, can be significantly improved through implementing these algorithms on higher-dimensional systolic networks. At the same time, the order of area complexity is kept constant. The technique can be applied to problems such as 1-D convolution, k-point discrete Fourier transform (DFT), finite-impulse response (FIR) filters, and matrix-vector multiplication. The k-point DFT algorithm example is illustrated along with other examples. Implementation issues of high-dimensional systolic networks on 2-D or 3-D VLSI are discussed.> Nam Ling, Magdy A. Bayoumi |
ICASSP | 2 |
| 1988 | A fault-tolerant bit-serial array structure for digital filtersabstractThe authors propose a bit-serial structure using systolic arrays in the implementation of a finite-impulse-response (FIR) algorithm for large sizes of data and higher-order filters. The bit cells are arranged in a 2-D array which enhances the extensibility and provides efficiency for high-precision data. Barrel shifters are used on the cell level which increases the throughput of the proposed pipelined structure. Moreover, the proposed architecture has distributed error-control features. The necessity of such distributed fault-tolerance in DSP (digital signal processing) architectures is due to its susceptibility to permanent and intermittent errors caused by the high complexity of these circuit structures.> Rajat Roy, Magdy A. Bayoumi |
ICASSP | 2 |
| 1988 | Algorithms for High Speed Multi-Dimensional Arithmetic and DSP Systolic Arrays
Nam Ling, Magdy A. Bayoumi |
ICPP (1) | 2 |
| 1987 | A quadratic residue processor for complex DSP applicationsabstractThe performance requirements for high speed real-time Digital Signal Processing (DSP) applications, coupled with the hardware complexity for manipulating data over complex fields can often exceed the capacity of the traditional signal processors. In this paper, a signal processor for complex DSP applications has been developed. It is based on using the Quadratic Residue Number System (QRNS) which establishes parallelism on the functional level and optimizes the required hardware. By employing QRNS, the interaction between the real and imaginary channels in complex arithmetic is eliminated, and two real multiplications are only required to perform complex multiplication. The processor design is optimised for efficient computation of the FFT and signal processing operations based on FFT such as FIR filtering, IFFT, convolution, correlation, and multiplication. This computational versatility is achieved through macroprogrammability. FIFO's (First-in First-out) are used for storing the input, intermediate, and output data (and coefficients). Customized look-up tables are employed for implementing the residue operations. The developed processor can be employed either as a stand-alone processor or as a peripheral processor (Co-processor). Magdy A. Bayoumi |
ICASSP | 1 |
| 1986 | A VLSI array for computing the DFT based on RNSabstractThe Discrete Fourier Transform (DFT) has been adopted in a wide spectrum of Digital Signal Processing (DSP) applications due to the advances in VLSI technology, One dimensional systolic arrays are employed to implement the DFT algorithms where N DFT points can be computed in O(N) time using O(N) area. Residue Number System (RNS) is used to achieve parallelism on the mathematical level, as the arithmetic operations are performed independently for each modulus. Modularity has been realized on both functional and layout levels. Two types of arrays are described. The first array offers higher speed performance, while the second requires less area and is more general. The proposed structures are based on bit parallel processing and lend themselves to pipelining. Magdy A. Bayoumi, Graham A. Jullien, William C. Miller |
ICASSP | 1 |
| 1986 | Lower bounds for VLSI implementation of residue number system architectures
Magdy A. Bayoumi |
Integr. | 1 |
| 1985 | A VLSI implementation of an FFT/NTT computational unitabstractThe coupling of Residue Number System (RNS) with the recent advances in VLSI technology leads to an efficient implementation of many digital signal processing algorithms. This paper discusses modularity in implementing RNS systems, as modularity is considered an important criterion for VLSI design. An NTT/FFT computational unit is implemented using two multi-look-up table modules as building block units. The layout can be optimized using a look-up table layout procedure which supports the custom design approach. The modularity has been achieved on both functional and layout levels where the interconnection area is minimum. Magdy A. Bayoumi, Graham A. Jullien, William C. Miller |
ICASSP | 1 |
| 1985 | An efficient VLSI adder for DSP architectures based on RNSabstractThe implementation of Residue Number System (RNS) architectures using the VLSI technology is discussed. An example of implementing an RNS adder is presented in this paper. Two approaches; the look-up table and the binary adder, have been analyzed in the scope of VLSI criteria where the performance measures are area and time. Two models have been developed, they are flexible, support any modulus, and they provide custom design capabilities. Within the context of this paper, it has been found that the look-up table approach is superior in both area and time up to 5 bits, while the binary adder approach offers better performance for larger moduli. Magdy A. Bayoumi, Graham A. Jullien, William C. Miller |
ICASSP | 1 |
| 1984 | A VLSI model for residue number system architectures
Magdy A. Bayoumi, Graham A. Jullien, William C. Miller |
Integr. | 1 |
| 1983 | Models for VLSI implementation of residue number system arithmetic modulesabstractThis paper discusses the implementation of RNS arithmetic modules using VLSI technology. The modules are based on the interconnection of read-only memory look-up tables. The paper first outlines a memory model for a single look-up table which allows the selection of the most efficient layout for memories which do not have power of 2 dimensions. The paper then discusses various examples of interconnected memory modules with associated optimizing layout algorithms. Finally, an example is given of the application of one of the modules to a large prime modulus multiplier. Magdy A. Bayoumi, Graham A. Jullien, William C. Miller |
IEEE Symposium on Computer Arithmetic | 1 |
| 1983 | An area-time efficient NMOS adder
Magdy A. Bayoumi, Graham A. Jullien, William C. Miller |
Integr. | 1 |