Ahmed M. Eltawil

dblp:73/816 · DBLP profile ↗
← Back
168ranked-venue papers
2as first author
86since 2021 · last 2026
0000-0003-1849-083XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 81 · 2 first-author · 45 since 2021Systems, architecture and hardware · 61 · 22 since 2021Software engineering, systems software and programming languages · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 DRL-based AoI Optimization for Energy Harvesting Human Body Communication ECG Networks
abstract
In energy harvesting (EH) Internet of Bodies (IoB) networks, maintaining information freshness is critical for time-sensitive applications such as electrocardiogram (ECG) monitoring. Timely transmissions are essential for reliable ECG analysis, which requires complete and synchronous data from all sensors, as each ECG lead is derived from potential differences between paired sensors. To this end, we leverage model-free deep reinforcement learning (DRL) to learn adaptive scheduling policies that minimize the age of information (AoI) in dynamic and intermittent energy arrival conditions. Specifically, in a human body communication (HBC)-enabled EH-ECG network, wearable ECG sensors transmit data to a central wearable hub acting as the scheduler. The hub implements the DRL scheduling algorithm that selects sensor transmissions based on their battery and AoI levels. Furthermore, we propose an ECG synthesis framework for a 3-lead ECG that mitigates the impact of EH scarcity by reconstructing missing leads from data transmitted by active sensors, thereby ensuring complete signal availability at the hub. Simulation results show that the proposed solution reduces AoI by up to 57% compared to the greedy myopic policy and up to 96% relative to a baseline learning policy.
Abeer Alamoudi, Abdulkadir Celik, Asmaa Abdallah, Ahmed M. Eltawil
ICC4
2026 A cGAN Empowered Physical Layer Authentication Against Malicious RIS Attacks
abstract
Reconfigurable intelligent surfaces (RIS) have emerged as a transformative technology for next-generation wireless networks, offering unprecedented control over radio propagation environments. However, their passive nature and ease of deployment introduces security vulnerabilities that remain largely unexplored. This paper investigates a spoofing attack where a malicious RIS strategically manipulates its reflection coefficients to impersonate a legitimate RIS, thereby deceiving the base station (BS) and gaining unauthorized network access. To counter this threat, we propose a novel authentication framework that formulates the detection problem as a data-driven binary classification task, leveraging conditional generative adversarial networks (cGAN). The framework employs a U-Net-based generator to synthesize realistic attack scenarios during training, while the discriminator serves as a lightweight authenticator enabling robust authentication without requiring apriori knowledge of attacker strategies. Through extensive simulations across diverse attack scenarios, including co-located and correlated configurations, we demonstrate that the trained discriminator achieves 96.4% detection accuracy against malicious RIS attackers positioned near the BS (co-located) and maintains 86.2% accuracy under correlated attack conditions.
Amira Bendaimi, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil, Hüseyin Arslan
ICC4
2026 Target-in-the-Loop Beam Tracking: Synergy of GPS and LSTM for Proactive mmWave V2V Networks
abstract
Although massive multi-input multi-output (mMIMO) systems offer higher directivity gains to mitigate propagation losses at millimeter wave (mmWave) frequencies, they present challenges in channel state information (CSI) acquisition and beam alignment, particularly in highly mobile vehicle-to-vehicle (V2V) scenarios with short channel coherence times. To address these issues, we propose a target-in-the-loop beam tracking approach that leverages GPS data and Long Short-Term Memory (LSTM) networks to select beams from predefined beamforming codebooks. By transforming GPS data into relative coordinates and extracting features such as relative velocity and orientation, our model predicts future beam states up to 500 ms in advance, enabling proactive beam selection and blockage avoidance. Using the DeepSense V2V dataset, our method achieves up to 7.2 dB power loss reduction and a 32% improvement in top-5 accuracy compared to a linear interpolation baseline. This approach highlights the potential of integrating GPS data and machine learning to enhance beam tracking in dynamic V2V mmWave networks.
Mattia Fabiani, Diego A. Silva, Asmaa Abdallah, Abdulkadir Celik, Davide Dardari, Ahmed M. Eltawil
ICC6
2026 Conditional Generative AoA/AoD Estimation: A Pilot-Free and System-Agnostic Approach
abstract
Accurate angle of arrival (AoA) and angle of departure (AoD) estimation underpins spatial precision, efficient beamforming, and the overall Quality of Service (QoS) and Quality of Experience (QoE) of integrated sensing and communication (ISAC). This paper introduces a generative AI (GenAI) framework that synthesizes the complete set of multipath 2D AoA/AoD parameters directly from transmitter (TX) and receiver (RX) locations, overcoming the limitations of traditional model-based and current data-driven learning methods. The proposed classifiers-guided conditional generative adversarial network (CG-CGAN) offers a system-agnostic approach that removes the need for pilot signals and operates effectively under challenging coherent multipath conditions. Its three-stage architecture integrates classification and conditional generative modeling to jointly infer line-of-sight (LoS) status, number of propagation paths, and angular parameters. Simulations on the DeepMIMO dataset demonstrate over 99.5% classification accuracy and 92% angle generation accuracy, significantly outperforming classical techniques while reducing computational complexity and enhancing the QoS/QoE of ISAC-enabled systems.
Bumin Kagan Yildirim, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil
ICC4
2026 A Synthesizable Mixed-Precision DCIM Macro with Parallel Write and Compute
Jinane Bazzi, Mohamed E. Fouda, Ahmed M. Eltawil
ISCAS3
2026 Sparse neural sampling mixers
Ahmed Elsheikh, Mohamed E. Fouda, Ahmed M. Eltawil
Neurocomputing3
2026 Analyzing URA Geometry for Enhanced Near-Field Beamfocusing and Spatial Degrees of Freedom
abstract
With the deployment of large antenna arrays at high-frequency bands, future wireless communication systems are likely to operate in the radiative near-field. Unlike far-field beam steering, near-field beams can be focused on a spatial region with a finite depth, enabling spatial multiplexing in the range dimension. Moreover, in the line-of-sight MIMO near-field, multiple spatial degrees of freedom (DoF) are accessible, akin to a scattering-rich environment. In this paper, we derive the beamdepth for a generalized uniform rectangular array (URA) and investigate how the array geometry influences near-field beamdepth and its limits. We define the effective beamfocusing Rayleigh distance (EBRD), to present a near-field boundary with respect to beamfocusing and spatial multiplexing gains for the generalized URA. Our results demonstrate that under a fixed element count constraint, the array geometry has a strong impact on beamdepth, whereas this effect diminishes under a fixed aperture length constraint. Moreover, compared to uniform square arrays, elongated configurations such as uniform linear arrays (ULAs) yield narrower beamdepth and extend the effective near-field region defined by the EBRD. Building on these insights, we design a polar codebook for compressed-sensing-based channel estimation that leverages our findings. Simulation results show that the proposed polar codebook achieves a 2 dB NMSE improvement over state-of-the-art methods. Additionally, we present an analytical expression to quantify the effective spatial DoF in the near-field, revealing that they are also constrained by the EBRD. Notably, the maximum spatial DoF is achieved with a ULA configuration, outperforming a square URA in this regard.
Ahmed Hussain 0001, Asmaa Abdallah, Abdulkadir Celik, Emil Björnson, Ahmed M. Eltawil
IEEE Trans. Commun.5
2026 ENWAR 2.0: An Agentic Multimodal Wireless LLM Framework With Reasoning, Situation-Aware Explainability and Beam Tracking
abstract
The evolution of next-generation wireless networks demands intelligent, adaptive, and explainable decision-making for robust communication in dynamic environments. This paper presentsEnwar 2.0, the first agentic large language model (LLM) framework integrating adaptive retrieval-augmented generation (RAG) and chain-of-thought (CoT) reasoning into situation-aware and explainable wireless network management.Enwar 2.0introduces two specialized agents: a transformer-fusion (TransFusion)-based beam prediction agent and an environment perception agent, both of which fuse multi-modal sensory inputs—including camera, LiDAR, radar, and GPS—from the DeepSense6G dataset. The beam prediction agent enables infrastructure-to-vehicle (I2V) target-in-the-loop beam tracking and real-time adaptation based on dynamic environmental conditions. In contrast, the environment perception agent provides situation-aware reasoning and justifications for beam decisions. Unlike its predecessor,Enwar 1.0, which relied on static knowledge bases (KBs) and text-only LLMs,Enwar 2.0is designed for CoT reasoning, leverages LLaMa3.2-3B/LLaMa3.1-8B/LLaMa3.3-70B for text-generation, the multi-modal capabilities of LLaMa 3.2, and employs LlamaIndex for fine-grained, dynamic context retrieval, eliminating retrieval ambiguities and enhancing response relevance. Numerical results show that the beam prediction agent achieves up to 90.0% Top-3 accuracy at$t+3$, effectively predicting optimal beam selections three time steps ahead. Overall,Enwar 2.0achieves state-of-the-art performance, with up to 89.7%/83.5% interpretation/perception correctness, 81.6%/80.9% faithfulness, and 89.9%/88.2% relevancy. In comparison, the baseline pretrained LLaMa3 models without adaptive RAG achieves up to 80.3%/77.3% correctness, and the baseline without RAG performs significantly worse at 67.1%/64.8%. Additionally,Enwar 2.0reduces processing time by over 100% relative to the baseline, while its adaptive RAG improves performance by up to 13.7% compared to static RAG.
Ahmad M. Nazar, Abdulkadir Celik, Mohamed Y. Selim, Asmaa Abdallah, Daji Qiao, Ahmed M. Eltawil
IEEE Trans. Mob. Comput.6
2026 Boosting Spectral Efficiency via Spatial Path Index Modulation in RIS-Aided mMIMO
abstract
Next generation wireless networks focus on improving spectral efficiency (SE) while reducing power consumption and hardware cost. Reconfigurable intelligent surfaces (RISs) offer a viable solution to meet these requirements. In order to enhance the SE, index modulation (IM) has been regarded as one of the enabling technologies via the transmission of additional information bits over the transmission media such as subcarriers, antennas and spatial paths. In this work, we explore the usage of spatial paths and introduce spatial path IM (SPIM) for RIS-aided massive multiple-input multiple-output (mMIMO) systems. Thus, the proposed framework improves the network efficiency and the coverage with the use of RIS while SPIM provides SE improvement. In order to perform SPIM, we exploit the spatial diversity of the millimeter wave channel and assign the index bits to the spatial patterns of the channel between the base station and the users through RIS. We introduce a low complexity approach for the design of hybrid beamformers, which are constructed by the steering vectors corresponding to the selected spatial path indices for SPIM-mMIMO. Furthermore, we conduct a theoretical analysis on the SE of the proposed SPIM approach, and derive the SE relationship between the SPIM-based hybrid beamforming and fully digital (FD) beamforming. Via numerical simulations, we validate our theoretical results and show that the proposed SPIM approach presents an improved SE performance, even higher than that of the use of FD beamformers while using a few RF chains.
Ahmet M. Elbir, Abdulkadir Celik, Asmaa Abdallah, Ahmed M. Eltawil
IEEE Trans. Wirel. Commun.4
2026 Digital Twin-Assisted Explainable AI for Robust Beam Prediction in mmWave MIMO Systems
abstract
In line with the AI-native 6G vision, explainability and robustness are crucial for building trust and ensuring reliable performance in millimeter-wave (mmWave) systems. Efficient beam alignment is essential for initial access, but deep learning (DL) solutions face challenges, including high data collection overhead, hardware constraints, lack of explainability, and susceptibility to adversarial attacks. This paper proposes a robust and explainable DL-based beam alignment engine (BAE) for mmWave multiple-input multiple-output (MIMO) systems. The BAE uses received signal strength indicator (RSSI) measurements from wide beams to predict the best narrow beam, reducing the overhead of exhaustive beam sweeping. To overcome the challenge of real-world data collection, this work leverages a site-specific digital twin (DT) to generate synthetic channel data closely resembling real-world environments. A model refinement via transfer learning is proposed to fine-tune the pre-trained model residing in the DT with minimal real-world data, effectively bridging mismatches between the digital replica and real-world environments. To reduce beam training overhead and enhance transparency, the framework uses deep Shapley additive explanations (SHAP) to rank input features by importance, prioritizing key spatial directions and minimizing beam sweeping. It also incorporates the Deep k-nearest neighbors (DkNN) algorithm, providing a credibility metric for detecting out-of-distribution inputs and ensuring robust, transparent decision-making. Experimental results show that the proposed framework reduces real-world data needs by 70%, beam training overhead by 62%, and improves outlier detection robustness by up to 8.5×, achieving near-optimal spectral efficiency and transparent decision making compared to traditional softmax based DL models.
Nasir Khan, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil, Sinem Coleri Ergen
IEEE Trans. Wirel. Commun.4
2026 RIS-Aided Protected Zone Formation for Physical Layer Security of In-Band Full Duplex Systems
abstract
The rapid evolution of mobile technologies presents a formidable security challenge, as traditional cryptographic methods struggle to keep pace. Integrating physical layer security (PLS) solutions with cutting-edge technologies such as in-band full-duplex (IBFD) and reconfigurable intelligent surfaces (RISs) holds promise for addressing these challenges effectively. This study introduces a novel RIS-driven protected zone (PZ) formation approach that employs artificial noise (AN) to safeguard legitimate users without requiringa prioriknowledge of eavesdropper locations, channels, or numbers. The proposed methodology partitions the RIS into two distinct segments: while the former segment enhances the achievable data rate for legitimate signal, the latter segment concurrently amplifies AN to jam illegitimate users within the PZ.We present formulations and solutions for maximizing secrecy capacity (SC) and minimizing power consumption through optimized transmit power allocation factors, RIS segmentation, and beams’ directions, all subject to stringent quality-of-service (QoS) constraints. Closed-form expressions are derived to facilitate efficient implementation and performance optimization. Simulation results validate closed-form solutions and demonstrate that the proposed scheme can significantly enhance SC compared to benchmarks where RIS and AN are used separately, with the proposed scheme achieving approximately 81% greater capacity than the “RIS-Only” approach and a substantial advantage over the “AN-Only” approach, which results in no secrecy. Additionally, this work includes an analysis of energy efficiency, emphasizing the critical importance of optimizing power consumption in practical applications. This dual focus on improving security while effectively managing energy resources underscores the scheme’s practical relevance and efficiency.
Hanadi Salman, Abdulkadir Celik, Sultangali Arzykulov, Ahmed M. Eltawil, Hüseyin Arslan
IEEE Trans. Wirel. Commun.4
2025 SoftmAP: Software-Hardware Co-Design for Integer-Only Softmax on Associative Processors
abstract
Recent research efforts focus on reducing the computational and memory overheads of Large Language Models (LLMs) to make them feasible on resource-constrained devices. Despite advancements in compression techniques, nonlinear operators like Softmax and Layernorm remain bottlenecks due to their sensitivity to quantization. We propose SoftmAP, a software-hardware co-design methodology that implements an integer-only low-precision Softmax using In-Memory Compute (IMC) hardware. Our method achieves up to three orders of magnitude improvement in the energy-delay product compared to A100 and RTX3090 GPUs, making LLMs more deployable without compromising performance.
Mariam Rakka, Jinhao Li 0006, Guohao Dai 0001, Ahmed M. Eltawil, Mohamed E. Fouda, Fadi J. Kurdahi
DATE4
2025 Deep Learning Framework for RSSI-Based Indoor Localization in RIS-Aided mmWave Systems
abstract
We present a novel deep learning (DL) framework for indoor localization in millimeter wave (mmWave) environments using received signal strength indicator (RSSI)-based measurements from a reconfigurable intelligent surface (RIS)-aided system. We address non-line-of-sight (NLoS) conditions and orientation variability through a dual-stream architecture that combines an orientation-gated convolutional neural network (OGCNN) with statistical feature extraction. Our approach processes RSSI matrices representing signal strengths across different user equipment (UE) and RIS beam patterns, and incorporates temporal modeling through bidirectional long short-term memory (BiLSTM) networks to capture user movement dynamics. The framework is validated using real-world experimental data collected in an indoor environment with a controlled RIS setup at 28 GHz, demonstrating superior performance with mean and median localization errors of 0.25 m and 0.19 m, respectively, outperforming the median error of classical baseline approaches by 74.7% and conventional DL baselines by 89.6%. The proposed solution maintains decimeter-level accuracy across various orientations and distances using only RSSI-based measurements, making it suitable for high-precision indoor positioning applications in next-generation wireless networks.
Varun S. Advani, Ahmed Nasser, Mohamed Y. Selim, Ahmed M. Eltawil
GLOBECOM4
2025 Near-Field Beam Prediction Using Far-Field Codebooks in Ultra-Massive MIMO Systems
abstract
Ultra-massive multiple-input multiple-output (UM-MIMO) technology is a key enabler for 6G networks, offering exceptional high data rates in millimeter-wave (mmWave) and Terahertz (THz) frequency bands. The deployment of large antenna arrays at high frequencies transitions wireless communication into the radiative near-field, where precise beam alignment becomes essential for accurate channel estimation. Unlike far-field systems, which rely on angular domain only, near-field necessitates beam search across both angle and distance dimensions, leading to substantially higher training overhead. To address this challenge, we propose a discrete Fourier transform (DFT) based beam alignment to mitigate the training overhead. We highlight that the reduced path loss at shorter distances can compensate for the beamforming losses typically associated with using far-field codebooks in near-field scenarios. Additionally, far-field beamforming in the near-field exhibits angular spread, with its width determined by the user's range and angle. Leveraging this relationship, we develop a correlation interferometry (CI) algorithm, termed CI-DFT, to efficiently estimate user angle and range parameters. Simulation results demonstrate that the proposed scheme achieves performance close to exhaustive search in terms of achievable rate while significantly reducing the training overhead by 87.5%.
Ahmed Hussain 0001, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil
ICC4
2025 Explainable and Robust Millimeter Wave Beam Alignment for AI-Native 6G Networks
abstract
Integrated artificial intelligence (AI) and communication has been recognized as a key pillar of 6 G and beyond networks. In line with AI-native 6 G vision, explainability and robustness in AI-driven systems are critical for establishing trust and ensuring reliable performance in diverse and evolving environments. This paper addresses these challenges by developing a robust and explainable deep learning (DL)-based beam alignment engine (BAE) for millimeter-wave (mmWave) multiple-input multiple-output (MIMO) systems. The proposed convolutional neural network (CNN)-based BAE utilizes received signal strength indicator (RSSI) measurements over a set of wide beams to accurately predict the best narrow beam for each UE, significantly reducing the overhead associated with exhaustive codebook-based narrow beam sweeping for initial access (IA) and data transmission. To ensure transparency and resilience, the Deep k-Nearest Neighbors (DkNN) algorithm is employed to assess the internal representations of the network via nearest neighbor approach, providing human-interpretable explanations and confidence metrics for detecting out-of-distribution inputs. Experimental results demonstrate that the proposed DL-based BAE exhibits robustness to measurement noise, reduces beam training overhead by 75 % compared to the exhaustive search while maintaining near-optimal performance in terms of spectral efficiency. Moreover, the proposed framework improves outlier detection robustness by up to$5 \times$and offers clearer insights into beam prediction decisions compared to traditional softmax-based classifiers.
Nasir Khan, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil, Sinem Coleri Ergen
ICC4
2025 RIS-Empowered Jamming for Protected Zone Formation in In-Band Full Duplex Systems
abstract
The rapid proliferation of mobile technologies introduces substantial security vulnerabilities, with conventional cryptographic approaches increasingly inadequate. Physical layer security (PLS) approaches, especially when paired with advanced methods like full-duplex (FD) communication and reconfigurable intelligent surfaces (RISs), offer substantial potential for enhanced security. This work proposes a novel RIS-aided inband FD (IBFD) framework that employs artificial noise (AN) to establish a protected zone (PZ) around the legitimate user, operating independently of any a priori knowledge regarding the number, positions, or channels of potential eavesdroppers. To achieve this, the RIS is functionally partitioned into two distinct segments: one enhances the legitimate user's data rate, while the other simultaneously intensifies AN to jam any unauthorized users within the PZ. We formulate a secrecy capacity (SC) maximization problem that optimizes both the legitimate user's transmission power and the RIS configuration, while meeting quality-of-service (QoS) requirements. Simulation results indicate the superior efficacy of the proposed technique in enhancing SC compared to benchmarks that employ RIS and AN independently. In particular, the proposed approach achieves an approximate SC improvement of 80% and 49% over the “RISonly” and “AN-only” approaches, respectively, underscoring its significant potential to enable robust PLS in next-generation communication networks.
Hanadi Salman, Abdulkadir Celik, Sultangali Arzykulov, Ahmed M. Eltawil, Hüseyin Arslan
ICC4
2025 Reconfigurable Precision INT4-8/FP8 Digital Compute-in-Memory Macro for AI Acceleration
abstract
Compute-in-memory (CIM) technology has emerged as a promising solution to address the computational demands of deep neural network (DNN) models, which require substantial multiply-accumulate (MAC) operations. However, there is a growing need for reconfigurable CIM architectures that can support both integer (INT) and floating-point (FP) operations within a single design. This flexibility is crucial for optimizing efficiency and resource utilization, especially in DNN applications involving mixed-precision computations. In this work, we propose a reconfigurable precision digital macro design that accelerates MAC computations while supporting INT4, INT8, and FP8 configurations within the same architecture. Both signed and unsigned operations are supported in INT mode. To enhance performance, the proposed design uses a parallel-input approach and a mantissa parallel-alignment technique in FP mode. The macro is implemented in 40nm CMOS technology. It achieves a peak throughput of 7123.48 GOPS in INT4 mode and 1187.25 GFLOPS in FP8 mode, with peak energy efficiencies of 367.45 TOPS/W and 23.14 TFLOPS/W, respectively.
Jinane Bazzi, Mohamed E. Fouda, Ahmed M. Eltawil
ISCAS3
2025 Towards Efficient IMC Accelerator Design Through Joint Hardware-Workload Co-optimization
abstract
Designing generalized in-memory computing (IMC) hardware that efficiently supports a variety of workloads requires extensive design space exploration, which is infeasible to perform manually. Optimizing hardware individually for each workload or solely for the largest workload often fails to yield the most efficient generalized solutions. To address this, we propose a joint hardware-workload optimization framework that identifies optimised IMC chip architecture parameters, enabling more efficient, workload-flexible hardware. We show that joint optimization achieves 36%, 36%, 20%, and 69% better energy-latency-area scores for VGG16, ResNet18, AlexNet, and MobileNetV3, respectively, compared to the separate architecture parameters search optimizing for a single largest workload. Additionally, we quantify the performance trade-offs and losses of the resulting generalized IMC hardware compared to workload-specific IMC designs.
Olga Krestinskaya, Mohamed E. Fouda, Ahmed M. Eltawil, Khaled N. Salama
ISCAS3
2025 Spatial Path Index Modulation for RIS-Aided Massive MIMO
abstract
The next generation wireless networks focus on improving the energy and spectral efficiency (SE/EE) of the communication systems in response to the demand for massive number of users and data rate. In this work, we aim to achieve the enhancement of SE and EE by employing index modulation (IM) techniques for reconfigurable intelligent surface (RIS)-aided communication systems. While RIS offers the network efficiency and improve the coverage, IM provides SE improvement by the transmission of additional index bits. In IM, we utilize the indices of the spatial paths between the base station and the user through the RIS. We introduce a low complexity approach for the design of hybrid beamformers, which are constructed by the steering vectors corresponding to the selected spatial path indices for IM. Via numerical experiments, we show that the proposed approach presents an improved SE performance, even higher than that of the use of fully-digital beamformers while using a few RF chains.
Ahmet M. Elbir, Abdulkadir Celik, Asmaa Abdallah, Ahmed M. Eltawil
PIMRC4
2025 Analyzing URA Geometry for Enhanced Spatial Multiplexing and Extended Near-Field Coverage
abstract
With the deployment of large antenna arrays at high-frequency bands, future wireless communication systems are likely to operate in the radiative near-field. Unlike far-field beam steering, near-field beams can be focused within a spatial region of finite depth, enabling spatial multiplexing in both the angular and range dimensions. This paper derives the beamdepth for a generalized uniform rectangular array (URA) and investigates how array geometry influences the near-field beamdepth and the limits where near-field beamfocusing is achievable. To characterize the near-field boundary in terms of beamfocusing and spatial multiplexing gains, we define the effective beamfocusing Rayleigh distance (EBRD) for a generalized URA. Our analysis reveals that while a square URA achieves the narrowest beamdepth, the EBRD is maximized for a wide or tall URA. However, despite its narrow beamdepth, a square URA may experience a reduction in multiuser sum rate due to its severely constrained EBRD. Simulation results confirm that a wide or tall URA achieves a sum rate of 3.5× more than that of a square URA, benefiting from the extended EBRD and improved spatial multiplexing capabilities.
Ahmed Hussain 0001, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil
PIMRC4
2025 Conditional Diffusion Model-Driven RSSI Generation in Cell-Free mmWave Systems
abstract
Millimeter-wave (mmWave) communication enables ultra-high data rates but suffers from severe path loss, especially under non-line-of-sight (NLoS) conditions. Cell-free networks with distributed base stations (BSs) mitigate this by increasing the chance of at least one BS maintaining a line-of-sight (LoS) link to each user equipment (UE). Integrated beamforming further boosts performance by focusing energy in desired directions. Conventional beam training uses exhaustive beam sweeping, testing all beam pair combinations from predefined codebooks via received signal strength indicator (RSSI) measurements. While effective, this method incurs significant latency and energy costs, particularly as the codebook size increases with the number of transmit and receive antennas. To address this, we propose a learning-based approach for cell-free mmWave systems that predicts RSSI matrices from UE and BS coordinates, eliminating the need for exhaustive sweeping. We train two conditional diffusion models—a denoising diffusion probabilistic model (DDPM) and a faster denoising diffusion implicit model (DDIM). Our method reduces beam training overhead by up to 97% while maintaining at least 94% of the peak rate achieved via exhaustive search. DDIM further enables simultaneous RSSI generation for 64 UEs at 60% of the time required by conventional methods, significantly cutting overhead with minimal performance loss.
Khalid Kanaan, Malak Alhulimi, Matteo Parsani, Ahmed M. Eltawil
PIMRC4
2025 Multimodal Sensing and DRL-Driven Beam Selection in RIS-Aided mmWave mMIMO Systems
abstract
The IMT-2030 vision emphasizes two key 6G directions: integrated sensing and communication (ISAC) alongside artificial intelligence (AI)-native frameworks, where multimodal sensory data inputs enhance situational awareness and adaptive decision-making of communication systems. Accordingly, this paper introduces a deep reinforcement learning (DRL)-based beam selection framework for downlink multi-user reconfigurable intelligent surface (RIS)-assisted millimeter-wave (mmWave) massive MIMO (mMIMO) systems. Targeting maximized sum rates under quality of service (QoS) and fairness constraints, the framework employs two primary sensing modalities: a stereo camera mounted on the RIS for user equipment (UE) detection and inertial measurement units (IMUs) on UEs to obtain 3D Cartesian coordinates, thus eliminating the need for channel state information (CSI) acquisition. The DRL framework combines two algorithms—double deep Q-network (DDQN) and proximal policy optimization (PPO)—to jointly optimize RIS phase shifts and UE receive beamformers through predefined codebooks and adaptive beam selection. A testbed was developed to validate the system, leveraging real-world data to train the DRL algorithms. Experimental results demonstrate that both agents achieve nearoptimal sum rates across diverse base station (BS) transmit power levels and QoS thresholds while reducing computational complexity by 95%, illustrating the framework’s potential for efficient and scalable beam alignment for AI-native wireless systems.
Khalid Kanaan, Ahmed Nasser, Abdulkadir Celik, Atif Shamim, Ahmed M. Eltawil
PIMRC6
2025 Optimizing Deployment and Partitioning Strategies for Aerial RIS-aided Uplink NOMA under Residual Hardware Impairments
abstract
The incorporation of reconfigurable intelligent surfaces (RISs) and unmanned aerial vehicles (UAVs) presents considerable potential for improving the functionality of ground-based terrestrial Internet of Things (IoT) networks. This paper introduces a novel UAV deployment and aerial RIS partitioning mechanism for the uplink non-orthogonal multiple access (NOMA)-based IoT networks under practical hardware constraints. The low-cost hardware of IoT devices results in imperfect successive interference cancellation (SIC) and distortion noise due to transceiver hardware impairments (T-HIs). Specifically, we have optimized the aerial RIS partitioning and its deployment to maximize the minimum rate under the impact of imperfect SIC and T-HIs. The numerical results show that the proposed scheme outperforms the benchmark. Through extensive simulations, we prove that the proposed UAV-based aerial RIS-aided uplink NOMA network significantly enhances the max-min fair rate.
Mohd Hamza Naim Shaikh, Abdulkadir Celik, Sultangali Arzykulov, Ahmed M. Eltawil, Galymzhan Nauryzbayev
PIMRC4
2025 From Lab to Digital Twin: Calibration of mmWave Ray-Tracing with RIS Reflections
abstract
Reconfigurable intelligent surfaces (RISs) are gaining significant attention as a key enabler of future wireless networks. However, its practical deployment is hindered by challenges in modeling and integration. Existing analytical approaches often depend on idealized assumptions, limiting their ability to reflect the complexity of real-world environments. In this work, we explore the integration of digital twin (DT) technology with ray tracing (RT), enabling a more accurate representation of practical scenarios and bridging the gap between theoretical models and the real-world implementation of RIS. RT allows accurate prediction of signal behavior, such as received signal strength indicator (RSSI) levels, without the need for expensive experimental measurements. We evaluate the performance of the DT with RT in three RIS-aided setups: a single RIS-aided, cascaded RISs-aided, and RIS partitioning. Our results show that the proposed DT model closely matches the experimental RSSI data with an error margin below 1 dB for single and partitioned RIS setups and under 2 dB for cascaded RISs system. These findings highlight the potential of DT with RT simulations as a practical tool for performance evaluation and optimization of RIS-assisted wireless systems.
Zhandos Zhakipov, Madi Makin, Ahmed Nasser, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil
PIMRC6
2025 Mitigating the Impact of ReRAM I-V Nonlinearity and IR Drop via Fast Offline Network Training
abstract
ReRAM crossbar arrays (RCAs) have the potential to provide extremely high efficiency for accelerating deep neural networks (DNNs). However, one crucial challenge for RCA-based DNN accelerators is functional inaccuracy due to nonidealities present in RCA hardware. While nonideality-aware training (NAT) could be used to mitigate the effect of nonidealities, with currently available methods it would take months to train even a medium size convolutional neural network (CNN). In this article we propose a nonideality prediction method that enables very fast training of RCA-based neural networks, and show its feasibility through NAT of DNNs. Our key ideas include 1) weight-centric nonideality modeling and 2) data-dependence elimination by tailored input randomization. Our experimental results using a multilayer perceptron and CNNs demonstrate that our method is very fast ($100\sim 15$$000\times $faster training speed) while achieving much better-crossbar-level accuracy ($2 \sim 90\times $lower-RMS error) and post-retraining validated accuracy than previous methods.
Sugil Lee, Mohamed E. Fouda, Chenghao Quan, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 Autoencoder-Based Transceivers for Multiple Access Human Body Communication Networks
abstract
The Internet of Bodies (IoB) represents a transformative technological innovation that merges the bio-physical and digital realms through networks of intelligent devices positioned in, on, and around the human body. Human Body Communication (HBC) offers a promising method for enabling IoB networks, using the human body as a communication channel for multiple wearable nodes. Despite the prevalence of HBC peer-to-peer communication methods in the literature, the challenge of implementing multiple access techniques for HBC at the physical layer remains largely unexplored. In this paper, we propose a new multiple access HBC (MA-HBC) system that leverages autoencoders to design and implement low-power and efficient transceivers sharing a common channel. The proposed MA-HBC system consists of multiple autoencoder-based transceivers trained jointly to optimize overall network performance. It supports various data rates ranging from 164 kbps to 5.25 Mbps, making it suitable for a wide range of IoB applications. To validate the design, a prototype implementation is presented. Additionally, to ensure suitability for wearable devices, a low-power hardware transceiver architecture is provided with an estimated energy efficiency of 105 pJ/b when implemented using TSMC 65nm LP technology. The results show that the proposed MA-HBC system outperforms traditional IEEE 802.15.6 based transceivers for two users with time Sharing, achieving a 3.9 dB improvement in Signal-to-Noise Ratio (SNR) at a block error rate of$10^{-2}$.
Abdelhay Ali, Amr N. Abdelrahman, Abdulkadir Celik, Ahmed M. Eltawil
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 Single-Cycle Independent Component Analysis Processor for In-Band Full Duplex Systems
abstract
This paper presents a single-cycle independent component analysis (ICA) algorithm and architecture for self-interference cancellation (SIC) in In-band Full-duplex (IBFD) communication systems. The proposed algorithm, AICA-EBM, incorporates the adaptive momentum (ADAM) approach with the entropy-bound estimation (EBM) to achieve rapid convergence. The simulation results show that the proposed AICA-EBM algorithm leads to a constant single-cycle processing time to achieve satisfactory SIC optimality in IBFD systems. The architecture of the ICA processor based on the proposed AICA-EBM algorithm is designed. Novel circuit structures are proposed to reduce the complexity of the ICA processor. The AICA-EBM processor is implemented by following an application-specific integrated circuit (ASIC) flow with the TSMC 90 nm process. The post-layout estimations show that Compared to previous ICA processors, the proposed design achieves a 2x improvement in processing throughput and leads to the best hardware efficiency.
Hao-Lun Weng, Chung-An Shen, Mohamed E. Fouda, Ahmed M. Eltawil
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 Multi-Agent DRL for Distributed Codebook Design in RIS-Aided Cell-Free Massive MIMO Networks
abstract
This paper proposes an innovative approach for enhancing network capacity and coverage by integrating cell-free massive multiple-input multiple-output (CF-mMIMO) networks with reconfigurable intelligent surfaces (RISs). A significant challenge in leveraging RIS-assisted CF-mMIMO lies in the cooperative beam training across multiple access points (APs) and RISs, complicated by the passive nature of reflective elements and the complexity channel state information (CSI) acquisition in millimeter wave mMIMO systems. To address these challenges, we develop a multi-agent deep reinforcement learning (MA-DRL) framework that jointly designs beamforming and reflection codebooks for distributed APs and RISs, eliminating the need for CSI and relying solely on received power measurements feedback. The joint beamforming and reflection codebook design problem is decomposed into two sub-problems: one for beam codebook design at APs and another for sequential reflection codebook design at RISs. We employ transfer learning to speed up learning convergence and reduce computational complexity for training multiple RISs. Additionally, we introduce an AP and RIS selection scheme that improves overall energy efficiency and reduces backhaul overhead. Extensive simulations demonstrate that our proposed MA-DRL approach curtails number of beams significantly, thereby outperforming the widely adopted discrete Fourier transform (DFT) codebooks by achieving an 84% reduction in beam training overhead. Our findings suggest that increasing the number of passive RISs allows putting more APs into idle mode, leading to substantial savings in hardware and energy costs.
Asmaa Abdallah, Abdulkadir Celik, Mohammad M. Mansour, Ahmed M. Eltawil
IEEE Trans. Commun.4
2025 Optimization of Hybrid Laser-Battery-Powered UAV-Assisted Backscatter Communications
abstract
This work considers a hybrid laser-battery powered uncrewed aerial vehicle (UAV) data collection system serving a passive Internet of Things deployment via monostatic backscatter communications. In this paper, we highlight the merits of the hybrid scheme over the laser only or battery only powered devices UAVs in terms of improved reach and durability. In addition, we study the laser energy consumption - destination battery level retention tradeoff optimization problem. Throughout this process, we optimize the single-rotor UAV’s three-dimensional trajectory, the UAV’s and the laser’s radiated power profiles, and the temporal battery usage profile while adopting path discretization. The resulting non-convex problem is solved via single-block successive convex approximation, for which novel bounds for the UAV propulsion energy, harvested energy, and collected data assuming a probabilistic line-of-sight channel model are derived. Finally, the simulation results show significant data collection gains, battery energy savings, and laser energy consumption reductions compared with a baseline scheme and highlight the complexity-optimality tradeoff.
Amr M. Abdelhady, Carles Diaz-Vilor, Mohammadreza Barzegaran, Hamid Jafarkhani, Ahmed M. Eltawil
IEEE Trans. Commun.5
2025 Explainable AI-Aided Feature Selection and Model Reduction for DRL-Based V2X Resource Allocation
abstract
Artificial intelligence (AI) is expected to significantly enhance radio resource management (RRM) in sixth-generation (6G) networks. However, the lack of explainability in complex deep learning (DL) models poses a challenge for practical implementation. This paper proposes a novel explainable AI (XAI)-based framework for feature selection and model complexity reduction in a model-agnostic manner. Applied to a multi-agent deep reinforcement learning (MADRL) setting, our approach addresses the joint sub-band assignment and power allocation problem in cellular vehicle-to-everything (V2X) communications. We propose a novel two-stage systematic explainability framework leveraging feature relevance-oriented XAI to simplify the DRL agents. While the former stage generates a state feature importance ranking of the trained models using Shapley additive explanations (SHAP)-based importance scores, the latter stage exploits these importance-based rankings to simplify the state space of the agents by removing the least important features from the model’s input. Simulation results demonstrate that the XAI-assisted methodology achieves ~97% of the original MADRL sum-rate performance while reducing optimal state features by ~28%, average training time by ~11%, and trainable weight parameters by ~46% in a network with eight vehicular pairs.
Nasir Khan, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil, Sinem Coleri Ergen
IEEE Trans. Commun.4
2025 Redefining Polar Boundaries for Near-Field Channel Estimation for Ultra-Massive MIMO Antenna Array
abstract
Ultra massive multiple-input-multiple-output (UM-MIMO) technology has emerged as a promising candidate for 6G networks, offering ultra-high spectral efficiency in wireless systems. The transition to large antenna arrays specially at high-frequency bands is fundamentally transforming wireless communication from the traditional far-field to the near-field realm. This transition poses a distinct challenge in channel estimation due to the associated pilot overhead from large antenna arrays and the absence of angular sparsity in near-field spherical wavefronts. However, polar-domain sparsity remains achievable, advocating the use of polar codebooks over traditional angle-based ones. Nevertheless, the size of the polar codebook presents a significant challenge, necessitating sampling of distance and angle points across the entire near-field. In this work, we investigate the near-field channel estimation techniques while identifying the boundaries of polar domain sparsity with minimal pilot overhead. We propose a novel polar codebook, which leverages our findings from sparsity analysis and exploits the beam-focusing properties of the near-field. Unlike existing work, the proposed polar codebook design is agnostic to user range information and has considerably reduced dimensions. Capitalizing on this new polar codebook, we introduce the beam focused simultaneous orthogonal matching pursuit (BF-SOMP) algorithm for efficient near-field channel estimation. To further improve the channel estimation accuracy, we then present a refinement procedure that iterates over off-grid angle and range samples to enhance the estimation accuracy. Simulation results demonstrate that the proposed polar codebook based algorithms outperform contemporary methods in terms of improved normalized mean square error (NMSE) and reduced computational complexity. When compared to existing channel estimation methods, the proposed algorithms achieve an NMSE improvement of 6 − 7 dB at low and high SNR values with 32 pilots, while utilizing a codebook nearly half the size of existing ones.
Ahmed Hussain 0001, Asmaa Abdallah, Ahmed M. Eltawil
IEEE Trans. Wirel. Commun.3
2024 End-to-End Learning of Beam Probing and RSSI-Based Multi-User Hybrid Precoding Design
abstract
This paper presents an end-to-end (E2E) autoencoder learning framework that relies on unsupervised deep learning for the joint design of millimeter wave (mmWave) probing beams and hybrid precoding matrices in multi-user communication systems. Our model utilizes prior channel observations to achieve two main objectives: designing a compact set of probing beams and predicting off-grid radio frequency (RF) beamforming vectors. The E2E learning framework optimizes probing beams in an unsupervised manner, concentrating sensing power on promising spatial directions based on the environment. To this aim, we develop a neural network architecture respecting RF chain constraints and model received signal strength (RSS) using complex-valued convolutional layers. The autoencoder is trained to directly produce RF beamforming vectors for hybrid architectures based on projected RSS indicators (RSSIs). Once RF beamforming vectors for multi-users are predicted, baseband digital precoders are designed by accounting for multi-user interference. The autoencoder neural network is trained E2E in an unsupervised manner with a customized loss function aimed at maximizing RSS. In a system with 64 antennas, 4 RF chains, and 4 users, our approach requires only 8 probing beams to design RF beamforming vectors, compared to the conventional predefined codebooks with 64 or 128 beams.
Asmaa Abdallah, Abdulkadir Celik, Ahmed Alkhateeb, Ahmed M. Eltawil
GLOBECOM4
2024 Joint Antenna and Spatial Path Index Modulation for THz Integrated Sensing and Communications
abstract
Beam-squint is a challenging issue in ultra-wideband systems, e.g., terahertz (THz) integrated sensing and communications (ISAC). In order to compensate for the loss due to beam-squint, this paper leverages index modulation in spatial domain, which enables the transmission of additional information bits to improve the spectral efficiency (SE). Specifically, a joint antenna and spatial path index modulation (JASPIM) technique is proposed by exploiting the spatial diversity of both antenna and path indices. We present a hybrid beamforming technique with JASPIM for ISAC, wherein the analog beamformers are designed in accordance with the radar targets and the communications user. Numerical simulations demonstrate that our JASPIM-ISAC approach exhibits a significant SE improvement even higher than that of the use of fully digital beamformers in the presence of beam-squint.
Ahmet M. Elbir, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil
GLOBECOM4
2024 Multiple Access Optimization for Multi-Modal STAR-RIS-assisted Full-Duplex Communication
abstract
This work investigates simultaneously transmitting and reflecting reconfigurable intelligent surfaces (STAR-RIS) in multiple-access (MA) bidirectional network architecture. The proposed optimization framework incorporates two practical operational protocols of STAR-RIS to optimize the channel gains for users and obviate the necessity for uplink power control while adhering to specified quality of service (QoS) constraints. The proposed approach undergoes thorough evaluation under key optimization problems: QoS feasible region, and energy efficiency. The closed-form solutions are validated through simulations, showcasing the notable advantages that STAR-RIS can provide to the considered multiple-access networks. Simulation findings revealed that in the context of the proposed system model, the mode switching approach can attain a superior QoS threshold rate compared to the energy splitting mode. The proposed framework has demonstrated its ability to meet bidirectional communication needs with a high degree of accuracy and precision.
Madi Makin, Abdulkadir Celik, Sultangali Arzykulov, Ahmed M. Eltawil, Galymzhan Nauryzbayev
GLOBECOM4
2024 Online DRL-based Beam Selection for RIS-Aided Physical Layer Security: An Experimental Study
abstract
The integration of reconfigurable intelligent surfaces (RIS) and artificial noise (AN) significantly enhances physical layer security (PLS) in wireless networks, provided that RIS’s phase shifts are precisely optimized to prevent security vulnerabilities. This paper introduces a reinforcement learning (RL)based algorithm designed to optimize the phase shifts in RIS-partitioning-aided PLS systems operating in the millimeter wave (mm-Wave), without requiring channel state information (CSI) for any users. The RL algorithm optimizes the phase shifts by efficiently selecting the best beam from a predefined codebook for different partitions, which simultaneously enhances the intended signal for legitimate users and increases the effectiveness of AN on eavesdroppers, thereby maximizing the system’s secrecy capacity (SC) and addressing the inherent non-convex challenges. Additionally, the paper details the development of an experimental testbed that provides essential data to refine the algorithm. The numerical results from the testbed highlight the significant impact of RIS partitioning in PLS, which can enhance the SC by an average of 55% over the full RIS scenario, and confirm the effectiveness of the RL-based algorithm in reducing computational complexity by approximately 80% compared to the exhaustive search algorithm.
Ahmed Nasser, Abdulkadir Celik, Asmaa Abdallah, David Lago-Cachón, Atif Shamim, Ahmed M. Eltawil
GLOBECOM8
2024 Secure and Efficient sEMG Signal Transmission Using Human Body Communication for Upper Limb Prostheses
abstract
In recent years, surface electromyography (sEMG) signals have emerged as a valuable tool for assisting individuals with physical disabilities. Traditional methods for transmitting sEMG signals often rely on radio frequency (RF) links, which are energy-intensive and lack robust security. This paper presents a new design for a Human Body Communication (HBC) transceiver tailored for sEMG sensors. The proposed HBC transceiver employs an end-to-end autoencoder approach, offering enhanced energy efficiency and security. The architecture supports the required data rates for sEMG applications and is optimized for low power consumption, making it suitable for wearable devices. The results show that the HBC transceiver design demonstrates a peak data rate of 62.5 Kbps and a block error rate of approximately$10^{-2}$at −5.17 dB of Ec/N0 when operating at a clock speed of 2 MHz. Moreover, the power consumption results indicate that the HBC transceiver consumes 287$\mu \mathrm{W}$resulting in an energy efficiency of 4.5 nJ/bit,
Abdelhay Ali, Abdulkadir Celik, Ahmed M. Eltawil
HealthCom3
2024 Reconfigurable Precision SRAM-based Analog In-memory-compute Macro Design
abstract
In-memory computing (IMC) is a promising approach for accelerating multiply and accumulate (MAC) operations, which are the primary calculations used in artificial intelligence (AI). The demand for flexible architectures supporting different bit precisions in MAC computations becomes evident. This flexibility balances adapting to specific model requirements and optimizing design performance efficiency. As such, in this paper, we propose a reconfigurable IMC macro design, utilizing 8T static random-access memory (SRAM) bit-cells in 65nm technology, to efficiently perform MAC operations while supporting three bit precisions: 2, 3, and 4 bits for each of the input, weight, and output. The proposed 64×180 macro achieves a normalized peak throughput of 13.82 TOPS, a normalized peak energy efficiency of 291.66 TOPS/W, and a normalized peak area efficiency of 165.98 TOPS/mm2.
Jinane Bazzi, Rachid Jamil, Dana El Hajj, Rouwaida Kanj, Mohamed E. Fouda, Ahmed M. Eltawil
ISCAS6
2024 Optimal Partitioning of Reconfigurable Intelligent Surfaces for Uplink NOMA Networks
abstract
In this work, we examine the potential of reconfigurable intelligent surfaces (RISs) to facilitate and enhance uplink (UL) transmissions in grant-free non-orthogonal multiple access (GF-NOMA) networks. The proposed RIS-assisted GF-NOMA approach employs virtual partitioning of RIS, with each partition tailored to optimize channel conditions for individual NOMA user equipment (UE). The resulting channel gain disparity bolsters the NOMA gain and obviates the necessity for UL power control of the grant-based NOMA schemes. Our approach is evaluated under three practical operational regimes: 1) quality-of-service (QoS) sufficient regime, 2) efficient RIS usage regime, and 3) max-min fair regime, all subject to UL-QoS constraints. We derive closed-form solutions to elucidate how optimal RIS partitioning can fulfill UL-QoS requirements across all three operational regimes. Comprehensive simulations are conducted to validate the precision of our analytical findings, demonstrating that the proposed approach substantially improves wireless communication system performance while mitigating signaling overhead and computational complexity.
Madi Makin, Abdulkadir Celik, Sultangali Arzykulov, Ahmed M. Eltawil, Galymzhan Nauryzbayev
PIMRC4
2024 A Multi-Armed Bandit Approach for User-Target Pairing in NOMA-Aided ISAC
abstract
In this paper, we propose a robust interference management approach for the integrated sensing and communication (ISAC) system that employs non-orthogonal multiple access (NOMA) for multiplexing. Our proposed approach effectively addresses interference challenges by optimizing the pairing of communication users (CUs) and radar targets (RTs) while simultaneously designing receiving beamformers. These optimizations aim to maximize the combined utility of communication rates and the radar estimation information rate (REIR), inherently constituting a challenging non-convex combinatorial problem. To tackle this intricate problem, we employ the upper confidence bound (UCB) algorithm, a powerful online learning technique rooted in multi-armed bandit (MAB) theory. Along with UCB, we harness zeroforcing beamforming to optimize the receiving beamformer. The numerical results underscore the importance of CU-RT pairing, with a $65 \%$ average performance improvement over traditional NOMA-ISAC and OMA-ISAC, close to the exhaustive search performance by only $2 \%$. It also substantially reduces complexity, with about $90 \%$ less computational complexity than exhaustive search.
Ahmed Nasser, Abdulkadir Celik, Ahmed M. Eltawil
PIMRC3
2024 Operation Optimization of Laser-Powered Aerial Data Harvesting for Passive IoT Networks
abstract
This paper investigates the maximization of har-vested data in a laser-powered uncrewed aerial vehicle (UAV) supporting Internet of Things (IoT) deployment. The system enables battery-free IoT devices to establish communication links with the UAV via bistatic backscattering with the aid of a power beacon source. Upon considering an unspecified flying time, we adopt path discretization and resort to the single-block successive convex approximation (SCA) to solve the data collection maximization problem. In addition to considering the UAV dynamics and power budget, two novel SCA-compatible bounds are introduced for the product of mixed convex/concave positive functions. Finally, the simulations conducted show that the proposed algorithm provides 90% increase in collected data under different operation conditions.
Amr M. Abdelhady, Abdulkadir Celik, Carles Diaz-Vilor, Hamid Jafarkhani, Ahmed M. Eltawil
WCNC5
2024 Spatial Path Index Modulation to Combat Beam-Squint Effect in THz-ISAC Systems
abstract
In terahertz (THz) wideband systems, beam-squint causes deviations in the generated beam directions at different subcarriers due to the use of subcarrier-independent analog beamformers. In order to combat the performance loss due to beam-squint effect, this work employs spatial path index modulation (SPIM) to improve the spectral efficiency (SE) performance of the overall system, thereby compensating the loss due to beam-squint. Specifically, SPIM allows the transmission of additional information bits to the receiver via modulating the indices of the spatial paths. The proposed approach is evaluated in a THz integrated sensing and communications (THz-ISAC) scenario, wherein the beamformer design allows generating multiple beams toward both radar targets and the communications user. Numerical simulations demonstrate that the proposed approach exhibits significant SE performance even higher than that of the use of fully digital beamformers without SPIM in the presence of beam-squint.
Ahmet M. Elbir, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil
WCNC4
2024 Near-Field Channel Estimation for Ultra-Massive MIMO Antenna Array with Hybrid Architecture
abstract
Ultra massive multiple-input-multiple-output (UM-MIMO) has emerged as a prospective capacity-enhancing technology for next generation (NG) networks. To harness its full potential, accurate channel estimation with low pilot overhead becomes paramount. The combination of large antenna arrays and high frequency bands transitions the wireless communication from the far-field to the near-field realm. This transition presents a unique challenge, as channel sparsity in the angular domain becomes unattainable in the presence of near-field spherical wavefronts. Nonetheless, polar-domain sparsity is achievable, allowing existing near-field channel estimation methods to use polar codebooks over the classical angle-based codebooks. However, a major challenge in utilizing polar domain sparsity is the size of the polar codebook, demanding sampling of both distance and angle points across the entire near-field. In this work, we investigate near-field channel response in the beam space domain to redefine the boundaries of polar domain sparsity. We propose a novel polar codebook leveraging our results of sparsity analysis and beamfocusing property of near-field. Unlike existing work, the proposed polar codebook design is agnostic to user range information and has considerably reduced dimensions. Exploiting the new polar codebook, we present the beam focused simultaneous orthogonal matching pursuit (BF-SOMP) algorithm for efficient near-field channel estimation. Simulation results demonstrate that the proposed algorithm surpasses contemporary methods in terms of normalized mean square error (NMSE) and reduced computational complexity.
Ahmed Hussain 0001, Asmaa Abdallah, Ahmed M. Eltawil
WCNC3
2024 Multi-Agent Deep Reinforcement Learning for Beam Codebook Design in RIS-Aided Systems
abstract
Reconfigurable intelligent surfaces (RISs) play a vital role in future wireless systems with the capability of enhancing propagation environments by intelligently reflecting the signals toward the target receivers. However, optimal tuning of the phase shifters at the RIS is challenging due to the passive nature of reflective elements and the high complexity of acquiring channel state information (CSI). Furthermore, the joint active beamforming and RIS reflection beam design is a tedious task due to the high computational complexity and the dynamic nature of the wireless environment. Today’s cellular networks establish data transmission by relying on pre-defined generic beamforming codebooks, which are neither site-specific nor adaptive to the changes in the wireless environment. Moreover, identifying the best beam is typically performed using an exhaustive search approach that prohibits the use of large codebook sizes due to the resulting high beam training overhead. Depending merely on the binary received signal strength, this work develops a multi-agent deep reinforcement learning (MA-DRL) framework that jointly designs the active and the passive reflection beam codebooks for the BS and the RIS, reflectively. To accelerate learning convergence and reduce the search space, the proposed model divides the RIS into multiple partitions and associates beam patterns to the surrounding environments with low computational complexity. Moreover, a hierarchical beam training solution is proposed to further reduce the beam training overhead of the single-beam training approach. Simulation results show that the proposed MA-DRL approach can provide a 97% beam training overhead reduction over the discrete Fourier transform (DFT) codebook.
Asmaa Abdallah, Abdulkadir Celik, Mohammad M. Mansour, Ahmed M. Eltawil
IEEE Trans. Wirel. Commun.4
2024 Sensing and Communication in UAV Cellular Networks: Design and Optimization
abstract
Recently, the use of uncrewed aerial vehicles (UAVs) in joint sensing and communication applications has received a lot of attention. However, integrating UAVs in current cellular systems presents major challenges related to trajectory optimization and interference management among others. This paper considers a multi-cell network including a UAV, which senses and forwards the sensory data from different events to the central base station. Particularly, the current manuscript covers how to design the UAV’s (i) 3D trajectory, (ii) power allocation, and (iii) sensing scheduling such that (a) a set of events are sensed, (b) interference to neighboring cells is kept at bay, and (c) the amount of energy required by the UAV is minimized. The resulting nonconvex optimization problem is tackled through a combination of (i) low-complexity binary optimization, (ii) successive convex approximation, and (iii) the Lagrangian method. Simulation results over a range of various key parameters have shown the merits of our approach, which consumes 33%-200% less energy compared to different benchmarks.
Carles Diaz-Vilor, Mojtaba Ahmadi Almasi, Amr M. Abdelhady, Abdulkadir Celik, Ahmed M. Eltawil, Hamid Jafarkhani
IEEE Trans. Wirel. Commun.5
2024 Multi-UAV Reinforcement Learning for Data Collection in Cellular MIMO Networks
abstract
Uncrewed Aerial Vehicles (UAVs) provide a compelling solution for data collection in Internet of Things (IoT) networks due to their mobility and adaptability. However, the line-of-sight dominance in their channels may result in severe interference to ground users during UAV operations. To address this, we present an optimization framework that concurrently optimizes UAV trajectories and transmit powers. Our approach efficiently results in the collection of data from a variety of IoT sensors while (a) minimizing the UAVs flying time and (b) mitigating interference with terrestrial networks. Given the complex nature of such an optimization problem, this paper leverages reinforcement learning, specifically the twin delayed deep deterministic policy gradient algorithm, where a distributed learning algorithm is presented. Experimental results validate the efficacy of our proposed approach, demonstrating its capability to significantly enhance data collection in IoT networks while minimizing UAV flight time and interference with ground user links.
Carles Diaz-Vilor, Amr M. Abdelhady, Ahmed M. Eltawil, Hamid Jafarkhani
IEEE Trans. Wirel. Commun.3
2024 Spatial Path Index Modulation in mmWave/THz Band Integrated Sensing and Communications
abstract
As the demand for wireless connectivity continues to soar, the fifth generation and beyond wireless networks are exploring new ways to efficiently utilize the wireless spectrum and reduce hardware costs. One such approach is the integration of sensing and communications (ISAC) paradigms to jointly access the spectrum. Recent ISAC studies have focused on upper millimeter-wave and low terahertz bands to exploit ultrawide bandwidths. At these frequencies, hybrid beamformers that employ fewer radio-frequency chains are employed to offset expensive hardware but at the cost of lower multiplexing gains. Wideband hybrid beamforming also suffers from the beam-split effect arising from the subcarrier-independent (SI) analog beamformers. To overcome these limitations, we introduce a spatial path index modulation (SPIM) ISAC architecture, which transmits additional information bits via modulating the spatial paths between the base station and communications users. We design the SPIM-ISAC beamformers by estimating both radar and communications parameters through our proposed beam-split-aware algorithms. We then develop a family of hybrid beamforming techniques – hybrid, SI, subcarrier-dependent analog-only, and beam-split-aware beamformers – for SPIM-ISAC. Numerical experiments demonstrate that the proposed approach exhibits significantly improved spectral efficiency performance in the presence of beam-split when compared with even fully digital non-SPIM beamformers.
Ahmet M. Elbir, Kumar Vijay Mishra, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil
IEEE Trans. Wirel. Commun.5
2024 Optimal RIS Partitioning and Power Control for Bidirectional NOMA Networks
abstract
This study delves into the capabilities of reconfigurable intelligent surfaces (RISs) in enhancing bidirectional non-orthogonal multiple access (NOMA) networks. The proposed approach partitions RIS to optimize the channel conditions for NOMA users, improving NOMA gain and eliminating the requirement for uplink (UL) power control. The proposed approach is rigorously evaluated under four practical operational regimes; 1) Quality-of-Service (QoS) sufficient regime, 2) RIS and power efficient regime, 3) max-min fair regime, and 4) maximum throughput regime, each subject to both UL and downlink (DL) QoS constraints. By leveraging decoupled nature of RIS portions and base station (BS) transmit power, closed-form solutions are derived to show how optimal RIS partitioning can meet UL-QoS requirements while optimal BS power control can ensure DL-QoS compliance. Analytical findings are validated by simulations, highlighting the significant benefits that RISs can bring to the NOMA networks in the aforementioned operational scenarios.
Madi Makin, Sultangali Arzykulov, Abdulkadir Celik, Ahmed M. Eltawil, Galymzhan Nauryzbayev
IEEE Trans. Wirel. Commun.4
2024 Grant-Free NOMA Through Optimal Partitioning and Cluster Assignment in STAR-RIS Networks
abstract
The integration of reconfigurable intelligent surfaces (RISs) and grant-free non-orthogonal multiple access (GF-NOMA) has emerged as a promising solution for enhancing spectral efficiency and massive connectivity in future wireless networks. This paper proposes a GF-NOMA communication network enabled by simultaneously transmitting and reflecting RISs (STAR-RIS). In the proposed GF-NOMA, all user equipments (UEs) have instantaneous access to resource blocks (RBs) without the need for grant acquisition and power control as in the traditional grant-based NOMA schemes. Specifically, we have considered two regimes of interest: 1) the max-min fair (MMF) regime and 2) the max-sum throughput (MST) regime. To achieve the required power disparity, a two-level power control mechanism is proposed; initially, the UEs are clustered according to their channel gains. Additionally, we introduce a multi-level GF-NOMA (MGF-NOMA) scheme that adjusts the transmit power levels for each UE in the cluster. The second level of power disparity is achieved through the assignment of STAR-RISs to the clusters and optimal partitioning of STAR-RIS to support each of the cluster members. Specifically, we have also derived the closed-form equations for the optimal partitioning of STAR-RIS within the clusters for both regimes of interest. Simulation results demonstrate that the proposed STAR-RIS-aided MGF-NOMA yields a gain of 60% and 20% in the MST regime with active and passive RIS realization, respectively. Furthermore, the active and passive RIS-based MGF-NOMA achieve nearly the equivalent fairness that can be obtained through optimal power control in the MMF regime. The finding emphasizes the potential of integrating STAR-RIS with GF-NOMA as a robust and promising solution for future wireless communication systems.
Mohd Hamza Naim Shaikh, Abdulkadir Celik, Ahmed M. Eltawil, Galymzhan Nauryzbayev
IEEE Trans. Wirel. Commun.3
2023 Deep Reinforcement Learning Based Beamforming Codebook Design for RIS-aided mmWave Systems
abstract
Reconfigurable intelligent surfaces (RISs) are envisioned to play a pivotal role in future wireless systems with the capability of enhancing propagation environments by intelligently reflecting the signals toward the target receivers. However, the optimal tuning of the phase shifters at the RIS is a challenging task due to the passive nature of reflective elements and the high complexity of acquiring channel state information (CSI). Conventionally, wireless systems rely on pre-defined reflection beamforming codebooks for both initial access and data transmission. However, these existing pre-defined codebooks are commonly not adaptive to the environments. Moreover, identifying the best beam is typically performed using an exhaustive search that leads to high beam training overhead. To address these issues, this paper develops a multi-agent deep reinforcement learning framework that learns how to jointly optimize the active beamforming from the BS and the RIS-reflection beam codebook relying only on the received power measurements. To accelerate learning convergence and reduce the search space, the proposed model divides the RIS into multiple partitions and associates beam patterns to the surrounding environments with low computational complexity. Simulation results show that the proposed learning framework can learn optimized active BS beamforming and RIS reflection codebook. For instance, the proposed MA-DRL approach with only 6 beams outperforms a 256-beam discrete Fourier transform (DFT) codebook with a 97% beam training overhead reduction.
Asmaa Abdallah, Abdulkadir Celik, Mohammad M. Mansour, Ahmed M. Eltawil
CCNC4
2023 Unsupervised Learning - Based Downlink Power Allocation for CF-mMIMO Networks
abstract
Cell-free massive MIMO (CF-mMIMO) is a transformative wireless network technology that surmounts conventional cellular network limitations concerning coverage, capacity, and interference management. Despite offering numerous benefits, CF-mMIMO also presents significant challenges, particularly in signal processing and power allocation. This paper introduces an unsupervised learning framework for downlink (DL) power allocation in CF-mMIMO networks, utilizing only large scaling fading coefficients instead of the hard-to-obtain exact user equipment (UE) locations or channel state information. We consider the sum spectral efficiency (sum-SE) optimization objective and investigate two distinct precoding schemes-maximum ratio (MR) and regularized zero-forcing (RZF)-for multi-antenna access points (APs). A custom loss function is formulated to maximize the sum-SE at each UE while accounting for pilot contamination and ensuring that power budget constraints are satisfied at each AP. The proposed unsupervised learning approach circumvents the arduous task of training data computations typically required in supervised learning methods, bypassing the use of conventional complex optimization methods and heuristic methodologies. The simulation results demonstrate that the proposed unsupervised learning approach outperforms existing methods in terms of SE, showcasing an improvement up to 20%. The proposed unsupervised neural network also approximates the optimal solutions generated by convex solvers while significantly reducing computational complexity.
Mattia Fabiani, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil
GLOBECOM4
2023 On the Optimization of Virtual RIS Partitioning for Grant-Free Non-Orthogonal Multiple Access
abstract
The integration of reconfigurable intelligent surfaces (RISs) and grant-free non-orthogonal multiple access (GF-NOMA) has emerged as a promising solution for enhancing spec-tral efficiency (SE) and massive connectivity in future wireless networks. This paper proposes a novel virtual RIS partitioning mechanism for GF-NOMA, where all user equipments (UEs) within a specific NOMA cluster have instantaneous access to resource blocks (RBs) without the need for grant acquisition and power control as in the traditional grant-based NOMA schemes. To achieve the required power disparity, RIS portions are allocated to the UEs in a manner that increases the reception power disparity. We derive closed-form equations for optimal RIS portions in two regimes of interest: 1) max-min fair regime and 2) maximum throughput regime. Simulation results demonstrate that the proposed RIS-assisted GF-NOMA yields a gain of 28% and 15% in terms of max-sum rate and max-min rate, respectively, outperforming existing grant-based approaches. The study highlights the potential of combining RIS with GF-NOMA as a powerful solution for future wireless communication systems.
Mohd Hamza Naim Shaikh, Abdulkadir Celik, Ahmed M. Eltawil, Galymzhan Nauryzbayev
GLOBECOM3
2023 High-Density FeFET-based CAM Cell Design Via Multi-Dimensional Encoding
abstract
Content addressable memory is one of the most frequently used technologies in Data-centric applications due to its exceptional search parallelism capability. SRAM cells were initially used to implement CAM designs. Recent innovations proposed using compact nonvolatile memories instead. FeFETs emerged as a multi-level NVM device with promising potential and 2T FeFET CAM designs were studied. In this paper, a new potential is discussed for increasing the density efficiency of FeFET CAM architectures by adapting higher-dimensional encoding using 3T and 4T CAM designs. We propose a scalable greedy search algorithm for maximizing encoding capabilities. We compare the density, latency, accuracy, and energy consumption of our designs to standard 2T architecture demonstrating a 4x and 8x decrease in fail probability with up to 16% and 26.5% increase in memory density (bits/unit-area) in the 3T and 4T designs respectively.
Hadi Noureddine, Omar Bekdache, Mohamad Al Tawil, Rouwaida Kanj, Ali Chehab, Mohamed E. Fouda, Ahmed M. Eltawil
ACM Great Lakes Symposium on VLSI7
2023 RIS-Assisted Grant-Free NOMA
abstract
This paper introduces a reconfigurable intelligent surface (RIS)-assisted grant-free non-orthogonal multiple access (GF-NOMA) scheme. To ensure the power reception disparity required by the power domain NOMA (PD-NOMA), we propose a joint user clustering and RIS assignment/alignment approach that maximizes the network sum rate by judiciously pairing user equipments (UEs) with distinct channel gains, assigning RISs to proper clusters, and aligning RIS phase shifts to the cluster members yielding the highest cluster sum rate. Once UEs are acknowledged with the cluster index, they are allowed to access their resource blocks (RBs) at any time requiring neither further grant acquisitions from the base station (BS) nor power control as all UEs are requested to transmit at the same power. In this way, the proposed approach performs an implicit over-the-air power control with minimal control signaling between the BS and UEs, which has shown to deliver up to 20% higher network sum rate than benchmark GF-NOMA and grant-based optimal (OPT) PD-NOMA schemes depending on the network parameters. The given numerical results also investigate the impact of UE density, RIS deployment, and RIS hardware specifications on the overall performance of the proposed RIS-aided GF-NOMA scheme.
Recep A. Tasci, Fatih Kilinc, Abdulkadir Celik, Asmaa Abdallah, Ahmed M. Eltawil, Ertugrul Basar
ICC5
2023 Low Precision Quantization-aware Training in Spiking Neural Networks with Differentiable Quantization Function
abstract
Deep neural networks have been proven to be highly effective tools in various domains, yet their computational and memory costs restrict them from being widely deployed on portable devices. The recent rapid increase of edge computing devices has led to an active search for techniques to address the abovementioned limitations of machine learning frameworks. The quantization of artificial neural networks (ANNs), which converts the full-precision synaptic weights into low-bit versions, emerged as one of the solutions. At the same time, spiking neural networks (SNNs) have become an attractive alternative to conventional ANNs due to their temporal information processing capability, energy efficiency, and high biological plausibility. Despite being driven by the same motivation, the simultaneous utilization of both concepts has yet to be thoroughly studied. Therefore, this work aims to bridge the gap between recent progress in quantized neural networks and SNNs. It presents an extensive study on the performance of the quantization function, represented as a linear combination of sigmoid functions, exploited in low-bit weight quantization in SNNs. The presented quantization function demonstrates the state-of-the-art performance on four popular benchmarks, CIFAR10-DVS, DVS128 Gesture, N-Caltech101, and N-MNIST, for binary networks (64.05%, 95.45%, 68.71%, and 99.43% respectively) with small accuracy drops and up to 31 × memory savings, which outperforms existing methods.
Ayan Shymyrbay, Mohamed E. Fouda, Ahmed M. Eltawil
IJCNN3
2023 Hardware Acceleration of DNA Pattern Matching with Binary Memristors
abstract
DNA pattern matching is a key technique applied in many bioinformatics applications. Recently, this technique has become very popular and is widely used for genetic disease diagnosis, where finding the number of consecutive repeats of a specific DNA pattern indicates the type and intensity of the patient's disorder. However, the remarkable growth of DNA data exacerbates the latency and power consumption required to perform DNA pattern matching. In this work, we propose a hardware accelerator design to detect the presence of different diseases efficiently using DNA pattern matching. We propose a novel CAM cell using binary memristors for reliable and robust data encoding. The proposed architecture consists of two main building blocks the Content-addressable memory (CAM) and pattern detector circuits in addition to the needed peripheral circuits for CAM read, write and match operation. CMOS PTM 45nm technology was used to design and simulate the full architecture. The evaluation of the proposed design shows$\sim 2\times$improvement in energy-delay-area product compared to the state-of-art work in the literature, in addition to robustness against noise and process variations.
Jinane Bazzi, Mohamed E. Fouda, Rouwaida Kanj, Ahmed M. Eltawil
ISCAS4
2023 Scalable Complementary FeFET CAM Design
abstract
CAMs are frequently employed for data-centric applications. They offer excellent parallelism. Traditionally, they were implemented using the area-consuming SRAM. Recent advancements suggest using compact nonvolatile memories (NVMs) to create CAM cells to reduce area. The ferroelectric field effect transistor (FeFET) has therefore emerged as an NVM device showing great potential in these memory architectures. In this work, we propose a novel multi-bit CAM architecture that utilizes p-type FeFETs – a topic yet to be explored in the literature – and we compare the latency, accuracy, and energy consumption of our design to other FeFET-based architectures demonstrating a 3-30× reduction in fail probability.
Omar Bekdache, Hadi Noureddine, Mohamad Al Tawil, Rouwaida Kanj, Mohamed E. Fouda, Ahmed M. Eltawil
ISCAS6
2023 Live Demonstration: Human Body Communication Health Monitoring System Using Flexible Substrate
abstract
We demonstrate the design and functionality of a flexible, miniaturized, ultra-low power, and affordable health monitoring system enabling continuous monitoring of individuals' health metrics, a.k.a. a wireless body area network (WBAN). To date, the most commonly-used means of communication for WBAN modules has been based on Radio Frequency (RF) communications. Though they significantly facilitated human healthcare monitoring, they require complex, power-hungry, RF front ends. As an alternative, we design our system to communicate utilizing human body communication (HBC), which has inherent physical layer security and enhanced overall energy efficiency. The live demo presents a vital signal monitoring system based on point-to-point HBC communication. Fig. 1 illustrates the transmitter and receiver diagrams and PCB boards respectively. The transmitter comprises a microcontroller, an oscillator, an on-off-keying (OOK) modulator, and signal/ground electrodes. A 3.7 V lithium-ion battery with a 3.3 V output low dropout regulator (LDO) powers the whole system. The receiver first detects the envelope of the received signal and slices it to binary digits by comparing the input signal with its average level extracted by low-pass filtering. Interested readers can find full details at [1], where the system has shown an energy efficiency of 8.3 nJ/b at a data rate of up to 1.3 Mbps.
Qi Huang 0002, Abeer Alamoudi, Abdulkadir Celik, Ahmed M. Eltawil
ISCAS4
2023 Cooperative Body Channel Communications for Energy-Efficient Internet of Bodies
abstract
The Internet of Bodies (IoB) is a network formed by wearable, implantable, ingestible, and injectable smart devices to collect physiological, behavioral, and structural information from the human body. Thus, the IoB technology can revolutionize the quality of human life by using these context-rich data in myriad smart-health applications. Radio frequency (RF) transceivers have been typically preferred due to their availability and maturity. However, for most RF standards (e.g., Bluetooth low energy), the highly radiative omnidirectional RF propagation (even at the lowest settings) reaches tens of meters of coverage, thereby reducing energy efficiency, causing interference and co-existence issues, and raising privacy and security concerns. On the other hand, body channel communication (BCC) confines low-power and low-frequency (10 kHz–100 MHz) signals to the human body, leading to more secure and efficient communications. Since energy efficiency is one of the critical design parameters of IoB networks, this article focuses on energy-efficient orthogonal body channel access (OBA) and non-OBA (NOBA) schemes with and without cooperation. To this aim, three main BCC topologies are presented: 1) point-to-point channel; 2) medium access channel; and 3) broadcast channel. These topologies are then used as building blocks to create IoB networks relying on OBA and NOBA schemes for downlink (DL) and uplink (UL) traffic. For all schemes and traffic directions, optimal transmit power and phase time allocations are derived in closed-form, which is essential to reduce energy consumption by eliminating computational power. The closed-form expressions are further leveraged to obtain maximum network size as a function of data rate requirement, bandwidth, and hardware parameters.
Abeer Alamoudi, Abdulkadir Celik, Ahmed M. Eltawil
IEEE Internet Things J.3
2023 High-Throughput Independent Component Analysis Processor for Full Duplex Systems
abstract
This paper presents the algorithm and very-large-scale integration (VLSI) architecture of a high-throughput and highly efficient independent component analysis (ICA) processor for self-interference cancellation (SIC) in in-band full-duplex (IBFD) systems. This is the first VLSI architecture reported in the literature based on the state-of-the-art entropy bound minimization (EBM) approach. A novel ICA algorithm is presented in this paper with momentum gradient descent optimization. Simulation results show that the number of iterations for the proposed algorithm is significantly reduced compared to the conventional ICA algorithms. Furthermore, a novel early-distribution estimation scheme is proposed in the designed ICA processor to compute multiple distribution functions with low latency and low complexity. The processing flow and the efficiency for the hardware utilization are specifically designed so that the processing speed is maximized with minimum employment of hardware components. The proposed ICA processor is designed and implemented based on the application-specific-integrated circuit (ASIC) flow. The post-layout estimations show that compared with the conventional EBM-based scheme, the proposed design improves the throughput and efficiency by 30x. In addition, compared to prior designs shown in the literature, the proposed ICA processor also demonstrates a significant enhancement in terms of throughput and efficiency.
Jen-Hao Cheng, Tien-Min Chang, Chung-An Shen, Mohamed E. Fouda, Ahmed M. Eltawil
IEEE J. Sel. Areas Commun.5
2023 Resistive Neural Hardware Accelerators
abstract
Deep neural networks (DNNs), as a subset of machine learning (ML) techniques, entail that real-world data can be learned, and decisions can be made in real time. However, their wide adoption is hindered by a number of software and hardware limitations. The existing general-purpose hardware platforms used to accelerate DNNs are facing new challenges associated with the growing amount of data and are exponentially increasing the complexity of computations. Emerging nonvolatile memory (NVM) devices and the compute-in-memory (CIM) paradigm are creating a new hardware architecture generation with increased computing and storage capabilities. In particular, the shift toward resistive random access memory (ReRAM)-based in-memory computing has great potential in the implementation of area- and power-efficient inference and in training large-scale neural network architectures. These can accelerate the process of IoT-enabled AI technologies entering our daily lives. In this survey, we review the state-of-the-art ReRAM-based DNN many-core accelerators, and their superiority compared to CMOS counterparts was shown. The review covers different aspects of hardware and software realization of DNN accelerators, their present limitations, and prospects. In particular, a comparison of the accelerators shows the need for the introduction of new performance metrics and benchmarking standards. In addition, the major concerns regarding the efficient design of accelerators include a lack of accuracy in simulation tools for software and hardware codesign.
Kamilya Smagulova, Mohamed E. Fouda, Fadi J. Kurdahi, Khaled N. Salama, Ahmed M. Eltawil
Proc. IEEE5
2023 Offline Training-Based Mitigation of IR Drop for ReRAM-Based Deep Neural Network Accelerators
abstract
Recently, resistive RAM (ReRAM)-based hardware accelerators showed unprecedented performance compared the digital accelerators. Technology scaling causes an inevitable increase in interconnect wire resistance, which leads to IR drops that could limit the performance of ReRAM-based accelerators. These IR drops deteriorate the signal integrity and quality, especially in the crossbar structures which are used to build high-density ReRAMs. Hence, finding a software solution, which can predict the effect of IR drop without involving expensive hardware or SPICE simulations, is very desirable. In this article, we propose two neural networks models to predict the impact of the IR drop problem. These models are used to evaluate the performance of the different deep neural network (DNN) models including binary and quantized neural networks showing similar performance (i.e., recognition accuracy) to the golden validation (i.e., SPICE-based DNN validation). In addition, these predication models are incorporated into the DNN training framework to efficiently retrain the DNN models and bridge the accuracy drop. To further enhance the validation accuracy, we propose incremental training methods. The DNN validation results, done through SPICE simulations, show very high improvement in performance close to the baseline performance, which demonstrates the efficacy of the proposed method even with challenging datasets, such as CIFAR10 and SVHN.
Sugil Lee, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 Training-Free Stuck-At Fault Mitigation for ReRAM-Based Deep Learning Accelerators
abstract
Although Resistive RAMs can support highly efficient matrix–vector multiplication, which is very useful for machine learning and other applications, the nonideal behavior of hardware, such as stuck-at fault (SAF) and IR drop is an important concern in making ReRAM crossbar array-based deep learning accelerators. Previous work has addressed the nonideality problem through either redundancy in hardware, which requires a permanent increase of hardware cost, or software retraining, which may be even more costly or unacceptable due to its need for a training dataset as well as high computation overhead. In this article, we propose a very lightweight method that can be applied on top of existing hardware or software solutions. Our method, called forward-parameter tuning (FPT), takes advantage of a certain statistical property existing in the activation data of neural network layers, and can mitigate the impact of mild nonidealities in ReRAM crossbar arrays (RCAs) for deep learning applications without using any hardware, a dataset, or gradient-based training. Our experimental results using MNIST, CIFAR-10, and CIFAR-100, and ImageNet datasets in binary and multibit networks demonstrate that our technique is very effective, both alone and together with previous methods, up to 20% fault rate, which is higher than even some of the previous remapping methods. We also evaluate our method in the presence of other nonidealities, such as variability and IR drop. Furthermore, we provide an analysis based on the concept of the effective fault rate (EFR), which not only demonstrates that EFR can be a useful tool to predict the accuracy of faulty RCA-based neural networks but also explains why mitigating the SAF problem is more difficult with multibit neural networks.
Chenghao Quan, Mohamed E. Fouda, Sugil Lee, Giju Jung, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2023 Architectural Trade-Off Analysis for Accelerating LSTM Network Using Radix-r OBC Scheme
abstract
This paper presents architectural trade-off analysis for accelerating two (Type I, II) fixed-point long short-term memory (LSTM) network based on circulant matrix-vector multiplications (MVMs) using radix-$r$offset binary coding (OBC) scheme. Type I MVM architecture rotates the weights with the proposed modulo-cum interleaver and uses partial product generators (PPGs) with a single generation unit across a column. It is hardware-optimized using a single adder tree through time-multiplexing. Meanwhile, Type II MVM architecture rotates the inputs with the proposed store-cum interleaver and uses single PPGs with a single generation unit across a row. It is time-optimized by unfolding shift-accumulate unit to a shift-add tree followed by pipelining. A new design for element-wise multiplication using radix-$r$PPG is also presented. Both the designs are extended to their block-circulant variants for certain accuracy requirements. Post-synthesis of Type I and II architectures for a different model, kernel, radix sizes and clock frequencies result in several efficient designs. Compared with the prior scheme, Type I architecture for$128 \times 128$with$r=2$on 28 nm FDSOI technology at 800 MHz occupies 32.27% lesser area, consumes 67.89% lesser power at the same throughput, while Type II architecture at the expense of area and power provides$40\times $higher throughput.
Mohd. Tasleem Khan, Hasan Erdem Yantir, Khaled N. Salama, Ahmed M. Eltawil
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 RIS-Aided mmWave MIMO Channel Estimation Using Deep Learning and Compressive Sensing
abstract
Reconfigurable intelligent surface (RIS) assisted wireless systems require accurate channel state information (CSI) to control wireless channels and improve both the bandwidth and energy efficiency. However, CSI acquisition is non-trivial for two reasons: 1) the passive nature of RIS does not allow transceiving and processing pilot signals, and 2) the dimensions of the cascaded channel between transceivers increases with the large number of RIS elements, which yields high training overhead and computational complexity. While prior art has mainly focused on frequency-flat channel estimation, this paper proposes novel data-driven and compressive sensing based approaches for estimating both frequency-flat and frequency-selective cascaded channels of RIS-assisted multi-user millimeter-wave large multiple input multiple output (MIMO) systems with limited training overhead. The proposed methods exploit the common sparsity property among the different subcarriers and the double-structured sparsity property of the angular cascaded channel matrices as different angular cascaded channels observed by different users share completely common non-zero rows and user-specific column supports. The proposed data-driven cascaded channel estimation approaches use denoising neural networks to accurately detect channel supports. Alternatively, when data-training capabilities are not available, the compressive sensing based orthogonal matching pursuit (OMP) approach relies on sparsity properties and applies simultaneous OMP to detect the channel supports. Simulation results show that the pilot overhead required by the proposed scheme is lower than existing schemes. When compared to other OMP approaches that achieve an NMSE gap of 5 to 6 dB with respect to the Oracle least square lower bound, the proposed algorithms reduce the lower bound gap to only 1 dB, while reducing complexity by more than two orders of magnitude.
Asmaa Abdallah, Abdulkadir Celik, Mohammad M. Mansour, Ahmed M. Eltawil
IEEE Trans. Wirel. Commun.4
2022 Wearable Vital Signal Monitoring Prototype Based on Capacitive Body Channel Communication
abstract
Wireless body area network (WBAN) provides a means for seamless individual health monitoring without imposing restrictive limitations on normal daily routines. To date, Radio Frequency (RF) transceivers have been the technology of choice, however, drawbacks such as vulnerability to body shadowing effects, higher power consumption due to omnidirectional radiation and security concerns, have prompted the adoption of transceivers that use the human body channel for communication. In this paper, a vital signal monitoring transceiver prototype based on the human body channel communication (HBC), using commercially available chipsets is presented. RF and HBC communications are briefly reviewed and compared, and different schemes of HBC are introduced. A circuit model that represents the human body channel is then discussed and simulations are presented to illustrate the influence of the return path capacitance and receiver terminations on the path loss. The architecture of the transceiver prototype is then introduced where it is designed at a 21 MHz IEEE 802.15.6 standard-compliant carrier frequency. Finally, the performance of the transceiver, including the bit error rate (BER) and power efficiency, are characterized. Path loss is measured for two different scenarios, where variations of up to 5 dB were observed due to environmental effects. Energy efficiency measured at a maximum data-rate of 1.3 Mbps was found to be 8.3 nJ/b.
Qi Huang 0002, Waseem Alkhayer, Mohamed E. Fouda, Abdulkadir Celik, Ahmed M. Eltawil
BSN5
2022 Deep-Learning Based Channel Estimation for RIS-Aided mmWave Systems with Beam Squint
abstract
Reconfigurable intelligent surface (RIS) assisted wireless systems require accurate channel state information (CSI) to control wireless channels and improve overall network performance. However, CSI acquisition is non-trivial due to the passive nature of RIS, and the dimensions of the cascaded channel between transceivers increase with the large number of RIS elements, which requires high training overhead. Prior art has considered frequency-selective channel estimation without considering the beam squint effect in wideband systems, severely degrading channel estimation performance. This paper proposes a novel data-driven approach for estimating wideband cascaded channels of RIS-assisted multi-user millimeter-wave massive multiple-input multiple-output (MIMO) systems with limited training overhead, explicitly considering the effect of beam squint. To circumvent the beam squint effect, the proposed method exploits the common sparsity property among the different subcarriers as well as the double-structured sparsity property of the users’ angular cascaded channel matrices. The proposed data-driven cascaded channel estimation approach exploits denoising neural networks to detect channel supports accurately. Compared to beam squint effect agnostic traditional orthogonal matching pursuit (OMP) approaches, the proposed data-driven approach achieves 5-6dB less normalized mean square error (NMSE) and reduces the lower bound gap to only 1dB for the oracle least-square benchmark.
Asmaa Abdallah, Abdulkadir Celik, Mohammad M. Mansour, Ahmed M. Eltawil
ICC4
2022 Accurate Prediction of ReRAM Crossbar Performance Under I-V Nonlinearity and IR Drop
abstract
Despite the promise of extremely efficient matrix-vector multiplication (MVM) by ReRAM crossbar arrays (RCAs), maintaining high accuracy has been challenging due to nonidealities such as wire resistance (also known as IR drop) and I-V nonlinearity (i.e., voltage-dependent conductance). For system architects, a fast method to accurately predict the MVM output of an RCA under nonidealities is highly desirable. While IR drop alone without I-V nonlinearity can be efficiently predicted, the existence of I-V nonlinearity makes the problem much harder. In this paper we propose a novel algorithm based on iterative refinement, which can predict with high accuracy the outcome of an MVM operation on an RCA in the presence of both I-V nonlinearity and IR drop. Our experiments using binary RCAs of different sizes demonstrate that our proposed method is order-of-magnitude more accurate than previous methods in terms of RMS error. We also present case studies predicting hardware-realistic accuracy of binarized neural networks on RCAs as well as nonideality-aware retraining, demonstrating the efficacy of our method for early design space exploration of ReRAM-based accelerators.
Sugil Lee, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi
ICCD4
2022 Enhancing Physical Layer Security in Large Intelligent Surface-aided Cooperative Networks
abstract
Intelligent surfaces have recently been presented as a revolutionary technique and recognized as one of the candidates for beyond fifth-generation wireless networks. This paper investigates the physical layer security of a large intelligent surface (LIS) aided wireless system over Nakagami-m channels. We propose a phase-based adaptive modulation scheme, where LIS’s phase-shift optimization process is effectively utilized to enhance the system’s security. Moreover, the effect of the Nakagami-m fading parameter (m), correlation parameter ($\rho$), and a number of passive LIS elements (M) on the system performance are examined. The significant improvement in confidentiality is shown while evaluating the bit error rate performance of the proposed scheme.
Madi Makin, Sultangali Arzykulov, Abdulkadir Celik, Ahmed M. Eltawil, Galymzhan Nauryzbayev
VTC Spring4
2022 Performance of RIS-empowered NOMA-based D2D Communication under Nakagami-m Fading
abstract
Reconfigurable intelligent surfaces (RISs) have sparked a renewed interest in the research community envisioning future wireless communication networks. In this study, we analyzed the performance of RIS-enabled non-orthogonal multiple access (NOMA) based device-to-device (D2D) wireless communication system, where the RIS is partitioned to serve a pair of D2D users. Specifically, closed-form expressions are derived for the upper and lower limits of spectral efficiency (SE) and energy efficiency (EE). In addition, the performance of the proposed NOMA-based system is also compared with its orthogonal counter-part. Extensive simulation is done to corroborate the analytical findings. The results demonstrate that RIS highly enhances the performance of a NOMA-based D2D network.
Mohd Hamza Naim Shaikh, Sultangali Arzykulov, Abdulkadir Celik, Ahmed M. Eltawil, Galymzhan Nauryzbayev
VTC Fall4
2022 Enhancing QoS Through Fluid Antenna Systems over Correlated Nakagami-m Fading Channels
abstract
Fluid antenna systems (FAS) enable mechanically flexible antennas that offer adaptability and flexibility for modern communication devices. In this work, we present a conceptual model for a single-antenna N-port (SANP) FAS over spatially correlated Nakagami-m fading channels and compare it with the traditional diversity schemes in terms of outage probability. The proposed FAS model switches to the best antenna port and resembles the operation of a selection combining (SC) diversity. FAS improves the quality of service (QoS) of the network through antenna port selection. The advantage of FAS is the ability to fit hundreds of antenna ports into a half-wavelength antenna size at the cost of spatial channel correlation. Simulation results demonstrate the superior outage probability performance of FAS at several tens of antenna ports compared to the traditional diversity schemes such as maximum ratio combining, equal gain combining, and SC. Moreover, the novel probability and cumulative density functions for the land mobile correlated Nakagami-m random variates are evaluated in this paper.
Leila Tlebaldiyeva, Galymzhan Nauryzbayev, Sultangali Arzykulov, Ahmed M. Eltawil, Theodoros A. Tsiftsis
WCNC4
2022 Enabling the Internet of Bodies Through Capacitive Body Channel Access Schemes
abstract
The Internet of Bodies (IoB) is an imminent extension of the vast Internet of Things (IoT) domain, where wearable, ingestible, injectable, and implantable smart objects form a network in, on, and around the human body. The highly radiative nature of radio-frequency (RF) IoB devices unnecessarily extends the coverage range beyond the human body, which reduces energy efficiency, causes co-existence and interference issues, and exposes sensitive personal data to security threats. Alternatively, capacitive body channel communication (BCC) confine signal transmission to the human body to reduce signal leakage, experience less propagation loss, and reach pJ/b energy efficiency levels. Therefore, capacitive BCC is a key enabler to reach the ultimate design goals of ultra low power, high throughput, and small form-factor IoB devices. Albeit these attractive features, the communication and networking aspects of the capacitive BCC are not thoroughly explored yet. Therefore, this article proposes orthogonal and nonorthogonal capacitive body channel access schemes with or without cooperation among the IoB nodes. In order to address the Quality of Service (QoS) demand scenarios of different IoB applications, we present and formulate the max–min rate, max-sum rate, and QoS sufficient operational regimes, and then provide closed-form and numerical solution optimal power and phase time allocations. Extensive numerical results are analyzed to compare the performance of orthogonal and nonorthogonal schemes with and without cooperation for various design parameters under prescribed QoS regimes. The obtained results show that capacitive body channel access schemes can provide several Mb/s rates even at low transmission powers ranging between −60 and −90 dBm. Moreover, the cooperative schemes are shown to be effective to avoid performance degradation caused by increasing network size, low transmission power, and poor channel quality.
Abdulkadir Celik, Ahmed M. Eltawil
IEEE Internet Things J.2
2022 The Internet of Bodies: A Systematic Survey on Propagation Characterization and Channel Modeling
abstract
The Internet of Bodies (IoBs) is an imminent extension to the vast Internet of Things domain, where interconnected devices (e.g., worn, implanted, embedded, swallowed, etc.) are located in-on-and-around the human body form a network. Thus, the IoB can enable a myriad of services and applications for a wide range of sectors, including medicine, safety, security, wellness, entertainment, to name but a few. Especially, considering the recent health and economic crisis caused by the novel coronavirus pandemic, also known as COVID-19, the IoB can revolutionize today’s public health and safety infrastructure. Nonetheless, reaping the full benefit of IoB is still subject to addressing related risks, concerns, and challenges. Hence, this survey first outlines the IoB requirements and related communication and networking standards. Considering the lossy and heterogeneous dielectric properties of the human body, one of the major technical challenges is characterizing the behavior of the communication links in-on-and-around the human body. Therefore, this article presents a systematic survey of channel modeling issues for various link types of human body communication (HBC) channels below 100 MHz, the narrowband (NB) channels between 400 and 2.5 GHz, and ultrawideband (UWB) channels from 3 to 10 GHz. After explaining bio-electromagnetics attributes of the human body, physical, and numerical body phantoms are presented along with electromagnetic propagation tool models. Then, the first-order and the second-order channel statistics for NB and UWB channels are covered with a special emphasis on body posture, mobility, and antenna effects. For capacitively, galvanically, and magnetically coupled HBC channels, four different channel modeling methods (i.e., analytical, numerical, circuit, and empirical) are investigated, and electrode effects are discussed. Finally, interested readers are provided with open research challenges and potential future research directions.
Abdulkadir Celik, Khaled N. Salama, Ahmed M. Eltawil
IEEE Internet Things J.3
2022 A hardware/software co-design methodology for in-memory processors
Hasan Erdem Yantir, Ahmed M. Eltawil, Khaled N. Salama
J. Parallel Distributed Comput.2
2022 Toward the Optimal Design and FPGA Implementation of Spiking Neural Networks
abstract
The performance of a biologically plausible spiking neural network (SNN) largely depends on the model parameters and neural dynamics. This article proposes a parameter optimization scheme for improving the performance of a biologically plausible SNN and a parallel on-field-programmable gate array (FPGA) online learning neuromorphic platform for the digital implementation based on two numerical methods, namely, the Euler and third-order Runge-Kutta (RK3) methods. The optimization scheme explores the impact of biological time constants on information transmission in the SNN and improves the convergence rate of the SNN on digit recognition with a suitable choice of the time constants. The parallel digital implementation leads to a significant speedup over software simulation on a general-purpose CPU. The parallel implementation with the Euler method enables around 180× ( 20× ) training (inference) speedup over a Pytorch-based SNN simulation on CPU. Moreover, compared with previous work, our parallel implementation shows more than 300× ( 240× ) improvement on speed and 180× ( 250× ) reduction in energy consumption for training (inference). In addition, due to the high-order accuracy, the RK3 method is demonstrated to gain 2× training speedup over the Euler method, which makes it suitable for online training in real-time applications.
Wenzhe Guo, Hasan Erdem Yantir, Mohamed E. Fouda, Ahmed M. Eltawil, Khaled N. Salama
IEEE Trans. Neural Networks Learn. Syst.4
2022 Efficient Neuromorphic Hardware Through Spiking Temporal Online Local Learning
abstract
Local learning schemes have shown promising performance in spiking neural networks (SNNs) training and are considered a step toward more biologically plausible learning. Despite many efforts to design high-performance neuromorphic systems, a fast and efficient on-chip training algorithm is still missing, which limits the deployment of neuromorphic systems in many real-time applications. This work proposes a scalable, fast, and efficient spiking neuromorphic hardware system with on-chip local learning capability. We introduce an effective hardware-friendly local training algorithm compatible with sparse temporal input coding and binary random classification weights. The algorithm is demonstrated to deliver competitive accuracy in different tasks. The proposed digital system explores spike sparsity in communication, parallelism in vector–matrix operations and process-level dataflow, and locality of training errors, which leads to low cost and fast training speed. The system is optimized under various performance metrics. Taking into consideration energy, speed, resources, and accuracy, the proposed method shows around$10\times $efficiency over a recent work with a direct feedback alignment (DFA) method and$4.5\times $efficiency over the spike-timing-dependent plasticity (STDP) method. Moreover, our hardware architecture can easily scale up with the network size at a linear rate. Thus, our method has demonstrated great potential for use in various applications, especially those demanding low latency.
Wenzhe Guo, Mohamed E. Fouda, Ahmed M. Eltawil, Khaled N. Salama
IEEE Trans. Very Large Scale Integr. Syst.3
2022 Configurable Independent Component Analysis Preprocessing Accelerator
abstract
An independent component analysis (ICA) has been used in many applications, including self-interference cancellation (SIC) for in-band full-duplex (IBFD) wireless systems and anomaly detection in industrial Internet of Things (IoT). This article presents a high-throughput and highly efficient configurable preprocessing accelerator for the ICA algorithm. The proposed ICA accelerator has three major blocks that perform data centering, covariance matrix for computation, and eigenvalue decomposition (EVD). Specifically, the proposed accelerator is based on a high-performance matrix multiplication array (MMA). The proposed MMA architecture uses time-multiplexed processing, so that the efficiency of hardware utilization is greatly enhanced. Furthermore, the processing flow utilizes parallel processing, such that the centering, the calculation of the covariance matrix, and the EVD are conducted simultaneously and are individually pipelined to maximize throughput. This article presents the architecture, circuit design, and performance estimates based on post-layout extraction of the proposed preprocessing ICA accelerator. The proposed design achieves a throughput of 40.7 kMatrices/s at a complexity of 73.3 kGE.
Hsi-Hung Lu, Chung-An Shen, Mohamed E. Fouda, Ahmed M. Eltawil
IEEE Trans. Very Large Scale Integr. Syst.4
2022 Deep Learning-Based Frequency-Selective Channel Estimation for Hybrid mmWave MIMO Systems
abstract
Millimeter wave (mmWave) massive multiple-input multiple-output (MIMO) systems typically employ hybrid mixed signal processing to avoid expensive hardware and high training overheads. However, the lack of fully digital beamforming at mmWave bands imposes additional challenges in channel estimation. Prior art on hybrid architectures has mainly focused on greedy optimization algorithms to estimate frequency-flat narrowband mmWave channels, despite the fact that in practice, the large bandwidth associated with mmWave channels results in frequency-selective channels. In this paper, we consider a frequency-selective wideband mmWave system and propose two deep learning (DL) compressive sensing (CS) based algorithms for channel estimation. The proposed algorithms learn critical apriori information from training data to provide highly accurate channel estimates with low training overhead. In the first approach, a DL-CS based algorithm simultaneously estimates the channel supports in the frequency domain, which are then used for channel reconstruction. The second approach exploits the estimated supports to apply a low-complexity multi-resolution fine-tuning method to further enhance the estimation performance. Simulation results demonstrate that the proposed DL-based schemes significantly outperform conventional orthogonal matching pursuit (OMP) techniques in terms of the normalized mean-squared error (NMSE), computational complexity, and spectral efficiency, particularly in the low signal-to-noise ratio regime. When compared to OMP approaches that achieve an NMSE gap of$\mathrm {\{4-10\}\,\,dB}$with respect to the Cramer Rao Lower Bound (CRLB), the proposed algorithms reduce the CRLB gap to only$\mathrm {\{1-1.5\}\,\,dB}$, while reducing complexity by two orders of magnitude.
Asmaa Abdallah, Abdulkadir Celik, Mohammad M. Mansour, Ahmed M. Eltawil
IEEE Trans. Wirel. Commun.4
2021 Cost- and Dataset-free Stuck-at Fault Mitigation for ReRAM-based Deep Learning Accelerators
abstract
Resistive RAMs can implement extremely efficient matrix vector multiplication, drawing much attention for deep learning accelerator research. However, high fault rate is one of the fundamental challenges of ReRAM crossbar array-based deep learning accelerators. In this paper we propose a dataset-free, cost-free method to mitigate the impact of stuck-at faults in ReRAM crossbar arrays for deep learning applications. Our technique exploits the statistical properties of deep learning applications, hence complementary to previous hardware or algorithmic methods. Our experimental results using MNIST and CIFAR-10 datasets in binary networks demonstrate that our technique is very effective, both alone and together with previous methods, up to 20 % fault rate, which is higher than the previous remapping methods. We also evaluate our method in the presence of other non-idealities such as variability and IR drop.
Giju Jung, Mohamed E. Fouda, Sugil Lee, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi
DATE5
2021 Energy Efficient Capacitive Body Channel Access Schemes for Internet of Bodies
abstract
The Internet of bodies is a network of wearable, ingestible, injectable, and implantable smart objects located in, on, and around the body. Although radio frequency (RF) systems are considered the default choice for implementing on-body communications, which need to be localized in the vicinity of the human body (typically < 5 cm), highly radiative RF propagations unnecessarily extend several meters beyond the human body. This intuitively degrades energy efficiency, leads to interference and co-existence issues, and exposes sensitive personal data to security threats. As an alternative, the capacitive body channel communication (BCC) couples the signal (between 10 kHz-100 MHz) to the human body, which is more conductive than air. Hence, BCC provides a lower propagation loss, better physical layer security, and nJ/bit to pJ/bit energy efficiency. Accordingly, this paper investigates orthogonal and non-orthogonal capacitive body channel access schemes for ultra-low-power IoB nodes. We present the optimal uplink and downlink power allocations in closed-form, which deliver better fairness and network lifetime than benchmark numerical solvers. For a given bandwidth and data rate requirement, we also derive the maximum affordable number of IoB nodes for both directions of orthogonal and non-orthogonal schemes.
Abeer Alamoudi, Abdulkadir Celik, Ahmed M. Eltawil
GLOBECOM3
2021 Coverage Analysis of CR-based Satellite-Terrestrial NOMA Networks with Practical System Impairments
abstract
In this paper, we investigate a non-orthogonal multiple access (NOMA) assisted cognitive satellite-terrestrial network under practical system conditions, such as transceiver hard-ware impairments, channel state information mismatch, imperfect successive interference cancellation and interference noises. Generalized coverage probability formulas for NOMA users in both primary and secondary networks are derived considering the impact of interference temperature constraint. Moreover, the numerical results demonstrate superior outperformance compared to the ones obtained for an orthogonal multiple access scheme. Finally, the derived analytical findings are fully supported by thorough Monte Carlo simulations.
Yerassyl Akhmetkaziyev, Galymzhan Nauryzbayev, Sultangali Arzykulov, Ahmed M. Eltawil, Theodoros A. Tsiftsis
ICC4
2021 Fast and Low-Cost Mitigation of ReRAM Variability for Deep Learning Applications
abstract
To overcome the programming variability (PV) of ReRAM crossbar arrays (RCAs), the most common method is program-verify, which, however, has high energy and latency overhead. In this paper we propose a very fast and low-cost method to mitigate the effect of PV and other variability for RCA-based DNN (Deep Neural Network) accelerators. Leveraging the statistical properties of DNN output, our method called Online Batch-Norm Correction (OBNC) can compensate for the effect of programming and other variability on RCA output without using on-chip training or an iterative procedure, and is thus very fast. Also our method does not require a nonideality model or a training dataset, hence very easy to apply. Our experimental results using ternary neural networks with binary and 4-bit activations demonstrate that our OBNC can recover the baseline performance in many variability settings and that our method outperforms a previously known method (VCAM) by large margins when input distribution is asymmetric or activation is multi-bit.
Sugil Lee, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi
ICCD4
2021 Underlay Hybrid Satellite-Terrestrial Relay Networks under Realistic Hardware and Channel Conditions
abstract
In this paper, we study a cognitive hybrid satellite-terrestrial relay network with imposed practical limitations, such as channel state information mismatch and transceiver-induced hardware impairments. Moreover, it is assumed that the network undergoes multiple independent and non-identically distributed interference noises arising from neighboring transmitters. Generalized closed-form expressions of the outage probability for a terrestrial user are obtained while taking into account the effect of an interference temperature constraint. Finally, analytical derivations are verified through Monte Carlo simulation and the impact of impairments is examined.
Yerassyl Akhmetkaziyev, Galymzhan Nauryzbayev, Sultangali Arzykulov, Khaled M. Rabie, Xingwang Li 0001, Ahmed M. Eltawil
VTC Fall6
2021 Outage Analysis of EH-based Cooperative NOMA Networks over Generalized Statistical Models
abstract
In this paper, the outage probability (OP) of the two-hop cooperative non-orthogonal multiple access network with an energy-constrained cooperative agent is evaluated over generalized α-µ and κ-µ fading models. The power splitting relaying protocol is implemented at the cooperative agent, acting as a relay in amplify-and-forward mode. The effect of hardware impairments (HIs) at the radio frequency front-ends of transceivers are incorporated during the performance evaluation. The results based on the derived analytical expressions suggest the importance of considering HIs for given network architecture and allow one to evaluate the OP over various statistical models.
Orken Omarov, Galymzhan Nauryzbayev, Sultangali Arzykulov, Ahmed M. Eltawil, Mohammad S. Hashmi
VTC Spring4
2021 Cognitive Non-ideal NOMA Satellite-Terrestrial Networks with Channel and Hardware Imperfections
abstract
This paper investigates a non-orthogonal multiple access (NOMA) assisted cognitive satellite-terrestrial network which is practically limited by interference noises, transceiver hardware impairments, imperfect successive interference cancellation, and channel state information mismatch. Generalized outage probability expressions for NOMA users in both primary and secondary networks are derived considering the impact of interference temperature constraint. Finally, obtained results are corroborated by Monte Carlo simulations and compared with the orthogonal multiple access to show the superior performance of the proposed network model.
Yerassyl Akhmetkaziyev, Galymzhan Nauryzbayev, Sultangali Arzykulov, Ahmed M. Eltawil, Khaled M. Rabie
WCNC4
2021 Capacity Analysis of Wireless Powered Cooperative NOMA Networks over Generalized Fading
abstract
This paper provides a performance evaluation of two-hop non-orthogonal multiple access (NOMA) architecture with energy harvesting (EH) cooperative agent and channel gains following the k-m fading, through analysis of ergodic capacity with respect to hardware impairments and channel conditions. The performance results of the distant user over two EH protocols, namely power splitting and time-switching relaying, are obtained and compared with simulation outcomes. The developed framework allows to evaluate the network under a range of external conditions and infers the importance of considering the hardware impairments.
Orken Omarov, Galymzhan Nauryzbayev, Sultangali Arzykulov, Mohammad S. Hashmi, Ahmed M. Eltawil
WCNC5
2021 IMCA: An Efficient In-Memory Convolution Accelerator
abstract
Traditional convolutional neural network (CNN) architectures suffer from two bottlenecks: computational complexity and memory access cost. In this study, an efficient in-memory convolution accelerator (IMCA) is proposed based on associative in-memory processing to alleviate these two problems directly. In the IMCA, the convolution operations are directly performed inside the memory as in-place operations. The proposed memory computational structure allows for a significant improvement in computational metrics, namely, TOPS/W. Furthermore, due to its unconventional computation style, the IMCA can take advantage of many potential opportunities, such as constant multiplication, bit-level sparsity, and dynamic approximate computing, which, while supported by traditional architectures, require extra overhead to exploit, thus reducing any potential gains. The proposed accelerator architecture exhibits a significant efficiency in terms of area and performance, achieving around 0.65 GOPS and 1.64 TOPS/W at 16-bit fixed-point precision with an area less than 0.25 mm2.
Hasan Erdem Yantir, Ahmed M. Eltawil, Khaled N. Salama
IEEE Trans. Very Large Scale Integr. Syst.2
2020 Learning to Predict IR Drop with Effective Training for ReRAM-based Neural Network Hardware
abstract
Due to the inevitability of the IR drop problem in passive ReRAM crossbar arrays, finding a software solution that can predict the effect of IR drop without the need of expensive SPICE simulations, is very desirable. In this paper, two simple neural networks are proposed as software solution to predict the effect of IR drop. These networks can be easily integrated in any deep neural network framework to incorporate the IR drop problem during training. As an example, the proposed solution is integrated in BinaryNet framework and the test validation results, done through SPICE simulations, show very high improvement in performance close to the baseline performance, which demonstrates the efficacy of the proposed method. In addition, the proposed solution outperforms the prior work on challenging datasets such as CIFAR10 and SVHN.
Sugil Lee, Giju Jung, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi
DAC5
2020 Throughput Characterization for Bluetooth Low Energy with Applications in Body Area Networks
abstract
Bluetooth Low Energy (BLE) has emerged as a technology of choice in many applications including body area networks (BAN). In this paper, we investigate the use of BLE in terms of throughput, power consumption and latency and evaluate its suitability for BAN applications. We compare the performance of different versions of the Bluetooth core specification using a theoretical model and an experimental setup based on nRF52840 chip by Nordic Semiconductor. We focus on Electrocardiography (EKG) and give the current consumption and battery lifetime estimation of an EKG BLE node for different BLE versions and configurations.
Michael A. Ayoub, Ahmed M. Eltawil
ISCAS2
2020 NOMA/OMA Mode Selection and Resource Allocation for Beyond 5G Networks
abstract
This paper considers hybridization of non-orthogonal multiple access (NOMA) and Orthogonal multiple access (OMA) schemes for next-generation cellular networks. The proposed hybrid multiple access (HMA) scheme considers a multi-cell environment that combines NOMA/OMA mode selection as well as channel and power allocation to improve the resource utilization and bandwidth efficiency. The NOMA/OMA modes are categorized into intra-cell and inter-cell OMA and NOMA modes based on an interference map. The HMA focuses on determining the best mode of operation between user pairs to improve the overall sum rate and quality of service (QoS). Results show that the proposed NOMA/OMA mode selection provides superior performance to the conventional OMA schemes without compromising the QoS demands.
Aysha Ebrahim, Abdulkadir Celik, Emad Alsusa, Ahmed M. Eltawil
PIMRC4
2020 A Machine Learning Approach for Structural Health Monitoring Using Noisy Data Sets
abstract
Continuous structural health monitoring of civil infrastructure can be achieved by deploying an Internet of Things network of distributed acceleration sensors in buildings to capture floor movement. Postdisaster damage levels can be computed based on the peak relative floor displacement as specified in government standards. This article uses machine learning approaches to identify the status of buildings postevent based on accelerometer traces. Prior work in the field assumed the use of high-quality accelerometers for displacement estimation. In this article, we focus on using lower quality and cheaper accelerometers, while accounting for noise effects by incorporating noisy data sets in machine learning approaches for classification. A labeled acceleration data set of buildings response to earthquakes was created, where each sample is labeled with its corresponding damage severity. Sensor noise is included in the data set to model nonideal sensors. Classification performance of machine learning algorithms, such as support vector machine, K-nearest neighbor, and convolutional neural network, is presented. Techniques for addressing noise levels are proposed, and the results are compared with regular noise cancellation techniques that adopt high-pass filtering. Note to Practitioners-This article presents a methodology for automatic estimation of buildings status in the aftermath of a natural disaster, such as an earthquake. It focuses on using low-cost inertial sensors, such as accelerometers, to sense buildings' vibrations and then applying machine learning algorithms to detect damage. Utilizing the convolutional network approach, the proposed methods detect the building damage state with high accuracy. Since this article focuses on using cheap sensors, the cost of deploying a sensor network to monitor buildings is reduced significantly. Deploying this network enables rescue and reconnaissance teams to have a clear view of the most vulnerable structures.
Ahmed Ibrahim 0007, Ahmed M. Eltawil, Yunsu Na, Sherif El-Tawil
IEEE Trans Autom. Sci. Eng.2
2019 Non-Stationary Polar Codes for Resistive Memories
abstract
Resistive memories are considered a promising memory technology enabling high storage densities. However, the readout reliability of resistive memories is impaired due to the inevitable existence of wire resistance, resulting in the sneak path problem. Motivated by this problem, we study polar coding over channels with different reliability levels, termed non-stationary polar codes, and we propose a technique improving the bit error rate (BER) performance. We then apply the framework of non-stationary polar codes to the crossbar array and evaluate its BER performance under two modeling approaches, namely binary symmetric channels and binary asymmetric channels. Finally, we propose a technique for biasing the proportion of high-resistance states in the crossbar array and show its advantage in reducing further the BER. Several simulations are carried out using a SPICE-like simulator, exhibiting significant reduction in BER.
Marwen Zorgui, Mohamed E. Fouda, Zhiying Wang 0001, Ahmed M. Eltawil, Fadi J. Kurdahi
GLOBECOM4
2019 Simple MOS Transistor-Based Realization of Fractional-Order Capacitors
abstract
A new second-order MOS transistor based circuit block approximating the behavior of a fractional-order capacitor is proposed. The circuit is modular and therefore the order of the approximation can be increased by more stages of the same circuit in cascade or in parallel. Simulation results using a TSMC 65nm CMOS technology are provided and show less than 2° of phase error in two decades around the center frequency of the approximation. Experimental results of realized fractional-order capacitors and of a fractional-order relaxation oscillator are also shown.
Mohamed E. Fouda, Ahmed AboBakr, Ahmed S. Elwakil, Ahmed Gomaa Radwan, Ahmed M. Eltawil
ISCAS5
2019 Testing Topology Adaptive Irrigation IoT with Circuits
abstract
There is a significant unrealized potential in developing state of the art electronics for agriculture, particularly, for irrigation systems. However, development velocity in the domain of Irrigation Internet of Things (IrIoT) is slowed due to lengthy and complex validation cycles that require multi domain integration testing. This paper proposes testing IrIoT distributed controllers on electrical circuit platform prior to final verification. This flexible testing environment utilizes wires as opposed to water lines for testing design shortcomings. To fully expose challenges associated with next generation IrIoT this paper discusses testing challenges of topology adaptive distributed wireless irrigation controllers. Here are presented contributions in topology adaptation method, software simulation tools, an intermediate testing step using circuits as opposed to directly conducting integration testing in the target setting. The proposed method establishes an isolated circuit sandbox where integral system components, in this case IrIoT controllers, and algorithms are tested prior to full integration testing.
Davit Hovhannisyan, Ahmed M. Eltawil, Fadi J. Kurdahi
ISCAS2
2019 Feasibility Study of Plant Health Monitoring
abstract
Continuous monitoring of crop is an essential task of agricultural practices for detection of diseases or pests, precision irrigation and fertilization. The state of the art monitoring and imaging systems use aerial imaging to obtain visual feedback and multi-spectral imagery to determine crop growth factors. The main idea is that the features can be automatically calculated and assessed after pre-processing the images. After pre-processing, then the parameters can be computed using image processing techniques. For example, key leaf function traits leaf life span, leaf mass per area can be calculated. Our findings indicate that plant health assessment could be moved from lab and expensive monitoring tools to ubiquitous silicon technology based cost effect solutions without much loss of accuracy.
Davit Hovhannisyan, Kareem Khalifeh, Peng Fei, Ahmed M. Eltawil, Fadi J. Kurdahi
ISCAS4
2019 Hybrid pyramid-DWT-SVD dual data hiding technique for videos ownership protection
Farhan A. Alenizi, Fadi J. Kurdahi, Ahmed M. Eltawil, Awad Kh. Al-Asmari
Multim. Tools Appl.3
2018 Rapid in-memory matrix multiplication using associative processor
abstract
Memory hierarchy latency is one of the main problems that prevents processors from achieving high performance. To eliminate the need of loading/storing large sets of data, Resistive Associative Processors (ReAP) have been proposed as a solution to the von Neumann bottleneck. In ReAPs, logic and memory structures are combined together to allow inmemory computations. In this paper, we propose a new algorithm to compute the matrix multiplication inside the memory that exploits the benefits of ReAP. The proposed approach is based on the Cannon algorithm and uses a series of rotations without duplicating the data. It runs in O(n), where n is the dimension of the matrix. The method also applies to a large set of row by column matrix-based applications. Experimental results show several orders of magnitude increase in performance and reduction in energy and area when compared to the latest FPGA and CPU implementations.
Mohamed A. Neggaz, Hasan Erdem Yantir, Smaïl Niar, Ahmed M. Eltawil, Fadi J. Kurdahi
DATE4
2018 Circuit Inspired Modeling Method for Irrigation
abstract
Precision irrigation systems promise to bring significant improvement in resource efficiency and crop yield by providing analytics and smart tools for the growers. While significant amounts of data can be collected in a sensor-rich system, there are no rigorously designed models that can provide actionable intelligence to the user. This paper proposes the integration of circuit-inspired modeling of natural phenomena and man-made artifacts to generate end-to-end irrigation system circuit models. Such models can take advantage of existing circuit design and simulation tools that have been perfected over the past decades to efficiently process large input sets. We show that circuit-inspired models are indeed qualitatively sound and quantitatively accurate in capturing both natural phenomena and engineered physical irrigation systems.
Davit Hovhannisyan, Ahmed M. Eltawil, Mohammad Abdullah Al Faruque, Fadi J. Kurdahi
DSD2
2018 Collision Tolerance and Throughput Gain in Full-Duplex IEEE 802.11 DCF
abstract
As WiFi networks become more prevalent, there is more demand to accommodate increasing data traffic over WiFi. Traditional methods have been heavily used to improve the performance of wireless systems, and enough improvements have been introduced to exhaust channel capacity close to the maximum theoretical limits. Thus, to meet the ever increasing demand, Full-Duplex (FD) communications are enabled by Self-Interference Cancellation (SIC) to theoretically double channel capacity. SIC is possible for WiFi signals due to the lower transmit power, which makes WiFi under IEEE 802.11 standard a strong candidate for FD techniques. In this paper, we provide matching analytical and simulation results to explore how packet collisions are reduced and how throughput increases when FD methods are implemented for IEEE 802.11. Additionally, a collision-free mode enabled by FD communications is explored for WiFi systems. Simulation results show that the proposed analytical FD framework for IEEE 802.11 is accurate even when randomness is established in simulation scenarios.
Murad Murad, Ahmed M. Eltawil
ICC2
2018 Power optimization techniques for associative processors
Hasan Erdem Yantir, Ahmed M. Eltawil, Smaïl Niar, Fadi J. Kurdahi
J. Syst. Archit.2
2018 A Two-Dimensional Associative Processor
Hasan Erdem Yantir, Ahmed M. Eltawil, Fadi J. Kurdahi
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Efficient pulsed-latch implementation for multiport register files: work-in-progress
abstract
In this paper, register file design using pulsed latches is presented. Having some advantages in performance, area and power, pulsed latches represent an attractive implementation of register files. In addition, a proposed multiport register file architecture is introduced using single physical read/write ports to virtualize additional ports for read and write. The initial results show huge savings in area and power in comparison to the traditional architectures.
Wael M. Elsharkasy, Hasan Erdem Yantir, Amin Khajeh, Ahmed M. Eltawil, Fadi J. Kurdahi
CASES4
2017 Frequency and Timing Synchronization for In-Band Full-Duplex OFDM System
abstract
This paper presents frequency and timing synchronization error compensation techniques for In-Band Full- Duplex (IBFD) communication systems that employ Orthogonal Frequency Division Multiplexing (OFDM). First, we describe a system model of a full-duplex base station receiving from a remote node, while transmitting to another node on the same frequency, and in same time slot. Synchronization issues between base station and remote node are analyzed. Impairments such as carrier frequency offset, sampling time offset, symbol timing offset are addressed considering both signal-of-interest and self-interferer components of the received composite signal. We then present the receive chain of the full-duplex OFDM system, and propose compensation techniques. The proposed receiver is simulated and overall performance degradation is measured to be within 0.3 to 0.7 dB of the ideal receiver, in various channel conditions. The proposed techniques are implemented and tested experimentally on a real time IBFD-OFDM system, using software defined radios. Overall performance degradation is measured to be within 0-1.5 dB as compared to a wired synchronized system, in indoor channel conditions.
Sergey Shaboyan, Elsayed Ahmed, Alireza Shahan Behbahani, Waleed Younis, Ahmed M. Eltawil
GLOBECOM5
2017 Low Latency Approximate Adder for Highly Correlated Input Streams
abstract
Approximate computing helps achieve better performance or energy efficiency by trading accuracy. Most approximate adders are composed of multiple sub-adders and long carry chains are split to reduce latency, thus benefiting from the fact that carry propagation across long carry chains is rare for uniformly distributed inputs. One key tradeoff of these approximate adders is between latency and error rate. The more prediction bits are used, the lower is the error rate, but the latency is longer. In this paper, we present a Correlation Aware Predictor (CAP) which utilizes spatial-temporal correlation information of input streams to predict carry-in value for sub-adders. CAP uses less prediction bits which help reduce adder latency significantly. For highly correlated input streams, we found that CAP can reduce adder latency by about 23% at the same error rate compared to prior work. We implemented a CAP-based approximate adder in Verilog and synthesized with TSMC 16nm library. Synthesis results show that CAP-based adder can reduce latency by 25% and save 13% in silicon area compared to state-of-the-art.
Ahmed M. Eltawil, Fadi J. Kurdahi
ICCD2
2017 A Simple Full-Duplex MAC Protocol Exploiting Asymmetric Traffic Loads in WiFi Systems
abstract
The pressing need to accommodate increasing wireless traffic demands coupled with exhausting many advancements in Modulation and Coding Schemes (MCS) requires alternative techniques. Full-Duplex (FD) communications can be a potential candidate since channel capacity is theoretically doubled by employing Simultaneous Transmission and Reception (STR) over a single channel. In this paper, we adopt a simple FD-MAC protocol to IEEE 802.11 Distributed Coordination Function (DCF) in order to improve the aggregate goodput of a typical WiFi system. Furthermore, we look at the concept of FD communications from an unconventional perspective to propose a mechanism that increases the symmetry between uplink (UL) and downlink (DL) traffic loads while preventing complete node starvation. Simulation results show an increase in the aggregate goodput of the system due to increasing UL/DL traffic symmetry. While the simple FD-MAC protocol alone improves the aggregate goodput by an average increase of ~85% compared to standard IEEE 802.11 DCF, introducing our proposed scheme in the system improves the aggregate goodput by an additional average factor of up to ~20%.
Murad Murad, Ahmed M. Eltawil
WCNC2
2017 Approximate Memristive In-memory Computing
abstract
The bottleneck between the processing elements and memory is the biggest issue contributing to the scalability problem in computing. In-memory computation is an alternative approach that combines memory and processor in the same location, and eliminates the potential memory bottlenecks. Associative processors are a promising candidate for in-memory computation, however the existing implementations have been deemed too costly and power hungry. Approximate computing is another promising approach for energy-efficient digital system designs where it sacrifices the accuracy for the sake of energy reduction and speedup in error-resilient applications. In this study, approximate in-memory computing is introduced in memristive associative processors. Two approximate computing methodologies are proposed; bit trimming and memristance scaling. Results show that the proposed methods not only reduce energy consumption of in-memory parallel computing but also improve their performance. As compared to other existing approximate computing methodologies on different architectures (e.g., CPU, GPU, and ASIC), approximate memristive in-memory computing exhibits better results in terms of energy reduction (up to 80x) and speedup (up to 20x) on a variety of benchmarks from different domains when quality degradation is limited to 10% and it confirms that memristive associative processors provide a highly-promising platform for approximate computing.
Hasan Erdem Yantir, Ahmed M. Eltawil, Fadi J. Kurdahi
ACM Trans. Embed. Comput. Syst.2
2016 Process variations-aware resistive associative processor design
abstract
Recent breakthroughs in memristive devices have demonstrated the potential of using resistive content addressable memories for associative processing. These architectures enable ultra-high density integrated circuits along with low-power computation. However, the reliability of memristive elements is limiting the widespread adoption of these architectures. In this study, we address the reliability issues that arise in high density, resistive associative processor architectures. We propose methods to design process variation immune resistive content addressable memories and minimize the error probabilities. According to SPICE-based circuit simulations, the reliability of the circuit increases significantly and thus positively influences the accuracy of arithmetic operations as well.
Hasan Erdem Yantir, Mohamed E. Fouda, Ahmed M. Eltawil, Fadi J. Kurdahi
ICCD3
2016 Performance analysis of full-duplex multiuser decode-and-forward relay networks with interference management
abstract
In this paper, a cooperative communications scheme with interference management is proposed for Full-Duplex (FD) multi-user, decode-and-forward (DF) relay networks. The scheme is based on a relay selection method that maximizes the received signal-to-noise ratio (SNR). We evaluate the performance of the system using two approaches for interference management, namely, the constellation real parts (CRP) of the modulated signals, and a power adjustment technique. We derive expressions for the average outage probability of the up-link (UL) and downlink (DL) for the proposed scheme, as well as for standard halfduplex (HD), and standard FD. Numerical results are provided to validate the analysis and the performance of the proposed scheme.
Aymen Omri, Alireza Shahan Behbahani, Ahmed M. Eltawil, Mazen Hasna
WCNC3
2015 Energy Aware Mapping for Reconfigurable Wireless MPSoCs
abstract
Energy management for multimode software defined radio systems remains a daunting challenge. This brief develops a high level framework that generates a multiprocessor systems on chip architecture from a library of heterogeneous processing resources that can be reconfigured to support various modes of operation. The framework proposes joint task and core mapping with system level floorplanning. With the objective of minimizing energy, we develop an analytical probabilistic model that considers static, dynamic, configuration, and communication energy components for multiple applications characterized by probabilities of execution. Finally, a fast energy aware joint task and core mapping heuristic is proposed and performance is demonstrated on realistic benchmarks.
Amr M. A. Hussien, Rahul Amin, Ahmed M. Eltawil, Jim Martin 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2015 On Phase Noise Suppression in Full-Duplex Systems
abstract
Oscillator phase noise has been shown to be one of the main performance limiting factors in full-duplex systems. In this paper, we consider the problem of self-interference cancellation with phase noise suppression in full-duplex systems. The feasibility of performing phase noise suppression in full-duplex systems in terms of both complexity and achieved gain is analytically and experimentally investigated. First, the effect of phase noise on full-duplex systems and the possibility of performing phase noise suppression are studied. Two different phase noise suppression techniques with a detailed complexity analysis are then proposed. For each suppression technique, both free-running and phase-locked loop-based oscillators are considered. Due to the fact that full-duplex system performance highly depends on hardware impairments that are difficult to fully model, experimental results in a typical indoor environment are presented. The experimental results performed on two different platforms confirm results obtained from numerical simulations. Finally, the tradeoff between the required complexity and the gain achieved using phase noise suppression is discussed.
Elsayed Ahmed, Ahmed M. Eltawil
IEEE Trans. Wirel. Commun.2
2015 All-Digital Self-Interference Cancellation Technique for Full-Duplex Systems
abstract
Full-duplex systems are expected to double the spectral efficiency compared to conventional half-duplex systems if the self-interference signal can be significantly mitigated. Digital cancellation is one of the lowest complexity self-interference cancellation techniques in full-duplex systems. However, its mitigation capability is very limited, mainly due to transmitter and receiver circuit's impairments (e.g., phase noise, nonlinear distortion, and quantization noise). In this paper, we propose a novel digital self-interference cancellation technique for full-duplex systems. The proposed technique is shown to significantly mitigate the self-interference signal as well as the associated transmitter and receiver impairments, more specifically, transceiver nonlinearities and phase noise. In the proposed technique, an auxiliary receiver chain is used to obtain a digital-domain copy of the transmitted Radio Frequency (RF) self-interference signal. The self-interference copy is then used in the digital-domain to cancel out both the self-interference signal and the associated transmitter impairments. Furthermore, to alleviate the receiver phase noise effect, a common oscillator is shared between the auxiliary and ordinary receiver chains. A thorough analytical and numerical analysis for the effect of the transmitter and receiver impairments on the cancellation capability of the proposed technique is presented. Finally, the overall performance is numerically investigated showing that using the proposed technique, the self-interference signal could be mitigated to ~3 dB higher than the receiver noise floor, which results in up to 76% rate improvement compared to conventional half-duplex systems at 20 dBm transmit power values.
Elsayed Ahmed, Ahmed M. Eltawil
IEEE Trans. Wirel. Commun.2
2015 Full-Duplex Systems Using Multireconfigurable Antennas
abstract
Full-duplex systems are expected to achieve 100% rate improvement over half-duplex systems if the self-interference signal can be significantly mitigated. In this paper, we propose the first full-duplex system utilizing multireconfigurable antenna (MRA) with ~90% rate improvement compared with half-duplex systems. MRA is a dynamically reconfigurable antenna structure that is capable of changing its properties according to certain input configurations. A comprehensive experimental analysis is conducted to characterize the system performance in typical indoor environments. The experiments are performed using a fabricated MRA that has 4096 configurable radiation patterns. The achieved MRA-based passive self-interference suppression is investigated, with detailed analysis for the MRA training overhead. In addition, a heuristic-based approach is proposed to reduce the MRA training overhead. The results show that at 1% training overhead, a total of 95 dB self-interference cancellation is achieved in typical indoor environments. The 95-dB self-interference cancellation is experimentally shown to be sufficient for 90% full-duplex rate improvement compared with half-duplex systems.
Elsayed Ahmed, Ahmed M. Eltawil, Zhouyuan Li, Bedri A. Cetiner
IEEE Trans. Wirel. Commun.2
2014 An interference cancellation strategy for broadcast in hierarchical cell structure
abstract
In this paper, a hierarchical cell structure is considered, where public safety broadcasting is fulfilled in a femtocell located within a macrocell. In the femtocell, also known as local cell, an access point broadcasts to each local node (LN) over an orthogonal frequency sub-band independently. Since the local cell shares the spectrum licensed to the macrocell, a given LN is interfered by transmissions of the macrocell user (MU) in the same sub-band. To improve the broadcast performance in the local cell, a novel scheme is proposed to mitigate the interference from the MU to the LN while achieving diversity gain. For the sake of performance evaluation, ergodic capacity of the proposed scheme is quantified and a corresponding closed-form expression is obtained. By comparing with the traditional scheme that suffers from the MU's interference, numerical results substantiate the advantage of the proposed scheme and provide a useful tool for the broadcast design in hierarchical cell systems.
Yuli Yang 0003, Sonia Aïssa, Ahmed M. Eltawil, Khaled N. Salama
GLOBECOM3
2014 State dependent statistical timing model for voltage scaled circuits
abstract
This paper presents a novel statistical state-dependent timing model for voltage over scaled (VoS) logic circuits that accurately and rapidly finds the timing distribution of output bits. Using this model erroneous VoS circuits can be represented as error-free circuits combined with an error-injector. A case study of a two point DFT unit employing the proposed model is presented and compared to HSPICE circuit simulation. Results show an accurate match, with significant speedup gains.
Aras Pirbadian, Muhammed S. Khairy, Ahmed M. Eltawil, Fadi J. Kurdahi
ISCAS3
2014 Low power reduced-complexity error-resilient MIMO detector
abstract
This paper presents a reduced-complexity low power error-resilient K-Best MIMO Detector. A novel tree-enumeration method is proposed such that the error-resilient detection processes a reduced search space and is more suitable for VLSI design. Moreover, a circuit-level optimization is employed to further simplify the complexity. Experimental results are given showing that the circuit-level optimization decreases the detector area by 15% and power consumption by 41%. Moreover, we show that the proposed error-resilient MIMO detector with reduced-voltage memory can achieve a total of 19% reduction in power consumption compared with the conventional scheme, while still maintaining close-to optimal PER performance.
Chung-An Shen, Muhammed S. Khairy, Ahmed M. Eltawil, Fadi J. Kurdahi
ISCAS3
2014 Multicopy Cache: A Highly Energy-Efficient Cache Architecture
abstract
Caches are known to consume a large part of total microprocessor energy. Traditionally, voltage scaling has been used to reduce both dynamic and leakage power in caches. However, aggressive voltage reduction causes process-variation-induced failures in cache SRAM arrays, thus compromising cache reliability. We present MultiCopy Cache (MC 2 ), a new cache architecture that achieves significant reduction in energy consumption through aggressive voltage scaling while maintaining high error resilience (reliability) by exploiting multiple copies of each data item in the cache. Unlike many previous approaches, MC 2 does not require any error map characterization and therefore is responsive to changing operating conditions (e.g., Vdd noise, temperature, and leakage) of the cache. MC 2 also incurs significantly lower overheads compared to other ECC-based caches. Our experimental results on embedded benchmarks demonstrate that MC 2 achieves up to 60% reduction in energy and energy-delay product (EDP) with only 3.5% reduction in IPC and no appreciable area overhead.
Arup Chakraborty, Houman Homayoun, Amin Khajeh, Nikil Dutt, Ahmed M. Eltawil, Fadi J. Kurdahi
ACM Trans. Embed. Comput. Syst.5
2014 Link adaptation for wireless systems
abstract
Copyright © 2012 John Wiley & Sons, Ltd. To improve the robustness and reliability of wireless transmissions, two complementary link adaptation techniques are employed: adaptive modulation and coding (AMC) at the physical layer and hybrid automatic retransmission request (HARQ) at the medium access control layer. Because of their effectiveness in combating errors induced by the wireless channel, AMC and HARQ are now integral components of most emerging broadband wireless system standards, for example, LTE and WiMAX. Spectral efficiency (SE) as measured in bit per second per Hertz is one important parameter used to characterize a wireless system for comparison between different systems or between different configurations of the same system. This work provides a holistic approach of cross-layer optimizations with the intent of maximizing SE by combining AMC and HARQ. It formulates closed-form equations for calculating the average SE for wireless systems with the Rayleigh fading channel model. A new online algorithm is developed to optimize SE for both Rayleigh and non-Rayleigh fading channel. Simulations using proven LTE model are performed to compare SE obtained from closed-form equations and the developed algorithm for different system configurations. With the developed algorithm to determine how many retransmissions required in addition to the initial transmission in advance depending on the current wireless channel condition, the latency can be reduced up to 24 ms when sending the initial transmission and all of its retransmissions sooner than waiting for retransmission requests as is done previously.
Sang V. Tran, Ahmed M. Eltawil
Wirel. Commun. Mob. Comput.2
2013 Heterogeneous memory management for 3D-DRAM and external DRAM with QoS
abstract
This paper presents an innovative memory management approach to utilize both 3D-DRAM and external DRAM (ex-DRAM). Our approach dynamically allocates and relocates memory blocks between the 3D-DRAM and the ex-DRAM to exploit the high memory bandwidth and the low memory latency of the 3D-DRAM as well as the high capacity and the low cost of the ex-DRAM. Our simulation shows that in workloads that are not memory intensive, our memory management technique transfers all active memory blocks to the 3D-DRAM which runs faster than the ex-DRAM. In memory intensive workloads, our memory management technique utilizes both the 3D-DRAM and the ex-DRAM to increase the memory bandwidth to alleviate bandwidth congestion. Our approach supports Quality of Service (QoS) for “latency sensitive”, “bandwidth sensitive”, and “insensitive” applications. To improve the performance and satisfy a certain level of QoS, memory blocks of different application types are allocated differently. Compared to the scratchpad memory management mechanism, the average memory access latency of our approach decreases by 19% and 23%, while performance improves by up to 5% and 12% in single threaded benchmarks and multi-threaded benchmarks respectively. Moreover, using our approach, applications do not need to manage memory explicitly like in the scratchpad case. Our memory block relocation comes with negligible performance overhead, particularly for applications which have high spatial memory locality.
Le-Nguyen Tran, Fadi J. Kurdahi, Ahmed M. Eltawil, Houman Homayoun
ASP-DAC3
2013 Self-interference cancellation with phase noise induced ICI suppression for full-duplex systems
abstract
One of the main bottlenecks in practical full-duplex systems is the oscillator phase noise, which bounds the possible cancellable self-interference power. In this paper, a digital-domain self-interference cancellation scheme for full-duplex orthogonal frequency division multiplexing systems is proposed. The proposed scheme increases the amount of cancellable self-interference power by suppressing the effect of both transmitter and receiver oscillator phase noise. The proposed scheme consists of two main phases, an estimation phase and a cancellation phase. In the estimation phase, the minimum mean square error estimator is used to jointly estimate the transmitter and receiver phase noise associated with the incoming self-interference signal. In the cancellation phase, the estimated phase noise is used to suppress the intercarrier interference caused by the phase noise associated with the incoming self-interference signal. The performance of the proposed scheme is numerically investigated under different operating conditions. It is demonstrated that the proposed scheme could achieve up to 9 dB more self-interference cancellation than the existing digital-domain cancellation schemes that ignore the intercarrier interference suppression.
Elsayed Ahmed, Ahmed M. Eltawil, Ashutosh Sabharwal
GLOBECOM2
2013 Balancing Spectral Efficiency, Energy Consumption, and Fairness in Future Heterogeneous Wireless Systems with Reconfigurable Devices
abstract
In this paper, we present an approach to managing resources in a large-scale heterogeneous wireless network that supports reconfigurable devices. The system under study embodies internetworking concepts requiring independent wireless networks to cooperate in order to provide a unified network to users. We propose a multi-attribute scheduling algorithm implemented by a central Global Resource Controller (GRC) that manages the resources of several different autonomous wireless systems. The attributes considered by the multi-attribute optimization function consist of system spectral efficiency, battery lifetime of each user (or overall energy consumption), and instantaneous and long-term fairness for each user in the system. To compute the relative importance of each attribute, we use the Analytical Hierarchy Process (AHP) that takes interview responses from wireless network providers as input and generates weight assignments for each attribute in our optimization problem. Through Matlab/CPLEX based simulations, we show an increase in a multi-attribute system utility measure of up to 57% for our algorithm compared to other widely studied resource allocation algorithms including Max-Sum Rate, Proportional Fair, Max-Min Fair and Min Power.
Rahul Amin, Jim Martin 0001, Juan D. Deaton, Luiz A. DaSilva, Amr M. A. Hussien, Ahmed M. Eltawil
IEEE J. Sel. Areas Commun.6
2013 Rate Gain Region and Design Tradeoffs for Full-Duplex Wireless Communications
abstract
In this paper, we analytically study the regime in which practical full-duplex systems can achieve larger rates than an equivalent half-duplex systems. The key challenge in practical full-duplex systems is uncancelled self-interference signal, which is caused by a combination of hardware and implementation imperfections. Thus, we first present a signal model which captures the effect of significant impairments such as oscillator phase noise, low-noise amplifier noise figure, mixer noise, and analog-to-digital converter quantization noise. Using the detailed signal model, we study the rate gain region, which is defined as the region of received signal-of-interest strength where full-duplex systems outperform half-duplex systems in terms of achievable rate. The rate gain region is derived as a piecewise linear approximation in log-domain, and numerical results show that the approximation closely matches the exact region. Our analysis shows that when phase noise dominates mixer and quantization noise, full-duplex systems can use either active analog cancellation or baseband digital cancellation to achieve near-identical rate gain regions. Finally, as a design example, we numerically investigate the full-duplex system performance and rate gain region in typical indoor environments for practical wireless applications.
Elsayed Ahmed, Ahmed M. Eltawil, Ashutosh Sabharwal
IEEE Trans. Wirel. Commun.2
2013 A Note on "Amplify-and-Forward Relay Networks under Received Power Constraint"
abstract
This letter is to correct the incorrect optimal relay coefficient at the k-th relay derived in [1] for an amplify-and-forward (AF) wireless relay network under received power constraints. While the trends remain the same, with the correct optimal relay coefficient in this paper, nonnegligible improvement, e.g., about 0.8 dB at BER = 10-7, can be achieved.
Kanghee Lee 0001, Hyuck M. Kwon, Alireza Shahan Behbahani, Ahmed M. Eltawil
IEEE Trans. Wirel. Commun.4
2012 Error resilient MIMO detector for memory-dominated wireless communication systems
abstract
In current broadband MIMO-OFDM systems such as 3GPP LTE, embedded buffering memories occupy a large portion of chip area and a significant amount of power consumption. Due to the dense structure of memories, they are especially vulnerable to scaling effects such as process variation. These effects (hardware errors) become more pronounced when aggressive voltage scaling is used due to the reduced voltage overhead. To address this issue, we present an error resilient MIMO detector. First, we derive a combined distribution of the received data in a MIMO-OFDM receiver that includes both the noise incurred by the wireless channel and errors introduced at the receiver buffering memory due to aggressive voltage scaling. Using the derived distribution, a modified MIMO detection algorithm based on the tree-searching structure is presented. A case study is presented showing that the proposed approach can achieve near-optimal performance in the presence of both channel noise and memory error, while 40% to 50% of memory power savings are realized.
Muhammed S. Khairy, Chung-An Shen, Ahmed M. Eltawil, Fadi J. Kurdahi
GLOBECOM3
2012 Fast error aware model for arithmetic and logic circuits
abstract
As a result of supply voltage reduction and process variations effects, the error free margin for dynamic voltage scaling has been drastically reduced. This paper presents an error aware model for arithmetic and logic circuits that accurately and rapidly estimates the propagation delays of the output bits in a digital block operating under voltage scaling to identify circuit-level failures (timing violations) within the block. Consequently, these failure models are then used to examine how circuit-level failures affect system-level reliability. A case study consisting of a CORDIC DSP unit employing the proposed model provides tradeoffs between power, performance and reliability.
Samy Zaynoun, Muhammed S. Khairy, Ahmed M. Eltawil, Fadi J. Kurdahi, Amin Khajeh
ICCD3
2012 Spectral efficiency and energy consumption tradeoffs for reconfigurable devices in heterogeneous wireless systems
abstract
The proliferation of wireless broadband usage over the last decade has led to the development and deployment of multiple broadband wireless radio access technologies (RATs) such as EVDO, WiMAX, HSPA and LTE. To support the ever-increasing wireless traffic demand, researchers have worked on the concept of an integrated heterogeneous wireless environment that encompasses several of these RATs which makes the resource allocation process more efficient by assigning each user in the system to the best RAT/RATs. In this paper, for such an integrated heterogeneous wireless system, we show the possible gains in spectral efficiency at the cost of increased energy consumption for an unbalanced heterogeneous wireless network deployment scenario. In prior work, based on the assumption that all cellular operators under study were equally well-provisioned, we showed an increase in spectral efficiency of up to 75%. In the research presented in this paper, we assume the coverage of each operator might differ significantly in a given area. With this `unbalanced' scenario, we show that an ideal, centralized allocation strategy provides an almost linear tradeoff between gain in spectral efficiency (554%) and worst-case increase in energy consumption (615%) for users supporting elastic traffic.
Rahul Amin, Jim Martin 0001, Ahmed M. Eltawil, Amr M. A. Hussien
WCNC3
2012 Multiuser communications using beam-tilting antennas
abstract
This paper proposes means to exploit the additional degrees of freedom that can be achieved by using beam-tilting antennas for the purposes of multiuser communication without substantially increasing investment in system resources. Specifically, by using such beam-tilting antennas the wireless channel can be used as a means to spread the users in space rather than in frequency as is currently the norm in spread spectrum systems. This technique thus allows multiuser communications without increasing spectrum usage. The performance of such a system is characterized in terms of average bit error rate for various number of users and for different amounts of correlation between the channels in different beam directions.
Chitaranjan P. Sukumar, Ahmed M. Eltawil
WCNC2
2012 Optimized scheduling algorithm for LTE downlink system
abstract
Orthogonal Frequency Division Multiple Access (OFDMA), Multiple Input Multiple Output (MIMO), and Adaptive Modulation and Coding (AMC) are advanced signal processing techniques introduced in the wireless 3GPP LTE standard. In MIMO-OFDMA system, by taking advantage of the spatial dimension and depending on the time-varying wireless channel condition, the most efficient modulation and coding with spatial multiplexing or spatial diversity is selected to increase system capacity or improve system reliability. Wireless resource is divided into resource blocks (RBs), and the central idea is to effectively allocate these RBs to maximize system capacity and satisfy all active users QoS requirements. Scheduling of the available RBs in an optimized manner is therefore a main thrust in the design of LTE system. In this paper, an optimized scheduling algorithm from a holistic view of the LTE downlink system taking into account the combination of frequency-time packet scheduling (FTPS) with added spatial dimension or MIMO and QoS awareness is developed. Simulation results are presented that confirm the performance improvement of the proposed technique.
Sang V. Tran, Ahmed M. Eltawil
WCNC2
2012 Error-Aware Algorithm/Architecture Coexploration for Video Over Wireless Applications
abstract
In this article, we propose a cross-layer algorithm/architecture coexploration for wireless multimedia systems to coordinate interactions among sublayer optimizers for improvements in energy/QoS/reliability. By exploiting the inherent redundancy in wireless multimedia systems, we generate an expanded design space over traditional layer-specific approaches. Specifically, we control the error resilient encoder at the application layer to provide awareness of architectural exploration at the physical layer allowing new design points with lower power consumption via aggressive voltage scaling. While trying to reduce energy consumption, the fault tolerant technique compensates the effect of the hardware and network errors due to aggressive voltage scaling and lossy transmission, respectively. Our experiments on H.263 video over a WCDMA communication system demonstrate that coexploration enlarges the feasible design space, which results in significant power savings of more than 20% in the WCDMA modem.
Amin Khajeh, Minyoung Kim 0002, Nikil Dutt, Ahmed M. Eltawil, Fadi J. Kurdahi
ACM Trans. Embed. Comput. Syst.4
2012 Variation Trained Drowsy Cache (VTD-Cache): A History Trained Variation Aware Drowsy Cache for Fine Grain Voltage Scaling
abstract
In this paper we present the “Variation Trained Drowsy Cache” (VTD-Cache) architecture. VTD-Cache allows for a significant reduction in power consumption while addressing reliability issues raised by memory cell process variability. By managing voltage scaling at a very fine granularity, each cache way can be sourced at a different voltage where the selection of voltage levels depends on both the vulnerability of the memory cells in that cache way to process variation and the likelihood of access to that cache location. After a short training period, the proposed architecture will micro-tune the cache, allowing significant power reduction with negligible increase in the number of misses. In addition, the proposed architecture actively monitors the access pattern and reconfigures the supply voltage setting to adapt to the execution pattern of the program. The novel and modular architecture of the VTD-Cache and its associated controller makes it easy to be implemented in memory compilers with a small area and power overhead. In a case study, the SimpleScalar simulation of the proposed 32 kB cache architecture reports over 57% reduction in power consumption over standard SPEC2000 integer benchmarks while incurring an area overhead of less than 4% and an execution time penalty smaller than 1%.
Avesta Sasan, Kiarash Amiri, Houman Homayoun, Ahmed M. Eltawil, Fadi J. Kurdahi
IEEE Trans. Very Large Scale Integr. Syst.4
2012 A Best-First Soft/Hard Decision Tree Searching MIMO Decoder for a 4 × 4 64-QAM System
abstract
This paper presents the algorithm and VLSI architecture of a configurable tree-searching approach that combines the features of classical depth-first and breadth-first methods. Based on this approach, techniques to reduce complexity while providing both hard and soft outputs decoding are presented. Furthermore, a single programmable parameter allows the user to tradeoff throughput versus BER performance. The proposed multiple-input-multiple-output decoder supports a 4 × 4 64-QAM system and was synthesized with 65-nm CMOS technology at 333 MHz clock frequency. For the hard output scheme the design can achieve an average throughput of 257.8 Mbps at 24 dB signal-to-noise ratio (SNR) with area equivalent to 54.2 Kgates and a power consumption of 7.26 mW. For the soft output scheme it achieves an average throughput of 83.3 Mbps across the SNR range of interest with an area equivalent to 64 Kgates and a power consumption of 11.5 mW.
Chung-An Shen, Ahmed M. Eltawil, Khaled N. Salama, Sudip Mondal
IEEE Trans. Very Large Scale Integr. Syst.2
2011 A Class of Low Power Error Compensation Iterative Decoders
abstract
Recent power reduction techniques aggressively modulate the supply voltage of embedded buffering memories allowing acceptable hardware errors to flow through the processing chain. In this paper, we introduce a class of modified Turbo and LDPC decoders that provide significant improvements over standard decoders in the presence of hardware noise. Simulation results show a consistent improvement in the BER performance of the modified decoders across all SNRs with very small area and power overheads as compared to the conventional decoders.
Amr M. A. Hussien, Muhammed S. Khairy, Amin Khajeh, Ahmed M. Eltawil, Fadi J. Kurdahi
GLOBECOM4
2011 Amplify-and-Forward Relay Networks under Received Power Constraint with Imperfect CSI
abstract
The mobility of relay stations within a cell creates an interesting scenario where multiple sources, transmitting correlated data, can co-operate to satisfy power constraints at the receiving node. While beneficial to the receive node, this approach creates un-intended interference for neighboring cells reusing the same frequency. Previously, a relay scheme was proposed to simultaneously maximize SNR and minimize MSE, for an amplify-and-forward (AF) relay network operating under a receive power constraint guaranteeing that the received signal power is bounded to control interference to neighboring cells. In this paper, we investigate the effect of channel uncertainties on system performance. A modified solution for the proposed scheme under imperfect channel state information (CSI) is provided. Furthermore, we investigate the diversity order of the proposed scheme under perfect CSI. Simulations are provided to verify the analysis for both perfect and imperfect CSI assumptions.
Alireza Shahan Behbahani, Ahmed M. Eltawil
ICC2
2011 Multiuser Sum MSE Minimization Relaying Strategy
abstract
In this paper, we design relay factors for multiuser amplify-and-forward (AF) relay networks where each node in the network is equipped with one antenna. The relay factors are chosen to minimize the sum of mean square error (sum-MSE) at the destinations without considering the destinations' noise. Subsequently, a filter is designed for each receiver to cancel out the effect of the destination's noise. A closed form solution is provided and shown to outperform the zero-forcing (ZF) scheme available in literature. For the case where the number of users increases, an approximate closed form solution is provided that can be implemented distributively. We investigate the average transmit power of each relay and show that as the number of users increases, relay transmit power is inversely proportional with the number of users in the network. Finally, a modified solution is provided to address the effect of channel estimation error. Simulations are provided to verify the analysis and present the performance of the proposed scheme.
Alireza Shahan Behbahani, Ahmed M. Eltawil
ICC2
2011 Using Reconfigurable Devices to Maximize Spectral Efficiency in Future Heterogeneous Wireless Systems
abstract
As broadband data further blends with cellular voice, mobile devices will become the dominant portals to the connected world. However current design practices still involve building independent networks that each make their own resource decisions. In spite of the tremendous amount of related research in this area, there are still several elemental questions that must be addressed. First, is it better to treat wireless systems as independent access networks requiring the user to handle aspects of roaming between disparate wireless networks or is an internet model better where independent autonomous wireless systems (AWS) cooperate to form a single, unified cloud to users, with network level resource allocation? Second, is it better to have dedicated, low power circuitry that supports a limited set of independent wireless Radio Access technologies (RATs) or is it better to build agile handsets that adapt (reconfigure) in real-time to operate over a large range of RAT technologies and operating modes? The results in this paper shed light on these questions. We present preliminary results from a MATLAB-based simulation study that highlights the increase in spectral efficiency as the modality of devices increase. Our analysis takes into account the cost of radio reconfiguration in terms of the temporary communications downtime and the surge of power that occurs with each reconfiguration operation. Our main result suggests that nomadic users benefit the most primarily due to their ability to route traffic over 'hotspot' type of RATs that tend to have high data rates at reduced coverage, and that this in turn helps increase the 3G or 4G bandwidth available to mobile users. All nodes in the system experience an increase in spectral efficiency ranging from 14% to 75% when compared to a similar scenario that assumes no network cooperation and static radios.
Jim Martin 0001, Rahul Amin, Ahmed M. Eltawil, Amr M. A. Hussien
ICCCN3
2011 Energy aware task mapping algorithm for heterogeneous MPSoC based architectures
abstract
Energy Management for multi-mode Software Defined Radio (SDR) systems remains a daunting challenge. In this paper, we focus on the issue of task allocation for multi-processor based systems with hybrid processing resources that can be reconfigured. With the objective of minimizing energy, we propose a fast, energy aware static task mapping heuristic to minimize the average overall energy consumption. Simulation results show that the proposed heuristic is capable of achieving results that are within 20% of the optimal solution while providing orders of magnitude speedup in processing time.
Amr M. A. Hussien, Ahmed M. Eltawil, Rahul Amin, Jim Martin 0001
ICCD2
2011 Reconfigurable filter implementation of a matched-filter based spectrum sensor for Cognitive Radio systems
abstract
Spectrum sensing is one of the most important features of Cognitive Radio (CR) systems. Matched-filter based spectrum sensing techniques provide optimum sensing performance given that a number of characteristics of the transmitted signal are known by the sensors. Assuming that the received signal pertains to one communication standard from a given set of wireless technologies, conventional spectrum sensors employ separate filters corresponding to each standard which gives rise to increased power consumption and ciruit size. A novel reconfigurable matched-filter based spectrum sensor to be deployed in CR systems is proposed in order to overcome the disadvantages of conventional design methods. This approach proposes a spectrum of design qualities which trade-off area for reconfiguration overhead. We will show that our approach is capable of designing reconfigurable filter for standards with widely varying filter characteristics.
Amir Hossein Gholamipour, Ali Gorcin, B. Ugur Töreyin, Mazen A. R. Saghir, Fadi J. Kurdahi, Ahmed M. Eltawil
ISCAS7
2011 Linear decentralized estimation of correlated data for wireless sensor networks
abstract
In this paper, we consider distributed estimation of an unknown random vector by using wireless sensors and a fusion center (FC). We adopt a linear model for distributed estimation of a vector source where both observation models and sensor operations are linear and the multiple access channel (MAC) is coherent. The sensors are designed to minimize the mean square error (MSE) at the fusion center without considering the noise at the fusion center. Subsequently, a filter is designed to cancel out the effect of the noise at the fusion center. We present a closed form solution. When the number of unknown parameters increases, an approximate closed form solution is provided that can be implemented distributively. Since there is no power constraint imposed on transmit power of each sensor, we investigate the average transmit power of each sensor. We show that as the number of unknown parameters increases, the sensor power is inversely proportional to the number of unknown parameters of interest. Finally, simulations are provided to verify the analysis and present the performance of the proposed scheme.
Alireza Shahan Behbahani, Ahmed M. Eltawil, Hamid Jafarkhani
SECON2
2011 Embedded Memories Fault-Tolerant Pre- and Post-Silicon Optimization
abstract
This paper proposes a structured method for scaling both the supply voltage as well as the body bias voltage for CMOS embedded static memory with the aim of achieving a controllable and dynamic probability of failure with minimum power consumption for each memory block. The target error probability is managed according to the time varying error tolerance attributes of the application using the memory at a certain instant in time. This approach enables system designers to abstract the concepts of power awareness, yield and reliability as design tradeoffs-that incorporate application knowledge-early in the design cycle. The paper develops a formal theoretical and practical foundation based on the underlying device statistics upon which both system and circuit designers can investigate error aware design.
Amin Khajeh, Ahmed M. Eltawil, Fadi J. Kurdahi
IEEE Trans. Very Large Scale Integr. Syst.2
2011 Inquisitive Defect Cache: A Means of Combating Manufacturing Induced Process Variation
abstract
This paper proposes a new fault tolerant cache organization capable of dynamically mapping the in-use defective locations in a processor cache to an auxiliary parallel memory, creating a defect-free view of the cache for the processor. While voltage scaling has a super-linear effect on reducing power, it exponentially increases the defect rate in memory. The ability of the proposed cache organization to tolerate a large number of defects makes it a perfect candidate for voltage-scalable architectures, especially in smaller geometries where manufacturing induced process variation (MIPV) is expected to rapidly increase. The introduced fault tolerant architecture consumes little energy and area overhead, but enables the system to operate correctly and boosts the system performance close to a defect-free system. Power savings of over 40% is reported on standard benchmarks while the performance degradation is maintained below 1%.
Avesta Sasan, Houman Homayoun, Ahmed M. Eltawil, Fadi J. Kurdahi
IEEE Trans. Very Large Scale Integr. Syst.3
2010 E < MC2: less energy through multi-copy cache
abstract
Caches are known to consume a large part of total microprocessor power. Traditionally, voltage scaling has been used to reduce both dynamic and leakage power in caches. However, aggressive voltage reduction causes process-variation-induced failures in cache SRAM arrays, which compromise cache reliability. We present Multi-Copy Cache (MC2), a new cache architecture that achieves significant reduction in energy consumption through aggressive voltage scaling, while maintaining high error resilience (reliability) by exploiting multiple copies of each data item in the cache. Unlike many previous approaches, MC2 does not require any error map characterization and therefore is responsive to changing operating conditions (e.g., Vdd-noise, temperature and leakage) of the cache. MC2 also incurs significantly lower overheads compared to other ECC-based caches. Our experimental results on embedded benchmarks demonstrate that MC2 achieves up to 60% reduction in energy and energy-delay product (EDP) with only 3.5% reduction in IPC and no appreciable area overhead.
Arup Chakraborty, Houman Homayoun, Amin Khajeh, Nikil Dutt, Ahmed M. Eltawil, Fadi J. Kurdahi
CASES5
2010 An Adaptive Reduced Complexity K-Best Decoding Algorithm with Early Termination
abstract
This paper presents a K-Best decoding algorithm that requires a much smaller K while preserving advantages of the sphere decoding algorithm such as branch pruning and an adaptively updated pruning threshold. The proposed approach results in examining a much smaller set of modulation points with a significantly reduced complexity. Simulations are presented that quantify the BER performance and complexity in terms of the number of visited nodes. The variability in the required operations (hence run-time) due to branch pruning is studied and compared with the sphere decoding algorithm.
Chung-An Shen, Ahmed M. Eltawil
CCNC2
2010 Exploiting Architectural Similarities and Mode Sequencing in Joint Cost Optimization of Multi-mode FIR Filters
abstract
We present a new approach to designing multi-mode FIR filters in FPGAs based on partitioning a design into shared, mode-independent blocks and reconfigurable, mode-specific regions. We also provide a theoretical formulation for finding an optimal sequence of modes that minimizes a joint cost function related to area and reconfiguration overhead. Our results show that for a group of template matching filters, appropriate mode sequencing can reduce area by up to 15% and reconfiguration overhead by as much as 26%.
Amir Hossein Gholamipour, Fadi J. Kurdahi, Ahmed M. Eltawil, Mazen A. R. Saghir
FPL3
2010 A Unified Hardware and Channel Noise Model for Communication Systems
abstract
This paper presents a single, scalable, unified statistical model that accurately reflects the impact of random embedded memory failures due to power management policies on the overall performance of a communication system. The proposed framework enables system designers to efficiently and accurately determine the effectiveness of novel power management techniques and algorithms that are designed to manage both hardware failure and communication channel noise, without the added cost of lengthy system simulations that are inherently limited and suffer from lack of scalability. Furthermore, the proposed framework facilitates performing both cross layer and intra layer tradeoffs where the faulty hardware can be treated as error-free hardware thus creating a much richer design space of power, performance and reliability.
Amin Khajeh, Kiarash Amiri, Muhammed S. Khairy, Ahmed M. Eltawil, Fadi J. Kurdahi
GLOBECOM4
2010 Effect of body biasing on embedded SRAM failure
abstract
This paper studies the tradeoffs when using body biasing as a power consumption modulator for Static Random Access Memories (SRAM) in term of reliability versus power consumption. We show that for fault tolerant applications such as wireless applications and multimedia, utilizing body biasing combined with voltage scaling can result in up to 47% power saving compared to the nominal case, while, for the same scenario, utilizing only voltage scaling will result in 20% power saving in the memory compared to nominal case.
Amin Khajeh, Ahmed M. Eltawil, Fadi J. Kurdahi
ISCAS2
2010 A best-first tree-searching approach for ML decoding in MIMO system
abstract
In MIMO communication systems maximum-likelihood (ML) decoding can be formulated as a tree-searching problem. This paper presents a tree-searching approach that combines the features of classical depth-first and breadth-first approaches to achieve close to ML performance while minimizing the number of visited nodes. A detailed outline of the algorithm is given, including the required storage. The effects of storage size on BER performance and complexity in terms of search space are also studied. Our result demonstrates that with a proper choice of storage size the proposed method visits 40% fewer nodes than a sphere decoding algorithm at signal to noise ratio (SNR) = 20dB and by an order of magnitude at 0 dB SNR.
Chung-An Shen, Ahmed M. Eltawil, Sudip Mondal, Khaled N. Salama
ISCAS2
2010 Low-Power Multimedia System Design by Aggressive Voltage Scaling
abstract
Mobile multimedia systems are growing in complexity and scalability and, correspondingly, in their implementation challenges. By design, these systems have built-in error resilience that has been exploited in many different compression and transmission schemes mainly as a quality tradeoff. This paper proposes a paradigm shift in utilizing error resilience in an application-aware method for reducing the power consumption of memories in such systems by aggressively scaling the supply voltage beyond what is currently considered as ¿safe¿ operating conditions while maintaining performance. Results on H.264 decoders show that power savings of more than 40% are possible.
Fadi J. Kurdahi, Ahmed M. Eltawil, Kang Yi, Stanley Cheng, Amin Khajeh
IEEE Trans. Very Large Scale Integr. Syst.2
2010 Design and Implementation of a Sort-Free K-Best Sphere Decoder
abstract
This paper describes the design and very-large-scale integration (VLSI) architecture for a 4 × 4 breadth-first K-best multiple-input-multiple-output (MIMO) decoder using a 64 quadrature-amplitude modulation (QAM) scheme. A novel sort-free approach to path extension, as well as quantized metrics result in a high-throughput VLSI architecture with lower power and area consumption compared to state-of-the-art published systems. Functionality is confirmed via a field-programmable gate array (FPGA) implementation on a Xilinx Virtex II Pro FPGA. Comparison of simulation and measurements are given, and FPGA utilization figures are provided. Finally, VLSI architectural tradeoffs are explored for a synthesized application-specific IC (ASIC) implementation in a 65-nm CMOS technology.
Sudip Mondal, Ahmed M. Eltawil, Chung-An Shen, Khaled N. Salama
IEEE Trans. Very Large Scale Integr. Syst.2
2010 Reduced Overhead Training for Multi Reconfigurable Antennas with Beam-Tilting Capability
abstract
This paper proposes low overhead training techniques for a wireless communication system equipped with a Multifunctional Reconfigurable Antenna (MRA) capable of dynamically changing beamwidth and beam directions. A novel microelectromechanical system (MEMS) MRA antenna is presented with radiation patterns (generated using complete electromagnetic full-wave analysis) which are used to quantify the communication link performance gains. In particular, it is shown that using the proposed Exhaustive Training at Reduced Frequency (ETRF) consistently results in a reduction in training overhead. It is also demonstrated that further reduction in training overhead is possible using statistical or MUSIC-based training schemes. Bit Error Rate (BER) and capacity simulations are carried out using an MRA, which can tilt its radiation beam into one of Ndir= 4 or 8 directions with variable beamwidth (≈2π/Ndir). The performance of each training scheme is quantified for OFDM systems operating in frequency selective channels with and without Line of Sight (LoS). We observe 6 dB of gain at BER = 10-4and 6 dB improvement in capacity (at capacity = 6 bits/sec/subcarrier) are achievable for an MRA with Ndir= 8 as compared to omni directional antennas using ETRF scheme in a LoS environment.
Hamid Eslami, Chitaranjan P. Sukumar, Daniel Rodrigo López, S. Mopidevi, Ahmed M. Eltawil, Lluis Jofre, Bedri A. Cetiner
IEEE Trans. Wirel. Commun.5
2009 A fault tolerant cache architecture for sub 500mV operation: resizable data composer cache (RDC-cache)
abstract
In this paper we introduce Resizable Data Composer-Cache (RDC-Cache). This novel cache architecture operates correctly at sub 500 mV in 65 nm technology tolerating large number of Manufacturing Process Variation induced defects. Based on a smart relocation methodology, RDC-Cache decomposes the data that is targeted for a defective cache way and relocates one or few word to a new location avoiding a write to defective bits. Upon a read request, the requested data is recomposed through an inverse operation. For the purpose of fault tolerance at low voltages the cache size is reduced, however, in this architecture the final cache size is considerably higher compared to previously suggested resizable cache organizations [2][3]. The following three features a) compaction of relocated words, b)ability to use defective words for fault tolerance and c) "linking" (relocating the defective word to any row in the next bank), allows this architecture to achieve far larger fault tolerance in comparison to [2][3]. In high voltage mode, the fault tolerant mechanism of RDC-Cache is turned-off with minimal (0.91%) latency overhead compared to a traditional cache.
Avesta Sasan, Houman Homayoun, Ahmed M. Eltawil, Fadi J. Kurdahi
CASES3
2009 TRAM: A tool for Temperature and Reliability Aware Memory Design
abstract
Memories are increasingly dominating Systems on Chip (SoC) designs and thus contribute a large percentage of the total system's power dissipation, area and reliability. In this paper, we present a tool which captures the effects of supply voltage Vddand temperature on memory performance and their interrelationships. We propose a Temperature- and Reliability- Aware Memory Design (TRAM) approach which allows designers to examine the effects of frequency, supply voltage, power dissipation, and temperature on reliability in a mutually interrelated manner. Our experimental results indicate that thermal unaware estimation of probability of error can be off by at least two orders of magnitude and up to five orders of magnitude from the realistic, temperature-aware cases. We also observed that thermal aware Vddselection using TRAM can reduce the total power dissipation by up to 2.5times while attaining an identical predefined limit on errors.
Amin Khajeh, Aseem Gupta, Nikil Dutt, Fadi J. Kurdahi, Ahmed M. Eltawil, Kamal S. Khouri, Magdy S. Abadir
DATE5
2009 Process Variation Aware SRAM/Cache for aggressive voltage-frequency scaling
abstract
This paper proposes a novel Process Variation Aware SRAM architecture designed to inherently support voltage scaling. The peripheral circuitry of the SRAM is modified to selectively allow overdriving a wordline which contains weak cell(s). This architecture allows reducing the power on the entire array; however it selectively trades power for correctness when rows containing weak cells are accessed. The cell sizing is designed to assure successful read operations. This avoids flipping the content of the cells when the wordline is overdriven. Our simulations report 23% to 30% improvement in cell access time and 31% to 51% improvement in cell write time in overdriven wordlines. Total area overhead is negligible (4%). Low voltage operation achieves more than 40% reduction in dynamic power consumption and approximately 50% reduction in leakage power consumption.
Avesta Sasan, Houman Homayoun, Ahmed M. Eltawil, Fadi J. Kurdahi
DATE3
2009 Size-Reconfiguration Delay Tradeoffs for a Class of DSP Blocks in Multi-mode Communication Systems
abstract
In this paper we propose a spectrum of designs for filters in multi-mode communication systems. The proposed designs lie in between generic filter to fully optimized coefficient specific filter. For each design a reconfigurable section and a static section are defined. We propose an algorithm to optimize the size of the reconfigurable section of each design independently. We also propose another algorithm that optimizes the reconfiguration time overhead for a given sequence of designs. The results of our experiments show the trade-off between area and reconfiguration delay in the design space.
Amir Hossein Gholamipour, Hamid Eslami, Ahmed M. Eltawil, Fadi J. Kurdahi
FCCM3
2009 Demonstration of highly programmable downlink OFDMA (WiMax) transceivers for SDR systems
abstract
In this paper, we present the architecture of a highly configurable multi-input multi-output (MIMO) orthogonal frequency division multiple access (OFDMA) platform. The platform is designed to support experimentation with various communication algorithms, thus allowing an intimate understanding of the performance of complex algorithms under real-life constraints. The hardware used is the wireless open access research platform (WARP) which facilitates rapid prototyping utilizing the FPGA and multi radio interfaces available.
Hamid Eslami, Gaurav Patel, Chitaranjan P. Sukumar, Sang V. Tran, Ahmed M. Eltawil, Raghu Mysore Rao, Chris Dick
MobiHoc5
2009 A Low Power JPEG2000 Encoder With Iterative and Fault Tolerant Error Concealment
abstract
This paper presents a novel approach to reduce power in multimedia devices. Specifically, we focus on JPEG2000 as a case study. This paper indicates that by utilizing the in-built error resiliency of multimedia content, and the disjoint nature of the encoding and decoding processes, ultra low power architectures that are hardware fault tolerant can be conceived. These architectures utilize aggressive voltage scaling to conserve power at the encoder side while incurring extra processing requirements at the decoder to blindly detect and correct for encoder hardware induced errors. Simulations indicate a reduction of up to 35% in encoder power depending on the choice of technology for a 65-nm CMOS process.
Avesta Sasan, Amin Khajeh, Ahmed M. Eltawil, Fadi J. Kurdahi
IEEE Trans. Very Large Scale Integr. Syst.3
2009 Amplify-and-Forward Relay Networks Under Received Power Constraint
abstract
Relay networks have received considerable attention recently, especially when limited size and power resources impose constraints on the number of antennas at each node. While fixed and mobile relays can cooperate to improve reception at the desired destination, they also contribute to unintended interference for neighboring cells reusing the same frequency. In this paper, we propose and analyze a relay scheme to simultaneously maximize SNR and minimize MSE, for an amplify-and-forward (AF) relay network operating under a receive power constraint guaranteeing that the received signal power is bounded to control interference to neighboring cells. If the intended destination lies at the periphery of the cell, then the proposed scheme guarantees that the total power leaking into neighboring cells is bounded. The optimal relay factors are provided for both correlated and uncorrelated noise at the relays. Simulation results are presented to verify the analysis.
Alireza Shahan Behbahani, Ahmed M. Eltawil
IEEE Trans. Wirel. Commun.2
2008 On Channel Estimation and Capacity for Amplify and Forward Relay Networks
abstract
Relay networks have received considerable attention recently, especially when limited size and power resources impose constraints on the number of antennas within a wireless sensor network. In this paper, we design and analyze a training based linear mean square error (LMMSE) channel estimator for time division multiplex amplify-and-forward (AF) relay networks. For the purpose of performance comparison we consider three distinct cases; In the first scenario, each relay estimates its backward and forward channels, in the second scenario each relay knows its backward and forward channels perfectly and finally in the third scenario relays have no knowledge of channels. Finally, we find a lower bound for the capacity considering the effect of training and estimation error.
Alireza Shahan Behbahani, Ahmed M. Eltawil
GLOBECOM2
2008 Joint Power Loading of Data and Pilots in OFDM Using Imperfect Channel State Information at the Transmitter
abstract
The search for optimality in the design of channel precoders and training symbols in block processing communication systems is one of paramount importance. Finding the best tradeoff in terms of power distribution between information and pilot symbols for frequency selective channels, when channel estimation via feedback is available, however, has not been fully addressed. In this paper, we solve the problem of finding the optimal power distribution between pilots and data symbols in the mean-square-error (MSE) sense when a delayless error-free channel feedback path is available to the transmitter. The novel approach adaptively designs the optimal precoders and training vectors based on the frequency domain estimates of the channel.
Chitaranjan P. Sukumar, Ricardo Merched, Ahmed M. Eltawil
GLOBECOM3
2007 Exploiting Fault Tolerance Towards Power Efficient Wireless Multimedia Applications
abstract
This paper exploits the inherent redundancy available in wireless multimedia systems to tradeoff system redundancy versus power consumption. It is shown that aggressive power management techniques result in hardware failures that are mostly localized to embedded memories. By isolating these errors and utilizing system level fault tolerance techniques, we show that a reduction in power of up to 38% is possible in an H.264 decoder. This is coupled with a reduction of up to 20% in the underlying 3GPP WCDMA modem. With the proliferation of wideband wireless systems, an inevitable result is increased demand on receiving high quality, high bandwidth video. Based on these trends, one can identify three main challenges facing mobile, multimedia and communication designers. The first and foremost is power consumption which is on the rise due to the complex algorithms necessary to enable broadband multimedia wireless communication in dispersive and highly mobile environments. The second challenge is technology related, where scaling is both an enabler and a limiter. It enables unprecedented integration, including the ability to integrate large memories on chip, with the downside being a penalty in leakage power as well as reliability. Finally, the third challenge, cost is typically a major restricting factor. Designers are faced with the daunting dilemma of generating high yielding architectures that integrate vast amounts of logic and memories in a minimum die size with minimum power consumption. In this paper, we present an alternative approach to designing systems that have built in inherent redundancy.
Amin Khajeh, Ahmed M. Eltawil, Fadi J. Kurdahi
CCNC2
2007 Error-Aware Design
abstract
The universal underlying assumption made today is that systems on chip must maintain 100% correctness regardless of the application. This work advocates the concept that some applications - by construction - are inherently error tolerant and therefore do not require this strict bound of 100% correctness. In such cases, it is possible to exploit this tolerance by aggressively reducing the supply voltage, thereby reducing power consumption significantly. This approach is demonstrated on several case studies in imaging, video and wireless communication fields.
Fadi J. Kurdahi, Ahmed M. Eltawil, Amin Khajeh, Avesta Sasan, Stanley Cheng
DSD2
2007 On Signal Processing Methods for MIMO Relay Architectures
abstract
Relay networks have received considerable attention recently, especially when limited size and power resources impose constraints on the number of antennas within a wireless sensor network. In this context, signal processing techniques play a fundamental role, and optimality within a given relay architecture can be achieved under several design criteria. In this paper, we extend recent optimal minimum-mean-square-error (MMSE) and SNR designs of relay networks to the corresponding multiple- input-multiple-output (MIMO) scenarios, whereby the source, relays and destination comprise multiple antennas. We shall investigate maximum SNR solutions subject to power constraints and zero-forcing (ZF) criteria, as well as approximate MMSE equalizers with specified target SNR and global power constraint.
Alireza Shahan Behbahani, Ricardo Merched, Ahmed M. Eltawil
GLOBECOM3
2007 Power Management for Cognitive Radio Platforms
abstract
This paper discusses how the cognitive radio concept can be extended to allow the system not only to manage shared resources such as spectrum, but to use this knowledge to optimize the overall system power consumption. We introduce a case study of video over wireless via a 3G WCDMA modem connected to an H.264 decoder. We show that by utilizing knowledge about the communication channel, a savings of more than 20% of the overall system power is possible while maintaining a required quality of service.
Amin Khajeh, Shih-Yang Cheng, Ahmed M. Eltawil, Fadi J. Kurdahi
GLOBECOM3
2007 A Scalable Wireless Channel Emulator for Broadband MIMO Systems
abstract
This paper addresses the issue of designing scalable prototypes for multi input multi output (MIMO) wireless channel emulation. To date, emulators are extending single input single output (SISO) time based filtering techniques to emulate MIMO channels, which in turn requires quadratic increase in computational resources as the array size increases. In this paper, a new scalable approach for MIMO wireless channel emulation is presented. It is shown that for a 2x2 MIMO system, by performing channel emulation in the frequency domain, a reduction of up to 43 % in computations per output sample can be achieved as compared to traditional techniques. For higher order arrays, further reduction is possible. It is also illustrated that a linear growth in computational resources is attained through this scheme. The architecture is presented and synthesis results using an FPGA platform are discussed.
Hamid Eslami, Ahmed M. Eltawil
ICC2
2007 Limits on voltage scaling for caches utilizing fault tolerant techniques
abstract
This paper proposes a new low power cache architecture that utilizes fault tolerance to allow aggressively reduced voltage levels. The fault tolerant overhead circuits consume little energy, but enable the system to operate correctly and boost the system performance to close to defect free operation. Overall, power savings of over 40% are reported on standard benchmarks.
Avesta Sasan, Amin Khajeh, Ahmed M. Eltawil, Fadi J. Kurdahi
ICCD3
2007 Fault Tolerant Approaches Targeting Ultra Low Power Communications System Design
abstract
This paper presents a new approach to co-designing communication systems and their respective hardware architectures. We show that by taking into account the specific needs and assumptions of the algorithms running on hardware, the circuit specifications targeted by ASIC engineers can be relaxed. This in turn leads to optimal designs in both power consumption and robustness. A case-study of a complete WCDMA modem incorporating this approach shows a savings of 23% in embedded memory power consumption and a total of 13% power savings for the whole system.
Amin Khajeh, Ahmed M. Eltawil, Fadi J. Kurdahi
VTC Spring2
2006 A Real-Time Wireless Channel Emulator for MIMO Systems
abstract
The improvement in channel capacity hailed by MIMO systems is directly related to intricate details of the wireless channel such as the degree of correlation between the multiple channels established due to a MIMO configuration. Capturing these details in real time emulation systems becomes absolutely essential to validate theoretical results in a controlled setting. This paper presents methods to achieve significant improvements in speed and complexity for real time channel emulation maintaining high resolutions in time and frequency. In particular, this paper presents a frequency domain emulation technique that reduces add/multiply operations by 54% compared to conventional FIR emulation. A comprehensive complexity analysis is performed to compare the complexity of different techniques.
Hamid Eslami, Ahmed M. Eltawil
VTC Fall2
2006 Implementation of a carrier frequency recovery loop for MIMO-CDMA systems
abstract
This paper presents simulation and implementation results of a fine frequency tracking loop optimized for multiple-input multiple-output (MIMO), code division multiple access (CDMA) receivers. The proposed tracking loop exploits spatial diversity and multi-threshold recovery schemes to improve the loop's robustness, convergence time and accuracy. Comprehensive simulation results in a frequency-selective Rayleigh fading channel are presented, as well as implementation results using a 0.18 um CMOS technology
Hamid Eslami, Ahmed M. Eltawil
WCNC2
2005 Wireless field trial results of a high hopping rate FHSS-FSK testbed
abstract
This paper presents a complete study and characterization of a real-time frequency-hopped, frequency shift-keyed testbed capable of transmitting data at 160 kb/s, with hopping rates of up to 80 Khops/s operating in the 900 MHz band. The system provides the highest hopping rate reported to date and sets a new trend for FHSS communications with superior low probability of interception/detection and anti-jamming (LPI/LPD/AJ) capabilities. The architecture features a direct digital frequency synthesizer to enable high-rate hopping, and a frequency correlator-based demodulator, plus all digital timing and frequency recovery algorithms to minimize complexity. Furthermore, single sideband modulation was used to achieve spectral efficiency. The testbed is software configured and provides the user with full control over the diversity combining techniques, symbol interleaving, packet structure, and acquisition protocols. A total of 5850 independent experiments were carried out under various receiver configurations and wireless environments. The results underscore the dramatic potential for a system that optimally combines high-rate hopping, interleaving, and equal gain combining to combat severe propagation conditions, including multipath fading and intentional jamming.
Danijela Cabric, Ahmed M. Eltawil, Hanli Zou, Sumit Mohan, Babak Daneshrad
IEEE J. Sel. Areas Commun.2
2003 Modified all digital timing tracking loop for wireless applications
abstract
In this paper, a novel architecture for a fine timing-tracking loop is presented. The loop is a second order feedback loop utilizing an interpolating filter to relax the requirements on the analog to digital section of the receiver. A complete analysis of the interpolating filter, loop filter along with a system simulation based on direct sequence spread spectrum are presented.
Ahmed M. Eltawil, Babak Daneshrad
ICC1
2002 Interpolation based direct digital frequency synthesis for wireless communications
abstract
In this paper, a compact architecture for direct digital frequency synthesis (DDFS) is presented. It uses a smaller lookup table for sine and cosine functions compared to existing architectures, with minimal hardware overhead. The computation of the sinusoidal values is performed by a parabolic interpolation structure, thus only interpolation coefficients need to be stored in the read-only memory (ROM). A DDFS with 64 dBc SFDR, 10-bit output resolution and 32 bit phase accumulator requires only 104 bits of ROM storage. The ROM size is consistently less than 1 Kbits for SFDR up to 85 dBc.
Ahmed M. Eltawil, Babak Daneshrad
WCNC1