EDBT 2026 Demo / reviewers in the wild / expert
Mohammad M. Mansour
dblp:59/3642
· DBLP profile ↗
60ranked-venue papers
17as first author
9since 2021 · last 2026
0000-0002-8316-1330ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 31 · 7 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-authorSystems, architecture and hardware · 8 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Partially Polarized Polar Codes: A New Design for 6G Control Channels
Arman Fazeli, Mohammad M. Mansour, Louay M. A. Jalloul |
ICC | 2 |
| 2025 | Multi-Agent DRL for Distributed Codebook Design in RIS-Aided Cell-Free Massive MIMO NetworksabstractThis paper proposes an innovative approach for enhancing network capacity and coverage by integrating cell-free massive multiple-input multiple-output (CF-mMIMO) networks with reconfigurable intelligent surfaces (RISs). A significant challenge in leveraging RIS-assisted CF-mMIMO lies in the cooperative beam training across multiple access points (APs) and RISs, complicated by the passive nature of reflective elements and the complexity channel state information (CSI) acquisition in millimeter wave mMIMO systems. To address these challenges, we develop a multi-agent deep reinforcement learning (MA-DRL) framework that jointly designs beamforming and reflection codebooks for distributed APs and RISs, eliminating the need for CSI and relying solely on received power measurements feedback. The joint beamforming and reflection codebook design problem is decomposed into two sub-problems: one for beam codebook design at APs and another for sequential reflection codebook design at RISs. We employ transfer learning to speed up learning convergence and reduce computational complexity for training multiple RISs. Additionally, we introduce an AP and RIS selection scheme that improves overall energy efficiency and reduces backhaul overhead. Extensive simulations demonstrate that our proposed MA-DRL approach curtails number of beams significantly, thereby outperforming the widely adopted discrete Fourier transform (DFT) codebooks by achieving an 84% reduction in beam training overhead. Our findings suggest that increasing the number of passive RISs allows putting more APs into idle mode, leading to substantial savings in hardware and energy costs. Asmaa Abdallah, Abdulkadir Celik, Mohammad M. Mansour, Ahmed M. Eltawil |
IEEE Trans. Commun. | 3 |
| 2024 | Multi-Agent Deep Reinforcement Learning for Beam Codebook Design in RIS-Aided SystemsabstractReconfigurable intelligent surfaces (RISs) play a vital role in future wireless systems with the capability of enhancing propagation environments by intelligently reflecting the signals toward the target receivers. However, optimal tuning of the phase shifters at the RIS is challenging due to the passive nature of reflective elements and the high complexity of acquiring channel state information (CSI). Furthermore, the joint active beamforming and RIS reflection beam design is a tedious task due to the high computational complexity and the dynamic nature of the wireless environment. Today’s cellular networks establish data transmission by relying on pre-defined generic beamforming codebooks, which are neither site-specific nor adaptive to the changes in the wireless environment. Moreover, identifying the best beam is typically performed using an exhaustive search approach that prohibits the use of large codebook sizes due to the resulting high beam training overhead. Depending merely on the binary received signal strength, this work develops a multi-agent deep reinforcement learning (MA-DRL) framework that jointly designs the active and the passive reflection beam codebooks for the BS and the RIS, reflectively. To accelerate learning convergence and reduce the search space, the proposed model divides the RIS into multiple partitions and associates beam patterns to the surrounding environments with low computational complexity. Moreover, a hierarchical beam training solution is proposed to further reduce the beam training overhead of the single-beam training approach. Simulation results show that the proposed MA-DRL approach can provide a 97% beam training overhead reduction over the discrete Fourier transform (DFT) codebook. Asmaa Abdallah, Abdulkadir Celik, Mohammad M. Mansour, Ahmed M. Eltawil |
IEEE Trans. Wirel. Commun. | 3 |
| 2024 | Lightweight and secure cipher scheme for multi-homed systems
Hassan N. Noura, Reem Melki, Mohammad M. Mansour, Ali Chehab |
Wirel. Networks | 3 |
| 2023 | Deep Reinforcement Learning Based Beamforming Codebook Design for RIS-aided mmWave SystemsabstractReconfigurable intelligent surfaces (RISs) are envisioned to play a pivotal role in future wireless systems with the capability of enhancing propagation environments by intelligently reflecting the signals toward the target receivers. However, the optimal tuning of the phase shifters at the RIS is a challenging task due to the passive nature of reflective elements and the high complexity of acquiring channel state information (CSI). Conventionally, wireless systems rely on pre-defined reflection beamforming codebooks for both initial access and data transmission. However, these existing pre-defined codebooks are commonly not adaptive to the environments. Moreover, identifying the best beam is typically performed using an exhaustive search that leads to high beam training overhead. To address these issues, this paper develops a multi-agent deep reinforcement learning framework that learns how to jointly optimize the active beamforming from the BS and the RIS-reflection beam codebook relying only on the received power measurements. To accelerate learning convergence and reduce the search space, the proposed model divides the RIS into multiple partitions and associates beam patterns to the surrounding environments with low computational complexity. Simulation results show that the proposed learning framework can learn optimized active BS beamforming and RIS reflection codebook. For instance, the proposed MA-DRL approach with only 6 beams outperforms a 256-beam discrete Fourier transform (DFT) codebook with a 97% beam training overhead reduction. Asmaa Abdallah, Abdulkadir Celik, Mohammad M. Mansour, Ahmed M. Eltawil |
CCNC | 3 |
| 2023 | RIS-Aided mmWave MIMO Channel Estimation Using Deep Learning and Compressive SensingabstractReconfigurable intelligent surface (RIS) assisted wireless systems require accurate channel state information (CSI) to control wireless channels and improve both the bandwidth and energy efficiency. However, CSI acquisition is non-trivial for two reasons: 1) the passive nature of RIS does not allow transceiving and processing pilot signals, and 2) the dimensions of the cascaded channel between transceivers increases with the large number of RIS elements, which yields high training overhead and computational complexity. While prior art has mainly focused on frequency-flat channel estimation, this paper proposes novel data-driven and compressive sensing based approaches for estimating both frequency-flat and frequency-selective cascaded channels of RIS-assisted multi-user millimeter-wave large multiple input multiple output (MIMO) systems with limited training overhead. The proposed methods exploit the common sparsity property among the different subcarriers and the double-structured sparsity property of the angular cascaded channel matrices as different angular cascaded channels observed by different users share completely common non-zero rows and user-specific column supports. The proposed data-driven cascaded channel estimation approaches use denoising neural networks to accurately detect channel supports. Alternatively, when data-training capabilities are not available, the compressive sensing based orthogonal matching pursuit (OMP) approach relies on sparsity properties and applies simultaneous OMP to detect the channel supports. Simulation results show that the pilot overhead required by the proposed scheme is lower than existing schemes. When compared to other OMP approaches that achieve an NMSE gap of 5 to 6 dB with respect to the Oracle least square lower bound, the proposed algorithms reduce the lower bound gap to only 1 dB, while reducing complexity by more than two orders of magnitude. Asmaa Abdallah, Abdulkadir Celik, Mohammad M. Mansour, Ahmed M. Eltawil |
IEEE Trans. Wirel. Commun. | 3 |
| 2022 | Deep-Learning Based Channel Estimation for RIS-Aided mmWave Systems with Beam SquintabstractReconfigurable intelligent surface (RIS) assisted wireless systems require accurate channel state information (CSI) to control wireless channels and improve overall network performance. However, CSI acquisition is non-trivial due to the passive nature of RIS, and the dimensions of the cascaded channel between transceivers increase with the large number of RIS elements, which requires high training overhead. Prior art has considered frequency-selective channel estimation without considering the beam squint effect in wideband systems, severely degrading channel estimation performance. This paper proposes a novel data-driven approach for estimating wideband cascaded channels of RIS-assisted multi-user millimeter-wave massive multiple-input multiple-output (MIMO) systems with limited training overhead, explicitly considering the effect of beam squint. To circumvent the beam squint effect, the proposed method exploits the common sparsity property among the different subcarriers as well as the double-structured sparsity property of the users’ angular cascaded channel matrices. The proposed data-driven cascaded channel estimation approach exploits denoising neural networks to detect channel supports accurately. Compared to beam squint effect agnostic traditional orthogonal matching pursuit (OMP) approaches, the proposed data-driven approach achieves 5-6dB less normalized mean square error (NMSE) and reduces the lower bound gap to only 1dB for the oracle least-square benchmark. Asmaa Abdallah, Abdulkadir Celik, Mohammad M. Mansour, Ahmed M. Eltawil |
ICC | 3 |
| 2022 | Deep Learning-Based Frequency-Selective Channel Estimation for Hybrid mmWave MIMO SystemsabstractMillimeter wave (mmWave) massive multiple-input multiple-output (MIMO) systems typically employ hybrid mixed signal processing to avoid expensive hardware and high training overheads. However, the lack of fully digital beamforming at mmWave bands imposes additional challenges in channel estimation. Prior art on hybrid architectures has mainly focused on greedy optimization algorithms to estimate frequency-flat narrowband mmWave channels, despite the fact that in practice, the large bandwidth associated with mmWave channels results in frequency-selective channels. In this paper, we consider a frequency-selective wideband mmWave system and propose two deep learning (DL) compressive sensing (CS) based algorithms for channel estimation. The proposed algorithms learn critical apriori information from training data to provide highly accurate channel estimates with low training overhead. In the first approach, a DL-CS based algorithm simultaneously estimates the channel supports in the frequency domain, which are then used for channel reconstruction. The second approach exploits the estimated supports to apply a low-complexity multi-resolution fine-tuning method to further enhance the estimation performance. Simulation results demonstrate that the proposed DL-based schemes significantly outperform conventional orthogonal matching pursuit (OMP) techniques in terms of the normalized mean-squared error (NMSE), computational complexity, and spectral efficiency, particularly in the low signal-to-noise ratio regime. When compared to OMP approaches that achieve an NMSE gap of$\mathrm {\{4-10\}\,\,dB}$with respect to the Cramer Rao Lower Bound (CRLB), the proposed algorithms reduce the CRLB gap to only$\mathrm {\{1-1.5\}\,\,dB}$, while reducing complexity by two orders of magnitude. Asmaa Abdallah, Abdulkadir Celik, Mohammad M. Mansour, Ahmed M. Eltawil |
IEEE Trans. Wirel. Commun. | 3 |
| 2021 | Low-Complexity Soft-Output MIMO Detectors Based on Optimal Channel PuncturingabstractChannel puncturing transforms a multiple-input multiple-output (MIMO) channel into a sparse lower-triangular form using the so-called WL decomposition scheme in order to reduce tree-based detection complexity. We propose computationally efficient soft-output detectors based on two forms of channel puncturing: augmented and two-sided. The augmented WL detector (AWLD) employs a punctured channel derived by triangularizing the true channel in augmented form, followed by left-sided Gaussian elimination. The two-sided WL detector (dubbed WLZ) employs right-sided reduction and left-sided elimination to puncture the channel. We prove that augmented channel puncturing is optimal in maximizing the lower-bound on the achievable information rate (AIR) based on a new mismatched detection model. We show that the AWLD decomposes into an MMSE prefilter and channel gain compensation stages, followed by a regular WL detector (WLD) that computes least-squares soft-decision estimates. Similarly, WLZ decomposes into a pre-processing reduction step followed by WLD. AWLD attains the same performance as the existing AIR-based partial marginalization (PM) detector, but with less computational complexity. We empirically show that WLZ attains the best complexity-performance tradeoff among tree-based detectors. Mohammad M. Mansour |
IEEE Trans. Wirel. Commun. | 1 |
| 2020 | Optimal Augmented-Channel Puncturing for Low-Complexity Soft-Output MIMO DetectorsabstractWe propose a computationally-efficient soft-output detector for multiple-input multiple-output channels based on augmented channel puncturing in order to reduce tree processing complexity. The proposed detector, dubbed augmented WL detector (AWLD), employs a punctured channel with a special structure derived by triangulizing the original channel in augmented form, followed by Gaussian elimination. We prove that these punctured channels are optimal in maximizing the lower-bound on the achievable information rate (AIR) based on a newly proposed mismatched detection model. We show that the AWLD decomposes into a minimum mean-square error (MMSE) prefilter and channel-gain compensation stages, followed by a regular unaugmented WL detector (WLD). It attains the same performance as the existing AIR partial marginalization (AIRPM) detector, but with much simpler processing. Mohammad M. Mansour |
ICC | 1 |
| 2020 | An Optimized VLSI Implementation of an IEEE 802.11n/ac/ax LDPC DecoderabstractThis paper proposes optimization techniques for multi-Gbps low-power VLSI implementation of IEEE 802.11n/ac/ax (WiFi) LDPC decoders. The IEEE 802.11n/ac/ax standard features Quasi-Cyclic LDPC (QC-LDPC) codes with modular decoder structure composed of arrays of memory blocks, barrel shifter networks, adders, and check-node units (CNUs). To achieve multi-Gbps throughput performance and high energy-efficiency while maintaining a small decoder footprint, careful implementation of these modules along with effective fixed-point analysis is required. This paper proposes techniques for optimized implementation and proper bit-width selection of these modules. These techniques are then employed to design a fully-pipelined IEEE 802.11n/ac/ax standard compliant LDPC decoder. The design is synthesized in a 40nm standard CMOS process. The synthesized decoder occupies an area of 0.71 mm2, runs at a frequency of 562 MHz, attains a peak throughput of 11.4 Gbps, and achieves an energy-efficiency of 12.5 pJ/bit. The presented decoder outperforms the best-reported decoders in the literature in terms of throughput/area and energy-efficiency, for IEEE 802.11n/ac/ax LDPC codes. Saleh Usman, Mohammad M. Mansour |
ISCAS | 2 |
| 2020 | Efficient Angle-Domain Processing for FDD-Based Cell-Free Massive MIMO SystemsabstractCell-free massive MIMO communications is an emerging network technology for 5G wireless communications wherein distributed multi-antenna access points (APs) serve many users simultaneously. Most prior work on cell-free massive MIMO systems assume time-division duplexing mode, although frequency-division duplexing (FDD) systems dominate current wireless standards. The key challenges in FDD massive MIMO systems are channel-state information (CSI) acquisition and feedback overhead. To address these challenges, we exploit the so-called angle reciprocity of multipath components in the uplink and downlink, so that the required CSI acquisition overhead scales only with the number of served users, and not the number of AP antennas nor APs. We propose a low complexity multipath component estimation technique and present linear angle-of-arrival (AoA)-based beamforming/combining schemes for FDD-based cell-free massive MIMO systems. We analyze the performance of these schemes by deriving closed-form expressions for the mean-square-error of the estimated multipath components, as well as expressions for the uplink and downlink spectral efficiency. Using semi-definite programming, we solve a max-min power allocation problem that maximizes the minimum user rate under per-user power constraints. Furthermore, we present a user-centric (UC) AP selection scheme in which each user chooses a subset of APs to improve the overall energy efficiency of the system. Simulation results demonstrate that the proposed multipath component estimation technique outperforms conventional subspace-based and gradient-descent based techniques. We also show that the proposed beamforming and combining techniques along with the proposed power control scheme substantially enhance the spectral and energy efficiencies with an adequate number of antennas at the APs. Asmaa Abdallah, Mohammad M. Mansour |
IEEE Trans. Commun. | 2 |
| 2020 | Physical layer security schemes for MIMO systems: an overview
Reem Melki, Hassan N. Noura, Mohammad M. Mansour, Ali Chehab |
Wirel. Networks | 3 |
| 2019 | Design and realization of efficient & secure multi-homed systems based on random linear network coding
Hassan N. Noura, Reem Melki, Mohammad M. Mansour, Ali Chehab |
Comput. Networks | 3 |
| 2019 | An Efficient OFDM-Based Encryption Scheme Using a Dynamic Key ApproachabstractPhysical layer (PHY) security has emerged as a promising methodology for securing current and future networks that employ orthogonal frequency-division multiplexing (OFDM) technology. OFDM is the basic building block for multicarrier modulation in most contemporary networks such as vehicular ad hoc networks, Internet of Things (IoT), as well as 4G/5G systems. Most existing OFDM-based security solutions lack the notion of secrecy and dynamicity when combining a secret key with random information extracted from the physical channel. Yet, some solutions perform encryption preinverse fast Fourier transform and some postinverse fast Fourier transform, without clear guidelines concerning the impact on performance and security. In this paper, OFDM-based encryption schemes at the PHY are investigated, analyzed, and weaknesses are identified. It is shown that encryption in the frequency domain slightly mitigates the effects of channel fading and improves the bit error-rate performance. On the other hand, time-domain encryption is shown to be more secure. Furthermore, a dynamic secret key approach that enhances the security level of OFDM-based encryption schemes, in addition to a new technique for updating cipher primitives for input OFDM symbols or frames, are proposed. These schemes are shown to strike a good balance between performance and security robustness as demonstrated through experimental simulations. Reem Melki, Hassan N. Noura, Mohammad M. Mansour, Ali Chehab |
IEEE Internet Things J. | 3 |
| 2019 | A Physical Encryption Scheme for Low-Power Wireless M2M Devices: a Dynamic Key Approach
Hassan N. Noura, Reem Melki, Ali Chehab, Mohammad M. Mansour |
Mob. Networks Appl. | 4 |
| 2019 | Lightweight, dynamic and efficient image encryption scheme
Hassan N. Noura, Ali Chehab, Mohamad Noura, Raphaël Couturier, Mohammad M. Mansour |
Multim. Tools Appl. | 5 |
| 2019 | Efficient and secure cipher scheme for multimedia contents
Hassan N. Noura, Mohamad Noura, Ali Chehab, Mohammad M. Mansour, Raphaël Couturier |
Multim. Tools Appl. | 4 |
| 2018 | Channel-Punctured Large MIMO DetectionabstractLow-complexity data detectors targeted for large multiple-input multiple-output (MIMO) systems are considered. By systematically puncturing the channel matrix to have a specific structure, the complexity of standard non-linear detectors can be significantly reduced. The performance of these detectors is characterized and analyzed mathematically, and bounds on the achievable diversity gain and probability of bit error are derived. It is shown that puncturing does not negatively impact the receive diversity gain in hard-output detectors. Moreover, in soft-output detection, significant performance gains are attainable by ordering the layer of interest to be at the root when puncturing the channel. The proposed schemes scale up efficiently both in the number of antennas and constellation size. Hadi Sarieddeen, Mohammad M. Mansour, Ali Chehab |
ISIT | 2 |
| 2018 | Efficient and Secure Physical Encryption Scheme for Low-Power Wireless M2M DevicesabstractRecently, physical layer security has emerged as a promising security scheme for wireless networks, in contrast to traditional solutions that mainly rely on upper network layers. As such, several physical layer encryption algorithms that benefit from the random characteristics of physical channels have appeared in the literature. However, the majority of these schemes lack the notion of secrecy and dynamicity. In this paper, we focus on enhancing the physical layer encryption for wireless machine-to-machine devices, which share the same channel, with the aim of striking a good balance between performance and security robustness. The main idea is to perform encryption at the physical layer after symbol modulation. The cipher scheme is based on one round and one operation that reduces the encryption overhead in terms of latency and required resources. Furthermore, we propose a dynamic key approach that combines a pre-shared/stored secret key with a dynamic nonce extracted from the channel information to generate a dynamic key. The main advantage of the dynamic key approach is that it achieves a high-security level with minimal overhead. The dynamic key can be changed frequently upon any change in channel parameters or upon starting a new session. In addition to data encryption, a preamble encryption scheme is also proposed to prevent unauthorized synchronization or channel estimation by illegitimate users. Finally, security and performance analyses are performed to demonstrate the validity, efficiency and robustness of the proposed approach. Hassan N. Noura, Reem Melki, Ali Chehab, Mohammad M. Mansour, Steven Martin 0001 |
IWCMC | 4 |
| 2018 | Joint channel allocation and power control for D2D communications using stochastic geometryabstractDevice-to-Device (D2D) communication is a viable network technology that can potentially enhance the spectral and energy efficiency of cellular networks. To exploit this benefit in D2D-underlaid cellular networks, the co-channel interference between D2D and cellular users should be properly managed. In this paper, we propose a joint channel allocation (CA) and power control (PC) scheme to mitigate interference in a D2D underlaid cellular system modeled as a random network using stochastic geometry. The novel aspect of the proposed CA scheme is that it enables D2D links to share resources with multiple cellular users as opposed to one as previously considered in the literature. The PC scheme compensates for large-scale path-loss effects by employing distance-dependent path-loss parameters with an estimation error margin. Closed-form expressions for the coverage probability of cellular links, D2D links, and the sum rate of the D2D links are derived in terms of the allocated power, density of the D2D links, and the path-loss exponent. Simulation results demonstrate an enhancement of 10%-40% for the cellular and D2D coverage probabilities, and 35% for spectral efficiency. Asmaa Abdallah, Mohammad M. Mansour, Ali Chehab |
WCNC | 2 |
| 2018 | A fairness-based congestion control algorithm for multipath TCPabstractMultipath TCP (MP-TCP) has been introduced as an extension to the legacy TCP transport protocol to support communication through multiple paths under a single connection session. The target is to improve both resource utilization and connection robustness. Several congestion control algorithms (CCAs) have emerged in the literature to adapt subflow rates to congestion conditions on the various paths without negatively impacting competing single-path TCP sources. The challenge is to provide a trade-off among three factors, namely, fairness, responsiveness, and window oscillation. In this paper, we propose a new fairness-based CCA (FCCA) based on the fluid model that improves fairness without degrading the other two metrics. The FCCA tracks the performance on each route and dynamically adapts the respective congestion windows, enhancing the overall performance. The proposed algorithm is implemented in a Linux kernel. Simulation results demonstrate that FCCA is capable of achieving almost maximal fairness (98%) while maintaining responsiveness, unlike existing CCAs. Reem Melki, Mohammad M. Mansour, Ali Chehab |
WCNC | 2 |
| 2018 | One round cipher algorithm for multimedia IoT devices
Hassan N. Noura, Ali Chehab, Lama Sleem, Mohamad Noura, Raphaël Couturier, Mohammad M. Mansour |
Multim. Tools Appl. | 6 |
| 2018 | A dynamic approach for a lightweight and secure cipher for medical images
Mohamad Noura, Hassan N. Noura, Ali Chehab, Mohammad M. Mansour, Lama Sleem, Raphaël Couturier |
Multim. Tools Appl. | 4 |
| 2018 | A new efficient lightweight and secure image cipher scheme
Hassan N. Noura, Lama Sleem, Mohamad Noura, Mohammad M. Mansour, Ali Chehab, Raphaël Couturier |
Multim. Tools Appl. | 4 |
| 2018 | Power Control and Channel Allocation for D2D Underlaid Cellular NetworksabstractDevice-to-Device (D2D) communications underlaying cellular networks is a viable network technology that can potentially increase spectral utilization and improve power efficiency for proximity-based wireless applications and services. However, a major challenge in such deployment scenarios is the interference caused by D2D links when sharing the same resources with cellular users. In this paper, we propose a channel allocation (CA) scheme together with a set of three power control (PC) schemes to mitigate interference in a D2D underlaid cellular system modeled as a random network using the mathematical tool of stochastic geometry. The novel aspect of the proposed CA scheme is that it enables D2D links to share resources with multiple cellular users as opposed to one as previously considered in the literature. Moreover, the accompanying distributed PC schemes further manage interference during link establishment and maintenance. The first two PC schemes compensate for large-scale path-loss effects and maximize the D2D sum rate by employing distance-dependent path-loss parameters of the D2D link and the base station, including an error estimation margin. The third scheme is an adaptive PC scheme based on a variable target signal-to-interference-plus-noise ratio, which limits the interference caused by D2D users and provides sufficient coverage probability for cellular users. Closed-form expressions for the coverage probability of cellular links, D2D links, and sum rate of D2D links are derived in terms of the allocated power, density of D2D links, and path-loss exponent. The impact of these key system parameters on network performance is analyzed and compared with previous work. Simulation results demonstrate an enhancement in cellular and D2D coverage probabilities, and an increase in spectral and power efficiency. Asmaa Abdallah, Mohammad M. Mansour, Ali Chehab |
IEEE Trans. Commun. | 2 |
| 2018 | Large MIMO Detection Schemes Based on Channel Puncturing: Performance and Complexity AnalysisabstractA family of low-complexity detection schemes based on channel matrix puncturing targeted for large multiple-input multiple-output (MIMO) systems is proposed. It is well known that the computational cost of MIMO detection based on QR decomposition is directly proportional to the number of nonzero entries involved in back-substitution and slicing operations in the triangularized channel matrix, which can be too high for low-latency applications involving large MIMO dimensions. By systematically puncturing the channel to have a specific structure, it is demonstrated that the detection process can be accelerated by employing standard schemes, such as chase detection, list detection, nulling-and-cancellation detection, and sub-space detection on the transformed matrix. The performance of these schemes is characterized and analyzed mathematically, and bounds on the achievable diversity gain and probability of bit error are derived. Surprisingly, it is shown that puncturing does not negatively impact the receive diversity gain in hard-output detectors. The analysis is extended to soft-output detection when computing per-layer bit log-likelihood ratios; it is shown that significant performance gains are attainable by ordering the layer of interest to be at the root when puncturing the channel. Simulations of coded and uncoded scenarios certify that the proposed schemes scale up efficiently both in the number of antennas and constellation size, as well as in the presence of correlated channels. In particular, soft-output per-layer sub-space detection is shown to achieve a 2.5 dB signal-to-noise ratio gain at 10-4bit error rate in 256-quadratic-amplitude modulation 16 × 16 MIMO, while saving 77% of nulling-and-cancellation computations. Hadi Sarieddeen, Mohammad M. Mansour, Ali Chehab |
IEEE Trans. Commun. | 2 |
| 2017 | Hard-output chase detectors for large MIMO: BER performance and complexity analysisabstractIn this paper, a family of cost-efficient hard-output detection algorithms for large multiple-input multiple-output (MIMO) systems is proposed. The schemes employ punctured QR decomposition (QRD) instead of regular QRD to reduce complexity. The bit error rate performance is studied analytically, where it is shown that channel matrix puncturing does not affect the diversity gain of the detectors. Through empirical simulations, the proposed schemes are shown to achieve significant reductions in computational complexity with graceful performance degradation. In particular, at an SNR cost of 4dB, 77% of complex multiplications in nulling and cancellation are saved in 16 × 16 MIMO, while 30% of multiplications are saved at a 2dB cost in 4×4 MIMO. The savings can reach 94% in 64×64 MIMO. Hadi Sarieddeen, Mohammad M. Mansour, Ali Chehab |
PIMRC | 2 |
| 2017 | A Distance-Based Power Control Scheme for D2D Communications Using Stochastic GeometryabstractDevice-to-Device (D2D) communication is a promising technology that can potentially enhance the spectral and energy efficiency of cellular networks. To exploit this benefit in D2D-underlaid cellular networks, the co-channel interference between D2D and cellular users should be properly managed. In this paper, we propose a distributed power control scheme to mitigate interference in a D2D underlaid cellular system modeled as a random network using the mathematical tool of stochastic geometry. The proposed PC scheme compensates for large-scale path-loss effects by employing distance-dependent path-loss parameters of the D2D link and the base station, including an estimation error margin. Closed-form expressions for the coverage probability of cellular links, D2D links, and the sum rate of D2D links are derived in terms of the allocated power, density of D2D links, and path-loss exponent. The coverage performance of both cellular and D2D users is analyzed, and the analytical results are validated through simulations. Experimental results demonstrate the efficacy and advantages of our proposed scheme over other schemes by an enhancement of 20%-30% for the cellular and D2D coverage probabilities, and an increase in spectral efficiency by 60%. Asmaa Abdallah, Mohammad M. Mansour, Ali Chehab |
VTC Fall | 2 |
| 2016 | Interlaced Column-Row Message-Passing Schedule for Decoding LDPC CodesabstractThis paper investigates efficient decoding algorithms for LDPC codes. Alternating column-row message-passing (ACRMP) and Interlaced column-row message-passing (I-CRMP) schedules for decoding of LDPC codes are proposed and investigated in this work. Existing serial scheduling schemes for LDPC decoding are based either on column message-passing (MP) or row MP, and roughly converge twice as fast as Gallager's flooding-based MP schedule at high signal-to-noise ratio (SNR). To further accelerate the convergence speed of serial decoders, hybrid column-row MP schedules that perform multiple message passes between check and variable nodes within or between iterations are proposed. Our proposed I-CRMP schedule converges in less than half the number of iterations compared to the best existing serial decoding schedules. Compared to column MP, the added complexity of this scheme is proportional only to check-node-degree times more additions at variable nodes. This increase in complexity is moderate compared to the convergence acceleration factor that the scheme achieves. Superior performance of the proposed I-CRMP scheme is confirmed by decoding randomly generated as well as IEEE 802.11n/ac LDPC codes. Saleh Usman, Mohammad M. Mansour, Ali Chehab |
GLOBECOM | 2 |
| 2016 | Efficient subspace detection for high-order MIMO systemsabstractIn this paper, low-complexity multiple-input multiple-output (MIMO) subspace detection schemes are studied, which decompose a channel into multiple decoupled streams to be detected disjointly. Existing schemes require a number of matrix decomposition operations equal to the number of detected streams, which is computationally complex, especially in high-order MIMO systems. We propose two computationally efficient detection algorithms, based on a preprocessing stage that consists of special layer ordering, followed by permutation-robust QR decomposition (QRD) and elementary matrix operations. The algorithms are illustrated in the context of a 4-layer MIMO system, and their complexity is studied. Simulations demonstrate that using the proposed scheme, the QRD overhead is reduced by almost 50% for very high order MIMO, without incurring any performance degradation. Hadi Sarieddeen, Mohammad M. Mansour, Ali Chehab |
ICASSP | 2 |
| 2016 | Efficient near optimal joint modulation classification and detection for MU-MIMO systemsabstractOptimum data detection schemes for dual layer multi-user multiple-input multiple-output (MU-MIMO) systems are studied. A joint maximum likelihood (ML) modulation classification (MC) of the co-scheduled user and data detection receiver is developed. By expanding the max-log-maximum-a-posteriori MC approach to include distances of counter ML hypothesis symbols, the decision metric for MC is shown to be an accumulation over a set of tones of Euclidean distance computations also used by the ML detector for bit log-likelihood ratio soft decision generation. With a small complexity overhead, the proposed approach achieves near-optimal performance. An efficient hardware architecture is presented for the proposed approach. Hadi Sarieddeen, Mohammad M. Mansour, Louay M. A. Jalloul, Ali Chehab |
ICASSP | 2 |
| 2016 | Enhanced low-complexity layer-ordering for MIMO sphere detectorsabstractIn this paper, optimum soft-output (SO) multiple-input multiple-output (MIMO) sphere detectors (SDs) are studied. Noting that ordering the channel matrix columns plays an important role in reducing the tree-search complexity of a SD, we propose an optimized layer-ordering scheme based on the minimum cumulative residual criterion. The proposed scheme is studied in the context of a 4 × 4 MIMO system, and a low-complexity dataflow architecture is proposed. The implementation employs a permutation-robust QR decomposition (PR-QRD) scheme, based on the modified Gram-Schmidt orthogonalization procedure. Simulations demonstrate that using the proposed scheme, the node count of a SO MIMO SD is reduced by one order of magnitude, while the QRD overhead is reduced by more than 25% in computations and 36% in time, without incurring any performance degradation. Hadi Sarieddeen, Mohammad M. Mansour |
ICC | 2 |
| 2016 | Efficient near-optimal 8×8 MIMO detectorabstractIn this paper, a low-complexity near-optimal detector for 8-layer MIMO systems is proposed. The detector employs subspace detection schemes, which decompose a spacially multiplexed MIMO channel into multiple decoupled streams to be detected separately. Several existing subspace detection algorithms are studied, all of which require a significant overhead for channel matrix decomposition. We propose computationally efficient schemes based on special layer ordering, followed by permutation-robust QR Decomposition (PR-QRD) using the modified Gram-Schmidt orthogonalization procedure, and elementary matrix operations. A hardware architecture is proposed, which allows building an 8-layer detector from 4-layer and 2-layer constituent detector blocks. Simulations demonstrate that using the proposed scheme, the QRD overhead is reduced by 30%, without incurring any performance degradation. Hadi Sarieddeen, Mohammad M. Mansour, Ali Chehab |
WCNC | 2 |
| 2016 | Low-complexity joint modulation classification and detection in MU-MIMOabstractIn this paper, dual-layer multi-user multiple-input multiple-output systems are studied. Building on the low-complexity layered orthogonal lattice detector (LC-LORD), an efficient sub-optimal joint modulation classification (MC) of the co-scheduled user and data detection receiver is developed. By adjusting the Max-Log-Maximum-a-Posteriori MC approach to the limitations of LC-LORD, and expanding it to include distances of counter maximum likelihood hypothesis symbols, the decision metric for MC is shown to be an accumulation over a set of tones of Euclidean distance computations also used by the LC-LORD detector for bit log-likelihood ratio soft decision generation. Simulations demonstrate that with a small complexity overhead, the proposed approaches achieve near interference-aware performance. An efficient hardware implementation scheme is presented. Hadi Sarieddeen, Mohammad M. Mansour, Louay M. A. Jalloul, Ali Chehab |
WCNC | 2 |
| 2016 | Inter-Frame Coding For Broadcast CommunicationabstractA novel inter-frame coding approach to the problem of varying channel-state conditions in broadcast wireless communication is developed in this paper; this problem causes the appropriate code-rate to vary across different transmitted frames and different receivers as well. The main aspect of the proposed approach is that it incorporates an iterative rate-matching process into the decoding of the received set of frames, such that: throughout inter-frame decoding, the code-rate of each frame is progressively lowered to or below the appropriate value, and prior to applying or re-applying conventional physical-layer channel decoding on it. This iterative rate-matching process is asymptotically analyzed in this paper. It is shown to be optimal, in the sense defined in the paper. Consequently, the data-rates achievable by the proposed scheme are derived. Overall, it is concluded that, compared to the existing solutions, inter-frame coding presents a better complexity versus data-rate tradeoff. In terms of complexity, the overhead of inter-frame decoding includes operations that are similar in type and scheduling to those employed in the relatively-simple iterative erasure decoding. In terms of data-rates, compared to the state-of-the-art two-stage scheme involving both error-correcting and erasure coding, inter-frame coding increases the data-rate by a factor that reaches up to 1.55×. Hady Zeineddine, Mohammad M. Mansour |
IEEE J. Sel. Areas Commun. | 2 |
| 2015 | A Low-Complexity PAPR Reduction Technique for LTE-Advanced Uplink with Carrier AggregationabstractIn LTE-Advanced, carrier aggregation (CA) is a key enabling technique to increase the peak data rates of users and enhance the mobility management in heterogeneous networks (HetNets) using dual connectivity solutions. Several CA schemes have been proposed with a maximum of five LTE Release 8 component carriers (CCs) where the NxSC-FDMA has been chosen as the bandwidth extension scheme for the uplink. This enables the extension of the bandwidth allocated to a user up to 100 MHz while maintaining backward compatibility with LTE release 8 legacy users. However, CA leads to a severe increase in the peak-to-average-power ratio (PAPR) of the aggregated time domain NxSC-FDMA signal at the user equipment (UE). This is an important issue as it affects the power amplifier (PA) efficiency and hence the coverage of transmissions. In this paper, we propose a low- complexity post-IFFT technique to reduce the PAPR of carrier aggregated NxSC-FDMA signals in the uplink of LTE-Advanced. Several case studies of CA were analyzed by using the proposed technique. Simulation results show that the PAPR improvement is 2.5 dB for the case of 2 CCs at only about 4% of the complexity required by the partial selective mapping (PSLM) technique to achieve the same PAPR reduction. Abdel-Karim Ajami, Hassan Artail, Mohammad M. Mansour |
GLOBECOM | 3 |
| 2015 | PAPR reduction in LTE-Advanced carrier aggregation using low-complexity joint interleaving techniqueabstractThe demand for high data rates in both the uplink and the downlink has motivated the use of carrier aggregation (CA) of several portions of the spectrum up to 100 MHz in LTE-Advanced, while maintaining backward compatibility with LTE release 8. One of the main practical challenges that comes with CA is the severe increase of peak-to-average-power-ratio (PAPR) of the corresponding generated time-domain OFDM signal, thus affecting the power amplifier (PA) efficiency, and hence the transmission coverage. This paper proposes a low-complexity joint interleaving technique to reduce the PAPR of carrier aggregated signals. Several CA scenarios were analyzed using our proposed technique. Simulation results demonstrate that our proposed technique can achieve the same PAPR reduction performance as that of the partial selective mapping (PSLM) technique with 66% reduction in terms of real multiplication and addition operations for the case of three aggregated component carriers (CCs). Abdel-Karim Ajami, Hassan Artail, Mohammad M. Mansour |
WCNC | 3 |
| 2015 | Soft-Output MIMO Detectors with Channel Estimation ErrorabstractNew expressions for the soft decision bit log-likelihood ratio (LLR) of a MIMO system using quadrature amplitude modulation (QAM) taking into account channel estimation error (CEE). The bit LLR for the maximum likelihood (ML) and the linear minimum mean-squared error (MMSE) receivers are derived, showing in both receivers explicit scaling of the LLR that is a function of the QAM symbol and the CEE variance. These new expressions for the LLRs are used to show that only modest improvements in the link performance are achieved relative to the LLRs that do not take into account the CEE. This indicates that separating the detector design from channel estimation does not significantly impact the system performance, which leads to simplifications in the overall receiver implementation. Louay M. A. Jalloul, Sam P. Alex, Mohammad M. Mansour |
IEEE Signal Process. Lett. | 3 |
| 2015 | A Near-ML MIMO Subspace Detection AlgorithmabstractA low-complexity MIMO detection scheme is presented that decomposes a MIMO channel into multiple decoupled subsets of streams that can be detected separately. The scheme employs QL decomposition followed by elementary matrix operations to transform the channel matrix into a generalized elementary structure matching the subsets of streams to be detected. The proposed scheme avoids matrix inversion operations, and allows subsets to overlap thus achieving better diversity gain. Simulations demonstrate that this approach performs to within a few tenths of a dB from the optimum detection algorithm. Mohammad M. Mansour |
IEEE Signal Process. Lett. | 1 |
| 2014 | Comments on "Soft Decision Metric Generation for QAM With Channel Estimation Error"abstractA log-max generalized bit log-likelihood ratio (GLLR) was derived in an earlier work of Wanget al.for soft decision decoding of quadrature amplitude modulation that takes into account channel estimation error. The results in the above paper show that the GLLR metric outperforms the conventional LLR metric that does not account for channel estimation error. In trying to reproduce the results of the above paper, our simulations indicate that their results were obtained without proper scaling of the transmit constellation. Conversely, through simulations, we demonstrate that by properly scaling the constellation to have unity average energy, the GLLR metric does not result in a noticeable advantage in the bit error rate over the conventional LLR metric for practically meaningful values of the channel estimation error. Louay M. A. Jalloul, Sam P. Alex, Mohammad M. Mansour |
IEEE Trans. Commun. | 3 |
| 2013 | On the contention-free and spread characteristics of serially-pruned interleaversabstractSerial pruning of turbo interleavers have been proposed in the literature as a simple scheme to provide more flexible codeword lengths. In this paper, we prove two important attributes about serially pruned interleavers. First, we show that serially pruned interleavers inherit the content-free property of their mother interleaver, and hence they remain parallelizable. An example serially-pruned QPP LTE interleaver is parallelized. Second, the minimum spread factor of a serially-pruned interleaver closely matches the spread factor of its mother interleaver for small pruning gaps with minimal impact on BER performance, and degrades gracefully with the pruning length. Simulation results of practical pruned LTE turbo interleavers demonstrate the graceful degradation of spread characteristics and BER performance of serially pruned interleavers. Mohammad M. Mansour |
ICASSP | 1 |
| 2013 | Fast Pruned InterleavingabstractIn this paper, computationally efficient schemes for enumerating the so-called inliers of a wide range of permutations employed in pruned variable-size (turbo) interleavers are proposed. The objective is to accelerate pruned interleaving time in turbo codes by computing a statistic known as the pruning gap that enables determining a permuted address under pruning without serially permuting all its predecessors. It is shown that for any linear or quadratic permutation, including variations such as dithered relative prime or almost regular, the pruning gap can be computed in logarithmic time. Moreover, it is shown that Dedekind sums form efficient building blocks for enumerating inliers of the widely adopted polynomial-based permutations. An efficient algorithm for computing such sums in vector form using integer operations is presented. The results are extended to 2D and higher dimensional interleavers that combine multiple permutations along all dimensions, and closed-form expressions for inliers are derived. It is shown that the inliers statistic is a linear combination of the constituent permutation inliers. A lower bound on the minimum spread of serially pruned interleavers using the inliers statistic is also derived. Moreover, it is shown that serially pruned interleavers inherit the content-free property of the mother interleaver, and hence they are parallelizable. Simulation results of practical pruned turbo interleavers demonstrate a speedup improvement of several orders of magnitude compared to serial interleaving. Mohammad M. Mansour |
IEEE Trans. Commun. | 1 |
| 2012 | A recursive algorithm for pruned bit-reversal permutationsabstractA fast recursive algorithm for pruned bit-reversal permutations is proposed. The algorithm is based on a computationally efficient scheme for evaluating a novel permutation statistic called permutation inliers that counts inlier addresses under pruning. This statistic is computed by evaluating a recursion using integer shift and add operations in logarithmic time complexity. Moreover, a parallel pruned interleaving algorithm based on computing multiple inliers in parallel is proposed. The advantages of the proposed algorithm are reduced latency and reduced memory requirements, which are describe in many signal processing and communication applications. Mohammad M. Mansour |
ICASSP | 1 |
| 2011 | Reconfigurable decoder architectures for Raptor codesabstractDecoder architectures for architecture-aware Raptor codes having regular message access-and-processing patterns are presented. Raptor codes are a class of concatenated codes composed of a fixed-rate precode and a Luby-Transform (LT) code that can be used as rate-less error-correcting codes over communication channels. In the proposed approach, the decoding procedure is mapped to row processing of a regular matrix, which adapts effectively to the code's randomness and degree-irregularity. This is achieved by 1) developing reconfigurable check node processors that attain a constant throughput while processing LT- and LDPC-nodes of varying degrees and numbers, 2) applying pseudo-random permutation on the communicated messages, and 3) computing bit-to-check messages in a serial, temporally distributed manner. A serial decoder for a rate-0.4 code implementing the proposed approach was synthesized in 65nm CMOS technology. Hardware simulations show that the decoder achieves a throughput of 22Mb/s at BER of 10-6, dissipates an average power of 222mW and occupies an area of 1.77mm2. A range of partially-parallel decoders with desired throughput can be designed by replicating the processing nodes of a serial decoder. Hady Zeineddine, Mohammad M. Mansour |
ICASSP | 2 |
| 2011 | A novel technique to measure data retention voltage of large SRAM arraysabstractThis paper presents a new technique to accurately measure the data retention voltage (DRV) of large SRAM arrays in the presence of process variations. The proposed technique relies on a built-in-self-test (BIST) unit along with a DC-DC converter. The BIST unit implements a modified version of the March C-test that accounts for data retention faults. Whereas, the DC-DC converter is used to scale down the supply voltage of the array as is done when the array is in data retention mode. The proposed technique can accurately measure the DRV to ensure the SRAM operates at its minimum energy point. The circuit was developed in 90nm technology and simulated using HSPICE. Monte-Carlo simulation of 100k samples determined the DRV as 150mV whereas the proposed technique showed that the DRV of the SRAM under test could be lowered to 80mV which would result in significant power savings. Farah B. Yahya, Mohammad M. Mansour, Ali Chehab |
ISCAS | 2 |
| 2011 | A design methodology for energy aware neural networksabstractThe increasing demand for mobile devices and high performance computing has made energy consumption a main issue in computer technology. Mobile devices require extended battery life, but the available technology still puts limits on the need for recharging the devices. High performance computing has a high price tag on energy for compute-intensive applications such as data mining. As a result, optimizations at various layers of the computer platform are becoming necessary to minimize energy usage or extend the time before a battery needs to be recharged. This paper focuses on back-propagation neural network algorithm, one of the popular compute-intensive data mining algorithms. The goal is to present a design methodology for developing an energy aware algorithm. The key idea revolves around identifying operations called kernels, which are frequently used in the algorithm, and that can be implemented in hardware. Optimizing these kernels for performance or energy would then lead to a major impact in these areas. These kernels are analyzed for their impact on the overall application energy using energy-based asymptotic analysis. The methodology then considers additional optimizations not related to kernels, but are specific to the back-propagation algorithm. Suggestions are provided to improve the performance and reduce energy consumption. Experiments show that there are significant potentials in energy reduction through the use of alternative lower energy kernels or through custom optimizations with tradeoffs in the accuracy of the results. Mehiar Dabbagh, Hazem M. Hajj, Ali Chehab, Wassim El-Hajj, Ayman I. Kayssi, Mohammad M. Mansour |
IWCMC | 6 |
| 2009 | Optimized Architecture for Computing Zadoff-Chu Sequences with Application to LTEabstractAn optimized algorithm and a corresponding reconfigurable architecture for computing Zadoff-Chu (ZC) complex sequence elements based on the CORDIC algorithm are proposed. The algorithm computes ZC-sequence elements both in time domain and frequency domain using a simple duality relationship. Algorithm transforms are employed to compute the elements recursively and eliminate multipliers with nonconstant terms. The algorithm is applied in a searcher block for detecting the physical random access channel (PRACH) in LTE PHY layer. The PRACH transmits a preamble constructed from ZC sequences to establish initial access along with uplink synchronization with the base station. A reconfigurable hardware architecture was implemented to generate these preambles on the fly with high accuracy based on the proposed algorithm, eliminating the need for storing a large number of long complex ZC sequence elements. Simulation results demonstrate that the proposed architecture is capable of achieving detection error rates for LTE PRACH that are close to ideal rates achieved using floating-point precision. Mohammad M. Mansour |
GLOBECOM | 1 |
| 2009 | A parallel architecture for 3GPP2/UMB turbo interleaversabstractIn this paper, an efficient architecture for a parallel pruned turbo interleaver for 3GPP2/UMB physical layer standard is presented. Turbo interleaving in UMB turbo codes is based on filling a 2D array row by row, interleaving each row using a linear congruential sequence, bit-reversing the order of the rows, and then reading the interleaved addresses column by column. Pruning creates a serial bottleneck since the interleaved address of a linear address x is a function of the number of pruned addresses up to x. An architecture based on the parallel lookahead pruned interleaving algorithm proposed in is presented. The algorithm breaks this dependency and interleaves any address in O(log2x) steps by enabling a parallel turbo interleaver design with a desired degree of parallelism. The architecture can be implemented efficiently in hardware using basic arithmetic building blocks. Mohammad M. Mansour |
ICASSP | 1 |
| 2009 | Parallel lookahead algorithms for pruned interleaversabstractIn this letter, the design of efficient parallel pruned channel and turbo interleavers for Ultra Mobile Broadband (UMB) physical layer standard [1] is addressed. Channel interleaving is based on a bit-reversal algorithm in which addresses are mapped from linear order into bit-reversed order. Turbo interleaving is based on filling a 2D array row by row, interleaving each row independently using a linear congruential sequence (LCS), bit-reversing the order of the rows, and then reading the interleaved addresses column by column. To accommodate for flexible codeword lengths L, interleaving is done using a mother interleaver of length M = 2n, where n is the smallest integer such that L ⩽ M, such that outlier interleaved addresses greater than L - 1 get pruned away. This pruning operation creates a serial bottleneck since the interleaved address of a linear address χ is now a function of the interleaving operation as well as the number of pruned addresses up to χ. A generic parallel lookahead pruned interleaving scheme that breaks this dependency is proposed. The efficiency of the proposed scheme is demonstrated in the context of both UMB interleavers. An iterative pruned bit-reversal algorithm that interleaves any address in O(log L) steps is presented. Moreover, an iterative pruned turbo interleaving algorithm based on LCSs that interleaves any address in O(log2L) steps is presented. Mohammad M. Mansour |
IEEE Trans. Commun. | 1 |
| 2009 | A Parallel Pruned Bit-Reversal InterleaverabstractA parallel algorithm and architecture for pruned bit-reversal interleaving (PBRI) are proposed. For a pruned interleaver of size$N$with mother interleaver size$M=2^{n} \geq N$, the proposed algorithm interleaves any number$x\in [0,N-1]$in at most$n-1$steps, as opposed to$x$steps using existing PBRI algorithms. A parallel architecture of the proposed algorithm employing simple logic gates and having a short critical path delay is presented. The proposed architecture is valuable in reducing (de-)interleaving latency in emerging wireless standards that employ PBRI channel (de-)interleaving in their PHY layer such as the 3GPP2 Ultra Mobile Broadband standard. Mohammad M. Mansour |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2004 | Analysis of MOS cross-coupled LC-tank oscillators using short-channel device equations
Makram M. Mansour, Mohammad M. Mansour, Amit Mehrotra |
ASP-DAC | 2 |
| 2004 | High-performance decoders for regular and irregular repeat-accumulate codesabstractThis paper investigates high-performance decoder design for regular and irregular repeat-accumulate (RA) codes of large block length. In order to achieve throughputs and bit-error rate performance that are inline with future trends in high-speed communications. high-throughput and low-power decoders of low complexity are needed. To meet such conflicting requirements for long codes, the concept of architecture-aware RA (AARA) code design is proposed. AARA code design decouples the complexity of the decoder from the owe structure by inducing structural regularity features that are amenable to efficient and scalable decoder implementations. Design methods of AARA codes with structured permuters for which an iterative decoding algorithm performs well under message-passing are analogous to those for AA LDPC codes. Algorithmic and architectural optimizations that address the latency, memory overhead, and complexity problems typical of iterative decoders for long RA codes are investigated, and a staggered decoding schedule is introduced. AARA decoders using the proposed schedule have substantial advantage over serial and parallel RA decoders. Mohammad M. Mansour |
GLOBECOM | 1 |
| 2003 | VLSI architectures for SISO-APP decodersabstractVery large scale integration (VLSI) design methodology and implementation complexities of high-speed, low-power soft-input soft-output (SISO) a posteriori probability (APP) decoders are considered. These decoders are used in iterative algorithms based on turbo codes and related concatenated codes and have shown significant advantage in error correction capability compared to conventional maximum likelihood decoders. This advantage, however, comes at the expense of increased computational complexity, decoding delay, and substantial memory overhead, all of which hinge primarily on the well-known recursion bottleneck of the SISO-APP algorithm. This paper provides a rigorous analysis of the requirements for computational hardware and memory at the architectural level based on a tile-graph approach that models the resource-time scheduling of the recursions of the algorithm. The problem of constructing the decoder architecture and optimizing it for high speed and low power is formulated in terms of the individual recursion patterns which together form a tile graph according to a tiling scheme. Using the tile-graph approach, optimized architectures are derived for the various forms of the sliding-window and parallel-window algorithms known in the literature. A proposed tiling scheme of the recursion patterns, called hybrid tiling, is shown to be particularly effective in reducing memory overhead of high-speed SISO-APP architectures. Simulations demonstrate that the proposed approach achieves savings in area and power in the range of 4.2%-53.1% over state of the art. Mohammad M. Mansour, Naresh R. Shanbhag |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2003 | High-throughput LDPC decodersabstractA high-throughput memory-efficient decoder architecture for low-density parity-check (LDPC) codes is proposed based on a novel turbo decoding algorithm. The architecture benefits from various optimizations performed at three levels of abstraction in system design-namely LDPC code design, decoding algorithm, and decoder architecture. First, the interconnect complexity problem of current decoder implementations is mitigated by designing architecture-aware LDPC codes having embedded structural regularity features that result in a regular and scalable message-transport network with reduced control overhead. Second, the memory overhead problem in current day decoders is reduced by more than 75% by employing a new turbo decoding algorithm for LDPC codes that removes the multiple checkto-bit message update bottleneck of the current algorithm. A new merged-schedule merge-passing algorithm is also proposed that reduces the memory overhead of the current algorithm for low to moderate-throughput decoders. Moreover, a parallel soft-input-soft-output (SISO) message update mechanism is proposed that implements the recursions of the Balh-Cocke-Jelinek-Raviv (BCJR) algorithm in terms of simple "max-quartet" operations that do not require lookup-tables and incur negligible loss in performance compared to the ideal case. Finally, an efficient programmable architecture coupled with a scalable and dynamic transport network for storing and routing messages is proposed, and a full-decoder architecture is presented. Simulations demonstrate that the proposed architecture attains a throughput of 1.92 Gb/s for a frame length of 2304 bits, and achieves savings of 89.13% and 69.83% in power consumption and silicon area over state-of-the-art, with a reduction of 60.5% in interconnect length. Mohammad M. Mansour, Naresh R. Shanbhag |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2002 | Turbo decoder architectures for low-density parity-check codesabstractTurbo decoding of low-density parity-check (LDPC) and generalized low-density (GLD) codes and the corresponding decoder architectures are considered. A regular (c, r)-LDPC code of length n is viewed as the intersection of c interleaved super-codes where each super-code is the direct sum of n/r independent single parity-check sub-codes. Extensions to GLD codes simply utilize more powerful sub-codes. The turbo decoding schedule is employed to decode LDPC and GLD codes using constituent soft-input soft-output (SISO) decoders that communicate through c interleavers. The proposed schedule exhibits a faster convergence behavior, and hence lower decoding latency, than the commonly employed two-phase schedule, and has a reduced memory requirement that is a function of the number of super-codes. The performance of the turbo decoding schedule is evaluated through simulations over an AWGN channel. Mohammad M. Mansour, Naresh R. Shanbhag |
GLOBECOM | 1 |
| 2002 | Design methodology for high-speed iterative decoder architecturesabstractWe propose a novel approach to the design and analysis of VLSI architectures for the soft-input soft-output a posteriori probability (SISO-APP) decoding algorithm used in iterative decoders such as turbo decoders. The approach is based on a tile-graph composed of recursion patterns that model the resource-time scheduling of the forward-backward recursion equations of the algorithm. The problem of constructing a SISO-APP architecture is formulated as a three-step process of constructing and counting the patterns needed and then tiling them. The problem of optimizing the architecture for high speed and low power reduces to optimizing the individual patterns and the tiling scheme for minimal delay and storage overhead. The various forms of the sliding and parallel-window (PW) architectures in the literature are instances of the proposed tile-graph. Using the tile-graph approach, a new PW architecture controlled by the window width r is proposed that achieves for r = 10 a 45%, a 71 %, a 51%, and a 25% reduction in decoding delay, state, input, and output metrics storage respectively, compared to a conventional architecture with a 10% increase in resources. Mohammad M. Mansour, Naresh R. Shanbhag |
ICASSP | 1 |
| 2002 | Low-power VLSI decoder architectures for LDPC codesabstractIterative decoding of low-density parity check codes (LDPC) using the message-passing algorithm have proved to be extraordinarily effective compared to conventional maximum-likelihood decoding. However, the lack of any structural regularity in these essentially random codes is a major challenge for building a practical low-power LDPC decoder. In this paper, we jointly design the code and the decoder to induce the structural regularity needed for a reduced complexity parallel decoder architecture. This interconnect-driven code design approach eliminates the need for a complex interconnection network while still retaining the algorithmic performance promised by random codes. Moreover, we propose a new approach for computing reliability metrics based on the BCJR algorithm that reduces the message switching activity in the decoder compared to existing approaches. Simulations show that the proposed approach results in power savings of up to 85.64% over conventional implementations. Mohammad M. Mansour, Naresh R. Shanbhag |
ISLPED | 1 |
| 2002 | A cloning approach to classifier trainingabstractThe Al-Alaoui algorithm is a weighted mean-square error (MSE) approach to pattern recognition. It employs cloning of the erroneously classified samples to increase the population of their corresponding classes. The algorithm was originally developed for linear classifiers. In this paper, the algorithm is extended to multilayer neural networks which may be used as nonlinear classifiers. It is also shown that the application of the Al-Alaoui algorithm to multilayer neural networks speeds up the convergence of the back-propagation algorithm. Mohamad Adnan Al-Alaoui, Rodolphe Mouci, Mohammad M. Mansour, Rony Ferzli |
IEEE Trans. Syst. Man Cybern. Part A | 3 |
| 1998 | FPGA-based Internet Protocol Version 6 routerabstractIn this paper, a novel hardware design for an Internet Protocol Version 6 router using field programmable gate arrays is proposed. A dataflow, parallel, pipelined and scalable architecture is presented that has the potential of matching the enormous communication bandwidths of transmission links. A ternary content addressable memory (CAM) in the form of cache is adopted as a routing table search engine. It can offer O(1) search time with just O(N) memory words. Adding a sorting (priority) mechanism by caching the routing table in CAM and using a modified form of sector mapping technique eliminates the slow insertion and deletion times without adding significant additional hardware costs. Mohammad M. Mansour, Ayman I. Kayssi |
ICCD | 1 |