Mostafa El-Khamy

dblp:00/4303 · DBLP profile ↗
← Back
50ranked-venue papers
16as first author
8since 2021 · last 2024
0000-0001-9421-6037ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 1 first-author · 7 since 2021Computer networks · 14 · 9 first-authorArtificial intelligence and machine learning · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-authorTheory of computation · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Knowledge Distillation for Tiny Speech Enhancement with Latent Feature Augmentation
Behnam Gholami, Mostafa El-Khamy, Kee-Bong Song
INTERSPEECH2
2023 Domain invariant regularization by disentangling content and style Features for visual domain generalization
abstract
In this paper, taking the advantage of multiple source domains, we propose a novel approach for visual Domain Generalization (DG). The three key ideas underlying our formulation are (1) leveraging disentangled representations of the images to define different factors of variations, (2) generating perturbed images by changing such factors composing the representations of the images, (3) enforcing the learner (classifier) to be invariant to such changes in the images. We demonstrate the effectiveness of our approach on several widely used datasets for the domain generalization problem, on all of which we achieve competitive results with state-of-the-art models.
Behnam Gholami, Mostafa El-Khamy, Kee-Bong Song
ICIP2
2023 Latent Feature Disentanglement for Visual Domain Generalization
abstract
Despite remarkable success in a variety of computer vision applications, it is well-known that deep learning can fail catastrophically when presented with out-of-distribution data, where there are usually style differences between the training and test images. Toward addressing this challenge, we consider the domain generalization problem, wherein predictors are trained using data drawn from a family of related training (source) domains and then evaluated on a distinct and unseen test domain. Naively training a model on the aggregate set of data (pooled from all source domains) has been shown to perform suboptimally, since the information learned by that model might be domain-specific and generalizes imperfectly to test domains. Data augmentation has been shown to be an effective approach to overcome this problem. However, its application has been limited to enforcing invariance to simple transformations like rotation, brightness change, etc. Such perturbations do not necessarily cover plausible real-world variations that preserve the semantics of the input (such as a change in the image style). In this paper, taking the advantage of multiple source domains, we propose a novel approach to express and formalize robustness to these kind of real-world image perturbations. The three key ideas underlying our formulation are (1) leveraging disentangled representations of the images to define different factors of variations, (2) generating perturbed images by changing such factors composing the representations of the images, (3) enforcing the learner (classifier) to be invariant to such changes in the images. We use image-to-image translation models to demonstrate the efficacy of this approach. Based on this, we propose a domain-invariant regularization (DIR) loss function that enforces invariant prediction of targets (class labels) across domains which yields improved generalization performance. We demonstrate the effectiveness of our approach on several widely used datasets for the domain generalization problem, on all of which our results are competitive with the state-of-the-art.
Behnam Gholami, Mostafa El-Khamy, Kee-Bong Song
IEEE Trans. Image Process.2
2023 Efficient frameworks for statistical seizure detection and prediction
Ali A. Khalil, Mostafa El-Khamy, Fatma E. Ibrahim, Ashraf A. M. Khalaf, Entessar Gemeay, Hossam Kasem, Salah Eldeen A. Khamis, Ghada M. El Banby, Walid El Shafai, S. El-Rabaie 0001, Adel S. El-Fishawy, Moawad I. Dessouky, Ibrahim M. Eldokany, Turky N. Alotaiby, Saleh Al-Shebeili, Fathi E. Abd El-Samie
J. Supercomput.2
2022 DeepGBASS: Deep Guided Boundary-Aware Semantic Segmentation
abstract
Image semantic segmentation is ubiquitously used in scene understanding applications, such as AI Camera, which require high accuracy and efficiency. Deep learning has significantly advanced the state-of-the-art in semantic segmentation. However, many of recent semantic segmentation works only consider class accuracy and ignore the accuracies at the boundaries between semantic classes. To improve the semantic boundary accuracy, we propose low complexity Deep Guided Decoder (DGD) networks, trained with a novel Semantic Boundary-Aware Learning (SBAL) strategy. Our ablation studies on Cityscapes and the ADE20K-32 confirm the effectiveness of our approach with network of different complexities. We show that our DeepGBASS approach significantly improves the mIoU by up to 11% relative gain and the mean boundary F1-score (mBF) by up to 39.4% when training MobileNetEdgeTPU DeepLab on ADE20K-32 dataset.
Qingfeng Liu, Hai Su, Mostafa El-Khamy, Kee-Bong Song
ICASSP3
2022 Panoptic-Deeplab-DVA: Improving Panoptic Deeplab with Dual Value Attention and Instance Boundary Aware Regression
abstract
Panoptic DeepLab is a state-of-the-art framework that has showed good tradeoff between performance and complexity. In this paper, we focus on improving it to increase wide deployment of panoptic segmentation on mobile devices with low complexity. Specifically, we first present a novel Dual Value Attention (DVA) module to enable context information exchange between the semantic segmentation branch and the instance segmentation branch. Second, we further propose a new instance Boundary Aware Regression (iBAR) loss that assigns more emphasis on the instance boundary during instance regression. To assess the effectiveness of our proposed approach, we evaluate the performance on MSCOCO dataset for panoptic segmentation task, to show that our approach can improve upon the state-of-the-art Panoptic DeepLab with both the light-weight backbone network MobileNetV3 and the heavy-weight backbone network HRNetV2.
Qingfeng Liu, Mostafa El-Khamy
ICIP2
2021 Zero-Shot Learning Of A Conditional Generative Adversarial Network For Data-Free Network Quantization
abstract
We propose a novel method for training a conditional generative adversarial network (CGAN) without the use of training data, called zero-shot learning of a CGAN (ZS-CGAN). Zero-shot learning of a conditional generator only needs a pre-trained discriminative (classification) model and does not need any training data. In particular, the conditional generator is trained to produce labeled synthetic samples whose characteristics mimic the original training data by using the statistics stored in the batch normalization layers of the pretrained model. We show the usefulness of ZS-CGAN in data-free quantization of deep neural networks. We achieved the state-of-the-art data-free network quantization of the ResNet and MobileNet classification models trained on the ImageNet dataset. Data-free quantization using ZS-CGAN showed a minimal loss in accuracy compared to that obtained by conventional data-dependent quantization.
Yoojin Choi, Mostafa El-Khamy
ICIP2
2021 HyperCon: Image-To-Video Model Transfer for Video-To-Video Translation Tasks
abstract
Video-to-video translation is more difficult than image-to-image translation due to the temporal consistency problem that, if unaddressed, leads to distracting flickering effects. Although video models designed from scratch produce temporally consistent results, training them to match the vast visual knowledge captured by image models requires an intractable number of videos. To combine the benefits of image and video models, we propose an image-to-video model transfer method called Hyperconsistency (HyperCon) that transforms any well-trained image model into a temporally consistent video model without fine-tuning. HyperCon works by translating a temporally interpolated video frame-wise and then aggregating over temporally localized windows on the interpolated video. It handles both masked and unmasked inputs, enabling support for even more video-to-video translation tasks than prior image-to-video model transfer techniques. We demonstrate HyperCon on video style transfer and inpainting, where it performs favorably compared to prior state-of-the-art methods without training on a single stylized or incomplete video. Our project website is available at ryanszeto.com/projects/hypercon.
Ryan Szeto, Mostafa El-Khamy, Jason J. Corso
WACV2
2020 CAD-AEC: Context-Aware Deep Acoustic Echo Cancellation
abstract
Deep-leaming based acoustic echo cancellation (AEC) methods have been shown to outperform the classical techniques. The main drawback of the learning-based AEC is its dependency on the training set, which limits its practical deployment in mobile devices and unconstrained environments. This paper proposes a context- aware deep AEC (CAD-AEC) by introducing two main components. The first component of the CAD-AEC borrows ideas from the classical AEC and performs frequency domain adaptive filtering of the microphone signal, to provide the deep AEC network with features that have less dependency on the development context. The second component is a deep contextual-attention module (CAM), inserted between the recurrent encoder and decoder architectures. The deep CAM adaptively scales the encoder output during inference with calculated attention weights that depend on the context. Experiments in both matched and mismatched training and testing environments, show that the proposed CAD-AEC can robustly achieve better echo return loss enhancement (ERLE) and perceptual speech quality compared to the previous classical and deep learning techniques.
Amin Fazel, Mostafa El-Khamy
ICASSP2
2020 T-GSA: Transformer with Gaussian-Weighted Self-Attention for Speech Enhancement
abstract
Transformer neural networks (TNN) demonstrated state-ofart performance on many natural language processing (NLP) tasks, replacing recurrent neural networks (RNNs), such as LSTMs or GRUs. However, TNNs did not perform well in speech enhancement, whose contextual nature is different than NLP tasks, like machine translation. Self-attention is a core building block of the Transformer, which not only enables parallelization of sequence computation, but also provides the constant path length between symbols that is essential to learning long-range dependencies. In this paper, we propose a Transformer with Gaussian-weighted self-attention (T-GSA), whose attention weights are attenuated according to the distance between target and context symbols. The experimental results show that the proposed T-GSA has significantly improved speech-enhancement performance, compared to the Transformer and RNNs.
Mostafa El-Khamy
ICASSP2
2020 Deep Monocular Video Depth Estimation Using Temporal Attention
abstract
Monocular video depth estimation (MVDE) plays a crucial role in 3D computer vision. In this paper, we propose an end-to-end monocular video depth estimation network based on temporal attention. Our network starts by a motion compensation module where the spatial temporal transformer network (STN) is utilized to warp the input frames using the estimated optical flow. Next, a temporal attention module is used to combine features from the warped frames, while emphasizing the temporal consistency. A monocular depth estimation network is used to estimate the depth from the temporally combined features. Experimental results demonstrate that our proposed framework achieves better performance compared to the state-of-the-art single image depth estimation (SIDE) networks, as well as existing MVDE methods.
Mostafa El-Khamy
ICASSP2
2020 GSANet: Semantic Segmentation With Global And Selective Attention
abstract
This paper proposes a novel deep learning architecture for semantic segmentation. The proposed Global and Selective Attention Network (GSANet) features Atrous Spatial Pyramid Pooling (ASPP) with a novel sparsemax global attention and a novel selective attention that deploys a condensation and diffusion mechanism to aggregate the multi-scale contextual information from the extracted deep features. A selective attention decoder is also proposed to process the GSA-ASPP outputs for optimizing the softmax volume. We are the first to benchmark the performance of semantic segmentation networks with the low-complexity feature extraction network (FXN) MobileNetEdge, that is optimized for low latency on edge devices. We show that GSANet can result in more accurate segmentation with MobileNetEdge, as well as with strong FXNs, such as Xception. GSANet improves the state-of-art semantic segmentation accuracy on both the ADE20k and the Cityscapes datasets.
Qingfeng Liu, Mostafa El-Khamy, Dongwoon Bai
ICIP2
2020 Stereo Disparity Estimation via Joint Supervised, Unsupervised, and Weakly Supervised Learning
abstract
We propose SUW-Stereo, a deep-learning framework with joint supervised learning (S), unsupervised learning (U), and weakly-supervised learning (W) for disparity estimation. The supervised learning module optimizes a disparity estimation network by knowledge based on the ground-truth disparity. In contrast, the unsupervised learning module has no knowledge of the ground-truth disparity, but optimizes the disparity estimation network by predicting the current right image from the left image and estimated disparity. The weakly supervised learning uses other expert information to derive some information about the disparity, without using ground-truth disparity knowledge either. SUW-Stereo trains the deep-learning networks end-to-end with joint optimization of the desired SUW objectives. We demonstrate that SUW-Stereo can even improve the accuracy of current state-of-the-art networks such as AMNet or GANet.
Mostafa El-Khamy
ICIP2
2019 Jointly Sparse Convolutional Neural Networks in Dual Spatial-winograd Domains
abstract
We consider the optimization of deep convolutional neural networks (CNNs) such that they provide good performance while having reduced complexity if deployed on either conventional systems with spatial-domain convolution or lower-complexity systems designed for Winograd convolution. The proposed framework produces one compressed model whose convolutional filters can be made sparse either in the spatial domain or in the Winograd domain. Hence, the compressed model can be deployed universally on any platform, without need for re-training on the deployed platform. To get a better compression ratio, the sparse model is compressed in the spatial domain that has a fewer number of parameters. From our experiments, we obtain 24.2× and 47.7× compressed models for ResNet-18 and AlexNet trained on the ImageNet dataset, while their computational cost is also reduced by 4.5× and 5.1×, respectively.
Yoojin Choi, Mostafa El-Khamy
ICASSP2
2019 Variable Rate Deep Image Compression With a Conditional Autoencoder
abstract
In this paper, we propose a novel variable-rate learned image compression framework with a conditional autoencoder. Previous learning-based image compression methods mostly require training separate networks for different compression rates so they can yield compressed images of varying quality. In contrast, we train and deploy only one variable-rate image compression network implemented with a conditional autoencoder. We provide two rate control parameters, i.e., the Lagrange multiplier and the quantization bin size, which are given as conditioning variables to the network. Coarse rate adaptation to a target is performed by changing the Lagrange multiplier, while the rate can be further fine-tuned by adjusting the bin size used in quantizing the encoded representation. Our experimental results show that the proposed scheme provides a better rate-distortion trade-off than the traditional variable-rate image compression codecs such as JPEG2000 and BPG. Our model also shows comparable and sometimes better performance than the state-of-the-art learned image compression models that deploy multiple networks trained for varying rates.
Yoojin Choi, Mostafa El-Khamy
ICCV2
2019 Multi-Task Learning of Depth from Tele and Wide Stereo Image Pairs
abstract
In this paper, we introduce the problem of estimating the real world depth of elements in a scene captured by two cameras with different field of views, where the first field of view (FOV) is a Wide FOV obtained by a lens with 1 × the optical zoom, and the second FOV is contained in the first FOV and corresponds to a tele zoom lens with 2 × the optical zoom. Traditional stereo matching techniques can estimate the stereo disparity, and hence the depth, in the overlapping FOV between both cameras only, which corresponds to the Tele FOV. We refer to the problem of estimating the disparity, or inverse depth, for the union of FOVs as `Tele-Wide disparity estimation'. We propose different deep learning solutions to establish baseline performances. We trained a single-image inverse-depth estimation (SIDE) network to estimate the inverse depth from the image corresponding to the Wide FOV only. We also trained a stereo image disparity estimation network to estimate the disparity for the overlapping Tele FOV only, and another tele-wide stereo matching network (TW-SMNet) for estimating the disparity for the union Wide FOV. We further propose an end-to-end multi-task tele-wide stereo matching deep neural network (MT-TW-SMNet) which attempts to do stereo matching in the overlapped Tele FOV and SIDE in the union Wide FOV. Experimental results on KITTI and the SceneFlow datasets establish baseline performances for the tele-wide stereo matching and demonstrate that multitask tele-wide stereo matching provides a reasonable solution to the Tele-Wide depth estimation problem.
Mostafa El-Khamy, Xianzhi Du
ICIP1
2019 Deep Multitask Acoustic Echo Cancellation
Amin Fazel, Mostafa El-Khamy
INTERSPEECH2
2018 DN-ResNet: Efficient Deep Residual Network for Image Denoising
Mostafa El-Khamy
ACCV (5)2
2018 Bridgenets: Student-Teacher Transfer Learning Based on Recursive Neural Networks and Its Application to Distant Speech Recognition
abstract
Despite the remarkable progress achieved on automatic speech recognition, recognizing far-field speeches mixed with various noise sources is still a challenging task. In this paper, we introduce novel student-teacher transfer learning, BridgeNet which can provide a solution to improve distant speech recognition. There are two key features in BridgeNet. First, BridgeNet extends traditional student-teacher frameworks by providing multiple hints from a teacher network. Hints are not limited to the soft labels from a teacher network. Teacher's intermediate feature representations can better guide a student network to learn how to denoise or dereverberate noisy input. Second, the proposed recursive architecture in the BridgeNet can iteratively improve denoising and recognition performance. The experimental results of BridgeNet showed significant improvements in tackling the distant speech recognition problem, where it achieved up to 13.24% relative WER reductions on AMI corpus compared to a baseline neural network without teacher's hints.
Mostafa El-Khamy
ICASSP2
2018 CT-SRCNN: Cascade Trained and Trimmed Deep Convolutional Neural Networks for Image Super Resolution
abstract
We propose methodologies to train highly accurate and efficient deep convolutional neural networks (CNNs) for image super resolution (SR). A cascade training approach to deep learning is proposed to improve the accuracy of the neural networks while gradually increasing the number of network layers. Next, we explore how to improve the SR efficiency by making the network slimmer. Two methodologies, the one-shot trimming and the cascade trimming, are proposed. With the cascade trimming, the network's size is gradually reduced layer by layer, without significant loss on its discriminative ability. Experiments on benchmark image datasets show that our proposed SR network achieves the state-of-the-art super resolution accuracy, while being more than 4 times faster compared to existing deep super resolution networks.
Mostafa El-Khamy
WACV2
2018 Circular Buffer Rate-Matched Polar Codes
abstract
A practical rate-matching system for constructing rate-compatible polar codes is proposed. The proposed polar code circular buffer rate-matching is suitable for transmissions on communication channels that support hybrid automatic repeat request communications, as well as for flexible resource-element rate-matching on single transmission channels. Our proposed circular buffer rate matching scheme also incorporates a bit-mapping scheme for transmission on bit-interleaved coded modulation (BICM) channels using higher order modulations. An interleaver is derived from a puncturing order obtained with a low complexity progressive puncturing search algorithm on a base code of short length, and has the flexibility to achieve any desired rate at the desired code length, through puncturing or repetition. The rate-matching scheme is implied by a two-stage polarization, for transmission at any desired code length, code rate, and modulation order, and is shown to achieve the symmetric capacity of BICM channels. Numerical results on AWGN and fast fading channels show that the rate-matched polar codes have a competitive performance when compared with the spatially-coupled quasi-cyclic LDPC codes or LTE turbo codes, while having similar rate-dematching storage and computational complexities.
Mostafa El-Khamy, Hsien-Ping Lin, Inyup Kang
IEEE Trans. Commun.1
2017 Towards the Limit of Network Quantization
Yoojin Choi, Mostafa El-Khamy
ICLR (Poster)2
2017 Residual LSTM: Design of a Deep Recurrent Architecture for Distant Speech Recognition
abstract
In this paper, a novel architecture for a deep recurrent neural network, residual LSTM is introduced.A plain LSTM has an internal memory cell that can learn long term dependencies of sequential data.It also provides a temporal shortcut path to avoid vanishing or exploding gradients in the temporal domain.The residual LSTM provides an additional spatial shortcut path from lower layers for efficient training of deep networks with multiple LSTM layers.Compared with the previous work, highway LSTM, residual LSTM separates a spatial shortcut path with temporal one by using output layers, which can help to avoid a conflict between spatial and temporal-domain gradient flows.Furthermore, residual LSTM reuses the output projection matrix and the output gate of LSTM to control the spatial information flow instead of additional gate networks, which effectively reduces more than 10% of network parameters.An experiment for distant speech recognition on the AMI SDM corpus shows that 10-layer plain and highway LSTM networks presented 13.7% and 6.2% increase in WER over 3-layer baselines, respectively.On the contrary, 10-layer residual LSTM networks provided the lowest WER 41.0%, which corresponds to 3.3% and 2.8% WER reduction over plain and highway LSTM networks, respectively.
Mostafa El-Khamy
INTERSPEECH2
2017 Fused DNN: A Deep Neural Network Fusion Approach to Fast and Robust Pedestrian Detection
abstract
We propose a deep neural network fusion architecture for fast and robust pedestrian detection. The proposed network fusion architecture allows for parallel processing of multiple networks for speed. A single shot deep convolutional network is trained as a object detector to generate all possible pedestrian candidates of different sizes and occlusions. This network outputs a large variety of pedestrian candidates to cover the majority of ground-truth pedestrians while also introducing a large number of false positives. Next, multiple deep neural networks are used in parallel for further refinement of these pedestrian candidates. We introduce a soft-rejection based network fusion method to fuse the soft metrics from all networks together to generate the final confidence scores. Our method performs better than existing state-of-the-arts, especially when detecting small-size and occluded pedestrians. Furthermore, we propose a method for integrating pixel-wise semantic segmentation network into the network fusion architecture as a reinforcement to the pedestrian detector. The approach outperforms state-of-the-art methods on most protocols on Caltech Pedestrian dataset, with significant boosts on several protocols. It is also faster than all other methods.
Xianzhi Du, Mostafa El-Khamy, Larry Davis 0001
WACV2
2017 Relaxed Polar Codes
abstract
Polar codes are the latest breakthrough in coding theory, as they are the first family of codes with explicit construction that provably achieve the symmetric capacity of binary-input discrete memoryless channels. Polar encoding and successive cancellation decoding have the complexities of N log N , for code length N. Although, the complexity bound of N log N is asymptotically favorable, we report in this work methods to further reduce the encoding and decoding complexities of polar coding. The crux is to relax the polarization of certain bit-channels without performance degradation. We consider schemes for relaxing the polarization of both very good and very bad bit-channels, in the process of channel polarization. Relaxed polar codes are proved to preserve the capacity achieving property of polar codes. Analytical bounds on the asymptotic and finite-length complexity reduction attainable by relaxed polarization are derived. For binary erasure channels, we show that the computation complexity can be reduced by a factor of six, while preserving the rate and error performance. We also show that relaxed polar codes can be decoded with significantly reduced latency. For additive white Gaussian noise channels with medium code lengths, we show that relaxed polar codes can have lower error probabilities than conventional polar codes, while having reduced encoding and decoding computation complexities.
Mostafa El-Khamy, Hessam Mahdavifar, Gennady Feygin, Inyup Kang
IEEE Trans. Inf. Theory1
2016 Finite-Length Algebraic Spatially-Coupled Quasi-Cyclic LDPC Codes
abstract
The replicate-and-mask (R&M) construction of finite-length spatially-coupled (SC) LDPC codes is proposed in this paper. The proposed R&M construction generalizes the conventional matrix unwrapping construction and contains it as a special case. The R&M construction of a class of algebraic spatially coupled (SC) quasi-cyclic (QC) LDPC codes over arbitrary finite fields is demonstrated. The girth, rank, and time-varying periodicity of the proposed R&M SC QC LDPC codes are analyzed. The error rate performance of finite-length nonbinary algebraic SC QC LDPC codes is investigated with window decoding. Compared to the conventional unwrapping construction, it is found through numerical simulations that the R&M construction resulted in SC QC LDPC codes with better block error rate performance and lower error floors. With a flooding schedule decoder, it is shown that the proposed R&M algebraic SC QC LDPC codes have better error performance than the corresponding LDPC block codes and random SC codes. The R&M construction of irregular SC QC LDPC codes is demonstrated. It is shown that low-complexity regular puncturing schemes can be deployed on these codes to construct families of rate-compatible irregular SC QC LDPC codes with good performance.
Keke Liu, Mostafa El-Khamy
IEEE J. Sel. Areas Commun.2
2016 Achieving the Uniform Rate Region of General Multiple Access Channels by Polar Coding
abstract
We consider the problem of polar coding for transmission over m-user multiple access channels. In the proposed scheme, all users encode their messages using a polar encoder, while a multiuser successive cancellation decoder is deployed at the receiver. The encoding is done separately across the users and is independent of the target achievable rate. For the code construction, the positions of information bits and frozen bits for each of the users are decided jointly. This is done by treating the polar transformations across all the m users as a single polar transformation with a certain polarization base. We characterize the resolution of achievable rates on the dominant face of the uniform rate region in terms of the number of users m and the length of the polarization base L. In particular, we prove that for any target rate on the dominant face, there exists an achievable rate, also on the dominant face, within the distance at most (m-1)√m/L from the target rate. We then prove that the proposed L MAC polar coding scheme achieves the whole uniform rate region with fine enough resolution by changing the decoding order in the multiuser successive cancellation decoder, as L and the code block length N grow large. The encoding and decoding complexities are O(N log N) and the asymptotic block error probability of O(2-N0.5-ϵ) is guaranteed. Examples of achievable rates for the 3-user multiple access channel are provided.
Hessam Mahdavifar, Mostafa El-Khamy, Inyup Kang
IEEE Trans. Commun.2
2015 HARQ Rate-Compatible Polar Codes for Wireless Channels
abstract
A design of rate-compatible polar codes suitable for HARQ communications is proposed in this paper. An important feature of the proposed design is that the puncturing order is chosen with low complexity on a base code of short length, which is then further polarized to the desired length. A practical rate-matching system that has the flexibility to choose any desired rate through puncturing or repetition while preserving the polarization is suggested. The proposed rate-matching system is combined with channel interleaving and a bit-mapping procedure that preserves the polarization of the rate-compatible polar code family over bit-interleaved coded modulation systems. Simulation results on AWGN and fast fading channels with different modulation orders show the robustness of the proposed rate-compatible polar code in both Chase combining and incremental redundancy HARQ communications.
Mostafa El-Khamy, Hsien-Ping Lin, Hessam Mahdavifar, Inyup Kang
GLOBECOM1
2015 Relaxed channel polarization for reduced complexity polar coding
abstract
Arıkan's polar codes are proven to be capacity-achieving error correcting codes while having explicit constructions. They are characterized to have encoding and decoding complexities of l log l, for code length l. In this work, we construct another family of capacity-achieving codes that have even lower encoding and decoding complexities, by relaxing the channel polarizations for certain bit-channels. We consider schemes for relaxing the polarization of both sufficiently good and sufficiently bad bit-channels, in the process of channel polarization. We prove that, similar to conventional polar codes, relaxed polar codes also achieve the capacity of binary memoryless symmetric channels. We analyze the complexity reductions achievable by relaxed polarization for asymptotic and finite-length codes, both numerically and analytically. We show that relaxed polar codes can have better bit error probabilities than conventional polar codes, while having reduced encoding and decoding complexities.
Mostafa El-Khamy, Hessam Mahdavifar, Gennady Feygin, Inyup Kang
WCNC1
2014 LLR optimization for iterative MIMO BICM receivers
abstract
Iterative detection and decoding (IDD) relies on passing useful extrinsic information between the detector and the decoder. Due to the sub-optimality of practical detector and/or decoder, the direct output LLRs from the detector or the decoder may not provide sufficient gains to each other. Proper scaling of the extrinsic LLRs based on certain optimality criteria may improve the performance of the IDD receiver. However, finding optimal scaling function for IDD receiver in general is still an open problem. In this paper, we investigate LLR scaling of the detector and the decoder output based on maximization of generalized mutual information.
Jinhong Wu, Mostafa El-Khamy, Inyup Kang
ICASSP2
2014 Non-binary algebraic spatially-coupled quasi-cyclic LDPC codes
abstract
This paper considers the algebraic construction and performance of non-binary spatially-coupled low density parity check (LDPC) codes. A replicate-and-mask approach is presented to construct finite-length algebraic quasi-cyclic (QC) spatially-coupled (SC) LDPC codes. Numerical results show the superiority of non-binary algebraic SC QC LDPC codes over the corresponding random non-binary (block and SC) LDPC codes. In this paper, it is demonstrated that the threshold saturation phenomenon, previously demonstrated for binary SC LDPC codes, also holds for non-binary SC LDPC codes over the binary-input AWGN channel with BPSK modulation.
Keke Liu, Mostafa El-Khamy, Inyup Kang, Arvind Yedla
ISIT2
2014 Wide-Band Cooperative Compressive Spectrum Sensing for Cognitive Radio Systems Using Distributed Sensing Matrix
abstract
In this paper, cooperative compressive spectrum sensing is considered to enable accurate sensing of the wide-band spectrum. The proposed algorithm is based on compressive sensing theory and aims to reduce the hardware complexity of the cognitive radio receiver by distributing the sensing work among groups of sensing nodes. The proposed algorithm classifies the cooperated sensing nodes into different sensing groups depending on the quality of the reporting channel between the sensing node and the fusion center (FC). To sense the wide- band analog signal and take a global decision about spectrum occupancy, each node uses its local sensing matrix, which is assigned to its sensing group and a part of a global sensing matrix at the FC. The size of the local sensing matrix of each sensing node,and consequently the contribution of this node in the overall measurement vector, depends on its sensing group. The FC classifies and rearranges the compressed data to formulate one global measurement vector which is used with a global sensing matrix to estimate the wide-band signal spectrum. The receiver operation characteristics (ROC) of the overall spectrum sensing system show that the proposed receiver provides more protection to primary users (higher detection probability) at the same secondary user throughput (probability of false alarm).
Mohammed Farrag, Osamu Muta, Mostafa El-Khamy, Hiroshi Furukawa, Mohamed El-Sharkawy 0001
VTC Fall3
2014 Online log-likelihood ratio scaling for robust turbo decoding
abstract
Optimal iterative log‐MAP decoding of turbo codes requires accurate knowledge of the operating signal‐to‐noise ratio (SNR). However, the SNR information, available at practical decoders for bit‐interleaved coded modulation systems, such as the third generation partnership project high‐speed packet access and long‐term evolution wireless cellular systems, may be inaccurate. In this study, two decoder architectures for improved turbo decoding in the presence of SNR mismatch are proposed. The SNR‐mismatch aware turbo decoder selects the decoder which is estimated to have the best performance at the current mismatch, according to the test criterion. The SNR‐mismatch compensated turbo decoder provides a more accurate estimation of the noise variance and concurrently scales the channel and the decoder log‐likelihood ratios (LLRs) to continue decoding. Two different methods are proposed to find the optimal scaling factors online, one on the symbol level and the other on the bit level. This study shows that online LLR scaling, without prior knowledge about the noise mismatch statistics, can result in near‐optimal turbo decoding regardless of the initial SNR mismatch.
Mostafa El-Khamy, Jinhong Wu, Inyup Kang
IET Commun.1
2014 Performance Limits and Practical Decoding of Interleaved Reed-Solomon Polar Concatenated Codes
abstract
A scheme for concatenating the recently invented polar codes with non-binary MDS codes, as Reed-Solomon codes, is considered. By concatenating binary polar codes with interleaved Reed-Solomon codes, we prove that the proposed concatenation scheme captures the capacity-achieving property of polar codes, while having a significantly better error-decay rate. We show that for any ε > 0, and total frame length N, the parameters of the scheme can be set such that the frame error probability is less than 2-N1-ε, while the scheme is still capacity achieving. This improves upon 2-N0.5-ε, the frame error probability of Arikan's polar codes. The proposed concatenated polar codes and Arikan's polar codes are also compared for transmission over channels with erasure bursts. We provide a sufficient condition on the length of erasure burst which guarantees failure of the polar decoder. On the other hand, it is shown that the parameters of the concatenated polar code can be set in such a way that the capacity-achieving properties of polar codes are preserved. We also propose decoding algorithms for concatenated polar codes, which significantly improve the error-rate performance at finite block lengths while preserving the low decoding complexity.
Hessam Mahdavifar, Mostafa El-Khamy, Inyup Kang
IEEE Trans. Commun.2
2013 Low complexity independent multi-view video coding
abstract
In 3D multi-view video coding (MVC), disparity estimation (DE) are used to exploit the correlation among different view sequences. The DE process greatly increases the computational complexity of the MVC. In this paper, a novel independent low complexity multi-view video coder (I-MVC) is introduced. In the proposed MVC, the coding complexity is shifted from the encoder side to the decoder side. Instead of disparity estimation, the proposed I-MVC deploys independent component analysis (ICA) on the video streams to remove the correlation between the view sequences. The correlated (dependent) video sequences are decomposed into uncorrelated (independent) sequences and a mixing matrix. Each independent sequence is independently encoded by the H.264/AVC video coder. Then the mixing matrix is used at decoder to jointly decode the received independent sequences. Our experimental results show that the proposed I-MVC has better coding efficiency than conventional 3D multi-view video coder. The I-MVC gives more than 21% savings in overall bit rate and reduces the MVC computational complexity by 49% with less than 0.2 dB loss in the video peak signal to noise ratio.
Hany S. Hussein, Mostafa El-Khamy, Farhad Mehdipour, Mohamed El-Sharkawy 0001
CCNC2
2013 Soft Turbo HARQ combining
abstract
In this paper, we consider a hybrid automatic repeat request (HARQ) system with bit-interleaved coded modulation over wireless channels. At higher-order modulations, bit-level combining of log-likelihood ratios (LLRs) is sub-optimal compared to symbol-level maximal ratio combining (MRC). Since symbol level combining is not always feasible, we propose novel bit-level LLR HARQ combining techniques that deploy a modified Turbo principle at the receiver. Our proposed combining methods make use of the information available from previous transmissions to improve the performance of both detection and decoding at the current transmission. We analyze the proposed combining schemes using modified information transfer charts. We show that significant coding and throughput gains can be achieved using our proposed HARQ combining schemes with minimal extra memory requirements.
Mostafa El-Khamy, Inyup Kang
ICC1
2013 On the construction and decoding of concatenated polar codes
abstract
A scheme for concatenating the recently invented polar codes with interleaved block codes is considered. By concatenating binary polar codes with interleaved Reed-Solomon codes, we prove that the proposed concatenation scheme captures the capacity-achieving property of polar codes, while having a significantly better error-decay rate. We show that for any ε > 0, and total frame length N, the parameters of the scheme can be set such that the frame error probability is less than 2-N 1-ε, while the scheme is still capacity achieving. This improves upon 2-N 0.5-ε, the frame error probability of Arikan's polar codes. We also propose decoding algorithms for concatenated polar codes, which significantly improve the error-rate performance at finite block lengths while preserving the low decoding complexity.
Hessam Mahdavifar, Mostafa El-Khamy, Inyup Kang
ISIT2
2013 BICM performance improvement via online LLR optimization
abstract
We consider bit interleaved coded modulation (BICM) receiver performance improvement based on the concept of generalized mutual information (GMI). Increasing achievable rates of BICM receiver with GMI maximization by proper scaling of the log likelihood ratio (LLR) is investigated. While it has been shown in the literature that look-up table based LLR scaling functions matched to each specific transmission scenario may provide close to optimal solutions, this method is difficult to adapt to time-varying channel conditions. To solve this problem, an online adaptive scaling factor searching algorithm is developed. Uniform scaling factors are applied to LLRs from different bit channels of each data frame by maximizing an approximate GMI that characterizes the transmission conditions of current data frame. Numerical analysis on effective achievable rates as well as link level simulation of realistic mobile transmission scenarios indicate that the proposed method is simple yet effective.
Jinhong Wu, Mostafa El-Khamy, Inyup Kang
WCNC2
2012 Near-optimal turbo decoding in presence of SNR estimation error
abstract
Optimal iterative log-MAP (LM) decoding of turbo codes requires accurate signal to noise ratio (SNR) information. In practice, there is SNR mismatch priori to decoding due to inaccurate SNR estimation. Although max-log-MAP turbo decoding avoids the detrimental effect of SNR mismatch, its performance is inferior to LM decoding at accurate SNR estimation. In this paper, we propose two architectures for improved turbo decoding in presence of SNR mismatch. The first architecture called “SNR-Mismatch Aware Turbo (SMAT) Decoder” selects the decoder with the best performance at any SNR mismatch. The second architecture called “SNR-Mismatch Compensated Turbo (SMCT) Decoder” performs accurate SNR-mismatch estimation and compensates for the mismatch while decoding. We provide symbol-based as well as bit-level LLR histogram-based approaches for SNR mismatch estimation. We show that the proposed SMCT decoder has near-optimal performance regardless of the initial SNR mismatch. We demonstrate the effectiveness of the proposed turbo decoding architectures by Monte Carlo simulations.
Mostafa El-Khamy, Jinhong Wu, Heejin Roh, Inyup Kang
GLOBECOM1
2012 Optimized dual relay deployment for LTE-Advanced cellular systems
abstract
LTE-Advanced adopts relay deployment to support higher data rates and better coverage, especially at the cell edges which suffer from inter-cell interference. The solution of relaying is attractive because of its low cost and easy deployment. In this paper, we show that deploying just two relays per sector can significantly improve system capacity and coverage. We optimize the geometric deployment of relays that maximizes the spectral efficiency of the worst users or minimize outage. We consider the effects of different backhauling techniques such as wirelessly connected relays or wired radio remote heads. We also investigate the effect of different downlink frame structures and their impact on the system capacity or outage. The paper shows that the proper choice of the geometric deployment, frame structure and backhauling technique of the relays can improve the performance of LTE-Advanced.
Ahmed Hamdi, Mostafa El-Khamy, Mohamed El-Sharkawy 0001
WCNC2
2011 Joint space-time-view error concealment algorithms for 3D multi-view video
abstract
Efficiently compressing 3D multi-view video, while maintaining a high quality of received 3D video, is very challenging. Error Concealment (EC) algorithms have the advantage of improving the received video quality without modifications in the transmission rate or in the encoder hardware or software. To improve the quality of reconstructed 3D multi-view video, we propose different algorithms to conceal the erroneous and lost blocks of intra-coded and inter-coded frames by exploiting the spatial, temporal and inter-view correlations between frames and views. A hybrid of Space Domain Error Concealment (SDEC) and Time Domain Error Concealment (TDEC) is introduced for concealment of intra-frames errors. Three EC modes are introduced for inter-frames, which are Time Domain Error Concealment (TDEC), Inter-view Domain Error Concealment (IVDEC) and joint Time and Interview Domain Error Concealment (TIVDEC). Our simulation results show that the proposed algorithms can significantly improve the objective and subjective quality of reconstructed 3D multi-view video sequences.
Walid El Shafai, Branislav Hrusovský, Mostafa El-Khamy, Mohamed El-Sharkawy 0001
ICIP3
2011 Optimized bases compressive spectrum sensing for wideband cognitive radio
abstract
Designing efficient spectrum sensing techniques with low power consumption is crucial for the success of cognitive radio (CR). This is particularly challenging when sensing a wideband spectrum due to the high sampling rate required. Compressive sensing (CS) theory states that a signal can be measured at a rate significantly lower than the Nyquist rate and consequently reconstructed from the measurements using an optimization process. The minimum number of measurements required for reliable reconstruction of the sampled signal depends on the sparsity of the measured signal. To improve the performance of compressive spectrum sensing, we optimize the sparsifying bases used to represent the measured spectrum. Entropy-based best basis selection algorithm of Coifman and Wickerhauser (CW) is deployed to find the optimum basis. Our simulation results show that the proposed compressive sampling technique can improve the spectrum estimation accuracy and enhance the detection and false-alarm probabilities of the CR system at the same sampling rates.
Mohammed Farrag, Mostafa El-Khamy, Mohamed El-Sharkawy 0001
PIMRC2
2009 Design of rate-compatible structured LDPC codes for hybrid ARQ applications
abstract
In this paper, families of rate-compatible protograph-based LDPC codes that are suitable for incrementalredundancy hybrid ARQ applications are constructed. A systematic technique to construct low-rate base codes from a higher rate code is presented. The base codes are designed to be robust against erasures while having a good performance on error channels. A progressive node puncturing algorithm is devised to construct a family of higher rate codes from the base code. The performance of this puncturing algorithm is compared to other puncturing schemes. Using the techniques in this paper, one can construct a rate-compatible family of codes with rates ranging from 0.1 to 0.9 that are within 1 dB from the channel capacity and have good error floors.
Mostafa El-Khamy, Jilei Hou, Naga Bhushan
IEEE J. Sel. Areas Commun.1
2009 Performance of sphere decoding of block codes
abstract
A sphere decoder searches for the closest lattice point within a certain search radius. The search radius provides a tradeoff between performance and complexity. We focus on analyzing the performance of sphere decoding of linear block codes. We analyze the performance of soft-decision sphere decoding on AWGN channels and a variety of modulation schemes. A hard-decision sphere decoder is a bounded distance decoder with the corresponding decoding radius. We analyze the performance of hard-decision sphere decoding on binary andq-ary symmetric channels. An upper bound on the performance of maximum-likelihood decoding of linear codes defined overFq(e.g. Reed- Solomon codes) and transmitted overq-ary symmetric channels is derived and used in the analysis. We then discuss sphere decoding of general block codes or lattices with arbitrary modulation schemes. The tradeoff between the performance and complexity of a sphere decoder is then discussed.
Mostafa El-Khamy, Haris Vikalo, Babak Hassibi, Robert J. McEliece
IEEE Trans. Commun.1
2006 H-ARQ Rate-Compatible Structured LDPC Codes
abstract
In this paper, we design families of rate-compatible structured LDPC codes suitable for hybrid ARQ applications with high throughput. We devise a systematic technique of low complexity for the design of structured low-rate LDPC codes from higher rate ones. These codes have a good performance on the AWGN channel and are robust against erasures and puncturing. The codes designed here are protograph-based codes and have fast encoding and decoding structures. These low rate codes are used as the parent codes of rate-compatible families. Then, we propose a number of algorithms for puncturing the codes in a rate compatible manner to construct codes of higher rates. The two most promising ones are the random puncturing search technique and progressive node puncturing. We show that using the techniques in this paper one could construct a high throughput rate compatible family with codes whose rates are in the range from 0.1 to 0.9 and which are within 1 dB from the channel capacity and have good error floors
Mostafa El-Khamy, Jilei Hou, Naga Bhushan
ISIT1
2006 On the Performance of Sphere Decoding of Block Codes
abstract
The performance of sphere decoding of block codes over a variety of channels is investigated. We derive a tight bound on the performance of maximum likelihood decoding of linear codes on q-ary symmetric channels. We use this result to bound the performance of q-ary hard decision sphere decoders. We also derive a tight bound on the performance of soft decision sphere decoders on the AWGN channel for BPSK and M-PSK modulated block codes. The performance of soft decision sphere decoding of arbitrary finite lattices or block codes is also analyzed
Mostafa El-Khamy, Haris Vikalo, Babak Hassibi, Robert J. McEliece
ISIT1
2006 Iterative algebraic soft-decision list decoding of Reed-Solomon codes
abstract
In this paper, we present an iterative soft-decision decoding algorithm for Reed-Solomon (RS) codes offering both complexity and performance advantages over previously known decoding algorithms. Our algorithm is a list decoding algorithm which combines two powerful soft-decision decoding techniques which were previously regarded in the literature as competitive, namely, the Koetter-Vardy algebraic soft-decision decoding algorithm and belief-propagation based on adaptive parity-check matrices, recently proposed by Jiang and Narayanan. Building on the Jiang-Narayanan algorithm, we present a belief-propagation-based algorithm with a significant reduction in computational complexity. We introduce the concept of using a belief-propagation-based decoder to enhance the soft-input information prior to decoding with an algebraic soft-decision decoder. Our algorithm can also be viewed as an interpolation multiplicity assignment scheme for algebraic soft-decision decoding of RS codes.
Mostafa El-Khamy, Robert J. McEliece
IEEE J. Sel. Areas Commun.1
2005 The partition weight enumerator of MDS codes and its applications
abstract
A closed form formula of the partition weight enumerator of maximum distance separable (MDS) codes is derived for an arbitrary number of partitions. Using this result, some properties of MDS codes are discussed. The results are extended for the average binary image of MDS codes in finite fields of characteristic two. As an application, we study the multiuser error probability of Reed Solomon codes
Mostafa El-Khamy, Robert J. McEliece
ISIT1
2005 Bounds on the performance of sphere decoding of linear block codes
abstract
A sphere decoder searches for the closest lattice point within a certain search radius. The search radius provides a tradeoff between performance and complexity. We derive tight upper bounds on the performance of sphere decoding of linear block codes. The performance of soft-decision sphere decoding on AWGN channels as well as that of hard-decision sphere decoding on binary symmetric channels is analyzed.
Mostafa El-Khamy, Haris Vikalo, Babak Hassibi
ITW1
2004 Performance enhancements for algebraic soft decision decoding of Reed-Solomon codes
abstract
In an attempt to determine the ultimate capabilities of the Sudan-Guruswami-Sudan-Kotter-Vardy algebraic soft decision decoding algorithm for Reed-Solomon codes, we present a new method, based on the Chernoff bound, for constructing multiplicity matrices. In many cases, this technique predicts that the potential performance of ASD decoding of RS codes is significantly better than previously thought.
Mostafa El-Khamy, Robert J. McEliece, Jonathan Harel
ISIT1