Rabab K. Ward

dblp:w/RababKreidiehWard · also Rabab Kreidieh Ward, Rabab Ward · DBLP profile ↗
← Back
216ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0002-2471-1902ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 152 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 40 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 2 since 2021Systems, architecture and hardware · 7 · 1 since 2021Computer networks · 6 · 1 since 2021Security and privacy · 5Databases, data management, data science and information retrieval · 5Human-computer interaction and ubiquitous computing · 3 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 ArchitectHead: Continuous Level of Detail Control for 3D Gaussian Head Avatars
abstract
3D Gaussian Splatting (3DGS) has enabled photorealistic and real-time rendering of 3D head avatars. Existing 3DGS-based avatars typically rely on tens of thousands of 3D Gaussian points (Gaussians), with the number of Gaussians fixed after training. However, many practical applications require adjustable levels of detail (LOD) to balance rendering efficiency and visual quality. In this work, we propose "ArchitectHead", the first framework for creating 3D Gaussian head avatars that support continuous control over LOD. Our key idea is to parameterize the Gaussians in a 2D UV feature space and propose a UV feature field composed of multi-level learnable feature maps to encode their latent features. A lightweight neural network-based decoder then transforms these latent features into 3D Gaussian attributes for rendering. ArchitectHead controls the number of Gaussians by dynamically resampling feature maps from the UV feature field at the desired resolutions. This method enables efficient and continuous control of LOD without retraining. Experimental results show that ArchitectHead achieves state-of-the-art (SOTA) quality in self and cross-identity reenactment tasks at the highest LOD, while maintaining near SOTA performance at lower LODs. At the lowest LOD, our method uses only 6.2% of the Gaussians while the quality degrades moderately (L1 Loss +7.9%, PSNR −0.97%, SSIM −0.6%, LPIPS Loss +24.1%), and the rendering speed nearly doubles. Project homepage: https://peizhiyan.github.io/docs/architect/.
Peizhi Yan, Rabab K. Ward, Qiang Tang 0002, Shan Du 0001
WACV2
2025 Estimating Virtual Camera FOV to Reduce Perspective Shape Distortion in 2D-to-3D Face Reconstruction
abstract
Existing image-based 3D face reconstruction methods rely on a virtual camera to project the reconstructed 3D face onto the 2D image plane for comparison with the input image, a crucial step for accurate results. To simplify the reconstruction process, these methods often use fixed camera intrinsics and assume minimal perspective distortion, overlooking the varying distortion levels in "in-the-wild" images and leading to inaccuracies in reconstructed 3D face shapes. To address this issue, we propose estimating the virtual camera’s optimal field-of-view (FOV) for a given image, enabling consistent 3D face reconstruction across varying distortion levels. We introduce two synthetic datasets: one to train our FOV estimation network (FOV-Net) and another to evaluate its performance and reconstruction accuracy. We use the FOV-Net predicted FOV to initialize the camera, which is used in the fitting-based reconstruction process. Experiments show that our approach significantly improves reconstruction consistency under different levels of perspective distortion.
Peizhi Yan, Rabab K. Ward, Qiang Tang 0002, Shan Du 0001
ICIP2
2025 Gaussian Déjà-vu: Creating Controllable 3D Gaussian Head-Avatars with Enhanced Generalization and Personalization Abilities
abstract
Recent advancements in 3D Gaussian Splatting (3DGS) have unlocked significant potential for modeling 3D head avatars, providing greater flexibility than mesh-based methods and more efficient rendering compared to NeRF-based approaches. Despite these advancements, the creation of controllable 3DGS-based head avatars remains time-intensive, often requiring tens of minutes to hours. To expedite this process, we here introduce the “Gaussian Déjà-vu” framework, which first obtains a generalized model of the head avatar and then personalizes the result. The generalized model is trained on large 2D (synthetic and real) image datasets. This model provides a well-initialized 3D Gaussian head that is further refined using a monocular video to achieve the personalized head avatar. For personalizing, we propose learnable expression-aware rectification blendmaps to correct the initial 3D Gaussians, ensuring rapid convergence without the reliance on neural networks. Experiments demonstrate that the proposed method meets its objectives. It outperforms state-of-the-art 3D Gaussian head avatars in terms of photorealistic quality as well as reduces training time consumption to at least a quarter of the existing methods, producing the avatar in minutes. Project homepage: https://peizhiyan.github.io/docs/dejavu
Peizhi Yan, Rabab K. Ward, Qiang Tang 0002, Shan Du 0001
WACV2
2025 Neural 3D Face Shape Stylization Based on Single Style Template via Weakly Supervised Learning
abstract
3D Face shape stylization refers to transforming a realistic 3D face shape into a different style, such as a cartoon face style. To solve this problem, this paper proposes modeling this task as a deformation transfer problem. This approach significantly reduces labor costs, as the artists would only need to create a single template for each face style. Realistic facial features of the original 3D face e.g. the nose or chin shape, would thus be automatically transferred to those in the style template. Deformation transfer methods, however, have two drawbacks. They are slow and they require re-optimization for every new input face. To address these weaknesses, we propose a neural network-based 3D face shape stylization method. This method is trained through weakly supervised learning, and its template's structure is preserved using our novel template-guided mesh smoothing regularization. Our method is the first learning-based deformation transfer method for 3D face shape stylization. Its employment offers the useful and practical benefit of not requiring paired training data. The experiments show that the quality of the stylized faces obtained by our method is comparable to that of the traditional deformation transfer method, achieving an average Chamfer Distance of approximately 0.01 mm. However, our approach significantly boosts the processing speed, achieving a rate approximately 3,000 times faster than the traditional deformation transfer.
Peizhi Yan, Rabab K. Ward, Qiang Tang 0002, Shan Du 0001
IEEE Trans. Vis. Comput. Graph.2
2024 PoseGen: Learning to Generate 3D Human Pose Dataset with NeRF
abstract
This paper proposes an end-to-end framework for generating 3D human pose datasets using Neural Radiance Fields (NeRF). Public datasets generally have limited diversity in terms of human poses and camera viewpoints, largely due to the resource-intensive nature of collecting 3D human pose data. As a result, pose estimators trained on public datasets significantly underperform when applied to unseen out-of-distribution samples. Previous works proposed augmenting public datasets by generating 2D-3D pose pairs or rendering a large amount of random data. Such approaches either overlook image rendering or result in suboptimal datasets for pre-trained models. Here we propose PoseGen, which learns to generate a dataset (human 3D poses and images) with a feedback loss from a given pre-trained pose estimator. In contrast to prior art, our generated data is optimized to improve the robustness of the pre-trained model. The objective of PoseGen is to learn a distribution of data that maximizes the prediction error of a given pre-trained model. As the learned data distribution contains OOD samples of the pre-trained model, sampling data from such a distribution for further fine-tuning a pre-trained model improves the generalizability of the model. This is the first work that proposes NeRFs for 3D human data generation. NeRFs are data-driven and do not require 3D scans of humans. Therefore, using NeRF for data generation is a new direction for convenient user-specific data generation. Our extensive experiments show that the proposed PoseGen improves two baseline models (SPIN and HybrIK) on four datasets with an average 6% relative improvement.
Mohsen Gholami, Rabab K. Ward, Z. Jane Wang 0001
AAAI2
2023 Learning Disentangled Features for Nerf-Based Face Reconstruction
abstract
The 3D-aware parametric face model named HeadNeRF achieved advantages in rendering photo-realistic face images. However, it has two limitations: (1) it uses single-image fitting reconstruction that is slow and prone to overfitting; (2) it lacks explicit 3D geometry information, making using semantic facial-parts-based loss challenging. This paper presents a 3D-aware face reconstruction learning framework tailored for HeadNeRF to address the limitations. We train a face encoder network that can directly learn the disentangled features for facial reconstruction to address the first limitation. For the second limitation, we introduce a lightweight semantic face segmentation network and facial-parts-based loss function to improve the reconstruction accuracy and quality. Our experiments show that the proposed method achieves a low reconstruction time consumption and enhanced reconstruction accuracy. Project page: https://peizhiyan.github.io/docs/headnerf+
Peizhi Yan, Rabab K. Ward, Dan Wang 0011, Qiang Tang 0002, Shan Du 0001
ICIP2
2023 Automatic labeling of Parkinson's Disease gait videos with weak supervision
Mohsen Gholami, Rabab K. Ward, Ravneet Mahal, Maryam S. Mirian, Kevin Yen, Kye Won Park, Martin J. McKeown, Z. Jane Wang 0001
Medical Image Anal.2
2022 NEO-3DF: Novel Editing-Oriented 3D Face Creation and Reconstruction
Peizhi Yan, James Gregson, Qiang Tang 0002, Rabab K. Ward, Shan Du 0001
ACCV (1)4
2022 AdaptPose: Cross-Dataset Adaptation for 3D Human Pose Estimation by Learnable Motion Generation
abstract
This paper addresses the problem of cross-dataset generalization of 3D human pose estimation models. Testing a pre-trained 3D pose estimator on a new dataset results in a major performance drop. Previous methods have mainly addressed this problem by improving the diversity of the training data. We argue that diversity alone is not sufficient and that the characteristics of the training data need to be adapted to those of the new dataset such as camera view-point, position, human actions, and body size. To this end, we propose AdaptPose, an end-to-end framework that generates synthetic 3D human motions from a source dataset and uses them to fine-tune a 3D pose estimator. AdaptPose follows an adversarial training scheme. From a source 3D pose the generator generates a sequence of 3D poses and a camera orientation that is used to project the generated poses to a novel view. Without any 3D labels or camera information AdaptPose successfully learns to create synthetic 3D poses from the target dataset while only being trained on 2D poses. In experiments on the Human3.6M, MPI-INF-3DHp, 3DPW, and Ski-Pose datasets our method outperforms previous work in cross-dataset evaluations by 14% and previous semi-supervised learning methods that use partial 3D annotations by 16%.
Mohsen Gholami, Bastian Wandt, Helge Rhodin, Rabab K. Ward, Z. Jane Wang 0001
CVPR4
2022 Self-supervised 3D human pose estimation from video
Mohsen Gholami, Ahmad Rezaei, Helge Rhodin, Rabab K. Ward, Z. Jane Wang 0001
Neurocomputing4
2021 Multi-view 3D Reconstruction with Transformers
abstract
Deep CNN-based methods have so far achieved the state of the art results in multi-view 3D object reconstruction. Despite the considerable progress, the two core modules of these methods - view feature extraction and multi-view fusion, are usually investigated separately, and the relations among multiple input views are rarely explored. Inspired by the recent great success in Transformer models, we reformulate the multi-view 3D reconstruction as a sequence-to-sequence prediction problem and propose a framework named 3D Volume Transformer. Unlike previous CNN-based methods using a separate design, we unify the feature extraction and view fusion in a single Transformer network. A natural advantage of our design lies in the exploration of view-to-view relationships using self-attention among multiple unordered inputs. On ShapeNet - a large-scale 3D reconstruction benchmark, our method achieves a new state-of-the-art accuracy in multi-view reconstruction with fewer parameters (70% less) than CNN-based methods. Experimental results also suggest the strong scaling capability of our method. Our code will be made publicly available.
Dan Wang 0011, Xinrui Cui, Xun Chen 0001, Zhengxia Zou, Tianyang Shi, Tim Salcudean, Z. Jane Wang 0001, Rabab K. Ward
ICCV8
2021 Towards Universal Physical Attacks On Cascaded Camera-Lidar 3d Object Detection Models
abstract
We propose a universal and physically realizable adversarial attack on a cascaded multi-modal deep learning network (DNN), in the context of self-driving cars. DNNs have achieved high performance in 3D object detection, but they are known to be vulnerable to adversarial attacks. These attacks have been heavily investigated in the RGB image domain and more recently in the point cloud domain, but rarely in both domains simultaneously - a gap to be filled in this paper. We use a single 3D mesh and differentiable rendering to explore how perturbing the mesh’s geometry and texture can reduce the robustness of DNNs to adversarial attacks. We attack a prominent cascaded multi-modal DNN, the Frustum-Pointnet model. Using the popular KITTI benchmark, we showed that the proposed universal multi-modal attack was successful in reducing the model’s ability to detect a car by nearly 73%. This work can aid in the understanding of what the cascaded RGB-point cloud DNN learns and its vulnerability to adversarial attacks.
Mazen Abdelfattah, Kaiwen Yuan, Z. Jane Wang 0001, Rabab K. Ward
ICIP4
2021 Interpolation of CT Projections by Exploiting Their Self-Similarity and Smoothness
abstract
As the medical usage of computed tomography (CT) grows, the radiation dose should remain at a low level to reduce the health risks. Therefore, there is a need for algorithms that can reconstruct high-quality images from low-dose scans. In this regard, most of the recent studies have focused on iterative reconstruction algorithms, and little attention has been paid to restoration of the projection measurements, i.e., the sinogram. In this paper, we propose a novel sinogram interpolation algorithm. The proposed algorithm exploits the self-similarity and smoothness of the sinogram. Sinogram self-similarity is modeled in terms of the similarity of small blocks extracted from stacked projections. The smoothness is modeled via second-order total variation. Experiments with simulated and real CT data show that sinogram interpolation with the proposed algorithm leads to a substantial improvement in the quality of the reconstructed image, especially on low-dose scans. The proposed method can result in a significant reduction in the number of projection measurements.
Davood Karimi, Rabab K. Ward
ICIP2
2021 Adversarial Attacks on Camera-LiDAR Models for 3D Car Detection
abstract
Most autonomous vehicles (AVs) rely on LiDAR and RGB camera sensors for perception. Using these point cloud and image data, perception models based on deep neural nets (DNNs) have achieved state-of-the-art performance in 3D detection. The vulnerability of DNNs to adversarial attacks have been heavily investigated in the RGB image domain and more recently in the point cloud domain, but rarely in both domains simultaneously. Multi-modal perception systems used in AVs can be divided into two broad types: cascaded models which use each modality independently, and fusion models which learn from different modalities simultaneously. We propose a universal and physically realizable adversarial attack for each type, and study and contrast their respective vulnerabilities to attacks. We place a single adversarial object with specific shape and texture on top of a car with the objective of making this car evade detection. Evaluating on the popular KITTI benchmark, our adversarial object made the host vehicle escape detection by each model type more than 50% of the time. The dense RGB input contributed more to the success of the adversarial attacks on both cascaded and fusion models.
Mazen Abdelfattah, Kaiwen Yuan, Z. Jane Wang 0001, Rabab K. Ward
IROS4
2021 Improving compression efficiency of HEVC using perceptual coding
Sima Valizadeh, Panos Nasiopoulos, Rabab K. Ward
Multim. Tools Appl.3
2021 Semi-dilated convolutional neural networks for epileptic seizure prediction
Ramy Hussein, Rabab K. Ward, Martin J. McKeown
Neural Networks3
2021 Perception matters: Exploring imperceptible and transferable anti-forensics for GAN-generated fake face imagery detection
Xin Ding 0004, Yixin Yang 0001, Rabab K. Ward, Z. Jane Wang 0001
Pattern Recognit. Lett.5
2021 Dual Pilot Scheme (DPS) and Its Application in Massive MIMO
abstract
The pilot scheme currently used in 5th generation (5G) cellular networks assigns the same set of orthogonal pilot signals to all cells. This results in inter-cell interference, also known as pilot contamination, which can significantly degrade performance, especially in massive multi-input multi-output (MIMO) systems. To mitigate this interference, we propose a novel Dual Pilot Scheme (DPS) that assigns a slightly modified set of nearly-orthogonal pilot signals. DPS is a general scheme that can be implemented in any wireless communication system, including 5G and beyond. We demonstrate the integration of DPS in a massive MIMO system in both microscopic and macroscopic levels and analytically prove that DPS enables more accurate estimates of the channel state information in the minimum mean-squared error sense, under the independent identically distributed (i.i.d.) and the correlated Rayleigh fading wireless communication channel models. We further validate and demonstrate the advantages of DPS over various channel models of massive MIMO 5G technology by extensive simulations.
A. Nasser Aljalai, Chen Feng 0001, Victor C. M. Leung, Rabab K. Ward
IEEE Trans. Commun.4
2021 Unifying Top-Down Views by Task-Specific Domain Adaptation
abstract
In this article, we aim to learn a unified representation of images from satellite/aerial/ground views by exploring their underlying correlations. Inspired by recent advances in domain adaptation (DA), we propose a novel task-specific DA method for this purpose. Different from traditional DA methods, this proposed method not only applies task-specific classifiers1but also introduces domain-specific tasks for different domains during the adaptation process. The experiments are conducted on two newly proposed ground-/satellite-to-aerial scene adaptation (GSSA) data sets. Since the semantic gap between the ground/satellite scenes and the aerial scenes is much larger than that between ground scenes, the DA task between these scenes is more challenging than traditional DA tasks. On GSSA data sets, we not only demonstrate the proposed unsupervised DA method but also explore the few-shot DA in the discussion section. The proposed method is easy to implement, and our method substantially outperforms the state-of-the-art methods on the studied data sets.We hope that the proposed method for the novel GSSA data sets can be a good baseline for future researchers. The related data sets/codes will be available online.
Jianzhe Lin, Tianze Yu, Lichao Mou, Xiao Xiang Zhu 0001, Rabab K. Ward, Z. Jane Wang 0001
IEEE Trans. Geosci. Remote. Sens.5
2021 Interpreting Bottom-Up Decision-Making of CNNs via Hierarchical Inference
abstract
With the great success of convolutional neural networks (CNNs), interpretation of their internal network mechanism has been increasingly critical, while the network decision-making logic is still an open issue. In the bottom-up hierarchical logic of neuroscience, the decision-making process can be deduced from a series of sub-decision-making processes from low to high levels. Inspired by this, we propose the Concept-harmonized HierArchical INference (CHAIN) interpretation scheme. In CHAIN, a network decision-making process from shallow to deep layers is interpreted by the hierarchical backward inference based on visual concepts from high to low semantic levels. Firstly, we learned a general hierarchical visual-concept representation in CNN layered feature space by concept harmonizing model on a large concept dataset. Secondly, for interpreting a specific network decision-making process, we conduct the concept-harmonized hierarchical inference backward from the highest to the lowest semantic level. Specifically, the network learning for a target concept at a deeper layer is disassembled into that for concepts at shallower layers. Finally, a specific network decision-making process is explained as a form of concept-harmonized hierarchical inference, which is intuitively comparable to the bottom-up hierarchical visual recognition way. Quantitative and qualitative experiments demonstrate the effectiveness of the proposed CHAIN at both instance and class levels.
Dan Wang 0011, Xinrui Cui, Xun Chen 0001, Rabab K. Ward, Z. Jane Wang 0001
IEEE Trans. Image Process.4
2020 Epileptic Seizure Prediction: A Semi-Dilated Convolutional Neural Network Architecture
abstract
Accurate prediction of epileptic seizures has remained elusive, despite the many advances in machine learning and time-series classification. In this work, we develop a convolutional network module that exploits Electroencephalogram (EEG) scalograms to distinguish between the pre-seizure and normal brain activities. Since these scalograms have rectangular image shapes with many more temporal bins than spectral bins, the presented module uses “semi-dilated convolutions” to create a proportional non-square receptive field. The proposed semi-dilated convolutions support exponential expansion of the receptive field over the long dimension (image width, i.e. time) while maintaining high resolution over the short dimension (image height, i.e., frequency). The proposed architecture comprises a set of co-operative semi-dilated convolutional blocks, each block has a stack of parallel semi-dilated convolutional modules with different dilation rates. Results show that our proposed solution outperforms the state-of-the-art methods, achieving seizure prediction sensitivity scores of 88.45% and 89.52% for the American Epilepsy Society and Melbourne University EEG datasets, respectively.
Ramy Hussein, Rabab K. Ward, Martin J. McKeown
ICPR3
2020 A binary water wave optimization for feature selection
Abdel-Monem M. Ibrahim, Mohamed A. Tawhid, Rabab K. Ward
Int. J. Approx. Reason.3
2020 Xnet: Task-specific attentional domain adaptation for satellite-to-aerial scene
Jianzhe Lin, Kaiwen Yuan, Rabab K. Ward, Z. Jane Wang 0001
Neurocomputing3
2020 DT-LET: Deep transfer learning by exploring where to transfer
Jianzhe Lin, Liang Zhao 0005, Qi Wang 0009, Rabab K. Ward, Z. Jane Wang 0001
Neurocomputing4
2019 Edge-based compression and classification for smart healthcare systems: Concept, implementation and evaluation
Alaa Awad, Carla Fabiana Chiasserini, Amr Mohamed 0001, Ali Jaoua, Rabab K. Ward
Expert Syst. Appl.6
2019 Medical Image Fusion via Convolutional Sparsity Based Morphological Component Analysis
abstract
In this letter, a sparse representation (SR) model named convolutional sparsity based morphological component analysis (CS-MCA) is introduced for pixel-level medical image fusion. Unlike the standard SR model, which is based on single image component and overlapping patches, the CS-MCA model can simultaneously achieve multi-component and global SRs of source images, by integrating MCA and convolutional sparse representation (CSR) into a unified optimization framework. For each source image, in the proposed fusion method, the CSRs of its cartoon and texture components are first obtained by the CS-MCA model using pre-learned dictionaries. Then, for each image component, the sparse coefficients of all the source images are merged and the fused component is accordingly reconstructed using the corresponding dictionary. Finally, the fused image is calculated as the superposition of the fused cartoon and texture components. Experimental results demonstrate that the proposed method can outperform some benchmarking and state-of-the-art SR-based fusion methods in terms of both visual perception and objective assessment.
Yu Liu 0023, Xun Chen 0001, Rabab K. Ward, Z. Jane Wang 0001
IEEE Signal Process. Lett.3
2019 Coarse-to-Fine Image DeHashing Using Deep Pyramidal Residual Learning
abstract
Image dehashing refers to the process of inferring images by inverting image hashes. Recently, image dehashing from real-valued image retrieval hashes is shown feasible using deep convolutional neural networks. However, the perceptual quality of dehashed images is challenged when real-valued hashes are quantized to less bits. Besides, the scalability to larger or color image dehashing is limited in the previous dehashing network. To this end, we propose a pyramidal long-range residual-learning network (PyLRR-Net). PyLRR-Net is a pyramidal image reconstruction network to dehash images in a progressive manner. At each image scale, we design and insert a long-range residual block to refine the coarse image reconstruction leveraging deep residual learning. Experiments on both grayscale and color image datasets show that the proposed PyLRR-Net outperforms previous work in terms of image dehashing quality, scalability, and flexibility for large and color image dehashing problems.
Rabab K. Ward, Z. Jane Wang 0001
IEEE Signal Process. Lett.2
2018 Robust Detection of Epileptic Seizures Using Deep Neural Networks
abstract
Robust detection of epileptic seizures in the presence of inevitable artifacts in Electroencephalogram (EEG) signals is addressed. The EEG dataset considered contains 300 signals recorded from 15 volunteers. Current seizure detection systems achieve good performance when the EEG data is entirely free of noise. However, their performance drastically decays with authentic EEG data polluted by real artifacts. We introduce a robust seizure detection method that can address clean and noisy data. The proposed method uses Long Short-Term Memory (LSTM) neural networks to extract the representative EEG features pertinent to seizures. Experimental results show that the proposed method beats existing methods by achieving 100% classification accuracy. Our method is also shown to be robust against the common EEG artifacts (e.g., muscle activities and eye-blinking) and white noise.
Ramy Hussein, Hamid Palangi, Z. Jane Wang 0001, Rabab K. Ward
ICASSP4
2018 Deep Transfer Learning for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) includes a vast quantities of samples, large number of bands, as well as randomly occurring redundancy. Classifying such complex data is challenging, and the classification performance generally is affected significantly by the amount of labeled training samples. Collecting such labeled training samples is labor and time consuming, motivating the idea of borrowing and reusing labeled samples from other preexisting related images. Therefore transfer learning, which can mitigate the semantic gap between existing and new HSI, has recently drawn increasing research attention. However, existing transfer learning methods for HSI which concentrated on how to overcome the divergence among images, may neglect the high level latent features during the transfer learning process. In this paper, we present two novel ideas based on this observation. We propose constructing and connecting higher level features for the source and target HSI data, to further overcome the cross-domain disparity. Different from existing methods, no priori knowledge on the target domain is needed for the proposed classification framework, and the proposed framework works for both homogeneous and heterogenous HSI data. Experimental results on real world hyperspectral images indicate the significance of the proposed method in HSI classification.
Jianzhe Lin, Rabab K. Ward, Z. Jane Wang 0001
MMSP2
2018 Robust detection of epileptic seizures based on L1-penalized robust regression of EEG signals
Ramy Hussein, Mohamed Elgendi, Z. Jane Wang 0001, Rabab K. Ward
Expert Syst. Appl.4
2018 Perceptual rate distortion optimization of 3D-HEVC using PSNR-HVS
Sima Valizadeh, Panos Nasiopoulos, Rabab K. Ward
Multim. Tools Appl.3
2017 Robust greedy deep dictionary learning for ECG arrhythmia classification
abstract
This work proposes a new deep learning method which we call robust deep dictionary learning RDDL. RDDL is suitable for learning representations from signals corrupted with sparse but large outliers such as artifacts and noise that are more heavy tailed than Gaussian distributions. Such outliers are common in biomedical signals e.g. EEG and ECG. RDDL learns multiple levels of non-linear dictionaries for representing the data. Instead of the standard Euclidean cost function that is usually employed in dictionary learning, we propose a robust l1-norm cost function. In order to achieve sparse representation, an l1-norm is imposed on the learned representation. The `depth' arises from the fact that multiple levels of dictionaries are learnt. The full formulation is solved in a greedy fashion, one layer at a time. To study the extent of usefulness of RDDL, we first benchmark it with two wellknown deep learning tools - the stacked denoising autoencoder and the deep belief network methods; experiments are carried out on benchmark deep learning datasets - MNIST, CIFAR-10 and SVHN. In all cases, our method yields the best results. Then the proposed method is used for learning representations of ECG data (containing arficacts) and for their classification using the MIT-BIH arrhythmia classification database. We compare it with traditional techniques as well as on deep learning tools. Our method yields the best results.
Angshul Majumdar, Rabab K. Ward
IJCNN2
2017 Eliminating Pilot Contamination Using Dual Pilot Sequences in Massive MIMO
abstract
The uplink transmission in a Massive MIMO system is studied. Pilot contamination during the uplink training is the main inherent limitation that degrades the performance of Massive MIMO. Current approaches are using the same pilot sequences for every cell, leading to the so-called inter-cell interference. In this paper, a novel method that employs dual pilot sequences is proposed where different cells are assigned with different pilot sequences (called them cells' IDs). In particular, each cell has the same set of orthogonal pilot sequences together with a unique pilot sequence (cell's ID). This dual structure mitigates pilot contamination, achieving better channel estimation at various signal-to-noise ratios.
A. Nasser Aljalai, Chen Feng 0001, Victor C. M. Leung, Rabab K. Ward
VTC Fall4
2017 Convolutional Deep Stacking Networks for distributed compressive sensing
Hamid Palangi, Rabab K. Ward, Li Deng 0001
Signal Process.2
2016 Robust dictionary learning: Application to signal disaggregation
abstract
It is well known that the Euclidean norm is sensitive to outliers; yet it is widely used for minimizing it is easy. Dictionary learning is no exception - the l2-norm allows for easy update of the basis/dictionary atoms. In this work, we propose a robust dictionary learning method that is based on minimizing the robust l1-norm. The ensuing optimization is solved using the Split Bregman approach. We apply the proposed technique to signal (energy and water) disaggregation and show that it excels over existing dictionary learning techniques (based on l2-norm).
Angshul Majumdar, Rabab K. Ward
ICASSP2
2016 Exploiting correlations among channels in distributed compressive sensing with convolutional deep stacking networks
abstract
This paper addresses the compressive sensing with Multiple Measurement Vectors (MMV) problem where the correlation amongst the different sparse vectors (channels) are used to improve the reconstruction performance. We propose the use of Convolutional Deep Stacking Networks (CDSN), where the correlations amongst the channels are captured by a moving window containing the "residuals" of different sparse vectors. We develop a greedy algorithm that exploits the structure captured by the CDSN to reconstruct the sparse vectors. Using a natural image dataset, we compare the performance of the proposed algorithm with two types of reconstruction algorithms: Simultaneous Orthogonal Matching Pursuit (SOMP) which is a greedy solver and the model-based Bayesian approaches that also exploit correlation among channels. We show experimentally that our proposed method outperforms these popular methods and is almost as fast as the greedy methods.
Hamid Palangi, Rabab K. Ward, Li Deng 0001
ICASSP2
2016 Neural Network Conditional Random Fields for Self-Paced Brain Computer Interfaces
abstract
The task of classifying EEG signals for self-paced Brain Computer Interface (BCI) applications is extremely challenging. This difficulty in classification of self-paced data stems from the fact that the system has no clue about the start time of a control task and the data contains a large number of periods during which the user has no intention to control the BCI. Therefore, to improve the performance of the BCI, it is imperative to exploit the characteristics of the EEG data as much as possible. For motor imagery based self-paced BCIs, during motor imagery task the EEG signal of each subject goes through several internal state changes. Applying appropriate classifiers that can exploit the temporal correlation in EEG data can enhance the performance of the BCI. In this paper, we propose an algorithm which is able to capture the temporal correlation of the EEG signal. We compare the performance of our algorithm that is based on neural network conditional random fields to two well-known dynamic classifiers, the Hidden Markov Models and Conditional Random Fields and to the static classifier, Support Vector Machines. We compare these methods using the data from SM2 dataset, and we show that our algorithm yields results that are considerably superior to the other approaches in terms of the Area Under the Curve (AUC) of the BCI system.
Hossein Bashashati, Rabab K. Ward, Ali Bashashati, Amr M. Mohamed
ICMLA2
2016 Energy Efficient EEG Monitoring System for Wireless Epileptic Seizure Detection
abstract
Wireless EEG monitoring systems have been successfully used for seizure detection outside clinical settings. The wireless EEG sensor nodes consume a considerable amount of battery energy to acquire, encode and transmit the data to the server side. In this paper, we introduce energy-efficient monitoring systems to increase the sensors' battery lifetime. Specifically, we propose a feature extraction method that is robust to artifacts and can effectively select the most discriminant features relevant to seizures. Second, we show how to use the missing at random (MAR) method to reduce the energy required at the sensor node for data transmission without compromising the seizure detection accuracy at the server side. Finally, we show how the expectation maximization (EM) method is used at the server side to accurately substitute the missing values. The performance of the proposed scheme is compared to those of the state-of-the art methods, and is shown to achieve less power consumption without compromising the seizure detection accuracy.
Ramy Hussein, Rabab K. Ward, Z. Jane Wang 0001, Amr Mohamed 0001
ICMLA2
2016 Class-wise deep dictionaries for EEG classification
abstract
In this work we propose a classification framework called class-wise deep dictionary learning (CWDDL). For each class, multiple levels of dictionaries are learnt using features from the previous level as inputs (for first level the input is the raw training sample). It is assumed that the cascaded dictionaries form a basis for expressing test samples for that class. Based on this assumption sparse representation based classification is employed. Benchmarking experiments have been carried out on some deep learning datasets (MNIST and its variations, CIFAR and SVHN); our proposed method has been compared with Deep Belief Network (DBN), Stacked Autoencoder, Convolutional Neural Net (CNN) and Label Consistent KSVD (dictionary learning). We find that our proposed method yields better results than these techniques and requires much smaller run-times. The technique is applied for Brain Computer Interface (BCI) classification problems using EEG signals. For this problem our method performs significantly better than Convolutional Deep Belief Network(CDBN).
Prerna Khurana, Angshul Majumdar, Rabab K. Ward
IJCNN3
2016 Real-time reconstruction of EEG signals from compressive measurements via deep learning
abstract
To elongate the battery life of sensors worn in wireless body area networks, recent studies have advocated compressing the acquired biological signals before transmitting them. The signals are compressed using compressive sensing (CS), by projecting them onto a lower dimension. The original signals are then recovered using CS recovery techniques at the base station, where the computational power is assumed to be abundant. This assumption however is not entirely true when a mobile phone acts as the base station. The computational capacity of a mobile phone is limited; therefore solving the CS recovery problem in the phone would be time consuming. In many cases (e,g. heart stroke detection or monitoring applications) this latency cannot be tolerated. In this work we propose a new technique to solve the inverse problem using stacked autoencoders. We show that the reconstruction of the proposed method can be done in real-time, and there is only a slight degradation in accuracy compared to CS based inversion methods.
Angshul Majumdar, Rabab K. Ward
IJCNN2
2016 Image Fusion With Convolutional Sparse Representation
abstract
As a popular signal modeling technique, sparse representation (SR) has achieved great success in image fusion over the last few years with a number of effective algorithms being proposed. However, due to the patch-based manner applied in sparse coding, most existing SR-based fusion methods suffer from two drawbacks, namely, limited ability in detail preservation and high sensitivity to misregistration, while these two issues are of great concern in image fusion. In this letter, we introduce a recently emerged signal decomposition model known as convolutional sparse representation (CSR) into image fusion to address this problem, which is motivated by the observation that the CSR model can effectively overcome the above two drawbacks. We propose a CSR-based image fusion framework, in which each source image is decomposed into a base layer and a detail layer, for multifocus image fusion and multimodal image fusion. Experimental results demonstrate that the proposed fusion methods clearly outperform the SR-based methods in terms of both objective assessment and visual quality.
Yu Liu 0023, Xun Chen 0001, Rabab K. Ward, Z. Jane Wang 0001
IEEE Signal Process. Lett.3
2016 Deep Sentence Embedding Using Long Short-Term Memory Networks: Analysis and Application to Information Retrieval
abstract
This paper develops a model that addresses sentence embedding, a hot topic in current natural language processing research, using recurrent neural networks (RNN) with Long Short-Term Memory (LSTM) cells. The proposed LSTM-RNN model sequentially takes each word in a sentence, extracts its information, and embeds it into a semantic vector. Due to its ability to capture long term memory, the LSTM-RNN accumulates increasingly richer information as it goes through the sentence, and when it reaches the last word, the hidden layer of the network provides a semantic representation of the whole sentence. In this paper, the LSTM-RNN is trained in a weakly supervised manner on user click-through data logged by a commercial web search engine. Visualization and analysis are performed to understand how the embedding process works. The model is found to automatically attenuate the unimportant words and detect the salient keywords in the sentence. Furthermore, these detected keywords are found to automatically activate different cells of the LSTM-RNN, where words belonging to a similar topic activate the same cell. As a semantic representation of the sentence, the embedding vector can be used in many different applications. These automatic keyword detection and topic allocation abilities enabled by the LSTM-RNN allow the network to perform document retrieval, a difficult language processing task, where the similarity between the query and documents can be measured by the distance between their corresponding sentence embedding vectors computed by the LSTM-RNN. On a web search task, the LSTM-RNN embedding is shown to significantly outperform several existing state of the art methods. We emphasize that the proposed model generates sentence embedding vectors that are specially useful for web document retrieval tasks. A comparison with a well known general sentence embedding method, the Paragraph Vector, is performed. The results show that the proposed method in this paper significantly outperforms Paragraph Vector method for web document retrieval task.
Hamid Palangi, Li Deng 0001, Yelong Shen, Jianfeng Gao 0001, Xiaodong He 0001, Jianshu Chen, Xinying Song, Rabab K. Ward
IEEE ACM Trans. Audio Speech Lang. Process.8
2015 Combining sparsity with rank-deficiency for energy efficient EEG sensing and transmission over Wireless Body Area Network
abstract
In Wireless Body Area Networks (WBAN) the energy consumption is dominated by sensing and communication. Previous techniques exploited the sparsity of the signal (in transform domains) to reduce communication costs for EEG transmission. For the first time, in this work, we propose to jointly exploit sparsity and rank-deficiency of the multi-channel signal ensemble in order to reduce both sensing and communication power consumptions. We test our method with state-of-the-art recovery techniques and find that the reconstruction accuracy from our method is considerably better and that too at lower energy consumption.
Angshul Majumdar, Ankita Shukla, Rabab K. Ward
ICASSP3
2015 Learning the sparsity basis in low-rank plus sparse model for dynamic MRI reconstruction
abstract
Modeling a temporal image sequence as a super-position of sparse and low-rank component stems from studies in principal component pursuit (PCP). Recently this technique was applied for dynamic MRI reconstruction with two modifications. First, unlike the original PCP, the problem was to recover the image sequence from under-sampled measurements. Second, the sparse component of the signal was not sparse in itself but in a transform domain. Recent studies in dynamic MRI reconstruction showed that, instead of using a fixed sparsity basis, better recovery results can be achieved when the sparsifying dictionary is adaptively learned from the data using Blind Compressed Sensing (BCS) framework. In this work, we demonstrate that learning the sparsity basis using BCS like techniques improve the recovery accuracy from PCP when applied to dynamic MRI reconstruction problems.
Angshul Majumdar, Rabab K. Ward
ICASSP2
2015 Angular upsampling of projection measurements in 3D computed tomography using a sparsity prior
abstract
We propose an algorithm for angular upsampling of the projections in 3D computed tomography (CT). The central assumption of the proposed method is that small blocks extracted from stacked projections have a sparse representation in an overcomplete dictionary. We present methods for fast solution of the optimization problems involved and apply the proposed algorithm on simulated and real projections. Our results show that upsampling of the projections with the proposed method can lead to a significant improvement in the quality of the reconstructed image.
Davood Karimi, Rabab K. Ward, Nancy L. Ford
ICIP2
2015 Learning space-time dictionaries for blind compressed sensing dynamic MRI reconstruction
abstract
This work addresses the problem of recovering dynamic MR sequences from their undersampled projections using the Blind Compressed Sensing (BCS) framework. In BCS, reconstructing the sparse signal and estimating the sparsifying basis proceeds simultaneously. Best results in CS based dynamic MRI reconstruction have been achieved when both spatial correlation and temporal redundancies were exploited. Prior studies in BCS dynamic MRI reconstruction either accounted for spatial redundancy or temporal correlation, but not both. In this work, we propose to jointly exploit spatio-temporal correlation within the BCS framework. The improvement in reconstruction results is significant compared to prior BCS techniques in dynamic MRI.
Angshul Majumdar, Rabab K. Ward
ICIP2
2015 Hidden Markov Support Vector Machines for Self-Paced Brain Computer Interfaces
abstract
Brain Computer Interfaces (BCI) aim at providing a means to control devices with brain signals. Self-paced BCIs, as opposed to synchronous ones, have the advantage of being operational at all times and not only at specific system-defined periods. Traditionally, in the BCI field, a sliding window over the brain signal is used to detect the intention of the user at a given time. This approach ignores the temporal correlations between the adjacent time windows. This paper proposes a novel approach to classify self-paced BCI data using structural support vector machines. Our proposed approach considers the history of the brain signals in the context of sequential supervised learning to better detect the intention of the user from his/her brain signals. We have compared our proposed model to the sliding window approach with Support Vector Machines (SVM) and Linear Discriminant Analysis (LDA) classifiers. Using data collected from 4 individuals form BCI competition IV, it is shown that the F1 score of our approach is significantly better than the sliding window approach. The average F1 score of our method across all subjects is 0.3 and 0.5 higher than the sliding window with SVM and LDA classifiers, respectively.
Hossein Bashashati, Rabab K. Ward, Ali Bashashati
ICMLA2
2015 Object-Based Multiple Foreground Video Co-Segmentation via Multi-State Selection Graph
abstract
We present a technique for multiple foreground video co-segmentation in a set of videos. This technique is based on category-independent object proposals. To identify the foreground objects in each frame, we examine the properties of the various regions that reflect the characteristics of foregrounds, considering the intra-video coherence of the foreground as well as the foreground consistency among the different videos in the set. Multiple foregrounds are handled via a multi-state selection graph in which a node representing a video frame can take multiple labels that correspond to different objects. In addition, our method incorporates an indicator matrix that for the first time allows accurate handling of cases with common foreground objects missing in some videos, thus preventing irrelevant regions from being misclassified as foreground objects. An iterative procedure is proposed to optimize our new objective function. As demonstrated through comprehensive experiments, this object-based multiple foreground video co-segmentation method compares well with related techniques that co-segment multiple foregrounds.
Huazhu Fu, Dong Xu 0001, Stephen Lin 0001, Rabab K. Ward
IEEE Trans. Image Process.5
2014 Learning sparse models for image quality assessment
abstract
Many successful image quality metrics rely on the structural information in an image to assess its perceptual quality. Extracting the structural information that is perceptually meaningful to our visual system, however, is a challenging task. This paper proposes a new quality assessment metric that relies on a sparse modeling approach to learn the inherent structures of the image. These structures are learnt as a set of basis vectors, such that any structure in the image can be represented by a linear combination of only a few of these basis vectors. This strategy is known to generate basis vectors that are qualitatively similar to the receptive field of the simple cells present in the mammalian primary visual cortex. The perceptual quality of the distorted image is estimated by comparing the structures of the reference and the distorted images in terms of the learnt basis vectors. Our approach is evaluated on five standard subject-rated image quality assessment datasets. The proposed metric exhibits high correlation with the subjective ratings outperforming several well established methods.
Tanaya Guha, Ehsan Nezhadarya, Rabab K. Ward
ICASSP3
2014 Improved MRI reconstruction via non-convex elastic net
abstract
This work proposes the use of an elastic-net to reconstruct Magnetic Resonance Images from their partially sampled K-space. The resulting elastic-net formulation of this problem is composed of two terms - the first term promotes sparsity and the other one promotes a grouping effect. The advantage of using an elastic-net for MRI reconstruction is that it can recover the hierarchically correlated sparse wavelet coefficients of the image. We develop two reconstruction methods via two elastic-net formulations - the synthesis prior and the analysis prior. We also impose non-convex sparsity penalties. There are no existing algorithms that solve such problems; hence we derive efficient algorithms for solving them. The experimental results show that our proposed analysis prior method outperforms state-of-the-art in MRI reconstruction.
Angshul Majumdar, Rabab K. Ward
ICASSP2
2014 A local fingerprinting approach for audio copy detection
Mani Malek 0001, Rabab K. Ward
Signal Process.2
2014 Sparse representation-based image quality assessment
Tanaya Guha, Ehsan Nezhadarya, Rabab K. Ward
Signal Process. Image Commun.3
2014 Performance Analysis of RFID Protocols: CDMA Versus the Standard EPC Gen-2
abstract
Radio frequency identification (RFID) is a ubiquitous wireless technology which allows objects to be identified automatically. An RFID tag is a small electronic device with an antenna and has a unique identification (ID) number. RFID tags can be categorized into passive and active tags. For passive tags, a standard communication protocol known as EPC-global Generation-2, or briefly EPC Gen-2, is currently in use. RFID systems are prone to transmission collisions due to the shared nature of the wireless channel used by tags. The EPC Gen-2 standard recommends using dynamic framed slotted ALOHA technique to solve the collision issue and to read the tag IDs successfully. Recently, some researchers have suggested to replace the dynamic framed slotted ALOHA technique used in the standard EPC Gen-2 protocol with the code division multiple access (CDMA) technique to reduce the number of collisions and to improve the tag identification procedure. In this paper, the standard EPC Gen-2 protocol and the CDMA-based tag identification schemes are modeled as absorbing Markov chain systems. Using the proposed Markov chain systems, the analytical formulae for the average number of queries and the total number of transmitted bits needed to identify all tags in an RFID system are derived for both the EPC Gen-2 protocol and the CDMA-based tag identification schemes. In the next step, the performance of the EPC Gen-2 protocol is compared with the CDMA-based tag identification schemes and it is shown that the standard EPC Gen-2 protocol outperforms the CDMA-based tag identification schemes in terms of the number of transmitted bits and the average time required to identify all tags in the system.
Ehsan Vahedi, Rabab K. Ward, Ian F. Blake
IEEE Trans Autom. Sci. Eng.2
2014 Image Similarity Using Sparse Representation and Compression Distance
abstract
A new line of research uses compression methods to measure the similarity between signals. Two signals are considered similar if one can be compressed significantly when the information of the other is known. The existing compression-based similarity methods, although successful in the discrete one dimensional domain, do not work well in the context of images. This paper proposes a sparse representation-based approach to encode the information content of an image using information from the other image, and uses the compactness (sparsity) of the representation as a measure of its compressibility (how much can the image be compressed) with respect to the other image. The sparser the representation of an image, the better it can be compressed and the more it is similar to the other image. The efficacy of the proposed measure is demonstrated through the high accuracies achieved in image clustering, retrieval and classification.
Tanaya Guha, Rabab K. Ward
IEEE Trans. Multim.2
2013 Compressed sensing and energy-aware independent component analysis for compression of EEG signals
abstract
In this paper, we propose the use of compressed sensing (CS) that is preceded by an energy-efficient, cross-product based independent component analysis (ICA) preprocessing method to efficiently compress electroencephalogram (EEG) signals in the context of a wireless body sensor network (WBSN). In WBSNs, the battery life puts a strict energy constraint at each sensor node. By providing a simple, nonadaptive compression scheme at the sensor nodes, CS offers an efficient solution to compress EEG signals in WBSNs. Through simulations, we demonstrate that our method requires less energy than other state-of-the-art methods using ICA, with a reduction in computations that can reach up to 94%. We also demonstrate that for a fixed compression ratio, the achievable reconstruction error is similar to the state-of-the-art method using ICA, and is much lower than when CS is used alone.
Simon Fauvel, Abhinav Agarwal, Rabab K. Ward
ICASSP3
2013 Image similarity measurement from sparse reconstruction errors
abstract
This paper presents a new approach to measuring the similarity between two images using sparse reconstruction. Our approach alleviates the difficulty of selecting and extracting suitable features from images which usually requires domain-specific knowledge. The proposed measure, the Sparse SNR (SSNR), does not use any prior knowledge about the data type or the application. SSNR is generic in the sense that it is applicable, without modification, to a variety of problems involving different types of images. Given a pair of images, a set of basis vectors (dictionary) is learnt for each image such that each image can be represented as a linear combination of a small number of its dictionary elements. Each image is reconstructed by two dictionaries - the one trained on the image itself and the second - trained on the other image. We develop a novel similarity measure based on the resulting reconstruction errors. To the best of our knowledge, this is the first attempt to develop a sparse reconstruction-based similarity measure. Excellent classification, clustering and retrieval results are achieved on benchmark datasets involving facial images and textures.
Tanaya Guha, Rabab K. Ward, Tyseer Aboulnasr
ICASSP2
2013 Exploiting sparsity and rank-deficiency in dynamic MRI reconstruction
abstract
This work addresses the problem of dynamic MRI reconstruction from partially sampled K-space. When the frames of the dynamic MRI sequences are stacked as columns of a matrix, the resultant matrix is both sparse (in a transform domain) and rank-deficient. The dynamic MRI sequence is reconstructed by solving an optimization problem that minimizes a sum of sparsity and rank-deficiency penalties subject to data constraints (K-space data acquisition model). In this work, we propose a non-convex optimization problem for dynamic MRI reconstruction where the sparsity penalty is an lp-norm and the rank-deficiency penalty is the Schatten-q norm (0p-norm and Schatten-q norm minimization problem; hence we derive a new algorithm based on the Majorization Minimization method. Our proposed method shows considerable improvement in reconstruction results over state-of-the-art techniques in dynamic MRI reconstruction.
Angshul Majumdar, Rabab K. Ward
ICASSP2
2013 Using deep stacking network to improve structured compressed sensing with Multiple Measurement Vectors
abstract
We study the MMV (Multiple Measurement Vectors) compressive sensing setting with a specific sparse structured support. The locations of the non-zero rows in the sparse matrix are not known. All that is known is that the locations of the non-zero rows have probabilities that vary from one group of rows to another. We propose two novel greedy algorithms for the exact recovery of the sparse matrix in this structured MMV compressive sensing problem. The first algorithm models the matrix sparse structure using a shallow non- linear neural network. The input of this network is the residual matrix after the prediction and the output is the sparse matrix to be recovered. The second algorithm improves the shallow neural network prediction by using the stacking operation to form a deep stacking network. Experimental evaluation demonstrates the superior performance of both new algorithms over existing MMV methods. Among all, the algorithm using the deep stacking network for modelling the structure in MMV compressive sensing performs the best.
Hamid Palangi, Rabab K. Ward, Li Deng 0001
ICASSP2
2013 Dynamic CT Reconstruction by Smoothed Rank Minimization
Angshul Majumdar, Rabab K. Ward
MICCAI (3)2
2013 Erratum: Dynamic CT Reconstruction by Smoothed Rank Minimization
Angshul Majumdar, Rabab K. Ward
MICCAI (3)2
2013 A Joint Multimodal Group Analysis Framework for Modeling Corticomuscular Activity
abstract
Corticomuscular coupling analysis based on multiple data sets such as electroencephalography (EEG) and electromyography (EMG) signals provides a useful tool for understanding human motor control systems. Two probably most popular methods are the pair-wise magnitude-squared coherence (MSC) between EEG and simultaneously-recorded EMG signals, and partial least square (PLS). Unfortunately, MSC and PLS generally deal with only two types of data sets at the same time, while we may need to analyze more than two types of data sets. Moreover, it is not straightforward to extend MSC to the group level for combining results across subjects. Also, PLS can have the information mixing problem since only the variations in one data set are used to predict the other data set. To address these concerns, we propose a joint multimodal analysis framework for corticomuscular coupling analysis. The proposed framework models multiple data spaces simultaneously in a multidirectional fashion. Furthermore, to address the inter-subject variability concern in real-world medical applications, we extend the proposed framework from the individual subject level to the group level to obtain common corticomuscular coupling patterns across subjects. We apply the proposed framework to concurrent EEG, EMG and behavior data collected in a Parkinson's disease (PD) study. The results reveal several highly correlated temporal patterns among the three types of signals and their corresponding spatial activation patterns. In PD subjects, there are enhanced connections between occipital region and other regions, which is consistent with the previous medical finding. The proposed framework is a promising technique for performing multi-subject and multi-modal data analysis.
Xun Chen 0001, Xiang Chen 0004, Rabab K. Ward, Z. Jane Wang 0001
IEEE Trans. Multim.3
2013 Visually Favorable Tone-Mapping With High Compression Performance in Bit-Depth Scalable Video Coding
abstract
In bit-depth scalable video coding, the tone-mapping scheme used to convert high-bit-depth to eight-bit videos is an essential yet very often ignored component. In this paper, we demonstrate that an appropriate choice of a tone-mapping operator can improve the coding efficiency of bit-depth scalable encoders. We present a new tone-mapping scheme that delivers superior compression efficiency while adhering to a predefined base layer perceptual quality. We develop numerical models that estimate the base layer bit-rate (Rb), the enhancement layer bitrate (Re), and the mismatch (QL) between the resulting low dynamic range (LDR) base-layer signal and the predefined base layer representation. Our proposed tone curve is given by the solution of an optimization problem which minimizes a weighted sum of Rb, Re, and QL. The problem formulation also considers the temporal effect of tone-mapping by adding a constraint to the optimization problem that suppresses flickering artifacts. We also propose a technique with which to tone-map a high-bit-depth video directly in a compression-friendly color space (e.g., one luma and two chroma channels) without converting to the RGB domain. Experimental results show that we can save up to 40% of the total bit-rate (or 3.5 dB PSNR improvement for the same bitrate), and, in general, about 20% bit-rate savings can be achieved.
Zicong Mai, Hassan Mansour, Panos Nasiopoulos, Rabab K. Ward
IEEE Trans. Multim.4
2012 Analytical modeling of RFID Generation-2 protocol using absorbing Markov chain theorem
abstract
Radio frequency identification (RFID) is a ubiquitous wireless technology which allows objects to be identified automatically. An RFID tag is a small electronic device with an antenna and has a unique identification number. RFID tags can be categorized into passive and active tags. For passive tags, there exists a standard communication protocol called EPC-global Generation-2, or briefly EPC Gen-2 [1]. In this paper, we investigate the EPC Gen-2 protocol and model it using an absorbing Markov chain. We formulate the proposed model and calculate the expected number of queries required to identify all tags in the system. Extensive simulations validate and confirm the accuracy of our proposed analytical model. Without this model, one has to run simulations and average the results to obtain the expected number of required queries for any given number of tags in the system. Using the mathematical formulations provided, there is no need to rely on simulations for studying the behavior of the EPC Gen-2 protocol and we are able to calculate the number of required queries directly. Our proposed analytical model is also useful in studying and comparing other RFID protocols, and in deploying better protocols for RFID systems.
Ehsan Vahedi, Rabab K. Ward, Ian F. Blake
GLOBECOM2
2012 A sparse reconstruction based algorithm for image and video classification
abstract
The success of sparse reconstruction based classification algorithms largely depends on the choice of overcomplete bases (dictionary). Existing methods either use the training samples as the dictionary elements or learn a dictionary by optimizing a cost function with an additional discriminating component. While the former method requires a good number of training samples per class and is not suitable for video signals, the later adds instability and more computational load. This paper presents a sparse reconstruction based classification algorithm that mitigates the above difficulties. We argue that learning class-specific dictionaries, one per class, is a natural approach to discrimination. We describe each training signal by an error vector consisting of the reconstruction errors the signal produces w.r.t each dictionary. This representation is robust to noise, occlusion and is also highly discriminative. The efficacy of the proposed method is demonstrated in terms of high accuracy for image-based Species and Face recognition and video-based Action recognition.
Tanaya Guha, Rabab K. Ward
ICASSP2
2012 Computationally efficient tone-mapping of high-bit-depth video in the YCbCr domain
abstract
High dynamic range (HDR) video content is able to provide superior picture quality. This is because the representation of HDR signals requires more bits than the 8-bit low dynamic range (LDR) video. Tone-mapping is the process that converts HDR to LDR signals. Most tone-mapping methods are derived only for the luminance component. This mapping function is then used in each of the R, G and B components to generate the LDR color image. This color tone mapping correction approach, however, cannot be directly applied to most videos since they are usually encoded in the YCbCr color space. This paper addresses this problem and proposes a tone-mapping method that is applied directly on the YCbCr signals. Experimental results show that the Cb and the Cr signals generated by our method are almost identical to those produced with the conventional pipeline up to round-off errors, with average PSNR at about 55 dB and average SSIM at 0.991. By avoiding all the round-off errors introduced in the conventional method, our approach provides a more accurate LDR picture. Moreover, the proposed solution has significantly lower complexity because it bypasses the processes such as color space transformation and up-sampling which are required by the conventional method.
Zicong Mai, Panos Nasiopoulos, Rabab K. Ward
ICASSP3
2012 Face recognition from video: An MMV recovery approach
abstract
In this paper we propose a new approach to video based face recognition. Our work is based on the Sparse Classification approach which assumes that each test sample can be formed by a linear combination of the training samples of the correct class. Based on this assumption, we formulate the classification problem as one of joint sparse recovery of Multiple Measurement Vectors (MMV). This requires solving an NP hard problem. This problem has not been solved earlier; thus we derive an algorithm for solving it. The experimental evaluation is carried on the VidTIMIT database. The proposed method is compared against an HMM based method for video based face recognition and the modified Sparse Classification method. The results show that the proposed method outperforms both these methods.
Angshul Majumdar, Rabab K. Ward
ICASSP2
2012 Synthesis and analysis prior algorithms for joint-sparse recovery
abstract
This paper proposes a Majorization-Minimization approach for solving the synthesis and analysis prior joint-sparse multiple measurement vector reconstruction problem. The proposed synthesis prior algorithm yielded the same results as the Spectral Projected Gradient (SPG) method. The analysis prior algorithm is the first to be proposed for this problem. It yielded considerably better results than the proposed synthesis prior algorithm. For problems of a given size, the run times for our proposed algorithms are fixed; unlike SPG where the reconstruction time also depends on the support size of the vectors.
Angshul Majumdar, Rabab K. Ward
ICASSP2
2012 A focuss based method for low rank matrix recovery
abstract
In this work, we address the problem of low-rank matrix recovery from its under-sampled projections. The recovery is formulated as a Schatten-p norm minimization problem. We proposed a novel algorithm to solve the Schatten-p norm minimization problem based on the FOCUSS (FOCally Under-determined System Solver) approach. We compared our proposed method with state-of-the-art solvers. Experimental evaluation was carried out on two problems - matrix completion and image inpainting. For matrix completion, our proposed method showed better recovery rate than other methods. In the image inpainting problem, our method yields 1.5 dB improvement over the nearest competing algorithm.
Angshul Majumdar, Rabab K. Ward, Tyseer Aboulnasr
ICIP2
2012 A P300-based BCI classification algorithm using median filtering and Bayesian feature extraction
abstract
A brain computer interface (BCI) system translates a person's brain activity into useful control or communication signals. In this paper, an effective P300-based BCI identification algorithm using median filtering and Bayesian classifier is proposed to improve the classification accuracy and computation efficiency of P300-based BCI. Median filtering is firstly applied to remove noises and Bayesian Linear Discriminant Analysis (BLDA) is then employed for classification. Testing on the P300 speller paradigm in dataset II of 2004 BCI Competition III, we show that a 90% average classification accuracy can be achieved and the highest accuracy is 100%. The proposed method is also computationally efficient and thus it represents a practical implementation for man-computer communication control, especially for on-line applications.
Xun Chen 0001, Rabab K. Ward
MMSP4
2012 A novel local audio fingerprinting algorithm
abstract
A local fingerprinting algorithm is proposed for the purpose of audio copy detection. The proposed algorithm is robust to noise as well as tempo and pitch modifications of the audio signal. The fingerprints are extracted from adaptively scaled patches of the time-chroma representation of the audio signal. The proposed time-chroma representation, converts tempo change and pitch shift attacks on an audio signal to scaling and circular shift attacks on images, respectively. The proposed algorithm is shown to outperform the state-of-the-art.
Mani Malek 0001, Rabab K. Ward
MMSP2
2012 Wavelet-based gradient transform and its applications
abstract
A wavelet-based image gradient transform is proposed. The proposed transform, called multi-scale gradient transform (MSGT), obtains the first order derivative of an image in terms of the wavelet detail coefficients. While traditional methods estimate the image gradients at each wavelet scale in terms of the horizontal and vertical wavelet coefficients only, the proposed transform obtains the gradients in terms of the diagonal, as well as the horizontal and vertical wavelet coefficients. The proposed MSGT is designed to be invertible, non-redundant and computationally efficient. We demonstrate the potential applications of the proposed transform in texture feature extraction, multi-scale edge detection, image quality assessment and image watermarking.
Ehsan Nezhadarya, Rabab K. Ward, Z. Jane Wang 0001
MMSP2
2012 A Fast Approximate Nearest Neighbor Search Algorithm in the Hamming Space
abstract
A fast approximate nearest neighbor search algorithm for the (binary) Hamming space is proposed. The proposed Error Weighted Hashing (EWH) algorithm is up to 20 times faster than the popular locality sensitive hashing (LSH) algorithm and works well even for large nearest neighbor distances where LSH fails. EWH significantly reduces the number of candidate nearest neighbors by weighing them based on the difference between their hash vectors. EWH can be used for multimedia retrieval and copy detection systems that are based on binary fingerprinting. On a fingerprint database with more than 1,000 videos, for a specific detection accuracy, we demonstrate that EWH is more than 10 times faster than LSH. For the same retrieval time, we show that EWH has a significantly better detection accuracy with a 15 times lower error rate.
Mani Malek 0001, Rabab K. Ward, Mehrdad Fatourechi
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 Learning Sparse Representations for Human Action Recognition
abstract
This paper explores the effectiveness of sparse representations obtained by learning a set of overcomplete basis (dictionary) in the context of action recognition in videos. Although this work concentrates on recognizing human movements-physical actions as well as facial expressions-the proposed approach is fairly general and can be used to address other classification problems. In order to model human actions, three overcomplete dictionary learning frameworks are investigated. An overcomplete dictionary is constructed using a set of spatio-temporal descriptors (extracted from the video sequences) in such a way that each descriptor is represented by some linear combination of a small number of dictionary elements. This leads to a more compact and richer representation of the video sequences compared to the existing methods that involve clustering and vector quantization. For each framework, a novel classification algorithm is proposed. Additionally, this work also presents the idea of a new local spatio-temporal feature that is distinctive, scale invariant, and fast to compute. The proposed approach repeatedly achieves state-of-the-art results on several public data sets containing various physical actions and facial expressions.
Tanaya Guha, Rabab K. Ward
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 On the choice of Compressed Sensing priors and sparsifying transforms for MR image reconstruction: An experimental study
Angshul Majumdar, Rabab K. Ward
Signal Process. Image Commun.2
2012 Compressed Sensing Based Real-Time Dynamic MRI Reconstruction
abstract
This work addresses the problem of real-time online reconstruction of dynamic magnetic resonance imaging sequences. The proposed method reconstructs the difference between the previous and the current image frames. This difference image is sparse. We recover the sparse difference image from its partial k-space scans by using a nonconvex compressed sensing algorithm. As there was no previous fast enough algorithm for real-time reconstruction, we derive a novel algorithm for this purpose. Our proposed method has been compared against state-of-the-art offline and online reconstruction methods. The accuracy of the proposed method is less than offline methods but noticeably higher than the online techniques. For real-time reconstruction we are also concerned about the reconstruction speed. Our method is capable of reconstructing 128 × 128 images at the rate of 6 frames/s, 180 × 180 images at the rate of 5 frames/s and 256 × 256 images at the rate of 2.5 frames/s.
Angshul Majumdar, Rabab K. Ward, Tyseer Aboulnasr
IEEE Trans. Medical Imaging2
2011 Action recognition by learnt class-specific overcomplete dictionaries
abstract
This paper presents a sparse signal representation based approach to address the problem of human action recognition in videos. For each action, a set of redundant basis (dictionary) is learnt by solving a sparse optimization problem. A dictionary is learnt using the image patches of its corresponding action, such that every patch vector is represented by some linear combination of a small number of basis vectors. By learning one dictionary per action, it is expected that each dictionary can efficiently represent one particular action. We show that such class-specific dictionaries - each representative of one action - provide a powerful means of action classification. Given a query sequence, the classifier seeks the dictionary that best approximates the query class. We have evaluated the proposed approach on the standard datasets. Experimental results demonstrate high accuracy and robustness against occlusion or viewpoint changes.
Tanaya Guha, Rabab K. Ward
FG2
2011 A new data hiding method using angle quantization index modulation in gradient domain
abstract
A robust data hiding scheme that embeds the watermark bits by quantizing the gradient directions of an image is proposed. By embedding the watermark in the angle, the watermark becomes robust to amplitude scaling attacks. To keep the watermark imperceptible and enhance its robustness, it is embedded in the significant gradient vectors of the image. The significant gradient vectors are obtained using discrete wavelet transform (DWT). Thus, the gradient vector at a pixel is first obtained in terms of the DWT coefficients. Then the gradient direction is quantized by modifying the DWT coefficients corresponding to the gradient vectors. Experimental results confirm that the proposed gradient direction watermarking (GDWM) method is robust to various types of attacks, specially amplitude scaling at tacks, and results in watermarked images of high fidelity.
Ehsan Nezhadarya, Z. Jane Wang 0001, Rabab K. Ward
ICASSP3
2011 Effect of brightness on the quality of visual 3D perception
abstract
Over the years, a consensus has been reached that the introduction of 3D entertainment can only be a lasting success if the perceived image quality and the viewing comfort are better than those of conventional 2D television. There are different factors that affect the perceived quality of 3D content. In this paper, our objective is to obtain a good understanding of the effect that brightness has on the visual quality of 3D videos and compare it to that of the 2D. We capture outdoor and indoor scenes with different exposures and we perform subjective evaluation to investigate how brightness affects the perceived quality of the 3D experience.
Mahsa T. Pourazad, Zicong Mai, Panos Nasiopoulos, Konstantinos N. Plataniotis, Rabab K. Ward
ICIP5
2011 Some empirical advances in matrix completion
Angshul Majumdar, Rabab K. Ward
Signal Process.2
2011 Probabilistic Analysis and Correction of Chen's Tag Estimate Method
abstract
Radio frequency identification (RFID) is a ubiquitous wireless technology which allows objects to be identified automatically. An RFID tag is a small electronic device with an antenna and has a unique serial number. For some RFID applications and in the ALOHA-based anticollision algorithms, the number of tags in the system needs to be estimated. In Trans. Autom. Sci. Eng., vol 6, no. 1, pp. 9-15, Jan. 2009, Chen, a probabilistic method for tag estimation in ALOHA-based RFID systems was proposed, based on the maximum a posteriori probability. Although this approach is novel and useful, it has a mathematical error in modeling the problem. In this short paper, we address this problem and provide the correct probabilistic model for the ALOHA-based RFID systems. Some consequences of correcting the error in Trans. Autom. Sci. Eng., vol 6, no. 1, pp. 9-15, Jan. 2009, Chen, are discussed and the model is validated via simulation. Using the correct model, the performance of the ALOHA-based anticollision algorithm can be improved.
Ehsan Vahedi, Vincent W. S. Wong 0001, Ian F. Blake, Rabab K. Ward
IEEE Trans Autom. Sci. Eng.4
2011 A Robust and Fast Video Copy Detection System Using Content-Based Fingerprinting
abstract
A video copy detection system that is based on content fingerprinting and can be used for video indexing and copyright applications is proposed. The system relies on a fingerprint extraction algorithm followed by a fast approximate search algorithm. The fingerprint extraction algorithm extracts compact content-based signatures from special images constructed from the video. Each such image represents a short segment of the video and contains temporal as well as spatial information about the video segment. These images are denoted by temporally informative representative images. To find whether a query video (or a part of it) is copied from a video in a video database, the fingerprints of all the videos in the database are extracted and stored in advance. The search algorithm searches the stored fingerprints to find close enough matches for the fingerprints of the query video. The proposed fast approximate search algorithm facilitates the online application of the system to a large video database of tens of millions of fingerprints, so that a match (if it exists) is found in a few seconds. The proposed system is tested on a database of 200 videos in the presence of different types of distortions such as noise, changes in brightness/contrast, frame loss, shift, rotation, and time shift. It yields a high average true positive rate of 97.6% and a low average false positive rate of 1.0%. These results emphasize the robustness and discrimination properties of the proposed copy detection system. As security of a fingerprinting system is important for certain applications such as copyright protections, a secure version of the system is also presented.
Mani Malek 0001, Mehrdad Fatourechi, Rabab K. Ward
IEEE Trans. Inf. Forensics Secur.3
2011 Robust Image Watermarking Based on Multiscale Gradient Direction Quantization
abstract
We propose a robust quantization-based image watermarking scheme, called the gradient direction watermarking (GDWM), based on the uniform quantization of the direction of gradient vectors. In GDWM, the watermark bits are embedded by quantizing the angles of significant gradient vectors at multiple wavelet scales. The proposed scheme has the following advantages: 1) increased invisibility of the embedded watermark because the watermark is embedded in significant gradient vectors, 2) robustness to amplitude scaling attacks because the watermark is embedded in the angles of the gradient vectors, and 3) increased watermarking capacity as the scheme uses multiple-scale embedding. The gradient vector at a pixel is expressed in terms of the discrete wavelet transform (DWT) coefficients. To quantize the gradient direction, the DWT coefficients are modified based on the derived relationship between the changes in the coefficients and the change in the gradient direction. Experimental results show that the proposed GDWM outperforms other watermarking methods and is robust to a wide range of attacks, e.g., Gaussian filtering, amplitude scaling, median filtering, sharpening, JPEG compression, Gaussian noise, salt & pepper noise, and scaling.
Ehsan Nezhadarya, Z. Jane Wang 0001, Rabab K. Ward
IEEE Trans. Inf. Forensics Secur.3
2011 Probabilistic Analysis of Blocking Attack in RFID Systems
abstract
Radio-frequency identification (RFID) is a ubiquitous wireless technology which allows objects to be identified automatically. An RFID tag is a small electronic device with an antenna and has a unique serial number. Using RFID tags can simplify many applications and provide many benefits. Meanwhile, the privacy of the customers should be taken into account. A potential threat for the privacy of a user is that of anonymous readers obtaining information about the tags in the system. The use of a blocker tag has been proposed as a solution to avoid unwanted tag interrogations. A blocker tag can simulate all or a portion of tag IDs in the system. This prevents the malicious readers from identifying the tags and obtaining information from the system. Although this solution is simple to implement and has a low cost, it may add another threat to the RFID system if used as a malicious tool to attack the system. A malicious blocker tag can deteriorate the performance of an RFID system by simulating fake tag IDs. In this paper, we study the use of blocker tags for malicious attacks that can prevent nearby legitimate readers from correctly receiving the reply messages from the tags. The blocker attack is a medium access control (MAC)-layer denial of service (DoS) threat and we propose a lower-layer solution for this attack. We mathematically model the blocker attack for RFID systems which operate based on the binary tree walking or ALOHA singulation techniques. Using the developed analytical framework, we propose a probabilistic blocker tag detection (P-BTD) algorithm to detect the presence of an attacker in the RFID system. The P-BTD algorithm can detect the existence of a blocker tag using the information extracted from the interrogations performed by the reader. Simulation results show that our proposed algorithm has a better performance than the threshold-based detection algorithm in terms of the number of required interrogations.
Ehsan Vahedi, Vahid Shah-Mansouri, Vincent W. S. Wong 0001, Ian F. Blake, Rabab K. Ward
IEEE Trans. Inf. Forensics Secur.5
2011 Optimizing a Tone Curve for Backward-Compatible High Dynamic Range Image and Video Compression
abstract
For backward compatible high dynamic range (HDR) video compression, the HDR sequence is reconstructed by inverse tone-mapping a compressed low dynamic range (LDR) version of the original HDR content. In this paper, we show that the appropriate choice of a tone-mapping operator (TMO) can significantly improve the reconstructed HDR quality. We develop a statistical model that approximates the distortion resulting from the combined processes of tone-mapping and compression. Using this model, we formulate a numerical optimization problem to find the tone-curve that minimizes the expected mean square error (MSE) in the reconstructed HDR sequence. We also develop a simplified model that reduces the computational complexity of the optimization problem to a closed-form solution. Performance evaluations show that the proposed methods provide superior performance in terms of HDR MSE and SSIM compared to existing tone-mapping schemes. It is also shown that the LDR image quality resulting from the proposed methods matches that produced by perceptually-based TMOs.
Zicong Mai, Hassan Mansour, Rafal Mantiuk, Panos Nasiopoulos, Rabab K. Ward, Wolfgang Heidrich
IEEE Trans. Image Process.5
2011 A New Scheme for Robust Gradient Vector Estimation in Color Images
abstract
Gradient estimators are mostly designed to yield accurate and robust estimates of the gradient magnitude, not the gradient direction. This paper proposes a method for the accurate and robust estimation of both the gradient magnitude and direction. It robustly estimates the gradient in the x- and y-directions. The robustness against noise is achieved by prefiltering and postfiltering of the gradient in each direction. To reduce edge blurring effects introduced by these filters, the gradient in a certain direction is obtained by applying the prefilter and postfilter in the perpendicular direction. The basic elements employed in each window are: highpass, lowpass and aggregation operators. The highpass operator is used as a gradient estimator, the lowpass operator is for prefiltering and postfiltering, and the aggregation operator is for aggregating the prefiltered and postfiltered gradients. Four different combinations of highpass, lowpass and aggregation operators are proposed: MVD-Median-Mean, MVD-Median-Max, RCMG-Median-Mean, and RCMG-Median-Max. Experimental results show that the RCMG-Median-Mean has the best performance in estimating the gradient and detecting the edges in noisy color images. It is computationally more efficient than the state-of-the-art gradient estimators and is able to accurately estimate the gradient direction as well as the gradient magnitude. Computer simulation results show that the proposed method outperforms other recently proposed color gradient estimators and edge detectors.
Ehsan Nezhadarya, Rabab K. Ward
IEEE Trans. Image Process.2
2010 A Matrix Completion Approach to Reduce Energy Consumption in Wireless Sensor Networks
abstract
The main challenge faced by wireless sensor networks today is the problem of power consumption at the sensor nodes. Over time, researchers have developed different strategies to address this issue. Such strategies are strongly model dependent and/or application specific. In this work, we take a fresh look at the problem of power consumption in wireless sensor networks from a signal processing perspective. The main idea is simple. Sample only a subset of all the sensor nodes at a given instant and transmit them (this reduces both sampling and communication cost for all the nodes combined). At the central unit (sink) use smart mathematical tools (matrix completion algorithms) to estimate the data for the entire network. We have showed that, if about 1% reconstruction error is allowed, only 20% of the sensors need to sample and transmit at a given instant. This means on an average the life of the network is increased 5-fold. If more error reconstruction error is allowed, even lesser number of sensors need to be active at a given instant leading to more prolonged life of the network.
Angshul Majumdar, Rabab K. Ward
DCC2
2010 Differential Radon Transform for gait recognition
abstract
Experimental studies have proved that high frequency components have the maximum contribution in silhouette-based gait recognition. The Radon Transform (RT), used in gait analysis for its ability to compute useful directional projections, fails to capture the necessary high frequency content of images. In this paper we present the Differential Radon Transform (DiffRT) - a novel adaptation of the standard RT designed to extract such high frequency information efficiently. The proposed transform is used to extract a set of features from gait silhouettes. We provide both theoretical and experimental evidence that DiffRT can indeed collect the important image information to facilitate gait-based human recognition. Averaged silhouettes from USF database are used for performance evaluation following the gait challenge framework. Our proposed method achieves high recognition accuracy and outperforms several state-of-the-art algorithms.
Tanaya Guha, Rabab K. Ward
ICASSP2
2010 Non-convex group sparsity: Application to color imaging
abstract
This work investigates a group-sparse solution to the under-determined system of linear equations b=Ax where the unknown x is formed of a group of vectors xi's. A group-sparse solution has only a few xivectors as non-zeroes while the rest are zeroes. To seek a group-sparse solution generally a convex optimization problem is solved. Such an optimization criterion is unsuitable when the system is highly under-determined or when some of the vector xi's are themselves sparse. For such cases, we propose an alternate non-convex optimization problem. Simulation results show that the proposed method yields significantly improved results (2 orders of magnitude) over the standard method. We also apply the proposed group-sparse optimization in a novel fashion to the problem of color imaging. The new method shows an improvement of more than 1dB over the standard method.
Angshul Majumdar, Rabab K. Ward
ICASSP2
2010 Color image desaturation using sparse reconstruction
abstract
In this paper, we propose an algorithm to estimate the true values of saturated pixels in color images. Pixel saturation occurs when at least one color channel is clipped at some value below the full dynamic range of the scene, resulting in a loss in image fidelity. The proposed algorithm is based on the assumptions that images are nearly sparse in an appropriate transform domain, and that saturated pixels can be inferred from the structure of non-saturated neighboring pixels. Consequently, we use a hierarchical windowing algorithm which selects image regions containing relatively few saturated pixels for processing. Starting with small sized regions, and progressively increasing the size, we solve a sparsity promoting constrained ℓ1minimization problem for each selected region to recover the saturated pixels. Moreover, we provide simulation results to show the effectiveness of our algorithm.
Hassan Mansour, Rayan Saab, Panos Nasiopoulos, Rabab K. Ward
ICASSP4
2010 A robust morphological gradient estimator and edge detector for color images
abstract
A new vector-wise scheme for the gradient estimation and edge detection in noisy color images is proposed. In color images, different types of noise may corrupt the image. To reduce the effects of noise in the gradient estimation, we introduce the RCMG-Median-Mean estimator. RCMG-Median-Mean is a combination of the robust color morphological gradient (RCMG), the median and the mean filters to accurately estimate both the gradient magnitude and the gradient direction at each pixel of a noisy color image. The simulation results show that the proposed method more accurately estimates the true gradient vector, has better corner detection and leads to better continuous edges with less computational complexity than the RCMG method.
Ehsan Nezhadarya, Rabab K. Ward
ICASSP2
2010 Visually-favorable tone-mapping with high compression performance
abstract
We develop a tone-mapping operator (TMO) that considers the perceptual quality of the tone-mapped image together with the compression efficiency. The proposed TMO is formulated as an optimization problem that incorporates statistical models of i) the quality of the tone-mapped image given a desired TMO, ii) the base layer bit-rate and iii) the enhancement layer bit-rate. The results show that our method achieves high coding gain while maintaining good quality tone-mapped images.
Zicong Mai, Hassan Mansour, Panos Nasiopoulos, Rabab K. Ward
ICIP4
2010 Compressive color imaging with group-sparsity on analysis prior
abstract
Compressed sensing (CS) of color images can be formulated as a group-sparsity promoting inverse problem. In the past, group-sparsity constraint was imposed on the CS synthesis prior formulation with an orthogonal transform to solve the inverse problem. The objective of this work is to empirically show that better results can be obtained if a group-sparsity constraint is imposed on the CS analysis prior formulation with a redundant transform. This problem requires solving a group-sparsity promoting inverse problem which has not been addressed earlier. Therefore we derive a new algorithm for solving it based on the Majorization-Minimization approach. Experimental results corroborate that analysis prior with a redundant transform gives far superior (about 1.5dB) improvement compared to synthesis prior with orthogonal transform.
Angshul Majumdar, Rabab K. Ward
ICIP2
2010 Watermark survival chance (WSC) concept for improving watermark robustness against JPEG compression
abstract
This paper presents the new concept of watermark survival chance (WSC) for improving watermark robustness. WSC provides a robustness measure for an image feature (e.g. a discrete wavelet transform (DWT) coefficient) when used for watermark embedding, and thus can provide the watermark designer with prior knowledge on robust image features. As an illustrative example, we study additive spread spectrum watermarking in the DWT domain and consider JPEG compression as the attack. WSC is obtained for each DWT coefficient/subband for different compression ratios. Based on the WSC table for JPEG compression distortion, we suggest that: Wavelet coefficients can be divided into two main categories: block boundary coefficients and block non-boundary coefficients; block boundary coefficients generally are more robust for watermark embedding than block non-boundary coefficients; larger scale wavelet coefficients are generally more robust than smaller scales; a vertical subband is slightly preferred at small and large scales, while a horizontal subband is preferred at a medium scale.
Ehsan Nezhadarya, Z. Jane Wang 0001, Rabab K. Ward
ICIP3
2010 HDR image construction from multi-exposed stereo LDR images
abstract
In this paper, we present an algorithm that generates high dynamic range (HDR) images from multi-exposed low dynamic range (LDR) stereo images. The vast majority of cameras in the market only capture a limited dynamic range of a scene. Our algorithm first computes the disparity map between the stereo images. The disparity map is used to compute the camera response function which in turn results in the scene radiance maps. A refinement step for the disparity map is then applied to eliminate edge artifacts in the final HDR image. Existing methods generate HDR images of good quality for still or slow motion scenes, but give defects when the motion is fast. Our algorithm can deal with images taken during fast motion scenes and tolerate saturation and radiometric changes better than other stereo matching algorithms.
Hassan Mansour, Rabab K. Ward
ICIP3
2010 A Simple Approach to Find the Best Wavelet Basis in Classification Problems
abstract
In this paper, we address the problem of finding the best wavelet basis in wavelet packet analysis for applications based on classification. We implement and evaluate our proposed method in the design of a self-paced 2-state mental task-based brain-computer interface (BCI) as one possible type of classification-based applications. The autoregressive coefficients of the best wavelet basis are concatenated to form the feature vector. The 2-stage classification process is based on quadratic discriminant analysis and majority voting. Seventeen wavelets from 2 different families are tested. A 5×5 cross-validation process is per-formed twice to do model selection and system performance evaluation. The results show that the proposed method can be well applied to BCI systems.
Farhad Faradji, Rabab K. Ward, Gary E. Birch
ICPR2
2010 Correcting unsynchronized zoom in 3D video
abstract
When capturing 3D video with a stereoscopic camera setup, it is important for the cameras to be precisely aligned and synchronized. This is particularly difficult in transitions such as zooming where the camera parameters must be changed in unison, or else the perceived 3D effect will be degraded. In this paper we study the problem of unsynchronized zooming in 3D video. First, we present a subjective study that shows that the perceived quality of stereo video is greatly reduced if the two views are zoomed by different amounts. Next, we present a method for correcting zoom mismatch by applying cropping and scaling to ones of the views. Our method involves finding matching points between the left and right views, and performing least-squares regressions to estimate the amount of scaling and cropping required to make the views consistent. Experiments were performed on videos with digitally introduced zoom mismatch and videos with optical unsynchronized zoom. In both cases the results show that our method is highly accurate and produces videos without size differences or vertical parallax between the two views.
Colin Doutre, Mahsa T. Pourazad, Alexis M. Tourapis, Panos Nasiopoulos, Rabab K. Ward
ISCAS5
2010 On-the-fly tone mapping for backward-compatible high dynamic range image/video compression
abstract
In this paper, we propose a real-time tone-mapping scheme for backward compatible high dynamic range (HDR) video compression. The appropriate choice of a tone-mapping operator (TMO) can significantly improve the HDR quality reconstructed from a low dynamic range (LDR) version. We develop a statistical model that approximates the mean square error (MSE) distortion resulting from the combined processes of tone-mapping and compression. Using this model, we formulate a numerical optimization problem to find the tone-curve that minimizes the expected MSE in the reconstructed HDR sequence. We then simplify the developed model in order to reduce the computational complexity of the optimization problem to a closed-form solution. Performance evaluations show that the proposed methods provide superior performance in terms of HDR MSE and SSIM compared to existing tone-mapping schemes. It is also shown that the LDR image quality resulting from the proposed methods matches that produced by perceptually-based TMOs.
Zicong Mai, Hassan Mansour, Rafal Mantiuk, Panos Nasiopoulos, Rabab K. Ward, Wolfgang Heidrich
ISCAS5
2010 Fast block-size partitioning using empirical rate-distortion models for MPEG-2 to H.264/AVC transcoding
abstract
We present an efficient H.264/AVC block-size partitioning prediction method, which is based on our proposed empirical rate and distortion models. Compared to other state-of-the-art transcoding methods, and for the same rate-distortion performance, our proposed algorithm requires the least computational complexity, reaching a 73% reduction in variable block-size motion estimation for SDTV sequences, and 71% reduction for CIF sequences.
Qiang Tang 0002, Panos Nasiopoulos, Rabab K. Ward
ISCAS3
2010 Under-Determined Non-cartesian MR Reconstruction with Non-convex Sparsity Promoting Analysis Prior
Angshul Majumdar, Rabab K. Ward
MICCAI (3)2
2010 Improved Group Sparse Classifier
Angshul Majumdar, Rabab K. Ward
Pattern Recognit. Lett.2
2010 Compressed sensing of color images
Angshul Majumdar, Rabab K. Ward
Signal Process.2
2010 Energy Optimization for Many-Core Platforms: Communication and PVT Aware Voltage-Island Formation and Voltage Selection Algorithm
abstract
In this paper, we propose a novel approach to voltage-island formation, for the energy optimization of many-core architectures, which mitigates the impact of process, voltage, and temperature (PVT) variations. The islands are created by balancing their shape constraints imposed by intra and inter-island communication with the desire to limit the spatial extent of each island to minimize PVT impact. In addition, to reduce the number of voltage levels in the design, we propose an efficient voltage selection approach that provides near optimal results, for a set of 33 examined cases, with more than a ten times speedup compared to the best-known previous methods. This run-time improvement is important, especially for large many-core platforms. Finally, we present an evaluation platform considering pre-fabrication and post-fabrication PVT scenarios where multiple applications with hundreds to thousands of tasks are mapped onto many-core platforms with hundreds to thousands of cores to evaluate the proposed techniques. Results show that the average energy savings for 33 test cases using the proposed methods are 37% compared to 16% obtained using previous methods.
Sohaib Majzoub, Res Saleh, Steve Wilton, Rabab K. Ward
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2010 Adaptive Region-Based Image Enhancement Method for Robust Face Recognition Under Variable Illumination Conditions
abstract
Variable illumination conditions, especially the side lighting effects in face images, form a main obstacle in face recognition systems. To deal with this problem, this paper presents a novel adaptive region-based image preprocessing scheme that enhances face images and facilitates the illumination invariant face recognition task. The proposed method first segments an image into different regions according to its different local illumination conditions, then both the contrast and the edges are enhanced regionally so as to alleviate the side lighting effect. Different from existing contrast enhancement methods, we apply the proposed adaptive region-based histogram equalization on the low-frequency coefficients to minimize illumination variations under different lighting conditions. Besides contrast enhancement, by observing that under poor illuminations the high-frequency features become more important in recognition, we propose enlarging the high-frequency coefficients to make face images more distinguishable. This procedure is called edge enhancement (EdgeE). The EdgeE is also region-based. Compared with existing image preprocessing methods, our method is shown to be more suitable for dealing with uneven illuminations in face images. Experimental results on the representative databases, the Yale B+Extended Yale B database and the Carnegie Mellon University-Pose, Illumination, and Expression database, show that the proposed method significantly improves the performance of face images with illumination variations. The proposed method does not require any modeling and model fitting steps and can be implemented easily. Moreover, it can be applied directly to any single image without using any lighting assumption, and any prior information on 3-D face geometry.
Shan Du 0001, Rabab K. Ward
IEEE Trans. Circuits Syst. Video Technol.2
2010 Segmentation and Classification of Polarimetric SAR Data Using Spectral Graph Partitioning
abstract
A new approach for segmentation and classification of polarimetric synthetic aperture radar (POLSAR) data is proposed based on spectral graph partitioning. Since automated analysis techniques are often challenged due to the noisy properties of POLSAR data, human experts are employed to aid in the interpretation of such data in an operational setting. Humans can improve the performance of segmentation and classification of POLSAR data, because their vision system can apply cognitive skills that are not easy to incorporate into an automated system. The motivation for this paper is to incorporate some of these human perceptual skills into the computer algorithms. A framework that has recently emerged in computer vision for solving grouping problems with perceptually plausible results-spectral graph partitioning-is customized for POLSAR data. Segmentation is performed using the contour information in a region-based setting with the aid of spatial proximity. This is followed by a classification step performed through graph partitioning based on similarities of the mean coherence matrices obtained for each segment. Using the proposed approach, the results achieved are superior to the Wishart classifier. Automated parameter selection procedures are under development. This framework also suggests a way to accommodate different representations of polarimetric data and combine them with other information sources (e.g., optical imagery and digital elevation models).
Kaan Ersahin, Ian G. Cumming, Rabab K. Ward
IEEE Trans. Geosci. Remote. Sens.3
2010 Robust Classifiers for Data Reduced via Random Projections
abstract
The computational cost for most classification algorithms is dependent on the dimensionality of the input samples. As the dimensionality could be high in many cases, particularly those associated with image classification, reducing the dimensionality of the data becomes a necessity. The traditional dimensionality reduction methods are data dependent, which poses certain practical problems. Random projection (RP) is an alternative dimensionality reduction method that is data independent and bypasses these problems. The nearest neighbor classifier has been used with the RP method in classification problems. To obtain higher recognition accuracy, this study looks at the robustness of RP dimensionality reduction for several recently proposed classifiers--sparse classifier (SC), group SC (along with their fast versions), and the nearest subspace classifier. Theoretical proofs are offered regarding the robustness of these classifiers to RP. The theoretical results are confirmed by experimental evaluations.
Angshul Majumdar, Rabab K. Ward
IEEE Trans. Syst. Man Cybern. Part B2
2009 Genetic programming based image segmentation with applications to biomedical object detection
abstract
Image segmentation is an essential process in many image analysis applications and is mainly used for automatic object recognition purposes. In this paper, we define a new genetic programming based image segmentation algorithm (GPIS). It uses a primitive image-operator based approach to produce linear sequences of MATLAB® code for image segmentation. We describe the evolutionary architecture of the approach and present results obtained after testing the algorithm on a biomedical image database for cell segmentation. We also compare our results with another EC-based image segmentation tool called GENIE Pro. We found the results obtained using GPIS were more accurate as compared to GENIE Pro. In addition, our approach is simpler to apply and evolved programs are available to anyone with access to MATLAB®.
Tarundeep Singh, Nawwaf Kharma, Mohmmad Daoud, Rabab K. Ward
GECCO4
2009 Component-wise pose normalization for pose-invariant face recognition
abstract
The pose variation involved in facial images significantly degrades the performance of face recognition systems. In this paper, a component-wise pose normalization method for facilitating pose-invariant face recognition is proposed. The main idea is to normalize a non-frontal facial image to a virtual frontal image component by component. In this method, we first partition the whole non-frontal facial image into different facial components and then the virtual frontal view for each component is estimated separately. The final virtual frontal image is generated by integrating the virtual frontal components. The proposed method relies only on 2D images, therefore complex 3D modeling is not needed. The experimental results using the CMU-PIE database demonstrate the advantages of the proposed method.
Shan Du 0001, Rabab K. Ward
ICASSP2
2009 A custom-designed mental task-based brain-computer interface
abstract
At present, brain-computer interfaces cannot be used in real-life applications mainly because of their high false activation rates. To achieve a zero false positive rate, a mental task-based brain-computer interface custom designed for each subject and each task is proposed. The most discriminatory mental task is determined for each subject. We used the EEG signals of four subjects recorded while they were performing five different mental tasks. Autoregressive modeling and stationary wavelet transform are used in the process of feature extraction. Classification is based on quadratic discriminant analysis. For the most discriminatory mental task of each subject, we achieved a false positive rate of zero value while the true positive rate obtained was above 60%.
Farhad Faradji, Rabab K. Ward, Gary E. Birch
ICASSP2
2009 Classification via group sparsity promoting regularization
abstract
Recently a new classification assumption was proposed in [1]. It assumed that the training samples of a particular class approximately form a linear basis for any test sample belonging to that class. The classification algorithm in [1] was based on the idea that all the correlated training samples belonging to the correct class are used to represent the test sample. The Lasso regularization was proposed to select the representative training samples from the entire training set (consisting of all the training samples). Lasso however tends to select a single sample from a group of correlated training samples and thus does not promote the representation of the test sample in terms of all the training samples from the correct group. To overcome this problem, we propose two alternate regularization methods, elastic net and sum-over-l2-norm. Both these regularization methods favor the selection of multiple correlated training samples to represent the test sample. Experimental results on benchmark datasets show that our regularization methods give better recognition results compared to [1].
Angshul Majumdar, Rabab K. Ward
ICASSP2
2009 Artifact removal in EEG using Morphological Component Analysis
abstract
To reduce the effects of artifacts in electroencephalography (EEG), we propose the use of morphological component analysis (MCA). Taking advantage of the sparse representation of data in overcomplete dictionaries, MCA decomposes EEG signals into parts that have different morphological characteristics. For denoising purpose, the parts related to artifacts are removed. An over complete dictionary is constructed using the discrete cosine transform, Daubechies wavelet basis, and Dirac basis. Movement-related potentials (MRP) and EEG signals contaminated by spikes, eye-blinks, and muscle artifacts caused by eye-brow raising are used to evaluate the performance of the method. The results demonstrate that MCA can be used to decompose the single-channel EEG signals into artifacts and MRP components. The correlation coefficient between the denoised MRP and the original MRP using MCA is significantly higher than that obtained using stationary wavelet transform.
Xinyi Yong, Rabab K. Ward, Gary E. Birch
ICASSP2
2009 An efficient method for robust gradient estimation of RGB color images
abstract
A new vector-wise scheme for the gradient estimation in noisy color images is proposed. In the color images, different types of noise may corrupt the image. To reduce the effect of noise in the gradient estimation, the distance weighted order statistic filter (DWOSF) is introduced, to be embedded in the gradient estimator. DWOSF is an order statistic filter in which the color outliers in a neighborhood are adaptively assigned smaller weights than the other pixels in that neighborhood. The simulation results show that the proposed method accurately estimates the true gradient vector in noisy RGB color images.
Ehsan Nezhadarya, Rabab K. Ward
ICIP2
2009 Image quality monitoring using spread spectrum watermarking
abstract
An improved blind image quality assessment scheme that is based on the Watson's just noticeable difference (JND) modulated spread spectrum watermarking, is proposed. For the purpose of quality monitoring, a watermark is embedded into the original image. The image quality is estimated based on the detected watermark at the receiver side. In terms of the peak signal-to-noise ratio (PSNR), the proposed method is shown to be more robust and less perceptible than simple spread spectrum watermarking. This is due to several factors: an optimum detector is used for watermark detection, the watermark is appropriately selected from a set and the detector parameter is adjusted accordingly to closely yield an empirical ideal quality curve. The proposed method was tested by finding the quality estimates of different images compressed with different quality factors. The results indicate that the method can accurately estimate the quality of the received images based on the detected watermark power.
Ehsan Nezhadarya, Z. Jane Wang 0001, Rabab K. Ward
ICIP3
2009 An efficient low random-access delay panorama-based multiview video coding scheme
abstract
We present an efficient low random delay scheme for multiview video coding (MVC). In the proposed scheme, inter-view prediction (disparity estimation), which introduces time-consuming computations and random access delay to MVC, is replaced with a residue-stream coding process. Our algorithm transforms the middle view to a panoramic view of the scene. Then the residue streams are created as the difference of the luma and chroma values of overlapping regions of each view and the panoramic view. Finally the panoramic stream and all residue streams are encoded separately (simulcast coding). The hierarchical B picture prediction structure is implemented for coding each stream. Performance evaluations show that our proposed coding method outperforms the recent multiview video coding standard by up to 2.13 dB PSNR and enhances the compression ratio by 24.6%, while reducing random-access delay by 50%.
Mahsa T. Pourazad, Panos Nasiopoulos, Rabab K. Ward
ICIP3
2009 Efficient motion vector re-estimation for MPEG-2 TO H.264/AVC transcoding with arbitrary down-sizing ratios
abstract
As for down-sizing MPEG-2 to H.264/AVC transcoding, an efficient algorithm of estimating initial H.264/AVC motion vectors is proposed. By using the estimated initial motion vectors, only a small range of motion vector refinement is sufficient to find the final motion vector for each partition. Experimental results show that our proposed algorithm achieves average 0.08 dB improvement (maximum 0.24 dB) in the picture quality compared to the other state-of-art method. At the same time, the computational complexity of estimating the initial motion vectors is less than that of the other state-of-art technique (average 36% reduction).
Qiang Tang 0002, Panos Nasiopoulos, Rabab K. Ward
ICIP3
2009 Video Copy Detection Using Temporally Informative Representative Images
abstract
Content-based video hashing was introduced recently to serve the purpose of video copy detection. A conventional approach to video hashing is to apply image hashing techniques to either every frame or to the selected key frames of a video sequence. Both approaches ignore the temporal information contained in a video sequence. This study proposes an approach for generating representative images of a video sequence that carry the temporal as well as the spatial information. These images are denoted as TIRIs, Temporally Informative Representative Images. Performance of the proposed approach is demonstrated by applying a simple image hashing technique on TIRIs of a video database. It is shown that the resulted video hashing algorithm is highly robust to noise, frame dropping, changes in brightness and contrast, as well as a range of geometric attacks. An average true positive rate of 99.2% and false positive rate of 0.4% of the proposed approach demonstrate the robustness and uniqueness of the generated hashes. It is demonstrated that the proposed approach is easy to implement and computationally more efficient than another state-of-the-art video hashing method.
Mani Malek 0001, Mehrdad Fatourechi, Rabab K. Ward
ICMLA3
2009 Converting H.264-Derived Motion Information into Depth Map
Mahsa T. Pourazad, Panos Nasiopoulos, Rabab K. Ward
MMM3
2009 Improved Face Representation by Nonuniform Multilevel Selection of Gabor Convolution Features
abstract
Gabor wavelets are widely employed in face representation to decompose face images into their spatial-frequency domains. The Gabor wavelet transform, however, introduces very high dimensional data. To reduce this dimensionality, uniform sampling of Gabor features has traditionally been used. Since uniform sampling equally treats all the features, it can lead to a loss of important features while retaining trivial ones. In this paper, we propose a new face representation method that employs nonuniform multilevel selection of Gabor features. The proposed method is based on the local statistics of the Gabor features and is implemented using a coarse-to-fine hierarchical strategy. Gabor features that correspond to important face regions are automatically selected and sampled finer than other features. The nonuniformly extracted Gabor features are then classified using principal component analysis and/or linear discriminant analysis for the purpose of face recognition. To verify the effectiveness of the proposed method, experiments have been conducted on benchmark face image databases where the images vary in illumination, expression, pose, and scale. Compared with the methods that use the original gray-scale image with 4096-dimensional data and uniform sampling with 2560-dimensional data, the proposed method results in a significantly higher recognition rate, with a substantial lower dimension of around 700. The experimental results also show that the proposed method works well not only when multiple sample images are available for training but also when only one sample image is available for each person. The proposed face representation method has the advantages of low complexity, low dimensionality, and high discriminance.
Shan Du 0001, Rabab K. Ward
IEEE Trans. Syst. Man Cybern. Part B2
2008 Pseudo-Fisherface method for single image per person face recognition
abstract
The problem of recognizing a face from a single sample available in a stored dataset is addressed. A new method of tackling this problem by using the Fisherface method on a generic dataset is explored. The recognition scheme is also extended to multiscale transform domains like wavelet, curvelet and contourlet. The proposed method in the transform domain shows better recognition errors than the SPCA algorithm and Eigenface selection method, both of which are specially tailored for recognizing faces from single samples.
Angshul Majumdar, Rabab K. Ward
ICASSP2
2008 Fast block size prediction for MPEG-2 to H.264/AVC transcoding
abstract
One objective in MPEG-2 to H.264 transcoding is to improve the H.264 compression ratio by using more accurate H.264 motion vectors. Motion re-estimation is by far the most time consuming process in video transcoding, and improving the searching speed is a challenging problem. We introduce a new transcoding scheme that uses the MPEG-2 DCT coefficients to predict the block size partitioning for H.264. Performance evaluations have shown that, for the same rate-distortion performance, our proposed scheme achieves an impressive reduction in the computational complexity of more than 82% compared to the full range motion estimation used by H.264.
Qiang Tang 0002, Panos Nasiopoulos, Rabab K. Ward
ICASSP3
2008 Sparse spatial filter optimization for EEG channel reduction in brain-computer interface
abstract
Spatial filters are useful in discriminating different classes of electroencephalogram (EEG) signals such as those corresponding to motor activities. In the case of discriminating two classes of signals, EEG signals are projected onto a space where one class of signals is maximally scattered and the other is minimally scattered. This paper finds a minimal number of electrodes that can achieve the discrimination. Applying many electrodes is tedious and time-consuming. To reduce the number of electrodes, we propose inducing sparsity in the spatial filter. We reformulate the optimization problem in Common Spatial Patterns by introducing an ℓ1-norm regularization term. Experimental results on five subjects show that the proposed method significantly reduces the number of electrodes while generating features with good discriminatory information. The number of electrodes on average, is reduced to 11% (of the 118 electrodes) while the average drop in the classification accuracy is only 3.8%.
Xinyi Yong, Rabab K. Ward, Gary E. Birch
ICASSP2
2008 Single image per person face recognition with images synthesized by non-linear approximation
abstract
This paper addresses the problem of identifying faces when the training face database consists of one face image of each person. It proposes a new approach that synthesizes new face samples of varying degrees of edge information; the synthesized images are generated from the original image and form non-linear approximations of the latter. The approximation is framed as anl1minimization problem in a transform domain. The paper also shows that a voting based approach to recognize faces from single available samples yields better results than previous works that only augmented the available database. The proposed approach yields considerably better results (about 6% increase in recognition accuracy) than the SPCA method, which was tailored for addressing this problem.
Angshul Majumdar, Rabab K. Ward
ICIP2
2008 On the security of singular value based watermarking
abstract
The advantage of the singular value (SV)-based image watermarking approach is its robustness to distortion attacks. In this paper, we provide a mathematical proof as to why SV-based watermarking algorithms are robust to small distortion attacks of any type. The same mathematical proof also confirms that such watermarking schemes unfortunately suffer from false watermark detection as long as the distortion attacks are small. Thus a false watermark (not the embedded one) can be detected from a watermarked image that is not severely distorted. SV-based watermarking methods cannot thus be used for protecting the image ownership.
Changzhen Xiong, Rabab K. Ward
ICIP2
2008 Comparison of Evaluation Metrics in Classification Applications with Imbalanced Datasets
abstract
A new framework is proposed for comparing evaluation metrics in classification applications with imbalanced datasets (i.e., the probability of one class vastly exceeds others). For model selection as well as testing the performance of a classifier, this framework finds the most suitable evaluation metric amongst a number of metrics. We apply this framework to compare two metrics: overall accuracy and Kappa coefficient. Simulation results demonstrate that Kappa coefficient is more suitable.
Mehrdad Fatourechi, Rabab K. Ward, Steven G. Mason, Jane E. Huggins, Alois Schlögl, Gary E. Birch
ICMLA2
2008 New narrowband active noise control systems with significantly less computational requirements
abstract
In a typical conventional narrowband ANC system, the discrete Fourier coefficients (DFC) for each frequency are estimated by a linear combiner. Each reference (cosine or sine) wave has to be filtered by an estimate of the secondary-path before it is fed to the LMS algorithm. We call this part x-filtering block. The number of x-filtering blocks is 2q where q is the number of targeted frequencies. For larger q and/or higher order (̂M) of estimated FIR-type secondary-path, the computational cost due to x-filtering may become a bottle-neck in real system implementation. Here, we propose a new narrowband ANC system structure which requires only two (2) x-filtering blocks regardless of q. All the cosine waves (or sine waves) are combined as an input to a x-filtering block. The output of this block is decomposed by an efficient bandpass filter bank into filtered-x cosine or sine waves for the FXLMS that follows. The computational cost of the new system is significantly reduced especially for large q and/or ̂M. The new structure is also modified to cope with the frequency mismatch (FM). Simulations demonstrate that the new systems present performance which is very similar to that of the conventional system, but enjoy great advantages in system implementation.
Yegui Xiao, Maha Shadaydeh, Rabab K. Ward
ISCAS3
2008 Compensation of Requantization and Interpolation Errors in MPEG-2 to H.264 Transcoding
abstract
Implementing MPEG-2 to H.264 transcoding schemes in the pixel domain introduces a high degree of computational complexity. In the transform domain, this transcoding is more computationally efficient, and several methods have been developed to address that approach. However, incompatibilities between the two standards, such as the mismatches between the MPEG-2 and H.264 motion compensation processes, cause several distortions that may affect the overall picture quality. In this study, we address the main distortions that result from requantization errors: luminance half-pixel and chrominance quarter/three-quarter interpolation errors. Then, we propose algorithms that compensate for these errors. The traditional requantization error compensation algorithm for DCT coefficients is updated so that it can be applied to the H.264 integer transform coefficients. Equations that compensate for the luminance half-pixel and chrominance quarter/three-quarter pixel interpolation errors are derived. To remove the interpolation errors, the previous H.264 frame is needed. Thus, the compensation scheme includes a closed-loop H.264 motion compensation process, which is implemented in the pixel domain. To evaluate the performance of the proposed compensation algorithms in terms of picture quality, our scheme is compared with two different cascaded pixel-domain transcoding structures. The first structure reuses the MPEG-2 motion vectors, and the other implements plusmn2 pixels motion vector refinement, but each one has an H.264 deblocking filter. The experimental results show that the proposed compensation algorithms achieve 5-dB quality improvement over the open-loop transform-domain-based transcoding and almost the same picture quality (0.3-0.6 dB) as the cascaded structures. An additional advantage is the reduction in computational complexity that ranges from 13% to 69% compared with the two cascaded methods.
Qiang Tang 0002, Panos Nasiopoulos, Rabab K. Ward
IEEE Trans. Circuits Syst. Video Technol.3
2008 A Video Watermarking Scheme Based on the Dual-Tree Complex Wavelet Transform
abstract
A watermarking scheme that discourages theater camcorder piracy through the enforcement of playback control is presented. In this method, the video is watermarked so that its display is not permitted if a compliant video player detects the watermark. A watermark that is robust to geometric distortions (rotation, scaling, cropping) and lossy compression is required in order to block access to media content that has been re-recorded with a camera inside a movie theater. We introduce a new video watermarking algorithm for playback control that takes advantage of the properties of the dual-tree complex wavelet transform. This transform offers the advantages of the regular and the complex wavelets (perfect reconstruction, shift invariance, and good directional selectivity). Our method relies on these characteristics to create a watermark that is robust to geometric distortions and lossy compression. The proposed scheme is simple to implement and outperforms comparable methods when tested against geometric distortions.
Lino Coria-Mendoza, Mark R. Pickering, Panos Nasiopoulos, Rabab K. Ward
IEEE Trans. Inf. Forensics Secur.4
2007 Applying a Hybrid Genetic Algorithm in the Design of a Self-Paced Brain Interface with a Low False Positive Rate
abstract
A new hybrid genetic algorithm (HGA) for optimization of a self-paced brain interface (SBI) is proposed. To identify intentional control (IC) commands in the noisy background EEG signal, the proposed SBI uses features extracted from three neurological phenomena - movement-related potentials as well as changes in the power of Mu and Beta rhythms. To identify the IC commands, for each neurological phenomenon, a multiple classifier system (MCS) is designed. Then a 2nd-stage MCS combines the outputs of the individual MCSs and generates the final decision. The HGA selects the optimal subset of features, the optimal parameter values of the classifiers, as well as the best configuration for combining the MCSs. Analysis of the data of four subjects shows an average TP = 56.18%, and an average FP = 0.14%, a significant improvement over our previous SBI design.
Mehrdad Fatourechi, Gary E. Birch, Rabab K. Ward
ICASSP (4)3
2007 Efficient Chrominance Compensation for MPEG2 to H.264 Transcoding
abstract
Although open-loop transcoding is known as the most computational efficient transcoding structure, it is also known to introduce many distortions in the transcoded video. This paper addresses the chrominance distortions resulting from the open-loop MPEG2 to H.264 transcoding structure and proposes algorithms to compensate for the chrominance distortions. The open-loop structure is replaced by a closed-loop transcoding structure, which provides high-quality video by removing the chrominance distortions, resulting in an average of 6 dB picture quality improvement.
Qiang Tang 0002, Panos Nasiopoulos, Rabab K. Ward
ICASSP (1)3
2007 A Robust Approach for Eye Localization Under Variable Illuminations
abstract
Illumination variation is a main obstacle in facial feature detection. This paper presents a novel automated approach that localizes eyes in gray-scale face images and that is robust to illumination changes. The approach does not require prior knowledge about face orientation and illumination strength. Other advantages are that no initialization and training process are needed. Based on an edge map obtained via multi-resolution wavelet transform, this approach first segments an image into different inhomogeneously illuminated regions. The illumination of every region is then adjusted so that the features' details are more pronounced. To locate the different facial features, for every region, Gabor-based image is constructed from the re-lit image. The eyes sub-regions are then identified using the edge map of the re-lit image. This method has been applied successfully to the images of the Yale B face database that have different illuminations.
Shan Du 0001, Rabab K. Ward
ICIP (1)2
2007 Segmentation of polarimetric SAR data using contour information via spectral graph partitioning
abstract
A new method for segmenting polarimetric Synthetic Aperture Radar (POLSAR) data is proposed. Image segmentation is formulated as a graph partitioning problem. Spectral graph partitioning - known to provide perceptually plausible image segmentation results using one or more cues (e.g., similarity, proximity, contour continuity) - is applied on POLSAR image data. The degree of similarities between pairs of pixels are calculated based on contour information. Graph partitioning is performed using the Multiclass Spectral Clustering method that minimizes the normalized cut cost function to ensure minimal similarity between partitions. The resulting segmentation is an approximation to the global optimal solution. C-band POLSAR data acquired by CV-580 are used for testing the performance. The results are found to closely agree with manual segmentations.
Kaan Ersahin, Ian G. Cumming, Rabab K. Ward
IGARSS3
2007 Fast RLS Fourier analyzers capable of accommodating frequency mismatch
Yegui Xiao, Liying Ma, Rabab K. Ward
Signal Process.3
2007 A New Orientation-Adaptive Interpolation Method
abstract
We propose an isophote-oriented, orientation-adaptive interpolation method. The proposed method employs an interpolation kernel that adapts to the local orientation of isophotes, and the pixel values are obtained through an oriented, bilinear interpolation. We show that, by doing so, the curvature of the interpolated isophotes is reduced, and, thus, zigzagging artifacts are largely suppressed. Analysis and experiments show that images interpolated using the proposed method are visually pleasing and almost artifact free.
Qing Wang 0048, Rabab K. Ward
IEEE Trans. Image Process.2
2006 A robust content-dependent algorithm for video watermarking
abstract
A watermarking method that relies on informed coding and informed embedding is presented. Our method uses a subset of various codewords to represent the 0 and 1 message bits to be embedded. We propose a codeword generation scheme that keeps control of the distance between codewords in order to secure fidelity and robustness of the watermark. When compared to existing video watermarking schemes, our method yields superior robustness to video compression.
Lino Coria-Mendoza, Panos Nasiopoulos, Rabab K. Ward
Digital Rights Management Workshop3
2006 Detection of Hand Extension Movements in the Context of a 3-State Asynchronous Brain Interface
abstract
The low-frequency asynchronous switch design (LF-ASD) is a direct brain interface (BI) that detects the presence of a specific finger movement in the ongoing EEG. Asynchronous interfaces have the advantage of being operational at all times and not only at specific system-defined periods. In this paper, we present the design of a 3-state asynchronous BI for the detection of two different movements from the ongoing EEG. The proposed 3-state asynchronous BI detects right and left hand extensions. Using data collected from two able-bodied individuals, it is shown that the error characteristics of the new system in detecting the presence of movement are significantly better than the 2-state LF-ASD, with true positive rate increases of up to 22.4% for false positive rates in the 1-2% range. An average performance of 61.5% was achieved in differentiating between left and right hand movements
Ali Bashashati, Rabab K. Ward, Gary E. Birch
ICASSP (5)2
2006 Obtaining LIP and Glottal Reflection Coefficients from Vowel Sounds
abstract
Knowledge about lip and glottal reflection coefficients during phonation is needed to eliminate their distortion effects on the estimates of vocal-tract area functions and glottal waves from vowel sounds. Direct measurements of these coefficients at human mouths are difficult. This paper presents a method for estimating them from vowel sounds. The estimation encounters an ill-defined inverse problem: the number of unknowns is greater than the number of constraints, and non-unique solutions exist for a sound. To overcome this problem, this paper uses a vowel sound produced by a subject whose vocal-tract area function (VTAF) for the sound is known. The estimates of the lip and the glottal reflection coefficients are determined as those that lead to a VTAF solution most similar to the known VTAF for the sound. The lip and the glottal reflection coefficients obtained for/a/and/i/ are presented.
Huiqun Deng, Rabab K. Ward, Michael P. Beddoes, Douglas D. O'Shaughnessy
ICASSP (1)2
2006 Adaptive Region-Based Image Enhancement Method for Face Recognition Under Varying Illumination Conditions
abstract
Illumination changes in face images form a main obstacle in face recognition systems. To deal with this problem, this study presents a novel adaptive region-based image preprocessing scheme that enhances face images and facilitates the face recognition task. This method enhances both the edges and the contrast in face images regionally so as to alleviate the side lighting effects. Compared with the conventional global histogram equalization method, our method is shown to be more suitable for dealing with uneven illuminations in face images. This method is evaluated on the Yale B face database. The experimental results show the advantages of the proposed method with an improvement of 16.1% on average over the histogram eualization method.
Shan Du 0001, Rabab K. Ward
ICASSP (2)2
2006 Using a Multiple Classifier System for Improving the Performance of Asynchronous Brain Interface Systems
abstract
To improve the performance of asynchronous brain interface (ABI) systems, a new classifier design is proposed. The spatial information of multiple EEG channels data is first used to create independent classifiers for different channels. A subset of these classifiers is then selected by a genetic algorithm to form a multiple classifier system (MCS) to decide whether a trial is an intended control or a no control signal. The analysis of the data from 4 subjects shows the effectiveness of the proposed method in improving the performance of an ABI system compared to the results obtained using only the best performing channel
Mehrdad Fatourechi, Gary E. Birch, Rabab K. Ward
ICASSP (5)3
2006 A Novel Algorithm For the Analysis Of Array Cgh Data
abstract
DNA copy number aberrations are common in cancer and other diseases. Newly developed array CGH technologies enable simultaneous measurement of DNA copy numbers for tens of thousands of sites within a genome. In array CGH experiments, DNA copy number of a test DNA sample relative to the DNA copy number of a reference DNA sample is measured. These relative measurements are mapped to their corresponding chromosomal locations. DNA copy gains and losses are then detected as deviations from the normal reference at specific chromosomal locations. In this paper, we introduce a novel algorithm to automatically identify the regions of DNA copy number gain and loss from array CGH data through a multi-scale edge detection algorithm. We demonstrate the method on two array CGH datasets. Our results show that this method can be successfully applied for the analysis of array CGH biological data.
Mehrnoush Khojasteh, Bradley P. Coe, Sohrab P. Shah, Rabab K. Ward, Wan L. Lam, Calum MacAulay
ICASSP (2)4
2006 An Efficient MPEG2 to H.264 Half-Pixel Motion Compensation Transcoding
abstract
An efficient MPEG2 to H.264 half-pixel motion compensation transcoding method is proposed. The Inter macroblock transcoding is implemented in the transform domain. An algorithm is designed to compensate for the errors which arise because of the different half-pixel interpolation procedures used by MPEG2 and H.264/AVC. The experimental results show that the PSNR values of the transcoded H.264 streams result in significant improvement (average 5.5 dB) after we reduce the half-pixel interpolation errors.
Qiang Tang 0002, Rabab K. Ward, Panos Nasiopoulos
ICIP2
2006 An Edge-based Image Interpolation Approach Using Symmetric Biorthogonal Wavelet Transform
abstract
Edge-based image interpolation often leads to an image with good quality because of the importance of sharp edges and smooth contours to the human vision. The wavelet-based image interpolation approach has good potential in producing interpolated images with high quality edges. Using the wavelet multiresolution analysis theory, such method is computationally demanding. In this paper, a new edge-based image interpolation method that uses symmetric biorthogonal wavelet transform is proposed. We form a list of ideal step edges and study the relationships between the wavelet approximation sub-image of each edge and its wavelet detail sub-images. Based on these relationships, a simple and efficient algorithm that predicts the edge information of high resolution images is proposed. For this method, the 9/7-M inverse wavelet transform is shown, experimentally, to yield better image interpolation performance compared to traditional image interpolation approaches
Weizhong Su, Rabab K. Ward
MMSP2
2006 Fast Image/Video Contrast Enhancement Based on WTHE
abstract
We present a fast and effective method for image contrast enhancement based on weighted and thresholded histogram equalization (WTHE). In our proposed method, the probability distribution function of an image is modified by weighting and thresholding before the histogram equalization (HE) is performed. We show that such an approach provides a convenient and effective mechanism to control the enhancement process while being adaptive to various types of images. We also discuss applications of the proposed method in video enhancement
Qing Wang 0048, Rabab K. Ward
MMSP2
2006 A new method for obtaining accurate estimates of vocal-tract filters and glottal waves from vowel sounds
abstract
Previously, estimating vocal-tract filters and glottal waves from vowel sounds imposed either the invalid assumption that glottal waves over closed glottal intervals are zero, or parametric models for glottal waves, resulting in biased vocal-tract-filter estimates and glottal-wave estimates lacking information over closed glottal intervals. We obtain unbiased vocal-tract-filter estimates from sustained vowel sounds, for which the glottal waveforms are periodically stationary random processes. Two equations are constructed each relating the vocal-tract filter to the sound signal and the glottal wave over one of two closed glottal intervals. By subtracting one equation from the other, the periodic components of the glottal wave are eliminated from the vocal-tract-filter estimation, and an unbiased vocal-tract-filter estimate is obtained. The average of many such estimates from different closed glottal intervals of the sound is the final estimate, which is used to obtain the glottal wave by inverse filtering the sound. The results obtained from vowel sounds /a/ produced by some subjects are presented. Over closed glottal phases, the glottal waves obtained are nonzero. During vocal-fold colliding, they increase; during vocal-fold parting, they decrease or even increase. The vocal-tract filters obtained yield vocal-tract area functions similar to that measured from an unknown subject's magnetic resonance image.
Huiqun Deng, Rabab K. Ward, Michael P. Beddoes, Murray Hodgson
IEEE Trans. Speech Audio Process.2
2006 Data Transmission Schemes for DVD-Like Interactive TV
abstract
Current interactive services for digital TV are limited. They basically display a Web page alongside the TV program, which enhances the viewer's experience by providing extra information about the TV program. We define new interactive services for digital TV, which provide DVD-like interactivity to TV viewers. These services enable viewers to control the content and final presentation of a TV program. Some of the attractive applications of our services include parental management, multilingual audio, multiangle video, video in video, etc. The challenge in implementing these services is in transmitting an extra audio or video stream (called incidental) along with the main streams of the TV program. In the first part of this paper, we present a framework for adding the incidental streams to the original transmission stream without increasing the required bandwidth, degrading the picture quality of the main streams, or violating the compatibility of the transmitted stream with standard TV receivers. In the second part of this paper, we explore the two basic mechanisms of the presented framework: traffic characterization and admission control. We present methods for implementing these mechanisms. Using our methods, one can determine whether a TV transmission network has the capability of sending an incidental stream or not. Simulations were conducted to test the validity of our method. The results verify that our method successfully transmits the incidental streams without any discrepancy and without affecting the quality of the main streams
Mehran Azimi, Panos Nasiopoulos, Rabab K. Ward
IEEE Trans. Multim.3
2005 Effects of Glottal and Lip Boundary Conditions on Vocal-Tract Area Function Estimates from Speech Signals
abstract
High-resolution vocal-tract area functions (VTAF) can be derived from vocal-tract filters (VTF) estimated from vowel sound signals with a wide bandwidth. However, the effects of open glottises and frequency-dependent lip reflection coefficients contained in the VTF estimates distort the VTAF estimates. Given VTF estimates obtained over closed glottal phases, we provide a method for eliminating the distortion effects of frequency-dependent lip reflection coefficients on the VTAF estimates. When the VTF estimates contain limited effects of incomplete glottal closures, this method can still obtain reasonable VTAF estimates if the vowel sounds are produced with large lip openings. The VTAF estimates obtained using our method from sounds /a/ produced by different subjects are very similar to that measured using the magnetic resonate imaging method. Theoretically, to eliminate both distortions caused by lip reflection coefficients and incomplete glottal closures in the VTAF estimates, lip-opening areas must be known.
Huiqun Deng, Rabab K. Ward, Michael P. Beddoes, Murray Hodgson
ICASSP (1)2
2005 Statistical Non-Uniform Sampling of Gabor Wavelet Coefficients for Face Recongnition
abstract
A statistics based, non-uniform sampling of the Gabor wavelet decomposition coefficients for face recognition is presented in this paper. Gabor wavelets are popularly used to decompose face images into their spatial/frequency domains. The derived Gabor coefficients generate an augmented vector, e.g., 40 times larger than the original gray-scale vector. To reduce the dimensionality, uniform sampling of the Gabor coefficients is normally used. In this paper, we propose a non-uniform sampling method of the Gabor coefficients such that the coefficients corresponding to important face features are sampled much finer than those of the other parts of the image. The non-uniform sampling is based on the local statistics of the Gabor coefficients obtained from a set of training images. This adaptation is implemented in a hierarchical fashion; a coarse-to-fine strategy results in multi-level sampling rates. After the samples are obtained, the traditional principal component analysis (PCA) is used to code the samples for the final classification. The experimental results show that the proposed non-uniform sampling of Gabor coefficients outperforms the uniform one and the popular eigenfaces method.
Shan Du 0001, Rabab K. Ward
ICASSP (2)2
2005 A hybrid genetic algorithm approach for improving the performance of the LF-ASD brain computer interface
abstract
An asynchronous brain computer interface (BCI) continuously monitors the brain signals and is activated only when a user intends control. Initial results from an asynchronous system, the LF-ASD, designed by our group have shown promise, but the reported error rates are still high for most practical applications. To improve its performance, we propose user customization. Since energy normalization of all channels' signals is shown to significantly improve the performance of the system, we choose to customize the parameters related to this normalization. We apply a hybrid genetic algorithm (a genetic algorithm followed by a local search) to customize the size of the energy normalization windows. This is shown to significantly improve the results. For a fixed false positive rate of 2%, the improvement in the true positive rate was raised from 65.7% to 76.9% in one subject and from 53.1% to 63.3% for another subject.
Mehrdad Fatourechi, Ali Bashashati, Rabab K. Ward, Gary E. Birch
ICASSP (5)3
2005 A robust watermarking scheme based on informed coding and informed embedding
abstract
A watermarking algorithm that relies on informed coding and informed embedding is presented. This method embeds one bit of the watermark in every image block, but offers higher watermarked image quality as well as higher robustness to image processing operation attacks than other known methods using informed coding and informed embedding. The method is shown to withstand high values of added Gaussian noise, valumetric scaling, low-pass filtering as well as lossy JPEG compression. Each bit (0 or 1) to be embedded is represented by a subset of codewords. For every image block, a vector is extracted. This vector is modified so that its correlation with the codewords related to the bit to be embedded in it has higher probability than those of the codewords representing the other bit even if the image is later modified by image processing operations. The vector modification is also carried so that the change in the image fidelity is minimal.
Lino Coria-Mendoza, Panos Nasiopoulos, Rabab K. Ward
ICIP (1)3
2005 Wavelet-based illumination normalization for face recognition
abstract
The appearance of a face image is severely affected by illumination conditions that hinder the automatic face recognition process. To recognize faces under varying illuminations, we propose a wavelet-based normalization method so as to normalize illuminations. This method enhances the contrast as well as the edges of face images simultaneously, in the frequency domain using the wavelet transform, to facilitate face recognition tasks. It outperforms the conventional illumination normalization method - the histogram equalization that only enhances image pixel gray-level contrast in the spatial domain. With this method, our face recognition system works effectively under a wide range of illumination conditions. The experimental results obtained by testing on the Yale face database B demonstrate the effectiveness of our method with 15.65% improvement, on average, in the face recognition system.
Shan Du 0001, Rabab K. Ward
ICIP (2)2
2005 Contrast enhancement for enlarged images based on edge sharpening
abstract
Blurring effect is among the major visual degradations resulting from digital image enlargement. We introduce a set of sigmoidal functions to sharpen edges of enlarged images. We show that these functions have properties that are especially desirable for sharpening expanded edges. We then develop an adaptive contrast enhancement scheme based on edge sharpening using these functions. It is shown that this scheme results in images of better contrast and natural-looking sharp edges.
Qing Wang 0048, Rabab K. Ward, Jiancheng Zou
ICIP (2)2
2005 A stepwise framework for the normalization of array CGH data
abstract
BACKGROUND: In two-channel competitive genomic hybridization microarray experiments, the ratio of the two fluorescent signal intensities at each spot on the microarray is commonly used to infer the relative amounts of the test and reference sample DNA levels. This ratio may be influenced by systematic measurement effects from non-biological sources that can introduce biases in the estimated ratios. These biases should be removed before drawing conclusions about the relative levels of DNA. The performance of existing gene expression microarray normalization strategies has not been evaluated for removing systematic biases encountered in array-based comparative genomic hybridization (CGH), which aims to detect single copy gains and losses typically in samples with heterogeneous cell populations resulting in only slight shifts in signal ratios. The purpose of this work is to establish a framework for correcting the systematic sources of variation in high density CGH array images, while maintaining the true biological variations. RESULTS: After an investigation of the systematic variations in the data from two array CGH platforms, SMRT (Sub Mega base Resolution Tiling) BAC arrays and cDNA arrays of Pollack et al., we have developed a stepwise normalization framework integrating novel and existing normalization methods in order to reduce intensity, spatial, plate and background biases. We used stringent measures to quantify the performance of this stepwise normalization using data derived from 5 sets of experiments representing self-self hybridizations, replicated experiments, detection of single copy changes, array CGH experiments which mimic cell population heterogeneity, and array CGH experiments simulating different levels of gene amplifications and deletions. Our results demonstrate that the three-step normalization procedure provides significant improvement in the sensitivity of detection of single copy changes compared to conventional single step normalization approaches in both SMRT BAC array and cDNA array platforms. CONCLUSION: The proposed stepwise normalization framework preserves the minute copy number changes while removing the observed systematic biases.
Mehrnoush Khojasteh, Wan L. Lam, Rabab K. Ward, Calum MacAulay
BMC Bioinform.3
2004 JasPer: a portable flexible open-source software tool kit for image coding/processing
abstract
JasPer, a portable, flexible, open-source, software tool kit for handling image data is described. This software provides a means for representing images, and facilitates the manipulation of image data as well as its import/export in various formats (such as JPEG 2000). JasPer is proving to be an extremely useful tool in a wide variety of applications, ranging from image coding/processing research to open-source and proprietary software development.
Michael D. Adams 0002, Rabab K. Ward
ICASSP (5)2
2004 A new signal model and identification algorithm for hidden semi-Markov signals
abstract
Markovian models form a powerful tool for modelling physical signals. In this approach, a signal generation model is employed, and its parameters are estimated from signal samples. We present a novel signal generation model for hidden semi-Markov models, HSMMs. Our model results in a significantly easier and more efficient parameter identification method. Instead of the constant probabilities presently used for modelling state transitions, we use state transition probabilities that are state-duration dependant. We then develop a parameter identification algorithm based on the maximum likelihood criterion. Our numerical results show that our parameter identification algorithm can successfully, and more efficiently, estimate the actual values of the model parameters of an HSMM signal.
Mehran Azimi, Panos Nasiopoulos, Rabab K. Ward
ICASSP (2)3
2004 Estimating vocal-tract area functions from vowel sound signals over closed glottal phases
abstract
Existing methods that estimate the vocal-tract area functions (VTAF) from vocal-tract filters (VTF) using speech signals suffer from inadequate elimination of the glottal wave, and the influence of non-ideal vocal-tract boundary conditions. To minimize these effects on the VTF estimation, we present a method that jointly estimates the glottal wave and the VTF corresponding to closed glottal phases. Experimental results show that our method yields better estimates. The VTAF obtained for /a/ and /i/ each produced by a female and a male subject show that more detailed and accurate VTAF are obtained using the speech signals corresponding to closed glottal phases.
Huiqun Deng, Rabab K. Ward, Michael P. Beddoes, Murray Hodgson
ICASSP (1)2
2004 A wavelet-based approach for the extraction of event related potentials from EEG
abstract
Event related potentials (ERPs) are of interest to many researchers seeking knowledge about the functions of the brain. ERPs are low-frequency events that are usually obscured in single trial analysis. To visualize these signals; most of the reliable solutions at the present time use the ensemble averages of many single trials. In this paper, a wavelet-based method called statistical coefficient selection (SCS) is used for the extraction of ERPs from EEG signals. Unlike other wavelet-based denoising methods, the current method does not focus on the wavelet coefficients of the signal itself. Instead, it selects the coefficients based on the statistical study of trials from training data sets. Simulation results show the superiority of the proposed SCS method in extracting ERPs in comparison with other filtering approaches.
Mehrdad Fatourechi, Steven G. Mason, Gary E. Birch, Rabab K. Ward
ICASSP (2)4
2004 The generalized Fibonacci transformations and application to image scrambling
abstract
This paper introduces a subfamily of the generalized Fibonacci sequence family, which we call the distinguished generalized Fibonacci sequence. Two members of this subfamily, the Fibonacci sequence and the Lucas sequence, are considered and two transformations, based on these sequences, are introduced. The applications of these transformations to image scrambling are studied in detail. It is found that these transformations have the desirable property of uniformity, that is, pixels that are equidistant in the original image remain equidistant after scrambling, albeit with different distance values. These transforms also spread adjacent pixels as far as possible. Besides totally decorrelating the image, these transformations also have the advantage of ease of implementation. This renders them useful for real-time and low cost implementations.
Jiancheng Zou, Rabab K. Ward, Dongxu Qi
ICASSP (3)2
2004 A new facial expression recognition technique using 2-D DCT and K-means algorithm
abstract
Facial expression recognition plays a vital role in realizing a highly intelligent human-machine interface, and has recently attracted much attention. In this paper, we propose a new facial expression recognition method that utilizes the 2D DCT, k-means algorithm and vector matching. This technique is based on two main intuitive ideas: (i) complicated facial expression categories such as "anger" and "sadness", may be divided into several subcategories with different subfeature spaces where the recognition task can be performed with higher accuracy and (ii) the k-means algorithm may be used to cluster these subcategories. A new image database with five facial expressions (neutral, smile, anger, sadness, surprise) of 60 women was constructed using a computationally efficient projection-based technique. Experimental results using the new database and an existing one (60 men) reveal that the new technique outperforms the standard vector matching technique and two recently developed methods using fixed-size and constructive one-hidden-layer neural networks. The mean recognition rate can be as high as 95% for the two databases.
Liying Ma, Yegui Xiao, Khashayar Khorasani, Rabab K. Ward
ICIP4
2003 A robust LMS-based Fourier analyzer capable of accommodating the frequency mismatch
abstract
The conventional LMS Fourier analyzer has been successfully used to analyze sinusoidal and periodic signals in additive noise. If the user provides the correct signal frequencies, the analyzer produces good estimates for the discrete Fourier coefficients (DFCs) of the signal. However, if the signal frequencies fed to the analyzer vary from the true signal frequencies, i.e., a frequency mismatch (FM) exists, the performance of the conventional LMS algorithm degenerates. We propose a new LMS-based Fourier analyzer that yields superior results to the conventional LMS one. The estimation of DFCs and the reduction of FM in the new algorithm are carried out simultaneously based on the least mean square and the least mean p-power error criteria, respectively. This new algorithm has a simple structure and shows a small increase in computations. Simulations as well as a real-life application to real signals generated by a large-scale factory cutting machine are provided to show the effectiveness of our new algorithm. For the latter, the performance improvement is as high as 8.3 [dB].
Yegio Xiao, Rabab K. Ward, Liying Ma
ICASSP (6)2
2003 A contour-preserving image interpolation method
abstract
We introduce a novel image interpolation method, which focuses on providing artifact-free contours. In our method, the image contours are divided into edges and ridges, and we estimate the orientation of them differently and apply directional interpolations on them. Our method gives visually pleasing and natural-looking images. Experiment results are shown and compared with other interpolations methods.
Qing Wang 0048, Rabab K. Ward
ICIP (3)2
2003 Estimating the vocal-tract area function and the derivative of the glottal wave from a speech signal
Huiqun Deng, Michael P. Beddoes, Rabab K. Ward, Murray Hodgson
INTERSPEECH3
2003 Wavelet packets-based digital watermarking for image verification and authentication
Alexandre H. Paquet, Rabab K. Ward, Ioannis Pitas
Signal Process.2
2003 Removing the blocking artifacts of block-based DCT compressed images
abstract
One of the major drawbacks of the block-based DCT compression methods is that they may result in visible artifacts at block boundaries due to coarse quantization of the coefficients. We propose an adaptive approach which performs blockiness reduction in both the DCT and spatial domains to reduce the block-to-block discontinuities. For smooth regions, our method takes advantage of the fact that the original pixel levels in the same block provide continuity and we use this property and the correlation between the neighboring blocks to reduce the discontinuity of the pixels across the boundaries. For texture and edge regions, we apply an edge-preserving smoothing filter. Simulation results show that the proposed algorithm significantly reduces the blocking artifacts of still and video images as judged by both objective and subjective measures.
Rabab K. Ward
IEEE Trans. Image Process.2
2002 Symmetry-preserving reversible integer-to-integer wavelet transforms
abstract
Studied are two lifting-based families of symmetry-preserving reversible integer-to-integer wavelet transforms. The transforms from both of these families are shown to be compatible with symmetric extension, which permits the treatment of arbitrary length signals in a nonexpansive manner. Throughout this work, particularly close attention is paid to rounding functions, and the properties that they must possess in various instances. Symmetric extension is also shown to be equivalent to constant per-lifting-step extension in certain circumstances.
Michael D. Adams 0002, Rabab K. Ward
ICASSP2
2002 Wavelet packets-based image retrieval
abstract
The ability to retrieve images from databases is of great importance for a wide range of applications. In this paper, we present a new method for image identification and retrieval that enables the recovery of original images even after size-conserving rotation or flipping operations. Our method uses the correlation of wavelet packets coefficients to create image signatures. It first computes a basic signature for the original image by summing the correlation values along all frequency bands. Possible image rotation/flipping cases are then used to generate additional short signatures that are added to the basis signature, to identify geometric transformations. Simulation results show that our method has image retrieval rates between 88.6% and 91.7%, and geometric transformation recognition of 99.57%.
Alexandre H. Paquet, Saif Zahir, Rabab K. Ward
ICASSP3
2002 A scheduling scheme for multiplexing of VBR sources in digital TV systems
abstract
Digital TV transmission systems allow a transmission channel to be shared by a number of sources. In order to improve the bandwidth utilization, variable bit rate encoding and statistical multiplexing techniques are usually used. However, the channel sharing requires a careful scheduling method for multiplexing. This is because the video and audio materials have to be presented at the receivers at specific points in time. In this paper, we present a novel scheduling scheme for statistical multiplexing of VBR sources. Our method is sensitive to the timing requirements of the sources and sends the packets as close to their transmission deadlines as possible. The advantages of our method are: (1) it decreases the broadcast deadline violation probability (or improves the bandwidth utilization), (2) it minimizes the delay and delay jitter of packets and (3) it generates a transport stream compliant with all the standard TV receivers. Simulations were conducted to compare our algorithm with the first-come-first-serve scheduling method. The results show that our algorithm significantly reduces both the percentage of dropped packets (by 35%-50%) and the average packet delay.
Mehran Azimi, Panos Nasiopoulos, Rabab K. Ward
ICIP (3)3
2002 Isophote estimation by cubic-spline interpolation
abstract
We apply the cubic-spline interpolation to estimate isophotes from sparsely sampled digital images. For any non-pixel, we interpolate it by cubic spline, and by solving the yielding cubic function analytically, we find positions of pixels with the same intensity value. Experiment results are given and discussed. This spreads some important light on the nature of interpolation in images and why the well-known zigzag effects are obtained when images are interpolated.
Qing Wang 0048, Rabab K. Ward, Hongjian Shi
ICIP (3)2
2002 A Rotation Invariant Rule-Based Thinning Algorithm for Character Recognition
abstract
This paper presents a novel rule-based system for thinning. The unique feature that distinguishes our thinning system is that it thins symbols to their central lines. This means that the shape of the symbol is preserved. It also means that the method is rotation invariant. The system has 20 rules in its inference engine. These rules are applied simultaneously to each pixel in the image. Therefore, the system has the advantages of symmetrical thinning and speed. The results show that the system is very efficient in preserving the topology of symbols and letters written in any language.
Maher Ahmed, Rabab K. Ward
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 Genetic Algorithms for Feature Selection and Weighting, A Review and Study
abstract
Our aim is: a) to present a comprehensive survey of previous attempts at using genetic algorithms (GA) for feature selection in pattern recognition applications, with a special focus on character recognition; and b) to report on work that uses GA to optimize the weights of the classification module of a character recognition system. The main purpose of feature selection is to reduce the number of features, by eliminating irrelevant and redundant features, while simultaneously maintaining or enhancing classification accuracy. Many search algorithms have been used for feature selection. Among those, GA have proven to be an effective computational method, especially in situations where the search space is uncharacterized (mathematically), not fully understood, or/and highly dimensional.
Faten Hussein, Rabab K. Ward, Nawwaf Kharma
ICDAR2
2001 New interactive services for digital TV
abstract
Current interactive TV services are limited. They mainly consist of accessing the World Wide Web. These services basically display a Web page beside a TV program that is related to the contents of TV program. We define new interactive services for digital TV, which can be used in many attractive applications such as parental management, multilingual audio and video in video. These services have also the advantages that they do not require an Internet return path and that they preserve the compatibility of the transmitted stream with conventional TV systems and receivers. In this paper we evaluate the feasibility of the proposed interactive services for digital TV and specify the challenges and problems in implementing these services. We present methods for overcoming these problems. We also specify the minimum decoder buffer required and the average random access delay for multilingual support in a typical standard definition TV program.
Mehran Azimi, Panos Nasiopoulos, Rabab K. Ward
ICIP (1)3
2001 Classification of homologous human chromosomes using mutual information maximization
abstract
Multi-feature analysis of human chromosome images is a major step towards classification of homologous chromosomes. An automatic quantitative classification method is proposed for homolog differentiation using multiple features. This method is based on mutual information maximization applied to an unsupervised neural network architecture. The neural network consists of separate modules which are trained to classify homologs using independent features. Mutual information is then maximized between the outputs of the modules forcing them to produce the same classification results, for a given chromosome. The proposed method was successfully applied to classify homologs of chromosome 16 with 100% accuracy.
Parvin Mousavi, Sidney S. Fels, Rabab K. Ward, Peter M. Lansdorp
ICIP (2)3
2001 A new edge-directed image expansion scheme
abstract
A new interpolation-based scheme for image expansion is introduced. In the proposed method, the fidelity and sharpness of edges in the expanded image is emphasized. This is achieved by applying edge-directed interpolation and edge sharpening operations. Compared with other methods, our method provides natural and sharp expanded images with light computational cost.
Qing Wang 0048, Rabab K. Ward
ICIP (3)2
2001 A novel invariant mapping applied to hand-written arabic character recognition
Nawwaf Kharma, Rabab K. Ward
Pattern Recognit.2
2001 A computation-distortion optimized framework for efficient DCT-based video coding
abstract
The rapidly expanding field of multimedia communications has fueled significant research and development work in the area of real-time video encoding. Dedicated hardware solutions have reached maturity and cost-efficient hardware encoders are being developed by several manufacturers. However, software solutions based on a general purpose processor or a programmable digital signal processor (DSP) have significant merits. Toward this objective, we have developed a flexible framework for video encoding that yields very good computation-performance tradeoffs. The proposed framework consists of a set of optimized core components: motion estimation (ME), the discrete cosine transform (DCT), quantization, and mode selection. Each of the components can be configured to achieve a desired computation-performance tradeoff. The components can be assembled to obtain encoders with varying degrees of computational complexity. Computation control has been implemented within the proposed framework to allow the resulting algorithms to adapt to the available computational resources. The proposed framework was applied to MPEG-2 and H.263 encoding using Intel's Pentium/MMX desktop processor. Excellent speed-performance tradeoffs were obtained.
Ismaeil R. Ismaeil, Alen Docef, Faouzi Kossentini, Rabab K. Ward
IEEE Trans. Multim.4
2000 Computation-performance control for DCT-based video coding
abstract
In this paper, we propose a flexible framework for DCT-based video encoding that yields very good computation performance tradeoffs. Each of the encoding components features a set of parameters that can be used to control its computational complexity and performance. A sequence of optimum parameter sets have been designed to obtain encoders with varying degrees of computational complexity. A computation control mechanism is proposed within the encoding framework to allow the encoding algorithm to adapt to the available computational resources. This will allow the encoder to run in real time on machines with different computing power levels, while also achieving the best possible reproduction quality. The proposed framework was applied to MPEG-2 and H.263 encoding. Our experimental results show that excellent speed-performance tradeoffs as well as accurate computation control can be obtained using the proposed method.
Ismaeil R. Ismaeil, Alen Docef, Faouzi Kossentini, Rabab K. Ward
ICASSP4
2000 An Efficient, Similarity-Based Error Concealment Method for Block-Based Coded Images
abstract
We propose an efficient, similarity-based error concealment method for block-based coded images. In a hierarchical matching procedure, the image is first searched at a lower resolution to find the best match of a layer of pixels around a missing block. Then, the search is performed on the full resolution image. A fast search algorithm which follows a diamond-shaped search area is employed in both resolutions. The missing block is replaced with the block connected to the layer that yields the best match. Moreover, the proposed matching criterion takes into account the geometrical structure extracted from the surrounding pixels of a lost block. The fast search matching method needs significantly less number of computations required by a full search matching method and achieves almost the same reconstruction quality.
Ismaeil R. Ismaeil, Shahram Shirani, Faouzi Kossentini, Rabab K. Ward
ICIP4
2000 Multi-Feature Analysis and Classification of Human Chromosome Images Using Centromere Segmentation Algorithms
abstract
Classification of homologous human chromosomes is essential to advanced studies of cancer genetics. This paper describes novel segmentation and classification algorithms to extract multiple features, from microscopy images of chromosomes, for classification purposes. Multicolour images of metaphase chromosomes prepared by applying PNA probes are used for this purpose. Centromeres are segmented using an iterative fuzzy algorithm as well as a gradient method. Moreover, telomere length measurements are performed on chromosome images and normalized for the image database. Multiple intensity features are calculated as a result of the developed algorithms. Heteromorphic chromosomes (such as 16 and 22) are then successfully classified into their parental homologues, based on the calculated multiple features, and used to verify the developed methods.
Parvin Mousavi, Rabab K. Ward, Peter M. Lansdorp, Sidney S. Fels
ICIP2
2000 A Simple and Effective Filter Based on the Rank Difference
abstract
We have developed a simple and effective filter for edge detection, called the rank difference filter. This filter is simple, employing only integer operations to implement and generates results comparable to or better than more complex edge detectors such as the Laplacian of Gaussian and the Canny (1986). For each pixel, we first apply two different rank filters. We then take the difference of the two rank filter results and assign it to that pixel. The parameters that determine the behavior of the rank difference filter are: (i) the alignment pixel location in the filter kernel, (ii) the kernel's shape, (iii) the kernel's size, and (iv) the values of the upper and lower rank numbers. This filter has properties that out-perform those of other filters especially when applied to images corrupted by uniform noise. In addition, the rank difference filter can also be used as a selective morphologic filter.
Steven S. S. Poon, Rabab K. Ward
ICIP2
2000 Interactive DVD Programming Using Next Generation Content-Based Encoded Multimedia Data
abstract
In this paper we propose a method of expanding the existing DVD-Video standard by incorporating MPEG-4 encoded content and its associated object-based interactivity into DVD. The proposed integration of MPEG-4 into DVD will enrich the already proven and successful DVD technology with the advanced interactive capabilities of MPEG-4, while maintaining backward capability with the existing DVD standard.
Katerina Pronina, Rabab K. Ward, Panos Nasiopoulos
ICIP2
2000 A Near Exact Image Expansion Scheme for Bi-Level Images
abstract
Exact bi-level image expansion techniques are required for a wide range of applications such as cartography, calligraphy, medical images, remote sensing, and satellite imagery. Among the methods proposed in the literature are (a) pixel replication; (b) area sizing; (c) interpolation and spline methods; and (d) DCT-based techniques. All these methods generate distortion and noticeable degradation in the quality of images especially around edges. We introduce a new image expansion scheme that produces significantly improved expanded and/or reduced images and maintains high quality edges. This scheme uses an elaborate look-up table that is based on look-ahead-and-back procedures for each pixel and maintains a memory of the pixels' chain code connectivity. The experimental simulation results show that the resized images of the proposed scheme are aesthetically and objectively much better than those of the other methods.
Saif Zahir, Abdul Karim Murad Agha, Rabab K. Ward
ICIP3
2000 A concealment method for video communications in an error-prone environment
abstract
In this paper, we propose a two-stage error-concealment method for block-based compressed video which was transmitted in an error-prone environment. In the first stage, we obtain initial estimates of the missing blocks. If the motion vectors associated with the missing blocks are available, motion compensation is used to provide good estimates. Otherwise, a novel algorithm which preserves image continuity is used to estimate the blocks. In the second stage, a maximum a posteriori (MAP) estimator, which employs an adaptive Markov random field (MRF) as the image a priori model is used to improve the video reconstruction quality. The adaptive model enables the estimation to incorporate information embedded not only in the immediate neighborhood pixels but also in a wider neighborhood into the reconstruction procedure without increasing the order of the MRF model. The proposed concealment method achieves very good computation-performance tradeoffs, as demonstrated via experimental results.
Shahram Shirani, Faouzi Kossentini, Rabab K. Ward
IEEE J. Sel. Areas Commun.3
2000 An expert system for general symbol recognition
Maher Ahmed, Rabab K. Ward
Pattern Recognit.2
2000 Reconstruction of baseline JPEG coded images in error prone environments
abstract
In this paper, a two-stage method for the reconstruction of missing data in the transmission of baseline JPEG coded images in error prone environments is proposed. In the first stage, we estimate the values of the missing DC coefficients. As effects of errors in estimating the missing DC values will appear as a number of stripes across the image, a technique for removing such stripes is also developed. In the second stage, the data of missing blocks is reconstructed by exploiting the correlation between adjacent blocks. Simulation results intricate that our reconstruction method performs very well. The two key contributions of our method are that it does not assume nondifferential encoding of the DC coefficients, and that it performs well in the reconstruction of diagonal edges.
Shahram Shirani, Faouzi Kossentini, Rabab K. Ward
IEEE Trans. Image Process.3
1999 Motion Estimation Using Long Term Motion Vector Prediction
abstract
Summary form only given. This paper presents a motion estimation technique for the coding of video sequences that is based on long-term temporal prediction. The motion vector of a moving object is tracked from one frame to another using a projection method. The traced motion vector is used as a starting point for the motion estimation search algorithm. The motion estimation algorithm used is based on an optimum fast block matching algorithm. Combinations of both spatial and temporal prediction are also used to obtain a more accurate estimate of the motion vector of the current macroblock. An inaccurately predicted motion vector can have a significant negative impact on the motion estimation algorithm. It can force the search algorithm to be trapped in a local minimum, or to spend unnecessary computations to find the optimum motion vector. The accuracy of the predicted motion vector is estimated using a reliability measure that allows the motion search algorithm to decide whether to use temporal motion vector predictor, spatial motion vector predictor, or both. As a reliability measure we used the mean squared error between the traced motion vectors in the current frame and in the previous frame. The reliability measure increases if the motion vector belongs to a moving object with constant speed. If the reliability measure is smaller than a certain threshold, then we abandon the temporal prediction and use spatial prediction. If both prediction methods fail, we abandon the fast motion search and perform full-search motion estimation in the low-resolution images. The experimental results show that long-term prediction reduces the number of computations performed by the motion search algorithm by up to 20%, while obtaining essentially the same quality.
Ismaeil R. Ismaeil, Alen Docef, Faouzi Kossentini, Rabab K. Ward
Data Compression Conference4
1999 Joint MPEG-2 coding for multi-program broadcasting of pre-recorded video
abstract
We developed a cost-effective operational system suitable for digital TV, video on demand, and high definition TV broadcast over satellite networks with limited bandwidth. This MPEG-2 based system is easy to implement and allows the joint video coding of multiple video programs. Compared to present broadcast operations and for the same level of picture quality, our system greatly increases the number of video streams transmitted in each channel. As a result, either a large number of transponders can be freed up to carry real-time broadcasting or the level of the transmitted picture quality can be significantly increased, By switching from tape storage to video server technology, the need for numerous (expensive) playback VTR systems at the headend is eliminated. In addition, the majority of the complete MPEG-2 encoders are replaced by much less complex MPEG-2 transcoders. All this means significant savings for the broadcast stations. In addition to the gain in bandwidth and the reduction in cost, our system speeds up the encoding process by six fold.
Irene Koo, Panos Nasiopoulos, Rabab K. Ward
ICASSP3
1999 Segmenting telomeres and chromosomes in cells
abstract
The very end of every chromosome is a region called the telomere. Telomeres are nucleo-protein complexes containing specific DNA repeat sequences whose lengths are strongly believed to give indications to aging and tumor progression. In order to study the role these repeat sequences play in the cell, we developed a fluorescence microscopy imaging system and associated image analysis methods to accurately measure these telomere lengths. To visualize the image of the tiny telomeres, we captured 2 spectrally different images of the same cell. One image contains only telomeres and the other contains only chromosomes. We next apply successful and novel methods to segment the telomere and chromosome images and then to link each chromosome with its telomeres. Our system is so far the only existing system available for this purpose and has already been in use in many research laboratories in Western Europe, North America, and Hong Kong.
Steven S. S. Poon, Rabab K. Ward, Peter M. Lansdorp
ICASSP2
1999 An adaptive Markov random field based error concealment method for video communication in an error prone environment
abstract
Loss of coded data during its transmission can affect a decoded video sequence to a large extent, making concealment of errors caused by data loss a serious issue. Previous work in spatial error concealment exploiting MRF models used a single pixel wide region around the erroneous area to achieve a reconstruction based on an optimality measure. This practically restricts the amount of available information that is used in a concealment procedure to a small region around the missing area. Incorporating more pixels usually means a higher order model and this is expensive as the complexity grows exponentially with the order of the MRF model. Using previously proposed approaches, the damaged area is reconstructed fairly well in very low frequency portions of the image. However, the reconstruction process yields blurry results with a significant loss of details in high frequency, or edge portions of the image. In our proposed approach, a MRF is used as the image a priori model. More available information is incorporated in the reconstruction procedure not by increasing the order of the model but instead by adaptively adjusting the model parameters. Adaptation is done based on the image characteristics determined in a large region around the damaged area. Thus, the reconstruction procedure can make use of information embedded in not only immediate neighborhood pixels but also in a wider neighborhood without a dramatic increase in computational complexity. The proposed method outperforms the previous methods in the reconstruction of missing edges.
Shahram Shirani, Faouzi Kossentini, Rabab K. Ward
ICASSP3
1999 A Simple Invariant Mapping Applied to Hand-written Pre-segmented Character Recognition
abstract
The paper describes an application of a novel (position, rotation, and size) invariant mapping, one that is intended for use in online handwritten character recognition. The mapping is as simple as any similar mapping we know of. This makes it computationally efficient and fast, which in turn makes it appropriate for online implementations. A recognition system utilising this mapping has been developed for handwritten Arabic characters. Character recognition is carried out via a decision tree-like classification system based mainly (but not solely) on simple features of the mapped pattern.
Nawwaf Kharma, Rabab K. Ward
ICDAR2
1999 Removal of Blocking Artifacts Using Random Pattern Filtering
abstract
Block edge artifacts is one of the most noticeable types of impairments associated with MPEG and other block based encoding techniques. Many postprocessing algorithms have been developed for the reduction of blocking artifacts, such as spatial averaging methods, DCT filtering, regularized image reconstruction techniques and wavelet filtering. Most of the times, these algorithms are either computationally complex, include multiple iterations, are not adaptive enough to remove different levels of blockiness severity or result in excessive smoothing of the image textures. In this paper, we introduce an original approach to the problem of removing blockiness-non-linear randomized displacement filtering. In this technique, blocking artifacts are removed using variable width non-linear filtering as well as randomized masking of geometric patterns associated with blockiness. Our algorithm is highly efficient in removing blockiness with different degrees of severity, does not result in over-smoothing of the image textures and is not computationally complex. Our technique is simple enough to be used in post-processing of real-time video. It is highly efficient in removing blocking artifacts and produces visual results that are much more pleasing to the eye than the results of other de-blocking algorithms of the same complexity level.
Ekaterina Barrykina, Rabab K. Ward
ICIP (2)2
1999 Efficient Motion Estimation Using Spatial and Temporal Motion Vector Prediction
abstract
This paper presents a motion estimation technique for the coding of moving video sequences that is based on spatial and temporal prediction. The motion vector of a moving object is tracked using spatial and temporal prediction and used as a starting point for the motion estimation search algorithm. The predicted motion vector is selected from several candidate motion vectors according to the block matching criterion. Experimental results show that spatio-temporal prediction reduces the number of computations performed by the motion search algorithm by 30% for MPEG-2 encoding and by 40% for H.263 encoding.
Ismaeil R. Ismaeil, Alen Docef, Faouzi Kossentini, Rabab K. Ward
ICIP (1)4
1999 Content-Based Retrieval of Video Sequences Under Partial Occlusion
abstract
In this paper, we present a spatio-temporal segmentation method for content-based video retrieval. Our method is effective in the presence of partial occlusion or when new areas become exposed. In addition, periodic spatial segmentation is not performed. Rather, it is invoked as needed depending on whether or not new regions have been detected. A discussion of how this method can be used for content-based retrieval is also provided. Finally, experimental results that demonstrate the performance of the proposed spatio-temporal segmentation method are presented.
Shahram Shirani, Ali Jerbi, Faouzi Kossentini, Rabab K. Ward, Q. M. Jonathan Wu
ICIP (3)4
1998 MPEG-2 video coding with image partitioning
abstract
In this paper we present an original motion compensation strategy based on frame partitioning. The proposed method uses different temporal resolutions within a frame to improve compression. We present a new bit allocation and rate control algorithm complementing our motion compensation technique. This unique approach to bit allocation ensures the consistency of quality throughout a single frame and a GOP. For the same picture quality, frame partitioning alone yields an additional increase of up to 20 percent or more of the encoding efficiency.
Ekaterina G. Barzykina, Panos Nasiopoulos, Rabab K. Ward
ICASSP3
1998 Reconstruction of Motion Vector Missing Macroblocks in263 Encoded Video Transmission over Lossy Networks
Shahram Shirani, Faouzi Kossentini, Rabab K. Ward
ICIP (3)3
1997 Prediction and search techniques for RD-optimized motion estimation in a very low bit rate video coding framework
abstract
Prediction and search techniques are introduced for efficient rate-distortion optimized motion estimation in a very low bit rate video coding framework. For prediction, three types of predictors are considered: mean, weighted mean, and median. Prediction allows us to constrain the motion vector search to a small diamond-shaped area whose center is the predicted motion vector. The size of the search area is further constrained by employing a probabilistic model. We evaluate two models, both of which permit the contraction or the expansion of the search area as a function of the local statistics of the motion flow. The proposed techniques are analyzed in the context of a very low bit rate DCT-based video coding framework, where a rate-distortion criterion is used for motion estimation as well as for 8/spl times/8 block coding mode selection. A particular resulting very low bit rate video coder is shown experimentally to outperform the H.263 TMN5 simulation model in terms of encoding speed and compression performance, simultaneously.
Yuen-Wen Lee, Faouzi Kossentini, Mark J. T. Smith, Rabab K. Ward
ICASSP4
1997 Efficient RD Optimized Macroblock Coding Mode Selection for MPEG-2 Video Encoding
abstract
The MPEG-2 bitstream syntax offers a great variety of coding mode options for encoding a macroblock. We present an MPEG-2 compliant interlaced video encoder that employs an efficient macroblock coding mode selection method. The macroblock coding mode is selected based on a rate-distortion criterion. However, a computational analysis as well as predictive and statistical modeling techniques are introduced that lead to substantially better computation-performance trade-offs. The proposed MPEG-2 video encoder is shown experimentally to be better than the MPEG-2 TM5 encoder in both compression performance and computation requirements, simultaneously.
Yuen-Wen Lee, Faouzi Kossentini, Rabab K. Ward
ICIP (2)3
1997 Predictive RD Optimized Motion Estimation for Very Low Bit-Rate Video Coding
abstract
Predictive rate-distortion (RD) optimized motion estimation techniques are studied and developed for very low bit-rate video coding. Four types of predictors are studied: mean, weighted mean, median, and statistical mean. The weighted mean is obtained using conventional linear prediction techniques. The statistical mean is obtained using a finite-state machine modeling method based on dynamic vector quantization. By employing prediction, the motion vector search can then be constrained to a small area. The effective search area is reduced further by varying its size based on the local statistics of the motion field, through using a Lagrangian as the search matching measure and imposing probabilistic models during the search process. The proposed motion estimation techniques are analyzed within a simple DCT-based video coding framework, where an RD criterion is used for alternating among three coding modes for each 8/spl times/8 block: motion only, motion-compensated prediction and DCT, and intra-DCT. Experimental results indicate that our techniques yield very good computation-performance tradeoffs. When such techniques are applied to an RD optimized H.263 framework at very low bit rates, the resulting H.263 compliant video coder is shown to outperform the H.263 TMN5 coder in terms of compression performance and computations simultaneously.
Faouzi Kossentini, Yuen-Wen Lee, Mark J. T. Smith, Rabab K. Ward
IEEE J. Sel. Areas Commun.4
1997 Towards MPEG4: An improved H.263-based video coder
Yuen-Wen Lee, Faouzi Kossentini, Rabab K. Ward, Mark J. T. Smith
Signal Process. Image Commun.3
1996 Improving the subjective quality of low bit rate subband/wavelet coded images
abstract
Distortion resulting from ringing and/or aliasing is the most objectionable subjective distortion normally appearing in low bit rate subband/wavelet coded images. This paper presents a method for significantly reducing both ringing and aliasing, while still achieving low bit rates and requiring low encoding/decoding complexity. The method employs a new signal dependent two-band filter bank system that employs short FIR filters, yet results in a relatively small loss of aliasing cancellation in its synthesis section. The signal dependent part of the filter bank is the statistical linear prediction used to decorrelate the subband signals. Experimental results indicate that this approach can achieve very good subjective quality at low bit rates.
Yuen-Wen Lee, Faouzi Kossentini, Rabab K. Ward
ICASSP3
1996 Very low rate DCT-based video coding using dynamic VQ
abstract
We introduce a low-complexity, high-performance DCT-based video coding algorithm that is designed for very low bit rates. The algorithm employs a fast statistical and rate-distortion-optimized motion estimation method with integer-pel accuracy. The searching is constrained to a small diamond-shaped area whose center is the statistical mean, which is obtained using a finite-state machine modeling method based on dynamic state vector quantization. The same modeling method is also employed for coding the motion vectors. A simple rate-distortion criterion is used for alternating between three modes of operation for each 8/spl times/8 block: motion-only coding, motion-compensated predictive coding, and intra coding. Experimental results indicate that the proposed algorithm can outperform current H.263-based video coders significantly at very low bit rates. But, more important is that our algorithm's rate, distortion, and complexity can be easily controlled, a feature desired in many very low bit rate video communication applications due to power and mobility constraints.
Yuen-Wen Lee, Rabab K. Ward, Faouzi Kossentini, Mark J. T. Smith
ICIP (1)2
1996 HDTV picture quality performance in the presence of random errors, analysis and measures for improvement
Panos Nasiopoulos, Rabab K. Ward, Dimitrios P. Bouras, P. Takis Mathiopoulos
Signal Process. Image Commun.2
1996 Automatic Assessment of Infants' Levels-of-Distress from the Cry Signals
abstract
It has long been an attractive prospect to monitor and assess infants' physical and emotional well-being by analyz- ing their cries. The feasibility of such an approach depends on the introduction of modern computer signal processing technology into infant cry research. In this paper, we first discuss the issue of quantifying the distress level of infants. We propose the use of a single item-the level of distress (L0D)-to describe parents' perceptual assess- ment of infants' distress situation from cry sounds. Then, we introduce the concept of cry modes and cry mode sequences to efficiently represent the time-frequency characteristics of infant cries. With the introduction of the cry mode representation scheme, a composite parameter-the H -value-is established as a descrip- tor of the LOD. The H-value can be derived from the cry sound signal. In our experiment, the H-value showed strong consistency with the parents' LOD rating. Finally, an automatic cry analysis system based on hidden Markov models (HMM's) is explained. This system automatically estimates the W-value from a cry sound signal. Testing with 58 different cries uttered by infants under various degrees of stimulation has shown that this HMM-based system is able to give assessments of the distress level of infants consistent with the perceptions of experienced parents. This work provides the basis for the design of automated devices for the monitoring and evaluation of cries of normal and sick infants.
Qiaobing Xie, Rabab K. Ward, Charles A. Laszlo
IEEE Trans. Speech Audio Process.2
1996 Reduction of boundary artifacts in image restoration
abstract
The abrupt boundary truncation of an image introduces artifacts in the restored image. The traditional solution is to smooth the image data using special window functions such as Hamming or trapezoidal windows. This is followed by zero-padding and linear convolution with the restoration filter. This method improves the results but still distorts the image, especially at the margins. Instead of the above method, we propose a different procedure. This procedure is simple and exploits the natural property of "circular" or periodic convolution of the discrete Fourier transform (DFT). Instead of padding the image by zeros, it is padded by a reflected version of it. This is followed by "circular" convolution with the restoration filter. This procedure is shown to lead to better restoration results than the windowing and linear convolution techniques. The computational effort is also improved since our method requires half the number of computations required by the conventional linear deconvolution method.
Farzin Aghdasi, Rabab K. Ward
IEEE Trans. Image Process.2
1995 A Hybrid Coding Method for Digital HDTV Signals
abstract
Digitally transmitted video images suffer from sudden degradation in picture quality at high channel noise levels. This is due to the variable length coding and the large synchronization blocks used. To remedy that, we propose encoding the DCT coefficients by a hybrid fixed and variable lengths compression scheme. This method improves the compression ratio by approximately 20% and increases the system's resistance to channel errors. We then combine this hybrid algorithm with an error protected synchronization method which uses small blocks of size 32/spl times/16 pixels. The resulting combined schemes 1) significantly improve the error resistance characteristic of the system, 2) eliminate the abrupt picture degradation, and 3) does not alter the original data transmission rate.
Panos Nasiopoulos, Rabab K. Ward
ISCAS2
1995 A high-quality fixed-length compression scheme for color images
abstract
We present a new compression method which compresses 8/spl times/8 picture blocks by fixed-length codewords. The compression operation is performed on the discrete cosine transforms, DCT, of each block. As a result, our method combines the distinct advantage of being fixed-length with the high image quality obtained by the DCT based compression methods. Our method has excellent error-resistance characteristics since it does not have the synchronization and error propagation problems inherent in variable-length coding methods.
Panos Nasiopoulos, Rabab K. Ward
IEEE Trans. Commun.2
1994 Improving the Picture Quality of Cable Television
abstract
Two major impairments which affect the cable television picture quality are the thermal noise and the composite triple beats (CTB). We introduce a novel CTB reducing scheme which reduces the visual effect of CTB significantly. This in turn allows cable operators to boost up signal levels so as to reduce the thermal noise, thus improve the picture quality of their cable systems.>
Pingnan Shi, Rabab K. Ward
ICIP (3)2
1993 Improving the HDTV picture performance under noisy transmission conditions
Panos Nasiopoulos, Rabab K. Ward
ICASSP (5)2
1993 Vector Quantization Technique for Nonparametric Classifier Design
abstract
An effective data reduction technique based on vector quantization is introduced for nonparametric classifier design. Two new nonparametric classifiers are developed, and their performance is evaluated using various examples. The new methods maintain a classification accuracy that is competitive with that of classical methods but, at the same time, yields very high data reduction rates.>
Qiaobing Xie, Charles A. Laszlo, Rabab K. Ward
IEEE Trans. Pattern Anal. Mach. Intell.3
1993 Restoration of differently blurred versions of an image with measurement errors in the PSF's
abstract
Restoration of an object from T observations is considered. Each image is distorted by a different deterministic blur and additive noise. The point spread function (PSF) for each observation is unknown; however, a noisy measurement of it is available. Taking the errors in measurements of the PSFs into consideration, the maximum-likelihood and Wiener filters are derived. It is shown that these filters give better results when the regression filter and the conventional Wiener filter, i.e., the one which ignores the presence of the noise in the PSFs. The consistency and the ill-conditioning characteristics of the filters are discussed. Regularized forms for these filers are obtained.
Rabab K. Ward
IEEE Trans. Image Process.1
1993 OSNet: a neural network implementation of order statistic filters
abstract
A dedicated neural network model called OSNet which finds the kth largest element in an array of real numbers is proposed. Its overall processing time is constant irrespective of the number of elements in the array and is four times the processing time of a single neuron. Networks of this kind may be used as building blocks for hardware implementation of order statistic filters. Examples of using OSNet for implementing various order statistic filters and for sorting are shown.
Pingnan Shi, Rabab K. Ward
IEEE Trans. Neural Networks2
1992 Restoration of randomly blurred images via the maximum a posteriori criterion
abstract
The maximum a posteriori (MAP) estimation technique is applied to the problem of restoring images distorted by noisy point spread functions and additive noise. The resulting MAP estimator is nonlinear and is obtained by numerically maximizing a conditional probability density function. The energy nonnegativity constraint is incorporated in the optimization process. Although the deblurring results are slightly inferior to those obtained by applying the Wiener criterion, the advantage of the MAP estimator lies in its significant suppression of noise.
Ling Guan, Rabab K. Ward
IEEE Trans. Image Process.2
1991 Semi-blind restoration from differently blurred versions of an image
abstract
Restoration of an object from K differently distorted versions in the presence of additive noise is considered. The point spread function (PSF) for each observation is unknown, however a noisy measurement of it is available. The blurring processes are assumed fixed and not random. The regression, the maximum likelihood, and the Wiener filters are derived. The consistency characteristics and the computation instabilities of these filters are discussed. Experimental results comparing the performance of these filters are presented.>
Rabab K. Ward, Edward Lam 0002
ICASSP1
1991 Adaptive compression coding
abstract
A compression technique which preserves edges in compressed pictures is developed. The proposed compression algorithm adapts itself to the local nature of the image. Smooth regions are represented by their averages and edges are preserved using quad trees. Textured regions are encoded using BTC (block truncation coding) and a modification of BTC using look-up tables. A threshold using a range which is the difference between the maximum and the minimum grey levels in a 4*4 pixel quadrant is used. At the recommended value of the threshold (equal to 18), the quality of the compressed texture regions is very high, the same as that of AMBTC (absolute moment block truncation coding), but the edge preservation quality is far superior to that of AMBTC. Compression levels below 0.5-0.8 b/pixel may be achieved.>
Panos Nasiopoulos, Rabab K. Ward, Daryl J. Morse
IEEE Trans. Commun.2
1990 The case for abandoning the biological resemblance restriction: an example of neural network solution of simultaneous equations
abstract
The authors construct, in an easy and straightforward way, a nonbiological neural network which solves any system of simultaneous equations. This network is inspired by the Hopfield network; however, its performance is far better than the Hopfield network. The authors show the potential of abandoning the restriction of biological resemblance in the computational approach and the essential difference between the biological and the computational approaches. It is concluded that the biological and the computational approaches have their own rightful places in the field of neural networks. While the biological approach may result in deeper understanding of the human brain, the computational approach may produce machines which are more powerful than the human brain, and each approach can benefit from the progress of the other
P. Shi, Rabab K. Ward
IJCNN2
1988 A maximum a posteriori approach to the restoration of randomly distorted signals
abstract
A maximum a posterior (MAP) approach to the restoration of randomly distorted signals is introduced. Since random distortion is due to the uncertainty in the impulse response of the physical system, it cannot be modelled as additive. The theoretical derivation of the MAP filter shows that it is necessary to solve a nonlinear optimization problem. The performance of the method is compared to that of a modified Wiener filter, and a numerical example shows that the MAP filter outperforms the modified Wiener filter.>
Ling Guan, Rabab K. Ward
ICASSP2
1987 On Determining the On-Line Minimax Linear Fit to a Discrete Point Set in the Plane
Claudio Rey, Rabab K. Ward
Inf. Process. Lett.2
1986 Parity Check Codes for Logic Processors
abstract
Parity check codes, applicable for error control for all bit-wise logical operations, are presented. These codes could be used for protecting information against errors in logic processors as well as protecting it during storage and transmission. They are useful for those cases when the Reed–Muller codes are not applicable and when the word lengths are not too large. Encoding and decoding of these codes are very simple.
Rabab K. Ward
Comput. J.1
1984 Error Correction and Detection, a Geometric Approach
abstract
This paper is basically a tutorial on error detection and correction. The presentation though is via a new approach. This approach has an inherent characteristic leading to parallel hardware implementation and is thus better suited for computer and parallel processing applications. Codes with the desired specifications are constructed by the computer. The design of the codes is within the geometric framework of the Karnaugh map. The mathematics involved is simpler and more intuitive than the traditionally highly mathematical approach. Original results on error detection and conditions for optimal codes are obtained, optimal in the sense of leading to minimum hardware time delays.
Rabab K. Ward, M. Tabandeh
Comput. J.1
1984 A development of a socioeconomic information system: A land reform model for Zimbabwe
abstract
The problem of redistribution of ownership of agricultural land in Zimbabwe is discussed. Based on the present circumstances and resources, a model is developed for maximizing the total production and the total income of the tribal population for the years 1980 to 1990 by determining the optimum split between the two kinds of land use.
Rabab K. Ward
IEEE Trans. Syst. Man Cybern.1