Lina J. Karam

dblp:k/LJKaram · DBLP profile ↗
← Back
102ranked-venue papers
10as first author
4since 2021 · last 2024
0000-0003-1870-1211ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 86 · 8 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4Computer networks · 3Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Trustworthy machine learning · 49% Image recognition and object detection · 30% Deep learning architectures and training · 13%
Computer graphics and multimedia
19 papers
Image and video coding · 60% Image and video processing · 23% Visualization and visual analytics · 12%

Topics — the 30 heaviest of 53, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video coding
image quality assessment
1.162021
Reduced-Reference Quality Assessment Based on the Entropy of DWT Coefficients of Locally Weighted Gradient Magnitudes · IEEE Trans. Image Process. 2016
A Locally Weighted Fixation Density-Based Metric for Assessing the Quality of Visual Saliency Predictions · IEEE Trans. Image Process. 2016
A No-Reference Texture Regularity Metric Based on Visual Saliency · IEEE Trans. Image Process. 2015
Machine learning › Trustworthy machine learning
robustness
0.822020
Defending Against Universal Attacks Through Selective Feature Regeneration · CVPR 2020
DeepCorrect: Correcting DNN Models Against Image Distortions · IEEE Trans. Image Process. 2019
Machine learning › Trustworthy machine learning › robustness
adversarial attack
0.612022
Frequency-Tuned Universal Adversarial Attacks on Texture Recognition · IEEE Trans. Image Process. 2022
Computer vision › Image recognition and object detection
texture classification
0.612022
Frequency-Tuned Universal Adversarial Attacks on Texture Recognition · IEEE Trans. Image Process. 2022
Machine learning › Trustworthy machine learning › robustness › adversarial examples
universal adversarial perturbation
0.612022
Frequency-Tuned Universal Adversarial Attacks on Texture Recognition · IEEE Trans. Image Process. 2022
Machine learning › Generative modeling
generative adversarial network
0.512021
It GAN Do Better: GAN-Based Detection of Objects on Images With Varying Quality · IEEE Trans. Image Process. 2021
Computer vision › Image recognition and object detection
object detection
0.512021
It GAN Do Better: GAN-Based Detection of Objects on Images With Varying Quality · IEEE Trans. Image Process. 2021
Computer vision › Image recognition and object detection › object detection
robust object detection
0.512021
It GAN Do Better: GAN-Based Detection of Objects on Images With Varying Quality · IEEE Trans. Image Process. 2021
Visualization and visual analytics
visual saliency
0.522016
A Locally Weighted Fixation Density-Based Metric for Assessing the Quality of Visual Saliency Predictions · IEEE Trans. Image Process. 2016
A No-Reference Texture Regularity Metric Based on Visual Saliency · IEEE Trans. Image Process. 2015
Machine learning › Trustworthy machine learning › adversarial machine learning
adversarial defense
0.412020
Defending Against Universal Attacks Through Selective Feature Regeneration · CVPR 2020
Machine learning › Trustworthy machine learning › adversarial machine learning › adversarial defense
universal perturbation defense
0.412020
Defending Against Universal Attacks Through Selective Feature Regeneration · CVPR 2020
Machine learning › Deep learning architectures and training
mixture of experts
0.422018
Quality Robust Mixtures of Deep Neural Networks · IEEE Trans. Image Process. 2018
Visual Saliency Prediction Using a Mixture of Deep Neural Networks · IEEE Trans. Image Process. 2018
Computer vision › Image recognition and object detection
image classification
0.412019
DeepCorrect: Correcting DNN Models Against Image Distortions · IEEE Trans. Image Process. 2019
Machine learning › Trustworthy machine learning › robustness › corruption robustness
image distortion robustness
0.412019
DeepCorrect: Correcting DNN Models Against Image Distortions · IEEE Trans. Image Process. 2019
Image and video processing
saliency detection
0.312018
Visual Saliency Prediction Using a Mixture of Deep Neural Networks · IEEE Trans. Image Process. 2018
Image and video coding › image quality assessment
blur detection
0.312017
Spatially-Varying Blur Detection Based on Multiscale Fused and Sorted Transform Coefficients of Gradient Magnitudes · CVPR 2017
Image and video coding › image quality assessment › objective image quality assessment
reduced-reference image quality assessment
0.212016
Reduced-Reference Quality Assessment Based on the Entropy of DWT Coefficients of Locally Weighted Gradient Magnitudes · IEEE Trans. Image Process. 2016
Image and video processing › edge detection
canny edge detection
0.212014
A Distributed Canny Edge Detector: Algorithm and FPGA Implementation · IEEE Trans. Image Process. 2014
Image and video processing
edge detection
0.212014
A Distributed Canny Edge Detector: Algorithm and FPGA Implementation · IEEE Trans. Image Process. 2014
Reconfigurable computing and FPGAs
FPGA accelerator
0.212014
A Distributed Canny Edge Detector: Algorithm and FPGA Implementation · IEEE Trans. Image Process. 2014
Image and video coding › image quality
image quality degradation
0.112021
It GAN Do Better: GAN-Based Detection of Objects on Images With Varying Quality · IEEE Trans. Image Process. 2021
Image and video coding › image compression
perceptual image compression
0.132006
JPEG2000 Encoding With Perceptual Distortion Control · IEEE Trans. Image Process. 2006
Adaptive image coding with perceptual distortion control · IEEE Trans. Image Process. 2002
Locally adaptive perceptual image coding · IEEE Trans. Image Process. 2000
Image and video processing › super-resolution
image super-resolution
0.112011
An Efficient Selective Perceptual-Based Super-Resolution Estimator · IEEE Trans. Image Process. 2011
Multimedia analysis and retrieval
image classification
0.112018
Quality Robust Mixtures of Deep Neural Networks · IEEE Trans. Image Process. 2018
Image and video coding › image quality assessment
no-reference image quality assessment
0.112009
A No-Reference Objective Image Sharpness Metric Based on the Notion of Just Noticeable Blur (JNB) · IEEE Trans. Image Process. 2009
Image and video coding › image quality assessment
sharpness metric
0.112009
A No-Reference Objective Image Sharpness Metric Based on the Notion of Just Noticeable Blur (JNB) · IEEE Trans. Image Process. 2009
Image and video coding
image compression
0.132002
Adaptive image coding with perceptual distortion control · IEEE Trans. Image Process. 2002
Locally adaptive perceptual image coding · IEEE Trans. Image Process. 2000
Image coding with robust channel-optimized trellis-coded quantization · IEEE J. Sel. Areas Commun. 2000
Visualization and visual analytics
visual attention
0.112016
A Locally Weighted Fixation Density-Based Metric for Assessing the Quality of Visual Saliency Predictions · IEEE Trans. Image Process. 2016
Image and video processing
wavelet transform
0.112016
Reduced-Reference Quality Assessment Based on the Entropy of DWT Coefficients of Locally Weighted Gradient Magnitudes · IEEE Trans. Image Process. 2016
Image and video coding › error resilience
error-resilient image coding
0.112007
Selective Error Detection for Error-Resilient Wavelet-Based Image Coding · IEEE Trans. Image Process. 2007

Methods — techniques the papers use, named apart from their topics

gating network · 1.3convolutional neural network · 1.0generative adversarial network · 1.0feature generation · 1.0bounding box regression · 1.0frequency-domain perturbation · 0.6adversarial attack · 0.6feature regeneration · 0.4adversarial training · 0.4fine-tuning · 0.4weight sharing · 0.3treenet · 0.3inverted treenet · 0.3end-to-end training · 0.3gradient magnitude transform · 0.3nonuniform gradient histogram · 0.2distributed block-level thresholding · 0.2probability summation model · 0.1
YearPublicationVenuePosition
2024 Reduced-complexity Convolutional Neural Network in the compressed domain
Hamdan Abdellatef, Lina J. Karam
Neural Networks2
2023 Panel: Faculty Leadership
abstract
This faculty leadership panel will be held jointly by the IEEE Education Society (IEEE EdSoc) and IEEE Educational Activities Board (EAB) Faculty Resources Committee (FRC). The panel intends to provide participants introductory knowledge and skills. Two moderators and four panelists will focus on discussing their various academic leadership experiences and practices. The panel will be held in an interactive way so that the participants will have enough time to ask questions or communicate with the panelists. It is expected that this joint panel may potentially help to build up networks and mentorships.
Lina J. Karam, Georges Zissis, Jason Yao, Jill K. Nelson, Minho Jo 0001, Rafal Sliz, Michele Nogueira Lima, Steve Watkins
FIE1
2022 Frequency-Tuned Universal Adversarial Attacks on Texture Recognition
abstract
Although deep neural networks (DNNs) have been shown to be susceptible to image-agnostic adversarial attacks on natural image classification problems, the effects of such attacks on DNN-based texture recognition have yet to be explored. As part of our work, we find that limiting the perturbation’s$l_{p}$norm in the spatial domain may not be a suitable way to restrict the perceptibility of universal adversarial perturbations for texture images. Based on the fact that human perception is affected by local visual frequency characteristics, we propose a frequency-tuned universal attack method to compute universal perturbations in the frequency domain. Our experiments indicate that our proposed method can produce less perceptible perturbations yet with a similar or higher white-box fooling rates on various DNN texture classifiers and texture datasets as compared to existing universal attack techniques. We also demonstrate that our approach can improve the attack robustness against defended models as well as the cross-dataset transferability for texture recognition problems.
Yingpeng Deng, Lina J. Karam
IEEE Trans. Image Process.2
2021 It GAN Do Better: GAN-Based Detection of Objects on Images With Varying Quality
abstract
In this paper, we propose a novel generative framework which uses Generative Adversarial Networks (GANs) to generate features that provide robustness for object detection on reduced-quality images. The proposed GAN-based Detection of Objects (GAN-DO) framework is not restricted to any particular architecture and can be generalized to several deep neural network (DNN) based architectures. The resulting deep neural network maintains the exact architecture as the selected baseline model without adding to the model parameter complexity or inference speed. We first evaluate the effect of image quality on both object classification and object bounding box regression. We then test the models resulting from our proposed GAN-DO framework, using two state-of-the-art object detection architectures as the baseline models. We also evaluate the effect of the number of re-trained parameters in the generator of GAN-DO on the accuracy of the final trained model. Performance results provided using GAN-DO on object detection datasets establish an improved robustness to varying image quality and a higher mAP compared to the existing approaches.
Charan D. Prakash, Lina J. Karam
IEEE Trans. Image Process.2
2020 Defending Against Universal Attacks Through Selective Feature Regeneration
abstract
Deep neural network (DNN) predictions have been shown to be vulnerable to carefully crafted adversarial perturbations. Specifically, image-agnostic (universal adversarial) perturbations added to any image can fool a target network into making erroneous predictions. Departing from existing defense strategies that work mostly in the image domain, we present a novel defense which operates in the DNN feature domain and effectively defends against such universal perturbations. Our approach identifies pre-trained convolutional features that are most vulnerable to adversarial noise and deploys trainable feature regeneration units which transform these DNN filter activations into resilient features that are robust to universal perturbations. Regenerating only the top 50% adversarially susceptible activations in at most 6 DNN layers and leaving all remaining DNN activations unchanged, we outperform existing defense strategies across different network architectures by more than 10% in restored accuracy. We show that without any additional modification, our defense trained on ImageNet with one type of universal attack examples effectively defends against other types of unseen universal attacks.
Tejas S. Borkar, Felix Heide, Lina J. Karam
CVPR3
2020 Universal Adversarial Attack Via Enhanced Projected Gradient Descent
abstract
It has been shown that there exist small and image-independent perturbations, called universal perturbations, that can fool deep-learning-based classifiers, resulting in a significant decrease in classification accuracy. In this paper, we propose a novel method to compute more effective universal perturbations via enhanced projected gradient descent on targeted classifiers. By maximizing the original loss function of the targeted model, we update the adversarial example with back-propagation and optimize the perturbation by accumulating small updates on perturbed images consecutively. We generate our attack for several modern CNN classifiers using ImageNet and compare the attack performance with other state-of-the-art universal adversarial attack methods. Performance results show that our proposed adversarial attack method can achieve much higher fooling rates as compared to state-of-the-art universal adversarial attack methods and can realize good generalization on cross-model evaluation.
Yingpeng Deng, Lina J. Karam
ICIP2
2019 Encoding and monitoring responsibility sensitive safety rules for automated vehicles in signal temporal logic
abstract
As Automated Vehicles (AV) get ready to hit the public roads unsupervised, many practical questions still remain open. For example, there is no commonly acceptable formal definition of what safe driving is. A formal definition of safe driving can be utilized in developing the vehicle behaviors as well as in certification and legal cases. Toward that goal, the Responsibility-Sensitive Safety (RSS) model was developed as a first step toward formalizing safe driving behavior upon which the broader AV community can expand. In this paper, we demonstrate that the RSS model can be encoded in Signal Temporal Logic (STL). Moreover, using the S-TaLiRo tools, we present a case study of monitoring RSS requirements on selected traffic scenarios from CommonRoad. We conclude that monitoring RSS rules encoded in STL is efficient even in heavy traffic scenarios. One interesting observation is that for the selected traffic data, vehicle parameters and response times, the RSS model violations are not frequent.
Mohammad Hekmatnejad, Shakiba Yaghoubi, Adel Dokhanchi, Heni Ben Amor, Aviral Shrivastava, Lina J. Karam, Georgios Fainekos
MEMOCODE6
2019 Human and DNN Classification Performance on Images With Quality Distortions: A Comparative Study
abstract
Image quality is an important practical challenge that is often overlooked in the design of machine vision systems. Commonly, machine vision systems are trained and tested on high-quality image datasets, yet in practical applications the input images cannot be assumed to be of high quality. Modern deep neural networks (DNNs) have been shown to perform poorly on images affected by blur or noise distortions. In this work, we investigate whether human subjects also perform poorly on distorted stimuli and provide a direct comparison with the performance of DNNs. Specifically, we study the effect of Gaussian blur and additive Gaussian noise on human and DNN classification performance. We perform two experiments: one crowd-sourced experiment with unlimited stimulus display time, and one lab experiment with 100ms display time. In both cases, we found that humans outperform neural networks on distorted stimuli, even when the networks are retrained with distorted data.
Samuel Fuller Dodge, Lina J. Karam
ACM Trans. Appl. Percept.2
2019 DeepCorrect: Correcting DNN Models Against Image Distortions
abstract
In recent years, the widespread use of deep neural networks (DNNs) has facilitated great improvements in performance for computer vision tasks like image classification and object recognition. In most realistic computer vision applications, an input image undergoes some form of image distortion such as blur and additive noise during image acquisition or transmission. Deep networks trained on pristine images perform poorly when tested on such distortions. In this paper, we evaluate the effect of image distortions like Gaussian blur and additive noise on the activations of pre-trained convolutional filters. We propose a metric to identify the most noise susceptible convolutional filters and rank them in order of the highest gain in classification accuracy upon correction. In our proposed approach called DeepCorrect, we apply small stacks of convolutional layers with residual connections at the output of these ranked filters and train them to correct the worst distortion affected filter activations, while leaving the rest of the pre-trained filter outputs in the network unchanged. Performance results show that applying DeepCorrect models for common vision tasks like image classification (ImageNet), object recognition (Caltech-101, Caltech-256), and scene classification (SUN-397), significantly improves the robustness of DNNs against distorted images and outperforms other alternative approaches.
Tejas S. Borkar, Lina J. Karam
IEEE Trans. Image Process.2
2018 Augmented Sparse Representation Classifier for Blurred Face Recognition
abstract
The sparse representation classifier (SRC) has been developed to offer a formulation to the face recognition problem under scene dependent conditions, such as illumination/pose variations, occlusion, and disguise. However, this method has not considered image quality degradations resulting from capture, such as blur, under the same scene variations. In this work, we explore the performance of the well-known face recognition framework SRC in the presence of Gaussian blur in both constrained and unconstrained environments. Finally, we propose an augmented SRC (ASRC) framework to improve the performance of the original SRC in the presence of Gaussian blur, while preserving its robustness to scene dependent variations.
Jinane Mounsef, Lina J. Karam
ICIP2
2018 Multifeature, Sparse-Based Approach for Defects Detection and Classification in Semiconductor Units
abstract
Automated inspection systems play an important role in manufacturing to guarantee higher quality and reduce production costs. In the semiconductor manufacturing industry, assembly and testing processes are getting more complex, resulting in a greater tendency of defects to impact the production process. These defects can cause field failures and can result in customer dissatisfactions and returns. Currently available defect detection and classification systems are customized and hard-wired to the detection of particular classes of defects and cannot deal with new unknown classes of defects. This issue is aggravated by the very small sample size of available anomalies for learning, by the data imbalance problem, since the number of defective samples is significantly much smaller than the number of normal samples, and by the presence of noise. This paper presents a novel multi-feature, sparse-based defect detection and classification approach that uses the stacking concept to enhance the classification accuracy. The stacking-based classifier is augmented with a novel adaptive over/downsampling technique to deal with the data imbalance problem. A new pruning technique is proposed to eliminate bad base learners. Shortage of defective units, similarities within different classes of defects, wide variation within the same defect class, and data imbalance are the basic challenges to deal with. Experimental results on real-world data from Intel show that the proposed approach results in a high classification accuracy as compared with the existing methods.
Bashar Haddad 0001, Lina J. Karam, Jieping Ye, Nital S. Patel, Martin W. Braun
IEEE Trans Autom. Sci. Eng.3
2018 Visual Saliency Prediction Using a Mixture of Deep Neural Networks
abstract
Visual saliency models have recently begun to incorporate deep learning to achieve predictive capacity much greater than previous unsupervised methods. However, most existing models predict saliency without explicit knowledge of global scene semantic information. We propose a model (MxSalNet) that incorporates global scene semantic information in addition to local information gathered by a convolutional neural network. Our model is formulated as a mixture of experts. Each expert network is trained to predict saliency for a set of closely related images. The final saliency map is computed as a weighted mixture of the expert networks' output, with weights determined by a separate gating network. This gating network is guided by global scene information to predict weights. The expert networks and the gating network are trained simultaneously in an end-toend manner. We show that our mixture formulation leads to improvement in performance over an otherwise identical nonmixture model that does not incorporate global scene information. Additionally, we show that our model achieves better performance than several other visual saliency models.
Samuel Fuller Dodge, Lina J. Karam
IEEE Trans. Image Process.2
2018 Quality Robust Mixtures of Deep Neural Networks
abstract
We study deep neural networks for classification of images with quality distortions. Deep network performance on poor quality images can be greatly improved if the network is fine-tuned with distorted data. However, it is difficult for a single fine-tuned network to perform well across multiple distortion types. We propose a mixture of experts based ensemble method, MixQualNet, that is robust to multiple different types of distortions. The "experts" in our model are trained on a particular type of distortion. The output of the model is a weighted sum of the expert models, where the weights are determined by a separate gating network. The gating network is trained to predict weights for a particular distortion type and level. During testing, the network is blind to the distortion level and type, yet can still assign appropriate weights to the expert models. In order to reduce the computational complexity, we introduce weight sharing into the MixQualNet. We utilize the TreeNet weight sharing architecture as well as introduce the Inverted TreeNet architecture. While both weight sharing architectures reduce memory requirements, our proposed Inverted TreeNet also achieves improved accuracy.
Samuel Fuller Dodge, Lina J. Karam
IEEE Trans. Image Process.2
2017 Spatially-Varying Blur Detection Based on Multiscale Fused and Sorted Transform Coefficients of Gradient Magnitudes
abstract
The detection of spatially-varying blur without having any information about the blur type is a challenging task. In this paper, we propose a novel effective approach to address this blur detection problem from a single image without requiring any knowledge about the blur type, level, or camera settings. Our approach computes blur detection maps based on a novel High-frequency multiscale Fusion and Sort Transform (HiFST) of gradient magnitudes. The evaluations of the proposed approach on a diverse set of blurry images with different blur types, levels, and contents demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods qualitatively and quantitatively.
S. Alireza Golestaneh, Lina J. Karam
CVPR2
2017 A Study and Comparison of Human and Deep Learning Recognition Performance under Visual Distortions
abstract
Deep neural networks (DNNs) achieve excellent performance on standard classification tasks. However, under image quality distortions such as blur and noise, classification accuracy becomes poor. In this work, we compare the performance of DNNs with human subjects on distorted images. We show that, although DNNs perform better than or on par with humans on good quality images, DNN performance is still much lower than human performance on distorted images. We additionally find that there is little correlation in errors between DNNs and human subjects. This could be an indication that the internal representation of images are different between DNNs and the human visual system. These comparisons with human performance could be used to guide future development of more robust DNNs.
Samuel Fuller Dodge, Lina J. Karam
ICCCN2
2016 Robust radial distortion correction based on alternate optimization
abstract
A robust radial distortion correction method that requires only a single image of the distorted imaged pattern is presented. The method is robust to the orientation of the imaged pattern and does not require the availability of the ideal reference regularly structured pattern which can be recreated from detected features in the distorted image. Radial distortion parameters, i.e., the Center of Distortion (CoD) and radial distortion coefficients (RDC), are optimized by minimizing corresponding cost functions in an alternate manner. Tests with synthetic and real data are presented in order to illustrate the performance robustness of the proposed algorithm.
Juan Andrade, Lina J. Karam
ICIP2
2016 Visual attention quality database for benchmarking performance evaluation metrics
abstract
With the increased focus on visual attention (VA) in the last decade, a large number of computational visual saliency methods have been developed. These models are evaluated by using performance evaluation metrics that measure how well a predicted map matches eye-tracking data obtained from human observers. Though there are a number of existing performance evaluation metrics, there is no clear consensus on which evaluation metric is the best. This work proposes a subjective study that uses ratings from human observers to evaluate saliency maps computed by existing VA models based on comparing the maps visually with ground-truth maps obtained from eye-tracking data. The subjective ratings are correlated with the scores obtained from existing as well as a proposed objective VA performance evaluation metric using several correlation measures. The correlation results show that the proposed objective VA metric outperforms the existing metrics.
Milind S. Gide, Samuel Fuller Dodge, Lina J. Karam
ICIP3
2016 Reduced-reference synthesized-texture quality assessment based on multi-scale spatial and statistical texture attributes
abstract
In this paper, we propose a reduced-reference (RR) objective quality assessment method that quantifies the perceived quality of synthesized textures. The proposed metric is based on measuring the granularity, regularity, and statistical attributes of the texture image. Furthermore, the proposed RR metric exhibits a significantly low overhead as compared to existing RR metrics by only requiring the transmission of 7 parameters as side information. Performance evaluations on two synthesized texture databases demonstrate that the proposed RR metric outperforms both full-reference (FR) and RR state-of-the-art quality metrics in predicting the perceived visual quality of the synthesized textures.
S. Alireza Golestaneh, Lina J. Karam
ICIP2
2016 Multi-feature sparse-based defect detection and classification in semiconductor units
abstract
Automated inspection systems have been used extensively for high-speed defect detection, gaging and quality control. In the semiconductor manufacturing industry, assembly and testing processes are getting more complex resulting in a greater tendency of defects to impact the production process. Currently available defect detection and classification systems are customized and hard-wired to the detection of particular classes of defects and cannot deal with new unknown classes. Shortage of defective units, similarities within different classes of defects, wide variations within the same defect class, and data imbalance are the basic challenges for this problem. This paper presents a novel stacking-based multi-feature, sparse-based defect detection and classification method that is robust to data imbalance and low number of training samples. Experimental results on real-world data from Intel show that the proposed approach results in a high classification accuracy as compared to existing methods.
Bashar Haddad 0001, Lina J. Karam, Jieping Ye, Nital S. Patel, Martin W. Braun
ICIP2
2016 General chair's welcome
abstract
Presents the introductory welcome message from the conference proceedings. May include the conference officers' congratulations to all involved with the conference event and publication of the proceedings record.
Lina J. Karam
ICIP1
2016 Efficient perceptual-based spatially varying out-of-focus blur detection
abstract
This paper proposes a blur detection algorithm that is capable of detecting and quantifying the level of spatially-varying blur by integrating directional edge spread calculation, Just Noticeable Blur (JNB) and local probability summation. The proposed method generates a blur map indicating the relative amount of perceived local blurriness. We compare the proposed method with six other state-of-the-art blur detection methods. Experimental results show that the proposed method performs the best both visually and quantitatively.
Lina J. Karam
ICIP2
2016 Understanding how image quality affects deep neural networks
abstract
Image quality is an important practical challenge that is often overlooked in the design of machine vision systems. Commonly, machine vision systems are trained and tested on high quality image datasets, yet in practical applications the input images can not be assumed to be of high quality. Recently, deep neural networks have obtained state-of-the-art performance on many machine vision tasks. In this paper we provide an evaluation of 4 state-of-the-art deep neural network models for image classification under quality distortions. We consider five types of quality distortions: blur, noise, contrast, JPEG, and JPEG2000 compression. We show that the existing networks are susceptible to these quality distortions, particularly to blur and noise. These results enable future work in developing deep neural networks that are more invariant to quality distortions.
Samuel Fuller Dodge, Lina J. Karam
QoMEX2
2016 Visual quality assessment of reconstructed background images
abstract
A clean background image is of great importance in multiple applications such as video surveillance, object tracking and context-based video encoding, but acquiring a clean background image in public areas is seldom possible. Many algorithms have been developed to initialize the background from videos and images. This paper presents a database consisting of 13 different scenes that can be used for benchmarking the performance of background initialization algorithms. We also conducted a subjective study on the perceptual quality of background images that are reconstructed using existing background initialization algorithms. The obtained subjective scores are used to evaluate existing image quality metrics and their capability in predicting the perceived quality of reconstructed background images.
Aditee Shrotre, Lina J. Karam
QoMEX2
2016 3D Blur Discrimination
abstract
Blur is an important attribute in the study and modeling of the human visual system. In the blur discrimination experiments, just-noticeable additional blur required to differentiate from the reference blur level is measured. The past studies on blur discrimination have measured the sensitivity of the human visual system to blur using two-dimensional (2D) test patterns. In this study, subjective tests are performed to measure blur discrimination thresholds using stereoscopic 3D test patterns. Specifically, how the binocular disparity affects the blur sensitivity is measured on a passive stereoscopic display. A passive stereoscopic display renders the left and right eye images in a row interleaved format. The subjects have to wear circularly polarized glasses to filter the appropriate images to the left and right eyes. Positive, negative, and zero disparity values are considered in these experiments. A positive disparity value projects the objects behind the display screen, a negative disparity value projects the objects in front of the display screen, and a zero disparity value projects the objects at the display plane. The blur discrimination thresholds are measured for both symmetric and asymmetric stereo viewing cases. In the symmetric viewing case, the same level of additional blur is applied to the left and right eye stimulus. In the asymmetric viewing case, different levels of additional blur are applied to the left and right eye stimuli. The results of this study indicate that, in the symmetric stereo viewing case, binocular disparity does not affect the blur discrimination thresholds for the selected 3D test patterns. As a consequence of these findings, we conclude that the models developed for 2D blur discrimination can be used for 3D blur discrimination. We also show that the Weber model provides a good fit to the blur discrimination threshold measurements for the symmetric stereo viewing case. In the asymmetric viewing case, the blur discrimination thresholds decreased, and the decrease in threshold values is found to be dominated by eye observing the higher blur.
Mahesh Subedar, Lina J. Karam
ACM Trans. Appl. Percept.2
2016 Stereo Vision Based Automated Solder Ball Height and Substrate Coplanarity Inspection
abstract
Solder ball height and substrate coplanarity inspection is essential to the detection of potential connectivity issues in semi-conductor units. Current ball height and substrate coplanarity inspection tools such as laser profiling, fringe projection, and confocal microscopy are expensive, require complicated setup and are slow, which makes them difficult to use in a real-time manufacturing setting. Therefore, a reliable, in-line ball height and substrate coplanarity measurement method is needed for inspecting units undergoing assembly.
Bonnie L. Bennett, Lina J. Karam, Jeffrey S. Pettinato
IEEE Trans Autom. Sci. Eng.3
2016 A Locally Weighted Fixation Density-Based Metric for Assessing the Quality of Visual Saliency Predictions
abstract
With the increased focus on visual attention (VA) in the last decade, a large number of computational visual saliency methods have been developed over the past few years. These models are traditionally evaluated by using performance evaluation metrics that quantify the match between predicted saliency and fixation data obtained from eye-tracking experiments on human observers. Though a considerable number of such metrics have been proposed in the literature, there are notable problems in them. In this paper, we discuss shortcomings in the existing metrics through illustrative examples and propose a new metric that uses local weights based on fixation density, which overcomes these flaws. To compare the performance of our proposed metric at assessing the quality of saliency prediction with other existing metrics, we construct a ground-truth subjective database in which saliency maps obtained from 17 different VA models are evaluated by 16 human observers on a five-point categorical scale in terms of their visual resemblance with corresponding ground-truth fixation density maps obtained from eye-tracking data. The metrics are evaluated by correlating metric scores with the human subjective ratings. The correlation results show that the proposed evaluation metric outperforms all other popular existing metrics. In addition, the constructed database and corresponding subjective ratings provide an insight into which of the existing metrics and future metrics are better at estimating the quality of saliency prediction and can be used as a benchmark.
Milind S. Gide, Lina J. Karam
IEEE Trans. Image Process.2
2016 Reduced-Reference Quality Assessment Based on the Entropy of DWT Coefficients of Locally Weighted Gradient Magnitudes
abstract
Perceptual image quality assessment (IQA) attempts to use computational models to estimate the image quality in accordance with subjective evaluations. Reduced-reference (RR) image quality assessment (IQA) methods make use of partial information or features extracted from the reference image for estimating the quality of distorted images. Finding a balance between the number of RR features and accuracy of the estimated image quality is essential and important in IQA. In this paper we propose a training-free low-cost RRIQA method that requires a very small number of RR features (6 RR features). The proposed RRIQA algorithm is based on the discrete wavelet transform (DWT) of locally weighted gradient magnitudes.We apply human visual system's contrast sensitivity and neighborhood gradient information to weight the gradient magnitudes in a locally adaptive manner. The RR features are computed by measuring the entropy of each DWT subband, for each scale, and pooling the subband entropies along all orientations, resulting in L RR features (one average entropy per scale) for an L-level DWT. Extensive experiments performed on seven large-scale benchmark databases demonstrate that the proposed RRIQA method delivers highly competitive performance as compared to the state-of-the-art RRIQA models as well as full reference ones for both natural and texture images. The MATLAB source code of REDLOG and the evaluation results are publicly available online at https://http://lab.engineering.asu.edu/ivulab/software/redlog/.
S. Alireza Golestaneh, Lina J. Karam
IEEE Trans. Image Process.2
2015 Reduced-reference quality assessment based on the entropy of DNT coefficients of locally weighted gradients
abstract
Reduced-reference (RR) image quality assessment (IQA) methods make use of partial information or features extracted from the reference image for estimating the quality of distorted images. Finding a balance between the number of RR features and image quality estimation accuracy is a difficult task. This paper presents a training-free low cost RRIQA method which requires a very small number of RR features (6 RR features). The proposed RRIQA algorithm is based on the divisive normalization transform (DNT) of locally weighted gradient magnitudes. The weighting of the gradient magnitudes is performed in a locally adaptive manner based on the human visual system's contrast sensitivity and neighborhood gradient information. The RR features are obtained by computing the entropy of each DNT subband and, for each scale, averaging the subband entropies along all orientations, resulting in L RR features (one average entropy per scale) for an L-level DNT. Performance evaluations on four large-scale benchmark databases demonstrates that the proposed RRIQA method delivers highly competitive performance as compared to the state-of-the-art RRIQA models as well as full reference ones.
S. Alireza Golestaneh, Lina J. Karam
ICIP2
2015 A no reference texture granularity index and application to visual media compression
abstract
Texture granularity is an important attribute to quantify the level of details present in the image. This work presents a no-reference texture granularity index and shows using subjective experiments that the proposed granularity index correlates well with the perceived granularity of textures. In addition, a subjective study is conducted to assess the effect of compression on textures with varying degrees of granularity. It is shown that a measure of texture granularity can predict the compression quality.
Mahesh Subedar, Lina J. Karam
ICIP2
2015 A No-Reference Texture Regularity Metric Based on Visual Saliency
abstract
This paper presents a no-reference perceptual metric that quantifies the degree of perceived regularity in textures. The metric is based on the similarity of visual attention (VA) of the textural primitives and the periodic spatial distribution of foveated fixation regions throughout the image. A ground-truth eye-tracking database for textures is also generated as part of this paper and is used to evaluate the performance of the most popular VA models. Using the saliency map generated by the best VA model, the proposed texture regularity metric is computed. It is shown through subjective testing that the proposed metric has a strong correlation with the mean opinion score for the perceived regularity of textures. The proposed texture regularity metric can be used to improve the quality and performance of many image processing applications like texture synthesis, texture compression, and content-based image retrieval.
Srenivas Varadarajan, Lina J. Karam
IEEE Trans. Image Process.2
2014 Change detection on SAR images using divisive normalization-based image representation
abstract
In the context of multi-temporal synthetic aperture radar (SAR) images for earth monitoring applications, one critical issue is the detection of changes occurring after a natural or an-thropic disaster. In this paper, we propose a new similarity measure for automatic change detection based on a divisive normalization image representation. The divisive normalization transform (DNT) has been recognized as a successful methodology to model the perceptual sensitivity of biological vision and a useful image representation that significantly reduces statistical dependence of natural images. In this work, we exploit the fact that the histogram of DNT coefficients within wavelet subbands can often be well fitted with a zero-mean Gaussian density function, which is a one-parameter function that allows efficient change detection of SAR images. The proposed change detector is compared to other recent modelbased approaches. Tests on real data show that our detector outperforms previously suggested methods in terms of the rate of false alarm rate and the total error rate.
Qian Xu 0009, Lina J. Karam
ICASSP2
2014 A reduced-reference perceptual quality metric for texture synthesis
abstract
This paper presents a reduced-reference quality metric that quantifies the perceptual quality of the synthesized textures. The metric is based on the change in perceived regularity between the original and the synthesized textures. The perceived regularity is quantified through a modified texture regularity metric based on visual attention. It is shown through subjective testing that the proposed metric has a strong correlation with the Mean Opinion Score for the fidelity of synthesized textures and outperforms the state-of-the-art full-reference quality metrics.
Srenivas Varadarajan, Lina J. Karam
ICIP2
2014 A Distributed Canny Edge Detector: Algorithm and FPGA Implementation
abstract
The Canny edge detector is one of the most widely used edge detection algorithms due to its superior performance. Unfortunately, not only is it computationally more intensive as compared with other edge detection algorithms, but it also has a higher latency because it is based on frame-level statistics. In this paper, we propose a mechanism to implement the Canny algorithm at the block level without any loss in edge detection performance compared with the original frame-level Canny algorithm. Directly applying the original Canny algorithm at the block-level leads to excessive edges in smooth regions and to loss of significant edges in high-detailed regions since the original Canny computes the high and low thresholds based on the frame-level statistics. To solve this problem, we present a distributed Canny edge detection algorithm that adaptively computes the edge detection thresholds based on the block type and the local distribution of the gradients in the image block. In addition, the new algorithm uses a nonuniform gradient magnitude histogram to compute block-based hysteresis thresholds. The resulting block-based algorithm has a significantly reduced latency and can be easily integrated with other block-based image codecs. It is capable of supporting fast edge detection of images and videos with high resolutions, including full-HD since the latency is now a function of the block size instead of the frame size. In addition, quantitative conformance evaluations and subjective tests show that the edge detection performance of the proposed algorithm is better than the original frame-based algorithm, especially when noise is present in the images. Finally, this algorithm is implemented using a 32 computing engine architecture and is synthesized on the Xilinx Virtex-5 FPGA. The synthesized architecture takes only 0.721 ms (including the SRAM READ/WRITE time and the computation time) to detect edges of 512 × 512 images in the USC SIPI database when clocked at 100 MHz and is faster than existing FPGA and GPU implementations.
Qian Xu 0009, Srenivas Varadarajan, Chaitali Chakrabarti, Lina J. Karam
IEEE Trans. Image Process.4
2013 A no-reference perceptual texture regularity metric
abstract
This paper presents a no reference perceptual metric that quantifies the degree of regularity in textures. The metric is based on the probability of visual attention at each pixel of the texture image, similarity of visual attention of the textural primitives and the periodic spatial distribution of foveated fixation regions throughout the image. It is shown through subjective testing that the proposed metric has a strong correlation with the Mean Opinion Score for the regularity of textures.
Srenivas Varadarajan, Lina J. Karam
ICASSP2
2013 Change detection on SAR images by a parametric estimation of the KL-divergence between Gaussian Mixture Models
abstract
In the context of multi-temporal synthetic aperture radar (SAR) images for earth monitoring applications, one critical issue is the detection of changes occurring after a natural or anthropic disaster. In this paper, we propose a new similarity measure for automatic change detection using a pair of SAR images acquired at different dates. This measure is based on the evolution of the local statistics of the image between two dates. The local statistics are modeled as a Gaussian Mixture Model (GMM), which approximates the probability density function in the neighborhood of each pixel in the image. The degree of evolution of the local statistics is measured using the Kullback-Leibler (KL) divergence. One analytical expression for approximating the KL divergence between GMMs is given and is compared with the Monte Carlo sampling method. The proposed change detector is compared to the classical mean ratio detector and also to other recent model-based approaches. Tests on the real data show that our detector outperforms previously suggested methods in terms of the rate of missed detections and the total error rates.
Qian Xu 0009, Lina J. Karam
ICASSP2
2012 Attentive gesture recognition
abstract
This paper presents a novel method for static gesture recognition based on visual attention. Our proposed method makes use of a visual attention model to automatically select points that correspond to fixation points of the human eye. Gesture recognition is then performed using the determined visual attention fixation points. For this purpose, shape context descriptors are used to compare the sparse fixation points of gestures for classification. Simulation results are presented in order to illustrate the performance of the proposed perceptual-based attentive gesture recognition method. The proposed method not only helps in the development of more natural user-centric interactive interfaces but is also able to achieve a 96.42% classification accuracy on the Triesch database of hand postures, which is superior to other methods presented in the literature.
Samuel Fuller Dodge, Lina J. Karam
ICIP2
2012 Camera calibration using adaptive segmentation and ellipse fitting for localizing control points
abstract
This paper presents a novel camera calibration method using a circular calibration pattern. The disadvantages and issues with existing state-of-the-art methods are discussed and are overcome in this work. In the proposed iterative method, the captured images of the circular pattern are undistorted and projected onto a fronto parallel plane. The control points are localized in this fronto parallel plane using a novel approach for more accurately localizing the control points in the images based on adaptive segmentation and ellipse fitting. The localized control points are then projected to their original planes using estimated camera calibration parameters that are refined by minimizing the reprojection error. Simulation results are presented to illustrate the performance of the proposed scheme. These results show that the proposed method reduces the error by up to 57% as compared to the state-of-the-art for high-resolution images, and that the proposed scheme is more robust to blur in the imaged calibration pattern.
Charan D. Prakash, Lina J. Karam
ICIP2
2011 Efficient Super-Resolution driven by saliency selectivity
abstract
This paper presents a low-complexity saliency detector targeted towards efficient selective Super-Resolution (SR). As a result, an improved efficient ATtentive-SELective Perceptual (AT-SELP) framework is presented. The proposed AT-SELP scheme results in a reduced computational complexity for iterative SR algorithms without any perceptible loss in the desired enhanced image/video quality. A perceptually significant set of active pixels is selected for processing by the SR algorithm based on a local contrast sensitivity threshold model and the proposed low complexity saliency detector. Simulation results show that the proposed AT-SELP scheme results in a 15-40% reduction in computational complexity over an efficient Selective Perceptual (SELP) SR scheme without degradation in the visual quality.
Nabil G. Sadaka, Lina J. Karam
ICIP2
2011 Automated Detection and Classification of Non-Wet Solder Joints
abstract
Non-wet solder joints in processor sockets are causing mother board failures. These board failures can escape to customers resulting in returns and dissatisfaction. The current process to identify these non-wets is to use a 2D or advanced X-ray tool with multidimension capability to image solder joints in processor sockets. The images are then examined by an operator who determines if each individual joint is good or bad. There can be an average of 150 images for an operator to examine for each socket. Each image contains more than 30 joints. These factors make the inspection process time consuming and the output variable depending on the skill and alertness of the operator. This paper presents an automatic defect identification and classification system for the detection of non-wet solder joints. The main components of the proposed system consist of region of interest (ROI) segmentation, feature extraction, reference-free classification, and automatic mapping. The ROI segmentation process is a noise-resilient segmentation method for the joint area. The centroids of the segmented joints (ROIs) are used as feature parameters to detect the suspect joints. The proposed reference-free classification can detect defective joints in the considered images with high accuracy without the need for training data or reference images. An automatic mapping procedure which maps the positions of all joints to a known Master Ball Grid Array file is used to get the precise label and location of the suspect joint for display to the operator and collection of non-wet statistics. The accuracy of the proposed system was determined to be 95.8% based on the examination of 56 sockets (76 496 joints). The false alarm rate is 1.1%. In comparison, the detection rate of a currently available advanced X-ray tool with multidimension capability is in the range of 43% to 75%. The proposed method reduces the operator effort to examine individual images by 89.6% (from looking at 154 images to 16 images) by presenting only images with suspect joints for inspection. When non-wet joints are missed, the presented system has been shown to identify the neighboring joints. This fact provides the operator with the capability to make 100% detection of all non-wets when utilizing a user interface that highlights the suspect joint area. The system works with a 2D X-ray imaging device, which saves cost over more expensive advanced X-ray tools with multidimension capability. The proposed scheme is relatively inexpensive to implement, easy to set up and can work with a variety of 2D X-ray tools.
Asaad F. Said, Bonnie L. Bennett, Lina J. Karam, Jeffrey S. Pettinato
IEEE Trans Autom. Sci. Eng.3
2011 An Efficient Selective Perceptual-Based Super-Resolution Estimator
abstract
In this paper, a selective perceptual-based (SELP) framework is presented to reduce the complexity of popular super-resolution (SR) algorithms while maintaining the desired quality of the enhanced images/video. A perceptual human visual system model is proposed to compute local contrast sensitivity thresholds. The obtained thresholds are used to select which pixels are super-resolved based on the perceived visibility of local edges. Processing only a set of perceptually significant pixels reduces significantly the computational complexity of SR algorithms without losing the achievable visual quality. The proposed SELP framework is integrated into a maximum-a posteriori-based SR algorithm as well as a fast two-stage fusion-restoration SR estimator. Simulation results show a significant reduction on average in computational complexity with comparable signal-to-noise ratio gains and visual quality.
Lina J. Karam, Nabil G. Sadaka, Rony Ferzli, Zoran A. Ivanovski
IEEE Trans. Image Process.1
2011 A No-Reference Image Blur Metric Based on the Cumulative Probability of Blur Detection (CPBD)
abstract
This paper presents a no-reference image blur metric that is based on the study of human blur perception for varying contrast values. The metric utilizes a probabilistic model to estimate the probability of detecting blur at each edge in the image, and then the information is pooled by computing the cumulative probability of blur detection (CPBD). The performance of the metric is demonstrated by comparing it with existing no-reference sharpness/blurriness metrics for various publicly available image databases.
Niranjan D. Narvekar, Lina J. Karam
IEEE Trans. Image Process.2
2010 Overview and Traffic Characterization of Coarse-Grain Quality Scalable (CGS) H.264 SVC Encoded Video
abstract
The scalable video coding extension (SVC) of the H.264/AVC standard is widely considered for IPTV. SVC supports a variety of scalability modes, including temporal, spatial as well as coarse-grain and medium-grain quality scalabilities. In this paper, we first give an overview of coarse-grain quality scalability (CGS). We generate traces of CGS encodings of long CIF resolution video sequences; the traces provide a simple yet effective characterization of CGS encoded video for performance evaluation of video transport systems, including IPTV systems. We conduct a detailed statistical analysis of the CGS video traces. We compare the bit rate-distortion (RD) and the bit rate variability-distortion (VD) performances of scalable CGS encodings with those of non-scalable SVC single layer encodings. We thus quantify the tradeoff between the rate adaptability afforded by CGS encoding and the cost in terms of RD efficiency compared to non-scalable single-layer video.
Akshay Pulipaka, Patrick Seeling, Martin Reisslein, Lina J. Karam
CCNC4
2010 Robust automatic void detection in solder balls
abstract
Voids in solder balls can cause board failures. The detection and assessment of voids in solder balls can help in reducing board yield issues caused by incorrect scrapping and rework. X-ray imaging machines make voids visible to the operator for manual inspection. Some existing x-ray inspection systems have void detection algorithms that require the use of intensive manual tuning operations that are time consuming, and inaccurate due to the inability to examine balls overshadowed with other components. In this paper, a robust automatic void detection algorithm is proposed. The proposed method is able to detect voids with different sizes inside the solder balls, including the ones that are overshadowed by board components and under different brightness conditions. Results show that the proposed method achieves a correlation squared in the range of 91% to 97% with ground truth data from a 3D x-ray scan. The proposed algorithm is fully automated and benefits the manufacturing process by reducing operator effort and by providing a cost effective solution to improve output quality.
Asaad F. Said, Bonnie L. Bennett, Lina J. Karam, Jeffrey S. Pettinato
ICASSP3
2010 A distributed psycho-visually motivated Canny edge detector
abstract
This paper proposes a distributed Canny edge detection algorithm which can be mapped onto multi-core architectures for high throughput applications. In contrast to the conventional Canny edge detection algorithm which makes use of the global image gradient histogram to determine the threshold for edge detection, the proposed algorithm adaptively computes the edge detection threshold based on the local distribution of the gradients in the considered image block. The efficacy of the distributed Canny in detecting psycho-visually important edges is validated using a visual sharpness metric. The proposed distributed Canny edge detection algorithm has the capacity to scale up the throughput adaptively, based on the number of computing engines. The algorithm achieves about 72 times speed up for a 16-core architecture, without any change in performance. Furthermore, the internal memory requirements are significantly reduced especially for smaller block sizes. For instance, if a 512×512 image is processed in 64×64 blocks using the proposed scheme, the memory is reduced by a factor of 70 as compared to the original Canny edge detector.
Srenivas Varadarajan, Chaitali Chakrabarti, Lina J. Karam, Judit Martinez Bauza
ICASSP3
2010 Super-resolution using a Wavelet-based Adaptive Wiener Filter
abstract
In this paper, a Wavelet-based Adaptive Wiener Filter (WAWF) super-resolution (SR) algorithm is presented. A redundant discrete dyadic wavelet transform (DDWT) is applied to the input sequence to classify the LR frames into subbands of similar contextual information. The similar subbands from each LR frame are registered using subpixel motion and merged on the same HR grid to form HR subbands of similar contextual and statistical information. Then an adaptive Wiener filter approach is applied locally to each subband in order to interpolate the missing information on the HR gird. The weights of the filter are designed using a parametric circularly symmetric statistical model that adapts to the statistics and the spatial proximity of the neighboring wavelet coefficients. Simulation results show that the proposed WAWF SR algorithm results in a superior performance as compared to existing recent SR schemes.
Nabil G. Sadaka, Lina J. Karam
ICIP2
2010 BLAST-DVC: BitpLAne SelecTive distributed video coding
Wei-Jung Chien, Lina J. Karam
Multim. Tools Appl.2
2009 Background recovery from video sequences using motion parameters
abstract
This paper presents a novel scheme for extracting a still background occluded by a number of foreground objects, moving in different directions and velocities in a video sequence, such that every background pixel is exposed in at least one of the frames. Each identified foreground object is decomposed into blocks. The proposed scheme is able to efficiently estimate, for each foreground block, a source frame from which the occluded background pixels can be extracted. The pixels of the identified source frames are used to populate the co-located occluded pixels in the initial frame. The efficacy and the simplicity of the algorithm lie in its capacity to recover the background directly from the estimated source frames instead of performing a foreground-background classification for every frame. The proposed algorithm is robust to variations in lighting and is effective in removing both rigid and deformable foreground objects. Simulation results are presented to illustrate the performance of the proposed scheme.
Srenivas Varadarajan, Lina J. Karam, Dinei A. F. Florêncio
ICASSP2
2009 AQT-DVC: Adaptive Quantization for transform-domain distributed video coding
abstract
This paper presents a rate-distortion based adaptive quantization (AQT) scheme for transform-domain distributed video coding (DVC). In the proposed DVC system, the Wyner-Ziv frame is divided into partitions and is adaptively quantized in the transform domain based on estimated local rate-distortion (R-D) characteristics. The R-D characteristics are estimated at the decoder using a correlation model between the original source information and the side information. Rate-distortion performance results and comparisons with existing DVC schemes and with H.264 are presented.
Wei-Jung Chien, Lina J. Karam
ICIP2
2009 An H.264/SVC memory architecture supporting spatial and course-grained quality scalabilities
abstract
The standardized scalable video coding (SVC) extension of H.264/AVC achieves significant improvements in coding efficiency relative to the scalable profiles of prior video coding standards, but its computational complexity and memory access requirements make the design of a low power hardware architecture a challenging task. This paper presents an SVC decoder architecture supporting spatial and coarse-grained quality scalability. The architecture optimizes the size of the on-chip memory and reduces the power-consuming and time-intensive external memory accesses.
Niranjan D. Narvekar, Bharatan Konnanath, Shalin M. Mehta, Santosh Chintalapati, Ismail AlKamal, Chaitali Chakrabarti, Lina J. Karam
ICIP7
2009 Efficient perceptual attentive super-resolution
abstract
An efficient perceptually attentive (PA) super-resolution method is proposed to significantly reduce the computational complexity of iterative super-resolution algorithms without loss of the desired perceptual quality. A perceptually significant constrained set of active pixels is selected for processing by the SR algorithm based on a just noticeable distortion threshold model. These selected active pixels are further reduced by using saliency information that is determined by a visual attention model. Furthermore, the active pixels lying in the attended regions are processed at a higher accuracy by the SR method relative to pixels in other regions. Simulation results are presented to show the preserved desired visual quality and a 30-40% reduction in complexity over a highly efficient selective perceptual fast two-step (SELP-FTS) scheme.
Nabil G. Sadaka, Lina J. Karam
ICIP2
2009 Transform-domain distributed video coding with rate-distortion-based adaptive quantisation
abstract
This study presents a transform-domain distributed video coding (DVC) system with a rate–distortion (R–D)-based Adaptive QuanTisation (AQT) scheme. In the proposed system, the transform-domain Wyner–Ziv frame is divided into partitions and is adaptively quantised based on estimated local R–D characteristics for each partition. The R–D estimation is performed based on a correlation model between the original source information and the side information and can be applied at the decoder without adding complexity to the encoder. Coding results and comparisons with existing DVC schemes and with H.264/AVC interframe and intraframe coding are presented to illustrate the performance of the proposed system.
Wei-Jung Chien, Lina J. Karam
IET Image Process.2
2009 A No-Reference Objective Image Sharpness Metric Based on the Notion of Just Noticeable Blur (JNB)
abstract
This work presents a perceptual-based no-reference objective image sharpness/blurriness metric by integrating the concept of just noticeable blur into a probability summation model. Unlike existing objective no-reference image sharpness/blurriness metrics, the proposed metric is able to predict the relative amount of blurriness in images with different content. Results are provided to illustrate the performance of the proposed perceptual-based sharpness metric. These results show that the proposed sharpness metric correlates well with the perceived sharpness being able to predict with high accuracy the relative amount of blurriness in images with different content.
Rony Ferzli, Lina J. Karam
IEEE Trans. Image Process.2
2008 Rate-distortion based selective decoding for pixel-domain distributed video coding
abstract
This paper presents a rate-distortion based selective decoding for pixel-domain distributed video compression (DVC) system. In the proposed system, the Wyner-Ziv frames are divided into several sub-images. Each of these sub-images is encoded and decoded independently. At the decoder, the rate-distortion ratio is estimated for each bitplane of the sub-images. Only the bitplanes with high rate-distortion ratios are Turbo decoded. A minimum-distance symbol reconstruction is proposed to estimate the rest of the bitplanes by using the side information and Turbo-decoded bitplanes. More accurate placement of the parity bits results in an improved system performance, especially for video sequences with a relatively large static background. Coding results and comparison with existing DVC schemes and with H.264 interframe coding are presented to illustrate the performance of the proposed system.
Wei-Jung Chien, Lina J. Karam
ICIP2
2008 An efficient, SELective, Perceptual-based super-resolution estimator
abstract
In this paper, a selective perceptual-based (SELP) scheme is presented to reduce the complexity of popular super-resolution (SR) algorithms while maintaining the desired quality of the enhanced images/video. A perceptual Human Visual System (HVS) model is proposed to compute the contrast sensitivity threshold for a given background intensity. The obtained thresholds are used to select which pixels are super-resolved based on the perceived visibility of local edges. This is accomplished by estimating the contrast sensitivity threshold locally over a block. Next, the absolute difference between each pixel and its neighbors is computed and compared to the threshold upon which a decision is made to include the pixel in the SR estimator for the next iteration or not. The perceptual model is integrated into a MAP-based SR algorithm as well as a fast ML estimator. Simulation results show up to 47% reduction on average in computational complexity with comparable SNR gains and visual quality.
Rony Ferzli, Zoran A. Ivanovski, Lina J. Karam
ICIP3
2008 A no-reference perceptual image sharpness metric based on saliency-weighted foveal pooling
abstract
A no-reference perceptual sharpness quality metric, inspired by visual attention information, is presented for a better simulation of the Human Visual System (HVS) response to blur distortions. Saliency information about a scene is used to accentuate blur distortions around edges present in conspicuous areas and attenuate those distortions present in the rest of the image. Simulation results are presented to illustrate the performance of the proposed metric.
Nabil G. Sadaka, Lina J. Karam, Rony Ferzli, Glen P. Abousleman
ICIP2
2008 An improved perception-based no-reference objective image sharpness metric using iterative edge refinement
abstract
The computation of existing sharpness/blurriness objective metrics involves measuring the spread of edge pixels in blurred images. However, in blurred images, many edges might go undetected causing the metrics to become inaccurate. In these scenarios, proper recovery of edge pixels can lead to a better correlation between the perceived sharpness and the sharpness metric. This paper presents an iterative edge refinement algorithm. The proposed edge refinement scheme is integrated into a perceptual- based no-reference sharpness metric resulting in an increased correlation with the perceived sharpness and, thus, in an increased sharpness/blurriness prediction accuracy. Results are presented to illustrate the performance of the proposed scheme and metric.
Srenivas Varadarajan, Lina J. Karam
ICIP2
2007 Suppression of Mosquito Noise by Recursive Epsilon-Filters
abstract
This work addresses the problem of mosquito noise (MN) reduction in compressed video sequences. A compression-blind approach is adopted; the advantage of such an approach is that it is independent of the particular compressor used and of its particular settings. A recursive filtering scheme is presented. It is shown how the filtering parameter ϵ can be adaptively selected to maximize the denoising performance by minimizing the number of outlier pixels in the filter's support. Simulation results show that the proposed blind MN-denoising scheme outperforms existing MN-denoising methods.
Houssam Abbas, Lina J. Karam
ICASSP (1)2
2007 Block-Adaptive Wyner-Ziv Coding for Transform-Domain Distributed Video Coding
abstract
This paper presents a transform-domain distributed video compression (DVC) system with block-adaptive Wyner-Ziv coding. In the proposed system, the source symbols are reformatted into several blocks. The parity bits are requested only for the blocks with decoding errors. More accurate placement of the parity bits results in an improved system performance that is much closer to the performance of H.263 interframe coding as compared to existing DCT-based DVC schemes. Coding results and comparison with existing DVC schemes and with H.263 are presented to illustrate the performance of the proposed system.
Wei-Jung Chien, Lina J. Karam, Glen P. Abousleman
ICASSP (1)2
2007 A No-Reference Objective Image Sharpness Metric Based on Just-Noticeable Blur and Probability Summation
abstract
This work presents a perceptual-based no-reference objective image sharpness/blurriness metric by integrating the concept of just noticeable blur (JNB) into a probability summation model. Unlike existing objective no-reference image sharpness/blurriness metrics, the proposed metric is able to predict the relative amount of blurriness in images with different content. Results are provided to illustrate the performance of the proposed perceptual-based sharpness metric. These results show that the proposed sharpness metric correlates well with the perceived sharpness.
Rony Ferzli, Lina J. Karam
ICIP (3)2
2007 Selective Error Detection for Error-Resilient Wavelet-Based Image Coding
abstract
This paper introduces the concept of a similarity check function for error-resilient multimedia data transmission. The proposed similarity check function provides information about the effects of corrupted data on the quality of the reconstructed image. The degree of data corruption is measured by the similarity check function at the receiver, without explicit knowledge of the original source data. The design of a perceptual similarity check function is presented for wavelet-based coders such as the JPEG2000 standard, and used with a proposed "progressive similarity-based ARQ" (ProS-ARQ) scheme to significantly decrease the retransmission rate of corrupted data while maintaining very good visual quality of images transmitted over noisy channels. Simulation results with JPEG2000-coded images transmitted over the Binary Symmetric Channel, show that the proposed ProS-ARQ scheme significantly reduces the number of retransmissions as compared to conventional ARQ-based schemes. The presented results also show that, for the same number of retransmitted data packets, the proposed ProS-ARQ scheme can achieve significantly higher PSNR and better visual quality as compared to the selective-repeat ARQ scheme.
Lina J. Karam, Tuyet-Trang Lam
IEEE Trans. Image Process.1
2006 Distributed Video Coding With Lossy Side Information
abstract
The paper presents a distributed video compression (DVC) system with improved rate-distortion performance that is much closer to the performance of H.263 interframe coding as compared to existing DCT-based DVC schemes. The system performance, in terms of compression rate and video quality, is affected by the difference between the source information and the generated side information at the decoder. To improve the accuracy of the side information, a modified three-dimensional recursive search block matching algorithm is proposed. The performance of the proposed DVC system is also investigated when the side information is estimated from lossy video frames that are compressed at different bit rates. Coding results and comparison with existing DVC schemes and with H.263 are presented to illustrate the performance of the proposed system.
Wei-Jung Chien, Lina J. Karam, Glen P. Abousleman
ICASSP (2)2
2006 DNA-Residual: A DNA Compression Algorithm using Forward Linear Prediction
abstract
This paper presents an efficient lossless DNA compression algorithm, DNA-residual, that significantly decreases the average bit-rate required to losslessly code correlated DNA sequences. The algorithm can be divided into two parts: modeling and coding. The modeling part consists of mapping the DNA bases into a binary representation and, then, a forward linear prediction filter is used to predict the current input from the previous ones. The prediction error is then transformed into a binary error sequence that is coded using an adaptive binary arithmetic coder. Compared to state-of-the-art compressors using benchmark DNA sequences, the proposed algorithm reveals a significantly higher compression ratio whenever correlation between bases is high
Rony Ferzli, Lina J. Karam
ICASSP (2)2
2006 Selective Error Detection for Error-Resilient JPEG2000 Coding
abstract
This paper presents a selective error detection (SED) scheme that is based on a similarity check function designed for the JPEG2000 compression algorithm. The presented scheme takes advantage of the error-resilience properties of JPEG-2000 to significantly reduce the number of corrupted data packets that need to be retransmitted in a typical ARQ scheme. The degree of corruption induced by a particular codeblock is measured by a similarity check function at the receiver, without using any explicit knowledge regarding the original source data. Accordingly, if the data is found to be badly corrupted based on a specified similarity criteria, it is considered unusable and retransmitted. On the other hand, if the corrupted data does not significantly affect the quality of the decoded image, it is simply processed by the decoder without retransmission. Simulation results comparing the proposed selective retransmission scheme to full retransmission of corrupted packets are provided to illustrate the performance of the proposed method.
Tuyet-Trang Lam, Lina J. Karam, Glen P. Abousleman
ICASSP (2)2
2006 Human Visual System Based No-Reference Objective Image Sharpness Metric
abstract
This work is motivated by the fact that existing no-reference sharpness metrics fail in predicting the correct amount of blurriness in images with different contexts. This paper presents a no-reference objective sharpness metric that can be applied to images with different contexts. The metric combines a human visual system (HVS)-based sharpness perception model as well as a local features extractor resulting in a content-invariant metric. The proposed HVS-based sharpness perception model is derived from performed subjective tests. Simulation results and comparison with existing no-reference sharpness metrics show that the proposed HVS-based sharpness metric correlates well with the perceived sharpness and is able to correctly predict the relative amount of blurriness in images with different contexts.
Rony Ferzli, Lina J. Karam
ICIP2
2006 Video Texture and Motion based Modeling of Rate Variability-Distortion (VD) Curves of I, P, and B Frames
abstract
We examine the bit rate variability-distortion (VD) curve of I, P, and B frames of MPEG-4 VBR encoded video sequences. We show that the concave VD curve shape at high compression ratios or large quantization scales, is influenced by both the texture and the motion information. We use linear and quadratic models for the texture and motion bits statistics and devise accurate VD curve models. The model parameters are obtained from statistics that are estimated from two encodings. This work extends our previous work on modeling the VD curve, which has applications for optimal statistical multiplexing of VBR streaming video
Geert Van der Auwera, Martin Reisslein, Lina J. Karam
ICME3
2006 A Captcha Based on the Human Visual Systems Masking Characteristics
abstract
In this paper, a CAPTCHA is presented based on the masking characteristics of the human visual system (HVS). Knowing that noise can be masked by high activity regions and showing that edges can be masked by noise for a human observer while still being detected by machines, the suggested CAPTCHA is composed of English alphabets that are picked randomly and written with a combination of texture and edges with added noise such as to deceive the machine by randomly changing the visibility of characters for humans. The proposed CAPTCHA is highly legible and robust to brute-force attacks and sophisticated object character recognition (OCR) segmentation algorithms
Rony Ferzli, Rida A. Bazzi, Lina J. Karam
ICME3
2006 Distributed video coding with 3D recursive search block matching
abstract
This paper presents a novel video compression system with a low complexity encoder. Based on the Slepian-Wolf theorem and the Wyner-Ziv theorem, a video is intraframe coded and interframe decoded in this system. The system performance, the compression rate and the video quality, is affected by the difference between the source information and the side information. To increase the accuracy of the side information, a modified three dimensional recursive search block matching is proposed. The resulting motion vectors preserve the true motion of objects because of the spatial and temporal consistency. The results obtained show a significant improvement over full search block matching algorithm in the term of the rate-distortion measure. The different quantizer design and puncturing table reduce the necessary parity bits for the decoding, and also increase the PSNR. The system performance is also much closer to H.263 interframe coding.
Wei-Jung Chien, Lina J. Karam, Glen P. Abousleman
ISCAS2
2006 JPEG2000 Encoding With Perceptual Distortion Control
abstract
In this paper, a new encoding approach is proposed to control the JPEG2000 encoding in order to reach a desired perceptual quality. The new method is based on a vision model that incorporates various masking effects of human visual perception and a perceptual distortion metric that takes spatial and spectral summation of individual quantization errors into account. Compared with the conventional rate-based distortion minimization JPEG2000 encoding, the new method provides a way to generate consistent quality images at a lower bit rate.
Lina J. Karam, Andrew B. Watson
IEEE Trans. Image Process.2
2005 A Motion-Augmented Super-Resolution Scheme for Very Low-Bit-Rate Video Enhancement
abstract
High compression ratios introduce visible compression artifacts and blurring. To enhance the quality of the compressed video, a super-resolution (SR) scheme is employed. However, the effectiveness of the SR process depends heavily on the amount of subpixel motion that exists between frames in a video sequence, and on the accuracy of the motion vectors (MV). To increase the effectiveness of the SR process, we introduce the concept of added motion prior to video compression. The parameters of this added motion are sent to the receiver with the compressed bit-stream. The proposed SR scheme also exploits existing motion between frames, by estimating MV for each pixel using the original (uncompressed) video. However, since transmitting the pixel-based MV is not practical due to the resulting significant increase in side information, the proposed scheme selects and sends only a very small set of MV based upon their significance in the super-resolution enhancement. The results obtained show a significant improvement in the visual quality of the video sequences, as well as improvement in PSNR.
Zoran A. Ivanovski, Lina J. Karam, Glen P. Abousleman
ICASSP (2)2
2005 Reduced-Delay Selective ARQ for Low Bit-Rate Image and Multimedia Data Transmission
abstract
The paper presents a reduced-delay selective-ARQ scheme based on a similarity check function that measures the degree of corruption in transmitted multimedia data packets. Accordingly, if the packet is found to be badly corrupted based on a specified similarity criterion, it can be considered as a lost packet and retransmitted. The degree of corruption contributed by a particular packet is measured by the proposed similarity check function at the receiver without explicit knowledge of the original source data. Simulation results and comparisons with a conventional ARQ scheme are provided to illustrate the performance of the proposed method.
Tuyet-Trang Lam, Lina J. Karam, Rida A. Bazzi, Glen P. Abousleman
ICASSP (2)2
2005 No-reference objective wavelet based noise immune image sharpness metric
abstract
This paper focuses on no-reference image sharpness/blurriness metrics due to their importance in image, video, and biomedical applications. Simulation results show that existing no-reference objective image sharpness metrics fail to predict correctly the sharpness of images in the presence of noise. A noise-immune wavelet-based sharpness metric is proposed based on the Lipschitz regularity for differentiating between edges and noise singularities. Comparison results reveal the superiority of the proposed method when dealing with a moderate noisy environment.
Rony Ferzli, Lina J. Karam
ICIP (1)2
2005 Mutual information-based analysis of JPEG2000 contexts
abstract
Context-based arithmetic coding has been widely adopted in image and video compression and is a key component of the new JPEG2000 image compression standard. In this paper, the contexts used in JPEG2000 are analyzed using the mutual information, which is closely related to the compression performance. We first show that, when combining the contexts, the mutual information between the contexts and the encoded data will decrease unless the conditional probability distributions of the combined contexts are the same. Given I, the initial number of contexts, and F, the final desired number of contexts, there are S(I, F) possible context classification schemes where S(I, F) is called the Stirling number of the second kind. The optimal classification scheme is the one that gives the maximum mutual information. Instead of using an exhaustive search, the optimal classification scheme can be obtained through a modified generalized Lloyd algorithm with the relative entropy as the distortion metric. For binary arithmetic coding, the search complexity can be reduced by using dynamic programming. Our experimental results show that the JPEG2000 contexts capture the correlations among the wavelet coefficients very well. At the same time, the number of contexts used as part of the standard can be reduced without loss in the coding performance.
Lina J. Karam
IEEE Trans. Image Process.2
2004 An embedded scaling-based arbitrary shape region-of-interest coding method for JPEG2000
abstract
The shape information of objects is perceptually very significant and can aid in the recognition of objects. Depending upon the bit budget, Internet browsing applications can take advantage of an arbitrary-shape region-of-interest (ROI) coding method by sending only the shape information first, and then progressively transmitting the ROI texture, followed by the less important non-ROI texture information. Unfortunately, the existing ROI coding methods, including the JPEG2000-based methods, do not support separate coding and decoding of the shape information. In particular, the maxshift and the scaling-based methods of JPEG2000 are limited in that the former does not allow the coding of any non-ROI bitplanes prior to the coding of the ROI bitplanes, and the latter supports only rectangular and circular ROIs. The paper presents a novel ROI-based coding method that extends the scaling-based method of JPEG2000, and allows the coding of arbitrary-shape ROIs, as well as the separate coding and transmission of the ROI shape information prior to the texture information. Coding results are presented to illustrate the performance of the proposed ROI-based coding method.
Mahesh Subedar, Lina J. Karam, Glen P. Abousleman
ICASSP (3)2
2004 Selective fec for error-resilient image coding and transmission using similarity check functions
abstract
This paper introduces the concept of similarity check functions that measure the degree of corruption in transmitted multimedia data packets. Rather than directly discarding or directly retransmitting erroneous packets, the degree of corruption contributed by a particular packet is measured by a similarity check function at the receiver, which does not require explicit knowledge of the source data. Accordingly, if the packet is found to be badly corrupted based on a specified similarity criteria, it can be considered as a lost packet and recovered using forward error correcting codes or retransmission. An example of a similarity check function design and a selective forward error correction (S-FEC) scheme that allocates channel protection bits only to those packets that are in most need of correction, is presented to illustrate the proposed concept.
Tuyet-Trang Lam, Lina J. Karam, Rida A. Bazzi, Glen P. Abousleman
ICIP2
2004 JPEG2000-based shape adaptive algorithm for the efficient coding of multiple regions-of interest
abstract
The JPEG2000 standard supports two methods to code regions-of-interest (ROIs) in an image-the maxshift method and the scaling-based method. Compared to the maxshift method, the scaling-based method is preferable in many applications because it allows the nonROI (background) region to be partially coded prior to coding all of the ROIs. This is accomplished by supporting arbitrary bitplane shift factors and by filling the most significant bit-planes of the nonROI region with zeros after the ROI bitplanes have been shifted up in value. In both of these methods, the embedded block coding with optimum truncation (EBCOT) algorithm is then applied to code the wavelet coefficients. However, when multiple regions-of-interest need to be coded, the performance of the scaling-based method degrades significantly since the EBCOT coder is no longer able to efficiently code the nonROI region, which is now less contiguous as compared to the single-ROI case. Also, the large header information that is required in JPEG2000 for each ROI degrades the coding performance. This paper presents an improved scaling-based method for the efficient coding of multiple, arbitrarily-shaped ROIs using a modified EBCOT algorithm. The proposed method utilizes the shape information of the different ROIs to generate a stripe mask, which is then used to optimize the coding performance. Coding results and comparisons with the JPEG2000 scaling-based method are presented to illustrate the improved performance of the proposed scheme.
Mahesh Subedar, Lina J. Karam, Glen P. Abousleman
ICIP2
2003 On-line laboratories for image and two-dimensional signal processing using 2D J-DSP
abstract
Java Digital Signal Processing (J-DSP) is a Java-based object-oriented programming environment that was developed at Arizona State University (ASU) for use in undergraduate- and graduate-level engineering classes. J-DSP is written as a platform-independent Java applet that resides on the Web and is thereby accessible by all students using a Web browser. We describe an innovative software extension to J-DSP, called 2D J-DSP, to accommodate on-line laboratories for two-dimensional digital signal processing. Two-dimensional DSP capabilities in J-DSP include: 2D signal generation; 2D FIR filter design and implementation; 2D transforms. Image processing capabilities include image restoration and enhancement. In order to illustrate 2D concepts graphically, contour (2D) and perspective (3D) plots have been incorporated in 2D J-DSP. Online laboratory exercises have been developed in the aforementioned areas for use in the graduate-level multidimensional signal processing and image processing courses at ASU, and are posted on the Website (http://jdsp.asu.edu). Statistical and qualitative evaluations that assess the learning experiences of the students that use 2D J-DSP are also presented.
Muhammad Yasin, Lina J. Karam, Andreas Spanias
ICASSP (3)2
2003 JPEG2000 encoding with perceptual distortion control
abstract
In this paper, a new JPEG2000-compliant encoding approach is proposed to control the JPEG2000 encoding in order to achieve a desired perceptual quality. Our method is based on a vision model that incorporates various masking effects of the human visual perception and on a perceptual distortion metric that takes spatial and spectral summation of individual quantization errors into account. Compared with the conventional JPEG2000 rate-based distortion minimization encoding, our method provides a way to generate consistent quality images with lower bit rate.
Lina J. Karam, Andrew B. Watson
ICIP (1)2
2003 Wavelet-based adaptive image denoising with edge preservation
abstract
This paper presents a state-of-the-art adaptive wavelet-based denoising method with edge preservation. More specifically, a redundant discrete dyadic wavelet transform (DDWT) is performed on the noisy image to get the wavelet frame decomposition at different scales. Based on the Lip-schitz regularity theory, correlation analysis across scales is performed to detect the significant coefficients from the signal and the insignificant coefficients from the noise for each subband. Different denoising techniques are applied to the significant coefficients and insignificant coefficients separately, based on different statistical models. Unlike most of the existing image denoising methods, the proposed method is able to not only shrink but also increase the magnitude of the noisy wavelet coefficients. Simulation results show that the proposed method has a remarkably superior ability to preserve the edge information and to achieve better visual quality.
Charles Q. Zhan, Lina J. Karam
ICIP (1)2
2002 Context formation by mutual information maximization
abstract
In this paper, the problem of how to form the contexts for context-based entropy coding is studied. The mutual information (MI) between the context and the encoded data is used to measure the context optimality. The MI decreases when contexts are combined together. Given a desired number of contexts, an algorithm is proposed for finding the set of contexts by iteratively combining the pairs that give the minimum MI reduction. The proposed algorithm is applied to form the contexts for the zero coding (ZC) primitive of the JPEG2000 image compression standard. Experimental results show that the number of contexts used as part of the standard can be reduced without loss in the coding performance.
Lina J. Karam
ICIP (3)2
2002 Robust hyperspectral image coding with channel-optimized trellis-coded quantization
abstract
This paper presents a wavelet-based hyperspectral image coder that is optimized for transmission over the binary symmetric channel (BSC). The proposed coder uses a robust channel-optimized trellis-coded quantization (COTCQ) stage that is designed to optimize the image coding based on the channel characteristics. This optimization is performed only at the level of the source encoder and does not include any channel coding for error protection. The robust nature of the coder increases the security level of the encoded bit stream, and provides a much higher quality decoded image. In the absence of channel noise, the proposed coder is shown to achieve a compression ratio greater than 70:1, with an average peak SNR of the coded hyperspectral sequence exceeding 40 dB. Additionally, the coder is shown to exhibit graceful degradation with increasing channel errors.
Glen P. Abousleman, Tuyet-Trang Lam, Lina J. Karam
IEEE Trans. Geosci. Remote. Sens.3
2002 Adaptive image coding with perceptual distortion control
abstract
This paper presents a discrete cosine transform (DCT)-based locally adaptive perceptual image coder, which discriminates between image components based on their perceptual relevance for achieving increased performance in terms of quality and bit rate. The new coder uses a locally adaptive perceptual quantization scheme based on a tractable perceptual distortion metric. Our strategy is to exploit human visual masking properties by deriving visual masking thresholds in a locally adaptive fashion. The derived masking thresholds are used in controlling the quantization stage by adapting the quantizer reconstruction levels in order to meet the desired target perceptual distortion. The proposed coding scheme is flexible in that it can be easily extended to work with any subband-based decomposition in addition to block-based transform methods. Compared to existing perceptual coding methods, the proposed perceptual coding method exhibits superior performance in terms of bit rate and distortion control. Coding results are presented to illustrate the performance of the presented coding scheme.
Ingo S. Hontsch, Lina J. Karam
IEEE Trans. Image Process.2
2001 Coding of digital imagery for transmission over multiple noisy channels
abstract
This paper presents a multiple description image coding scheme that facilitates the transmission of digital imagery over multiple noisy channels. The proposed scheme divides the image into smaller parts that are transmitted over the individual channels of an inverse multiplexing system. The division or splitting is done in such a fashion that it facilitates the interpolation of lost coefficients in the case of one or more channel failures. At the receiver, the image is reconstructed by proper assembly of the data received from each channel. In case of channel failure, the missing coefficients are estimated from the available data with the use of a novel post-processing scheme. For operation over four noisy channels with various bit error probabilities, we investigate the quantitative and subjective performance of the proposed system for the case of multiple channel failures.
Sumohana S. Channappayya, Glen P. Abousleman, Lina J. Karam
ICASSP3
2001 Image coding for transmission over multiple noisy channels using punctured convolutional codes and trellis-coded quantization
abstract
This paper presents a multiple description image coding scheme that facilitates the transmission of digital imagery over multiple noisy channels. The proposed scheme divides the image into smaller parts that are transmitted over the individual channels of an inverse multiplexing system. The division or splitting is done in such a fashion that it facilitates the interpolation of lost coefficients in the case of one or more channel failures. To combat the effects of channel noise, a novel joint source-channel coding scheme is employed, which uses punctured convolutional channel codes and trellis-coded quantization. At the receiver, the image is reconstructed by proper assembly of the data received from each channel. In case of channel failure, the missing coefficients are estimated from the available data with the use of a novel post processing scheme. For operation over four noisy channels with various bit error probabilities, we investigate the quantitative and subjective performance with and without channel failures.
Glen P. Abousleman, Sumohana S. Channappayya, Lina J. Karam
ICIP (1)3
2001 Canonic signed digit Chebyshev FIR filter design
abstract
A new algorithm is presented for the design of fixed point finite impulse response (FIR) digital filters that can be realized using a minimal number of adders while satisfying both the desired hardware implementation constraints and the frequency domain design constraints. The proposed algorithm makes use of a two-stage search method with a new soft-switching optimization criterion, minimizing the Chebyshev error norm in a discrete canonic signed digit (CSD) space. The proposed design method is implemented as a Matlab-based software package (CSDFIR) that allows the design of both Chebyshev optimal floating point and fixed point CSD FIR filters.
Yassin M. Y. Hasan, Lina J. Karam, Matt Falkinburg, Art Helwig, Matt Ronning
IEEE Signal Process. Lett.2
2000 Multiple description channel-optimized trellis-coded quantization
abstract
This paper presents the design of multiple description trellis-coded quantizers that are optimized for transmitting data over multiple noisy channels. A multiple description channel-optimized trellis-coded quantization (MDCOTCQ) scheme is developed whereby all trellis-coded quantizers (side and central) are designed to operate over the binary symmetric channel with a given bit error rate (BER). Using the tensor product of trellises, an expression of the multi-channel transition probability matrix is derived. This formulation enables the construction of multiple description channel-optimized trellis-coded quantizers for arbitrary bit rates and arbitrary channel bit error rates. The proposed MDCOTCQ system enables the transmission of compressed data over noisy diversity or packet-based communication systems, and enables outstanding robustness to channel or packet loss. Examples are presented to illustrate the performance of the proposed MDCOTCQ-based source coder.
Tuyet-Trang Lam, Glen P. Abousleman, Lina J. Karam
ICASSP3
2000 An Efficient Embedded Zerotree Wavelet Image Codec Based on Intraband Partitioning
abstract
Embedded zerotree wavelet (EZW) coding, introduced by J. M. Shapiro (1993) and improved in the SPIHT codec of A. Said and W.A. Pearlman (1996), is a very effective wavelet image coding method. One main concept in EZW and SPIHT is the exploitation of self-similarity across different scales of the image wavelet transform to effectively compress the image data. This paper presents an efficient and fully embedded zerotree wavelet image codec that achieves a comparable and even better performance than SPIHT by exploiting only the self-similarity within the scales (intraband self-similarity). The proposed algorithm results in a simpler implementation and in a more efficient representation of the ordering information than SPIHT. Coding results are presented to illustrate the performance of the proposed embedded wavelet image codec.
Lina J. Karam
ICIP2
2000 Error-resilient video coding with channel-optimized trellis-coded quantization
abstract
We present an error-resilient channel-optimized coder for the transmission of video over noisy channels. The proposed coder uses a robust channel-optimized trellis-coded quantization (COTCQ) stage that is designed to optimize the video coding based on the channel characteristics. The resilience to channel errors is obtained only at the level of the source encoder, with no explicit use of channel coding. The robust nature of the coder increases the security level of the encoded bit stream, and eliminates impulsive artifacts induced by channel errors. A novel adaptive classification scheme is employed, which eliminates the need for entropy coding and motion compensation. Consequently, the proposed channel-optimized video coder is especially suitable for wireless video transmission due to its reduced complexity, its robustness to nonstationary signals and channels, and its increased security level. Simulation results are presented to illustrate the performance of the proposed video coder for a wide variety of channel conditions, without the use of error concealment techniques.
Glen P. Abousleman, Lina J. Karam
WCNC3
2000 Image coding with robust channel-optimized trellis-coded quantization
abstract
This paper presents a wavelet-based image coder that is optimized for transmission over the binary symmetric channel (BSC). The proposed coder uses a robust channel-optimized trellis-coded quantization (COTCQ) stage that is designed to optimize the image coding based on the channel characteristics. A phase scrambling stage is also used to further increase the coding performance and robustness to nonstationary signals and channels. The resilience to channel errors is obtained by optimizing the coder performance only at the level of the source encoder with no explicit channel coding for error protection. For the considered TCQ trellis structure, a general expression is derived for the transition probability matrix. In terms of the TCQ encoding rat and the channel bit error rate, and is used to design the COTCQ stage of the image coder. The robust nature of the coder also increases the security level of the encoded bit stream and provides a much more visually pleasing rendition of the decoded image. Examples are presented to illustrate the performance of the proposed robust image coder.
Tuyet-Trang Lam, Glen P. Abousleman, Lina J. Karam
IEEE J. Sel. Areas Commun.3
2000 Morphological Reversible Contour Representation
abstract
In this paper, a novel morphological reversible contour representation of discrete binary images is proposed. A binary image is represented by a set of nonoverlapping multilevel contours and a residual image. In this proposed representation, the total number of pixels representing an image is far less than the total number of pixels obtained by the seed-based morphological contour-skeleton lossless representation. The proposed contour representation is simple, unique, and general without restrictions on the binary image to be represented. The resulting multicontour image component is robust to noise. An efficient differential chain contour coding scheme is employed to further compress the represented image. The proposed method yields very low bit rates compared to the existing morphological techniques. An automatic filling procedure, which properly fills a proper multicontour image according to its topological structure without need of seed points, is proposed.
Yassin M. Y. Hasan, Lina J. Karam
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 Morphological text extraction from images
abstract
This paper presents a morphological technique for text extraction from images. The proposed morphological technique is insensitive to noise, skew and text orientation. It is also free from artifacts that are usually introduced by both fixed/optimal global thresholding and fixed-size block-based local thresholding. Examples are presented to illustrate the performance of the proposed method.
Yassin M. Y. Hasan, Lina J. Karam
IEEE Trans. Image Process.2
2000 Locally adaptive perceptual image coding
abstract
Most existing efforts in image and video compression have focused on developing methods to minimize not perceptual but rather mathematically tractable, easy to measure, distortion metrics. While nonperceptual distortion measures were found to be reasonably reliable for higher bit rates (high-quality applications), they do not correlate well with the perceived quality at lower bit rates and they fail to guarantee preservation of important perceptual qualities in the reconstructed images despite the potential for a good signal-to-noise ratio (SNR). This paper presents a perceptual-based image coder, which discriminates between image components based on their perceptual relevance for achieving increased performance in terms of quality and bit rate. The new coder is based on a locally adaptive perceptual quantization scheme for compressing the visual data. Our strategy is to exploit human visual masking properties by deriving visual masking thresholds in a locally adaptive fashion based on a subband decomposition. The derived masking thresholds are used in controlling the quantization stage by adapting the quantizer reconstruction levels to the local amount of masking present at the level of each subband transform coefficient. Compared to the existing non-locally adaptive perceptual quantization methods, the new locally adaptive algorithm exhibits superior performance and does not require additional side information. This is accomplished by estimating the amount of available masking from the already quantized data and linear prediction of the coefficient under consideration. By virtue of the local adaptation, the proposed quantization scheme is able to remove a large amount of perceptually redundant information. Since the algorithm does not require additional side information, it yields a low entropy representation of the image and is well suited for perceptually lossless image compression.
Ingo S. Hontsch, Lina J. Karam
IEEE Trans. Image Process.2
1999 Wavelet-based image coder with channel-optimized trellis-coded quantization
abstract
This paper presents a wavelet-based image coder optimized for transmission over binary symmetric channels (BSC). The proposed coder uses a channel-optimized trellis-coded quantization (COTCQ) stage that is designed to optimize the image coding based on the channel characteristics. This optimization is performed only at the level of the source encoder, and does not include any channel coding for error protection. Consequently, the proposed channel-optimized image coder is especially suitable for wireless transmission due to its reduced complexity. Furthermore, the improvement over TCQ-based image coders is significant. Examples are presented to illustrate the performance of the proposed COTCQ-based image coder.
Tuyet-Trang Lam, Glen P. Abousleman, Lina J. Karam
ICASSP3
1999 A Perceptually Tuned Image Coder with Channel-Optimized Trellis-Coded Quantization
abstract
A system is presented for the compression of digital imagery over binary symmetric channels (BSC). The objective is to minimize the perceived distortion for a desired target bit rate. The proposed system utilizes wavelet decomposition, channel-optimized trellis-coded quantization (COTCQ), and a perceptually-tuned rate allocation scheme to control the quantization. Because of the inherent noise-robust properties of the COTCQ stage, no channel coding is employed. Additionally, the COTCQ design does not utilize entropy coding. Consequently, the proposed image coder is especially suitable for wireless transmission due to its reduced complexity and robustness to channel errors. Examples are presented to illustrate the performance of the proposed image coding system.
Tuyet-Trang Lam, Glen P. Abousleman, Lina J. Karam
ICIP (1)3
1999 Chebyshev digital FIR filter design
Lina J. Karam, James H. McClellan
Signal Process.1
1998 Locally-adaptive image coding based on a perceptual target distortion
abstract
This paper presents a perceptual-based image coder, which discriminates between image components based on their perceptual relevance for achieving increased performance in terms of quality and bit rate. The new coder uses a locally-adaptive perceptual quantization scheme based on a tractable perceptual distortion metric. Our strategy is to exploit human visual masking properties by deriving visual masking thresholds in a locally-adaptive fashion. The derived masking thresholds are used in controlling the quantization stage by adapting the quantizer reconstruction levels in order to meet a desired target perceptual distortion. The proposed coding scheme is flexible in that it works with any subband-based decomposition and with block-based transform methods. Compared to the existing perceptual transform-based and block-based methods, the proposed perceptual coding method exhibits superior performance in terms of the bit rate and distortion control. Coding results are presented to illustrate the performance of the presented coding scheme.
Ingo S. Hontsch, Lina J. Karam
ICASSP2
1997 On the design of multidimensional FIR filters by transformation
abstract
This paper studies the applicability and limitations of the McClellan transformation method and, as a result, extends this method so that new types of one-dimensional filters can be transformed and new types of multi-dimensional filters can be designed. For this purpose, a new expression for the frequency response of an arbitrary one-dimensional filter is derived in terms of Chebyshev polynomials and other introduced polynomials satisfying recurrence formulae. The main objective is to identify which prototype filters can be transformed, determine what types of symmetry can be designed, and present procedures for transforming the new identified prototypes as well as rules for achieving the possible symmetries.
Lina J. Karam
ICASSP1
1997 APIC: Adaptive Perceptual Image Coding Based on Sub-Band Decomposition with Locally Adaptive Perceptual Weighting
abstract
The perceptual subband image coder (PIC) introduced by Safranek and Johnston (1989), selects a noise target level for each subband based on an empirically derived perceptual masking measure. These noise target levels are used to set the quantization level in the DPCM quantizer for every particular subband. It achieves high quality output at bit rates from 0.1 to 0.9 bits/pixel (bpp) depending on the complexity of the image. In this paper, we present an algorithm that locally adapts the quantizer step size at each pixel according to an estimate of the masking measure. This estimate is based on the already coded pixels and predictions of the not yet coded pixels. Compared to the PIC, the proposed method does not require any additional side information. In fact, it eliminates the need to transmit the quantizer step size for each subband. For comparable perceptual quality, the proposed method achieves compression gains up to 40 percent. Typical values are in the order of 20 to 30 percent, depending on the nature of the image. Our algorithm has also better performance for supra-threshold image compression since the perceptual error is distributed more evenly and is not concentrated in the most sensitive regions.
Ingo S. Hontsch, Lina J. Karam
ICIP (1)2
1997 A Perceptually Tuned Embedded Zerotree Image Coder
abstract
Embedded zerotree wavelet coding (EZW), introduced by Shapiro (1993) and improved by Said and Pearlman (1996) using an algorithm based on set partitioning in hierarchical trees (SPIHT), is a computationally inexpensive image coding technique and has proven to be very effective at minimizing the mean-square error (MSE) distortion measure. However, minimizing MSE does not guarantee preservation of good perceptual qualities in the decoded image, especially at low bit rates. In this paper we propose a perceptually-tuned embedded zerotree image codec (PEZ) that introduces a perceptual weighting to the wavelet transform coefficients prior to EZW encoding. In this coder the EZW minimizes a perceptually based distortion measure instead of MSE. The perceptual weights for all subbands are computed based on the just noticeable distortion (JND) thresholds for uniform noise. The new perceptually tuned codec has the same complexity as the original EZW/SPIHT techniques and results in a comparable or superior coding performance. Coding results are presented to illustrate the performance of the new coder.
Ingo S. Hontsch, Lina J. Karam, Robert J. Safranek
ICIP (1)2
1996 Efficient design of families of FIR filters by transformation
abstract
This paper introduces an efficient technique for designing a parameterized family of one- and multi-dimensional filters by regarding it as a single multi-dimensional filter. The new design technique is very fast since it reduces the design of a large set of filters to the design of a complex or real one-dimensional filter which is then mapped to produce the desired set of filters. It has the additional advantage that the one- and multi-dimensional filters in the set can be designed and realized efficiently without having to explicitly compute the coefficients of the separate filters. Among the types of filters which can be designed with this technique are seismic migration filters and linear-phase filters.
Lina J. Karam, James H. McClellan
ICASSP1
1996 Design of complex multi-dimensional FIR filters by transformation
abstract
This paper extends the transformation method by developing two new efficient procedures for the design of complex and real, positive- and negative-symmetric, multi-dimensional filters by transforming even-length 1-D prototype filters with complex (or real) coefficients. The first procedure is used to design complex multi-dimensional filters with a rectangular region of support having odd-length sides, while the second procedure is used for filters with a rectangular region of support having even-length sides. The designed filters can be implemented efficiently using special Chebyshev structures. Design examples are presented to illustrate the performance of the design procedures.
Lina J. Karam
ICIP (1)1
1995 Chroma coding for video at very low bit rates
abstract
Color pictures are usually compressed in a luminance-chrominance coordinate space. We consider the problem of encoding the chrominance information for very low bit rate video coding systems aimed at bit rates in the range 8 to 40 kbps. The challenge is that the chrominance components typically get less than 10 to 20% of the total very low bit rate allocated for the video data. We found that it is sufficient to encode the chrominance information at 1/8 of the luminance resolution in both the horizontal and vertical directions. While, for many of the previous coding methods, the compression is performed independently for the luminance and chrominance coordinates, we propose a coding scheme which exploits the coded luminance data in coding and retrieving the chrominance components. The proposed video coder is an improved extension of an existing luminance-only coder so that color motion video can be coded at very low bit rates under fixed frame and bit rate constraints. It is based on a hybrid waveform coding technique with an implicit model-based component. Very good results were obtained for head-and-shoulders sequences even with chroma rates of less than 7% of the total very low bit rate. In addition, subjective tests indicate that the coded chrominance information improves the visual perception of noisy image features.
Lina J. Karam, Christine Podilchuk
ICIP1
1994 A Multiple Exchange Remez Algorithm for Complex FIR Filter Design in the Chebyshev Sense
abstract
The alternation theorem is at the core of the efficient real Chebyshev approximation algorithms. In this paper, the alternation theorem is extended from the real-only to the complex case. A new efficient algorithm is described for designing FIR filters that best approximate in the Chebyshev sense a desired complex-valued function. This algorithm is based on an ascent Remez exchange method applied to a transformation of the complex Chebyshev error, and is basically a generalization of the Parks-McClellan algorithm to the complex case. Numerical examples are presented to illustrate the performance of the proposed algorithm.>
Lina J. Karam, James H. McClellan
ISCAS1