Nasser M. Nasrabadi

dblp:45/4884 · DBLP profile ↗
← Back
244ranked-venue papers
28as first author
47since 2021 · last 2025
0000-0001-8730-627XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 170 · 15 first-author · 38 since 2021Artificial intelligence and machine learning · 82 · 7 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 21 · 2 first-author · 15 since 2021Security and privacy · 18 · 16 since 2021Computer networks · 9 · 4 first-authorDatabases, data management, data science and information retrieval · 4Systems, architecture and hardware · 2 · 1 first-author
YearPublicationVenuePosition
2025 GIF: Generative Inspiration for Face Recognition at Scale
abstract
Aiming to reduce the computational cost of Softmax in massive label space of Face Recognition (FR) benchmarks, recent studies estimate the output using a subset of identities. Although promising, the association between the computation cost and the number of identities in the dataset remains linear only with a reduced ratio. A shared characteristic among available FR methods is the employment of atomic scalar labels during training. Consequently, the input to label matching is through a dot product between the feature vector of the input and the Softmax centroids. Inspired by generative modeling, we present a simple yet effective method that substitutes scalar labels with structured identity code, i.e., a sequence of integers. Specifically, we propose a tokenization scheme that transforms atomic scalar labels into structured identity codes. Then, we train an FR backbone to predict the code for each input instead of its scalar label. As a result, the associated computational cost becomes logarithmic w.r.t. number of identities. We demonstrate the benefits of the proposed method by conducting experiments. In particular, our method outperforms its competitors by 1.52%, and 0.6% at TAR@FAR= le — 4 on IJB- B and IJB-C, respectively, while transforming the association between computational cost and the number of identities from linear to logarithmic. Code
Saeed Ebrahimi, Sahar Rahimi Malakshan, Ali Dabouei, Srinjoy Das, Jeremy M. Dawson, Nasser M. Nasrabadi
CVPR6
2025 CLFace: A Scalable and Resource-Efficient Continual Learning Framework for Lifelong Face Recognition
abstract
An important aspect of deploying face recognition (FR) algorithms in real-world applications is their ability to learn new face identities from a continuous data stream. However, the online training of existing deep neural network-based FR algorithms, which are pre-trained offline on large-scale stationary datasets, encounter two major challenges: (I) catastrophic forgetting of previously learned identities, and (II) the need to store past data for complete retraining from scratch, leading to significant storage constraints and privacy concerns. In this paper, we introduce CLFace, a continual learning framework designed to preserve and incrementally extend the learned knowledge. CLFace eliminates the classification layer, resulting in a resource-efficient FR model that remains fixed throughout lifelong learning and provides label-free supervision to a student model, making it suitable for open-set face recog-nition during incremental steps. We introduce an objective function that employs feature-level distillation to reduce drift between feature maps of the student and teacher models across multiple stages. Additionally, it incorpo-rates a geometry-preserving distillation scheme to maintain the orientation of the teacher model's feature embedding. Furthermore, a contrastive knowledge distillation is incor-porated to continually enhance the discriminative power of the feature representation by matching similarities between new identities. Experiments on several benchmark FR datasets demonstrate that CLFace outperforms baseline approaches and state-of-the-art methods on unseen identities using both in-domain and out-of-domain datasets.
Md Mahedi Hasan, Shoaib Meraj Sami, Nasser M. Nasrabadi
WACV3
2025 Decomposed Distribution Matching in Dataset Condensation
abstract
Dataset Condensation (DC) aims to reduce deep neural networks training efforts by synthesizing a small dataset such that it will be as effective as the original large dataset. Conventionally, DC relies on a costly bi-level optimization which prohibits its practicality. Recent research formulates DC as a distribution matching problem which circumvents the costly bi-level optimization. However, this efficiency sacrifices the DC performance. To investigate this performance degradation, we decomposed the dataset distribution into content and style. Our observations indicate two major shortcomings of: 1) style discrepancy between original and condensed data, and 2) limited intra-class diversity of condensed dataset. We present a simple yet effective method to match the style information between original and condensed data, employing statistical moments of feature maps as well-established style indicators. Moreover, we enhance the intra-class diversity by maximizing the Kullback-Leibler divergence within each synthetic class, i.e., content. We demonstrate the efficacy of our method through experiments on diverse datasets of varying size and resolution, achieving improvements of up to 4.1% on CI-FAR 10, 4.2% on CIFAR100, 4.3% on TinyImageNet, 2.0% on ImageNet-1K, 3.3% on ImageWoof, 2.5% on ImageNette, and 5.5% in continual learning accuracy. Code
Sahar Rahimi Malakshan, Mohammad Saeed Ebrahimi Saadabadi, Ali Dabouei, Nasser M. Nasrabadi
WACV4
2024 Laplacian-guided Entropy Model in Neural Codec with Blur-dissipated Synthesis
abstract
While replacing Gaussian decoders with a conditional diffusion model enhances the perceptual quality of reconstructions in neural image compression, their lack of in-ductive bias for image data restricts their ability to achieve state-of-the-art perceptual levels. To address this limitation, we adopt a non-isotropic diffusion model at the de-coder side. This model imposes an inductive bias aimed at distinguishing between frequency contents, thereby fa-cilitating the generation of high-quality images. Moreover, our framework is equipped with a novel entropy model that accurately models the probability distribution of la-tent representation by exploiting spatio-channel correlations in latent space, while accelerating the entropy de-coding step. This channel-wise entropy model leverages both local and global spatial contexts within each channel chunk. The global spatial context is built upon the Trans-former, which is specifically designed for image compression tasks. The designed Transformer employs a Laplacian-shaped positional encoding, the learnable parameters of which are adaptively adjusted for each channel cluster. Our experiments demonstrate that our proposed frame-work yields better perceptual quality compared to cutting-edge generative-based codecs, and the proposed entropy model contributes to notable bitrate savings. The code is available at https://github.com/Atefeh-Khoshtinat/Blur-dissipated-compression.
Atefeh Khoshkhahtinat, Piyush M. Mehta, Nasser M. Nasrabadi
CVPR4
2024 Hyperspherical Classification with Dynamic Label-to-Prototype Assignment
abstract
Aiming to enhance the utilization of metric space by the parametric softmax classifier, recent studies suggest replacing it with a non-parametric alternative. Although a nonparametric classifier may provide better metric space utilization, it introduces the challenge of capturing inter-class relationships. A shared characteristic among prior nonparametric classifiers is the static assignment of labels to prototypes during the training, i.e., each prototype consistently represents a class throughout the training course. Orthogonal to previous works, we present a simple yet effective method to optimize the category assigned to each prototype (label-to-prototype assignment) during the training. To this aim, we formalize the problem as a two-step optimization objective over network parameters and label-to-prototype assignment mapping. We solve this optimization using a sequential combination of gradient descent and Bi-partide matching. We demonstrate the benefits of the proposed approach by conducting experiments on balanced and long-tail classification problems using different backbone network architectures. In particular, our method outperforms its competitors by 1.22% accuracy on CIFAR-100, and 2.15% on ImageNet-200 using a metric space dimension half of the size of its competitors. Code
Mohammad Saeed Ebrahimi Saadabadi, Ali Dabouei, Sahar Rahimi Malakshan, Nasser M. Nasrabadi
CVPR4
2024 ARoFace: Alignment Robustness to Improve Low-Quality Face Recognition
Mohammad Saeed Ebrahimi Saadabadi, Sahar Rahimi Malakshan, Ali Dabouei, Nasser M. Nasrabadi
ECCV (33)4
2024 Identity-Preserving GAN for Cross Spectral Iris Recognition
abstract
Cross spectral iris recognition has been shown to cause a degradation in iris matching scenarios due to the inherent differences between the NIR and visible spectra. This led us to explore methods of iris domain translation, allowing us to generate images between the NIR and visible domains using generative adversarial networks (GANs). We train a GAN network with an additional classifier component to act as an identity-preserving module allowing the generator to produce not only high quality, but identity-specific images. We apply this method on three cross-spectral iris datasets, namely, the Cross-eyed-cross-spectral iris database, the PolyU bi-spectral database and the WVU Multispectral database collected from our lab. We implement image enhancement techniques on the cropped iris images and unrolled, normalized iris images, allowing for the generator to learn the iris texture with minimal noise surrounding the iris and to show the performance of the generated images in different matching scenarios. We show the performance of our model by matching the generated iris images against the true iris images in their translated domain. We show that applying this image translation technique as a preprocessing step increases the matching performance when applied to iris matching software, such as Neurotechnology’s commercial iris recognition software, VeriEye and an open-source iris recognition software, OSIRIS. Lastly, we perform an ablation study for each set of experiments by removing the classifier component and comparing the results with our model, showing that the competition between the generator and classifier has an important role in learning identity-specific features.
Hannah Anderson, Moktari Mostofa, Nasser M. Nasrabadi, Jeremy M. Dawson
IJCB3
2024 UFQA: Utility guided Fingerphoto Quality Assessment
abstract
Quality assessment of fingerprints captured using digital cameras and smartphones, also called fingerphotos, is a challenging problem in biometric recognition systems. As contactless biometric modalities are gaining more attention, their reliability should also be improved. Many factors, such as illumination, image contrast, camera angle, etc., in fingerphoto acquisition introduce various types of distortion that may render the samples useless. Current quality estimation methods developed for fingerprints collected using contact-based sensors are inadequate for fingerphotos. We propose Utility guided Fingerphoto Quality Assessment (UFQA), a self-supervised dual encoder framework to learn meaningful feature representations to assess fingerphoto quality. A quality prediction model is trained to assess fingerphoto quality with additional supervision of quality maps. The quality metric is a predictor of the utility of fingerphotos in matching scenarios. Therefore, we use a holistic approach by including fingerphoto utility and local quality when labeling the training data. Experimental results verify that our approach performs better than the widely used fingerprint quality metric NFIQ2.2 and state-of-the-art image quality assessment algorithms on multiple publicly available fingerphoto datasets.
Amol S. Joshi, Ali Dabouei, Jeremy M. Dawson, Nasser M. Nasrabadi
IJCB4
2024 FDWST: Fingerphoto Deblurring using Wavelet Style Transfer
abstract
The challenge of deblurring fingerphoto images, or generating a sharp fingerphoto from a given blurry one, is a significant problem in the realm of computer vision. To address this problem, we propose a fingerphoto deblurring architecture referred to as Fingerphoto Deblurring using Wavelet Style Transfer (FDWST), which aims to utilize the information transmission of Style Transfer techniques to deblur fingerphotos. Additionally, we incorporate the Discrete Wavelet Transform (DWT) for its ability to split images into different frequency bands. By combining these two techniques, we can perform Style Transfer over a wide array of wavelet frequency bands, thereby increasing the quality and variety of sharpness information transferred from sharp to blurry images. Using this technique, our model was able to drastically increase the quality of the generated fingerphotos compared to their originals, and achieve a peak matching accuracy of 0.9907 when tasked with matching a deblurred fingerphoto to its sharp counterpart, outperforming multiple other state-of-the-art deblurring and style transfer techniques.
David Keaton, Amol S. Joshi, Jeremy M. Dawson, Nasser M. Nasrabadi
IJCB4
2024 Boosting Unconstrained Face Recognition with Targeted Style Adversary
abstract
While deep face recognition models have demonstrated remarkable performance, they often struggle on the inputs from domains beyond their training data. Recent attempts aim to expand the training set by relying on computationally expensive and inherently challenging image-space augmentation of image generation modules. In an orthogonal direction, we present a simple yet effective method to expand the training data by interpolating between instance-level feature statistics across labeled and unlabeled sets. Our method, dubbed Targeted Style Adversary (TSA), is motivated by two observations: (i) the input domain is reflected in feature statistics, and (ii) face recognition model performance is influenced by style information. Shifting towards an unlabeled style implicitly synthesizes challenging training instances. We devise a recognizability metric to constraint our framework to preserve the inherent identity-related information of labeled instances. The efficacy of our method is demonstrated through evaluations on unconstrained benchmarks, outperforming or being on par with its competitors while offering nearly a 70% improvement in training speed and 40% less memory consumption.
Mohammad Saeed Ebrahimi Saadabadi, Sahar Rahimi Malakshan, Seyed Rasoul Hosseini, Nasser M. Nasrabadi
IJCB4
2024 Text-Guided Face Recognition using Multi-Granularity Cross-Modal Contrastive Learning
abstract
State-of-the-art face recognition (FR) models often experience a significant performance drop when dealing with facial images in surveillance scenarios where images are in low quality and often corrupted with noise. Leveraging facial characteristics, such as freckles, scars, gender, and ethnicity, becomes highly beneficial in improving FR performance in such scenarios. In this paper, we introduce text-guided face recognition (TGFR) to analyze the impact of integrating facial attributes in the form of natural language descriptions. We hypothesize that adding semantic information into the loop can significantly improve the image understanding capability of an FR algorithm compared to other soft biometrics. However, learning a discriminative joint embedding within the multimodal space poses a considerable challenge due to the semantic gap in the unaligned image-text representations, along with the complexities arising from ambiguous and incoherent textual descriptions of the face. To address these challenges, we introduce a face-caption alignment module (FCAM), which incorporates cross-modal contrastive losses across multiple granularities to maximize the mutual information between local and global features of the face-caption pair. Within FCAM, we refine both facial and textual features for learning aligned and discriminative features. We also design a face-caption fusion module (FCFM) that applies fine-grained interactions and coarse-grained associations among cross-modal features. Through extensive experiments conducted on three face-caption datasets, proposed TGFR demonstrates remarkable improvements, particularly on low-quality images, over existing FR models and outperforms other related methods and benchmarks.
Md Mahedi Hasan, Shoaib Meraj Sami, Nasser M. Nasrabadi
WACV3
2023 Improving Face Recognition from Caption Supervision with Multi-Granular Contextual Feature Aggregation
abstract
We introduce caption-guided face recognition (CGFR) as a new framework to improve the performance of commercial-off-the-shelf (COTS) face recognition (FR) systems. In contrast to combining soft biometrics (e.g., facial marks, gender, and age) with face images, in this work, we use facial descriptions provided by face examiners as a piece of auxiliary information. However, due to the heterogeneity of the modalities, improving the performance by directly fusing the textual and facial features is very challenging, as both lie in different embedding spaces. In this paper, we propose a contextual feature aggregation module (CFAM) that addresses this issue by effectively exploiting the fine-grained word-region interaction and global image-caption association. Specifically, CFAM adopts a self-attention and a cross-attention scheme for improving the intra-modality and inter-modality relationship between the image and textual features, respectively. Additionally, we design a textual feature refinement module (TFRM) that refines the textual features of the pre-trained BERT encoder by updating the contextual embeddings. This module enhances the discriminative power of textual features with a cross-modal projection loss and realigns the word and caption embeddings with visual features by incorporating a visual-semantic alignment loss. We implemented the proposed CGFR framework on two face recognition models (ArcFace and AdaFace) and evaluated its performance on the Multi-Modal CelebA-HQ dataset. Our framework significantly improves the performance of ArcFace in both 1:1 verification and 1:N identification protocol.
Md Mahedi Hasan, Nasser M. Nasrabadi
IJCB2
2023 Towards Generalizable Morph Attack Detection with Consistency Regularization
abstract
Though recent studies have made significant progress in morph attack detection by virtue of deep neural networks, they often fail to generalize well to unseen morph attacks. With numerous morph attacks emerging frequently, generalizable morph attack detection has gained significant attention. This paper focuses on enhancing the generalization capability of morph attack detection from the perspective of consistency regularization. Consistency regularization operates under the premise that generalizable morph attack detection should output consistent predictions irrespective of the possible variations that may occur in the input space. In this work, to reach this objective, two simple yet effective morph-wise augmentations are proposed to explore a wide space of realistic morph transformations in our consistency regularization. Then, the model is regularized to learn consistently at the logit as well as embedding levels across a wide range of morph-wise augmented images. The proposed consistency regularization aligns the abstraction in the hidden layers of our model across the morph attack images which are generated from diverse domains in the wild. Experimental results demonstrate the superior generalization and robustness performance of our proposed method compared to the state-of-the-art studies.
Hossein Kashiyani, Niloufar Alipour Talemi, Mohammad Saeed Ebrahimi Saadabadi, Nasser M. Nasrabadi
IJCB4
2023 Deep Boosting Multi-Modal Ensemble Face Recognition with Sample-Level Weighting
abstract
Deep convolutional neural networks have achieved remarkable success in face recognition (FR), partly due to the abundant data availability. However, the current training benchmarks exhibit an imbalanced quality distribution; most images are of high quality. This poses issues for generalization on hard samples since they are underrepresented during training. In this work, we employ the multi-model boosting technique to deal with this issue. Inspired by the well-known AdaBoost, we propose a sample-level weighting approach to incorporate the importance of different samples into the FR loss. Individual models of the proposed framework are experts at distinct levels of sample hardness. Therefore, the combination of models leads to a robust feature extractor without losing the discriminability on the easy samples. Also, for incorporating the sample hardness into the training criterion, we analytically show the effect of sample mining on the important aspects of current angular margin loss functions, i.e., margin and scale. The proposed method shows superior performance in comparison with the state-of-the-art algorithms in extensive experiments on the CFP-FP, LFW, CPLFW, CALFW, AgeDB, TinyFace, IJB-B, and IJB-C evaluation datasets.
Sahar Rahimi Malakshan, Mohammad Saeed Ebrahimi Saadabadi, Nima Najafzadeh, Nasser M. Nasrabadi
IJCB4
2023 CCFace: Classification Consistency for Low-Resolution Face Recognition
abstract
In recent years, deep face recognition methods have demonstrated impressive results on in-the-wild datasets. However, these methods have shown a significant decline in performance when applied to real-world low-resolution benchmarks like TinyFace or SCFace. To address this challenge, we propose a novel classification consistency knowledge distillation approach that transfers the learned classifier from a high-resolution model to a low-resolution network. This approach helps in finding discriminative representations for low-resolution instances. To further improve the performance, we designed a knowledge distillation loss using the adaptive angular penalty inspired by the success of the popular angular margin loss function. The adaptive penalty reduces overfitting on low-resolution samples and alleviates the convergence issue of the model integrated with data augmentation. Additionally, we utilize an asymmetric cross-resolution learning approach based on the state-of-the-art semi-supervised representation learning paradigm to improve discriminability on low-resolution instances and prevent them from forming a cluster. Our proposed method outperforms state-of-the-art approaches on low-resolution benchmarks, with a three percent improvement on TinyFace while maintaining performance on high-resolution benchmarks.
Mohammad Saeed Ebrahimi Saadabadi, Sahar Rahimi Malakshan, Hossein Kashiyani, Nasser M. Nasrabadi
IJCB4
2023 AAFACE: Attribute-Aware Attentional Network for Face Recognition
abstract
In this paper, we present a new multi-branch neural network that simultaneously performs soft biometric (SB) prediction as an auxiliary modality and face recognition (FR) as the main task. Our proposed network named AAFace utilizes SB attributes to enhance the discriminative ability of FR representation. To achieve this goal, we propose an attribute-aware attentional integration (AAI) module to perform weighted integration of FR with SB feature maps. Our proposed AAI module is not only fully context-aware but also capable of learning complex relationships between input features by means of the sequential multi-scale channel and spatial sub-modules. Experimental results verify the superiority of our proposed network compared with the state-of-the-art (SoTA) SB prediction and FR methods.
Niloufar Alipour Talemi, Hossein Kashiyani, Sahar Rahimi Malakshan, Mohammad Saeed Ebrahimi Saadabadi, Nima Najafzadeh, Mohammad Akyash, Nasser M. Nasrabadi
ICIP7
2023 Frequency Disentangled Features in Neural Image Compression
abstract
The design of a neural image compression network is governed by how well the entropy model matches the true distribution of the latent code. Apart from the model capacity, this ability is indirectly under the effect of how close the relaxed quantization is to the actual hard quantization. Optimizing the parameters of a rate-distortion variational autoencoder (R-D VAE) is ruled by this approximated quantization scheme. In this paper, we propose a feature-level frequency disentanglement to help the relaxed scalar quantization achieve lower bit rates by guiding the high entropy latent features to include most of the low-frequency texture of the image. In addition, to strengthen the de-correlating power of the transformer-based analysis/synthesis transform, an augmented self-attention score calculation based on the Hadamard product is utilized during both encoding and decoding. Channel-wise autoregressive entropy modeling takes advantage of the proposed frequency separation as it inherently directs high-informational low-frequency channels to the first chunks and conditions the future chunks on it. The proposed network not only outperforms hand-engineered codecs, but also neural network-based codecs built on computation-heavy spatially autoregressive entropy models.
Atefeh Khoshkhahtinat, Piyush M. Mehta, Mohammad Saeed Ebrahimi Saadabadi, Mohammad Akyash, Nasser M. Nasrabadi
ICIP6
2023 Trading-Off Mutual Information on Feature Aggregation for Face Recognition
abstract
Despite the advances in the field of Face Recognition (FR), the precision of these methods is not yet sufficient. To improve the FR performance, this paper proposes a technique to aggregate the outputs of two state-of-the-art (SOTA) deep FR models, namely ArcFace and AdaFace. In our approach, we leverage the transformer attention mechanism to exploit the relationship between different parts of two feature maps. By doing so, we aim to enhance the overall discriminative power of the FR system. One of the challenges in feature aggregation is the effective modeling of both local and global dependencies. Conventional transformers are known for their ability to capture long-range dependencies, but they often struggle with modeling local dependencies accurately. To address this limitation, we augment the self-attention mechanism to capture both local and global dependencies effectively. This allows our model to take advantage of the overlapping receptive fields present in corresponding locations of the feature maps. However, fusing two feature maps from different FR models might introduce redundancies to the face embedding. Since these models often share identical backbone architectures, the resulting feature maps may contain overlapping information, which can mislead the training process. To overcome this problem, we leverage the principle of Information Bottleneck to obtain a maximally informative facial representation. This ensures that the aggregated features retain the most relevant and discriminative information while minimizing redundant or misleading details. To evaluate the effectiveness of our proposed method, we conducted experiments on popular benchmarks and compared our results with state-of-the-art algorithms. The consistent improvement we observed in these benchmarks demonstrates the efficacy of our approach in enhancing FR performance. Moreover, our model aggregation framework offers a novel perspective on model fusion and establishes a powerful paradigm for feature aggregation using transformer-based attention mechanisms.
Mohammad Akyash, Nasser M. Nasrabadi
ICMLA3
2023 Multi-Context Dual Hyper-Prior Neural Image Compression
abstract
Transform and entropy models are the two core components in deep image compression neural networks. Most existing learning-based image compression methods utilize convolutional-based transform, which lacks the ability to model long-range dependencies, primarily due to the limited receptive field of the convolution operation. To address this limitation, we propose a Transformer-based nonlinear transform. This transform has the remarkable ability to efficiently capture both local and global information from the input image, leading to a more decorrelated latent representation. In addition, we introduce a novel entropy model that incorporates two different hyperpriors to model cross-channel and spatial dependencies of the latent representation. To further improve the entropy model, we add a global context that leverages distant relationships to predict the current latent more accurately. This global context employs a causal attention mechanism to extract long-range information in a content-dependent manner. Our experiments show that our proposed framework performs better than the state-of-the-art methods in terms of rate-distortion performance.
Atefeh Khoshkhahtinat, Piyush M. Mehta, Mohammad Akyash, Hossein Kashiyani, Nasser M. Nasrabadi
ICMLA6
2023 Context-Aware Neural Video Compression on Solar Dynamics Observatory
abstract
NASA's Solar Dynamics Observatory (SDO) mission collects large data volumes of the Sun's daily activity. Data compression is crucial for space missions to reduce data storage and video bandwidth requirements by eliminating redundancies in the data. In this paper, we present a novel neural Transformer-based video compression approach specifically designed for the SDO images. Our primary objective is to efficiently exploit the temporal and spatial redundancies inherent in solar images to obtain a high compression ratio. Our proposed architecture benefits from a novel Transformer block called Fused Local-aware Window (FLa Win), which incorporates window-based self-attention modules and an efficient fused local-aware feed-forward (FLaFF) network. This architectural design allows us to simultaneously capture short-range and long-range information while facilitating the extraction of rich and diverse contextual representations. Moreover, this design choice results in reduced computational complexity. Experimental results demonstrate the significant contribution of the FLaWin Transformer block to the compression performance, outperforming conventional hand-engineered video codecs such as H.264 and H.265 in terms of rate-distortion trade-off.
Atefeh Khoshkhahtinat, Piyush M. Mehta, Nasser M. Nasrabadi, Barbara J. Thompson, Michael S. F. Kirk
ICMLA4
2023 Multi-Spectral Entropy Constrained Neural Compression of Solar Imagery
abstract
Missions studying the dynamic behaviour of the Sun are defined to capture multi-spectral images of the sun and transmit them to the ground station in a daily basis. To make transmission efficient and feasible, image compression systems need to be exploited. Recently successful end-to-end optimized neural network-based image compression systems have shown great potential to be used in an ad-hoc manner. In this work we have proposed a transformer-based multi-spectral neural image compressor to efficiently capture redundancies both intra/inter-wavelength. To unleash the locality of window-based self attention mechanism, we propose an inter-window aggregated token multi head self attention. Additionally to make the neural compressor autoencoder shift invariant, a randomly shifted window attention mechanism is used which makes the transformer blocks insensitive to translations in their input domain. We demonstrate that the proposed approach not only outperforms the conventional compression algorithms but also it is able to better decorrelates images along the multiple wavelengths compared to single spectral compression.
Atefeh Khoshkhahtinat, Piyush M. Mehta, Nasser M. Nasrabadi, Barbara J. Thompson, Michael S. F. Kirk
ICMLA4
2023 A Quality Aware Sample-to-Sample Comparison for Face Recognition
abstract
Currently available face datasets mainly consist of a large number of high-quality and a small number of low-quality samples. As a result, a Face Recognition (FR) network fails to learn the distribution of low-quality samples since they are less frequent during training (underrepresented). Moreover, current state-of-the-art FR training paradigms are based on the sample-to-center comparison (i.e., Softmax-based classifier), which results in a lack of uniformity between train and test metrics. This work integrates a quality-aware learning process at the sample level into the classification training paradigm (QAFace). In this regard, Softmax centers are adaptively guided to pay more attention to low-quality samples by using a quality-aware function. Accordingly, QAFace adds a quality-based adjustment to the updating procedure of the Softmax-based classifier to improve the performance on the underrepresented low-quality samples. Our method adaptively finds and assigns more attention to the recognizable low-quality samples in the training datasets. In addition, QAFace ignores the unrecognizable low-quality samples using the feature magnitude as a proxy for quality. As a result, QAFace prevents class centers from getting distracted from the optimal direction. The proposed method is superior to the state-of-the-art algorithms in extensive experimental results on the CFP-FP, LFW, CPLFW, CALFW, AgeDB, IJB-B, and IJB-C datasets.
Mohammad Saeed Ebrahimi Saadabadi, Sahar Rahimi Malakshan, Moktari Mostofa, Nasser M. Nasrabadi
WACV5
2022 Revisiting Outer Optimization in Adversarial Training
Ali Dabouei, Fariborz Taherkhani, Sobhan Soleymani, Nasser M. Nasrabadi
ECCV (5)4
2022 Ubiquitous Physiological Prediction of SUD Patients' Wellness State Using Memory-Based Convolutional Models
abstract
The prevalence of substance use disorder (SUD) and rates of overdose in the United States have reached epidemic levels. Despite availability of effective evidence-based treatments for SUD, the rates of treatment attrition remain elevated. We have designed a cloud-based continuous physiological sensing for longitudinal SUD patient monitoring. Using wearable sensors, we aim to evaluate the impact of changes in heart rate (HR) and heart rate variability (HRV) signals on SUD wellness development using long-term and ubiquitous monitoring and machine learning and collected data from 10 subjects over an extended period of time. We designed a signal processing recipe and employed several recurrent neural network (RNN)-based architectures to track the temporal and spectral behavior of HR and HRV signals to predict the patients’ wellness state. In addition, we have designed an architecture that combines RNN architectures and Time Scattered convolutional neural networks (TS-CNNs), where CNNs objectify the underlying features in the temporal dimensions within the signals (TS). The goal is to quantitatively and qualitatively evaluate the contribution of TS-CNNs in ubiquitous wellness prediction. The experimental results demonstrate that the best architecture configuration achieves 90.21% accuracy in predicting the wellness state of the SUD patients.
Omid Dehzangi, Paria Jeihouni, Jad Ramadan, Victor S. Finomore Jr., Nasser M. Nasrabadi, Ali Rezai
ICASSP5
2022 Superresolution and Segmentation of OCT Scans Using Multi-Stage Adversarial Guided Attention Training
abstract
Optical coherence tomography (OCT) is one of the noninvasive and easy-to-acquire biomarkers (the thickness of the retinal layers, which is detectable within OCT scans) being investigated to diagnose Alzheimer’s disease (AD). This work aims to segment the OCT images automatically; however, it is a challenging task due to various issues such as the speckle noise, small target region, and unfavorable imaging conditions. In our previous work, we have proposed the multi-stage & multi-discriminatory generative adversarial network (MultiSDGAN) [1] to translate OCT scans in high-resolution segmentation labels. In this investigation, we aim to evaluate and compare various combinations of channel and spatial attention to the MultiSDGAN architecture to extract more powerful feature maps by capturing rich contextual relationships to improve segmentation performance. Moreover, we developed and evaluated a guided mutli-stage attention framework where we incorporated a guided attention mechanism by forcing an L-1 loss between a specifically designed binary mask and the generated attention maps. Our ablation study results on the WVU-OCT data-set in five-fold cross-validation (5-CV) suggest that the proposed MultiSDGAN with a serial attention module provides the most competitive performance, and guiding the spatial attention feature maps by binary masks further improves the performance in our proposed network. Comparing the baseline model with adding the guided-attention, our results demonstrated relative improvements of 21.44% and 19.45% on the Dice coefficient and SSIM, respectively.
Paria Jeihouni, Omid Dehzangi, Annahita Amireskandari, Ali Dabouei, Ali Rezai, Nasser M. Nasrabadi
ICASSP6
2022 Robust Ensemble Morph Detection with Domain Generalization
abstract
Although a substantial amount of studies is dedicated to morph detection, most of them fail to generalize for morph faces outside of their training paradigm. Moreover, recent morph detection methods are highly vulnerable to adversarial attacks. In this paper, we intend to learn a morph detection model with high generalization to a wide range of morphing attacks and high robustness against different adversarial attacks. To this aim, we develop an ensemble of convolutional neural networks (CNNs) and Transformer models to benefit from their capabilities simultaneously. To improve the robust accuracy of the ensemble model, we employ multi-perturbation adversarial training and generate adversarial examples with high transferability for several single models. Our exhaustive evaluations demonstrate that the proposed robust ensemble model generalizes to several morphing attacks and face datasets. In addition, we validate that our robust ensemble model gain better robustness against several adversarial attacks while outperforming the state-of-the-art studies.
Hossein Kashiyani, Shoaib Meraj Sami, Sobhan Soleymani, Nasser M. Nasrabadi
IJCB4
2022 Pose Attention-Guided Profile-to-Frontal Face Recognition
abstract
In recent years, face recognition systems have achieved exceptional success due to promising advances in deep learning architectures. However, they still fail to achieve expected accuracy when matching profile images against a gallery of frontal images. Current approaches either perform pose normalization (i.e., frontalization) or disentangle pose information for face recognition. We instead propose a new approach to utilize pose as an auxiliary information via an attention mechanism. In this paper, we hypothesize that pose attended information using an attention mechanism can guide contextual and distinctive feature extraction from profile faces, which further benefits a better representation learning in an embedded domain. To achieve this, first, we design a unified coupled profile-to-frontal face recognition network. It learns the mapping from faces to a compact em-bedding subspace via a class-specific contrastive loss. Second, we develop a novel pose attention block (PAB) to specially guide the pose-agnostic feature extraction from profile faces. To be more specific, PAB is designed to explicitly help the network to focus on important features along both “channel” and “spatial” dimension while learning discriminative yet pose-invariant features in an embedding subspace. To validate the effectiveness of our proposed method, we conduct experiments on both controlled and in-the-wild benchmarks including Multi-PIE, CFP, IJB-C, and show superiority over the state-of-the-arts.
Moktari Mostofa, Mohammad Saeed Ebrahimi Saadabadi, Sahar Rahimi Malakshan, Nasser M. Nasrabadi
IJCB4
2022 Identical Twins Face Morph Database Generation
abstract
By combining two or more face images of look-alikes, morphed face images are generated to fool Facial Recognition Systems (FRS) into falsely accepting multiple people, leading to failures in security systems. Despite several attempts in the literature, finding pairs of bona fide faces to generate the morphed images is still a challenging problem. In this paper, we morph identical twin pairs to generate extremely difficult morphs for FRS. We first explore three methods of morphed face generation, GAN-based, landmark-based, and a wavelet-based morphing approach. We leverage these methods to generate morphs from the identical twin pairs that retain high similarity to both subjects while resulting in minimal artifacts in the visual domain. To further improve the difficulty of recognizing morphed face images, we perform an ablation study to apply adversarial perturbation to the morphs such that they cannot be detected by trained morph classifiers. The evaluation of the generated identical twin morphed dataset is performed in terms of vulnerability analysis and presentation attack error rates.
Kelsey O'Haire, Sobhan Soleymani, Baaria Chaudhary, Jeremy M. Dawson, Nasser M. Nasrabadi
IJCB5
2022 Landmark Enforcement and Style Manipulation for Generative Morphing
abstract
Morph images threaten Facial Recognition Systems (FRS) by presenting as multiple individuals, allowing an adversary to swap identities with another subject. Morph generation using generative adversarial networks (GANs) results in high-quality morphs unaffected by the spatial artifacts caused by landmark-based methods, but there is an apparent loss in identity with standard GAN-based morphing methods. In this paper, we propose a novel StyleGAN morph generation technique by introducing a landmark enforcement method to resolve this issue. Considering this method, we aim to enforce the landmarks of the morph image to represent the spatial average of the landmarks of the bona fide faces and subsequently the morph images to inherit the geometric identity of both bona fide faces. Exploration of the latent space of our model is conducted using Principal Component Analysis (PCA) to accentuate the effect of both the bona fide faces on the morphed latent representation and address the identity loss issue with latent domain averaging. Additionally, to improve high frequency reconstruction in the morphs, we study the train-ability of the noise input for the StyleGAN2 model.
Samuel Price, Sobhan Soleymani, Nasser M. Nasrabadi
IJCB3
2022 Information Maximization for Extreme Pose Face Recognition
abstract
In this paper, we seek to draw connections between the frontal and profile face images in an abstract embedding space. We exploit this connection using a coupled-encoder network to project frontal/profile face images into a common latent embedding space. The proposed model forces the similarity of representations in the embedding space by maximizing the mutual information between two views of the face. The proposed coupled-encoder benefits from three contributions for matching faces with extreme pose disparities. First, we leverage our pose-aware contrastive learning to maximize the mutual information between frontal and profile representations of identities. Second, a memory buffer, which consists of latent representations accumulated over past iterations, is integrated into the model so it can refer to relatively much more instances than the mini-batch size. Third, a novel pose-aware adversarial domain adaptation method forces the model to learn an asymmetric mapping from profile to frontal representation. In our framework, the coupled-encoder learns to enlarge the margin between the distribution of genuine and imposter faces, which results in high mutual information between different views of the same identity. The effectiveness of the proposed model is investigated through extensive experiments, evaluations, and ablation studies on four benchmark datasets, and comparison with the compelling state-of-the-art algorithms.
Mohammad Saeed Ebrahimi Saadabadi, Sahar Rahimi Malakshan, Sobhan Soleymani, Moktari Mostofa, Nasser M. Nasrabadi
IJCB5
2022 Attention-Based Generative Neural Image Compression on Solar Dynamics Observatory
abstract
NASA’s Solar Dynamics Observatory (SDO) mission gathers 1.4 terabytes of data each day from its geosynchronous orbit in space. SDO data includes images of the Sun captured at different wavelengths, with the primary scientific goal of understanding the dynamic processes governing the Sun. Recently, end-to-end optimized artificial neural networks (ANN) have shown great potential in performing image compression. ANN-based compression schemes have outperformed conventional hand-engineered algorithms for lossy and lossless image compression. We have designed an ad-hoc ANN-based image compression scheme to reduce the amount of data needed to be stored and retrieved on space missions studying solar dynamics. In this work, we propose an attention module to make use of both local and non-local attention mechanisms in an adversarially trained neural image compression network. We have also demonstrated the superior perceptual quality of this neural image compressor. Our proposed algorithm for compressing images downloaded from the SDO spacecraft performs better in rate-distortion tradeoff than the popular currently-in-use image compression codecs such as JPEG and JPEG2000. In addition we have shown that the proposed method outperforms state-of-the art lossy transform coding compression codec, i.e., BPG.
Atefeh Khoshkhahtinat, Piyush M. Mehta, Nasser M. Nasrabadi, Barbara J. Thompson, Michael S. F. Kirk
ICMLA4
2022 Ortho-Shot: Low Displacement Rank Regularization with Data Augmentation for Few-Shot Learning
abstract
In few-shot classification, the primary goal is to learn representations from a few samples that generalize well for novel classes. In this paper, we propose an efficient low displacement rank (LDR) regularization strategy termed Ortho-Shot; a technique that imposes orthogonal regularization on the convolutional layers of a few-shot classifier, which is based on the doubly-block toeplitz (DBT) matrix structure. The regularized convolutional layers of the few-shot classifier enhances model generalization and intra-class feature embeddings that are crucial for few-shot learning. Overfitting is a typical issue for few-shot models, the lack of data diversity inhibits proper model inference which weakens the classification accuracy of few-shot learners to novel classes. In this regard, we broke down the pipeline of the few-shot classifier and established that the support, query and task data augmentation collectively alleviates overfitting in networks. With compelling results, we demonstrated that combining a DBT-based low-rank orthogonal regularizer with data augmentation strategies, significantly boosts the performance of a few-shot classifier. We perform our experiments on the miniImagenet, CIFAR- FS and Stanford datasets with performance values of about 5% when compared to state-of-the-art.
Uche Osahor, Nasser M. Nasrabadi
WACV2
2022 Attribute-Based Deep Periocular Recognition: Leveraging Soft Biometrics to Improve Periocular Recognition
abstract
In recent years, periocular recognition has been developed as a valuable biometric identification approach, especially in wild environments (for example, masked faces due to COVID-19 pandemic) where facial recognition may not be applicable. This paper presents a new deep periocular recognition framework called attribute-based deep periocular recognition (ADPR), which predicts soft biometrics and incorporates the prediction into a periocular recognition algorithm to determine identity from periocular images with high accuracy. We propose an end-to-end framework, which uses several shared convolutional neural network (CNN) layers (a common network) whose output feeds two separate dedicated branches (modality dedicated layers); the first branch classifies periocular images while the second branch predicts soft biometrics. Next, the features from these two branches are fused together for a final periocular recognition. The proposed method is different from existing methods as it not only uses a shared CNN feature space to train these two tasks jointly, but it also fuses predicted soft biometric features with the periocular features in the training step to improve the overall periocular recognition performance. Our proposed model is extensively evaluated using four different publicly available datasets. Experimental results indicate that our soft biometric based periocular recognition approach outperforms other state-of-the-art methods for periocular recognition in wild environments.
Veeru Talreja, Nasser M. Nasrabadi, Matthew C. Valenti
WACV2
2022 Unsupervised Pixel-Wise Hyperspectral Anomaly Detection via Autoencoding Adversarial Networks
abstract
We propose a completely unsupervised pixel-wise anomaly detection (AD) method for hyperspectral images (HSIs). The proposed method consists of three steps called data preparation, reconstruction, and detection. In the data preparation step, we apply a background purification to train the deep network in an unsupervised manner. In the reconstruction step, we propose to use three different deep autoencoding adversarial network (AEAN) models including 1-D-AEAN, 2-D-AEAN, and 3-D-AEAN which are developed for working on spectral, spatial, and joint spectral–spatial domains, respectively. The goal of the AEAN models is to generate synthesized HSIs which are close to real ones. A reconstruction error map (REM) is calculated between the original and the synthesized image pixels. In the detection step, we propose to use a weighted RX (WRX) -based detector in which the pixel weights are obtained according to REM. We compare our proposed method with the classical Reed–Xiaoli (RX), WRX, support vector data description (SVDD)-based, collaborative representation-based detector (CRD), adaptive weight deep belief network (AW-DBN) detector, and deep autoencoder AD (DAEAD) method on real hyperspectral data sets. The experimental results show that the proposed approach outperforms other detectors in the benchmark.
Sertac Arisoy, Nasser M. Nasrabadi, Koray Kayabol
IEEE Geosci. Remote. Sens. Lett.2
2022 MultiSDGAN: Translation of OCT Images to Superresolved Segmentation Labels Using Multi-Discriminators in Multi-Stages
abstract
Optical coherence tomography (OCT) has been identified as a non-invasive and inexpensive imaging modality to discover potential biomarkers for Alzheimer's diagnosis and progress determination. Current hypotheses presume the thickness of the retinal layers, which are analyzable within OCT scans, as an effective biomarker for the presence of Alzheimer's. As a logical first step, this work concentrates on the accurate segmentation of retinal layers to isolate the layers for further analysis. This paper proposes a generative adversarial network (GAN) that concurrently learns to increase the image resolution for higher clarity and then segment the retinal layers. We propose a multi-stage and multi-discriminatory generative adversarial network (MultiSDGAN) specifically for superresolution and segmentation of OCT scans of the retinal layer. The resulting generator is adversarially trained against multiple discriminator networks at multiple stages. We aim to avoid early saturation of generator model training leading to poor segmentation accuracies and enhance the process of OCT domain translation by satisfying all the discriminators in multiple scales. We also investigated incorporating the Dice loss and Structured Similarity Index Measure (SSIM) as additional loss functions to specifically target and improve our proposed GAN architecture's segmentation and superresolution performance, respectively. The ablation study results conducted on our data set suggest that the proposed MultiSDGAN with ten-fold cross-validation (10-CV) provides a reduced equal error rate with 44.24% and 34.09% relative improvements, respectively (p-values of the improvement level tests .01). Furthermore, our experimental results also demonstrate that the addition of the new terms to the loss function improves the segmentation results significantly by relative improvements of 31.33% (p-value .01).
Paria Jeihouni, Omid Dehzangi, Annahita Amireskandari, Ali Rezai, Nasser M. Nasrabadi
IEEE J. Biomed. Health Informatics5
2021 Human Age Estimation from Gene Expression Data using Artificial Neural Networks
abstract
The study of signatures of aging in terms of genomic biomarkers can be uniquely helpful in understanding the mechanisms of aging and developing models to accurately predict the age. Prior studies have employed gene expression and DNA methylation data aiming at accurate prediction of age. In this line, we propose a new framework for human age estimation using information from human dermal fibroblast gene expression data. First, we propose a new spatial representation as well as a data augmentation approach for gene expression data. Next in order to predict the age, we design an architecture of neural network and apply it to this new representation of the original and augmented data, as an ensemble classification approach. Our experimental results suggest the superiority of the proposed framework over state-of-the-art age estimation methods using DNA methylation and gene expression data.
Salman Mohamadi, Nasser M. Nasrabadi, Gianfranco Doretto, Donald A. Adjeroh
BIBM2
2021 Quality Map Fusion for Adversarial Learning
Uche Osahor, Nasser M. Nasrabadi
BMVC2
2021 SuperMix: Supervising the Mixing Data Augmentation
abstract
This paper presents a supervised mixing augmentation method termed SuperMix, which exploits the salient regions within input images to construct mixed training samples. SuperMix is designed to obtain mixed images rich in visual features and complying with realistic image priors. To enhance the efficiency of the algorithm, we develop a variant of the Newton iterative method, 65× faster than gradient descent on this problem. We validate the effectiveness of SuperMix through extensive evaluations and ablation studies on two tasks of object classification and knowledge distillation. On the classification task, SuperMix provides comparable performance to the advanced augmentation methods, such as AutoAugment and RandAugment. In particular, combining SuperMix with RandAugment achieves 78.2% top-1 accuracy on ImageNet with ResNet50. On the distillation task, solely classifying images mixed using the teacher’s knowledge achieves comparable performance to the state-of-the-art distillation methods. Furthermore, on average, incorporating mixed images into the distillation objective improves the performance by 3.4% and 3.1% on CIFAR-100 and ImageNet, respectively. The code is available at https://github.com/alldbi/SuperMix.
Ali Dabouei, Sobhan Soleymani, Fariborz Taherkhani, Nasser M. Nasrabadi
CVPR4
2021 Self-Supervised Wasserstein Pseudo-Labeling for Semi-Supervised Image Classification
abstract
The goal is to use Wasserstein metric to provide pseudo labels for the unlabeled images to train a Convolutional Neural Networks (CNN) in a Semi-Supervised Learning (SSL) manner for the classification task. The basic premise in our method is that the discrepancy between two discrete empirical measures (e.g., clusters) which come from the same or similar distribution is expected to be less than the case where these measures come from completely two different distributions. In our proposed method, we first pre-train our CNN using a self-supervised learning method to make a cluster assumption on the unlabeled images. Next, inspired by the Wasserstein metric which considers the geometry of the metric space to provide a natural notion of similarity between discrete empirical measures, we leverage it to cluster the unlabeled images and then match the clusters to their similar class of labeled images to provide a pseudo label for the data within each cluster. We have evaluated and compared our method with state-of-the-art SSL methods on the standard datasets to demonstrate its effectiveness.
Fariborz Taherkhani, Ali Dabouei, Sobhan Soleymani, Jeremy M. Dawson, Nasser M. Nasrabadi
CVPR5
2021 Adversarially Perturbed Wavelet-based Morphed Face Generation
abstract
Morphing is the process of combining two or more subjects in an image in order to create a new identity which contains features of both individuals. Morphed images can fool Facial Recognition Systems (FRS) into falsely accepting multiple people, leading to failures in national security. As morphed image synthesis becomes easier, it is vital to expand the research community's available data to help combat this dilemma. In this paper, we explore combination of two methods for morphed image generation, those of geometric transformation (warping and blending to create morphed images) and photometric perturbation. We leverage both methods to generate high-quality adversarially perturbed morphs from the FERET, FRGC, and FRLL datasets. The final images retain high similarity to both input subjects while resulting in minimal artifacts in the visual domain. Images are synthesized by fusing the wavelet sub-bands from the two look-alike subjects, and then adversarially perturbed to create highly convincing imagery to deceive both humans and deep morph detectors.
Kelsey O'Haire, Sobhan Soleymani, Baaria Chaudhary, Poorya Aghdaie, Jeremy M. Dawson, Nasser M. Nasrabadi
FG6
2021 Synthesis-Guided Feature Learning for Cross-Spectral Periocular Recognition
abstract
A common yet challenging scenario in periocular biometrics is cross-spectral matching - in particular, the matching of visible wavelength against near-infrared (NIR) periocular images. We propose a novel approach to cross-spectral periocular verification that primarily focuses on learning a mapping from visible and NIR periocular images to a shared latent representational subspace, and supports this effort by simultaneously learning intra-spectral image reconstruction. We show the auxiliary image reconstruction task (and in particular the reconstruction of high-level, semantic features) results in learning a more discriminative, domain-invariant subspace compared to the baseline while incurring no additional computational or memory costs at test-time. The proposed Coupled Conditional Generative Adversarial Network (CoGAN) architecture uses paired generator networks (one operating on visible images and the other on NIR) composed of U-Nets with ResNet-18 encoders trained for feature learning via contrastive loss and for intra-spectral image reconstruction with adversarial, pixel-based, and perceptual reconstruction losses. Moreover, the proposed CoGAN model beats the current state-of-art (SotA) in cross-spectral periocular recognition. On the Hong Kong PolyU benchmark dataset, we achieve 98.65% AUC and 5.14% EER compared to the SotA EER of 8.02%. On the Cross-Eyed dataset, we achieve 99.31 % AUC and 3.99% EER versus SotA EER of 4.39%.
Domenick Poster, Nasser M. Nasrabadi
FG2
2021 Attention Aware Wavelet-based Detection of Morphed Face Images
abstract
Morphed images have exploited loopholes in the face recognition checkpoints, e.g., Credential Authentication Technology (CAT), used by Transportation Security Administration (TSA), which is a non-trivial security concern. To overcome the risks incurred due to morphed presentations, we propose a wavelet-based morph detection methodology which adopts an end-to-end trainable soft attention mechanism. Our attention-based deep neural network (DNN) focuses on the salient Regions of Interest (ROI) which have the most spatial support for morph detector decision function, i.e, morph class binary softmax output. A retrospective of morph synthesizing procedure aids us to speculate the ROI as regions around facial landmarks, particularly for the case of landmark-based morphing techniques. Moreover, our attention-based DNN is adapted to the wavelet space, where inputs of the network are coarse-to-fine spectral representations, 48 stacked wavelet sub-bands to be exact. We evaluate performance of the proposed framework using three datasets, VISAPP17, LMA, and MorGAN. In addition, as attention maps can be a robust indicator whether a probe image under investigation is genuine or counterfeit, we analyze the estimated attention maps for both a bona fide image and its corresponding morphed image. Finally, we present an ablation study on the efficacy of utilizing attention mechanism for the sake of morph detection.
Poorya Aghdaie, Baaria Chaudhary, Sobhan Soleymani, Jeremy M. Dawson, Nasser M. Nasrabadi
IJCB5
2021 FDeblur-GAN: Fingerprint Deblurring using Generative Adversarial Network
abstract
While working with fingerprint images acquired from crime scenes, mobile cameras, or low-quality sensors, it becomes difficult for automated identification systems to verify the identity due to image blur and distortion. We propose a fingerprint deblurring model FDeblur-GAN, based on the conditional Generative Adversarial Networks (cGANs) and multi-stage framework of the stack GAN. Additionally, we integrate two auxiliary sub-networks into the model for the deblurring task. The first sub-network is a ridge extractor model. It is added to generate ridge maps to ensure that fingerprint information and minutiae are preserved in the deblurring process and prevent the model from generating erroneous minutiae. The second sub-network is a verifier that helps the generator to preserve the ID information during the generation process. Using a database of blurred fingerprints and corresponding ridge maps, the deep network learns to deblur from the input blurry samples. We evaluate the proposed method in combination with two different fingerprint matching algorithms. We achieved an accuracy of 95.18% on our fingerprint database for the task of matching deblurred and ground truth fingerprints.
Amol S. Joshi, Ali Dabouei, Jeremy M. Dawson, Nasser M. Nasrabadi
IJCB4
2021 Gan-Based Super-Resolution and Segmentation of Retinal Layers in Optical Coherence Tomography Scans
abstract
In this paper, we design a Generative Adversarial Network (GAN)-based solution for super-resolution and segmentation of optical coherence tomography (OCT) scans of the retinal layers. OCT has been identified as a non-invasive and inexpensive modality of imaging to discover potential biomarkers for the diagnosis and progress determination of neurodegenerative diseases, such as Alzheimer’s Disease (AD). Current hypotheses presume the thickness of the retinal layers, which are analyzable within OCT scans, can be effective biomarkers. As a logical first step, this work concentrates on the challenging task of retinal layer segmentation and also super-resolution for higher clarity and accuracy. We propose a GAN-based segmentation model and evaluate incorporating popular networks, namely, U-Net and ResNet, in the GAN architecture with additional blocks of transposed convolution and sub-pixel convolution for the task of upscaling OCT images from low to high resolution by a factor of four. We also incorporate the Dice loss as an additional reconstruction loss term to improve the performance of this joint optimization task. Our best model configuration empirically achieved the Dice coefficient of 0.867 and mIOU of 0.765.
Paria Jeihouni, Omid Dehzangi, Annahita Amireskandari, Ali Rezai, Nasser M. Nasrabadi
ICIP5
2021 A Large-Scale, Time-Synchronized Visible and Thermal Face Dataset
abstract
Thermal face imagery, which captures the naturally emitted heat from the face, is limited in availability compared to face imagery in the visible spectrum. To help address this scarcity of thermal face imagery for research and algorithm development, we present the DEVCOM Army Research Laboratory Visible-Thermal Face Dataset (ARL-VTF). With over 500,000 images from 395 subjects, the ARL-VTF dataset represents, to the best of our knowledge, the largest collection of paired visible and thermal face images to date. The data was captured using a modern long wave infrared (LWIR) camera mounted alongside a stereo setup of three visible spectrum cameras. Variability in expressions, pose, and eyewear has been systematically recorded. The dataset has been curated with extensive annotations, metadata, and standardized protocols for evaluation. Furthermore, this paper presents extensive benchmark results and analysis on thermal face landmark detection and thermal-to-visible face verification by evaluating state-of-the-art models on the ARL-VTF dataset.
Domenick Poster, Matthew Thielke, Robert Nguyen, Srinivasan Rajaraman, Xing Di, Cedric Nimpa Fondje, Vishal M. Patel, Nathan J. Short, Benjamin S. Riggan, Nasser M. Nasrabadi, Shuowen Hu
WACV10
2021 Mutual Information Maximization on Disentangled Representations for Differential Morph Detection
abstract
In this paper, we present a novel differential morph detection framework, utilizing landmark and appearance disentanglement. In our framework, the face image is represented in the embedding domain using two disentangled but complementary representations. The network is trained by triplets of face images, in which the intermediate image inherits the landmarks from one image and the appearance from the other image. This initially trained network is further trained for each dataset using contrastive representations. We demonstrate that, by employing appearance and landmark disentanglement, the proposed frame-work can provide state-of-the-art differential morph detection performance. This functionality is achieved by the using distances in landmark, appearance, and ID domains. The performance of the proposed framework is evaluated using three morph datasets generated with different methodologies.
Sobhan Soleymani, Ali Dabouei, Fariborz Taherkhani, Jeremy M. Dawson, Nasser M. Nasrabadi
WACV5
2021 Deep Hashing for Secure Multimodal Biometrics
abstract
When compared to unimodal systems, multimodal biometric systems have several advantages, including lower error rate, higher accuracy, and larger population coverage. However, multimodal systems have an increased demand for integrity and privacy because they must store multiple biometric traits associated with each user. In this paper, we present a deep learning framework for feature-level fusion that generates a secure multimodal template from each user's face and iris biometrics. We integrate a deep hashing (binarization) technique into the fusion architecture to generate a robust binary multimodal shared latent representation. Further, we employ a hybrid secure architecture by combining cancelable biometrics with secure sketch techniques and integrate it with a deep hashing framework, which makes it computationally prohibitive to forge a combination of multiple biometrics that passes the authentication. The efficacy of the proposed approach is shown using a multimodal database of face and iris and it is observed that the matching performance is improved due to the fusion of multiple biometrics. Furthermore, the proposed approach also provides cancelability and unlinkability of the templates along with improved privacy of the biometric data. Additionally, we also test the proposed hashing function for an image retrieval application using a benchmark dataset. The main goal of this paper is to develop a method for integrating multimodal fusion, deep hashing, and biometric security, with an emphasis on structural data from modalities like face and iris. The proposed approach is in no way a general biometrics security framework that can be applied to all biometrics modalities, as further research is needed to extend the proposed framework to other unconstrained biometric modalities.
Veeru Talreja, Matthew C. Valenti, Nasser M. Nasrabadi
IEEE Trans. Inf. Forensics Secur.3
2020 Attribute Adaptive Margin Softmax Loss using Privileged Information
Seyed Mehdi Iranmanesh, Ali Dabouei, Nasser M. Nasrabadi
BMVC3
2020 Exploiting Joint Robustness to Adversarial Perturbations
abstract
Recently, ensemble models have demonstrated empirical capabilities to alleviate the adversarial vulnerability. In this paper, we exploit first-order interactions within ensembles to formalize a reliable and practical defense. We introduce a scenario of interactions that certifiably improves the robustness according to the size of the ensemble, the diversity of the gradient directions, and the balance of the member's contribution to the robustness. We present a joint gradient phase and magnitude regularization (GPMR) as a vigorous approach to impose the desired scenario of interactions among members of the ensemble. Through extensive experiments, including gradient-based and gradient-free evaluations on several datasets and network architectures, we validate the practical effectiveness of the proposed approach compared to the previous methods. Furthermore, we demonstrate that GPMR is orthogonal to other defense strategies developed for single classifiers and their combination can further improve the robustness of ensembles.
Ali Dabouei, Sobhan Soleymani, Fariborz Taherkhani, Jeremy M. Dawson, Nasser M. Nasrabadi
CVPR5
2020 Transporting Labels via Hierarchical Optimal Transport for Semi-Supervised Learning
Fariborz Taherkhani, Ali Dabouei, Sobhan Soleymani, Jeremy M. Dawson, Nasser M. Nasrabadi
ECCV (4)5
2020 Cross-Spectral Iris Matching Using Conditional Coupled GAN
abstract
Cross-spectral iris recognition is emerging as a promising biometric approach to authenticating the identity of individuals. However, matching iris images acquired at different spectral bands shows significant performance degradation when compared to single-band near-infrared (NIR) matching due to the spectral gap between iris images obtained in the NIR and visual-light (VIS) spectra. Although researchers have recently focused on deep-learning-based approaches to recover invariant representative features for more accurate recognition performance, the existing methods cannot achieve the expected accuracy required for commercial applications. Hence, in this paper, we propose a conditional coupled generative adversarial network (CpGAN) architecture for cross-spectral iris recognition by projecting the VIS and NIR iris images into a low-dimensional embedding domain to explore the hidden relationship between them. The conditional CpGAN framework consists of a pair of GAN-based networks, one responsible for retrieving images in the visible domain and other responsible for retrieving images in the NIR domain. Both networks try to map the data into a common embedding subspace to ensure maximum pair-wise similarity between the feature vectors from the two iris modalities of the same subject. To prove the usefulness of our proposed approach, extensive experimental results obtained on the PolyU dataset are compared to existing state-of-the-art cross-spectral recognition methods.
Moktari Mostofa, Fariborz Taherkhani, Jeremy M. Dawson, Nasser M. Nasrabadi
IJCB4
2020 PF -cpGAN: Profile to Frontal Coupled GAN for Face Recognition in the Wild
abstract
In recent years, due to the emergence of deep learning, face recognition has achieved exceptional success. However, many of these deep face recognition models perform relatively poorly in handling profile faces compared to frontal faces. The major reason for this poor performance is that it is inherently difficult to learn large pose invariant deep representations that are useful for profile face recognition. In this paper, we hypothesize that the profile face domain possesses a gradual connection with the frontal face domain in the deep feature space. We look to exploit this connection by projecting the profile faces and frontal faces into a common latent space and perform verification or retrieval in the latent domain. We leverage a coupled generative adversarial network (cpGAN) structure to find the hidden relationship between the profile and frontal images in a latent common embedding subspace. Specifically, the cp-GAN framework consists of two GAN-based sub-networks, one dedicated to the frontal domain and the other dedicated to the profile domain. Each sub-network tends to find a projection that maximizes the pair-wise correlation between two feature domains in a common embedding feature subspace. The efficacy of our approach compared with the state-of-the-art is demonstrated using the CFp, CMU Multi-PIE, IJB-A, and IJB-C datasets.
Fariborz Taherkhani, Veeru Talreja, Jeremy M. Dawson, Matthew C. Valenti, Nasser M. Nasrabadi
IJCB5
2020 Efficient Oct Image Segmentation Using Neural Architecture Search
abstract
In this work, we propose a Neural Architecture Search (NAS) for retinal layer segmentation in Optical Coherence Tomography (OCT) scans. We incorporate the Unet architecture in the NAS framework as its backbone for the segmentation of the retinal layers in our collected and preprocessed OCT image dataset. At the pre-processing stage, we conduct super resolution and image processing techniques on the raw OCT scans to improve the quality of the raw images. For our search strategy, different primitive operations are suggested to find the down- & up-sampling cell blocks, and the binary gate method is applied to make the search strategy practical for the task in hand. We empirically evaluated our method on our in-house OCT dataset. The experimental results demonstrate that the self-adapting NAS-Unet architecture substantially outperformed the competitive human-designed architecture by achieving 95.4% in mean Intersection over Union metric and 78.7% in Dice similarity coefficient.
Saba Heidari Gheshlaghi, Omid Dehzangi, Ali Dabouei, Annahita Amireskandari, Ali Rezai, Nasser M. Nasrabadi
ICIP6
2020 OCT Image Segmentation Using Neural Architecture Search and SRGAN
abstract
Medical image segmentation is a critical field in the domain of computer vision and with the growing acclaim of deep learning based models, research in this field is constantly expanding. Optical coherence tomography (OCT) is a non-invasive method that scans the human's retina with depth. It has been hypothesized that the thickness of the retinal layers extracted from OCTs could be an efficient and effective biomarker for early diagnosis of AD. In this work, we aim to design a self-training model architecture for the task of segmenting the retinal layers in OCT scans. Neural architecture search (NAS) is a subfield of AutoML domain, which has a significant impact on improving the accuracy of machine vision tasks. We integrate the NAS algorithm with a Unet auto-encoder architecture as its backbone. Then, we employ our proposed model to segment the retinal nerve fiber layer in our preprocessed OCT images with the aim of AD diagnosis. In this work, we trained a super-resolution generative adversarial network on the raw OCT scans to improve the quality of the images before the modeling stage. In our architecture search strategy, different primitive operations suggested to find down- & up-sampling Unet cell blocks and the binary gate method has been applied to make the search strategy more practical. Our architecture search method is empirically evaluated by training on the Unet and NAS-Unet from scratch. Specifically, the proposed NAS-Unet training significantly outperforms the baseline human-designed architecture by achieving 95.1% in the mean Intersection over Union metric and 79.1% in the Dice similarity coefficient.
Omid Dehzangi, Saba Heidari Gheshlaghi, Annahita Amireskandari, Nasser M. Nasrabadi, Ali Rezai
ICPR4
2020 XGBoost to Interpret the Opioid Patients' State Based on Cognitive and Physiological Measures
abstract
Dealing with opioid addiction and its long-term consequences is of great importance, as the addiction to opioids is emerged gradually, and established strongly in a given patient's body. Based on recent research, quitting the opioid requires clinicians to arrange a gradual plan for the patients who deal with the difficulties of overcoming addiction. This, in turn, necessitates observing the patients' wellness periodically, which is conventionally made by setting clinical appointments. With the advent of wearable sensors continuous patient monitoring becomes possible. However, the data collected through the sensors is pervasively noisy, where using sensors with different sampling frequency challenges the data processing. In this work, we handle this problem by using data from cognitive tests, along with heart rate (HR) and heart rate variability (HRV). The proposed recipe enables us to interpret the data as a feature space, where we can predict the wellness of the opioid patients by employing extreme gradient boosting (XGBoost), which results in 96.12% average accuracy of prediction as the best achieved performance.
Omid Dehzangi, Arash Shokouhmand, Paria Jeihouni, Jad Ramadan, Victor S. Finomore Jr., Nasser M. Nasrabadi, Ali Rezai
ICPR6
2020 Super-resolution Guided Pore Detection for Fingerprint Recognition
abstract
Performance of fingerprint recognition algorithms substantially rely on fine features extracted from fingerprints. Apart from minutiae and ridge patterns, pore features have proven to be usable for fingerprint recognition. Although features from minutiae and ridge patterns are quite attainable from low-resolution images, using pore features is practical only if the fingerprint image is of high resolution which necessitates a model that enhances the image quality of the conventional 500 ppi legacy fingerprints preserving the fine details. To find a solution for recovering pore information from low-resolution fingerprints, we adopt a joint learning-based approach that combines both super-resolution and pore detection networks. Our modified single image Super-Resolution Generative Adversarial Network (SRGAN) framework helps to reliably reconstruct high-resolution fingerprint samples from low-resolution ones assisting the pore detection network to identify pores with a high accuracy. The network jointly learns a distinctive feature representation from a real low-resolution fingerprint sample and successfully synthesizes a high-resolution sample from it. To add discriminative information and uniqueness for all the subjects, we have integrated features extracted from a deep fingerprint verifier with the SRGAN quality discriminator. We also add ridge reconstruction loss, utilizing ridge patterns to make the best use of extracted features. Our proposed method solves the recognition problem by improving the quality of fingerprint images. High recognition accuracy of the synthesized samples that is close to the accuracy achieved using the original high-resolution images validate the effectiveness of our proposed model.
Syeda Nyma Ferdous, Ali Dabouei, Jeremy M. Dawson, Nasser M. Nasrabadi
ICPR4
2020 SmoothFool: An Efficient Framework for Computing Smooth Adversarial Perturbations
abstract
Deep neural networks are susceptible to adversarial manipulations in the input domain. The extent of vulnerability has been explored intensively in cases of ℓp-bounded and ℓp-minimal adversarial perturbations. However, the vulnerability of DNNs to adversarial perturbations with specific statistical properties or frequency-domain characteristics has not been sufficiently explored. In this paper, we study the smoothness of perturbations and propose Smooth-Fool, a general and computationally efficient framework for computing smooth adversarial perturbations. Through extensive experiments, we validate the efficacy of the proposed method for both the white-box and black-box attack scenarios. In particular, we demonstrate that: (i) there exist extremely smooth adversarial perturbations for well-established and widely used network architectures, (ii) smoothness significantly enhances the robustness of perturbations against state-of-the-art defense mechanisms, (iii) smoothness improves the transferability of adversarial perturbations across both data points and network architectures, and (iv) class categories exhibit a variable range of susceptibility to smooth perturbations. Our results suggest that smooth APs can play a significant role in exploring the vulnerability extent of DNNs to adversarial examples. The code is available at https://github.com/alldbi/SmoothFool.
Ali Dabouei, Sobhan Soleymani, Fariborz Taherkhani, Jeremy M. Dawson, Nasser M. Nasrabadi
WACV5
2020 Boosting Deep Face Recognition via Disentangling Appearance and Geometry
abstract
In this paper, we propose a framework for disentangling the appearance and geometry representations in the face recognition task. To provide supervision for this aim, we generate geometrically identical faces by incorporating spatial transformations. We demonstrate that the proposed approach enhances the performance of deep face recognition models by assisting the training process in two ways. First, it enforces the early and intermediate convolutional layers to learn more representative features that satisfy the properties of disentangled embeddings. Second, it augments the training set by altering faces geometrically. Through extensive experiments, we demonstrate that integrating the proposed approach into state-of-the-art face recognition methods effectively improves their performance on challenging datasets, such as LFW, YTF, and MegaFace. Both theoretical and practical aspects of the method are analyzed rigorously by concerning ablation studies and knowledge transfer tasks. Furthermore, we show that the knowledge leaned by the proposed method can favor other face-related tasks, such as attribute prediction.
Ali Dabouei, Fariborz Taherkhani, Sobhan Soleymani, Jeremy M. Dawson, Nasser M. Nasrabadi
WACV5
2020 Robust Facial Landmark Detection via Aggregation on Geometrically Manipulated Faces
abstract
In this work, we present a practical approach to the problem of facial landmark detection. The proposed method can deal with large shape and appearance variations under the rich shape deformation. To handle the shape variations we equip our method with the aggregation of manipulated face images. The proposed framework generates different manipulated faces using only one given face image. The approach utilizes the fact that small but carefully crafted geometric manipulation in the input domain can fool deep face recognition models. We propose three different approaches to generate manipulated faces in which two of them perform the manipulations via adversarial attacks and the other one uses known transformations. Aggregating the manipulated faces provides a more robust landmark detection approach which is able to capture more important deformations and variations of the face shapes. Our approach is demonstrated its superiority compared to the state-of-the-art method on benchmark datasets AFLW, 300-W, and COFW.
Seyed Mehdi Iranmanesh, Ali Dabouei, Sobhan Soleymani, Hadi Kazemi, Nasser M. Nasrabadi
WACV5
2020 Preference-Based Image Generation
abstract
Deep generative models are a set of promising methods, that are able to model complex data and generate new samples. In principle, they learn to map a random latent code sampled from a prior distribution into a high dimensional data space, such as image space. However, these models have limited utilities as the user has minimal control over what the network produces. Despite the success of some recent work in learning an interpretable latent code, the field still lacks a coherent framework to learn a fully interpretable latent code, without any random part for sample diversity. Consequently, it is generally hard, if not impossible, for a non-expert user to produce a desired image by tuning the random and interpretable parts of the latent code. In this paper, we introduce the Preference-Based Image Generation (PbIG), a new method to retrieve the corresponding latent code of the user's mental image. We propose to adopt preference-based reinforcement learning, which learns from a user's judgment of the generated images by a pre-trained generative model. Since the proposed method is completely decoupled from the training stage of the underlying generative models, it can easily be adopted by any method, such as GANs and VAEs. We evaluate the effectiveness of PbIG framework using a set of experiments on baseline datasets using a pretraind StackGAN++.
Hadi Kazemi, Fariborz Taherkhani, Nasser M. Nasrabadi
WACV3
2020 Coupled generative adversarial network for heterogeneous face recognition
Seyed Mehdi Iranmanesh, Benjamin S. Riggan, Shuowen Hu, Nasser M. Nasrabadi
Image Vis. Comput.4
2020 Supervised Deep Sparse Coding Networks for Image Classification
abstract
In this paper, we propose a novel deep sparse coding network (SCN) capable of efficiently adapting its own regularization parameters for a given application. The network is trained end-to-end with a supervised task-driven learning algorithm via error backpropagation. During training, the network learns both the dictionaries and the regularization parameters of each sparse coding layer so that the reconstructive dictionaries are smoothly transformed into increasingly discriminative representations. In addition, the adaptive regularization also offers the network more flexibility to adjust sparsity levels. Furthermore, we have devised a sparse coding layer utilizing a 'skinny' dictionary. Integral to computational efficiency, these skinny dictionaries compress the high dimensional sparse codes into lower dimensional structures. The adaptivity and discriminability of our fifteen-layer sparse coding network are demonstrated on five benchmark datasets, namely Cifar-10, Cifar-100, STL-10, SVHN and MNIST, most of which are considered difficult for sparse coding models. Experimental results show that our architecture overwhelmingly outperforms traditional one-layer sparse coding architectures while using much fewer parameters. Moreover, our multilayer architecture exploits the benefits of depth with sparse coding's characteristic ability to operate on smaller datasets. In such data-constrained scenarios, our technique demonstrates highly competitive performance compared to the deep neural networks.
Xiaoxia Sun, Nasser M. Nasrabadi, Trac D. Tran
IEEE Trans. Image Process.2
2019 Matrix Completion for Graph-Based Deep Semi-Supervised Learning
abstract
Convolutional Neural Networks (CNNs) have provided promising achievements for image classification problems. However, training a CNN model relies on a large number of labeled data. Considering the vast amount of unlabeled data available on the web, it is important to make use of these data in conjunction with a small set of labeled data to train a deep learning model. In this paper, we introduce a new iterative Graph-based Semi-Supervised Learning (GSSL) method to train a CNN-based classifier using a large amount of unlabeled data and a small amount of labeled data. In this method, we first construct a similarity graph in which the nodes represent the CNN features corresponding to data points (labeled and unlabeled) while the edges tend to connect the data points with the same class label. In this graph, the missing label of unsupervised nodes is predicted by using a matrix completion method based on rank minimization criterion. In the next step, we use the constructed graph to calculate triplet regularization loss which is added to the supervised loss obtained by initially labeled data to update the CNN network parameters.
Fariborz Taherkhani, Hadi Kazemi, Nasser M. Nasrabadi
AAAI3
2019 Learning to Authenticate with Deep Multibiometric Hashing and Neural Network Decoding
abstract
In this paper, we propose a novel multimodal deep hashing neural decoder (MDHND) architecture, which integrates a deep hashing framework with a neural network decoder (NND) to create an effective multibiometric authentication system. The MDHND consists of two separate modules: a multimodal deep hashing (MDH) module, which is used for feature-level fusion and binarization of multiple biometrics, and a neural network decoder (NND) module, which is used to refine the intermediate binary codes generated by the MDH and compensate for the difference between enrollment and probe biometrics (variations in pose, illumination, etc.). Use of NND helps to improve the performance of the overall multimodal authentication system. The MDHND framework is trained in 3 steps using joint optimization of the two modules. In Step 1, the MDH parameters are trained and learned to generate a shared multimodal latent code; in Step 2, the latent codes from Step 1 are passed through a conventional error-correcting code (ECC) decoder to generate the ground truth to train a neural network decoder (NND); in Step 3, the NND decoder is trained using the ground truth from Step 2 and the MDH and NND are jointly optimized. Experimental results on a standard multimodal dataset demonstrate the superiority of our method relative to other current multimodal authentication systems.
Veeru Talreja, Sobhan Soleymani, Matthew C. Valenti, Nasser M. Nasrabadi
ICC4
2019 A Weakly Supervised Fine Label Classifier Enhanced by Coarse Supervision
abstract
Objects are usually organized in a hierarchical structure in which each coarse category (e.g., big cat) corresponds to a superclass of several fine categories (e.g., cheetah, leopard). The objects grouped within the same coarse category, but in different fine categories, usually share a set of global visual features; however, these objects have distinctive local properties that characterize them at a fine level. This paper addresses the challenge of fine image classification in a weakly supervised fashion, whereby a subset of images is tagged by fine labels, while the remaining are tagged by coarse labels. We propose a new deep model that leverages coarse images to improve the classification performance of fine images within the coarse category. Our model is an end to end framework consisting of a Convolutional Neural Network (CNN) which uses both fine and coarse images to tune its parameters. The CNN outputs are then fanned out into two separate branches such that the first branch uses a supervised low rank self expressive layer to project the CNN outputs to the low rank subspaces to capture the global structures for the coarse classification, while the other branch uses a supervised sparse self expressive layer to project them to the sparse subspaces to capture the local structures for the fine classification. Our deep model uses coarse images in conjunction with fine images to jointly explore the low rank and sparse subspaces by sharing the parameters during the training which causes the data points obtained by the CNN to be well-projected to both sparse and low rank subspaces for classification.
Fariborz Taherkhani, Hadi Kazemi, Ali Dabouei, Jeremy M. Dawson, Nasser M. Nasrabadi
ICCV5
2019 Fast Geometrically-Perturbed Adversarial Faces
abstract
The state-of-the-art performance of deep learning algorithms has led to a considerable increase in the utilization of machine learning in security-sensitive and critical applications. However, it has recently been shown that a small and carefully crafted perturbation in the input space can completely fool a deep model. In this study, we explore the extent to which face recognition systems are vulnerable to geometrically-perturbed adversarial faces. We propose a fast landmark manipulation method for generating adversarial faces, which is approximately 200 times faster than the previous geometric attacks and obtains 99.86% success rate on the state-of-the-art face recognition models. To further force the generated samples to be natural, we introduce a second attack constrained on the semantic structure of the face which has the half speed of the first attack with the success rate of 99.96%. Both attacks are extremely robust against the state-of-the-art defense methods with the success rate of equal or greater than 53.59%. Code is available at https://github.com/alldbi/FLM.
Ali Dabouei, Sobhan Soleymani, Jeremy M. Dawson, Nasser M. Nasrabadi
WACV4
2019 Style and Content Disentanglement in Generative Adversarial Networks
abstract
Disentangling factors of variation within data has become a very challenging problem for image generation tasks. Current frameworks for training a Generative Adversarial Network (GAN), learn to disentangle the representations of the data in an unsupervised fashion and capture the most significant factors of the data variations. However, these approaches ignore the principle of content and style disentanglement in image generation, which means their learned latent code may alter the content and style of the generated images at the same time. This paper describes the Style and Content Disentangled GAN (SC-GAN), a new unsupervised algorithm for training GANs that learns disentangled style and content representations of the data. We assume that the representation of an image can be decomposed into a content code that represents the geometrical information of the data, and a style code that captures textural properties. Consequently, by fixing the style portion of the latent representation, we can generate diverse images in a particular style. Reversely, we can set the content code and generate a specific scene in a variety of styles. The proposed SC-GAN has two components: a content code which is the input to the generator, and a style code which modifies the scene style through modification of the Adaptive Instance Normalization (AdaIN) layers' parameters. We evaluate the proposed SC-GAN framework on a set of baseline datasets.
Hadi Kazemi, Seyed Mehdi Iranmanesh, Nasser M. Nasrabadi
WACV3
2018 Convolutional Neural Networks for Aerial Multi-Label Pedestrian Detection
abstract
The low resolution of objects of interest in aerial images makes pedestrian detection and action detection extremely challenging tasks. Furthermore, using deep convolutional neural networks to process large images can be demanding in terms of computational requirements. In order to alleviate these challenges, we propose a two-step, yes and no question answering framework to find specific individuals doing one or multiple specific actions in aerial images. First, a deep object detector, Single Shot Multibox Detector (SSD), is used to generate object proposals from small aerial images. Second, another deep network, is used to learn a latent common sub-space which associates the high resolution aerial imagery and the pedestrian action labels that are provided by the human-based sources.
Amir Soleimani, Nasser M. Nasrabadi
FUSION2
2018 Generalized Bilinear Deep Convolutional Neural Networks for Multimodal Biometric Identification
abstract
In this paper, we propose to employ a bank of modality-dedicated Convolutional Neural Networks (CNNs), fuse, train, and optimize them together for person classification tasks. A modality-dedicated CNN is used for each modality to extract modality-specific features. We demonstrate that, rather than spatial fusion at the convolutional layers, the fusion can be performed on the outputs of the fully-connected layers of the modality-specific CNNs without any loss of performance and with significant reduction in the number of parameters. We show that, using multiple CNNs with multimodal fusion at the feature-level, we significantly outperform systems that use unimodal representation. We study weighted feature, bilinear, and compact bilinear feature-level fusion algorithms for multimodal biometric person identification. Finally, We propose generalized compact bilinear fusion algorithm to deploy both the weighted feature fusion and compact bilinear schemes. We provide the results for the proposed algorithms on three challenging databases: CMU Multi-PIE, BioCop, and BIOMDATA.
Sobhan Soleymani, Amirsina Torfi, Jeremy M. Dawson, Nasser M. Nasrabadi
ICIP4
2018 Supervised Deep Sparse Coding Networks
abstract
In this paper, we present the deep sparse coding network (DSCN) - a novel deep learning framework that encodes intermediate representations with nonnegative sparse coding. DSCN is constructed from a cascade of bottleneck modules, each of which consists of two sparse coding layers with relatively wide and slim dictionaries that are specialized to produce high dimensional discriminative features and low dimensional clustered representations, respectively. During training, all dictionaries at all depth levels along with all regularization parameters are optimized jointly with an end-to-end supervised learning algorithm based on multilevel optimization. The effectiveness of the proposed DSCN with seven bottleneck modules11Consisting 14 sparse coding layers. is verified on several popular benchmark datasets Remarkably, with few parameters to learn, our SCN achieves 5.81 % and 19.93% classification error rate on CIFAR-10 and CIFAR-100, respectively.
Xiaoxia Sun, Nasser M. Nasrabadi, Trac D. Tran
ICIP2
2018 Text-Independent Speaker Verification Using 3D Convolutional Neural Networks
abstract
In this paper, a novel method using 3D Convolutional Neural Network (3D-CNN) architecture has been proposed for speaker verification in the text-independent setting. One of the main challenges is the creation of the speaker models. Most of the previously-reported approaches create speaker models based on averaging the extracted features from utterances of the speaker, which is known as the d-vector system. In our paper, we propose an adaptive feature learning by utilizing the 3D-CNN s for direct speaker model creation in which, for both development and enrollment phases, an identical number of spoken utterances per speaker is fed to the network for representing the speakers' utterances and creation of the speaker model. This leads to simultaneously capturing the speaker-related information and building a more robust system to cope with within-speaker variation. We demonstrate that the proposed method significantly outperforms the traditional d-vector verification system. Moreover, the proposed system can also be an alternative to the traditional d-vector system which is a one-shot speaker modeling system by utilizing 3D-CNNs.
Amirsina Torfi, Jeremy M. Dawson, Nasser M. Nasrabadi
ICME3
2018 Multi-Level Feature Abstraction from Convolutional Neural Networks for Multimodal Biometric Identification
abstract
In this paper, we propose a deep multimodal fusion network to fuse multiple modalities (face, iris, and fingerprint) for person identification. The proposed deep multimodal fusion algorithm consists of multiple streams of modality-specific Convolutional Neural Networks (CNNs), which are jointly optimized at multiple feature abstraction levels. Multiple features are extracted at several different convolutional layers from each modality-specific CNN for joint feature fusion, optimization, and classification. Features extracted at different convolutional layers of a modality-specific CNN represent the input at several different levels of abstract representations. We demonstrate that an efficient multimodal classification can be accomplished with a significant reduction in the number of network parameters by exploiting these multi-level abstract representations extracted from all the modality-specific CNNs. We demonstrate an increase in multimodal person identification performance by utilizing the proposed multi-level feature abstract representations in our multimodal fusion, rather than using only the features from the last layer of each modality-specific CNNs. We show that our deep multi-modal CNNs with multimodal fusion at several different feature level abstraction can significantly outperform the unimodal representation accuracy. We also demonstrate that the joint optimization of all the modality-specific CNNs excels the score and decision level fusions of independently optimized CNNs.
Sobhan Soleymani, Ali Dabouei, Hadi Kazemi, Jeremy M. Dawson, Nasser M. Nasrabadi
ICPR5
2018 Unsupervised Image-to-Image Translation Using Domain-Specific Variational Information Bound
abstract
Unsupervised image-to-image translation is a class of computer vision problems which aims at modeling conditional distribution of images in the target domain, given a set of unpaired images in the source and target domains. An image in the source domain might have multiple representations in the target domain. Therefore, ambiguity in modeling of the conditional distribution arises, specially when the images in the source and target domains come from different modalities. Current approaches mostly rely on simplifying assumptions to map both domains into a shared-latent space. Consequently, they are only able to model the domain-invariant information between the two modalities. These approaches cannot model domain-specific information which has no representation in the target domain. In this work, we propose an unsupervised image-to-image translation framework which maximizes a domain-specific variational information bound and learns the target domain-invariant representation of the two domain. The proposed framework makes it possible to map a single source image into multiple images in the target domain, utilizing several target domain-specific codes sampled randomly from the prior distribution, or extracted from reference images.
Hadi Kazemi, Sobhan Soleymani, Fariborz Taherkhani, Seyed Mehdi Iranmanesh, Nasser M. Nasrabadi
NeurIPS5
2018 An Order Preserving Bilinear Model for Person Detection in Multi-Modal Data
abstract
We propose a new order preserving bilinear framework that exploits low-resolution video for person detection in a multi-modal setting using deep neural networks. In this setting cameras are strategically placed such that less robust sensors, e.g. geophones that monitor seismic activity, are located within the field of views (FOVs) of cameras. The primary challenge is being able to leverage sufficient information from videos where there are less than 40 pixels on targets, while also taking advantage of less discriminative information from other modalities, e.g. seismic. Unlike state-of-the-art methods, our bilinear framework retains spatio-temporal order when computing the vector outer products between pairs of features. Despite the high dimensionality of these outer products, we demonstrate that our order preserving bilinear framework yields better performance than recent orderless bilinear models and alternative fusion methods. Code is available at https://github.com/oulutan/OP-Bilinear-Model.
Oytun Ulutan, Benjamin S. Riggan, Nasser M. Nasrabadi, B. S. Manjunath
WACV3
2018 Semi-supervised Deep Domain Adaptation via Coupled Neural Networks
abstract
Domain adaptation is a promising technique when addressing limited or no labeled target data by borrowing well-labeled knowledge from the auxiliary source data. Recently, researchers have exploited multi-layer structures for discriminative feature learning to reduce the domain discrepancy. However, there are limited research efforts on simultaneously building a deep structure and a discriminative classifier over both labeled source and unlabeled target. In this paper, we propose a semi-supervised deep domain adaptation framework, in which the multi-layer feature extractor and a multi-class classifier are jointly learned to benefit from each other. Specifically, we develop a novel semi-supervised class-wise adaptation manner to fight off the conditional distribution mismatch between two domains by assigning a probabilistic label to each target sample, i.e., multiple class labels with different probabilities. Furthermore, a multi-class classifier is simultaneously trained on labeled source and unlabeled target samples in a semi-supervised fashion. In this way, the deep structure can formally alleviate the domain divergence and enhance the feature transferability. Experimental evaluations on several standard cross-domain benchmarks verify the superiority of our proposed approach.
Zhengming Ding, Nasser M. Nasrabadi, Yun Fu 0001
IEEE Trans. Image Process.2
2017 Accurate and Timely Situation Awareness Retrieval from a Bandwidth Constrained Camera Network
abstract
Wireless cameras can be used to gather situation awareness information (e.g., humans in distress) in disaster recovery scenarios. However, blindly sending raw video streams from such cameras, to an operations center or controller can be prohibitive in terms of bandwidth. Further, these raw streams could contain either redundant or irrelevant information. Thus, we ask "how do we extract accurate situation awareness information from such camera nodes and send it in a timely manner, back to the operations center?" Towards this, we design ACTION, a framework that (a) detects objects of interest (e.g., humans) from the video streams, (b) combines these streams intelligently to eliminate redundancies and (c) transmits only parts of the feeds that are sufficient in achieving a desired detection accuracy to the controller. ACTION uses small amounts of metadata to determine if the objects from different camera feeds are the same. A resource-aware greedy algorithm is used to select a subset of video feeds that are associated with the same object, so as to provide a desired accuracy, for being sent to the operations center. Our evaluations show that ACTION helps reduce the network usage up to threefold, and yet achieves a high detection accuracy of ≈ 90%.
Tuan Dao, Amit K. Roy-Chowdhury, Nasser M. Nasrabadi, Srikanth V. Krishnamurthy, Prasant Mohapatra, Lance M. Kaplan
MASS3
2016 Learning a Mixture of Deep Networks for Single Image Super-Resolution
Ding Liu 0001, Nasser M. Nasrabadi, Thomas S. Huang
ACCV (3)3
2016 Task-driven deep transfer learning for image classification
abstract
Transfer learning tends to be a powerful tool that can mitigate the divergence across different domains through knowledge transfer. Recent research efforts on transfer learning have exploited deep neural network (NN) structures for discriminative feature representation to better tackle cross-domain disparity. However, few of these techniques are able to jointly learn deep features and train a classifier in a unified transfer learning framework. To this end, we design a task-driven deep transfer learning framework for image classification, where the deep feature and classifier are obtained simultaneously for optimal classification performance. Therefore, the proposed deep structure can generate more discriminative features by using the classifier performance as a guide. Furthermore, the classifier performance is increased since it is optimized on a more discriminative deep feature. The developed supervised formulation is a task-driven scheme, which will provide better learned features for the classification task. By giving pseudo labels for target data, we can facilitate the knowledge transfer from source to target through the deep structures. Experimental results witness the superiority of our proposed algorithm by comparing with other ones.
Zhengming Ding, Nasser M. Nasrabadi, Yun Fu 0001
ICASSP2
2016 Sparse coding with fast image alignment via large displacement optical flow
abstract
Sparse representation-based classifiers have shown outstanding accuracy and robustness in image classification tasks even with the presence of intense noise and occlusion. However, it has been discovered that the performance degrades significantly either when test image is not aligned with the dictionary atoms or the dictionary atoms themselves are not aligned with each other, in which cases the sparse linear representation assumption fails. In this paper, having both training and test images misaligned, we introduce a novel sparse coding framework that is able to efficiently adapt the dictionary atoms to the test image via large displacement optical flow. In the proposed algorithm, every dictionary atom is automatically aligned with the input image and the sparse code is then recovered using the adapted dictionary atoms. A corresponding supervised dictionary learning algorithm is also developed for the proposed framework. Experimental results on digit datasets recognition verify the efficacy and robustness of the proposed algorithm.
Xiaoxia Sun, Nasser M. Nasrabadi, Trac D. Tran
ICASSP2
2016 Multimodal Task-Driven Dictionary Learning for Image Classification
abstract
Dictionary learning algorithms have been successfully used for both reconstructive and discriminative tasks, where an input signal is represented with a sparse linear combination of dictionary atoms. While these methods are mostly developed for single-modality scenarios, recent studies have demonstrated the advantages of feature-level fusion based on the joint sparse representation of the multimodal inputs. In this paper, we propose a multimodal task-driven dictionary learning algorithm under the joint sparsity constraint (prior) to enforce collaborations among multiple homogeneous/heterogeneous sources of information. In this task-driven formulation, the multimodal dictionaries are learned simultaneously with their corresponding classifiers. The resulting multimodal dictionaries can generate discriminative latent features (sparse codes) from the data that are optimized for a given task such as binary or multiclass classification. Moreover, we present an extension of the proposed formulation using a mixed joint and independent sparsity prior, which facilitates more flexible fusion of the modalities at feature level. The efficacy of the proposed algorithms for multimodal classification is illustrated on four different applications--multimodal face recognition, multi-view face recognition, multi-view action recognition, and multimodal biometric recognition. It is also shown that, compared with the counterpart reconstructive-based dictionary learning algorithms, the task-driven formulations are more computationally efficient in the sense that they can be equipped with more compact dictionaries and still achieve superior performance.
Soheil Bahrampour, Nasser M. Nasrabadi, Asok Ray, W. Kenneth Jenkins
IEEE Trans. Image Process.2
2015 Incremental Dictionary Learning for Unsupervised Domain Adaptation
abstract
Domain adaptation (DA) methods attempt to solve the domain mismatch problem between source and target data. In this paper, we propose an incremental dictionary learning method where some target data called supportive samples are selected to assist adaptation. Supportive samples are close to the source domain and have two properties: first, their predicted class labels are reliable and can be used for building more discriminative classification models; second, they act as a bridge to connect the two domains and reduce the domain mismatch. Theoretical analysis shows that both properties are important for adaptation, enabling the idea of adding supportive samples to the source domain. A stopping criterion is designed to guarantee that the domain mismatch decreases monotonically during adaptation. Experimental results on several widely used visual datasets show that the proposed approach performs better than many state-of-the-art methods.
Boyu Lu, Rama Chellappa, Nasser M. Nasrabadi
BMVC3
2015 Kernel task-driven dictionary learning for hyperspectral image classification
abstract
Dictionary learning algorithms have been successfully used in both reconstructive and discriminative tasks, where the input signal is represented by a linear combination of a few dictionary atoms. While these methods are usually developed under ℓ1sparsity constrain (prior) in the input domain, recent studies have demonstrated the advantages of sparse representation using structured sparsity priors in the kernel domain. In this paper, we propose a supervised dictionary learning algorithm in the kernel domain for hyperspectral image classification. In the proposed formulation, the dictionary and classifier are obtained jointly for optimal classification performance. The supervised formulation is task-driven and provides learned features from the hyperspectral data that are well suited for the classification task. Moreover, the proposed algorithm uses a joint (ℓ12) sparsity prior to enforce collaboration among the neighboring pixels. The simulation results illustrate the efficiency of the proposed dictionary learning algorithm.
Soheil Bahrampour, Nasser M. Nasrabadi, Asok Ray, W. Kenneth Jenkins
ICASSP2
2015 Multi-sensor classification via sparsity-based representation with low-rank interference
abstract
In this paper, we propose a general collaborative sparse representation framework for multi-sensor classification which exploits correlation as well as complementary information among homogeneous and heterogeneous sensors while simultaneously extracting the low-rank interference term. Specifically, we observe that incorporating the noise or interfered signal as a low-rank component is essential in a multi-sensor problem when multiple co-located sources/sensors simultaneously record the same physical event. We further extend our frameworks to kernelized models which rely on sparsely representing a test sample in terms of all the training samples in a feature space induced by a kernel function. A fast and efficient algorithm based on alternative direction method is proposed where its convergence to optimal solution is guaranteed. Extensive experiments are conducted on a real data set for a multi-sensor classification problem focusing on discriminating between human and animal footsteps. Results are compared with the conventional classifiers and existing sparsity-based representation methods to verify the effectiveness of our proposed models.
Minh Dao, Nasser M. Nasrabadi, Trac D. Tran
ICASSP2
2015 Semi-supervised multi-sensor classification via consensus-based Multi-View Maximum Entropy Discrimination
abstract
In this paper, we consider multi-sensor classification when there is a large number of unlabeled samples. The problem is formulated under the multi-view learning framework and a Consensus-based Multi-View Maximum Entropy Discrimination (CMV-MED) algorithm is proposed. By iteratively maximizing the stochastic agreement between multiple classifiers on the unlabeled dataset, the algorithm simultaneously learns multiple high accuracy classifiers. We demonstrate that our proposed method can yield improved performance over previous multi-view learning approaches by comparing performance on three real multi-sensor data sets.
Tianpei Xie, Nasser M. Nasrabadi, Alfred O. Hero III
ICASSP2
2015 Multichannel transient acoustic signal classification using task-driven dictionary with joint sparsity and beamforming
abstract
We are interested in a multichannel transient acoustic signal classification task which suffers from additive/convolutionary noise corruption. To address this problem, we propose a double-scheme classifier that takes the advantage of multichannel data to improve noise robustness. Both schemes adopt task-driven dictionary learning as the basic framework, and exploit multichannel data at different levels - scheme 1 imposes joint sparsity constraint while learning the dictionary and classifier; scheme 2 adopts beamforming at signal formation level. In addition, matched filter and robust ceptral coefficients are applied to improve noise robustness of the input feature. Experiments show that the proposed classifier significantly outperforms the baseline algorithms.
Yang Zhang 0001, Nasser M. Nasrabadi, Mark Hasegawa-Johnson
ICASSP2
2015 Task-driven dictionary learning with different Laplacian priors for hyperspectral image classification
abstract
Task-driven dictionary learning (TDDL) has shown great success in many classification applications. However, the performance of TDDL is limited by the challenging properties of hyperspectral images (HSI). Fortunately, previous research has made significant progress in HSI classification by enforcing various structured sparsity constraints (priors) on the TDDL-based model. In this paper, we extend some previous work by relaxing the structured sparsity priors and make the model become more flexible and powerful. Specifically, we add class label Laplacian sparsity constraints in two different places, either on the sparse code or on the classifier outputs. Experimental results on widely used datasets shown improvement in performance compared to current state-of-the-art approaches.
Boyu Lu, Nasser M. Nasrabadi
IGARSS2
2015 Graph-Based Sensor Fusion for Classification of Transient Acoustic Signals
abstract
Advances in acoustic sensing have enabled the simultaneous acquisition of multiple measurements of the same physical event via co-located acoustic sensors. We exploit the inherent correlation among such multiple measurements for acoustic signal classification, to identify the launch/impact of munition (i.e., rockets, mortars). Specifically, we propose a probabilistic graphical model framework that can explicitly learn the class conditional correlations between the cepstral features extracted from these different measurements. Additionally, we employ symbolic dynamic filtering-based features, which offer improvements over the traditional cepstral features in terms of robustness to signal distortions. Experiments on real acoustic data sets show that our proposed algorithm outperforms conventional classifiers as well as the recently proposed joint sparsity models for multisensor acoustic classification. Additionally our proposed algorithm is less sensitive to insufficiency in training samples compared to competing approaches.
Umamahesh Srinivas, Nasser M. Nasrabadi, Vishal Monga
IEEE Trans. Cybern.2
2015 Task-Driven Dictionary Learning for Hyperspectral Image Classification With Structured Sparsity Constraints
abstract
Sparse representation models a signal as a linear combination of a small number of dictionary atoms. As a generative model, it requires the dictionary to be highly redundant in order to ensure both a stable high sparsity level and a low reconstruction error for the signal. However, in practice, this requirement is usually impaired by the lack of labeled training samples. Fortunately, previous research has shown that the requirement for a redundant dictionary can be less rigorous if simultaneous sparse approximation is employed, which can be carried out by enforcing various structured sparsity constraints on the sparse codes of the neighboring pixels. In addition, numerous works have shown that applying a variety of dictionary learning methods for the sparse representation model can also improve the classification performance. In this paper, we highlight the task-driven dictionary learning (TDDL) algorithm, which is a general framework for the supervised dictionary learning method. We propose to enforce structured sparsity priors on the TDDL method in order to improve the performance of the hyperspectral classification. Our approach is able to benefit from both the advantages of the simultaneous sparse representation and those of the supervised dictionary learning. We enforce two different structured sparsity priors, the joint and Laplacian sparsities, on the TDDL method and provide the details of the corresponding optimization algorithms. Experiments on numerous popular hyperspectral images demonstrate that the classification performance of our approach is superior to that of the sparse representation classifier with structured priors or the TDDL method.
Xiaoxia Sun, Nasser M. Nasrabadi, Trac D. Tran
IEEE Trans. Geosci. Remote. Sens.2
2015 Semisupervised Hyperspectral Classification Using Task-Driven Dictionary Learning With Laplacian Regularization
abstract
We present a semisupervised method for single-pixel classification of hyperspectral images. The proposed method is designed to address the special problematic characteristics of hyperspectral images, namely, high dimensionality of hyperspectral pixels, lack of labeled samples, and spatial variability of spectral signatures. To alleviate these problems, the proposed method features the following components. First, being a semisupervised approach, it exploits the wealth of unlabeled samples in the image by evaluating the confidence probability of the predicted labels, for each unlabeled sample. Second, we propose to jointly optimize the classifier parameters and the dictionary atoms by a task-driven formulation, to ensure that the learned features (sparse codes) are optimal for the trained classifier. Finally, it incorporates spatial information through adding a Laplacian smoothness regularization to the output of the classifier, rather than the sparse codes, making the spatial constraint more flexible. The proposed method is compared with a few comparable methods for classification of several popular data sets, and it produces significantly better classification results.
Zhangyang Wang, Nasser M. Nasrabadi, Thomas S. Huang
IEEE Trans. Geosci. Remote. Sens.2
2014 Quality-Based Multimodal Classification Using Tree-Structured Sparsity
abstract
Recent studies have demonstrated advantages of information fusion based on sparsity models for multimodal classification. Among several sparsity models, tree-structured sparsity provides a flexible framework for extraction of cross-correlated information from different sources and for enforcing group sparsity at multiple granularities. However, the existing algorithm only solves an approximated version of the cost functional and the resulting solution is not necessarily sparse at group levels. This paper reformulates the tree-structured sparse model for multimodal classification task. An accelerated proximal algorithm is proposed to solve the optimization problem, which is an efficient tool for feature-level fusion among either homogeneous or heterogeneous sources of information. In addition, a (fuzzy-set-theoretic) possibilistic scheme is proposed to weight the available modalities, based on their respective reliability, in a joint optimization problem for finding the sparsity codes. This approach provides a general framework for quality-based fusion that offers added robustness to several sparsity-based multimodal classification algorithms. To demonstrate their efficacy, the proposed methods are evaluated on three different applications - multiview face recognition, multimodal face recognition, and target classification.
Soheil Bahrampour, Asok Ray, Nasser M. Nasrabadi, W. Kenneth Jenkins
CVPR3
2014 Subspace vertex pursuit for separable non-negative matrix factorization in hyperspectral unmixing
abstract
Recently, the separability assumption turns the nonnegative matrix factorization (NMF) into a tractable problem. The assumption coincides with the pixel purity assumption and provides new insights for the hyperspectral unmixing problem. In this paper, we present a quasi-greedy algorithm for solving the problem by employing a back-tracking strategy. Unlike the current greedy methods, the proposed method can refresh the endmember index set in every iteration. Therefore, our method has two important characteristics: (i) low computational complexity comparable to state-of-the-art greedy methods but (ii) empirically enhanced robustness against noise. Finally, computer simulations on synthetic hyperspectral data demonstrate the effectiveness of the proposed method.
Qing Qu 0001, Xiaoxia Sun, Nasser M. Nasrabadi, Trac D. Tran
ICASSP3
2014 Learning to classify with possible sensor failures
abstract
In this paper, we propose an efficient algorithm to train a robust large-margin classifier, when corrupt measurements caused by sensor failure might be present in the training set. By incorporating a non-parametric prior based on the empirical distribution of the training data, we propose a Geometric-Entropy-Minimization regularized Maximum Entropy Discrimination (GEM-MED) method to perform classification and anomaly detection in a joint manner. We demonstrate that our proposed method can yield improved performance over previous robust classification methods in terms of both classification accuracy and anomaly detection rate using simulated data and real footstep data.
Tianpei Xie, Nasser M. Nasrabadi, Alfred O. Hero III
ICASSP2
2014 Coupled dictionaries for thermal to visible face recognition
abstract
Thermal to visible face recognition is the problem of identifying a thermal infrared (IR) face image given a gallery of visible light face images. We attempt to solve this problem by learning coupled dictionaries to represent the two domains. The dictionaries provide a sparse representation which transforms the data into a single, domain-independent, latent space. We formulate the dictionary learning problem as a bi-level optimization problem and perform a stochastic gradient descent on the dictionaries to solve it. We present experimental results demonstrating the effectiveness of our approach.
Christopher Reale, Nasser M. Nasrabadi, Rama Chellappa
ICIP2
2014 Task-driven dictionary learning for hyperspectral image classification with structured sparsity priors
abstract
In hyperspectral pixel classification, previous research have shown that the sparse representation classifier can achieve a better performance when exploiting the neighboring test pixels through enforcing different structured sparsity priors. In this paper, we propose a supervised sparse-representation-based dictionary learning method with joint or Laplacian s-parsity priors. The proposed method has numerous advantages over the existing dictionary learning techniques. It uses a structured sparsity and provides a more robust and stable sparse coefficients. Besides, it is capable of reducing the classification error by jointly optimizing the dictionary and the classifier's parameters during the dictionary training stage.
Xiaoxia Sun, Nasser M. Nasrabadi, Trac D. Tran
ICIP2
2014 Structured Priors for Sparse-Representation-Based Hyperspectral Image Classification
abstract
Pixelwise classification, where each pixel is assigned to a predefined class, is one of the most important procedures in hyperspectral image (HSI) analysis. By representing a test pixel as a linear combination of a small subset of labeled pixels, a sparse representation classifier (SRC) gives rather plausible results compared with that of traditional classifiers such as the support vector machine. Recently, by incorporating additional structured sparsity priors, the second-generation SRCs have appeared in the literature and are reported to further improve the performance of HSI. These priors are based on exploiting the spatial dependences between the neighboring pixels, the inherent structure of the dictionary, or both. In this letter, we review and compare several structured priors for sparse-representation-based HSI classification. We also propose a new structured prior called the low-rank (LR) group prior, which can be considered as a modification of the LR prior. Furthermore, we will investigate how different structured priors improve the result for the HSI classification.
Xiaoxia Sun, Qing Qu 0001, Nasser M. Nasrabadi, Trac D. Tran
IEEE Geosci. Remote. Sens. Lett.3
2014 Joint Sparse Representation for Robust Multimodal Biometrics Recognition
abstract
Traditional biometric recognition systems rely on a single biometric signature for authentication. While the advantage of using multiple sources of information for establishing the identity has been widely recognized, computational models for multimodal biometrics recognition have only recently received attention. We propose a multimodal sparse representation method, which represents the test data by a sparse linear combination of training data, while constraining the observations from different modalities of the test subject to share their sparse representations. Thus, we simultaneously take into account correlations as well as coupling information among biometric modalities. A multimodal quality measure is also proposed to weigh each modality as it gets fused. Furthermore, we also kernelize the algorithm to handle nonlinearity in data. The optimization problem is solved using an efficient alternative direction method. Various experiments show that the proposed method compares favorably with competing fusion-based methods.
Vishal M. Patel, Nasser M. Nasrabadi, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Separated Component-Based Restoration of Speckled SAR Images
abstract
Many coherent imaging modalities such as synthetic aperture radar suffer from a multiplicative noise, commonly referred to as speckle, which often makes the interpretation of data difficult. An effective strategy for speckle reduction is to use a dictionary that can sparsely represent the features in the speckled image. However, such approaches fail to capture important salient features such as texture. In this paper, we present a speckle reduction algorithm that handles this issue by formulating the restoration problem so that the structure and texture components can be separately estimated with different dictionaries. To solve this formulation, an iterative algorithm based on surrogate functionals is proposed. Experiments indicate the proposed method performs favorably compared to state-of-the-art speckle reduction methods.
Vishal M. Patel, Glenn R. Easley, Rama Chellappa, Nasser M. Nasrabadi
IEEE Trans. Geosci. Remote. Sens.4
2014 Abundance Estimation for Bilinear Mixture Models via Joint Sparse and Low-Rank Representation
abstract
Sparsity-based unmixing algorithms, exploiting the sparseness property of the abundances, have recently been proposed with promising performances. However, these algorithms are developed for the linear mixture model (LMM), which cannot effectively handle the nonlinear effects. In this paper, we extend the current sparse regression methods for the LMM to bilinear mixture models (BMMs), where the BMMs introduce additional bilinear terms in the LMM in order to model second-order photon scattering effects. To solve the abundance estimation problem for the BMMs, we propose to perform a sparsity-based abundance estimation by using two dictionaries: a linear dictionary containing all the pure endmembers and a bilinear dictionary consisting of all the possible second-order endmember interaction components. Then, the abundance values can be estimated from the sparse codes associated with the linear dictionary. Moreover, to exploit the spatial data structure where the adjacent pixels are usually homogeneous and are often mixtures of the same materials, we first employ the joint-sparsity (row-sparsity) model to enforce structured sparsity on the abundance coefficients. However, the joint-sparsity model is often a strict assumption, which might cause some aliasing artifacts for the pixels that lie on the boundaries of different materials. To deal with this problem, the low-rank-representation model, which seeks the lowest rank representation of the data, is further introduced to better capture the spatial data structure. Our simulation results demonstrate that the proposed algorithms provide much enhanced performance compared with state-of-the-art algorithms.
Qing Qu 0001, Nasser M. Nasrabadi, Trac D. Tran
IEEE Trans. Geosci. Remote. Sens.2
2014 Spatial-Spectral Classification of Hyperspectral Images Using Discriminative Dictionary Designed by Learning Vector Quantization
abstract
In this paper, a novel discriminative dictionary learning method is proposed for sparse-representation-based classification (SRC) to label highly dimensional hyperspectral imagery (HSI). In SRC, a dictionary is conventionally constructed using all of the training pixels, which is not only inefficient due to the large size of typical HSI images but also ineffective in capturing class-discriminative information crucial for classification. We address the dictionary design problem with the inspiration from the learning vector quantization technique and propose a hinge loss function that is directly related to the classification task as the objective function for dictionary learning. The resulting online learning procedure systematically “pulls” and “pushes” dictionary atoms so that they become better adapted to distinguish between different classes. In addition, the spatial context for a test pixel within its local neighborhood is modeled using a Bayesian graph model and is incorporated with the sparse representation of a single test pixel in a unified probabilistic framework, which enables further refinement of our dictionary to capture the spatial class dependence that complements the spectral information. Experiments on different HSI images demonstrate that the dictionaries optimized using our method can achieve higher classification accuracy with substantially reduced dictionary size than using the whole training set. The proposed method also outperforms existing dictionary learning methods and attains the state-of-the-art results in both the spectral-only and spatial-spectral settings.
Nasser M. Nasrabadi, Thomas S. Huang
IEEE Trans. Geosci. Remote. Sens.2
2013 Hyperspectral abundance estimation for the generalized bilinear model with joint sparsity constraint
abstract
In this paper, we present a novel abundance estimation method for the generalized bilinear model (GBM) via sparse representation for hyperspectral imagery. Because the GBM generalizes the linear mixture model (LMM) by introducing an additional bilinear term, our sparsity-based abundance estimation is performed by utilizing two dictionaries-a linear dictionary containing all the pure endmembers and a bilinear dictionary consisting of all the possible bilinear interaction components. Because the components within the bilinear term are also linearly combined, by employing a composite dictionary made up by the concatenation of the linear and bilinear dictionaries we can reformulate the bilinear problem in a linear sparse regression framework. In this way, the abundance values are estimated from the sparse codes only associated with the linear dictionary. To further improve the estimation performance, we incorporate the joint-sparsity model to exploit the spatial information in the data. The experiments demonstrate the effectiveness of the proposed algorithms on both synthetic and real data.
Qing Qu 0001, Nasser M. Nasrabadi, Trac D. Tran
ICASSP2
2013 Graph-based multi-sensor fusion for acoustic signal classification
abstract
Advances in acoustic sensing have enabled the simultaneous acquisition of multiple measurements of the same physical event via co-located acoustic sensors. We exploit the inherent correlation among such multiple measurements for acoustic signal classification, to identify the launch/impact of munition (i.e. rockets, mortars). Specifically, we propose a probabilistic graphical model framework that can explicitly learn the class conditional correlations between the cepstral features extracted from these different measurements. Additionally, we employ symbolic dynamic filtering-based features, which offer improvements over the traditional cepstral features. Experiments on real acoustic data sets show that our proposed algorithm outperforms conventional classifiers as well as recently proposed joint sparsity models for multi-sensor acoustic signal classification.
Umamahesh Srinivas, Nasser M. Nasrabadi, Vishal Monga
ICASSP2
2013 Discriminative and compact dictionary design for Hyperspectral Image classification using learning VQ framework
abstract
Sparse representation provides an efficient description for high-dimensional Hyperspectral Imagery (HSI) and also encodes discriminative information useful for classification. However, due to the large size of typical HSI images, the naive way to construct a dictionary with all training pixels is neither efficient nor practical. In this paper, a novel approach is proposed to design compact dictionary for Sparse Representation-based Classification (SRC). Inspired by Learning Vector Quantization (LVQ) techniques, we use a hinge loss function directly related to classification task as our objective function, and optimize the dictionary by exploiting the differentiable parts of sparse codes. The resultant dictionary updating procedure adapts the “push” and “pull” actions in LVQ to SRC, which is therefore named as Learning Sparse Representation-based Classification (LSRC). Experiments on different HSI images demonstrate that our LSRC approach can achieve higher classification accuracy with substantially smaller dictionary size than using the whole training set, and also outperforms existing dictionary learning methods.
Nasser M. Nasrabadi, Thomas S. Huang
ICASSP2
2013 A Max-Margin Perspective on Sparse Representation-Based Classification
abstract
Sparse Representation-based Classification (SRC) is a powerful tool in distinguishing signal categories which lie on different subspaces. Despite its wide application to visual recognition tasks, current understanding of SRC is solely based on a reconstructive perspective, which neither offers any guarantee on its classification performance nor provides any insight on how to design a discriminative dictionary for SRC. In this paper, we present a novel perspective towards SRC and interpret it as a margin classifier. The decision boundary and margin of SRC are analyzed in local regions where the support of sparse code is stable. Based on the derived margin, we propose a hinge loss function as the gauge for the classification performance of SRC. A stochastic gradient descent algorithm is implemented to maximize the margin of SRC and obtain more discriminative dictionaries. Experiments validate the effectiveness of the proposed approach in predicting classification performance and improving dictionary quality over reconstructive ones. Classification results competitive with other state-of-the-art sparse coding methods are reported on several data sets.
Jianchao Yang, Nasser M. Nasrabadi, Thomas S. Huang
ICCV3
2013 Opportunistic sensing for object recognition - A unified formulation for dynamic sensor selection and feature extraction
abstract
A novel problem of object recognition with dynamically allocated sensing resources is considered in this paper. We call this problem opportunistic sensing since prior knowledge about the correlation between class label and signal distribution is exploited as early as in data acquisition. Two forms of sensing parameters — discrete sensor index and continuous linear measurement vector — are optimized within the same maximum negative entropy framework. The computationally intractable expected entropy is approximated using unscented transform for Gaussian models, and we solve the problem using a gradient-based method. Our formulation is theoretically shown to be closely related to the maximum mutual information criterion for sensor selection and linear feature extraction techniques such as PCA, LDA, and CCA. The proposed approach is validated on multi-view vehicle classification and face recognition datasets, and remarkable improvement over baseline methods is demonstrated in the experiments.
Jianchao Yang, Nasser M. Nasrabadi, Jiangping Wang, Thomas S. Huang
ICME3
2013 Exploiting Sparsity in Hyperspectral Image Classification via Graphical Models
abstract
A significant recent advance in hyperspectral image (HSI) classification relies on the observation that the spectral signature of a pixel can be represented by a sparse linear combination of training spectra from an overcomplete dictionary. A spatiospectral notion of sparsity is further captured by developing a joint sparsity model, wherein spectral signatures of pixels in a local spatial neighborhood (of the pixel of interest) are constrained to be represented by a common collection of training spectra, albeit with different weights. A challenging open problem is to effectively capture the class conditional correlations between these multiple sparse representations corresponding to different pixels in the spatial neighborhood. We propose a probabilistic graphical model framework to explicitly mine the conditional dependences between these distinct sparse features. Our graphical models are synthesized using simple tree structures which can be discriminatively learnt (even with limited training samples) for classification. Experiments on benchmark HSI data sets reveal significant improvements over existing approaches in classification rates as well as robustness to choice of training.
Umamahesh Srinivas, Yi Chen 0014, Vishal Monga, Nasser M. Nasrabadi, Trac D. Tran
IEEE Geosci. Remote. Sens. Lett.4
2013 Performance comparison of feature extraction algorithms for target detection and classification
Soheil Bahrampour, Asok Ray, Soumalya Sarkar, Thyagaraju Damarla, Nasser M. Nasrabadi
Pattern Recognit. Lett.5
2013 Hyperspectral Image Classification via Kernel Sparse Representation
abstract
In this paper, a novel nonlinear technique for hyperspectral image (HSI) classification is proposed. Our approach relies on sparsely representing a test sample in terms of all of the training samples in a feature space induced by a kernel function. For each test pixel in the feature space, a sparse representation vector is obtained by decomposing the test pixel over a training dictionary, also in the same feature space, by using a kernel-based greedy pursuit algorithm. The recovered sparse representation vector is then used directly to determine the class label of the test pixel. Projecting the samples into a high-dimensional feature space and kernelizing the sparse representation improve the data separability between different classes, providing a higher classification accuracy compared to the more conventional linear sparsity-based classification algorithms. Moreover, the spatial coherency across neighboring pixels is also incorporated through a kernelized joint sparsity model, where all of the pixels within a small neighborhood are jointly represented in the feature space by selecting a few common training samples. Kernel greedy optimization algorithms are suggested in this paper to solve the kernel versions of the single-pixel and multi-pixel joint sparsity-based recovery problems. Experimental results on several HSIs show that the proposed technique outperforms the linear sparsity-based classification technique, as well as the classical support vector machines and sparse kernel logistic regression classifiers.
Yi Chen 0014, Nasser M. Nasrabadi, Trac D. Tran
IEEE Trans. Geosci. Remote. Sens.2
2013 Design of Non-Linear Kernel Dictionaries for Object Recognition
abstract
In this paper, we present dictionary learning methods for sparse signal representations in a high dimensional feature space. Using the kernel method, we describe how the well known dictionary learning approaches, such as the method of optimal directions and KSVD, can be made nonlinear. We analyze their kernel constructions and demonstrate their effectiveness through several experiments on classification problems. It is shown that nonlinear dictionary learning approaches can provide significantly better performance compared with their linear counterparts and kernel principal component analysis, especially when the data is corrupted by different types of degradations.
Hien Van Nguyen, Vishal M. Patel, Nasser M. Nasrabadi, Rama Chellappa
IEEE Trans. Image Process.3
2012 Sparse Embedding: A Framework for Sparsity Promoting Dimensionality Reduction
Hien Van Nguyen, Vishal M. Patel, Nasser M. Nasrabadi, Rama Chellappa
ECCV (6)3
2012 Kernel dictionary learning
abstract
In this paper, we present dictionary learning methods for sparse and redundant signal representations in high dimensional feature space. Using the kernel method, we describe how the well-known dictionary learning approaches such as the method of optimal directions and K-SVD can be made nonlinear. We analyze these constructions and demonstrate their improved performance through several experiments on classification problems. It is shown that nonlinear dictionary learning approaches can provide better discrimination compared to their linear counterparts and kernel PCA, especially when the data is corrupted by noise.
Hien Van Nguyen, Vishal M. Patel, Nasser M. Nasrabadi, Rama Chellappa
ICASSP3
2012 Kernel multi-metric learning for multi-channel transient acoustic signal classification
abstract
In this paper, we propose a kernel multi-metric learning algorithm for multi-channel transient acoustic signal classification. The proposed method learns a set of metrics jointly for multi-channel transient acoustic signals in a kernel-induced feature space to exploit the non-linearity of the data for improving the classification performance. An effective algorithm is developed for the task of learning multiple metrics in the kernel space. By learning the multiple metrics jointly within a single unified optimization framework, we can learn better metrics to integrate the multiple channels of the signal for a joint classification. Experimental results compared with classical as well as recent algorithms on real-world acoustic datasets verified the effectiveness of the proposed method.
Haichao Zhang 0001, Yanning Zhang 0001, Nasser M. Nasrabadi, Thomas S. Huang
ICASSP3
2012 Joint dynamic sparse learning and its application to multi-view face recognition
Haichao Zhang 0001, Yanning Zhang 0001, Nasser M. Nasrabadi, Thomas S. Huang
ICPR3
2012 Kernel sparse representation for hyperspectral target detection
abstract
In this paper, we present a nonlinear kernel-based target detection algorithm for hyperspectral images. The proposed approach relies on the sparse representation of an unknown sample with respect to both background and target training samples in a high-dimensional feature space induced by a kernel function. The sparse representation vector can be recovered via a kernelized greedy algorithm, where the kernel trick is used to avoid explicit evaluations of the data in the feature space. The spatial smoothness in hyperspectral images is also taken into account through a kernelized joint sparsity model. The detection decision is then made by comparing the reconstruction accuracy in terms of the background and target sub-dictionaries. The detection algorithm in a high-dimensional feature space implicitly exploits the higher-order structure (correlations) within the data which cannot be captured by a linear model. Therefore, projecting the pixels into a kernel feature space and kernelizing the linear sparse representation model improves the separability between the background and target classes, leading to a more accurate detection performance.
Yi Chen 0014, Nasser M. Nasrabadi, Trac D. Tran
IGARSS2
2012 Covariance trace for polarimetric anomaly detection
abstract
We propose a new method for autonomous manmade object detection, which is solely based on the use of second order statistics from two polarization components (0 and 90 deg) of polarimetric imagery. Using the approach, manmade objects can be detected as anomalies in scenes spatially dominated by natural objects. The approach exploits a key discovery: manmade objects are separable from natural objects in the (0 and 90 deg) variance-covariance space, holding invariant to diurnal cycle variation and geometry of illumination. Testing real imagery acquired outdoor (0.55 km sensor-to-target range) showed that the approach significantly outperforms the classical use of Stokes vector and DOLP (degree of linear polarization) during a full diurnal cycle.
João M. Romano, Dalton S. Rosario, Nasser M. Nasrabadi
IGARSS3
2012 Discriminative graphical models for sparsity-based hyperspectral target detection
abstract
The inherent discriminative capability of sparse representations has been exploited recently for hyperspectral target detection. This approach relies on the observation that the spectral signature of a pixel can be represented as a linear combination of a few training spectra drawn from both target and background classes. The sparse representation corresponding to a given test spectrum captures class-specific discriminative information crucial for detection tasks. Spatio-spectral information has also been introduced into this framework via a joint sparsity model that simultaneously solves for the sparse features for a group of spatially local pixels, since such pixels are highly likely to have similar spectral characteristics. In this paper, we propose a probabilistic graphical model framework that can explicitly learn the class conditional correlations between these distinct sparse representations corresponding to different pixels in a spatial neighborhood. Simulation results show that the proposed algorithm outperforms classical hyperspectral target detection algorithms as well as support vector machines.
Umamahesh Srinivas, Yi Chen 0014, Vishal Monga, Nasser M. Nasrabadi, Trac D. Tran
IGARSS4
2012 Joint dynamic sparse representation for multi-view face recognition
Haichao Zhang 0001, Nasser M. Nasrabadi, Yanning Zhang 0001, Thomas S. Huang
Pattern Recognit.2
2012 Joint-Structured-Sparsity-Based Classification for Multiple-Measurement Transient Acoustic Signals
abstract
This paper investigates the joint-structured-sparsity-based methods for transient acoustic signal classification with multiple measurements. By joint structured sparsity, we not only use the sparsity prior for each measurement but we also exploit the structural information across the sparse representation vectors of multiple measurements. Several different sparse prior models are investigated in this paper to exploit the correlations among the multiple measurements with the notion of the joint structured sparsity for improving the classification accuracy. Specifically, we propose models with the joint structured sparsity under different assumptions: same sparse code model, common sparse pattern model, and a newly proposed joint dynamic sparse model. For the joint dynamic sparse model, we also develop an efficient greedy algorithm to solve it. Extensive experiments are carried out on real acoustic data sets, and the results are compared with the conventional discriminative classifiers in order to verify the effectiveness of the proposed method.
Haichao Zhang 0001, Yanning Zhang 0001, Nasser M. Nasrabadi, Thomas S. Huang
IEEE Trans. Syst. Man Cybern. Part B3
2011 Robust multi-sensor classification via joint sparse representation
Nam H. Nguyen, Nasser M. Nasrabadi, Trac D. Tran
FUSION2
2011 Heterogeneous multi-metric learning for multi-sensor fusion
Haichao Zhang 0001, Thomas S. Huang, Nasser M. Nasrabadi, Yanning Zhang 0001
FUSION3
2011 Transient acoustic signal classification using joint sparse representation
abstract
In this paper, we present a novel joint sparse representation based method for acoustic signal classification with multiple measurements. The proposed method exploits the correlations among the multiple measurements with the notion of joint sparsity for improving the classification accuracy. Extensive experiments are carried out on real acoustic data sets and the results are compared with the conventional discriminative classifiers in order to verify the effectiveness of the proposed method.
Haichao Zhang 0001, Nasser M. Nasrabadi, Thomas S. Huang, Yanning Zhang 0001
ICASSP2
2011 Multi-observation visual recognition via joint dynamic sparse representation
abstract
We address the problem of visual recognition from multiple observations of the same physical object, which can be generated under different conditions, such as frames at different time instances or snapshots from different viewpoints. We formulate the multi-observation visual recognition task as a joint sparse representation model and take advantage of the correlations among the multiple observations for classification using a novel joint dynamic sparsity prior. The proposed joint dynamic sparsity prior promotes shared joint sparsity pattern among the multiple sparse representation vectors at class-level, while allowing distinct sparsity patterns at atom-level within each class in order to facilitate a flexible representation. The proposed method can handle both homogenous as well as heterogenous data within the same framework. Extensive experiments on various visual classification tasks including face recognition and generic object classification demonstrate that the proposed method outperforms existing state-of-the-art methods
Haichao Zhang 0001, Nasser M. Nasrabadi, Yanning Zhang 0001, Thomas S. Huang
ICCV2
2011 Close the loop: Joint blind image restoration and recognition with sparse representation prior
abstract
Most previous visual recognition systems simply assume ideal inputs without real-world degradations, such as low resolution, motion blur and out-of-focus blur. In presence of such unknown degradations, the conventional approach first resorts to blind image restoration and then feeds the restored image into a classifier. Treating restoration and recognition separately, such a straightforward approach, however, suffers greatly from the defective output of the ill-posed blind image restoration. In this paper, we present a joint blind image restoration and recognition method based on the sparse representation prior to handle the challenging problem of face recognition from low-quality images, where the degradation model is realistic and totally unknown. The sparse representation prior states that the degraded input image, if correctly restored, will have a good sparse representation in terms of the training set, which indicates the identity of the test image. The proposed algorithm achieves simultaneous restoration and recognition by iteratively solving the blind image restoration in pursuit of the sparest representation for recognition. Based on such a sparse representation prior, we demonstrate that the image restoration task and the recognition task can benefit greatly from each other. Extensive experiments on face datasets under various degradations are carried out and the results of our joint model shows significant improvements over conventional methods of treating the two tasks independently.
Haichao Zhang 0001, Jianchao Yang, Yanning Zhang 0001, Nasser M. Nasrabadi, Thomas S. Huang
ICCV4
2011 Hyperspectral image classification via kernel sparse representation
abstract
In this paper, a new technique for hyperspectral image classification is proposed. Our approach relies on the sparse representation of a test sample with respect to all training samples in a feature space induced by a kernel function. Projecting the samples into the feature space and kernelizing the sparse representation improves the separability of the data and thus yields higher classification accuracy compared to the more conventional linear sparsity-based classification algorithm. Moreover, the spatial coherence across neighboring pixels is also incorporated through a kernelized joint sparsity model, where all of the pixels within a small neighborhood are sparsely represented in the feature space by selecting a few common training samples. Two greedy algorithms are also provided in this paper to solve the kernel versions of the pixel-wise and jointly sparse recovery problems. Experimental results show that the proposed technique outperforms the linear sparsity-based classification technique and the classical Support Vector Machine classifiers.
Yi Chen 0014, Nasser M. Nasrabadi, Trac D. Tran
ICIP2
2011 Multi-view face recognition via joint dynamic sparse representation
abstract
We consider the problem of automatically recognizing a human face from its multi-view images with unconstrained poses and illuminations. We formulate the multi-view face recognition problem as that of classifying among several multi-input (views) regression models by using a novel joint dynamic sparse representation method which exploits jointly the inter-correlation among all the multi-view images in order to make a decision. Extensive experiments on CMU Multi-PIE face database are conducted to verify the efficacy of the proposed method.
Haichao Zhang 0001, Nasser M. Nasrabadi, Thomas S. Huang, Yanning Zhang 0001
ICIP2
2011 Robust Lasso with missing and grossly corrupted observations
abstract
This paper studies the problem of accurately recovering a sparse vector $\beta^{\star}$ from highly corrupted linear measurements $y = X \beta^{\star} + e^{\star} + w$ where $e^{\star}$ is a sparse error vector whose nonzero entries may be unbounded and $w$ is a bounded noise. We propose a so-called extended Lasso optimization which takes into consideration sparse prior information of both $\beta^{\star}$ and $e^{\star}$. Our first result shows that the extended Lasso can faithfully recover both the regression and the corruption vectors. Our analysis is relied on a notion of extended restricted eigenvalue for the design matrix $X$. Our second set of results applies to a general class of Gaussian design matrix $X$ with i.i.d rows $\oper N(0, \Sigma)$, for which we provide a surprising phenomenon: the extended Lasso can recover exact signed supports of both $\beta^{\star}$ and $e^{\star}$ from only $\Omega(k \log p \log n)$ observations, even the fraction of corruption is arbitrarily close to one. Our analysis also shows that this amount of observations required to achieve exact signed support is optimal.
Nam H. Nguyen, Nasser M. Nasrabadi, Trac D. Tran
NIPS2
2011 Simultaneous Joint Sparsity Model for Target Detection in Hyperspectral Imagery
abstract
This letter proposes a simultaneous joint sparsity model for target detection in hyperspectral imagery (HSI). The key innovative idea here is that hyperspectral pixels within a small neighborhood in the test image can be simultaneously represented by a linear combination of a few common training samples but weighted with a different set of coefficients for each pixel. The joint sparsity model automatically incorporates the interpixel correlation within the HSI by assuming that neighboring pixels usually consist of similar materials. The sparse representations of the neighboring pixels are obtained by simultaneously decomposing the pixels over a given dictionary consisting of training samples of both the target and background classes. The recovered sparse coefficient vectors are then directly used for determining the label of the test pixels. Simulation results show that the proposed algorithm outperforms the classical hyperspectral target detection algorithms, such as the popular spectral matched filters, matched subspace detectors, and adaptive subspace detectors, as well as binary classifiers such as support vector machines.
Yi Chen 0014, Nasser M. Nasrabadi, Trac D. Tran
IEEE Geosci. Remote. Sens. Lett.2
2011 Hyperspectral Image Classification Using Dictionary-Based Sparse Representation
abstract
A new sparsity-based algorithm for the classification of hyperspectral imagery is proposed in this paper. The proposed algorithm relies on the observation that a hyperspectral pixel can be sparsely represented by a linear combination of a few training samples from a structured dictionary. The sparse representation of an unknown pixel is expressed as a sparse vector whose nonzero entries correspond to the weights of the selected training samples. The sparse vector is recovered by solving a sparsity-constrained optimization problem, and it can directly determine the class label of the test sample. Two different approaches are proposed to incorporate the contextual information into the sparse recovery optimization problem in order to improve the classification performance. In the first approach, an explicit smoothing constraint is imposed on the problem formulation by forcing the vector Laplacian of the reconstructed image to become zero. In this approach, the reconstructed pixel of interest has similar spectral characteristics to its four nearest neighbors. The second approach is via a joint sparsity model where hyperspectral pixels in a small neighborhood around the test pixel are simultaneously represented by linear combinations of a few common training samples, which are weighted with a different set of coefficients for each pixel. The proposed sparsity-based algorithm is applied to several real hyperspectral images for classification. Experimental results show that our algorithm outperforms the classical supervised classifier support vector machines in most cases.
Yi Chen 0014, Nasser M. Nasrabadi, Trac D. Tran
IEEE Trans. Geosci. Remote. Sens.2
2010 Automatic target recognition based on simultaneous sparse representation
abstract
In this paper, an automatic target recognition algorithm is presented based on a framework for learning dictionaries for simultaneous sparse signal representation and feature extraction. The dictionary learning algorithm is based on class supervised simultaneous orthogonal matching pursuit while a matching pursuit-based similarity measure is used for classification. We show how the proposed framework can be helpful for efficient utilization of data, with the possibility of developing real-time, robust target classification. We verify the efficacy of the proposed algorithm using confusion matrices on the well known Comanche forward-looking infrared data set consisting of ten different military targets at different orientations.
Vishal M. Patel, Nasser M. Nasrabadi, Rama Chellappa
ICIP2
2010 Anomaly Detection for Longwave FLIR Imagery Using Kernel Wavelet-RX
abstract
This paper describes a new kernel wavelet-based anomaly detection technique for long-wave (LW) Forward Looking Infrared (FLIR) imagery. The proposed approach called kernel wavelet-RX algorithm is essentially an extension of the wavelet-RX algorithm (combination of wavelet transform and RX anomaly detector) to a high dimensional feature space (possibly infinite) via a certain nonlinear mapping function of the input data. The wavelet-RX algorithm in this high dimensional feature space can easily be implemented in terms of kernels that implicitly compute dot products in the feature space (kernelizing the wavelet-RX algorithm). In our kernel wavelet-RX algorithm, a 2-D wavelet transform is first applied to decompose the input image into uniform subbands. A number of significant subbands (high energy subbands) are concatenated together to form a subband-image cube. The kernel RX algorithm is then applied to these subband-image cubes obtained from wavelet decomposition of the LW database images. Experimental results are presented for the proposed kernel wavelet-RX, wavelet-RX and the classical CFAR algorithm for detecting anomalies (targets) in a large database of LW imagery. The ROC plots show that the proposed kernel wavelet-RX algorithm outperforms the wavelet-RX as well as the classical CFAR detector.
Asif Mehmood, Nasser M. Nasrabadi
ICPR2
2010 Sparsity-based classification of hyperspectral imagery
abstract
In this paper, a new sparsity-based classification algorithm for hyperspectral imagery is proposed. This algorithm is based on the concept that a pixel in hyperspectral imagery lies in a low-dimensional subspace and thus can be represented by a sparse linear combination of the training samples. The sparse representation (a sparse vector representing the selected training samples) of a test sample can be recovered by solving a constrained optimization problem. Once the sparse vector is obtained, the class of the test sample can be directly determined by the behavior of the vector on reconstruction. In addition to the constraints on sparsity and reconstruction accuracy, we also exploit the fact that hyperspectral images are usually smooth within a neighborhood. In our proposed algorithm, a smoothness constraint is imposed by forcing the Laplacian of the reconstructed image to be minimum in the optimization process. The proposed sparsity-based algorithm is applied to several hyperspectral imagery to classify the pixels into target and background classes. Simulation results show that our algorithm outperforms the classical hyperspectral target detection algorithms, such as the popular spectral matched filters, matched subspace detectors, and adaptive subspace detectors.
Yi Chen 0014, Nasser M. Nasrabadi, Trac D. Tran
IGARSS2
2009 Wiener Prediction-based Change Detection for Locating Mines in Multilook SAR Imagery
abstract
In this paper, we present a Wiener-based change detection method and compare its performance with several other methods for a pair of multi-look synthetic aperture radar (SAR) images of the same scene. We implement and compare several techniques which vary in complexity. Among the simple methods that are implemented are differencing, Euclidean distance, and image ratioing. These methods require minimal processing time, with little computational complexity, and incorporate no statistical information. We also implemented methods which incorporate second order statistic calculations in making a change decision in efforts to mitigate false alarms arising from the speckle noise, misregistration errors, and nonlinear variations in SAR images. These methods include a Wiener prediction-based method, Mahalanobis distance measure and subspace projection method. We compare the performance of these methods using multi-look SAR images containing several targets (mines). We present results in the form of receiver operating characteristics (ROC) curves.
Nasser M. Nasrabadi
IGARSS (2)1
2009 Block Wiener-based image registration for moving target indication
Lance M. Kaplan, Nasser M. Nasrabadi
Image Vis. Comput.2
2009 Automated Hyperspectral Cueing for Civilian Search and Rescue
abstract
Hyperspectral remote sensing provides information related to surface material characteristics that can be exploited to perform automated detection of targets of interest and has been applied to a variety of remote sensing applications. This paper explores the application to civilian search and rescue, using the Airborne Real-time Cueing Hyperspectral Enhanced Reconnaissance (ARCHER) system developed for the Civil Air Patrol as a key example of how evolving hyperspectral technology can be employed to support these operations. ARCHER combines a visible/near-infrared hyperspectral imaging system, a high-resolution visible panchromatic imaging sensor, and an integrated geopositioning and inertial navigation unit with onboard real-time processing for data acquisition and correction, precision image georegistration, and target detection and cueing. Processing for detecting downed aircraft wreckage and other related objects employs real-time adaptive anomaly detection and matched filtering algorithms, and a non-real-time change detection mode to provide further false alarm reduction in some instances. This paper describes the system technology, with an emphasis on the current and evolving automated target detection methods, and summarizes the operational experience in the airborne employment against civilian search and rescue missions.
Michael T. Eismann, Alan D. Stocker, Nasser M. Nasrabadi
Proc. IEEE3
2008 A nonlinear kernel-based joint fusion/detection of anomalies using Hyperspectral and SAR imagery
abstract
In this paper a new nonlinear joint fusion and detection algorithm is proposed for locating anomalies from two different types of sensor data (synthetic aperture radar (SAR) and hyperspectral sensor (HS) data). The proposed approach jointly exploits the nonlinear correlation or dependencies between the two sensors in order to simultaneously fuse and detect the objects of interest (mines). A well-known anomaly detector, so called RX algorithm is extended to perform fusion and detection simultaneously at the pixel level by appropriately concatenating the information from the two sensors. This approach is then extended to its nonlinear version using the idea of kernel learning which explicitly exploits the higher order dependencies (nonlinear correlations) between the two sensor data through an appropriate kernel.
Nasser M. Nasrabadi
ICIP1
2008 Regularized Spectral Matched Filter for Target Recognition in Hyperspectral Imagery
abstract
This letter extends the idea of regularization to spectral matched filters. It incorporates a quadratic penalization term in the design of spectral matched filters in order to restrict the possible matched filters (models) to a subset which are more stable and have better performance than the non-regularized adaptive spectral matched filters. The effect of regularization depends on the form of the regularization term and the amount of regularization which is controlled by a parameter so-called the regularization coefficient. In this letter, the sum-of-squares of the filter coefficients is used as the regularization term, and different values for the regularization coefficient are tested. A Bayesian-based derivation of the regularized matched filter is also described which provides a procedure for choosing the regularization coefficient. Experimental results for detecting targets in hyperspectral imagery are presented for regularized and non-regularized spectral matched filters.
Nasser M. Nasrabadi
IEEE Signal Process. Lett.1
2007 Regularized Spectral Matched Filter for Target Detection in Hyperspectral Imagery
abstract
This paper describes a new adaptive spectral matched filter that incorporates the idea of regularization (shrinkage) to penalize and shrink the filter coefficients to a range of values. The regularization has the effect of restricting the possible matched filters (models) to a subset which are more stable and have better performance than the non-regularized adaptive spectral matched filters. The effect of regularization depends on the form of the regularization term and the amount of regularization is controlled by so called the regularization coefficient. Experimental results for detecting targets in hyperspectral imagery are presented for regularized and non-regularized spectral matched filters.
Nasser M. Nasrabadi
ICIP (4)1
2007 Penalized spectral matched filter for target detection in hyperspectral imagery
abstract
This paper describes a new adaptive spectral matched filter that incorporates the idea of regularization (shrinkage) to penalize and shrink the filter coefficients to a range of values. The regularization has the effect of restricting the possible matched filters (models) to a subset which are more stable and have better performance than the non-regularized adaptive spectral matched filters. The effect of regularization depends on the form of the regularization term and the amount of regularization is controlled by so called regularization coefficient. In this paper the sum-of-squares of the filter coefficients is used as the regularization term and several different values for the regularization coefficient are tested. Experimental results for detecting targets in hyperspectral imagery are presented for regularized and non-regularized spectral matched filters.
Nasser M. Nasrabadi
IGARSS1
2007 Kernel Spectral Matched Filter for Hyperspectral Imagery
Heesung Kwon, Nasser M. Nasrabadi
Int. J. Comput. Vis.2
2007 Kernel Eigenspace Separation Transform for Subspace Anomaly Detection in Hyperspectral Imagery
abstract
This letter proposes a nonlinear version of the eigenspace separation transform (EST) for subspace anomaly detection in hyperspectral imaging. The EST is defined in terms of the eigenvectors of the difference correlation matrix (DCOR) obtained using the data from the two classes. Using ideas found in the machine learning literature (i.e., the kernel trick), a nonlinear version-kernel EST (KEST)-is achieved by expressing the DCOR in terms of dot products in feature space and replacing all dot products with a Mercer kernel function that is defined in terms of input data space. Experimental results indicate that KEST outperforms many other commonly used subspace anomaly detection algorithms.
Hirsh R. Goldberg, Heesung Kwon, Nasser M. Nasrabadi
IEEE Geosci. Remote. Sens. Lett.3
2006 Kernel adaptive subspace detector for hyperspectral imagery
abstract
In this letter, we present a kernel-based nonlinear version of the adaptive subspace detector (ASD) that implicitly detects signals of interest in a high-dimensional (possibly infinite) feature space associated with a particular nonlinear mapping. In order to address the high dimensionality of the feature space, ASD is first implicitly formulated in the feature space, which is then converted into an expression in terms of kernel functions via the kernel trick property of the Mercer kernels. Experimental results based on simulated data and real hyperspectral imagery show that the proposed kernel-based ASD outperforms the conventional ASD and a nonlinear anomaly detector so called the kernel RX-algorithm.
Heesung Kwon, Nasser M. Nasrabadi
IEEE Geosci. Remote. Sens. Lett.2
2006 Kernel Matched Subspace Detectors for Hyperspectral Target Detection
abstract
In this paper, we present a kernel realization of a matched subspace detector (MSD) that is based on a subspace mixture model defined in a high-dimensional feature space associated with a kernel function. The linear subspace mixture model for the MSD is first reformulated in a high-dimensional feature space and then the corresponding expression for the generalized likelihood ratio test (GLRT) is obtained for this model. The subspace mixture model in the feature space and its corresponding GLRT expression are equivalent to a nonlinear subspace mixture model with a corresponding nonlinear GLRT expression in the original input space. In order to address the intractability of the GLRT in the feature space, we kernelize the GLRT expression using the kernel eigenvector representations as well as the kernel trick where dot products in the feature space are implicitly computed by kernels. The proposed kernel-based nonlinear detector, so-called kernel matched subspace detector (KMSD), is applied to several hyperspectral images to detect targets of interest. KMSD showed superior detection performance over the conventional MSD when tested on several synthetic data and real hyperspectral imagery.
Heesung Kwon, Nasser M. Nasrabadi
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Kernel adaptive subspace detector for hyperspectral target detection
abstract
In this paper, we present a kernel-based nonlinear version of the adaptive subspace detector (ASD) that detects signals of interest in a high dimensional (possibly infinite) feature space associated with a certain nonlinear mapping. In order to address the high dimensionality of the feature space, ASD is first implicitly formulated in the feature space which is then converted into an expression in terms of kernel functions via the kernel trick of the Mercer kernels. The proposed kernel-based ASD (KASD) exploits the nonlinear correlations between the spectral bands that is ignored by the conventional ASD. Experimental results based on the given hyperspectral image show that the proposed KASD outperforms the conventional ASD.
Heesung Kwon, Nasser M. Nasrabadi
ICASSP (4)2
2005 Kernel spectral matched filter for hyperspectral target detection
abstract
In this paper a kernel-based nonlinear spectral matched filter is introduced for target detection in hyperspectral imagery. The proposed spectral matched filter is defined in a kernel feature space which is equivalent to a nonlinear matched filter in the original input space. This nonlinear spectral matched filter is based on the notion that performing matched filtering in the high dimensional feature space increases the separability of spectral data mainly because it exploits the higher order correlation between the spectral bands. It is also shown that the nonlinear spectral matched filter can easily be implemented in terms of kernel functions using the so called kernel trick property of the Mercer kernels. The kernel version of the nonlinear spectral matched filter is implemented and simulation results on hyperspectral imagery are shown to outperform the linear version.
Nasser M. Nasrabadi, Heesung Kwon
ICASSP (4)1
2005 Hyperspectral target detection using kernel orthogonal subspace projection
abstract
In this paper, a kernel-based nonlinear version of the orthogonal subspace projection (OSP) classifier is defined in terms of kernel functions. Input data is implicitly mapped into a high dimensional kernel feature space by a nonlinear mapping which is associated with a kernel function. The OSP expression is then derived in the feature space which is kernelized in terms of kernel functions in order to avoid explicit computation in the high dimensional feature space. The resulting kernelized OSP algorithm is equivalent to a nonlinear OSP in the original input space. Experimental results are presented for target detection in hyperspectral imagery and it is shown that the kernel OSP outperforms the conventional OSP classifier.
Heesung Kwon, Nasser M. Nasrabadi
ICIP (2)2
2005 Kernel RX-algorithm: a nonlinear anomaly detector for hyperspectral imagery
abstract
We present a nonlinear version of the well-known anomaly detection method referred to as the RX-algorithm. Extending this algorithm to a feature space associated with the original input space via a certain nonlinear mapping function can provide a nonlinear version of the RX-algorithm. This nonlinear RX-algorithm, referred to as the kernel RX-algorithm, is basically intractable mainly due to the high dimensionality of the feature space produced by the nonlinear mapping function. However, in this paper it is shown that the kernel RX-algorithm can easily be implemented by kernelizing the RX-algorithm in the feature space in terms of kernels that implicitly compute dot products in the feature space. Improved performance of the kernel RX-algorithm over the conventional RX-algorithm is shown by testing several hyperspectral imagery for military target and mine detection.
Heesung Kwon, Nasser M. Nasrabadi
IEEE Trans. Geosci. Remote. Sens.2
2005 Kernel orthogonal subspace projection for hyperspectral signal classification
abstract
In this paper, a kernel-based nonlinear version of the orthogonal subspace projection (OSP) operator is defined in terms of kernel functions. Input data are implicitly mapped into a high-dimensional kernel feature space by a nonlinear mapping, which is associated with a kernel function. The OSP expression is then derived in the feature space, which is kernelized in terms of the kernel functions in order to avoid explicit computation in the high-dimensional feature space. The resulting kernelized OSP algorithm is equivalent to a nonlinear OSP in the original input space. Experimental results are presented for detection of roads, roof tops, mines, and targets in hyperspectral imagery, and it is shown that the kernelized OSP method outperforms the conventional OSP approach.
Heesung Kwon, Nasser M. Nasrabadi
IEEE Trans. Geosci. Remote. Sens.2
2005 A joint compression-discrimination neural transformation applied to target detection
abstract
Many image recognition algorithms based on data-learning perform dimensionality reduction before the actual learning and classification because the high dimensionality of raw imagery would require enormous training sets to achieve satisfactory performance. A potential problem with this approach is that most dimensionality reduction techniques, such as principal component analysis (PCA), seek to maximize the representation of data variation into a small number of PCA components, without considering interclass discriminability. This paper presents a neural-network-based transformation that simultaneously seeks to provide dimensionality reduction and a high degree of discriminability by combining together the learning mechanism of a neural-network-based PCA and a backpropagation learning algorithm. The joint discrimination-compression algorithm is applied to infrared imagery to detect military vehicles.
LipChen Alex Chan, Sandor Z. Der, Nasser M. Nasrabadi
IEEE Trans. Syst. Man Cybern. Part B3
2004 Hyperspectral target detection using kernel matched subspace detector
abstract
In this paper we present a nonlinear realization of a subspace signal detection approach based on the generalized likelihood ratio test (GLRT) - so called matched subspace detectors (MSD). The linear model for MSD is first extended to a high, possibly infinite, dimensional feature space and then the corresponding nonlinear GLRT expression is obtained. In order to address the intractability of the GLRT in the nonlinear feature space we kernelize the nonlinear GLRT using kernel eigenvector representations as well as the kernel trick where dot products in the nonlinear feature space are implicitly computed by kernels. The proposed kernel-based nonlinear detector, so called kernel matched subspace detector (KMSD), is applied to a given hyperspectral imagery - HYDICE (hyperspectral digital imagery collection experiment) images - to detect targets of interest. KMSD showed superior detection performance over MSD for the HYDICE images tested in this paper.
Heesung Kwon, Nasser M. Nasrabadi
ICIP2
2004 Hyperspectral anomaly detection using kernel rx-algorithm
abstract
In this paper we present a nonlinear version of the well-known anomaly detection method, referred to as the RX-algorithm, by extending this algorithm in a feature space associated with the original input space via a certain nonlinear mapping function. An expression for the nonlinear form of the RX-algorithm is derived which is basically intractable mainly due to the high dimensionality of the feature space. We convert the nonlinear RX expression into kernels, which implicitly compute dot products in the nonlinear domain. The proposed kernel RX-algorithm is applied to hyperspectral images for anomaly detection. Improved performance of the kernel RX over the conventional RX is shown for the HYDICE (hyperspectral digital imagery collection experiment) images tested.
Heesung Kwon, Nasser M. Nasrabadi
ICIP2
2004 Kernel-based subpixel target detection in hyperspectral images
abstract
In This work we present a nonlinear realization of a signal detection approach that uses the generalized likelihood ratio tests (GLRTs). It is based on converting the linear mixture subspace model, so called matched subspace detector (MSD) into its corresponding nonlinear subspace model. The linear model for the GLRT of MSD is first extended to a high dimensional feature space (equivalent to a non-linear space in the input domain) and then the corresponding nonlinear GLRT expression is obtained. In order to address the intractability of the GLRT in the feature space we kernelize the nonlinear GLRT using kernel eigenvector representations as well as the kernel trick where dot products in the feature space are implicitly computed by kernels. The proposed kernel-based nonlinear detector, so called kernel matched subspace detector (KMSD), is applied to a given hyperspectral imagery - HYDICE (hyperspectral digital imagery collection experiment) images - to detect targets of interest. KMSD showed superior detection performance over MSD for the HYDICE images tested in this paper.
Heesung Kwon, Nasser M. Nasrabadi
IJCNN2
2003 Improved target detector for FLIR imagery
abstract
Algorithms are considered for searching wide area forward-looking infrared imagery for military vehicles. Wide area search has typically been handled by using a simple detection algorithm with low computational cost to search the entire image or set of images, followed by a clutter rejection algorithm that analyzes only those portions of the image that are marked by the detection algorithm. We start with a feature based detector and a eigen-neural based clutter rejecter, and examine a number of architectures for combining these modules to maximize joint performance. The architectures considered include a clutter rejection threshold method and a nonlinear learning-based combination. The performance of the architectures are compared using a set of several thousand real images.
LipChen Alex Chan, Sandor Z. Der, Nasser M. Nasrabadi
ICASSP (2)3
2003 Time series processing of FLIR imagery for MTI and change detection
abstract
The paper addresses the problem of calibrating FLIR images of a scene that are acquired at different time points to construct information for moving target indication (MTI) and change detection. A signal model is developed to identify variations and imperfections of an FLIR sensor in time. This model is utilized to compensate for relatively slow variations of a bias in the FLIR sensor via a Fourier-based processing. Furthermore, a two-dimensional adaptive filtering method is developed to compensate for variations of the image point response (IPR) of a FLIR sensor as well as sub-pixel changes in the relative coordinates of the sensor-target over time. Results with time series FLIR, data of a scene with an airborne helicopter and a ground target are provided.
Mehrdad Soumekh, Susan S. Young, Nasser M. Nasrabadi
ICASSP (5)3
2003 Projection-based adaptive anomaly detection for hyperspectral imagery
abstract
Adaptive anomaly detectors that find any materials whose spectral characteristics are out of context with those of the neighboring materials are proposed. We use a dual rectangular window that separates the local area into two regions- the inner window region (IWR) and outer window region (OWR). The statistical differences between the IWR and OWR is exploited by generating projection vectors onto which the IWR and OWR vectors are projected. Anomalies are detected if the projection separation between the IWR and OWR vectors is greater than a predefined threshold. Four different methods are used to produce the projection vectors. The proposed anomaly detectors have been applied to HYDICE (HYper-spectral Digital Imagery Collection Experiment) images and detection performance for each method has been measured.
Heesung Kwon, Sandor Z. Der, Nasser M. Nasrabadi
ICIP (1)3
2003 Clutter rejection in FLIR imagery using spatially-varying adaptive filtering
abstract
This paper introduces an adaptive method for rejecting clutter in forward-looking infra-red (FLIR) imagery. In this approach, a spatially-varying two-dimensional adaptive filtering method is developed to identify a metric distance (error energy) between a test scene and a library of reference man-made targets (trucks, tanks, etc.). The function of the 2D spatially-varying adaptive filter is to compensate for: a) variations of image point response (IPR) of a FLIR sensor; b) variations of heat distribution on the test and reference targets; c) subpixel shifts in the relative coordinates of these targets; and d) subtle scaling and rotations. A statistic is constructed that is the least square energy of the error between the test target and its projection into each member of the library of reference target chips based on the above-mentioned spatially-varying 2D adaptive filtering; this database is then used to identify the test chip as a man-made target or clutter. Results with FLIR data of scenes composed of trucks, tanks, and APCs at various angles in different types of clutter environment will be provided.
Susan S. Young, Nasser M. Nasrabadi, Mehrdad Soumekh
ICIP (1)2
2002 Infrared-image classification using support vector machines
abstract
A target recognition classifier for forward-looking infrared (FUR) imagery is developed. A target class is defined as a set of contiguous target-sensor orientations (aspects) for which the associated FLIR imagery is stationary. We designed four sets of templates for each target class, to represent the overall image as well as three class-dependent subcomponents. The templates are designed by using expansion matching (EXM) filters and the Karhunen-Loeve transform (KLT). The feature vectors obtained with these eigen templates are used in the context of a support vector machine (SVM). The performance of the SVM classifier is presented and compared with other competitive classifiers.
Shaorong Chang, Nasser M. Nasrabadi, Lawrence Carin
ICASSP2
2002 A modular clutter rejection technique for FLIR imagery using region-based principal component analysis
Syed A. Rizvi, Nasser M. Nasrabadi
Pattern Recognit.2
2001 Joint compression and discrimination algorithm for clutter rejection
abstract
Many pattern recognition (ATR) systems perform dimensionality reduction on the input imagery to reduce the requirement for a large number of training samples and high computational cost. Most of the dimensionality reduction techniques seek to optimize only the overall data compression, but not the interclass discriminability. We present a neural-network-based algorithm that simultaneously achieves data compression and target discriminability by adjusting the pre-trained base components to maximize separability between classes. This allows the classifiers to operate at a higher level of efficiency and generalization capability on low-dimensional data. We have applied this technique to the problem of automatic detection of military vehicles in infrared imagery.
Sandor Z. Der, LipChen Alex Chan, Nasser M. Nasrabadi
ICIP (1)3
2001 An adaptive segmentation algorithm using iterative local feature extraction for hyperspectral imagery
abstract
We present an adaptive segmentation algorithm based on the iterative use of a modified minimum-distance classifier. Local adaptivity is achieved by gradually updating each class centroid over a local region whose size is reduced progressively during a segmentation process. The proposed method provides improved segmentation performance over template matching segmentation techniques because it adapts to the local context. The proposed algorithm can be applied to virtually any hyperspectral image regardless of size, dimensionality, and spectral sensitivity. Experimental results on a set of visible to near-infrared hyperspectral images using both the proposed algorithm and a standard template matching technique are presented.
Heesung Kwon, Nasser M. Nasrabadi
ICIP (1)2
2001 Unsupervised segmentation algorithm based on an iterative spectral dissimilarity measure for hyperspectral imagery
Heesung Kwon, Sandor Z. Der, Nasser M. Nasrabadi
VCIP3
2001 Experimental Evaluation of FLIR ATR Approaches - A Comparative Study
Baoxin Li, Rama Chellappa, Qinfen Zheng, Sandor Z. Der, Nasser M. Nasrabadi, LipChen Alex Chan, Lin-Cheng Wang
Comput. Vis. Image Underst.5
2000 Automatic target detection using dualband infrared imagery
abstract
An automatic target detector often produces too many false alarms that could bog down the performance of a subsequent target classifier. Therefore, we need a good clutter rejector to remove as many clutterers as possible, before feeding the most likely target detections to the classifier. We investigate the benefits of using dual-band forward-looking infrared images to improve the performance of an eigen-neural based clutter rejector. With individual or combined bands as input, we use either principal component analysis or the eigenspace separation transform to perform feature extraction and dimensionality reduction. The transformed data is then fed to a properly trained multilayer perceptron that predicts the identity of the input, which is either a target or clutter. Experimental results are presented on a dataset of real dualband images.
LipChen Alex Chan, Sandor Z. Der, Nasser M. Nasrabadi
ICASSP3
2000 Dual-Band Passive Infrared Imagery for Automatic Clutter Rejection
abstract
In a typical automatic target recognition (ATR) system, the target detection module often produces too many false alarms, which may severely inhibit the effectiveness of the subsequent target classifier module. An effective clutter rejector is therefore needed between these two modules to reduce these false alarms. We explore the potential benefits of using dual-band infrared imagery to improve the performance of an eigenneural-based clutter rejector. Individual or combined bands of images are first compressed through eigenspace transformations, such as principal component analysis. The transformed data are then fed to a neural network that decides whether the input is a target or clutter. A huge and realistic set of dual-band passive infrared images was used in a series of experiments.
LipChen Alex Chan, Sandor Z. Der, Nasser M. Nasrabadi
ICIP3
2000 An Adaptive Hierarchical Segmentation Algorithm Based on Quadtree Decomposition for Hyperspectral Imagery
abstract
We present an adaptive hierarchical segmentation algorithm based on quadtree decomposition and a modified minimum-distance classifier. The proposed algorithm uses quadtree decomposition because this technique can adapt to the local characteristics of the hyperspectral data. A feature vector (i.e., a class centroid) for each material type is recursively estimated and updated, so that increasingly accurate segmentation results are achieved as the decomposition proceeds. The proposed method provides improved segmentation performance over standard template-matching segmentation techniques because it adapts to the local context. It also imposes a spatial smoothness constraint on the pixel classification that provides spatial continuity during the segmentation process. Both the proposed algorithm and a standard template-matching technique were applied to a set of visible to near-infrared hyperspectral images results are presented.
Heesung Kwon, Sandor Z. Der, Nasser M. Nasrabadi
ICIP3
2000 A Modular Clutter Rejection Technique for FLIR Imagery Using Region-Based Principal Component Analysis
abstract
In this paper, a modular clutter rejection technique using region-based principal component analysis (PCA) is proposed. Our modular clutter rejection system uses dynamic ROI (region of interest) extraction to overcome the problem of poorly centered targets. In dynamic ROI extraction, a representative ROI is moved in several directions with respect to the center of the potential target image to extract a number of ROIs. Each module in the proposed system applies region-based PCA to generate the feature vectors, which are subsequently used to decide about the identity of the potential target. We also present experimental results using real-life data evaluating and comparing the performance of the clutter rejection systems with static and dynamic ROI extraction.
Syed A. Rizvi, Tarek N. Saadawi, Nasser M. Nasrabadi
ICIP3
2000 A clutter rejection technique for FLIR imagery using region based principal component analysis
Syed A. Rizvi, Tarek N. Saadawi, Nasser M. Nasrabadi
Pattern Recognit.3
1999 An admission control framework to support media-streaming over packet-switched networks
abstract
This paper presents an admission control framework to support 'streaming' of audio-video data over packet-switched networks. To ensure smooth functioning, the network provides a degree of assurance to the applications in the form of soft QOS guarantees. A measurement based admission control algorithm is developed to enable the network to provide the required service assurances. Sources are characterized using a statistical bounding model, which is fairly universal. A measurement procedure is first developed for online estimation of model parameters. This measurement procedure is then used in the admission control framework. Simulation results are presented and the performance of the algorithms are evaluated.
Mahesh Venkatraman, Nasser M. Nasrabadi
ICC2
1999 Bipolar Elgenspace Separation Transformation for Automatic Clutter Rejection
abstract
A major problem for a detection algorithm is the vast amount of false alarms normally generated. This amount of false alarms has to be substantially reduced so that a typical target classifier in the subsequent stage may work reasonably. We use the bipolar eigenspace separation transformation (BEST) and neural network techniques to improve the clutter rejection performance of an automatic target detector. Experiments have been conducted on huge and realistic datasets of forward looking infrared (FLIR) imagery. Compared to the performance of the unipolar EST and principal component analysis (PCA) with the same datasets, significant improvement in clutter rejection rates has been achieved with BEST.
LipChen Alex Chan, Nasser M. Nasrabadi, Don J. Torrieri
ICIP (1)2
1999 Lossless Image Compression Using Modular Differential Pulse Code Modulation
abstract
This paper presents a new lossless image compression technique called modular differential pulse code modulation (MDPCM). The proposed technique consists of a VQ classifier and several neural network class predictors. The classifier uses the four previously encoded pixels to identify the class of the current pixel (the pixel to be predicted). The current pixel is then predicted by the corresponding class predictor. Experimental results demonstrate that the proposed technique reduces the bit rate by as much as 10 percent when compared to the lossless JPEG.
Syed A. Rizvi, Nasser M. Nasrabadi
ICIP (1)2
1999 A Clutter Rejection Technique for FLIR Imagery Using Region-Based Principal Component Analysis
abstract
The preprocessing stage of an automatic target recognition system extracts areas containing potential targets from a battlefield scene. These potential target images are then sent to the classification stage to identify the targets. It is highly desirable at the preprocessing stage to minimize the incorrect rejection rate. This, however, results in a high false alarm rate. The high false alarm rate, in turn, makes subsequent target classification decisions unreliable. We present a new technique to reject false alarms (clutter images) produced by the preprocessing stage. Our technique, which we call region-based principal component analysis (PCA), uses topological features of the targets to reject false alarms. In this technique a potential target is divided into several regions and a PCA is performed on each region to extract regional feature vectors. We propose to use regional feature vectors of arbitrary shapes and dimensions that are optimized for the topology of a target in a particular region. These regional feature vectors are then used by a two-class classifier based on the learning vector quantization to decide whether a potential target is a false alarm or a real target.
Syed A. Rizvi, Nasser M. Nasrabadi, Sandor Z. Der
ICIP (4)2
1999 Selectively optimized networks for automatic clutter rejection
abstract
An effective clutter rejection scheme is needed to distinguish between clutter and targets in a high-performance automatic target recognition (ATR) system. We present a clutter rejection scheme that consists of an eigenspace transformation and a multilayer perceptron (MLP). We use either principal component analysis (PCA) or the eigenspace separation transform (EST) to perform feature extraction and dimensionality reduction. The transformed data is then fed to an MLP that predicts the identity of the input, which is either a target or clutter. We devise an MLP training algorithm that seeks to maximize the class separation at a given false-alarm rate, which does not necessarily minimize the average deviation of the MLP outputs from their target valves. Experimental results are presented on a huge and realistic dataset of forward-looking infrared imagery.
LipChen Alex Chan, Nasser M. Nasrabadi, Don J. Torrieri
IJCNN2
1999 Rate-constrained modular predictive residual vector quantization of digital images
abstract
A novel modular coding paradigm is investigated using residual vector quantization (RVQ) with memory that incorporates a modular neural network vector predictor in the feedback loop. A modular neural network predictor consists of several expert networks that are optimized for predicting a particular class of data. The predictor also consists of an integrating unit that mixes the outputs of the expert networks to form the final output of the prediction system. The vector quantizer also has a modular structure. The proposed modular predictive RVQ (modular PRVQ) is designed by imposing a constraint on the output rate of the system. Experimental results show that the modular PRVQ outperforms simple PRVQ by as much as 1 dB at low bit rates. Furthermore, for the same peak signal-to-noise ratio (PSNR), the modular PRVQ reduces the bit rate by more than a half when compared to the JPEG algorithm.
Syed A. Rizvi, Lin-Cheng Wang, Nasser M. Nasrabadi
IEEE Signal Process. Lett.3
1998 Very-Low-Bit-Rate Video Coding using Quadtree Decomposition and Cache-based Vector Quantization
Heesung Kwon, Mahesh Venkatraman, Nasser M. Nasrabadi
ICIP (3)3
1998 Guest Editorial Applications Of Artificial Neural Networks To Image Processing
Rama Chellappa, Kunihiko Fukushima, Aggelos K. Katsaggelos, Sun-Yuan Kung, Yann LeCun, Nasser M. Nasrabadi, Tomaso A. Poggio
IEEE Trans. Image Process.6
1998 Automatic target recognition using a feature-decomposition and data-decomposition modular neural network
abstract
A modular neural network classifier has been applied to the problem of automatic target recognition using forward-looking infrared (FLIR) imagery. The classifier consists of several independently trained neural networks. Each neural network makes a decision based on local features extracted from a specific portion of a target image. The classification decisions of the individual networks are combined to determine the final classification. Experiments show that decomposition of the input features results in performance superior to a fully connected network in terms of both network complexity and probability of classification. Performance of the classifier is further improved by the use of multiresolution features and by the introduction of a higher level neural network on the top of the individual networks, a method known as stacked generalization. In addition to feature decomposition, we implemented a data-decomposition classifier network and demonstrated improved performance. Experimental results are reported on a large set of real FLIR images.
Lin-Cheng Wang, Sandor Z. Der, Nasser M. Nasrabadi
IEEE Trans. Image Process.3
1998 A modular neural network vector predictor for predictive image coding
abstract
In this paper, we present a modular neural network vector predictor that improves the predictive component of a predictive vector quantization (PVQ) scheme. The proposed vector prediction technique consists of five dedicated predictors (experts), where each expert predictor is optimized for a particular class of input vectors. An input vector is classified into one of five classes, based on its directional variances. One expert predictor is optimized for stationary blocks, and each of the other four expert predictors are optimized to predict horizontal, vertical, 45 degrees , and 135 degrees diagonally oriented edge-blocks, respectively. An integrating unit is then used to select or combine the outputs of the experts in order to form the final output of the modular network. Therefore, no side information is transmitted to the receiver about the selected predictor or the integration of the predictors. Experimental results show that the proposed scheme gives an improvement of 1.7 dB over a single multilayer perceptron (MLP) predictor. Furthermore, if the information about the predictor selection is sent to the receiver, the improvement could be up to 3 dB over a single MLP predictor. The perceptual quality of the predicted images is also significantly improved.
Lin-Cheng Wang, Syed A. Rizvi, Nasser M. Nasrabadi
IEEE Trans. Image Process.3
1997 A predictive residual VQ using modular neural network vector predictor
abstract
This paper presents a predictive residual vector quantization (PRVQ) scheme using a modular neural network vector predictor. The proposed PRVQ scheme takes the advantage of the high prediction gain and the improved edge fidelity of a modular neural network vector predictor in order to implement a high performance vector quantization (VQ) scheme with low search complexity and a high perceptual quality. Simulation results show that the proposed PRVQ with modular vector predictor outperforms the equivalent PRVQ with general vector predictor (operating at the same bit rate) by more than 1 dB. Furthermore, the perceptual quality of the reconstructed image is also improved.
Lin-Cheng Wang, Syed A. Rizvi, Nasser M. Nasrabadi
ICASSP3
1997 Rate-Constrained Modular Predictive Residual Vector Quantization
abstract
This paper investigates a novel modular image coding paradigm using residual vector quantization (RVQ) with memory that incorporates a modular neural network vector predictor in the feedback loop. A modular neural network predictor consists of several expert networks, where each expert network is optimized for predicting a particular class of data, and an integrating unit that mixes the outputs of the expert networks in order to form the final output of the prediction system. The vector quantizer also has a modular structure. The proposed modular predictive RVQ (MPRVQ) is designed by imposing a constraint on the output rate of the system. Experimental results show that the modular PRVQ outperforms simple PRVQ by as much as 1 dB at low bit rates. Furthermore, for the same PSNR, the modular PRVQ reduces the bit rate by more than a half when compared to the JPEG algorithm.
Syed A. Rizvi, Lin-Cheng Wang, Nasser M. Nasrabadi
ICIP (3)3
1997 Combination of Two Learning Algorithms for Automatic Target Recognition
abstract
Composite classifiers consisting of a number of component classifiers have been designed and evaluated on the problem of automatic target recognition (ATR) using a large set of real forward-looking infrared (FLIR) imagery. Two existing classifiers are used as the building blocks for our composite classifiers. The performance of the proposed composite classifiers are compared based on their classification ability and computational complexity. It is demonstrated that the composite classifier based on a cascade architecture greatly reduces the computational complexity with a statistically insignificant decrease in performance in comparison to standard classifier fusion algorithms.
Lin-Cheng Wang, LipChen Alex Chan, Nasser M. Nasrabadi, Sandor Z. Der
ICIP (1)3
1997 Very Low Bit-Rate Video Coding Using Variable Block-Size Entropy-Constrained Residual Vector Quantizers
abstract
We present a practical video coding algorithm for use at very low bit rates. For efficient coding at very low bit rates, it is important to intelligently allocate bits within a frame, and so a powerful variable-rate algorithm is required. We use vector quantization to encode the motion-compensated residue signal in an H.263-like framework. For a given complexity, it is well understood that structured vector quantizers perform better than unstructured and unconstrained vector quantizers. A combination of structured vector quantizers is used in our work to encode the video sequences. The proposed codec is a multistage residual vector quantizer, with transform vector quantizers in the initial stages. The transform-VQ captures the low-frequency information, using only a small portion of the bit budget, while the later stage residual VQ captures the high-frequency information, using the remaining bits. We used a strategy to adaptively refine only areas of high activity, using recursive decomposition and selective refinement in the later stages. An entropy constraint was used to modify the codebooks to allow better entropy coding of the indexes. We evaluate the performance of the proposed codec, and compare this data with the performance of the H.263-based codec. Experimental results show that the proposed codec delivered significantly better perceptual quality along with better quantitative performance.
Heesung Kwon, Mahesh Venkatraman, Nasser M. Nasrabadi
IEEE J. Sel. Areas Commun.3
1997 Finite-state residual vector quantization using a tree-structured competitive neural network
abstract
Finite-state vector quantization (FSVQ) is known to give better performance than the memoryless vector quantization (VQ). This paper presents a new FSVQ scheme, called finite-state residual vector quantization (FSRVQ), in which each state uses a residual vector quantizer (RVQ) to encode the input vector. This scheme differs from the conventional FSVQ in that the state-RVQ codebooks encode the residual vectors instead of the original vectors. A neural network predictor estimates the current block based on the four previously encoded blocks. The predicted vector is then used to identify the current state as well as to generate a residual vector (the difference between the current vector and the predicted vector). This residual vector is encoded using the current state-RVQ codebooks. A major task in designing our proposed FSRVQ is the joint optimization of the next-state codebook and the state-RVQ codebooks. This is achieved by introducing a novel tree-structured competitive neural network in which the first layer implements the next-state function, and each branch of the tree implements the corresponding state-RVQ. A joint training algorithm is also developed that mutually optimizes the next-state and the state-RVQ codebooks for the proposed FSBVQ. Joint optimization of the next-state function and the state-RVQ codebooks eliminates a large number of redundant states in the conventional FSVQ design; consequently, the memory requirements are substantially reduced in the proposed FSRVQ scheme. The proposed FSRVQ can be designed for high bit rates due to its very low memory requirements and the low search complexity of the state-RVQ's. Simulation results show that the proposed FSRVQ scheme outperforms conventional FSVQ schemes both in terms of memory requirements and the visual quality of the reconstructed image. The proposed FSRVQ scheme also outperforms JPEG (the current standard for still image compression) at low bit rates.
Syed A. Rizvi, Nasser M. Nasrabadi
IEEE Trans. Circuits Syst. Video Technol.2
1997 Nonlinear vector prediction using feed-forward neural networks
abstract
The performance of a classical linear vector predictor is limited by its ability to exploit only the linear correlation between the blocks. However, a nonlinear predictor exploits the higher order correlations among the neighboring blocks, and can predict edge blocks with increased accuracy. We have investigated several neural network architectures that can be used to implement a nonlinear vector predictor, including the multilayer perceptron (MLP), the functional link (FL) network, and the radial basis function (RBF) network. Our experimental results show that a neural network predictor can predict the blocks containing edges with a higher accuracy than a linear predictor.
Syed A. Rizvi, Lin-Cheng Wang, Nasser M. Nasrabadi
IEEE Trans. Image Process.3
1997 Object recognition using multilayer Hopfield neural network
abstract
An object recognition approach based on concurrent coarse-and-fine matching using a multilayer Hopfield neural network is presented. The proposed network consists of several cascaded single-layer Hopfield networks, each encoding object features at a distinct resolution, with bidirectional interconnections linking adjacent layers. The interconnection weights between nodes associating adjacent layers are structured to favor node pairs for which model translation and rotation, when viewed at the two corresponding resolutions, are consistent. This interlayer feedback feature of the algorithm reinforces the usual intralayer matching process in the conventional single-layer Hopfield network in order to compute the most consistent model-object match across several resolution levels. The performance of the algorithm is demonstrated for test images containing single objects, and multiple occluded objects. These results are compared with recognition results obtained using a single-layer Hopfield network.
Susan S. Young, Peter D. Scott, Nasser M. Nasrabadi
IEEE Trans. Image Process.3
1996 Multi-Stage Target Recognition Using Modular Vector Quantizers and Multilayer Perceptrons
abstract
An automatic target recognition (ATR) classifier is proposed that uses modularly cascaded vector quantizers (VQs) and multilayer perceptrons (MLPs). A dedicated VQ codebook is constructed for each target class at a specific range of aspects, which is trained with the K-means algorithm and a modified learning vector quantization (LVQ) algorithm. Each final codebook is expected to give the lowest mean squared error (MSE) for its correct target class at a given range of aspects. These MSEs are then processed by an array of window MLPs and a target MLP consecutively. In the spatial domain, target recognition rates of 90.3 and 65.3 percent are achieved for moderately and highly cluttered test sets, respectively. Using the wavelet decomposition with an adaptive and independent codebook per sub-band, the VQs alone have produced recognition rates of 98.7 and 69.0 percent on more challenging training and test sets, respectively.
LipChen Alex Chan, Nasser M. Nasrabadi, Vincent Mirelli
CVPR2
1996 Automatic target recognition using modularly cascaded vector quantizers and multilayer perceptrons
abstract
An automatic target recognition classifier is constructed of a set of vector quantizers (VQs) and multilayer perceptrons (MLPs) that are modularly cascaded. A dedicated VQ codebook is constructed for each target at a specific range of aspects. Each codebook is a set of block feature templates that are iteratively adapted to represent a particular target at a specific range of aspects. These templates are further trained by a modified learning vector quantization (LVQ) algorithm that enhances their discriminatory power. The mean squared errors resulting from matching the input image with the block templates in each each codebook are input to an array of window MLPs (WMLPs). Each WMLP is trained to recognize its intended-target at a specific range of aspects. The outputs of the WMLPs are manipulated and fed into a target MLP (TMLP) that produces the final recognition results. A recognition rate of 65.3 percent is achieved on a highly cluttered test set.
LipChen Alex Chan, Nasser M. Nasrabadi, Vincent Mirelli
ICASSP2
1996 Segmentation based wavelet coding of digital images
abstract
In this paper, we present a segmentation based wavelet coding scheme, in which an image is segmented into two regions: stationary areas (background) and the areas containing edge information (foreground). These regions are then encoded independently using two dedicated encoders that are optimized for each region. A 2-D edge operator is used for segmenting the image. We use the embedded zerotree wavelet (EZW) algorithm for encoding the background due to its good performance on stationary areas. The foreground area is, however, encoded using a predictive residual vector quantizer (PRVQ). Experimental results show that the proposed technique improves the quality of the reconstructed images, both numerically (in terms of mean square error) and perceptually when compared to EZW at the same bit rate.
Euee S. Jang, Heesung Kwon, Lin-Cheng Wang, Syed A. Rizvi, Nasser M. Nasrabadi
ICASSP5
1996 Subband image coding using block-zero tree coding and vector quantization
abstract
The need for developing effective coding techniques for various multimedia services is increasing in order to meet the demand for image and video data. In this paper, a block-zero tree coding (BZTC) algorithm is proposed for multilayer coding and progressive transmission of images. The BZTC algorithm is constructed as an embedded coding so that the encoding and decoding process can be terminated at any point and allowing reasonable image quality. Some features of our BZTC algorithm are (1) prediction of the significance of blocks across scale, (2) state transition rules for representing the significance block map, and (3) block coding by vector quantization using a multiband codebook consisting of several subcodebooks dedicated for each band at a given threshold.
Sahng H. Park, Hyeon J. Moon, Nasser M. Nasrabadi
ICASSP3
1996 A modular neural network vector predictor for predictive VQ
abstract
In this paper, we present a modular neural network vector predictor in order to improve the predictive component of a predictive vector quantization (PVQ) scheme. The proposed vector prediction technique consists of five dedicated predictors (experts), where each expert predictor is optimized for a particular class of input vectors. An input vector is classified into one of five classes based on its directional variances. One expert is optimized for stationary blocks, and each of the other four experts are optimized to predict horizontal, vertical, 45/spl deg/, and 135/spl deg/ diagonally oriented edge-blocks, respectively. An integrating unit is then used to select or combine the outputs of the experts in order to form the final output of the modular network. Therefore, no side information is required to transmit to the receiver about the predictor selection. Experimental results show that the proposed scheme gives an improvement of 1 to 1.5 dB better than a single multilayer perceptron (MLP) predictor. However, if the information about the predictor selection is sent to the receiver, the improvement could be up to 3 dB over the single MLP predictor. The perceptual quality of the predicted images is also significantly improved.
Lin-Cheng Wang, Syed A. Rizvi, Nasser M. Nasrabadi
ICIP (3)3
1996 Large Vocabulary Recognition of On-Line Handwritten Cursive Words
abstract
This paper presents a writer independent system for large vocabulary recognition of on-line handwritten cursive words. The system first uses a filtering module, based on simple letter features, to quickly reduce a large reference dictionary (lexicon) to a more manageable size; the reduced lexicon is subsequently fed to a recognition module. The recognition module uses a temporal representation of the input, instead of a static two-dimensional image, thereby preserving the sequential nature of the data and enabling the use of a Time-Delay Neural Network (TDNN); such networks have been previously successful in the continuous speech recognition domain. Explicit segmentation of the input words into characters is avoided by sequentially presenting the input word representation to the neural network-based recognizer. The outputs of the recognition module are collected and converted into a string of characters that is matched against the reduced lexicon using an extended Damerau-Levenshtein function. Trained on 2,443 unconstrained word images (11 k characters) from 55 writers and using a 21 k lexicon we reached a 97.9% and 82.4% top-5 word recognition rate on a writer-dependent and writer-independent test, respectively.
Giovanni Seni, Rohini K. Srihari, Nasser M. Nasrabadi
IEEE Trans. Pattern Anal. Mach. Intell.3
1996 Neural network architectures for vector prediction
abstract
A vector predictor is an integral part of a predictive vector quantization coding scheme. The conventional techniques for designing a nonlinear predictor are extremely complex and suboptimal due to the absence of a suitable model for the source data. We investigated several neural network architectures that can be used to implement a nonlinear vector predictor, including the multilayer perceptron, the functional link network and the radial basis function network. We also evaluated and compared the performance of these neural network predictors with that of a linear vector predictor. Our experimental results show that a neural network predictor can predict the blocks containing edges with a higher accuracy than a linear predictor. However, the performance of a neural network predictor is comparable to that of a linear predictor for predicting the stationary and shade blocks.
Syed A. Rizvi, Lin-Cheng Wang, Nasser M. Nasrabadi
Proc. IEEE3
1996 Advances in residual vector quantization: a review
abstract
Advances in residual vector quantization (RVQ) are surveyed. Definitions of joint encoder optimality and joint decoder optimality are discussed. Design techniques for RVQs with large numbers of stages and generally different encoder and decoder codebooks are elaborated and extended. Fixed-rate RVQs, and variable-rate RVQs that employ entropy coding are examined. Predictive and finite state RVQs designed and integrated into neural-network based source coding structures are revisited. Successive approximation RVQs that achieve embedded and refinable coding are reviewed. A new type of successive approximation RVQ that varies the instantaneous block rate by using different numbers of stages on different blocks is introduced and applied to image waveforms, and a scalar version of the new residual quantizer is applied to image subbands in an embedded wavelet transform coding system.
Christopher F. Barnes, Syed A. Rizvi, Nasser M. Nasrabadi
IEEE Trans. Image Process.3
1996 Scalar-vector quantization of medical images
abstract
A new coding scheme based on the scalar-vector quantizer (SVQ) is developed for compression of medical images. The SVQ is a fixed rate encoder and its rate-distortion performance is close to that of optimal entropy-constrained scalar quantizers (ECSQs) for memoryless sources. The use of a fixed-rate quantizer is expected to eliminate some of the complexity of using variable-length scalar quantizers. When transmission of images over noisy channels is considered, our coding scheme does not suffer from error propagation that is typical of coding schemes using variable-length codes. For a set of magnetic resonance (MR) images, coding results obtained from SVQ and ECSQ at low bit rates are indistinguishable. Furthermore, our encoded images are perceptually indistinguishable from the original when displayed on a monitor. This makes our SVQ-based coder an attractive compression scheme for picture archiving and communication systems (PACS). PACS are currently under study for use in an all-digital radiology environment in hospitals, where reliable transmission, storage, and high fidelity reconstruction of images are desired.
Nader Mohsenian, Homayoun Shahri, Nasser M. Nasrabadi
IEEE Trans. Image Process.3
1995 A hashing-based scheme for organizing vector quantization codebook
abstract
One of the problems in vector quantization (VQ) is its relatively long encoding time especially when an exhaustive search is made for the codevector. This paper presents a hashing-based technique to organize the codebook so that the search time can be significantly reduced. Hashing gives the speed advantages of a direct search, while maintaining a codebook of reasonable size. Experiments show that hashing-based VQ sustained image quality as the encoding time was reduced, while full search VQ suffered greatly. For example, for 2/spl times/2 vectors and with 1024 codebook entries, encoding time was reduced by a factor of 10 without significant loss of image quality.
Chang Y. Choo, Erik Kristenson, Nasser M. Nasrabadi, Xiaonong Ran
ICASSP3
1995 Finite state residual vector quantization using tree-structured competitive neural network
abstract
The performance of an ordinary vector quantizer (VQ) can be improved by incorporating memory in the VQ scheme. A VQ scheme with finite memory known as finite state vector quantization has been shown to give better performance than the ordinary VQ. The major problems with the FSVQ are the lack of accurate prediction of the current state, the state codebook design, and the amount of memory required to store all the state codebooks. The paper presents a new FSVQ scheme called finite-state residual vector quantization (FSRVQ) in which a neural network based state prediction is used. Furthermore, a novel tree-structured competitive neural network is used to jointly design the next-state and the state codebooks for the proposed FSRVQ. Simulation results show that the new scheme gives better performance with significant reduction in the memory requirement when compared to the conventional FSVQ schemes.
Syed A. Rizvi, Nasser M. Nasrabadi
ICASSP2
1995 SAR moving target detection and identification using stochastic gradient techniques
abstract
This paper presents methods for detecting and identifying moving targets in a synthetic aperture radar (SAR) scene. An analytical expression is derived for the coherent SAR signature of a target. SAR system model of a moving target is developed. These principles are then used to construct a SAR signal statistic (energy function) in a parameter space which is defined by the target's coordinates, speed, and coherent SAR signature. Stochastic gradient techniques are used to search for the maximum point of this energy function which is located at the desired target's parameters.
Susan S. Young, Nasser M. Nasrabadi, Mehrdad Soumekh
ICASSP2
1995 Entropy-constrained predictive residual vector quantization of digital images
abstract
A major problem with a VQ based image compression scheme is its codebook search complexity. Recently, a new VQ scheme called predictive residual vector quantizer (PRVQ) was proposed by Rizvi and Nasrabadi (see Proc. IEEE Int. Conf. Image Processing (Austin), vol.1, p.608-12, Nov. 13-16, 1994) which has a performance very close to that of the predictive vector quantizer (PVQ) with very low search complexity. This paper presents a new variable-rate VQ scheme called entropy-constrained PRVQ (EC-PRVQ), which is designed by imposing a constraint on the output entropy of the PRVQ. The proposed EC-PRVQ is found to give a good rate-distortion performance and clearly outperforms the state-of-the-art image compression algorithm developed by the Joint Photographic Experts Group (JPEG). The robustness of EC-PRVQ is demonstrated by encoding several test images taken from outside the training data.
Syed A. Rizvi, Nasser M. Nasrabadi, Lin-Cheng Wang
ICIP (3)2
1995 Neural network vector predictors with application to image coding
abstract
A vector predictor is an integral part of the predictive vector quantization (PVQ) scheme. The performance of a predictor deteriorates as the vector dimension (block size) is increased. This makes it necessary to investigate new design techniques in order to design a vector predictor which gives better performance when compared to a conventional vector predictor. This paper investigates several neural network configurations which can be employed in order to design a vector predictor. The following architectures are investigated: (a) multilayer perceptron, (b) functional link network, and (c) radial basis function network. The performance of the above mentioned neural network vector predictors is evaluated and compared with that of a linear vector predictor.
Syed A. Rizvi, Lin-Cheng Wang, Nasser M. Nasrabadi
ICIP (3)3
1995 Variable-rate predictive residual vector quantizer
abstract
A major problem with a VQ-based image compression scheme is its codebook search complexity. Recently, a predictive residual vector quantizer (PRVQ) was proposed by Rizvi and Nasrabadi (see IEEE Int. Conf. Image Processing, Austin, vol.1, p.608-612, Nov. 13-16, 1994). This scheme has a very low search complexity, and its performance is very close to that of the predictive vector quantizer (PVQ). The article presents a new VQ scheme called variable-rate PRVQ (VR-PRVQ), which is designed by imposing a constraint on the output entropy of the PRVQ. The proposed VR-PRVQ is found to give an excellent rate-distortion performance and clearly outperforms the state-of-the-art image compression algorithm developed by the Joint Photographic Experts Group (JPEG).>
Syed A. Rizvi, Nasser M. Nasrabadi
IEEE Signal Process. Lett.2
1995 Subband coding with multistage VQ for wireless image communication
abstract
A subband multistage vector quantization (SB-MSVQ) coding technique is proposed for wireless image communication. The Digital European Cordless Telecommunications (DECT) system has been used as a wireless communication environment. The impact of channel fading on the SB-MSVQ coding technique is investigated and its performance is compared with that of the baseline JPEG standard. Simulation results show that the proposed technique, SB-MSVQ, is very robust to channel fading errors even at a very low channel power, whereas the baseline JPEG suffers from channel fading errors which are visually very annoying. Due to the inherent structure of SB-MSVQ, the important low frequency information of an image is packed into a small portion of the encoded data. If this portion is received intact by the SB-MSVQ decoder, the overall quality degradation of the reconstructed image is marginal.>
Euee S. Jang, Nasser M. Nasrabadi
IEEE Trans. Circuits Syst. Video Technol.2
1995 An efficient Euclidean distance computation for vector quantization using a truncated look-up table
abstract
Vector quantizer (VQ) encoders generally use Euclidean distance measure to encode the vectors. The major computation in the Euclidean distance is square (multiplication operation) of the difference between the vector components. This article explores Euclidean distance computation and introduces a new technique which uses a truncated look-up table (LUT) to store a small set of repeatedly generated scalars. Specifically, for numbers represented by m bits, this technique requires to store only 2/sup m/ product terms instead of 2/sup m//spl times/2/sup m/ product terms needed to store in a conventional LUT.>
Syed A. Rizvi, Nasser M. Nasrabadi
IEEE Trans. Circuits Syst. Video Technol.2
1995 Next-state functions for finite-state vector quantization
Nasser M. Nasrabadi, Syed A. Rizvi
IEEE Trans. Image Process.1
1995 Predictive residual vector quantization [image coding]
abstract
This paper presents a new vector quantization technique called predictive residual vector quantization (PRVQ). It combines the concepts of predictive vector quantization (PVQ) and residual vector quantization (RVQ) to implement a high performance VQ scheme with low search complexity. The proposed PRVQ consists of a vector predictor, designed by a multilayer perceptron, and an RVQ that is designed by a multilayer competitive neural network. A major task in our proposed PRVQ design is the joint optimization of the vector predictor and the RVQ codebooks. In order to achieve this, a new design based on the neural network learning algorithm is introduced. This technique is basically a nonlinear constrained optimization where each constituent component of the PRVQ scheme is optimized by minimizing an appropriate stage error function with a constraint on the overall error. This technique makes use of a Lagrangian formulation and iteratively solves a Lagrangian error function to obtain a locally optimal solution. This approach is then compared to a jointly designed and a closed-loop design approach. In the jointly designed approach, the predictor and quantizers are jointly optimized by minimizing only the overall error. In the closed-loop design, however, a predictor is first implemented; then the stage quantizers are optimized for this predictor in a stage-by-stage fashion. Simulation results show that the proposed PRVQ scheme outperforms the equivalent RVQ (operating at the same bit rate) and the unconstrained VQ by 2 and 1.7 dB, respectively. Furthermore, the proposed PRVQ outperforms the PVQ in the rate-distortion sense with significantly lower codebook search complexity.
Syed A. Rizvi, Nasser M. Nasrabadi
IEEE Trans. Image Process.2
1994 An on-line cursive word recognition system
abstract
This paper presents a system for large vocabulary recognition of on-line handwritten cursive words. The system first uses a filtering module, based on simple letter features, to quickly reduce a large reference dictionary to a smaller number of candidates; the reduced lexicon along with the original input is subsequently fed to a recognition module. In order to exploit the sequential nature of the temporal data, we employ a TDNN-style network architecture which has been successfully used in the speech recognition domain. Explicit segmentation of the input words into characters is avoided by using a sliding window concept where the input word representation (a set of frames) is presented to the neural network-based recognizer sequentially. The outputs of the recognition module are collected and converted into a string of characters that can be matched with the candidate words. A description of the complete system and its components is given.>
Giovanni Seni, Nasser M. Nasrabadi, Rohini K. Srihari
CVPR2
1994 Object recognition using multi-layer Hopfield neural network
abstract
An object recognition approach based on concurrent coarse-and-fine matching using a multi-layer Hopfield neural network is presented. The proposed network consists of several cascaded single layer Hopfield networks, each encoding object features at a distinct resolution, with bidirectional interconnections linking adjacent layers. The interconnection weights between nodes associating adjacent layers are structured to favor node pairs for which model translation and rotation, when viewed at the two corresponding resolutions, are consistent. This inter-layer feedback feature of the algorithm reinforces the usual intra-layer matching process in conventional single layer Hopfield nets in order to compute the model-object match which is most consistent across several resolution levels. The performance of the algorithm is demonstrated in cases of images containing single and multiple occluded objects. These results are compared with recognition results obtained using a single layer Hopfield network.>
Susan S. Young, Peter D. Scott, Nasser M. Nasrabadi
CVPR3
1994 Next-state functions for finite-state vector quantization
abstract
A finite-state vector quantizer called dynamic finite-state vector quantization (DFSVQ) is investigated with regard to its subcodebook construction. In DFSVQ each input vector is encoded by a small codebook called the subcodebook which is created from a much larger codebook called the supercodebook by selecting (reordering procedure) a set of appropriate codevectors. The performance of the DFSVQ depends on this reordering procedure. In the paper, several reordering procedures including the conditional histogram, address prediction, vector prediction, nearest neighbor design and the frequency usage of codevectors are introduced and their performance are evaluated by comparing their hit ratios (the number of blocks encoded by the subcodebook) and their computational complexity.>
Nasser M. Nasrabadi, Syed A. Rizvi
ICASSP (5)1
1994 Evaluation of Design Parameters for a Cache Vector Quantization System
abstract
Observation of the usage of codevectors in vector quantization (VQ) reveals the locality of the reference, both temporally and spatially. Cache VQ takes advantage of such property by providing a large main codebook for quality concerns and a much smaller cache codebook for compression and computational concerns. We evaluate some design parameters for the cache VQ system. Our experimental results show that for various combinations of cache sizes, distortion thresholds, and line sizes, bit ratios ranging from 60% to 97% can be achieved by the cache VQ, with PSNRs between 33 dB and 38 dB and with codevector dimensions of 4 and 16. The corresponding compression ratio without entropy coding ranges from 6:1 to 24:1. These results clearly indicate that a very good performance in terms of time as well as fidelity and bit rate can be realized with the cache VQ.>
Chang Y. Choo, Nasser M. Nasrabadi
ICIP (1)2
1994 Predictive Residual Vector Quantization
abstract
Presents a new vector quantization technique, called predictive residual vector quantization (PRVQ), which combines the concepts of predictive vector quantization (PVQ) and residual vector quantization (RVQ) to implement a high performance VQ scheme with low search complexity. A major task in the PRVQ design is the joint optimization of the vector predictor and the RVQ codebooks. In order to achieve this, a constrained optimization technique is introduced which is compared with a jointly designed technique and a closed loop design technique. Simulation results show the superiority of the proposed PRVQ scheme over the equivalent RVQ, PVQ and an unconstrained VQ scheme. The proposed PRVQ scheme gives the best performance when the predictor and all the stage quantizers are jointly optimized.>
Syed A. Rizvi, Nasser M. Nasrabadi
ICIP (1)2
1994 Constrained Gradient Descent Algorithm for Residual Vector Quantizer Design
abstract
Residual vector quantizers have been proposed to overcome the search complexity of regular single stage vector quantizers. We present a design algorithm for residual vector quantizer codebooks. An attempt is made to make full use of the sequential search nature of the encoding process. An error energy is formulated based on the distortion criterion with a constraint imposed to optimize for the sequential search. The codebook design is formulated as multidimensional minimization problem, where the error energy is minimized to obtain the required codebooks. The proposed algorithm is based on the gradient descent algorithm.>
Mahesh Venkatraman, Nasser M. Nasrabadi
ICIP (1)2
1994 Predictive residual vector quantization
abstract
This paper presents a new vector quantization technique, called predictive residual vector quantization (PRVQ), which combines the concepts of predictive vector quantization (PVQ) and residual vector quantization (RVQ) to implement a high performance VQ scheme with low search complexity. A major task in the PRVQ design is the joint optimization of the vector predictor and the RVQ codebooks. In order to achieve this, a constrained optimization technique is introduced, which is then compared with a jointly designed technique and a closed loop design technique. Simulation results show the superiority of the proposed PRVQ scheme over the equivalent RVQ, PVQ and an unconstrained VQ scheme. The proposed PRVQ scheme gives the best performance when the predictor and all the stage quantizers are jointly optimized.
Syed A. Rizvi, Nasser M. Nasrabadi
ICPR (3)2
1994 Residual vector quantization using a multilayer competitive neural network
abstract
This paper presents a new technique for designing a jointly optimized residual vector quantizer (RVQ). In conventional stage-by-stage design procedure, each stage codebook is optimized for that particular stage distortion and does not consider the distortion from the subsequent stages. However, the overall performance can be improved if each stage codebook is optimized by minimizing the distortion from the subsequent stage quantizers as well as the distortion from the previous stage quantizers. This can only be achieved when stage codebooks are jointly designed for each other. In this paper, the proposed codebook design procedure is based on a multilayer competitive neural network where each layer of this network represents one stage of the RVQ. The weight connecting these layers form the corresponding stage codebooks of the RVQ. The joint design problem of the RVQ's codebooks (weights of the multilayer competitive neural network) is formulated as a nonlinearly constrained optimization task which is based on a Lagrangian error function. This Lagrangian error function includes all the constraints that are imposed by the joint optimization of the codebooks. The proposed procedure seeks a locally optimal solution by iteratively solving the equations for this Lagrangian error function. Simulation results show an improvement in the performance of an RVQ when designed using the proposed joint optimization technique as compared to the stage-by-stage design, where both generalized Lloyd algorithm (GLA) and the Kohonen learning algorithm (KLA) were used to design each stage codebook independently, as well as the conventional joint-optimization technique.>
Syed A. Rizvi, Nasser M. Nasrabadi
IEEE J. Sel. Areas Commun.2
1994 A New Image and Video Compression Technique in an Asynchronous Transfer Mode Network Environment
Peter P. Polit, Nasser M. Nasrabadi
J. Vis. Commun. Image Represent.2
1994 Predictive vector quantizer using constrained optimization
abstract
A joint optimization technique is developed for designing the predictor and quantizer of a predictive vector quantizer (PVQ). The proposed technique is based on a constrained optimization technique that makes use of a Lagrangian formulation and iteratively solves the Lagrangian error function to obtain a locally optimal solution for the predictor and quantizer. Simulation results show that the proposed PVQ design outperforms the conventional PVQ schemes, such as the closed-loop design and the jointly-optimized technique.>
Syed A. Rizvi, Nasser M. Nasrabadi
IEEE Signal Process. Lett.2
1994 Dynamic finite-state vector quantization of digital images
abstract
A vector quantization (VQ) scheme with finite memory called dynamic finite-state vector quantization (DFSVQ) is presented. The encoder consists of a large codebook, so called super-codebook, where for each input vector a fixed number of its codevectors are chosen to generate a much smaller codebook (sub-codebook). This sub-codebook represents the best matching codevectors that could be found in the super-codebook for encoding the current input vector. The choice for the codevectors in the sub-codebook is based on the information obtained from the previously encoded blocks where directional conditional block probability (histogram) matrices are used in the selection of the codevectors. The index of the best matching codevector in the sub-codebook is transmitted to the receiver. An adaptive DFSVQ scheme is also proposed in which, when encoding an input vector, first the sub-codebook is searched for a matching codevector to satisfy a pre-specified waveform distortion. If such a codevector is not found in tile current sub-codebook then the whole super-codebook is checked for a better match. If a better match is found then a signaling flag along with the corresponding index of the codevector is transmitted to the receiver. Both the DFSVQ encoder and its adaptive version are implemented. Experimental results for several monochrome images with a super-codebook size of 256 or 512 and different sub-codebook sizes are presented.>
Nasser M. Nasrabadi, Chang Y. Choo, Yushu Feng
IEEE Trans. Commun.1
1994 Edge-based subband VQ techniques for images and video
abstract
A key issue in subband coding is the efficient compression of the less informative but perceptually important upper frequency bands of the decomposed image. A new approach capable of effectively encoding the upper-bands is described. An intra-band vector quantization (VQ) technique is employed for compression of the base-band while the upper frequency bands are encoded by a hierarchical inter-band VQ method. The proposed inter-band vector quantization scheme exploits the redundancies that exist between pels of significant perceptual importance across the upper frequency bands of the same resolution. Such pels are observed to be at or around the edge-locations, displaying similar discontinuity behavior across the subbands. Therefore, an edge-detector is applied on the reconstructed base-band to extract the edge locations, and as a result no overhead information is transmitted to identify the position of these pels in the upper-bands. Furthermore, a residual subband, being the difference between the original base-band of each layer and its encoded version, is incorporated in the proposed inter-band VQ model. Thus, the proposed interband VQ scheme is simply the vector quantization of pixels across the upper-bands together with their corresponding residual band. Compression results are presented for both digital images and video sequences which demonstrate high subjective visual qualities.>
Nader Mohsenian, Nasser M. Nasrabadi
IEEE Trans. Circuits Syst. Video Technol.2
1993 Predictive vector quantization using a neural network
Nader Mohsenian, Nasser M. Nasrabadi
ICASSP (5)2
1993 Next-state functions for finite-state vector quantization
abstract
In this paper, a finite-state vector quantizer called Dynamic Finite-State Vector Quantization (DFSVQ) is investigated with regard to its subcodebook construction. In DFSVQ each input vector encoded by a small codebook, called subcodebook, is created from a much larger codebook called supercodebook. The subcodebook is constructed by selecting (reordering procedure) a set of appropriate codevectors from the supercodebook. The performance of the DFSVQ depends on this reordering procedure, therefore, several reordering procedures are introduced and their performances are evaluated in this paper. The reordering procedures that are investigated are the conditional histogram, address prediction, vector prediction, nearest neighbor design, and the frequency usage of codevectors. The performance of the reordering procedures are evaluated by comparing their hit ratios (the number of blocks encoded by the subcodebook) and their computational complexity. Experimental results are presented for both still images and video. It is found that for still images the conditional histogram performs the best and for video the nearest neighbor design performs the best.
Nasser M. Nasrabadi, Nader Mohsenian, Hon-Tung Mak, Syed A. Rizvi
VCIP1
1993 Invariant Object Recognition Based on A Neural Network of Cascaded RCE Nets
abstract
A neural network of cascaded Restricted Coulomb Energy (RCE) nets is constructed for the recognition of two-dimensional objects. A number of RCE nets are cascaded together to form a classifier where the overlapping decision regions are progressively resolved by a set of cascaded networks. Similarities among objects which have complex decision boundaries in the feature space are resolved by this multi-net approach. The generalization ability of an RCE net recognition system, referring to the ability of the system to correctly recognize a new pattern even when the number of learning exemplars is small, is increased by the proposed coarse-to-fine learning strategy. A feature extraction technique is used to map the geometrical shape information of an object into an ordered feature vector of fixed length. This feature vector is then used as an input to the neural network. The feature vector is invariant to object changes such as positional shift, rotation, scaling, illumination variance, variation of camera setup, perspective distortion, and noise distortion. Experimental results for recognition of several objects are also presented. A correct recognition rate of 100% was achieved for both the training and the testing input patterns.
Wei Li 0047, Nasser M. Nasrabadi
Int. J. Pattern Recognit. Artif. Intell.2
1992 An interframe dynamic FSVQ codec for video sequence coding
abstract
An interframe vector quantization scheme with finite memory called dynamic finite-state vector quantization (DFSVQ) is presented. The encoding of a current input vector consists of a subcodebook, representing the best matchable code vectors, selected from a larger codebook, i.e., a supercodebook. The choice for the entries in the supercodebook is based on the information obtained from the previously encoded blocks where directional conditional block probability matrices are used in the selection of the entries. An adaptive DFSVQ scheme is also proposed in which, when encoding an input vector, first the supercodebook is searched for a matching code vector to satisfy a prespecified waveform distortion. If such a code vector is not found then the whole supercodebook is checked for a better match and a signaling flag along with the corresponding address of the code vector is transmitted to the receiver. The performances of the DFSVQ encoder and its adaptive version are evaluated for several video sequences.>
Nasser M. Nasrabadi, Nader Mohsenian
ICASSP2
1992 Subband coding of video using an edge-based vector quantization technique for compression of the upper bands
abstract
A new subband video coding technique is introduced which utilizes several memoryless vector quantizing (VQ) schemes for encoding the various layers. A set of quadrature mirror filter banks was applied to the motion compensated frame differences (MCFDs) of the moving sequences to produce seven nonuniform bands. The upper bands displayed a significant amount of information at edge locations, these being the positions where the baseband of each layer was observed to have a similar behavior. Therefore, an edge-detecting operator, e.g., Laplacian or a Gaussian, was incorporated into the video compression model to extract the perceptually important locations of the upper bands from their correspondingly encoded baseband, thus eliminating the need for transmission of their addresses. Promising results were obtained which are suitable for low-bit-rate video applications where videophone and videoconferencing systems may be realized.>
Nader Mohsenian, Nasser M. Nasrabadi
ICASSP2
1992 Interframe Hierarchical Address-Vector Quantization
abstract
A new interframe coding technique called interframe hierarchical address-vector quantization (IHA-VQ) is presented. It exploits the local characteristics of a moving image's motion-compensated (via block matching) difference signals by using quadtree segmentation to divide each image into large, uniform regions and smaller, highly detailed regions. The detailed regions are encoded by vector quantization (VQ), and the larger, low-detail regions are replenished from the previous frame, IHA-VQ also exploits the correlation between the small blocks by encoding the addresses of neighboring vectors by using several codebooks of address code-vectors.>
Nasser M. Nasrabadi, Chang Y. Choo, Jennifer U. Roy
IEEE J. Sel. Areas Commun.1
1992 A Stereo Vision Technique Using Curve-Segments and Relaxation Matching
abstract
A multichannel feature-based stereo vision technique where curve segments are used as feature primitives in the matching process is described. The left image and the right image are filtered by using several Laplacian-of-Gaussian operators of different widths (channels). Curve segments are extracted by a tracking algorithm, and their centroids are obtained. At each channel, the generalized Hough transform of each curve segment in the left and the right image is evaluated. The epipolar constraint on the centroids of the curve segment and the channel size is used to limit the searching space in the right image. To resolve the ambiguity of the false targets (multiple matches), a relaxation technique is used where the initial scores of the node assignments are updated by the compatibility measures between the centroids of the curve segments. The node assignments with the highest score are chosen as the matching curve segments.>
Nasser M. Nasrabadi
IEEE Trans. Pattern Anal. Mach. Intell.1
1992 Hopfield network for stereo vision correspondence
abstract
An optimization approach is used to solve the correspondence problem for a set of features extracted from a pair of stereo images. A cost function is defined to represent the constraints on the solution, which is then mapped onto a two-dimensional Hopfield neural network for minimization. Each neuron in the network represents a possible match between a feature in the left image and one in the right image. Correspondence is achieved by initializing (exciting) each neuron that represents a possible match and then allowing the network to settle down into a stable state. The network uses the initial inputs and the compatibility measures between the matched points to find a stable state.
Nasser M. Nasrabadi, Chang Y. Choo
IEEE Trans. Neural Networks1
1991 A New Transform Domain Vector Quantization Technique for Image Data Compression in an Asynchronous Transfer Mode Network
abstract
This coding technique combines vector quantization with a discrete cosine transform. It requires little or no priority layer, is computationally inexpensive to implement, and is robust to packet loss.>
Peter P. Polit, Nasser M. Nasrabadi
Data Compression Conference2
1991 A non-linear predictor for differential pulse-code encoder (DPCM) using artificial neural networks
abstract
A nonlinear predictor is designed for a DPCM encoder using artificial neural networks (ANN). The predictor is based on a multilayer perceptron with three input nodes, 30 hidden nodes and one output node. The back-propagation learning algorithm is used for the training of the network. Simulation results are presented to evaluate and compare the performance of the neural net based predictor (nonlinear) with that of an optimized linear predictor. Success in the use of the nonlinear predictor is demonstrated through the reduction in the entropy of the differential error signal as compared to that of a linear predictor. Also it is shown that the ANN predictor is much more robust for encoding noisy images compared to that of a linear predictor.>
Sohail A. Dianat, Nasser M. Nasrabadi, S. Venkataraman
ICASSP2
1991 An efficient terrain acquisition algorithm for a mobile robot
abstract
A terrain acquisition algorithm for an autonomous mobile robot to learn and model its surrounding terrain efficiently with respect to travel distance, and sensing operations is presented. It is shown that this algorithm enables a mobile robot to visit relatively few obstacle vertices to construct a map of a planar terrain occupied by polygonal obstacles. According to this algorithm, a mobile robot travels to a vertex of the obstacles only when additional terrain information can be obtained from there. The decision to move to a vertex is thus based on whether any incident edge of a vertex is missing or occluded.>
Chang Y. Choo, John M. Smith, Nasser M. Nasrabadi
ICRA3
1991 A self-organizing adaptive vector quantization technique
Yushu Feng, Nasser M. Nasrabadi, Chang Y. Choo
J. Vis. Commun. Image Represent.2
1991 Object recognition by a Hopfield neural network
abstract
A two-dimensional model-based object recognition technique is introduced to identify and locate isolated or overlapping 2-D objects in any position and orientation. A cooperative feature-matching technique is proposed that is implemented by a Hopfield neural network. The proposed matching technique uses the parallelism of the neural network to globally match all the objects in the input scene against all the object models in the model-database at the same time. A global model graph representing all the object models is constructed where each node in the graph represents a feature that has a numerical feature value and is connected to other nodes by an arc representing the relationship or compatibility between them. Object recognition is formulated as matching this global model with an input scene graph representing a single object or several overlapping objects. The performance of the proposed technique is compared with that of a relaxation technique.>
Nasser M. Nasrabadi, Wei Li 0047
IEEE Trans. Syst. Man Cybern.1
1990 A dynamic finite-state vector quantization scheme
abstract
A vector quantization (VQ) scheme with memory called dynamic finite-state vector quantization (DFSVQ) is described. In the proposed DFSVQ system, a super-codebook is designed using the generalized Lloyd algorithm. During encoding, an input vector subcodebook is dynamically generated by reordering the positions of the codevectors in the super-codebook. The codevectors that are the most probable approximations of the input vector are therefore moved to the top of the super-codebook. The first N codevectors of the super-codebook form the subcodebook for that input vector. An adaptive DFSVQ scheme is proposed in which the subcodebook is searched for a matching codevector satisfying a prespecified waveform distortion when encoding an input vector. If such a codevector is not found in the current subcodebook, the whole super-codebook is checked for the best match and a signaling flag along with the corresponding address of the best-matching codevector is transmitted to the receiver.>
Nasser M. Nasrabadi, Yushu Feng
ICASSP1
1990 Object recognition by a Hopfield neural network
abstract
A model-based recognition method is introduced which is formulated as an optimization problem. An energy function is derived which represents the constraints on the best solution in order to find the best match. A two-dimensional binary Hopfield neural network is implemented to minimize the energy function. The state of each neuron in the Hopfield network represents the possibility of a match between a node in the model graph and a node in the scene graph. >
Nasser M. Nasrabadi, Wei Li 0047, Chang Y. Choo
ICCV1
1990 Invariant object recognition based on a neural network of cascaded RCE nets
abstract
A neural network of cascaded restricted Coulomb energy (RCE) networks is constructed for object recognition. A number of RCE networks are cascaded together to form a classifier where the overlapping decision regions in a previously learned network are solved by the next network. The similarities among objects which have complex decision boundaries in the feature space are resolved by this multinetworks approach. The generalization ability of a RCE network recognition system, referring to the ability of the system to correctly recognize a new pattern even when the number of learning exemplars is small, is increased by the proposed coarse-to-fine learning strategy. A new feature extraction technique is proposed for mapping the geometrical shape information of an object into an ordered feature vector of fixed length which is the required form for input to this neural network
Wei Li 0047, Nasser M. Nasrabadi
IJCNN2
1990 Interframe hierarchical address-vector quantization
abstract
A new interframe coding technique is presented in this paper called Interframe Hierarchical Address- Vector Quantization (IHA-VQ). It exploits the local characteristics of a moving image's motion compensated (via block matching) difference signals by using quadtree segmentation to divide each signal into large, uniform regions and smaller, highly detailed regions. The detailed regions are encoded by Vector Quantization (VQ), and the larger, low detail regions are replenished from the previous image. IHA-VQ also exploits the correlation between the small blocks by encoding the addresses of neighboring vectors using several codebooks of address-codevectors. The IHA-VQ technique is applied to a test sequence of images in a computer simulation. The sequence was encoded at an average bit rate of 0.607 bits per pixel (bpp). The Signal-to- Noise Ratio (SNR) averaged 39.5 dB. By changing the quadtree segmentation parameter a much lower bit rate, 0.237 bpp, is achieved for a lower SNR, 37.73 dB.
Nasser M. Nasrabadi
VCIP1
1990 Image compression using address-vector quantization
abstract
A novel vector quantization scheme, called the address-vector quantizer (A-VQ), is proposed. It is based on exploiting the interblock correlation by encoding a group of blocks together using an address-codebook. The address-codebook consists of a set of address-codevectors where each codevector represents a combination of addresses (indexes). Each element of this codevector is an address of an entry in the LBG-codebook, representing a vector quantized block. The address-codebook consists of two regions: one is the active (addressable) region, and the other is the inactive (nonaddressable) region. During the encoding process the codevectors in the address-codebook are reordered adaptively in order to bring the most probable address-codevectors into the active region. When encoding an address-codevector, the active region of the address-codebook is checked, and if such an address combination exist its index is transmitted to the receiver. Otherwise, the address of each block is transmitted individually. The quality (SNR value) of the images encoded by the proposed A-VQ method is the same as that of a memoryless vector quantizer, but the bit rate would be reduced by a factor of approximately two when compared to a memoryless vector quantizer.>
Nasser M. Nasrabadi, Yushu Feng
IEEE Trans. Commun.1
1989 A dynamic address-vector quantization algorithm based on inter-block and inter-color correlation for color image coding
abstract
The authors extend the A-VQ (address-vector quantization) coding technique to encode RGB (red, green, blue) color images by utilizing the interblock and intercolor correlation. A modified A-VQ technique called dynamic address-vector quantization (DA-VQ) is introduced for encoding of color images. A multilayered address vector quantizer is also introduced, where the code vectors at each layer represent the most probable address combinations of the entries of the address codebook at the previous lower layer. Experimental results show reconstructed images at bit rates of 0.5-0.6 bit per pixel with SNR (signal/noise ratio)=28-31 dB. The resulting bit rate is less than that of a standard VQ coding technique by at least a factor of two.>
Yushu Feng, Nasser M. Nasrabadi
ICASSP2
1989 Interframe hierarchical vector quantization
abstract
An interframe coding scheme for the transmission of image sequences at bit rates below 0.2 bits per pixel is presented. The proposed scheme is called an interframe hierarchical (quadtree) vector quantizer. A regular decomposition quadtree method is used to segment the interframe differential signal into homogeneous regions of different block size. Small blocks representing high-contrast moving boundaries known as the impulsive component of the difference signal are vector-quantized. Large blocks typically representing smooth regions of the image are encoded by the local sample mean of the region. Experimental results are presented, confirming that the system can encode image sequence scenes at bit rates below 0.2 bit per pixel per frame. A comparison of the interframe hierarchical vector quantization system and an interframe mean reconstructed quadtree method is presented.>
Nasser M. Nasrabadi, Shihkuan E. Lin, Yushu Feng
ICASSP1
1989 Use of Hopfield network for stereo vision correspondence
abstract
An optimization approach is used to solve the correspondence problem for a set of features extracted from a pair of stereo images. A cost function is defined to represent the constraints on the solution which is then mapped onto a 2-D neural network for minimization. Each neuron in the network represents a possible match between a feature in the left image and one in the right image. Correspondence is achieved by initializing all the neurons that represent the possible matches and allowing the network to use the compatibility measures between the matched points to settle down into a stable state. >
Nasser M. Nasrabadi, Wei Li 0047, Bradley G. Epranian, Charles A. Butkus
SMC1
1989 Stereo vision correspondence using a multichannel graph matching technique
Nasser M. Nasrabadi
Image Vis. Comput.1
1988 A stereo vision technique using curve-segments and relaxation matching
abstract
A multichannel feature-based stereo vision technique is described in which curve segments are used as the feature primitives in the matching process. The left and right images are first filtered by using several Laplacian of Gaussian operators of different widths (channel). Curve segments are extracted by a tracking algorithm, and their centroids are obtained. At each channel, the generalized Hough transform of each curve-segment in images is evaluated. The R-table is used as a local feature vector in representing the distinctive characteristics of a segment. The epipolar constraint on the centroids of the curve segment and the channel size is used to limit the search space in the right image. To resolve the ambiguity of the false targets (multiple matches), a relaxation technique is used in which the initial scores of the node assignments are updated by the compatibility measures between the centroids of the curve segments. The node assignments with the highest score are chosen as the matching segments.>
Nasser M. Nasrabadi, Jen-Lai Chiang
ICPR1
1988 Stereo vision correspondence using a multichannel graph matching technique
abstract
A multichannel feature-based stereo vision technique is described wherein curve segments are used as the feature primitives in the matching process. Curve segments are extracted by tracking the zero-crossings of the left and right images. The generalized Hough transform of each curve and the length of the segment in the left image are used as a local feature vector to represent the distinctive characteristics of the segment. The feature vector of each segment is used as a constraint to find an instance of the same segment in the right image. The epipolar constraint on the centroids of the curve segment is used to limit the searching space in the right image. A relational graph is formed from the left image by treating the centroids as the nodes of the graph. The local features of the segments are used to represent the local properties of the nodes, and the relationship between the nodes represents the structural properties of the object in the scene. A similar graph is formed from the right image curve segments. A graph isomorphism is then formed between the two graphs.>
Nasser M. Nasrabadi, Jen-hi Chiang
ICRA1
1988 Integration of stereo-vision and optical flow using markov random fields
Sandra P. Clifford, Nasser M. Nasrabadi
Neural Networks2
1988 Vector quantization of images based upon the kohonen self-organization feature maps
Nasser M. Nasrabadi, Yushu Feng
Neural Networks1
1988 Image coding using vector quantization: a review
abstract
A review of vector quantization techniques used for encoding digital images is presented. First, the concept of vector quantization is introduced, then its application to digital images is explained. Spatial, predictive, transform, hybrid, binary, and subband vector quantizers are reviewed. The emphasis is on the usefulness of the vector quantization when it is combined with conventional image coding techniques, or when it is used in different domains.>
Nasser M. Nasrabadi, Robert A. King
IEEE Trans. Commun.1
1985 Use of vector quantizers in image coding
abstract
This paper presents a review of the vector quantization techniques in image coding. The application of vector quantization to code images in the spatial and the frequency domain are discussed, residual vector quantizers, predictive vector quantizers, sub-band vector quantizer as well as binary vector quantizers are also mentioned. Finally, this paper is believed to give a short review of the current vector quantization techniques.
Nasser M. Nasrabadi
ICASSP1
1984 Complex number theoretic transform in p-adic field
abstract
Transforms in the field of p-adic numbers and its extensions such as quadratic extensions are introduced. It is shown that rational integers can be represented exactly and error free convolution can be done with a larger dynamic range compared to number theoretic transforms.
Nasser M. Nasrabadi, Robert A. King
ICASSP1
1984 A new image coding technique using transforms vector quantization
abstract
A new interframe coding technique is proposed. Where a two-dimensional Hadamard transform is applied on each sub-block of successive frames, and an adaptive vector quantisation scheme is applied along the transformed blocks of the successive frames. The performance of the algorithm is evaluated by computer simulation on sequence of moving images.
Nasser M. Nasrabadi, Robert A. King
ICASSP1
1983 Image coding using vector quantization in the transform domain
Robert A. King, Nasser M. Nasrabadi
Pattern Recognit. Lett.2