Vishal Monga

dblp:73/1467 · DBLP profile ↗
← Back
94ranked-venue papers
11as first author
18since 2021 · last 2025
0000-0002-5100-2263ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 67 · 9 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 10 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Security and privacy · 4 · 2 first-author
YearPublicationVenuePosition
2025 Regularized Weighted Descent: Model-Based Learner for Multi-Target Radar Waveform Design
abstract
This study focuses on multiple target detection in the presence of signal-dependent clutter using a Multiple-Input Multiple-Output (MIMO) radar system. The problem is formulated as worst-case SINR maximization (max-min optimization), which is a function of the MIMO waveform, under the hardware-inspired constant modulus constraint (CMC). While existing approaches invariably rely on computationally expensive iterative optimization over the waveform variable, we develop a model-based deep learning algorithm that shifts the computational burden to the neural network training, yielding fast inference. We utilize a surrogate cost function – the sum of SINR-Reciprocals (SRs) – that enables converting the max-min problem into the minimization of the sum of SRs. Our model-based learner unrolls an iterative optimization method that utilizes the SR descent vectors but with novel inter-target and inter-step parameters. The inter-target parameter weighs the SR descent vectors so that the net descent direction is dominated by the vector associated with the largest SR, thereby focusing on the worst-case SINR. The inter-step parameter ensures the update between the steps encourages a monotonic decrease in the cost function. To effectively guide the parameter learning, we introduce regularizers aligned with the learning goals, and consequently, we term the proposed method Regularized Weighted Descent (RWD). We demonstrate that the RWD achieves a larger worst-case SINR value (superior solution quality) in a shorter time (lower computational complexity) compared to the state-of-the-art alternatives.
Junho Kweon, Fulvio Gini, Maria Greco 0001, Muralidhar Rangaswamy, Vishal Monga
ICASSP5
2025 CLIF-Net: Intersection-Guided Cross-View Fusion Network for Infection Detection From Cranial Ultrasound
abstract
This paper addresses the problem of detecting possible serious bacterial infection (pSBI) of infancy, i.e. a clinical presentation consistent with bacterial sepsis in newborn infants using cranial ultrasound (cUS) images. The captured image set for each patient enables multi-view imagery: coronal and sagittal, with geometric overlap. To exploit this geometric relation, we develop a new learning framework, called the intersection-guided Cross-view Local- and Image-level Fusion Network (CLIF-Net). Our technique employs two distinct convolutional neural network branches to extract features from coronal and sagittal images with newly developed multi-level fusion blocks. Specifically, we leverage the spatial position of these images to locate the intersecting region. We then identify and enhance the semantic features from this region across multiple levels using cross-attention modules, facilitating the acquisition of mutually beneficial and more representative features from both views. The final enhanced features from the two views are then integrated and projected through the image-level fusion layer, outputting pSBI and non-pSBI class probabilities. We contend that our method of exploiting multi-view cUS images enables a first of its kind, robust 3D representation tailored for pSBI detection. When evaluated on a dataset of 302 cUS scans from Mbale Regional Referral Hospital in Uganda, CLIF-Net demonstrates substantially enhanced performance, surpassing the prevailing state-of-the-art infection detection techniques.
Mingzhao Yu, Mallory R. Peterson, Kathy Burgoine, Thaddeus Harbaugh, Peter Olupot-Olupot, Melissa Gladstone, Cornelia F. Hagmann, Frances M. Cowan, Andrew Weeks, Sarah U. Morton, Ronald Mulondo, Edith Mbabazi Kabachelor, Steven J. Schiff, Vishal Monga
IEEE Trans. Medical Imaging14
2024 Fast and Physically Enriched Deep Network for Joint Low-Light Enhancement and Image Deblurring
abstract
Joint low-light enhancement and deblurring is a challenging imaging inverse problem that estimates clean images from photography corrupted by both low-light and blurring artifacts. To address this task, we propose FELI, a Fast and physically Enriched deep neural network for joint Low-light enhancement and Image deblurring. In a departure from recently proposed end-to-end networks, FELI employs a learnable Decomposer during training based on Retinex theory that helps with low-light scene recovery. FELI’s encoded features are further enriched by an input reconstruction task cognizant of the blur model leading to effective deblurring. We introduce a new customized contrastive regularization (CCR) term that pulls the restored clean image closer to the ground truth while pushing it far away from both the input and reconstructed input. Experiments performed on challenging synthetic and real-world datasets demonstrate that FELI outperforms state-of-the-art methods at a lower computational cost.
Trung Hoang, Jon S. McElvain, Vishal Monga
ICASSP3
2024 Progressive Diffusion Autofocus for Synthetic Aperture Sonar Imagery
abstract
Autofocus algorithms for synthetic aperture sonar (SAS) enhance the autonomy of uncrewed underwater vehicles by counteracting environmental factors that often compromise image quality, thus preserving the vehicles’ capacity for accurate perception. Often, these errors are systematic resulting from misestimation of environmental parameters (e.g., sound speed) or sonar position. Traditional autofocusing techniques aimed at correcting these distortions may result in degrading well-focused imagery when applied repeatedly, while single-attempt focusing methods may not fully correct the errors. In this work, we introduce a novel SAS autofocus approach named Progressive Diffusion Autofocus (PDA), which incrementally refines the focus over several iterations, drawing inspiration from recent deep-learning-centric diffusion processes. Our method iteratively improves focus with each iteration, thereby enhancing the image progressively. We evaluate our approach against established methods, such as optimization-based sharpness metrics and a recently introduced deep learning method, Deep Adaptive Phase Learning (DAPL). Our results indicate that our PDA method not only consistently improves image quality, but also minimizes the degradation of images that are already well-focused.
Balkan V. Bingol, Isaac Gerg, Vishal Monga
IGARSS3
2023 Maturity-Aware Active Learning for Semantic Segmentation with Hierarchically-Adaptive Sample Assessment
Amirsaeed Yazdani, Xuelu Li, Vishal Monga
BMVC3
2023 Interpretable, Unrolled Deep Radar Beampattern Design
abstract
Optimizing a transmit MIMO radar waveform subject to the non-convex constant modulus constraint remains a problem of enduring interest. The past decade has seen a variety of tailored iterative approaches with various performance-complexity trade-offs. Despite promising work, iterative algorithms have a speed handicap and require meticulous parameter tuning. Once trained, a deep network can quickly regress the desired waveform coefficients, but it is a black box and may excel only when generous training is available. We present a fast, learned, and - for the first time - interpretable (FLI) deep learning approach by unrolling a state-of-the-art iterative optimization approach. We particularly leverage the recently proposed projection, descent, and retraction (PDR) algorithm and design a deep network where each PDR step is mapped to a layer in the neural network while preserving the non-convex constant modulus constraint. FLI breaks the trade-off between complexity and performance. It is near real-time with boosted performance – fidelity to the desired beampattern – compared to the state-of-the-art alternatives.
Kareem M. Metwaly, Junho Kweon, Khaled Alhujaili, Maria Greco 0001, Fulvio Gini, Vishal Monga
ICASSP6
2023 A Convergent Neural Network for Non-Blind Image Deblurring
abstract
In recent years, algorithm unrolling has emerged as a powerful technique for designing interpretable neural networks based on iterative algorithms. Imaging inverse problems have particularly benefited from unrolling based deep network design since many traditional model-based approaches rely on iterative optimization. Despite exciting progress, typical unrolling approaches heuristically design layer-specific convolution weights to improve performance. Crucially, convergence properties of the underlying iterative algorithm are lost once layer specific parameters are learned from training data. In this paper, we propose a neural network architecture that breaks the trade-off between retaining algorithm properties while simultaneously enhancing performance. We focus on non-blind image deblurring problem and unroll the widely-applied Half-Quadratic Splitting (HQS) algorithm. We develop a new parameterization scheme that enforces the layer-specific parameters to asymptotically approach certain fixed points, a new result that we analytically establish. Experimental results show that our approach outperforms many state of the art non-blind deblurring techniques on benchmark datasets, while enabling convergence and interpretability.
Yanan Zhao 0003, Haichuan Zhang 0001, Vishal Monga, Yonina C. Eldar
ICIP4
2023 Deep Multi-Look Sequence Processing for Synthetic Aperture Sonar Image Segmentation
abstract
Deep learning has enabled significant improvements in semantic image segmentation, especially in underwater imaging domains such as side scan sonar (SSS). In this work, we apply deep learning to synthetic aperture sonar (SAS) imagery, which has an advantage over traditional SSS in which SAS produces coherent high- and constant-resolution imagery. Despite the successes of deep learning, one drawback is the need for abundant labeled training data to enable success. Such abundant labeled data are not always available as in the case of SAS where collections are expensive and obtaining quality ground-truth labels may require diver intervention. To overcome these challenges, we propose a domain-specific deep learning network architecture utilizing a unique property to complex-valued SAS imagery: the ability to resolve angle-of-arrival (AoA) of acoustic returns through$k$-space processing. By sweeping through consecutive incrementally advanced AoA bandpass filters (a process sometimes referred to as multi-look processing), this technique generates a sequence of images emphasizing angle-dependent seafloor scattering and motion from biologics along the seafloor or in the water column. Our proposal, which we call multi-look sequence processing network (MLSP-Net), is a domain-enriched deep neural network architecture that models the multi-look image sequence using a recurrent neural network (RNN) to extract robust features suitable for semantic segmentation of the seafloor without the need for abundant training data. Unlike previous segmentation works in SAS, our model ingests a complex-valued SAS image and affords the ability to learn the AoA filters in$k$-space as part of the training procedure. We show the results on a challenging real-world SAS database, and despite the lack of abundant training data, our proposed method shows superior results over state-of-the-art techniques.
Isaac Gerg, Vishal Monga
IEEE Trans. Geosci. Remote. Sens.2
2023 Resonant Scattering-Inspired Deep Networks for Munition Detection in 3D Sonar Imagery
abstract
Underwater sites affected by unexploded ordnance (UXO) pose significant risk to both human safety and environmental well-being. Sonar imaging is commonly employed to investigate such sites and aid in UXO remediation efforts. However, manually identifying and classifying potential targets in sonar data is challenging and time-consuming. Previous research has explored the use of machine learning models to recognize and categorize targets; however, many of these approaches lack transparency and fail to consider the underlying physical acoustics. Additionally, acquiring sufficient training data for these models can often be problematic. In this study, we present a novel approach by designing neural networks that explicitly account for the unique physics involved in the problem domain. UXOs examined using low-frequency sound frequently exhibit resonant behavior, where the sound is re-radiated after initial geometric scattering, owing to the elastic and compressional properties of the objects. Moreover, such resonant effects are typically absent in clutter objects, making them advantageous in discriminating UXOs from non-UXOs. Consequently, we propose several neural network architectures that leverage these resonant effects, utilizing 3D data obtained from a synthetic aperture sonar (SAS) imaging sonar. Our first proposal incorporates a recurrent neural network to model the physics-based correlation among adjacent time/spatial slices, originating from the resonant phenomena. For our second proposal, we employ intensity imagery of orthogonal projections of the 3D data cube, which capture shape-specific resonant scattering mechanisms unique to specific types of UXOs. To evaluate the effectiveness of our methods, we compare them against recent state-of-the-art algorithms using a real-world 3D SAS dataset. Remarkably, even when confronted with limited training data, our approaches consistently demonstrate superior results. Our findings highlight the significant potential of incorporating physical acoustics into neural network designs for UXO detection and classification, offering improved accuracy and efficiency in underwater remediation operations.
Trung Hoang, Kyle S. Dalton, Isaac Gerg, Thomas E. Blanford, Daniel C. Brown, Vishal Monga
IEEE Trans. Geosci. Remote. Sens.6
2022 GlideNet: Global, Local and Intrinsic based Dense Embedding NETwork for Multi-category Attributes Prediction
abstract
Attaching attributes (such as color, shape, state, action) to object categories is an important computer vision problem. Attribute prediction has seen exciting recent progress and is often formulated as a multi-label classification problem. Yet significant challenges remain in: 1) predicting a large number of attributes over multiple object categories, 2) modeling category-dependence of attributes, 3) methodically capturing both global and local scene context, and 4) robustly predicting attributes of objects with low pixel-count. To address these issues, we propose a novel multi-category attribute prediction deep architecture named GlideNet, which contains three distinct feature extractors. A global feature extractor recognizes what objects are present in a scene, whereas a local one focuses on the area surrounding the object of interest. Meanwhile, an intrinsic feature extractor uses an extension of standard convolution dubbed Informed Convolution to retrieve features of objects with low pixel-count utilizing its binary mask. GlideNet then uses gating mechanisms with binary masks and its self-learned category embedding to combine the dense embeddings. Collectively, the Global-Local-Intrinsic blocks comprehend the scene's global context while attending to the characteristics of the local object of interest. The architecture adapts the feature composition based on the category via category embedding. Finally, using the combined features, an interpreter predicts the attributes, and the length of the output is determined by the category, thereby removing unnecessary attributes. GlideNet can achieve compelling results on two recent and challenging datasets - VAW and CAR -for large-scale attribute prediction. For instance, it obtains more than 5% gain over state of the art in the mean recall (mR) metric. GlideNet's advantages are especially apparent when predicting attributes of objects with low pixel counts as well as attributes that demand global context understanding. Finally, we show that GlideNet excels in training starved real-world scenarios.
Kareem M. Metwaly, Aerin Kim, Elliot Branson, Vishal Monga
CVPR4
2022 Structural Prior Models for 3-D Deep Vessel Segmentation
abstract
We address the problem of 3-D blood vessel segmentation with a deep learning method that incorporates domain information via priors and regularizers on vessel structure and morphology. Inspired by the observation that 3-D vessel structures project onto 2-D image slices with distinctive edges that can aid 3-D vessel segmentation, we propose a novel multi-task learning architecture comprising a shared encoder and two decoders that respectively predict vessel segmentation maps and edge profiles. 3-D features from the two branches are concatenated to facilitate edge-guidance when learning segmentation maps. We introduce new regularization terms that encourage local homogeneity of 3-D blood vessel volumes brought about by biomarkers, as well as sparsity of edge pixels. Experiments on benchmark datasets demonstrate superior performance of our method over the state-of-the-art, especially when training data is limited.
Xuelu Li, Raja Bala, Vishal Monga
ICASSP3
2022 Synthetic Aperture Sonar Image Segmentation Using Adaptive, Learned Beam Steering
abstract
Synthetic aperture sonar (SAS) produces high res-olution imagery of the seafloor. Automatic segmentation of this imagery is of interest to a variety of ocean-studying communities. In this work, we propose a novel supervised image segmentation method using adaptive, learned beam steering in the k-space domain in order to exploit the aspect dependent information available in the single-look-complex (SLC) SAS image which is overlooked in current SAS segmentation methods. Our proposal is a deep neural network architecture designed around the domain-specific beam steering geometry conveyed by the k-space domain. Results on a small real-world SAS dataset demonstrate that our proposal gives improved results for a variety of seafloor classes over a state-of-the-art (SOTA) U-Net.
Isaac Gerg, Vishal Monga
IGARSS2
2022 Domain Enriched Deep Networks for Munition Detection in Underwater 3D Sonar Imagery
abstract
Underwater sites impacted by unexploded ordnance (UXO) may pose an unacceptable risk to human and environmen-tal health. Sonar imaging is commonly used to interrogate such sites during UXO remediation, however manually identifying and classifying potential targets is difficult and time intensive. Previous work has explored training machine learning models to recognize and classify targets, however many of these “black-box” approaches fail to model the underlying physical acoustics and require abundant training data which is often hard to obtain. Specifically, UXOs interrogated with low frequency sound often exhibit resonant behavior which re-radiates the sound after the initial scattering due to elastic and compressional properties of the object. Such effects are usually not present in clutter objects, making them advantageous in discriminating UXO from non-UXO. In this work, we propose two neural networks which specifically model resonant scattering effects in order to find UXO from a 3D synthetic aperture sonar (SAS) imaging sonar. We do this by utilizing sequence models which are efficient at modeling the spatially correlated nature of the resonant scattering features. We compare our proposal to two recent state-of-the-art algorithms on a real-world 3D SAS dataset and show superior results even when limited training data is available.
Trung Hoang, Kyle S. Dalton, Isaac Gerg, Thomas E. Blanford, Daniel C. Brown, Vishal Monga
IGARSS6
2022 Structural Prior Driven Regularized Deep Learning for Sonar Image Classification
abstract
Deep learning has been recently shown to improve performance in the domain of synthetic aperture sonar (SAS) image classification. Given the constant resolution with a range of SAS, it is no surprise that deep learning techniques perform so well. Despite deep learning’s recent success, there are still compelling open challenges in reducing the high false alarm rate and enabling success when training imagery is limited, which is a practical challenge that distinguishes the SAS classification problem from standard image classification set-ups where training imagery may be abundant. We address these challenges by exploiting prior knowledge that humans use to grasp the scene. These include unconscious elimination of the image speckle and localization of objects in the scene. We introduce a new deep learning architecture that incorporates these priors with the goal of improving automatic target recognition (ATR) from SAS imagery. Our proposal—called SPDRDL, structural prior driven regularized deep learning—incorporates the previously mentioned priors in a multitask convolutional neural network (CNN) and requires no additional training data when compared to traditional SAS ATR methods. Two structural priors are enforced via regularization terms in the learning of the network: 1) structural similarity prior—enhanced imagery (often through despeckling) aids human interpretation and is semantically similar to the original imagery and 2) structural scene context priors—learned features ideally encapsulate target centering information; hence learning may be enhanced via a regularization that encourages fidelity against known ground truth target shifts (relative target position from scene center). Experiments on a challenging real-world data set reveal that SPDRDL outperforms state-of-the-art deep learning and other competing methods for SAS image classification.
Isaac Gerg, Vishal Monga
IEEE Trans. Geosci. Remote. Sens.2
2022 Iterative, Deep Synthetic Aperture Sonar Image Segmentation
abstract
Synthetic aperture sonar (SAS) systems produce high-resolution images of the seabed environment. Moreover, deep learning has demonstrated superior ability in finding robust features for automating imagery analysis. However, the success of deep learning is conditioned on having lots of labeled training data but obtaining generous pixel-level annotations of SAS imagery is often practically infeasible. This challenge has thus far limited the adoption of deep learning methods for SAS segmentation. Algorithms exist to segment SAS imagery in an unsupervised manner, but they lack the benefit of state-of-the-art learning methods and the results present significant room for improvement. In view of the above, we propose a new iterative algorithm for unsupervised SAS image segmentation combining superpixel formation, deep learning, and traditional clustering methods. We call our method iterative deep unsupervised segmentation (IDUS). IDUS is an unsupervised learning framework that can be divided into four main steps: 1) a deep network estimates class assignments; 2) low-level image features from the deep network are clustered into superpixels; 3) superpixels are clustered into class assignments (which we call pseudo-labels) using$k$-means; and 4) resulting pseudo-labels are used for loss backpropagation of the deep network prediction. These four steps are performed iteratively until convergence. A comparison of IDUS to current state-of-the-art methods on a realistic benchmark dataset for SAS image segmentation demonstrates the benefits of our proposal even as the IDUS incurs a much lower computational burden during inference (actual labeling of a test image). Because our design combines merits of classical superpixel methods with deep learning, practically we demonstrate a very significant benefit in terms of reduced selection bias, i.e., IDUS shows markedly improved robustness against the choice of training images. Finally, we also develop a semi-supervised (SS) extension of IDUS called Iterative Deep SS Segmentation (IDSS) and demonstrate experimentally that it can further enhance performance while outperforming supervised alternatives that exploit the same labeled training imagery.
Yung-Chen Sun, Isaac Gerg, Vishal Monga
IEEE Trans. Geosci. Remote. Sens.3
2022 Robust Deep 3D Blood Vessel Segmentation Using Structural Priors
abstract
Deep learning has enabled significant improvements in the accuracy of 3D blood vessel segmentation. Open challenges remain in scenarios where labeled 3D segmentation maps for training are severely limited, as is often the case in practice, and in ensuring robustness to noise. Inspired by the observation that 3D vessel structures project onto 2D image slices with informative and unique edge profiles, we propose a novel deep 3D vessel segmentation network guided by edge profiles. Our network architecture comprises a shared encoder and two decoders that learn segmentation maps and edge profiles jointly. 3D context is mined in both the segmentation and edge prediction branches by employing bidirectional convolutional long-short term memory (BCLSTM) modules. 3D features from the two branches are concatenated to facilitate learning of the segmentation map. As a key contribution, we introduce new regularization terms that: a) capture the local homogeneity of 3D blood vessel volumes in the presence of biomarkers; and b) ensure performance robustness to domain-specific noise by suppressing false positive responses. Experiments on benchmark datasets with ground truth labels reveal that the proposed approach outperforms state-of-the-art techniques on standard measures such as DICE overlap and mean Intersection-over-Union. The performance gains of our method are even more pronounced when training is limited. Furthermore, the computational cost of our network inference is among the lowest compared with state-of-the-art.
Xuelu Li, Raja Bala, Vishal Monga
IEEE Trans. Image Process.3
2021 Real-Time, Deep Synthetic Aperture Sonar (SAS) Autofocus
abstract
Synthetic aperture sonar (SAS) requires precise time-of-flight measurements of the transmitted/received waveform to produce well-focused imagery. It is not uncommon for errors in these measurements to be present resulting in image defocusing. To overcome this, an autofocus algorithm is employed as a post-processing step after image reconstruction to improve image focus. A particular class of these algorithms can be framed as a sharpness/contrast metric-based optimization. To improve convergence, a hand-crafted weighting function to remove “bad” areas of the image is sometimes applied to the image-under-test before the optimization procedure. Additionally, dozens of iterations are necessary for convergence which is a large compute burden for low size, weight, and power (SWaP) systems. We propose a deep learning technique to overcome these limitations and implicitly learn the weighting function in a data-driven manner. Our proposed method, which we call Deep Autofocus, uses features from the single-look-complex (SLC) to estimate the phase correction which is applied in k-space. Furthermore, we train our algorithm on batches of training imagery so that during deployment, only a single iteration of our method is sufficient to autofocus. We show results demonstrating the robustness of our technique by comparing our results to four commonly used image sharpness metrics. Our results demonstrate Deep Autofocus can produce imagery perceptually better than common iterative techniques but at a lower computational cost. We conclude that Deep Autofocus can provide a more favorable cost-quality tradeoff than alternatives with significant potential of future research.
Isaac Gerg, Vishal Monga
IGARSS2
2021 Simultaneous Denoising and Localization Network for Photoacoustic Target Localization
abstract
A significant research problem of recent interest is the localization of targets like vessels, surgical needles, and tumors in photoacoustic (PA) images.To achieve accurate localization, a high photoacoustic signal-to-noise ratio (SNR) is required. However, this is not guaranteed for deep targets, as optical scattering causes an exponential decay in optical fluence with respect to tissue depth. To address this, we develop a novel deep learning method designed to explicitly exhibit robustness to noise present in photoacoustic radio-frequency (RF) data. More precisely, we describe and evaluate a deep neural network architecture consisting of a shared encoder and two parallel decoders. One decoder extracts the target coordinates from the input RF data while the other boosts the SNR and estimates clean RF data. The joint optimization of the shared encoder and dual decoders lends significant noise robustness to the features extracted by the encoder, which in turn enables the network to contain detailed information about deep targets that may be obscured by noise. Additional custom layers and newly proposed regularizers in the training loss function (designed based on observed RF data signal and noise behavior) serve to increase the SNR in the cleaned RF output and improve model performance. To account for depth-dependent strong optical scattering, our network was trained with simulated photoacoustic datasets of targets embedded at different depths inside tissue media of different scattering levels. The network trained on this novel dataset accurately locates targets in experimental PA data that is clinically relevant with respect to the localization of vessels, needles, or brachytherapy seeds. We verify the merits of the proposed architecture by outperforming the state of the art on both simulated and experimental datasets.
Amirsaeed Yazdani, Sumit Agrawal, Kerrick Johnstonbaugh, Sri-Rajasekhar Kothapalli, Vishal Monga
IEEE Trans. Medical Imaging5
2020 Reinforced Depth-Aware Deep Learning for Single Image Dehazing
abstract
Image dehazing continues to be one of the most challenging inverse problems. Deep learning methods have emerged to complement traditional model-based methods and have helped define a new state of the art in achievable dehazed image quality. However, most deep learning-based methods usually design a regression network as a black-box tool to either estimate the dehazed image and/or the physical parameters in the haze model, i.e. ambient light (A) and transmission map (t). The inverse haze model may then be used to estimate the dehazed image. In this work, we proposed a Depth-aware Dehazing using Reinforcement Learning system, denoted as DDRL. DDRL generates the dehazed image in a near-to-far progressive manner by utilizing the depth-information from the scene. This contrasts with the most recent learning-based methods that estimate these parameters in one pass. In particular, DDRL exploits the fact that the haze is less dense near the camera and gets increasingly denser as the scene moves farther away from the camera. DDRL consists of a policy network and a dehazing (regression) network. The policy network estimates the current depth for the dehazing network to use. A novel policy regularization term is introduced for the policy network to generate the policy sequence following the near-to-far order. Based on extensive tests over three benchmark test sets, DDRL demonstrates vastly enhanced dehazing results, particularly when training is limited.
Tiantong Guo, Vishal Monga
ICASSP2
2020 Attention-Mask Dense Merger (Attendense) Deep HDR for Ghost Removal
abstract
High Dynamic Range (HDR) reconstruction is the process of producing an HDR image from a set of Standard Dynamic Range (SDR) images with different exposure times. This is a particularly challenging problem when relative camera or object motion exists between the available SDR images. Recently, deep learning methods, specifically those based on convolutional neural networks (CNNs) have been developed for HDR and shown to achieve unprecedented quality gains. Invariably an image alignment phase precedes the CNN mapping and merging. In practice, this alignment step greatly increases the computational burden of deep HDR methods often rendering them unsuitable for real-time composition. We propose a new deep HDR technique that does not need any explicit alignment of SDR images. Instead, a novel attention mask is developed that enables the network to focus on parts of the scene with considerable motion. Further, a dense merger is proposed that leads to an economical network. Evaluation over benchmark databases reveals that the proposed AttenDense network achieves high quality HDR results with significantly reduced computation time than state of the art. Further, the incorporation of domain knowledge (development of a custom attention mask) allows a more graceful decay in performance in the face of limited training.
Kareem M. Metwaly, Vishal Monga
ICASSP2
2020 Data Adaptive Image Enhancement and Classification for Synthetic Aperture Sonar
abstract
Deep learning has been recently shown to improve performance in the domain of synthetic aperture sonar (SAS) image classification over existing shallow learning solutions. Given the constant resolution with range of a SAS, it is no surprise that deep learning techniques perform so well; the image of the seafloor produced by a SAS system is almost photographic in quality. Despite the image quality benefits of SAS, there is still room for classification improvement particularly in reducing the number of false alarms. This work addresses this by tackling one facet of the classification pipeline: image enhancement. Specifically, we ask and address the following question: Can we train a deep neural network to simultaneously enhance and classify a SAS image? We will respond in the affirmative as we introduce a new deep learning architecture tackling the problem, Data Adaptive Enhancement and Classification Network (DA-ECNet). DA-ECNet is a deep learning architecture which combines image enhancement as part of the classification procedure eliminating the need for a fixed state-of-the-art despeckling algorithm or enhancement module. Additionally, we train both image enhancement and classification jointly resulting in data adaptive image enhancement. Experiments on a challenging real-world dataset reveal that the proposed DA-ECNet outperforms state of the art deep learning as well as traditional feature based methods for SAS image classification.
Isaac Gerg, David P. Williams, Vishal Monga
IGARSS3
2020 Multiview Automatic Target Recognition for Infrared Imagery Using Collaborative Sparse Priors
abstract
The low resolution of infrared (IR) images makes feature extraction for classification of a challenging work. Learning-based methods, therefore, are preferred to be used on such raw imagery. In this article, in order to avoid difficulties in feature extraction, a novel multitask extension of the widely used sparse-representation-classification (SRC) method is proposed in both single and multiview set-ups. That is, the test sample could be a single IR image or images from different views. In both single-view and multiview scenarios, we try to employ collaborative spike and slab priors. This is because the traditional sparsity-inducing measures such as the l0-row pseudonorm makes it hard to capture the sparse structure of the coefficient matrix when expanded in terms of a training dictionary, and the priors are proved to be able to capture fairly general sparse structures. Furthermore, a joint prior and sparse coefficient estimation method (JPCEM) is proposed for the first time in this article in order to alleviate the need to handpick prior parameters required before classification. Multiple experiments are conducted on a synthetic Comanche Forward Looking IR (FLIR) Automatic Target Recognition (ATR) database collected by Army Research Lab and a challenging mid-wave IR (MWIR) image ATR database made available by the U.S. Army Night Vision and Electronic Sensors Directorate. The final results substantiate the merits of the proposed JPCEM through comparisons with other state-of-the-art methods, including both the ones based on SRC and the ones constructed using deep learning frameworks.
Xuelu Li, Vishal Monga, Abhijit Mahalanobis
IEEE Trans. Geosci. Remote. Sens.2
2020 Deep Retinal Image Segmentation With Regularization Under Geometric Priors
abstract
Vessel segmentation of retinal images is a key diagnostic capability in ophthalmology. This problem faces several challenges including low contrast, variable vessel size and thickness, and presence of interfering pathology such as micro-aneurysms and hemorrhages. Early approaches addressing this problem employed hand-crafted filters to capture vessel structures, accompanied by morphological post-processing. More recently, deep learning techniques have been employed with significantly enhanced segmentation accuracy. We propose a novel domain enriched deep network that consists of two components: 1) a representation network that learns geometric features specific to retinal images, and 2) a custom designed computationally efficient residual task network that utilizes the features obtained from the representation layer to perform pixel-level segmentation. The representation and task networks are jointly learned for any given training set. To obtain physically meaningful and practically effective representation filters, we propose two new constraints that are inspired by expected prior structure on these filters: 1) orientation constraint that promotes geometric diversity of curvilinear features, and 2) a data adaptive noise regularizer that penalizes false positives. Multi-scale extensions are developed to enable accurate detection of thin vessels. Experiments performed on three challenging benchmark databases under a variety of training scenarios show that the proposed prior guided deep network outperforms state of the art alternatives as measured by common evaluation metrics, while being more economical in network size and inference time.
Venkateswararao Cherukuri, Raja Bala, Vishal Monga
IEEE Trans. Image Process.4
2020 Deep MR Brain Image Super-Resolution Using Spatio-Structural Priors
abstract
High resolution Magnetic Resonance (MR) images are desired for accurate diagnostics. In practice, image resolution is restricted by factors like hardware and processing constraints. Recently, deep learning methods have been shown to produce compelling state-of-the-art results for image enhancement/super-resolution. Paying particular attention to desired hi-resolution MR image structure, we propose a new regularized network that exploits image priors, namely a low-rank structure and a sharpness prior to enhance deep MR image super-resolution (SR). Our contributions are then incorporating these priors in an analytically tractable fashion as well as towards a novel prior guided network architecture that accomplishes the super-resolution task. This is particularly challenging for the low rank prior since the rank is not a differentiable function of the image matrix (and hence the network parameters), an issue we address by pursuing differentiable approximations of the rank. Sharpness is emphasized by the variance of the Laplacian which we show can be implemented by a fixed feedback layer at the output of the network. As a key extension, we modify the fixed feedback (Laplacian) layer by learning a new set of training data driven filters that are optimized for enhanced sharpness. Experiments performed on publicly available MR brain image databases and comparisons against existing state-of-the-art methods show that the proposed prior guided network offers significant practical gains in terms of improved SNR/image quality measures. Because our priors are on output images, the proposed method is versatile and can be combined with a wide variety of existing network architectures to further enhance their performance.
Venkateswararao Cherukuri, Tiantong Guo, Steven J. Schiff, Vishal Monga
IEEE Trans. Image Process.4
2019 Group Based Deep Shared Feature Learning for Fine-grained Image Classification
Xuelu Li, Vishal Monga
BMVC2
2019 An Algorithm Unrolling Approach to Deep Image Deblurring
abstract
While neural networks have achieved vastly enhanced performance over traditional iterative methods in many cases, they are generally empirically designed and the underlying structures are difficult to interpret. The algorithm unrolling approach has helped connect iterative algorithms to neural network architectures. However, such connections have not been made yet for blind image deblurring. In this paper, we propose a neural network architecture that advances this idea. We first present an iterative algorithm that may be considered a generalization of the traditional total-variation regularization method on the gradient domain, and subsequently unroll the half-quadratic splitting algorithm to construct a neural network. Our proposed deep network achieves significant practical performance gains while enjoying interpretability at the same time. Experimental results show that our approach outperforms many state-of-the-art methods.
Mohammad Tofighi, Vishal Monga, Yonina C. Eldar
ICASSP3
2019 Multi-Scale Regularized Deep Network for Retinal Vessel Segmentation
abstract
Vessel segmentation of retinal images is a key diagnostic capability in ophthalmology. Early approaches addressing this problem employed hand-crafted filters to capture vessel structures, accompanied by morphological processing. More recently, deep learning techniques have been employed to significantly enhance segmentation accuracy. We propose a novel domain enriched deep network that consists of two components: 1) a representation network which learns geometric (specifically curvilinear) features that are tailored to retinal images, followed by 2) a task network that utilizes the features obtained from the representation layer to perform pixel-level segmentation. The representation and task networks are learned jointly for any given training set. To obtain effective representation filters, we develop a new orientation constraint that enables geometric diversity of curvilinear features. A multi-scale extension is further developed to enhance segmentation of thin vessels. Experiments performed on two challenging benchmark databases reveal that the proposed regularized deep network can outperform state of the art alternatives as measured by common evaluation metrics. Further, the proposed method exhibits a more graceful decay in performance as training data is reduced.
Venkateswararao Cherukuri, Raja Bala, Vishal Monga
ICIP4
2019 Adaptive Transform Domain Image Super-Resolution via Orthogonally Regularized Deep Networks
abstract
Deep learning methods, in particular, trained Convolutional Neural Networks (CNN) have recently been shown to produce compelling results for single image Super-Resolution (SR). Invariably, a CNN is learned to map the Low Resolution (LR) image to its corresponding High Resolution (HR) version in the spatial domain. We propose a novel network structure for learning the SR mapping function in an image transform domain, specifically the Discrete Cosine Transform (DCT). As the first contribution, we show that DCT can be integrated into the network structure as a Convolutional DCT (CDCT) layer. With the CDCT layer, we construct the DCT Deep SR (DCT-DSR) network. We further extend the DCT-DSR to allow the CDCT layer to become trainable (i.e., optimizable). Because this layer represents an image transform, we enforce pairwise orthogonality constraints and newly formulated complexity order constraints on the individual basis functions/filters. This Orthogonally Regularized Deep SR network (ORDSR) simplifies the SR task by taking advantage of image transform domain while adapting the design of transform basis to the training image set. Experimental results show ORDSR achieves state-of-the-art SR image quality with fewer parameters than most of the deep CNN methods. A particular success of ORDSR is in overcoming the artifacts introduced by bicubic interpolation. A key burden of deep SR has been identified as the requirement of generous training LR and HR image pairs; ORSDR exhibits a much more graceful degradation as training size is reduced with significant benefits in the regime of limited training. Analysis of memory and computation requirements confirms that ORDSR can allow for a more efficient network with faster inference.
Tiantong Guo, Hojjat Seyed Mousavi, Vishal Monga
IEEE Trans. Image Process.3
2019 Robust Alignment for Panoramic Stitching Via an Exact Rank Constraint
abstract
We study the problem of image alignment for panoramic stitching. Unlike most existing approaches that are feature-based, our algorithm works on pixels directly, and accounts for errors across the whole images globally. Technically, we formulate the alignment problem as rank-1 and sparse matrix decomposition over transformed images, and develop an efficient algorithm for solving this challenging non-convex optimization problem. The algorithm reduces to solving a sequence of subproblems, where we analytically establish exact recovery conditions, convergence and optimality, together with convergence rate and complexity. We generalize it to simultaneously align multiple images and recover multiple homographies, extending its application scope toward vast majority of practical scenarios. The experimental results demonstrate that the proposed algorithm is capable of more accurately aligning the images and generating higher quality stitched images than the state-of-the-art methods.
Mohammad Tofighi, Vishal Monga
IEEE Trans. Image Process.3
2019 Prior Information Guided Regularized Deep Learning for Cell Nucleus Detection
abstract
Cell nuclei detection is a challenging research topic because of limitations in cellular image quality and diversity of nuclear morphology, i.e., varying nuclei shapes, sizes, and overlaps between multiple cell nuclei. This has been a topic of enduring interest with promising recent success shown by deep learning methods. These methods train convolutional neural networks (CNNs) with a training set of input images and known, labeled nuclei locations. Many such methods are supplemented by spatial or morphological processing. Using a set of canonical cell nuclei shapes, prepared with the help of a domain expert, we develop a new approach that we call shape priors (SPs) with CNNs (SPs-CNN). We further extend the network to introduce an SP layer and then allowing it to become trainable (i.e., optimizable). We call this network as tunable SP-CNN (TSP-CNN). In summary, we present new network structures that can incorporate "expected behavior" of nucleus shapes via two components: learnable layers that perform the nucleus detection and a fixed processing part that guides the learning with prior information. Analytically, we formulate two new regularization terms that are targeted at: 1) learning the shapes and 2) reducing false positives while simultaneously encouraging detection inside the cell nucleus boundary. Experimental results on two challenging datasets reveal that the proposed SP-CNN and TSP-CNN can outperform the state-of-the-art alternatives.
Mohammad Tofighi, Tiantong Guo, Jairam K. P. Vanamala, Vishal Monga
IEEE Trans. Medical Imaging4
2018 Orthogonally Regularized Deep Networks for Image Super-Resolution
abstract
Deep learning methods, in particular trained Convolutional Neural Networks (CNNs) have recently been shown to produce compelling state-of-the-art results for single image Super-Resolution (SR). Invariably, a CNN is learned to map the low resolution (LR) image to its corresponding high resolution (HR) version in the spatial domain. Aiming for faster inference and more efficient solutions than solving the SR problem in the spatial domain, we propose a novel network structure for learning the SR mapping function in an image transform domain, specifically the Discrete Cosine Transform (DCT). As a first contribution, we show that DCT can be integrated into the network structure as a Convolutional DCT (CDCT) layer. We further extend the network to allow the CDCT layer to become trainable (i.e. optimizable). Because this layer represents an image transform, we enforce pairwise orthogonality constraints on the individual basis functions/filters. This Orthogonally Regularized Deep SR network (ORDSR) simplifies the SR task by taking advantage of image transform domain while adapting the design of transform basis to the training image set. Experimental results show ORDSR achieves state-of-the-art SR image quality with fewer parameters than most of the deep CNN methods.
Tiantong Guo, Hojjat Seyed Mousavi, Vishal Monga
ICASSP3
2018 Deep Image Super Resolution via Natural Image Priors
abstract
Single image super-resolution (SR) via deep learning has recently gained significant attention in the literature. Convolutional neural networks (CNNs) are typically learned to represent the mapping between low-resolution (LR) and high-resolution (HR) images/patches with the help of training examples. Most existing deep networks for SR produce high quality results when training data is abundant. However, their performance degrades sharply when training is limited. We propose to regularize deep structures with prior knowledge about the images so that they can capture more structural information from the same limited data. In particular, we incorporate in a tractable fashion within the CNN framework, natural image priors which have shown to have much recent success in imaging and vision inverse problems. Experimental results show that the proposed deep network with natural image priors is particularly effective in training starved regimes.
Hojjat Seyed Mousavi, Tiantong Guo, Vishal Monga
ICASSP3
2018 Deep Mr Image Super-Resolution Using Structural Priors
abstract
High resolution magnetic resonance (MR) images are desired for accurate diagnostics. In practice, image resolution is restricted by factors like hardware, cost and processing constraints. Recently, deep learning methods have been shown to produce compelling state of the art results for image super-resolution. Paying particular attention to desired hi-resolution MR image structure, we propose a new regularized network that exploits image priors, namely a low-rank structure and a sharpness prior to enhance deep MR image superresolution. Our contributions are then incorporating these priors in an analytically tractable fashion in the learning of a convolutional neural network (CNN) that accomplishes the super-resolution task. This is particularly challenging for the low rank prior, since the rank is not a differentiable function of the image matrix (and hence the network parameters), an issue we address by pursuing differentiable approximations of the rank. Sharpness is emphasized by the variance of the Laplacian which we show can be implemented by a fixed feedback layer at the output of the network. Experiments performed on two publicly available MR brain image databases exhibit promising results particularly when training imagery is limited.
Venkateswararao Cherukuri, Tiantong Guo, Steven J. Schiff, Vishal Monga
ICIP4
2018 Deep Networks with Shape Priors for Nucleus Detection
abstract
Detection of cell nuclei in microscopic images is a challenging research topic, because of limitations in cellular image quality and diversity of nuclear morphology, i.e. varying nuclei shapes, sizes, and overlaps between multiple cell nuclei. This has been a topic of enduring interest with promising recent success shown by deep learning methods. These methods train for example convolutional neural networks (CNNs) with a training set of input images and known, labeled nuclei locations. Many of these methods are supplemented by spatial or morphological processing. We develop a new approach that we call Shape Priors with Convolutional Neural Networks (SP-CNN) to perform significantly enhanced nuclei detection. A set of canonical shapes is prepared with the help of a domain expert. Subsequently, we present a new network structure that can incorporate `expected behavior' of nucleus shapes via two components: learnable layers that perform the nucleus detection and a fixed processing part that guides the learning with prior information. Analytically, we formulate a new regularization term that is targeted at penalizing false positives while simultaneously encouraging detection inside cell nucleus boundary. Experimental results on a challenging dataset reveal that SP-CNN is competitive with or outperforms several state-of-the-art methods.
Mohammad Tofighi, Tiantong Guo, Jairam K. P. Vanamala, Vishal Monga
ICIP4
2018 Collaborative Sparse Priors for Infrared Image Multi-View ATR
abstract
Feature extraction from infrared (IR) images remains a challenging task. Learning based methods that can work on raw imagery/patches have therefore assumed significance. We propose a novel multi-task extension of the widely used sparse-representation-classification (SRC) method in both single and multi-view set-ups. That is, the test sample could be a single IR image or images from different views. When expanded in terms of a training dictionary, the coefficient matrix in a multi-view scenario admits a sparse structure that is not easily captured by traditional sparsity-inducing measures such as the l0-row pseudo norm. To that end, we employ collaborative spike and slab priors on the coefficient matrix, which can capture fairly general sparse structures. Our work involves joint parameter and sparse coefficient estimation (JPCEM) which alleviates the need to handpick prior parameters before classification. The experimental merits of JPCEM are substantiated through comparisons with other state-of-art methods on a challenging mid-wave IR image (MWIR) ATR database made available by the US Army Night Vision and Electronic Sensors Directorate.
Xuelu Li, Vishal Monga
IGARSS2
2018 Bridging The GAP: Simultaneous Fine Tuning for Data Re-Balancing
abstract
There are many real-world classification problems wherein the issue of data imbalance (the case when a data set contains substantially more samples for one/many classes than the rest) is unavoidable. While under-sampling the problematic classes is a common solution, this is not a compelling option when the large data class is itself diverse and/or the limited data class is especially small. We suggest a strategy based on recent work concerning limited data problems which utilizes a supplemental set of images with similar properties to the limited data class to aid in the training of a neural network. We show results for our model against other typical methods on a real-world synthetic aperture sonar data set. Code can be found at github.com/JohnMcKay/dataImbalance.
John McKay 0002, Isaac Gerg, Vishal Monga
IGARSS3
2018 Blind Image Deblurring Using Row-Column Sparse Representations
abstract
Blind image deblurring is a particularly challenging inverse problem where the blur kernel is unknown and must be estimated en route to recover the deblurred image. The problem is of strong practical relevance since many imaging devices, such as cellphone cameras, must rely on deblurring algorithms to yield satisfactory image quality. Despite significant research effort, handling large motions remains an open problem. In this letter, we develop a new method called blind image deblurring using row-column sparsity (BD-RCS) to address this issue. Specifically, we model the outer product of kernel and image coefficients in certain transformation domains as a rank-one matrix, and recover it by solving a rank minimization problem. Our central contribution then includes solving two new optimization problems involving RCS to automatically determine blur kernel and image support sequentially. The kernel and image can then be recovered through a singular value decomposition. Experimental results on linear motion deblurring demonstrate that BD-RCS can yield better results than state of the art, particularly when the blur is caused by large motion. This is confirmed both visually and through quantitative measures.
Mohammad Tofighi, Vishal Monga
IEEE Signal Process. Lett.3
2017 Adaptive matching pursuit for sparse signal recovery
abstract
Spike and Slab priors have been of much recent interest in signal processing as a means of inducing sparsity in Bayesian inference. Applications domains that benefit from the use of these priors include sparse recovery, regression and classification. It is well-known that solving for the sparse coefficient vector to maximize these priors results in a hard non-convex and mixed integer programming problem. Most existing solutions to this optimization problem either involve simplifying assumptions/relaxations or are computationally expensive. We propose a new greedy and adaptive matching pursuit (AMP) algorithm to directly solve this hard problem. Essentially, in each step of the algorithm, the set of active elements would be updated by either adding or removing one index, whichever results in better improvement. In addition, the intermediate steps of the algorithm are calculated via an inexpensive Cholesky decomposition which makes the algorithm much faster. Results on simulated data sets as well as real-world image recovery challenges confirm the benefits of the proposed AMP, particularly in providing a superior cost-quality trade-off over existing alternatives.
Tiep Huu Vu, Hojjat Seyed Mousavi, Vishal Monga
ICASSP3
2017 Robust Sonar ATR Through Bayesian Pose-Corrected Sparse Classification
abstract
Sonar imaging has seen vast improvements over the last few decades due in part to advances in synthetic aperture sonar. Sophisticated classification techniques can now be used in sonar automatic target recognition (ATR) to locate mines and other threatening objects. Among the most promising of these methods is sparse reconstruction-based classification (SRC), which has shown an impressive resiliency to noise, blur, and occlusion. We present a coherent strategy for expanding upon SRC for sonar ATR that retains SRC's robustness while also being able to handle targets with diverse geometric arrangements, bothersome Rayleigh noise, and unavoidable background clutter. Our method, pose-corrected sparsity (PCS), incorporates a novel interpretation of a spike and slab probability distribution toward use as a Bayesian prior for class-specific discrimination in combination with a dictionary learning scheme for localized patch extractions. Additionally, PCS offers the potential for anomaly detection in order to avoid false identifications of tested objects from outside the training set with no additional training required. Compelling results are shown using a database provided by the U.S. Naval Surface Warfare Center.
John McKay 0002, Vishal Monga, Raghu G. Raj
IEEE Trans. Geosci. Remote. Sens.2
2017 A Maximum a Posteriori Estimation Framework for Robust High Dynamic Range Video Synthesis
abstract
High dynamic range (HDR) image synthesis from multiple low dynamic range exposures continues to be actively researched. The extension to HDR video synthesis is a topic of significant current interest due to potential cost benefits. For HDR video, a stiff practical challenge presents itself in the form of accurate correspondence estimation of objects between video frames. In particular, loss of data resulting from poor exposures and varying intensity makes conventional optical flow methods highly inaccurate. We avoid exact correspondence estimation by proposing a statistical approach via maximum a posterior estimation, and under appropriate statistical assumptions and choice of priors and models, we reduce it to an optimization problem of solving for the foreground and background of the target frame. We obtain the background through rank minimization and estimate the foreground via a novel multiscale adaptive kernel regression technique, which implicitly captures local structure and temporal motion by solving an unconstrained optimization problem. Extensive experimental results on both real and synthetic data sets demonstrate that our algorithm is more capable of delivering high-quality HDR videos than current state-of-the-art methods, under both subjective and objective assessments. Furthermore, a thorough complexity analysis reveals that our algorithm achieves better complexity-performance tradeoff than conventional methods.
Chul Lee, Vishal Monga
IEEE Trans. Image Process.3
2017 Sparsity-Based Color Image Super Resolution via Exploiting Cross Channel Constraints
abstract
Sparsity constrained single image super-resolution (SR) has been of much recent interest. A typical approach involves sparsely representing patches in a low-resolution (LR) input image via a dictionary of example LR patches, and then using the coefficients of this representation to generate the high-resolution (HR) output via an analogous HR dictionary. However, most existing sparse representation methods for SR focus on the luminance channel information and do not capture interactions between color channels. In this paper, we extend sparsity-based SR to multiple color channels by taking the color information into account. Edge similarities amongst RGB color bands are exploited as cross channel correlation constraints. These additional constraints lead to a new optimization problem, which is not easily solvable; however, a tractable solution is proposed to solve it efficiently. Moreover, to fully exploit the complementary information among color channels, a dictionary learning method is also proposed specifically to learn color dictionaries that encourage edge similarities. Merits of the proposed method over state of the art are demonstrated both visually and quantitatively using image quality metrics.
Hojjat Seyed Mousavi, Vishal Monga
IEEE Trans. Image Process.2
2017 Fast Low-Rank Shared Dictionary Learning for Image Classification
abstract
Despite the fact that different objects possess distinct class-specific features, they also usually share common patterns. This observation has been exploited partially in a recently proposed dictionary learning framework by separating the particularity and the commonality (COPAR). Inspired by this, we propose a novel method to explicitly and simultaneously learn a set of common patterns as well as class-specific features for classification with more intuitive constraints. Our dictionary learning framework is hence characterized by both a shared dictionary and particular (class-specific) dictionaries. For the shared dictionary, we enforce a low-rank constraint, i.e., claim that its spanning subspace should have low dimension and the coefficients corresponding to this dictionary should be similar. For the particular dictionaries, we impose on them the well-known constraints stated in the Fisher discrimination dictionary learning (FDDL). Furthermore, we develop new fast and accurate algorithms to solve the subproblems in the learning step, accelerating its convergence. The said algorithms could also be applied to FDDL and its extensions. The efficiencies of these algorithms are theoretically and experimentally verified by comparing their complexities and running time with those of other well-known dictionary learning methods. Experimental results on widely used image data sets establish the advantages of our method over the state-of-the-art dictionary learning methods.
Tiep Huu Vu, Vishal Monga
IEEE Trans. Image Process.2
2016 SIASM: Sparsity-based image alignment and stitching method for robust image mosaicking
abstract
Image alignment and stitching continue to be the topics of great interest. Image mosaicking is a key application that involves both alignment and stitching of multiple images. Despite significant previous effort, existing methods have limited robustness in dealing with occlusions and local object motion in different captures. To address this issue, we investigate the potential of applying sparsity-based methods to the task of image alignment and stitching. We formulate the alignment problem as a low-rank and sparse matrix decomposition problem under incomplete observations (multiple parts of a scene), and the stitching problem as a multiple labeling problem which utilizes the sparse components. Additionally we develop efficient algorithms for solving them. Unlike typical pairwise alignment manners in classical image alignment algorithms, our algorithm is capable of simultaneously aligning multiple images, making full use of inter-frame relationships among all images. Experimental results demonstrate that the proposed algorithm is capable of generating artifact-free stitched image mosaics that are robust against occlusions and object motion.
Vishal Monga
ICIP2
2016 Sparsity based super resolution using color channel constraints
abstract
Sparsity constrained single image super-resolution has been of much recent interest. A typical approach involves sparsely representing patches in a low-resolution (LR) input image via a dictionary of example LR patches, and then use the coefficients of this representation to generate the high-resolution (HR) output via an analogous HR dictionary. However, most existing sparse representation methods for super resolution only use luminance channel information and do not use any information from other color channels. In this work, we extend sparsity based super-resolution to multiple color channels. Edge similarities amongst color bands are exploited as cross channel correlation constraints. These additional constraints lead to a new optimization problem which is not easily solvable; however, a tractable solution is proposed to solve it efficiently. Experimental results shows the merits of our proposed method both visually and quantitatively.
Hojjat Seyed Mousavi, Vishal Monga
ICIP2
2016 Learning a low-rank shared dictionary for object classification
abstract
Despite the fact that different objects possess distinct class-specific features, they also usually share common patterns. Inspired by this observation, we propose a novel method to explicitly and simultaneously learn a set of common patterns as well as class-specific features for classification. Our dictionary learning framework is hence characterized by both a shared dictionary and particular (class-specific) dictionaries. For the shared dictionary, we enforce a low-rank constraint, i.e. claim that its spanning subspace should have low dimension and the coefficients corresponding to this dictionary should be similar. For the particular dictionaries, we impose on them the well-known constraints stated in the Fisher discrimination dictionary learning (FDDL). Further, we propose a new fast and accurate algorithm to solve the sparse coding problems in the learning step, accelerating its convergence. The said algorithm could also be applied to FDDL and its extensions. Experimental results on widely used image databases establish the advantages of our method over state-of-the-art dictionary learning methods.
Tiep Huu Vu, Vishal Monga
ICIP2
2016 Localized dictionary design for geometrically robust sonar ATR
abstract
Advancements in Sonar image capture have opened the door to powerful classification schemes for automatic target recognition (ATR). Recent work has particularly seen the application of sparse reconstruction-based classification (SRC) to sonar ATR, which provides compelling accuracy rates even in the presence of noise and blur. However, existing sparsity based sonar ATR techniques assume that the test images exhibit geometric pose that is consistent with respect to the training set. This work addresses the outstanding open challenge of handling inconsistently posed Sonar images relative to training. We develop a new localized block-based dictionary design that can enable geometric robustness. Further, a dictionary learning method is incorporated to increase performance and efficiency. The proposed SRC with Localized Pose Management (LPM), is shown to outperform the state of the art SIFT feature and SVM approach, due to its power to discern background clutter in Sonar images.
John McKay 0002, Vishal Monga, Raghu G. Raj
IGARSS2
2016 Power-Constrained RGB-to-RGBW Conversion for Emissive Displays: Optimization-Based Approaches
abstract
We propose an optimization-based power-constrained red-green-blue (RGB)-to-red-green-blue-white (RGBW) conversion algorithm for emissive RGBW displays. We measure the perceived color distortion using a color difference model in a perceptually uniform color space, and compute the power consumption for displaying an RGBW pixel on an emissive display. The central contribution of this paper is to formulate the optimization problem to minimize the color distortion subject to a constraint on the power consumption. Subsequently, we solve the optimization problem efficiently to convert an image in real time. Furthermore, based on the properties of the human visual system, we extend the proposed algorithm to image-dependent conversion that can preserve spatial detail in an input image. The simulation results show that the proposed algorithm provides a significantly less color distortion than the conventional methods, while providing a graceful tradeoff with the amount of power consumed. Specifically, it is shown that the power consumption can be reduced by up to 20%, while providing about 50% less color distortion than the conventional algorithms. In addition, a subjective evaluation on a real RGBW display is performed, which reveals the merits of the proposed image-dependent conversion for improving the perceptual quality over state-of-the-art techniques.
Chul Lee, Vishal Monga
IEEE Trans. Circuits Syst. Video Technol.2
2016 Histopathological Image Classification Using Discriminative Feature-Oriented Dictionary Learning
abstract
In histopathological image analysis, feature extraction for classification is a challenging task due to the diversity of histology features suitable for each problem as well as presence of rich geometrical structures. In this paper, we propose an automatic feature discovery framework via learning class-specific dictionaries and present a low-complexity method for classification and disease grading in histopathology. Essentially, our Discriminative Feature-oriented Dictionary Learning (DFDL) method learns class-specific dictionaries such that under a sparsity constraint, the learned dictionaries allow representing a new image sample parsimoniously via the dictionary corresponding to the class identity of the sample. At the same time, the dictionary is designed to be poorly capable of representing samples from other classes. Experiments on three challenging real-world image databases: 1) histopathological images of intraductal breast lesions, 2) mammalian kidney, lung and spleen images provided by the Animal Diagnostics Lab (ADL) at Pennsylvania State University, and 3) brain tumor images from The Cancer Genome Atlas (TCGA) database, reveal the merits of our proposal over state-of-the-art alternatives. Moreover, we demonstrate that DFDL exhibits a more graceful decay in classification accuracy against the number of training images which is highly desirable in practice where generous training is often not available.
Tiep Huu Vu, Hojjat Seyed Mousavi, Vishal Monga, Ganesh Rao, Arvind U. K. Rao
IEEE Trans. Medical Imaging3
2015 A map estimation framework for HDR video synthesis
abstract
High dynamic range (HDR) image synthesis from multiple low dynamic range (LDR) exposures continues to be a topic of great interest. The extension to HDR video comprises a stiff challenge due to significant motion. In particular, loss of data due to poor exposures introduces great difficulty in exact motion estimation, and under such circumstances conventional optical flow calculation techniques usually fail. We propose a maximum a posterior (MAP) estimation framework for HDR video synthesis algorithm free of explicit optical flow calculation. We formulate HDR video synthesis as a MAP estimation problem, which subsequently can be reduced to an optimization problem based on meaningful statistical assumptions on foreground and background regions of the input video. In the background regions the underlying scenes are static, while in the foreground regions motion information is captured implicitly by a modified 3D steering kernel regression (3D SKR) approach. Solution to the optimization problem provides us with temporally coherent HDR video sequences without noticeable artifacts. Experimental results on challenging LDR video sets demonstrate that our proposed algorithm can achieve HDR video quality that is competitive with or better than state of the art alternatives.
Chul Lee, Vishal Monga
ICIP3
2015 Iterative Convex Refinement for Sparse Recovery
abstract
In this letter, we address sparse signal recovery in a Bayesian framework where sparsity is enforced on reconstruction coefficients via probabilistic priors. In particular, we focus on the setup of Yenwho employ a variant of spike and slab prior to encourage sparsity. The optimization problem resulting from this model has broad applicability in recovery and regression problems and is known to be a hard non-convex problem whose existing solutions involve simplifying assumptions and/or relaxations. We propose an approach called Iterative Convex Refinement (ICR) that aims to solve the aforementioned optimization problem directly allowing for greater generality in the sparse structure. Essentially, ICR solves a sequence of convex optimization problems such that sequence of solutions converges to a sub-optimal solution of the original hard optimization problem. We propose two versions of our algorithm: a.) an unconstrained version, and b.) with a non-negativity constraint on sparse coefficients, which may be required in some real-world problems. Experimental validation is performed on both synthetic data and for a real-world image recovery problem, which illustrates merits of ICR over state of the art alternatives.
Hojjat Seyed Mousavi, Vishal Monga, Trac D. Tran
IEEE Signal Process. Lett.2
2015 Graph-Based Sensor Fusion for Classification of Transient Acoustic Signals
abstract
Advances in acoustic sensing have enabled the simultaneous acquisition of multiple measurements of the same physical event via co-located acoustic sensors. We exploit the inherent correlation among such multiple measurements for acoustic signal classification, to identify the launch/impact of munition (i.e., rockets, mortars). Specifically, we propose a probabilistic graphical model framework that can explicitly learn the class conditional correlations between the cepstral features extracted from these different measurements. Additionally, we employ symbolic dynamic filtering-based features, which offer improvements over the traditional cepstral features in terms of robustness to signal distortions. Experiments on real acoustic data sets show that our proposed algorithm outperforms conventional classifiers as well as the recently proposed joint sparsity models for multisensor acoustic classification. Additionally our proposed algorithm is less sensitive to insufficiency in training samples compared to competing approaches.
Umamahesh Srinivas, Nasser M. Nasrabadi, Vishal Monga
IEEE Trans. Cybern.3
2015 Twofold Video Hashing With Automatic Synchronization
abstract
As a robust media representation technique, video hashing is frequently used in near-duplicate detection, video authentication, and antipiracy search. Distortions to a video may include spatial modifications to each frame, temporal de-synchronization, and joint spatio-temporal attacks. To address the increasingly difficult case of finding videos under spatio-temporal modifications, we propose a new framework called two-stage video hashing. First, an efficient automatic synchronization is achieved using dynamic time warping (DTW) and a complementary video comparison measure is developed based on flow hashing (FH), which is extracted from the synchronized videos. Next, a fusion mechanism called distance boosting is proposed to fuse the information extracted by DTW and FH in a future-proof manner in the sense whenever model retraining is needed, the existing hash vectors do not need to be regenerated. Experiments on real video collections show that such a hash extraction and fusion method enables unprecedented robustness under both spatial and temporal attacks.
Mu Li 0002, Vishal Monga
IEEE Trans. Inf. Forensics Secur.2
2015 Structured Sparse Priors for Image Classification
abstract
Model-based compressive sensing (CS) exploits the structure inherent in sparse signals for the design of better signal recovery algorithms. This information about structure is often captured in the form of a prior on the sparse coefficients, with the Laplacian being the most common such choice (leading to l1 -norm minimization). Recent work has exploited the discriminative capability of sparse representations for image classification by employing class-specific dictionaries in the CS framework. Our contribution is a logical extension of these ideas into structured sparsity for classification. We introduce the notion of discriminative class-specific priors in conjunction with class specific dictionaries, specifically the spike-and-slab prior widely applied in Bayesian sparse regression. Significantly, the proposed framework takes the burden off the demand for abundant training image samples necessary for the success of sparsity-based classification schemes. We demonstrate this practical benefit of our approach in important applications, such as face recognition and object categorization.
Umamahesh Srinivas, Yuanming Suo, Minh Dao, Vishal Monga, Trac D. Tran
IEEE Trans. Image Process.4
2014 Power-constrained RGB-to-RGBW conversion for emissive displays
abstract
We propose a novel power-constrained RGB-to-RGBW conversion algorithm for emissive RGBW displays. We measure the perceived color distortion using a color difference model in a perceptually uniform color space, and compute the power consumption for displaying an RGBW pixel on an emissive display. The main contribution is to formulate the optimization problem to minimize the color distortion subject to the constraint on the power consumption. Then, we solve it efficiently to convert an image in real time. Simulation results show that the proposed algorithm provides significantly less color distortion than the conventional methods while providing a graceful trade-off with the amount of power consumed.
Chul Lee, Vishal Monga
ICASSP2
2014 Low rank sparsity prior for robust video anomaly detection
abstract
Recently, sparsity based classification has been applied to video anomaly detection. A linear model is assumed over video features (e.g. trajectories) such that the feature representation of a new event is written as a sparse linear combination of existing feature representations in the dictionary. Sparsity based video anomaly detection shows promise but open challenges remain in that existing methods assume object specific and class specific event dictionaries making them applicable mostly in highly structured scenarios. Second, using conventional sparsity models on matrices/vectors, the computational burden is often high. In this work, we advocate a more general and practical sparsity model using a low-rank structure on the matrix of sparse coefficients. We find that enforcing a low-rank structure can ease the rigidity of traditional row-sparse constraints on sparse coefficient vectors/matrices. Because low-rank matrices are of course not always sparse, an additional l1regularization term is added. Further, if rank is substituted by its convex nuclear norm alternative, then significant computational benefits can be obtained over existing methods in sparsity based video anomaly detection. Experimental evaluation on benchmark video datasets reveal, our method is competitive with state-of-the art while providing robustness benefits under occlusion.
Xuan Mo, Vishal Monga, Raja Bala, Zhigang Fan 0001, Aaron M. Burry
ICASSP2
2014 Twofold video hashing with automatic synchronization
abstract
Video hashing finds a wide array of applications in content authentication, robust retrieval and anti-piracy search. While much of the existing research has focused on extracting robust and secure content descriptors, a significant open challenge still remains: Most existing video hashing methods are fallible to temporal desynchronization. That is, when the query video results by deleting or inserting some frames from the reference video, most existing methods assume the positions of the deleted (or inserted) frames are either perfectly known or reliably estimated. This assumption may be okay under typical transcoding and frame-rate changes but is highly inappropriate in adversarial scenarios such as anti-piracy video search. For example, an illegal uploader will try to bypass the `piracy check' mechanism of YouTube/Dailymotion etc by performing a cleverly designed non-uniform resampling of the video. We present a new solution based on dynamic time warping (DTW), which can implement automatic synchronization and can be used together with existing video hashing methods. The second contribution of this paper is to propose a new robust feature extraction method called flow hashing (FH), based on frame averaging and optical flow descriptors. Finally, a fusion mechanism called distance boosting is proposed to combine the information extracted by DTW and FH. Experiments on real video collections show that such a hash extraction and comparison enables unprecedented robustness under both spatial and temporal attacks.
Mu Li 0002, Vishal Monga
ICIP2
2014 Simultaneous sparsity model for multi-perspective video anomaly detection
abstract
Recently, sparsity based classification has been applied to video anomaly detection. A linear model is assumed over video features (e.g. trajectories) such that the feature representation of a new event is written as a sparse linear combination of existing feature representations in the dictionary. Sparsity based video anomaly detection has shown promise over alternate video anomaly detection methods in that the sparse representations exhibit excellent robustness under noise (common in surveillance videos) and missing or corrupted features, e.g. vehicle occlusion in transportation videos. One limitation of existing sparsity based video anomaly detection techniques is that they are based on only a single feature representation (known formally as video event encoding). One can easily envision that different event representations such as object trajectories and spatio-temporal volumes often contain correlated yet complementary information. In this paper, we propose to extend sparsity models based on single feature representations to simultaneous sparse representations of multiple feature representations. In this model, the matrix of sparse coefficients does not confirm to the commonly seen row-sparsity and a modified greedy heuristic approach that extends simultaneous orthogonal matching pursuit (SOMP) is needed to solve the resulting optimization problem. Experiments on two benchmark video datasets reveal that our method significantly outperforms state-of-the art approaches that utilize only a single-perspective or event encoding.
Xuan Mo, Vishal Monga, Raja Bala
ICIP2
2014 Multi-task image classification via collaborative, hierarchical spike-and-slab priors
abstract
Promising results have been achieved in image classification problems by exploiting the discriminative power of sparse representations for classification (SRC). Recently, it has been shown that the use of class-specific spike-and-slab priors in conjunction with the class-specific dictionaries from SRC is particularly effective in low training scenarios. As a logical extension, we build on this framework for multitask scenarios, wherein multiple representations of the same physical phenomena are available. We experimentally demonstrate the benefits of mining joint information from different camera views for multi-view face recognition.
Hojjat Seyed Mousavi, Umamahesh Srinivas, Vishal Monga, Yuanming Suo, Minh Dao, Trac D. Tran
ICIP3
2014 Group structured dirty dictionary learning for classification
abstract
Dictionary learning techniques have gained tremendous success in many classification problems. Inspired by the dirty model for multi-task regression problems, we proposed a novel method called group-structured dirty dictionary learning (GDDL) that incorporates the group structure (for each task) with the dirty model (across tasks) in the dictionary training process. Its benefits are two-fold: 1) the group structure enforces implicitly the label consistency needed between dictionary atoms and training data for classification; and 2) for each class, the dirty model separates the sparse coefficients into ones with shared support and unique support, with the first set being more discriminative. We use proximal operators and block coordinate decent to solve the optimization problem. GDDL has been shown to give state-of-art result on both synthetic simulation and two face recognition datasets.
Yuanming Suo, Minh Dao, Trac D. Tran, Hojjat Seyed Mousavi, Umamahesh Srinivas, Vishal Monga
ICIP6
2014 Ghost-Free High Dynamic Range Imaging via Rank Minimization
abstract
We propose a ghost-free high dynamic range (HDR) image synthesis algorithm using a low-rank matrix completion framework, which we call RM-HDR. Based on the assumption that irradiance maps are linearly related to low dynamic range (LDR) image exposures, we formulate ghost region detection as a rank minimization problem. We incorporate constraints on moving objects, i.e., sparsity, connectivity, and priors on under- and over-exposed regions into the framework. Experiments on real image collections show that the RM-HDR can often provide significant gains in synthesized HDR image quality over state-of-the-art approaches. Additionally, a complexity analysis is performed which reveals computational merits of RM-HDR over recent advances in deghosting for HDR.
Chul Lee, Vishal Monga
IEEE Signal Process. Lett.3
2014 Adaptive Sparse Representations for Video Anomaly Detection
abstract
Video anomaly detection can be used in the transportation domain to identify unusual patterns such as traffic violations, accidents, unsafe driver behavior, street crime, and other suspicious activities. A common class of approaches relies on object tracking and trajectory analysis. Very recently, sparse reconstruction techniques have been employed in video anomaly detection. The fundamental underlying assumption of these methods is that any new feature representation of a normal/anomalous event can be approximately modeled as a (sparse) linear combination prelabeled feature representations (of previously observed events) in a training dictionary. Sparsity can be a powerful prior on model coefficients but challenges remain in the detection of anomalies involving multiple objects and the ability of the linear sparsity model to effectively allow for class separation. The proposed research addresses both these issues. First, we develop a new joint sparsity model for anomaly detection that enables the detection of joint anomalies involving multiple objects. This extension is highly nontrivial since it leads to a new simultaneous sparsity problem that we solve using a greedy pursuit technique. Second, we introduce nonlinearity into, that is, kernelize. The linear sparsity model to enable superior class separability and hence anomaly detection. We extensively test on several real world video datasets involving both single and multiple object anomalies. Results show marked improvements in detection of anomalies in both supervised and unsupervised scenarios when using the proposed sparsity models.
Xuan Mo, Vishal Monga, Raja Bala, Zhigang Fan 0001
IEEE Trans. Circuits Syst. Video Technol.2
2014 Simultaneous Sparsity Model for Histopathological Image Representation and Classification
abstract
The multi-channel nature of digital histopathological images presents an opportunity to exploit the correlated color channel information for better image modeling. Inspired by recent work in sparsity for single channel image classification, we propose a new simultaneous sparsity model for multi-channel histopathological image representation and classification (SHIRC). Essentially, we represent a histopathological image as a sparse linear combination of training examples under suitable channel-wise constraints. Classification is performed by solving a newly formulated simultaneous sparsity-based optimization problem. A practical challenge is the correspondence of image objects (cellular and nuclear structures) at different spatial locations in the image. We propose a robust locally adaptive variant of SHIRC (LA-SHIRC) to tackle this issue. Experiments on two challenging real-world image data sets: 1) mammalian tissue images acquired by pathologists of the animal diagnostics lab (ADL) at Pennsylvania State University, and 2) human intraductal breast lesions, reveal the merits of our proposal over state-of-the-art alternatives. Further, we demonstrate that LA-SHIRC exhibits a more graceful decay in classification accuracy against the number of training images which is highly desirable in practice where generous training per class is often not available.
Umamahesh Srinivas, Hojjat Seyed Mousavi, Vishal Monga, Arthur Hattel, Bhushan Jayarao
IEEE Trans. Medical Imaging3
2013 Graph-based multi-sensor fusion for acoustic signal classification
abstract
Advances in acoustic sensing have enabled the simultaneous acquisition of multiple measurements of the same physical event via co-located acoustic sensors. We exploit the inherent correlation among such multiple measurements for acoustic signal classification, to identify the launch/impact of munition (i.e. rockets, mortars). Specifically, we propose a probabilistic graphical model framework that can explicitly learn the class conditional correlations between the cepstral features extracted from these different measurements. Additionally, we employ symbolic dynamic filtering-based features, which offer improvements over the traditional cepstral features. Experiments on real acoustic data sets show that our proposed algorithm outperforms conventional classifiers as well as recently proposed joint sparsity models for multi-sensor acoustic signal classification.
Umamahesh Srinivas, Nasser M. Nasrabadi, Vishal Monga
ICASSP3
2013 Hierarchical sparse modeling using Spike and Slab priors
abstract
Sparse modeling has demonstrated its superior performances in many applications. Compared to optimization based approaches, Bayesian sparse modeling generally provides a more sparse result with a knowledge of confidence. Using the Spike and Slab priors, we propose the hierarchical sparse models for the scenario of single task and multitask - Hi-BCS and CHi-BCS. We draw the connections of these two methods to their optimization based counterparts and use expectation propagation for inference. The experiment results using synthetic and real data demonstrate that the performance of Hi-BCS and Chi-BCS are comparable or better than their optimization based counterparts.
Yuanming Suo, Minh Dao, Trac D. Tran, Umamahesh Srinivas, Vishal Monga
ICASSP5
2013 Structured sparse priors for image classification
abstract
Model-based compressive sensing (CS) exploits the structure inherent in sparse signals for the design of better signal recovery algorithms. This information about structure is often captured in the form of a prior on the sparse coefficients, the Laplacian being the most common such choice (leading to l1-norm minimization). The recent seminal contribution by Wright et al. exploits the discriminative capability of sparse representations for image classification, specifically face recognition. Their approach employs the analytical framework of CS with class-specific dictionaries. Our contribution is a logical extension of these ideas into structured sparsity for classification. We use class-specific dictionaries in conjunction with discriminative class-specific priors, specifically the spike-and-slab prior widely applied in Bayesian regression. Significantly, the proposed framework takes the burden off the demand for abundant training image samples necessary for the success of sparsity-based classification schemes.
Umamahesh Srinivas, Yuanming Suo, Minh Dao, Vishal Monga, Trac D. Tran
ICIP4
2013 Exploiting Sparsity in Hyperspectral Image Classification via Graphical Models
abstract
A significant recent advance in hyperspectral image (HSI) classification relies on the observation that the spectral signature of a pixel can be represented by a sparse linear combination of training spectra from an overcomplete dictionary. A spatiospectral notion of sparsity is further captured by developing a joint sparsity model, wherein spectral signatures of pixels in a local spatial neighborhood (of the pixel of interest) are constrained to be represented by a common collection of training spectra, albeit with different weights. A challenging open problem is to effectively capture the class conditional correlations between these multiple sparse representations corresponding to different pixels in the spatial neighborhood. We propose a probabilistic graphical model framework to explicitly mine the conditional dependences between these distinct sparse features. Our graphical models are synthesized using simple tree structures which can be discriminatively learnt (even with limited training samples) for classification. Experiments on benchmark HSI data sets reveal significant improvements over existing approaches in classification rates as well as robustness to choice of training.
Umamahesh Srinivas, Yi Chen 0014, Vishal Monga, Nasser M. Nasrabadi, Trac D. Tran
IEEE Geosci. Remote. Sens. Lett.3
2013 Robust Extrema Features for Time-Series Data Analysis
abstract
The extraction of robust features for comparing and analyzing time series is a fundamentally important problem. Research efforts in this area encompass dimensionality reduction using popular signal analysis tools such as the discrete Fourier and wavelet transforms, various distance metrics, and the extraction of interest points from time series. Recently, extrema features for analysis of time-series data have assumed increasing significance because of their natural robustness under a variety of practical distortions, their economy of representation, and their computational benefits. Invariably, the process of encoding extrema features is preceded by filtering of the time series with an intuitively motivated filter (e.g., for smoothing), and subsequent thresholding to identify robust extrema. We define the properties of robustness, uniqueness, and cardinality as a means to identify the design choices available in each step of the feature generation process. Unlike existing methods, which utilize filters "inspired" from either domain knowledge or intuition, we explicitly optimize the filter based on training time series to optimize robustness of the extracted extrema features. We demonstrate further that the underlying filter optimization problem reduces to an eigenvalue problem and has a tractable solution. An encoding technique that enhances control over cardinality and uniqueness is also presented. Experimental results obtained for the problem of time series subsequence matching establish the merits of the proposed algorithm.
Pramod K. Vemulapalli, Vishal Monga, Sean Brennan 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Compact Video Fingerprinting via Structural Graphical Models
abstract
Much previous work in video fingerprinting has focused on robustness and security issues, but the compactness requirement, i.e., the hash should be of a short length with acceptable robustness and discriminability, continues to be a significant practical challenge. In this paper, we propose a video fingerprinting method with explicit attention on compactness. First, we develop a new graphical representation of the video which reduces temporal redundancies and makes robust feature extraction much more economical. Second, a randomized adaptive quantizer is proposed to further decrease the final hash length while maintaining acceptable detection performance in terms of receiver operating characteristics (ROCs). Experimental results reveal that the proposed method offers a more favorable robustness versus discriminability tradeoff over the state of the art particularly when the bit budget of the video fingerprint is low.
Mu Li 0002, Vishal Monga
IEEE Trans. Inf. Forensics Secur.2
2012 Robust video fingerprinting via structural graphical models
abstract
Applications of video fingerprinting range from traditional video retrieval and authentication to the more recent problem of anti-piracy search brought about by the emergence of video websites such as Youtube. Video fingerprints offer the potential of identifying in a robust and scalable manner - illegal or undesirable uploads of copyrighted video content. The principal challenge in video fingerprinting is to extract reduced dimensionality descriptors that can withstand incidental spatial and temporal distortions to the video while still allowing the discrimination of distinct videos. To address this fundamental problem, we propose to first represent a video as a graphical structure which can encode temporal relationships between video shots that are crucial to uniquely identifying the video. Next, we leverage ideas from graph theory, namely the normalized cuts graph partitioning method to divide the video representation into sub-graphs. Robust dimensionality reduction applied to these sub-graphs yields the final video hash/fingerprint. Experimental results in the form of receiver operating characteristic (ROC) curves on video databases acquired from YouTube reveal that the proposed video fingerprinting can enable a much more favorable robustness vs. discriminability trade-off over state-of-the art algorithms in video hashing.
Mu Li 0002, Vishal Monga
ICIP2
2012 Discriminative graphical models for sparsity-based hyperspectral target detection
abstract
The inherent discriminative capability of sparse representations has been exploited recently for hyperspectral target detection. This approach relies on the observation that the spectral signature of a pixel can be represented as a linear combination of a few training spectra drawn from both target and background classes. The sparse representation corresponding to a given test spectrum captures class-specific discriminative information crucial for detection tasks. Spatio-spectral information has also been introduced into this framework via a joint sparsity model that simultaneously solves for the sparse features for a group of spatially local pixels, since such pixels are highly likely to have similar spectral characteristics. In this paper, we propose a probabilistic graphical model framework that can explicitly learn the class conditional correlations between these distinct sparse representations corresponding to different pixels in a spatial neighborhood. Simulation results show that the proposed algorithm outperforms classical hyperspectral target detection algorithms as well as support vector machines.
Umamahesh Srinivas, Yi Chen 0014, Vishal Monga, Nasser M. Nasrabadi, Trac D. Tran
IGARSS3
2012 Capacity Analysis For Orthogonal Halftone Orientation Modulation Channels
abstract
Halftone dot orientation modulation has recently been proposed as a method for data hiding in printed images. Extraction of data embedded with halftone orientation modulation is accomplished by computing, from the scanned hardcopy image, detection statistics that uniquely identify the embedded orientation. From a communications perspective, this data hiding setup forms an interesting class of channels with dot orientation as input and a vector of statistics as the output. This paper derives capacity expressions for these channels that allow for numerical evaluation of the capacity. Results provide significant insight for orientation modulation based print-scan resilient data hiding: the capacity varies significantly as a function of the image graylevel and experimentally observed error free data rates closely mirror the variation in capacity.
Orhan Bulan, Vishal Monga, Gaurav Sharma 0001
IEEE Trans. Image Process.2
2012 Robust Video Hashing via Multilinear Subspace Projections
abstract
The goal of video hashing is to design hash functions that summarize videos by short fingerprints or hashes. While traditional applications of video hashing lie in database searches and content authentication, the emergence of websites such as YouTube and DailyMotion poses a challenging problem of anti-piracy video search. That is, hashes or fingerprints of an original video (provided to YouTube by the content owner) must be matched against those uploaded to YouTube by users to identify instances of "illegal" or undesirable uploads. Because the uploaded videos invariably differ from the original in their digital representation (owing to incidental or malicious distortions), robust video hashes are desired. We model videos as order-3 tensors and use multilinear subspace projections, such as a reduced rank parallel factor analysis (PARAFAC) to construct video hashes. We observe that, unlike most standard descriptors of video content, tensor-based subspace projections can offer excellent robustness while effectively capturing the spatio-temporal essence of the video for discriminability. We introduce randomization in the hash function by dividing the video into (secret key based) pseudo-randomly selected overlapping sub-cubes to prevent against intentional guessing and forgery. Detection theoretic analysis of the proposed hash-based video identification is presented, where we derive analytical approximations for error probabilities. Remarkably, these theoretic error estimates closely mimic empirically observed error probability for our hash algorithm. Furthermore, experimental receiver operating characteristic (ROC) curves reveal that the proposed tensor-based video hash exhibits enhanced robustness against both spatial and temporal video distortions over state-of-the-art video hashing techniques.
Mu Li 0002, Vishal Monga
IEEE Trans. Image Process.2
2012 Design and Optimization of Color Lookup Tables on a Simplex Topology
abstract
An important computational problem in color imaging is the design of color transforms that map color between devices or from a device-dependent space (e.g., RGB/CMYK) to a device-independent space (e.g., CIELAB) and vice versa. Real-time processing constraints entail that such nonlinear color transforms be implemented using multidimensional lookup tables (LUTs). Furthermore, relatively sparse LUTs (with efficient interpolation) are employed in practice because of storage and memory constraints. This paper presents a principled design methodology rooted in constrained convex optimization to design color LUTs on a simplex topology. The use of n simplexes, i.e., simplexes in n dimensions, as opposed to traditional lattices, recently has been of great interest in color LUT design for simplex topologies that allow both more analytically tractable formulations and greater efficiency in the LUT. In this framework of n-simplex interpolation, our central contribution is to develop an elegant iterative algorithm that jointly optimizes the placement of nodes of the color LUT and the output values at those nodes to minimize interpolation error in an expected sense. This is in contrast to existing work, which exclusively designs either node locations or the output values. We also develop new analytical results for the problem of node location optimization, which reduces to constrained optimization of a large but sparse interpolation matrix in our framework. We evaluate our n -simplex color LUTs against the state-of-the-art lattice (e.g., International Color Consortium profiles) and simplex-based techniques for approximating two representative multidimensional color transforms that characterize a CMYK xerographic printer and an RGB scanner, respectively. The results show that color LUTs designed on simplexes offer very significant benefits over traditional lattice-based alternatives in improving color transform accuracy even with a much smaller number of nodes.
Vishal Monga, Raja Bala, Xuan Mo
IEEE Trans. Image Process.1
2011 Scalable robust hypothesis tests using graphical models
abstract
Traditional binary hypothesis testing relies on the precise knowledge of the probability density of an observed random vector conditioned on each hypothesis. However, for many applications, these densities can only be approximated due to limited training data or dynamic changes affecting the observed signal. A classical approach to handle such scenarios of imprecise knowledge is via minimax robust hypothesis testing (RHT), where a test is designed to minimize the worst case performance for all models in the vicinity of the approximated imprecise density. Despite the promise of RHT for robust classification problems, its applications have remained rather limited because RHT in its native form does not scale gracefully with the dimension of the observed random vector. In this paper, we use approximations via probabilistic graphical models, in particular block-tree graphs, to enable computationally tractable algorithms for realizing RHT on high-dimensional data. We quantify the reductions in computational complexity. Experimental results on simulated data and a target recognition problem show minimal loss over a true RHT.
Divyanshu Vats, Vishal Monga, Umamahesh Srinivas, José M. F. Moura
ICASSP2
2011 Automatic target recognition using discriminative graphical models
abstract
Of recent interest in automatic target recognition (ATR) is the problem of combining the merits of multiple classifiers. This is commonly done by “fusing” the soft-outputs of several classifiers into making a single decision. We observe that the improvement in recognition rates afforded by these approaches is due to the complementary yet correlated information captured by different features/signal representations that these individual classifiers employ. We present the use of probabilistic graphical models in modeling and capturing feature dependencies that are crucial for target classification. In particular, we develop a two-stage target recognition framework that combines the merits of distinct and sparse signal representations with discriminatively learnt graphical models. The first stage designs multiple projections yielding M >; 1 sparse representations, while the second stage models each individual representation using graphs and combines these initially disjoint and simple graphical models into a thicker probabilistic graphical model. Experimental results show that our approach outperforms state-of-the art target classification techniques in terms of recognition rates. The use of graphical models is particularly meritorious when feature dimensionality is high and training is limited - a commonly observed constraint in synthetic aperture radar (SAR) imagery based target recognition.
Umamahesh Srinivas, Vishal Monga, Raghu G. Raj
ICIP2
2011 Desynchronization resilient video fingerprinting via randomized, low-rank tensor approximations
abstract
The problem of summarizing videos by short fingerprints or hashes has garnered significant attention recently. While traditional applications of video hashing lie in database search and content authentication, the emergence of websites such as YouTube and DailyMotion poses a challenging problem of anti-piracy video search. That is, hashes or fingerprints of an original video (provided to YouTube by the content owner) must be matched against those uploaded to YouTube by users to identify instances of “illegal” or undesirable uploads. Because the uploaded videos invariably differ from the original in their digital representation (owing to incidental or malicious distortions), robust video hashes are desired. In this paper, we model videos as order-3 tensors and use multilinear subspace projections, such as a reduced rank parallel factor analysis (PARAFAC) to construct video hashes. We observe that unlike most standard descriptors of video content, tensor based subspace projections can offer excellent robustness while effectively capturing the spatio-temporal essence of the video for discriminability. We further randomize the construction of the hash by dividing the video into randomly selected overlapping sub-cubes to prevent against intentional guessing and forgery. The most significant gains are seen for the difficult attacks of spatial (e.g. geometric) as well as temporal (random frame dropping) desynchronization. Experimental validation is provided in the form of ROC curves and we further perform detection-theoretic analysis which closely mimics empirically observed probability of error.
Mu Li 0002, Vishal Monga
MMSP2
2010 Algorithms for color look-up-table (LUT) design via joint optimization of node locations and output values
abstract
Real-time processing constraints entail that non-linear color transforms be implemented using multi-dimensional look-up-tables (LUT). Further, relatively sparse LUTs (with efficient interpolation) are employed in practice because of storage and memory constraints. Much research has been devoted towards optimizing “nodes” (or equivalently partitioning the input color space) of this color LUT based on the curvature of the color transform to be processed through the LUT. Likewise, for a given LUT structure, the optimization of transform output values has been suggested so as to minimize interpolation error in an expected sense even if the values stored in the LUT do not agree with true transform output values. This paper presents a principled algorithmic approach to combine the merits of these two complementary techniques. The error (cost) function does not exhibit joint convexity over the multidimensional variable sets of node locations and corresponding output values which makes this optimization particularly challenging. The paper makes two significant contributions: 1.) for the case of simplex interpolation, a cost function is formulated that exhibits separable convexity in its arguments and enables an efficient alternating convex optimization algorithm, and 2.) in the aforementioned framework, for fixed node outputs, the optimization of node locations is split into a primary and an auxiliary optimization, which greatly improves the quality of the solution over traditional alternatives where node locations are directly optimized. Preliminary experiments show remarkable improvements in color transform accuracy over what is obtained by individually optimizing just the node locations or output values.
Vishal Monga, Raja Bala
ICASSP1
2010 A learning framework for robust hashing of face images
abstract
Robust image hashing has been actively researched over the last decade with varied applications in image content authentication and identification under distortions. In the existing literature on robust image hashing, hash algorithms are ignorant of the class of images being hashed. There are however significant application domains such as that of face image hashing where apriori knowledge of the image class as well as permissible distortions can benefit hash algorithm design. In this paper, we present a two stage cascade of dimensionality reduction constructs for face image hashing. The first stage aims to project the face image to a space where geometric distortions manifest approximately as additive noise. For this purpose, we use the non-negative matrix approximations based hash vector developed by Monga et al. which is known to possess excellent geometric attack robustness. In the second stage, we employ oriented principal component analysis (OPCA) based on estimating signal as well as noise statistics in a learning phase and deriving a projection that mitigates the effect of noise. We obtain both experimentally based ROC curves as well as analytical ones via a detection theoretic analysis of the proposed framework. The ROC curves reveal clearly that incorporating such a learning phase greatly reduces error probabilities.
Kamil Senel, Mehmet Kivanç Mihçak, Vishal Monga
ICIP3
2010 Orientation Modulation for Data Hiding in Clustered-Dot Halftone Prints
abstract
We present a new framework for data hiding in images printed with clustered dot halftones. Our application scenario, like other hardcopy embedding methods, encounters fundamental challenges due to extreme bilevel quantization inherent in halftoning, the stringent requirements of image fidelity, and other unavoidable printing and scanning distortions. To overcome these challenges, while still allowing for automated extraction of the embedded data and a high embedding capacity, we propose a number of innovations. First, we perform the embedding jointly with the halftoning by employing an analytical halftone threshold function that allows steering of the halftone spot orientation within each halftone cell based upon embedded data. In this process, image fidelity is emphasized and, if necessary, the capability to recover individual data values is sacrificed resulting in unavoidable erasures and errors. To overcome these and other sources of errors, we propose a suitable data detection and error control methodology based upon a statistical representation for the print-scan channel that effectively models the channel dependence upon the cover image gray-level. To combat the geometric distortion inherent in the print-scan process, we exploit the periodic halftone structure to recover from global scaling and rotation and propose a novel decision directed synchronization technique that counters locally varying printing distortion. Experimental results demonstrate the power of the proposed framework: we achieve high operational rates while preserving halftone image quality.
Orhan Bulan, Gaurav Sharma 0001, Vishal Monga
IEEE Trans. Image Process.3
2009 Meta-classifiers for multimodal document classification
abstract
This paper proposes learning algorithms for the problem of multimodal document classification. Specifically, we develop classifiers that automatically assign documents to categories by exploiting features from both text as well as image content. In particular, we use meta-classifiers that combine state-of-the-art text and image based classifiers into making joint decisions. The two meta classifiers we choose are based on support vector machines and Adaboost. Experiments on real-world databases from Wikipedia demonstrate the benefits of a joint exploitation of these modalities.
Scott Deeann Chen, Vishal Monga, Pierre Moulin
MMSP2
2008 On the capacity of orientation modulation halftone channels
abstract
Clustered-dot halftones are extensively utilized in hardcopy printing. Modulation of the dot orientation in these halftones offers an avenue for data embedding which has been exploited in a number of different methods. We consider the capacity of these channels, modeling them as binary orientation input channels with vector valued output detection statistics. We derive upper bounds on the capacity for three channel conditional distributions corresponding to sub-Gaussian, Gaussian and super-Gaussian distributions. Using experimentally estimated channel parameters our bounds reveal that channel capacity has noticeable variations as a function of gray level. Highlights, shadows and mid-tones offer negligible capacity, on the contrary the regions between highlights and mid-tones or shadows and mid-tones offer high capacity for data embedding.
Orhan Bulan, Gaurav Sharma 0001, Vishal Monga
ICASSP3
2008 Adaptive decoding for halftone orientation-based data hiding
abstract
Halftone image watermarking techniques that allow automated extraction of the embedded watermark data are useful in a variety of document security and workflow applications. The print-scan process inherent in these applications introduces distortions whose characteristics exhibit a strong spatial dependence on the cover image in which the data is embedded. In this paper, we demonstrate that a characterization of this channel dependence and the use of error correction coding that exploits this dependence via image adaptive decoding offers a significant performance gain. We show this advantage in the specific context of a high rate data embedding method that utilizes orientation modulation for data embedding in clustered dot halftones and moment based detection at the receiver. Channel coding for this scenario utilizing convolutional codes and Repeat Accumulate (RA) codes highlights the advantage of the proposed adaptive decoding methodology. The performance for the RA codes with adaptive decoding also reveals that the resulting system has significantly higher operational rates than prior schemes.
Orhan Bulan, Gaurav Sharma 0001, Vishal Monga
ICIP3
2007 Robust and Secure Image Hashing via Non-Negative Matrix Factorizations
abstract
Abstract—In this paper, we propose the use of non-negative matrix factorization (NMF) for image hashing. In particular, we view images as matrices and the goal of hashing as a randomized dimensionality reduction that retains the essence of the original image matrix while preventing intentional attacks of guessing and forgery. Our work is motivated by the fact that standard-rank reduction techniques, such as QR and singular value decomposition, produce low-rank bases which do not respect the structure (i.e., non-negativity for images) of the original data. We observe that NMFs have two very desirable properties for secure image hashing applications: 1) The additivity property resulting from the non-negativity constraints results in bases that capture local components of the image, thereby significantly reducing misclassification and 2) the effect of geometric attacks on images in the spatial domain manifests (approximately) as independent identically distributed noise on NMF vectors, allowing the design of detectors that are both computationally simple and, at the same time, optimal in the sense of minimizing error probabilities. Receiver operating characteristics analysis over a large image database reveals that the proposed algorithms significantly outperform existing approaches for image hashing. Index Terms—Authentication, matrix approximations, robust hashing. I.
Vishal Monga, Mehmet Kivanç Mihçak
IEEE Trans. Inf. Forensics Secur.1
2007 Design of Tone-Dependent Color-Error Diffusion Halftoning Systems
abstract
Grayscale error diffusion introduces nonlinear distortion (directional artifacts and false textures), linear distortion (sharpening), and additive noise. Tone-dependent error diffusion (TDED) reduces these artifacts by controlling the diffusion of quantization errors based on the input graylevel. We present an extension of TDED to color. In color-error diffusion, which color to render becomes a major concern in addition to finding optimal dot patterns. We propose a visually meaningful scheme to train input-level (or tone-) dependent color-error filters. Our design approach employs a Neugebauer printer model and a color human visual system model that takes into account spatial considerations in color reproduction. The resulting halftones overcome several traditional error-diffusion artifacts and achieve significantly greater accuracy in color rendition.
Vishal Monga, Niranjan Damera-Venkata, Brian L. Evans
IEEE Trans. Image Process.1
2006 Robust Image Hashing Via Non-Negative Matrix Factorizations
abstract
In this paper, we propose the use of non-negative matrix factorization (NMF) for robust image hashing. In particular, we view images as matrices and the goal of hashing as a randomized dimensionality reduction that retains the essence of the original image matrix while preventing against intentional attacks of guessing and forgery. Our work is motivated by the fact that standard-rank reduction techniques such as the QR, and singular value decomposition (SVD), produce low rank bases which do not respect the structure (i.e. non-negativity for images) of the original data. We observe that NMFs have two very desirable properties for secure image hashing applications: 1) The additivity property resulting from the non-negativity constraints results in bases that capture local characteristics of the image, thereby significantly reducing misclassification, and 2) the effect of geometric attacks on images in the spatial domain manifests (approximately) as independent identically distributed noise on NMF vectors, allowing design of detectors that are both computationally simple and at the same time optimal in the sense of minimizing error probabilities. ROC (receiver operating characteristics) analysis over a large image database reveals that the proposed algorithms significantly outperform existing approaches for robust image hashing
Vishal Monga, Mehmet Kivanç Mihçak
ICASSP (2)1
2006 Geometrically Invariant Image Watermarking via Robust Perceptual Hashes
abstract
We consider the problem of blind watermark detection from an image, when the image has undergone a geometric attack. Existing schemes for solving this problem involve insertion of periodic marks or templates, invariant properties of transforms, and feature point based techniques. A major deficiency of these approaches is either a security leak (owing to redundancy in the mark) and/or poor watermark detection performance esp. under lossy geometric attacks, e.g. cropping. In this paper, we introduce a novel watermark synchronization paradigm by using perceptual image hashes. We exploit the fact that a perceptual hash acts as a robust as well as randomized digest of image content, and use it to embed a synchronization mark into the image. Unlike existing synchronization schemes, our synchronization mark has favorable security and robustness to adversarial attacks and hence, cannot be easily removed. Experimental results confirm that our proposed algorithm has significantly lower watermark detection error probabilities than existing methods.
Oztan Harmanci, Vishal Monga, Mehmet Kivanç Mihçak
ICIP2
2006 A clustering based approach to perceptual image hashing
abstract
A perceptual image hash function maps an image to a short binary string based on an image's appearance to the human eye. Perceptual image hashing is useful in image databases, watermarking, and authentication. In this paper, we decouple image hashing into feature extraction (intermediate hash) followed by data clustering (final hash). For any perceptually significant feature extractor, we propose a polynomial-time heuristic clustering algorithm that automatically determines the final hash length needed to satisfy a specified distortion. We prove that the decision version of our clustering problem is NP complete. Based on the proposed algorithm, we develop two variations to facilitate perceptual robustness versus fragility tradeoffs. We validate the perceptual significance of our hash by testing under Stirmark attacks. Finally, we develop randomized clustering algorithms for the purposes of secure image hashing.
Vishal Monga, Arindam Banerjee 0001, Brian L. Evans
IEEE Trans. Inf. Forensics Secur.1
2006 Perceptual Image Hashing Via Feature Points: Performance Evaluation and Tradeoffs
abstract
We propose an image hashing paradigm using visually significant feature points. The feature points should be largely invariant under perceptually insignificant distortions. To satisfy this, we propose an iterative feature detector to extract significant geometry preserving feature points. We apply probabilistic quantization on the derived features to introduce randomness, which, in turn, reduces vulnerability to adversarial attacks. The proposed hash algorithm withstands standard benchmark (e.g., Stirmark) attacks, including compression, geometric distortions of scaling and small-angle rotation, and common signal-processing operations. Content changing (malicious) manipulations of image data are also accurately detected. Detailed statistical analysis in the form of receiver operating characteristic (ROC) curves is presented and reveals the success of the proposed scheme in achieving perceptual robustness while avoiding misclassification.
Vishal Monga, Brian L. Evans
IEEE Trans. Image Process.1
2005 Image Authentication Under Geometric Attacks Via Structure Matching
abstract
Surviving geometric attacks in image authentication is considered to be of great importance. This is because of the vulnerability of classical watermarking and digital signature based schemes to geometric image manipulations, particularly local geometric attacks. In this paper, we present a general framework for image content authentication using salient feature points. We first develop an iterative feature detector based on an explicit modeling of the human visual system. Then, we compare features from two images by developing a generalized Hausdorff distance measure. The use of such a distance measure is crucial to the robustness of the scheme, and accounts for feature detector failure or occlusion, which previously proposed methods do not address. The proposed algorithm withstands standard benchmark (e.g. Stirmark) attacks including compression, common signal processing operations, global as well as local geometric transformations, and even hard to model distortions such as print and scan. Content changing (malicious) manipulations of image data are also accurately detected
Vishal Monga, Divyanshu Vats, Brian L. Evans
ICME1
2005 Two-dimensional transforms for device color correction and calibration
abstract
Color device calibration is traditionally performed using one-dimensional (1-D) per-channel tone-response corrections (TRCs). While 1-D TRCs are attractive in view of their low implementation complexity and efficient real-time processing of color images, their use severely restricts the degree of control that can be exercised along various device axes. A typical example is that per separation (or per-channel), TRCs in a printer can be used to either ensure gray balance along the C = M = Y axis or to provide a linear response in delta-E units along each of the individual (C, M, and Y) axis, but not both. This paper proposes a novel two-dimensional color correction architecture that enables much greater control over the device color gamut with a modest increase in implementation cost. Results show significant improvement in calibration accuracy and stability when compared to traditional 1-D calibration. Superior cost quality tradeoffs (over 1-D methods) are also achieved for emulation of one color device on another.
Raja Bala, Gaurav Sharma 0001, Vishal Monga, Jean-Pierre Van de Capelle
IEEE Trans. Image Process.3
2005 Hardcopy image barcodes via block-error diffusion
abstract
Error diffusion halftoning is a popular method of producing frequency modulated (FM) halftones for printing and display. FM halftoning fixes the dot size (e.g., to one pixel in conventional error diffusion) and varies the dot frequency according to the intensity of the original grayscale image. We generalize error diffusion to produce FM halftones with user-controlled dot size and shape by using block quantization and block filtering. As a key application, we show how block-error diffusion may be applied to embed information in hardcopy using dot shape modulation. We enable the encoding and subsequent decoding of information embedded in the hardcopy version of continuous-tone base images. The encoding-decoding process is modeled by robust data transmission through a noisy print-scan channel that is explicitly modeled. We refer to the encoded printed version as an image barcode due to its high information capacity that differentiates it from common hardcopy watermarks. The encoding/halftoning strategy is based on a modified version of block-error diffusion. Encoder stability, image quality versus information capacity tradeoffs, and decoding issues with and without explicit knowledge of the base image are discussed.
Niranjan Damera-Venkata, Jonathan Yen, Vishal Monga, Brian L. Evans
IEEE Trans. Image Process.3
2004 Tone dependent color error diffusion
abstract
Conventional grayscale error diffusion halftoning produces worms and other objectionable artifacts. Tone dependent error diffusion (Li, P. and Allebach, J.P. Proc. SPIE Color Imaging, vol.4663, p.310-21, 2002) reduces these artifacts by controlling the diffusion of quantization errors based on the input graylevel. Li and Allebach designed error filter weights and thresholds for each (input) graylevel with optimization based on a human visual system (HVS) model. We extend tone dependent error diffusion to color. In color error diffusion, what color to render becomes a major concern in addition to finding optimal dot patterns. We present a visually optimum design approach for input level (tone) dependent error filters (for each color plane). The resulting halftones reduce traditional error diffusion artifacts and achieve greater accuracy in color rendition.
Vishal Monga, Brian L. Evans
ICASSP (3)1
2004 Robust perceptual image hashing using feature points
abstract
Perceptual image hashing maps an image to a fixed length binary string based on the image's appearance to the human eye, and has applications in image indexing, authentication, and watermarking. We present a general framework for perceptual image hashing using feature points. The feature points should be largely invariant under perceptually insignificant distortions. To satisfy this, we propose an iterative feature detector to extract significant geometry preserving feature points. We apply probabilistic quantization on the derived features to enhance perceptual robustness further. The proposed hash algorithm withstands standard benchmark (e.g. Stirmark) attacks including compression, geometric distortions of scaling and small angle rotation, and common signal processing operations. Content changing (malicious) manipulations of image data are also accurately detected.
Vishal Monga, Brian L. Evans
ICIP1
2003 Linear color-separable human visual system models for vector error diffusion halftoning
abstract
Image halftoning converts a high-resolution image to a low-resolution image, e.g., a 24-bit color image to a three-bit color image, for printing and display. Vector error diffusion captures correlation among color planes by using an error filter with matrix-valued coefficients. In optimizing vector error filters, Damera-Venkata and Evans (see IEEE Trans. Image Processing, vol.10, p.1552-65, Oct. 2001) transform the error image into an opponent color space where Euclidean distance has perceptual meaning. This letter evaluates color spaces for vector error filter optimization. In order of increasing quality, the color spaces are YIQ, YUV, opponent (by Poirson and Wandell, 1993), and linearized CIELab (by Flohr, Kolpatzik, Balasubramanian, Carrara, Bouman, and Allebach, 1993).
Vishal Monga, Wilson S. Geisler, Brian L. Evans
IEEE Signal Process. Lett.1