VLDB 2026 Research / reviewers in the wild / expert
Gaurav Sharma 0001
dblp:s/GauravSharma1
· DBLP profile ↗
106ranked-venue papers
14as first author
13since 2021 · last 2026
0000-0001-9735-9519ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 90 · 12 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Security and privacy · 4 · 1 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Computer networks · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Image-Text Knowledge Modeling for Unsupervised Multi-Scenario Person Re-IdentificationabstractWe propose unsupervised multi-scenario (UMS) person re-identification (ReID) as a new task that expands ReID across diverse scenarios (cross-resolution, clothing change, etc.) within a single coherent framework. To tackle UMS-ReID, we introduce image-text knowledge modeling (ITKM) -- a three-stage framework that effectively exploits the representational power of vision-language models. We start with a pre-trained CLIP model with an image encoder and a text encoder. In Stage I, we introduce a scenario embedding in the image encoder and fine-tune the encoder to adaptively leverage knowledge from multiple scenarios. In Stage II, we optimize a set of learned text embeddings to associate with pseudo-labels from Stage I and introduce a multi-scenario separation loss to increase the divergence between inter-scenario text representations. In Stage III, we first introduce cluster-level and instance-level heterogeneous matching modules to obtain reliable heterogeneous positive pairs (e.g., a visible image and an infrared image of the same person) within each scenario. Next, we propose a dynamic text representation update strategy to maintain consistency between text and image supervision signals. Experimental results across multiple scenarios demonstrate the superiority and generalizability of ITKM; it not only outperforms existing scenario-specific methods but also enhances overall performance by integrating knowledge from multiple scenarios. Zhiqi Pang, Lingling Zhao, Yang Liu 0006, Chunyu Wang 0002, Gaurav Sharma 0001 |
AAAI | 5 |
| 2025 | DHAG-DTA: Dynamic Hierarchical Affinity Graph Model for Drug-Target Binding Affinity PredictionabstractComputational methods for predicting drug-target binding affinity (DTA) are critical for large-scale screening of prospective therapeutic compounds during drug discovery. Deep neural networks (DNNs) have recently shown significant promise for DTA prediction. By leveraging available data for training, DNNs can expand the use of DTA prediction to situations where only sequence information is available for potential drug molecules and their targets, and there is no prior knowledge regarding the molecular geometric conformations. We propose DHAG-DTA, a general dynamic hierarchical affinity graph DNN approach, for DTA prediction using molecular sequence information and already known drug-target interactions. DHAG-DTA introduces a two-level hierarchical graph structure: at the upper level, interactions between drug and target molecules are represented via an affinity graph and at the lower level, embedded molecular graphs represent interactions within the individual molecules. This allows for integration of information from both inter and intra molecular interactions for DTA prediction, which has also been addressed in other recent independent work. The fundamental innovations introduced by DHAG-DTA include: (a) a single overall hierarchical graph that allows better assimilation of information during the learning process compared with loosely-coupled individual graphs, (b) dynamic determination of the affinity graph structure via the introduction of unlabeled edges and a maximum entropy criterion for active edge selection, (c) skip connections in the DNN for fusing intra and inter molecular information, and (d) fusion of both model-based and similarity-based feature embeddings to get robust embeddings of unseen molecules. Experimental results on two common benchmark datasets demonstrate that DHAG-DTA outperforms other existing models on multiple evaluation metrics, achieving state-of-the-art performance. Cheng Wang 0049, Yang Liu 0006, Shitao Song, Gaurav Sharma 0001, Maozu Guo 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2025 | Joint Augmentation and Part Learning for Unsupervised Clothing Change Person Re-IdentificationabstractClothing change person re-identification (CC-ReID) is a crucial task in intelligent surveillance, aiming to match images of the same person wearing different clothing. Promising performance in existing CC-ReID methods is achieved at the cost of labor-intensive manual annotation of identity labels. While some researchers have explored unsupervised CC-ReID, these methods still depend on additional deep learning models for preprocessing. To eliminate the need for additional models and improve performance, we propose a joint augmentation and part learning (JAPL) framework that obtains clothing change positive pairs in an unsupervised fashion by synergistically combining augmentation-based invariant learning (AugIL) and part-based invariant learning (ParIL). AugIL first constructs clothing change pseudo-positive pairs and then encourages the model to focus on clothing-invariant information by enhancing feature consistency between the pseudo-positive pairs. ParIL beneficially encourages high similarity between inter-cluster clothing change positive pair using part images and a prediction sharpening loss. PartIL also introduces a soft consistency loss that promotes clothing-invariant feature learning by encouraging consistency of class vectors between the real features actually used for CC-ReID and the part features. Experimental results on multiple ReID datasets demonstrate that the proposed JAPL not only surpasses existing unsupervised methods but also achieves competitive performance compared to some supervised CC-ReID methods. Zhiqi Pang, Lingling Zhao, Yang Liu 0006, Gaurav Sharma 0001, Chunyu Wang 0002 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Robust Labeling and Invariance Modeling for Unsupervised Cross-Resolution Person Re-IdentificationabstractCross-resolution person re-identification (CR-ReID) aims to match low-resolution (LR) and high-resolution (HR) images of the same individual. To reduce the cost of manual annotation, existing unsupervised CR-ReID methods typically rely on cross-resolution fusion to obtain pseudo-labels and resolution-invariant features. However, the fusion process requires two encoders and a fusion module, which significantly increases computational complexity and reduces efficiency. To address this issue, we propose a robust labeling and invariance modeling (RLIM) framework, which utilizes a single encoder to tackle the unsupervised CR-ReID problem. To obtain pseudo-labels robust to resolution gaps, we develop cross-resolution robust labeling (CRL), which utilizes two clustering criteria to encourage cross-resolution positive pairs to cluster together and exploit the reliable relationships between images. We also introduce random texture augmentation (TexA) to enhance the model's robustness to noisy textures related to artifacts and backgrounds by randomly adjusting texture strength. During the optimization process, we introduce the resolution-cluster consistency loss, which promotes resolution-invariant feature learning by aligning inter-resolution distances with intra-cluster distances. Experimental results on multiple datasets demonstrate that RLIM not only surpasses existing unsupervised methods, but also achieves performance close to some supervised CR-ReID methods. Code is available at https://github.com/zqpang/RLIM. Zhiqi Pang, Lingling Zhao, Yang Liu 0006, Chunyu Wang 0002, Gaurav Sharma 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | SelfBC: Self Behavior Cloning for Offline Reinforcement LearningabstractPolicy constraint methods in offline reinforcement learning employ additional regularization techniques to constrain the discrepancy between the learned policy and the offline dataset. However, these methods tend to result in overly conservative policies that resemble the behavior policy, thus limiting their performance. We investigate this limitation and attribute it to the static nature of traditional constraints. In this paper, we propose a novel dynamic policy constraint that restricts the learned policy on the samples generated by the exponential moving average of previously learned policies. By integrating this self-constraint mechanism into off-policy methods, our method facilitates the learning of non-conservative policies while avoiding policy collapse in the offline setting. Theoretical results show that our approach results in a nearly monotonically improved reference policy. Extensive experiments on the D4RL MuJoCo domain demonstrate that our proposed method achieves state-of-the-art performance among the policy constraint methods. Shirong Liu, Chenjia Bai, Zixian Guo, Hao Zhang 0128, Gaurav Sharma 0001, Yang Liu 0006 |
ECAI | 5 |
| 2024 | Computational Trichromacy Reconstruction: Empowering the Color-Vision Deficient to Recognize Colors Using Augmented RealityabstractWe propose an assistive technology that helps individuals with Color Vision Deficiencies (CVD) to recognize/name colors. A dichromat’s color perception is a reduced two-dimensional (2D) subset of a normal trichromat’s three dimensional color (3D) perception, leading to confusion when visual stimuli that appear identical to the dichromat are referred to by different color names. Using our proposed system, CVD individuals can interactively induce distinct perceptual changes to originally confusing colors via a computational color space transformation. By combining their original 2D precepts for colors with the discriminative changes, a three dimensional color space is reconstructed, where the dichromat can learn to resolve color name confusions and accurately recognize colors. Our system is implemented as an Augmented Reality (AR) interface on smartphones, where users interactively control the rotation through swipe gestures and observe the induced color shifts in the camera view or in a displayed image. Through psychophysical experiments and a longitudinal user study, we demonstrate that such rotational color shifts have discriminative power (initially confusing colors become distinct under rotation) and exhibit structured perceptual shifts dichromats can learn with modest training. The AR App is also evaluated in two real-world scenarios (building with lego blocks and interpreting artistic works); users all report positive experience in using the App to recognize object colors that they otherwise could not. Yuhao Zhu 0001, Ethan Chen, Colin Hascup, Yukang Yan, Gaurav Sharma 0001 |
UIST | 5 |
| 2024 | Cross-Modality Hierarchical Clustering and Refinement for Unsupervised Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a challenging cross-modality image retrieval task. Compared to visible modality person re-identification that handles only the intra-modality discrepancy, VI-ReID suffers from an additional modality gap. Most existing VI-ReID methods achieve promising accuracy in a supervised setting, but the high annotation cost limits their scalability to real-world scenarios. Although a few unsupervised VI-ReID methods already exist, they typically rely on intra-modality initialization and cross-modality instance selection, despite the additional computational time required for intra-modality initialization. In this paper, we study the fully unsupervised VI-ReID problem and propose a novel cross-modality hierarchical clustering and refinement (CHCR) method by promoting modality-invariant feature learning and improving the reliability of pseudo-labels. Unlike conventional VI-ReID methods, CHCR does not rely on any manual identity annotation and intra-modality initialization. First, we design a simple and effective cross-modality clustering baseline that clusters between modalities. Then, to provide sufficient inter-modality positive sample pairs for modality-invariant feature learning, we propose a cross-modality hierarchical clustering algorithm to promote the clustering of inter-modality positive samples into the same cluster. In addition, we develop an inter-channel pseudo-label refinement algorithm to eliminate unreliable pseudo-labels by checking the clustering results of three channels in the visible modality. Extensive experiments demonstrate that CHCR outperforms state-of-the-art unsupervised methods and achieves performance competitive with many supervised methods. Zhiqi Pang, Chunyu Wang 0002, Lingling Zhao, Yang Liu 0006, Gaurav Sharma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Inter-Modality Similarity Learning for Unsupervised Multi-Modality Person Re-IdentificationabstractRGB (visible), near-infrared (NI), and thermal infrared (TI) imaging modalities are commonly combined for round-the-clock surveillance. We introduce a novel unsupervised multi-modality person re-identification (MM-ReID) task, which, based on an individual’s image in any one modality, seeks to identify matches in the other two modalities. Compared to prior MM-ReID problem formulations, unsupervised MM-ReID significantly reduces labeling cost and imaging constraints. To address the unsupervised MM-ReID task, we propose a novel inter-modality similarity learning (IMSL) framework consisting of four synergistic interconnected modules: modality mean clustering (MMC), multi-modality reliability estimation (MMRE), shape-based mutual reinforcement (SMR), and modality-aware invariant learning (MIL). MMC iterates with SMR and MIL in a mutually beneficial manner to provide pseudo-labels that are robust to modality gap. MMRE normalizes sample weights, mitigating the impact of noisy labels in the multi-modality setting. SMR emphasizes shape information to implicitly enhance the model’s robustness to the modality gap and is additionally guided by pseudo-labels provided by MMC to attend to identity-related details. MIL explicitly encourages learning of modality-invariant and identity-related features via contrastive feedback for the MMC module. Extensive experimental results on the multi-modality and cross-modality datasets demonstrate that IMSL provides substantial performance gains over existing methods. Code is made available at https://github.com/zqpang/IMSL. Zhiqi Pang, Lingling Zhao, Yang Liu 0006, Gaurav Sharma 0001, Chunyu Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Sentence Bag Graph Formulation for Biomedical Distant Supervision Relation ExtractionabstractWe introduce a novel graph-based framework for alleviating key challenges in distantly-supervised relation extraction and demonstrate its effectiveness in the challenging and important domain of biomedical data. Specifically, we propose a graph view of sentence bags referring to an entity pair, which enables message-passing based aggregation of information related to the entity pair over the sentence bag. The proposed framework alleviates the common problem of noisy labeling in distantly supervised relation extraction and also effectively incorporates inter-dependencies between sentences within a bag. Extensive experiments on two large-scale biomedical relation datasets and the widely utilized NYT dataset demonstrate that our proposed framework significantly outperforms the state-of-the-art methods for biomedical distant supervision relation extraction while also providing excellent performance for relation extraction in the general text mining domain. Hao Zhang 0128, Yang Liu 0006, Tianming Liang, Gaurav Sharma 0001, Maozu Guo 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | Optimized Modulation and Coding for Dual Modulated QR CodesabstractWe present optimized modulation and coding for the recently introduced dual modulated QR (DMQR) codes that extend traditional QR codes to carry additional secondary data in the orientation of elliptical dots that replace black modules in the barcode images. By dynamically adjusting the dot size, we realize gains in embedding strength for both the intensity modulation and the orientation modulation that carry the primary and secondary data, respectively. Furthermore, we develop a model for the coding channel for the secondary data that enables soft-decoding via 5G NR (new radio) codes already supported by mobile devices. The performance gains for the proposed optimized designs are characterized via theoretical analysis, simulations, and actual experiments using smartphone devices. The theoretical analysis and simulations inform our design choices for the modulation and coding, and the experiments characterize the overall improvement in performance for the optimized design over the prior unoptimized designs. Importantly, the optimized designs significantly increase usability of DMQR codes with commonly used QR code beautification that cannibalizes a portion of the barcode image area for the insertion of a logo or image. In experiments with a capture distance of 15 inches, the optimized designs increase the decoding success rates between 10% and 32% for the secondary data while also providing gains for primary data decoding at larger capture distances. When used with beautification in typical settings, the secondary message is decoded with a high success rate for the proposed optimized designs, whereas it invariably fails for the prior unoptimized designs. Irving Barron, Gaurav Sharma 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Distantly-Supervised Long-Tailed Relation Extraction Using Constraint GraphsabstractLabel noise and long-tailed distributions are two major challenges in distantly supervised relation extraction. Recent studies have shown great progress on denoising, but paid little attention to the problem of long-tailed relations. In this paper, we introduce a constraint graph to model the dependencies between relation labels. On top of that, we further propose a novel constraint graph-based relation extraction framework(CGRE) to handle the two challenges simultaneously. CGRE employs graph convolution networks to propagate information from data-rich relation nodes to data-poor relation nodes, and thus boosts the representation learning of long-tailed relations. To further improve the noise immunity, a constraint-aware attention module is designed in CGRE to integrate the constraint information. Extensive experimental results indicate that CGRE achieves significant improvements over the previous methods for both denoising and long-tailed relation extraction. Tianming Liang, Yang Liu 0006, Hao Zhang 0128, Gaurav Sharma 0001, Maozu Guo 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2021 | Dual Modulated QR Codes for Proximal Privacy and SecurityabstractThe ubiquitous presence of surveillance cameras severely compromises the security of private information (e.g. passwords) entered via a conventional keyboard interface in public places. We address this problem by proposing dual modulated QR (DMQR) codes, a novel QR code extension via which users can securely communicate private information in public places using their smartphones and a camera interface. Dual modulated QR codes use the same synchronization patterns and module geometry as conventional monochrome QR codes. Within each module, primary data is embedded using intensity modulation compatible with conventional QR code decoding. Specifically, depending on the bit to be embedded, a module is either left white or an elliptical black dot is placed within it. Additionally, for each module containing an elliptical dot, secondary data is embedded by orientation modulation; that is, by using different orientations for the elliptical dots. Because the orientation of the elliptical dots can only be reliably assessed when the barcodes are captured from a close distance, the secondary data provides "proximal privacy" and can be effectively used to communicate private information securely in public settings. Tests conducted using several alternative parameter settings demonstrate that the proposed DMQR codes are effective in meeting their objective- the secondary data can be accurately decoded for short capture distances (6 in.) but cannot be recovered from images captured over long distances (>12 in.). Furthermore, the proximal privacy can be adapted to application needs by varying the eccentricity of the elliptical dots used. Irving Barron, Hsin Jui Yeh, Karthik Dinesh, Gaurav Sharma 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Weakly-Supervised Vessel Detection in Ultra-Widefield Fundus Photography via Iterative Multi-Modal Registration and LearningabstractWe propose a deep-learning based annotation-efficient framework for vessel detection in ultra-widefield (UWF) fundus photography (FP) that does not require de novo labeled UWF FP vessel maps. Our approach utilizes concurrently captured UWF fluorescein angiography (FA) images, for which effective deep learning approaches have recently become available, and iterates between a multi-modal registration step and a weakly-supervised learning step. In the registration step, the UWF FA vessel maps detected with a pre-trained deep neural network (DNN) are registered with the UWF FP via parametric chamfer alignment. The warped vessel maps can be used as the tentative training data but inevitably contain incorrect (noisy) labels due to the differences between FA and FP modalities and the errors in the registration. In the learning step, a robust learning method is proposed to train DNNs with noisy labels. The detected FP vessel maps are used for the registration in the following iteration. The registration and the vessel detection benefit from each other and are progressively improved. Once trained, the UWF FP vessel detection DNN from the proposed approach allows FP vessel detection without requiring concurrently captured UWF FA images. We validate the proposed framework on a new UWF FP dataset, PRIME-FP20, and on existing narrow-field FP datasets. Experimental evaluation, using both pixel-wise metrics and the CAL metrics designed to provide better agreement with human assessment, shows that the proposed approach provides accurate vessel detection, without requiring manually labeled UWF FP training data. Li Ding 0009, Ajay Kuriyan, Rajeev S. Ramchandran, Charles C. Wykoff, Gaurav Sharma 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2020 | A Survey of Healthcare Internet of Things (HIoT): A Clinical PerspectiveabstractIn combination with current sociological trends, the maturing development of IoT devices is projected to revolutionize healthcare. A network of body-worn sensors, each with a unique ID, can collect health data that is orders-of-magnitude richer than what is available today from sporadic observations in clinical/hospital environments. When databased, analyzed, and compared against information from other individuals using data analytics, HIoT data enables the personalization and modernization of care with radical improvements in outcomes and reductions in cost. In this paper, we survey existing and emerging technologies that can enable this vision for the future of healthcare, particularly in the clinical practice of healthcare. Three main technology areas underlie the development of this field: (a) sensing, where there is an increased drive for miniaturization and power efficiency; (b) communications, where the enabling factors are ubiquitous connectivity, standardized protocols, and the wide availability of cloud infrastructure, and (c) data analytics and inference, where the availability of large amounts of data and computational resources is revolutionizing algorithms for individualizing inference and actions in health management. Throughout the paper, we use a case study to concretely illustrate the impact of these trends. We conclude our paper with a discussion of the emerging directions, open issues, and challenges. Hadi Habibzadeh, Karthik Dinesh, Omid Rajabi Shishvan, Andrew Boggio-Dandry, Gaurav Sharma 0001, Tolga Soyata |
IEEE Internet Things J. | 5 |
| 2020 | A Novel Deep Learning Pipeline for Retinal Vessel Detection In Fluorescein AngiographyabstractWhile recent advances in deep learning have significantly advanced the state of the art for vessel detection in color fundus (CF) images, the success for detecting vessels in fluorescein angiography (FA) has been stymied due to the lack of labeled ground truth datasets. We propose a novel pipeline to detect retinal vessels in FA images using deep neural networks (DNNs) that reduces the effort required for generating labeled ground truth data by combining two key components: cross-modality transfer and human-in-the-loop learning. The cross-modality transfer exploits concurrently captured CF and fundus FA images. Binary vessels maps are first detected from CF images with a pre-trained neural network and then are geometrically registered with and transferred to FA images via robust parametric chamfer alignment to a preliminary FA vessel detection obtained with an unsupervised technique. Using the transferred vessels as initial ground truth labels for deep learning, the human-in-the-loop approach progressively improves the quality of the ground truth labeling by iterating between deep-learning and labeling. The approach significantly reduces manual labeling effort while increasing engagement. We highlight several important considerations for the proposed methodology and validate the performance on three datasets. Experimental results demonstrate that the proposed pipeline significantly reduces the annotation effort and the resulting deep learning methods outperform prior existing FA vessel detection methods by a significant margin. A new public dataset, RECOVERY-FA19, is introduced that includes high-resolution ultra-widefield images and accurately labeled ground truth binary vessel maps. Li Ding 0009, Mohammad H. Bawany, Ajay Kuriyan, Rajeev S. Ramchandran, Charles C. Wykoff, Gaurav Sharma 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Color Control Functions for Multiprimary Displays - I: Robustness Analysis and Optimization FormulationsabstractColor management for a multiprimary display requires, as a fundamental step, the determination of a color control function (CCF) that specifies control values for reproducing each color in the display's gamut. Multiprimary displays offer alternative choices of control values for reproducing a color in the interior of the gamut and accordingly alternative choices of CCFs. Under ideal conditions, alternative CCFs render colors identically. However, deviations in the spectral distributions of the primaries and the diversity of cone sensitivities among observers impact alternative CCFs differently, and, in particular, make some CCFs prone to artifacts in rendered images. We develop a framework for analyzing robustness of CCFs for multiprimary displays against primary and observer variations, incorporating a common model of human color perception. Using the framework, we propose analytical and numerical approaches for determining robust CCFs. First, via analytical development, we: (a) demonstrate that linearity of the CCF in tristimulus space endows it with resilience to variations, particularly, linearity can ensure invariance of the gray axis, (b) construct an axially linear CCF that is defined by the property of linearity over constant chromaticity loci, and (c) obtain an analytical form for the axially linear CCF that demonstrates it is continuous but suffers from the limitation that it does not have continuous derivatives. Second, to overcome the limitation of the axially linear CCF, we motivate and develop two variational objective functions for optimization of multiprimary CCFs, the first aims to preserve color transitions in the presence of primary/observer variations and the second combines this objective with desirable invariance along the gray axis, by incorporating the axially linear CCF. A companion Part II paper, presents an algorithmic approach for numerically computing optimal CCFs for the two alternative variational objective functions proposed here and presents results comparing alternative CCFs for several different 4,5, and 6 primary designs. Carlos Eduardo Rodríguez-Pardo, Gaurav Sharma 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Color Control Functions for Multiprimary Displays - II: Variational Robustness OptimizationabstractIn a companion Part I paper, we presented a framework for analyzing robustness of color control functions (CCFs) for multiprimary displays against primary and observer variations and proposed a variational minimization for obtaining robust CCFs. The objective function proposed in the Part I paper combines two nonnegative terms that serve as useful figures of merit for quantitatively characterizing CCFs. The first term measures lack of smoothness of the CCFs and characterizes how well transitions in perceptual color space are preserved in the presence of the primary/observer variations. The second term measures deviation of the CCF, in the vicinity of the gray axis, from a specific axially linear CCF that provides perceptual invariance to the variations along the gray axis. In this paper, using calculus of variations, we develop an algorithm for numerically computing optimal CCFs under the proposed variational formulation. Using the proposed algorithm, we determine optimal CCFs for a several multiprimary display designs and assess and compare their performance against alternative approaches. The variationally optimal CCFs obtained using the proposed approach offer improvements over the alternatives, as assessed visually and via quantitative metrics measuring smoothness and invariance in the presence of primary variations. The relative improvements provided by the proposed CCF increase with increasing number of primaries. Carlos Eduardo Rodríguez-Pardo, Gaurav Sharma 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Creating a Multitrack Classical Music Performance Dataset for Multimodal Music Analysis: Challenges, Insights, and ApplicationsabstractWe introduce a dataset for facilitating audio-visual analysis of music performances. The dataset comprises 44 simple multi-instrument classical music pieces assembled from coordinated but separately recorded performances of individual tracks. For each piece, we provide the musical score in MIDI format, the audio recordings of the individual tracks, the audio and video recording of the assembled mixture, and ground-truth annotation files including frame-level and note-level transcriptions. We describe our methodology for the creation of the dataset, particularly highlighting our approaches to address the challenges involved in maintaining synchronization and expressiveness. We demonstrate the high quality of synchronization achieved with our proposed approach by comparing the dataset with existing widely used music audio datasets. We anticipate that the dataset will be useful for the development and evaluation of existing music information retrieval (MIR) tasks, as well as for novel multimodal tasks. We benchmark two existing MIR tasks (multipitch analysis and score-informed source separation) on the dataset and compare them with other existing music audio datasets. In addition, we consider two novel multimodal MIR tasks (visually informed multipitch analysis and polyphonic vibrato analysis) enabled by the dataset and provide evaluation measurements and baseline systems for future comparisons (from our recent work). Finally, we propose several emerging research directions that the dataset enables. Bochen Li, Xinzhao Liu, Karthik Dinesh, Zhiyao Duan, Gaurav Sharma 0001 |
IEEE Trans. Multim. | 5 |
| 2018 | Quantification of Longitudinal Changes in Retinal Vasculature from Wide-Field Fluorescein Angiography via a Novel Registration and Change Detection ApproachabstractWide-field fluorescein angiography (FA) images are commonly used in ophthalmology to assess longitudinal changes in retinal vasculature, specifically, non-perfusion. Current practice relies on manual qualitative comparisons between images taken at successive clinic visits, a few months apart. Objective quantitative assessments, although desirable for evaluating disease progression and treatment, are impractical to perform manually and challenging for image analysis because of the changes in the capture viewpoints and temporal imaging variations seen as the FA dye injection perfuses the retina. We propose a methodology for quantifying retinal non-perfusion by automated analysis of the FA images captured during successive clinical visits. Blood vessels are first detected in the image from each visit. The vascular networks in FA images are then precisely registered to obtain a co-aligned pair via parametric chamfer matching under polynomial transformation, a process that explicitly allows for increase or decrease in perfusion. Changes in perfusion are then quantified by identifying the common and distinct regions in co-aligned image pairs. The proposed framework is tested on FA images that are manually annotated by an ophthalmologist to provide ground truth binary vessel masks and to identify vasculature changes. Results indicate that the proposed method provides assessments of vasculature changes that are in good agreement with the ophthalmologist-provided annotations. Li Ding 0009, Ajay Kuriyan, Rajeev S. Ramchandran, Gaurav Sharma 0001 |
ICASSP | 4 |
| 2018 | Retinal Vessel Detection in Wide-Field Fluorescein Angiography with Deep Neural Networks: A Novel Training Data Generation ApproachabstractRetinal blood vessel detection is a crucial step in automatic retinal image analysis. Recently, deep neural networks have significantly advanced the state of the art for retinal blood vessel detection in color fundus (CF) images. Thus far, similar gains have not been seen in fluorescein angiography (FA) because the FA modality is entirely different from CF and annotated training data has not been available for FA imagery. We address retinal vessel detection in wide-field FA images with generative adversarial networks (GAN) via a novel approach for generating training data. Using a publicly available dataset that contains concurrently acquired pairs of CF and fundus FA images, vessel maps are detected in CF images via a pre-trained neural network and registered with fundus FA images via parametric chamfer matching to a preliminary FA vessel detection map. The co-aligned pairs of vessel maps (detected from CF images) and fundus FA images are used as ground truth labeled data for de novo training of a deep neural network for FA vessel detection. Specifically, we utilize adversarial learning to train a GAN where the generator learns to map FA images to binary vessel maps and the discriminator attempts to distinguish generated vs. ground-truth vessel maps. We highlight several important considerations for the proposed data generation methodology. The proposed method is validated on VAMpIRE dataset that contains high-resolution wide-field FA images and manual annotation of vessel segments. Experimental results demonstrate that the proposed method achieves an estimated ROC AUC of 0.9758. Li Ding 0009, Ajay Kuriyan, Rajeev S. Ramchandran, Gaurav Sharma 0001 |
ICIP | 4 |
| 2018 | Vehicle Tracking in Wide Area Motion Imagery via Stochastic Progressive Association Across Multiple FramesabstractVehicle tracking in Wide Area Motion Imagery (WAMI) relies on associating vehicle detections across multiple WAMI frames to form tracks corresponding to individual vehicles. The temporal window length, i.e., the number M of sequential frames, over which associations are collectively estimated poses a trade-off between accuracy and computational complexity. A larger M improves performance because the increased temporal context enables the use of motion models and allows occlusions and spurious detections to be handled better. The number of total hypotheses tracks, on the other hand, grows exponentially with increasing M, making larger values of M computationally challenging to tackle. In this paper, we introduce SPAAM an iterative approach that progressively grows M with each iteration to improve estimated tracks by exploiting the enlarged temporal context while keeping computation manageable through two novel approaches for pruning association hypotheses. First, guided by a road network, accurately co-registered to the WAMI frames, we disregard unlikely associations that do not agree with the road network. Second, as M is progressively enlarged at each iteration, the related increase in association hypotheses is limited by revisiting only the subset of association possibilities rendered open by stochastically determined dis-associations for the previous iteration. The stochastic disassociation at each iteration maintains each estimated association according to an estimated probability for confidence, obtained via a probabilistic model. Associations at each iteration are then estimated globally over the M frames by (approximately) solving a binary integer programming problem for selecting a set of compatible tracks. Vehicle tracking results obtained over test WAMI datasets indicate that our proposed approach provides significant performance improvements over state of the art alternatives. Ahmed S. Elliethy, Gaurav Sharma 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Visually informed multi-pitch analysis of string ensemblesabstractMulti-pitch analysis of polyphonic music requires estimating concurrent pitches (estimation) and organizing them into temporal streams according to their sound sources (streaming). This is challenging for approaches based on audio alone due to the polyphonic nature of the audio signals. Video of the performance, when available, can be useful to alleviate some of the difficulties. In this paper, we propose to detect the play/non-play (P/NP) activities from musical performance videos using optical flow analysis to help with audio-based multi-pitch analysis. Specifically, the detected P/NP activity provides a more accurate estimate of the instantaneous polyphony (i.e., the number of pitches at a time instant), and also helps with assigning pitch estimates to only active sound sources. As the first attempt towards audio-visual multi-pitch analysis of multi-instrument musical performances, we demonstrate the concept on 11 string ensembles. Experiments show a high overall P/NP detection accuracy of 85.3%, and a statistically significant improvement on both the multi-pitch estimation and streaming accuracy, under paired t-tests at a significance level of 0:01 in most cases. Karthik Dinesh, Bochen Li, Xinzhao Liu, Zhiyao Duan, Gaurav Sharma 0001 |
ICASSP | 5 |
| 2017 | Fusing structure from motion and lidar for dense accurate depth map estimationabstractWe present a novel framework for precisely estimating dense depth maps by combining 3D lidar scans with a set of uncalibrated camera RGB color images for the same scene. Rough estimates for 3D structure obtained using structure from motion (SfM) on the uncalibrated images are first co-registered with the lidar scan and then a precise alignment between the datasets is estimated by identifying correspondences between the captured images and reprojected images for individual cameras from the 3D lidar point clouds. The precise alignment is used to update both the camera geometry parameters for the images and the individual camera radial distortion estimates, thereby providing a 3D-to-2D transformation that accurately maps the 3D lidar scan onto the 2D image planes. The 3D to 2D map is then utilized to estimate a dense depth map for each image. Experimental results on two datasets that include independently acquired high-resolution color images and 3D point cloud datasets indicate the utility of the framework. The proposed approach offers significant improvements on results obtained with SfM alone. Li Ding 0009, Gaurav Sharma 0001 |
ICASSP | 2 |
| 2017 | Vehicle tracking in Wide area motion imagery: A facility location motivated combinatorial approachabstractWe propose a practical combinatorial approach for addressing the challenging task of vehicle tracking in Wide area motion imagery (WAMI) by leveraging a pixel accurate co-registered vector road-map. Specifically, guided by the co-registered road network, we obtain a sparse trellis graph linking each vehicle detection (VD) in a WAMI frame with a road reachable VD in the next frame, which then allows us to enumerate all possible hypotheses tracks. The globally optimal selection of tracks over a K frame window is then formulated as a minimum cost K capacitated facility location problem, where each hypothesis track is allocated VDs in each frame and assigned a cost that models desirable properties of vehicle track. Computation of optimized combinations of jointly feasible tracks that minimize the total cost for the tracks becomes feasible in our formulation for moderate values of K by utilizing available solvers for the facility location problem. The approach automatically selects the optimal number of tracks and provides flexibility in defining costs for tracks globally across the K frames. Vehicle tracking results obtained over test WAMI datasets indicate that our proposed method provides significant better performance than two other state of the art alternatives. Ahmed S. Elliethy, Gaurav Sharma 0001 |
ICASSP | 2 |
| 2017 | See and listen: Score-informed association of sound tracks to players in chamber music performance videosabstractBoth audio and visual aspects of a musical performance, especially their association, are important for expressing players' ideas and for engaging the audience. In this paper, we present a framework for combining audio and video analyses of multi-instrument chamber music performances to associate players in the video to the individual separated instrument sources from the audio, in a score-informed fashion. The instrument sources are first separated using a score-informed source separation techniques. The individual sources are then associated with different players in the video by correlating the onset instants of their aligned score tracks with the players' motion detected using optical flow. Experiments on 19 musical pieces with varying polyphony show that the proposed method obtains the correct association for 17 pieces, and an accuracy of 89.2% of the association of all individual tracks. The approach enables novel music enjoyment experiences by allowing users to target an audio source by clicking on the player in the video to separate/enhance it. Bochen Li, Karthik Dinesh, Zhiyao Duan, Gaurav Sharma 0001 |
ICASSP | 4 |
| 2017 | In-situ calibration of accelerometers in body-worn sensors using quiescent gravityabstractAs the cost, size and power required by sensor devices decrease, an increasing range of applications are possible. We focus on the application of tracking accelerometer data from body worn sensors over long durations in health monitoring applications. Body worn sensors must be compact for the convenience of the patient and low power to support extended operation, but these benefits can come at the cost of accuracy. To mitigate this loss of accuracy we first examine the ability of calibration to improve accuracy. Next we propose an in-situ method for calibrating the accelerometers using the quiescent acceleration due to gravity as a calibration signal. This in-situ calibration uses only the stored data from the duration the sensor is worn by the patient and does not require any extra procedures or measurements from the physical sensor devices. Compared to a manual three-axis calibration technique, the proposed calibration can be applied to data without the need for specific calibration procedures or even access to the original sensor and provides comparable or better accuracy. Andrew Nadeau, Karthik Dinesh, Gaurav Sharma 0001, Mulin Xiong |
ICASSP | 3 |
| 2017 | 3D georegistration of wide area motion imagery by combining SFM and chamfer alignment of vehicle detections to vector roadmapsabstractWe propose a novel framework for accurate 3D georegistration of wide area motion imagery (WAMI), which is a challenging problem because parametric transformations are insufficient for aligning WAMI image frames to a georeferenced coordinate system in urban areas containing tall buildings and 3D structures. Using structure from motion (SfM) we estimate a 3D point cloud for the scene. Independently, we also compute a precise alignment between the roads in the WAMI frames and a georeferenced vector roadmap by detecting locations of moving vehicles and aligning these locations with the roads in the vector roadmap via parametric chamfer matching. The aligned vector roadmap then identifies corresponding pixels in the WAMI frames, which can be triangulated using the SfM camera parameters to obtain a set of sparse but georeferenced points in the SfM 3D coordinate frame that directly enable georegistration of the complete 3D scene point cloud via a similarity transform. The proposed methodology enables 3D georegistration of a sequence of WAMI frames using only georeferenced vector roadmaps, which are readily available, and without requiring independent georeferenced lidar scans that have been used in prior work. Our framework is validated on WAMI dataset including high resolution WAMI frames for the downtown Rochester, NY region. Experimental results demonstrate that the proposed framework produces an accurate georeferenced point cloud representation for the scene. Li Ding 0009, Ahmed S. Elliethy, Gaurav Sharma 0001 |
ICIP | 3 |
| 2017 | HazeRD: An outdoor scene dataset and benchmark for single image dehazingabstractIn this paper, a new dataset, HazeRD, is proposed for benchmarking dehazing algorithms under more realistic haze conditions. HazeRD contains fifteen real outdoor scenes, for each of which five different weather conditions are simulated. As opposed to prior datasets that made use of synthetically generated images or indoor images with unrealistic parameters for haze simulation, our outdoor dataset allows for more realistic simulation of haze with parameters that are physically realistic and justified by scattering theory. All images are of high resolution, typically six to eight megapixels. We test the performance of several state-of-the-art dehazing techniques on HazeRD. The results exhibit a significant difference among algorithms across the different datasets, reiterating the need for more realistic datasets such as ours and for more careful benchmarking of the methods. Yanfu Zhang, Li Ding 0009, Gaurav Sharma 0001 |
ICIP | 3 |
| 2017 | An Audio Watermark Designed for Efficient and Robust Resynchronization After Analog PlaybackabstractWe propose a spread spectrum (SS) audio watermark designed to withstand analog playback, including desynchronization caused by small differences between playback and recording rates. Desynchronization robustness relies on detecting short blocks of a magnitude-only watermark embedded in the frequency domain where the resolution of the SS chips can be reduced. Lost spreading gain due to the lower number of SS chips is compensated using blind dynamic time warping (DTW) detection (does not access original signal). DTW aligns sequences of blocks to improve robustness to interference while mitigating the vulnerabilities of long SS sequences to desynchronization. Results demonstrate that the proposed watermark survives analog playback and warping up to ±2%. Additionally, compared with a recent baseline scheme that uses brute force resampling to search for resynchronization, the proposed watermark is 300 times more computationally efficient, and does not compromise robustness to either desynchronization (e.g., jitter, resampling, time warping, frequency scaling) or non-desynchronizing modifications (e.g., AAC compression and additive noise). Andrew Nadeau, Gaurav Sharma 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2016 | A joint approach to vector road map registration and vehicle tracking for wide area motion imageryabstractModern aerial imaging platforms provide wide-area motion imagery (WAMI) at high spatial and moderate temporal resolutions making feasible a range of new applications. We consider the dual tasks of registering WAMI frames to geo-referenced vector road-maps and tracking vehicles through the progression of WAMI frames. We present a novel algorithm that performs these tasks jointly and offers improvements in both by exploiting the synergy between the tasks. Tracking for the large number of vehicles seen in urban-area WAMI is improved by auxiliary information that registration to the vector road-map provides by localizing roads within the scene. Similarly, registration of the WAMI frames to the vector map is improved by formulating the registration as a chamfer minimization between the vehicular trajectories and the road network, an approach that resolves challenges for registration posed by the fundamentally different data modalities between the aerial images and the vector road maps. Results obtained over our test datasets show the effectiveness of the proposed joint methodology. For both road network alignment and vehicle tracking, the proposed method offers a very significant improvement over available alternatives: the proposed approach yields better numerical metrics for quantification of registration accuracy and fewer false identification switches for tracked vehicles. Ahmed S. Elliethy, Gaurav Sharma 0001 |
ICASSP | 2 |
| 2016 | Per-channel color barcodes for displaysabstractWe present a per-channel framework for extending monochrome barcodes to color for display applications, offering increased data rates and capacity. Data is independently encoded into barcodes, incorporated within a color image as the red (R), green (G), and blue (B) channels and decoded from the corresponding R, G, and B channels in the image of the displayed barcode captured with a smartphone. Using a physical model for the display and capture processes, we show that the cross-channel interference in this situation is appropriately modeled by a linear mixing relation. Via experiments done over multiple smartphones and displays, we demonstrate that the impact of the interference is small and can be further mitigated through the use of interference cancellation. Two methods are presented for estimating and canceling the interference: either using pilot blocks in coherent designs where the synchronization marker patterns align between the codes or using expectation maximization. Results are presented for both QR and Aztec color codes in this framework, highlighting its versatility. Overall bit error rates are shown to be below 1.23% and are further reduced by interference cancellation. These low error rates are well within the correction capacity of the error correction codes used for typical color barcodes. Karthik Dinesh, Gaurav Sharma 0001 |
ICIP | 2 |
| 2016 | Comparative analysis of homologous buildings using range imagingabstractThis paper reports on a novel application of computer vision and image processing technologies to an interdisciplinary project in architectural history that seeks to help identify and visualize differences between homologous buildings constructed to a common template design. By identifying the mutations in homologous buildings, we assist humanists in giving voice to the contributions of the myriad additional “authors” for these buildings beyond their primary designers. We develop a framework for comparing 3D point cloud representations of homologous buildings captured using lidar: focusing on identifying similarities and differences, both among 3D scans of different buildings and between the 3D scans and the design specifications of architectural drawings. The framework addresses global and local alignment for highlighting gross differences as well as differences in individual structural elements and provides methods for readily highlighting the differences via suitable visualizations. The framework is demonstrated on pairs of homologous buildings selected from the Canadian and Ottoman rail networks. Results demonstrate the utility of the framework confirming differences already apparent to the humanist researchers and also revealing new differences that were not previously observed. Li Ding 0009, Ahmed S. Elliethy, Eitan Freedenberg, S. Alana Wolf-Johnson, Joshua Romphf, Peter Christensen, Gaurav Sharma 0001 |
ICIP | 7 |
| 2016 | Image anonymization for PRNU forensics: A set theoretic framework addressing compression resilienceabstractImage forensics using sensor photo-response nonuniformity (PRNU) provides a powerful method for associating an image with the camera that captured the image. To preserve privacy despite the availability of this powerful tool, we present a new framework for image anonymization. We formulate anonymization as a feasibility problem subject to multiple constraints that seek to ensure non-detectability of the PRNU fingerprint, visual fidelity to the original image, and compatibility to compression. A feasible anonymized image is then obtained via the method of projections onto convex sets using the inherent convexity of several constraints and convex approximations for the others. We demonstrate the effectiveness of our framework by benchmarking it over a publicly available dataset of images from multiple cameras and comparing against a recently presented alternative method. In the process we also highlight a key failing of several prior methods that fail to account for quantization in the compression process and suffer from catastrophic loss of anonymity when the anonymized image is stored in a compressed format, as would commonly be the case in realistic applications. We demonstrate specifically that the compression compatibility constraints we introduce help ensure that our method does not encounter this common pitfall. Ahmed S. Elliethy, Gaurav Sharma 0001 |
ICIP | 2 |
| 2016 | Automatic Registration of Wide Area Motion Imagery to Vector Road Maps by Exploiting Vehicle DetectionsabstractTo enrich large-scale visual analytics applications enabled by aerial wide area motion imagery (WAMI), we propose a novel methodology for accurately registering a geo-referenced vector roadmap to WAMI by using the locations of detected vehicles and determining a parametric transform that aligns these locations with the network of roads in the roadmap. Specifically, the problem is formulated in a probabilistic framework, explicitly allowing for spurious detections that do not correspond to on-road vehicles. The registration is estimated via the expectation-maximization (EM) algorithm as the planar homography that minimizes the sum of weighted squared distances between the homography-mapped detection locations and the corresponding closest point on the road network, where the weights are estimated posterior probabilities of detections being on-road vehicles. The weighted distance minimization is efficiently performed using the distance transform with the Levenberg-Marquardt nonlinear least-squares minimization procedure, and the fraction of spurious detections is estimated within the EM framework. The proposed method effectively sidesteps the challenges of feature correspondence estimation, applies directly to different imaging modalities, is robust to spurious detections, and is also more appropriate than feature matching for a planar homography. Results over three WAMI data sets captured by both visual and infrared sensors indicate the effectiveness of the proposed methodology: both visual comparison and numerical metrics for the registration accuracy are significantly better for the proposed method as compared with the existing alternatives. Ahmed S. Elliethy, Gaurav Sharma 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Improved specular regions localization and optical-flow based motion estimation via joint processingabstractFor specular regions (SRs), the assumption of brightness (or other) constancy between images corresponding to multiple views of a scene breaks down. As a consequence, optical-flow (OF) based motion-estimation (ME) algorithms that rely on constancy assumptions fail for specular regions. At the same time estimation of SRs in an image is also prone to errors, particularly to false positives from bright regions in the scene. In this paper, motivated by the fact that specular regions are typically encountered in image regions corresponding to portions of relatively smooth 3D surfaces, we propose an algorithm for improving ME and SRs localization via joint processing. Initial estimates of OF and of the SRs are obtained by conventional methods. The estimate of the SRs is updated using inconsistency of the OF with respect to the neighboring region to reinforce true positives and to reject false positives. The OF is then re-computed with a modified energy functional that, in effect, emphasizes regularization in a spatially adaptive neighborhood of the SRs to improve the estimated OF. Experimental results on synthetic and real image pairs demonstrate that the proposed algorithm offers a significant improvement in both SRs localization and ME over recently proposed methods for tackling these problems. Ahmed S. Elliethy, Gaurav Sharma 0001 |
ICIP | 2 |
| 2014 | State-of-charge estimation for supercapacitors: A Kalman filtering formulationabstractSupercapacitors are an attractive option for energy buffering because of their high efficiency, durability, and low environmental impact. For energy-aware applications, it is desirable to accurately estimate the buffered energy. Under conditions of varying energy supply and demand, estimation of buffered energy by using only the supercapacitor terminal voltage is inaccurate because this does not fully comprehend the physical state of charge. To address this problem, we present a Kalman filtering formulation, using the accepted three-branch circuit model for supercapacitors. Compared with an ideal capacitor, the physically-motivated three-branch model provides a much more accurate representation of the state of charge via three internal state voltages associated with short, medium, and long term charging constants. The proposed Kalman formulation tracks these unobservable internal states. This methodology demonstrates a significantly more accurate estimate of the buffered energy as compared with the alternative models of ideal capacitance or a recursive computation of the stored energy. Simulations conducted with variations that approximate recorded solar intensity profiles, our proposed approach has an error of 1% compared with 31% and 85% for the respective alternative models. Andrew Nadeau, Gaurav Sharma 0001, Tolga Soyata |
ICASSP | 2 |
| 2014 | A Regularized Model-Based Optimization Framework for Pan-SharpeningabstractPan-sharpening is a common postprocessing operation for captured multispectral satellite imagery, where the spatial resolution of images gathered in various spectral bands is enhanced by fusing them with a panchromatic image captured at a higher resolution. In this paper, pan-sharpening is formulated as the problem of jointly estimating the high-resolution (HR) multispectral images to minimize an objective function comprised of the sum of squared residual errors in physically motivated observation models of the low-resolution (LR) multispectral and the HR panchromatic images and a correlation-dependent regularization term. The objective function differs from and improves upon previously reported model-based optimization approaches to pan-sharpening in two major aspects: 1) a new regularization term is introduced and 2) a highpass filter, complementary to the lowpass filter for the LR spectral observations, is introduced for the residual error corresponding to the panchromatic observation model. To obtain pan-sharpened images, an iterative algorithm is developed to solve the proposed joint minimization. The proposed algorithm is compared with previously proposed methods both visually and using established quantitative measures of SNR, spectral angle mapper, relative dimensionless global error in synthesis, Q, and Q4 indices. Both the quantitative results and visual evaluation demonstrate that the proposed joint formulation provides superior results compared with pre-existing methods. A software implementation is provided. Hussein A. Aly 0002, Gaurav Sharma 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Joint multichannel pansharpening for multispectral imageryabstractPan-sharpening is a common post-processing operation for captured multispectral satellite imagery, where the spatial resolution of images gathered in various spectral bands is enhanced by fusing them with a panchromatic image captured at a higher resolution. Previously proposed pan-sharpening techniques operate on a per-channel basis, sharpening each multispectral band independently based on the panchromatic image, often in an ad hoc manner. In contrast with most prior techniques, we formulate pan-sharpening as the problem of jointly estimating the high resolution multispectral images to minimize the combined squared residual error in physically motivated observation models of the low resolution multispectral and the high resolution panchromatic images. To realize pan-sharpening using our proposed formulation, we develop an iterative algorithm to solve the joint minimization resulting in an overall algorithm with modest computational complexity. We evaluate our proposed algorithm and benchmark it against previously proposed methods using established quantitative measures of SNR, SAM, ERGAS, Q, and Q4 indices. Both the quantitative results and visual evaluation demonstrate that the proposed joint formulation provides superior results compared with pre-existing methods. Hussein A. Aly 0002, Gaurav Sharma 0001 |
ICASSP | 2 |
| 2013 | Two dimensional color calibration for four primary displaysabstractThe process to ensure a fixed and desired response from a color display generally consists of a per-channel calibration transform combined with a multi-dimensional characterization transformation. In this paper We focus on the former, i.e., the channel calibration. Conventional one-dimensional channel calibration strategies are inadequate for high resolution LCD displays because of inter-channel crosstalk. We address this problem by developing a color calibration strategy for a four primary LCD display, based on a two dimensional structure for channel calibration, which allows for simultaneously meeting the dual objectives of perceptual linearization of individual channels and gray balance along the device gray axes, despite inter-channel crosstalk. The two-dimensional nature of the transform represents a good balance between the dual objectives of low complexity and accurate control of key attributes of the displayed colors via channel calibration and experimental results demonstrate that the proposed scheme accomplishes its objectives offering a significant improvement over the per channel calibration for our four primary display system. Carlos Eduardo Rodríguez-Pardo, Gaurav Sharma 0001, Xiao-Fan Feng |
ICASSP | 2 |
| 2013 | Improved Low-Density Parity Check Accumulate (LDPCA) CodesabstractWe present improved constructions for Low-Density Parity-Check Accumulate (LDPCA) codes, which are rate-adaptive codes commonly used for distributed source coding (DSC) applications. Our proposed constructions mirror the traditional LDPCA approach; higher rate codes are obtained by splitting the check nodes in the decoding graph of lower rate codes, beginning with a lowest rate mother code. In a departure from the uniform splitting strategy adopted by prior LDPCA codes, however, the proposed constructions introduce non-uniform splitting of the check nodes at higher rates. Codes are designed by a global minimization of the average rate gap between the code operating rates and the corresponding theoretical lower bounds evaluated by density-evolution. In the process of formulating the design framework, the paper also contributes a formal definition of LDPCA codes. Performance improvements provided by the proposed non-uniform splitting strategy over the conventional uniform splitting approach used in prior work are substantiated via density evolution based analysis and DSC codec simulations. Optimized designs for our proposed constructions yield codes with a lower average rate gap than conventional designs and alleviate the trade-off between the performance at different rates inherent in conventional designs. A software implementation is provided for the codec developed. Gaurav Sharma 0001 |
IEEE Trans. Commun. | 2 |
| 2013 | Per-Colorant-Channel Color Barcodes for Mobile Applications: An Interference Cancellation FrameworkabstractWe propose a color barcode framework for mobile phone applications by exploiting the spectral diversity afforded by the cyan (C), magenta (M), and yellow (Y) print colorant channels commonly used for color printing and the complementary red (R), green (G), and blue (B) channels, respectively, used for capturing color images. Specifically, we exploit this spectral diversity to realize a three-fold increase in the data rate by encoding independent data in the C, M, and Y print colorant channels and decoding the data from the complementary R, G, and B channels captured via a mobile phone camera. To mitigate the effect of cross-channel interference among the print-colorant and capture color channels, we develop an algorithm for interference cancellation based on a physically-motivated mathematical model for the print and capture processes. To estimate the model parameters required for cross-channel interference cancellation, we propose two alternative methodologies: a pilot block approach that uses suitable selections of colors for the synchronization blocks and an expectation maximization approach that estimates the parameters from regions encoding the data itself. We evaluate the performance of the proposed framework using specific implementations of the framework for two of the most commonly used barcodes in mobile applications, QR and Aztec codes. Experimental results show that the proposed framework successfully overcomes the impact of the color interference, providing a low bit error rate and a high decoding rate for each of the colorant channels when used with a corresponding error correction scheme. Henryk Blasinski, Orhan Bulan, Gaurav Sharma 0001 |
IEEE Trans. Image Process. | 3 |
| 2012 | Correcting illumination variations in photomicrograph mosaics of daguerreotypesabstractThe Cincinnati Waterfront Daguerreotype Panorama, owned by the Public Library of Cincinnati and Hamilton County, was brought to George Eastman House in 2007 for conservation treatment. At George Eastman House the plates were imaged under a microscope to capture details at nearly their full resolution and composited mosaic images comprised of 756 images per plate were created. Although the mosaic yields an otherwise unattainable contiguous high resolution and detailed view of the Cincinnati waterfront at the time of capture, non-uniformity of the capture illumination introduced a patterning that impairs the visual and analytic value of the mosaic. In this paper, we develop a mathematical model for the mosaic image capture to analyze the degradation caused by the illumination non-uniformity. Motivated by the analysis, we develop a simple signal processing solution to correct for the lighting variation by utilizing suitably tuned frequency domain filtering. Using the proposed method, we demonstrate that the patterning due to non-uniformity of illumination is corrected, resulting in a significantly improved mosaic. Orhan Bulan, Robert Buckley, Ralph Wiegandt, Gaurav Sharma 0001 |
ICASSP | 4 |
| 2012 | Improved color barcodes via Expectation Maximization style interference cancellationabstractEncoding data independently in cyan, magenta, and yellow (CMY) print colorant channels with detection in complementary Red, green, and blue (RGB) image capture channels offers an attractive framework for extending monochrome barcodes to color with increased data rates. The undesired absorption of colorants in regions of spectral sensitivity of the noncomplementary capture channels, however, gives rise to cross-channel color interference that significantly deteriorates the performance of the color barcode system. In this paper, we propose an Expectation Maximization (EM) style algorithm to estimate and cancel this color interference and improve the overall performance of the barcode system. Our method utilizes a physical model for print-capture process where the model parameters vary depending on printer, capture device, and illumination. We estimate the model parameters using an iterative EM-style approach and obtain an estimate of CMY colorant channels from the scanned RGB barcode by using the estimated model parameters. Our experimental results show that the proposed method mitigates the effect of color interference and significantly reduces the bit error rates for the recovered data. Orhan Bulan, Gaurav Sharma 0001 |
ICASSP | 2 |
| 2012 | Technical program chairs' overviewabstractWelcome to the 2012 IEEE International Conference on Image Processing (ICIP), the 19th in the series of ICIPs. We hope you find ICIP inspiring and rewarding both for the knowledge you gain of advances in Image and Video processing through the conference technical presentations, and for the equally valuable “side information” that comes from personal interactions and hallway conversations with colleagues and friends, old and new, that you meet here. Gaurav Sharma 0001, Sheila S. Hemami |
ICIP | 1 |
| 2012 | High contrast stochastic screenwatermarks for color halftone printsabstractEmbedded watermarks in printed halftone images, which can subsequently be detected using an visual aid or using a watermark detection algorithm on a scan of the image, are of interest in wide range of applications. For black and white halftone printing using stochastic screens, digital watermarks that are embedded as correlations in the halftone screen have been previously proposed. Here we present a novel extension of these watermarks to color that produces a high contrast watermark by using the colorant separations coherently with a single watermarked stochastic screen and performing detection coherently across the color separations. Compared with independent watermarking of the halftone separations, the resulting watermark offers significantly higher contrast in the detected image. Gaurav Sharma 0001, Shen-Ge Wang |
ICIP | 1 |
| 2012 | Capacity Analysis For Orthogonal Halftone Orientation Modulation ChannelsabstractHalftone dot orientation modulation has recently been proposed as a method for data hiding in printed images. Extraction of data embedded with halftone orientation modulation is accomplished by computing, from the scanned hardcopy image, detection statistics that uniquely identify the embedded orientation. From a communications perspective, this data hiding setup forms an interesting class of channels with dot orientation as input and a vector of statistics as the output. This paper derives capacity expressions for these channels that allow for numerical evaluation of the capacity. Results provide significant insight for orientation modulation based print-scan resilient data hiding: the capacity varies significantly as a function of the image graylevel and experimentally observed error free data rates closely mirror the variation in capacity. Orhan Bulan, Vishal Monga, Gaurav Sharma 0001 |
IEEE Trans. Image Process. | 3 |
| 2011 | Iterative estimation of structures of multiple RNA homologs: TurbofoldabstractTurboFold, an iterative algorithm for estimating the common secondary structures of multiple RNA homologs, is presented. The algorithm is motivated by and has structure and attributes analogous to the turbo decoding algorithm in communications. Instead of solving the joint problem of aligning and folding multiple RNA sequences, TurboFold uses an iterative process to fold a collection of RNA homologs. Beneficial information from inter-sequence comparisons is incorporated by using feedback from iteration to iteration in the form of pseudo-prior probabilities for base pairing which are incorporated in the computation of base pairing probabilities. As a result Turbo Fold retains several of the advantages of join alignment and folding while maintaining a per iteration computational complexity comparable to single sequence RNA folding. Experimental evaluation of the algorithm, performed over six ncRNA families, demonstrates that TurboFold achieves high accuracy, offering better performance than available alternatives for estimating RNA base pairing probabilities. Gaurav Sharma 0001, Arif Ozgun Harmanci, David H. Mathews |
ICASSP | 1 |
| 2011 | TurboFold: Iterative probabilistic estimation of secondary structures for multiple RNA sequencesabstractBACKGROUND: The prediction of secondary structure, i.e. the set of canonical base pairs between nucleotides, is a first step in developing an understanding of the function of an RNA sequence. The most accurate computational methods predict conserved structures for a set of homologous RNA sequences. These methods usually suffer from high computational complexity. In this paper, TurboFold, a novel and efficient method for secondary structure prediction for multiple RNA sequences, is presented. RESULTS: TurboFold takes, as input, a set of homologous RNA sequences and outputs estimates of the base pairing probabilities for each sequence. The base pairing probabilities for a sequence are estimated by combining intrinsic information, derived from the sequence itself via the nearest neighbor thermodynamic model, with extrinsic information, derived from the other sequences in the input set. For a given sequence, the extrinsic information is computed by using pairwise-sequence-alignment-based probabilities for co-incidence with each of the other sequences, along with estimated base pairing probabilities, from the previous iteration, for the other sequences. The extrinsic information is introduced as free energy modifications for base pairing in a partition function computation based on the nearest neighbor thermodynamic model. This process yields updated estimates of base pairing probability. The updated base pairing probabilities in turn are used to recompute extrinsic information, resulting in the overall iterative estimation procedure that defines TurboFold.TurboFold is benchmarked on a number of ncRNA datasets and compared against alternative secondary structure prediction methods. The iterative procedure in TurboFold is shown to improve estimates of base pairing probability with each iteration, though only small gains are obtained beyond three iterations. Secondary structures composed of base pairs with estimated probabilities higher than a significance threshold are shown to be more accurate for TurboFold than for alternative methods that estimate base pairing probabilities. TurboFold-MEA, which uses base pairing probabilities from TurboFold in a maximum expected accuracy algorithm for secondary structure prediction, has accuracy comparable to the best performing secondary structure prediction methods. The computational and memory requirements for TurboFold are modest and, in terms of sequence length and number of sequences, scale much more favorably than joint alignment and folding algorithms. CONCLUSIONS: TurboFold is an iterative probabilistic method for predicting secondary structures for multiple RNA sequences that efficiently and accurately combines the information from the comparative analysis between sequences with the thermodynamic folding model. Unlike most other multi-sequence structure prediction methods, TurboFold does not enforce strict commonality of structures and is therefore useful for predicting structures for homologous sequences that have diverged significantly. TurboFold can be downloaded as part of the RNAstructure package at http://rna.urmc.rochester.edu. Arif Ozgun Harmanci, Gaurav Sharma 0001, David H. Mathews |
BMC Bioinform. | 2 |
| 2011 | High Capacity Color Barcodes: Per Channel Data Encoding via Orientation Modulation in Elliptical Dot ArraysabstractWe present a new high capacity color barcode. The barcode we propose uses the cyan, magenta, and yellow (C,M,Y) colorant separations available in color printers and enables high capacity by independently encoding data in each of these separations. In each colorant channel, payload data is conveyed by using a periodic array of elliptically shaped dots whose individual orientations are modulated to encode the data. The orientation based data encoding provides beneficial robustness against printer and scanner tone variations. The overall color barcode is obtained when these color separations are printed in overlay as is common in color printing. A reader recovers the barcode data from a conventional color scan of the barcode, using red, green, and blue (R,G,B) channels complementary, respectively, to the print C, M, and Y channels. For each channel, first the periodic arrangement of dots is exploited at the reader to enable synchronization by compensating for both global rotation/scaling in scanning and local distortion in printing. To overcome the color interference resulting from colorant absorptions in noncomplementary scanner channels, we propose a novel interference minimizing data encoding approach and a statistical channel model (at the reader) that captures the characteristics of the interference, enabling more accurate data recovery. We also employ an error correction methodology that effectively utilizes the channel model. The experimental results show that the proposed method works well, offering (error-free) operational rates that are comparable to or better than the highest capacity barcodes known in the literature. Orhan Bulan, Gaurav Sharma 0001 |
IEEE Trans. Image Process. | 2 |
| 2010 | Swift: Scalable weighted iterative sampling for flow cytometry clusteringabstractFlow cytometry (FC) is a powerful technology for rapid multivariate analysis and functional discrimination of cells. Current FC platforms generate large, high-dimensional datasets which pose a significant challenge for traditional manual bivariate analysis. Automated multivariate clustering, though highly desirable, is also stymied by the critical requirement of identifying rare populations that form rather small clusters, in addition to the computational challenges posed by the large size and dimensionality of the datasets. In this paper, we address these twin challenges by developing a two-stage scalable multivariate parametric clustering algorithm. In the first stage, we model the data as a mixture of Gaussians and use an iterative weighted sampling technique to estimate the mixture components successively in order of decreasing size. In the second stage, we apply a graph-based hierarchical merging technique to combine Gaussian components with significant overlaps into the final number of desired clusters. The resulting algorithm offers a reduction in complexity over conventional mixture modeling while simultaneously allowing for better detection of small populations. We demonstrate the effectiveness of our method both on simulated data and actual flow cytometry datasets. Iftekhar Naim, Suprakash Datta, Gaurav Sharma 0001, James S. Cavenaugh, Tim R. Mosmann |
ICASSP | 3 |
| 2010 | Multiplexed clustered-dot halftone watermarks using bi-directional phase modulation and detectionabstractWe present a method for embedding and detection of visual watermark patterns in printed images that use clustered-dot halftones in the printing process. The method allows two independent watermark patterns to be multiplexed, i.e. embedded in the same region of the printed image, thereby offering an improvement over prior techniques. The watermark patterns are embedded via phase modulation along the two periodicity directions of the halftone. Each embedded pattern can be visually detected, without interference from the other watermark, when the printed image, or a scan thereof, is overlaid with a decoder mask consisting of periodic lines oriented along the corresponding halftone periodicity direction. We experimentally demonstrate that our proposed multiplexing technique is effective. We also present analysis that demonstrates that the embedding has minimal visual impact and explains the pattern observed in the watermark detection process. Basak Oztan, Gaurav Sharma 0001 |
ICIP | 2 |
| 2010 | Efficient Classification of Scanned Media Using Spatial StatisticsabstractPhotography, lithography, xerography, and inkjet printing are the dominant technologies for color printing. Images produced on these "different media" are often scanned either for the purpose of copying or creating an electronic representation. For an improved color calibration during scanning, a media identification from the scanned image data is desirable. In this paper, we propose an efficient algorithm for automated classification of input media into four major classes corresponding to photographic, lithographic, xerographic and inkjet. Our technique exploits the strong correlation between the type of input media and the spatial statistics of corresponding images, which are observed in the scanned images. We adopt ideas from spatial statistics literature, and design two spatial statistical measures of dispersion and periodicity, which are computed over spatial point patterns generated from blocks of the scanned image, and whose distributions provide the features for making a decision. We utilize extensive training data and determined well separated decision regions to classify the input media. We validate and tested our classification technique results over an independent extensive data set. The results demonstrate that the proposed method is able to distinguish between the different media with high reliability. Gozde Unal, Gaurav Sharma 0001, Reiner Eschbach |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2010 | Orientation Modulation for Data Hiding in Clustered-Dot Halftone PrintsabstractWe present a new framework for data hiding in images printed with clustered dot halftones. Our application scenario, like other hardcopy embedding methods, encounters fundamental challenges due to extreme bilevel quantization inherent in halftoning, the stringent requirements of image fidelity, and other unavoidable printing and scanning distortions. To overcome these challenges, while still allowing for automated extraction of the embedded data and a high embedding capacity, we propose a number of innovations. First, we perform the embedding jointly with the halftoning by employing an analytical halftone threshold function that allows steering of the halftone spot orientation within each halftone cell based upon embedded data. In this process, image fidelity is emphasized and, if necessary, the capability to recover individual data values is sacrificed resulting in unavoidable erasures and errors. To overcome these and other sources of errors, we propose a suitable data detection and error control methodology based upon a statistical representation for the print-scan channel that effectively models the channel dependence upon the cover image gray-level. To combat the geometric distortion inherent in the print-scan process, we exploit the periodic halftone structure to recover from global scaling and rotation and propose a novel decision directed synchronization technique that counters locally varying printing distortion. Experimental results demonstrate the power of the proposed framework: we achieve high operational rates while preserving halftone image quality. Orhan Bulan, Gaurav Sharma 0001, Vishal Monga |
IEEE Trans. Image Process. | 2 |
| 2010 | Adaptive Sensing and Optimal Power Allocation for Wireless Video Sensors With Sigma-Delta ImagerabstractWe consider optimal power allocation for wireless video sensors (WVSs), including the image sensor subsystem in the system analysis. By assigning a power-rate-distortion (P-R-D) characteristic for the image sensor, we build a comprehensive P-R-D optimization framework for WVSs. For a WVS node operating under a power budget, we propose power allocation among the image sensor, compression, and transmission modules, in order to minimize the distortion of the video reconstructed at the receiver. To demonstrate the proposed optimization method, we establish a P-R-D model for an image sensor based upon a pixel level sigma-delta (Σ∆) image sensor design that allows investigation of the tradeoff between the bit depth of the captured images and spatio-temporal characteristics of the video sequence under the power constraint. The optimization results obtained in this setting confirm that including the image sensor in the system optimization procedure can improve the overall video quality under power constraint and prolong the lifetime of the WVSs. In particular, when the available power budget for a WVS node falls below a threshold, adaptive sensing becomes necessary to ensure that the node communicates useful information about the video content while meeting its power budget. Malisa Marijan, Ilker Demirkol, Danijel Maricic, Gaurav Sharma 0001, Zeljko Ignjatovic |
IEEE Trans. Image Process. | 4 |
| 2010 | Camera Scheduling and Energy Allocation for Lifetime Maximization in User-Centric Visual Sensor NetworksabstractWe explore camera scheduling and energy allocation strategies for lifetime optimization in image sensor networks. For the application scenarios that we consider, visual coverage over a monitored region is obtained by deploying wireless, battery-powered image sensors. Each sensor camera provides coverage over a part of the monitored region and a central processor coordinates the sensors in order to gather required visual data. For the purpose of maximizing the network operational lifetime, we consider two problems in this setting: a) camera scheduling, i.e., the selection, among available possibilities, of a set of cameras providing the desired coverage at each time instance, and b) energy allocation, i.e., the distribution of total available energy between the camera sensor nodes. We model the network lifetime as a stochastic random variable that depends upon the coverage geometry for the sensors and the distribution of data requests over the monitored region, two key characteristics that distinguish our problem from other wireless sensor network applications. By suitably abstracting this model of network lifetime and utilizing asymptotic analysis, we propose lifetime-maximizing camera scheduling and energy allocation strategies. The effectiveness of the proposed camera scheduling and energy allocation strategies is validated by simulations. Gaurav Sharma 0001 |
IEEE Trans. Image Process. | 2 |
| 2009 | Geometric distortion signatures for printer identificationabstractWe present a forensic technique for analyzing a printed image in order to trace the originating printer. Our method, which is applicable for commonly used electrophotographic (EP) printers, operates by exploiting the geometric distortion that these devices inevitably introduce in the printing process. In the proposed method, first a geometric distortion signature is estimated for an EP printer. This estimate is obtained using only the images printed on the printer and without access to the internal printer controls. Once a database of printer signatures is available, the printer utilized to print a test image is identified by computing the geometric distortion signature from test image and correlating the computes signatures against the printer signatures in the database. Experiments conducted over a corpus of EP printers demonstrate that the geometric distortion signatures of test documents exhibit high correlation with the corresponding printer signatures and a low correlation with other printer signatures. The method is therefore quite promising for forensic printer identification applications. We highlight several of the capabilities and challenges for the method. Orhan Bulan, Junwen Mao, Gaurav Sharma 0001 |
ICASSP | 3 |
| 2009 | Q-SIFT: Efficient feature descriptors for distributed camera calibrationabstractWe consider camera self-calibration, i.e. the estimation of parameters for camera sensors, in the setting of a visual sensor network where the sensors are distributed and energy-constrained. With the objective of reducing the communication burden and thereby maximizing network lifetime, we propose an energy-efficient approach for self-calibration where feature points are extracted locally at the cameras and efficient descriptions for these features are transmitted to a central processor that performs the self-calibration. Specifically, in this work we use reduced-dimensionality quantized approximations as efficient feature descriptors. The effectiveness of the proposed technique is validated through feature matching, and epipolar geometry estimation which enable self-calibration of the network. Gaurav Sharma 0001 |
ICASSP | 2 |
| 2009 | Device temporal forensics: An information theoretic approachabstractBy formulating the problem of ordering the outputs observed from a device over time, we pose a new problem in forensics and propose a framework for addressing this problem of device temporal forensics. Our proposed framework is based on a two-stage approach wherein time-dependent device parameters are first estimated from observed outputs and the resulting estimates are then temporally ordered by employing a Markov model for the temporal evolution of device parameters and exploiting the data processing inequality in information theory. We demonstrate and evaluate a simple realization of the framework for digital camera forensics based on photo-response non-uniformity. Results obtained over a database of online images indicate that the method provides accurate temporal ordering. Junwen Mao, Orhan Bulan, Gaurav Sharma 0001, Suprakash Datta |
ICIP | 3 |
| 2009 | Optimized energy allocation in battery powered image sensor networksabstractWe investigate energy allocation strategies in image sensor networks for the purpose of maximizing the network operational lifetime. For the application scenarios that we consider, visual coverage over a monitored region is obtained by deploying wireless, battery-powered image sensors. Each sensor camera provides coverage over a part of the monitored region and a central processor coordinates the sensors in order to gather required visual data. We characterize the network lifetime as a stochastic random variable that depends upon the coverage geometry for the sensors and the distribution of data requests over the monitored region. Using this characterization we consider optimized strategies for energy allocation among the sensors that maximize the expected network lifetime. The formulation naturally leads to a max-min optimization problem that aims to maximize the duration of coverage for the most critical region for which the available energy is the least. We transform this problem into an equivalent linear programming problem, leading to a computationally efficient solution. The effectiveness of the proposed energy allocation strategy is validated by simulations. Gaurav Sharma 0001 |
ICIP | 2 |
| 2009 | Optimal Spread Spectrum Watermark Embedding via a Multistep Feasibility FormulationabstractWe consider optimal formulations of spread spectrum watermark embedding where the common requirements of watermarking, such as perceptual closeness of the watermarked image to the cover and detectability of the watermark in the presence of noise and compression, are posed as constraints while one metric pertaining to these requirements is optimized. We propose an algorithmic framework for solving these optimal embedding problems via a multistep feasibility approach that combines projections onto convex sets (POCS) based feasibility watermarking with a bisection parameter search for determining the optimum value of the objective function and the optimum watermarked image. The framework is general and can handle optimal watermark embedding problems with convex and quasi-convex formulations of watermark requirements with assured convergence to the global optimum. The proposed scheme is a natural extension of set-theoretic watermark design and provides a link between convex feasibility and optimization formulations for watermark embedding. We demonstrate a number of optimal watermark embeddings in the proposed framework corresponding to maximal robustness to additive noise, maximal robustness to compression, minimal frequency weighted perceptual distortion, and minimal watermark texture visibility. Experimental results demonstrate that the framework is effective in optimizing the desired characteristic while meeting the constraints. The results also highlight both anticipated and unanticipated competition between the common requirements for watermark embedding. Oktay Altun, Adem Orsdemir, Gaurav Sharma 0001, Mark F. Bocko |
IEEE Trans. Image Process. | 3 |
| 2009 | Continuous Phase-Modulated HalftonesabstractA generalization of periodic clustered-dot halftones is proposed, wherein the phase of the halftone spots is modulated using a secondary signal. The process is accomplished by using an analytic halftone threshold function that allows halftones to be generated with controlled phase variation in different regions of the printed page. The method can also be used to modulate the screen frequency, albeit with additional constraints. Visible artifacts are minimized/eliminated by ensuring the continuity of the modulation in phase. Limitations and capabilities of the method are analyzed through a quantitative model. The technique can be exploited for two applications that are presented in this paper: a) embedding watermarks in the halftone image by encoding information in phase or in frequency and b) modulating the screen frequency according to the frequency content of the continuous tone image in order to improve spatial and tonal rendering. Experimental performance is demonstrated for both applications. Basak Oztan, Gaurav Sharma 0001 |
IEEE Trans. Image Process. | 2 |
| 2008 | On the capacity of orientation modulation halftone channelsabstractClustered-dot halftones are extensively utilized in hardcopy printing. Modulation of the dot orientation in these halftones offers an avenue for data embedding which has been exploited in a number of different methods. We consider the capacity of these channels, modeling them as binary orientation input channels with vector valued output detection statistics. We derive upper bounds on the capacity for three channel conditional distributions corresponding to sub-Gaussian, Gaussian and super-Gaussian distributions. Using experimentally estimated channel parameters our bounds reveal that channel capacity has noticeable variations as a function of gray level. Highlights, shadows and mid-tones offer negligible capacity, on the contrary the regions between highlights and mid-tones or shadows and mid-tones offer high capacity for data embedding. Orhan Bulan, Gaurav Sharma 0001, Vishal Monga |
ICASSP | 2 |
| 2008 | Probabilistic structural alignment of RNA sequencesabstractWe propose an algorithm for estimating the common secondary structure, alignment, and posterior base pairing probabilities for two RNA sequences. A definition of structural alignment is presented based on a novel concept of matched helical regions that generalizes the common secondary structure and alignment constraints used in prior work. A probabilistic framework for scoring structural alignments is developed based on a pseudo free energy model. Utilizing the model, maximum a posteriori probability estimates of secondary structure and alignment, and a posteriori probabilities for base pairing are computed using an efficient dynamic programming algorithm. Experimental results demonstrate that the proposed method offers significant improvements in structure and alignment prediction accuracy in comparison with single sequence thermodynamic methods for secondary structure prediction and purely sequence based alignment. Arif Ozgun Harmanci, Gaurav Sharma 0001, David H. Mathews |
ICASSP | 2 |
| 2008 | Adaptive decoding for halftone orientation-based data hidingabstractHalftone image watermarking techniques that allow automated extraction of the embedded watermark data are useful in a variety of document security and workflow applications. The print-scan process inherent in these applications introduces distortions whose characteristics exhibit a strong spatial dependence on the cover image in which the data is embedded. In this paper, we demonstrate that a characterization of this channel dependence and the use of error correction coding that exploits this dependence via image adaptive decoding offers a significant performance gain. We show this advantage in the specific context of a high rate data embedding method that utilizes orientation modulation for data embedding in clustered dot halftones and moment based detection at the receiver. Channel coding for this scenario utilizing convolutional codes and Repeat Accumulate (RA) codes highlights the advantage of the proposed adaptive decoding methodology. The performance for the RA codes with adaptive decoding also reveals that the resulting system has significantly higher operational rates than prior schemes. Orhan Bulan, Gaurav Sharma 0001, Vishal Monga |
ICIP | 2 |
| 2008 | Insertion, Deletion Codes With Feature-Based Embedding: A New Paradigm for Watermark Synchronization With Applications to Speech WatermarkingabstractA framework is proposed for synchronization in feature-based data embedding systems that is tolerant of errors in estimated features. The method combines feature-based embedding with codes capable of simultaneous synchronization and error correction, thereby allowing recovery from both desynchronization caused by feature estimation discrepancies between the embedder and receiver; and alterations in estimated symbols arising from other channel perturbations. A speech watermark is presented that constitutes a realization of the framework for 1-D signals. The speech watermark employs pitch modification for data embedding and Davey and Mackay's insertion, deletion, and substitution (IDS) codes for synchronization and error recovery. Experimental results demonstrate that the system indeed allows watermark data recovery, despite feature desynchronization. The performance of the speech watermark is optimized by estimating the channel parameters required for the IDS decoding at the receiver via the expectation-maximization algorithm. In addition, acceptable watermark power levels (i.e., the range of pitch modification that is perceptually tolerable) are determined from psychophysical tests. The proposed watermark demonstrates robustness to low-bit-rate speech coding channels (Global System for Mobile Communications at 13 kb/s and AMR at 5.1 kb/s), which have posed a serious challenge for prior speech watermarks. Thus, the watermark presented in this paper not only highlights the utility of the proposed framework but also represents a significant advance in speech watermarking. Issues in extending the proposed framework to 2-D and 3-D signals and different application scenarios are identified. David J. Coumou, Gaurav Sharma 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2007 | Toward Turbo Decoding of RNA Secondary StructureabstractWe propose an iterative probabilistic algorithm for estimation of RNA secondary structure using sequence data from two homologous sequences. The method is intended to exploit intersequence correlations "encoded" in the form of probabilistic models for alignment and for common secondary structure. In analogy with turbo-decoding in digital communications, we formulate a maximum a posteriori probability objective function for joint structural prediction and sequence alignment using iterations over individual structural and sequential alignment models with soft-input soft-output estimators. As a preliminary step toward realizing this methodology, we present results obtained from incorporating (hard) constraints based on posterior sequence alignment probabilities in joint secondary structure prediction. Through experimental evaluations over available databases of known secondary structure, we demonstrate that this results in a significant decrease in computation time while simultaneously providing a marginal increase in structural prediction accuracy. Arif Ozgun Harmanci, Gaurav Sharma 0001, David H. Mathews |
ICASSP (1) | 2 |
| 2007 | Collusion Resilient Fingerprint Design by Alternating ProjectionsabstractDigital fingerprinting techniques aim to embed unique identification information into digital content distributed to individual users in order to track unauthorized use of multimedia files. Fraudulent users may not only attempt to remove the embedded signatures but also may form coalitions in order to remove the embedded fingerprint and disable tracking. This makes the design of fingerprints challenging. An effective fingerprint should not only carry the assigned users information but also guard against the possibility of falsely implicating an innocent user. Furthermore, in possible collusion scenarios, the colluded copies should identify each of the colluders. The embedded fingerprints should be imperceptible to maintain the commercial value of the content and preferably the fingerprint-based identification should survive content preserving signal processing. In this paper we give a precise description of each of these requirements and give a solution framework to obtain a set of fingerprinted images meeting these requirements. Oktay Altun, Gaurav Sharma 0001, Adem Orsdemir, Mark F. Bocko |
ICIP (4) | 2 |
| 2007 | Conditions for Color Misregistration Sensitivity in Clustered-dot HalftonesabstractMisregistration between the color separations of a printed image, which is often inevitable, can cause objectionable color shifts in average color. We analyze the impact of inter-separation misregistration on clustered-dot halftones using Fourier analysis in a lattice framework. Our analysis provides a complete characterization of the conditions under which the average color is invariant to displacement misregistration. In addition to known conditions on colorant spectra and periodicity of the halftones, the work reveals that invariance can also be obtained when these conditions are violated for suitable dot shapes and displacements. Examples for these conditions are included, as is the consideration of traditional halftone configurations. Basak Oztan, Gaurav Sharma 0001, Robert P. Loce |
ICIP (4) | 2 |
| 2007 | Lifetime-Distortion Trade-off in Image Sensor NetworksabstractWe examine the trade-off between lifetime and distortion in image sensor networks deployed for gathering visual information over a monitored region. Users navigate over the monitored region by specifying a viewpoint that varies with time, and the network attempts to meet the user requirement by synthesizing the desired view using a selection of cameras. We compare two camera selection methods, the first maximizes PSNR for the user's view without considering the cameras' available energy whereas the second uses knowledge of available energy at the cameras to maximize network lifetime. Our simulation results demonstrate a clear trade-off between the two strategies, with selection based on energy alone performing up to 3 dB worse initially than the PSNR based selection yet providing significantly higher coverage lifetime. In addition, we observe that under a fixed total energy constraint, more cameras with lower energy per node are preferable over fewer cameras with higher energy per node. Our results suggest that a hybrid or adaptive camera selection algorithm may provide the optimal lifetime-distortion trade-off. Stanislava Soro, Gaurav Sharma 0001, Wendi B. Heinzelman |
ICIP (5) | 3 |
| 2007 | Efficient pairwise RNA structure prediction using probabilistic alignment constraints in DynalignabstractBACKGROUND: Joint alignment and secondary structure prediction of two RNA sequences can significantly improve the accuracy of the structural predictions. Methods addressing this problem, however, are forced to employ constraints that reduce computation by restricting the alignments and/or structures (i.e. folds) that are permissible. In this paper, a new methodology is presented for the purpose of establishing alignment constraints based on nucleotide alignment and insertion posterior probabilities. Using a hidden Markov model, posterior probabilities of alignment and insertion are computed for all possible pairings of nucleotide positions from the two sequences. These alignment and insertion posterior probabilities are additively combined to obtain probabilities of co-incidence for nucleotide position pairs. A suitable alignment constraint is obtained by thresholding the co-incidence probabilities. The constraint is integrated with Dynalign, a free energy minimization algorithm for joint alignment and secondary structure prediction. The resulting method is benchmarked against the previous version of Dynalign and against other programs for pairwise RNA structure prediction. RESULTS: The proposed technique eliminates manual parameter selection in Dynalign and provides significant computational time savings in comparison to prior constraints in Dynalign while simultaneously providing a small improvement in the structural prediction accuracy. Savings are also realized in memory. In experiments over a 5S RNA dataset with average sequence length of approximately 120 nucleotides, the method reduces computation by a factor of 2. The method performs favorably in comparison to other programs for pairwise RNA structure prediction: yielding better accuracy, on average, and requiring significantly lesser computational resources. CONCLUSION: Probabilistic analysis can be utilized in order to automate the determination of alignment constraints for pairwise RNA structure prediction methods in a principled fashion. These constraints can reduce the computational and memory requirements of these methods while maintaining or improving their accuracy of structural prediction. This extends the practical reach of these methods to longer length sequences. The revised Dynalign code is freely available for download. Arif Ozgun Harmanci, Gaurav Sharma 0001, David H. Mathews |
BMC Bioinform. | 2 |
| 2006 | Set Theoretic Quantization Index Modulation WatermarkingabstractWe introduce a set theoretic framework for quantization index modulation (QIM) watermarking and illustrate its potency by designing a semi-fragile watermark that is both visually adaptive and tolerant to compression. We determine the watermarked image to satisfy the multiple constraints of watermark detectability, imperceptibility and robustness to compression using the method of projections onto convex sets (POCS). Mark embedding is performed through implicit quantization of statistical features, specifically the mean, of randomly selected pixel locations from the image. This is accomplished by defining a detectability constraint set that imposes the quantization constraint. We present experimental results demonstrating the efficacy of the technique in the presence of JPEG compression Oktay Altun, Gaurav Sharma 0001, Mark F. Bocko |
ICASSP (2) | 2 |
| 2006 | Continuous Phase Modulated Halftones and Their Application to Halftone Data EmbeddingabstractWe propose a generalization of periodic clustered-dot halftones, wherein the phase of the halftones is modulated using a secondary signal. The process is accomplished using an analytic halftone threshold function and allows halftones to be generated with different phase variation in different regions of the printed page. We demonstrate that ensuring continuity of the phase assures that the resulting halftone images are free from visible artifacts despite the modulation in phase. We present, applications of the proposed method to halftone data embedding, wherein the changes in phase or in frequency encode the embedded information. For the frequency embedding, using continuous phase modulation, we also consider the limitations on signals that are embedded within our framework. For both applications, we demonstrate how the embedded signals may be recovered from the printed halftones either using the image as a self-reference which reveals the watermark when it is shifted and overlaid with itself or by employing a separate transparency mask Basak Oztan, Gaurav Sharma 0001 |
ICASSP (2) | 2 |
| 2006 | Optimum Watermark Design by Vector Space ProjectionsabstractWe introduce an optimum watermark embedding technique that satisfies common watermarking requirements such as visual fidelity, sufficient embedding rate, robustness against noise and tolerance to benign signal processing, while optimizing one of these requirements. The algorithm distinguishes itself from other watermark optimization techniques in its flexibility for incorporating constraints and in assuring convergence to the globally optimum point when all constraints are convex. The proposed scheme is a natural extension of set-theoretic watermark design and inherits all the advantages of it. A watermarked image is first obtained by POCS, then optimum watermarked image is determined by an iterative search procedure based on a simple bi-section method. Experimental results are presented to illustrate the effectiveness of the method. Oktay Altun, Gaurav Sharma 0001, Mark F. Bocko |
ICIP | 2 |
| 2006 | Multi-View Image Registration for Wide-Baseline Visual Sensor NetworksabstractWe present a new dense multi-view registration technique for wide-baseline video/images that integrates a parametric optical flow-based approach with a sparse set of feature correspondences, based on a locally planar approximation of a nonplanar scene. The proposed method can deal with illuminance variations between the views, which is critically important for wide-baseline applications. It differs from existing work on wide-baseline image registration in that it requires only image information and provides dense matching without computing any camera calibration matrices or performing any prior scene segmentation. These characteristics render the method suitable for practical deployment in visual sensor networks, towards which the current work is directed. We demonstrate the performance of the proposed method on simulated multi-view images of a virtual 3D world composed of piece-wise smooth textured surfaces, as well as real wide-baseline images of nonplanar textured surfaces. Gulcin Caner, A. Murat Tekalp, Gaurav Sharma 0001, Wendi B. Heinzelman |
ICIP | 3 |
| 2006 | Self-Modulated HalftonesabstractWe propose an analytic method to overcome the trade-off between the spatial and tonal resolution of traditional clustered dot halftones. Continuous phase modulated halftones that allow variations in screen frequency in different regions of the printed image are employed and halftone screen frequency is varied according to the frequency content of the image to be halftoned. The method, which we term self-modulated halftoning, has a computational complexity similar to screening and is significantly lower than that of other adaptive methods that have previously been used to address the same problem. We demonstrate the experimental performance of self-modulated halftones and discuss its capabilities and limitations. Basak Oztan, Gaurav Sharma 0001 |
ICIP | 2 |
| 2006 | Watermark Synchronization for Feature-Based Embedding: Application to SpeechabstractWe propose a novel framework for synchronization in feature-based data embedding systems. The framework is tolerant to de-synchronizing errors in feature estimates, which have hitherto crippled feature-based embedding methods. The method uses a concatenated coding system comprising of an outer q-ary LDPC code and an inner insertion-deletion code to recover from both de-synchronization caused by feature estimation discrepancies between the transmitter and receiver; and errors in estimated symbols arising from other channel perturbations. We illustrate the framework in a speech watermarking application employing pitch modification for data-embedding. We show that the method indeed allows recovery of watermark data even in the presence of de-synchronization errors in the underlying pitch-based embedding. The resilience of the method is also demonstrated over channels employing low bit rate speech encoders David J. Coumou, Gaurav Sharma 0001 |
ICME | 2 |
| 2006 | A Set Theoretic Framework for Watermarking and Its Application to Semifragile Tamper DetectionabstractWe introduce a set theoretic framework for watermarking. Multiple requirements, such as watermark embedding strength, imperceptibility, robustness to benign signal processing, and fragility under malicious attacks are described as constraint sets and a watermarked image is determined as a feasible solution satisfying these constraints. We illustrate that several constraints can be formulated as convex sets and develop a watermarking algorithm based on the method of projections onto convex sets. The framework allows flexible incorporation of different constraints, including embedding strength requirements for multiple watermarks that share the same spatial context and different imperceptibility requirements based on frequency-weighted error and local texture perceptual models. We illustrate the effectiveness of the framework by designing a hierarchical semifragile watermark that is tolerant to mild compression, allows tamper localization, and is fragile under aggressive compression. Using a quad-tree representation, a spatial resolution hierarchy is established on the image and a watermark is embedded corresponding to each node of the hierarchy. The spatial hierarchy of watermarks provides a graceful tradeoff between robustness and localization under mild JPEG compression, where watermarks at coarser levels demonstrate progressively higher immunity to JPEG compression. Under aggressive compression, watermarks at all hierarchy levels vanish, indicating a lack of trust in the image data. The constraints implicitly partition watermark power in the resolution hierarchy as well as among image regions based on robustness and invisibility requirements. Experimental results illustrate the flexibility and effectiveness of the method Oktay Altun, Gaurav Sharma 0001, Mehmet Utku Celik, Mark F. Bocko |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2006 | Local Image Registration by Adaptive FilteringabstractWe propose a new adaptive filtering framework for local image registration, which compensates for the effect of local distortions/displacements without explicitly estimating a distortion/displacement field. To this effect, we formulate local image registration as a two-dimensional (2-D) system identification problem with spatially varying system parameters. We utilize a 2-D adaptive filtering framework to identify the locally varying system parameters, where a new block adaptive filtering scheme is introduced. We discuss the conditions under which the adaptive filter coefficients conform to a local displacement vector at each pixel. Experimental results demonstrate that the proposed 2-D adaptive filtering framework is very successful in modeling and compensation of both local distortions, such as Stirmark attacks, and local motion, such as in the presence of a parallax field. In particular, we show that the proposed method can provide image registration to: a) enable reliable detection of watermarks following a Stirmark attack in nonblind detection scenarios, b) compensate for lens distortions, and c) align multiview images with nonparametric local motion. Gulcin Caner, A. Murat Tekalp, Gaurav Sharma 0001, Wendi B. Heinzelman |
IEEE Trans. Image Process. | 3 |
| 2006 | Lossless watermarking for image authentication: a new framework and an implementationabstractWe present a novel framework for lossless (invertible) authentication watermarking, which enables zero-distortion reconstruction of the un-watermarked images upon verification. As opposed to earlier lossless authentication methods that required reconstruction of the original image prior to validation, the new framework allows validation of the watermarked images before recovery of the original image. This reduces computational requirements in situations when either the verification step fails or the zero-distortion reconstruction is not needed. For verified images, integrity of the reconstructed image is ensured by the uniqueness of the reconstruction procedure. The framework also enables public(-key) authentication without granting access to the perfect original and allows for efficient tamper localization. Effectiveness of the framework is demonstrated by implementing the framework using hierarchical image authentication along with lossless generalized-least significant bit data embedding. Mehmet Utku Celik, Gaurav Sharma 0001, A. Murat Tekalp |
IEEE Trans. Image Process. | 2 |
| 2005 | Morphological Steganalysis of Audio Signals and the Principle of Diminishing Marginal DistortionsabstractSteganographic methods attempt to insert data in multimedia signals in an undetectable fashion. However, these methods often disrupt the underlying signal characteristics, thereby allowing detection under careful steganalysis. Under repeated embedding, disruption of the signal characteristics is the highest for the first embedding and decreases subsequently. That is, the marginal distortions due to repeated embeddings decrease monotonically. We name this general principle as the principle of diminishing marginal distortions (DMD) and illustrate its validity in the audio domain using a morphological distortion metric. The principle of DMD is used to derive a steganalysis tool that detects the presence of hidden messages in uncompressed audio files. Detailed analysis and experimental results are provided for the detection of spread spectrum watermarking and stochastic modulation steganography. Oktay Altun, Gaurav Sharma 0001, Mehmet Utku Celik, Mark Sterling, Edward L. Titlebaum, Mark F. Bocko |
ICASSP (2) | 2 |
| 2005 | An adaptive filtering framework for image registrationabstractImage registration is a fundamental task in both image processing and computer vision. We present a novel method for local image registration based on adaptive filtering techniques. We utilize an adaptive filter to estimate and track correspondences among multiple images containing overlapping views of common scene regions. Image pixels are traversed in an order established by space-filling curves, to preserve the contiguity and hence track locally varying registration changes. The algorithm differs from pre-existing work on image registration in that it requires only local information and relatively low computational effort. These characteristics render the method suitable for deployment in imaging sensor networks, toward which the current work is directed. We evaluate the performance of the proposed algorithm using images captured with a digital camera in various real-world scenarios. Experimental results show that the proposed method can significantly improve accuracy and robustness over a global 2D parametric registration and can also outperform the local registration algorithm based on the Lucas-Kanade optical flow technique (Lucas, B. and Kanade, T., 1981). Gulcin Caner, A. Murat Tekalp, Gaurav Sharma 0001, Wendi B. Heinzelman |
ICASSP (2) | 3 |
| 2005 | Pitch and Duration Modification for Speech WatermarkingabstractWe propose a speech watermarking algorithm based on the modification of the pitch (fundamental frequency) and duration of the quasi-periodic speech segments. Natural variability of these speech features allows watermarking modifications to be imperceptible to the human observer. On the other hand, the significance of these features makes the system robust to common signal processing operations and low data-rate source excitation based speech coders. This class of coders is particularly obstructive for conventional audio watermarking algorithms when applied to speech signals. A pitch synchronous overlap and add (PSOLA) algorithm is used for pitch and duration modifications in the watermark embedding phase. Experiments with multiple speech codecs show very good robustness with low data-rate (5-8 kbps) speech coders. Mehmet Utku Celik, Gaurav Sharma 0001, A. Murat Tekalp |
ICASSP (2) | 2 |
| 2005 | Semifragile hierarchical watermarking in a set theoretic frameworkabstractWe introduce a set theoretic framework for watermarking and illustrate its effectiveness by designing a hierarchical semi-fragile watermark that is tolerant to compression and allows tamper localization. Using a quad-tree representation, a spatial resolution hierarchy is established on the image and a watermark is embedded corresponding to each node of the hierarchy. The watermarked image is determined so as to jointly satisfy the multiple constraints of watermark detectability, imperceptibility, and robustness to compression using the method of projections onto convex sets. The spatial hierarchy of watermarks provides a graceful trade-off between robustness and localization under JPEG compression: mild JPEG compression preserves watermarks at all levels of the hierarchy allowing fine localization of malicious changes while aggressive JPEG compression preserves watermarks at coarser levels of the hierarchy still assuring overall image integrity but giving up the capability for localization. Experimental results are presented to illustrate the effectiveness of the method. Oktay Altun, Gaurav Sharma 0001, Mehmet Utku Celik, Mark F. Bocko |
ICIP (1) | 2 |
| 2005 | Two-dimensional transforms for device color correction and calibrationabstractColor device calibration is traditionally performed using one-dimensional (1-D) per-channel tone-response corrections (TRCs). While 1-D TRCs are attractive in view of their low implementation complexity and efficient real-time processing of color images, their use severely restricts the degree of control that can be exercised along various device axes. A typical example is that per separation (or per-channel), TRCs in a printer can be used to either ensure gray balance along the C = M = Y axis or to provide a linear response in delta-E units along each of the individual (C, M, and Y) axis, but not both. This paper proposes a novel two-dimensional color correction architecture that enables much greater control over the device color gamut with a modest increase in implementation cost. Results show significant improvement in calibration accuracy and stability when compared to traditional 1-D calibration. Superior cost quality tradeoffs (over 1-D methods) are also achieved for emulation of one color device on another. Raja Bala, Gaurav Sharma 0001, Vishal Monga, Jean-Pierre Van de Capelle |
IEEE Trans. Image Process. | 2 |
| 2005 | Lossless generalized-LSB data embeddingabstractWe present a novel lossless (reversible) data-embedding technique, which enables the exact recovery of the original host signal upon extraction of the embedded information. A generalization of the well-known least significant bit (LSB) modification is proposed as the data-embedding method, which introduces additional operating points on the capacity-distortion curve. Lossless recovery of the original is achieved by compressing portions of the signal that are susceptible to embedding distortion and transmitting these compressed descriptions as a part of the embedded payload. A prediction-based conditional entropy coder which utilizes unaltered portions of the host signal as side-information improves the compression efficiency and, thus, the lossless data-embedding capacity. Mehmet Utku Celik, Gaurav Sharma 0001, A. Murat Tekalp, Eli Saber |
IEEE Trans. Image Process. | 2 |
| 2004 | Efficient classification of scanned media using spatial statisticsabstractWe address the automatic classification of scanned input media in order to improve color calibration. Since scanner responses vary significantly according to the type of input, a media dependent color calibration for a scanner is desirable for accurately mapping scanner responses to a standard color space. To assist such media dependent calibration, we propose an efficient algorithm for automated classification of input media into four major classes corresponding to photographic, lithographic, xerographic and inkjet. Our technique exploits the strong correlation between the type of input medium and the spatial statistics of corresponding images, which may be observed in the scanned images. Adopting two spatial statistical measures of dispersion and periodicity and utilizing extensive training data, we determine well separated decision regions to classify the input medium with a high confidence level. Experimental results over an independent test data set validate the results. Gozde Unal, Gaurav Sharma 0001, Reiner Eschbach |
ICIP | 2 |
| 2004 | Collusion-resilient fingerprinting by random pre-warpingabstractFingerprinting of audio-visual content using digital watermarks is an effective means of determining originators of unauthorized/pirated copies. Watermarks embedded in content can trace the traitor responsible for piracy. Multiple users may, however, collude and collectively escape identification by creating an average of their individually watermarked copies that appears unwatermarked. We propose a novel collusion-resilience mechanism, wherein the host signal is warped randomly prior to watermarking. As each copy undergoes a distinctive warp, collusion through averaging either yields low-quality results or requires substantial computational resources to undo random warps. The method is independent of the watermarking scheme used and imposes no restrictions on the watermark signal. We demonstrate the effectiveness of this approach on digital images. Mehmet Utku Celik, Gaurav Sharma 0001, A. Murat Tekalp |
IEEE Signal Process. Lett. | 2 |
| 2003 | Level-embedded lossless image compressionabstractA level-embedded lossless compression method for continuous-tone still images is presented. Level (bit-plane) scalability is achieved by separating the image into two layers before compression and excellent compression performance is obtained by exploiting both spatial and inter-level correlations. A comparison of the proposed scheme with a number of scalable and non-scalable lossless image compression algorithms is performed to benchmark its performance. The results indicate that the level-embedded compression incurs only a small penalty in compression efficiency. Mehmet Utku Celik, A. Murat Tekalp, Gaurav Sharma 0001 |
ICASSP (3) | 3 |
| 2003 | Collusion-resilient fingerprinting using random prewarpingabstractFingerprinting of audio-visual content using digital watermarks is an effective means of determining the originators of unauthorized copies and fighting piracy in digital distribution networks. In particular, watermarks embedded within the content help trace the traitor responsible for the piracy. A group of users may, however, collude and collectively escape identification by creating an average of their individually watermarked copies that appears unwatermarked. We propose a novel collusion-resilience mechanism, wherein the host signal is warped randomly prior to watermarking. As each copy undergoes a distinctive warp, collusion through averaging either yields low-quality results or requires substantial computational resources to undo random warps. The proposed method is independent of the watermarking scheme used and does not impose any restrictions on the watermark signal that are required by some collusion resistant watermarking schemes. We demonstrate the effectiveness of this approach on digital images. Mehmet Utku Celik, Gaurav Sharma 0001, A. Murat Tekalp |
ICIP (1) | 2 |
| 2003 | Gray-level-embedded lossless image compression
Mehmet Utku Celik, Gaurav Sharma 0001, A. Murat Tekalp |
Signal Process. Image Commun. | 2 |
| 2002 | Reversible data hidingabstractWe present a novel reversible (lossless) data hiding (embedding) technique, which enables the exact recovery of the original host signal upon extraction of the embedded information. A generalization of the well-known LSB (least significant bit) modification is proposed as the data embedding method, which introduces additional operating points on the capacity-distortion curve. Lossless recovery of the original is achieved by compressing portions of the signal that are susceptible to embedding distortion, and transmitting these compressed descriptions as a part of the embedded payload. A prediction-based conditional entropy coder which utilizes static portions of the host as side-information improves the compression efficiency, and thus the lossless data embedding capacity. Mehmet Utku Celik, Gaurav Sharma 0001, Eli Saber, A. Murat Tekalp |
ICIP (2) | 2 |
| 2002 | Hierarchical watermarking for secure image authentication with localizationabstractSeveral fragile watermarking schemes presented in the literature are either vulnerable to vector quantization (VQ) counterfeiting attacks or sacrifice localization accuracy to improve security. Using a hierarchical structure, we propose a method that thwarts the VQ attack while sustaining the superior localization properties of blockwise independent watermarking methods. In particular, we propose dividing the image into blocks in a multilevel hierarchy and calculating block signatures in this hierarchy. While signatures of small blocks on the lowest level of the hierarchy ensure superior accuracy of tamper localization, higher level block signatures provide increasing resistance to VQ attacks. At the top level, a signature calculated using the whole image completely thwarts the counterfeiting attack. Moreover, "sliding window" searches through the hierarchy enable the verification of untampered regions after an image has been cropped. We provide experimental results to demonstrate the effectiveness of our method. Mehmet Utku Celik, Gaurav Sharma 0001, Eli Saber, A. Murat Tekalp |
IEEE Trans. Image Process. | 2 |
| 2001 | A hierarchical image authentication watermark with improved localization and securityabstractSeveral fragile watermarking schemes presented in the literature are either vulnerable to vector quantization (VQ) counterfeiting attacks or sacrifice localization accuracy to improve security. Using a hierarchical structure, we propose a method that thwarts the VQ attack while sustaining the superior localization properties of blockwise independent watermarking methods. In particular, we propose dividing the image into blocks in a multi-level hierarchy and calculating block signatures in this hierarchy. While signatures of small blocks on the lowest level of the hierarchy ensure superior accuracy of tamper localization, higher level block signatures provide increasing resistance to VQ attacks. At the top level, a signature calculated using the whole image completely thwarts the counterfeiting attack. Moreover, "sliding window" searches through the hierarchy enable the verification of untampered regions after an image has been cropped. Mehmet Utku Celik, Gaurav Sharma 0001, Eli Saber, A. Murat Tekalp |
ICIP (2) | 2 |
| 2001 | Show-through cancellation in scans of duplex printed documentsabstractIn scanning pages with double-sided printing, often the printing on the back-side shows through in the scan of the front-side because the paper is not completely opaque. This show-through is an undesirable artifact that one would like to remove. In this paper, the phenomenon of show-through is analyzed using first physical principles to obtain a simplified mathematical model. The model is linearized using suitable transformations and simplifying approximations. Based on the linearized model, an adaptive linear filtering scheme is developed for the electronic removal of show-through using scans of both sides of the document. Experimental results demonstrating the effectiveness of the method developed are presented. Gaurav Sharma 0001 |
IEEE Trans. Image Process. | 1 |
| 2000 | Cancellation of Show-Through in Duplex ScanningabstractWhen scanning a page with printing on both sides, the printing on the back-side often shows through in the scan of the front-side because the page is not completely opaque. This phenomenon of show-through is analyzed and an image processing method for removal of this commonly encountered degradation is developed. A simplified mathematical model is obtained from first physical principles. The model is linearized using suitable transformations and simplifying approximations. Based on the linearized model, an adaptive linear-filtering scheme is developed for the electronic removal of show-through using scans of both sides of the document. Experimental results demonstrating the effectiveness of the method developed are presented. Gaurav Sharma 0001 |
ICIP | 1 |
| 1999 | End-to-end color printer calibration by total least squares regressionabstractNeugebauer modeling plays an important role in obtaining end-to-end device characterization profiles for halftone color printer calibration. This paper proposes total least square (TLS) regression methods to estimate the parameters of various Neugebauer models. Compared to the traditional least squares (LS) based methods, the TLS approach is physically more appropriate for the printer modeling problem because it accounts for errors in the measured reflectance of both the primaries and the modeled samples. A TLS method based on print measurements from single-colorant step-wedges is first developed. The method is then extended to incorporate multicolorant print measurements using an iterative algorithm. The LS and TLS techniques are compared through tests performed on two color printers, one employing conventional rotated halftone screens and the other using a dot-on-dot halftone screen configuration. Our experiments indicate that the TLS methods yield a consistent and significant improvement over the LS-based techniques for model parameter estimation. The gains from the TLS method are particularly significant when the number of patches for which measured data is available is limited. Minghui Xia, Eli Saber, Gaurav Sharma 0001, A. Murat Tekalp |
IEEE Trans. Image Process. | 3 |
| 1998 | Total Least Square Techniques in Color Printer CharacterizationabstractThe Neugebauer model is a powerful tool in obtaining end-to-end device characterization profiles for halftone color printer calibration. In this paper, we propose total least square (TLS) regression methods to estimate the parameters of Neugebauer models. Compared to the traditional least squares (LS) based methods, the TLS approach is a physically more appropriate procedure, because it accounts for errors in the measured reflectance of both the selected primaries and the modeled reflectance. The proposed TLS techniques are tested on a Xerox color printer with rotated halftone screen, and the results are compared with the LS based algorithms. Our experiments indicate that the TLS methods yield a significant improvement over the LS based techniques for model parameter estimation. Minghui Xia, Eli Saber, Gaurav Sharma 0001, A. Murat Tekalp, Ronald Sosinski |
ICIP (2) | 3 |
| 1998 | Color imaging for multimediaabstractTo a significant degree, multimedia applications derive their effectiveness from the use of color graphics, images, and video. However, the requirements for accurate color reproduction and for the preservation of this information across display and print devices that have very different characteristics and may be geographically apart are often not clearly understood. This paper describes the basics of color science, color input and output devices, color management, and calibration that help in defining and meeting these requirements. Gaurav Sharma 0001, Michael J. Vrhel, H. Joel Trussell |
Proc. IEEE | 1 |
| 1998 | Performance evaluation of burst-error-correcting codes on a Gilbert-Elliott channelabstractThe performance of single burst-error-correcting (BEC) codes used over bursty channels is evaluated. The channel is represented by the Gilbert-Elliott (1960, 1963) model, which has been used by numerous authors to evaluate the performance of random-error-correcting (REC) codes over bursty channels. Recursive expressions are derived, which are used in evaluating the probability of a codeword error. These expressions and an approximate closed-form expression are applied to the performance of a single (23,12) BEC code. Gaurav Sharma 0001, Amer A. Hassan, Ajay Dholakia |
IEEE Trans. Commun. | 1 |
| 1998 | Optimal nonnegative color scanning filtersabstractIn this correspondence, the problem of designing color scanning filters for multi-illuminant color recording is considered. The filter transmittances are determined from a minimum-mean-squared orthogonal tristimulus error criterion that minimizes the color error in estimates obtained from noisy recorded data. Nonnegativity constraints essential for physical realizability are imposed on the filter transmittances. In order to demonstrate the significant improvements obtained, the resulting filters are compared with suboptimal filters reported in earlier literature. Gaurav Sharma 0001, H. Joel Trussell, Michael J. Vrhel |
IEEE Trans. Image Process. | 1 |
| 1997 | Digital color imagingabstractThis paper surveys current technology and research in the area of digital color imaging. In order to establish the background and lay down terminology, fundamental concepts of color perception and measurement are first presented using vector-space notation and terminology. Present-day color recording and reproduction systems are reviewed along with the common mathematical models used for representing these devices. Algorithms for processing color images for display and communication are surveyed, and a forecast of research trends is attempted. An extensive bibliography is provided. Gaurav Sharma 0001, H. Joel Trussell |
IEEE Trans. Image Process. | 1 |
| 1997 | Figures of merit for color scannersabstractIn the design and evaluation of color scanners and cameras, it is useful to have a single figure of merit that closely agrees with perceived color accuracy. In the past, several measures of goodness for color scanning filters have been proposed to fulfil such a requirement. Most of the proposed measures have had shortcomings in that they are either based on error metrics in color spaces that are not perceptually uniform, or in that they do not take into account the effects of measurement noise. An extension of the most promising measure, based on linearized CIELAB space, is proposed to obtain a new figure of merit that has a high degree of perceptual relevance and also accounts for the varying noise performance of different filters. The paper also provides a common framework for the different figures of merit and a comprehensive comparison of their computational complexity and reliability. Gaurav Sharma 0001, H. Joel Trussell |
IEEE Trans. Image Process. | 1 |
| 1997 | Set theoretic signal restoration using an error in variables criterionabstractThe restoration of a signal degraded by a stochastic impulse response is formulated as a problem with uncertainties in both the measurements and the impulse response. The method of total least squares, and variants thereof, are effective techniques for solving this class of problems. However, unlike set theoretic estimation schemes, these methods do not allow the incorporation of other a priori information in the estimate. In this correspondence, two new sets motivated by total least squares are introduced for set theoretic estimation. The convexity of these sets is established and the projection operators onto these sets are given. Through simulations, the advantages of the new technique over conventional and older set theoretic schemes for restoration are demonstrated. Gaurav Sharma 0001, H. Joel Trussell |
IEEE Trans. Image Process. | 1 |
| 1996 | Restoration of uncertain blurs using an error in variables criterionabstractIn image-restoration problems involving a known blur, set-theoretic estimation provides an effective framework for incorporation of both noise statistics and a priori information. However, the sets formulated for use with known blurs are not suitable for situations involving unknown stochastic blurs. In this paper, new sets based on an error in variables criterion are developed, which are more appropriate for restoration of uncertain blurs. The conditions for convexity of these sets are rigorously established. Through 1-D simulations, the set-theoretic restoration scheme utilizing these sets is compared with the stochastic MMSE filter and with set-theoretic schemes for stochastic blurs developed earlier. Gaurav Sharma 0001, H. Joel Trussell |
ICIP (3) | 1 |
| 1994 | Decomposition of Fluorescent Illuminant Spectra for Accurate ColorimetryabstractFluorescent lamps are widely used in the office environment and in common desktop scanners. These lamps are characterized by the nearly monochromatic emission lines present in their radiant spectrum. Due to the presence of spectral lines special care is required in the computation of tristimulus values under these illuminants. Accurate computation of the tristimulus values can be performed digitally by modeling the radiance as the sum of a continuous spectrum emitted by the phosphors and weighted impulses corresponding to the emission lines. The paper considers a scheme for estimating the continuous spectrum and the strengths of the impulses from measurements of the spectrum with band limited sensors. The impact of noise in measured data on the accuracy of the final color computation is also considered.> Gaurav Sharma 0001, H. Joel Trussell |
ICIP (2) | 1 |
| 1994 | A DFT based alternating projection algorithm for parameter estimation of superimposed complex sinusoids
Gaurav Sharma 0001, V. U. Reddy |
Signal Process. | 1 |