Prakash Ishwar

dblp:61/5637 · DBLP profile ↗
← Back
80ranked-venue papers
12as first author
9since 2021 · last 2024
0000-0002-2621-1549ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 8 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 first-authorArtificial intelligence and machine learning · 10 · 4 since 2021Theory of computation · 10 · 1 first-authorComputer networks · 2 · 1 first-authorSecurity and privacy · 2Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Supervised Contrastive Learning with Hard Negative Samples
abstract
Through minimization of an appropriate loss function such as the InfoNCE loss, contrastive learning (CL) learns a useful representation function by pulling positive samples close to each other while pushing negative samples far apart in the embedding space. The positive samples are typically created using "label-preserving" augmentations, i.e., domain-specific transformations of a given datum or anchor. In absence of class information, in unsupervised CL (UCL), the negative samples are typically chosen randomly and independently of the anchor from a preset negative sampling distribution over the entire dataset. This leads to class-collisions in UCL. Supervised CL (SCL), avoids this class collision by conditioning the negative sampling distribution to samples having labels different from that of the anchor. In hard-UCL (H-UCL), which has been shown to be an effective method to further enhance UCL, the negative sampling distribution is conditionally tilted, by means of a hardening function, towards samples that are closer to the anchor. Motivated by this, in this paper we propose hard-SCL (H-SCL) wherein the class conditional negative sampling distribution is tilted via a hardening function. Our simulation results confirm the utility of H-SCL over SCL with significant performance gains in downstream classification tasks. Analytically, we show that in the limit of infinite negative samples per anchor and a suitable assumption, the H-SCL loss is upper bounded by the H-UCL loss, thereby justifying the utility of H-UCL for controlling the H-SCL loss in the absence of label information. Through experiments on several datasets, we verify the assumption as well as the claimed inequality between H-UCL and H-SCL losses. We also provide a plausible scenario where H-SCL loss is lower bounded by UCL loss, indicating the limited utility of UCL in controlling the H-SCL loss.1
Ruijie Jiang, Thuan Nguyen 0001, Prakash Ishwar, Shuchin Aeron
IJCNN3
2023 A Principled Approach to Model Validation in Domain Generalization
abstract
Domain generalization aims to learn a model with good generalization ability, that is, the learned model should not only perform well on several seen domains but also on unseen domains with different data distributions. State-of-the-art domain generalization methods typically train a representation function followed by a classifier jointly to minimize both the classification risk and the domain discrepancy. However, when it comes to model selection, most of these methods rely on traditional validation routines that select models solely based on the lowest classification risk on the validation set. In this paper, we theoretically demonstrate a trade-off between minimizing classification risk and mitigating domain discrepancy, i.e., it is impossible to achieve the minimum of these two objectives simultaneously. Motivated by this theoretical result, we propose a novel model selection method suggesting that the validation process should account for both the classification risk and the domain discrepancy. We validate the effectiveness of the proposed method by numerical results on several domain generalization datasets.
Boyang Lyu, Thuan Nguyen 0001, Matthias Scheutz, Prakash Ishwar, Shuchin Aeron
ICASSP4
2023 Hard Negative Sampling via Regularized Optimal Transport for Contrastive Representation Learning
abstract
We study the problem of designing hard negative sampling distributions for unsupervised contrastive representation learning. We propose and analyze a novel min-max framework that seeks a representation which minimizes the maximum (worst-case) generalized contrastive learning loss over all couplings (joint distributions between positive and negative samples subject to marginal constraints) and prove that the resulting min-max optimum representation will be degenerate. This provides the first theoretical justification for incorporating additional regularization constraints on the couplings. We re-interpret the min-max problem through the lens of Optimal Transport (OT) theory and utilize regularized transport couplings to control the degree of hardness of negative examples. Through experiments we demonstrate that the negative samples generated from our designed negative distribution are more similar to the anchor than those generated from the baseline negative distribution. We also demonstrate that entropic regularization yields negative sampling distributions with parametric form similar to that in a recent state-of-the-art negative sampling design and has similar performance in multiple datasets. Utilizing the uncovered connection with OT, we propose a new ground cost for designing the negative distribution and show improved performance of the learned representation on downstream tasks compared to the representation learned when using squared Euclidean cost.11Our code is publicly available at https://github.com/rjiang03/HCL-OT.
Ruijie Jiang, Prakash Ishwar, Shuchin Aeron
IJCNN2
2022 FRIDA: Fisheye Re-Identification Dataset with Annotations
abstract
Person re-identification (PRID) from side-mounted rectilinear-lens cameras is a well-studied problem. On the other hand, PRID from overhead fisheye cameras is new and largely unstudied, primarily due to the lack of suitable image datasets. To fill this void, we introduce the “Fisheye Re-IDentification Dataset with Annotations” (FRIDA)1, with 240k+ bounding-box annotations of people, captured by 3 time-synchronized, ceiling-mounted fisheye cameras in a large indoor space. Due to a field-of-view overlap, PRID in this case differs from a typical PRID problem, which we discuss in depth. We also evaluate the performance of 10 state-of-the-art PRID algorithms on FRIDA. We show that for 6 CNN-based algorithms, training on FRIDA boosts the performance by up to 11.64% points in mAP compared to training on a common rectilinear-camera PRID dataset.1vip. bu.edu/frida
Mertcan Cokbas, John Bolognino, Janusz Konrad, Prakash Ishwar
AVSS4
2022 Trade-off between reconstruction loss and feature alignment for domain generalization
abstract
Domain generalization (DG) is a branch of transfer learning that aims to train the learning models on several seen domains and subsequently apply these pre-trained models to other unseen (unknown but related) domains. To deal with challenging settings in DG where both data and label of the unseen domain are not available at training time, the most common approach is to design the classifiers based on the domain-invariant representation features, i.e., the latent representations that are unchanged and transferable between domains. Contrary to popular belief, we show that designing classifiers based on invariant representation features alone is necessary but insufficient in DG. Our analysis indicates the necessity of imposing a constraint on the reconstruction loss induced by representation functions to preserve most of the relevant information about the label in the latent space. More importantly, we point out the trade-off between minimizing the reconstruction loss and achieving domain alignment in DG. Our theoretical results motivate a new DG framework that jointly optimizes the reconstruction loss and the domain discrepancy. Both theoretical and numerical results are provided to justify our approach.
Thuan Nguyen 0001, Boyang Lyu, Prakash Ishwar, Matthias Scheutz, Shuchin Aeron
ICMLA3
2022 Conditional entropy minimization principle for learning domain invariant representation features
abstract
Invariance-principle-based methods such as Invariant Risk Minimization (IRM), have recently emerged as promising approaches for Domain Generalization (DG). Despite promising theory, such approaches fail in common classification tasks due to mixing of true invariant features and spurious invariant features1. To address this, we propose a framework based on the conditional entropy minimization (CEM) principle to filter-out the spurious invariant features leading to a new algorithm with a better generalization capability. We show that our proposed approach is closely related to the well-known Information Bottleneck (IB) framework and prove that under certain assumptions, entropy minimization can exactly recover the true invariant features. Our approach provides competitive classification accuracy compared to recent theoretically-principled state-of-the-art alternatives across several DG datasets.
Thuan Nguyen 0001, Boyang Lyu, Prakash Ishwar, Matthias Scheutz, Shuchin Aeron
ICPR3
2022 WEPDTOF: A Dataset and Benchmark Algorithms for In-the-Wild People Detection and Tracking from Overhead Fisheye Cameras
abstract
Owing to their large field of view, overhead fisheye cameras are becoming a surveillance modality of choice for large indoor spaces. However, traditional people detection and tracking algorithms developed for side-mounted, rectilinear-lens cameras do not work well on images from overhead fisheye cameras due to their viewpoint and unique optics. While several people-detection algorithms have been recently developed for such cameras, they have all been tested on datasets consisting of "staged" recordings with a limited variety of people, scenes and challenges. Clearly, the performance of these algorithms "in the wild", i.e., on recordings with real-world challenges, remains un-known. In this paper, we introduce a new benchmark dataset of in-the-Wild Events for People Detection and Tracking from Overhead Fisheye cameras (WEPDTOF)1. The dataset features 14 YouTube videos captured in a wide range of scenes, 188 distinct person identities consistently labeled across time, and real-world challenges such as extreme occlusions and camouflage. Also, we propose 3 spatiotemporal extensions2of a state-of-the-art people-detection algorithm to enhance the coherence of detections across time. Compared to top-performing algorithms, that are purely spatial, the new algorithms offer a significant performance improvement on the new dataset. Finally, we compare the people tracking performance of these algorithms on WEPDTOF.
Mustafa Ozan Tezcan, Zhihao Duan, Mertcan Cokbas, Prakash Ishwar, Janusz Konrad
WACV4
2021 Geometry-Based Person Re-Identification in Fisheye Stereo
abstract
Person re-identification using rectilinear cameras has been thoroughly researched to date. However, the topic has received little attention for fisheye cameras and the few developed methods are appearance-based. We propose a geometry-based approach to re-identification for overhead fisheye cameras with overlapping fields of view. The main idea is that a person visible in two camera views is uniquely located in the view of one camera given their height and location in the other camera’s view. We develop a height-dependent mathematical relationship between these locations using the unified spherical model for omnidirectional cameras. We also propose a new fisheye-camera calibration method and a novel automated approach to calibration-data collection. Finally, we propose four re-identification algorithms that leverage geometric constraints and demonstrate their excellent accuracy, which vastly exceeds that of a state-of-the-art appearance-based method, on a fisheye-camera dataset we collected.
Joshua Bone, Mertcan Cokbas, Mustafa Ozan Tezcan, Janusz Konrad, Prakash Ishwar
AVSS5
2021 PETS2021: Through-foliage detection and tracking challenge and evaluation
abstract
This paper presents the outcomes of the PETS2021 challenge held in conjunction with AVSS2021 and sponsored by the EU FOLDOUT project. The challenge comprises the publication of a novel video surveillance dataset on through-foliage detection, the defined challenges addressing person detection and tracking in fragmented occlusion scenarios, and quantitative and qualitative performance evaluation of challenge results submitted by six worldwide participants. The results show that while several detection and tracking methods achieve overall good results, through-foliage detection and tracking remains a challenging task for surveillance systems especially as it serves as the input to behaviour (threat) recognition.
Jose Luis Patino, Jonathan N. Boyle, James M. Ferryman, Jonas Auer, Julian Pegoraro, Roman P. Pflugfelder, Mertcan Cokbas, Janusz Konrad, Prakash Ishwar, Giulia Slavic, Lucio Marcenaro, Yifan Jiang 0002, Youngsaeng Jin, Hanseok Ko, Guangliang Zhao, Guy Ben-Yosef, Jianwei Qiu
AVSS9
2020 Multi-Label and Multilingual News Framing Analysis
abstract
News framing refers to the practice in which aspects of specific issues are highlighted in the news to promote a particular interpretation.In NLP, although recent works have studied framing in English news, few have studied how the analysis can be extended to other languages and in a multi-label setting.In this work, we explore multilingual transfer learning to detect multiple frames from just the news headline in a genuinely low-resource context where there are few/no frame annotations in the target language.We propose a novel method that can leverage elementary resources consisting of a dictionary and few annotations to detect frames in the target language.Our method performs comparably or better than translating the entire target language headline to the source language for which we have annotated data.This work opens up an exciting new capability of scaling up frame analysis to many languages, even those without existing translation technologies.Lastly, we apply our method to detect frames on the issue of U.S. gun violence in multiple languages and obtain exciting insights on the relationship between different frames of the same problem across different countries with different languages.
Afra Feyza Akyürek, Lei Guo 0017, Randa I. Elanwar, Prakash Ishwar, Margrit Betke, Derry Wijaya
ACL4
2020 BSUV-Net: A Fully-Convolutional Neural Network for Background Subtraction of Unseen Videos
abstract
Background subtraction is a basic task in computer vision and video processing often applied as a pre-processing step for object tracking, people recognition, etc. Recently, a number of successful background-subtraction algorithms have been proposed, however nearly all of the top-performing ones are supervised. Crucially, their success relies upon the availability of some annotated frames of the test video during training. Consequently, their performance on completely "unseen" videos is undocumented in the literature. In this work, we propose a new, supervised, background-subtraction algorithm for unseen videos (BSUV-Net) based on a fully-convolutional neural network. The input to our network consists of the current frame and two background frames captured at different time scales along with their semantic segmentation maps. In order to reduce the chance of overfitting, we also introduce a new data-augmentation technique which mitigates the impact of illumination difference between the background frames and the current frame. On the CDNet-2014 dataset, BSUV-Net outperforms stateof-the-art algorithms evaluated on unseen videos in terms of several metrics including F-measure, recall and precision.
Mustafa Ozan Tezcan, Prakash Ishwar, Janusz Konrad
WACV2
2019 Supervised People Counting Using An Overhead Fisheye Camera
abstract
We propose two supervised methods for people counting using an overhead fisheye camera. As opposed to standard cameras, fisheye cameras offer a large field of view and, when mounted overhead, reduce occlusions. However, methods developed for standard cameras perform poorly on fisheye images since they do not account for the radial image geometry. Furthermore, no large-scale fisheye-image datasets with radially-aligned bounding box annotations are available for training. We adapt YOLOv3 trained on standard images for people counting in fisheye images. In one method, YOLOv3 is applied to 24 rotated, overlapping windows and the results are post-processed to produce a people count. In another method, YOLOv3 is applied to windows of interest extracted by background subtraction. For evaluation, we collected and annotated an indoor fisheye-image dataset that we make public. Experiments on this dataset show that our methods reduce the people counting MAE of two natural benchmarks by over 60%.
Shengye Li, Mustafa Ozan Tezcan, Prakash Ishwar, Janusz Konrad
AVSS3
2019 CNN-Based Indoor Occupant Localization via Active Scene Illumination
abstract
We propose and study a data-driven approach to indoor occupant localization using a network of single-pixel light sensors and modulated LED light sources. Locations are estimated by processing sensor data using a simple convolutional neural network (CNN). Unlike previous model-based methods, the proposed approach does not require knowledge of room dimensions, locations of LEDs and sensors, and assumptions about material properties and object heights. We quantitatively validate the performance of our approach in simulated and real-world environments in private and public scenarios. In Unity3D simulations, compared to the best-performing benchmark method, our approach reduces the average localization error by 47.69% in private scenarios and by 46.99% in public scenarios. Similarly, in a real testbed the error is reduced by 36.54% and 11.46% in private and public scenarios respectively.
Jinyuan Zhao, Natalia Frumkin, Prakash Ishwar, Janusz Konrad
ICIP3
2017 Node embedding for network community discovery
abstract
Neural node embedding has been recently developed as a powerful representation for supervised tasks with graph data. We leverage this recent advance and propose a novel approach for unsupervised community discovery in graphs. Through extensive experimental studies on simulated and real-world data, we demonstrate consistent improvement of the proposed approach over the current state-of-the-arts. Specifically, our approach empirically attains the information theoretic limits under the benchmark Stochastic Block Models and exhibits better stability and accuracy over the best known algorithms in the community recovery limits.
Christy Lin, Prakash Ishwar, Weicong Ding
ICASSP2
2017 On methods for privacy-preserving energy disaggregation
abstract
Household energy monitoring via smart-meters motivates the problem of disaggregating the total energy usage signal into the component energy usage and operating patterns of individual appliances. While energy disaggregation enables useful analytics, it also raises privacy concerns because sensitive household information may also be revealed. Our goal is to preserve analytical utility while mitigating privacy concerns by processing the total energy usage signal. We consider processing methods that attempt to remove the contribution of a set of sensitive appliances from the total energy signal. We show that while a simple model-based approach is effective against an adversary making the same model assumptions, it is much less effective against a stronger adversary employing neural networks in an inference attack. We also investigate the performance of employing neural networks to estimate and remove the energy usage of sensitive appliances. The experiments used the publicly available UK-DALE dataset that was collected from actual households.
Ye Wang 0001, Nisarg Raval, Prakash Ishwar, Mitsuhiro Hattori, Takato Hirano, Nori Matsuda, Rina Shimizu
ICASSP3
2017 Privacy-preserving indoor localization via light transport analysis
abstract
We propose a system for indoor localization using intensity-controllable LED light fixtures and light sensors mounted on the ceiling. While providing accurate location estimates, our approach preserves user privacy and is robust to ambient light conditions. We develop a LASSO algorithm and a localized ridge regression algorithm for locating a single object. In synthetic experiments, our localized ridge regression algorithm achieves an average localization error ranging from 0.24in to 1.39in, for different object sizes, in a 7×12-foot room. The localized ridge regression algorithm also shows the ability to locate multiple objects in experiments with a real-world occupancy scenario.
Jinyuan Zhao, Prakash Ishwar, Janusz Konrad
ICASSP2
2017 Semi-Coupled Two-Stream Fusion ConvNets for Action Recognition at Extremely Low Resolutions
abstract
Deep convolutional neural networks (ConvNets) have been recently shown to attain state-of-the-art performance for action recognition on standard-resolution videos. However, less attention has been paid to recognition performance at extremely low resolutions (eLR) (e.g., 16 12 pixels). Reliable action recognition using eLR cameras would address privacy concerns in various application environments such as private homes, hospitals, nursing/rehabilitation facilities, etc. In this paper, we propose a semi-coupled, filter-sharing network that leverages high-resolution (HR) videos during training in order to assist an eLR ConvNet. We also study methods for fusing spatial and temporal ConvNets customized for eLR videos in order to take advantage of appearance and motion information. Our method outperforms state-of-the-art methods at extremely low resolutions on IXMAS (93:7%) and HMDB (29:2%) datasets.
Jiawei Chen 0006, Jon Wu, Janusz Konrad, Prakash Ishwar
WACV4
2016 Privacy-preserving, indoor occupant localization using a network of single-pixel sensors
abstract
We propose an approach to indoor occupant localization using a network of single-pixel, visible-light sensors. In addition to preserving privacy, our approach vastly reduces data transmission rate and is agnostic to eavesdropping. We develop two purely data-driven localization algorithms and study their performance using a network of 6 such sensors. In one algorithm, we divide the monitored floor area (2.37m×2.72m) into a 3×3 grid of cells and classify location of a single person as belonging to one of the 9 cells using a support vector machine classifier. In the second algorithm, we estimate person's coordinates using support vector regression. In cross-validation tests in public (e.g., conference room) and private (e.g., home) scenarios, we obtain 67–72% correct classification rate for cells and 0.31–0.35m mean absolute distance error within the monitored space. Given the simplicity of sensors and processing, these are encouraging results and can lead to useful applications today.
Douglas Roeper, Jiawei Chen 0006, Janusz Konrad, Prakash Ishwar
AVSS4
2015 A Topic Modeling Approach to Ranking
abstract
We propose a topic modeling approach to the prediction of preferences in pairwise comparisons. We develop a new generative model for pairwise comparisons that accounts for multiple shared latent rankings that are prevalent in a population of users. This new model also captures inconsistent user behavior in a natural way. We show how the estimation of latent rankings in the new generative model can be formally reduced to the estimation of topics in a statistically equivalent topic modeling problem. We leverage recent advances in the topic modeling literature to develop an algorithm that can learn shared latent rankings with provable consistency as well as sample and computational complexity guarantees. We demonstrate that the new approach is empirically competitive with the current state-of-the-art approaches in predicting preferences on some semi-synthetic and real world datasets.
Weicong Ding, Prakash Ishwar, Venkatesh Saligrama
AISTATS2
2015 Learning shared rankings from mixtures of noisy pairwise comparisons
abstract
We propose a novel model for rank aggregation from pairwise comparisons which accounts for a heterogeneous population of inconsistent users whose preferences are different mixtures of multiple shared ranking schemes. By connecting this problem to recent advances in the non-negative matrix factorization (NMF) literature, we develop an algorithm that can learn the underlying shared rankings with provable statistical and computational efficiency guarantees. We validate the approach using semi-synthetic and real world datasets.
Weicong Ding, Prakash Ishwar, Venkatesh Saligrama
ICASSP2
2015 Towards privacy-preserving recognition of human activities
abstract
A smart room of the future is expected to facilitate intelligent interaction with its occupants while respecting their privacy. Although standard video cameras can be used to learn where the occupants are and what they do, they raise privacy concerns. While this can be mitigated by severely reducing camera resolution, it will also impact the utility of the camera network. This work investigates and quantifies the tradeoff between camera resolution and action recognition accuracy. Rather than building a physical testbed to carry out this study, we use a graphics engine to simulate a room with 5 cameras, and to animate avatars using skeletal movements of real users captured by a Kinect v2 camera. We study resolutions from 100×100 pixels down to 1×1 using a state-of-the-art action recognition method at higher resolutions and we propose a new approach at ultra-low resolutions. In extensive simulations, we conclude that on a dataset of 12 individuals performing 4 actions our algorithm applied to single-pixel data performs very close to the state-of-the-art method applied to 100×100 data, suggesting that reliable action recognition can be achieved without compromising occupant's identity.
Ji Dai, Behrouz Saghafi, Jon Wu, Janusz Konrad, Prakash Ishwar
ICIP5
2015 Leveraging shape and depth in user authentication from in-air hand gestures
abstract
Depth-sensors, such as the Kinect, have predominately been used as a gesture recognition device. Recent works, however, have proposed using these sensors for user authentication using biometric modalities such as: face, speech, gait and gesture. The last of these modalities - gestures, used in the context of full-body and hand-based gestures, is relatively new but has shown promising authentication performance. In this paper, we focus on hand-based gestures that are performed in-air. We present a novel approach to user authentication from such gestures by leveraging a temporal hierarchy of depth-aware silhouette covariances. Further, we investigate the usefulness of shape and depth information in this modality, as well as the importance of hand movement when performing a gesture. By exploiting both shape and depth information our method attains an average 1.92% Equal Error Rate (EER) on a dataset of 21 users across 4 predefined hand-gestures. Our method consistently outperforms related methods on this dataset.
Jon Wu, James Christianson, Janusz Konrad, Prakash Ishwar
ICIP4
2015 Learning-Based Object Identification and Segmentation Using Dual-Energy CT Images for Security
abstract
In recent years, baggage screening at airports has included the use of dual-energy X-ray computed tomography (DECT), an advanced technology for nondestructive evaluation. The main challenge remains to reliably find and identify threat objects in the bag from DECT data. This task is particularly hard due to the wide variety of objects, the high clutter, and the presence of metal, which causes streaks and shading in the scanner images. Image noise and artifacts are generally much more severe than in medical CT and can lead to splitting of objects and inaccurate object labeling. The conventional approach performs object segmentation and material identification in two decoupled processes. Dual-energy information is typically not used for the segmentation, and object localization is not explicitly used to stabilize the material parameter estimates. We propose a novel learning-based framework for joint segmentation and identification of objects directly from volumetric DECT images, which is robust to streaks, noise and variability due to clutter. We focus on segmenting and identifying a small set of objects of interest with characteristics that are learned from training images, and consider everything else as background. We include data weighting to mitigate metal artifacts and incorporate an object boundary field to reduce object splitting. The overall formulation is posed as a multilabel discrete optimization problem and solved using an efficient graph-cut algorithm. We test the method on real data and show its potential for producing accurate labels of the objects of interest without splits in the presence of metal and clutter.
Limor Martin, Ahmet Tuysuzoglu, W. Clem Karl, Prakash Ishwar
IEEE Trans. Image Process.4
2014 Efficient Distributed Topic Modeling with Provable Guarantees
abstract
Topic modeling for large-scale distributed web-collections requires distributed techniques that account for both computational and communication costs. We consider topic modeling under the separability assumption and develop novel computationally efficient methods that provably achieve the statistical performance of the state-of-the-art centralized approaches while requiring insignificant communication between the distributed document collections. We achieve tradeoffs between communication and computation without actually transmitting the documents. Our scheme is based on exploiting the geometry of normalized word-word co-occurrence matrix and viewing each row of this matrix as a vector in a high-dimensional space. We relate the solid angle subtended by extreme points of the convex hull of these vectors to topic identities and construct distributed schemes to identify topics.
Weicong Ding, Mohammad H. Rohban, Prakash Ishwar, Venkatesh Saligrama
AISTATS3
2014 Silhouettes versus skeletons in gesture-based authentication with Kinect
abstract
Since its release, the Kinect has been successfully used in gesture recognition. Recent work has extended Kinect's use towards biometric user authentication based on face, speech, gait, and gestures. Our work expands on the last of these modalities - gestures, which have yielded promising authentication results in prior work. This paper aims to gain insight into how authentication methods that are based on silhouette features compare against those that are based on skeletal features in terms of trade-offs between authentication performance and robustness against some real-world degradations. On a dataset of 40 users that contains two types of degradations namely, user-memory and personal-effects (heavy coats, bags, etc.), we found that for user-defined gestures, skeletal features outperform silhouettes on average by 4.89% in terms of the Equal Error Rate (EER).
Jon Wu, Prakash Ishwar, Janusz Konrad
AVSS2
2014 Sensing-aware kernel SVM
abstract
We propose a novel approach for designing kernels for support vector machines (SVMs) when the class label is linked to the observation through a latent state and the likelihood function of the observation given the state (the sensing model) is available. We show that the Bayes-optimum decision boundary is a hyperplane under a mapping defined by the likelihood function. Combining this with the maximum margin principle yields kernels for SVMs that leverage knowledge of the sensing model in an optimal way. We derive the optimum kernel for the bag-of-words (BoWs) sensing model and demonstrate its superior performance over other kernels in document and image classification tasks. These results indicate that such optimum sensing-aware kernel SVMs can match the performance of rather sophisticated state-of-the-art approaches.
Weicong Ding, Prakash Ishwar, Venkatesh Saligrama, W. Clem Karl
ICASSP2
2014 Structure-preserving dual-energy CT for luggage screening
abstract
We propose a new structure-preserving dual-energy (SPDE) CT inversion technique for luggage screening, which can mitigate metal artifacts and provide precise object localization. Such artifact reduction can increase material identification accuracy in security applications. Our main objective is formation of enhanced photoelectric and Compton pixel property images from dual-energy X-ray tomographic data. We achieve this aim by incorporating three important elements in a single unified framework. First, we generate our images as the solution of a joint optimization problem, which explicitly models the projection process. Second, we include metal aware data weighting to reduce streaks and metal artifacts. Third, we estimate a regularized joint boundary field and apply it to both the photoelectric and Compton images in order to improve object localization as well as smoothing inside the objects. We evaluate the performance of the method using real dual-energy data. We demonstrate a significant reduction in noise and metal artifacts.
Limor Martin, W. Clem Karl, Prakash Ishwar
ICASSP3
2014 The value of posture, build and dynamics in gesture-based user authentication
abstract
User authentication based on biometrics such as fingerprint, iris, face, speech or gait has been around for many years. Recently, intentional user gestures have been shown to be a promising modality for user authentication. However, it is unclear how much of the performance can be attributed to pure biometric information that a user has no control over, such as individual limb lengths, and how much to the gesture dynamics, that a user can fully control. A related question is: How easy is it to copy these dynamics? In this paper, we propose a framework to decompose a gesture into three components: initial posture, limb proportions, and gesture dynamics. We then study the impact of each component and various component combinations on the performance of gesture-based user authentication using a dataset of 36 users performing 3 gestures of varying complexity. We also study spoof attacks using the same dataset and show, somewhat surprisingly, that amateurs are unable to copy gestures with sufficient accuracy so as to significantly degrade the overall authentication performance even when they are trained on users that they are closest to. While training certainly improves an attacker's ability to copy gesture dynamics, it seems that the unique limb proportions (which cannot be altered) and the initial posture (which amateurs attackers fail to pay attention to), more than make up for the loss due to compromised dynamics (which can always be renewed).
Jon Wu, Prakash Ishwar, Janusz Konrad
IJCB2
2014 An elementary completeness proof for secure two-party computation primitives
Ye Wang 0001, Prakash Ishwar, Shantanu Rane
ITW2
2014 A Novel Video Dataset for Change Detection Benchmarking
abstract
Change detection is one of the most commonly encountered low-level tasks in computer vision and video processing. A plethora of algorithms have been developed to date, yet no widely accepted, realistic, large-scale video data set exists for benchmarking different methods. Presented here is a unique change detection video data set consisting of nearly 90 000 frames in 31 video sequences representing six categories selected to cover a wide range of challenges in two modalities (color and thermal infrared). A distinguishing characteristic of this benchmark video data set is that each frame is meticulously annotated by hand for ground-truth foreground, background, and shadow area boundaries-an effort that goes much beyond a simple binary label denoting the presence of change. This enables objective and precise quantitative comparison and ranking of video-based change detection algorithms. This paper discusses various aspects of the new data set, quantitative performance metrics used, and comparative results for over two dozen change detection algorithms. It draws important conclusions on solved and remaining issues in change detection, and describes future challenges for the scientific community. The data set, evaluation tools, and algorithm rankings are available to the public on a website and will be updated with feedback from academia and industry in the future.
Nil Goyette, Pierre-Marc Jodoin, Fatih Porikli, Janusz Konrad, Prakash Ishwar
IEEE Trans. Image Process.5
2013 A new geometric approach to latent topic modeling and discovery
abstract
A new geometrically-motivated algorithm for topic modeling is developed and applied to the discovery of latent “topics” in text and image “document” corpora. The algorithm is based on robustly finding and clustering extreme-points of empirical cross-document word-frequencies that correspond to novel words unique to each topic. In contrast to related approaches that are based on solving non-convex optimization problems using suboptimal approximations, locally-optimal methods, or heuristics, the new algorithm is convex, has polynomial complexity, and has competitive qualitative and quantitative performance compared to the current state- of-the-art approaches on synthetic and real-world datasets.
Weicong Ding, Mohammad H. Rohban, Prakash Ishwar, Venkatesh Saligrama
ICASSP3
2013 Dynamic time warping for gesture-based user identification and authentication with Kinect
abstract
The Kinect has primarily been used as a gesture-driven device for motion-based controls. To date, Kinect-based research has predominantly focused on improving tracking and gesture recognition across a wide base of users. In this paper, we propose to use the Kinect for biometrics; rather than accommodating a wide range of users we exploit each user's uniqueness in terms of gestures. Unlike pure biometrics, such as iris scanners, face detectors, and fingerprint recognition which depend on irrevocable biometric data, the Kinect can provide additional revocable gesture information. We propose a dynamic time-warping (DTW) based framework applied to the Kinect's skeletal information for user access control. Our approach is validated in two scenarios: user identification, and user authentication on a dataset of 20 individuals performing 8 unique gestures. We obtain an overall 4.14%, and 1.89% Equal Error Rate (EER) in user identification, and user authentication, respectively, for a gesture and consistently outperform related work on this dataset. Given the natural noise present in the real-time depth sensor this yields promising results.
Jon Wu, Janusz Konrad, Prakash Ishwar
ICASSP3
2013 Topic Discovery through Data Dependent and Random Projections
abstract
We present algorithms for topic modeling based on the geometry of cross-document word-frequency patterns. This perspective gains significance under the so called separability condition. This is a condition on existence of novel-words that are unique to each topic. We present a suite of highly efficient algorithms with provable guarantees based on data-dependent and random projections to identify novel words and associated topics. Our key insight here is that the maximum and minimum values of cross-document frequency patterns projected along any direction are associated with novel words. While our sample complexity bounds for topic recovery are similar to the state-of-art, the computational complexity of our random projection scheme scales linearly with the number of documents and the number of words per document. We present several experiments on synthetic and realworld datasets to demonstrate qualitative and quantitative merits of our scheme.
Weicong Ding, Mohammad H. Rohban, Prakash Ishwar, Venkatesh Saligrama
ICML (3)3
2013 Information-theoretically secure three-party computation with One corrupted party
abstract
The problem in which one of three pairwise interacting parties is required to securely compute a function of the inputs held by the other two, when one party may arbitrarily deviate from the computation protocol (active behavioral model), is studied. An information-theoretic characterization of unconditionally secure computation protocols under the active behavioral model is provided. A protocol for Hamming distance computation is provided and shown to be unconditionally secure under both active and passive behavioral models using the information-theoretic characterization. The difference between the notions of security under the active and passive behavioral models is illustrated by examining a protocol for computing quadratic and Hamming distances that is secure under the passive model, but is insecure under the active model.
Ye Wang 0001, Prakash Ishwar, Shantanu Rane
ISIT2
2013 An impossibility result for high dimensional supervised learning
abstract
We study high-dimensional asymptotic performance limits of binary supervised classification problems where the class conditional densities are Gaussian with unknown means and covariances and the number of signal dimensions scales faster than the number of labeled training samples. We show that the Bayes error, namely the minimum attainable error probability with complete distributional knowledge and equally likely classes, can be arbitrarily close to zero and yet the limiting minimax error probability of every supervised learning algorithm is no better than a random coin toss. In contrast to related studies where the classification difficulty (Bayes error) is made to vanish, we hold it constant when taking high-dimensional limits. In contrast to VC-dimension based minimax lower bounds that consider the worst case error probability over all distributions that have a fixed Bayes error, our worst case is over the family of Gaussian distributions with constant Bayes error. We also show that a nontrivial asymptotic minimax error probability can only be attained for parametric subsets of zero measure (in a suitable measure space). These results expose the fundamental importance of prior knowledge and suggest that unless we impose strong structural constraints, such as sparsity, on the parametric space, supervised learning may be ineffective in high dimensional small sample settings.
Mohammad H. Rohban, Prakash Ishwar, Burkay Orten, W. Clem Karl, Venkatesh Saligrama
ITW2
2013 Action Recognition From Video Using Feature Covariance Matrices
abstract
We propose a general framework for fast and accurate recognition of actions in video using empirical covariance matrices of features. A dense set of spatio-temporal feature vectors are computed from video to provide a localized description of the action, and subsequently aggregated in an empirical covariance matrix to compactly represent the action. Two supervised learning methods for action recognition are developed using feature covariance matrices. Common to both methods is the transformation of the classification problem in the closed convex cone of covariance matrices into an equivalent problem in the vector space of symmetric matrices via the matrix logarithm. The first method applies nearest-neighbor classification using a suitable Riemannian metric for covariance matrices. The second method approximates the logarithm of a query covariance matrix by a sparse linear combination of the logarithms of training covariance matrices. The action label is then determined from the sparse coefficients. Both methods achieve state-of-the-art classification performance on several datasets, and are robust to action variability, viewpoint changes, and low object resolution. The proposed framework is conceptually simple and has low storage and computational requirements making it attractive for real-time implementation.
Prakash Ishwar, Janusz Konrad
IEEE Trans. Image Process.2
2013 Learning-Based, Automatic 2D-to-3D Image and Video Conversion
abstract
Despite a significant growth in the last few years, the availability of 3D content is still dwarfed by that of its 2D counterpart. To close this gap, many 2D-to-3D image and video conversion methods have been proposed. Methods involving human operators have been most successful but also time-consuming and costly. Automatic methods, which typically make use of a deterministic 3D scene model, have not yet achieved the same level of quality for they rely on assumptions that are often violated in practice. In this paper, we propose a new class of methods that are based on the radically different approach of learning the 2D-to-3D conversion from examples. We develop two types of methods. The first is based on learning a point mapping from local image/video attributes, such as color, spatial position, and, in the case of video, motion at each pixel, to scene-depth at that pixel using a regression type idea. The second method is based on globally estimating the entire depth map of a query image directly from a repository of 3D images ( image+depth pairs or stereopairs) using a nearest-neighbor regression type idea. We demonstrate both the efficacy and the computational efficiency of our methods on numerous 2D images and discuss their drawbacks and benefits. Although far from perfect, our results demonstrate that repositories of 3D content can be used for effective 2D-to-3D image conversion. An extension to video is immediate by enforcing temporal continuity of computed depth maps.
Janusz Konrad, Prakash Ishwar, Debargha Mukherjee
IEEE Trans. Image Process.3
2013 The Infinite-Message Limit of Two-Terminal Interactive Source Coding
abstract
A two-terminal interactive function computation problem with alternating messages is studied within the framework of distributed block source coding theory. For any finite number of messages, a single-letter characterization of the sum-rate-distortion function was established in previous works using standard information-theoretic techniques. This, however, does not provide a satisfactory characterization of the infinite-message limit, which is a new, unexplored dimension for asymptotic analysis in distributed block source coding involving potentially an infinite number of infinitesimal-rate messages. In this paper, the infinite-message sum-rate-distortion function, viewed as a functional of the joint source distribution and the distortion levels, is characterized as the least element of a partially ordered family of functionals having certain convex-geometric properties. The new characterization does not involve evaluating the infinite-message limit of a finite-message sum-rate-distortion expression. This characterization leads to a family of lower bounds for the infinite-message sum-rate-distortion expression and a simple criterion to test the optimality of any achievable infinite-message sum-rate-distortion expression. The new convex-geometric characterization is used to develop an iterative algorithm for evaluating any finite-message sum-rate-distortion function. It is also used to construct the first examples which demonstrate that for lossy source reproduction, two messages can strictly improve the one-message Wyner-Ziv rate-distortion function settling an unresolved question from a 1985 paper. It is shown that a single backward message of arbitrarily small rate can lead to an arbitrarily large gain in the sum-rate.
Prakash Ishwar
IEEE Trans. Inf. Theory2
2012 Towards Gesture-Based User Authentication
abstract
Video cameras are extensively used in modern surveillance systems to detect, track, and recognize, objects, people, and anomalies. Their use in user authentication, however, has been limited primarily to close-range face recognition systems. In this paper, we explore user authentication based on gestures captured by a video camera. Unlike pure biometrics, such as fingerprints, iris scans, and faces, gesture-based authentication combines irrevocable biometric information, such as the shapes and relative sizes of body parts, with voluntary movements which can be revoked. Our authentication method applies the empirical feature covariance matrix framework that has previously been used for tracking, face localization, and action recognition, to features extracted from body silhouettes. We have tested the performance of our algorithm in both user classification and user authentication on a database of 20 individuals performing 8 different gestures. We have obtained a 93-99% Correct Classification Rate (CCR) for user classification and a 5-6% Equal Error Rate (EER) for user authentication on single gestures from this dataset. This is a very encouraging result suggesting that gesture-based user authentication may be feasible in scenarios with a limited number of users.
Kam Lai, Janusz Konrad, Prakash Ishwar
AVSS3
2012 Coherent image selection using a fast approximation to the generalized traveling salesman problem
abstract
Searching for images on-line using keywords returns results that are often difficult to interpret. This becomes even more complicated if one attempts to compare image search output for several keywords with a common theme. We focus on the latter problem and propose a method to efficiently compare sets of images in order to find representative images, one from each set, that are coherent in certain sense. However, the search for an optimal set of representative images is very complex even for as few as 10 sets of 20 images each since all possible combinations of 10 images need to be considered. Therefore, we formulate our problem as the Generalized Traveling Salesman Problem (GTSP) and propose an efficient approximation algorithm to solve it. Our approximate GTSP algorithm is faster than other well-known approximations and is also more likely to reach the exact solution for large-scale inputs. We present a number of experimental results using the proposed algorithm and conclude that it can be a useful, almost real-time tool for on-line search.
Prakash Ishwar, Janusz Konrad, Cenk Gazen, Rohit Saboo
ACM Multimedia2
2012 A Theoretical Analysis of Authentication, Privacy, and Reusability Across Secure Biometric Systems
abstract
We present a theoretical framework for the analysis of privacy and security trade-offs in secure biometric authentication systems. We use this framework to conduct a comparative information-theoretic analysis of two biometric systems that are based on linear error correction codes, namely fuzzy commitment and secure sketches. We derive upper bounds for the probability of false rejection$(P_{FR})$and false acceptance$(P_{FA})$for these systems. We use mutual information to quantify the information leaked about a user's biometric identity, in the scenario where one or multiple biometric enrollments of the user are fully or partially compromised. We also quantify the probability of successful attack$(P_{SA})$based on the compromised information. Our analysis reveals that fuzzy commitment and secure sketch systems have identical$P_{FR}$,$P_{FA}$,$P_{SA}$, and information leakage, but secure sketch systems have lower storage requirements. We analyze both single-factor (keyless) and two-factor (key-based) variants of secure biometrics, and consider the most general scenarios in which a single user may provide noisy biometric enrollments at several access control devices, some of which may be subsequently compromised by an attacker. Our analysis highlights the revocability and reusability properties of key-based systems and exposes a subtle design trade-off between reducing information leakage from compromised systems and preventing successful attacks on systems whose data have not been compromised.
Ye Wang 0001, Shantanu Rane, Stark C. Draper, Prakash Ishwar
IEEE Trans. Inf. Forensics Secur.4
2012 Erratum to "On Delayed Sequential Coding of Correlated Sources"
abstract
Mehdi Torbatian kindly pointed out a difficulty in the proof of the unnumbered lemma in the above titled paper, Sec. III-C, p. 3768. The issue is addressed by the authors of the original paper.
Prakash Ishwar
IEEE Trans. Inf. Theory2
2012 Interactive Source Coding for Function Computation in Collocated Networks
abstract
A problem of interactive function computation in a collocated network is studied in a distributed block source coding framework. With the goal of computing samples of a desired function of sources at the sink, the source nodes exchange messages through a sequence of error-free broadcasts. For any function of independent sources, a computable characterization of the set of all feasible message coding rates—the rate region—is derived in terms of single-letter information measures. In the limit as the number of messages tends to infinity, the infinite-message minimum sum rate, viewed as a functional of the joint source probability mass function, is characterized as the least element of a partially ordered family of functionals having certain convex-geometric properties. This characterization leads to a family of lower bounds for the infinite-message minimum sum rate and a simple criterion to test the optimality of any achievable infinite-message sum rate. An iterative algorithm for evaluating the infinite-message minimum sum-rate functional is proposed and is demonstrated through an example of computing the minimum function of three Bernoulli sources. Based on the characterizations of the rate regions, it is shown that when computing symmetric functions of binary sources, the sink will inevitably learn certain additional information that is not demanded in computing the function. This conceptual understanding leads to new improved bounds for the minimum sum rate. The new bounds are shown to be orderwise better than those based on cut-sets as the network scales. The scaling law of the minimum sum rate is explored for different classes of symmetric functions and source parameters.
Prakash Ishwar
IEEE Trans. Inf. Theory2
2011 Image saliency: From intrinsic to extrinsic context
abstract
We propose a novel framework for automatic saliency estimation in natural images. We consider saliency to be an anomaly with respect to a given context that can be global or local. In the case of global context, we estimate saliency in the whole image relative to a large dictionary of images. Unlike in some prior methods, this dictionary is not annotated, i.e., saliency is assumed unknown. In the case of local context, we partition the image into patches and estimate saliency in each patch relative to a large dictionary of un-annotated patches from the rest of the image. We propose a unified framework that applies to both cases in three steps. First, given an input (image or patch) we extract k nearest neighbors from the dictionary. Then, we geometrically warp each neighbor to match the input. Finally, we derive the saliency map from the mean absolute error between the input and all its warped neighbors. This algorithm is not only easy to implement but also outperforms state-of-the-art methods.
Janusz Konrad, Prakash Ishwar, Kevin Jing, Henry A. Rowley
CVPR3
2011 A learning-based approach to explosives detection using Multi-Energy X-Ray Computed Tomography
abstract
In this paper we consider the task of classifying materials into explosives and non-explosives according to features obtainable from Multi-Energy X-ray Computed Tomography (MECT) measurements. The discriminative ability of MECT derives from its sensitivity to the attenuation versus energy curves of materials. Thus we focus on the fundamental information available in these curves and features extracted from them. We study the dimensionality and span of these curves for a set of explosive and non-explosive compounds and show that their space is larger than two-dimensional, as is typically assumed. In addition, we build support vector machine classifiers with different feature sets and find superior classification performance when using more than two features and when using features different than the standard photoelectric and Compton coefficients. These results suggest the potential for improved detection performance relative to conventional dual-energy X-ray systems.
Limor Eger, Synho Do, Prakash Ishwar, W. Clem Karl, Homer H. Pien
ICASSP3
2011 Sensing-aware classification with high-dimensional data
abstract
In many applications decisions must be made about the state of an object based on indirect noisy observation of high-dimensional data. An example is the determination of the presence or absence of stroke from tomographic projections. Conventionally, the sensing process is inverted and a classifier is built in the reconstructed domain, which requires complete knowledge of the sensing mechanism. Alternatively, a direct data domain classifier might be constructed, but the constraints imposed by the sensing process are then lost. In this work we study the behavior of a third path we term “sensing-aware classification.” Our aim is to contribute to the development of a rigorous theory for such challenging problems. To this end, we consider an abstracted binary classification problem with very high dimensional observations, a restricting sensing configuration, and unknown statistical models of noise and object which must be learned from constrained training data. We analyze the impact of different levels of prior knowledge concerning the sensing mechanism for various classification strategies. In particular we prove that the strategies based on the naive estimation of all model elements results in a classification performance asymptotically no better than guessing whereas sensing-aware, projection-based classification rules attain Bayes-optimal risk. Simulation results are also provided.
Burkay Orten, Prakash Ishwar, W. Clem Karl, Venkatesh Saligrama, Homer H. Pien
ICASSP2
2011 On unconditionally secure multi-party sampling from scratch1
abstract
In the problem of secure multi-party sampling, n parties wish to securely sample an n-variate joint distribution, with each party receiving a sample of one of the correlated variables. The objective is to correctly produce the samples using a distributed message passing protocol, while maintaining privacy against a coalition of passively cheating parties. In the two-party case, we fully characterize the joint distributions that can be securely sampled under perfect correctness and privacy requirements as well as under weakened correctness and privacy requirements. Furthermore, we show that the distributions that can be securely sampled can be produced with a protocol that only uses one round of unidirectional communication. For the n-party case, any distribution can be securely sampled with privacy against a strict minority coalition, due to well-known results in secure multi-party computation. However, when privacy against a majority coalition is required, not all distributions can be securely sampled. We give necessary conditions and sufficient conditions for distributions that can be securely sampled. However, the exact characterization of the distributions that can be securely sampled remains open.
Ye Wang 0001, Prakash Ishwar
ISIT2
2011 High-Resolution Distributed Sampling of Bandlimited Fields With Low-Precision Sensors
abstract
The problem of sampling a discrete-time sequence of spatially bandlimited fields, with a bounded dynamic range, in a distributed, communication-constrained processing environment is studied. A central unit having access to the data gathered by a dense network of low-precision sensors, is required to reconstruct the field snapshots to maximum accuracy. Both deterministic and stochastic field models are considered. For stochastic fields, results are established in the almost-sure sense. The feasibility of having a flexible tradeoff between the oversampling rate (sensor density) and the analog-to-digital converter (ADC) precision, while achieving an exponential accuracy in the number of bits per Nyquist-interval per snapshot is demonstrated. This exposes an underlying “conservation of bits” principle: the bit-budget per Nyquist-interval per snapshot (the rate) can be distributed along the amplitude axis (sensor-precision) and space (sensor density) in an almost arbitrary discrete-valued manner, while retaining the same (exponential) distortion-rate characteristics.
Animesh Kumar, Prakash Ishwar, Kannan Ramchandran
IEEE Trans. Inf. Theory2
2011 On Delayed Sequential Coding of Correlated Sources
abstract
Motivated by video coding applications, the problem of sequential coding of correlated sources with encoding and/or decoding frame-delays is studied. The fundamental tradeoffs between individual frame rates, individual frame distortions, and encoding/decoding frame-delays are derived in terms of a single-letter information-theoretic characterization of the rate-distortion region for general interframe source correlations and certain types of potentially frame specific and coupled single-letter fidelity criteria. The sum-rate-distortion region is characterized in terms of generalized directed information measures highlighting their role in delayed sequential source coding problems. For video sources which are spatially stationary memoryless and temporally Gauss–Markov, MSE frame distortions, and a sum-rate constraint, our results expose the optimality of idealized differential predictive coding among all causal sequential coders, when the encoder uses a positive rate to describe each frame. Somewhat surprisingly, causal sequential encoding with one-frame-delayed noncausal sequential decoding can exactly match the sum-rate-MSE performance of joint coding for all nontrivial MSE-tuples satisfying certain positive semidefiniteness conditions. Thus, even a single frame-delay holds potential for yielding significant performance improvements. Generalizations to higher order Markov sources are also presented and discussed. A rate-distortion performance equivalence between, causal sequential encoding with delayed noncausal sequential decoding, and delayed noncausal sequential encoding with causal sequential decoding, is also established.
Prakash Ishwar
IEEE Trans. Inf. Theory2
2011 Some Results on Distributed Source Coding for Interactive Function Computation
abstract
A two-terminal interactive distributed source coding problem with alternating messages for function computation at both locations is studied. For any number of messages, a computable characterization of the rate region is provided in terms of single-letter information measures. While interaction is useless in terms of the minimum sum-rate for lossless source reproduction at one or both locations, the gains can be arbitrarily large for function computation even when the sources are independent. For a class of sources and functions, interaction is shown to be useless, even with infinite messages, when a function has to be computed at only one location, but is shown to be useful, if functions have to be computed at both locations. For computing the Boolean AND function of two independent Bernoulli sources at both locations, an achievable infinite-message sum-rate with infinitesimal-rate messages is derived in terms of a 2-D definite integral and a rate-allocation curve. The benefit of interaction is highlighted in multiterminal function computation problem through examples. For networks with a star topology, multiple rounds of interactive coding is shown to decrease the scaling law of the total network rate by an order of magnitude as the network grows.
Prakash Ishwar
IEEE Trans. Inf. Theory2
2010 Action Recognition Using Sparse Representation on Covariance Manifolds of Optical Flow
abstract
A novel approach to action recognition in video based on the analysis of optical flow is presented. Properties of optical flow useful for action recognition are captured using only the empirical covariance matrix of a bag of features such as flow velocity, gradient, and divergence. The feature covariance matrix is a low-dimensional representation of video dynamics that belongs to a Riemannian manifold. The Riemannian manifold of covariance matrices is transformed into the vector space of symmetric matrices under the matrix logarithm mapping. The log-covariance matrix of a test action segment is approximated by a sparse linear combination of the log-covariance matrices of training action segments using a linear program and the coefficients of the sparse linear representation are used to recognize actions. This approach based on the unique blend of a logcovariance-descriptor and a sparse linear representation is tested on the Weizmann and KTH datasets. The proposed approach attains leave-one-out cross validation scores of 94.4% correct classification rate for the Weizmann dataset and 98.5% for the KTH dataset. Furthermore, the method is computationally efficient and easy to implement.
Prakash Ishwar, Janusz Konrad
AVSS2
2010 Action change detection in video by covariance matching of silhouette tunnels
abstract
Action recognition is an important but challenging problem in video analytics with a number of solutions proposed to date. However, even if a reliable model for action representation is identified and an accurate metric for comparing actions is developed, it is still unclear to how many video frames should the representation and comparison apply. In this paper, we develop a method to detect when actions change, i.e., the temporal boundaries of actions, without classifying the actions. We use a silhouette-based framework for action representation and comparison, both centered around dimensionality reduction using covariance descriptors. We use a nonparametric statistical framework to learn the distribution of the distance between covariance descriptors and detect action changes as covariance-distance outliers. Experimental results on ground-truth data show 1.64% false negative error and 0.19% false positive error, while those for surveillance video agree 100% with manual annotations.
Prakash Ishwar, Janusz Konrad
ICASSP2
2010 Interaction strictly improves the Wyner-Ziv rate-distortion function
abstract
In 1985 Kaspi provided a single-letter characterization of the sum-rate-distortion function for a two-way lossy source coding problem in which two terminals send multiple messages back and forth with the goal of reproducing each other's sources. Yet, the question remained whether more messages can strictly improve the sum-rate-distortion function. Viewing the sum-rate as a functional of the distortions and the joint source distribution and leveraging its convex-geometric properties, we construct an example which shows that two messages can strictly improve the one-message (Wyner-Ziv) rate-distortion function. The example also shows that the ratio of the one-message rate to the two-message sum-rate can be arbitrarily large and simultaneously the ratio of the backward rate to the forward rate in the two-message sum-rate can be arbitrarily small.
Prakash Ishwar
ISIT2
2010 Infinite-message interactive function computation in collocated networks
abstract
An interactive function computation problem in a collocated network is studied in a distributed block source coding framework. With the goal of computing a desired function at the sink, the source nodes exchange messages through a sequence of error-free broadcasts. The infinite-message minimum sum-rate is viewed as a functional of the joint source PMF and is characterized as the least element in a partially ordered family of functionals having certain convex-geometric properties. This characterization leads to a family of lower bounds for the infinite-message minimum sum-rate and a simple optimality test for any achievable infinite-message sum-rate. An iterative algorithm for evaluating the infinite-message minimum sum-rate functional is proposed and is demonstrated through an example of computing the minimum function of three Bernoulli sources.
Prakash Ishwar
ISIT2
2009 Information-theoretic bounds for multiround function computation in collocated networks
abstract
We study the limits of communication efficiency for function computation in collocated networks within the framework of multi-terminal block source coding theory. With the goal of computing a desired function of sources at a sink, nodes interact with each other through a sequence of error-free, network-wide broadcasts of finite-rate messages. For any function of independent sources, we derive a computable characterization of the set of all feasible message coding rates - the rate region -in terms of single-letter information measures. We show that when computing symmetric functions of binary sources, the sink will inevitably learn certain additional information which is not demanded in computing the function. This conceptual understanding leads to new improved bounds for the minimum sum-rate. The new bounds are shown to be orderwise better than those based on cut-sets as the network scales. The scaling law of the minimum sum-rate is explored for different classes of symmetric functions and source parameters.
Prakash Ishwar
ISIT2
2009 Bootstrapped oblivious transfer and secure two-party function computation
abstract
We propose an information theoretic framework for the secure two-party function computation (SFC) problem and introduce the notion of SFC capacity. We study and extend string oblivious transfer (OT) tosample-wiseOT. We propose an efficient,perfectlyprivateOT protocol utilizing the binary erasure channel or source. We also propose thebootstrapstring OT protocol which provides disjoint (weakened) privacy while achieving a multiplicative increase in rate, thus trading off security for rate. Finally, leveraging our OT protocol, we construct a protocol for SFC and establish a general lower bound on SFC capacity of the binary erasure channel and source.
Ye Wang 0001, Prakash Ishwar
ISIT2
2009 Video Condensation by Ribbon Carving
abstract
Efficient browsing of long video sequences is a key tool in visual surveillance, e.g., for postevent video forensics, but can also be used for fast review of motion pictures and home videos. While frame skipping (fixed or adaptive) is straightforward to implement, its performance is quite limited. Although more efficient techniques have been developed, such as video summarization and video montage, they lose either the temporal or semantic context of events. A recently proposed method called video synopsis deals with some of these issues but involves multiple processing stages and is fairly complex. Video condensation, that we propose here, is novel in the way information is removed from the space-time video volume, is conceptually simple and relatively easy to implement. We introduce the concept of a video ribbon inspired by that of a seam recently proposed for image resizing. We recursively carve ribbons out by minimizing an activity-aware cost function using dynamic programming. The ribbon model we develop is flexible and permits an easy adjustment of the compromise between temporal condensation ratio and anachronism of events. We also propose sliding-window ribbon carving to handle streaming video and demonstrate the method's efficiency on motor and pedestrian traffic data.
Prakash Ishwar, Janusz Konrad
IEEE Trans. Image Process.2
2009 Field estimation from randomly located binary noisy sensors
abstract
The estimation of bounded multivariate fields from 1-bit quantized and dithered noisy observations is considered. We consider two models for random sensor deployment based on regular Monte Carlo (simple random sampling) and stratified sampling. We propose linear estimators, and for both sensors deployment methods we establish exact expressions for the bias and variance of the estimates (including integrated mean-square errors). We show in particular that estimates of the field on the basis of stratified sensor locations always outperform estimates based on regular Monte Carlo sensor locations. For both estimation schemes, we also establish central limit theorems which can be used to compute the probability of events involving the estimates including confidence intervals.
Elias Masry, Prakash Ishwar
IEEE Trans. Inf. Theory2
2008 Two-terminal distributed source coding with alternating messages for function computation
abstract
A two-terminal interactive distributed source coding problem with alternating messages is studied. The focus is on function computation at both locations with a probability which tends to one as the blocklength tends to infinity. A single-letter characterization of the rate region is provided. It is observed that interaction is useless (in terms of the minimum sum-rate) if the goal is pure source reproduction at one or both locations but the gains can be arbitrarily large for (general) function computation. For doubly symmetric binary sources and any function, interaction is useless with even infinite messages, when computation is desired at only one location, but is useful, when desired at both locations. For independent Bernoulli sources and the Boolean AND function computation at both locations, an interesting achievable infinite-message sum-rate is derived. This sum-rate is expressed, in analytic closed-form, in terms of a two-dimensional definite integral with an infinitesimal rate for each message.
Prakash Ishwar
ISIT2
2007 Fundamental Redundancy Versus Power Trade-Off in Standby SRAM
abstract
We study the problem of reducing power during data-retention in a standby static random access memory (SRAM). For successful data-retention, the supply voltage of an SRAM cell should be greater than a critical data retention voltage (DRV). Due to circuit parameter variations, the DRV for different cells on the same chip exhibits variation with a distribution having diminishing tail. For reliable data retention, the existing low-power design uses a worst-case technique in which a standby supply voltage that is larger than the highest DRV among all cells in an SRAM is used. Instead, our approach uses aggressive voltage reduction and counters the ensuing unreliability through a fault-tolerant memory architecture. The main results of this work are as follows: (i) We establish fundamental bounds on the power reduction in terms of the DRV-distribution using techniques from information theory. For the DRV-distribution of test-chip in (Qin, H, et al., 2006), we show that 49% power reduction with respect to (w.r.t.) the worst-case is a fundamental lower bound while 40% power reduction w.r.t. the worst-case is achievable with a practical combinatorial scheme, (ii) We study the power reduction as a function of the block-length for low-latency codes since most applications using SRAM are latency constrained. We propose a reliable memory architecture based on the Hamming code for the next test-chip implementation with a predicted power reduction of 33% while accounting for coding overheads.
Animesh Kumar, Huifang Qin, Prakash Ishwar, Jan M. Rabaey, Kannan Ramchandran
ICASSP (2)3
2007 Fundamental Bounds on Power Reduction during Data-Retention in Standby SRAM
abstract
The authors study leakage-power reduction in standby random access memories (SRAMs) during data-retention. An SRAM cell requires a minimum critical supply voltage (DRV) above which it preserves the stored-bit reliably. Due to process-variations, the intra-chip DRV exhibits variation with a distribution having a diminishing tail. In order to minimize leakage power while preserving data reliably, existing low-power design methods use a worst-case standby supply voltage. This worst-case voltage is larger than the highest DRV among all cells in an SRAM. In contrast, the approach uses aggressive voltage reduction and counters the ensuing unreliability by an error-control code based memory architecture. Using this approach, we explore fundamental trade-offs between power reduction and redundancy present in the SRAM. The authors establish fundamental bounds on the power reduction in terms of the DRV-distribution using techniques from information theory and algebraic coding theory. For an experimental test-chip DRV-distribution in the 90nm CMOS technology, the authors show that 49% power reduction with respect to (w.r.t.) the worst-case is a fundamental lower bound while 40% power reduction w.r.t. the worst-case is achievable by using a practical algebraic coding scheme. The authors also study the power reduction as a function of the block-length for low-latency codes since most applications using SRAM are latency constrained. The authors propose a reliable low-power memory architecture based on the Hamming code for the next test-chip implementation with a predicted power reduction of 33% while accounting for coding overheads
Animesh Kumar, Huifang Qin, Prakash Ishwar, Jan M. Rabaey, Kannan Ramchandran
ISCAS3
2007 The Value of Frame-Delays in the Sequential Coding of Correlated Sources
abstract
The problem of sequential coding of correlated sources with causal encoding and noncausal decoding frame- delays is studied. The fundamental tradeoffs between individual frame rates, individual frame distortions, and decoding frame-delays are stated in terms of a single-letter information-theoretic characterization of the rate-distortion region for general inter- frame source correlations and certain types of (potentially frame- specific and coupled) single-letter fidelity criteria. For sources which are spatially stationary memoryless and temporally Gauss-Markov, mean squared error (MSE) frame distortions, and a sum-rate constraint, it is shown that causal sequential encoding with one-step delayed noncausal sequential decoding exactly matches the sum-rate-MSE performance of joint coding for all nontrivial MSE-tuples satisfying certain positive semi-definiteness conditions. Generalizations to multiple frames and arbitrary frame-delays are also presented and discussed.
Prakash Ishwar
ISIT2
2007 Benefit of Delay on the Diversity-Multiplexing Tradeoffs of MIMO Channels with Partial CSI
abstract
This paper re-examines the well-known fundamental tradeoffs between rate and reliability for the multi-antenna, block Rayleigh fading channel in the high signal to noise ratio (SNR) regime when (i) the transmitter has access to (noiseless) one bit per coherence-interval of causal channel state information (CSI) and (it) soft decoding delays together with worst-case delay guarantees are acceptable. A key finding of this work is that substantial improvements in reliability can be realized with a very short expected delay and a slightly longer (but bounded) worst-case decoding delay guarantee in communication systems where the transmitter has access to even one bit per coherence interval of causal CSI. While similar in spirit to the recent work on communication systems based on automatic repeat requests (ARQ) where decoding failure is known at the transmitter and leads to re-transmission, here transmit side-information is purely based on CSI. The findings reported here also lend further support to an emerging understanding that decoding delay (related to throughput) and codeword blocklength (related to coding complexity and delays) are distinctly different design parameters which can be tuned to control reliability.
Masoud Sharif, Prakash Ishwar
ISIT2
2007 On Non-Parametric Field Estimation using Randomly Deployed, Noisy, Binary Sensors
abstract
We consider the problem of reconstructing a deterministic data field from binary quantized noisy observations of sensors randomly deployed over the field domain. Our focus is on the extremes of lack of control in the sensor deployment, arbitrariness and lack of knowledge of the noise distribution, and low-precision and unreliability in the sensors. These adverse conditions are motivated by possible real-world scenarios where a large collection of low-cost, crudely manufactured sensors are mass-deployed in an environment where little can be assumed about the ambient noise. We propose a simple estimator that reconstructs the entire data field from these unreliable, binary quantized, noisy observations. Under the assumption of a bounded amplitude field, we prove almost sure and mean-square convergence of the estimator to the actual field as the number of sensors tends to infinity. For fields with bounded-variation, Sobolev differentiable, or finite-dimensionality properties, we derive specific mean squared error (MSE) decay rates. The analysis techniques used herein expose the effects of field "smoothness" properties, location randomness, and noise on the MSE scaling behavior.
Ye Wang 0001, Prakash Ishwar
ISIT2
2006 On Source Encoding with Side-information Under Ambiguous State of Nature
abstract
In this paper, we address the case of a single remote sensing unit (encoder), a central processing unit (decoder), and a finite bitrate constraint as an abstraction of a bandwidth-limited channel between the encoder and decoder. The goal is to spend this bit budget in an optimal sense, in terms of classification or estimation performance (minimize the probability of classification error or minimize the probability that the parameter estimation error exceeds a desired tolerance) while ensuring the reconstruction of the raw data with maximum fidelity (in a rate-distortion sense)
Prakash Ishwar, Vinod M. Prabhakaran, Kannan Ramchandran
ISIT1
2005 Complexity/performance trade-offs for robust distributed video coding
abstract
In this work, we analytically study the complexity-performance trade-offs associated with video codecs based on the principle of source coding with side information at the decoder. We address three important aspects. First, we quantify the theoretical performance gains attained with side-information based codecs over prediction-based coders like MPEG under a lossy transmission scenario when there is drift. Secondly, we show that it is possible to closely approach MPEG's compression performance using side-information video coding principles with accurate DFD (displaced frame difference) modeling even without sophisticated channel codes. Thirdly, we analytically show that the value of accurately estimating DFD statistics diminishes as the channel gets noisier.
Abhik Majumdar, Rohit Puri, Prakash Ishwar, Kannan Ramchandran
ICIP (2)3
2005 On rate-constrained distributed estimation in unreliable sensor networks
abstract
We study the problem of estimating a physical process at a central processing unit (CPU) based on noisy measurements collected from a distributed, bandwidth-constrained, unreliable, network of sensors, modeled as an erasure network of unreliable "bit-pipes" between each sensor and the CPU. The CPU is guaranteed to receive data from a minimum fraction of the sensors and is tasked with optimally estimating the physical process under a specified distortion criterion. We study the noncollaborative (i.e., fully distributed) sensor network regime, and derive an information-theoretic achievable rate-distortion region for this network based on distributed source-coding insights. Specializing these results to the Gaussian setting and the mean-squared-error (MSE) distortion criterion reveals interesting robust-optimality properties of the solution. We also study the regime of clusters of collaborative sensors, where we address the important question: given a communication rate constraint between the sensor clusters and the CPU, should these clusters transmit their "raw data" or some low-dimensional "local estimates"? For a broad set of distortion criteria and sensor correlation statistics, we derive conditions under which rate-distortion-optimal compression of correlated cluster-observations separates into the tasks of dimension-reducing local estimation followed by optimal distributed compression of the local estimates.
Prakash Ishwar, Rohit Puri, Kannan Ramchandran, S. Sandeep Pradhan
IEEE J. Sel. Areas Commun.1
2005 On the Existence and Characterization of the Maxent Distribution Under General Moment Inequality Constraints
abstract
A broad set of sufficient conditions that guarantees the existence of the maximum entropy (maxent) distribution consistent with specified bounds on certain generalized moments is derived. Most results in the literature are either focused on the minimum cross-entropy distribution or apply only to distributions with a bounded-volume support or address only equality constraints. The results of this work hold for general moment inequality constraints for probability distributions with possibly unbounded support, and the technical conditions are explicitly on the underlying generalized moment functions. An analytical characterization of the maxent distribution is also derived using results from the theory of constrained optimization in infinite-dimensional normed linear spaces. Several auxiliary results of independent interest pertaining to certain properties of convex coercive functions are also presented.
Prakash Ishwar, Pierre Moulin
IEEE Trans. Inf. Theory1
2004 On distributed sampling of bandlimited and non-bandlimited sensor fields
abstract
Distributed sampling and reconstruction of a physical field using an array of sensors is a problem of increasing interest in environmental monitoring applications of sensor networks. We show, using a dither-based scheme, that it is possible to reconstruct non-bandlimited fields with a reconstruction accuracy that depends on the available bitrate R and the spectral decay characteristics of the sensor field - we study exponentially decaying spectra as an illustration. For bandlimited fields f(t), the maximum pointwise error D/sub f/ decays as D/sub f/ /spl sim/ 2/sup -a1R/, i.e. exponentially with rate R. For the non-bandlimited case, we show that for fields u(t) with exponentially decaying spectral tails, i.e., |U(/spl omega/)|/spl sim/e/sup -a|/spl omega/|/, the maximum pointwise error D/sub u/ decays as D/sub u//spl sim/e/sup -a2/spl radic/R/(1+o(R)) with spatial bit rate R bits/metre. We also show that it is possible to trade off the number of sensors with their precision, while maintaining a similar reconstruction accuracy - a phenomenon that may be dubbed as a "bit-conservation" principle underlying the sampling framework.
Animesh Kumar, Prakash Ishwar, Kannan Ramchandran
ICASSP (3)2
2004 On decoder-latency versus performance tradeoffs in differential predictive coding
abstract
Theoretical analysis of differential predictive coding (DPC) has almost exclusively focused on scalar quantizers and the high-rate regime for tractability reasons. As a result, the role of noncausal decoding in improving the quality has been largely ignored in the literature. In this work we conduct a rigorous performance analysis of DPC-based schemes under a simple independent, vector-Gaussian, AR-1 source model and large-block (as opposed to high-rate) asymptotics. This analysis reveals that noncausal decoding can offer a significant relative improvement in the mean squared error (by as much as 3 dB) at medium to low rates (0.1-0.5 bit per sample) for sources having strong temporal correlation. Furthermore, most of this relative improvement can be attained with a modest decoder-latency. At very high and very low rates, the gains are negligible.
Prakash Ishwar, Kannan Ramchandran
ICIP1
2004 On distributed sampling of smooth non-bandlimited fields
abstract
Distributed sampling and reconstruction of a physical field using an array of sensors is a problem of considerable interest in environmental monitoring applications of sensor networks. Our recent work has focused on the sampling of bandlimited sensor fields. However, sensor fields are not perfectly bandlimited but typically have rapidly decaying spectra. In a classical sampling set-up it is possible to precede the A/D sampling operation with an appropriate analog anti-aliasing filter. However, in the case of sensor networks, this is infeasible since sampling must precede filtering. We show that even though the effects of aliasing on the reconstruction cannot be prevented due to the "filter-less" sampling constraint, they can be suitably controlled by oversampling and carefully reconstructing the field from the samples. We show using a dither-based scheme that it is possible to estimate non-bandlimited fields with a precision that depends on how fast the spectral content of the field decays. We develop a framework for analyzing non-bandlimited fields that lead to upper bounds on the maximum pointwise error for a spatial bit rate of R bits/meter. We present results for fields with exponentially decaying spectra as an illustration. In particular, we show that for fields f(t) with exponential tails; i.e., F(ω) ‹ παε–αω, the maximum pointwise error decays as c2e–α1√R+c3 1 over √R e >–2α1√R with spatial bit rate R bits/meter. Finally, we show that for fields with spectra that have a finite second moment, the distortion decreases as O((1 overN)2 over 3) as the density of sensors, N, scales up to infinity . We show that if D is the targeted non-zero distortion, then the required (finite) rate R scales as O (1 over √ overD log 1 over D).
Animesh Kumar, Prakash Ishwar, Kannan Ramchandran
IPSN2
2004 Compressing encrypted sources using side-information coding
abstract
When transmitting a source over an insecure and bandwidth-limited channel, compression (lossy/lossless) precedes encryption. We show that, through the use of side-information coding principles, the order of these operations can be reversed without loss of Wyner-sense perfect secrecy and often with a significant compression ratio. Further, when the source has to be recovered perfectly (with high probability) or is Gaussian (with the mean-squared error fidelity criterion), there is no loss of compression efficiency and the proposed system requires no more randomness in the encryption key compared to systems where compression precedes encryption.
Prakash Ishwar, Vinod M. Prabhakaran, Kannan Ramchandran
ISIT1
2003 Towards a theory for video coding using distributed compression principles
abstract
This paper presents an information-theoretic study of video codecs that are based on the principle of source coding with side information at the decoder. In contrast to the classical Wyner-Ziv side-information source coding problem (1976), in this work we address the situation where the source and side-information are connected through a state of nature that is unknown to both the encoder and the decoder. We dub this framework as source encoding with side-information under ambiguous state of nature (SEASON). Our objective is to compare the achievable rate-distortion (R/D) performance of conventional video codecs designed under the motion-compensated predictive coding (MCPC) framework and video codecs designed under the SEASON framework. Our analysis shows that under appropriate motion models and for Gaussian displaced frame difference (DFD) statistics, the R/D performance of a classical MCPC-based video codec is matched by that of our proposed SEASON-based video codec, with the hitter being characterized by the novel concept of moving the motion compensation task from the encoder to the decoder.
Prakash Ishwar, Vinod M. Prabhakaran, Kannan Ramchandran
ICIP (2)1
2000 Fundamental equivalences between set-theoretic and maximum-entropy methods in multiple-domain image restoration
abstract
Several powerful, but heuristic techniques in the image denoising literature have used overcomplete image representations. A general framework for incorporating information from multiple representations based on fundamental statistical estimation principles was presented in Ishwar and Moulin (1999) where, information about image attributes from multiple wavelet transforms was incorporated as moment constraints on the underlying image prior. In this paper we explore the fundamental equivalence between the stochastic setting of multiple-domain restoration in Ishwar and Moulin and its deterministic set-theoretic counterpart. The main technical tool is the Lagrange multiplier theory of constrained optimization. The insights gained by this analysis allow us to derive a state-of-the-art denoising algorithm.
Prakash Ishwar, Pierre Moulin
ICASSP1
2000 Shift Invariant Restoration - an Overcomplete Maxent Map Framework
abstract
Translation-invariant denoising was introduced by Coifman and Donoho (1995) to overcome Gibbs-type phenomena produced by transform-domain shrinkage estimators in the vicinity of signal discontinuities. Shrinkage estimators are in general not shift-invariant. Shift-invariant denoising consists of a simple averaging of the shrinkage estimates over a family of cyclic spatial-shifts of the image. Shift-invariant denoising is denoising in an overcomplete basis, and work in this area has been devoted towards finding a best basis in the overcomplete family. This paper presents a maximum a posteriori (MAP) framework for shift-invariant restoration of images using the maximum-entropy prior consistent with moment constraints on the transform coefficients in different subbands. The simple averaging of estimates in the classical shift-invariant denoising can then be shown to be a certain limiting case within this framework.
Prakash Ishwar, Pierre Moulin
ICIP1
2000 On spatial adaptation of motion-field smoothness in video coding
abstract
Most motion-compensation methods dealt with in the literature make strong assumptions about the smoothness of the underlying motion field. For instance, block-matching algorithms assume a blockwise-constant motion field and are adequate for translational motion models; control-grid interpolation assumes a blockwise bilinear motion field and captures zooming and warping fairly well. Time-varying imagery, however often contains both types of motion (as well as others), and hence exhibits a high degree of spatial variability of its motion-field smoothness properties. We develop a simple method to spatially adapt the smoothness of the motion field. The proposed method demonstrates substantial improvements in video quality over a wide range of bit rates. To this end, we introduce the notion of a motion field that is characterized by a set of labels. The labels provide the flexibility to adaptively switch between two different motion models locally. The individual motion models have very different smoothness properties. The switched framework for motion compensation performs significantly better than each of its constituent motion models, in terms of both visual quality and signal-to-noise ratio (0.3-0.7 dB on the average). Finally, we develop an extension of this method that enhances the overlapped block motion compensation scheme by allowing spatial adaptation of the window function.
Prakash Ishwar, Pierre Moulin
IEEE Trans. Circuits Syst. Video Technol.1
1999 Multiple-Domain Image Modeling and Restoration
abstract
Several powerful, but heuristic techniques in recent image denoising literature have used multiple (typically overcomplete) image representations. This paper presents a framework for multiple-domain image modeling and restoration, based on fundamental statistical estimation principles. Information about image attributes from multiple wavelet transforms is incorporated as moment constraints on the underlying image prior. Our method constructs the maximum entropy distribution consistent with these moment constraints. A maximum a posteriori probability (MAP) image restoration algorithm based on this maximum entropy prior is developed. Unlike previous multiple-domain algorithms, ours satisfies certain desirable optimality properties and provides an information-theoretic figure of merit for the choice of domains. Simulation results show that the estimator is vastly superior to single-domain image restoration both in terms of mean squared error and perceptual quality.
Prakash Ishwar, Pierre Moulin
ICIP (1)1
1999 Segmentation Based Denoising Using Multiple Compaction Domains
abstract
In this paper, we propose a novel segmentation based denoising algorithm. Segmentation yields intrinsically homogeneous and extrinsically heterogeneous regions. A denoising algorithm that uses Multiple Compaction Domains (MCD) is then applied on each of the resulting segments. Such a scheme retains important perceptual information in the segment boundaries while the denoising algorithm operates only on homogeneous segments. Further, the MCD algorithm is demonstrably superior to the classical denoising algorithms using transform domain thresholding. Our algorithm yields better perceptual quality and superior PSNR as compared to MATLAB's adaptive Wiener filter.
Maneesh Kumar Singh 0001, Prakash Ishwar, Krishna Ratakonda, Narendra Ahuja
ICIP (1)2
1998 Image denoising using multiple compaction domains
abstract
We present a novel framework for denoising signals from their compact representation in multiple domains. Each domain captures, uniquely, certain signal characteristics better than others. We define confidence sets around data in each domain and find sparse estimates that lie in the intersection of these sets, using a POCS algorithm. Simulations demonstrate the superior nature of the reconstruction (both in terms of mean-square error and perceptual quality) in comparison to the adaptive Wiener filter.
Prakash Ishwar, Krishna Ratakonda, Pierre Moulin, Narendra Ahuja
ICASSP1
1997 Switched control grid interpolation for motion compensated video coding
abstract
In this paper, fundamental limitations of the control grid interpolation and block matching schemes for motion compensation are remedied by a novel idea which embodies positive features of both schemes. The enhanced scheme performs both consistently and significantly better than either method does individually.
Prakash Ishwar, Pierre Moulin
ICIP (3)1