EDBT 2026 Demo / reviewers in the wild / expert
Sudeep Sarkar
dblp:72/3470
· DBLP profile ↗
124ranked-venue papers
20as first author
14since 2021 · last 2024
0000-0001-7332-4207ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 95 · 18 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 56 · 8 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 1 since 2021Systems, architecture and hardware · 4Security and privacy · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | BirdCollect: A Comprehensive Benchmark for Analyzing Dense Bird Flock AttributesabstractAutomatic recognition of bird behavior from long-term, un controlled outdoor imagery can contribute to conservation efforts by enabling large-scale monitoring of bird populations. Current techniques in AI-based wildlife monitoring have focused on short-term tracking and monitoring birds individually rather than in species-rich flocks. We present Bird-Collect, a comprehensive benchmark dataset for monitoring dense bird flock attributes. It includes a unique collection of more than 6,000 high-resolution images of Demoiselle Cranes (Anthropoides virgo) feeding and nesting in the vicinity of Khichan region of Rajasthan. Particularly, each image contains an average of 190 individual birds, illustrating the complex dynamics of densely populated bird flocks on a scale that has not previously been studied. In addition, a total of 433 distinct pictures captured at Keoladeo National Park, Bharatpur provide a comprehensive representation of 34 distinct bird species belonging to various taxonomic groups. These images offer details into the diversity and the behaviour of birds in vital natural ecosystem along the migratory flyways. Additionally, we provide a set of 2,500 point-annotated samples which serve as ground truth for benchmarking various computer vision tasks like crowd counting, density estimation, segmentation, and species classification. The benchmark performance for these tasks highlight the need for tailored approaches for specific wildlife applications, which include varied conditions including views, illumination, and resolutions. With around 46.2 GBs in size encompassing data collected from two distinct nesting ground sets, it is the largest birds dataset containing detailed annotations, showcasing a substantial leap in bird research possibilities. We intend to publicly release the dataset to the research community. The database is available at: https://iab-rubric.org/resources/wildlife-dataset/birdcollect Kshitiz, Sonu Sreshtha, Bikash Dutta, Muskan Dosi, Mayank Vatsa, Richa Singh 0001, Saket Anand, Sudeep Sarkar, Sevaram Mali Parihar |
AAAI | 8 |
| 2024 | Unveiling Gender Effects in Gait Recognition Using Conditional-Matched Bootstrap AnalysisabstractWhile biases such as gender, race, and age have been closely examined in biometric recognition, especially in face and fingerprint traits, their exploration in gait-based recognition is lacking, except for one study. We formulate conditional-matched bootstrap analysis to control for confounding covariates like clothing style, height, and walking speed. The goal is to isolate genuine gender effects on gait recognition. We delve into gender-based disparities in gait recognition by using several state-of-the-art gait recognition methodologies - GaitSet, GaitPart, and GaitGL. For our analysis, the widely-referenced OU-MVLP dataset served as our foundation, which we enhanced with annotations about clothing style, body height, and walking speed. The results were illuminating. We observed a disparity in recognition performance across genders on the original dataset, with recognition for females higher than for males. However, after controlling for covariate distributions using conditional-matched bootstrap analysis, the gap was reduced, with clothing type emerging as the most significant contributor. Code available at https://github.com/azimIbragimov/gait-gender Azim Ibragimov, Maurício Pamplona Segundo, Sudeep Sarkar, Kevin W. Bowyer |
FG | 3 |
| 2024 | Predictive Attractor ModelsabstractSequential memory, the ability to form and accurately recall a sequence of events or stimuli in the correct order, is a fundamental prerequisite for biological and artificial intelligence as it underpins numerous cognitive functions (e.g., language comprehension, planning, episodic memory formation, etc.) However, existing methods of sequential memory suffer from catastrophic forgetting, limited capacity, slow iterative learning procedures, low-order Markov memory, and, most importantly, the inability to represent and generate multiple valid future possibilities stemming from the same context. Inspired by biologically plausible neuroscience theories of cognition, we propose Predictive Attractor Models (PAM), a novel sequence memory architecture with desirable generative properties. PAM is a streaming model that learns a sequence in an online, continuous manner by observing each input only once. Additionally, we find that PAM avoids catastrophic forgetting by uniquely representing past context through lateral inhibition in cortical minicolumns, which prevents new memories from overwriting previously learned knowledge. PAM generates future predictions by sampling from a union set of predicted possibilities; this generative ability is realized through an attractor model trained alongside the predictor. We show that PAM is trained with local computations through Hebbian plasticity rules in a biologically plausible framework. Other desirable traits (e.g., noise tolerance, CPU-based learning, capacity scaling) are discussed throughout the paper. Our findings suggest that PAM represents a significant step forward in the pursuit of biologically plausible and computationally efficient sequential memory models, with broad implications for cognitive science and artificial intelligence research. Illustration videos and code are available on our project page: https://ramymounir.com/publications/pam. Ramy Mounir, Sudeep Sarkar |
NeurIPS | 2 |
| 2023 | DOERS: Distant Observation Enhancement and Recognition SystemabstractIn order to recognize people across long distances and from elevated viewpoints, biometric systems must handle the challenges of imaging through atmospheric turbulence and non-frontal presentations, in addition to the traditional A-PIE challenges of aging, pose, illumination, and expression. While individual biometric modalities such as facial appearance, gait, and whole body appearance each have a role to play, no single modality can address all of these challenges. This paper describes a novel multi-modal biometric recognition system that addresses the challenges of atmospheric turbulence, occlusions, and elevated viewpoints by combining these modalities. We demonstrate our system on both $R G B$ video-based identity verification and both open and closed-world search. Dawei Du, Cole Hill, Gabriel Bertocco, Maurício Pamplona Segundo, Wes Robbins, Brandon RichardWebster, Roderic Collins, Sudeep Sarkar, Terrance E. Boult, Scott McCloskey |
IJCB | 8 |
| 2023 | Long-term Monitoring of Bird Flocks in the WildabstractMonitoring and analysis of wildlife are key to conservation planning and conflict management. The widespread use of camera traps coupled with AI-based analysis tools serves as an excellent example of successful and non-invasive use of technology for design, planning, and evaluation of conservation policies. As opposed to the typical use of camera traps that capture still images or short videos, in this project, we propose to analyze longer term videos monitoring a large flock of birds. This project, which is part of the NSF-TIH Indo-US joint R&D partnership, focuses on solving challenges associated with the analysis of long-term videos captured at feeding grounds and nesting sites, among other such locations that host large flocks of migratory birds. We foresee that the objectives of this project would lead to datasets and benchmarking tools as well as novel algorithms that would be instrumental in developing automated video analysis tools that could in turn help understand individual and social behavior of birds. The first of the key outcomes of this research will include the curation of challenging, real-world datasets for benchmarking various image and video analytics algorithms for tasks such as counting, detection, segmentation, and tracking. Our recent efforts towards this outcome is a curated dataset of 812 high-resolution, point-annotated, images (4K - 32MP) of a flock of Demoiselle cranes (Anthropoides virgo) taken from their feeding site at Khichan, Rajasthan, India. The average number of birds in each image is about 207, with a maximum count of 1500. The benchmark experiments show that state-of-the-art vision techniques struggle with tasks such as segmentation, detection, localization, and density estimation for the proposed dataset. Over the execution of this open science research, we will be scaling this dataset for segmentation and tracking in videos, as well as developing novel techniques for video analytics for wildlife monitoring. Kshitiz, Sonu Sreshtha, Ramy Mounir, Mayank Vatsa, Richa Singh 0001, Saket Anand, Sudeep Sarkar, Sevaram Mali Parihar |
IJCAI | 7 |
| 2023 | STREAMER: Streaming Representation Learning and Event Segmentation in a Hierarchical MannerabstractWe present a novel self-supervised approach for hierarchical representation learning and segmentation of perceptual inputs in a streaming fashion. Our research addresses how to semantically group streaming inputs into chunks at various levels of a hierarchy while simultaneously learning, for each chunk, robust global representations throughout the domain. To achieve this, we propose STREAMER, an architecture that is trained layer-by-layer, adapting to the complexity of the input domain. In our approach, each layer is trained with two primary objectives: making accurate predictions into the future and providing necessary information to other levels for achieving the same objective. The event hierarchy is constructed by detecting prediction error peaks at different levels, where a detected boundary triggers a bottom-up information flow. At an event boundary, the encoded representation of inputs at one layer becomes the input to a higher-level layer. Additionally, we design a communication module that facilitates top-down and bottom-up exchange of information during the prediction process. Notably, our model is fully self-supervised and trained in a streaming manner, enabling a single pass on the training data. This means that the model encounters each input only once and does not store the data. We evaluate the performance of our model on the egocentric EPIC-KITCHENS dataset, specifically focusing on temporal event segmentation. Furthermore, we conduct event retrieval experiments using the learned representations to demonstrate the high quality of our video event representations. Illustration videos and code are available on our project page: https://ramymounir.com/publications/streamer Ramy Mounir, Sujal Vijayaraghavan, Sudeep Sarkar |
NeurIPS | 3 |
| 2023 | Towards Automated Ethogramming: Cognitively-Inspired Event Segmentation for Streaming Wildlife Video MonitoringabstractAbstract Advances in visual perceptual tasks have been mainly driven by the amount, and types, of annotations of large-scale datasets. Researchers have focused on fully-supervised settings to train models using offline epoch-based schemes. Despite the evident advancements, limitations and cost of manually annotated datasets have hindered further development for event perceptual tasks, such as detection and localization of objects and events in videos. The problem is more apparent in zoological applications due to the scarcity of annotations and length of videos-most videos are at most ten minutes long. Inspired by cognitive theories, we present a self-supervised perceptual prediction framework to tackle the problem of temporal event segmentation by building a stable representation of event-related objects. The approach is simple but effective. We rely on LSTM predictions of high-level features computed by a standard deep learning backbone. For spatial segmentation, the stable representation of the object is used by an attention mechanism to filter the input features before the prediction step. The self-learned attention maps effectively localize the object as a side effect of perceptual prediction. We demonstrate our approach on long videos from continuous wildlife video monitoring, spanning multiple days at 25 FPS. We aim to facilitate automated ethogramming by detecting and localizing events without the need for labels. Our approach is trained in an online manner on streaming input and requires only a single pass through the video, with no separate training set. Given the lack of long and realistic (includes real-world challenges) datasets, we introduce a new wildlife video dataset–nest monitoring of the Kagu (a flightless bird from New Caledonia)–to benchmark our approach. Our dataset features a video from 10 days (over 23 million frames) of continuous monitoring of the Kagu in its natural habitat. We annotate every frame with bounding boxes and event labels. Additionally, each frame is annotated with time-of-day and illumination conditions. We will make the dataset, which is the first of its kind, and the code available to the research community. We find that the approach significantly outperforms other self-supervised, traditional (e.g., Optical Flow, Background Subtraction) and NN-based (e.g., PA-DPC, DINO, iBOT), baselines and performs on par with supervised boundary detection approaches (i.e., PC). At a recall rate of 80%, our best performing model detects one false positive activity every 50 min of training. On average, we at least double the performance of self-supervised approaches for spatial segmentation. Additionally, we show that our approach is robust to various environmental conditions (e.g., moving shadows). We also benchmark the framework on other datasets (i.e., Kinetics-GEBD, TAPOS) from different domains to demonstrate its generalizability. The data and code are available on our project page: https://aix.eng.usf.edu/research_automated_ethogramming.html Ramy Mounir, Ahmed Shahabaz, Roman Gula, Jörn Theuerkauf, Sudeep Sarkar |
Int. J. Comput. Vis. | 5 |
| 2023 | Correction: Towards Automated Ethogramming: Cognitively-Inspired Event Segmentation for Streaming Wildlife Video Monitoring
Ramy Mounir, Ahmed Shahabaz, Roman Gula, Jörn Theuerkauf, Sudeep Sarkar |
Int. J. Comput. Vis. | 5 |
| 2023 | A systematic literature review on object detection using near infrared and thermal images
Nicolas Bustos, Mehrsa Mashhadi, Susana K. Lai-Yuen, Sudeep Sarkar, Tapas K. Das |
Neurocomputing | 4 |
| 2023 | Leveraging Symbolic Knowledge Bases for Commonsense Natural Language Inference Using Pattern TheoryabstractThe commonsense natural language inference (CNLI) tasks aim to select the most likely follow-up statement to a contextual description of ordinary, everyday events and facts. Current approaches to transfer learning of CNLI models across tasks require many labeled data from the new task. This paper presents a way to reduce this need for additional annotated training data from the new task by leveraging symbolic knowledge bases, such as ConceptNet. We formulate a teacher-student framework for mixed symbolic-neural reasoning, with the large-scale symbolic knowledge base serving as the teacher and a trained CNLI model as the student. This hybrid distillation process involves two steps. The first step is a symbolic reasoning process. Given a collection of unlabeled data, we use an abductive reasoning framework based on Grenander's pattern theory to create weakly labeled data. Pattern theory is an energy-based graphical probabilistic framework for reasoning among random variables with varying dependency structures. In the second step, the weakly labeled data, along with a fraction of the labeled data, is used to transfer-learn the CNLI model into the new task. The goal is to reduce the fraction of labeled data required. We demonstrate the efficacy of our approach by using three publicly available datasets (OpenBookQA, SWAG, and HellaSWAG) and evaluating three CNLI models (BERT, LSTM, and ESIM) that represent different tasks. We show that, on average, we achieve 63% of the top performance of a fully supervised BERT model with no labeled data. With only 1,000 labeled samples, we can improve this performance to 72%. Interestingly, without training, the teacher mechanism itself has significant inference power. The pattern theory framework achieves 32.7% accuracy on OpenBookQA, outperforming transformer-based models such as GPT (26.6%), GPT-2 (30.2%), and BERT (27.1%) by a significant margin. We demonstrate that the framework can be generalized to successfully train neural CNLI models using knowledge distillation under unsupervised and semi-supervised learning settings. Our results show that it outperforms all unsupervised and weakly supervised baselines and some early supervised approaches, while offering competitive performance with fully supervised baselines. Additionally, we show that the abductive learning framework can be adapted for other downstream tasks, such as unsupervised semantic textual similarity, unsupervised sentiment classification, and zero-shot text classification, without significant modification to the framework. Finally, user studies show that the generated interpretations enhance its explainability by providing key insights into its reasoning mechanism. Sathyanarayanan N. Aakur, Sudeep Sarkar |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Actor-Centered Representations for Action Localization in Streaming Videos
Sathyanarayanan N. Aakur, Sudeep Sarkar |
ECCV (38) | 2 |
| 2022 | Bayesian Tracking of Video Graphs Using Joint Kalman Smoothing and Registration
Aditi Basu Bal, Ramy Mounir, Sathyanarayanan N. Aakur, Sudeep Sarkar, Anuj Srivastava |
ECCV (35) | 4 |
| 2021 | Virtual special issue on novel data-representation and classification techniques
José Arturo Olvera-López, Joaquín Salas, Jesús Ariel Carrasco-Ochoa, José Fco. Martínez-Trinidad, Sudeep Sarkar |
Pattern Recognit. Lett. | 5 |
| 2021 | Measuring Human and Economic Activity From Satellite Imagery to Support City-Scale Decision-Making During COVID-19 PandemicabstractThe COVID-19 outbreak forced governments worldwide to impose lockdowns and quarantines to prevent virus transmission. As a consequence, there are disruptions in human and economic activities all over the globe. The recovery process is also expected to be rough. Economic activities impact social behaviors, which leave signatures in satellite images that can be automatically detected and classified. Satellite imagery can support the decision-making of analysts and policymakers by providing a different kind of visibility into the unfolding economic changes. In this article, we use a deep learning approach that combines strategic location sampling and an ensemble of lightweight convolutional neural networks (CNNs) to recognize specific elements in satellite images that could be used to compute economic indicators based on it, automatically. This CNN ensemble framework ranked third place in the US Department of Defense xView challenge, the most advanced benchmark for object detection in satellite images. We show the potential of our framework for temporal analysis using the US IARPA Function Map of the World (fMoW) dataset. We also show results on real examples of different sites before and after the COVID-19 outbreak to illustrate different measurable indicators. Our code and annotated high-resolution aerial scenes before and after the outbreak are available on GitHub.1.https://github.com/maups/covid19-satellite-analysis. Rodrigo Minetto, Maurício Pamplona Segundo, Gilbert Rotich, Sudeep Sarkar |
IEEE Trans. Big Data | 4 |
| 2020 | Action Localization Through Continual Predictive Learning
Sathyanarayanan N. Aakur, Sudeep Sarkar |
ECCV (14) | 2 |
| 2019 | A Perceptual Prediction Framework for Self Supervised Event SegmentationabstractTemporal segmentation of long videos is an important problem, that has largely been tackled through supervised learning, often requiring large amounts of annotated training data. In this paper, we tackle the problem of self-supervised temporal segmentation that alleviates the need for any supervision in the form of labels (full supervision) or temporal ordering (weak supervision). We introduce a self-supervised, predictive learning framework that draws inspiration from cognitive psychology to segment long, visually complex videos into constituent events. Learning involves only a single pass through the training data. We also introduce a new adaptive learning paradigm that helps reduce the effect of catastrophic forgetting in recurrent neural networks. Extensive experiments on three publicly available datasets - Breakfast Actions, 50 Salads, and INRIA Instructional Videos datasets show the efficacy of the proposed approach. We show that the proposed approach outperforms weakly-supervised and unsupervised baselines by up to 24% and achieves competitive segmentation results compared to fully supervised baselines with only a single pass through the training data. Finally, we show that the proposed self-supervised learning paradigm learns highly discriminating features to improve action recognition. Sathyanarayanan N. Aakur, Sudeep Sarkar |
CVPR | 2 |
| 2019 | Going Deeper With Semantics: Video Activity Interpretation Using Semantic ContextualizationabstractA deeper understanding of video activities extends beyond recognition of underlying concepts such as actions and objects: constructing deep semantic representations requires reasoning about the semantic relationships among these concepts, often beyond what is directly observed in the data. To this end, we propose an energy minimization framework that leverages large-scale commonsense knowledge bases, such as ConceptNet, to provide contextual cues to establish semantic relationships among entities directly hypothesized from video. We mathematically express this using the language of Grenander's canonical pattern generator theory. We show that the use of prior encoded commonsense knowledge alleviate the need for large annotated training datasets and help tackle imbalance in training through prior knowledge. Using three different publicly available datasets - Charades, Microsoft Visual Description Corpus and Breakfast Actions datasets, we show that the proposed model can generate video interpretations whose quality is better than those reported by state-of-the-art approaches, which have substantial training needs. Through extensive experiments, we show that the use of commonsense knowledge from ConceptNet allows the proposed approach to handling various challenges such as training data imbalance, weak features and complex semantic relationships and visual scenes. Sathyanarayanan N. Aakur, Fillipe D. M. de Souza, Sudeep Sarkar |
WACV | 3 |
| 2019 | Hydra: An Ensemble of Convolutional Neural Networks for Geospatial Land ClassificationabstractIn this paper, we describe Hydra, an ensemble of convolutional neural networks (CNNs) for geospatial land classification. The idea behind Hydra is to create an initial CNN that is coarsely optimized but provides a good starting pointing for further optimization, which will serve as the Hydra's body. Then, the obtained weights are fine-tuned multiple times with different augmentation techniques, crop styles, and classes weights to form an ensemble of CNNs that represent the Hydra's heads. By doing so, we prompt convergence to different endpoints, which is a desirable aspect for ensembles. With this framework, we were able to reduce the training time while maintaining the classification performance of the ensemble. We created ensembles for our experiments using two state-of-the-art CNN architectures, residual network (ResNet), and dense convolutional networks (DenseNet). We have demonstrated the application of our Hydra framework in two data sets, functional map of world (FMOW) and NWPU-RESISC45, achieving results comparable to the state-of-the-art for the former and the best-reported performance so far for the latter. Code and CNN models are available at https://github.com/maups/hydra-fmow. Rodrigo Minetto, Maurício Pamplona Segundo, Sudeep Sarkar |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Evaluation of Algorithms for Orientation Invariant Inertial Gait MatchingabstractWith the prevalent use of smart phones in sensitive applications, unobtrusive methods for continuously verifying the identity of the user have become critical. The embedded inertial sensors in these devices provide an opportunity to develop authentication processes based on behavioral biometrics such as gait. However, one major obstacle is that the orientation of the device relative to the user is hard to control and difficult to determine reliably. This paper presents five methods: magnitude (MAG), principal component analysis (PCA), vector cross product (VCP), reduced gait dynamics image (rGDI), and Kabsch alignment (KAB) that make the authentication process independent of device orientation and hence improve the performance. The five methods are evaluated and compared on two large, publicly available, inertial gait datasets. The baseline (orientation dependent) average equal error rate (EER) when the device was freely oriented is 26.4%. The MAG, PCA, VCP, and rGDI methods reduce the average EER to approximately 23%. The Kabsch (KAB) method is more effective and reduces the average EER to 20.2%. Ravi Subramanian, Sudeep Sarkar |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | A Graph-based Approach for Static Ensemble Selection in Remote Sensing Image AnalysisabstractMany works in the literature have used machine learning techniques to solve their classification problems in different knowledge areas, e.g., medicine, agriculture, and remote sensing. Since there is no a single machine learning technique that achieves the best results for all kind of applications, a good alternative is the fusion of classification techniques, also known as multiple classifier systems (MCS). A common challenge in MCS is the selection of a few classifiers among many classifiers that are available in the literature; using all possible classifiers is not a feasible alternative. The choice of the classifiers becomes an essential factor, i.e., we need an ensemble selection approach. In this work, we propose a novel graph-based approach for static ensemble selection (GASES) to find or choose the best classifier set for remote sensing image classification. Experiments demonstrate that GASES improves performance by up to 70% over different baseline approaches when fusing classifiers. It decreases the number of classifiers used while retaining the effectiveness of using all of the classifiers. Furthermore, our proposed method is a more straightforward and intuitive technique for static ensemble selection scheme than other baseline approaches such as Consensus and Kendall. Fábio Augusto Faria, Sudeep Sarkar |
ICPR | 2 |
| 2018 | Distance metric learning for pattern recognition
Jiwen Lu, Ruiping Wang 0001, Ajmal Mian, Sudeep Sarkar |
Pattern Recognit. | 5 |
| 2017 | Spatially Coherent Interpretations of Videos Using Pattern Theory
Fillipe D. M. de Souza, Sudeep Sarkar, Anuj Srivastava, Jingyong Su |
Int. J. Comput. Vis. | 2 |
| 2016 | Learning Camera Viewpoint Using CNN to Improve 3D Body Pose EstimationabstractThe objective of this work is to estimate 3D human pose from a single RGB image. Extracting image representations which incorporate both spatial relation of body parts and their relative depth plays an essential role in accurate3D pose reconstruction. In this paper, for the first time, we show that camera viewpoint in combination to 2D joint locations significantly improves 3D pose accuracy without the explicit use of perspective geometry mathematical models. To this end, we train a deep Convolutional Neural Net-work (CNN) to learn categorical camera viewpoint. To make the network robust against clothing and body shape of the subject in the image, we utilized 3D computer rendering to synthesize additional training images. We test our framework on the largest 3D pose estimation bench-mark, Human3.6m, and achieve up to 20% error reduction on standing-pose activities compared to the state-of-the-art approaches that do not use body part segmentation. Mona Fathollahi Ghezelghieh, Rangachar Kasturi, Sudeep Sarkar |
3DV | 3 |
| 2016 | Building semantic understanding beyond deep learning from sound and visionabstractDeep learning-based models have recently been widely successful at outperforming traditional approaches in several computer vision applications such as image classification, object recognition and action recognition. However, those models are not naturally designed to learn structural information that can be important to tasks such as human pose estimation and structured semantic interpretation of video events. In this paper, we demonstrate how to build structured semantic understanding of audio-video events by reasoning on multiple-label decisions of deep visual models and auditory models using Grenander's structures for imposing semantic consistency. The proposed structured model does not require joint training of the structural semantic dependencies and deep models. Instead they are independent components linked by Grenander's structures. Furthermore, we exploited Grenander's structures as a means to facilitate and enrich the model with fusion of multimodal sensory data; in particular, auditory features with visual features. Overall, we observed improvements in the quality of semantic interpretations using deep models and auditory features in combination with Grenander's structures, reflecting as numerical improvements of up to 11.5% and 12.3% in precision and recall, respectively. Fillipe D. M. de Souza, Sudeep Sarkar, Guillermo Cámara Chávez |
ICPR | 2 |
| 2016 | Pattern theory for representation and inference of semantic structures in videos
Fillipe D. M. de Souza, Sudeep Sarkar, Anuj Srivastava, Jingyong Su |
Pattern Recognit. Lett. | 2 |
| 2015 | Temporally coherent interpretations for long videos using pattern theoryabstractGraph-theoretical methods have successfully provided semantic and structural interpretations of images and videos. A recent paper introduced a pattern-theoretic approach that allows construction of flexible graphs for representing interactions of actors with objects and inference is accomplished by an efficient annealing algorithm. Actions and objects are termed generators and their interactions are termed bonds; together they form high-probability configurations, or interpretations, of observed scenes. This work and other structural methods have generally been limited to analyzing short videos involving isolated actions. Here we provide an extension that uses additional temporal bonds across individual actions to enable semantic interpretations of longer videos. Longer temporal connections improve scene interpretations as they help discard (temporally) local solutions in favor of globally superior ones. Using this extension, we demonstrate improvements in understanding longer videos, compared to individual interpretations of non-overlapping time segments. We verified the success of our approach by generating interpretations for more than 700 video segments from the YouCook data set, with intricate videos that exhibit cluttered background, scenarios of occlusion, viewpoint variations and changing conditions of illumination. Interpretations for long video segments were able to yield performance increases of about 70% and, in addition, proved to be more robust to different severe scenarios of classification errors. Fillipe D. M. de Souza, Sudeep Sarkar, Anuj Srivastava, Jingyong Su |
CVPR | 2 |
| 2015 | Conditional distance based matching for one-shot gesture recognition
RaviKiran Krishnan, Sudeep Sarkar |
Pattern Recognit. | 2 |
| 2014 | Rate-Invariant Analysis of Trajectories on Riemannian Manifolds with Application in Visual Speech RecognitionabstractIn statistical analysis of video sequences for speech recognition, and more generally activity recognition, it is natural to treat temporal evolutions of features as trajectories on Riemannian manifolds. However, different evolution patterns result in arbitrary parameterizations of these trajectories. We investigate a recent framework from statistics literature that handles this nuisance variability using a cost function/distance for temporal registration and statistical summarization & modeling of trajectories. It is based on a mathematical representation of trajectories, termed transported square-root vector field (TSRVF), and the L2 norm on the space of TSRVFs. We apply this framework to the problem of speech recognition using both audio and visual components. In each case, we extract features, form trajectories on corresponding manifolds, and compute parametrization-invariant distances using TSRVFs for speech classification. On the OuluVS database the classification performance under metric increases significantly, by nearly 100% under both modalities and for all choices of features. We obtained speaker-dependent classification rate of 70% and 96% for visual and audio components, respectively. Jingyong Su, Anuj Srivastava, Fillipe D. M. de Souza, Sudeep Sarkar |
CVPR | 4 |
| 2014 | Optical Flow Based Expression Suppression in VideoabstractIn this paper we propose a novel method for suppressing facial expressions in video sequences based on analysis of apparent strain in the face. The proposed method performs continuous optical strain analysis on a target face for all frames in a video. This analysis is used in conjunction with strain maps for every frame to counteract the elastic deformation of the facial tissue and keep the expression neutral. The method is capable of suppressing out all expression types without the need of training for specific expressions and is demonstrated on a publicly available data set. In addition, we test our method using publicly available expression (smile) recognition program that is included with OpenCV to quantify the suppression results. Our method shows an average reduction of expression detection confidence, the confidence of an expression present in a single frame, of nearly 90%. Jesse Brizzi, Dmitry B. Goldgof, Sudeep Sarkar, Matthew Shreve |
ICPR | 3 |
| 2014 | Pattern Theory-Based Interpretation of ActivitiesabstractWe present a novel framework, based on Germander's pattern theoretic concepts, for high-level interpretation of video activities. This framework allows us to elegantly integrate ontological constraints and machine learning classifiers in one formalism to construct high-level semantic interpretations that describe video activity. The unit of analysis is a generator that could represent either an ontological label as well as a group of features from a video. These generators are linked using bonds with different constraints. An interpretation of a video is a configuration of these connected generators, which results in a graph structure that is richer than conventional graphs used in computer vision. The quality of the interpretation is quantified by an energy function that is optimized using Markov Chain Monte Carlo based simulated annealing. We demonstrate the superiority of our approach over a purely machine learning based approach (SVM) using more than 650 video shots from the You Cook dataset. This dataset is very challenging in terms of complexity of background, presence of camera motion, object occlusion, clutter, and actor variability. We find significantly improved performance in nearly all cases. Our results show that the pattern theory inference process is able to construct the correct interpretation by leveraging the ontological constraints even when the machine learning classifier is poor and the most confident labels are wrong. Fillipe D. M. de Souza, Sudeep Sarkar, Anuj Srivastava, Jingyong Su |
ICPR | 2 |
| 2014 | A novel telerobotic method for human-in-the-loop assisted grasping based on intention recognitionabstractIn this work, we present a methodology for enabling a robot to identify an object and grasp configuration of interest and assist the human teleoperating the robot, to grasp the object. The identification is carried out in real-time by detecting the motion intention of the human as they are teleoperating the remote robotic arm towards the object and the grasp configuration. Simultaneously, depending on the detected object and grasp configuration, the human user is assisted to translate and orient the remote arm gripper in order to preshape and grasp the object. The complete process occurs with the human teleoperating the arm, and without them having to interact with another interface. Motion intention recognition is carried out by using Hidden Markov Models (HMMs), trained offline by preshape trials performed by a skilled teleoperator. The environment is unstructured and comprises of a number of objects, each with multiple grasp configurations. Experimental tests on healthy human subjects have validated our intention recognition based assistance method. They show that the method allows objects to be grasped and placed 48% faster, and with much ease compared to unassisted teleoperation. Moreover, we have proved that the model for intention recognition, trained by a skilled teleoperator, can be used by novice users to efficiently execute a grasping task in teleoperation. Karan Khokar, Redwan Alqasemi, Sudeep Sarkar, Kyle B. Reed, Rajiv V. Dubey |
ICRA | 3 |
| 2014 | Automatic expression spotting in videos
Matthew Shreve, Jesse Brizzi, Sergiy Fefilatyev, Timur Luguev, Dmitry B. Goldgof, Sudeep Sarkar |
Image Vis. Comput. | 6 |
| 2014 | Orthogonal projection images for 3D face detection
Maurício Pamplona Segundo, Luciano Silva, Olga R. P. Bellon, Sudeep Sarkar |
Pattern Recognit. Lett. | 4 |
| 2014 | High-resolution 3D surface strain magnitude using 2D camera and low-resolution depth sensor
Matthew Shreve, Maurício Pamplona Segundo, Timur Luguev, Dmitry B. Goldgof, Sudeep Sarkar |
Pattern Recognit. Lett. | 5 |
| 2014 | SIBGRAPI 25th: Advances in Pattern Recognition and Computer Vision
Luciano Silva, Sudeep Sarkar, Carla M. D. S. Freitas, Roberto Scopigno |
Pattern Recognit. Lett. | 2 |
| 2013 | Multi-scale superquadric fitting for efficient shape and pose recovery of unknown objectsabstractRapidly acquiring the shape and pose information of unknown objects is an essential characteristic of modern robotic systems in order to perform efficient manipulation tasks. In this work, we present a framework for 3D geometric shape recovery and pose estimation from unorganized point cloud data. We propose a low latency multi-scale voxelization strategy that rapidly fits superquadrics to single view 3D point clouds. As a result, we are able to quickly and accurately estimate the shape and pose parameters of relevant objects in a scene. We evaluate our approach on two datasets of common household objects collected using Microsoft's Kinect sensor. We also compare our work to the state of the art and achieve comparable results in less computational time. Our experimental results demonstrate the efficacy of our approach. Kester Duncan, Sudeep Sarkar, Redwan Alqasemi, Rajiv V. Dubey |
ICRA | 2 |
| 2013 | Hop-Diffusion Monte Carlo for Epipolar Geometry Estimation between Very Wide-Baseline ImagesabstractWe present a Monte Carlo approach for epipolar geometry estimation that efficiently searches for minimal sets of inlier correspondences in the presence of many outliers in the putative correspondence set, a condition that is prevalent when we have wide baselines, significant scale changes, rotations in depth, occlusion, and repeated patterns. The proposed Monte Carlo algorithm uses Balanced LOcal and Global Search (BLOGS) to find the best minimal set of correspondences. The local search is a diffusion process using Joint Feature Distributions that captures the dependencies among the correspondences. And, the global search is a hopping search process across the minimal set space controlled by photometric properties. Using a novel experimental protocol that involves computing errors for manually marked ground truth points and images with outlier rates as high as 90 percent, we find that BLOGS is better than related approaches such as MAPSAC, NAPSAC, and BEEM. BLOGS results are of similar quality as other approaches, but BLOGS generate them in 10 times fewer iterations. The time per iteration for BLOGS is also the lowest among the ones we studied. Aveek Shankar Brahmachari, Sudeep Sarkar |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Relational entropy-based saliency detection in images and videosabstractSalient regions in an image facilitate the non-uniform allocation of computational resources to just the interesting parts of an image. In this paper, we present a saliency detection mechanism using relational distributions that capture geometric statistics based on distance and gradient direction relationships between pixels. The entropy of these normalized distributions is related to saliency. We employ an efficient technique for calculating the Rényi entropy of the probabilistic relational distributions using Parzen window weighted samples, thus eliminating the need for constructing intermediate histogram representations. We quantitatively demonstrate the biological plausibility of our method by showing how the saliency maps produced strongly correlate to human fixations in still images and to dominant objects in video. We find that our approach is better than six other saliency models. Kester Duncan, Sudeep Sarkar |
ICIP | 2 |
| 2012 | Finding recurrent patterns from continuous sign language sentences for automated extraction of signs
Sunita Nayak, Kester Duncan, Sudeep Sarkar, Barbara L. Loeding |
J. Mach. Learn. Res. | 3 |
| 2011 | Evaluation of Facial Reconstructive Surgery on Patients with Facial Palsy Using Optical Strain
Matthew Shreve, Neeha Jain, Dmitry B. Goldgof, Sudeep Sarkar, Walter G. Kropatsch, Chieh-Han John Tzou, Manfred Frey |
CAIP (1) | 4 |
| 2011 | Macro- and micro-expression spotting in long videos using spatio-temporal strainabstractWe propose a method for the automatic spotting (temporal segmentation) of facial expressions in long videos comprising of macro- and micro-expressions. The method utilizes the strain impacted on the facial skin due to the non-rigid motion caused during expressions. The strain magnitude is calculated using the central difference method over the robust and dense optical flow field observed in several regions (chin, mouth, cheek, forehead) on each subject's face. This new approach is able to successfully detect and distinguish between large expressions (macro) and rapid and localized expressions (micro). Extensive testing was completed on a dataset containing 181 macro-expressions and 124 micro-expressions. The dataset consists of 56 videos collected at USF, 6 videos from the Canal-9 political debates, and 3 low quality videos found on the internet. A spotting accuracy of 85% was achieved for macro-expressions and 74% of all micro-expressions were spotted. Matthew Shreve, Sridhar Godavarthy, Dmitry B. Goldgof, Sudeep Sarkar |
FG | 4 |
| 2011 | QCAPro - An error-power estimation tool for QCA circuit designabstractIn this work we present a novel probabilistic modeling tool (QCAPro) to estimate polarization error and non- adiabatic switching power loss in Quantum-dot Cellular Automata (QCA) circuits. The tool uses a fast approximation based technique to estimate highly erroneous cells in QCA circuit design. QCAPro also provides an estimate of power loss in a QCA circuit for clocks with sharp transitions, which result in non-adiabatic operations and provides an upper bound of power expended. QCAPro can be used to estimate average power loss, maximum and minimum power loss in a QCA circuit during an input switching operation. This work will provide a good platform for researchers who wish to study polarization error and power dissipation related issues in QCA circuits. Saket Srivastava, Arjun Asthana, Sanjukta Bhanja, Sudeep Sarkar |
ISCAS | 4 |
| 2011 | Fast detection of noisy GPS and magnetometer tags in wide-baseline multi-viewsabstractWe propose an algorithm for detection of noisy GPS and magnetometer tags in wide-baseline camera views. Our algorithm neither needs densely sampled views nor does it need a single visually connected path through all the views in the dataset. We use vision-based estimates of mutual rotation and translation between cameras to compute a measure of confidence on the correctness of the associated GPS and magnetometer tags. The vision algorithm can find the epipolar geometry between two wide-baseline images without needing pre-specified correspondences. We have two versions of our approach; one that requires geometric pose estimation between all pairs of images and a faster version that uses a pre-filter based on photometric comparison of images to quickly reject non-overlapping views from further geometric consideration. We show qualitative and quantitative results on the Nokia Grand Challenge 2010 Dataset. We find that magnetometer readings are more accurate than GPS readings. Aveek Shankar Brahmachari, Sudeep Sarkar |
ACM Multimedia | 2 |
| 2010 | Detecting Group Turn Patterns in Conversations Using Audio-Video Change Scale-SpaceabstractAutomatic analysis of conversations is important for extracting high-level descriptions of meetings. In this work, as an alternative to linguistic approaches, we develop a novel, purely bottom-up representation, constructed from both audio and video signals that help us characterize and build a rich description of the content at multiple temporal scales. We consider the evolution of the detected change, using Bayesian Information Criterion (BIC) at multiple temporal scales to build an audio-visual change scale space. Peaks detected in this representation, yields group-turn based conversational changes at different temporal scales. Conversation overlaps, changes and their inferred models offer an intermediate-level description of meeting videos that can be useful in summarization and indexing of meetings. Results on NIST meeting room dataset showed a true positive rate of 88%. RaviKiran Krishnan, Sudeep Sarkar |
ICPR | 2 |
| 2010 | Modeling Facial Skin Motion Properties in Video and Its Application to Matching Faces across ExpressionsabstractIn this paper, we propose a method to model the material constants (Young's modulus) of the skin in subregions of the face from the motion observed in multiple facial expressions and present its relevance to an image analysis task such as face verification. On a public database consisting of 40 subjects undergoing some set of facial motions associated with anger, disgust, fear, happy, sad, and surprise expressions, we present an expression invariant strategy to matching faces using the Young's modulus of the skin. Results show that it is indeed possible to match faces across expressions using the material properties of their skin. Vasant Manohar, Matthew Shreve, Dmitry B. Goldgof, Sudeep Sarkar |
ICPR | 4 |
| 2010 | Handling Movement Epenthesis and Hand Segmentation Ambiguities in Continuous Sign Language Recognition Using Nested Dynamic ProgrammingabstractWe consider two crucial problems in continuous sign language recognition from unaided video sequences. At the sentence level, we consider the movement epenthesis (me) problem and at the feature level, we consider the problem of hand segmentation and grouping. We construct a framework that can handle both of these problems based on an enhanced, nested version of the dynamic programming approach. To address movement epenthesis, a dynamic programming (DP) process employs a virtual me option that does not need explicit models. We call this the enhanced level building (eLB) algorithm. This formulation also allows the incorporation of grammar models. Nested within this eLB is another DP that handles the problem of selecting among multiple hand candidates. We demonstrate our ideas on four American Sign Language data sets with simple background, with the signer wearing short sleeves, with complex background, and across signers. We compared the performance with Conditional Random Fields (CRF) and Latent Dynamic-CRF-based approaches. The experiments show more than 40 percent improvement over CRF or LDCRF approaches in terms of the frame labeling rate. We show the flexibility of our approach when handling a changing context. We also find a 70 percent improvement in sign recognition rate over the unenhanced DP matching algorithm that does not accommodate the me effect. Ruiduo Yang, Sudeep Sarkar, Barbara L. Loeding |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Vision-IMU Integration Using a Slow-Frame-Rate Monocular Vision System in an Actual Roadway SettingabstractWe present results of an effort where position and orientation data from vision and inertial sensors are integrated and validated using data from an actual roadway. Information from a sequence of images, which were captured by a monocular camera attached to a survey vehicle at a maximum frequency of 3 frames/s, is fused with position and orientation estimates from the inertial system to correct for inherent error accumulation in such integral-based systems. The rotations and translations are estimated from point correspondences tracked through a sequence of images. To reduce unsuitable correspondences, we used constraints such asepipolar linesandcorrespondence flow directions. The vision algorithm automatically operates and involves the identification of point correspondences, the pruning of correspondences, and the estimation of motion parameters. To simply obtain the geodetic coordinates, i.e., latitude, longitude, and altitude, from the translation-direction estimates from the vision sensor, we expand the Kalman filter space to incorporate distance. Hence, it was possible to extract the translational vector from the available translational direction estimate of the vision system. Finally, a decentralized Kalman filter is used to integrate the position estimates based on the vision sensor with those of the inertial system. The fusion of the two sensors was carried out at the system level in the model. The comparison of integrated vision–inertial-measuring-unit (IMU) position estimates with those from inertial–GPS system output and actual survey demonstrates that vision sensing can be used to reduce errors in inertial measurements during potential GPS outages. Duminda I. B. Randeniya, Sudeep Sarkar, Manjriker Gunaratne |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2009 | Automated extraction of signs from continuous sign language sentences using Iterated Conditional ModesabstractRecognition of signs in sentences requires a training set constructed out of signs found in continuous sentences. Currently, this is done manually, which is a tedious process. In this work, we consider a framework where the modeler just provides multiple video sequences of sign language sentences, constructed to contain the vocabulary of interest. We learn the models of the recurring signs, automatically. Specifically, we automatically extract the parts of the signs that are present in most occurrences of the sign in context. These parts of the signs that is stable with respect to adjacent signs, are referred to as signemes. Each video is first transformed into a multidimensional time series representation, capturing the motion and shape aspects of the sign. We then extract signemes from multiple sentences, concurrently, using Iterated Conditional Modes (ICM). We show results by learning multiple instances of 10 different signs from a set of 136 sign language sentences. We classify the extracted signemes as correct, partially correct or incorrect depending on whether both the start and end locations are correct, only one of them is correct or both are incorrect, respectively. Out of the 136 extracted video signemes, 98 were correct, 20 were partially correct and 18 were incorrect. To demonstrate the generality of the unsupervised modeling idea, we also show the ability to automatically extract common spoken words in audio. We consider the English glosses (spoken) corresponding to the sign language sentences and extract the audio counterparts of the signs. Of the 136 such instances, we recovered 127 correct, 8 partially correct, and 1 incorrect representation of the words. Sunita Nayak, Sudeep Sarkar, Barbara L. Loeding |
CVPR | 2 |
| 2009 | BLOGS: Balanced local and global search for non-degenerate two view epipolar geometryabstractThis work considers the problem of estimating the epipolar geometry between two cameras without needing a prespecified set of correspondences. It is capable of resolving the epipolar geometry for cases when the views differ significantly in terms of baseline and rotation, resulting in a large number features in one image that have no correspondence in the other image. We do conditional characterization of the probability space of correspondences based on Joint Feature Distributions (JFD). We seek to maximize the probabilistic support of the putative correspondence set over a number of MCMC iterations, guided by proposal distributions based on similarity or JFD. Similarity based guidance provides large movements (global) through correspondence space and JFD based guidance provides small movements (local) around the best known epipolar geometry the algorithm has found so far. We also propose a simple and novel method to rule out, at each iteration, correspondences that lead to degenerate configurations, thus speeding up convergence. We compare our algorithm with LO-RANSAC, NAPSAC, MAPSAC and BEEM, which are the current state of the art competing methods, on a dataset that has significantly more change in baseline, rotation, and scale than those used in the current literature. We quantitatively benchmark the performance using manually specified ground truth corresponding point pairs. We find that our approach can achieve results of similar quality as the current state of art in 10 times lesser number of iterations. We are also able to tolerate upto 90% outlier correspondences. Aveek Shankar Brahmachari, Sudeep Sarkar |
ICCV | 2 |
| 2009 | Towards macro- and micro-expression spotting in video using strain patternsabstractThis paper presents a novel method for automatic spotting (temporal segmentation) of facial expressions in long videos comprising of continuous and changing expressions. The method utilizes the strain impacted on the facial skin due to the non-rigid motion caused during expressions. The strain magnitude is calculated using the central difference method over the robust and dense optical flow field of each subjects face. Testing has been done on 2 datasets (which includes 100 macro-expressions) and promising results have been obtained. The method is robust to several common drawbacks found in automatic facial expression segmentation including moderate in-plane and out-of-plane motion. Additionally, the method has also been modified to work with videos containing micro-expressions. Micro-expressions are detected utilizing their smaller spatial and temporal extent. A subject's face is divided in to sub-regions (mouth, cheeks, forehead, and eyes) and facial strain is calculated for each of these regions. Strain patterns in individual regions are used to identify subtle changes which facilitate the detection of micro-expressions. Matthew Shreve, Sridhar Godavarthy, Vasant Manohar, Dmitry B. Goldgof, Sudeep Sarkar |
WACV | 5 |
| 2009 | Coupled grouping and matching for sign and gesture recognition
Ruiduo Yang, Sudeep Sarkar |
Comput. Vis. Image Underst. | 2 |
| 2009 | Distribution-Based Dimensionality Reduction Applied to Articulated Motion RecognitionabstractSome articulated motion representations rely on frame-wise abstractions of the statistical distribution of low-level features such as orientation, color, or relational distributions. As configuration among parts changes with articulated motion, the distribution changes, tracing a trajectory in the latent space of distributions, which we call the configuration space. These trajectories can then be used for recognition using standard techniques such as dynamic time warping. The core theory in this paper concerns embedding the frame-wise distributions, which can be looked upon as probability functions, into a low-dimensional space so that we can estimate various meaningful probabilistic distances such as the Chernoff, Bhattacharya, Matusita, Kullback-Leibler (KL) or symmetric-KL distances based on dot products between points in this space. Apart from computational advantages, this representation also affords speed-normalized matching of motion signatures. Speed normalized representations can be formed by interpolating the configuration trajectories along their arc lengths, without using any knowledge of the temporal scale variations between the sequences. We experiment with five different probabilistic distance measures and show the usefulness of the representation in three different contexts-sign recognition (with large number of possible classes), gesture recognition (with person variations), and classification of human-human interaction sequences (with segmentation problems). We find the importance of using the right distance measure for each situation. The low-dimensional embedding makes matching two to three times faster, while achieving recognition accuracies that are close to those obtained without using a low-dimensional embedding. We also empirically establish the robustness of the representation with respect to low-level parameters, embedding parameters, and temporal-scale parameters. Sunita Nayak, Sudeep Sarkar, Barbara L. Loeding |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Robust outdoor text detection using text intensity and shape featuresabstractRecognizing texts from camera images is a known hard problem because of the difficulties in text segmentation from the varied and complicated backgrounds. In this paper, we propose an algorithm that employs two novel filters and a basic component-based text detection framework. The framework uses the Niblack algorithm to threshold images and groups components into regions with commonly used geometry features. The intensity filter considers the overlap between the intensity histogram of a component and that of its adjoining area. For non-text regions, we have found that this overlap is large, and so we can prune out components with large values of this measure. The shape filter, on the other hand, deletes regions whose constituent components come from a same object, as most words consist of different characters. The proposed method is evaluated with the text locating database with 249 images used in the ICDAR2003 robust reading competition. The result shows that the algorithm is robust to both indoor images and outdoor images, even for the images of complex background, which usually is a hard factor to overcome for traditional component-based algorithms. In terms of performance statistics, we tested the algorithm on the ICDAR 2003 challenge experiment, and the algorithm achieves 66% precision rate (p), 46% recall rate (r), and 54% the combined rate ( f ), which is the best reported in the literature on this dataset. Sudeep Sarkar |
ICPR | 2 |
| 2008 | Finite element modeling of facial deformation in videos for computing strain patternabstractWe present a finite element modeling based approach to compute strain patterns caused by facial deformation during expressions in videos. A sparse motion field computed through a robust optical flow method drives the FE model. While the geometry of the model is generic, the material constants associated with an individualpsilas facial skin are learned at a coarse level sufficient for accurate strain map computation. Experimental results using the computational strategy presented in this paper emphasize the uniqueness and stability of strain maps across adverse data conditions (shadow lighting and face camouflage) making it a promising feature for image analysis tasks that can benefit from such auxiliary information. Vasant Manohar, Matthew Shreve, Dmitry B. Goldgof, Sudeep Sarkar |
ICPR | 4 |
| 2008 | Clip retrieval using multi-modal biometrics in meeting archivesabstractWe present a system to retrieve all clips from a meeting archive that show a particular individual speaking, using a single face or voice sample as the query. The system incorporates three novel ideas. One, rather than match the query to each individual sample in the archive, samples within a meeting are grouped first, generating a cluster of samples per individual. The query is then matched to the cluster, taking advantage of multiple samples to yield a robust decision. Two, automatic audio-visual association is performed which allows a bi-modal retrieval of clips, even when the query is uni-modal. Three, the biometric recognition uses individual-specific score distributions learnt from the clusters, in a likelihood ratio based decision framework that obviates the need for explicit normalization or modality weighting. The resulting system, which is completely automated, performs with 92.6% precision at 90% recall on a dataset of 16 real meetings spanning a total of 13 hours. Himanshu Vajaria, Sudeep Sarkar, Rangachar Kasturi |
ICPR | 2 |
| 2008 | Exploring Co-Occurence Between Speech and Body Movement for Audio-Guided Video LocalizationabstractThis paper presents a bottom-up approach that combines audio and video to simultaneously locate individual speakers in the video (2D source localization) and segment their speech (speaker diarization), in meetings recorded by a single stationary camera and a single microphone. The novelty lies in using motion information from the entire body rather than just the face to perform these tasks, which permits processing nonfrontal views, unlike previous work. Since body movements do not exhibit instantaneous signal-level synchrony with speech, the approach targets long term co-occurrences between audio and video subspaces. First, temporal clustering of the audio produces a large number of intermediate clusters, each containing speech from only a single speaker. Then, spatial clustering is performed in the video frames of each cluster by a novel eigen-analysis method to find the region of dominant motion. This region is associated with the speech assuming that a speaker exhibits more movement than the listeners. Thus, partial diarization and localization is obtained from the intermediate clusters. Speech from an intermediate cluster is modeled by a mixture of Gaussians and the speaker's location is represented by an eigen-blob model. In the ensuing iterative clustering stage, the diarization and localization results are progressively refined by merging the closest pair of clusters and updating the models until a stop criterion is met. Ideally, each final cluster contains all the speech from a single speaker and the corresponding eigen-blob model localizes the speaker in the image. Experiments conducted on 21 h of real data indicate that the proposed localization approach leads to a relative improvement of 40% over mutual information-based localization and that speaker diarization improves by 16% by incorporating visual information. The proposed approach does not require training and does not rely onapriorihand/face/person detection. Himanshu Vajaria, Sudeep Sarkar, Rangachar Kasturi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Subspace Approximation of Face Recognition Algorithms: An Empirical StudyabstractWe present a theory for constructing linear subspace approximations to face-recognition algorithms and empirically demonstrate that a surprisingly diverse set of face-recognition approaches can be approximated well by using a linear model. A linear model, built using a training set of face images, is specified in terms of a linear subspace spanned by, possibly nonorthogonal vectors. We divide the linear transformation used to project face images into this linear subspace into two parts: 1) a rigid transformation obtained through principal component analysis, followed by a nonrigid, affine transformation. The construction of the affine subspace involves embedding of a training set of face images constrained by the distances between them, as computed by the face-recognition algorithm being approximated. We accomplish this embedding by iterative majorization, initialized by classical MDS. Any new face image is projected into this embedded space using an affine transformation. We empirically demonstrate the adequacy of the linear model using six different face-recognition algorithms, spanning template-based and feature-based approaches, with a complete separation of the training and test sets. A subset of the face-recognition grand challenge training set is used to model the algorithms and the performance of the proposed modeling scheme is evaluated on the facial recognition technology (FERET) data set. The experimental results show that the average error in modeling for six algorithms is 6.3% at 0.001 false acceptance rate for the FERET fafb probe set which has 1195 subjects, the most among all of the FERET experiments. The built subspace approximation not only matches the recognition rate for the original approach, but the local manifold structure, as measured by the similarity of identity of nearest neighbors, is also modeled well. We found, on average, 87% similarity of the local neighborhood. We also demonstrate the usefulness of the linear model for algorithm-dependent indexing of face databases and find that it results in more than 20 times reduction in face comparisons for Bayesian, elastic bunch graph matching, and one proprietary algorithm. Pranab K. Mohanty, Sudeep Sarkar, Rangachar Kasturi, P. Jonathon Phillips |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2008 | Thermal Switching Error Versus Delay Tradeoffs in Clocked QCA CircuitsabstractThe quantum-dot cellular automata (QCA) model offers a novel nano-domain computing architecture by mapping the intended logic onto the lowest energy configuration of a collection of QCA cells, each with two possible ground states. A four-phased clocking scheme has been suggested to keep the computations at the ground state throughout the circuit. This clocking scheme, however, induces latency or delay in the transmission of information from input to output. In this paper, we study the interplay of computing error behavior with delay or latency of computation induced by the clocking scheme. Computing errors in QCA circuits can arise due to the failure of the clocking scheme to switch portions of the circuit to the ground state with change in input. Some of these non-ground states will result in output errors and some will not. The larger the size of each clocking zone, i.e., the greater the number of cells in each zone, the more the probability of computing errors. However, larger clocking zones imply faster propagation of information from input to output, i.e., reduced delay. Current QCA simulators compute just the ground state configuration of a QCA arrangement. In this paper, we offer an efficient method to compute the$N$-lowest energy modes of a clocked QCA circuit. We model the QCA cell arrangement in each zone using a graph-based probabilistic model, which is then transformed into a Markov tree structure defined over subsets of QCA cells. This tree structure allows us to compute the$N$-lowest energy configurations in an efficient manner by local message passing. We analyze the complexity of the model and show it to be polynomial in terms of the number of cells, assuming a finite neighborhood of influence for each QCA cell, which is usually the case. The overall low-energy spectrum of multiple clocking zones is constructed by concatenating the low-energy spectra of the individual clocking zones. We demonstrate how the model can be used to study the tradeoff between switching errors and clocking zones. Sanjukta Bhanja, Sudeep Sarkar |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2007 | Enhanced Level Building Algorithm for the Movement Epenthesis Problem in Sign Language RecognitionabstractOne of the hard problems in automated sign language recognition is the movement epenthesis (me) problem. Movement epenthesis is the gesture movement that bridges two consecutive signs. This effect can be over a long duration and involve variations in hand shape, position, and movement, making it hard to explicitly model these intervening segments. This creates a problem when trying to match individual signs to full sign sentences since for many chunks of the sentence, corresponding to these mes, we do not have models. We present an approach based on version of a dynamic programming framework, called Level Building, to simultaneously segment and match signs to continuous sign language sentences in the presence of movement epenthesis (me). We enhance the classical Level Building framework so that it can accommodate me labels for which we do not have explicit models. This enhanced Level Building algorithm is then coupled with a trigram grammar model to optimally segment and label sign language sentences. We demonstrate the efficiency of the algorithm using a single view video dataset of continuous sign language sentences. We obtain 83% word level recognition rate with the enhanced Level Building approach, as opposed to a 20% recognition rate using a classical Level Building framework on the same dataset. The proposed approach is novel since it does not need explicit models for movement epenthesis. Ruiduo Yang, Sudeep Sarkar, Barbara L. Loeding |
CVPR | 2 |
| 2007 | Facial Strain Pattern as a Soft Forensic EvidenceabstractThe success of forensic identification largely depends on the availability of strong evidence or traces that substantiate the prosecution hypothesis that a certain person is guilty of crime. In light of this, extracting subtle evidences which the criminals leave behind at the crime scene will be of valuable help to investigators. We propose a novel method of using strain pattern extracted from changing facial expressions in video as an auxiliary evidence for person identification. The strength of strain evidence is analyzed based on the increase in likelihood ratio it provides in a suspect population. Results show that strain pattern can be used as a supplementary biometric evidence in adverse operational conditions such as shadow lighting and face camouflage where pure intensity-based face recognition algorithms will fail Vasant Manohar, Dmitry B. Goldgof, Sudeep Sarkar, Yong Zhang 0017 |
WACV | 3 |
| 2007 | 3D Finite Element Modeling of Nonrigid Breast Deformation for Feature Registration in -ray and MR ImagesabstractRegistering features in multiple mammographic views is an important technique to improve breast cancer detection rate. However, nonrigid breast deformation during X-ray imaging poses a severe challenge to the conventional 2D registration methods. We present a method that utilizes a 3D model to facilitate two-view registration by predicting breast deformation. At first, a finite element model of a breast is constructed using its MRIs. The model is capable of simulating both compression and decompression. Feature registration is then accomplished through a series of projections and compression-decompression operations. Experiments using real patient data demonstrate that a mammographic feature can be successfully registered from one view to another. Yong Zhang 0017, Dmitry B. Goldgof, Sudeep Sarkar, Lihua Li 0002 |
WACV | 4 |
| 2007 | Outdoor recognition at a distance by fusing gait and face
Sudeep Sarkar |
Image Vis. Comput. | 2 |
| 2007 | A sensitivity analysis method and its application in physics-based nonrigid motion modeling
Yong Zhang 0017, Dmitry B. Goldgof, Sudeep Sarkar, Leonid V. Tsap |
Image Vis. Comput. | 3 |
| 2007 | From Scores to Face Templates: A Model-Based ApproachabstractRegeneration of templates from match scores has security and privacy implications related to any biometric authentication system. We propose a novel paradigm to reconstruct face templates from match scores using a linear approach. It proceeds by first modeling the behavior of the given face recognition algorithm by an affine transformation. The goal of the modeling is to approximate the distances computed by a face recognition algorithm between two faces by distances between points, representing these faces, in an affine space. Given this space, templates from an independent image set (break-in) are matched only once with the enrolled template of the targeted subject and match scores are recorded. These scores are then used to embed the targeted subject in the approximating affine (non-orthogonal) space. Given the coordinates of the targeted subject in the affine space, the original template of the targeted subject is reconstructed using the inverse of the affine transformation. We demonstrate our ideas using three, fundamentally different, face recognition algorithms: Principal Component Analysis (PCA) with Mahalanobis cosine distance measure, Bayesian intra-extrapersonal classifier (BIC), and a feature-based commercial algorithm. To demonstrate the independence of the break-in set with the gallery set, we select face templates from two different databases: Face Recognition Grand Challenge (FRGC) and Facial Recognition Technology (FERET) Database (FERET). With an operational point set at 1 percent False Acceptance Rate (FAR) and 99 percent True Acceptance Rate (TAR) for 1,196 enrollments (FERET gallery), we show that at most 600 attempts (score computations) are required to achieve a 73 percent chance of breaking in as a randomly chosen target subject for the commercial face recognition system. With similar operational set up, we achieve a 72 percent and 100 percent chance of breaking in for the Bayesian and PCA based face recognition systems, respectively. With three different levels of score quantization, we achieve 69 percent, 68 percent and 49 percent probability of break-in, indicating the robustness of our proposed scheme to score quantization. We also show that the proposed reconstruction scheme has 47 percent more probability of breaking in as a randomly chosen target subject for the commercial system as compared to a hill climbing approach with the same number of attempts. Given that the proposed template reconstruction method uses distinct face templates to reconstruct faces, this work exposes a more severe form of vulnerability than a hill climbing kind of attack where incrementally different versions of the same face are used. Also, the ability of the proposed approach to reconstruct actual face templates of the users increases privacy concerns in biometric systems. Pranab K. Mohanty, Sudeep Sarkar, Rangachar Kasturi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Evaluation and analysis of a face and voice outdoor multi-biometric system
Himanshu Vajaria, Tanmoy Islam, Pranab K. Mohanty, Sudeep Sarkar, Ravi Sankar, Rangachar Kasturi |
Pattern Recognit. Lett. | 4 |
| 2006 | Gesture Recognition using Hidden Markov Models from Fragmented ObservationsabstractWe consider the problem of computing the likelihood of a gesture from regular, unaided video sequences, without relying on perfect segmentation of the scene. Instead of requiring that low-and mid-level processes produce near-perfect segmentation of relevant body parts such as hands, we take into account that such processes can only produce uncertain information. The hands can only be detected as fragmented regions along with clutter. To address this problem, we propose an extension of the HMM formalism, which we call the frag-HMM, to allow for reasoning based on fragmented observations, via the use of an intermediate grouping process. In this formulation, we do not match the frag- HMMto one observation sequence, but rather to a sequence of observation sets, where each observation set is a collection of groups of fragmented observations. Based on the developed model, we show how to perform three kinds of computations. The first one is to decide on the best observation group for each frame, given a sequence of observation groups for the past frames. This allows us to incrementally compute the best segmentation of the hand for each frame, given the model. The second one involves the computation of likelihood of a sequence, averaged over all possible states sequences and possible groupings. The third is the computation of the likelihood of a sequence, maximized over all possible state sequences and group sequences. This can give us the best possible groupings for each frame, as well. We demonstrate our ideas using a publicly available hand gesture dataset that spans different subjects, is against complex background, and involves hand occlusions. The recognition performance is within 2% of that obtained with manually segmented hands and about 10% better than that obtained with segmentations that use the prior knowledge of the hand color. Ruiduo Yang, Sudeep Sarkar |
CVPR (1) | 2 |
| 2006 | Efficient Generation of Large Amounts of Training Data for Sign Language Recognition: A Semi-automatic Tool
Ruiduo Yang, Sudeep Sarkar, Barbara L. Loeding, Arthur I. Karshmer |
ICCHP | 2 |
| 2006 | Improved Gait Recognition by Gait Dynamics NormalizationabstractPotential sources for gait biometrics can be seen to derive from two aspects: gait shape and gait dynamics. We show that improved gait recognition can be achieved after normalization of dynamics and focusing on the shape information. We normalize for gait dynamics using a generic walking model, as captured by a population Hidden Markov Model (pHMM) defined for a set of individuals. The states of this pHMM represent gait stances over one gait cycle and the observations are the silhouettes of the corresponding gait stances. For each sequence, we first use Viterbi decoding of the gait dynamics to arrive at one dynamics-normalized, averaged, gait cycle of fixed length. The distance between two sequences is the distance between the two corresponding dynamics-normalized gait cycles, which we quantify by the sum of the distances between the corresponding gait stances. Distances between two silhouettes from the same generic gait stance are computed in the linear discriminant analysis space so as to maximize the discrimination between persons, while minimizing the variations of the same subject under different conditions. The distance computation is constructed so that it is invariant to dilations and erosions of the silhouettes. This helps us handle variations in silhouette shape that can occur with changing imaging conditions. We present results on three different, publicly available, data sets. First, we consider the HumanlD Gait Challenge data set, which is the largest gait benchmarking data set that is available (122 subjects), exercising five different factors, i.e., viewpoint, shoe, surface, carrying condition, and time. We significantly improve the performance across the hard experiments involving surface change and briefcase carrying conditions. Second, we also show improved performance on the UMD gait data set that exercises time variations for 55 subjects. Third, on the CMU Mobo data set, we show results for matching across different walking speeds. It is worth noting that there was no separate training for the UMD and CMU data sets. Sudeep Sarkar |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | A constrained genetic approach for computing material property of elastic objectsabstractThis paper presents a constrained genetic approach for reconstructing the material properties of elastic objects. The considered reconstruction problem is ill-posed and must be constrained properly so that a unique and stable numerical solution can be obtained. Qualitative prior information is incorporated using a rank-based scheme to constrain the admissible solutions. Experiments show that the proposed approach is robust when presented with noisy data and can reconstruct the elastic property accurately and reliably. In a comparison study with the deterministic Gauss-Newton methods, the constrained genetic approach also shows very consistent performance. Yong Zhang 0017, Lawrence O. Hall, Dmitry B. Goldgof, Sudeep Sarkar |
IEEE Trans. Evol. Comput. | 4 |
| 2005 | The HumanID Gait Challenge Problem: Data Sets, Performance, and AnalysisabstractIdentification of people by analysis of gait patterns extracted from video has recently become a popular research problem. However, the conditions under which the problem is "solvable" are not understood or characterized. To provide a means for measuring progress and characterizing the properties of gait recognition, we introduce the HumanID Gait Challenge Problem. The challenge problem consists of a baseline algorithm, a set of 12 experiments, and a large data set. The baseline algorithm estimates silhouettes by background subtraction and performs recognition by temporal correlation of silhouettes. The 12 experiments are of increasing difficulty, as measured by the baseline algorithm, and examine the effects of five covariates on performance. The covariates are: change in viewing angle, change in shoe type, change in walking surface, carrying or not carrying a briefcase, and elapsed time between sequences being compared. Identification rates for the 12 experiments range from 78 percent on the easiest experiment to 3 percent on the hardest. All five covariates had statistically significant effects on performance, with walking surface and time difference having the greatest impact. The data set consists of 1,870 sequences from 122 subjects spanning five covariates (1.2 Gigabytes of data). The gait data, the source code of the baseline algorithm, and scripts to run, score, and analyze the challenge experiments are available at http://www.GaitChallenge.org. This infrastructure supports further development of gait recognition algorithms and additional experiments to understand the strengths and weaknesses of new algorithms. The more detailed the experimental results presented, the more detailed is the possible meta-analysis and greater is the understanding. It is this potential from the adoption of this challenge problem that represents a radical departure from traditional computer vision research methodology. Sudeep Sarkar, P. Jonathon Phillips, Isidro Robledo Vega, Patrick Grother, Kevin W. Bowyer |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Effect of silhouette quality on hard problems in gait recognitionabstractGait as a behavioral biometric has been the subject of recent investigations. However, understanding the limits of gait-based recognition and the quantitative study of the factors effecting gait have been confounded by errors in the extracted silhouettes, upon which most recognition algorithms are based. To enable us to study this effect on a large population of subjects, we present a novel model based silhouette reconstruction strategy, based on a population based hidden Markov model (HMM), coupled with an eigen-stance model, to correct for common errors in silhouette detection arising from shadows and background subtraction. The model is trained and benchmarked using manually specified silhouettes for 71 subjects from the recently formulated HumanID Gait Challenge database. Unlike other essentially pixel-level silhouette cleaning methods, this method can remove shadows, especially between feet for the legs-apart stance, and remove parts due to any objects being carried, such as briefcase or a walking cane. After quantitatively establishing the improved quality of the silhouette over simple background subtraction, we show on the 122 subjects HumanID Gait Challenge Dataset and using two gait recognition algorithms that the observed poor performance of gait recognition for hard problems involving matching across factors such as surface, time, and shoe are not due to poor silhouette quality, beyond what is available from statistical background subtraction based methods. Sudeep Sarkar |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2004 | Studies on Silhouette Quality and Gait Recognition
Laura Malave, Sudeep Sarkar |
CVPR (2) | 3 |
| 2004 | Progress in Automated Computer Recognition of Sign Language
Barbara L. Loeding, Sudeep Sarkar, Ayush Parashar, Arthur I. Karshmer |
ICCHP | 2 |
| 2004 | A modeling approach for burn scar assessment using natural features and elastic propertyabstractA modeling approach is presented for quantitative burn scar assessment. Emphases are given to: 1) constructing a finite-element model from natural image features with an adaptive mesh and 2) quantifying the Young's modulus of scars using the finite-element model and regularization method. A set of natural point features is extracted from the images of burn patients. A Delaunay triangle mesh is then generated that adapts to the point features. A three-dimensional finite-element model is built on top of the mesh with the aid of range images providing the depth information. The Young's modulus of scars is quantified with a simplified regularization functional, assuming that the knowledge of the scar's geometry is available. The consistency between the relative elasticity index and the physician's rating based on the Vancouver scale (a relative scale used to rate burn scars) indicates that the proposed modeling approach has high potential for image-based quantitative burn scar assessment. Yong Zhang 0017, Dmitry B. Goldgof, Sudeep Sarkar, Leonid V. Tsap |
IEEE Trans. Medical Imaging | 3 |
| 2004 | A methodology for extracting objective color from imagesabstractWe present a methodology for correcting color images taken in practical indoor environments, such as laboratories, factories, and studios, that explicitly models illuminant location, surface reflectance and geometry, and camera responsivity. We explicitly model surfaces by taking our color images with corresponding registered three-dimensional (3-D) range images, which provide surface orientation and location information for every point in the scene. We automatically detect regions where color correction should not be applied, such as specularities, coarse texture regions, and jump edges. This correction results in objective color measures of the imaged surfaces. This kind of integrated, comprehensive system of color correction has not existed until now. i.e., it is the first of its kind in computer vision. We demonstrate results of applying this methodology to real images for applications in photorealistic rerendering, skin lesion detection, burn scar color measurement, and general color image enhancement. We also have tested the method under different lighting configurations and with three different range scanners. Mark W. Powell, Sudeep Sarkar, Dmitry B. Goldgof, Krassimir Ivanov |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2003 | Comparison and Combination of Ear and Face Images in Appearance-Based BiometricsabstractResearchers have suggested that the ear may have advantages over the face for biometric recognition. Our previous experiments with ear and face recognition, using the standard principal component analysis approach, showed lower recognition performance using ear images. We report results of similar experiments on larger data sets that are more rigorously controlled for relative quality of face and ear images. We find that recognition performance is not significantly different between the face and the ear, for example, 70.5 percent versus 71.6 percent, respectively, in one experiment. We also find that multimodal recognition using both the ear and face results in statistically significant improvement over either individual biometric, for example, 90.9 percent in the analogous experiment. Kyong I. Chang, Kevin W. Bowyer, Sudeep Sarkar, Barnabas Victor |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2003 | An In-Depth Study of Graph Partitioning Measures for Perceptual OrganizationabstractIn recent years, one of the effective engines for perceptual organization of low-level image features is based on the partitioning of a graph representation that captures Gestalt inspired local structures, such as similarity, proximity, continuity, parallelism, and perpendicularity, over the low-level image features. Mainly motivated by computational efficiency considerations, this graph partitioning process is usually implemented as a recursive bipartitioning process, where, at each step, the graph is broken into two parts based on a partitioning measure. We focus on three such measures, namely, the minimum, average, and normalized cuts. The minimum cut partition seeks to minimize the total link weights cut. The average cut measure is proportional to the total link weight cut, normalized by the sizes of the partitions. The normalized cut measure is normalized by the product of the total connectivity (valencies) of the nodes in each partition. We provide theoretical and empirical insight into the nature of the three partitioning measures in terms of the underlying image statistics. In particular, we consider for what kinds of image statistics would optimizing a measure, irrespective of the particular algorithm used, result in correct partitioning. Are the quality of the groups significantly different for each cut measure? Are there classes of images for which grouping by partitioning does not work well? Also, can the recursive bipartitioning strategy separate out groups corresponding to K objects from each other? In the analysis, we draw from probability theory and the rich body of work on stochastic ordering of random variables. Our major conclusion is that optimization of none of the three measures is guaranteed to result in the correct partitioning of K objects, in the strict stochastic order sense, for all image statistics. Padmanabhan Soundararajan, Sudeep Sarkar |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | Statistical Motion Model Based on the Change of Feature Relationships: Human Gait-Based RecognitionabstractWe offer a novel representation scheme for view-based motion analysis using just the change in the relational statistics among the detected image features, without the need for object models, perfect segmentation, or part-level tracking. We model the relational statistics using the probability that a random group of features in an image would exhibit a particular relation. To reduce the representational combinatorics of these relational distributions, we represent them in a Space of Probability Functions (SoPF), where the Euclidean distance is related to the Bhattacharya distance between probability functions. Different motion types sweep out different traces in this space. We demonstrate and evaluate the effectiveness of this representation in the context of recognizing persons from gait. In particular, on outdoor sequences: (1) we demonstrate the possibility of recognizing persons from not only walking gait, but running and jogging gaits as well; (2) we study recognition robustness with respect to view-point variation; and (3) we benchmark the recognition performance on a database of 71 subjects walking on soft grass surface, where we achieve around 90 percent recognition rates in the presence of viewpoint variation. Isidro Robledo Vega, Sudeep Sarkar |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | A constrained genetic approach for reconstructing Young's modulus of elastic objects from boundary displacement measurementsabstractThis paper presents a constrained genetic approach (CGA) for reconstructing the Young's modulus of elastic objects. Qualitative a priori information is incorporated using a rank based scheme to constrain the admissible solutions. Balance between the fitness function (adhesion to the measurement data) and the penalty function (fidelity to a priori knowledge) is achieved by a stochastic sort algorithm. The over-smoothing of Young's modulus discontinuity is avoided without the need of computing a deterministic weight coefficient. The experiment on synthetic data indicates that the proposed method not only reconstructed reliable Young's modulus from noisy data, but also expedited the convergence process significantly. Yong Zhang 0017, Lawrence O. Hall, Dmitry B. Goldgof, Sudeep Sarkar |
IEEE Congress on Evolutionary Computation | 4 |
| 2002 | Perceptual Organization Based Computational Model for Robust Segmentation of Moving Objects
Sudeep Sarkar, Daniel Majchrzak, Kishore Korimilli |
Comput. Vis. Image Underst. | 1 |
| 2001 | Discrimination of Motion Based on Traces in the Space of Probability Functions over Feature RelationsabstractIn this paper we demonstrate that it is possible to discriminate between high level motion types such as walking, jogging, or running based on just the change in the relational statistics among the detected image features, without the need for object models, perfect segmentation, or tracking. Instead of the statistics of the feature attributes themselves, we consider the distribution of the statistics of the relations among the features. We represent the observed distribution of feature relations in an image as a point in a space where the Euclidean distance is related to the Bhattacharya distance between probability functions. Different motion types sweep out different traces in this Space of Probability Functions (SoPF). We demonstrate the effectiveness of this representation on image sequences of human in motion, gathered using a digital video camera. We show that it is not only possible to distinguish between motion types but also to discriminate between persons based on the SoPF traces. Sudeep Sarkar, Isidro Robledo Vega |
CVPR (1) | 1 |
| 2001 | Investigation of Measures for Grouping by Graph PartitioningabstractGrouping by graph partitioning is an effective engine for perceptual organization. This graph partitioning process, mainly motivated by computational efficiency considerations, is usually implemented as recursive bi-partitioning, where at each step the graph is broken into two parts based on a partitioning measure. We study four such measures, namely, the minimum cut, average cut, Shi-Malik normalized cut, and a variation of the Shi-Malik normalized cut. Using probabilistic analysis we show that the minimization of the average cut and the normalized cut measure, using recursive bi-partitioning will, on an average, result in the correct segmentation. The minimum cut and the variation of the normalized cut will, on an average, not result in the correct segmentation and we can precisely express the conditions. Based on a rigorous empirical evaluation, we also show that, in practice, the quality of the groups generated using minimum, average or normalized cuts are statistically equivalent for object recognition, i.e. the best, the mean, and the variation of the qualities are statistically equivalent. We also find that for certain image classes, such as aerial and scenes with man-made objects in man-made surroundings, the performance of grouping by partitioning is the worst, irrespective of the cut measure. Padmanabhan Soundararajan, Sudeep Sarkar |
CVPR (1) | 2 |
| 2001 | A Simple Strategy for Calibrating the Geometry of Light SourcesabstractWe present a methodology for calibrating multiple light source locations in 3D from images. The procedure involves the use of a novel calibration object that consists of three spheres at known relative positions. The process uses intensity images to find the positions of the light sources. We conducted experiments to locate light sources in 51 different positions in a laboratory setting. Our data shows that the vector from a point in the scene to a light source can be measured to within 2.7/spl plusmn/4/spl deg/ at /spl alpha/=.05 (6 percent relative) of its true direction and within 0.13/spl plusmn/.02 m at /spl alpha/=.05 (9 percent relative) of its true magnitude compared to empirically measured ground truth. Finally, we demonstrate how light source information is used for color correction. Mark W. Powell, Sudeep Sarkar, Dmitry B. Goldgof |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | Fusion of physically-based registration and deformation modeling for nonrigid motion analysisabstractIn our previous work, we used finite element models to determine nonrigid motion parameters and recover unknown local properties of objects given correspondence data recovered with snakes or other tracking models. In this paper, we present a novel multiscale approach to recovery of nonrigid motion from sequences of registered intensity and range images. The main idea of our approach is that a finite element (FEM) model incorporating material properties of the object can naturally handle both registration and deformation modeling using a single model-driving strategy. The method includes a multiscale iterative algorithm based on analysis of the undirected Hausdorff distance to recover correspondences. The method is evaluated with respect to speed and accuracy. Noise sensitivity issues are addressed. Advantages of the proposed approach are demonstrated using man-made elastic materials and human skin motion. Experiments with regular grid features are used for performance comparison with a conventional approach (separate snakes and FEM models). It is shown, however, that the new method does not require a sampling/correspondence template and can adapt the model to available object features. Usefulness of the method is presented not only in the context of tracking and motion analysis, but also for a burn scar detection application. Leonid V. Tsap, Dmitry B. Goldgof, Sudeep Sarkar |
IEEE Trans. Image Process. | 3 |
| 2000 | Calibration of Light SourcesabstractWe present a methodology for calibrating multiple light source locations in 3D from images. The procedure involves the use of a novel calibration object that consists of either 2 or 3 spheres at known relative positions. There are two variants of the process: one which uses range and intensity imaging to find the positions of the light sources, and one that uses only the intensity image to locate the illuminants. We conducted experiments using both variations of the technique to locate light sources in 51 different positions in a laboratory setting. Our data shows that the vector from a point in the scene to a light source can be measured to within 3/spl deg/(6%) of its tote direction and within 0.13 m (9%) of its true magnitude compared to empirically measured ground truth. Finally, we demonstrate how light source information can be applied to burn scar color correction and color segmentation. Mark W. Powell, Sudeep Sarkar, Dmitry B. Goldgof |
CVPR | 2 |
| 2000 | Multiscale Combination of Physically-Based Registration and Deformation ModelingabstractIn this paper we present a novel multiscale approach to recovery of nonrigid motion from sequences of registered intensity and range images. The main idea of our approach is that a finite element (FEM) model can naturally handle both registration and deformation modeling using a single model-driving strategy. The method includes a multiscale iterative algorithm based on analysis of the undirected Hausdorff distance to recover correspondences. The method is evaluated with respect to speed, accuracy, and noise sensitivity. Advantages of the proposed approach are demonstrated using man-made elastic materials and human skin motion. Experiments with regular grid features are used for performance comparison with a conventional approach (separate snakes and FEM models). It is shown that the new method does not require a grid and can adapt the model to available object features. Leonid V. Tsap, Dmitry B. Goldgof, Sudeep Sarkar |
CVPR | 3 |
| 2000 | Motion Segmentation Based on Perceptual Organization of Spatio-Temporal VolumesabstractThe role of perceptual organization in motion analysis has heretofore been minimal. In this work we demonstrate that the use of perceptual organization principles of temporal coherence (common fate) and spatial proximity can result in a robust motion segmentation algorithm that is able to handle drastic illumination changes, occlusion events, and multiple moving objects, without the use of object models. The adopted algorithm does not employ the traditional frame by frame motion analysis, but rather treats the image sequence as a single 3D spatio-temporal block of data. We describe motion using spatio-temporal surfaces, which we, in turn, describe as compositions of finite planar patches. These planar patches, referred to as temporal envelopes, capture the local nature of the motions. We detect these temporal envelopes using 3D-edge detection followed by Hough transform, and represent them with convex hulls. We present a graph-based method to group these temporal envelopes arising from one object based on Gestalt organizational principles. A probabilistic Bayesian network quantifies the saliencies of the relationships between temporal envelopes. We present results on sequences with multiple moving persons, significant occlusions, and scene illumination changes. Kishore Korimilli, Sudeep Sarkar |
ICPR | 2 |
| 2000 | Motion Detection from Temporally Integrated ImagesabstractMotion blur arises when motion is fast relative to the shutter time of a camera. Unlike most work on motion blur, which considers the streaks due to motion blur to be noisy artifacts. In this paper we introduce a new method to extract motion information from these streaks. Previous methods with similar goals first extract an optic flow field from local information in the motion streaks and then infer global motion parameters. On the contrary, we adopt a more direct feature-based approach and extract global motion parameters from the motion streaks. We first extract edges in the motion blurred images, which we then group to determine the foci of expansion, the center of rotation, or motion parallel to the image plane. Furthermore, we determine the direction of motion. We present results on real images from a mobile robot in cluttered environments. Daniel Majchrzak, Sudeep Sarkar, Barry Sheppard, Robin R. Murphy |
ICPR | 2 |
| 2000 | Model-Based Nonrigid Motion Analysis Using Natural Feature Adaptive MeshabstractThe success of nonrigid motion analysis using physical finite element model is dependent on the mesh that characterizes the object's geometric structure. We suggest a deformable mesh adapted to the natural features of images. The adaptive mesh requires much fewer number of nodes than the fixed mesh which was used in the work by Tsap et al. (1998). We demonstrate the higher efficiency of the adaptive mesh in the context of estimating burn scare elasticity relative to normal skin elasticity using the observed 2D image sequence. Our results show that the scar assessment method based on the physical model using natural feature adaptive mesh can be applied to images which do not have artificial markers. Yong Zhang 0017, Dmitry B. Goldgof, Sudeep Sarkar, Leonid V. Tsap |
ICPR | 3 |
| 2000 | Modeling Parameter Space Behavior of Vision Systems Using Bayesian Networks
Sudeep Sarkar, Srikanth Chavali |
Comput. Vis. Image Underst. | 1 |
| 2000 | A Method for Increasing Precision and Reliability of Elasticity Analysis in Complicated Burn Scar CasesabstractIn this paper we propose a method for increasing precision and reliability of elasticity analysis in complicated burn scar cases. The need for a technique that would help physicians by objectively assessing elastic properties of scars, motivated our original algorithm. This algorithm successfully employed active contours for tracking and finite element models for strain analysis. However, the previous approach considered only one normal area and one abnormal area within the region of interest, and scar shapes which were somewhat simplified. Most burn scars have rather complicated shapes and may include multiple regions with different elastic properties. Hence, we need a method capable of adequately addressing these characteristics. The new method can split the region into more than two localities with different material properties, select and quantify abnormal areas, and apply different forces if it is necessary for a better shape description of the scar. The method also demonstrates the application of scale and mesh refinement techniques in this important domain. It is accomplished by increasing the number of Finite Element Method (FEM) areas as well as the number of elements within the area. The method is successfully applied to elastic materials and real burn scar cases. We demonstrate all of the proposed techniques and investigate the behavior of elasticity function in a 3-D space. Recovered properties of elastic materials are compared with those obtained by a conventional mechanics-based approach. Scar ratings achieved with the method are correlated against the judgments of physicians. Leonid V. Tsap, Dmitry B. Goldgof, Sudeep Sarkar, Pauline S. Powers |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2000 | Supervised Learning of Large Perceptual Organization: Graph Spectral Partitioning and Learning AutomataabstractPerceptual organization offers an elegant framework to group low-level features that are likely to come from a single object. We offer a novel strategy to adapt this grouping process to objects in a domain. Given a set of training images of objects in context, the associated learning process decides on the relative importance of the basic salient relationships such as proximity, parallelness, continuity, junctions, and common region toward segregating the objects from the background. The parameters of the grouping process are cast as probabilistic specifications of Bayesian networks that need to be learned. This learning is accomplished using a team of stochastic automata in an N-player cooperative game framework. The grouping process, which is based on graph partitioning is able to form large groups from relationships defined over a small set of primitives and is fast. We statistically demonstrate the robust performance of the grouping and the learning frameworks on a variety of real images. Among the interesting conclusions is the significant role of photometric attributes in grouping and the ability to form large salient groups from a set of local relations, each defined over a small number of primitives. Sudeep Sarkar, Padmanabhan Soundararajan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | Nonrigid Motion Analysis Based on Dynamic Refinement of Finite Element ModelsabstractWe propose new algorithms for accurate nonrigid motion tracking. Given an initial model representing general knowledge of the object, a set of sparse correspondences, and incomplete or missing information about geometry or material properties, we can recover dense motion vectors using finite element models. The method is based on the iterative analysis of the differences between the actual and predicted behaviors. Unknown parameters are recovered using an iterative descent search for the best nonlinear finite element model that approximates nonrigid motion of the given object. During this search process, we not only estimate material properties, but also infer dense point correspondences from our initial set of sparse correspondences. Thus, during tracking, the model is refined which, in turn, improves tracking quality. Experimental results demonstrate the success of the proposed algorithm. Our work demonstrates the possibility of accurate quantitative analysis of nonrigid motion in range image sequences with objects consisting of multiple materials and 3D volumes. Leonid V. Tsap, Dmitry B. Goldgof, Sudeep Sarkar |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1999 | Guest Editors' Introduction: Perceptual Organization in Computer Vision: Status, Challenges, and Potential
Kim L. Boyer, Sudeep Sarkar |
Comput. Vis. Image Underst. | 2 |
| 1999 | Model-based force-driven nonrigid motion recovery from sequences of range images without point correspondences
Leonid V. Tsap, Dmitry B. Goldgof, Sudeep Sarkar |
Image Vis. Comput. | 3 |
| 1998 | Learning to Form Large Groups of Salient Image FeaturesabstractWe offer a novel strategy to adapt the perceptual organization process to an object and its contest in a scene. Given a set of training images of an object in context, a learning process decides on the relative importance of the basic Gestalt relationships such as proximity, parallelness, similarity, symmetry, closure, and common region towards segregating the object from the background. This learning is accomplished using a team of stochastic automata in a N-player cooperative game framework. The grouping process which is based on graph partitioning is able to form large groups from relationships defined over a small set of primitives and is fast. We demonstrate the robust performance of the growing system on a variety of real images. Among the interesting conclusions is the significant role of photometric attributes in grouping and the ability to perform figure-ground segmentation from a set of local relations, each defined over a small number of primitives. Sudeep Sarkar |
CVPR | 1 |
| 1998 | Nonrigid Motion Analysis Based on Dynamic Refinement of Finite Element ModelsabstractIn this paper we propose new algorithms for accurate nonrigid motion tracking. Given only a set of sparse correspondences and incomplete or missing information about geometry or material properties, we recover dense motion vectors using nonlinear finite element models. The method is based on the iterative analysis of the differences between the actual and predicted behavior. Large differences indicate that an object's properties are not captured properly by the model. Feedback from the images during the motion allows the refinement of the model by minimizing the error between the expected and true position of the object's points. Unknown parameters are recovered using an iterative descent search for the best model that approximates nonrigid motion of the given object. Thus, during tracking the model is refined which, in turn, improves tracking quality. The method was applied successfully to man-made elastic materials and human skin to recover unknown elasticity, to complex 3-D objects to find details of their geometry, and to a hand motion analysis application. Leonid V. Tsap, Dmitry B. Goldgof, Sudeep Sarkar |
CVPR | 3 |
| 1998 | An Approximate Algorithm for Structural Matching of ImagesabstractWe provide a novel strategy to structurally match two images. The matching process takes into account minor structural variations between the two images. Given two images, we first construct random parametric structural descriptions (RPSDs) to capture their acceptable variations. A RPSD induces a joint probability distribution over a hypergraph that captures the image structure using image primitives and the relations among them. The RSPDs from the two images are matched using an information theoretic entropy measure which quantify the similarity of the joint structural probability distributions modeled by the RPSDs. The computation of this entropic similarity measure requires the matching of the hypergraphs underlying the RPSDs which, as we know, is NP-hard. So, we propose an approximate strategy which finds the solution by generating a sequence of possible mapping functions between the primitives from the two images. We present experimental validation of the strategy on a variety of images. Sriprakash Sainath, Sudeep Sarkar |
ICIP (1) | 2 |
| 1998 | Model-based Nonrigid Motion Recovery from Sequences of Range Images without Point CorrespondencesabstractWe propose a new method for accurate nonrigid motion analysis when point correspondence data is not available. We construct nonlinear finite element models by integrating range data and prior knowledge about an object's properties. We attempt to recover the motion sequence given an initial alignment of the model with the first frame of the sequence. The main idea of the method is to find the forces that are responsible for the motion or shape deformation of the given object. The task is broken into subtasks of finding the forces for each frame. Both absolute values and directions of these forces are taken into consideration and iteratively varied not only for each frame, but also between the frames. Our work demonstrates the possibility of accurate nonrigid motion analysis and force recovery from range image sequences containing nonrigid objects and large motion without interframe point correspondences. Leonid V. Tsap, Dmitry B. Goldgof, Sudeep Sarkar |
ICIP (2) | 3 |
| 1998 | The Effect of Edge Strength on Object Recognition from Edge Images
Kevin W. Bowyer, Thomas A. Sanocki, Sudeep Sarkar |
ICIP (3) | 4 |
| 1998 | Comparison of Edge Detectors: A Methodology and Initial Study
Michael D. Heath, Sudeep Sarkar, Thomas A. Sanocki, Kevin W. Bowyer |
Comput. Vis. Image Underst. | 2 |
| 1998 | Quantitative Measures of Change Based on Feature Organization: Eigenvalues and Eigenvectors
Sudeep Sarkar, Kim L. Boyer |
Comput. Vis. Image Underst. | 1 |
| 1998 | Efficient Nonlinear Finite Element Modeling of Nonrigid Objects via Optimization of Mesh Models
Leonid V. Tsap, Dmitry B. Goldgof, Sudeep Sarkar, Wen-Chen Huang |
Comput. Vis. Image Underst. | 3 |
| 1998 | Integrating Image Computation in Undergraduate Level Data-Structure EducationabstractThere is a growing need for expertise both in image analysis and in software engineering. To date, these two areas have been taught separately in an undergraduate computer and information science curriculum. However, we have found that introduction to image analysis can be easily integrated in data-structure courses without detracting from the original goal of teaching data structures. Some of the image processing tasks offer a natural way to introduce basic data structures such as arrays, queues, stacks, trees and hash tables. Not only does this integrated strategy expose the students to image related manipulations at an early stage of the curriculum but it also imparts cohesiveness to the data-structure assignments and brings them closer to real life. In this paper we present a set of programming assignments that integrates undergraduate data-structure education with image processing tasks. These assignments can be incorporated in existing data-structure courses with low time and software overheads. We have used these assignment sets thrice: once in a 10-week duration data-structure course at the University of California, Santa Barbara and the other two times in 15-week duration courses at the University of South Florida, Tampa. Sudeep Sarkar, Dmitry B. Goldgof |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1998 | A Vision-Based Technique for Objective Assessment of Burn ScarsabstractIn this paper a method for the objective assessment of burn scars is proposed. The quantitative measures developed in this research provide an objective way to calculate elastic properties of burn scars relative to the surrounding areas. The approach combines range data and the mechanics and motion dynamics of human tissues. Active contours are employed to locate regions of interest and to find displacements of feature points using automatically established correspondences. Changes in strain distribution over time are evaluated. Given images at two time instances and their corresponding features, the finite element method is used to synthesize strain distributions of the underlying tissues. This results in a physically based framework for motion and strain analysis. Relative elasticity of the burn scar is then recovered using iterative descent search for the best nonlinear finite element model that approximates stretching behavior of the region containing the burn scar. The results from the skin elasticity experiments illustrate the ability to objectively detect differences in elasticity between normal and abnormal tissue. These estimated differences in elasticity are correlated against the subjective judgments of physicians that are presently the practice. Leonid V. Tsap, Dmitry B. Goldgof, Sudeep Sarkar, Pauline S. Powers |
IEEE Trans. Medical Imaging | 3 |
| 1997 | Experimental Performance Evaluation of Feature Grouping ModulesabstractWe present five performance measures to evaluate grouping modules in the context of constrained search and indexing based object recognition. Using these measures, we demonstrate a sound experimental framework based on statistical ANOVA tests to compare and contrast three edge based organization modules, namely those of A. Etemadi et al. (1991), D.W. Jacobs (1996), and S. Sarkar and K.L. Boyer (1993) in the domain of aerial objects using 50 images. With adapted parameters, the Jacobs module is overall the best choice for constraint based recognition. For fixed parameters, the Sarkar-Boyer module is the best in terms of recognition accuracy and indexing speedup. Etemadi et al.'s module performs equally well with fixed and adapted parameters while the Jacobs module is most sensitive to fixed and adapted parameter choices. The overall performance ranking of the modules is Jacobs, Sakar-Boyer, and Etemadi et al. Sudhir Borra, Sudeep Sarkar |
CVPR | 2 |
| 1997 | A Framework for Performance Characterization of Intermediate-Level Grouping ModulesabstractWe present five performance measures to evaluate grouping modules in the context of constrained search and indexing based object recognition. Using these measures, we demonstrate a sound experimental framework, based on statistical ANOVA tests, to compare and contrast three edge based organization modules, namely, those of Etemadi et al. (1991), Jacobs (1996), and Sarkar-Boyer (1993) in the domain of aerial objects using 50 images. With adapted parameters, the Jacobs module performs overall the best for constraint based recognition. For fixed parameters, the Sarkar-Boyer module is the best in terms of recognition accuracy and indexing speedup. Etemadi et al.'s module performs equally well with fixed and adapted parameters while the Jacobs module is most sensitive to fixed and adapted parameter choices. The overall performance ranking of the modules is Jacobs, Sarkar-Boyer, and Etemadi et al. Sudhir Borra, Sudeep Sarkar |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1997 | Robust Visual Method for Assessing the Relative Performance of Edge-Detection AlgorithmsabstractA new method for evaluating edge detection algorithms is presented and applied to measure the relative performance of algorithms by Canny, Nalwa-Binford, Iverson-Zucker, Bergholm, and Rothwell. The basic measure of performance is a visual rating score which indicates the perceived quality of the edges for identifying an object. The process of evaluating edge detection algorithms with this performance measure requires the collection of a set of gray-scale images, optimizing the input parameters for each algorithm, conducting visual evaluation experiments and applying statistical analysis methods. The novel aspect of this work is the use of a visual task and real images of complex scenes in evaluating edge detectors. The method is appealing because, by definition, the results agree with visual evaluations of the edge images. Michael D. Heath, Sudeep Sarkar, Thomas A. Sanocki, Kevin W. Bowyer |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1996 | Comparison of Edge Detectors: A Methodology and Initial StudyabstractThe purpose of this paper is to describe a new (to computer vision) experimental framework which allows us to make quantitative comparisons using subjective ratings made by people. This approach avoids the issue of pixel-level ground truth. As a result, it does not allow us to make statements about the frequency of false positive and false negative errors at the pixel level. Instead, using experimental design and statistical techniques borrowed from Psychology, we make statements about whether the outputs of one edge detector are rated statistically significantly higher than the outputs of another. This approach offers itself as a nice complement to signal-based quantitative measures. Also, the evaluation paradigm in this paper is goal oriented; in particular, we consider edge detection in the context of object recognition. The human judges rate the edge, detectors based on how well the capture the salient features of real objects. So far, edge detection modules have been designed and evaluated in isolation, except for the recent work by Ramesh and Haralick (1992). The only prior work (that we are aware of) which also uses humans to rate image algorithms is that of Reeves and Higdon (1995). They use human ratings to decide on regularization parameters of image restoration. Fram and Deutch (1975) also used human subjects, however, the focus was on human versus machine performance rather than using human ratings to compare different edge detectors. The use of human judges to rate image outputs mist be approached systematically. Experiments must be designed and conducted carefully, and results interpreted with appropriate statistical tools. The use of statistical analysis in vision system performance characterization has been rare. The only prior work in the area that we are aware of is that of Nair et al. (1995), who used statistical ranking procedures to compare neural network based object recognition systems. Michael D. Heath, Sudeep Sarkar, Thomas A. Sanocki, Kevin W. Bowyer |
CVPR | 2 |
| 1996 | Quantitative Measures of Change based on Feature Organization: Eigenvalues and EigenvectorsabstractWe propose four measures of image organizational change which can be used to monitor construction activity. The measures are based on the thesis that the progress of construction will see a change in the individual image feature attributes as well as an evolution in the relationships among these features. This change in the relationship is captured by the eigenvalues and eigenvectors of the relation graph embodying the organization among the image features. We demonstrate the ability of the measures to differentiate between no development, the onset of construction, and full development, on the available real test image set. Sudeep Sarkar, Kim L. Boyer |
CVPR | 1 |
| 1995 | Using Perceptual Inference Networks to Manage Vision Processes
Sudeep Sarkar, Kim L. Boyer |
Comput. Vis. Image Underst. | 1 |
| 1994 | Automated design of Bayesian perceptual inference networksabstractWe previously presented (Sarkar and Boyer, 1993) the Perceptual Inference Network (PIN), a formalism based on Bayesian Networks, to reason among a set of object or feature hypotheses and to integrate multiple sources of information in the context of perceptual organization. The design of a PIN requires knowledge of the dependency structure among the organizations of interest and the specification of the conditional probabilities. This design was done manually with large doses of tedium and guesswork. In this paper we present an algorithm based on structural entropic measures and random parametric structural descriptions (RPSDs) to design a PIN automatically and in a (more) theoretically sound fashion. Experimental results present evidence of the robustness of the algorithm and make performance comparisons on real image data with a manually structured PIN. Since PINs are a form of Bayesian Network, we hope that this work will also prove useful towards structuring Bayesian Networks in other computer vision contexts.> Sudeep Sarkar, Kim L. Boyer |
CVPR | 1 |
| 1994 | Using perceptual inference networks to manage vision processesabstractThe aim is to generate a hierarchical description of the scene using preattentive and attentive modules. The preattentive module provides evidence in terms of primitive organizations like parallelism, continuity, closure, and strands. The attentive organization integrates this preattentive evidence to hypothesize more complex organizations such as parallelograms, circles, ellipses, and ribbons. This attentive part is realized by the perceptual inference network (PIN) which is a form of Bayesian network. The output set of hypotheses of the PIN is large and redundant. A set of lines is described as a parallelogram and/or ellipse and/or circle. There is considerable ambiguity in such a description. The strategy is to use special-purpose modules to resolve the ambiguous hypotheses and to generate a comprehensive scene description. These special purpose modules tend to be computationally expensive and have limited applicability. Therefore, we want to apply them only when and where we expect the greatest amount of information gain per unit computational resource. Sudeep Sarkar, Kim L. Boyer |
ICPR (1) | 1 |
| 1994 | "On the localization performance measure and optimal edge detection"abstractTagare and deFigueiredo (see ibid., vol. 12, p. 1186-1189, 1990 ) present a localization performance measure for edge detectors. They correctly point out a flaw in Canny's formulation of the localization criterion, which was subsequently adopted by Sarkar and Boyer (1991). They motivate their form of the localization criterion along a different line of reasoning. In this comment, the authors show that although Canny's derivation was in error, the final form of his criterion is adequate and can, in fact, be derived from Tagare and deFigueiredo's formulation of the problem. The authors also point out some problems with Tagare and deFigueiredo's localization criterion.> Kim L. Boyer, Sudeep Sarkar |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1994 | A Computational Structure for Preattentive Perceptual Organization: Graphical Enumeration and Voting MethodsabstractPresents an efficient computational structure for preattentive perceptual organization. By perceptual organization the authors refer to the ability of a vision system to organize features detected in images based on viewpoint consistency and other Gestaltic perceptual phenomena. This usually has two components, a primarily bottom up preattentive part and a top down attentive part, with meaningful features emerging in a synergistic fashion from the original set of (very) primitive features. In this work the authors advance a computational structure for preattentive perceptual organization. The authors propose a hierarchical approach, using voting methods to build associations through consensus and relational graphs to represent the organization at each level. The voting method is very efficient in terms of time and space and performs impressively for a wide range of organizations. The graphical representation allows the ready extraction of higher order features, or perceptual tokens, because the relational information is rendered explicit.> Sudeep Sarkar, Kim L. Boyer |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 1993 | Integration, Inference, and Management of Spatial Information Using Bayesian Networks: Perceptual OrganizationabstractThe formalism of Bayesian networks provides a very elegant solution, in a probabilistic framework, to the problem of integrating top-down and bottom-up visual processes, as well serving as a knowledge base. The formalism is modified to handle spatial data, and thus the application of Bayesian networks is extended to visual processing. The modified form is called the perceptual inference network (PIN). The theoretical background of a PIN is presented, and its viability is demonstrated in the context of perceptual organization. Perceptual organization imparts robustness, efficiency, and a qualitative and holistic nature to vision. Thus far, the approaches to the problem of perceptual organization have been purely bottom up, without much top-down knowledge-base influence, and are therefore entirely dependent on the inputs, which are obviously imperfect. The knowledge base, besides coping with such input imperfection, also makes it possible to integrate multiple organizations and form a composite organization hypothesis. The PIN imparts an active inferential and integrating nature to perceptual organization in an elegant probabilistic framework.> Sudeep Sarkar, Kim L. Boyer |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1993 | Perceptual organization in computer vision: a review and a proposal for a classificatory structureabstractThe role of perceptual organization in computer vision systems is explored. This is done from four vantage points. A brief history of perceptual organization research in both humans and computer vision is offered. A classificatory structure in which to cast perceptual organization research to clarify both the nomenclature and the relationships among the many contributions is proposed. The perceptual organization work in computer vision in the context of this classificatory structure is reviewed. The array of computational techniques applied to perceptual organization problems in computer vision is surveyed.> Sudeep Sarkar, Kim L. Boyer |
IEEE Trans. Syst. Man Cybern. | 1 |
| 1992 | Perceptual organization using Bayesian networksabstractIt is shown that the formalism of Bayesian networks provides an elegant solution, in a probabilistic framework, to the problem of integrating top-down and bottom-up visual processes as well serving as a knowledge base. The formalism is modified to handle spatial data and thus extends the applicability of Bayesian networks to visual processing. The modified form is called the perceptual inference network (PIN). The theoretical background of a PIN is presented, and its viability is demonstrated in the context of perceptual organization. The PIN imparts an active inferential and integrating nature to perceptual organization.> Sudeep Sarkar, Kim L. Boyer |
CVPR | 1 |
| 1992 | Computing perceptual organization using voting methods and graphical enumerationabstractPresents an efficient hierarchical computational paradigm for perceptual organization. Organization at each level of the hierarchy is done by graph enumeration on a set of Gestalt graphs. Efficient voting methods are proposed. The authors develop the method in detail and analyze its computational efficiency, considering both time and space. The theoretical and practical results are very encouraging. They strongly advocate the idea of an organizational hierarchy, constructed using graph enumeration. Graph theoretic representations enable one to extract various structures with considerable ease. They evaluated the performance using real images, with good results.> Sudeep Sarkar, Kim L. Boyer |
ICPR (1) | 1 |
| 1991 | Optimal infinite impulse response zero crossing based edge detectors
Sudeep Sarkar, Kim L. Boyer |
CVGIP Image Underst. | 1 |
| 1991 | On Optimal Infinite Impulse Response Edge Detection FiltersabstractThe authors outline the design of an optimal, computationally efficient, infinite impulse response edge detection filter. The optimal filter is computed based on Canny's high signal to noise ratio, good localization criteria, and a criterion on the spurious response of the filter to noise. An expression for the width of the filter, which is appropriate for infinite-length filters, is incorporated directly in the expression for spurious responses. The three criteria are maximized using the variational method and nonlinear constrained optimization. The optimal filter parameters are tabulated for various values of the filter performance criteria. A complete methodology for implementing the optimal filter using approximating recursive digital filtering is presented. The approximating recursive digital filter is separable into two linear filters, operating in two orthogonal directions. The implementation is very simple and computationally efficient. has a constant time of execution for different sizes of the operator, and is readily amenable to real-time hardware implementation.> Sudeep Sarkar, Kim L. Boyer |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1991 | Dynamic edge warping: an experimental system for recovering disparity maps in weakly constrained systemsabstractDynamic edge warping (DEW), a technique for recovering reasonably accurate disparity maps from uncalibrated stereo image pairs, is presented. No precise knowledge of the epipolar camera geometry is assumed. The technique is embedded in a system including structural stereopsis on the front end and robust estimation in digital photogrammetry on the other for the purpose of self-calibrating stereo image pairs. Once the relative camera orientation is known, the epipolar geometry is computed and the system can use this information to refine its representation of the object space. Such a system will find application in the autonomous extraction of terrain maps from stereo aerial photographs, for which camera position and orientation are unknown a priori, and for online autonomous calibration maintenance for robotic vision applications, in which the cameras are subject to vibration and other physical disturbances after calibration. This work thus forms a component of an intelligent system that begins with a pair of images and, having only vague knowledge of the conditions under which they were acquired, produces an accurate, dense, relative depth map. The resulting disparity map can also be used directly in some high-level applications involving qualitative scene analysis, spatial reasoning, and perceptual organization of the object space. The system as a whole substitutes high-level information and constraints for precise geometric knowledge in driving and constraining the early correspondence process.> Kim L. Boyer, Daniel M. Wuescher, Sudeep Sarkar |
IEEE Trans. Syst. Man Cybern. | 3 |
| 1990 | Dynamic edge warping: experiments in disparity estimation under weak constraintsabstractA technique, dynamic edge warping, for recovering reasonable disparity maps from uncalibrated stereo image pairs is presented. No precise knowledge of the epipolar camera geometry is assumed. The technique is part of a system including structural stereopsis and digital photogrammetry for self-calibrating stereo image pairs with application in autonomous extraction of terrain maps from stereo aerial photographs and autonomous calibration maintenance in robotic vision. The system substitutes high-level information and constraints for precise geometric knowledge, in constraining the matching process.> Kim L. Boyer, Daniel M. Wuescher, Sudeep Sarkar |
ICCV | 3 |
| 1990 | Optimal, efficient, recursive edge detection filtersabstractThe design of an optimal, efficient, infinite-impulse-response (IIR) edge detection filter is described. J. Canny (1986) approached the problem by formulating three criteria designed in any edge detection filter: good detection, good localization, and low spurious response. He maximized the product of the first two criteria while keeping the spurious response criterion constant. Using the variational approach, he derived a set of finite extent step edge detection filters corresponding to various values of the spurious response criterion, approximating the filters by the first derivative of a Gaussian. A more direct approach is described in this paper. The three criteria are formulated as appropriate for a filter of infinite impulse response, and the calculus of variations is used to optimize the composite criteria. Although the filter derived is also well approximated by first derivative of a Gaussian, a superior recursively implemented approximation is achieved directly. The approximating filter is separable into two linear filters operating in two orthogonal directions allowing for parallel edge detection processing. The implementation is very simple and computationally efficient.> Sudeep Sarkar, Kim L. Boyer |
ICPR (1) | 1 |