EDBT 2026 Demo / reviewers in the wild / expert
Shishir Shah 0001
dblp:85/8613 · also Shishir K. Shah
· DBLP profile ↗
80ranked-venue papers
14as first author
11since 2021 · last 2025
0000-0003-4093-6906ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 59 · 9 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 56 · 10 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 since 2021Security and privacy · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Occlusion-aware appearance and shape learning for occluded cloth-changing person re-identification
Vuong D. Nguyen, Pranav Mantini, Shishir Shah 0001 |
Pattern Anal. Appl. | 3 |
| 2024 | Cross-Modality Complementary Learning for Video-Based Cloth-Changing Person Re-identification
Vuong D. Nguyen, Pranav Mantini, Shishir Shah 0001 |
ACCV (1) | 3 |
| 2024 | CrossViT-ReID: Cross-Attention Vision Transformer for Occluded Cloth-Changing Person Re-Identification
Vuong D. Nguyen, Pranav Mantini, Shishir Shah 0001 |
ACCV (1) | 3 |
| 2024 | Occluded Cloth-Changing Person Re-Identification via Occlusion-aware Appearance and Shape ReasoningabstractExisting methods in Person Re-Identification (ReID) often fail when simultaneously confronted with occlusions and clothing changes. In this paper, we introduce a challenging yet practical task called Occluded Cloth-Changing Re-ID (OCCRe-ID/). We propose Occlusion-aware Appearance and Shape Reasoning, the first framework OCCReID. We first propose an occlusion synthesis strategy to expose the model to real-world occlusion variations. We mitigate clothing changes by coupling silhouette-based body shape information with appearance. Unlike previous works that directly leverage unreliable features extracted from occluded images by off-the-shelf backbones, we propose an occlusion-awareness strategy to handle occlusions for ReID. An occlusion detection module is elaborately designed to generate occlusion-aware feature, which is then used to guide the framework to reason robust appearance and shape features. Extensive experiments demonstrate the superiority of our framework over both cloth-changing Re-ID and occluded Re-ID methods. Vuong D. Nguyen, Pranav Mantini, Shishir Shah 0001 |
AVSS | 3 |
| 2024 | Occlusion-aware Cross-Attention Fusion for Video-based Occluded Cloth-Changing Person Re-IdentificationabstractVideo-based Person Re-Identification (Re-ID) is an important task in video surveillance analysis. Real-world video-based Re-ID commonly suffers from clothing changes and occlusions, which severely degenerates performance of traditional Re-ID methods. In this paper, we introduce a challenging yet practical task called Video-based Occluded Cloth-Changing Re-ID (VOCCRe-ID). To tackle occlusions, we propose an occlusion synthesis strategy to expose the model to real-world occlusion variations. To mitigate unreliable appearance caused by clothing changes, we couple body shape information from the normalized silhouette sequence. Then, we propose a cross-attention fusion mechanism to capture the complementary relationships between appearance and shape under occlusions, thus enhancing Re-ID robustness. In addition, since there are no dataset for VOCCRe-ID, we build the large-scale Occluded-VCCR dataset which explicitly presents occlusions and contains the most clothing variations. Extensive experiments show that we achieve SOTA performance over previous methods. Vuong D. Nguyen, Pranav Mantini, Shishir Shah 0001 |
IJCB | 3 |
| 2024 | ACML: Attention-Based Cross-Modality Learning For Cloth-Changing and Occluded Person Re-IdentificationabstractPerson Re-Identification (Re-ID) aims at matching a person captured by a non-overlapping camera system. Real-world Re-ID presents challenges like clothing changes and occlusions, which limits the applicability of traditional appearance-based methods. Cloth-Changing Re-ID (CCRe-ID) methods that rely on cloth-invariant modalities, such as shape, gait, etc., ignore occlusions and fail to mine the complementary relationship across modalities. Meanwhile, methods that explicitly focus on occlusion management struggle with cloth-changing scenarios. To address these, we propose ACML: Attention-based Cross-Modality Learning, the first framework to tackle both clothing changes and occlusion in Re-ID. Our lightweight framework comprises a unified network with cascaded Cross-Attention Blocks that extracts appearance and shape features collaboratively, enhancing robustness under clothing changes, viewpoint variations, and poor illumination conditions. Inputs to the network are produced by our novel occlusion synthesis module, which not only helps exposing the model to occlusions but also guides the model to adaptively attend to informative cues and reduce noise. Experiments demonstrate the effectiveness of ACML on both CCRe-ID and occluded Re-ID datasets. Vuong D. Nguyen, Pranav Mantini, Shishir Shah 0001 |
ICIP | 3 |
| 2024 | Recall-Based Knowledge Distillation for Data Distribution Based Catastrophic Forgetting in Semantic Segmentation
Samiha Mirza, Apurva Gala, Pandu Devarakota, Vuong D. Nguyen, Pranav Mantini, Shishir Shah 0001 |
ICPR (23) | 6 |
| 2024 | Contrastive Viewpoint-aware Shape Learning for Long-term Person Re-IdentificationabstractTraditional approaches for Person Re-identification (ReID) rely heavily on modeling the appearance of persons. This measure is unreliable over longer durations due to the possibility for changes in clothing or biometric information. Furthermore, viewpoint changes significantly degrade the matching ability of these methods. In this paper, we propose "Contrastive Viewpoint-aware Shape Learning for Long-term Person Re-Identification" (CVSL) to address these challenges. Our method robustly extracts local and global texture-invariant human body shape cues from 2D pose using the Relational Shape Embedding branch, which consists of a pose estimator and a shape encoder built on a Graph Attention Network. To enhance the discriminability of the shape and appearance of identities under viewpoint variations, we propose Contrastive Viewpoint-aware Losses (CVL). CVL leverages contrastive learning to simultaneously minimize the intra-class gap under different viewpoints and maximize the inter-class gap under the same viewpoint. Extensive experiments demonstrate that our proposed framework outperforms state-of-the-art methods on long-term person Re-ID benchmarks. Vuong D. Nguyen, Khadija Khaldi, Pranav Mantini, Shishir Shah 0001 |
WACV | 5 |
| 2023 | From Perception to Precision: Navigating Perceptual Loss in MRI Super-ResolutionabstractIn the field of MRI super-resolution, training an image upscaling network under a pixel-oriented cost function (e.g., Mean-Intensity-Error) has proven to boost the signal-to-noise ratio. However, these types of cost functions tend to miss high-frequency details and fail to achieve an ideal sharpness, which is a pivotal image property for clinical applications to make diagnoses. To address this issue, the cost function of these upscaling networks typically includes a perceptual loss function, which is well recognized for the reconstruction of textures and enhancing sharpness, in addition to a pixel-oriented one. In this paper, we investigate the effect of perceptual loss on several MRI super-resolution metrics. We train UNet architecture under two loss function scenarios: One only including a pixel-oriented loss function, and the other a fusion of pixel-oriented and perceptual losses. We then employ an ablation study using a mixed effect model on a comprehensive set of evaluation criteria to measure the significance of change upon the inclusion of perceptual loss. Our results show that even though perceptual loss substantially shifts the networks towards outputting sharper images, it only causes negligible performance degradation in the accuracy of the reconstructed regions of interest, which can be alleviated using proper hyperparameter tuning. Mohammad Javadi, Panagiotis Tsiamyrtzis, Shishir Shah 0001, Ernst L. Leiss, Nikolaos V. Tsekos |
BIBE | 4 |
| 2023 | Identification of Visual Objects in Lecture Videos with Color and Keypoints AnalysisabstractRecorded lecture videos are an increasingly important learning resource. However, traditional video format does not allow quick navigation to the desired content of interest. Recent research has enhanced navigation by dividing lecture videos into chapters and creating a summary of each chapter. The visual content on lecture video frames represents a valuable source of information for identifying topic boundaries as well as summarizing content. The focus of the research presented in this paper is to accurately identify visual objects in lecture video frames. The methods developed for camera videos are not directly applicable here as the visual content includes charts, graphs, and illustrations intermingled with text. A common approach based on locating regions with continuous pixel changes has a key limitation that logically consistent visual objects can have modest size gaps inside them. The result is over-segmentation, where a logical object is split into multiple objects if the gap threshold is too low, or under-segmentation, where adjacent objects are recognized as a single large object if the gap threshold is too high. This paper introduces a novel approach that exploits the observation that components of logical objects often have color and geometrical similarity. In our methodology, first a relatively large number of visual elements are identified with a small gap threshold. Subsequently, these visual elements are selectively combined using gap along with color and geometrical similarity. An evaluation was conducted with a suite of 170 lecture video frames from STEM coursework. The results demonstrate the significant impact of color and geometry in improving the accuracy of visual object identification in lecture video frames. Dipayan Biswas, Shishir Shah 0001, Jaspal Subhlok |
ISM | 2 |
| 2021 | A Graph-Based Approach for Making Consensus-Based Decisions in Image Search and Person Re-IdentificationabstractImage matching and retrieval is the underlying problem in various directions of computer vision research, such as image search, biometrics, and person re-identification. The problem involves searching for the closest match to a query image in a database of images. This work presents a method for generating a consensus amongst multiple algorithms for image matching and retrieval. The proposed algorithm, Shortest Hamiltonian Path Estimation (SHaPE), maps the process of ranking candidates based on a set of scores to a graph-theoretic problem. This mapping is extended to incorporate results from multiple sets of scores obtained from different matching algorithms. The problem of consensus-based decision-making is solved by searching for a suitable path in the graph under specified constraints using a two-step process. First, a greedy algorithm is employed to generate an approximate solution. In the second step, the graph is extended and the problem is solved by applying Ant Colony Optimization. Experiments are performed for image search and person re-identification to illustrate the efficiency of SHaPE in image matching and retrieval. Although SHaPE is presented in the context of image retrieval, it can be applied, in general, to any problem involving the ranking of candidates based on multiple sets of scores. Arko Barman, Shishir Shah 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | A Day on Campus - An Anomaly Detection Dataset for Events in a Single Camera
Pranav Mantini, Zhenggang Li, Shishir Shah 0001 |
ACCV (6) | 3 |
| 2020 | CQNN: Convolutional Quadratic Neural NetworksabstractImage classification is a fundamental task in computer vision. A variety of deep learning models based on the Convolutional Neural Network (CNN) architecture have proven to be an efficient solution. Numerous improvements have been proposed over the years, where broader, deeper, and denser networks have been constructed. However, the atomic operation for these models has remained a linear unit (single neuron). In this work, we pursue an alternative dimension by hypothesizing the atomic operation to be performed by a quadratic unit. We construct convolutional layers using quadratic neurons for feature extraction and subsequently use dense layers for classification. We perform analysis to quantify the implication of replacing linear neurons with quadratic units. Results show a keen improvement in classification accuracy with quadratic neurons over linear neurons. Pranav Mantini, Shishir Shah 0001 |
ICPR | 2 |
| 2020 | Visual Summarization of Lecture Video Segments for Enhanced NavigationabstractThe following topics are dealt with: video signal processing; video streaming; feature extraction; deep learning (artificial intelligence); video coding; learning (artificial intelligence); quality of experience; computer aided instruction; convolutional neural nets; image segmentation. Mohammad Rajiur Rahman, Shishir Shah 0001, Jaspal Subhlok |
ISM | 2 |
| 2019 | UHCTD: A Comprehensive Dataset for Camera Tampering DetectionabstractAn unauthorized or an accidental change in the view of a surveillance camera is called a tampering. Algorithms that detect tampering by analyzing the video are referred to as camera tampering detection algorithms. Most evaluations on camera tampering detection methods are presented based on individually collected datasets. One of the major challenges in the area of camera tamper detection is the absence of a public dataset with sufficient size and variations for an extensive performance evaluation. We propose a large scale synthetic dataset called University of Houston Camera Tampering Detection dataset (UHCTD) for development and testing of camera tampering detection methods. The dataset consists of a total 576 tampers with over 288 hours of video captured from two surveillance cameras. To establish an initial benchmark, we cast camera tampering detection as a classification problem. We train and evaluate three different deep architectures that have shown promise in scene classification, Alexnet, Resnet, and Densenet. Results are presented to show how the dataset can be used to train and classify images as normal, and tampered within and across cameras. Pranav Mantini, Shishir Shah 0001 |
AVSS | 2 |
| 2018 | A Generalized Optimization Framework for Score Aggregation in Person Re-identification SystemsabstractPerson re-identification is the problem of identifying a person over multiple cameras in video-based surveillance. In this paper, we propose a novel generalized optimization framework for combining results from different methods for person re-identification to significantly improve re-identification rates. The proposed framework evaluates the similarity of score distributions by means of Bhattacharyya distance to arrive at an optimum solution that minimizes a defined cost function. Using our framework, we employ similarity scores from existing algorithms to generate score aggregates, which are then used for ranking the gallery images based on their "closeness" to a given probe image. The optimization problem is solved using Genetic Algorithm, a heuristic optimization algorithm. Our results show significant improvement in performance for person re-identification using existing algorithms on challenging datasets - VIPeR, CUHK01, CUHK03 and QMUL-GRID. Arko Barman, Shishir Shah 0001 |
AVSS | 2 |
| 2018 | GoDP: Globally Optimized Dual Pathway deep network architecture for facial landmark localization in-the-wild
Yuhang Wu 0002, Shishir Shah 0001, Ioannis A. Kakadiaris |
Image Vis. Comput. | 2 |
| 2018 | Annotated face model-based alignment: a robust landmark-free pose estimation approach for 3D model registration
Yuhang Wu 0002, Shishir Shah 0001, Ioannis A. Kakadiaris |
Mach. Vis. Appl. | 2 |
| 2018 | Monocular 3D facial shape reconstruction from a single 2D image with coupled-dictionary learning and sparse coding
Pengfei Dou, Yuhang Wu 0002, Shishir Shah 0001, Ioannis A. Kakadiaris |
Pattern Recognit. | 3 |
| 2017 | A signal detection theory approach for camera tamper detectionabstractCamera tamper detection is the ability to detect faults and operational failures in video surveillance cameras by analyzing the video. Researchers have increasingly focused on such techniques attributing to the ubiquitous deployment of large scale surveillance systems. In this paper, a signal detection theory approach is proposed to quantitatively analyze the information being captured by the camera and to detect tampers. Signal activity is used as a feature to measure the amount of information in the image. The distribution of features representing the normal operation of a camera are modeled as a Gaussian mixture model (GMM). The GMM is trained using synthetic data. To reduce the effects of noise, a Kalman filter is used to model changes in signal activity in the video. Experimental results show that the proposed approach out performed the state-of-the-art [13] in detecting tampered images with higher accuracy while generating lower false alarms. Pranav Mantini, Shishir Shah 0001 |
AVSS | 2 |
| 2017 | End-to-End 3D Face Reconstruction with Deep Neural NetworksabstractMonocular 3D facial shape reconstruction from a single 2D facial image has been an active research area due to its wide applications. Inspired by the success of deep neural networks (DNN), we propose a DNN-based approach for End-to-End 3D FAce Reconstruction (UH-E2FAR) from a single 2D image. Different from recent works that reconstruct and refine the 3D face in an iterative manner using both an RGB image and an initial 3D facial shape rendering, our DNN model is end-to-end, and thus the complicated 3D rendering process can be avoided. Moreover, we integrate in the DNN architecture two components, namely a multi-task loss function and a fusion convolutional neural network (CNN) to improve facial expression reconstruction. With the multi-task loss function, 3D face reconstruction is divided into neutral 3D facial shape reconstruction and expressive 3D facial shape reconstruction. The neutral 3D facial shape is class-specific. Therefore, higher layer features are useful. In comparison, the expressive 3D facial shape favors lower or intermediate layer features. With the fusion-CNN, features from different intermediate layers are fused and transformed for predicting the 3D expressive facial shape. Through extensive experiments, we demonstrate the superiority of our end-to-end framework in improving the accuracy of 3D face reconstruction. Pengfei Dou, Shishir Shah 0001, Ioannis A. Kakadiaris |
CVPR | 2 |
| 2017 | SHaPE: A Novel Graph Theoretic Algorithm for Making Consensus-Based Decisions in Person Re-identification SystemsabstractPerson re-identification is a challenge in video-based surveillance where the goal is to identify the same person in different camera views. In recent years, many algorithms have been proposed that approach this problem by designing suitable feature representations for images of persons or by training appropriate distance metrics that learn to distinguish between images of different persons. Aggregating the results from multiple algorithms for person re-identification is a relatively less-explored area of research. In this paper, we formulate an algorithm that maps the ranking process in a person re-identification algorithm to a problem in graph theory. We then extend this formulation to allow for the use of results from multiple algorithms to make a consensus-based decision for the person re-identification problem. The algorithm is unsupervised and takes into account only the matching scores generated by multiple algorithms for creating a consensus of results. Further, we show how the graph theoretic problem can be solved by a two-step process. First, we obtain a rough estimate of the solution using a greedy algorithm. Then, we extend the construction of the proposed graph so that the problem can be efficiently solved by means of Ant Colony Optimization, a heuristic path-searching algorithm for complex graphs. While we present the algorithm in the context of person reidentification, it can potentially be applied to the general problem of ranking items based on a consensus of multiple sets of scores or metric values. Arko Barman, Shishir Shah 0001 |
ICCV | 2 |
| 2017 | 3D-2D face recognition with pose and illumination normalization
Ioannis A. Kakadiaris, George Toderici, Georgios Evangelopoulos, Georgios Passalis, Dat Chu, Xi Zhao 0001, Shishir Shah 0001, Theoharis Theoharis |
Comput. Vis. Image Underst. | 7 |
| 2017 | Hierarchical Multi-label Classification using Fully Associative Ensemble Learning
Lingfeng Zhang 0001, Shishir Shah 0001, Ioannis A. Kakadiaris |
Pattern Recognit. | 2 |
| 2014 | Robust 3D Face Shape Reconstruction from Single Images via Two-Fold Coupled Structure Learning and Off-the-Shelf Landmark Detectors
Pengfei Dou, Yuhang Wu 0002, Shishir Shah 0001, Ioannis A. Kakadiaris |
BMVC | 3 |
| 2014 | Fully Associative Ensemble Learning for Hierarchical Multi-Label Classification
Lingfeng Zhang 0001, Shishir Shah 0001, Ioannis A. Kakadiaris |
BMVC | 2 |
| 2014 | What Do I See? Modeling Human Visual Perception for Multi-person Tracking
Xu Yan 0003, Ioannis A. Kakadiaris, Shishir Shah 0001 |
ECCV (2) | 3 |
| 2014 | Benchmarking 3D Pose Estimation for Face Recognitionabstract3D-Model-Aided 2D face recognition (MaFR) has attracted a lot of attention in recent years. By registering a 3D model, facial textures of the gallery and the probe can be lifted and aligned in a common space, thus alleviating the challenge of pose variations. One obstacle preventing accurate registration is the 3D-2D pose estimation, which is easily affected by landmarks. In this work, we present the performance that state-of-the-art pose estimation algorithms could reach using state-of-the-art automatic landmark localization methods. We generated an application-specific dataset with more than 59,000 synthetic face images and ground truth camera pose and landmarks, covering 45 poses and six illumination conditions. Our experiments compared four recently proposed pose estimation algorithms using 2D landmarks detected by two automatic methods. Our results highlight one near-real-time landmark detection method and a highly accurate pose estimation algorithm, which would potentially boost the 3D-Model-Aided 2D face recognition performance. Pengfei Dou, Yuhang Wu 0002, Shishir Shah 0001, Ioannis A. Kakadiaris |
ICPR | 3 |
| 2014 | Regularized Multi-view Multi-metric Learning for Action RecognitionabstractAlthough multi-view datasets have become more accessible in the real-world applications, most state-of-the-art action recognition methods applied to those datasets rely on simple view agreement when combining local information from various views together. This leads to deteriorated performance in situations with view insufficiency and view disagreements. In this paper, we propose a novel framework for boosting action recognition performance by quantifying the connection between the viewpoint and an action. The proposed approach searches for the best combination of multiple views based on a co-learning strategy that simultaneously learns a local distance metric related to each action class and the relationships between each viewpoint and the action category. Consequently, the spatio-temporal representation of each action class in different viewpoints plays a key role in shaping the local distance metric space. We test our method on the IXMAS dataset and shows competitive performance compared to other state-of-the-art methods. Xuqing Wu 0001, Shishir Shah 0001 |
ICPR | 2 |
| 2014 | Hierarchical Group Structures in Multi-person TrackingabstractThis paper presents a novel approach for improving multi-person tracking using hierarchical group structures. The groups are identified by a bottom-up social group discovery method. The inter- and intra-group structures are modeled as a two-layer graph and tracking is posed as optimization of the integrated structure. The target appearance is modeled using HOG features, and the tracking solution is obtained via dynamic programming. The group structures are updated continuously and re-initialized intermittently using collected tracking evidence. We test our method on videos from four challenging datasets and evaluate it against state-of-the-art trackers. The significant performance improvement shows the importance of modeling the intra-group relationships and the advantage of the two-layer graph structure. Xu Yan 0003, Anil M. Cheriyadat, Shishir Shah 0001 |
ICPR | 3 |
| 2014 | A survey of approaches and trends in person re-identification
Apurva Gala, Shishir Shah 0001 |
Image Vis. Comput. | 2 |
| 2014 | Modeling local behavior for predicting social interactions towards human tracking
Xu Yan 0003, Ioannis A. Kakadiaris, Shishir Shah 0001 |
Pattern Recognit. | 3 |
| 2014 | Activity analysis in crowded environments using social cues for group discovery and human interaction modeling
Khai N. Tran, Apurva Gala, Ioannis A. Kakadiaris, Shishir Shah 0001 |
Pattern Recognit. Lett. | 4 |
| 2014 | Minimizing Illumination Differences for 3D to 2D Face Recognition Using Lighting MapsabstractAsymmetric 3D to 2D face recognition has gained attention from the research community since the real-world application of 3D to 3D recognition is limited by the unavailability of inexpensive 3D data acquisition equipment. A 3D to 2D face recognition system explicitly relies on 3D facial data to account for uncontrolled image conditions related to head pose or illumination. We build upon such a system, which matches relit gallery textures with pose-normalized probe images, using the gallery facial meshes. The relighting process, however, is based on an assumption of indoor lighting conditions and limits recognition performance on outdoor images. In this paper, we propose a novel method for minimizing illumination difference by unlighting a 3D face texture via albedo estimation using lighting maps. The algorithm is evaluated on challenging databases (UHDB30, UHDB11, FRGC v2.0) with drastic lighting and pose variations. The experimental results demonstrate the robustness of our method for estimating the albedo from both indoor and outdoor captured images, and the effectiveness and efficiency for illumination normalization in face recognition. Xi Zhao 0001, Georgios Evangelopoulos, Dat Chu, Shishir Shah 0001, Ioannis A. Kakadiaris |
IEEE Trans. Cybern. | 4 |
| 2013 | Change Detection in Dynamic Scenes using Local Adaptive TransformabstractIn this paper, we propose a framework that can be used for detecting relevant changes in highly dynamic scenes, where the background has several changing elements. To establish a clear distinction between what is relevant and what is not is a very challenging task. Therefore, we first categorize the changes into two main classes called ordinary changes and relevant changes. Detected changes are considered as irrelevant if they are recurrent elements and changes pertaining on the dynamic background of the scene. The proposed framework makes use of a set of orthogonal linear transforms to capture spatiotemporal signatures of local ordinary change patterns and subsequently employ them in the detection of relevant changes. The use of this framework is demonstrated in a variety of videos with highly dynamic backgrounds including lakes, pools, and roads. Compared to existing methods reported on the same test videos, the proposed framework detects the relevant changes more accurately. Hakan Haberdar, Shishir Shah 0001 |
BMVC | 2 |
| 2013 | Longitudinal Characterization of Breast Morphology during Reconstructive SurgeryabstractQuantitative analysis of breast morphology facilitates pre-operative planning and post-operative outcome assessments in breast reconstruction. Our project is developing algorithms to quantify changes in local breast morphology occurring over time. The project encompasses three topics: (1) Three-dimensional (3D) images registration, (2) Breast contour detection, and (3) Quantitative analysis of local breast morphology changes. We developed a semi-automated 3D image registration algorithm. We have also developed an approach to directly compute breast contour on 3D images. In the future, we will improve existing and develop additional algorithms to fulfill our project goals. Lijuan Zhao, Shishir Shah 0001, Fatima A. Merchant |
ISM | 2 |
| 2013 | 3D Face Discriminant Analysis Using Gauss-Markov Posterior MarginalsabstractWe present a Markov Random Field model for the analysis of lattices (e.g., images or 3D meshes) in terms of the discriminative information of their vertices. The proposed method provides a measure field that estimates the probability of each vertex being "discriminative" or "nondiscriminative" for a given classification task. To illustrate the applicability and generality of our framework, we use the estimated probabilities as feature scoring to define compact signatures for three different classification tasks: 1) 3D Face Recognition, 2) 3D Facial Expression Recognition, and 3) Ethnicity-based Subject Retrieval, obtaining very competitive results. The main contribution of this work lies in the development of a novel framework for feature selection in scenaria in which the most discriminative information is smoothly distributed along a lattice. Omar Ocegueda, Tianhong Fang, Shishir Shah 0001, Ioannis A. Kakadiaris |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Towards Quality Aware Collaborative Video Analytic CloudabstractAs cloud diversifies into different application fields, understanding and characterizing the specific work load sand application requirements play important roles in the design of efficient cloud infrastructure and system software support. Video analytic is a rapidly advancing field and it is widely used in many application domains (i.e., health, medical care, surveillance, and defense). To support video analytic applications efficiently in cloud, one has to overcome many challenges such as lack of understanding of the relationship and trade off between analytic performance metrics and resource requirements. Furthermore, cloud computing has grown from the early model of resource sharing to data sharing and workflow sharing. To address the challenges and to lever age emerging trends, we propose and experiment with a domain specific cloud environment for video analytic applications. We design a cloud infrastructure framework for sharing video data, analytic software, and workflow. In addition, we create a video analytic quality aware resource plan model to guarantee users QoS and optimize usage of resources based on predictive knowledge of video analytic softwares performance metrics and a resource planning model that optimizes the overall analytic service quality under users constraints (i.e., time and cost).The predictive knowledge is represented as input and analytic software specific predictors. The experimental results show that the video analytic quality aware resource planning model can balance the tradeoff between analytic quality and resource requirements, and achieve optimal or near-optimal planning for video analytic workloads with constraints in a resource shared environment. Simulation studies show that resource planning results using ground truth and video analytic performance predictions are very similar, which indicates that our analytic quality/resource predictors are very accurate. Jong-Hyuk Lee, Tao Feng 0011, Larry Shi, Apurva Gala, Shishir Shah 0001, Hanako Yoshida |
IEEE CLOUD | 5 |
| 2012 | To Track or To Detect? An Ensemble Framework for Optimal Selection
Xu Yan 0003, Xuqing Wu 0001, Ioannis A. Kakadiaris, Shishir Shah 0001 |
ECCV (5) | 4 |
| 2012 | Development and evaluation of indexed captioned searchable videos for STEM courseworkabstractVideos of classroom lectures have proven to be a popular and versatile learning resource. This paper reports on videos featuring Indexing, Captioning, and Search capability (ICS Videos). The goal is to allow a user to rapidly search and access a topic of interest, a key shortcoming of the standard video format. A lecture is automatically divided into logical indexed video segments by analyzing video frames. Text is automatically identified with OCR technology enhanced with image transformations to drive keyword search. Captions can be added to videos. The ICS video player integrates indexing, search, and captioning in video playback and has been used by dozens of courses and 1000s of students. This paper reports on the development and evaluation of ICS videos framework and assessment of its value as an academic learning resource. Tayfun Tuna, Jaspal Subhlok, Lecia Jane Barker, Varun Varghese, Olin G. Johnson, Shishir Shah 0001 |
SIGCSE | 6 |
| 2012 | 3D/4D facial expression analysis: An advanced annotated face model approach
Tianhong Fang, Xi Zhao 0001, Omar Ocegueda, Shishir Shah 0001, Ioannis A. Kakadiaris |
Image Vis. Comput. | 4 |
| 2012 | Profile-based 3D-aided face recognition
Boris A. Efraty, Emil Bilgazyev, Shishir Shah 0001, Ioannis A. Kakadiaris |
Pattern Recognit. | 3 |
| 2012 | Part-based motion descriptor image for human action recognition
Khai N. Tran, Ioannis A. Kakadiaris, Shishir Shah 0001 |
Pattern Recognit. | 3 |
| 2012 | Part-based spatio-temporal model for multi-person re-identification
Apurva Gala, Shishir Shah 0001 |
Pattern Recognit. Lett. | 2 |
| 2011 | Sparse Representation-Based Super Resolution for Face Recognition At a DistanceabstractFace recognition is a challenging task, especially when low-resolution images or image sequences are used. A decrease in image resolution results in a loss of facial high frequency components leading to a decrease in recognition rates. In this paper, we propose a new method for super-resolution by building a dictionary of high-frequency components in the facial data, which are added to a low-resolution input image to create a super-resolved image. Our method is different from existing methods as we estimate the high-frequency components, rather than studying the direct relationship between the high- and low-resolution images. Quantitative and qualitative results are reported for both synthetic and surveillance facial image databases. Emil Bilgazyev, Boris A. Efraty, Shishir Shah 0001, Ioannis A. Kakadiaris |
BMVC | 3 |
| 2011 | Modeling Motion of Body Parts for Action RecognitionabstractThis paper presents a simple and computationally efficient framework for human action recognition based on modeling the motion of human body parts. Intuitively, a collective understanding of human part movements can lead to better understanding and representation of any human action. In this paper, we propose a generative representation of the motion of the human body parts to learn and classify human actions. The proposed representation combines the advantages of both local and global representations, encoding the relevant motion information as well as being robust to local appearance changes. Our work is motivated by the pictorial structures model and the framework of sparse representations for recognition. Human part movements are represented efficiently through quantization in the polar space. The key discrimination within each action is efficiently encoded by sparse representation to perform classification. The proposed method is evaluated on both the KTH and the UCF action datasets and the results are compared against other state-of-the-art methods. Khai N. Tran, Ioannis A. Kakadiaris, Shishir Shah 0001 |
BMVC | 3 |
| 2011 | Predicting Social Interactions for Visual TrackingabstractHuman interaction dynamics are known to play an important role in the development of robust pedestrian trackers that are applicable to a variety of applications in video surveillance. Traditional approaches to pedestrian tracking assume that each pedestrian walks independently and the tracker predicts the location based on an underlying motion model, such as a constant velocity or autoregressive model. Recent approaches have begun to leverage interaction, especially by modeling the repulsion force, among pedestrians to improve motion predictions. However, human interaction is more complex and is influenced by both repulsion and attraction effects. This motivates the use of a more complex human interaction model for pedestrian tracking. In this paper, we propose a novel visual tracking method by leveraging complex social interactions. We present an algorithm that decomposes social interactions into multiple potential interaction modes. We integrate these multiple social interaction modes into an interactive Markov Chain Monte Carlo tracker. We demonstrate how the developed method translates into a more informed motion prediction, resulting in a robust tracking performance. We test our method on videos from unconstrained outdoor environments and compare it against popular multi-object trackers. Xu Yan 0003, Ioannis A. Kakadiaris, Shishir Shah 0001 |
BMVC | 3 |
| 2011 | Which parts of the face give out your identity?abstractWe present a Markov Random Field model for the analysis of lattices (e.g., images or 3D meshes) in terms of the discriminative information of their vertices. The proposed method provides a measure field that estimates the probability of each vertex to be “discriminative” or “non-discriminative”. As an application of the proposed framework, we present a method for the selection of compact and robust features for 3D face recognition. The resulting signature consists of 360 coefficients, based on which we are able to build a classifier yielding better recognition rates than currently reported in the literature. The main contribution of this work lies in the development of a novel framework for feature selection in scenarios in which the most discriminative information is known to be concentrated along piece-wise smooth regions of a lattice. Omar Ocegueda, Shishir Shah 0001, Ioannis A. Kakadiaris |
CVPR | 2 |
| 2011 | Comparative evaluation of wavelet-based super-resolution from video for face recognition at a distanceabstractFace recognition is a challenging problem, especially when low resolution images or image sequences are used for the task. Many methods have been proposed that can combine multiple low resolution images to realize a higher resolution or super-resolved image. Nonetheless, their utility and limitations for use in face recognition are not well understood. In this paper, we present a quantitative and comparative evaluation of wavelet transform based methods for image super-resolution. We evaluate different basis functions, varying levels of decomposition, and multiple methods for coefficient fusion to maximize the benefit of the super-resolved image for the task of face recognition. We have used a Discrete Wavelet Transform and the shift-invariant Dual-Tree Complex Wavelet Transform. Results are reported across both manually generated datasets and data from a surveillance system. Emil Bilgazyev, Shishir Shah 0001, Ioannis A. Kakadiaris |
FG | 2 |
| 2011 | Facial component-landmark detectionabstractLandmark detection has proven to be a very challenging task in biometrics. In this paper, we address the task of facial component-landmark detection. By “component” we refer to a rectangular subregion of the face, containing an anatomical component (e.g., “eye”). We present a fully-automated system for facial component-landmark detection based on multi-resolution isotropic analysis and adaptive bag-of-words descriptors incorporated into a cascade of boosted classifiers. Specifically, first each component-landmark detector is applied independently and then the information obtained is used to make inferences for the localization of multiple components. The advantage of our approach is that it has robustness to pose as well as illumination. Our method has a failure rate lower than that of commercial software. Additionally, we demonstrate that using our method for the initialization of a point landmark detector results in performance comparable with that of state-of-the-art methods. All of our experiments are carried out using data from a publicly available database. Boris A. Efraty, Emmanuel Papadakis 0001, Adam Profitt, Shishir Shah 0001, Ioannis A. Kakadiaris |
FG | 4 |
| 2011 | 3D facial expression recognition: A perspective on promises and challengesabstractThis survey focuses on discrete expression classification and facial action unit recognition performed using 3D face data, possibly including a corresponding 2D texture image. Research trends to date are summarized and the limitations of current methods are discussed. The challenges towards the development of more accurate and automated 3D facial expression recognition methods are identified. We also call for standardized experimental protocols in order to draw fair and meaningful comparisons between different systems. Tianhong Fang, Xi Zhao 0001, Omar Ocegueda, Shishir Shah 0001, Ioannis A. Kakadiaris |
FG | 4 |
| 2011 | Improved face recognition using super-resolutionabstractFace recognition is a challenging task, especially when low-resolution images or image sequences are used. A de crease in image resolution typically results in loss of facial component details leading to a decrease in recognition rates. In this paper, we propose a new method for super resolution by first learning the high-frequency components in the facial data that can be added to a low-resolution in put image to create a super-resolved image. Our method is different from conventional methods as we estimate the high-frequency components, that are not used in other methods, to reconstruct a higher-resolution image, rather than studying the direct relationship between the high and low resolution images. Quantitative and qualitative results are reported for both synthetic and surveillance facial image databases. Emil Bilgazyev, Boris A. Efraty, Shishir Shah 0001, Ioannis A. Kakadiaris |
IJCB | 3 |
| 2011 | Facial landmark detection in uncontrolled conditionsabstractFacial landmark detection is a fundamental step for many tasks in computer vision such as expression recognition and face alignment. In this paper, we focus on the detection of landmarks under realistic scenarios that include pose, illumination and expression challenges as well as blur and low-resolution input. In our approach, an n-point shape of point-landmarks is represented as a union of simpler polygonal sub-shapes. The core idea of our method is to find the sequence of deformation parameters simultaneously for all sub-shapes that transform each point-landmark into its target landmark location. To accomplish this task, we introduce an agglomerate of fern regressors. To optimize the convergence speed and accuracy we take advantage of search localization using component-landmark detectors, multi-scale analysis and learning of point cloud dynamics. Results from extensive experiments on facial images from several challenging publicly available databases demonstrate that our method (ACFeR) can reliably detect landmarks with accuracy comparable to commercial soft ware and other state-of-the-art methods. Boris A. Efraty, Chengwei Huang, Shishir Shah 0001, Ioannis A. Kakadiaris |
IJCB | 3 |
| 2011 | UR3D-C: Linear dimensionality reduction for efficient 3D face recognitionabstractWe present a novel approach for computing a compact and highly discriminant biometric signature for 3D face recognition using linear dimensionality reduction techniques. Initially, a geometry-image representation is used to effectively resample the raw 3D data. Subsequently, a wavelet transform is applied and a biometric signature composed of 7,200 wavelet coefficients is extracted. Finally, we apply a second linear dimensionality reduction step to the wavelet coefficients using Linear Discriminant Analysis and compute a compact biometric signature. Although this biometric signature consists of just 57 coefficients, it is highly discriminant. Our approach, UR3D-C, is experimentally validated using four publicly available databases (FRGC vl, FRGC v2, Bosphorus and BU-3DFE). State-of-the-art performance is reported in all of the above databases. Omar Ocegueda, Georgios Passalis, Theoharis Theoharis, Shishir Shah 0001, Ioannis A. Kakadiaris |
IJCB | 4 |
| 2011 | Twins 3D face recognition challengeabstractExisting 3D face recognition algorithms have achieved high enough performances against public datasets like FRGC v2, that it is difficult to achieve further significant increases in recognition performance. However, the 3D TEC dataset is a more challenging dataset which consists of 3D scans of 107 pairs of twins that were acquired in a single session, with each subject having a scan of a neutral expression and a smiling expression. The combination of factors related to the facial similarity of identical twins and the variation in facial expression makes this a challenging dataset. We conduct experiments using state of the art face recognition algorithms and present the results. Our results indicate that 3D face recognition of identical twins in the presence of varying facial expressions is far from a solved problem, but that good performance is possible. Vipin Vijayan, Kevin W. Bowyer, Patrick J. Flynn, Di Huang 0001, Liming Chen 0002, Mark Hansen, Omar Ocegueda, Shishir Shah 0001, Ioannis A. Kakadiaris |
IJCB | 8 |
| 2011 | Pose invariant facial component-landmark detectionabstractFacial landmark detection has proved to be a very challenging task in biometrics due to the numerous sources of variation. In this work, we present an algorithm for robust detection of facial component-landmarks. Specifically, we address the variation due to extreme pose and illumination. To achieve robust detection for extreme poses, we use a set of independent pose and landmark specific detectors. Each component-landmark detector is applied independently and the information obtained is used to make inferences about the layout of multiple components. In addition, we incorporate a multi-view representation based on an aspect graph approach. The performance of our algorithm is assessed using data from a publicly available database. The failure rate of our method is lower than that of commercially available software. Boris A. Efraty, Emmanuel Papadakis 0001, Adam Profitt, Shishir Shah 0001, Ioannis A. Kakadiaris |
ICIP | 4 |
| 2010 | Level Set with Embedded Conditional Random Fields and Shape Priors for Segmentation of Overlapping Objects
Xuqing Wu 0001, Shishir Shah 0001 |
ACCV (2) | 2 |
| 2010 | Personalized 3D-Aided 2D Facial Landmark Localization
Zhihong Zeng, Tianhong Fang, Shishir Shah 0001, Ioannis A. Kakadiaris |
ACCV (2) | 3 |
| 2010 | Joint Modeling of Algorithm Behavior and Image Quality for Algorithm Performance PredictionabstractEstimation of an algorithm’s performance given a particular input image/video is difficult for a computer but quite easy for a human observer. Humans can assess the ability of an algorithm to generate a positive outcome for a given input based on small number of observations since they can learn high level cognitive association from input-output observations. Simulation of this process can lead to automation of algorithm performance assessment and thus eventual prediction of algorithm performance given an input image. In any computer vision system, predicting the performance of a vision algorithm can prove valuable in optimizing the overall success of the application. In this paper, we propose a framework for predicting the performance of a vision algorithm given the input image or video so as to maximize the algorithm’s ability to provide the desired output. This is achieved by modeling the performance prediction process as one that accounts for the algorithm’s behavioral properties as well as the quality of the algorithm’s input. A prototype system is designed for optimal prediction of vision algorithm’s performance given the inputs image/video quality and its application to an optimal algorithm selection process is demonstrated. This system can be considered an intelligent system that combines algorithm’s input’s quality with knowledge based prediction of algorithm performance. Performance evaluation is used to obtain knowledge about algorithm’s behavioral properties [3] and intelligent system design is used to apply that knowledge. The quality of a frame is expressed as a function or combination of functions of image features that represent degradations present in the video frames. Algorithm’s performance is characterized by performance metrics that capture its ability to provide a desired result. The framework shown in figure 1 contains three main modules, quality extractor, performance evaluator, and predictor. The quality extractor is meant to quantify the degradations in the input image that affect the vision algorithm’s performance. The performance evaluator, evaluates the algorithm’s performance on a set of training input data with varying levels of image quality in order to capture its behavioral properties. In essence, the performance evaluator simulates the algorithm’s perception of the input image. Finally, the predictor combines information from quality extractor and performance evaluator. It provides a mechanism to automatically acquire, store and utilize knowledge about algorithm’s perception of image quality. The framework is prototyped using 3 object tracking algorithms used in video surveillance applications. During the training phase, the input video is processed using each algorithm and performance evaluation on the results provides the performance metrics. The performance metric represents the quality of algorithm’s performance associated with the input. Image quality measures are also calculated for every input frame in the video. Both image quality measure and performance measures are used to design the predictor. This allows for representation and capture of needed knowledge. When the predictor encounters a new input video, the knowledge learnt during training and the image quality are used to predict each algorithm’s ability to succeed, without actual algorithm execution. The one expected to achieve maximum success is selected. The signal activity measure proposed in [4], edge entropy and structural similarity index metric (SSIM) proposed in [5] are used to quantify image quality. Multiple Object Tracker performance evaluation is used to quantify a tracker’s performance in terms of performance metrics. A systematic and objective performance evaluation of the tracker’s characteristics proposed in [2] is used for this purpose. The tracking algorithms used are Uniform Motion Connected Component tracking, Mean Shift tracking, and Particle filter tracking. In order to observe and learn the effect of the degradations in input data on the algorithm’s performance, we use 12 videos from the INRIA data set [1]. Table 1 shows the performance of each individual tracker and the performance of the tracker selection based on prediction on the 12 test sequences. Labels CC, MS and PF represent the Connected Component Figure 1: Performance Prediction Learning Framework. Apurva Gala, Shishir Shah 0001 |
BMVC | 2 |
| 2010 | A shape-driven MRF model for the segmentation of organs in medical imagesabstractIn this paper, we present a knowledge-driven Markov Random Field (MRF) model for the segmentation of organs in medical images with particular emphasis on the incorporation of shape constraints into the segmentation problem. We cast the problem of image segmentation as the Maximum A Posteriori (MAP) estimation of a Markov Random Field which, in essence, is equivalent to the minimization of the corresponding Gibbs energy function. We then incorporate a set of constraints into the Gibbs energy function that collectively force the resulting segmentation contour/surface to have a shape similar to that of a given shape template. In particular, we introduce a flux-maximization constraint and a generalized template-based star-shape constraint that are encoded into the first- and second-order clique potentials of the Gibbs energy function, respectively. Our main contribution is in the translation of a set of global notions about the shape of the desired segmentation contour into a set of local measures that can be conveniently encoded into the Gibbs energy function and used in combination with other traditionally used constraints derived from image information. In our experiments, we demonstrate the application of the proposed method to the challenging problem of heart segmentation in non-contrast computed tomography (CT) data. Deepak Roy Chittajallu, Shishir Shah 0001, Ioannis A. Kakadiaris |
CVPR | 2 |
| 2010 | Disparity Map Refinement for Video Based Scene Change Detection Using a Mobile Stereo Camera PlatformabstractThis paper presents a novel disparity map refinement method and vision based surveillance framework for the task of detecting objects of interest in dynamic outdoor environments from two stereo video sequences taken at different times and from different viewing angles by a mobile camera platform. The proposed framework includes several steps, the first of which computes disparity maps of the same scene in two video sequences. Preliminary disparity images are refined based on estimated disparities in neighboring frames. Segmentation is performed to estimate ground planes, which in turn are used for establishing spatial registration between the two video sequences. Finally, the regions of change are detected using the combination of texture and intensity gradient features. We present experiments on detection of objects of different sizes and textures in real videos. Hakan Haberdar, Shishir Shah 0001 |
ICPR | 2 |
| 2009 | 3D-aided profile-based face recognitionabstractThe silhouette of the face profile is a well-known biometric that is already in use in face recognition research. One of the challenges for successful employment of this biometric is the sensitivity of its geometry to face rotation. In this paper, we introduce a new method that improves robustness to rotation. We achieve this by exploring the feature space of profiles under various rotations with the aid of a 3D face model. Based on fiducial points on the profile silhouette, we extract a set of rotation-, translation- and scale-invariant features which are used to design and train a hierarchical pose-identity classifier. In our experiments the classifier is used for the identification of a driver using his/her side-view image. We present our results on a publicly available database. Boris A. Efraty, Dat Chu, Emil Ismailov, Shishir Shah 0001, Ioannis A. Kakadiaris |
ICIP | 4 |
| 2008 | Commentary Paper on "Person Tracking With Audio-Visual Cues Using the Iterative Decoding Framework"abstractThis paper presents an information theoretic approach to the problem of fusing information from multiple disparate sources for the problem of person tracking. Specifically, the approach is presented for the use of audio-visual cues in tracking people in indoor environments. Shishir Shah 0001 |
AVSS | 1 |
| 2008 | Comparative analysis of cell segmentation using absorption and color images in fine needle aspiration cytologyabstractSegmentation of cytological smears plays a critical role in the automated analysis of histological abnormalities by fine needle aspiration cytology. However, smears obtained from fine needle aspiration biopsy are often contaminated with blood. Segmentation of such an image is not a trivial task and the false positive rate could be high if the blood cells cannot be correctly separated from the rest of the sample. Moreover, the fine textured nature of the cell chromatin gives it a non-uniform intensity appearance in both color and gray images. In this paper, we propose an enhanced watershed approach to remove background noise by using short wavelength spectral image and the computed absorption image to improve segmentation accuracy. We also demonstrate a color image segmentation method by applying watershed to the minima imposed aggregation image. Results of segmentation on 20 images of cytological smears are presented and the accuracy compared for the two methods. Xuqing Wu 0001, Shishir Shah 0001 |
SMC | 2 |
| 2008 | Performance Modeling and Algorithm Characterization for Robust Image Segmentation
Shishir Shah 0001 |
Int. J. Comput. Vis. | 1 |
| 2007 | Segmenting Biological Particles in Multispectral Microscopy ImagesabstractThis paper presents a methodology and results for segmentation of biological particles in multispectral images by learning disparate models from each spectra for pixel classification coupled with contour evolution based on the use of level set theory. Traditional contour models have some limitations on the segmentation of complicated images whose sub-regions consist of multiple components. The segmentation of multispectral images is even a more difficult problem. Our proposed model overcomes these limitations and uses multiple classifiers, each of which solves the problem independently based on its input observations. Each classifier module is trained to detect distinct regions and a higher order decision integrator collects evidence from each of the modules to delineate a final region Shishir Shah 0001 |
WACV | 1 |
| 2005 | Multispectral Integration for Segmentation of Chromosome Images
Shishir Shah 0001 |
CAIP | 1 |
| 1999 | Hierarchical Multifeature Integration for Automatic Object Recognition in Forward Looking Infrared Images
Shishir Shah 0001, Jake K. Aggarwal |
IEA/AIE | 1 |
| 1998 | Bayesian Paradigm for Recognition of Objects - Innovative Applications
Jake K. Aggarwal, Shishir Shah 0001 |
ACCV (2) | 2 |
| 1998 | Multiple Feature Integration for Robust Object LocalizationabstractThis paper presents a methodology for localization of manmade objects in complex scenes by learning multiple feature models in images. The methodology is based on a modular structure consisting of multiple classi#ers, each of which solves the problem independently based on its input observations. Each classi- #er module is trained to detect manmade object regions and a higher order decision integrator collects evidencefrom each of the modules to delineate a #nal region of interest. The proposed framework is applied to the problem of Automatic Manmade Object Localization #Detection. Results obtained on the detection of vehicles in color visual and infrared imagery are presented in this paper. 1 Introduction This paper addresses the problem of object localization in complex scenes imaged by a single sensor or registered multiple sensors. The general topic of determining region of interest #ROI# and object detection is a critical step in all existing paradigms for object recognition. In... Shishir Shah 0001, Jake K. Aggarwal |
CVPR | 1 |
| 1998 | Partial Face Recognition Using Radial Basis Function Networks
Kiminori Sato, Shishir Shah 0001, Jake K. Aggarwal |
FG | 2 |
| 1998 | A hybrid architecture for performance reasoning in classification systemsabstractThis paper presents a unified methodology for reasoning in classification systems. The methodology is based on a two-stage structure that incorporates both neural and Bayesian formulations in the first stage and a rule-based system created by extracting rules from both the classifiers in the second stage. The rule-based system provides a measure of the cause-effect relationship between the inputs and the outputs. This is a novel and useful method for reasoning about the performance of classifier systems and for representing qualitative knowledge about the causal relationship in decision-making systems. The proposed system is tested and results are reported for the problem of automatic target detection. Shishir Shah 0001, Jake K. Aggarwal |
ICPR | 1 |
| 1998 | Robust automatic target detection/recognition system for second generation FLIR imageryabstractAutomatic target detection and recognition (ATD/R) is of crucial interest to the defense community. We present a robust ATD/R system developed at the CVRC at UT-Austin for recognition in second generation forward looking infrared (FLIR) images. An experiment conducted on 1930 FLIR images shows that this ATR system can achieve recognition with a high degree of accuracy and a low false alarm rate. This demo first presents a brief overview of the whole methodology, then shows the detailed procedures and temporary outputs step by step, by running this ATR system on typical low-contrast FLIR images. Results and examples are presented at the end of the demonstration. Huaibin Zhao, Shishir Shah 0001, Jae Hun Choi, Dinesh Nair, Dinesh K. Aggarwal |
WACV | 2 |
| 1997 | A Bayesian Segmentation Framework for Textured Visual ImagesabstractThis paper presents a new framework for segmentation of textured visual imagery. The proposed method consists of a Bayesian formulation for labeling similar regions. Similarity is defined via texture features obtained by Gabor Wavelets. Multivariate Gaussian distributions are employed to model the feature class-conditional densities, while the Markov process is used to characterize the distributions of the region labeling due to each feature. A coarse nearest neighbor clustering is performed over the feature space to estimate the initial labelings. An iterative solution to the Maximum A Posteriori (MAP) estimation is developed, where the parameters of the prior distribution of region labels are estimated using the Expectation-Maximization (EM) algorithm. Finally, for man-made object segmentation, a region-growing procedure is used to analyze the classified texture regions by incorporating measures of local shape characteristics to obtain smooth boundaries and region homogeneity. Results of the developed algorithm on real scene images are presented. Shishir Shah 0001, Jake K. Aggarwal |
CVPR | 1 |
| 1997 | Multisensor Integration for Scene Classifiction: An Experiment in Human Form DetectionabstractThis paper presents a system for classification of scenes using a multisensor integration framework. Indoor scenes are imaged using a visual and an infrared sensor and the images processed in three stages to perform classification of sensed objects into two classes: human and background. Finally, information from individual classifiers is integrated in order to obtain an improved classification performance. Details of feature extraction and classification using neural network combining a multi-Bayesian framework are presented. Segmentation of the imaged scene is performed using existing techniques such as texture analysis and histogram modeling. Classification results on real-world data are presented. The system represents a first step in the development of improved, robust classifiers based on the concepts of neural networks and multisensor integration. Shishir Shah 0001, Jake K. Aggarwal, Jayan Eledath, Joydeep Ghosh |
ICIP (2) | 1 |
| 1997 | Mobile robot navigation and scene modeling using stereo fish-eye lens system
Shishir Shah 0001, Jake K. Aggarwal |
Mach. Vis. Appl. | 1 |
| 1996 | Intrinsic parameter calibration procedure for a (high-distortion) fish-eye lens camera with distortion model and accuracy estimation
Shishir Shah 0001, Jake K. Aggarwal |
Pattern Recognit. | 1 |
| 1995 | Modeling Structured Environments Using Robot Vision
Shishir Shah 0001, Jake K. Aggarwal |
ACCV | 1 |
| 1994 | Depth Estimation using Stereo Fish-Eye LensesabstractThis paper presents the estimation of depth in an indoor, structured environment based on a stereo setup consisting of two fish-eye lenses, with parallel optical axes, mounted on a robot platform. The use of fish-eye lenses provides for a large field of view to estimate better the depth of features very close to the lens. To extract significant information from the fish-eye lens images, we first correct for the distortion before using a special line detector, based on vanishing points, to extract significant features. We use a relaxation procedure to achieve correspondence between features in the left and right images. The process of prediction and recursive verification of the hypotheses is utilized to find a one-to-one correspondence. Experimental results obtained on several stereo images are presented, and an accuracy analysis is performed. Further, the algorithm is tested using a pair of wide-angle lenses, and the accuracy and difference in the spatial information obtained are compared.> Shishir Shah 0001, Jake K. Aggarwal |
ICIP (2) | 1 |
| 1994 | A Simple Calibration Procedure for Fish-Eye (High-Distortion) Lens CameraabstractPresents a new algorithm for the geometric camera calibration of a fish-eye lens (a high distortion lens) mounted on a CCD TV camera. The algorithm determines a mapping between points in the world coordinate system and their corresponding point locations in the image plane. The parameters to be calibrated are effective focal length, one-pixel width on the image plane, image distortion center, and distortion coefficients. A simple calibration pattern consisting of equally spaced dots is introduced as a reference for calibration. Some parameters to be calibrated are eliminated by setting up the calibration pattern precisely and assuming negligible distortion at the image distortion center. Thus, the number of unknown parameters to be calibrated is drastically reduced, enabling simple and useful calibration. The method employs a polynomial transformation between points in the world coordinate system and their corresponding image plane locations. The coefficients of the polynomial are determined using the Lagrangian estimation. Furthermore, the effectiveness of the proposed calibration method is confirmed by experimentation.> Shishir Shah 0001, Jake K. Aggarwal |
ICRA | 1 |