Bir Bhanu

dblp:b/BirBhanu · DBLP profile ↗
← Back
279ranked-venue papers
45as first author
34since 2021 · last 2026
0000-0001-8971-6416ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 165 · 35 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 148 · 14 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 4 since 2021Human-computer interaction and ubiquitous computing · 14 · 2 first-author · 1 since 2021Systems, architecture and hardware · 8 · 5 first-authorSecurity and privacy · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3Computer networks · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Reliable image super-resolution using dual-teacher knowledge distillation
Zhan Li 0004, Weijun Yuan, Boyang Yao, Yihang Chen 0005, Bir Bhanu, Kehuan Zhang
Knowl. Based Syst.5
2026 Lightweight and effective crowd counting and localization in diverse low-visibility conditions
Boyang Yao, Zhan Li 0004, Weijun Yuan, Bir Bhanu, Ruikai Ke, Zhanglu Chen
Pattern Recognit.4
2026 TSMnet: Two-step separation pipeline based on threshold shrinkage memory network for weakly-supervised video anomaly detection
Qun Li 0002, Xinping Gao, Bir Bhanu
Pattern Recognit. Lett.4
2026 PE-ViT: Parameter-efficient vision transformer with dimension-adaptive experts and economical attention
Qun Li 0002, Jiru He, Tiancheng Guo, Xinping Gao, Bir Bhanu
Pattern Recognit. Lett.5
2026 Latent Diffusion-Guided Feature Inpainting for Occluded Person Re-Identification With Hybrid Re-Ranking
abstract
Occlusion remains a persistent challenge in person re-identification (ReID). Existing approaches either rely on human pose estimation and semantic parsing, or attempt to make models resistant to occlusion through real or synthetic occlusion augmentations. While these methods suppress the effect of occlusion, they do notsolvethe underlying problem, since the corrupted feature representation itself remains occluded. To directly address this gap, we formulate occlusion as afeature-level distortionand propose aLatent Diffusion guided De-Occluder (DDO)that learns to inpaint corrupted feature embeddings and reconstruct clean, identity-preserving representations in the latent space. By providing the downstream ReID model with occlusion-free priors, our method eliminates the need for the backbone to model occlusion explicitly, leading to inherently more robust retrieval. Furthermore, recognizing that person ReID is conventionally a closed-set problem where gallery identities are known at test time, we introduce aHybrid Re-Ranking (HRR)scheme. HRR directly addresses the limitations of standard re-ranking by leveraging centroid-based identity anchors to refine k-reciprocal re-ranking, thus boosting retrieval precision while suppressing noise. Extensive experiments on standard and occlusion-focused benchmarks confirm that our approach not only overcomes the shortcomings of existing occlusion handling strategies but also achieves new state-of-the-art performance, validating the effectiveness of de-occluding features rather than merely resisting occlusion.
Pratyay Dutta, Bir Bhanu
IEEE Trans. Circuits Syst. Video Technol.2
2026 Toward Generative Understanding: Incremental Few-Shot Semantic Segmentation With Diffusion Models
abstract
Incremental Few-shot Semantic Segmentation (iFSS) aims to learn novel classes with limited samples while preserving segmentation capability for base classes, addressing the challenge of continual learning of novel classes and catastrophic forgetting of previously seen classes. Existing methods mainly rely on techniques such as knowledge distillation and background learning, which, while partially effective, still suffer from issues such as feature drift and limited generalization to real-world novel classes, primarily due to a bidirectional coupling bottleneck between the learning of base classes and novel classes. To address these challenges, we propose, for the first time, a diffusion-based generative framework for iFSS. Specifically, we bridge the gap between generative and discriminative tasks through an innovative binary-to-RGB mask mapping mechanism, enabling pre-trained diffusion models to focus on target regions via class-specific semantic embedding optimization while sharpening foreground-background contrast with color embeddings. A lightweight post-processor then refines the generated images into high-quality binary masks. Crucially, by leveraging diffusion priors, our framework avoids complex training strategies. The optimization of class-specific semantic embeddings decouples the embedding spaces of base and novel classes, inherently preventing feature drift, mitigating catastrophic forgetting, and enabling rapid novel-class adaptation. Experimental results show that our method achieves state-of-the-art performance on the PASCAL- $5^{i}$ and COCO- $20^{i}$ datasets using much less data than other methods, and exhibiting competitive results in cross-domain few-shot segmentation tasks. Project page: https://ifss-diff.github.io/.
Qun Li 0002, Fu Xiao 0001, Na Zhao 0004, Bir Bhanu
IEEE Trans. Image Process.5
2025 PS-CoT-Adapter: adapting plan-and-solve chain-of-thought for ScienceQA
Qun Li 0002, Fu Xiao 0001, Yiming Wang 0007, Xinping Gao, Bir Bhanu
Sci. China Inf. Sci.6
2025 Spatial-temporal multi-scale interaction for few-shot video summarization
Qun Li 0002, Zhuxi Zhan, Yanchao Li 0001, Bir Bhanu
Eng. Appl. Artif. Intell.4
2025 Masked Graph Attention network for classification of facial micro-expression
Ankith Jain Rakesh Kumar, Bir Bhanu
Image Vis. Comput.2
2025 Low-Visibility Scene Enhancement by Isomorphic Dual-Branch Framework With Attention Learning
abstract
Vision-based intelligent systems are extensively used in autonomous driving, traffic monitoring, and transportation surveillance due to their high performance, low cost, and ease of installation. However, their effectiveness is often compromised by adverse conditions such as haze, fog, low light, motion blur, and low resolution, leading to reduced visibility and increased safety risks. Additionally, the prevalence of high-definition imaging in embedded and mobile devices presents challenges related to the conflict between large image sizes and limited computing resources. To address these issues and enhance visual perception for intelligent systems operating under adverse conditions, this study proposes an all-in-one isomorphic dual-branch (IDB) framework consisting of two branches with identical structures for different functions, a loss-attention (LA) learning strategy, and feature fusion super-resolution (FFSR) module. The versatile IDB network employs a simple and effective encoder-decoder structure as the backbone for both branches, which can be replaced with task-specific tailored backbones. The plug-in LA strategy differentiates the functions of the two branches, adapting them to various tasks without increasing computational demands during inference. The FFSR module concatenates multi-scale features and restores details progressively in downsampled images, producing outputs with improved visibility, brightness, edge sharpness, and color fidelity. Extensive experimental results demonstrate that the proposed framework outperforms several state-of-the-art methods for image dehazing, low-light enhancement, image deblurring, and super-resolution image reconstruction while maintaining low computational overhead. The associated code is publicly available athttps://github.com/lizhangray/IDBall.
Zhan Li 0004, Wenqing Kuang, Bir Bhanu, Yihang Chen 0005
IEEE Trans. Intell. Transp. Syst.3
2024 Co-GZSL: Feature Contrastive Optimization for Generalized Zero-Shot Learning
abstract
Abstract Generalized Zero-Shot Learning (GZSL) learns from only labeled seen classes during training but discriminates both seen and unseen classes during testing. In GZSL tasks, most of the existing methods commonly utilize visual and semantic features for training. Due to the lack of visual features for unseen classes, recent works generate real-like visual features by using semantic features. However, the synthesized features in the original feature space lack discriminative information. It is important that the synthesized visual features should be similar to the ones in the same class, but different from the other classes. One way to solve this problem is to introduce the embedding space after generating visual features. Following this situation, the embedded features from the embedding space can be inconsistent with the original semantic features. For another way, some recent methods constrain the representation by reconstructing the semantic features using the original visual features and the synthesized visual features. In this paper, we propose a hybrid GZSL model, named feature Contrastive optimization for GZSL (Co-GZSL), to reconstruct the semantic features from the embedded features, which ensures that the embedded features are close to the original semantic features indirectly by comparing reconstructed semantic features with original semantic features. In addition, to settle the problem that the synthesized features lack discrimination and semantic consistency, we introduce a Feature Contrastive Optimization Module (FCOM) and jointly utilize contrastive and semantic cycle-consistency losses in the FCOM to strengthen the intra-class compactness and the inter-class separability and to encourage the model to generate semantically consistent and discriminative visual features. By combining the generative module, the embedding module, and the FCOM, we achieve Co-GZSL. We evaluate the proposed Co-GZSL model on four benchmarks, and the experimental results indicate that our model is superior over current methods. Code is available at: https://github.com/zhanzhuxi/Co-GZSL .
Qun Li 0002, Zhuxi Zhan, Yaying Shen, Bir Bhanu
Neural Process. Lett.4
2024 RepSGG: Novel Representations of Entities and Relationships for Scene Graph Generation
abstract
Scene Graph Generation (SGG) has achieved significant progress recently. However, most previous works rely heavily on fixed-size entity representations based on bounding box proposals, anchors, or learnable queries. As each representation's cardinality has different trade-offs between performance and computation overhead, extracting highly representative features efficiently and dynamically is both challenging and crucial for SGG. In this work, a novel architecture called RepSGG is proposed to address the aforementioned challenges, formulating a subject as queries, an object as keys, and their relationship as the maximum attention weight between pairwise queries and keys. With more fine-grained and flexible representation power for entities and relationships, RepSGG learns to sample semantically discriminative and representative points for relationship inference. Moreover, the long-tailed distribution also poses a significant challenge for generalization of SGG. A run-time performance-guided logit adjustment (PGLA) strategy is proposed such that the relationship logits are modified via affine transformations based on run-time performance during training. This strategy encourages a more balanced performance between dominant and rare classes. Experimental results show that RepSGG achieves the state-of-the-art or comparable performance on the Visual Genome and Open Images V6 datasets with fast inference speed, demonstrating the efficacy and efficiency of the proposed methods.
Hengyue Liu, Bir Bhanu
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Energy-Motion Features Aggregation Network for Players' Fine-Grained Action Analysis in Soccer Videos
abstract
Rich and complex events in sports have led to the development of a wide-variety of techniques for interpreting content of sports videos in terms of players’ actions, poses, gait, performance, etc. This is due to the requirements from coaches, trainers and players who expect to analyze actions in top sports events, as well as sports fans who practice to imitate professional playing skills, e.g., dribbling, shooting, etc. However, this poses two key challenges for automated sports analysis community. Firstly, there are extremely limited public sports datasets. Secondly, recent advances in interpretations of sports activities, e.g., soccer, are predominantly made through analyzing coarse-grained contents. Players’ fine-grained skills analysis still remains under-explored. To alleviate these problems, this paper (a) collects the dataset of highlight videos of soccer players, including two coarse-grained action types of soccer players and six fine-grained actions of players. Detailed annotations are provided for the collected dataset, in terms of action classes, bounding boxes, segmentation maps, and body keypoints of soccer players, and positions of a soccer ball in a game. (b) leverages the understanding of complex highlight videos by proposing an energy-motion features aggregation network-EMANet to fully exploit energy-based representation of soccer players movements in video sequences and explicit motion dynamics of soccer players in videos for soccer players’ fine-grained action analysis. Experimental results and ablation studies validate the proposed approach in recognizing soccer players actions using the collected soccer highlight video datasets.
Runze Li 0003, Bir Bhanu
IEEE Trans. Circuits Syst. Video Technol.2
2024 Searching for Life: End-to-End Automated Detection and Characterization of Ediacaran Biosignatures
abstract
With state-of-the-art imaging and analytical tools on board the NASA Perseverance Rover mission, geological information for remote astrobiological analysis is readily available and more widespread than ever. For analysis of such data, there is a need for automated, remote assessment tools, capable of translational and scalable research. To search for evidence of life in the universe, one of the goals of astrobiological research, we must first recognize robust indicators of life on Earth. One method is the study of Earth’s early fossil record, dominated by simple organisms similar to that predicted to have existed on Mars - if life ever did develop. We identify a potential morphological biosignature, known as double-rippled bedforms (DRBs), from the Ediacaran Member (Rawnsley Quartzite, 550-555 Ma) of South Australia. We utilize these DRBs to develop a new tool for astrobiological investigation via an improved end-to-end remote, automated biosignature assessment tool, “Scene-aware Perception Automation using Composite Embedding for Segmentation" (SPACESeg2.0). The proposed tool accurately detects and quantitatively characterizes DRBs in imagery from varied fields-of-view, with practical artifact-ridden data, even when the DRBs constitute less than 1% of the total image. We establish the efficacy of SPACESeg2.0 in detecting desired structures compared with other deep-learning techniques.
Padmaja Jonnalagedda, Rachel Surprenant, Mary Droser, Bir Bhanu
IEEE Trans. Geosci. Remote. Sens.4
2023 ESSL: Enhanced Spatio-Temporal Self-Selective Learning Framework for Unsupervised Video Anomaly Detection
abstract
Unsupervised Video Anomaly Detection (UVAD) utilizes completely unlabeled videos for training without any human intervention. Due to the existence of unlabeled abnormal videos in the training data, the performance of UVAD has a large gap compared with semi-supervised VAD, which only uses normal videos for training. To address the problem of insufficient ability of the existing UVAD methods to learn normality and reduce the negative impact of abnormal events, this paper proposes a novel Enhanced Spatio-temporal Self-selective Learning (ESSL) framework for UVAD. This framework is designed for capturing both the appearance and motion features through effective network structures by solving the spatial and temporal jigsaw puzzles. Specially, we develop a Self-selective Learning Module (SLM) for UVAD, which prevents the model learning abnormal features and enhances the model by selecting normal features. Experimental results on three benchmark datasets show that the proposed method not only surpasses the state-of-the-art UVAD works, but also achieves the performance comparable to the classic semi-supervised methods for video anomaly detection that needs normal videos selected manually. Code is available at: https://github.com/xusuger/ESSL.
Qun Li 0002, Xubei Pan, Fu Xiao 0001, Bir Bhanu
ECAI4
2023 A Depth-Guided Attention Strategy for Crowd Counting
Zhan Li 0004, Bir Bhanu, Dongping Lu, Xuming Han
ICANN (10)3
2023 Novel Body Biometric for Long-Range Recognition Under Extreme Conditions
abstract
The task of video-based human recognition is complicated by many factors such as imaging distortions, imaging range, lack of frames, arbitrary pose, occlusions, air turbulence, and changing clothes. This work presents the first study that utilizes single-frame binary silhouettes and their auxiliary representations for human recognition under extreme distortions. The proposed representation is compact, modular, and robust to distortions, allowing for easy deployability for long-range recognition. Quantitative metrics are reported on long-range dataset such as Briar, demonstrating the robustness of the proposed approach to common challenges of video-based recognition in the wild. The proposed single-frame method is compared against gait techniques using limited frames, outperforming in most cases. Performance is also compared to grayscale images with varying ranges, environments, and changing clothes, where the proposed model outperforms grayscale images. Under consistent conditions, the proposed model still augments the performance of the baseline grayscale model by over 15%.
Padmaja Jonnalagedda, Bir Bhanu
IJCB2
2023 Lite-FENet: Lightweight multi-scale feature enrichment network for few-shot segmentation
Qun Li 0002, Baoquan Sun, Bir Bhanu
Knowl. Based Syst.3
2023 Triplet-Net Classification of Contiguous Stem Cell Microscopy Images
abstract
Cellular microscopy imaging is a common form of data acquisition for biological experimentation. Observation of gray-level morphological features allows for the inference of useful biological information such as cellular health and growth status. Cellular colonies can contain multiple cell types, making colony level classification very difficult. Additionally, cell types growing in a hierarchical, downstream fashion, can often look visually similar, although biologically distinct. In this paper, it is determined empirically that traditional deep Convolutional Neural Networks (CNN) and classical object recognition techniques are not sufficient to distinguish between these subtle visual differences, resulting in misclassifications. Instead, Triplet-net CNN learning is employed in a hierarchical classification scheme to improve the ability of the model to discern distinct, fine-grain features of two commonly confused morphological image-patch classes, namely Dense and Spread colonies. The Triplet-net method improves classification accuracy over a four-class deep neural network by ∼ 3 %, a value that was determined to be statistically significant, as well as existing state-of-the-art image patch classification approaches and standard template matching. These findings allow for the accurate classification of multi-class cell colonies with contiguous boundaries, and increased reliability and efficiency of automated, high-throughput experimental quantification using non-invasive microscopy.
Adam Witmer, Rajkumar Theagarajan, Bir Bhanu
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 MonoIndoor++: Towards Better Practice of Self-Supervised Monocular Depth Estimation for Indoor Environments
abstract
Self-supervised monocular depth estimation has seen significant progress in recent years, especially in outdoor environments, i.e., autonomous driving scenes. However, depth prediction results are not satisfying in indoor scenes where most of the existing data are captured with hand-held devices. As compared to outdoor environments, estimating depth of monocular videos for indoor environments, using self-supervised methods, results in two additional challenges: (i) the depth range of indoor video sequences varies a lot across different frames, making it difficult for the depth network to induce consistent depth cues for training, whereas the maximum distance in outdoor scenes mostly stays the same as the camera usually sees the sky; (ii) the indoor sequences recorded with handheld devices often contain much more rotational motions, which cause difficulties for the pose network to predict accurate relative camera poses, while the motions of outdoor sequences are pre-dominantly translational, especially for street-scene driving datasets such as KITTI. In this work, we propose a novel framework-MonoIndoor++ by giving special considerations to those challenges and consolidating a set of good practices for improving the performance of self-supervised monocular depth estimation for indoor environments. First, a depth factorization module with transformer-based scale regression network is proposed to estimate a global depth scale factor explicitly, and the predicted scale factor can indicate the maximum depth values. Second, rather than using a single-stage pose estimation strategy as in previous methods, we propose to utilize a residual pose estimation module to estimate relative camera poses across consecutive frames iteratively. Third, to incorporate extensive coordinates guidance for our residual pose estimation module, we propose to perform coordinate convolutional encoding directly over the inputs to pose networks. The proposed method is validated on a variety of benchmark indoor datasets, i.e., EuRoC MAV, NYUv2, ScanNet and 7-Scenes, demonstrating the state-of-the-art performance. In addition, the effectiveness of each module is shown through a carefully conducted ablation study and the good generalization and universality of our trained model is also demonstrated, specifically on ScanNet and 7-Scenes datasets.
Runze Li 0003, Pan Ji, Yi Xu 0002, Bir Bhanu
IEEE Trans. Circuits Syst. Video Technol.4
2022 MTKDSR: Multi-Teacher Knowledge Distillation for Super Resolution Image Reconstruction
abstract
In recent years, the performance of single image super-resolution (SISR) methods based on deep neural networks has significantly improved. However, large model sizes and high computational costs are common problems for most SR networks. Meanwhile, a trade-off exists between higher reconstruction fidelity and improved perceptual quality in solving the SISR problem. In this paper, we propose a multi-teacher knowledge distillation approach for SR (MTKDSR) tasks that can train a balanced, lightweight, and efficient student network using different types of teacher models that are proficient in terms of reconstruction fidelity or perceptual quality. In addition, to generate more realistic and learnable textures, we propose an edge-guided SR network, EdgeSRN, as a perceptual teacher used in the MTKDSR framework. In our experiments, EdgeSRN was superior to the models based on adversarial learning in terms of the ability of effective knowledge transfer. Extensive experiments show that the student trained by MTKDSR exhibit superior performance compared to those of state-of-the-art lightweight SR networks in terms of perceptual quality with a smaller model size and fewer computations. Our code is available at https://github.com/lizhangray/MTKDSR.
Gengqi Yao, Zhan Li 0004, Bir Bhanu, Zhiqing Kang, Ziyi Zhong
ICPR3
2022 Dite-HRNet: Dynamic Lightweight High-Resolution Network for Human Pose Estimation
abstract
A high-resolution network exhibits remarkable capability in extracting multi-scale features for human pose estimation, but fails to capture long-range interactions between joints and has high computational complexity. To address these problems, we present a Dynamic lightweight High-Resolution Network (Dite-HRNet), which can efficiently extract multi-scale contextual information and model long-range spatial dependency for human pose estimation. Specifically, we propose two methods, dynamic split convolution and adaptive context modeling, and embed them into two novel lightweight blocks, which are named dynamic multi-scale context block and dynamic global context block. These two blocks, as the basic component units of our Dite-HRNet, are specially designed for the high-resolution networks to make full use of the parallel multi-resolution architecture. Experimental results show that the proposed network achieves superior performance on both COCO and MPII human pose estimation datasets, surpassing the state-of-the-art lightweight networks. Code is available at: https://github.com/ZiyiZhang27/Dite-HRNet.
Qun Li 0002, Ziyi Zhang 0001, Fu Xiao 0001, Bir Bhanu
IJCAI5
2022 Attention-based anomaly detection in multi-view surveillance videos
Qun Li 0002, Fu Xiao 0001, Bir Bhanu
Knowl. Based Syst.4
2022 Dynamically throttleable neural networks
Hengyue Liu, Samyak Parajuli, Jesse Hostetler, Sek M. Chai, Bir Bhanu
Mach. Vis. Appl.5
2022 Privacy Preserving Defense For Black Box Classifiers Against On-Line Adversarial Attacks
abstract
Deep learning models have been shown to be vulnerable to adversarial attacks. Adversarial attacks are imperceptible perturbations added to an image such that the deep learning model misclassifies the image with a high confidence. Existing adversarial defenses validate their performance using only the classification accuracy. However, classification accuracy by itself is not a reliable metric to determine if the resulting image is "adversarial-free". This is a foundational problem for online image recognition applications where the ground-truth of the incoming image is not known and hence we cannot compute the accuracy of the classifier or validate if the image is "adversarial-free" or not. This paper proposes a novel privacy preserving framework for defending Black box classifiers from adversarial attacks using an ensemble of iterative adversarial image purifiers whose performance is continuously validated in a loop using Bayesian uncertainties. The proposed approach can convert a single-step black box adversarial defense into an iterative defense and proposes three novel privacy preserving Knowledge Distillation (KD) approaches that use prior meta-information from various datasets to mimic the performance of the Black box classifier. Additionally, this paper proves the existence of an optimal distribution for the purified images that can reach a theoretical lower bound, beyond which the image can no longer be purified. Experimental results on six public benchmark datasets namely: 1) Fashion-MNIST, 2) CIFAR-10, 3) GTSRB, 4) MIO-TCD, 5) Tiny-ImageNet, and 6) MS-Celeb show that the proposed approach can consistently detect adversarial examples and purify or reject them against a variety of adversarial attacks.
Rajkumar Theagarajan, Bir Bhanu
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 JEDE: Universal Jersey Number Detector for Sports
abstract
The rapid progress in deep learning-based computer vision has opened unprecedented possibilities in computing various high-level analytics for sports. Artificial intelligence techniques such as predictive analysis, automatic highlight generation, and assistant coaching have been applied to improve performance and decision-making for teams and players. To perform any high-level analysis from a game match, collecting the locations (where) and identities (who) of players is crucial and challenging. In this paper, a universal JErsey number DEtector (JEDE) for player identification is presented that predicts players’ bounding boxes and keypoints, along with bounding boxes and classes of jersey digits and numbers in an end-to-end manner. Instead of generating digit proposals from pre-defined anchors, JEDE predicts more robust proposals guided by players’ features and pose estimation. Moreover, a dataset is collected from soccer and basketball matches with annotations on players’ bounding boxes and body keypoints, and jersey digits’ bounding boxes and labels. Extensive experimental results and ablation studies on the collected dataset show that the proposed method outperforms the state-of-the-art methods by a large margin. Both quantitative and qualitative results also demonstrate JEDE’s superior practicality and generalizability over different sports.
Hengyue Liu, Bir Bhanu
IEEE Trans. Circuits Syst. Video Technol.2
2022 Data Assimilation Network for Generalizable Person Re-Identification
abstract
In this paper, a data assimilation network is proposed to tackle the challenges of domain generalization for person re-identification (ReID). Most of the existing research efforts only focus on single-dataset issues, and the trained models are difficult to generalize to unseen scenarios. This paper presents a distinctive idea to improve the generality of the model by assimilating three types of images: style-variant images, misaligned images and unlabeled images. The latter two are often ignored in the previous domain generalization ReID studies. In this paper, a non-local convolutional block attention module is designed for assimilating the misaligned images, and an attention adversary network is introduced to correct it. A progressive augmented memory is designed for assimilating the unlabeled images by progressive learning. Moreover, we propose an attention adversary difference loss for attention correction, and a labeling-guide discriminative embedding loss for progressive learning. Rather than designing a specific feature extractor that is robust to style shift as in most previous domain generalization work, we propose a data assimilation meta-learning procedure to train the proposed network, so that it learns to assimilate style-variant images. It is worth mentioning that we add an unlabeled augmented dataset to the source domain to tackle the domain generalization ReID tasks. Extensive experiments demonstrate that our approach significantly outperforms the state-of-the-art domain generalization methods.
Yixiu Liu, Yunzhou Zhang, Bir Bhanu, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Circuits Syst. Video Technol.3
2022 Inner Knowledge-based Img2Doc Scheme for Visual Question Answering
abstract
Visual Question Answering (VQA) is a research topic of significant interest at the intersection of computer vision and natural language understanding. Recent research indicates that attributes and knowledge can effectively improve performance for both image captioning and VQA. In this article, an inner knowledge-based Img2Doc algorithm for VQA is presented. The inner knowledge is characterized as the inner attribute relationship in visual images. In addition to using an attribute network for inner knowledge-based image representation, VQA scheme is associated with a question-guided Doc2Vec method for question–answering. The attribute network generates inner knowledge-based features for visual images, while a novel question-guided Doc2Vec method aims at converting natural language text to vector features. After the vector features are extracted, they are combined with visual image features into a classifier to provide an answer. Based on our model, the VQA problem is resolved by textual question answering. The experimental results demonstrate that the proposed method achieves superior performance on multiple benchmark datasets.
Qun Li 0002, Fu Xiao 0001, Bir Bhanu, Biyun Sheng, Richang Hong
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Learning Local Recurrent Models for Human Mesh Recovery
abstract
We consider the problem of estimating frame-level full human body meshes given a video of a person with natural motion dynamics. While much progress in this field has been in single image-based mesh estimation, there has been a recent uptick in efforts to infer mesh dynamics from video given its role in alleviating issues such as depth ambiguity and occlusions. However, a key limitation of existing work is the assumption that all the observed motion dynamics can be modeled using one dynamical/recurrent model. While this may work well in cases with relatively simplistic dynamics, inference with in-the-wild videos presents many challenges. In particular, it is typically the case that different body parts of a person undergo different dynamics in the video, e.g., legs may move in a way that may be dynamically different from hands (e.g., a person dancing). To address these issues, we present a new method for video mesh recovery that divides the human mesh into several local parts following the standard skeletal model. We then model the dynamics of each local part with separate recurrent models, with each model conditioned appropriately based on the known kinematic structure of the human body. This results in a structure-informed local recurrent learning architecture that can be trained in an end-to-end fashion with available annotations. We conduct a variety of experiments on standard video mesh recovery benchmark datasets such as Human3.6M, MPI-INF-3DHP, and 3DPW, demonstrating the efficacy of our design of modeling local dynamics as well as establishing state-of-the-art results based on standard evaluation metrics.
Runze Li 0003, Srikrishna Karanam, Terrence Chen, Bir Bhanu, Ziyan Wu 0001
3DV5
2021 Fully Convolutional Scene Graph Generation
abstract
This paper presents a fully convolutional scene graph generation (FCSGG) model that detects objects and relations simultaneously. Most of the scene graph generation frameworks use a pre-trained two-stage object detector, like Faster R-CNN, and build scene graphs using bounding box features. Such pipeline usually has a large number of parameters and low inference speed. Unlike these approaches, FCSGG is a conceptually elegant and efficient bottom-up approach that encodes objects as bounding box center points, and relationships as 2D vector fields which are named as Relation Affinity Fields (RAFs). RAFs encode both semantic and spatial features, and explicitly represent the relationship between a pair of objects by the integral on a sub-region that points from subject to object. FCSGG only utilizes visual features and still generates strong results for scene graph generation. Comprehensive experiments on the Visual Genome dataset demonstrate the efficacy, efficiency, and generalizability of the proposed method. FCSGG achieves highly competitive results on recall and zeroshot recall with significantly reduced inference time.
Hengyue Liu, Masood S. Mortazavi, Bir Bhanu
CVPR4
2021 MonoIndoor: Towards Good Practice of Self-Supervised Monocular Depth Estimation for Indoor Environments
abstract
Self-supervised depth estimation for indoor environments is more challenging than its outdoor counterpart in at least the following two aspects: (i) the depth range of indoor sequences varies a lot across different frames, making it difficult for the depth network to induce consistent depth cues, whereas the maximum distance in outdoor scenes mostly stays the same as the camera usually sees the sky; (ii) the indoor sequences contain much more rotational motions, which cause difficulties for the pose network, while the motions of outdoor sequences are pre-dominantly translational, especially for driving datasets such as KITTI. In this paper, special considerations are given to those challenges and a set of good practices are consolidated for improving the performance of self-supervised monocular depth estimation in indoor environments. The proposed method mainly consists of two novel modules, i.e., a depth factorization module and a residual pose estimation module, each of which is designed to respectively tackle the aforementioned challenges. The effectiveness of each module is shown through a carefully conducted ablation study and the demonstration of the state-of-the-art performance on three indoor datasets, i.e., EuRoC, NYUv2 and 7-Scenes.
Pan Ji, Runze Li 0003, Bir Bhanu, Yi Xu 0002
ICCV3
2021 Ada-VSR: Adaptive Video Super-Resolution with Meta-Learning
abstract
Most of the existing works in supervised spatio-temporal video super-resolution (STVSR) heavily rely on a large-scale external dataset consisting of paired low-resolution low-frame rate (LR-LFR) and high-resolution high-frame-rate (HR-HFR) videos. Despite their remarkable performance, these methods make a prior assumption that the low-resolution video is obtained by down-scaling the high-resolution video using a known degradation kernel, which does not hold in practical settings. Another problem with these methods is that they cannot exploit instance-specific internal information of a video at testing time. Recently, deep internal learning approaches have gained attention due to their ability to utilize the instance-specific statistics of a video. However, these methods have a large inference time as they require thousands of gradient updates to learn the intrinsic structure of the data. In this work, we present Adaptive VideoSuper-Resolution (Ada-VSR) which leverages external, as well as internal, information through meta-transfer learning and internal learning, respectively. Specifically, meta-learning is employed to obtain adaptive parameters, using a large-scale external dataset, that can adapt quickly to the novel condition (degradation model) of the given test video during the internal learning task, thereby exploiting external and internal information of a video for super-resolution. The model trained using our approach can quickly adapt to a specific video condition with only a few gradient updates, which reduces the inference time significantly. Extensive experiments on standard datasets demonstrate that our method performs favorably against various state-of-the-art approaches.
Akash Gupta 0001, Padmaja Jonnalagedda, Bir Bhanu, Amit K. Roy-Chowdhury
ACM Multimedia3
2021 Multi-level cross-view consistent feature learning for person re-identification
Yixiu Liu, Yunzhou Zhang, Bir Bhanu, Sonya A. Coleman, Dermot Kerr
Neurocomputing3
2021 An Automated System for Generating Tactical Performance Statistics for Individual Soccer Players From Videos
abstract
The world of sports intrinsically involves fast and complex events that are difficult for coaches, trainers and players to analyze, and also for audiences to follow. In fast paced team sports such as soccer, keeping track of all the players and analyzing their performance after every match are very challenging. Current scenarios for identifying the best talents in soccer involve word-of-mouth and coaches/recruiters scouring through hours of manually annotated videos. This is a very expensive and laborious process and also biased by the nature of the recruiters. To alleviate these problems, this paper proposes an automated system that can detect, track, classify the teams of multiple players and identify the player controlling the ball in a video. The system generates three very important tactical statistics for a player: 1) duration of ball possession, 2) number of successful passes and 3) number of successful steals. This is done by training Convolutional Neural Networks (CNNs) to (a) localize and track the players on the field, (b) classify the team of a detected player, (c) identify the player controlling the ball and (d) pooling all the information extracted from (a), (b), and (c) to generate the statistics of players. To overcome the problem that the features learned from specific soccer matches do not necessarily generalize across different soccer matches, the paper proposes minimal amount of match-specific annotation and data augmentation, using a variant of Deep Convolutional Generative Adversarial Networks (DCGAN) to improve the accuracy. Experimental results and ablation studies show that the proposed approach outperforms the state-of-the-art approaches in terms of accuracy and processing speed.
Rajkumar Theagarajan, Bir Bhanu
IEEE Trans. Circuits Syst. Video Technol.2
2020 Towards Visually Explaining Variational Autoencoders
abstract
Recent advances in Convolutional Neural Network (CNN) model interpretability have led to impressive progress in visualizing and understanding model predictions. In particular, gradient-based visual attention methods have driven much recent effort in using visual attention maps as a means for visual explanations. A key problem, however, is these methods are designed for classification and categorization tasks, and their extension to explaining generative models, e.g., variational autoencoders (VAE) is not trivial. In this work, we take a step towards bridging this crucial gap, proposing the first technique to visually explain VAEs by means of gradient-based attention. We present methods to generate visual attention from the learned latent space, and also demonstrate such attention explanations serve more than just explaining VAE predictions. We show how these attention maps can be used to localize anomalies in images, demonstrating state-of-the-art performance on the MVTec-AD dataset. We also show how they can be infused into model training, helping bootstrap the VAE into learning improved latent space disentanglement, demonstrated on the Dsprites dataset.
WenQian Liu, Runze Li 0003, Meng Zheng 0002, Srikrishna Karanam, Ziyan Wu 0001, Bir Bhanu, Richard J. Radke, Octavia I. Camps
CVPR6
2020 Early Wildfire Smoke Detection in Videos
abstract
Recent advances in unmanned aerial vehicles and camera technology have proven useful for the detection of smoke that emerges above the trees during a forest fire. Automatic detection of smoke in videos is of great interest to Fire department. To date, in most parts of the world, the fire is not detected in its early stage and generally it turns catastrophic. This paper introduces a novel technique that integrates spatial and temporal features in a deep learning framework using semi-supervised spatio-temporal video object segmentation and dense optical flow. However, detecting this smoke in the presence of haze and without the labeled data is difficult. Considering the visibility of haze in the sky, a dark channel pre-processing method is used that reduces the amount of haze in video frames and consequently improves the detection results. Online training is performed on a video at the time of testing that reduces the need for ground-truth data. Tests using the publicly available video datasets show that the proposed algorithms outperform previous work and they are robust across different wildfire-threatened locations.
Taanya Gupta, Hengyue Liu, Bir Bhanu
ICPR3
2020 SAGE: Sequential Attribute Generator for Analyzing Glioblastomas Using Limited Dataset
abstract
While deep learning approaches have shown remarkable performance in many imaging tasks, most of these methods rely on the availability of large quantities of data. Medical imaging data, however, are scarce and fragmented. Generative Adversarial Networks (GANs) have recently been very effective in handling such datasets by generating more data. If the datasets are very small, however, GANs cannot learn the data distribution properly, resulting in less diverse or low-quality results. One such limited dataset is that for the concurrent gain of 19/20 chromosomes (19/20 co-gain), a mutation with positive prognostic value in Glioblastomas (GBM). In this paper, imaging biomarkers are detected for the mutation to streamline the extensive and invasive prognosis pipeline. Since this mutation is relatively rare, i.e. small dataset, a novel generative framework - the Sequential Attribute GEnerator (SAGE), is proposed, that generates detailed tumor imaging features while learning from a limited dataset. Experiments show that not only does SAGE generate high quality tumors when compared to Progressively Growing GAN (PGGAN), Wasserstein GAN with Gradient Penalty (WGAN-GP) and Deep Convolutional-GAN (DC-GAN), but also captures the imaging biomarkers accurately.
Padmaja Jonnalagedda, Brent D. Weinberg, Jason Allen, Taejin L. Min, Shiv Bhanu, Bir Bhanu
ICPR6
2020 Depth Videos for the Classification of Micro-Expressions
abstract
Facial micro-expressions are spontaneous, subtle, involuntary muscle movements occurring briefly on the face. The spotting and recognition of these expressions are difficult due to the subtle behavior, and the time duration of these expressions is about half a second, which makes it difficult for humans to identify them. These micro-expressions have many applications in our daily life, such as in the field of online learning, game playing, lie detection, and therapy sessions. Traditionally, researchers use RGB images/videos to spot and classify these micro-expressions, which pose challenging problems, such as illumination, privacy concerns and pose variation. The use of depth videos solves these issues to some extent, as the depth videos are not susceptible to the variation in illumination. This paper describes the collection of a first RGB-D dataset for the classification of facial micro-expressions into 6 universal expressions: Anger, Happy, Sad, Fear, Disgust, and Surprise. This paper shows the comparison between the RGB and Depth videos for the classification of facial micro-expressions. Further, a comparison of results shows that depth videos alone can be used to classify facial micro-expressions correctly in a decision tree structure by using the traditional and deep learning approaches with good classification accuracy. The dataset will be released to the public in the near future.
Ankith Jain Rakesh Kumar, Bir Bhanu, Christopher Casey, Sierra Grace Cheung, Aaron R. Seitz
ICPR2
2020 The Role of Cycle Consistency for Generating Better Human Action Videos from a Single Frame
abstract
This paper addresses the challenging problem of generating videos with human action semantics. Unlike previous work which predict future frames in a single forward pass, this paper introduces the cycle constraints in both forward and backward passes in the generation of human actions. This is achieved by enforcing the appearance and motion consistency across a sequence of frames generated in the future. The approach consists of two stages. In the first stage, the pose of a human body is generated. In the second stage, an image generator is used to generate future frames by using (a) generated human poses in the future from the first stage, (b) the single observed human pose, and (c) the single corresponding future frame. The experiments are performed on three datasets: Weizmann dataset involving simple human actions, Penn Action dataset and UCF-101 dataset containing complicated human actions, especially in sports. The results from these experiments demonstrate the effectiveness of the proposed approach.
Runze Li 0003, Bir Bhanu
ICPR2
2020 Fast Region-Adaptive Defogging and Enhancement for Outdoor Images Containing Sky
abstract
Inclement weather, haze, and fog severely decrease the performance of outdoor imaging systems. Due to a large range of the depth-of-field, most image dehazing or enhancement methods suffer from color distortions and halo artifacts when applied to real-world hazy outdoor scenes, especially those with the sky. To effectively recover details in both distant and nearby regions as well as to preserve color fidelity of the sky, in this study, we propose a novel image defogging and enhancement approach based on a replaceable plug-in segmentation module and region-adaptive processing. First, regions of the grayish sky, pure white objects, and other parts are separated. Second, a luminance-inverted multi-scale Retinex with color restoration (MSRCR) and region-ratio-based adaptive Gamma correction are applied to non-grayish and non-white areas. Finally, the enhanced regions are stitched seamlessly by using a mean-filtered region mask. Extensive experiments show that the proposed approach not only outperforms several state-of-the-art defogging methods in terms of both visibility and color fidelity, but also provides enhanced outputs with fewer artifacts and halos, particularly in sky regions.
Zhan Li 0004, Xiaopeng Zheng, Bir Bhanu, Shun Long, Zhenghao Huang
ICPR3
2020 A new patch selection method based on parsing and saliency detection for person re-identification
Yixiu Liu, Yunzhou Zhang, Sonya A. Coleman, Bir Bhanu, Shuangwei Liu
Neurocomputing4
2020 Physical Features and Deep Learning-based Appearance Features for Vehicle Classification from Rear View Videos
abstract
Currently, there are many approaches for vehicle classification, but there is no specific study on automated, rear view, and video-based robust vehicle classification. The rear view is important for intelligent transportation systems since not all states in the United States require a frontal license plate on a vehicle. The classification of vehicles, from their rear views, is challenging since vehicles have only subtle appearance differences and there are changing illumination conditions and the presence of moving shadows. In this paper, we present a novel multi-class vehicle classification system that classifies a vehicle into one of four possible classes (sedan, minivan, SUV, and a pickup truck) from its rear view video, using physical and visual features. For a given geometric setup of the camera on highways, we make physical measurements on a vehicle. These measurements include visual rear ground clearance, the height of the vehicle, and the distance between the license plate and the rear bumper. We call these distances as the physical features. The visual features, also called appearance-based features, are extracted using convolutional neural networks from the input images. We achieve a classification accuracy of 93.22% and 91.52% using physical and visual features, respectively. Furthermore, we achieve a higher classification accuracy of 94.81% by fusing both the features together. The results are shown on a dataset consisting of 1831 rear view videos of vehicles and they are compared with various approaches, including deep learning techniques.
Rajkumar Theagarajan, Ninad Thakoor, Bir Bhanu
IEEE Trans. Intell. Transp. Syst.3
2019 ShieldNets: Defending Against Adversarial Attacks Using Probabilistic Adversarial Robustness
abstract
Defending adversarial attack is a critical step towards reliable deployment of deep learning empowered solutions for industrial applications. Probabilistic adversarial robustness (PAR), as a theoretical framework, is introduced to neutralize adversarial attacks by concentrating sample probability to adversarial-free zones. Distinct to most of the existing defense mechanisms that require modifying the architecture/training of the target classifier which is not feasible in the real-world scenario, e.g., when a model has already been deployed, PAR is designed in the first place to provide proactive protection to an existing fixed model. ShieldNet is implemented as a demonstration of PAR in this work by using PixelCNN. Experimental results show that this approach is generalizable, robust against adversarial transferability and resistant to a wide variety of attacks on the Fashion-MNIST and CIFAR10 datasets, respectively.
Rajkumar Theagarajan, Ming Chen 0018, Bir Bhanu
CVPR3
2018 MVPNets: Multi-viewing Path Deep Learning Neural Networks for Magnification Invariant Diagnosis in Breast Cancer
abstract
Breast cancer diagnosis requires a pathologist to analyze the histology slides under various magnifications. An automated diagnosis method to aid pathologists that is magnification independent will significantly save time, reduce cost and mitigate subjectivity and errors in current histopathological diagnosis procedures. This paper presents a deep learning network, called MVPNet and a customized data augmentation technique, called NuView, for magnification independent diagnosis. MVPNet is tailored to tackle the most common issues (diversity, relatively small size of datasets and manifestation of diagnostic biomarkers at various magnification levels) with breast cancer histology data to perform the classification. The network simultaneously analyzes local and global features of a given tissue image. It does so by viewing the tissue at varying levels of relative nuclei sizes. MVPNet has significantly less parameters than standard transfer learning deep models with comparable performance and it combines and processes local and global features simulatenously for effective diagnosis. Additionally, NuView extracts tumor nuclei location and points the attention of MVPNet to the informative region specifically. The method gives an average magnification independent classification accuracy of 92.2% as compared to 83% reported in literature on the BreaKHis database.
Padmaja Jonnalagedda, Daniel Schmolze, Bir Bhanu
BIBE3
2018 Patch Based Latent Fingerprint Matching Using Deep Learning
abstract
Latent fingerprints are fingerprint impressions unintentionally left on surfaces at a crime scene. Such fingerprints are usually incomplete or partial, making it challenging to match them to full fingerprints registered in fingerprint databases. Latent fingerprints may contain few minutiae and no singular structures. Matching algorithms that entirely rely on minutiae or alignment of singular structures fail when those structures are missing. This paper presents an approach for matching latent to rolled fingerprints using the (a) similarity of learned representations of patches and (b) the minutiae on the correlated patches. A deep learning network is used to learn optimized representations of image patches. Similarity scores between patches from the latent and reference fingerprints are determined using a distance metric learned with a convolutional neural network. The matching score is obtained by fusing the patch and minutiae similarity scores. The proposed system was tested by matching fingerprints segmented from the 258 latent fingerprints in the NIST SD27 database against a database of 2,257 rolled fingerprints from NIST SD27 and SD4 databases. Experimental results show a rank-1 identification rate of 81.35% and highlights the promise of our proposed approach.
Jude Ezeobiejesi, Bir Bhanu
ICIP2
2018 Deepagent: An Algorithm Integration Approach for Person Re-Identification
abstract
Person re-identification(RE-ID) has played a significant role in the fields of image processing and computer vision because of its potential value in practical applications. Researchers are striving to design new algorithms to improve the performance of RE-ID but ignore the advantages of existing approaches. In this paper, motivated by deep reinforcement learning, we propose a Deep Agent which can integrate existing algorithms and enable them to complement each other. Two Deep Agents are designed to integrate algorithms for data augmentation and feature extraction parts separately for RE-ID. Experiment results demonstrate that the integrated algorithms can achieve a better accuracy than using each one of them alone.
Fulong Jiao, Bir Bhanu
ICIP2
2018 3D Reconstruction of Phase Contrast Images Using Focus Measures
abstract
In this paper, we present an approach for 3D phase-contrast microscopy using focus measure features. By using fluorescence data from the same location as the phase contrast data, we can train supervised regression algorithms to compute a depth map indicating the height of objects in imaged volume. From these depth maps, a 3D reconstruction of phase contrast images can be generated. This paper has shown the ability to 3D reconstruct phase contrast images using a variance metric inspired by all-in-focus methods. The proposed method has been used on A549 lung epithelial cells.
Vincent On, Atena Zahedi, Bir Bhanu
ICIP3
2018 Dyfusion: Dynamic IR/RGB Fusion for Maritime Vessel Recognition
abstract
We propose a novel multi-sensor data fusion approach called DyFusion for maritime vessel recognition using long-wave infrared and visible images. DyFusion consists of a decision-level fusion of convolutional networks using a probabilistic model that can adapt to changes in the scene. The probabilistic model avails of contextual clues from each sensor decision pipeline to maximize accuracy and to update probabilities given to each sensor pipeline. Additional sensors are simulated by applying simple transformations on visible images. Evaluation is presented on the VAIS dataset, demonstrating the effectiveness and robustness of DyFusion with a reliable accuracy of up to 88% in hard scenarios.
Cassio E. Santos, Bir Bhanu
ICIP2
2018 HESCNET: A Synthetically Pre-Trained Convolutional Neural Network for Human Embryonic Stem Cell Colony Classification
abstract
This paper proposes a method for improving the results of deep convolutional neural network classification using synthetic image samples. Generative adversarial networks are used to generate synthetic images from a dataset of phase-contrast, human embryonic stem cell (hESC) microscopy images. hESCnet, a deep convolutional neural network is trained, and the results are shown on various combinations of synthetic and real images in order to improve the classification results with minimal data.
Adam Witmer, Bir Bhanu
ICIP2
2018 An Unbiased Temporal Representation for Video-Based Person Re-Identification
abstract
Person re-identification (re-id) aims to associate pedestrians across different camera views. As compared to the still image-based re-id, video-based re-id provides not only the spatial information but also the temporal dependency among frames. Most of the existing works apply the convolutional neural networks as a spatial feature extractor and then use backpropagation through time (BPTT) to train recurrent neural networks for temporal information. However, the long-term dependency is very hard to learn in RNNs via BPTT due to gradient vanishing or exploding. In the re-id task, the long-term dependency is quite common since the key information (iden-tity of the pedestrian) exists most of the time along the given sequence. Thus, the importance of a frame should not be determined by its position in a sequence, which is usually biased in state-of-the-art models with RNNs. In this paper, we argue that long-term dependency can be very important and propose an unbiased siamese recurrent convolutional neural network architecture to model and associate pedestrians in a video. Experimental results on two public datasets demonstrate the effectiveness of the proposed method.
Bir Bhanu
ICIP2
2018 DeepDriver: Automated System For measuring Valence and Arousal in Car Driver Videos
abstract
We develop an automated system for analyzing facial expressions using valence and arousal measurements of a car driver. This information is used by Motor Trends magazine to provide car manufacturers a report on how the drivers felt at each moment on the race track. The reason for this is that, the drivers remember only a brief description of the emotions they felt after test driving a car. Our approach is a data driven approach and does not include any pre-processing done to the faces of the drivers. The motivation of this paper is to show that with large amount of data, deep learning networks can extract better and more robust facial features compared to state-of-the-art hand crafted features. The network was trained on just the raw facial images and achieves better results compared to state-of-the-art methods. Our system incorporates Convolutional Neural Networks (CNN) for detecting the face and extracting the facial features, and a Long Short Term Memory (LSTM) for modelling the changes in CNN features with respect to time. The system was evaluated on videos from the Motor Trend Magazines Best Driver Car of the Year 2014-16 and the AFEW-VA dataset. We compared our approach with state-of-the-art methods and show that our approach achieves the better results compared to seven other methods.
Rajkumar Theagarajan, Bir Bhanu, Albert C. Cruz
ICPR2
2018 DeephESC: An Automated System for Generating and Classification of Human Embryonic Stem Cells
abstract
Human Embryonic Stem Cells (hESC's) are promising for the treatment of many diseases such as cancer, Parkinsons, Huntingtons, diabetes mellitus etc. and for toxicological testing. Automated detection and classification of human embryonic stem cell (hESC) videos is of great interest among biologists for quantified analysis of various states of hESC in experimental work. To date, the biologists who study hESC's have to analyze stem cell videos manually. In this paper we introduce a hierarchical classification system consisting of Convolutional Neural Networks (CNN) and Triplet CNN's to classify hESC images into six different classes. We also design an ensemble of Generative Adversarial Networks (GAN) for generating synthetic images of hESC's. We validate the quality of the generated hESC images by training all of our CNN's exclusively on the synthetic images generated by the GAN's and evaluating them on the original hESC images. Experimental results shows that we classify the original hESC images, with an accuracy of 85.67% using the CNN alone, 91.38% accuracy using the CNN and Triplet CNN and 94.11% accuracy by fusing the outputs of the CNN and Triplet CNN's, out performing existing state-of-the-art approaches.
Rajkumar Theagarajan, Benjamin Xueqi Guan, Bir Bhanu
ICPR3
2018 Multi-label Classification of Stem Cell Microscopy Images Using Deep Learning
abstract
This paper develops a pattern recognition and machine learning system to localize cell colony subtypes in multi-label, phase-contrast microscopy images. A convolutional neural network is trained to recognize homogeneous cell colonies, and is used in a sliding-window patch based testing method to localize these homogeneous cell types within heterogeneous, multi-label images. The method is used to determine the effects of nicotine on induced pluripotent stem cells expressing the Huntington's disease phenotype. The results of the network are compared to those of an ECOC classifier trained on texture features. The ability of the network to localize cell phenotypes within heterogeneous colonies is visualized and the temporal behavior of stem cells is analyzed.
Adam Witmer, Bir Bhanu
ICPR2
2018 Words alignment based on association rules for cross-domain sentiment classification
abstract
Automatic classification of sentiment data (e.g., reviews, blogs) has many applications in enterprise user management systems, and can help us understand people’s attitudes about products or services. However, it is difficult to train an accurate sentiment classifier for different domains. One of the major reasons is that people often use different words to express the same sentiment in different domains, and we cannot easily find a direct mapping relationship between them to reduce the differences between domains. So, the accuracy of the sentiment classifier will decline sharply when we apply a classifier trained in one domain to other domains. In this paper, we propose a novel approach called words alignment based on association rules (WAAR) for cross-domain sentiment classification, which can establish an indirect mapping relationship between domain-specific words in different domains by learning the strong association rules between domain-shared words and domain-specific words in the same domain. In this way, the differences between the source domain and target domain can be reduced to some extent, and a more accurate cross-domain classifier can be trained. Experimental results on Amazon® datasets show the effectiveness of our approach on improving the performance of cross-domain sentiment classification.
Xibin Jia, Ya Jin, Xing Su 0001, Barry Cardiff, Bir Bhanu
Frontiers Inf. Technol. Electron. Eng.6
2017 Attributes co-occurrence pattern mining for video-based person re-identification
abstract
Person re-identification has received considerable attention in the image processing, computer vision and pattern recognition communities because of its huge potential for video-based surveillance applications and the challenges it presents due to illumination, pose and viewpoint changes among non-overlapping cameras. Being different from the widely used low-level descriptors, visual attributes (e.g., hair and shirt color) offer a human understandable way to recognize people. In this paper, a new way to take advantage of them is proposed. First, convolutional neural networks are adopted to detect the attributes. Second, the dependencies among attributes are obtained by mining association rules, and they are used to refine the attributes classification results. Third, metric learning technique is used to transfer the attribute learning task to person re-identification. Finally, the approach is integrated into an appearance-based method for video-based person re-identification. Experimental results on two benchmark datasets indicate that attributes can provide improvements both in accuracy and generalization capabilities.
Federico Pala, Bir Bhanu
AVSS3
2017 Foreword
Bir Bhanu, Abdenour Hadid, Mark Nixon, Vitomir Struc
FG1
2017 On the accuracy and robustness of deep triplet embedding for fingerprint liveness detection
abstract
Liveness detection is an anti-spoofing technique for dealing with presentation attacks on biometrics authentication systems. Since biometrics are usually visible to everyone, they can be easily captured by a malignant user and replicated to steal someone's identity. In particular, fingerprints can be easily reproduced by using gummy materials and attached to the impostor's fingertips, making the attack go unnoticed by security personnel and camera networks. In this paper, the classical binary classification formulation (live/fake) is substituted by a deep metric learning framework that can generate a representation of real and artificial fingerprints and explicitly models the underlying factors that explain their inter-and intra-class variations. The framework is based on a deep triplet network architecture and consists of a variation of the original triplet loss function. Experiments show that the approach can perform liveness detection in real-time outperforming the state-of-the-art on several benchmark datasets.
Federico Pala, Bir Bhanu
ICIP2
2017 Novel representation for driver emotion recognition in motor vehicle videos
abstract
A novel feature representation of human facial expressions for emotion recognition is developed. The representation leveraged the background texture removal ability of Anisotropic Inhibited Gabor Filtering (AIGF) with the compact representation of spatiotemporal local binary patterns. The emotion recognition system incorporated face detection and registration followed by the proposed feature representation: Local Anisotropic Inhibited Binary Patterns in Three Orthogonal Planes (LAIBP-TOP) and classification. The system is evaluated on videos from Motor Trend Magazine's Best Driver Car of the Year 2014-2016. The results showed improved performance compared to other state-of-the-art feature representations.
Rajkumar Theagarajan, Bir Bhanu, Albert C. Cruz, Belinda Le, Asongu L. Tambo
ICIP2
2017 A dense flow-based framework for real-time object registration under compound motion
Songfan Yang, Yinjie Lei, Mingyang Li 0001, Ninad Thakoor, Bir Bhanu, Yiguang Liu
Pattern Recognit.6
2017 Integrating Social Grouping for Multitarget Tracking Across Cameras in a CRF Model
abstract
Tracking multiple targets across nonoverlapping cameras aims at estimating the trajectories of all targets, and maintaining their identity labels consistent while they move from one camera to another. Matching targets from different cameras can be very challenging, as there might be significant appearance variation and the blind area between cameras makes the target’s motion less predictable. Unlike most of the existing methods that only focus on modeling the appearance and spatiotemporal cues for inter-camera tracking, this paper presents a novel online learning approach that considers integrating high-level contextual information into the tracking system. The tracking problem is formulated using an online learned conditional random field (CRF) model that minimizes a global energy cost. Besides low-level information, social grouping behavior is explored in order to maintain targets’ identities as they move across cameras. In the proposed method, pairwise grouping behavior of targets is first learned within each camera. During inter-camera tracking, track associations that maintain single camera grouping consistencies are preferred. In addition, we introduce an iterative algorithm to find a good solution for the CRF model. Comparison experiments on several challenging real-world multicamera video sequences show that the proposed method is effective and outperforms the state-of-the-art approaches.
Bir Bhanu
IEEE Trans. Circuits Syst. Video Technol.2
2017 Group Structure Preserving Pedestrian Tracking in a Multicamera Video Network
abstract
Pedestrian tracking in video has been a popular research topic with many practical applications. In order to improve tracking performance, many ideas have been proposed, among which the use of geometric information is one of the most popular directions in recent research. In this paper, we propose a novel multicamera pedestrian tracking framework, which incorporates the structural information of pedestrian groups in the crowd. In this framework, first, a new cross-camera model is proposed, which enables the fusion of the confidence information from all camera views. Second, the group structures on the ground plane provide extra constraints between pedestrians. Third, the structured support vector machine is adopted to update the cross-camera model for each pedestrian according to the most recent tracked location. The experiments and detailed analysis are conducted on challenging data. The results demonstrate that the improvement in tracking performance is significant when a group structure is integrated.
Zhixing Jin, Bir Bhanu
IEEE Trans. Circuits Syst. Video Technol.3
2016 Selective experience replay in reinforcement learning for reidentification
abstract
Person reidentification is a problem of recognizing a person across non-overlapping camera views. Pose variations, illumination conditions, low resolution images, and occlusion are the main challenges encountered in reidentification. Due to the uncontrolled environment in which the videos are captured, people could appear in different poses and due to which the appearance of a person could vary significantly. The walking direction of a person can provide a good estimation of their pose. Therefore, in this paper, we propose a reidentification system which adaptively selects an appropriate distance metric based on context of walking direction using reinforcement learning. Though experiments, we show that such a dynamic strategy outperforms static strategy learned or designed offline.
Ninad Thakoor, Bir Bhanu
ICIP2
2016 Local Invariance Representation Learning Algorithm with Multi-layer Extreme Learning Machine
Xibin Jia, Hua Du, Bir Bhanu
ICONIP (4)4
2016 Spatio-temporal pattern recognition of dendritic spines and protein dynamics using live multichannel fluorescence microscopy
abstract
Actin-regulating proteins, such as cofilin, are essential in regulating the shape of dendritic spines, and synaptic plasticity in both neuronal functionality as well as in neurodegeneration related to aging. The analysis of the motility of cofilin in fluorescence video-microscopy allows the discovery of its effects on cell functions. However, the flow of cofilin has not been analyzed to date by automatic means. This paper presents a novel automated pattern recognition system to analyze protein trafficking in neurons. Using spatio-temporal information present in multichannel fluorescence videos, the system generates a temporal maximum intensity projection that enhances the signal-to-noise ratio of important biological structures, segments and tracks dendritic spines, and quantifies the flux and density of proteins in spines. The temporal dynamics of spines is used to generate spine energy images which are used to automatically classify the shape of dendritic spines as stubby, mushroom, or thin. By tracking these spines over time and using their intensity profiles, the system is able to analyze the flux patterns of cofilin and other fluorescently stained proteins. The cofilin flux patterns is found to be correlated with the dynamically changing dendritic spine shapes. The results are presented using multichannel fluorescence videos.
Vincent On, Atena Zahedi, Iryna Ethell, Bir Bhanu
ICPR4
2016 Temporal dynamics of tip fluorescence predict cell growth behavior in pollen tubes
abstract
In the sexual reproductive life cycle of flowering plants, the growth of the pollen tube plays a vital role. The pollen tube grows towards the ovary of the flower where it delivers male reproductive material. This growth often involves twists and turns as the pollen tube navigates towards the ovary. Current growth models are a collection of mathematical equations to explain observable linear growth behavior in pollen tubes. However, there are few studies on the relationship between the fluorescence signal at the tip of the cell and the growth behavior (straight vs. turning). In this paper, we propose a method of extracting features from the tip fluorescence signal which will be used to distinguishing between straight vs. turning growth behavior. The tip signal is obtained as a ratio of the average membrane-to-cytoplasm fluorescence values over time. A two-stage scheme is used to automatically detect individual growth intervals/cycles from the tip signal and split the experimental video into growth segments. In each growth segment, we extract relevant features. An initial classification uses structure-based features to distinguish between straight vs. turning growth cycles. The signal-based features are then used to train a Naive Bayes classifier to refine the miss-classifications of the initial classification. Our results show that this two-stage process yields good classification results.
Asongu L. Tambo, Bir Bhanu
ICPR2
2016 Sparse representation matching for person re-identification
Songfan Yang, Bir Bhanu
Inf. Sci.4
2016 Semantic Concept Co-Occurrence Patterns for Image Annotation and Retrieval
abstract
Describing visual image contents by semantic concepts is an effective and straightforward way to facilitate various high level applications. Inferring semantic concepts from low-level pictorial feature analysis is challenging due to the semantic gap problem, while manually labeling concepts is unwise because of a large number of images in both online and offline collections. In this paper, we present a novel approach to automatically generate intermediate image descriptors by exploiting concept co-occurrence patterns in the pre-labeled training set that renders it possible to depict complex scene images semantically. Our work is motivated by the fact that multiple concepts that frequently co-occur across images form patterns which could provide contextual cues for individual concept inference. We discover the co-occurrence patterns as hierarchical communities by graph modularity maximization in a network with nodes and edges representing concepts and co-occurrence relationships separately. A random walk process working on the inferred concept probabilities with the discovered co-occurrence patterns is applied to acquire the refined concept signature representation. Through experiments in automatic image annotation and semantic image retrieval on several challenging datasets, we demonstrate the effectiveness of the proposed concept co-occurrence patterns as well as the concept signature representation in comparison with state-of-the-art approaches.
Linan Feng, Bir Bhanu
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 A software system for automated identification and retrieval of moth images based on wing attributes
Linan Feng, Bir Bhanu, John Heraty
Pattern Recognit.2
2016 Understanding pollen tube growth dynamics using the Unscented Kalman Filter
Asongu L. Tambo, Bir Bhanu, Nolan Ung, Ninad Thakoor, Nan Luo, Zhenbiao Yang
Pattern Recognit. Lett.2
2016 Modeling and Classifying Tip Dynamics of Growing Cells in Video
abstract
Plant biologists study pollen tubes to discover the functions of many proteins/ions and map the complex network of pathways that lead to an observable growth behavior. Many growth models have been proposed that address parts of the growth process: internal dynamics and cell wall dynamics, but they do not distinguish between the two types of growth segments: straight versus turning behavior. We propose a method of classifying segments of experimental videos by extracting features from the growth process during each interval. We use a stress-strain relationship to measure the extensibility in the tip region. A biologically relevant three-component Gaussian is used to model spatial distribution of tip extensibility and a second-order damping system is used to explain the temporal dynamics. Feature-based classification shows that the location of maximum tip extensibility is the most distinguishing feature between straight versus turning behavior.
Asongu L. Tambo, Bir Bhanu
IEEE Signal Process. Lett.2
2016 Extraction of Blebs in Human Embryonic Stem Cell Videos
abstract
Blebbing is an important biological indicator in determining the health of human embryonic stem cells (hESC). Especially, areas of a bleb sequence in a video are often used to distinguish two cell blebbing behaviors in hESC: dynamic and apoptotic blebbings. This paper analyzes various segmentation methods for bleb extraction in hESC videos and introduces a bio-inspired score function to improve the performance in bleb extraction. Full bleb formation consists of bleb expansion and retraction. Blebs change their size and image properties dynamically in both processes and between frames. Therefore, adaptive parameters are needed for each segmentation method. A score function derived from the change of bleb area and orientation between consecutive frames is proposed which provides adaptive parameters for bleb extraction in videos. In comparison to manual analysis, the proposed method provides an automated fast and accurate approach for bleb sequence extraction.
Benjamin Xueqi Guan, Bir Bhanu, Prudence Talbot, Nikki Jo-Hao Weng
IEEE ACM Trans. Comput. Biol. Bioinform.2
2016 Person Reidentification With Reference Descriptor
abstract
Person identification across nonoverlapping cameras, also known as person reidentification, aims to match people at different times and locations. Reidentifying people is of great importance in crucial applications such as wide-area surveillance and visual tracking. Due to the appearance variations in pose, illumination, and occlusion in different camera views, person reidentification is inherently difficult. To address these challenges, a reference-based method is proposed for person reidentification across different cameras. Instead of directly matching people by their appearance, the matching is conducted in a reference space where the descriptor for a person is translated from the original color or texture descriptors to similarity measures between this person and the exemplars in the reference set. A subspace is first learned in which the correlations of the reference data from different cameras are maximized using regularized canonical correlation analysis (RCCA). For reidentification, the gallery data and the probe data are projected onto this RCCA subspace and the reference descriptors (RDs) of the gallery and probe are generated by computing the similarity between them and the reference data. The identity of a probe is determined by comparing the RD of the probe and the RDs of the gallery. A reranking step is added to further improve the results using a saliency-based matching scheme. Experiments on publicly available datasets show that the proposed method outperforms most of the state-of-the-art approaches.
Mehran Kafai, Songfan Yang, Bir Bhanu
IEEE Trans. Circuits Syst. Video Technol.4
2016 Multiperson Tracking by Online Learned Grouping Model With Nonlinear Motion Context
abstract
An online approach to learn elementary groups containing only two targets, i.e., pedestrians, for inferring high-level context is introduced to improve multiperson tracking. In most existing data association-based tracking approaches, only low-level information (e.g., time, appearance, and motion) is used to build the affinity model, and each target is considered as an independent agent. Unlike those previous methods, in this paper, an online learned social grouping behavior model is used to provide more robust tracklet affinities. A disjoint grouping graph is used to encode social grouping behavior of pairwise targets, where each node represents an elementary group of two targets, and two nodes are connected if they share a common target. Probabilities of the uncertain target in two connected nodes being the same person are inferred from each edge of the grouping graph. Relationships between elementary groups are discovered by group tracking, and a nonlinear motion map is used for explaining nonlinear motion pattern between elementary groups. The proposed method is efficient, able to handle group split and merge, and can be easily integrated into any basic affinity model. The approach is evaluated on four data sets, and it shows significant improvements compared with state-of-the-art methods.
Zhen Qin 0001, Bir Bhanu
IEEE Trans. Circuits Syst. Video Technol.4
2016 Segmentation of Pollen Tube Growth Videos Using Dynamic Bi-Modal Fusion and Seam Carving
abstract
The growth of pollen tubes is of significant interest in plant cell biology, as it provides an understanding of internal cell dynamics that affect observable structural characteristics such as cell diameter, length, and growth rate. However, these parameters can only be measured in experimental videos if the complete shape of the cell is known. The challenge is to accurately obtain the cell boundary in noisy video images. Usually, these measurements are performed by a scientist who manually draws regions-of-interest on the images displayed on a computer screen. In this paper, a new automated technique is presented for boundary detection by fusing fluorescence and brightfield images, and a new efficient method of obtaining the final cell boundary through the process of Seam Carving is proposed. This approach takes advantage of the nature of the fusion process and also the shape of the pollen tube to efficiently search for the optimal cell boundary. In video segmentation, the first two frames are used to initialize the segmentation process by creating a search space based on a parametric model of the cell shape. Updates to the search space are performed based on the location of past segmentations and a prediction of the next segmentation.Experimental results show comparable accuracy to a previous method, but significant decrease in processing time. This has the potential for real time applications in pollen tube microscopy.
Asongu L. Tambo, Bir Bhanu
IEEE Trans. Image Process.2
2015 Dynamic bi-modal fusion of images for the segmentation of pollen tubes in video
abstract
Biologists study pollen tube growth to understand how internal cell dynamics affect observable structural characteristics like cell diameter, length, and growth rate. Fluorescence microscopy is used to study the dynamics of internal proteins and ions, but this often produces images with missing parts of the pollen tube. Brightfield microscopy provides a low-cost way of obtaining structural information about the pollen tube, but the images are crowded with false edges. We propose a dynamic segmentation fusion scheme that uses both Bright-field and Fluorescence images of growing pollen tubes to get a unified segmentation. Knowledge of the image formation process is used to create an initial estimate of the location of the cell boundary. Fusing this estimate with an edge indicator function amplifies desired edges and attenuates undesired edges. The cell boundary is obtained using Level Set evolution on the fused edge indicator function. Experimental testing shows that this fusion produces significantly better results than those obtained without it.
Asongu L. Tambo, Bir Bhanu
ICIP2
2015 Tracking People by Evolving Social Groups: An Approach with Social Network Perspective
abstract
We address the problem of multi-people tracking in unconstrained and semi-crowded scenes. People typically walk in groups that split and merge over time. The evolving or dynamic social group property embodies pedestrians' connections and interactions during walking which we attempt to identify and exploit in this paper. To this end, instead of seeking more robust appearance or motion models to track each person as an isolated moving entity, we pose the multi-people tracking problem as a group-based tracklets association problem using the discovered social groups of track lets as the contextual cues. We formulate tracking the evolution of social groups of tracklets as detecting closely connected communities in a "tracklet interaction network" (TIN) with nodes standing for the tracklets and edges denoting the spatio-temporal co-occurrence correlations measured by the edge weights. We incorporate the detected social groups in the tracklet interaction network to improve multi-people tracking performance. We evaluate our approach against state-of-the-art and show improvements on three real-world datasets.
Linan Feng, Bir Bhanu
WACV2
2015 Editorial introduction to the special issue on "Image Understanding for Real-World Distributed Video Networks" - Computer Vision and Image Understanding Journal
Bir Bhanu, Andrea Prati 0001, Faisal Z. Qureshi
Comput. Vis. Image Underst.1
2015 Analysis-by-synthesis: Pedestrian tracking with crowd simulation models in a multi-camera video network
Zhixing Jin, Bir Bhanu
Comput. Vis. Image Underst.2
2015 Efficient smile detection by Extreme Learning Machine
Songfan Yang, Bir Bhanu
Neurocomputing3
2015 Background suppressing Gabor energy filtering
Albert C. Cruz, Bir Bhanu, Ninad Thakoor
Pattern Recognit. Lett.2
2015 Person Re-Identification by Robust Canonical Correlation Analysis
abstract
Person re-identification is the task to match people in surveillance cameras at different time and location. Due to significant view and pose change across non-overlapping cameras, directly matching data from different views is a challenging issue to solve. In this letter, we propose a robust canonical correlation analysis (ROCCA) to match people from different views in a coherent subspace. Given a small training set as in most re-identification problems, direct application of canonical correlation analysis (CCA) may lead to poor performance due to the inaccuracy in estimating the data covariance matrices. The proposed ROCCA with shrinkage estimation and smoothing technique is simple to implement and can robustly estimate the data covariance matrices with limited training samples. Experimental results on two publicly available datasets show that the proposed ROCCA outperforms regularized CCA (RCCA), and achieves state-of-the-art matching results for person re-identification as compared to the most recent methods.
Songfan Yang, Bir Bhanu
IEEE Signal Process. Lett.3
2014 An Online Learned Elementary Grouping Model for Multi-target Tracking
abstract
We introduce an online approach to learn possible elementary groups (groups that contain only two targets) for inferring high level context that can be used to improve multi-target tracking in a data-association based framework. Unlike most existing association-based tracking approaches that use only low level information (e.g., time, appearance, and motion) to build the affinity model and consider each target as an independent agent, we online learn social grouping behavior to provide additional information for producing more robust tracklets affinities. Social grouping behavior of pairwise targets is first learned from confident tracklets and encoded in a disjoint grouping graph. The grouping graph is further completed with the help of group tracking. The proposed method is efficient, handles group merge and split, and can be easily integrated into any basic affinity model. We evaluate our approach on two public datasets, and show significant improvements compared with state-of-the-art methods.
Zhen Qin 0001, Bir Bhanu
CVPR4
2014 One shot emotion scores for facial emotion recognition
abstract
Facial emotion recognition in unconstrained settings is a difficult task. They key problems are that people express their emotions in ways that are different from other people, and, for large datasets, there are not enough examples of a specific person to model his/her emotion. A model for predicting emotions will not generalize well to predicting the emotions of a person who has not been encountered during the training. We propose a system that addresses these issues by matching a face video to references of emotion. It does not require examples from the person in the video being queried. We compute the matching scores without requiring fine registration. The method is called one-shot emotion score. We improve classification rate of interdataset experiments over a baseline system by 23% when training on MMI and testing on CK+.
Albert C. Cruz, Bir Bhanu, Ninad Thakoor
ICIP2
2014 Comparison of texture features for human embryonic stem cells with bio-inspired multi-class support vector machine
abstract
Determining the meaningful texture features for human embryonic stem cells (hESC) is important in the development of online hESC classification system. This paper proposes the use of novel support vector machine with bio-inspired one-against-all (OAA) multi-class structural and statistical Gabor descriptors for hESC classification. It investigates the statistical histogram information at four different orientations and two different window sizes of the Gabor filter. It demonstrates that statistical Gabor features are more accurate and reliable than a conventional histogram based features.
Benjamin Xueqi Guan, Bir Bhanu, Prudence Talbot, Sabrina Lin, Nikki Jo-Hao Weng
ICIP2
2014 Efficient alignment for vehicle make and model recognition
abstract
This paper presents a make and model recognition system for passenger vehicles. We propose a two-step efficient alignment mechanism to account for view point changes. The 2D alignment problem is solved as two separate one dimensional shortest path problems. To avoid the alignment of the query with the entire database, reference views are used. These views are generated iteratively from the database. To improve the alignment performance further, use of two references is proposed: a universal view and type specific showcase views. The query is aligned with universal view first and compared with the database to find the type of the query. Then the query is aligned with type specific showcase view and compared with the database to achieve the final make and model recognition. We report results on database of 1500 vehicles with more than 250 makes and models.
Ninad Thakoor, Bir Bhanu
ICIP2
2014 Soft Biometrics Integrated Multi-target Tracking
abstract
In this paper, we present a soft biometrics based appearance model for multi-target tracking in a single camera. Track lets, the short-term tracking results, are generated by linking detections in consecutive frames based on conservative constraints. Our goal is to "re-stitching" the adjacent track lets that contain the same target so that robust long-term tracking results can be achieved. As the appearance of the same target may change greatly due to heavy occlusion, pose variations and changing lighting conditions, a discriminative appearance model is crucial for association-based tracking. Unlike most previous methods which simply use the similarity of color histograms or other low level features to construct the appearance model, we propose to use the fusion of soft biometrics generated from sub-track lets to learn a discriminative appearance model in an online manner. Compared to low level features, soft biometrics are robust against appearance variation. The experimental results demonstrate that our method is robust and greatly improves the tracking performance over the state-of-the-art method.
Bir Bhanu
ICPR2
2014 Integrated Model for Understanding Pollen Tube Growth in Video
abstract
Pollen tube growth is an essential part of the sexual reproductive process in plants. It is the result of a complex interaction of cytoplasmic contents (proteins, ions, cellular structures, etc.). Existing pollen tube models use differential equations to represent these complex intra-cellular interactions that lead to growth. As a result of this complex nature, these models are not used to verify the shape and growth behavior observed in living cells. We present a method of analyzing the growth behavior of pollen tubes in experimental videos through affine transformations on the detected cell tip. The method relies on underlying biological knowledge about the growth process and leverages these processes to determine tip morphology. Experimental results on videos of growing pollen tube cells show that our method is superior to the current method of treating cell tip morphology as well as adaptive active appearance models.
Asongu L. Tambo, Bir Bhanu, Nan Luo, Geoffrey Harlow, Zhenbiao Yang
ICPR2
2014 Automated detection of brain abnormalities in neonatal hypoxia ischemic injury from MR images
Nirmalya Ghosh, Yu Sun 0007, Bir Bhanu, Stephen Ashwal, Andre Obenaus
Medical Image Anal.3
2014 Predictive models for multibiometric systems
Suresh Kumar Ramachandran Nair, Bir Bhanu, Subir Ghosh, Ninad Thakoor
Pattern Recognit.2
2014 Face image super-resolution using 2D CCA
Bir Bhanu
Signal Process.2
2014 Vision and Attention Theory Based Sampling for Continuous Facial Emotion Recognition
abstract
Affective computing-the emergent field in which computers detect emotions and project appropriate expressions of their own-has reached a bottleneck where algorithms are not able to infer a person's emotions from natural and spontaneous facial expressions captured in video. While the field of emotion recognition has seen many advances in the past decade, a facial emotion recognition approach has not yet been revealed which performs well in unconstrained settings. In this paper, we propose a principled method which addresses the temporal dynamics of facial emotions and expressions in video with a sampling approach inspired from human perceptual psychology. We test the efficacy of the method on the Audio/Visual Emotion Challenge 2011 and 2012, CohnKanade and the MMI Facial Expression Database. The method shows an average improvement of 9.8 percent over the baseline for weighted accuracy on the Audio/Visual Emotion Challenge 2011 video-based frame-level subchallenge testing set.
Albert C. Cruz, Bir Bhanu, Ninad Thakoor
IEEE Trans. Affect. Comput.2
2014 Zapping Index: Using Smile to Measure Advertisement Zapping Likelihood
abstract
In marketing and advertising research, “zapping” is defined as the action when a viewer stops watching a commercial. Researchers analyze users' behavior in order to prevent zapping which helps advertisers to design effective commercials. Since emotions can be used to engage consumers, in this paper, we leverage automated facial expression analysis to understand consumers' zapping behavior. Firstly, we provide an accurate moment-to-moment smile detection algorithm. Secondly, we formulate a binary classification problem (zapping/non-zapping) based on real-world scenarios, and adopt smile response as the feature to predict zapping. Thirdly, to cope with the lack of a metric in advertising evaluation, we propose a new metric called Zapping Index (ZI). ZI is a moment-to-moment measurement of a user's zapping probability. It gauges not only the reaction of a user, but also the preference of a user to commercials. Finally, extensive experiments are performed to provide insights and we make recommendations that will be useful to both advertisers and advertisement publishers.
Songfan Yang, Mehran Kafai, Bir Bhanu
IEEE Trans. Affect. Comput.4
2014 Bio-Driven Cell Region Detection in Human Embryonic Stem Cell Assay
abstract
This paper proposes a bio-driven algorithm that detects cell regions automatically in the human embryonic stem cell (hESC) images obtained using a phase contrast microscope. The algorithm uses both statistical intensity distributions of foreground/hESCs and background/substrate as well as cell property for cell region detection. The intensity distributions of foreground/hESCs and background/substrate are modeled as a mixture of two Gaussians. The cell property is translated into local spatial information. The algorithm is optimized by parameters of the modeled distributions and cell regions evolve with the local cell property. The paper validates the method with various videos acquired using different microscope objectives. In comparison with the state-of-the-art methods, the proposed method is able to detect the entire cell region instead of fragmented cell regions. It also yields high marks on measures such as Jacard similarity, Dice coefficient, sensitivity and specificity. Automated detection by the proposed method has the potential to enable fast quantifiable analysis of hESCs using large data sets which are needed to understand dynamic cell behaviors.
Benjamin Xueqi Guan, Bir Bhanu, Prudence Talbot, Sabrina Lin
IEEE ACM Trans. Comput. Biol. Bioinform.2
2014 Reference Face Graph for Face Recognition
abstract
Face recognition has been studied extensively; however, real-world face recognition still remains a challenging task. The demand for unconstrained practical face recognition is rising with the explosion of online multimedia such as social networks, and video surveillance footage where face analysis is of significant importance. In this paper, we approach face recognition in the context of graph theory. We recognize an unknown face using an external reference face graph (RFG). An RFG is generated and recognition of a given face is achieved by comparing it to the faces in the constructed RFG. Centrality measures are utilized to identify distinctive faces in the reference face graph. The proposed RFG-based face recognition algorithm is robust to the changes in pose and it is also alignment free. The RFG recognition is used in conjunction with DCT locality sensitive hashing for efficient retrieval to ensure scalability. Experiments are conducted on several publicly available databases and the results show that the proposed approach outperforms the state-of-the-art methods without any preprocessing necessities such as face alignment. Due to the richness in the reference set construction, the proposed method can also handle illumination and expression variation.
Mehran Kafai, Bir Bhanu
IEEE Trans. Inf. Forensics Secur.3
2014 Evolving Bayesian Graph for Three-Dimensional Vehicle Model Building From Video
abstract
Traffic videos often capture slowly changing views of moving vehicles. These different and incrementally related views provide visual cues for 3-D perception of the vehicles from 2-D videos. This paper focuses on 3-D model building ofmultiplevehicles with different shapes from asinglegeneric 3-D vehicle model by incrementally accumulating evidences in streaming traffic videos collected from a single static uncalibrated camera. When we do not knowa priorithe class of the following vehicle to be seen (which is true in a real traffic scenario), a flexible and evolvable Bayesian graphical model (BGM) is required, where the number of nodes, the structure of links between them, and the associated conditional probability distributions can change on the fly. Current BGMs fail to provide such online flexibility. We propose a novel BGM, which is called structure-modifiable adaptive reason-building temporal Bayesian graph (SmartBG), that self-modifies in a data-driven way to model uncertainty propagation in 3-D vehicle model building from 2-D video features, where only a subset of the 2-D vehicle features is visible at any time point, e.g., out of field-of-view (entry/exit) and self-occlusion. Uncertainties are used as relative weights to fuse evidences and to compute the overall reliability of the generated models. Results for different vehicles from several traffic videos and two different viewpoints demonstrate the performance of the proposed method.
Nirmalya Ghosh, Bir Bhanu
IEEE Trans. Intell. Transp. Syst.2
2014 Visual and Contextual Modeling for the Detection of Repeated Mild Traumatic Brain Injury
abstract
Currently, there is a lack of computational methods for the evaluation of mild traumatic brain injury (mTBI) from magnetic resonance imaging (MRI). Further, the development of automated analyses has been hindered by the subtle nature of mTBI abnormalities, which appear as low contrast MR regions. This paper proposes an approach that is able to detect mTBI lesions by combining both the high-level context and low-level visual information. The contextual model estimates the progression of the disease using subject information, such as the time since injury and the knowledge about the location of mTBI. The visual model utilizes texture features in MRI along with a probabilistic support vector machine to maximize the discrimination in unimodal MR images. These two models are fused to obtain a final estimate of the locations of the mTBI lesion. The models are tested using a novel rodent model of repeated mTBI dataset. The experimental results demonstrate that the fusion of both contextual and visual textural features outperforms other state-of-the-art approaches. Clinically, our approach has the potential to benefit both clinicians by speeding diagnosis and patients by improving clinical care.
Anthony C. Bianchi, Bir Bhanu, Virginia Donovan, Andre Obenaus
IEEE Trans. Medical Imaging2
2014 Discrete Cosine Transform Locality-Sensitive Hashes for Face Retrieval
abstract
Descriptors such as local binary patterns perform well for face recognition. Searching large databases using such descriptors has been problematic due to the cost of the linear search, and the inadequate performance of existing indexing methods. We present Discrete Cosine Transform (DCT) hashing for creating index structures for face descriptors. Hashes play the role of keywords: an index is created, and queried to find the images most similar to the query image. Common hash suppression is used to improve retrieval efficiency and accuracy. Results are shown on a combination of six publicly available face databases (LFW, FERET, FEI, BioID, Multi-PIE, and RaFD). It is shown that DCT hashing has significantly better retrieval accuracy and it is more efficient compared to other popular state-of-the-art hash algorithms.
Mehran Kafai, Kave Eshghi, Bir Bhanu
IEEE Trans. Multim.3
2013 Reference-based person re-identification
abstract
Person re-identification refers to recognizing people across non-overlapping cameras at different times and locations. Due to the variations in pose, illumination condition, background, and occlusion, person re-identification is inherently difficult. In this paper, we propose a reference-based method for across camera person re-identification. In the training, we learn a subspace in which the correlations of the reference data from different cameras are maximized using Regularized Canonical Correlation Analysis (RCCA). For re-identification, the gallery data and the probe data are projected into the RCCA subspace and the reference descriptors (RDs) of the gallery and probe are constructed by measuring the similarity between them and the reference data. The identity of the probe is determined by comparing the RD of the probe and the RDs of the gallery. Experiments on benchmark dataset show that the proposed method outperforms the state-of-the-art approaches.
Mehran Kafai, Songfan Yang, Bir Bhanu
AVSS4
2013 Detecting mild traumatic brain injury using dynamic low level context
abstract
Mild traumatic brain injury is difficult to detect in standard magnetic resonance (MR) images due to the low contrast appearance of lesions. In this paper a discriminative approach is presented, using a classifier to directly estimates the posterior probability of lesion at every voxel using low-level context learned from previous classifiers. Both visual features including multiple texture measures, and context features, which include novel features such as proximity, directional distance, and posterior marginal edge distance, are used. The context is also taken from previous time points, so the system automatically captures the dynamics of the injury progression. The approach is tested on an mTBI rat model using MR imaging at multiple time points. Our results show an improved performance in both the dice score and convergence rate compared to other approaches.
Anthony C. Bianchi, Bir Bhanu, Virginia Donovan, Andre Obenaus
ICIP2
2013 Improving large-scale face image retrieval using multi-level features
abstract
In recent years, extensive efforts have been made for face recognition and retrieval systems. However, there remain several challenging tasks for face image retrieval in unconstrained databases where the face images were captured with varying poses, lighting conditions, etc. In addition, the databases are often large-scale, which demand efficient retrieval algorithms that have the merit of scalability. To improve the retrieval accuracy of the face images with different poses and imaging characteristics, we introduce a novel feature extraction method to bag-of-words (BoW) based face image retrieval system. It employs various scales of features simultaneously to encode different texture information and emphasizes image patches that are more discriminative as parts of the face. Moreover, the overlapping image patches at different scales compensate for the pose variation and face misalignment. Experiments conducted on a large-scale public face database demonstrate the superior performance of the proposed approach compared to the state-of-the-art method.
Bir Bhanu
ICIP3
2013 Facial emotion recognition with anisotropic inhibited Gabor energy histograms
abstract
State-of-the-art approaches have yet to deliver a feature representation for facial emotion recognition that can be applied to non-trivial unconstrained, continuous video data sets. Initially, research advanced with the use of Gabor energy filters. However, in recent work more attention has been given to other features. Gabor energy filters lack generalization needed in unconstrained situations. Additionally, they result in an undesirably high feature vector dimensionality. Nontrivial data sets have millions of samples; feature vectors must be as low dimensional as possible. We propose a novel texture feature based on Gabor energy filters that offers generalization with a background texture suppression component and is as compact as possible due to a maximal response representation and local histograms. We improve performance on the non-trivial Audio/Visual Emotion Challenge 2012 grandchallenge data set.
Albert C. Cruz, Bir Bhanu, Ninad Thakoor
ICIP2
2013 Automated identification and retrieval of moth images with semantically related visual attributes on the wings
abstract
A new automated identification and retrieval system is proposed that aims to provide entomologists, who manage insect specimen images, with fast computer-based processing and analyzing techniques. Several relevant image attributes were designed, such as the so-called semantically-related visual (SRV) attributes detected from the insect wings and the co-occurrence patterns of the SRV attributes which are uncovered from manually labeled training samples. A joint probabilistic model is used as SRV attribute detector working on image visual contents. The identification and retrieval of moth species are conducted by comparing the similarity of SRV attributes and their co-occurrence patterns. The prototype system used moth images while it can be generalized to any insect species with wing structures. The system performed with good stability and the accuracy reached 85% for species identification and 71% for content-based image retrieval on a entomology database.
Linan Feng, Bir Bhanu
ICIP2
2013 Optimizing crowd simulation based on real video data
abstract
Tracking of individuals and groups in video is an active topic of research in image processing and analyzing. This paper proposes an approach for the purpose of guiding a crowd simulation algorithm to mimic the trajectories of individuals in crowds as observed in real videos, which can be further used in image processing and computer vision research extensively. This is achieved by tuning the parameters used in the simulation automatically. It is required because the result of crowd simulation is very sensitive to the parameters. In our experiment, the simulation trajectories are generated by the RVO2 library and the real trajectories are extracted from the UCSD crowd video dataset. The Edit Distance on Real sequence (EDR) between the simulated and real trajectories are calculated. A genetic algorithm is applied to find the parameters that minimize the distances. The experimental results demonstrate that the trajectory distances between simulation and reality are significantly reduced after tuning the parameters of the simulator.
Zhixing Jin, Bir Bhanu
ICIP2
2013 Representative reference-set and betweenness centrality for scene image categorization
abstract
Reference-based image classification approach introduces a reference-set for both image representation and dictionary learning. It significantly reduces the dimensionality of represented images and shows outstanding performance even with randomly selected reference images and simple distance measure. In this paper, we improve upon existing work with two major contributions. First, we show that a more representative reference-set contributes to better classification accuracy. To this end, we carefully adapt the K-means clustering algorithm in the feature space to select a distinguished reference-set. Second, in the image classification process, we propose to represent each image by measuring its betweenness centrality in a social network composed of the representative reference-set in each class, leading to a more coherent distance measure that considers the overall connectivity between the probe image and the reference-set. Extensive experiment results demonstrate that our proposed scheme achieves better performance than existing methods.
Qun Li 0002, Zhen Qin 0001, Lunshao Chai, Honggang Zhang 0002, Jun Guo 0002, Bir Bhanu
ICIP6
2013 Learning small gallery size for prediction of recognition performance on large populations
Rong Wang 0007, Bir Bhanu, Ninad Thakoor
Pattern Recognit.2
2013 Reference-Based Scheme Combined With K-SVD for Scene Image Categorization
abstract
A reference-based algorithm for scene image categorization is presented in this letter. In addition to using a reference-set for images representation, we also associate the reference-set with training data in sparse codes during the dictionary learning process. The reference-set is combined with the reconstruction error to form a unified objective function. The optimal solution is efficiently obtained using the K-SVD algorithm. After dictionaries are constructed, Locality-constrained Linear Coding (LLC) features of images are extracted. Then, we represent each image feature vector using the similarities between the image and the reference-set, leading to a significant reduction of the dimensionality in the feature space. Experimental results demonstrate that our method achieves outstanding performance.
Qun Li 0002, Honggang Zhang 0002, Jun Guo 0002, Bir Bhanu
IEEE Signal Process. Lett.4
2013 Structural Signatures for Passenger Vehicle Classification in Video
abstract
This paper focuses on a challenging pattern recognition problem of significant industrial impact, i.e., classifying vehicles from their rear videos as observed by a camera mounted on top of a highway with vehicles traveling at high speed. To solve this problem, this paper presents a novel feature called structural signature. From a rear-view video, a structural signature recovers the vehicle side profile information, which is crucial in its classification. As a vehicle moves away from a camera, its surfaces deform differently based on their relative orientation to the camera. This information is used to extract the structure of a vehicle, which captures the relative orientation of vehicle surfaces and the road surface. This paper presents a complete system that computes structural signatures and uses them for classification of passenger vehicles into sedans, pickups, and minivans/sport utility vehicles in highway videos. It analyzes the performance of the proposed system on a large video data set.
Ninad Thakoor, Bir Bhanu
IEEE Trans. Intell. Transp. Syst.2
2012 Boosting Face Recognition in Real-World Surveillance Videos
abstract
Face recognition becomes a challenging problem in real-world surveillance videos where the low-resolution probe frames exhibit variations in pose, lighting condition, and facial expressions. This is in contrast with the gallery images which are generally frontal view faces acquired under controlled environments. A direct matching of probe images with gallery data often leads to poor recognition accuracy due to the significant discrepancy between the two kinds of data. In addition, the artifacts such as low resolution, blurriness and noise further enlarge this discrepancy. In this paper, we propose a video based face recognition framework using a novel image representation called warped average face (WAF). The WAFs are generated in two stages: in-sequence warping and frontal view warping. The WAFs can be easily used with various feature descriptors or classifiers. As compared to the original probe data, the image quality of the WAFs is significantly better and the appearance difference between the WAFs and the gallery data is suppressed. Given a probe sequence, only a few WAFs need to be generated for the recognition purpose. We test the proposed method on the ChokePoint dataset and our in-house dataset of surveillance quality. Experiments show that with the new image representation, the recognition accuracy can be boosted significantly.
Bir Bhanu, Songfan Yang
AVSS2
2012 Real-Time Pedestrian Tracking with Bacterial Foraging Optimization
abstract
In this paper, we present swarm intelligence algorithms for pedestrian tracking. In particular, we present a modified Bacterial Foraging Optimization (BFO) algorithm and show that it outperforms PSO in a number of important metrics for pedestrian tracking. In our experiments, we show that BFO's search strategy is inherently more efficient than PSO under a range of variables with regard to the number of fitness evaluations which need to be performed when tracking. We also compare the proposed BFO approach with other commonly-used trackers and present experimental results on the CAVIAR dataset as well as on the difficult PETS2010 S2.L3 crowd video.
Hoang Thanh Nguyen 0001, Bir Bhanu
AVSS2
2012 Image super-resolution by extreme learning machine
abstract
Image super-resolution is the process to generate high-resolution images from low-resolution inputs. In this paper, an efficient image super-resolution approach based on the recent development of extreme learning machine (ELM) is proposed. We aim at reconstructing the high-frequency components containing details and fine structures that are missing from the low-resolution images. In the training step, high-frequency components from the original high-resolution images as the target values and image features from low-resolution images are fed to ELM to learn a model. Given a low-resolution image, the high-frequency components are generated via the learned model and added to the initially interpolated low-resolution image. Experiments show that with simple image features our algorithm performs better in terms of accuracy and efficiency with different magnification factors compared to the state-of-the-art methods.
Bir Bhanu
ICIP2
2012 Vehicle logo super-resolution by canonical correlation analysis
abstract
Recognition of a vehicle make is of interest in the fields of law enforcement and surveillance. In this paper, we develop a canonical correlation analysis (CCA) based method for vehicle logo super-resolution to facilitate the recognition of the vehicle make. From a limited number of high-resolution logos, we populate the training dataset for each make using gamma transformations. Given a vehicle logo from low-resolution source (i.e., surveillance or traffic camera recordings), the learned models yield super-resolved results. By matching the low-resolution image and the generated high-resolution images, we select the final output that is closest to the low-resolution image in the histogram of oriented gradients (HOG) feature space. Experimental results show that our approach outperforms the state-of-the-art super-resolution methods in qualitative and quantitative measures. Furthermore, the super-resolved logos help to improve the accuracy in the subsequent recognition tasks significantly.
Ninad Thakoor, Bir Bhanu
ICIP3
2012 Contextual and visual modeling for detection of mild traumatic brain injury in MRI
abstract
Mild traumatic brain injury (mTBI) is difficult to detect as the current tools are qualitative, which can lead to poor diagnosis and treatment. The low contrast appearance of mTBI abnormalities on magnetic resonance (MR) images makes quantification problematic for image processing and analysis techniques. To overcome these difficulties, an algorithm is proposed that takes advantage of subject information and texture information from MR images. A contextual model is developed to simulate the progression of the disease using multiple inputs, such as the time post-injury and the location of injury. Textural features are used along with feature selection for a single MR modality. Results from a probabilistic support vector machine using textural features are fused with the contextual model to obtain a robust estimation of abnormal tissue. A novel rat temporal dataset demonstrates the ability of our approach to outperform other state of the art approaches.
Anthony C. Bianchi, Bir Bhanu, Virginia Donovan, Andre Obenaus
ICIP2
2012 A biologically inspired approach for fusing facial expression and appearance for emotion recognition
abstract
Facial emotion recognition from video is an exemplar case where both humans and computers underperform. In recent emotion recognition competitions, top approaches were using either geometric relationships that best captured facial dynamics or an accurate registration technique to develop appearance features. These two methods capture two different types of facial information similarly to how the human visual system divides information when perceiving faces. In this paper, we propose a biologically-inspired fusion approach that emulates this process. The efficacy of the approach is tested with the Audio/Visual Emotion Challenge 2011 data set, a non-trivial data set where state-of-the-art approaches perform under chance. The proposed approach increases classification rates by 18.5% on publicly available data.
Albert C. Cruz, Bir Bhanu
ICIP2
2012 Semantic-visual concept relatedness and co-occurrences for image retrieval
abstract
This paper introduces a novel approach that allows the retrieval of complex images by integrating visual and semantic concepts. The basic idea consists of three aspects. First, we measure the relatedness of semantic and visual concepts and select the visually separable semantic concepts as elements in the proposed image signature representation. Second, we demonstrate the existence of concept co-occurrence patterns. We propose to uncover those underlying patterns by detecting the communities in a network structure. Third, we leverage the visual and semantic correspondence and the co-occurrence patterns to improve the accuracy and efficiency for image retrieval. We perform experiments on two popular datasets that confirm the effectiveness of our approach.
Linan Feng, Bir Bhanu
ICIP2
2012 Detection of non-dynamic blebbing single unattached Human Embryonic Stem Cells
abstract
Human Embryonic Stem Cells (HESCs) are promising for the treatment of many diseases and for toxicological testing. There is a great interest among biologists to automatically determine the number of various types of cells in a population of mixed morphologies. This study addresses quantification of non-dynamic blebbing single unattached human embryonic stem cells (NDBSU-HESCs) that are in suspension and do not show evidence of blebbing. Current image processing methods are inadequate for detecting these cells in real time. In this paper, we propose a method for NDBSU-HESC detection by using multiple trained classifiers where each classifier eliminates cells with properties unmatched to NDBSU-HESCs. The paper validates the method with many videos captured with live stem cells.
Benjamin Xueqi Guan, Bir Bhanu, Prudence Talbot, Sabrina Lin
ICIP2
2012 Codebook optimization using word activation forces for scene categorization
abstract
Visual codebook based quantization of robust appearance descriptors extracted from local image patches is an effective means of capturing image statistics for texture analysis and natural scene classification. In this paper, based on the newly proposed statistics of word activation forces (WAFs), we optimize the codebook. Currently, codebooks are typically created from a set of training images using a clustering algorithm. However, these codebooks are often functionally limited due to redundancy. We show that WAFs can remove the redundancy efficiently. In the experiment, the proposed method achieved the state-of-the-art performance on the Caltech-101, fifteen natural scene categories and VOC2007 databases. The optimization method also offers insights into the success of several recently proposed images classification approaches, including vector quantization (VQ) coding in the Spatial Pyramid Matching (SPM), sparse coding SPM (ScSPM), and Locality-constrained Linear Coding (LLC).
Qun Li 0002, Honggang Zhang 0002, Jun Guo 0002, Bir Bhanu
ICIP5
2012 Facial emotion recognition with expression energy
abstract
Facial emotion recognition, the inference of an emotion from apparent facial expressions, in unconstrained settings is a typical case where algorithms perform poorly. A property of the AVEC2012 data set is that individuals in testing data are not encountered in training data. In these situations, conventional approaches suffer because models developed from training data cannot properly discriminate unforeseen testing samples. Additional information beyond the feature vectors is required for successful detection of emotions. We propose two similarity metrics that address the problems of a conventional approach: neutral similarity, measuring the intensity of an expression; and temporal similarity, measuring changes in an expression over time. These similarities are taken to be the energy of facial expressions, measured with a SIFT-based warping process. Our method improves correlation by 35.5% over the baseline approach on the frame-level sub-challenge.
Albert C. Cruz, Bir Bhanu, Ninad Thakoor
ICMI2
2012 Multiple local kernel integrated feature selection for image classification
Yu Sun 0007, Bir Bhanu
ICPR2
2012 Face recognition in multi-camera surveillance videos
Bir Bhanu, Songfan Yang
ICPR2
2012 Facial emotion recognition in continuous video
Albert C. Cruz, Bir Bhanu, Ninad Thakoor
ICPR2
2012 Utilizing co-occurrence patterns for semantic concept detection in images
Linan Feng, Bir Bhanu
ICPR2
2012 Single camera multi-person tracking based on crowd simulation
Zhixing Jin, Bir Bhanu
ICPR2
2012 Cluster-Classification Bayesian Networks for head pose estimation
Mehran Kafai, Bir Bhanu
ICPR2
2012 Camera pan/tilt control with multiple trackers
Yiming Li 0007, Bir Bhanu
ICPR2
2012 Zombie Survival Optimization: A swarm intelligence algorithm inspired by zombie foraging
Hoang Thanh Nguyen 0001, Bir Bhanu
ICPR2
2012 Integrated personalized video summarization and retrieval
Hessamoddin Shafeian, Bir Bhanu
ICPR2
2012 Structural signatures for passenger vehicle classification in video
Ninad Thakoor, Bir Bhanu
ICPR2
2012 Reflection Symmetry-Integrated Image Segmentation
abstract
This paper presents a new symmetry-integrated region-based image segmentation method. The method is developed to obtain improved image segmentation by exploiting image symmetry. It is realized by constructing a symmetry token that can be flexibly embedded into segmentation cues. Interesting points are initially extracted from an image by the SIFT operator and they are further refined for detecting the global bilateral symmetry. A symmetry affinity matrix is then computed using the symmetry axis and it is used explicitly as a constraint in a region growing algorithm in order to refine the symmetry of the segmented regions. A multi-objective genetic search finds the segmentation result with the highest performance for both segmentation and symmetry, which is close to the global optimum. The method has been investigated experimentally in challenging natural images and images containing man-made objects. It is shown that the proposed method outperforms current segmentation methods both with and without exploiting symmetry. A thorough experimental analysis indicates that symmetry plays an important role as a segmentation cue, in conjunction with other attributes like color and texture.
Yu Sun 0007, Bir Bhanu
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 Dynamic Bayesian Networks for Vehicle Classification in Video
abstract
Vehicle classification has evolved into a significant subject of study due to its importance in autonomous navigation, traffic analysis, surveillance and security systems, and transportation management. While numerous approaches have been introduced for this purpose, no specific study has been conducted to provide a robust and complete video-based vehicle classification system based on the rear-side view where the camera's field of view is directly behind the vehicle. In this paper, we present a stochastic multiclass vehicle classification system which classifies a vehicle (given its direct rear-side view) into one of four classes: sedan, pickup truck, SUV/minivan, and unknown. A feature set of tail light and vehicle dimensions is extracted which feeds a feature selection algorithm to define a low-dimensional feature vector. The feature vector is then processed by a hybrid dynamic Bayesian network to classify each vehicle. Results are shown on a database of 169 videos for four classes.
Mehran Kafai, Bir Bhanu
IEEE Trans. Ind. Informatics2
2012 Understanding Discrete Facial Expressions in Video Using an Emotion Avatar Image
abstract
Existing video-based facial expression recognition techniques analyze the geometry-based and appearance-based information in every frame as well as explore the temporal relation among frames. On the contrary, we present a new image-based representation and an associated reference image called the emotion avatar image (EAI), and the avatar reference, respectively. This representation leverages the out-of-plane head rotation. It is not only robust to outliers but also provides a method to aggregate dynamic information from expressions with various lengths. The approach to facial expression analysis consists of the following steps: 1) face detection; 2) face registration of video frames with the avatar reference to form the EAI representation; 3) computation of features from EAIs using both local binary patterns and local phase quantization; and 4) the classification of the feature as one of the emotion type by using a linear support vector machine classifier. Our system is tested on the Facial Expression Recognition and Analysis Challenge (FERA2011) data, i.e., the Geneva Multimodal Emotion Portrayal-Facial Expression Recognition and Analysis Challenge (GEMEP-FERA) data set. The experimental results demonstrate that the information captured in an EAI for a facial expression is a very strong cue for emotion inference. Moreover, our method suppresses the person-specific information for emotion and performs well on unseen data.
Songfan Yang, Bir Bhanu
IEEE Trans. Syst. Man Cybern. Part B2
2011 A Psychologically-Inspired Match-Score Fusion Model for Video-Based Facial Expression Recognition
Albert C. Cruz, Bir Bhanu, Songfan Yang
ACII (2)2
2011 Tracking pedestrians with bacterial foraging optimization swarms
abstract
Pedestrian tracking is an important problem with many practical applications in fields such as security, animation, and human computer interaction (HCI). In this paper, we introduce a previously-unexplored swarm intelligence approach to multi-object monocular tracking by using Bacterial Foraging Optimization (BFO) swarms to drive a novel part-based pedestrian appearance tracker. We show that tracking a pedestrian by segmenting the body into parts outperforms popular blob based methods and that using BFO can improve performance over traditional Particle Swarm Optimization and Particle Filter methods.
Hoang Thanh Nguyen 0001, Bir Bhanu
IEEE Congress on Evolutionary Computation2
2011 Facial expression recognition using emotion avatar image
abstract
Existing facial expression recognition techniques analyze the spatial and temporal information for every single frame in a human emotion video. On the contrary, we create the Emotion Avatar Image (EAI) as a single good representation for each video or image sequence for emotion recognition. In this paper, we adopt the recently introduced SIFT flow algorithm to register every frame with respect to an Avatar reference face model. Then, an iterative algorithm is used not only to super-resolve the EAI representation for each video and the Avatar reference, but also to improve the recognition performance. Subsequently, we extract the features from EAIs using both Local Binary Pattern (LBP) and Local Phase Quantization (LPQ). Then the results from both texture descriptors are tested on the Facial Expression Recognition and Analysis Challenge (FERA2011) data, GEMEP-FERA dataset. To evaluate this simple yet powerful idea, we train our algorithm only using the given 155 videos of training data from GEMEP-FERA dataset. The result shows that our algorithm eliminates the person-specific information for emotion and performs well on unseen data.
Songfan Yang, Bir Bhanu
FG2
2011 Prediction and validation of indexing performance for biometrics
abstract
The performance of a recognition system is usually experimentally determined. Therefore, one cannot predict the performance of a recognition system a priori for a new dataset. In this paper, a statistical model to predict the value of k in the rank-k identification rate for a given bio- metric system is presented. Thus, one needs to search only the topmost k match scores to locate the true match object. A geometrical probability distribution is used to model the number of non match scores present in the set of similarity scores. The model is tested in simulation and by using a public dataset. The model is also indirectly validated against the previously published results. The actual results obtained using publicly available database are very close to the predicted results which validates the proposed model.
R. Suresh Kumar, Bir Bhanu, Subir Ghosh, Ninad Thakoor
IJCB2
2011 Improved image super-resolution by Support Vector Regression
abstract
Support Vector Machine (SVM) can construct a hyperplane in a high or infinite dimensional space which can be used for classification. Its regression version, Support Vector Regression (SVR) has been used in various image processing tasks. In this paper, we develop an image super-resolution algorithm based on SVR. Experiments demonstrated that our proposed method with limited training samples outperforms some of the state-of-the-art approaches and during the super-resolution process the model learned by SVR is robust to reconstruct edges and fine details in various testing images.
Bir Bhanu
IJCNN2
2011 Concept Learning with Co-occurrence Network for Image Retrieval
abstract
This paper addresses the problem of concept learning for semantic image retrieval. Two types of semantic concepts are introduced in our system: the individual concept and the scene concept. The individual concepts are explicitly provided in a vocabulary of semantic words, which are the labels or annotations in an image database. Scene concepts are higher level concepts which are defined as potential patterns of co occurrence of individual concepts. Scene concepts exist since some of the individual concepts co-occur frequently across different images. This is similar to human learning where understanding of simpler ideas is generally useful prior to developing more sophisticated ones. Scene concepts can have more discriminative power compared to individual concepts but methods are needed to find them. A novel method for deriving scene concepts is presented. It is based on a weighted concept co-occurrence network (graph) with detected community structure property. An image similarity comparison and retrieval framework is described with the proposed individual and scene concept signature as the image semantic descriptors. Extensive experiments are conducted on a publicly available dataset to demonstrate the effectiveness of our concept learning and semantic image retrieval framework.
Linan Feng, Bir Bhanu
ISM2
2010 Auction protocol for camera active control
abstract
In this paper, we apply the auction-based theories in economics to camera networks. We develop a set of auction protocols to do camera active control (pan/tilt/zoom) intelligently. Unlike the economic auction, the bid price in our case is formulated to have a vector representation, such that when a camera is available to follow multiple objects, we consider the “willingness” of this camera to track a particular object. Most of the computation is decentralized by computing the bid price locally while the final decision is made by a virtual auctioneer based on all the available bids, which is analogous to a real auction in economics. Thus, we can take the advantages of distributed/centralized computation and avoid their pitfalls. The experimental results show that the proposed approach is effective and efficient for dynamically active control based on user defined performance metrics.
Yiming Li 0007, Bir Bhanu
ICIP2
2010 Image retrieval with feature selection and relevance feedback
abstract
This paper proposes a new content based image retrieval (CBIR) system combined with relevance feedback and the online feature selection procedures. A measure of inconsistency from relevance feedback is explicitly used as a new semantic criterion to guide the feature selection. By integrating the user feedback information, the feature selection is able to bridge the gap between low-level visual features and high-level semantic information, leading to the improved image retrieval accuracy. Experimental results show that the proposed method obtains higher retrieval accuracy than a commonly used approach.
Yu Sun 0007, Bir Bhanu
ICIP2
2010 On the Performance of Handoff and Tracking in a Camera Network
abstract
Camera handoff is an important problem when using multiple cameras to follow a number of objects in a video network. However, almost all the handoff techniques rely on a robust tracker. State-of-the-art techniques used to evaluate the performance of camera handoff use either annotated videos or simulated data, and the handoff performance is evaluated in conjunction with a tracker. This does not allow a deeper understanding into the performance of a tracker and a handoff technique separately in the real-world settings. In this paper, we evaluate three camera handoff techniques, two different color-based trackers in seven real-life cases, with varying numbers of cameras, number of objects and the changing environmental conditions. We also perform experiments on annotated videos to provide the ground-truth for all the scenarios. This evaluation of performance isolates the effect of tracking and handoff techniques and clarifies their role in a video network.
Yiming Li 0007, Bir Bhanu
ICPR2
2010 3D Filtering for Injury Detection in Brain MRI
abstract
This paper introduces a brain injury detection approach, using 3D filtering technique, for the images acquired by the magnetic resonance imaging (MRI) technique. The proposed method uses the symmetry property of brain MRI on both 2D images and 3D volumetric information of the MRI sequences. The approach consists of two key steps: (1) each slice of a brain image is segmented into different parts using a region growing algorithm, and a symmetry affinity matrix is computed, (2) non-symmetric regions are extracted, and they are further used to detect brain injury. The Kalman filter is explicitly used in step (2) to filter out the non-injury regions in 3D. Experiments are carried out to indicate the high efficiency of the method to detect the brain injuries.
Yu Sun 0007, Bir Bhanu
ICPR2
2010 3D Human Body Modeling Using Range Data
abstract
For the 3D modeling of walking humans the determination of body pose and extraction of body parts, from the sensed 3D range data, are challenging image processing problems. Real body data may have holes because of self-occlusions and grazing angle views. Most of the existing modeling methods rely on direct fitting a 3D model into the data without considering the fact that the parts in an image are indeed the human body parts. In this paper, we present a method for 3D human body modeling using range data that attempts to overcome these problems. In our approach the entire human body is first decomposed into major body parts by a parts-based image segmentation method, and then a kinematics model is fitted to the segmented body parts in an optimized manner. The fitted model is adjusted by the iterative closest point (ICP) algorithm to resolve the gaps in the body data. Experimental results and comparisons demonstrate the effectiveness of our approach.
Koichiro Yamauchi 0002, Bir Bhanu, Hideo Saito 0001
ICPR2
2010 Age Classification Base on Gait Using HMM
abstract
In this paper we propose a new framework for age classification based on human gait using Hidden Markov Model (HMM). A gait database including young people and elderly people is built. To extract appropriate gait features, we consider a contour related method in terms of shape variations during human walking. Then the image feature is transformed to a lower-dimensional space by using the Frame to Exemplar (FED) distance. A HMM is trained on the FED vector sequences. Thus, the framework provides flexibility in the selection of gait feature representation. In addition, the framework is robust for classification due to the statistical nature of HMM. The experimental results show that video-based automatic age classification from human gait is feasible and reliable.
Yunhong Wang 0001, Bir Bhanu
ICPR3
2010 Incremental Unsupervised Three-Dimensional Vehicle Model Learning From Video
abstract
In this paper, we present a new generic model-based approach for building 3-D models of vehicles from color video from a single uncalibrated traffic-surveillance camera. We propose a novel directional template method that uses trigonometric relations of the 2-D features and geometric relations of a single 3-D generic vehicle model to map 2-D features to 3-D in the face of projection and foreshortening effects. We use novel hierarchical structural similarity measures to evaluate these single-frame-based 3-D estimates with respect to the generic vehicle model. Using these similarities, we adopt a weighted clustering technique to build a 3-D model of the vehicle for the current frame. The 3-D features are then adaptively clustered again over the frame sequence to generate an incremental 3-D model of the vehicle. Results are shown for several simulated and real traffic videos in an uncontrolled setup. Finally, the results are evaluated by the same structural performance measure, underscoring the usefulness of incremental learning. The performance of the proposed method for several types of vehicles in two considerably different traffic spots is very promising to encourage its applicability in 3-D reconstruction of other rigid objects in video.
Nirmalya Ghosh, Bir Bhanu
IEEE Trans. Intell. Transp. Syst.2
2009 Symmetry integrated region-based image segmentation
abstract
Symmetry is an important cue for machine perception that involves high-level knowledge of image components. Unlike most of the previous research that only computes symmetry in an image, this paper integrates symmetry with image segmentation to improve the segmentation performance. The symmetry integration is used to optimize both the segmentation and the symmetry of regions simultaneously. Interesting points are initially extracted from an image and they are further refined for detecting symmetry axis. A symmetry affinity matrix is used explicitly as a constraint in a region growing algorithm in order to refine the symmetry of segmented regions. Experimental results and comparisons from a wide domain of images indicate a promising improvement by symmetry integrated image segmentation compared to other image segmentation methods that do not exploit symmetry.
Yu Sun 0007, Bir Bhanu
CVPR2
2009 Tracking multiple objects in non-stationary video
abstract
One of the key problems in computer vision and pattern recognition is tracking. Multiple objects, occlusion, and tracking moving objects using a moving camera are some of the challenges that one may face in developing an ef-fective approach for tracking. While there are numerous algorithms and approaches to the tracking problem with their own shortcomings, a less-studied approach considers swarm intelligence. Swarm intelligence algorithms are often suited for optimization problems, but require advancements for tracking objects in video. This paper presents an im-proved algorithm based on Bacterial Foraging Optimization in order to track multiple objects in real-time video exposed to full and partial occlusion, using video from both fixed and moving cameras. A comparison with various algorithms is provided.
Hoang Thanh Nguyen 0001, Bir Bhanu
GECCO2
2009 Task-oriented camera assignment in a video network
abstract
Camera assignment and hand-off are some of the key image processing problems in a video network. In this paper, we propose a new approach for camera assignment and hand-off in a video network. The camera assignment problem is modeled as a weakly acyclic game which allows the design of utility functions based on different user-supplied criteria. A theoretical and experimental comparison of the proposed approach with the two recently proposed approaches based on potential game theory and constraint satisfaction problem is provided. This comparison shows that the proposed approach is theoretically more general and computationally more efficient than the other approaches.
Yiming Li 0007, Bir Bhanu
ICIP2
2009 Multi-object tracking in non-stationary video using bacterial foraging swarms
abstract
One of the key problems in the field of image processing is object tracking in video. Multiple objects, occlusion, and non-stationary video are some of the challenges that one may face in developing an effective approach. A less-studied approach considers swarm intelligence. This paper presents a new and improved algorithm based on Bacterial Foraging Optimization in order to track multiple objects in real-time video exposed to full and partial occlusion, using video from a moving camera. A comparison with various algorithms is provided.
Hoang Thanh Nguyen 0001, Bir Bhanu
ICIP2
2009 Symmetry-integrated injury detection for brain MRI
abstract
This paper presents a new brain injury detection approach in images acquired by magnetic resonance imaging (MRI). The proposed approach is based on the fact that the anatomical structure of a 2D brain is highly symmetric, while most of the injury in the brain generally indicates asymmetry. The approach starts from symmetry integrated region growing segmentation of the brain images using the symmetry affinity matrix, and candidate asymmetric regions are initially extracted using kurtosis and skewness of symmetry affinity matrix. An expectation maximum classifier with Gaussian mixture model is used explicitly to classify asymmetric regions into injury and non-injury. Experimental results are carried out to demonstrate the efficacy of the approach for injury detection.
Yu Sun 0007, Bir Bhanu, Shiv Bhanu
ICIP2
2009 Efficient Recognition of Highly Similar 3D Objects in Range Images
abstract
Most existing work in 3D object recognition in computer vision has been on recognizing dissimilar objects using a small database. For rapid indexing and recognition of highly similar objects, this paper proposes a novel method which combines the feature embedding for the fast retrieval of surface descriptors, novel similarity measures for correspondence and a support vector machine (SVM)-based learning technique for ranking the hypotheses. The local surface patch (LSP) representation is used to find the correspondences between a model-test pair. Due to its high dimensionality, an embedding algorithm is used that maps the feature vectors to a low-dimensional space where distance relationships are preserved. By searching the nearest neighbors in low dimensions, the similarity between a model-test pair is computed using the novel features. The similarities for all model-test pairs are ranked using the learning algorithm to generate a short list of candidate models for verification. The verification is performed by aligning a model with the test object. The experimental results, on the UND dataset (302 subjects with 604 images) and the UCR dataset (155 subjects with 902 images) that contain 3D human ears, are presented and compared with the geometric hashing technique to demonstrate the efficiency and effectiveness of the proposed approach.
Hui Chen 0019, Bir Bhanu
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Super-Resolution of Facial Images in Video with Expression Changes
abstract
Super-resolution (SR) of facial images from video suffers from facial expression changes. Most of the existing SR algorithms for facial images make an unrealistic assumption that the ¿perfect¿ registration has been done prior to the SR process. However, the registration is a challenging task for SR with expression changes. This paper proposes a new method for enhancing the resolution of low-resolution (LR) facial image by handling the facial image in a non-rigid manner. It consists of global tracking, local alignment for precise registration and SR algorithms. A B-spline based resolution aware incremental free form deformation (RAIFFD) model is used to recover a dense local non-rigid flow field. In this scheme, low-resolution image model is explicitly embedded in the optimization function formulation to simulate the formation of low resolution image. The results achieved by the proposed approach are significantly better as compared to the SR approaches applied on the whole face image without considering local deformations. The results are also compared with two state-of-the-art SR algorithms to show the effectiveness of the approach in super-resolving facial images with local expression changes.
Jiangang Yu, Bir Bhanu
AVSS2
2008 Bayesian based 3D shape reconstruction from video
abstract
In a video sequence with a 3D rigid object moving, changing shapes of the 2D projections provide interrelated spatio-temporal cues for incremental 3D shape reconstruction. This paper describes a probabilistic approach for intelligent view-integration to build 3D model of vehicles from traffic videos collected from an uncalibrated static camera. The proposed Bayesian net framework allows the handling of uncertainties in a systematic manner. The performance is verified with several types of vehicles in different videos.
Nirmalya Ghosh, Bir Bhanu
ICIP2
2008 Super-resolution of deformed facial images in video
abstract
Super-resolution (SR) of facial images from video suffers from facial expression changes. Most of the existing SR algorithms for facial images make an unrealistic assumption that the “perfect” registration has been done prior to the SR process. However, the registration is a challenging task for SR with expression changes. This paper proposes a new method for enhancing the resolution of low-resolution (LR) facial image by handling the facial image in a non-rigid manner. It consists of global tracking, local alignment for precise registration and SR algorithms. A B-spline based Resolution Aware Incremental Free Form Deformation (RAIFFD) model is used to recover a dense local non-rigid flow field. In this scheme, low-resolution image model is explicitly embedded in the optimization function formulation to simulate the formation of low resolution image. The results achieved by the proposed approach are significantly better as compared to the SR approaches applied on the whole face image without considering local deformations.
Jiangang Yu, Bir Bhanu
ICIP2
2008 Anomalous activity classification in the distributed camera network
abstract
Unlike existing methods that used the human actions or trajectories to analyze the human activity in overlapping field-of-views, this paper proposes the appearance and travel time-based human activity classification in the camera network of non-overlapping field-of-views. The mixture of Gaussian-based appearance similarity model incorporates the appearance variance between different cameras to address changes in varying lighting conditions. To address the problem of limited labeled training data, we propose the use of semi-supervised expectation-maximization algorithm for activity classification. The human activities observed in a simulated camera network with nine cameras and twenty-five nodes are classified into one normal and three anomalous classes. A similar camera network is built and tested in real-life experiments, in which the proposed approach achieves satisfactory performance.
Xiaotao Zou, Bir Bhanu
ICIP2
2008 How current BNs fail to represent evolvable pattern recognition problems and a proposed solution
abstract
In the real world, systems/processes often evolve without fixed and predictable dynamic models. To represent such applications we need uncertainty models, like Bayesian Nets (BN) that are formed online and in a self-evolving data-driven way. But current BN frameworks cannot handle simultaneous scalability in the model structure and causal relations. We show how current BNs fail in different applications from several fields, ranging from computer vision to database retrieval to medical diagnostics. We propose a novel Structure Modifiable Adaptive Reason-building Temporal Bayesian Networks (SmartBN) that has scalability for uncertainty in both, structures and causal relations. We evaluate its performance for a 3D model building application for vehicles in traffic video.
Nirmalya Ghosh, Bir Bhanu
ICPR2
2008 Feature fusion of side face and gait for video-based human identification
Bir Bhanu
Pattern Recognit.2
2008 Long-Term Cross-Session Relevance Feedback Using Virtual Features
abstract
Relevance feedback (RF) is an iterative process, which refines the retrievals by utilizing the user's feedback on previously retrieved results. Traditional RF techniques solely use the short-term learning experience and do not exploit the knowledge created during cross sessions with multiple users. In this paper, we propose a novel RF framework, which facilitates the combination of short-term and long-term learning processes by integrating the traditional methods with a new technique called the virtual feature. The feedback history with all the users is digested by the system and is represented in a very efficient form as a virtual feature of the images. As such, the dissimilarity measure can dynamically be adapted, depending on the estimate of the semantic relevance derived from the virtual features. In addition, with a dynamic database, the user's subject concepts may transit from one to another. By monitoring the changes in retrieval performance, the proposed system can automatically adapt the concepts according to the new subject concepts. The experiments are conducted on a real image database. The results manifest that the proposed framework outperforms the traditional within-session and log-based long-term RF techniques.
Peng-Yeng Yin, Bir Bhanu, Kuang-Cheng Chang, Anlei Dong
IEEE Trans. Knowl. Data Eng.2
2008 Feature synthesized EM algorithm for image retrieval
abstract
As a commonly used unsupervised learning algorithm in Content-Based Image Retrieval (CBIR), Expectation-Maximization (EM) algorithm has several limitations, including the curse of dimensionality and the convergence at a local maximum. In this article, we propose a novel learning approach, namely Coevolutionary Feature Synthesized Expectation-Maximization (CFS-EM), to address the above problems. The CFS-EM is a hybrid of coevolutionary genetic programming (CGP) and EM algorithm applied on partially labeled data. CFS-EM is especially suitable for image retrieval because the images can be searched in the synthesized low-dimensional feature space, while a kernel-based method has to make classification computation in the original high-dimensional space. Experiments on real image databases show that CFS-EM outperforms Radial Basis Function Support Vector Machine (RBF-SVM), CGP, Discriminant-EM (D-EM) and Transductive-SVM (TSVM) in the sense of classification performance and it is computationally more efficient than RBF-SVM in the query phase.
Rui Li 0083, Bir Bhanu, Anlei Dong
ACM Trans. Multim. Comput. Commun. Appl.2
2007 On the Performance Prediction and Validation for Multisensor Fusion
abstract
Multiple sensors are commonly fused to improve the detection and recognition performance of computer vision and pattern recognition systems. The traditional approach to determine the optimal sensor combination is to try all possible sensor combinations by performing exhaustive experiments. In this paper, we present a theoretical approach that predicts the performance of sensor fusion that allows us to select the optimal combination. We start with the characteristics of each sensor by computing the match score and non-match score distributions of objects to be recognized. These distributions are modeled as a mixture of Gaussians. Then, we use an explicit Φ transformation that maps a receiver operating characteristic (ROC) curve to a straight line in 2-D space whose axes are related to the false alarm rate (FAR) and the Hit rate (Hit). Finally, using this representation, we derive a set of metrics to evaluate the sensor fusion performance and find the optimal sensor combination. We verify our prediction approach on the publicly available XM2VTS database as well as other databases.
Rong Wang 0007, Bir Bhanu
CVPR2
2007 Hybrid coevolutionary algorithms vs. SVM algorithms
abstract
As a learning method support vector machine is regarded as one of the best classifiers with a strong mathematical foundation. On the other hand, evolutionary computational technique is characterized as a soft computing learning method with its roots in the theory of evolution. During the past decade, SVM has been commonly used as a classifier for various applications. The evolutionary computation has also attracted a lot of attention in pattern recognition and has shown significant performance improvement on a variety of applications. However, there has been no comparison of the two methods. In this paper, first we propose an improvement of a coevolutionary computational classification algorithm, called Improved Coevolutionary Feature Synthesized EM (I-CFS-EM) algorithm. It is a hybrid of coevolutionary genetic programming and EM algorithm applied on partially labeled data. It requires less labeled data and it makes the test in a lower dimension, which speeds up the testing. Then, we provide a comprehensive comparison between SVM with different kernel functions and I-CFS-EM on several real datasets. This comparison shows that I-CFS-EM outperforms SVM in the sense of both the classification performance and the computational efficiency in the testing phase. We also give an intensive analysis of the pros and cons of both approaches.
Rui Li 0083, Bir Bhanu, Krzysztof Krawiec
GECCO2
2007 On the number of subpopulations in coevolutionary computation: a database application
abstract
Among the existing feature selection/synthesis approaches, Coevolutionary Feature Synthesis (CFS) based on Coevolutionary Genetic Programming (CGP) has shown good performance on a variety of applications. In this paper, we propose an MDL-based fitness function to help pick a reasonable number of synthesized features which is equal to the number of subpopulations. It naturally balances the feature transformation complexity and classification performance. Experiments on a real image database show that the new fitness function solves the problem quite well.al.
Rui Li 0083, Bir Bhanu, Krzysztof Krawiec
GECCO2
2007 Super-Resolved Facial Texture Under Changing Pose and Illumination
abstract
In this paper, we propose a method to incrementally super-resolve 3D facial texture by integrating information frame by frame from a video captured under changing poses and illuminations. First, we recover illumination, 3D motion and shape parameters from our tracking algorithm. This information is then used to super-resolve 3D texture using iterative back-projection (IBP) method. Finally, the super-resolved texture is fed back to the tracking part to improve the estimation of illumination and motion parameters. This closed-loop process continues to refine the texture as new frames come in. We also propose a local-region based scheme to handle non-rigidity of the human face. Experiments demonstrate that our framework not only incrementally super-resolves facial images, but recovers the detailed expression changes in high quality.
Jiangang Yu, Bir Bhanu, Yilei Xu, Amit K. Roy-Chowdhury
ICIP (3)2
2007 Determining Topology in a Distributed Camera Network
abstract
Recently, 'entry/exit' events of objects in the field-of-views of cameras were used to learn the topology of the camera network. The integration of object appearance was also proposed to employ the visual information provided by the imaging sensors. A problem with these methods is the lack of robustness to appearance changes. This paper integrates face recognition in the statistical model to better estimate the correspondence in the time-varying network. The statistical dependence between the entry and exit nodes indicates the connectivity and traffic patterns of the camera network, which are represented by a weighted directed graph and transition time distributions. A nine-camera network with 25 nodes is analyzed both in simulation and in real-life experiments, and compared with the previous approaches.
Xiaotao Zou, Bir Bhanu, Bi Song, Amit K. Roy-Chowdhury
ICIP (5)2
2007 Human Ear Recognition in 3D
abstract
Human ear is a new class of relatively stable biometrics that has drawn researchers' attention recently. In this paper, we propose a complete human recognition system using 3D ear biometrics. The system consists of 3D ear detection, 3D ear identification, and 3D ear verification. For ear detection, we propose a new approach which uses a single reference 3D ear shape model and locates the ear helix and the antihelix parts in registered 2D color and 3D range images. For ear identification and verification using range images, two new representations are proposed. These include the ear helix/antihelix representation obtained from the detection algorithm and the local surface patch (LSP) representation computed at feature points. A local surface descriptor is characterized by a centroid, a local surface type, and a 2D histogram. The 2D histogram shows the frequency of occurrence of shape index values versus the angles between the normal of reference feature point and that of its neighbors. Both shape representations are used to estimate the initial rigid transformation between a gallery-probe pair. This transformation is applied to selected locations of ears in the gallery set and a modified Iterative Closest Point (ICP) algorithm is used to iteratively refine the transformation to bring the gallery ear and probe ear into the best alignment in the sense of the least root mean square error. The experimental results on the UCR data set of 155 subjects with 902 images under pose variations and the University of Notre Dame data set of 302 subjects with time-lapse gallery-probe pairs are presented to compare and demonstrate the effectiveness of the proposed algorithms and the system.
Hui Chen 0019, Bir Bhanu
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Fusion of color and infrared video for moving human detection
Ju Han, Bir Bhanu
Pattern Recognit.2
2007 3D free-form object recognition in range images using local surface patches
Hui Chen 0019, Bir Bhanu
Pattern Recognit. Lett.2
2007 Predicting fingerprint biometrics performance from a small gallery
Rong Wang 0007, Bir Bhanu
Pattern Recognit. Lett.2
2007 Visual Learning by Evolutionary and Coevolutionary Feature Synthesis
abstract
In this paper, we present a novel method for learning complex concepts/hypotheses directly from raw training data. The task addressed here concerns data-driven synthesis of recognition procedures for real-world object recognition. The method uses linear genetic programming to encode potential solutions expressed in terms of elementary operations, and handles the complexity of the learning task by applying cooperative coevolution to decompose the problem automatically at the genotype level. The training coevolves feature extraction procedures, each being a sequence of elementary image processing and computer vision operations applied to input images. Extensive experimental results show that the approach attains competitive performance for three-dimensional object recognition in real synthetic aperture radar imagery.
Krzysztof Krawiec, Bir Bhanu
IEEE Trans. Evol. Comput.2
2007 Guest Editorial: Special Issue on Human Detection and Recognition
abstract
The 12 regular papers and three correspondences in this special issue focus on human detection and recognition. The papers represent gait, face (3-D, 2-D, video), iris, palmprint, cardiac sounds, and vulnerability of biometrics and protection against the spoof attacks.
Bir Bhanu, Nalini K. Ratha, B. V. K. Vijaya Kumar, Rama Chellappa, Josef Bigün
IEEE Trans. Inf. Forensics Secur.1
2007 Integrating Face and Gait for Human Recognition at a Distance in Video
abstract
This paper introduces a new video-based recognition method to recognize noncooperating individuals at a distance in video who expose side views to the camera. Information from two biometrics sources, side face and gait, is utilized and integrated for recognition. For side face, an enhanced side-face image (ESFI), a higher resolution image compared with the image directly obtained from a single video frame, is constructed, which integrates face information from multiple video frames. For gait, the gait energy image (GEI), a spatio-temporal compact representation of gait in video, is used to characterize human-walking properties. The features of face and gait are obtained separately using the principal component analysis and multiple discriminant analysis combined method from ESFI and GEI, respectively. They are then integrated at the match score level by using different fusion strategies. The approach is tested on a database of video sequences, corresponding to 45 people, which are collected over seven months. The different fusion methods are compared and analyzed. The experimental results show that: 1) the idea of constructing ESFI from multiple frames is promising for human recognition in video, and better face features are extracted from ESFI compared to those from the original side-face images (OSFIs); 2) the synchronization of face and gait is not necessary for face template ESFI and gait template GEI; the synthetic match scores combine information from them; and 3) an integrated information from side face and gait is effective for human recognition in video.
Bir Bhanu
IEEE Trans. Syst. Man Cybern. Part B2
2006 Individual Recognition Using Gait Energy Image
abstract
In this paper, we propose a new spatio-temporal gait representation, called Gait Energy Image (GEI), to characterize human walking properties for individual recognition by gait. To address the problem of the lack of training templates, we also propose a novel approach for human recognition by combining statistical gait features from real and synthetic templates. We directly compute the real templates from training silhouette sequences, while we generate the synthetic templates from training sequences by simulating silhouette distortion. We use a statistical approach for learning effective features from real and synthetic templates. We compare the proposed GEI-based gait recognition approach with other gait recognition approaches on USF HumanID Database. Experimental results show that the proposed GEI is an effective and efficient gait representation for individual recognition, and the proposed approach achieves highly competitive performance with respect to the published gait recognition approaches.
Ju Han, Bir Bhanu
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 Fingerprint matching by genetic algorithms
Xuejun Tan, Bir Bhanu
Pattern Recognit.2
2006 Evolutionary feature synthesis for facial expression recognition
Jiangang Yu, Bir Bhanu
Pattern Recognit. Lett.2
2005 Hierarchical multi-sensor image registration using evolutionary computation
abstract
Image registration between multi-sensor imagery is a challenging problem due to the difficulties associated with finding a correspondence between pixels from images taken by the two sensors. However, the moving people in a static scene provide cues to address this problem. In this paper, we propose a hierarchical approach to automatically find the correspondence between the preliminary human silhouettes extracted from synchronous color and infrared (IR) image sequences for image registration using evolutionary computation. The proposed approach reduces the overall computational load without decreasing the final estimation accuracy. Experimental results show that the proposed approach achieves good results for image registration between color and IR imagery.
Ju Han, Bir Bhanu
GECCO2
2005 Learning Models for Predicting Recognition Performance
abstract
This paper addresses one of the fundamental problems encountered in performance prediction for object recognition. In particular we address the problems related to estimation of small gallery size that can give good error estimates and their confidences on large probe sets and populations. We use a generalized two-dimensional prediction model that integrates a hypergeometric probability distribution model with a binomial model explicitly and considers the distortion problem in large populations. We incorporate learning in the prediction process in order to find the optimal small gallery size and to improve its performance. The Chernoff and Chebychev inequalities are used as a guide to obtain the small gallery size. During the prediction we use the expectation-maximum (EM) algorithm to learn the match score and the non-match score distributions (the number of components, their weights, means and covariances) that are represented as Gaussian mixtures. By learning we find the optimal size of small gallery and at the same time provide the upper bound and the lower bound for the prediction on large populations. Results are shown using real-world databases
Rong Wang 0007, Bir Bhanu
ICCV2
2005 A study on view-insensitive gait recognition
abstract
Most gait recognition approaches only study human walking frontoparallel to the image plane which is not realistic in video surveillance applications. Human gait appearance depends on various factors including locations of the camera and the person, the camera axis and the walking direction. By analyzing these factors, we propose a statistical approach for view-insensitive gait recognition. The proposed approach recognizes human using a single camera, and avoids the difficulties of recovering the human body structure and camera calibration. Experimental results show that the proposed approach achieves good performance in recognizing individuals walking along different directions.
Ju Han, Bir Bhanu, Amit K. Roy-Chowdhury
ICIP (3)2
2005 Coevolutionary feature synthesized EM algorithm for image retrieval
abstract
As a commonly used unsupervised learning algorithm in Content-Based Image Retrieval (CBIR), Expectation-Maximization (EM) algorithm has several limitations, especially in high dimensional feature spaces where the data are limited and the computational cost varies exponentially with the number of feature dimensions. Moreover, the convergence is guaranteed only at a local maximum. In this paper, we propose a unified framework of a novel learning approach, namely Coevolutionary Feature Synthesized Expectation-Maximization (CFS-EM), to achieve satisfactory learning in spite of these difficulties. The CFS-EM is a hybrid of coevolutionary genetic programming (CGP) and EM algorithm. The advantages of CFS-EM are: 1) it synthesizes low-dimensional features based on CGP algorithm, which yields near optimal nonlinear transformation and classification precision comparable to kernel methods such as the support vector machine (SVM); 2) the explicitness of feature transformation is especially suitable for image retrieval because the images can be searched in the synthesized low-dimensional space, while kernel-based methods have to make classification computation in the original high-dimensional space; 3) the unlabeled data can be boosted with the help of the class distribution learning using CGP feature synthesis approach. Experimental results show that CFS-EM outperforms pure EM and CGP alone, and is comparable to SVM in the sense of classification. It is computationally more efficient than SVM in query phase. Moreover, it has a high likelihood that it will jump out of a local maximum to provide near optimal results and a better estimation of parameters.
Rui Li 0083, Bir Bhanu, Anlei Dong
ACM Multimedia2
2005 Integrating Relevance Feedback Techniques for Image Retrieval Using Reinforcement Learning
abstract
Relevance feedback (RF) is an interactive process which refines the retrievals to a particular query by utilizing the user's feedback on previously retrieved results. Most researchers strive to develop new RF techniques and ignore the advantages of existing ones. In this paper, we propose an image relevance reinforcement learning (IRRL) model for integrating existing RF techniques in a content-based image retrieval system. Various integration schemes are presented and a long-term shared memory is used to exploit the retrieval experience from multiple users. Also, a concept digesting method is proposed to reduce the complexity of storage demand. The experimental results manifest that the integration of multiple RF approaches gives better retrieval performance than using one RF technique alone, and that the sharing of relevance knowledge between multiple query sessions significantly improves the performance. Further, the storage demand is significantly reduced by the concept digesting technique. This shows the scalability of the proposed model with the increasing-size of database.
Peng-Yeng Yin, Bir Bhanu, Kuang-Cheng Chang, Anlei Dong
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Performance prediction for individual recognition by gait
Ju Han, Bir Bhanu
Pattern Recognit. Lett.2
2005 Active concept learning in image databases
abstract
Concept learning in content-based image retrieval systems is a challenging task. This paper presents an active concept learning approach based on the mixture model to deal with the two basic aspects of a database system: the changing (image insertion or removal) nature of a database and user queries. To achieve concept learning, we a) propose a new user directed semi-supervised expectation-maximization algorithm for mixture parameter estimation, and b) develop a novel model selection method based on Bayesian analysis that evaluates the consistency of hypothesized models with the available information. The analysis of exploitation versus exploration in the search space helps to find the optimal model efficiently. Our concept knowledge transduction approach is able to deal with the cases of image insertion and query images being outside the database. The system handles the situation where users may mislabel images during relevance feedback. Experimental results on Corel database show the efficacy of our active concept learning approach and the improvement in retrieval performance by concept transduction.
Anlei Dong, Bir Bhanu
IEEE Trans. Syst. Man Cybern. Part B2
2005 Visual learning by coevolutionary feature synthesis
abstract
In this paper, a novel genetically inspired visual learning method is proposed. Given the training raster images, this general approach induces a sophisticated feature-based recognition system. It employs the paradigm of cooperative coevolution to handle the computational difficulty of this task. To represent the feature extraction agents, the linear genetic programming is used. The paper describes the learning algorithm and provides a firm rationale for its design. Different architectures of recognition systems are considered that employ the proposed feature synthesis method. An extensive experimental evaluation on the demanding real-world task of object recognition in synthetic aperture radar (SAR) imagery shows the ability of the proposed approach to attain high recognition performance in different operating conditions.
Krzysztof Krawiec, Bir Bhanu
IEEE Trans. Syst. Man Cybern. Part B2
2005 Object detection via feature synthesis using MDL-based genetic programming
abstract
In this paper, we use genetic programming (GP) to synthesize composite operators and composite features from combinations of primitive operations and primitive features for object detection. The motivation for using GP is to overcome the human experts' limitations of focusing only on conventional combinations of primitive image processing operations in the feature synthesis. GP attempts many unconventional combinations that in some cases yield exceptionally good results. To improve the efficiency of GP and prevent its well-known code bloat problem without imposing severe restriction on the GP search, we design a new fitness function based on minimum description length principle to incorporate both the pixel labeling error and the size of a composite operator into the fitness evaluation process. To further improve the efficiency of GP, smart crossover, smart mutation and a public library ideas are incorporated to identify and keep the effective components of composite operators. Our experiments, which are performed on selected training regions of a training image to reduce the training time, show that compared to normal GP, our GP algorithm finds effective composite operators more quickly and the learned composite operators can be applied to the whole training image and other similar testing images. Also, compared to a traditional region-of-interest extraction algorithm, the composite operators learned by GP are more effective and efficient for object detection.
Yingqiang Lin, Bir Bhanu
IEEE Trans. Syst. Man Cybern. Part B2
2005 Evolutionary feature synthesis for object recognition
abstract
Features represent the characteristics of objects and selecting or synthesizing effective composite features are the key to the performance of object recognition. In this paper, we propose a coevolutionary genetic programming (CGP) approach to learn composite features for object recognition. The knowledge about the problem domain is incorporated in primitive features that are used in the synthesis of composite features by CGP using domain-independent primitive operators. The motivation for using CGP is to overcome the limitations of human experts who consider only a small number of conventional combinations of primitive features during synthesis. CGP, on the other hand, can try a very large number of unconventional combinations and these unconventional combinations yield exceptionally good results in some cases. Our experimental results with real synthetic aperture radar (SAR) images show that CGP can discover good composite features to distinguish objects from clutter and to distinguish among objects belonging to several classes. The comparison with other classical classification algorithms is favorable to the CGP-based approach proposed in this paper.
Yingqiang Lin, Bir Bhanu
IEEE Trans. Syst. Man Cybern. Part C2
2005 Fingerprint classification based on learned features
abstract
In this paper, we present a fingerprint classification approach based on a novel feature-learning algorithm. Unlike current research for fingerprint classification that generally uses well defined meaningful features, our approach is based on Genetic Programming (GP), which learns to discover composite operators and features that are evolved from combinations of primitive image processing operations. Our experimental results show that our approach can find good composite operators to effectively extract useful features. Using a Bayesian classifier, without rejecting any fingerprints from the NIST-4 database, the correct rates for 4- and 5-class classification are 93.3% and 91.6%, respectively, which compare favorably with other published research and are one of the best results published to date.
Xuejun Tan, Bir Bhanu, Yingqiang Lin
IEEE Trans. Syst. Man Cybern. Part C2
2004 Statistical Feature Fusion for Gait-Based Human Recognition
Ju Han, Bir Bhanu
CVPR (2)2
2004 Feature Synthesis Using Genetic Programming for Face Expression Recognition
Bir Bhanu, Jiangang Yu, Xuejun Tan, Yingqiang Lin
GECCO (2)1
2004 Cooperative Coevolution Fusion for Moving Object Detection
Sohail Nadimi, Bir Bhanu
GECCO (1)2
2004 Physical Models for Moving Shadow and Object Detection in Video
abstract
Current moving object detection systems typically detect shadows cast by the moving object as part of the moving object. In this paper, the problem of separating moving cast shadows from the moving objects in an outdoor environment is addressed. Unlike previous work, we present an approach that does not rely on any geometrical assumptions such as camera location and ground surface/object geometry. The approach is based on a new spatio-temporal albedo test and dichromatic reflection model and accounts for both the sun and the sky illuminations. Results are presented for several video sequences representing a variety of ground materials when the shadows are cast on different surface types. These results show that our approach is robust to widely different background and foreground materials, and illuminations.
Sohail Nadimi, Bir Bhanu
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Functional template-based SAR image segmentation
Bir Bhanu, Stephanie Fonder
Pattern Recognit.1
2004 Synthesizing feature agents using evolutionary computation
Bir Bhanu, Yingqiang Lin
Pattern Recognit. Lett.1
2003 Fingerprint Identification: Classification vs. Indexing
abstract
We present a comparison of two key approaches for fingerprint identification. These approaches are based on (a) classification followed by verification, and (b) indexing followed by verification. The fingerprint classification approach is based on a novel feature-learning algorithm. It learns to discover composite operators and features that are evolved from combinations of primitive image processing operations. These features are then used for classification of fingerprints into five classes. The indexing approach is based on novel triplets of minutiae. The verification algorithm, based on least square minimization over each of the possible minutiae triplet pairs, is used for identification in both cases. On the NIST-4 fingerprint database, the comparison shows that, although correct classification rate can be as high as 92.8% for 5-class problems, the indexing approach performs better, based on the size of the search space and identification results.
Xuejun Tan, Bir Bhanu, Yingqiang Lin
AVSS2
2003 A New Semi-Supervised EM Algorithm for Image Retrieval
abstract
One of the main tasks in content-based image retrieval (CBIR) is to reduce the gap between low-level visual features and high-level human concepts. This paper presents a new semi-supervised EM algorithm (NSSEM), where the image distribution in feature space is modeled as a mixture of Gaussian densities. Due to the statistical mechanism of accumulating and processing meta knowledge, the NSS-EM algorithm with long term learning of mixture model parameters can deal with the cases where users may mislabel images during relevance feedback. Our approach that integrates mixture model of the data, relevance feedback and long term learning helps to improve retrieval performance. The concept learning is incrementally refined with increased retrieval experiences. Experiment results on Corel database show the efficacy of our proposed concept learning approach.
Anlei Dong, Bir Bhanu
CVPR (2)2
2003 On The Fundamental Performance For Fingerprint Matching
abstract
Fingerprints have long been used for person authentication. However, there is not enough scientific research to explain the probability that two fingerprints, which are impressions of different fingers, may be taken as the same one. In this paper, we propose a formal framework to estimate the fundamental algorithm independent error rate of fingerprint matching. Unlike a previous work, which assumes that there is no overlap between any two minutiae uncertainty areas and only measures minutiae's positions and orientations. In our model, we do not make this assumption and measure the relations, i.e., ridge counts between different minutiae as well as minutiae's positions and orientations. The error rates of fingerprint matching obtained by our approach is significantly lower than that of previously published research. Results are shown using NIST-4 fingerprint database. These results contribute toward making fingerprint matching a science and settling the legal challenges to fingerprints.
Xuejun Tan, Bir Bhanu
CVPR (2)2
2003 Coevolution and Linear Genetic Programming for Visual Learning
Krzysztof Krawiec, Bir Bhanu
GECCO2
2003 Learning Features for Object Recognition
Yingqiang Lin, Bir Bhanu
GECCO2
2003 Active Concept Learning for Image Retrieval in Dynamic Databases
abstract
Concept learning in content-based image retrieval (CBIR) systems is a challenging task. We present an active concept learning approach based on mixture model to deal with the two basic aspects of a database system: changing (image insertion or removal) nature of a database and user queries. To achieve concept learning, we develop a novel model selection method based on Bayesian analysis that evaluates the consistency of hypothesized models with the available information. The analysis of exploitation vs. exploration in the search space helps to find optimal model efficiently. Experimental results on Corel database show the efficacy of our approach.
Anlei Dong, Bir Bhanu
ICCV2
2003 Reinforcement Learning for Combining Relevance Feedback Techniques
abstract
Relevance feedback (RF) is an interactive process which refines the retrievals by utilizing user's feedback history. Most researchers strive to develop new RF techniques and ignore the advantages of existing ones. We propose an image relevance reinforcement learning (IRRL) model for integrating existing RF techniques. Various integration schemes are presented and a long-term shared memory is used to exploit the retrieval experience from multiple users. Also, a concept digesting method is proposed to reduce the complexity of storage demand. The experimental results manifest that the integration of multiple RF approaches gives better retrieval performance than using one RF technique alone, and that the sharing of relevance knowledge between multiple query sessions also provides significant contributions for improvement. Further, the storage demand is significantly reduced by the concept digesting technique. This shows the scalability of the proposed model against a growing-size database.
Peng-Yeng Yin, Bir Bhanu, Kuang-Cheng Chang, Anlei Dong
ICCV2
2003 Concept learning and transplantation for dynamic image databases
abstract
The task of a content-based image retrieval (CBIR) system is to cater to users who expect to get relevant images with high precision and efficiency in response to query images. This paper presents a concept learning approach that integrates a mixture model of the data, relevance feedback and long-term continuous learning. The concepts are incrementally refined with increased retrieval experiences. The concept knowledge can be immediately transplanted to deal with the dynamic database situations such as insertion of new images, removal of existing images and query images, which are outside the database. Experimental results on Corel database show the efficacy of our approach.
Anlei Dong, Bir Bhanu
ICME2
2003 Visual Learning by Evolutionary Feature Synthesis
Krzysztof Krawiec, Bir Bhanu
ICML2
2003 Probabilistic Spatial Database Operations
Jinfeng Ni, Chinya V. Ravishankar, Bir Bhanu
SSTD3
2003 Genetic algorithm based feature selection for target detection in SAR images
Bir Bhanu, Yingqiang Lin
Image Vis. Comput.1
2003 Guest editorial: Special issue on computer vision beyond the visible spectrum
Ioannis Pavlidis, Bir Bhanu
Image Vis. Comput.2
2003 Fingerprint Indexing Based on Novel Features of Minutiae Triplets
abstract
We are concerned with accurate and efficient indexing of fingerprint images. We present a model-based approach, which efficiently retrieves correct hypotheses using novel features of triangles formed by the triplets of minutiae as the basic representation unit. The triangle features that we use are its angles, handedness, type, direction, and maximum side. Geometric constraints based on other characteristics of minutiae are used to eliminate false correspondences. Experimental results on live-scan fingerprint images of varying quality and NIST special database 4 (NIST-4) show that our indexing approach efficiently narrows down the number of candidate hypotheses in the presence of translation, rotation, scale, shear, occlusion, and clutter. We also perform scientific experiments to compare the performance of our approach with another prominent indexing approach and show that the performance of our approach is better for both the live scan database and the ink based database NIST-4.
Bir Bhanu, Xuejun Tan
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Stochastic models for recognition of occluded targets
Bir Bhanu, Yingqiang Lin
Pattern Recognit.1
2003 A robust two step approach for fingerprint identification
Xuejun Tan, Bir Bhanu
Pattern Recognit. Lett.2
2002 Learning Composite Operators For Object Detection
Bir Bhanu, Yingqiang Lin
GECCO1
2002 Robust fingerprint identification
abstract
Due to the complex distortions involved in two impressions of the same finger, fingerprint identification is still a challenging problem for person authentication. In this paper, we propose a fingerprint identification approach based on the triplets of minutiae. The features that we use to find the potential corresponding triangles include angles, triangle orientation, triangle direction, maximum side, minutiae density and ridge counts. False corresponding triangles are eliminated by applying constraints to the transformation between two potential corresponding triangles. The experimental results on National Institute of Standards and Technology special fingerprint database 4, NIST-4, show that, as compared to the linear search, the proposed approach provides a reduction by a factor of 200 for the number of the hypotheses that need to be considered and it can achieve good performance even when a large portion of fingerprints in the database are of poor quality.
Xuejun Tan, Bir Bhanu
ICIP (1)2
2002 Kinematic based Human Motion Analysis in Infrared Sequences
abstract
In an infrared (IR) image sequence of human walking, the human silhouette can be reliably extracted from the background regardless of lighting conditions and colors of the human surfaces and backgrounds in most cases. Moreover, some important regions containing skin, such as face and hands, can be accurately detected in IR image sequences. In this paper, we propose a kinematic-based approach for automatic human motion analysis from IR image sequences. The proposed approach estimates 3D human walking parameters by performing a modified least squares fit of the 3D kinematic model to the 2D silhouette extracted from a monocular IR image sequence, where continuity and symmetry of human walking and detected hand regions are also considered in the optimization function. Experimental results show that the proposed approach achieves good performance in gait analysis with different view angles With respect to the walking direction, and is promising for further gait recognition.
Bir Bhanu, Ju Han
WACV1
2002 Fingerprint Verification Using Genetic Algorithms
abstract
Fingerprint matching is still a challenging problem for reliable person authentication because of the complex distortions involved in two impressions of the same finger. In this paper, we propose a fingerprint matching approach based on Genetic Algorithms (GA), which finds the optimal global transformation between two different fingerprints. In order to deal with low quality fingerprint images, which introduce significant occlusion and clutter of minutiae features, we design the fitness function based on the local properties of each triplet of minutiae. The experimental results on National Institute of Standards and Technology fingerprint database, NIST-4, not only show that the proposed approach can achieve good performance even when a large portion of fingerprints in the database are of poor quality, but also show that the proposed approach is better than another approach, which is based on mean-squared error estimation.
Xuejun Tan, Bir Bhanu
WACV2
2002 Introduction to the Special Issue on Innovative Applications of Computer Vision
Bir Bhanu, Terrance E. Boult, Alok Gupta, David Michael
Mach. Vis. Appl.1
2002 Model-based recognition of articulated objects
Joon Ahn, Bir Bhanu
Pattern Recognit. Lett.2
2001 Learned Templates for Feature Extraction in Fingerprint Images
abstract
Most current techniques for minutiae extraction in fingerprint images utilize complex preprocessing and postprocessing. In this paper, we propose a new technique, based on the use of learned templates, which statistically characterize the minutiae. Templates are teamed from examples by optimizing a criterion function using Lagrange's method. To detect the presence of minutiae in test images, templates are applied with appropriate orientations to the binary image only at selected potential minutia locations. Several performance measures, which evaluate the quality and quantity of extracted features and their impact on identification, are used to evaluate the significance of learned templates. The performance of the proposed approach is evaluated on two sets of fingerprint images: one is collected by an optical scanner and the other one is chosen from NIST special fingerprint database 4. The experimental results show that learned templates can improve both the features and the performance of the identification system.
Bir Bhanu, Xuejun Tan
CVPR (2)1
2001 Real Time Robot Learning
abstract
This paper presents the design, implementation and testing of a real-time system using computer vision and machine learning techniques to demonstrate learning behavior in a miniature mobile robot. The miniature robot, through environmental sensing, learns to navigate a maze choosing the optimum route. Several reinforcement learning based algorithms, such as the Q-learning, Q(/spl lambda/)-learning, fast online Q(/spl lambda/)-learning and DYNA structure, are considered. Experimental results based on simulation and an integrated real-time system are presented for varying density of obstacles in a 15/spl times/15 maze.
Bir Bhanu, Pat Leang, Chris Cowden, Yingqiang Lin, Mark Patterson
ICRA1
2001 Recognizing articulated objects in SAR images
Grinnell Jones, Bir Bhanu
Pattern Recognit.2
2001 Local discriminative learning for pattern recognition
Bir Bhanu
Pattern Recognit.2
2001 Independent feature analysis for image retrieval
Bir Bhanu
Pattern Recognit. Lett.2
2000 Logical Templates for Feature Extraction in Fingerprint Images
abstract
We present an approach for extraction of minutiae features from fingerprint images. The proposed approach is based on the use of logical templates for minutiae extraction in the presence of data distortion. A logical template is an expression that is applied to the binary ridge (valley) image at selected potential locations to detect the presence of minutiae at these locations. It is adapted to local ridge orientation and frequency. We discuss the proposed technique in detail, and present experimental results on low-resolution images of various qualities.
Bir Bhanu, Michael Boshra, Xuejun Tan
ICPR1
2000 Learning Based Interactive Image Segmentation
abstract
In this paper we present an approach to image segmentation in which user selected sets of examples and counter-examples supply information about the specific segmentation problem. Image segmentation is guided by a genetic algorithm which learns the appropriate subset and spatial combination of a collection of discriminating functions, associated with image features. The genetic algorithm encodes discriminating functions into a functional template representation, which can be applied to the input image to produce a candidate segmentation. The quality of each segmentation is evaluated within the genetic algorithm, by a comparison of two physics-based techniques for region growing and edge detection. Experimental results on real SAR imagery demonstrate that evolved segmentations are consistently better than segmentations derived from the Bayesian best single feature.
Bir Bhanu, Stephanie Fonder
ICPR1
2000 Adaptive target recognition
Bir Bhanu, Yingqiang Lin, Grinnell Jones
Mach. Vis. Appl.1
2000 Special issue on computer vision beyond the visible spectrum
Bir Bhanu, Ioannis Pavlidis, Robert A. Hummel
Mach. Vis. Appl.1
2000 Predicting Performance of Object Recognition
abstract
We present a method for predicting fundamental performance of object recognition. We assume that both scene data and model objects are represented by 2D point features and a data/model match is evaluated using a vote-based criterion. The proposed method considers data distortion factors such as uncertainty, occlusion, and clutter, in addition to model similarity. This is unlike previous approaches, which consider only a subset of these factors. Performance is predicted in two stages. In the first stage, the similarity between every pair of model objects is captured by comparing their structures as a function of the relative transformation between them. In the second stage, the similarity information is used along with statistical models of the data-distortion factors to determine an upper bound on the probability of recognition error. This bound is directly used to determine a lower bound on the probability of correct recognition. The validity of the method is experimentally demonstrated using real synthetic aperture radar (SAR) data.
Michael Boshra, Bir Bhanu
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 Adaptive integrated image segmentation and object recognition
abstract
The paper presents a general approach to image segmentation and object recognition that can adapt the image segmentation algorithm parameters to the changing environmental conditions. Segmentation parameters are represented by a team of generalized stochastic learning automata and learned using connectionist reinforcement learning techniques. The edge-border coincidence measure is first used as reinforcement for segmentation evaluation to reduce computational expenses associated with model matching during the early stage of adaptation. This measure alone, however, cannot reliably predict the outcome of object recognition. Therefore, it is used in conjunction with model matching where the matching confidence is used as a reinforcement signal to provide optimal segmentation evaluation in a closed-loop object recognition system. The adaptation alternates between global and local segmentation processes in order to achieve optimal recognition performance. Results are presented for both indoor and outdoor color images where the performance improvement over time is shown for both image segmentation and object recognition.
Bir Bhanu
IEEE Trans. Syst. Man Cybern. Part C1
1999 Performance Prediction and Validation for Object Recognition
abstract
This paper addresses the problem of predicting fundamental performance of vote-based object recognition using 2-D point features. It presents a method for predicting a tight lower bound on performance. Unlike previous approaches, the proposed method considers data-distortion factors, namely uncertainty, occlusion, and clutter, in addition to model similarity, simultaneously. The similarity between every pair of model objects is captured by comparing their structures as a function of the relative transformation between them. This information is used along with statistical models of the data-distortion factors to determine an upper bound on the probability of recognition error. This bound is directly used to determine a lower bound on the probability of correct recognition. The validity of the method is experimentally demonstrated using synthetic aperture radar (SAR) data obtained under different depression angles and target configurations.
Michael Boshra, Bir Bhanu
CVPR2
1999 Probabilistic Feature Relevance Learning for Content-Based Image Retrieval
Bir Bhanu, Shan Qing
Comput. Vis. Image Underst.2
1999 Recognition of Articulated and Occluded Objects
abstract
A model-based automatic target recognition system is developed to recognize articulated and occluded objects in synthetic aperture radar (SAR) images, based on invariant features of the objects. Characteristics of SAR target image scattering centers, azimuth variation, and articulation invariants are presented. The basic elements of the new recognition system are described and performance results are given for articulated, occluded and occluded articulated objects, and they are related to the target articulation invariance and percent unoccluded.
Grinnell Jones, Bir Bhanu
IEEE Trans. Pattern Anal. Mach. Intell.2
1998 Learning Integrated Online Indexing for Image Databases
abstract
Most of the current image retrieval systems use "one-shot" queries to a database to retrieve similar images. Typically a K-NN (nearest neighbor) kind of algorithms is used where weights measuring feature importance along input dimensions remain fixed (or manually tweaked by the user) in the computation of a given similarity metric. However, the similarity does not vary with equal strength or in the same proportion in all directions in the feature space emanating from the query image. The manual adjustment of these weights is time consuming and exhausting. Moreover, it requires a very sophisticated user. We present a novel method that enables image retrieval procedures to continuously learn feature relevance based on user's feedback, and which is highly adaptive to query locations. Experimental results are presented that provide the objective evaluation of learning behaviour of the method for image retrieval.
Bir Bhanu, Shan Qing
ICIP (2)1
1998 Predicting Object Recognition Performance under Data Uncertainty Occlusion and Clutter
Michael Boshra, Bir Bhanu
ICIP (3)2
1998 A system for model-based recognition of articulated objects
abstract
This paper presents a model-based matching technique for recognition of articulated objects (with two parts) and the poses of these parts in synthetic aperture radar (SAR) images. Using articulation invariants as features, the recognition system first hypothesizes the pose of the larger part and then the pose of the smaller part. Geometric reasoning is carried out to correct identification errors. The thresholds for the quality of match are determined dynamically by minimizing the probability of a random match. Results are presented using SAR images of three articulated objects. The system performance is evaluated with respect to identification performance, accuracy of estimates for the poses of the object parts and noise.
Bir Bhanu, Joon Ahn
ICPR1
1998 Local reinforcement learning for object recognition
abstract
Current computer vision systems, whose basic methodology is open-loop or filter type, typically use image segmentation followed by object recognition algorithms. These systems are not robust for most real-world applications. In contrast, the system presented here achieves robust performance by using local reinforcement learning to induce a highly adaptive mapping from input images to segmentation strategies. This is accomplished by using the confidence level of model matching as reinforcement to drive learning. The system is verified through experiments on a large set of real images.
Bir Bhanu
ICPR2
1998 Closed-Loop Object Recognition Using Reinforcement Learning
abstract
Current computer vision systems whose basic methodology is open-loop or filter type typically use image segmentation followed by object recognition algorithms. These systems are not robust for most real-world applications. In contrast, the system presented here achieves robust performance by using reinforcement learning to induce a mapping from input images to corresponding segmentation parameters. This is accomplished by using the confidence level of model matching as a reinforcement signal for a team of learning automata to search for segmentation parameters during training. The use of the recognition algorithm as part of the evaluation function for image segmentation gives rise to significant improvement of the system performance by automatic generation of recognition strategies. The system is verified through experiments on sequences of indoor and outdoor color images with varying external conditions.
Bir Bhanu
IEEE Trans. Pattern Anal. Mach. Intell.2
1998 A system for model-based object recognition in perspective aerial images
Subhodev Das, Bir Bhanu
Pattern Recognit.2
1998 Delayed reinforcement learning for adaptive image segmentation and feature extraction
abstract
Object recognition is a multilevel process requiring a sequence of algorithms at low, intermediate, and high levels. Generally, such systems are open loop with no feedback between levels and assuring their robustness is a key challenge in computer vision and pattern recognition research. A robust closed-loop system based on "delayed" reinforcement learning is introduced. The parameters of a multilevel system employed for model-based object recognition are learned. The method improves recognition results over time by using the output at the highest level as feedback for the learning system. It has been experimentally validated by learning the parameters of image segmentation and feature extraction and thereby recognizing 2D objects. The approach systematically controls feedback in a multilevel vision system and shows promise in approaching a long-standing problem in the field of computer vision and pattern recognition.
Bir Bhanu
IEEE Trans. Syst. Man Cybern. Part C2
1997 Stochastic Models for Recognition of Articulated Objects
abstract
We present a hidden Markov modeling (HMM) based approach for recognition of articulated objects in synthetic aperture radar (SAR) images. We develop multiple models for a given SAR image of an object and integrate these models synergistically using their probabilistic estimates for recognition and estimates of invariance of features as a result of articulation. The models are based on sequentialization of scattering centers extracted from SAR images. Experimental results are presented using 1440 training images and 2520 testing images for 4 classes.
Bir Bhanu, Bing Tian
ICIP (2)1
1997 Oracle: An Integrated Learning Approach for Object Recognition
abstract
Model-based object recognition has become a popular paradigm in computer vision research. In most of the current model-based vision systems, the object models used for recognition are generally a priori given (e.g. obtained using a CAD model). For many object recognition applications, it is not realistic to utilize a fixed object model database with static model features. Rather, it is desirable to have a recognition system capable of performing automated object model acquisition and refinement. In order to achieve these capabilities, we have developed a system called ORACLE: Object Recognition Accomplished through Consolidated Learning Expertise. It uses two machine learning techniques known as Explanation-Based Learning (EBL) and Structured Conceptual Clustering (SCC) combined in a synergistic manner. As compared to systems which learn from numerous positive and negative examples, EBL allows the generalization of object model descriptions from a single example. Using these generalized descriptions, SCC constructs an efficient classification tree which is incremently built and modified over time. Learning from experience is used to dynamically update the specific feature values of each object. These capabilities provide a dynamic object model database which allows the system to exhibit improved performance over time. We provide an overview of the ORACLE system and present experimental results using a database of thirty aircraft models.
John C. Ming, Bir Bhanu
Int. J. Pattern Recognit. Artif. Intell.2
1997 Analysis of terrain using multispectral images
Bir Bhanu, Peter Symosek, Subhodev Das
Pattern Recognit.1
1997 Guest Editorial Introduction To The Special Issue On Automatic Target Detection And Recognition
abstract
Automatic target recognition (ATR) generally refers to the autonomous or aided target detection and recognition by computer processing of data from a variety of sensors such as forward looking infrared (FLIR), synthetic aperture radar(SAR), inverse synthetic aperture radar (ISAR), laser radar (LADAR), millimeter wave (MMW) radar, multispectral/hyperspectral sensors, low-light television (LLTV), video, ete. It is an extremely important capability for targeting and surveillance missions of defense weapon systems operating from a variety of platforms.
Bir Bhanu, Dan E. Dudgeon, Edmund G. Zelnio, Azriel Rosenfeld, David P. Casasent, Irving S. Reed
IEEE Trans. Image Process.1
1997 Gabor wavelet representation for 3-D object recognition
abstract
This paper presents a model-based object recognition approach that uses a Gabor wavelet representation. The key idea is to use magnitude, phase, and frequency measures of the Gabor wavelet representation in an innovative flexible matching approach that can provide robust recognition. The Gabor grid, a topology-preserving map, efficiently encodes both signal energy and structural information of an object in a sparse multiresolution representation. The Gabor grid subsamples the Gabor wavelet decomposition of an object model and is deformed to allow the indexed object model match with similar representation obtained using image data. Flexible matching between the model and the image minimizes a cost function based on local similarity and geometric distortion of the Gabor grid. Grid erosion and repairing is performed whenever a collapsed grid, due to object occlusion, is detected. The results on infrared imagery are presented, where objects undergo rotation, translation, scale, occlusion, and aspect variations under changing environmental conditions.
Bir Bhanu
IEEE Trans. Image Process.2
1996 Closed-Loop Object Recognition Using Reinforcement Learning
abstract
Current computer vision systems whose basic methodology is open-loop or filter type typically use image segmentation followed by object recognition algorithms. These systems are not robust for most real-world applications. In contrast, the system presented here achieves robust performance by using reinforcement learning to induce a mapping from input images to corresponding segmentation parameters. This is accomplished by using the confidence level of model matching as a reinforcement signal for a team of learning automata to search for segmentation parameters during training. The use of the recognition algorithm as part of the evaluation function for image segmentation gives rise to significant improvement of the system performance by automatic generation of recognition strategies. The system is verified through experiments on sequences of color images with varying external conditions.
Bir Bhanu
CVPR2
1996 Modeling Clutter and Context for Target Detection in Infrared Images
abstract
In order to reduce false alarms and to improve the target detection performance of an automatic target detection and recognition system operating in a cluttered environment, it is important to develop the models not only for man-made targets bat also of natural background clutters. Because of the high complexity of natural clutters, this clutter model can only be reliably built through learning from real examples. If available, contextual information that characterizes each training example can be used to further improve the learned clutter model. In this paper, we present such a clutter model aided target detection system. Emphases are placed on two topics: (1) learning the background clutter model from sensory data through a self-organizing process, (2) reinforcing the learned clutter model using contextual information.
Songnian Rong, Bir Bhanu
CVPR2
1996 Delayed reinforcement learning for closed-loop object recognition
abstract
Object recognition is a multi-level process requiring a sequence of algorithms at low, intermediate and high levels. Generally, such systems are open loop with no feedback between levels and assuring their robustness is a key challenge in computer vision research. A robust closed-loop system based on "delayed" reinforcement learning is introduced in this paper. The parameters of a multi-level system employed for model-based object recognition are learned. The method improves recognition results over time by using the output at the highest level as feedback for the learning system. It has been experimentally validated by learning the parameters of image segmentation and feature extraction and thereby recognizing 2D objects. The approach systematically controls feedback in a multi-level vision system and provides a potential solution to a long-standing problem in the field of computer vision.
Bir Bhanu
ICPR2
1996 Automatic model construction for object recognition using ISAR images
abstract
This paper discusses model construction for object recognition using ISAR images. A learning-from-examples approach is used to construct recognition models of the objects from their ISAR data. Given a set of ISAR data of an object of interest, structural features are extracted from the images. Statistical analysis and geometrical reasoning are then used to analyze the features to find spatial and statistical invariance so that a structural model of the object suitable for object recognition can be constructed. Results of experiments using the automatically constructed model in object recognition are presented.
Bir Bhanu
ICPR2
1996 Adaptive object detection based on modified Hebbian learning
abstract
This paper focuses on the issue of developing self-adapting automatic object detection systems for improving their performance. Two general methodologies for performance improvement are first introduced. They are based on parameter optimizing and input adapting. Different modified Hebbian learning rules are developed to build adaptive, feature extractors which transform the input data into a desired form for a given algorithm. To show its feasibility, an input adaptor for object detection is designed as an example and tested using multisensor data (optical, SAR, and FLIR). Test results are presented and discussed in the paper.
Yong-Jian Zheng, Bir Bhanu
ICPR2
1996 Generic object recognition using multiple representations
Subhodev Das, Bir Bhanu, Chih-Cheng Ho
Image Vis. Comput.2
1996 Target indexing in SAR images using scattering centers and the Hausdorff distance
June-Ho Yi, Bir Bhanu
Pattern Recognit. Lett.2
1995 Gabor Wavelets for 3-D Object Recognition
abstract
This paper presents a model-based object recognition approach that uses a hierarchical Gabor wavelet representation. The key idea is to use magnitude, phase and frequency measures of Gabor wavelet representation in an innovative flexible matching approach that can provide robust recognition. A Gabor grid a topology-preserving map, efficiently encodes both signal energy and structural information of an object in a sparse multi-resolution representation. The Gabor grid subsamples the Gabor wavelet decomposition of an object model and is deformed to allow the indexed object model match with the image data. Flexible matching between the model and the image minimizes a cost function based on local similarity and geometric distortion of the Gabor grid. Grid erosion and repairing is performed whenever a collapsed grid, due to object occlusion, is detected. The results on infrared imagery are presented. Where objects undergo rotation, translation, scale, occlusion and aspect variations under changing environmental conditions.>
Bir Bhanu
ICCV2
1995 Error bound for multi-stage synthesis of narrow bandwidth Gabor filters
abstract
This paper develops an error bound for narrow bandwidth Gabor filters synthesized using multiple stages. It is shown that the error introduced by approximating narrow bandwidth Gabor (1946) kernels by a weighted sum of spatially offset, separable kernels is a function of the frequency offset and the reduction in bandwidth of the desired kernel compared to the basis values, as well as the spatial subsampling rate between filter stages. This error bound should prove useful in the design of a general basis filter set for multi-stage filtering because the maximum frequency offset is largely determined by the spacing of the basis filters.
R. Neil Braithwaite, Bir Bhanu
ICIP2
1995 Composite phase and phase-based Gabor element aggregation
abstract
This paper describes how the phase, obtained by Gabor filtering an image, can be used to aggregate related Gabor elements (simple features identified by peaks in the Gabor magnitude). This phase-based feature grouping simplifies the perennial problem of target/background segmentation because we need to only determine if the aggregate feature is target or background, rather than determine the status of each feature independently. Since the phase from a single quadrature Gabor output cannot tolerate large changes in orientation, a new local measure, referred to as the composite phase, has been developed. It is a combination of the filter responses from multiple orientations which allows the phase to follow contours with large changes in orientation. A constant composite phase contour is used to connect related Gabor elements that would otherwise appear separated within the magnitude response.
R. Neil Braithwaite, Bir Bhanu
ICIP2
1995 Adaptive image segmentation using a genetic algorithm
abstract
Image segmentation is an old and difficult problem. One of the fundamental weaknesses of current computer vision systems to be used in practical applications is their inability to adapt the segmentation process as real-world changes occur in the image. We present the first closed loop image segmentation system which incorporates a genetic algorithm to adapt the segmentation process to changes in image characteristics caused by variable environmental conditions such as time of day, time of year, clouds, etc. The segmentation problem is formulated as an optimization problem and the genetic algorithm efficiently searches the hyperspace of segmentation parameter combinations to determine the parameter set which maximizes the segmentation quality criteria. The goals of our adaptive image segmentation system are to provide continuous adaptation to normal environmental variations, to exhibit learning capabilities, and to provide robust performance when interacting with a dynamic environment. We present experimental results which demonstrate learning and the ability to adapt the segmentation performance in outdoor color imagery.
Bir Bhanu, Sungkee Lee, John C. Ming
IEEE Trans. Syst. Man Cybern.1
1994 Hierarchical Gabor filters for object detection in infrared images
abstract
This paper presents a new representation called "hierarchical Gabor filters" and associated novel local measures which are used to detect potential objects of interest in images. The "first stage" of the approach uses a wavelet set of wide-bandwidth separable Gabor filters to extract local measures from an image. The "second stage" makes certain spatial groupings explicit by creating small-bandwidth, non-separable Gabor filters that are tuned to elongated contours or periodic patterns. The non-separable filter responses are obtained from a weighted combination of the separable basis filters, which preserves the computational efficiency of separable filters while providing the distinctiveness required to discriminate objects from clutter. This technique is demonstrated on images obtained from a forward looking infrared (FLIR) sensor.>
R. Neil Braithwaite, Bir Bhanu
CVPR2
1994 A Geometric Constraint Method for Estimating 3-D Camera Motion
abstract
We investigate the problem of estimating the camera motion parameters from a 2D image sequence that has been obtained under combined 3D camera translation and rotation. Geometrical constraints in 2D are used to iteratively narrow down regions of possible values in the space of 3D motion parameters. The approach is based on a new concept called "FOE-feasibility" of an image region, for which an efficient algorithm has been implemented. Results are shown on real images.>
Wilhelm Burger, Bir Bhanu
ICRA2
1994 A system for aircraft recognition in perspective aerial images
abstract
Recognition of aircraft in complex, perspective aerial imagery has to be accomplished in presence of clutter, occlusion, shadow, and various forms of image degradation. This paper presents a system for aircraft recognition under real-world conditions that is based on the use of a hierarchical database of object models. The particular approach involves three key processes: (a) The qualitative object recognition process performs model-based symbolic feature extraction and generic object recognition; (b) The refocused matching and evaluation process refines the extracted features for more specific classification with input from (a); and (c) The primitive feature extraction process regulates the extracted features based on their saliency and interacts with (a) and (b). Experimental results showing the qualitative recognition of aircraft in perspective, aerial images are presented.>
Subhodev Das, Bir Bhanu, R. Neil Braithwaite
WACV2
1994 Introduction to the Special Section on Learning in Computer Vision
Bir Bhanu, Tomaso A. Poggio
IEEE Trans. Pattern Anal. Mach. Intell.1
1992 A system for obstacle detection during rotorcraft low-altitude flight
abstract
Airborne vehicles such as rotorcraft must avoid obstacles such as antennas, towers, poles, fences, tree branches, and wires strung across the flight path. The paper analyzes the requirements of an obstacle detection system for rotorcrafts in low-altitude Nap-of-the-Earth flight based on various rotorcraft motion constraints. It argues that an automated obstacle detection system for the rotorcraft scenario should include both passive and active sensors. Consequently, it introduces a maximally passive system which involves the use of passive sensors (TV, FLIR) as well as the selective use of an active (laser) sensor. The passive component is concerned with estimating range using optical flow-based motion analysis and binocular stereo in conjunction with inertial navigation system information. Experimental results obtained using land vehicle data illustrate the particular approach to motion analysis.>
Bir Bhanu, Barry A. Roberts, David Duncan, Subhodev Das
WACV1
1991 Closed-loop adaptive image segmentation
abstract
A closed-loop image segmentation system that incorporates a genetic algorithm to adapt the segmentation process to changes in image characteristics caused by variable environmental conditions is presented. The genetic algorithm efficiently searches the hyperspace of segmentation parameter combinations to determine the parameter set which maximizes the segmentation quality criteria. A summary of the experimental results that demonstrates the ability to perform adaptive image segmentation and to learn from experience using a collection of outdoor color imagery is given.>
Bir Bhanu, John C. Ming, Sungkee Lee
CVPR1
1991 A qualitative approach to dynamic scene understanding
Bir Bhanu, Wilhelm Burger
CVGIP Image Underst.1
1990 Inertial navigation sensor integrated motion analysis for obstacle detection
abstract
A maximally passive approach to obstacle detection is described, and the details of an inertial sensor integrated optical flow analysis technique are discussed. The optical flow algorithm has been used to generate range samples using both synthetic data and real data (imagery and inertial navigation system information) obtained from a moving vehicle. The conditions under which the data were created/collected are described, and images illustrating the results of the major steps in the optical flow algorithm are provided.>
Bir Bhanu, Barry A. Roberts, John C. Ming
ICRA1
1990 Estimating 3D Egomotion from Perspective Image Sequence
abstract
The computation of sensor motion from sets of displacement vectors obtained from consecutive pairs of images is discussed. The problem is investigated with emphasis on its application to autonomous robots and land vehicles. The effects of 3D camera rotation and translation upon the observed image are discussed, particularly the concept of the focus of expansion (FOE). It is shown that locating the FOE precisely is difficult when displacement vectors are corrupted by noise and errors. A more robust performance can be achieved by computing a 2D region of possible FOE locations (termed the fuzzy FOE) instead of looking for a single-point FOE. The shape of this FOE region is an explicit indicator of the accuracy of the result. It has been shown elsewhere that given the fuzzy FOE, a number of powerful inferences about the 3D sense structure and motion become possible. Aspects of computing the fuzzy FOE are emphasized, and the performance of a particular algorithm on real motion sequences taken from a moving autonomous land vehicle is shown.>
Wilhelm Burger, Bir Bhanu
IEEE Trans. Pattern Anal. Mach. Intell.2
1989 On computing a 'fuzzy' focus of expansion for autonomous navigation
abstract
The focus of expansion (FOE) is an important concept in dynamic scene analysis, particularly where translational motion is dominant, as in mobile robot applications. In practice, it is difficult to determine the exact location of the FOE from a given set of displacement vectors due to the effects of camera rotation, digitization, and noise. Instead of a single image location, the authors propose to compute a connected region, termed fuzzy FOE, that marks the approximate direction of camera heading. The fuzzy FOE provides numerous clues about 3D scene structure and independent object motion. The main problems of the classic FOE approach are discussed, concentrating on the details of computing the fuzzy FOE for a camera undergoing translation and rotation in 3D space. Results on real outdoor images are presented. Experiments show that even erroneous point matches in the given image displacement vector field can be tolerated, which is a necessity in a fully automated process.>
W. Burger, Bir Bhanu
CVPR2
1989 VLSI design and implementation of a real-time image segmentation processor
Bir Bhanu, Brad L. Hutchings, Kent F. Smith
Mach. Vis. Appl.1
1989 Recognition of 3-D objects in range images using a butterfly multiprocessor
Bir Bhanu, Lawrence A. Nuttall
Pattern Recognit.1
1988 Dynamic scene understanding for autonomous mobile robots
abstract
A new approach to the dynamic scene analysis is presented which departs from previous work by emphasizing a qualitative strategy of reasoning and modeling. Instead of refining a single quantitative description of the observed environment over time, multiple qualitative interpretations are maintained simultaneous-ly. This offers superior robustness and flexibility over traditional numerical techniques which are often ill-conditioned and noise-sensitive. The main tasks of our approach are (a) to detect and to classify the motion of individual objects in the scene, (b) to estimate the robot's egomotion, and (c) to derive the 3-D struc-ture of the stationary environment. These three tasks strongly depend on each other. First, the direction of heading (i.e. trans-lation) and rotation of the robot are estimated with respect to stationary locations in the scene. The focus of expansion (FOE) is not determined as particular image location, but as a region of possible FOE-locations called the Fuzzy FOE. From this infor-mation, a rule-based system constructs and maintains a Qualitative Scene Model. Results of this approach from real and synthetic imagery are presented. 1.
Wilhelm Burger, Bir Bhanu
CVPR2
1988 Landmark recognition for autonomous mobile robots
abstract
A novel approach for landmark recognition based on the perception, reasoning, action, and expectation (PREACTE) paradigm is presented for the navigation of autonomous mobile robots. PREACTE uses expectations to predict the appearance and disappearance of objects, thereby reducing computational complexity and locational uncertainty. It uses an innovative concept called dynamic model matching (DMM), which is based on the automatic generation of landmark description at different ranges and aspect angles and uses explicit knowledge about maps and landmarks. Map information is used to generate an expected site model (ESM) for search delimitation, given the location and velocity of the mobile robot. The landmark recognition vision system generates 2-D and 3-D scene models from the observed scene. The ESM hypotheses are verified by matching them to the image model. Experimental results that verify the performance of the PREACTE and DMM algorithms for real imagery are also presented.>
Hatem Nasr, Bir Bhanu
ICRA2
1988 Approximation of displacement fields using wavefront region growing
Bir Bhanu, Wilhelm Burger
Comput. Vis. Graph. Image Process.1
1987 CAOS: A hierarchical robot control system
abstract
Control systems which enable robots to behave intelligently is a major issue in todays process of automating factories. This paper presents a hierarchical robot control system, termed CAOS for Control using Action Oriented Schemas, with ideas taken from the neurosciences. We are using action oriented schemas (called neuroschemas) as the basic building blocks in a hierarchical control structure which is being implemented on a BBN Butterfly Parallel Processor. Serial versions in C and LISP are presented with examples showing how CAOS achieves the goals of recognizing three-dimensional polyhedral objects. We also describe a simulation on how it manipulates objects. Moreover, the ongoing implementation of a parallel version of the system is discussed.
Bir Bhanu, Nils Thune, Mari Thune
ICRA1
1987 CAD-based robotics
abstract
We describe an approach which facilitates and makes explicit the organization of the knowledge necessary to map robotic system requirements onto an appropriate assembly of algorlthms, processors, sensors, and actuators. In order to achieve this mapping several kinds of knowledge are needed. In this paper, we describe a system under development which exploits the Computer Aided Design (CAD) database in order to synthesize. •recognition Code for vision systems (both 2-D and 3-D), •grasping sites for simple parallel grippers, and •manipulation strategies for dextrous manipulation. We use an object-based approach and give an example application of the system to CAD-based 2-D vision.
Thomas C. Henderson, Eliot Weitz, Charles D. Hansen, Roderic A. Grupen, C. C. Ho, Bir Bhanu
ICRA6
1987 Qualitative Motion Understanding
Wilhelm Burger, Bir Bhanu
IJCAI2
1987 Recognition of occluded objects: A cluster-structure algorithm
Bir Bhanu, John C. Ming
Pattern Recognit.1
1987 Segmentation of natural scenes
Bir Bhanu, Bahram Parvin
Pattern Recognit.1
1987 3-D model building for computer vision
Bir Bhanu, Chih-Cheng Ho, Tom Henderson
Pattern Recognit. Lett.1
1986 Recognition of occluded objects: A cluster structure paradigm
abstract
Clustering techniques have been used to perform image segmentation, to detect lines and curves in the images and to solve several other problems in pattern recognition and image analysis. In this paper we apply clustering methods to a new problem domain and present a new method based on a cluster-structure paradigm for the recognition of 2-D partially occluded objects. The cluster-structure paradigm entails the application of clustering concepts in a hierarchical manner. The amount of computational effort decreases as the recognition algorithm progresses. As compared to some of the earlier methods, which identify an object based on only one sequence of matched segments, the new technique allows the identification of all parts of the model which match with the apparent object. Also the method is able to tolerate a moderate change in scale and a significant amount of shape distortion arising as a result of segmentation and/or the polygonal approximation of the boundary of the object. The method has been evaluated with respect to a large number of examples where several objects partially occlude one another. A summary of the results is presented.
Bir Bhanu, John C. Ming
ICRA1
1985 CAGD based 3-D vision
abstract
This paper presents the initial results of our work in using the Computer Aided Geometric Design (CAGD) representations and models as a basis for the visual recognition of 3-D objects for robotic applications. We describe some techniques and algorithms which allow the generation of computer representations and geometric models of complicated realizable 3-D objects in a systematic manner. These representations and models are obtained using (a) available CAGD techniques, and (b) data acquired from various range finding techniques. As compared to previous work in machine vision, multiple hierarchical representations of an object obtained from geometric models can be used for finding orientation and position information.
Bir Bhanu, Thomas C. Henderson
ICRA1
1985 A Framework for Distributed Sensing and Control
Tom Henderson, Charles D. Hansen, Bir Bhanu
IJCAI3
1985 Overview of the computer vision and robotics programme at the University of Utah
Thomas C. Henderson, Bir Bhanu
Image Vis. Comput.2
1985 Intrinsic characteristics as the interface between CAD and machine vision systems
Thomas C. Henderson, Bir Bhanu
Pattern Recognit. Lett.2
1984 Representation and Shape Matching of 3-D Objects
abstract
A three-dimensional scene analysis system for the shape matching of real world 3-D objects is presented. Various issues related to representation and modeling of 3-D objects are addressed. A new method for the approximation of 3-D objects by a set of planar faces is discussed. The major advantage of this method is that it is applicable to a complete object and not restricted to single range view which was the limitation of the previous work in 3-D scene analysis. The method is a sequential region growing algorithm. It is not applied to range images, but rather to a set of 3-D points. The 3-D model of an object is obtained by combining the object points from a sequence of range data images corresponding to various views of the object, applying the necessary transformations and then approximating the surface by polygons. A stochastic labeling technique is used to do the shape matching of 3-D objects. The technique matches the faces of an unknown view against the faces of the model. It explicitly maximizes a criterion function based on the ambiguity and inconsistency of classification. It is hierarchical and uses results obtained at low levels to speed up and improve the accuracy of results at higher levels. The objective here is to match the individual views of the object taken from any vantage point. Details of the algorithm are presented and the results are shown on several unknown views of a complicated automobile casting.
Bir Bhanu
IEEE Trans. Pattern Anal. Mach. Intell.1
1984 Shape Matching of Two-Dimensional Objects
abstract
In this paper we present results in the areas of shape matching of nonoccluded and occluded two-dimensional objects. Shape matching is viewed as a ``segment matching'' problem. Unlike the previous work, the technique is based on a stochastic labeling procedure which explicitly maximizes a criterion function based on the ambiguity and inconsistency of classification. To reduce the computation time, the technique is hierarchical and uses results obtained at low levels to speed up and improve the accuracy of results at higher levels. This basic technique has been extended to the situation where various objects partially occlude each other to form an apparent object and our interest is to find all the objects participating in the occlusion. In such a case several hierarchical processes are executed in parallel for every object participating in the occlusion and are coordinated in such a way that the same segment of the apparent object is not matched to the segments of different actual objects. These techniques have been applied to two-dimensional simple closed curves represented by polygons and the power of the techniques is demonstrated by the examples taken from synthetic, aerial, industrial and biological images where the matching is done after using the actual segmentation methods.
Bir Bhanu, Olivier D. Faugeras
IEEE Trans. Pattern Anal. Mach. Intell.1
1983 Recognition of Occluded Objects
Bir Bhanu
IJCAI1
1982 Computation of two-dimensional complex cepstrum
abstract
A technique based on fitting splines to the phase derivative curve is presented for the efficient and reliable computation of the two-dimensional complex cepstrum. The technique is an adaptive numerical integration scheme and makes use of several computational strategies within the Tribolet's phase unwrapping algorithm. An application of the complex cepstrum in testing the stability of two-dimensional recursive digital filters is considered. Susceptibility of the computation of complex cepstrum to slight changes in the coefficients of a two-dimensional array is studied. Several examples of stable and unstable two-dimensional quarter-plane and non-symmetric half-plane recursive digital filters are presented.
Bir Bhanu
ICASSP1
1982 Segmentation of Images Having Unimodal Distributions
abstract
A gradient relaxation method based on maximizing a criterion function is studied and compared to the nonlinear probabilistic relaxation method for the purpose of segmentation of images having unimodal distributions. Although both methods provide comparable segmentation results, the gradient method has the additional advantage of providing control over the relaxation process by choosing three parameters which can be tuned to obtain the desired segmentation results at a faster rate. Examples are given on two different types of scenes.
Bir Bhanu, Olivier D. Faugeras
IEEE Trans. Pattern Anal. Mach. Intell.1
1982 Correction to "Segmentation of Images Having Unimodal Distributions"
Bir Bhanu, Olivier D. Faugeras
IEEE Trans. Pattern Anal. Mach. Intell.1