Rama Chellappa

dblp:c/RamaChellappa · also Ramalingam Chellappa · DBLP profile ↗
← Back
681ranked-venue papers
42as first author
86since 2021 · last 2026
0000-0002-7638-1650ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 490 · 13 first-author · 63 since 2021Artificial intelligence and machine learning · 367 · 27 first-author · 63 since 2021Security and privacy · 16 · 8 since 2021Human-computer interaction and ubiquitous computing · 16 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 3 first-author · 4 since 2021Theory of computation · 7Systems, architecture and hardware · 6 · 3 since 2021Computer networks · 2
YearPublicationVenuePosition
2026 Mesh-Gait: A Unified Framework for Gait Recognition Through Multi-Modal Representation Learning from 2D Silhouettes
Zhao-Yang Wang, Jieneng Chen, Yuxiang Guo 0001, Jiang Liu 0014, Rama Chellappa
FG5
2026 DiffRegCD: Integrated Registration and Change Detection with Diffusion Features
abstract
Change detection (CD) is critical in computer vision and remote sensing, with applications in monitoring, disaster response, and urban analysis. Most CD models assume co-registered inputs, but real imagery often suffers from parallax, viewpoint shifts, or long temporal gaps, leading to severe misalignment. Conventional register-then-detect pipelines and recent joint frameworks (e.g., BiFA, ChangeRD) remain limited: they rely on regression-only flow, global homographies, or synthetic perturbations that fail under large displacements. We propose DiffRegCD, an integrated framework that couples dense registration and change detection. DiffRegCD reformulates correspondence as a Gaussian-smoothed classification task, delivering sub-pixel accuracy and stable training. It builds on frozen multi-scale features from a pretrained denoising diffusion model, which provide invariance to viewpoint and illumination variation. Supervision is enabled by controlled affine perturbations applied to standard CD datasets, yielding paired ground truth for both flow and change detection without pseudo-labels. Experiments on aerial (LEVIR-CD, DSIFN-CD, WHU-CD, SYSU-CD) and ground-level (VL-CMU-CD) datasets show that DiffRegCD outperforms recent baselines and remains robust under wide temporal and viewpoint variation, establishing diffusion features and classification-based correspondence as a strong foundation for integrated CD. The code is available at GitHub.
Seyedehanita Madani, Rama Chellappa, Vishal M. Patel
WACV2
2026 VRAgent: Self-Refining Agent for Zero-Shot Multimodal Video Retrieval
abstract
Recent advances in Vision-Language Models (VLMs) and Large Language Models (LLMs) have demonstrated remarkable zero-shot capabilities while becoming increasingly accessible. They have enabled zero-shot text-to-video search, yet they still struggle on real-world queries that demand temporally aligned reasoning across vision and speech. We present VRAgent, an agentic retrieval framework that leverages a central LLM as a planner which (i) decomposes a free-form user query into a tool-instruction set spanning visual, dialogue and other modality-specific retrievers, and (ii) iteratively self-refines this plan by scoring the outputs and rewriting the next tool-instruction set. The resulting closed-loop optimization acts at test time and requires no additional training data or gradient updates.VRAgent is modular by design—adding a new modality is as simple as adding a corresponding foundation model into the toolbox. On our newly proposed MM-MSRVTT and TVR-1200 multimodal benchmarks, VRAgent improves average recall by +8.3% and +1.7% over the best zero-shot baselines, while on single-modality MSR-VTT and DiDeMo it obtains consistent gains of +3.7% and +4.1%. An interactive variant that asks the user up to two multiple-choice questions pushes average recall to 79.7% on MSR-VTT, underscoring the value of on-the-fly human feedback.
Ketul Shah, Pankaj Nathani, Rama Chellappa, Fabian Caba Heilbron
WACV3
2026 Multi-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting
Guangxing Han, Long Chen 0016, Jiawei Ma, Shiyuan Huang 0001, Rama Chellappa, Shih-Fu Chang
Int. J. Comput. Vis.5
2026 DiffProtect: Generative adversarial examples using diffusion models for facial privacy protection
abstract
The increasingly pervasive facial recognition (FR) systems raise serious concerns about personal privacy, especially for billions of users who have publicly shared their photos on social media. To address this challenge, several adversarial attack methods have been proposed to protect individuals from being identified by unauthorized FR systems with perturbed facial images. However, these approaches suffer from poor visual quality or low attack success rates, which limit their practical utility. Recently, diffusion models have achieved tremendous success in image generation. In this work, we ask: can diffusion models be used to generate adversarial examples against FR systems to improve both visual quality and attack performance? We propose DiffProtect, a novel method leveraging a diffusion autoencoder to generate semantically meaningful perturbations on FR systems. Extensive experiments demonstrate that DiffProtect produces more natural-looking encrypted images than state-of-the-art methods while achieving significantly higher attack success rates, e.g. , 24.5 % and 25.1 % absolute improvements on the CelebA-HQ and FFHQ datasets. We further evaluate the effectiveness of DiffProtect in the real world using a commercial FR API and validate its usefulness in practice through a user study. Our code is available at https://github.com/joellliu/DiffProtect .
Jiang Liu 0014, Chun Pong Lau 0001, Zhongliang Guo 0001, Yuxiang Guo 0001, Zhao-Yang Wang, Rama Chellappa
Pattern Recognit.6
2026 Recovering Pulse Waves From Video Using Deep Unrolling and Deep Equilibrium Models
abstract
Camera-based contactless monitoring of vital signs, also known as imaging photoplethysmography (iPPG), has seen applications in driver-monitoring, perfusion assessment, affective computing, and more. iPPG involves sensing the underlying cardiac pulse from video of the skin and estimating vital signs such as the pulse rate or a full pulse waveform. Some previous iPPG methods impose model-based sparse priors on the pulse signals and use iterative optimization for pulse wave recovery, while others use end-to-end black-box deep learning methods. In contrast, we introduce methods that combine signal processing and deep learning methods in an inverse problem framework. Our methods estimate the underlying pulse signal, pulse rate, and pulse rate variability from facial video by learning deep-network-based denoising operators that leverage deep algorithm unfolding and deep equilibrium models. Experiments show that our methods can denoise an acquired signal from the face and infer the correct underlying pulse rate and pulse rate variability, achieving pulse rate estimation performance consistent with the state-of-the-art on well-known benchmarks, all with less than one-fifth the number of learnable parameters as the closest competing method.
Vineet R. Shenoy, Suhas Lohit, Hassan Mansour, Rama Chellappa, Tim K. Marks
IEEE Trans. Image Process.4
2025 Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval
abstract
In this work, we tackle the problem of text-to-video retrieval (T2VR). Inspired by the success of late interaction techniques in text-document, text-image, and text-video retrieval, our approach, Video-ColBERT, introduces a simple and efficient mechanism for fine-grained similarity assessment between queries and videos. Video-ColBERT is built upon three main components: a fine-grained spatial and temporal token-wise interaction, query and visual expansions, and a dual sigmoid loss during training. We find that this interaction and training paradigm leads to strong individual, yet compatible, representations for encoding video content. These representations lead to increases in performance on common text-to-video retrieval benchmarks compared to other bi-encoder methods.
Arun V. Reddy, Alexander Martin 0006, Eugene Yang 0001, Andrew Yates, Kate Sanders 0002, Kenton Murray, Reno Kriz, Celso de Melo, Benjamin Van Durme, Rama Chellappa
CVPR10
2025 Polysemantic Dropout: Conformal OOD Detection for Specialized LLMs
abstract
We propose a novel inference-time out-ofdomain (OOD) detection algorithm for specialized large language models (LLMs).Despite achieving state-of-the-art performance on in-domain tasks through fine-tuning, specialized LLMs remain vulnerable to incorrect or unreliable outputs when presented with OOD inputs, posing risks in critical applications.Our method leverages the Inductive Conformal Anomaly Detection (ICAD) framework, using a new non-conformity measure based on the model's dropout tolerance.Motivated by recent findings on polysemanticity and redundancy in LLMs, we hypothesize that in-domain inputs exhibit higher dropout tolerance than OOD inputs.We aggregate dropout tolerance across multiple layers via a valid ensemble approach, improving detection while maintaining theoretical false alarm bounds from ICAD.Experiments with medical-specialized LLMs show that our approach detects OOD inputs better than baseline methods, with AUROC improvements of 2% to 37% when treating OOD datapoints as positives and in-domain test datapoints as negatives.
Ayush Gupta 0001, Ramneet Kaur, Adam D. Cobb, Rama Chellappa, Susmit Jha
EMNLP5
2025 Improved Representation Learning for Unconstrained Face Recognition
abstract
Face recognition is a widely studied problem where the aim is to design a robust network that assigns higher similarity to the same face and reduces similarity between dissimilar faces. Previous research utilizing margin-based loss functions has achieved near-perfect accuracies on high-quality face recognition datasets. However, the same networks fail to perform well on low-quality images due to the degradation of facial attributes necessary for distinguishing different faces. In this paper, we tackle the problem of low-quality face recognition. We base our analysis on an observation that the change of loss functions produce marginal changes in performance for low-quality face recognition. Hence, rather than following the traditional approach of defining problem-specific regularized functions, we take a closer look at the nature of data in low resolution datasets and redefine paradigms in terms of model choice, data input pipeline and fine-tuning schemes. With the accumulated effect of all our design choices, we achieve state-of-the-art results in medium-quality benchmarks (IJB-B, IJB-C) as well as multiple challenging benchmarks for unconstrained face recognition (Tinyface, IJB-S and BRIAR), thereby opening up a new avenue of research in the area. The pretrained model are publically available in https://github.com/ Kartik-3004/PETALface
Nithin Gopalakrishnan Nair, Kartik Narayan, Maitreya Suin, Ram Prabhakar Kathirvel, Soraya Stevens, Joshua Gleason, Nathan Shnidman, Rama Chellappa, Vishal M. Patel
FG9
2025 Multi-Domain Biometric Recognition using Body Embeddings
abstract
Biometric recognition becomes increasingly challenging as we move away from the visible spectrum to infrared imagery, where domain discrepancies significantly impact identification performance. In this paper, we show that body embeddings perform better than face embeddings for cross-spectral person identification in medium-wave infrared (MWIR) and long-wave infrared (LWIR) domains. Due to the lack of multi-domain datasets, previous research on cross-spectral body identification - also known as Visible-Infrared Person Re-Identification (VI-ReID) - has primarily focused on individual infrared bands, such as near-infrared (NIR) or LWIR, separately. We address the multi-domain body recognition problem using the IARPA Janus Benchmark Multi-Domain Face (IJB-MDF) dataset, which requires the matching of shortwave infrared (SWIR), MWIR, and LWIR images against RGB (VIS) images. We leverage a vision transformer architecture to establish benchmark results on the IJB-MDF dataset and, through extensive experiments, provide valuable insights into the interrelation of infrared domains, the adaptability of VIS-pretrained models, the role of local semantic features in body-embeddings, and effective training strategies for small datasets. Additionally, we show that finetuning a body model, pretrained exclusively on VIS data, with a simple combination of cross-entropy and triplet losses achieves state-of-the-art mAP scores on the LLCM dataset.
Anirudh Nanduri, Siyuan Huang 0005, Rama Chellappa
FG3
2025 AeroGen: Ground-to-Air Generalization for Action Recognition
abstract
We address the problem of action recognition from aerial views using only ground-based videos for training. Due to the viewpoint-induced domain shift, models trained solely on ground videos exhibit significant performance degradation when naively applied to aerial videos. To mitigate this performance gap, we introduce a domain generalization technique that addresses the viewpoint-induced domain shift. Our method uses available real ground videos to generate additional synthetic training data from both ground and air viewpoints, for improving generalization to aerial video-based action recognition. Specifically, we perform 3D human mesh estimation from the ground videos, then render synthetic videos from alternate viewpoints with additional appearance randomizations. To further align the ground-air and syntheticreal domains, we propose a Dual Domain Alignment loss by enforcing consistency in predictions between the original ground videos and augmented videos from each domain. In order to facilitate research on the problem of ground-to-air generalization for human action recognition, we also create a benchmark by combining parts of the NTU-60, UAV-Human and NEC-DRONE datasets. We demonstrate the effectiveness of our approach on these new benchmarks along with the existing RoCoG-Ground $\rightarrow$ RoCoG-Air benchmark, and also perform extensive ablations.
Ketul Shah, Anshul Shah 0001, Arun V. Reddy, Aniket Roy, Celso de Melo, Rama Chellappa
FG6
2025 UniGait: A Unified Transformer-based Multitask Framework for Gait Analysis in the Wild
abstract
Gait recognition is a rapidly emerging and significant area of biometrics, leveraging the unique walking patterns of individuals to perform personal identification and facilitate healthcare monitoring, such as elderly care, fall detection, etc. While existing gait recognition methods perform well in indoor, or short-range environments, their effectiveness diminishes significantly when applied to unconstrained outdoor scenarios. Challenges such as environmental turbulence, occlusion, varying viewing angles contribute to this performance drop. To address these challenges and enhance gait recognition accuracy in real-world settings, while also expanding the functionality of gait features for healthcare applications, we propose a unified multitask framework called UniGait. UniGait is designed to perform a comprehensive range of gait analysis tasks, including gait recognition and estimation of gait-related human attributes. UniGait is built upon a transformer-based architecture, which leverages the power of a cross-attention mechanism to simultaneously process multiple sub-tasks. This multitask learning approach allows the model to extract more robust gait features by jointly learning gait recognition and human attribute estimation, leading to improved overall performance. We report the results of extensive experiments and analysis on large-scale, real-world datasets collected under challenging conditions, including long-range (up to 1000 meters) and high-pitch angles (including UAV-based data). The results demonstrated state-of-the-art performance, highlighting the potential of UniGait for deployment in real-world applications, making it a valuable tool for a range of biometric and healthcare monitoring scenarios.
Zhao-Yang Wang, Jiang Liu 0014, Yuxiang Guo 0001, Jieneng Chen, Rama Chellappa
FG5
2025 Mind the Gap: Bridging Occlusion in Gait Recognition via Residual Gap Correction
abstract
Gait is becoming popular as a method of person reidentification because of its ability to identify people at a distance. However, most current works in gait recognition do not address the practical problem of occlusions. Among those which do, some require paired tuples of occluded and holistic sequences, which are impractical to collect in the real world. Further, these approaches work on occlusions but fail to retain performance on holistic inputs. To address these challenges, we propose RG-Gait, a method for residual correction for occluded gait recognition with holistic retention. We model the problem as a residual learning task, conceptualizing the occluded gait signature as a residual deviation from the holistic gait representation. Our proposed network adaptively integrates the learned residual, significantly improving performance on occluded gait sequences without compromising the holistic recognition accuracy. We evaluate our approach on the challenging Gait3D, GREW and BRIAR datasets and show that learning the residual can be an effective technique to tackle occluded gait recognition with holistic retention. We release our code publicly at https://github.com/Ayush-00/rg-gait.
Ayush Gupta 0001, Siyuan Huang 0005, Rama Chellappa
IJCB3
2025 A Quantitative Evaluation of the Expressivity of BMI, Pose and Gender in Body Embeddings for Recognition and Identification
abstract
Person Re-identification (ReID) systems that match individuals across images or video frames are essential in many real-world applications. However, existing methods are often influenced by attributes such as gender, pose, and body mass index (BMI), which vary in unconstrained settings and raise concerns related to fairness and generalization. To address this, we extend the notion of expressivity, defined as the mutual information between learned features and specific attributes, using a secondary neural network to quantify how strongly attributes are encoded. Applying this framework to three ReID models, we find that BMI consistently shows the highest expressivity in the final layers, indicating its dominant role in recognition. In the last attention layer, attributes are ranked as BMI > Pitch > Gender > Yaw, revealing their relative influences in representation learning. Expressivity values also evolve across layers and training epochs, reflecting a dynamic encoding of attributes. These findings demonstrate the central role of body attributes in ReID and establish a principled approach for uncovering attribute driven correlations.
Basudha Pal, Siyuan Huang 0005, Rama Chellappa
IJCB3
2025 TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision
abstract
We address the problem of video question answering (video QA) with temporal grounding in a weakly supervised setup, without any temporal annotations. Given a video and a question, we generate an open-ended answer grounded with the start and end time. For this task, we propose TOGA: a vision-language model for Temporally Grounded Open-Ended Video QA with Weak Supervision. We instruct-tune TOGA to jointly generate the answer and the temporal grounding. We operate in a weakly supervised setup where the temporal grounding annotations are not available. We generate pseudo labels for temporal grounding and ensure the validity of these labels by imposing a consistency constraint between the question of a grounding response and the response generated by a question referring to the same temporal segment. We notice that jointly generating the answers with the grounding improves performance on question answering as well as grounding. We evaluate TOGA on grounded QA and open-ended QA tasks. For grounded QA, we consider the NExT-GQA benchmark which is designed to evaluate weakly supervised grounded question answering. For open-ended QA, we consider the MSVD-QA and ActivityNet-QA benchmarks. We achieve state-of-the-art performance for both tasks on these benchmarks.
Ayush Gupta 0001, Rama Chellappa, Nathaniel D. Bastian, Alvaro Velasquez, Susmit Jha
ICCV3
2025 FaceXFormer: A Unified Transformer for Facial Analysis
abstract
In this work, we introduce FaceXFormer, an end-to-end unified transformer model capable of performing ten facial analysis tasks within a single framework. These tasks include face parsing, landmark detection, head pose estimation, attribute prediction, age, gender, and race estimation, facial expression recognition, face recognition, and face visibility. Traditional face analysis approaches rely on task-specific architectures and pre-processing techniques, limiting scalability and integration. In contrast, FaceXFormer employs a transformer-based encoder-decoder architecture, where each task is represented as a learnable token, enabling seamless multi-task processing within a unified model. To enhance efficiency, we introduce FaceX, a lightweight decoder with a novel bi-directional cross-attention mechanism, which jointly processes face and task tokens to learn robust and generalized facial representations. We train FaceXFormer on ten diverse face perception datasets and evaluate it against both specialized and multi-task models across multiple benchmarks, demonstrating state-of-the-art or competitive performance. Additionally, we analyze the impact of various components of FaceXFormer on performance, assess real-world robustness in "in-the-wild" settings, and conduct a computational performance evaluation. To the best of our knowledge, FaceXFormer is the first model capable of handling ten facial analysis tasks while maintaining real-time performance at 33.21 FPS. Code: https://github.com/Kartik-3004/facexformer
Kartik Narayan, Vibashan VS, Rama Chellappa, Vishal M. Patel
ICCV3
2025 Enrich and Detect: Video Temporal Grounding With Multimodal Llms
abstract
We introduce ED-VTG, a method for fine-grained video temporal grounding utilizing multi-modal large language models. Our approach harnesses the capabilities of multimodal LLMs to jointly process text and video, in order to effectively localize natural language queries in videos through a two-stage process. Rather than being directly grounded, language queries are initially transformed into enriched sentences that incorporate missing details and cues to aid in grounding. In the second stage, these enriched queries are grounded, using a lightweight decoder, which specializes at predicting accurate boundaries conditioned on contextualized representations of the enriched queries. To mitigate noise and reduce the impact of hallucinations, our model is trained with a multiple-instance-learning objective that dynamically selects the optimal version of the query for each training sample. We demonstrate state-of-the-art results across various benchmarks in temporal video grounding and paragraph grounding settings. Experiments reveal that our method significantly outperforms all previously proposed LLM-based temporal grounding approaches and is either superior or comparable to specialized models, while maintaining a clear advantage against them in zero-shot evaluation scenarios.
Shraman Pramanick, Effrosyni Mavroudi, Yale Song, Rama Chellappa, Lorenzo Torresani, Triantafyllos Afouras
ICCV4
2025 DuoLoRA: Cycle-Consistent and Rank-Disentangled Content-Style Personalization
Aniket Roy, Shubhankar Borse, Shreya Kadambi, Debasmit Das, Shweta Mahajan, Risheek Garrepalli, Hyojin Park 0004, Ankita Nayak, Rama Chellappa, Munawar Hayat, Fatih Porikli
ICCV9
2025 ViT-Linearizer: Distilling Quadratic Knowledge into Linear-Time Vision Models
abstract
Vision Transformers (ViTs) have delivered remarkable progress through global self-attention, yet their quadratic complexity can become prohibitive for high-resolution inputs. In this work, we present ViT-Linearizer, a cross-architecture distillation framework that transfers rich ViT representations into a linear-time, recurrent-style model. Our approach leverages 1) activation matching, an intermediate constraint that encourages student to align its token-wise dependencies with those produced by the teacher, and 2) masked prediction, a contextual reconstruction objective that requires the student to predict the teacher's representations for unseen (masked) tokens, to effectively distill the quadratic self-attention knowledge into the student while maintaining efficient complexity. Empirically, our method provides notable speedups particularly for high-resolution tasks, significantly addressing the hardware challenges in inference. Additionally, it also elevates Mamba-based architectures' performance on standard vision benchmarks, achieving a competitive 84.3% top-1 accuracy on ImageNet with a base-sized model. Our results underscore the good potential of RNN-based solutions for large-scale visual tasks, bridging the gap between theoretical efficiency and real-world practice.
Guoyizhe Wei, Rama Chellappa
ICCV2
2025 Medical World Model
Zhao-Yang Wang, Qiuping Liu, Shuwen Sun, Kang Wang 0016, Rama Chellappa, Zongwei Zhou, Alan L. Yuille, Lei Zhu 0003, Jieneng Chen
ICCV6
2025 METAREG: Robust Camera Parameter Estimation by Leveraging Noisy Camera Extrinsics
abstract
Novel view synthesis methods, such as Neural Radiance Fields (NeRFs) and 3D Gaussian Splatting (3DGS), rely on Structure-from-Motion (SfM) pipelines like COLMAP for camera parameter estimation. However, these pipelines are prone to errors due to factors like doppelgangers and perspective distortion. While edge devices (e.g., mobile phones) capture images with embedded GPS and IMU data, this meta-data is often noisy. We introduce MetaReg, a robust pipeline that improves camera parameter estimation by leveraging noisy GPS metadata. MetaReg enhances COLMAP with pre-and post-processing steps: the preprocessing stage estimates image overlap using metadata to reduce unnecessary image matching, while the post-processing stage aligns estimated camera coordinates with the world coordinate system using noisy GPS priors. Experiments on challenging datasets demonstrate that MetaReg significantly improves camera parameter estimation, enhancing robustness and accuracy.
Chandrakanth Gudavalli, Tajuddin Manhar Mohammed, Ananth Vishnu Bhaskar, Elliot Staudt, Cheng Peng 0008, Abhay Yadav, Rama Chellappa, Shivkumar Chandrasekaran, B. S. Manjunath
ICIP7
2025 DIFFUSE2ADAPT: Controlled Diffusion for Synthetic-to-Real Domain Adaptation
abstract
Synthetic data generated from graphics engines has been shown to be effective for learning, while also being a cost-effective alternative to annotating real-world data. However, models trained on synthetic data often suffer from performance degradation when applied to real-world data due to the domain gap. In this paper, we propose Diffuse2Adapt, a novel unsupervised domain adaptation (UDA) approach that leverages controlled diffusion models to bridge the synthetic-to-real domain gap. Our method utilizes text-to-image generative models to translate synthetic images to the target domain while preserving class semantics. We introduce two methods to reduce the domain gap: (1) incorporating target domain context extracted from multimodal language models, and (2) capturing target domain style via learned textual tokens. Extensive experiments are performed on three synthetic-to-real domain adaptation benchmarks, VisDA-2017, S2RDA-49, and S2RDA-MS-39. Diffuse2Adapt outperforms state-of-the-art methods by +3.00% on VisDA-2017, +2.55% on S2RDA-49 and +1.01% on S2RDA-MS-39. Code will be released.
Ketul Shah, Arushi Sinha, Arun V. Reddy, Aniket Roy, Rama Chellappa
ICIP5
2025 ConceptAgent: LLM-Driven Precondition Grounding and Tree Search for Robust Task Planning and Execution
abstract
Robotic planning and execution in open-world environments is a complex problem due to the vast state spaces and high variability of task embodiment. Recent advances in perception algorithms, combined with Large Language Models (LLMs) for planning, offer promising solutions to these challenges, as the common sense reasoning capabilities of LLMs provide a strong heuristic for efficiently searching the action space. However, prior work fails to address the possibility of hallucinations from LLMs, which results in failures to execute the planned actions largely due to logical fallacies at high-or low-levels. To contend with automation failure due to such hallucinations, we introduce ConceptAgent, a natural language-driven robotic platform designed for task execution in unstructured environments. With a focus on scalability and reliability of LLM-based planning in complex state and action spaces, we present innovations designed to limit these shortcomings, including 1) Predicate Grounding to prevent and recover from infeasible actions, and 2) an embodied version of LLM-guided Monte Carlo Tree Search with self reflection. ConceptAgent combines these planning enhancements with dynamic language aligned 3d scene graphs, and large multi-modal pretrained models to perceive, localize, and interact with its environment, enabling reliable task completion. In simulation experiments, ConceptAgent achieved a 19% task completion rate across three room layouts and 30 easy level embodied tasks outperforming other state-of-the-art LLM-driven reasoning baselines that scored 10.26% and 8.11% on the same benchmark. Additionally, ablation studies on moderate to hard embodied tasks revealed a 20% increase in task completion from the baseline agent to the fully enhanced ConceptAgent, highlighting the individual and combined contributions of Predicate Grounding and LLM-guided Tree Search to enable more robust automation in complex state and action spaces. Additionally, in real-world mobile manipulation trials, conducted in randomized, low-clutter environments, a ConceptAgent-driven Spot robot achieved a 40% task completion rate, demonstrating the performance of our perception system in real-world scenarios.
Corban Rivera, Grayson Byrd, William Paul, Tyler Feldman, Meghan Booker, Emma Holmes, David Handelman, Bethany Kemp, Andrew Badger, Aurora Schmidt, Krishna Murthy Jatavallabhula, Celso de Melo, Seenivasan Lalithkumar, Mathias Unberath, Rama Chellappa
ICRA15
2025 MS-GS: Multi-Appearance Sparse-View 3D Gaussian Splatting in the Wild
abstract
In-the-wild photo collections often contain limited volumes of imagery and exhibit multiple appearances, e.g., taken at different times of day or seasons, posing significant challenges to scene reconstruction and novel view synthesis. Although recent adaptations of Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) have improved in these areas, they tend to oversmooth and are prone to overfitting. In this paper, we present MS-GS, a novel framework designed with \textbf{M}ulti-appearance capabilities in \textbf{S}parse-view scenarios using 3D\textbf{GS}. To address the lack of support due to sparse initializations, our approach is built on the geometric priors elicited from monocular depth estimations. The key lies in extracting and utilizing local semantic regions with a Structure-from-Motion (SfM) points anchored algorithm for reliable alignment and geometry cues. Then, to introduce multi-view constraints, we propose a series of geometry-guided supervision steps at virtual views in pixel and feature levels to encourage 3D consistency and reduce overfitting. We also introduce a dataset and an in-the-wild experiment setting to set up more realistic benchmarks. We demonstrate that MS-GS achieves photorealistic renderings under various challenging sparse-view and multi-appearance conditions, and outperforms existing approaches significantly across different datasets.
Deming Li, Yutao Tang, Ravi Ramamoorthi, Rama Chellappa, Cheng Peng 0008
NeurIPS5
2025 GaitContour: Efficient Gait Recognition Based on a Contour-Pose Representation
abstract
Gait recognition holds the promise to robustly identify subjects based on walking patterns instead of appearance information. In recent years, this field has been dominated by learning methods based on two input formats: silhouette images and sparse keypoints. Compared to image-based approaches, keypoint-based methods can achieve significantly higher efficiency due to their sparsity. However, sparsity also results in information loss, thereby reducing performance. In this work, we propose a novel, keypoint-based Contour-Pose representation, which compactly encodes both body shape and parts information. We further propose a local-to-global architecture, called GaitContour, to leverage this novel representation and efficiently compute subject embedding in two stages. The first stage consists of a local transformer that extracts features from five different body regions. The second stage then aggregates the regional features to estimate a global human gait representation. Such a design significantly reduces the complexity of the attention operation and improves both efficiency and performance. Through large scale experiments, Gait-Contour is shown to perform significantly better than previous keypoint-based methods. Furthermore, the ContourPose representation also achieves new SoTA performances on fusion-based gait recognition methods.
Yuxiang Guo 0001, Anshul Shah 0001, Jiang Liu 0014, Ayush Gupta 0001, Rama Chellappa, Cheng Peng 0008
WACV5
2025 MimicGait: A Model Agnostic approach for Occluded Gait Recognition Using Correlational Knowledge Distillation
abstract
Gait recognition is an important biometric technique over large distances. State-of-the-art gait recognition systems perform very well in controlled environments at close range. Recently, there has been an increased interest in gait recognition in the wild prompted by the collection of outdoor, more challenging datasets containing variations in terms of illumination, pitch angles and distances. An important problem in these environments is that of occlusion, where the subject is partially blocked from camera view. While important, this problem has received little attention. Thus, we propose MimicGail, a model-agnostic approach for gait recognition in the presence of occlusions. We train the network using a multi-instance correlational distillation loss to capture both inter-sequence and intra-sequence correlations in the occluded gait patterns of a subject, utilizing an auxiliary Visibility Estimation Network to guide the training of the proposed mimic network. We demonstrate the effectiveness of our approach on challenging real-world datasets like GREW, Gait3D and BRIAR. We release the code in https://github.com/Ayush-00/mimicgait.
Ayush Gupta 0001, Rama Chellappa
WACV2
2025 VILLS: Video-Image Learning to Learn Semantics for Person Re-Identification
abstract
Person Re-identification is a research area with significant real world applications. Despite recent progress, existing methods face challenges in robust re-identification in the wild, e.g., by focusing only on a particular modality and on unreliable patterns such as clothing. A generalized method is highly desired, but remains elusive to achieve due to issues such as the trade-off between spatial and temporal resolution and inaccurate feature extraction. We propose VILLS (Video-Image Learning to Learn Semantics), a self-supervised method that jointly learns spatial and temporal features from images and videos. VILLS first designs a local semantic extraction module that adaptively extracts semantically consistent and robust spatial features. Then, VILLS designs a unified feature learning and adaptation module to represent image and video modalities in a consistent feature space. By Leveraging self-supervised, large-scale pre-training, VILLS establishes a new State-of-The-Art that significantly outperforms existing image and video-based methods.
Siyuan Huang 0005, Ram Prabhakar, Yuxiang Guo 0001, Rama Chellappa, Cheng Peng 0008
WACV4
2025 PETALface: Parameter Efficient Transfer Learning for Low-Resolution Face Recognition
abstract
Pre-training on large-scale datasets and utilizing margin-based loss functions have been highly successful in training models for high-resolution face recognition. However, these models struggle with low-resolution face datasets, in which the faces lack the facial attributes necessary for distinguishing different faces. Full fine-tuning on low-resolution datasets, a naive method for adapting the model, yields inferior performance due to catastrophic for-getting of pre-trained knowledge. Additionally the domain difference between high-resolution (HR) gallery images and low-resolution (LR) probe images in low resolution datasets leads to poor convergence for a single model to adapt to both gallery and probe after fine-tuning. To this end, we propose PETALface, a Parameter-Efficient Transfer Learning approach for low-resolution face recognition. Through PETALface, we attempt to solve both the aforementioned problems. (1) We solve catastrophic forgetting by leveraging the power of parameter efficient fine-tuning(PEFT). (2) We introduce two low-rank adaptation modules to the back-bone, with weights adjusted based on the input image quality to account for the difference in quality for the gallery and probe images. To the best of our knowledge, PETALface is the first work leveraging the powers of PEFT for low resolution face recognition. Extensive experiments demonstrate that the proposed method outperforms full fine-tuning on low-resolution datasets while preserving performance on high-resolution and mixed-quality datasets, all while using only 0.48% of the parameters.
Kartik Narayan, Nithin Gopalakrishnan Nair, Rama Chellappa, Vishal M. Patel
WACV4
2025 CEMIL: Contextual Attention Based Efficient Weakly Supervised Approach for Histopathology Image Classification
abstract
Multiple Instance Learning (MIL) has shown potential for analyzing Whole Slide Images (WSIs) in digital pathology, but it faces challenges related to redundant information learning and generalization due to limited supervision and the computational complexity of Gigapixel WSIs. Many MIL-based methods apply a small weight matrix to all WSI patches. In this study, we focus on developing computationally efficient models that improve MIL-based WSI classification by processing fewer patches while improving performance. We propose an attention-based approach using knowledge distillation, where a compute-intensive “instructor” model analyzes all WSI patches to train a resourceefficient “learner” model, which considers only a subset of patches. Comprehensive evaluations on four cancer subtype datasets—TCGA-BRCA, TCGA-NSCLC, TCGA-RCC, and PANDA—demonstrate that an “observe-everything” instructor can effectively train an “observe-minimally” learner network. Overall, our proposed learner network enhances performance by 4% compared to the state-of-the-art, while reducing inference time by 45% and FLOPs by approximately 88%.
Tawsifur Rahman, Alexander S. Baras, Rama Chellappa
WACV3
2025 Cap2Aug: Caption Guided Image data Augmentation
abstract
Visual recognition in a low-data regime is challenging and often prone to overfitting. To mitigate this issue, several data augmentation strategies have been proposed. However, standard transformations, e.g., rotation, cropping, and flip-ping provide limited semantic variations. To this end, we propose Cap2Aug, an image-to-image diffusion model-based data augmentation strategy using image captions to condition the image synthesis step. We generate a caption for an image and use this caption as an additional input for an image-to-image diffusion model. This increases the semantic diversity of the augmented images due to caption conditioning compared to the usual data augmentation techniques. We show that Cap2Aug is particularly effective where only a few samples are available for an object class. However, naively generating the synthetic images is not adequate due to the domain gap between real and synthetic images. Thus, we employ a maximum mean discrepancy loss to align the synthetic images to the real images to minimize the domain gap. We evaluate our method on few-shot classification and image classification with long-tail class distribution tasks. Cap2Aug achieves state-of-the-art performance on both tasks while evaluated on eleven benchmarks. Code: https://github.com/aniket004/Cap_2_Aug.git
Aniket Roy, Anshul Shah 0001, Ketul Shah, Rama Chellappa
WACV5
2025 VM-Gait: Multi-Modal 3D Representation Based on Virtual Marker for Gait Recognition
abstract
Gait recognition plays a vital role in biometric applications by analyzing the unique characteristics of an individ-ual's walking pattern. Methods based on 2D representations, such as silhouettes and skeletons, are increasingly being developed to learn the shape features and joint dy-namic movements. Nevertheless, the effectiveness of 2D representation-based methods is impeded by factors such as changes in viewpoint, partial occlusion, and noisy en-vironments. 3D representation-based methods can complement 2D representation-based approaches by providing more precise dynamic body shapes and motion information, along with increased robustness against changes in viewpoint and partial occlusion. However, the complex-ity of acquiring accurate 3D representations and the chal-lenges associated with extracting dynamic topological features from sequences of 3D representations hinder the de-velopment of 3D representations-based methods. In this pa-per, we present VM-Gait, a novel multi-modal gait recognition framework that harnesses the advantages of integrating both 2D and 3D representations. Furthermore, we in-troduce a new 3D representation, Virtual Marker, into gait recognition to efficiently learn topological features from 3D representations, avoiding the computational complexi-ties inherent in directly learning from 3D representations like 3D meshes or 3D point clouds. Extensive experiments demonstrate that the proposed framework effectively learns and fuses discriminative information from different gait modalities, enhancing gait recognition performance.
Zhao-Yang Wang, Jiang Liu 0014, Jieneng Chen, Rama Chellappa
WACV4
2025 StimuVAR: Spatiotemporal Stimuli-Aware Video Affective Reasoning with Multimodal Large Language Models
Yuxiang Guo 0001, Yang Zhao 0024, Rama Chellappa, Shao-Yuan Lo
Int. J. Comput. Vis.4
2024 Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model
abstract
The ability of large language models (LLMs) to process visual inputs has given rise to general-purpose vision systems, unifying various vision-language (VL) tasks by instruction tuning. However, due to the enormous diversity in input-output formats in the vision domain, existing general-purpose models fail to successfully integrate segmentation and multi-image inputs with coarse-level tasks into a single framework. In this work, we introduce VistaLLM, a powerful visual system that addresses coarse- and fine-grained VL tasks over single and multiple input images using a unified framework. VistaLLM utilizes an instruction-guided image tokenizer that filters global embeddings using task descriptions to extract compressed and refined features from numerous images. Moreover, VistaLLM employs a gradient-aware adaptive sampling technique to represent binary segmentation masks as sequences, significantly improving over previously used uniform sampling. To bolster the desired capability of VistaLLM, we curate CoinIt, a comprehensive coarse-to-fine instruction tuning dataset with 6.8M samples. We also address the lack of multi-image grounding datasets by introducing a novel task, AttCoSeg (Attribute-level Co-Segmentation), which boosts the model's reasoning and grounding capability over multiple input images. Extensive experiments on a wide range of V- and VL tasks demonstrate the effectiveness of VistaLLM by achieving consistent state-of-the-art performance over strong baselines across many downstream tasks. Our project page can be found at https://shramanpramanick.github.io/VistaLLM/.
Shraman Pramanick, Guangxing Han, Sayan Nag, Ser-Nam Lim, Nicolas Ballas, Qifan Wang 0001, Rama Chellappa, Amjad Almahairi
CVPR8
2024 Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-Training
abstract
In this work, we tackle the problem of unsupervised domain adaptation (UDA) for video action recognition. Our approach, which we call UNITE, uses an image teacher model to adapt a video student model to the target domain. UNITE first employs self-supervised pretraining to promote discriminative feature learning on target domain videos using a teacher-guided masked distillation objective. We then perform self-training on masked target data, using the video student model and image teacher model together to generate improved pseudolabels for unlabeled target videos. Our self-training process successfully leverages the strengths of both models to achieve strong transfer performance across domains. We evaluate our approach on multiple video domain adaptation benchmarks and observe significant improvements upon previously reported results.
Arun V. Reddy, William Paul, Corban Rivera, Ketul Shah, Celso de Melo, Rama Chellappa
CVPR6
2024 GAMMA-FACE: GAussian Mixture Models Amend Diffusion Models for Bias Mitigation in Face Images
Basudha Pal, Arunkumar Kannan, Ram Prabhakar Kathirvel, Alice J. O'Toole, Rama Chellappa
ECCV (67)5
2024 BAGS: Blur Agnostic Gaussian Splatting Through Multi-scale Kernel Modeling
Cheng Peng 0008, Yutao Tang, Nengyu Wang, Xijun Liu, Deming Li, Rama Chellappa
ECCV (80)7
2024 Identifying Attack-Specific Signatures in Adversarial Examples
abstract
The adversarial attack literature contains numerous algorithms for crafting perturbations which manipulate neural network predictions. Many of these adversarial attacks optimize inputs with the same constraints and have similar downstream impact on the models they attack. In this work, we first show how to reconstruct an adversarial perturbation, namely the difference between an adversarial example and the original natural image, from an adversarial example. Then, we classify reconstructed adversarial perturbations based on the algorithm that generated them. This pipeline, REDRL, can detect the attack algorithm used to generate a sample from only the sample itself. The ability to determine which algorithm generated an example implies that different attack algorithms actually produce unique signatures in their adversarial examples.
Hossein Souri, Pirazh Khorramshahi, Chun Pong Lau 0001, Micah Goldblum, Rama Chellappa
ICASSP5
2024 Distillation-guided Representation Learning for Unconstrained Gait Recognition
abstract
Gait recognition holds the promise of robustly identifying subjects based on walking patterns instead of appearance information. While previous approaches have performed well for curated indoor data, they tend to underperform in unconstrained situations, e.g. in outdoor, long distance scenes, etc. We propose a framework, termed GAit DEtection and Recognition (GADER), for human authentication in challenging outdoor scenarios. Specifically, GADER leverages a Double Helical Signature to detect segments that contain human movement and builds discriminative features through a novel gait recognition method, where only frames containing gait information are used. To further enhance robustness, GADER encodes viewpoint information in its architecture, and distills representation from an auxiliary RGB recognition model, which enables GADER to learn from silhouette and RGB data at training time. At test time, GADER only infers from the silhouette modality. We evaluate our method on multiple State-of-The-Arts(SoTA) gait baselines and demonstrate consistent improvements on indoor and outdoor datasets, especially with a significant 25.2% improvement on unconstrained, remote gait data.
Yuxiang Guo 0001, Siyuan Huang 0005, Ram Prabhakar, Chun Pong Lau 0001, Rama Chellappa, Cheng Peng 0008
IJCB5
2024 Template-based Multi-Domain Face Recognition
abstract
Despite the remarkable performance of deep neural networks for face detection and recognition tasks in the visible spectrum, their performance on more challenging non-visible domains is comparatively still lacking. While significant research has been done in the fields of domain adaptation and domain generalization, in this paper we tackle scenarios in which these methods have limited applicability owing to the lack of training data from target domains. We focus on the problem of single-source (visible) and multi-target (SWIR, long-range/remote, surveillance, and body-worn) face recognition task. We show through experiments that a good template generation algorithm becomes crucial as the complexity of the target domain increases. In this context, we introduce a template generation algorithm called Norm Pooling (and a variant known as Sparse Pooling) and show that it outperforms average pooling across different domains and networks, on the IARPA JANUS Benchmark Multi-domain Face (IJB-MDF) dataset.
Anirudh Nanduri, Rama Chellappa
IJCB2
2024 DiversiNet: Mitigating Bias in Deep Classification Networks across Sensitive Attributes through Diffusion-Generated Data
abstract
Deep learning models trained on sensitive data often show biases towards certain demographics, posing fairness challenges, especially with limited datasets. Diffusion generated data effectively supplement the underrepresented dataset, serving as a regularization technique to enhance feature learning. In addition to the original balanced dataset, we incorporate synthetic data generated by the diffusion model to train classifiers and subsequently assess their performance. Experimental results demonstrate a reduction in bias across all target attributes along with an increase in overall accuracy. For instance, for gender classification in the FFHQ dataset, the overall accuracy rises to 94.44% from 93.92% after including data generated from a diffusion model. Simultaneously, the bias, measured as the absolute difference between the true positive rates of young and old individuals, decreases from 0.0340 to 0.0204 (reduction of 40%). Moreover, we extend our analysis to multi-attribute scenarios, successfully mitigating bias with respect to multiple sensitive attributes simultaneously in sensitive attribute classification as well as in other downstream tasks. To the best of our knowledge, this study introduces a novel approach to bias mitigation, highlighting the versatility of diffusion-based data augmentation in addressing biases concerning age, gender, and race.
Basudha Pal, Aniket Roy, Ram Prabhakar Kathirvel, Alice J. O'Toole, Rama Chellappa
IJCB5
2024 HyperGait: A Video-based Multitask Network for Gait Recognition and Human Attribute Estimation at Range and Altitude
abstract
Gait recognition is one of the mainstream approaches for identifying individuals when face information is not available. Most previous methods achieve good performance on structured indoor walking sequences with silhouettes provided. However, when these methods are applied to unconstrained outdoor sequences, a significant reduction in performance is inevitably observed due to factors such as turbulence, occlusion, view angle, and oversized clothing. To make gait recognition methods stable and effective for real-world settings, we extend gait-only-based approaches by introducing more useful biometric information such as gender, age, height, weight, and body mass index to cooperatively work with the gait recognition module. In this paper, we propose a video-based multitasking network for gait recognition and human attribute prediction at ranges of up to 1000 meters and high-pitch angles to mutually improve the robustness and accuracy of each task. Through a series of experiments on OU-MVLP and BRIAR datasets, we show that our multitasking network outperforms previous methods and provides more useful biometric information for human identification tasks.
Zhao-Yang Wang, Jiang Liu 0014, Ram Prabhakar Kathirvel, Chun Pong Lau 0001, Rama Chellappa
IJCB5
2024 Robust Feature Space Organization with Distillation for Few-Shot Object Detection
Vineet R. Shenoy, Rama Chellappa
ICPR (7)2
2024 ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning
abstract
For robots to perform a wide variety of tasks, they require a 3D representation of the world that is semantically rich, yet compact and efficient for task-driven perception and planning. Recent approaches have attempted to leverage features from large vision-language models to encode semantics in 3D representations. However, these approaches tend to produce maps with per-point feature vectors, which do not scale well in larger environments, nor do they contain semantic spatial relationships between entities in the environment, which are useful for downstream planning. In this work, we propose ConceptGraphs, an open-vocabulary graph-structured representation for 3D scenes. ConceptGraphs is built by leveraging 2D foundation models and fusing their output to 3D by multi-view association. The resulting representations generalize to novel semantic classes, without the need to collect large 3D datasets or finetune models. We demonstrate the utility of this representation through a number of downstream planning tasks that are specified through abstract (language) prompts and require complex reasoning over spatial and semantic concepts. To explore the full scope of our experiments and results, we encourage readers to visit our project webpage.
Qiao Gu, Alihusein Kuwajerwala, Sacha Morin, Krishna Murthy Jatavallabhula, Bipasha Sen, Aditya Agarwal, Corban Rivera, William Paul, Kirsty Ellis, Rama Chellappa, Chuang Gan 0001, Celso de Melo, Josh Tenenbaum, Antonio Torralba 0001, Florian Shkurti, Liam Paull
ICRA10
2024 CLR-Face: Conditional Latent Refinement for Blind Face Restoration Using Score-Based Diffusion Models
Maitreya Suin, Rama Chellappa
IJCAI2
2024 SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers
abstract
Seeking answers to questions within long scientific research articles is a crucial area of study that aids readers in quickly addressing their inquiries. However, existing question-answering (QA) datasets based on scientific papers are limited in scale and focus solely on textual content. We introduce SPIQA (Scientific Paper Image Question Answering), the first large-scale QA dataset specifically designed to interpret complex figures and tables within the context of scientific research articles across various domains of computer science. Leveraging the breadth of expertise and ability of multimodal large language models (MLLMs) to understand figures, we employ automatic and manual curation to create the dataset. We craft an information-seeking task on interleaved images and text that involves multiple images covering a wide variety of plots, charts, tables, schematic diagrams, and result visualizations. SPIQA comprises 270K questions divided into training, validation, and three different evaluation splits. Through extensive experiments with 12 prominent foundational models, we evaluate the ability of current multimodal systems to comprehend the nuanced aspects of research articles. Additionally, we propose a Chain-of-Thought (CoT) evaluation strategy with in-context retrieval that allows fine-grained, step-by-step assessment and improves model performance. We further explore the upper bounds of performance enhancement with additional textual information, highlighting its promising potential for future research and the dataset’s impact on revolutionizing how we interact with scientific literature.
Shraman Pramanick, Rama Chellappa, Subhashini Venugopalan
NeurIPS2
2024 LP-3DGS: Learning to Prune 3D Gaussian Splatting
abstract
Recently, 3D Gaussian Splatting (3DGS) has become one of the mainstream methodologies for novel view synthesis (NVS) due to its high quality and fast rendering speed. However, as a point-based scene representation, 3DGS potentially generates a large number of Gaussians to fit the scene, leading to high memory usage. Improvements that have been proposed require either an empirical pre-set pruning ratio or importance score threshold to prune the point cloud. Such hyperparameters require multiple rounds of training to optimize and achieve the maximum pruning ratio while maintaining the rendering quality for each scene. In this work, we propose learning-to-prune 3DGS (LP-3DGS), where a trainable binary mask is applied to the importance score to automatically find a favorable pruning ratio. Instead of using the traditional straight-through estimator (STE) method to approximate the binary mask gradient, we redesign the masking function to leverage the Gumbel-Sigmoid method, making it differentiable and compatible with the existing training process of 3DGS. Extensive experiments have shown that LP-3DGS consistently achieves a good balance between efficiency and high quality.
Zhaoliang Zhang, Tianchen Song, Li Yang 0009, Cheng Peng 0008, Rama Chellappa, Deliang Fan
NeurIPS6
2024 You Can Run but not Hide: Improving Gait Recognition with Intrinsic Occlusion Type Awareness
abstract
While gait recognition has seen many advances in recent years, the occlusion problem has largely been ignored. This problem is especially important for gait recognition from uncontrolled outdoor sequences at range - since any small obstruction can affect the recognition system. Most current methods assume the availability of complete body information while extracting the gait features. When parts of the body are occluded, these methods may hallucinate and output a corrupted gait signature as they try to look for body parts which are not present in the input at all. To address this, we exploit the learned occlusion type while extracting identity features from videos. Thus, in this work, we propose an occlusion aware gait recognition method which can be used to model intrinsic occlusion awareness into potentially any state-of-the-art gait recognition method. Our experiments on the challenging GREW and BRIAR datasets show that networks enhanced with this occlusion awareness perform better at recognition tasks than their counterparts trained on similar occlusions.
Ayush Gupta 0001, Rama Chellappa
WACV2
2024 Diffuse and Restore: A Region-Adaptive Diffusion Model for Identity-Preserving Blind Face Restoration
abstract
Blind face restoration (BFR) from severely degraded face images in the wild is a highly ill-posed problem. Due to the complex unknown degradation, existing generative works typically struggle to restore realistic details when the input is of poor quality. Recently, diffusion-based approaches were successfully used for high-quality image synthesis. But, for BFR, maintaining a balance between the fidelity of the restored image and the reconstructed identity information is important. Minor changes in certain facial regions may alter the identity or degrade the perceptual quality. With this observation, we present a conditional diffusion-based framework for BFR. We alleviate the drawbacks of existing diffusion-based approaches and design a region-adaptive strategy. Specifically, we use an identity preserving conditioner network to recover the identity information from the input image as much as possible and use that to guide the reverse diffusion process, specifically for important facial locations that contribute the most to the identity. This leads to a significant improvement in perceptual quality as well as face-recognition scores over existing GAN and diffusion-based restoration models. Our approach achieves superior results to prior art on a range of real and synthetic datasets, particularly for severely degraded face images.
Maitreya Suin, Nithin Gopalakrishnan Nair, Chun Pong Lau 0001, Vishal M. Patel, Rama Chellappa
WACV5
2024 Guest Editorial: Special Issue on ACCV 2022
Lei Wang 0001, Juergen Gall, Tat-Jun Chin, Imari Sato, Rama Chellappa
Int. J. Comput. Vis.5
2023 PDRF: Progressively Deblurring Radiance Field for Fast Scene Reconstruction from Blurry Images
abstract
We present Progressively Deblurring Radiance Field (PDRF), a novel approach to efficiently reconstruct high quality radiance fields from blurry images. While current State-of-The-Art (SoTA) scene reconstruction methods achieve photo-realistic renderings from clean source views, their performances suffer when the source views are affected by blur, which is commonly observed in the wild. Previous deblurring methods either do not account for 3D geometry, or are computationally intense. To addresses these issues, PDRF uses a progressively deblurring scheme for radiance field modeling, which can accurately model blur with 3D scene context. PDRF further uses an efficient importance sampling scheme that results in fast scene optimization. We perform extensive experiments and show that PDRF is 15X faster than previous SoTA while achieving better performance on both synthetic and real scenes.
Cheng Peng 0008, Rama Chellappa
AAAI2
2023 HaLP: Hallucinating Latent Positives for Skeleton-based Self-Supervised Learning of Actions
abstract
Supervised learning of skeleton sequence encoders for action recognition has received significant attention in recent times. However, learning such encoders without labels continues to be a challenging problem. While prior works have shown promising results by applying contrastive learning to pose sequences, the quality of the learned representations is often observed to be closely tied to data augmentations that are used to craft the positives. However, augmenting pose sequences is a difficult task as the geometric constraints among the skeleton joints need to be enforced to make the augmentations realistic for that action. In this work, we propose a new contrastive learning approach to train models for skeleton-based action recognition without labels. Our key contribution is a simple module, HaLP - to Hallucinate Latent Positives for contrastive learning. Specifically, HaLP explores the latent space of poses in suitable directions to generate new positives. To this end, we present a novel optimization formulation to solve for the synthetic positives with an explicit control on their hardness. We propose approximations to the objective, making them solvable in closed form with minimal overhead. We show via experiments that using these generated positives within a standard contrastive learning framework leads to consistent improvements across benchmarks such as NTU-60, NTU-120, and PKU-II on tasks like linear evaluation, transfer learning, and kNN evaluation. Our code can be found at https://github.com/anshulbshah/HaLP.
Anshul Shah 0001, Aniket Roy, Ketul Shah, Shlok Kumar Mishra, David Jacobs 0001, Anoop Cherian, Rama Chellappa
CVPR7
2023 Multi-Modal Human Authentication Using Silhouettes, Gait and RGB
abstract
Whole-body-based human authentication is a promising approach for remote biometrics scenarios. Current literature focuses on either body recognition based on RGB images or gait recognition based on body shapes and walking patterns; both have their advantages and drawbacks. In this work, we propose Dual-Modal Ensemble (DME), which combines both RGB and silhouette data to achieve more robust performances for indoor and outdoor whole-body based recognition. Within DME, we propose GaitPattern, which is inspired by the double helical gait pattern used in traditional gait analysis. The GaitPattern contributes to robust identification performance over a large range of viewing angles. Extensive experimental results on the CASIA-B dataset demonstrate that the proposed method outperforms state-of-the-art recognition systems. We also provide experimental results using the newly collected BRIAR dataset.
Yuxiang Guo 0001, Cheng Peng 0008, Chun Pong Lau 0001, Rama Chellappa
FG4
2023 ATDetect: Face Detection and Keypoint Extraction at Range and Altitude
abstract
Face detection and alignment are the crucial preprocessing steps in face recognition. While face detection works well in ideal situations, the performance deteriorates significantly when the image is degraded, due to factors such as blur, deformation, low resolution, and extreme headpose. However, there are very few works on face detection with realistic data captured from a long range (100m to 500m) and high altitude (30° to 50° pitch angle). We first evaluated several state-of-the-art methods on data collected at ranges of 100-500 meters and large pitch angles. One challenge is videos captured from long ranges usually lack bounding boxes and keypoint annotations, needed for training deep networks. This motivates us to develop a face detection and alignment algorithm that could perform effectively on videos captured from a long range and high altitude without groundtruth annotations. Moreover, meta information such as age, gender, and headpose of the subject could help face recognition. Therefore, we propose a single-stage face localization model ATDetect, which detects face bounding boxes, keypoints, and meta information simultaneously with realistic video captured at range and altitude.
Chun Pong Lau 0001, Maitreya Suin, Rama Chellappa
IJCB3
2023 EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone
abstract
Video-language pre-training (VLP) has become increasingly important due to its ability to generalize to various vision and language tasks. However, existing egocentric VLP frameworks utilize separate video and language encoders and learn task-specific cross-modal information only during fine-tuning, limiting the development of a unified system. In this work, we introduce the second generation of egocentric video-language pre-training (EgoVLPv2), a significant improvement from the previous generation, by incorporating cross-modal fusion directly into the video and language backbones. EgoVLPv2 learns strong video-text representation during pre-training and reuses the cross-modal attention modules to support different downstream tasks in a flexible and efficient manner, reducing fine-tuning costs. Moreover, our proposed fusion in the backbone strategy is more lightweight and compute-efficient than stacking additional fusion-specific layers. Extensive experiments on a wide range of VL tasks demonstrate the effectiveness of EgoVLPv2 by achieving consistent state-of-the-art performance over strong baselines across all downstream. Our project page can be found at https://shramanpramanick.github.io/EgoVLPv2/.
Shraman Pramanick, Yale Song, Sayan Nag, Qinghong Lin, Hardik Shah, Zheng Shou 0001, Rama Chellappa, Pengchuan Zhang
ICCV7
2023 MOST: Multiple Object localization with Self-supervised Transformers for object discovery
abstract
We tackle the challenging task of unsupervised object localization in this work. Recently, transformers trained with self-supervised learning have been shown to exhibit object localization properties without being trained for this task. In this work, we present Multiple Object localization with Self-supervised Transformers (MOST) that uses features of transformers trained using self-supervised learning to localize multiple objects in real world images. MOST analyzes the similarity maps of the features using box counting; a fractal analysis tool to identify tokens lying on foreground patches. The identified tokens are then clustered together, and tokens of each cluster are used to generate bounding boxes on foreground regions. Unlike recent state-of-the-art object localization methods, MOST can localize multiple objects per image and outperforms SOTA algorithms on several object localization and discovery benchmarks on PASCAL-VOC 07, 12 and COCO20k datasets. Additionally, we show that MOST can be used for self-supervised pretraining of object detectors, and yields consistent improvements on fully, semi-supervised object detection and unsupervised region proposal generation.Our project is publicly available at rssaketh.github.io/most.
Sai Saketh Rambhatla, Ishan Misra, Rama Chellappa, Abhinav Shrivastava
ICCV3
2023 STEPs: Self-Supervised Key Step Extraction and Localization from Unlabeled Procedural Videos
abstract
We address the problem of extracting key steps from un-labeled procedural videos, motivated by the potential of Augmented Reality (AR) headsets to revolutionize job training and performance. We decompose the problem into two steps: representation learning and key steps extraction. We propose a training objective, Bootstrapped Multi-Cue Contrastive (BMC2) loss to learn discriminative representations for various steps without any labels. Different from prior works, we develop techniques to train a light-weight temporal module which uses off-the-shelf features for self supervision. Our approach can seamlessly leverage information from multiple cues like optical flow, depth or gaze to learn discriminative features for key-steps, making it amenable for AR applications. We finally extract key steps via a tunable algorithm that clusters the representations and samples. We show significant improvements over prior works for the task of key step localization and phase classification. Qualitative results demonstrate that the extracted key steps are meaningful and succinctly represent various steps of the procedural tasks. Our code can be found at https://github.com/anshulbshah/STEPs.
Anshul Shah 0001, Ben Lundell, Harpreet Sawhney, Rama Chellappa
ICCV4
2023 SparseDet: Improving Sparsely Annotated Object Detection with Pseudo-positive Mining
abstract
Training with sparse annotations is known to reduce the performance of object detectors. Previous methods have focused on proxies for missing ground truth annotations in the form of pseudo-labels for unlabeled boxes. We observe that existing methods suffer at higher levels of sparsity in the data due to noisy pseudo-labels. To prevent this, we propose an end-to-end system that learns to separate the proposals into labeled and unlabeled regions using Pseudo-positive mining. While the labeled regions are processed as usual, self-supervised learning is used to process the unlabeled regions thereby preventing the negative effects of noisy pseudo-labels. This novel approach has multiple advantages such as improved robustness to higher sparsity when compared to existing methods. We conduct exhaustive experiments on five splits on the PASCAL-VOC and COCO datasets achieving state-of-the-art performance. We also unify various splits used across literature for this task and present a standardized benchmark. On average, we improve by 2.6, 3.9 and 9.6 mAP over previous state-of-the-art methods on three splits of increasing sparsity on COCO. Our project is publicly available at cs.umd.edu/~sakshams/SparseDet.
Saksham Suri, Sai Saketh Rambhatla, Rama Chellappa, Abhinav Shrivastava
ICCV3
2023 Synthetic-to-Real Domain Adaptation for Action Recognition: A Dataset and Baseline Performances
abstract
Human action recognition is a challenging problem, particularly when there is high variability in factors such as subject appearance, backgrounds and viewpoint. While deep neural networks (DNNs) have been shown to perform well on action recognition tasks, they typically require large amounts of high-quality labeled data to achieve robust performance across a variety of conditions. Synthetic data has shown promise as a way to avoid the substantial costs and potential ethical concerns associated with collecting and labeling enormous amounts of data in the real-world. However, synthetic data may differ from real data in important ways. This phenomenon, known as domain shift, can limit the utility of synthetic data in robotics applications. To mitigate the effects of domain shift, substantial effort is being dedicated to the development of domain adaptation (DA) techniques. Yet, much remains to be understood about how best to develop these techniques. In this paper, we introduce a new dataset called Robot Control Gestures (RoCoG-v2). The dataset is composed of both real and synthetic videos from seven gesture classes, and is intended to support the study of synthetic-to-real domain shift for video-based action recognition. Our work expands upon existing datasets by focusing the action classes on gestures for human-robot teaming, as well as by enabling investigation of domain shift in both ground and aerial views. We present baseline results using state-of-the-art action recognition and domain adaptation algorithms and offer initial insight on tackling the synthetic-to-real and ground-to-air domain shifts. Instructions on accessing the dataset can be found at https://github.com/reddyav1/RoCoG-v2.
Arun V. Reddy, Ketul Shah, William Paul, Rohita Mocharla, Judy Hoffman, Kapil D. Katyal, Dinesh Manocha, Celso de Melo, Rama Chellappa
ICRA9
2023 Certified Robustness via Dynamic Margin Maximization and Improved Lipschitz Regularization
abstract
To improve the robustness of deep classifiers against adversarial perturbations, many approaches have been proposed, such as designing new architectures with better robustness properties (e.g., Lipschitz-capped networks), or modifying the training process itself (e.g., min-max optimization, constrained learning, or regularization). These approaches, however, might not be effective at increasing the margin in the input (feature) space. In this paper, we propose a differentiable regularizer that is a lower bound on the distance of the data points to the classification boundary. The proposed regularizer requires knowledge of the model's Lipschitz constant along certain directions. To this end, we develop a scalable method for calculating guaranteed differentiable upper bounds on the Lipschitz constant of neural networks accurately and efficiently. The relative accuracy of the bounds prevents excessive regularization and allows for more direct manipulation of the decision boundary. Furthermore, our Lipschitz bounding algorithm exploits the monotonicity and Lipschitz continuity of the activation layers, and the resulting bounds can be used to design new layers with controllable bounds on their Lipschitz constant. Experiments on the MNIST, CIFAR-10, and Tiny-ImageNet data sets verify that our proposed algorithm obtains competitively improved results compared to the state-of-the-art.
Mahyar Fazlyab, Taha Entesari, Aniket Roy, Rama Chellappa
NeurIPS4
2023 Battle of the Backbones: A Large-Scale Comparison of Pretrained Models across Computer Vision Tasks
abstract
Neural network based computer vision systems are typically built on a backbone, a pretrained or randomly initialized feature extractor. Several years ago, the default option was an ImageNet-trained convolutional neural network. However, the recent past has seen the emergence of countless backbones pretrained using various algorithms and datasets. While this abundance of choice has led to performance increases for a range of systems, it is difficult for practitioners to make informed decisions about which backbone to choose. Battle of the Backbones (BoB) makes this choice easier by benchmarking a diverse suite of pretrained models, including vision-language models, those trained via self-supervised learning, and the Stable Diffusion backbone, across a diverse set of computer vision tasks ranging from classification to object detection to OOD generalization and more. Furthermore, BoB sheds light on promising directions for the research community to advance computer vision by illuminating strengths and weakness of existing approaches through a comprehensive analysis conducted on more than 1500 training runs. While vision transformers (ViTs) and self-supervised learning (SSL) are increasingly popular, we find that convolutional neural networks pretrained in a supervised fashion on large training sets still perform best on most tasks among the models we consider. Moreover, in apples-to-apples comparisons on the same architectures and similarly sized pretraining datasets, we find that SSL backbones are highly competitive, indicating that future works should perform SSL pretraining with advanced architectures and larger pretraining datasets. We release the raw results of our experiments along with code that allows researchers to put their own backbones through the gauntlet here: https://github.com/hsouri/Battle-of-the-Backbones.
Micah Goldblum, Hossein Souri, Renkun Ni, Manli Shu, Viraj Prabhu, Gowthami Somepalli, Prithvijit Chattopadhyay, Mark Ibrahim, Adrien Bardes, Judy Hoffman, Rama Chellappa, Andrew Gordon Wilson, Tom Goldstein
NeurIPS11
2023 Multi-View Action Recognition using Contrastive Learning
abstract
In this work, we present a method for RGB-based action recognition using multi-view videos. We present a supervised contrastive learning framework to learn a feature embedding robust to changes in viewpoint, by effectively leveraging multi-view data. We use an improved supervised contrastive loss and augment the positives with those coming from synchronized viewpoints. We also propose a new approach to use classifier probabilities to guide the selection of hard negatives in the contrastive loss, to learn a more discriminative representation. Negative samples from confusing classes based on posterior are weighted higher. We also show that our method leads to better domain generalization compared to the standard supervised training based on synthetic multi-view data. Extensive experiments on real (NTU-60, NTU-120, NUMA) and synthetic (RoCoG) data demonstrate the effectiveness of our approach.
Ketul Shah, Anshul Shah 0001, Chun Pong Lau 0001, Celso de Melo, Rama Chellappa
WACV5
2023 Interpolated Joint Space Adversarial Training for Robust and Generalizable Defenses
abstract
Adversarial training (AT) is considered to be one of the most reliable defenses against adversarial attacks. However, models trained with AT sacrifice standard accuracy and do not generalize well to unseen attacks. Recent works show generalization improvement with adversarial samples under unseen threat models such as on-manifold threat model or neural perceptual threat model. However, the former requires exact manifold information while the latter requires algorithm relaxation. Motivated by these considerations, we propose a novel threat model called Joint Space Threat Model (JSTM), which exploits the underlying manifold information with Normalizing Flow, ensuring that the exact manifold assumption holds. Under JSTM, we develop novel adversarial attacks and defenses. Specifically, we propose the Robust Mixup strategy in which we maximize the adversity of the interpolated images and gain robustness and prevent overfitting. Our experiments show that Interpolated Joint Space Adversarial Training (IJSAT) achieves good performance in standard accuracy, robustness, and generalization. IJSAT is also flexible and can be used as a data augmentation method to improve standard accuracy and combined with many existing AT approaches to improve robustness. We demonstrate the effectiveness of our approach on three benchmark datasets, CIFAR-10/100, OM-ImageNet and CIFAR-10-C.
Chun Pong Lau 0001, Jiang Liu 0014, Hossein Souri, Wei-An Lin, Soheil Feizi, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Max-Margin Contrastive Learning
abstract
Standard contrastive learning approaches usually require a large number of negatives for effective unsupervised learning and often exhibit slow convergence. We suspect this behavior is due to the suboptimal selection of negatives used for offering contrast to the positives. We counter this difficulty by taking inspiration from support vector machines (SVMs) to present max-margin contrastive learning (MMCL). Our approach selects negatives as the sparse support vectors obtained via a quadratic optimization problem, and contrastiveness is enforced by maximizing the decision margin. As SVM optimization can be computationally demanding, especially in an end-to-end setting, we present simplifications that alleviate the computational burden. We validate our approach on standard vision benchmark datasets, demonstrating better performance in unsupervised representation learning over state-of-the-art, while having better empirical convergence properties.
Anshul Shah 0001, Suvrit Sra, Rama Chellappa, Anoop Cherian
AAAI3
2022 EyePAD++: A Distillation-based approach for joint Eye Authentication and Presentation Attack Detection using Periocular Images
abstract
A practical eye authentication (EA) system targeted for edge devices needs to perform authentication and be robust to presentation attacks, all while remaining compute and latency efficient. However, existing eye-based frameworks a) perform authentication and Presentation Attack Detection (PAD) independently and b) involve significant pre-processing steps to extract the iris region. Here, we introduce a joint framework for EA and PAD using periocular images. While a deep Multitask Learning (MTL) network can perform both the tasks, MTL suffers from the forgetting effect since the training datasets for EA and PAD are disjoint. To overcome this, we propose Eye Authentication with PAD (EyePAD), a distillation-based method that trains a single network for EA and PAD while reducing the effect of forgetting. To further improve the EA performance, we introduce a novel approach called EyePAD++ that includes training an MTL network on both EA and PAD data, while distilling the ‘versatility’ of the EyePAD network through an additional distillation step. Our proposed methods outperform the SOTA in PAD and obtain near-SOTA performance in eye-to-eye verification, without any pre-processing. We also demonstrate the efficacy of EyePAD and EyePAD++ in user-to-user verification with PAD across network backbones and image quality.
Prithviraj Dhar, Amit Kumar 0013, Kirsten Kaplan, Khushi Gupta, Rama Chellappa
CVPR6
2022 Segment and Complete: Defending Object Detectors against Adversarial Patch Attacks with Robust Patch Detection
abstract
Object detection plays a key role in many security-critical systems. Adversarial patch attacks, which are easy to implement in the physical world, pose a serious threat to state-of-the-art object detectors. Developing reliable defenses for object detectors against patch attacks is critical but severely understudied. In this paper, we propose Segment and Complete defense (SAC), a general framework for defending object detectors against patch attacks through detection and removal of adversarial patches. We first train a patch segmenter that outputs patch masks which provide pixel-level localization of adversarial patches. We then propose a self adversarial training algorithm to robustify the patch segmenter. In addition, we design a robust shape completion algorithm, which is guaranteed to remove the entire patch from the images if the outputs of the patch segmenter are within a certain Hamming distance of the ground-truth patch masks. Our experiments on COCO and xView datasets demonstrate that SAC achieves superior robustness even under strong adaptive attacks with no reduction in performance on clean images, and generalizes well to unseen patch shapes, attack budgets, and unseen attack methods. Furthermore, we present the APRICOT-Mask dataset, which augments the APRICOT dataset with pixel-level annotations of adversarial patches. We show SAC can significantly reduce the targeted attack success rate of physical patch attacks. Our code is available at https://github.com/joellliu/SegmentAndComplete.
Jiang Liu 0014, Alexander Levine 0001, Chun Pong Lau 0001, Rama Chellappa, Soheil Feizi
CVPR4
2022 HyperSegNAS: Bridging One-Shot Neural Architecture Search with 3D Medical Image Segmentation using HyperNet
abstract
Semantic segmentation of 3D medical images is a challenging task due to the high variability of the shape and pattern of objects (such as organs or tumors). Given the recent success of deep learning in medical image segmentation, Neural Architecture Search (NAS) has been introduced to find high-performance 3D segmentation network architectures. However, because of the massive computational requirements of 3D data and the discrete optimization nature of architecture search, previous NAS methods require a long search time or necessary continuous relaxation, and commonly lead to sub-optimal network architectures. While one-shot NAS can potentially address these disadvantages, its application in the segmentation domain has not been well studied in the expansive multi-scale multi-path search space. To enable one-shot NAS for medical image segmentation, our method, named HyperSegNAS, introduces a HyperNet to assist super-net training by incorporating architecture topology information. Such a HyperNet can be removed once the super-net is trained and introduces no overhead during architecture search. We show that HyperSegNAS yields better performing and more intuitive architectures compared to the previous state-of-the-art (SOTA) segmentation networks; furthermore, it can quickly and accurately find good architecture candidates under different computing constraints. Our method is evaluated on public datasets from the Medical Segmentation Decathlon (MSD) challenge, and achieves SOTA performances.
Cheng Peng 0008, Andriy Myronenko, Ali Hatamizadeh, Vishwesh Nath, Md Mahfuzur Rahman Siddiquee, Yufan He, Daguang Xu, Rama Chellappa, Dong Yang 0005
CVPR8
2022 Where in the World Is This Image? Transformer-Based Geo-localization in the Wild
Shraman Pramanick, Ewa Magdalena Nowara, Joshua Gleason, Carlos Domingo Castillo, Rama Chellappa
ECCV (38)5
2022 An Empirical Analysis of Boosting Deep Networks
abstract
Boosting is a method for finding a highly accurate classifier by linearly combining many “weak” classifiers, each of which may be only moderately accurate. Thus, boosting is a method for learning an ensemble of classifiers. While boosting has been shown to be very effective for decision trees, its impact on neural networks has not been extensively studied. Using standard object recognition datasets, we verify experimentally the well-known result that a boosted ensemble of decision trees usually generalizes much better on testing data than a single decision tree with the same number of parameters. In contrast, using the same datasets and boosting algorithms, our experiments show the opposite to be true when using neural networks (both convolutional neural networks (CNNs) and multilayer perceptrons (MLPs)). We find that a single neural network usually generalizes better than a boosted ensemble of smaller neural networks with the same total number of parameters. While this is an experimental investigation, more theoretical research is warranted to understand the role of boosting in deep learning-based classifiers.
Sai Saketh Rambhatla, Michael J. Jones 0001, Rama Chellappa
IJCNN3
2022 Towards Performant and Reliable Undersampled MR Reconstruction via Diffusion Model Sampling
Cheng Peng 0008, Shaohua Kevin Zhou, Vishal M. Patel, Rama Chellappa
MICCAI (6)5
2022 FeLMi : Few shot Learning with hard Mixup
abstract
Learning from a few examples is a challenging computer vision task. Traditionally,meta-learning-based methods have shown promise towards solving this problem.Recent approaches show benefits by learning a feature extractor on the abundantbase examples and transferring these to the fewer novel examples. However, thefinetuning stage is often prone to overfitting due to the small size of the noveldataset. To this end, we propose Few shot Learning with hard Mixup (FeLMi)using manifold mixup to synthetically generate samples that helps in mitigatingthe data scarcity issue. Different from a naïve mixup, our approach selects the hardmixup samples using an uncertainty-based criteria. To the best of our knowledge,we are the first to use hard-mixup for the few-shot learning problem. Our approachallows better use of the pseudo-labeled base examples through base-novel mixupand entropy-based filtering. We evaluate our approach on several common few-shotbenchmarks - FC-100, CIFAR-FS, miniImageNet and tieredImageNet and obtainimprovements in both 1-shot and 5-shot settings. Additionally, we experimented onthe cross-domain few-shot setting (miniImageNet → CUB) and obtain significantimprovements.
Aniket Roy, Anshul Shah 0001, Ketul Shah, Prithviraj Dhar, Anoop Cherian, Rama Chellappa
NeurIPS6
2022 Sleeper Agent: Scalable Hidden Trigger Backdoors for Neural Networks Trained from Scratch
abstract
As the curation of data for machine learning becomes increasingly automated, dataset tampering is a mounting threat. Backdoor attackers tamper with training data to embed a vulnerability in models that are trained on that data. This vulnerability is then activated at inference time by placing a "trigger'' into the model's input. Typical backdoor attacks insert the trigger directly into the training data, although the presence of such an attack may be visible upon inspection. In contrast, the Hidden Trigger Backdoor Attack achieves poisoning without placing a trigger into the training data at all. However, this hidden trigger attack is ineffective at poisoning neural networks trained from scratch. We develop a new hidden trigger attack, Sleeper Agent, which employs gradient matching, data selection, and target model re-training during the crafting process. Sleeper Agent is the first hidden trigger backdoor attack to be effective against neural networks trained from scratch. We demonstrate its effectiveness on ImageNet and in black-box settings. Our implementation code can be found at: https://github.com/hsouri/Sleeper-Agent.
Hossein Souri, Liam Fowl, Rama Chellappa, Micah Goldblum, Tom Goldstein
NeurIPS3
2022 Pose and Joint-Aware Action Recognition
abstract
Recent progress on action recognition has mainly focused on RGB and optical flow features. In this paper, we approach the problem of joint-based action recognition. Unlike other modalities, constellation of joints and their motion generate models with succinct human motion information for activity recognition. We present a new model for joint-based action recognition, which first extracts motion features from each joint separately through a shared motion encoder before performing collective reasoning. Our joint selector module re-weights the joint information to select the most discriminative joints for the task. We also propose a novel joint-contrastive loss that pulls together groups of joint features which convey the same action. We strengthen the joint-based representations by using a geometry-aware data augmentation technique which jitters pose heatmaps while retaining the dynamics of the action. We show large improvements over the current state-of-the-art joint-based approaches on JHMDB, HMDB, Charades, AVA action recognition datasets. A late fusion with RGB and Flow-based approaches yields additional improvements. Our model also outperforms the existing baseline on Mimetics, a dataset with out-of-context actions.
Anshul Shah 0001, Shlok Kumar Mishra, Ankan Bansal, Jun-Cheng Chen, Rama Chellappa, Abhinav Shrivastava
WACV5
2022 Mutual Adversarial Training: Learning Together is Better Than Going Alone
abstract
Recent studies have shown that robustness to adversarial attacks can be transferred across deep neural networks. In other words, we can make a weak model more robust with the help of a strong teacher model. In this paper, we ask if models can “learn together” and “teach each other” to achieve better robustness instead of learning from a static teacher.We study how interactions among models affect robustness via knowledge distillation. We propose mutual adversarial training (MAT), in which multiple models are trained together and share the knowledge of adversarial examples to achieve improved robustness. MAT allows robust models to explore a larger space of adversarial samples and find more robust feature spaces and decision boundaries. Through extensive experiments on the CIFAR-10, CIFAR-100, and mini-ImageNet datasets, we demonstrate that MAT can effectively improve model robustness and outperform state-of-the-art methods under white-box attacks. In addition, we show that MAT can also mitigate the robustness trade-off among different perturbation types. Specially, we train specialist models that learn to defend a specific perturbation type and a generalist model that learns to defend multiple perturbation types by learning from the specialists, which brings as much as 13.4% accuracy gain to AT baselines against the union ofl∞,l2, andl1 attacks. Our results show the superiority of the proposed method and demonstrate that collaborative learning is an effective strategy for designing robust models.
Jiang Liu 0014, Chun Pong Lau 0001, Hossein Souri, Soheil Feizi, Rama Chellappa
IEEE Trans. Inf. Forensics Secur.5
2021 XraySyn: Realistic View Synthesis From a Single Radiograph Through CT Priors
abstract
A radiograph visualizes the internal anatomy of a patient through the use of X-ray, which projects 3D information onto a 2D plane. Hence, radiograph analysis naturally requires physicians to relate their prior knowledge about 3D human anatomy to 2D radiographs. Synthesizing novel radiographic views in a small range can assist physicians in interpreting anatomy more reliably; however, radiograph view synthesis is heavily ill-posed, lacking in paired data, and lacking in differentiable operations to leverage learning-based approaches. To address these problems, we use Computed Tomography (CT) for radiograph simulation and design a differentiable projection algorithm, which enables us to achieve geometrically consistent transformations between the radiography and CT domains. Our method, XraySyn, can synthesize novel views on real radiographs through a combination of realistic simulation and finetuning on real radiographs. To the best of our knowledge, this is the first work on radiograph view synthesis. We show that by gaining an understanding of radiography in 3D space, our method can be applied to radiograph bone extraction and suppression without requiring groundtruth bone labels.
Cheng Peng 0008, Haofu Liao, Gina Wong, Jiebo Luo 0001, Shaohua Kevin Zhou, Rama Chellappa
AAAI6
2021 Hierarchical Video Prediction Using Relational Layouts for Human-Object Interactions
abstract
Learning to model and predict how humans interact with objects while performing an action is challenging, and most of the existing video prediction models are ineffective in modeling complicated human-object interactions. Our work builds on hierarchical video prediction models, which disentangle the video generation process into two stages: predicting a high-level representation, such as pose sequence, and then learning a pose-to-pixels translation model for pixel generation. An action sequence for a human-object interaction task is typically very complicated, involving the evolution of pose, person’s appearance, object locations, and object appearances over time. To this end, we propose a Hierarchical Video Prediction model using Relational Layouts. In the first stage, we learn to predict a sequence of layouts. A layout is a high-level representation of the video containing both pose and objects’ information for every frame. The layout sequence is learned by modeling the relationships between the pose and objects using relational reasoning and recurrent neural networks. The layout sequence acts as a strong structure prior to the second stage that learns to map the layouts into pixel space. Experimental evaluation of our method on two datasets, UMD-HOI and Bimanual, shows significant improvements in standard video evaluation metrics such as LPIPS, PSNR, and SSIM. We also perform a detailed qualitative analysis of our model to demonstrate various generalizations.
Navaneeth Bodla, Gaurav Shrivastava, Rama Chellappa, Abhinav Shrivastava
CVPR3
2021 A Synthesis-Based Approach for Thermal-to-Visible Face Verification
abstract
In recent years, visible-spectrum face verification systems have been shown to match the performance of experienced forensic examiners. However, such systems are ineffective in low-light and nighttime conditions. Thermal face imagery, which captures body heat emissions, effectively augments the visible spectrum, capturing discriminative facial features in scenes with limited illumination. Due to the increased cost and difficulty of obtaining diverse, paired thermal and visible spectrum datasets, not many algorithms and large-scale benchmarks for low-light recognition are available. This paper presents an algorithm that achieves state-of-the-art performance on both the ARL-VTF and TUFTS multi-spectral face datasets. Importantly, we study the impact of face alignment, pixel-level correspondence, and identity classification with label smoothing for multi-spectral face synthesis and verification. We show that our proposed method is widely applicable, robust, and highly effective. In addition, we show that the proposed method significantly outperforms face frontalization methods on profile-to-frontal verification. Finally, we present MILAB-VTF(B), a challenging multi-spectral face dataset that is composed of paired thermal and visible videos. To the best of our knowledge, with face data from 400 subjects, this dataset represents the most extensive collection of publicly available indoor and long-range outdoor thermal-visible face imagery. Lastly, we show that our end-to-end thermal-to-visible face verification system provides strong performance on the MILAB-VTF(B) dataset.
Neehar Peri, Joshua Gleason, Carlos Domingo Castillo, Thirimachos Bourlai, Vishal M. Patel, Rama Chellappa
FG6
2021 The 5th Recognizing Families in the Wild Data Challenge: Predicting Kinship from Faces
abstract
Recognizing Families In the Wild (RFIW), held as a data challenge in conjunction with the 16thIEEE International Conference on Automatic Face and Gesture Recognition (FG), is a large-scale, multi-track visual kinship recognition evaluation. For the fifth edition of RFIW, we continue to attract scholars, bring together professionals, publish new work, and discuss prospects. In this paper, we summarize submissions for the three tasks of this year's RFIW: specifically, we review the results for kinship verification, tri-subject verification, and family member search and retrieval. We look at the RFIW problem, share current efforts, and make recommendations for promising future directions.
Joseph P. Robinson, Can Qin, Ming Shao, Matthew Turk 0001, Rama Chellappa, Yun Fu 0001
FG5
2021 PASS: Protected Attribute Suppression System for Mitigating Bias in Face Recognition
abstract
Face recognition networks encode information about sensitive attributes while being trained for identity classification. Such encoding has two major issues: (a) it makes the face representations susceptible to privacy leakage (b) it appears to contribute to bias in face recognition. However, existing bias mitigation approaches generally require end-to-end training and are unable to achieve high verification accuracy. Therefore, we present a descriptor-based adversarial de-biasing approach called ‘Protected Attribute Suppression System (PASS)’. PASS can be trained on top of descriptors obtained from any previously trained high-performing network to classify identities and simultaneously reduce encoding of sensitive attributes. This eliminates the need for end-to-end training. As a component of PASS, we present a novel discriminator training strategy that discourages a network from encoding protected attribute information. We show the efficacy of PASS to reduce gender and skintone information in descriptors from SOTA face recognition networks like Arcface. As a result, PASS descriptors outperform existing baselines in reducing gender and skintone bias on the IJB-C dataset, while maintaining a high verification accuracy.
Prithviraj Dhar, Joshua Gleason, Aniket Roy, Carlos Domingo Castillo, Rama Chellappa
ICCV5
2021 The Pursuit of Knowledge: Discovering and Localizing Novel Categories using Dual Memory
abstract
We tackle object category discovery, which is the problem of discovering and localizing novel objects in a large unlabeled dataset. While existing methods show results on datasets with less cluttered scenes and fewer object in-stances per image, we present our results on the challenging COCO dataset. Moreover, we argue that, rather than discovering new categories from scratch, discovery algorithms can benefit from identifying what is already known and focusing their attention on the unknown. We propose a method that exploits prior knowledge about certain object types to discover new categories by leveraging two memory modules, namely Working and Semantic memory. We show the performance of our detector on the COCO minival dataset to demonstrate its in-the-wild capabilities.
Sai Saketh Rambhatla, Rama Chellappa, Abhinav Shrivastava
ICCV2
2021 DA-VSR: Domain Adaptable Volumetric Super-Resolution for Medical Images
Cheng Peng 0008, Shaohua Kevin Zhou, Rama Chellappa
MICCAI (6)3
2021 A Multi-Class Hinge Loss for Conditional GANs
abstract
We propose a new algorithm to incorporate class conditional information into the critic of GANs via a multi-class generalization of the commonly used Hinge loss that is compatible with both supervised and semi-supervised settings. We study the compromise between training a state of the art generator and an accurate classifier simultaneously, and propose a way to use our algorithm to measure the degree to which a generator and critic are class conditional. We show the trade-off between a generator-critic pair respecting class conditioning inputs and generating the highest quality images. With our multi-hinge loss modification we are able to improve Inception Scores and Frechet Inception Distance on the Imagenet dataset.
Ilya Kavalerov, Wojciech Czaja, Rama Chellappa
WACV3
2021 Guest Editorial: Adversarial Deep Learning in Biometrics & Forensics
Rama Chellappa, Diego Gragnaniello, Chang-Tsun Li, Francesco Marra, Richa Singh 0001
Comput. Vis. Image Underst.1
2021 Guest Editorial: Special Issue on Deep Learning for Video Analysis and Compression
Dong Xu 0001, Rama Chellappa, Luc Van Gool, Guo Lu
Int. J. Comput. Vis.2
2021 Deep Regionlets: Blended Representation and Deep Learning for Generic Object Detection
abstract
In this article, we propose a novel object detection algorithm named "Deep Regionlets" by integrating deep neural networks and a conventional detection schema for accurate generic object detection. Motivated by the effectiveness of regionlets for modeling object deformations and multiple aspect ratios, we incorporate regionlets into an end-to-end trainable deep learning framework. The deep regionlets framework consists of a region selection network and a deep regionlet learning module. Specifically, given a detection bounding box proposal, the region selection network provides guidance on where to select sub-regions from which features can be learned from. An object proposal typically contains three - 16 sub-regions. The regionlet learning module focuses on local feature selection and transformations to alleviate the effects of appearance variations. To this end, we first realize non-rectangular region selection within the detection framework to accommodate variations in object appearance. Moreover, we design a "gating network" within the regionlet leaning module to enable instance dependent soft feature selection and pooling. The Deep Regionlets framework is trained end-to-end without additional efforts. We present ablation studies and extensive experiments on the PASCAL VOC dataset and the Microsoft COCO dataset. The proposed method yields competitive performance over state-of-the-art algorithms, such as RetinaNet and Mask R-CNN, even without additional segmentation labels.
Hongyu Xu, Xutao Lv, Xiaoyu Wang 0002, Zhou Ren, Navaneeth Bodla, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.6
2021 Advances in Machine Learning and Deep Neural Networks
abstract
We are currently experiencing the dawn of what is known as the fourth industrial revolution. At the center of this historical happening, as one of the key enabling technologies, lies a discipline that deals with data and whose goal is to extract information and related knowledge that is hidden in it, in order to make predictions and, subsequently, take decisions. Machine learning (ML) is the name that is used as an umbrella to cover a wide range of theories, methods, algorithms, and architectures that are used to this end. The articles in this special issue cover promising developments in the related areas of machine learning and deep neural networks and offers possible paths for the future.
Rama Chellappa, Sergios Theodoridis, André van Schaik
Proc. IEEE1
2021 3-D Fourier Scattering Transform and Classification of Hyperspectral Images
abstract
Recent developments in machine learning and signal processing have resulted in many new techniques that are able to effectively capture the intrinsic yet complex properties of hyperspectral imagery (HSI). Tasks ranging from anomaly detection to classification can now be solved by taking advantage of very efficient algorithms which have their roots in representation theory and computational approximation. Time–frequency methods are one example of such techniques. They provide means to analyze and extract the spectral content from data. On the other hand, hierarchical methods such as neural networks (NNs) incorporate spatial information across scales and model multiple levels of dependencies between spectral features. Both of these approaches have recently been proven to provide significant advances in the spectral-spatial classification of HSI. The 3-D Fourier scattering transform, which is introduced in this article, is an amalgamation of time–frequency representations with NN architectures. It leverages the benefits provided by the short-time Fourier transform with the numerical efficiency of deep learning network structures. We test the proposed method on several standard hyperspectral data sets, and we present results that indicate that the 3-D Fourier scattering transform is highly effective at representing spectral content when compared with other state-of-the-art spectral-spatial classification methods.
Ilya Kavalerov, Wojciech Czaja, Rama Chellappa
IEEE Trans. Geosci. Remote. Sens.4
2020 Detecting Human-Object Interactions via Functional Generalization
abstract
We present an approach for detecting human-object interactions (HOIs) in images, based on the idea that humans interact with functionally similar objects in a similar manner. The proposed model is simple and efficiently uses the data, visual features of the human, relative spatial orientation of the human and the object, and the knowledge that functionally similar objects take part in similar interactions with humans. We provide extensive experimental validation for our approach and demonstrate state-of-the-art results for HOI detection. On the HICO-Det dataset our method achieves a gain of over 2.5% absolute points in mean average precision (mAP) over state-of-the-art. We also show that our approach leads to significant performance gains for zero-shot HOI detection in the seen object setting. We further demonstrate that using a generic object detector, our model can generalize to interactions involving previously unseen objects.
Ankan Bansal, Sai Saketh Rambhatla, Abhinav Shrivastava, Rama Chellappa
AAAI4
2020 3DRegNet: A Deep Neural Network for 3D Point Registration
abstract
We present 3DRegNet, a novel deep learning architecture for the registration of 3D scans. Given a set of 3D point correspondences, we build a deep neural network to address the following two challenges: (i) classification of the point correspondences into inliers/outliers, and (ii) regression of the motion parameters that align the scans into a common reference frame. With regard to regression, we present two alternative approaches: (i) a Deep Neural Network (DNN) registration and (ii) a Procrustes approach using SVD to estimate the transformation. Our correspondence-based approach achieves a higher speedup compared to competing baselines. We further propose the use of a refinement network, which consists of a smaller 3DRegNet as a refinement to improve the accuracy of the registration. Extensive experiments on two challenging datasets demonstrate that we outperform other methods and achieve state-of-the-art results.
Gonçalo Dias Pais, Srikumar Ramalingam, Venu Madhav Govindu, Jacinto C. Nascimento, Rama Chellappa, Pedro Miraldo
CVPR5
2020 SAINT: Spatially Aware Interpolation NeTwork for Medical Slice Synthesis
abstract
Deep learning-based single image super-resolution (SISR) methods face various challenges when applied to 3D medical volumetric data (i.e., CT and MR images) due to the high memory cost and anisotropic resolution, which adversely affect their performance. Furthermore, mainstream SISR methods are designed to work over specific upsampling factors, which makes them ineffective in clinical practice. In this paper, we introduce a Spatially Aware Interpolation NeTwork (SAINT) for medical slice synthesis to alleviate the memory constraint that volumetric data poses. Compared to other super-resolution methods, SAINT utilizes voxel spacing information to provide desirable levels of details, and allows for the upsampling factor to be determined on the fly. Our evaluations based on 853 CT scans from four datasets that contain liver, colon, hepatic vessels, and kidneys show that SAINT consistently outperforms other SISR methods in terms of medical slice synthesis quality, while using only a single model to deal with different upsampling factors.
Cheng Peng 0008, Wei-An Lin, Haofu Liao, Rama Chellappa, Shaohua Kevin Zhou
CVPR4
2020 Visual Question Answering on Image Sets
Ankan Bansal, Rama Chellappa
ECCV (21)3
2020 The Devil Is in the Details: Self-supervised Attention for Vehicle Re-identification
Pirazh Khorramshahi, Neehar Peri, Jun-Cheng Chen, Rama Chellappa
ECCV (14)4
2020 ATFaceGAN: Single Face Image Restoration and Recognition from Atmospheric Turbulence
abstract
Image degradation due to atmospheric turbulence is common while capturing images at long ranges. To mitigate the degradation due to turbulence which includes deformation and blur, we propose a generative single frame restoration algorithm which disentangles the blur and deformation due to turbulence and reconstructs a restored image. The disentanglement is achieved by decomposing the distortion due to turbulence into blur and deformation components using deblur generator and deformation correction generator respectively. Two paths of restoration are implemented to regularize the disentanglement and generate two restored images from one degraded image. A fusion function combines the features of the restored images to reconstruct a sharp image with rich details. Adversarial and perceptual losses are added to reconstruct a sharp image and suppress the artifacts respectively. Extensive experiments demonstrate the effectiveness of the proposed restoration algorithm, which achieves satisfactory performance in face restoration and face recognition.
Chun Pong Lau 0001, Hossein Souri, Rama Chellappa
FG3
2020 How are attributes expressed in face DCNNs?
abstract
As deep networks become increasingly accurate at recognizing faces, it is vital to understand how these networks process faces. While these networks are solely trained to recognize identities, they also contain face related information such as sex, age, and pose of the face even when the networks are not trained to learn these attributes. We introduce expressivity as a measure of how much a feature vector informs us about an attribute, where a feature vector can be from internal or final layers of a network. Expressivity is computed by a second neural network whose inputs are features and attributes. The output of the second neural network approximates the mutual information between feature vectors and an attribute. We investigate the expressivity for two different deep convolutional neural network (DCNN) architectures: a Resnet-101 and an Inception Resnet v2. In the final fully connected layer of the networks, we found the order of expressivity for facial attributes to be Age > Sex > Yaw. Additionally, we studied the changes in the encoding of facial attributes over training iterations. We found that as training progresses, expressivities of yaw, sex, and age decrease. Our technique can be a tool for investigating the sources of bias in a network and a step towards explaining the network's identity decisions.
Prithviraj Dhar, Ankan Bansal, Carlos Domingo Castillo, Joshua Gleason, P. Jonathon Phillips, Rama Chellappa
FG6
2020 Recognizing Families In the Wild (RFIW): The 4th Edition
abstract
Recognizing Families In the Wild (RFIW)- an annual large-scale, multi-track automatic kinship recognition evaluation- supports various visual kin-based problems on scales much higher than ever before. Organized in conjunction with the as a Challenge, RFIW provides a platform for publishing original work and the gathering of experts for a discussion of the next steps. This paper summarizes the supported tasks (i.e., kinship verification, tri-subject verification, and search & retrieval of missing children) in the evaluation protocols, which include the practical motivation, technical background, data splits, metrics, and benchmark results. Furthermore, top submissions (i.e., leader-board stats) are listed and reviewed as a high-level analysis on the state of the problem. In the end, the purpose of this paper is to describe the 2020 RFIW challenge, end-to-end, along with forecasts in promising future directions.
Joseph P. Robinson, Yu Yin 0001, Zaid Khan 0001, Ming Shao, Si-Yu Xia, Michael Stopa, Samson Timoner, Matthew Turk 0001, Rama Chellappa, Yun Fu 0001
FG9
2020 Occlusion-Adaptive Deep Network for Robust Facial Expression Recognition
abstract
Recognizing the expressions of partially occluded faces is a challenging computer vision problem. Previous expression recognition methods, either overlooked this issue or resolved it using unrealistic assumptions. Motivated by the fact that the human visual system is adept at ignoring the occlusions and focus on non-occluded facial areas, we propose a landmark-guided attention branch to find and discard corrupted features from occluded regions so that they are not used for recognition. An attention map is first generated to indicate if a specific facial part is occluded and guide our model to attend to non-occluded regions. To further improve robustness, we propose a facial region branch to partition the feature maps into non-overlapping facial blocks and task each block to predict the expression independently. This results in more diverse and discriminative features, enabling the expression recognition system to re-cover even though the face is partially occluded. Depending on the synergistic effects of the two branches, our occlusion-adaptive deep network significantly outperforms state-of-the-art methods on two challenging in-the-wild benchmark datasets and three real-world occluded expression datasets.
Hui Ding 0002, Peng Zhou 0003, Rama Chellappa
IJCB3
2020 Robust Optimal Transport with Applications in Generative Modeling and Domain Adaptation
abstract
Optimal Transport (OT) distances such as Wasserstein have been used in several areas such as GANs and domain adaptation. OT, however, is very sensitive to outliers (samples with large noise) in the data since in its objective function, every sample, including outliers, is weighed similarly due to the marginal constraints. To remedy this issue, robust formulations of OT with unbalanced marginal constraints have previously been proposed. However, employing these methods in deep learning problems such as GANs and domain adaptation is challenging due to the instability of their dual optimization solvers. In this paper, we resolve these issues by deriving a computationally-efficient dual form of the robust OT optimization that is amenable to modern deep learning applications. We demonstrate the effectiveness of our formulation in two applications of GANs and domain adaptation. Our approach can train state-of-the-art GAN models on noisy datasets corrupted with outlier distributions. In particular, the proposed optimization method computes weights for training samples reflecting how difficult it is for those samples to be generated in the model. In domain adaptation, our robust OT formulation leads to improved accuracy compared to the standard adversarial adaptation methods. Our code is available at https://github.com/yogeshbalaji/robustOT.
Yogesh Balaji, Rama Chellappa, Soheil Feizi
NeurIPS2
2020 Dual Manifold Adversarial Robustness: Defense against Lp and non-Lp Adversarial Attacks
abstract
Adversarial training is a popular defense strategy against attack threat models with bounded Lp norms. However, it often degrades the model performance on normal images and more importantly, the defense does not generalize well to novel attacks. Given the success of deep generative models such as GANs and VAEs in characterizing the underlying manifold of images, we investigate whether or not the aforementioned deficiencies of adversarial training can be remedied by exploiting the underlying manifold information. To partially answer this question, we consider the scenario when the manifold information of the underlying data is available. We use a subset of ImageNet natural images where an approximate underlying manifold is learned using StyleGAN. We also construct an ``On-Manifold ImageNet'' (OM-ImageNet) dataset by projecting the ImageNet samples onto the learned manifold. For OM-ImageNet, the underlying manifold information is exact. Using OM-ImageNet, we first show that on-manifold adversarial training improves both standard accuracy and robustness to on-manifold attacks. However, since no out-of-manifold perturbations are realized, the defense can be broken by Lp adversarial attacks. We further propose Dual Manifold Adversarial Training (DMAT) where adversarial perturbations in both latent and image spaces are used in robustifying the model. Our DMAT improves performance on normal images, and achieves comparable robustness to the standard adversarial training against Lp attacks. In addition, we observe that models defended by DMAT achieve improved robustness against novel attacks which manipulate images by global color shifts or various types of image filtering. Interestingly, similar improvements are also achieved when the defended models are tested on (out-of-manifold) natural images. These results demonstrate the potential benefits of using manifold information in enhancing robustness of deep learning models against various types of novel adversarial attacks.
Wei-An Lin, Chun Pong Lau 0001, Alexander Levine 0001, Rama Chellappa, Soheil Feizi
NeurIPS4
2020 Impact of injection attacks on sensor-based continuous authentication for smartphones
abstract
Given the relevance of smartphones for accessing personalized services in smart cities, Continuous Authentication (CA) mechanisms are attracting attention to avoid impersonation attacks. Some of them leverage Data Stream Mining (DSM) techniques applied over sensorial information. Injection attacks can undermine the effectiveness of DSM-based CA by fabricating artificial sensorial readings.The goal of this paper is to study the impact of injection attacks in terms of accuracy and immediacy to illustrate the time the adversary remains unnoticed. Two well-known DSM techniques (K-Nearest Neighbours and Hoeffding Adaptive Trees) and three data sources (location, gyroscope and accelerometer) are considered due to their widespread usage Results show that even if the attacker does not previously know anything about the victim, a significant attack surface arises – 1.35 min are needed, in the best case, to detect the attack on gyroscope and accelerometer and 7.27 min on location data. Moreover, we show that the type of sensor at stake and configuration settings may have a dramatic effect on countering this threat.
Lorena González-Manzano, Upal Mahbub, José María de Fuentes, Rama Chellappa
Comput. Commun.4
2020 Segment-Based Methods for Facial Attribute Detection from Partial Faces
abstract
State-of-the-art methods of attribute detection from faces almost always assume the presence of a full, unoccluded face. Hence, their performance degrades for partially visible and occluded faces. In this paper, we introduce SPLITFACE, a deep convolutional neural network-based method that is explicitly designed to perform attribute detection in partially occluded faces. Taking several facial segments and the full face as input, the proposed method takes a data driven approach to determine which attributes are localized in which facial segments. The unique architecture of the network allows each attribute to be predicted by multiple segments, which permits the implementation of committee machine techniques for combining local and global decisions to boost performance. With access to segment-based predictions, SPLITFACE can predict well those attributes which are localized in the visible parts of the face, without having to rely on the presence of the whole face. We use the CelebA and LFWA facial attribute datasets for standard evaluations. We also modify both datasets, to occlude the faces, so that we can evaluate the performance of attribute detection algorithms on partial faces. Our evaluation shows that SPLITFACE significantly outperforms other recent methods especially for partial faces.
Upal Mahbub, Sayantan Sarkar, Rama Chellappa
IEEE Trans. Affect. Comput.3
2019 Soft Sampling for Robust Object Detection
Zhe Wu 0001, Navaneeth Bodla, Mahyar Najibi, Rama Chellappa, Larry Davis 0001
BMVC5
2019 Learning Without Memorizing
abstract
Incremental learning (IL) is an important task aimed at increasing the capability of a trained model, in terms of the number of classes recognizable by the model. The key problem in this task is the requirement of storing data (e.g. images) associated with existing classes, while teaching the classifier to learn new classes. However, this is impractical as it increases the memory requirement at every incremental step, which makes it impossible to implement IL algorithms on edge devices with limited memory. Hence, we propose a novel approach, called `Learning without Memorizing (LwM)', to preserve the information about existing (base) classes, without storing any of their data, while making the classifier progressively learn the new classes. In LwM, we present an information preserving penalty: Attention Distillation Loss (L_{AD}), and demonstrate that penalizing the changes in classifiers' attention maps helps to retain information of the base classes, as new classes are added. We show that adding L_{AD} to the distillation loss which is an existing information preserving loss consistently outperforms the state-of-the-art performance in the iILSVRC-small and iCIFAR-100 datasets in terms of the overall accuracy of base and incrementally learned classes.
Prithviraj Dhar, Rajat Vikram Singh, Kuan-Chuan Peng, Ziyan Wu 0001, Rama Chellappa
CVPR5
2019 DuDoNet: Dual Domain Network for CT Metal Artifact Reduction
abstract
Computed tomography (CT) is an imaging modality widely used for medical diagnosis and treatment. CT images are often corrupted by undesirable artifacts when metallic implants are carried by patients, which creates the problem of metal artifact reduction (MAR). Existing methods for reducing the artifacts due to metallic implants are inadequate for two main reasons. First, metal artifacts are structured and non-local so that simple image domain enhancement approaches would not suffice. Second, the MAR approaches which attempt to reduce metal artifacts in the X-ray projection (sinogram) domain inevitably lead to severe secondary artifact due to sinogram inconsistency. To overcome these difficulties, we propose an end-to-end trainable Dual Domain Network (DuDoNet) to simultaneously restore sinogram consistency and enhance CT images. The linkage between the sigogram and image domains is a novel Radon inversion layer that allows the gradients to back-propagate from the image domain to the sinogram domain during training. Extensive experiments show that our method achieves significant improvements over other single domain MAR approaches. To the best of our knowledge, it is the first end-to-end dual-domain network for MAR.
Wei-An Lin, Haofu Liao, Cheng Peng 0008, Xiaohang Sun, Jingdan Zhang, Jiebo Luo 0001, Rama Chellappa, Shaohua Kevin Zhou
CVPR7
2019 Unsupervised Domain-Specific Deblurring via Disentangled Representations
abstract
Image deblurring aims to restore the latent sharp images from the corresponding blurred ones. In this paper, we present an unsupervised method for domain-specific, single-image deblurring based on disentangled representations. The disentanglement is achieved by splitting the content and blur features in a blurred image using content encoders and blur encoders. We enforce a KL divergence loss to regularize the distribution range of extracted blur attributes such that little content information is contained. Meanwhile, to handle the unpaired training data, a blurring branch and the cycle-consistency loss are added to guarantee that the content structures of the deblurred results match the original images. We also add an adversarial loss on deblurred results to generate visually realistic images and a perceptual loss to further mitigate the artifacts. We perform extensive experiments on the tasks of face and text deblurring using both synthetic datasets and real images, and achieve improved results compared to recent state-of-the-art deblurring methods.
Boyu Lu, Jun-Cheng Chen, Rama Chellappa
CVPR3
2019 Normalized Wasserstein for Mixture Distributions With Applications in Adversarial Learning and Domain Adaptation
abstract
Understanding proper distance measures between distributions is at the core of several learning tasks such as generative models, domain adaptation, clustering, etc. In this work, we focus on mixture distributions that arise naturally in several application domains where the data contains different sub-populations. For mixture distributions, established distance measures such as the Wasserstein distance do not take into account imbalanced mixture proportions. Thus, even if two mixture distributions have identical mixture components but different mixture proportions, the Wasserstein distance between them will be large. This often leads to undesired results in distance-based learning methods for mixture distributions. In this paper, we resolve this issue by introducing the Normalized Wasserstein measure. The key idea is to introduce mixture proportions as optimization variables, effectively normalizing mixture proportions in the Wasserstein formulation. Using the proposed normalized Wasserstein measure leads to significant performance gains for mixture distributions with imbalanced mixture proportions compared to the vanilla Wasserstein distance. We demonstrate the effectiveness of the proposed measure in GANs, domain adaptation and adversarial clustering in several benchmark datasets.
Yogesh Balaji, Rama Chellappa, Soheil Feizi
ICCV2
2019 A Dual-Path Model With Adaptive Attention for Vehicle Re-Identification
abstract
In recent years, attention models have been extensively used for person and vehicle re-identification. Most re-identification methods are designed to focus attention on key-point locations. However, depending on the orientation, the contribution of each key-point varies. In this paper, we present a novel dual-path adaptive attention model for vehicle re-identification (AAVER). The global appearance path captures macroscopic vehicle features while the orientation conditioned part appearance path learns to capture localized discriminative features by focusing attention on the most informative key-points. Through extensive experimentation, we show that the proposed AAVER method is able to accurately re-identify vehicles in unconstrained scenarios, yielding state of the art results on the challenging dataset VeRi-776. As a byproduct, the proposed system is also able to accurately predict vehicle key-points and shows an improvement of more than 7% over state of the art. The code for key-point estimation model is available at https://github.com/Pirazh/Vehicle_Key_ Point_Orientation_Estimation.
Pirazh Khorramshahi, Amit Kumar 0013, Neehar Peri, Sai Saketh Rambhatla, Jun-Cheng Chen, Rama Chellappa
ICCV6
2019 Uncertainty Modeling of Contextual-Connections Between Tracklets for Unconstrained Video-Based Face Recognition
abstract
Unconstrained video-based face recognition is a challenging problem due to significant within-video variations caused by pose, occlusion and blur. To tackle this problem, an effective idea is to propagate the identity from high-quality faces to low-quality ones through contextual connections, which are constructed based on context such as body appearance. However, previous methods have often propagated erroneous information due to lack of uncertainty modeling of the noisy contextual connections. In this paper, we propose the Uncertainty-Gated Graph (UGG), which conducts graph-based identity propagation between tracklets, which are represented by nodes in a graph. UGG explicitly models the uncertainty of the contextual connections by adaptively updating the weights of the edge gates according to the identity distributions of the nodes during inference. UGG is a generic graphical model that can be applied at only inference time or with end-to-end training. We demonstrate the effectiveness of UGG with state-of-the-art results in the recently released challenging Cast Search in Movies and IARPA Janus Surveillance Video Benchmark dataset.
Jingxiao Zheng, Ruichi Yu, Jun-Cheng Chen, Boyu Lu, Carlos Domingo Castillo, Rama Chellappa
ICCV6
2019 Entropic GANs meet VAEs: A Statistical Approach to Compute Sample Likelihoods in GANs
abstract
Building on the success of deep learning, two modern approaches to learn a probability model from the data are Generative Adversarial Networks (GANs) and Variational AutoEncoders (VAEs). VAEs consider an explicit probability model for the data and compute a generative distribution by maximizing a variational lower-bound on the log-likelihood function. GANs, however, compute a generative model by minimizing a distance between observed and generated probability distributions without considering an explicit model for the observed data. The lack of having explicit probability models in GANs prohibits computation of sample likelihoods in their frameworks and limits their use in statistical inference problems. In this work, we resolve this issue by constructing an explicit probability model that can be used to compute sample likelihood statistics in GANs. In particular, we prove that under this probability model, a family of Wasserstein GANs with an entropy regularization can be viewed as a generative model that maximizes a variational lower-bound on average sample log likelihoods, an approach that VAEs are based on. This result makes a principled connection between two modern generative models, namely GANs and VAEs. In addition to the aforementioned theoretical results, we compute likelihood statistics for GANs trained on Gaussian, MNIST, SVHN, CIFAR-10 and LSUN datasets. Our numerical results validate the proposed theory.
Yogesh Balaji, Seyed Hamed Hassani, Rama Chellappa, Soheil Feizi
ICML3
2019 Unsupervised Super-Resolution of Satellite Imagery for High Fidelity Material Label Transfer
abstract
Urban material recognition in remote sensing imagery is a challenging problem due to the difficulty of obtaining human annotations, especially on low resolution satellite images. To this end, we propose an unsupervised domain adaptation-based approach using adversarial learning. We aim to harvest information from smaller quantities of high resolution data (source domain) and utilize the same to super-resolve low resolution imagery (target domain). This can potentially aid in semantic as well as material label transfer from a richly annotated source to a target domain.
Arthita Ghosh, Max Ehrlich, Larry Davis 0001, Rama Chellappa
IGARSS4
2019 Conditional GAN with Discriminative Filter Generation for Text-to-Video Synthesis
abstract
Developing conditional generative models for text-to-video synthesis is an extremely challenging yet an important topic of research in machine learning. In this work, we address this problem by introducing Text-Filter conditioning Generative Adversarial Network (TFGAN), a conditional GAN model with a novel multi-scale text-conditioning scheme that improves text-video associations. By combining the proposed conditioning scheme with a deep GAN architecture, TFGAN generates high quality videos from text on challenging real-world video datasets. In addition, we construct a synthetic dataset of text-conditioned moving shapes to systematically evaluate our conditioning scheme. Extensive experiments demonstrate that TFGAN significantly outperforms existing approaches, and can also generate videos of novel categories not seen during training.
Yogesh Balaji, Martin Renqiang Min, Rama Chellappa, Hans Peter Graf
IJCAI4
2019 On Measuring the Iconicity of a Face
abstract
For a given identity in a face dataset, there are certain iconic images which are more representative of the subject than others. In this paper, we explore the problem of computing the iconicity of a face. The premise of the proposed approach is as follows: For an identity containing a mixture of iconic and non iconic images, if a given face cannot be successfully matched with any other face of the same identity, then the iconicity of the face image is low. Using this information, we train a Siamese Multi-Layer Perceptron network, such that each of its twins predict iconicity scores of the image feature pair, fed in as input. We observe the variation of the obtained scores with respect to covariates such as blur, yaw, pitch, roll and occlusion to demonstrate that they effectively predict the quality of the image and compare it with other existing metrics. Furthermore, we use these scores to weight features for template-based face verification and compare it with media averaging of features.
Prithviraj Dhar, Carlos Domingo Castillo, Rama Chellappa
WACV3
2019 A Proposal-Based Solution to Spatio-Temporal Action Detection in Untrimmed Videos
abstract
Existing approaches for spatio-temporal action detection in videos are limited by the spatial extent and temporal duration of the actions. In this paper, we present a modular system for spatio-temporal action detection in untrimmed surveillance videos. We propose a two stage approach. The first stage generates dense spatio-temporal proposals using hierarchical clustering and temporal jittering techniques on frame-wise object detections. The second stage is a Temporal Refinement I3D (TRI-3D) network that performs action classification and temporal refinement on the generated proposals. The object detection-based proposal generation step helps in detecting actions occurring in a small spatial region of a video frame, while temporal jittering and refinement helps in detecting actions of variable lengths. Experimental results on an unconstrained surveillance action detection dataset - DIVA - show the effectiveness of our system. For comparison, the performance of our system is also evaluated on the THUMOS'14 temporal action detection dataset.
Joshua Gleason, Rajeev Ranjan 0003, Steven Schwarcz, Carlos Domingo Castillo, Jun-Cheng Chen, Rama Chellappa
WACV6
2019 From BoW to CNN: Two Decades of Texture Representation for Texture Classification
abstract
Texture is a fundamental characteristic of many types of images, and texture representation is one of the essential and challenging problems in computer vision and pattern recognition which has attracted extensive research attention over several decades. Since 2000, texture representations based on Bag of Words and on Convolutional Neural Networks have been extensively studied with impressive performance. Given this period of remarkable evolution, this paper aims to present a comprehensive survey of advances in texture representation over the last two decades. More than 250 major publications are cited in this survey covering different aspects of the research, including benchmark datasets and state of the art results. In retrospect of what has been achieved so far, the survey discusses open challenges and directions for future research.
Li Liu 0002, Jie Chen 0001, Paul W. Fieguth, Guoying Zhao 0001, Rama Chellappa, Matti Pietikäinen
Int. J. Comput. Vis.5
2019 Editorial: Special Issue on Deep Learning for Face Analysis
Chen Change Loy, Xiaoming Liu 0002, Tae-Kyun Kim 0001, Fernando De la Torre, Rama Chellappa
Int. J. Comput. Vis.5
2019 Partial face detection in the mobile domain
Upal Mahbub, Sayantan Sarkar, Rama Chellappa
Image Vis. Comput.3
2019 Guest Editors' Introduction to the Special Section on Compact and Efficient Feature Representation and Learning in Computer Vision
abstract
The papers in this special section examine compact and efficient feature representation and learning in computer vision.
Li Liu 0002, Matti Pietikäinen, Jie Chen 0001, Guoying Zhao 0001, Xiaogang Wang 0001, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.6
2019 HyperFace: A Deep Multi-Task Learning Framework for Face Detection, Landmark Localization, Pose Estimation, and Gender Recognition
abstract
We present an algorithm for simultaneous face detection, landmarks localization, pose estimation and gender recognition using deep convolutional neural networks (CNN). The proposed method called, HyperFace, fuses the intermediate layers of a deep CNN using a separate CNN followed by a multi-task learning algorithm that operates on the fused features. It exploits the synergy among the tasks which boosts up their individual performances. Additionally, we propose two variants of HyperFace: (1) HyperFace-ResNet that builds on the ResNet-101 model and achieves significant improvement in performance, and (2) Fast-HyperFace that uses a high recall fast face detector for generating region proposals to improve the speed of the algorithm. Extensive experiments show that the proposed models are able to capture both global and local information in faces and performs significantly better than many competitive algorithms for each of these four tasks.
Rajeev Ranjan 0003, Vishal M. Patel, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 ExprGAN: Facial Expression Editing With Controllable Expression Intensity
abstract
Facial expression editing is a challenging task as it needs a high-level semantic understanding of the input face image. In conventional methods, either paired training data is required or the synthetic face’s resolution is low. Moreover,only the categories of facial expression can be changed. To address these limitations, we propose an Expression Generative Adversarial Network (ExprGAN) for photo-realistic facial expression editing with controllable expression intensity. An expression controller module is specially designed to learn an expressive and compact expression code in addition to the encoder-decoder network. This novel architecture enables the expression intensity to be continuously adjusted from low to high. We further show that our ExprGAN can be applied for other tasks, such as expression transfer, image retrieval, and data augmentation for training improved face expression recognition models. To tackle the small size of the training database, an effective incremental learning scheme is proposed. Quantitative and qualitative evaluations on the widely used Oulu-CASIA dataset demonstrate the effectiveness of ExprGAN.
Hui Ding 0002, Kumar Sricharan, Rama Chellappa
AAAI3
2018 A Deep Cascade Network for Unaligned Face Attribute Classification
abstract
Humans focus attention on different face regions when recognizing face attributes. Most existing face attribute classification methods use the whole image as input. Moreover, some of these methods rely on fiducial landmarks to provide defined face parts. In this paper, we propose a cascade network that simultaneously learns to localize face regions specific to attributes and performs attribute classification without alignment. First, a weakly-supervised face region localization network is designed to automatically detect regions (or parts) specific to attributes. Then multiple part-based networks and a whole-image-based network are separately constructed and combined together by the region switch layer and attribute relation layer for final attribute classification. A multi-net learning method and hint-based model compression is further proposed to get an effective localization model and a compact classification model, respectively. Our approach achieves significantly better performance than state-of-the-art methods on unaligned CelebA dataset, reducing the classification error by 30.9%.
Hui Ding 0002, Shaohua Kevin Zhou, Rama Chellappa
AAAI4
2018 Doing the Best We Can With What We Have: Multi-Label Balancing With Selective Learning for Attribute Prediction
abstract
Attributes are human describable features, which have been used successfully for face, object, and activity recognition. Facial attributes are intuitive descriptions of faces and have proven to be very useful in face recognition and verification. Despite their usefulness, to date there is only one large-scale facial attribute dataset, CelebA. Impressive results have been achieved on this dataset, but it exhibits a variety of very significant biases. As CelebA contains mostly frontal idealized images of celebrities, it is difficult to generalize a model trained on this data for use on another dataset (of non celebrities). A typical approach to dealing with imbalanced data involves sampling the data in order to balance the positive and negative labels, however, with a multi-label problem this becomes a non-trivial task. By sampling to balance one label, we affect the distribution of other labels in the data. To address this problem, we introduce a novel Selective Learning method for deep networks which adaptively balances the data in each batch according to the desired distribution for each label. The bias in CelebA can be corrected for in this way, allowing the network to learn a more robust attribute model. We argue that without this multi-label balancing, the network cannot learn to accurately predict attributes that are poorly represented in CelebA. We demonstrate the effectiveness of our method on the problem of facial attribute prediction on CelebA, LFWA, and the new University of Maryland Attribute Evaluation Dataset (UMD-AED), outperforming the state-of-the-art on each dataset.
Emily Morgan Hand, Carlos Domingo Castillo, Rama Chellappa
AAAI3
2018 Task-Aware Compressed Sensing With Generative Adversarial Networks
abstract
In recent years, neural network approaches have been widely adopted for machine learning tasks, with applications in computer vision. More recently, unsupervised generative models based on neural networks have been successfully applied to model data distributions via low-dimensional latent spaces. In this paper, we use Generative Adversarial Networks (GANs) to impose structure in compressed sensing problems, replacing the usual sparsity constraint. We propose to train the GANs in a task-aware fashion, specifically for reconstruction tasks. We also show that it is possible to train our model without using any (or much) non-compressed data. Finally, we show that the latent space of the GAN carries discriminative information and can further be regularized to generate input features for general inference tasks. We demonstrate the effectiveness of our method on a variety of reconstruction and classification problems.
Maya Kabkab, Pouya Samangouei, Rama Chellappa
AAAI3
2018 Regularizing Deep Networks Using Efficient Layerwise Adversarial Training
abstract
Adversarial training has been shown to regularize deep neural networks in addition to increasing their robustness to adversarial examples. However, the regularization effect on very deep state of the art networks has not been fully investigated. In this paper, we present a novel approach to regularize deep neural networks by perturbing intermediate layer activations in an efficient manner. We use these perturbations to train very deep models such as ResNets and WideResNets and show improvement in performance across datasets of different sizes such as CIFAR-10, CIFAR-100 and ImageNet. Our ablative experiments show that the proposed approach not only provides stronger regularization compared to Dropout but also improves adversarial robustness comparable to traditional adversarial training approaches.
Swami Sankaranarayanan, Rama Chellappa, Ser-Nam Lim
AAAI3
2018 Nonlinear Subspace Feature Enhancement for Image Set Classification
Mohammed E. Fathy 0001, Azadeh Alavi, Rama Chellappa
ACCV (4)3
2018 Disentangling 3D Pose in a Dendritic CNN for Unconstrained 2D Face Alignment
abstract
Heatmap regression has been used for landmark localization for quite a while now. Most of the methods use a very deep stack of bottleneck modules for heatmap classification stage, followed by heatmap regression to extract the keypoints. In this paper, we present a single dendritic CNN, termed as Pose Conditioned Dendritic Convolution Neural Network (PCD-CNN), where a classification network is followed by a second and modular classification network, trained in an end to end fashion to obtain accurate landmark points. Following a Bayesian formulation, we disentangle the 3D pose of a face image explicitly by conditioning the landmark estimation on pose, making it different from multi-tasking approaches. Extensive experimentation shows that conditioning on pose reduces the localization error by making it agnostic to face pose. The proposed model can be extended to yield variable number of landmark points and hence broadening its applicability to other datasets. Instead of increasing depth or width of the network, we train the CNN efficiently with Mask-Softmax Loss and hard sample mining to achieve upto 15% reduction in error compared to state-of-the-art methods for extreme and medium pose face images from challenging datasets including AFLW, AFW, COFW and IBUG.
Amit Kumar 0013, Rama Chellappa
CVPR2
2018 Deep Density Clustering of Unconstrained Faces
abstract
In this paper, we consider the problem of grouping a collection of unconstrained face images in which the number of subjects is not known. We propose an unsupervised clustering algorithm called Deep Density Clustering (DDC) which is based on measuring density affinities between local neighborhoods in the feature space. By learning the minimal covering sphere for each neighborhood, information about the underlying structure is encapsulated. The encapsulation is also capable of locating high-density region of the neighborhood, which aids in measuring the neighborhood similarity. We theoretically show that the encapsulation asymptotically converges to a Parzen window density estimator. Our experiments show that DDC is a superior candidate for clustering unconstrained faces when the number of subjects is unknown. Unlike conventional linkage and density-based methods that are sensitive to the selection operating points, DDC attains more consistent and improved performance. Furthermore, the density-aware property reduces the difficulty in finding appropriate operating points.
Wei-An Lin, Jun-Cheng Chen, Carlos Domingo Castillo, Rama Chellappa
CVPR4
2018 Learning From Synthetic Data: Addressing Domain Shift for Semantic Segmentation
abstract
Visual Domain Adaptation is a problem of immense importance in computer vision. Previous approaches showcase the inability of even deep neural networks to learn informative representations across domain shift. This problem is more severe for tasks where acquiring hand labeled data is extremely hard and tedious. In this work, we focus on adapting the representations learned by segmentation networks across synthetic and real domains. Contrary to previous approaches that use a simple adversarial objective or superpixel information to aid the process, we propose an approach based on Generative Adversarial Networks (GANs) that brings the embeddings closer in the learned feature space. To showcase the generality and scalability of our approach, we show that we can achieve state of the art results on two challenging scenarios of synthetic to real domain adaptation. Additional exploratory experiments show that our approach: (1) generalizes to unseen domains and (2) results in improved alignment of source and target distributions.
Swami Sankaranarayanan, Yogesh Balaji, Ser-Nam Lim, Rama Chellappa
CVPR5
2018 Generate to Adapt: Aligning Domains Using Generative Adversarial Networks
abstract
Domain Adaptation is an actively researched problem in Computer Vision. In this work, we propose an approach that leverages unsupervised data to bring the source and target distributions closer in a learned joint feature space. We accomplish this by inducing a symbiotic relationship between the learned embedding and a generative adversarial network. This is in contrast to methods which use the adversarial framework for realistic data generation and retraining deep models with such data. We demonstrate the strength and generality of our approach by performing experiments on three different tasks with varying levels of difficulty: (1) Digit classification (MNIST, SVHN and USPS datasets) (2) Object recognition using OFFICE dataset and (3) Domain adaptation from synthetic to real data. Our method achieves state-of-the art performance in most experimental settings and by far the only GAN-based method that has been shown to work well across different datasets such as OFFICE and DIGITS.
Swami Sankaranarayanan, Yogesh Balaji, Carlos Domingo Castillo, Rama Chellappa
CVPR4
2018 Zero-Shot Object Detection
Ankan Bansal, Karan Sikka, Gaurav Sharma 0004, Rama Chellappa, Ajay Divakaran
ECCV (1)4
2018 Semi-supervised FusedGAN for Conditional Image Generation
Navaneeth Bodla, Gang Hua 0001, Rama Chellappa
ECCV (5)3
2018 Deep Regionlets for Object Detection
Hongyu Xu, Xutao Lv, Xiaoyu Wang 0002, Zhou Ren, Navaneeth Bodla, Rama Chellappa
ECCV (11)6
2018 A Real-Time Multi-Task Single Shot Face Detector
abstract
Face, fiducial detection, and 3D head pose estimation are important face preprocessing modules for face recognition which are usually performed separately and loosely coupled. In this paper, we propose a unifying framework to simultaneously detect face, fiducial points, and head pose in real-time. In addition, since no single dataset contains all the required and best annotations, we develop a progressive training strategy to overcome the annotation discrepancy across different datasets. Extensive experiments on face detection, fiducial detection, and pose estimation benchmarks demonstrate the proposed approach can achieve comparable performance to a state-of-the-art system [1] but runs 60 times faster. (i.e., 20 frames per second).
Jun-Cheng Chen, Wei-An Lin, Jingxiao Zheng, Rama Chellappa
ICIP4
2018 Defense-GAN: Protecting Classifiers Against Adversarial Attacks Using Generative Models
Pouya Samangouei, Maya Kabkab, Rama Chellappa
ICLR (Poster)3
2018 MetaReg: Towards Domain Generalization using Meta-Regularization
abstract
Training models that generalize to new domains at test time is a problem of fundamental importance in machine learning. In this work, we encode this notion of domain generalization using a novel regularization function. We pose the problem of finding such a regularization function in a Learning to Learn (or) meta-learning framework. The objective of domain generalization is explicitly modeled by learning a regularizer that makes the model trained on one domain to perform well on another domain. Experimental validations on computer vision and natural language datasets indicate that our method can learn regularizers that achieve good cross-domain generalization.
Yogesh Balaji, Swami Sankaranarayanan, Rama Chellappa
NeurIPS3
2018 Predicting Facial Attributes in Video Using Temporal Coherence and Motion-Attention
abstract
Recent research progress in facial attribute recognition has been dominated by small improvements on the only large-scale publicly available benchmark dataset, CelebA [18]. We propose to extend attribute prediction research to unconstrained videos. Applying attribute models trained on CelebA - a still image dataset - to video data highlights several major problems with current models, including the lack of consideration for both time and motion. Many facial attributes (e.g. gender, hair color) should be consistent throughout a video, however, current models do not produce consistent results. We introduce two methods to increase the consistency and accuracy of attribute responses in videos: a temporal coherence constraint, and a motionattention mechanism. Both methods work on weakly labeled data, requiring attribute labels for only one frame in a sequence, which we call the anchor frame. The temporal coherence constraint moves the network responses of non-anchor frames toward the responses of anchor frames for each sequence, resulting in more stable and accurate attribute predictions. We use the motion between anchor and non-anchor video frames as an attention mechanism, discarding the information from parts of the non-anchor frame where no motion occurred. This motion-attention focuses the network on the moving parts of the non-anchor frames (i.e. the face). Since there is no large-scale video dataset labeled with attributes, it is essential for attribute models to be able to learn from weakly labeled data. We demonstrate the effectiveness of the proposed methods by evaluating them on the challenging YouTube Faces video dataset [31]. The proposed motion-attention and temporal coherence methods outperform attribute models trained on CelebA, as well as those fine-tuned on video data. To the best of our knowledge, this paper is the first to address the problem of facial attribute prediction in video.
Emily Morgan Hand, Carlos Domingo Castillo, Rama Chellappa
WACV3
2018 Face-MagNet: Magnifying Feature Maps to Detect Small Faces
abstract
In this paper, we introduce the Face Magnifier Network (Face-MageNet), a face detector based on the Faster-RCNN framework which enables the flow of discriminative information of small scale faces to the classifier without any skip or residual connections. To achieve this, Face-MagNet deploys a set of ConvTranspose, also known as deconvolution, layers in the Region Proposal Network (RPN) and another set before the Region of Interest (RoI) pooling layer to facilitate detection of finer faces. In addition, we also design, train, and evaluate three other well-tuned architectures that represent the conventional solutions to the scale problem: context pooling, skip connections, and scale partitioning. Each of these three networks achieves comparable results to the state-of-the-art face detectors. With extensive experiments, we show that Face-MagNet based on a VGG16 architecture achieves better results than the recently proposed ResNet101-based HR [7] method on the task of face detection on WIDER [25] dataset and also achieves similar results on the hard set as our other method SSH [17].
Pouya Samangouei, Rama Chellappa, Mahyar Najibi, Larry Davis 0001
WACV2
2018 Unconstrained Still/Video-Based Face Verification with Deep Convolutional Neural Networks
Jun-Cheng Chen, Rajeev Ranjan 0003, Swami Sankaranarayanan, Amit Kumar 0013, Ching-Hui Chen, Vishal M. Patel, Carlos Domingo Castillo, Rama Chellappa
Int. J. Comput. Vis.8
2018 KEPLER: Simultaneous estimation of keypoints and 3D pose of unconstrained faces in a unified framework by learning efficient H-CNN regressors
Amit Kumar 0013, Azadeh Alavi, Rama Chellappa
Image Vis. Comput.3
2018 Proximity-Aware Hierarchical Clustering of unconstrained faces
Wei-An Lin, Jun-Cheng Chen, Rajeev Ranjan 0003, Ankan Bansal, Swami Sankaranarayanan, Carlos Domingo Castillo, Rama Chellappa
Image Vis. Comput.7
2018 Learning from Ambiguously Labeled Face Images
abstract
Learning a classifier from ambiguously labeled face images is challenging since training images are not always explicitly-labeled. For instance, face images of two persons in a news photo are not explicitly labeled by their names in the caption. We propose a Matrix Completion for Ambiguity Resolution (MCar) method for predicting the actual labels from ambiguously labeled images. This step is followed by learning a standard supervised classifier from the disambiguated labels to classify new images. To prevent the majority labels from dominating the result of MCar, we generalize MCar to a weighted MCar (WMCar) that handles label imbalance. Since WMCar outputs a soft labeling vector of reduced ambiguity for each instance, we can iteratively refine it by feeding it as the input to WMCar. Nevertheless, such an iterative implementation can be affected by the noisy soft labeling vectors, and thus the performance may degrade. Our proposed Iterative Candidate Elimination (ICE) procedure makes the iterative ambiguity resolution possible by gradually eliminating a portion of least likely candidates in ambiguously labeled faces. We further extend MCar to incorporate the labeling constraints among instances when such prior knowledge is available. Compared to existing methods, our approach demonstrates improvements on several ambiguously labeled datasets.
Ching-Hui Chen, Vishal M. Patel, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 Learning Common and Feature-Specific Patterns: A Novel Multiple-Sparse-Representation-Based Tracker
abstract
The use of multiple features has been shown to be an effective strategy for visual tracking because of their complementary contributions to appearance modeling. The key problem is how to learn a fused representation from multiple features for appearance modeling. Different features extracted from the same object should share some commonalities in their representations while each feature should also have some feature-specific representation patterns which reflect its complementarity in appearance modeling. Different from existing multi-feature sparse trackers which only consider the commonalities among the sparsity patterns of multiple features, this paper proposes a novel multiple sparse representation framework for visual tracking which jointly exploits the shared and feature-specific properties of different features by decomposing multiple sparsity patterns. Moreover, we introduce a novel online multiple metric learning to efficiently and adaptively incorporate the appearance proximity constraint, which ensures that the learned commonalities of multiple features are more representative. Experimental results on tracking benchmark videos and other challenging videos demonstrate the effectiveness of the proposed tracker.
Xiangyuan Lan, Shengping Zhang, Pong C. Yuen, Rama Chellappa
IEEE Trans. Image Process.4
2017 Attributes for Improved Attributes: A Multi-Task Network Utilizing Implicit and Explicit Relationships for Facial Attribute Classification
abstract
Attributes, or mid-level semantic features, have gained popularity in the past few years in domains ranging from activity recognition to face verification. Improving the accuracy of attribute classifiers is an important first step in any application which uses these attributes. In most works to date, attributes have been considered independent of each other. However, attributes can be strongly related, such as heavy makeup and wearing lipstick as well as male and goatee and many others. We propose a multi-task deep convolutional neural network (MCNN) with an auxiliary network at the top (AUX) which takes advantage of attribute relationships for improved classification. We call our final network MCNN-AUX. MCNN-AUX uses attribute relationships in three ways: by sharing the lowest layers for all attributes, by sharing the higher layers for spatially-related attributes, and by feeding the attribute scores from MCNN into the AUX network to find score-level relationships. Using MCNN-AUX rather than individual attribute classifiers, we are able to reduce the number of parameters in the network from 64 million to fewer than 16 million and reduce the training time by a factor of 16. We demonstrate the effectiveness of our method by producing results on two challenging publicly available datasets achieving state-of-the-art performance on many attributes.
Emily Morgan Hand, Rama Chellappa
AAAI2
2017 Robust MIL-Based Feature Template Learning for Object Tracking
abstract
Because of appearance variations, training samples of the tracked targets collected by the online tracker are required for updating the tracking model. However, this often leads to tracking drift problem because of potentially corrupted samples: 1) contaminated/outlier samples resulting from large variations (e.g. occlusion, illumination), and 2) misaligned samples caused by tracking inaccuracy. Therefore, in order to reduce the tracking drift while maintaining the adaptability of a visual tracker, how to alleviate these two issues via an effective model learning (updating) strategy is a key problem to be solved. To address these issues, this paper proposes a novel and optimal model learning (updating) scheme which aims to simultaneously eliminate the negative effects from these two issues mentioned above in a unified robust feature template learning framework. Particularly, the proposed feature template learning framework is capable of: 1) adaptively learning uncontaminated feature templates by separating out contaminated samples, and 2) resolving label ambiguities caused by misaligned samples via a probabilistic multiple instance learning (MIL) model. Experiments on challenging video sequences show that the proposed tracker performs favourably against several state-of-the-art trackers.
Xiangyuan Lan, Pong C. Yuen, Rama Chellappa
AAAI3
2017 Hierarchical Multimodal Metric Learning for Multimodal Classification
abstract
Multimodal classification arises in many computer vision tasks such as object classification and image retrieval. The idea is to utilize multiple sources (modalities) measuring the same instance to improve the overall performance compared to using a single source (modality). The varying characteristics exhibited by multiple modalities make it necessary to simultaneously learn the corresponding metrics. In this paper, we propose a multiple metrics learning algorithm for multimodal data. Metric of each modality is a product of two matrices: one matrix is modality specific, the other is enforced to be shared by all the modalities. The learned metrics can improve multimodal classification accuracy and experimental results on four datasets show that the proposed algorithm outperforms existing learning algorithms based on multiple metrics as well as other approaches tested on these datasets. Specifically, we report 95.0% object instance recognition accuracy, 89.2% object category recognition accuracy on the multi-view RGB-D dataset and 52.3% scene category recognition accuracy on SUN RGB-D dataset.
Heng Zhang 0003, Vishal M. Patel, Rama Chellappa
CVPR3
2017 Video-Based Face Association and Identification
abstract
In this paper, we present a new video-based face identification algorithm, where the target (i.e., person of interest) in the probe video is only annotated once with a face bounding box in a frame and the video may consist of multiple shots. Most video face identification techniques assume that the video is of single shot, and thus the bounding boxes of the target face can be extracted by tracking a face across the video frames. Nevertheless, such automatic annotation is vulnerable to the drifting of the face tracker, and the face tracking algorithm is inadequate to associate the face images of the target across multiple shots. In this paper, we propose a target face association (TFA) technique that retrieves a set of representative face images in a given video that are likely to have the same identity as the target face. These face images are then utilized to construct a robust face representation of the target face for searching the corresponding subject in the gallery. Since two faces that appear in the same video frame cannot belong to the same person, such cannot-link constraints are utilized for learning a target-specific linear classifier for establishing the intra/inter-shot face association of the target. Experimental results on the newly released JANUS challenge set 3 (JANUS CS3) dataset show that our method generates robust representations from target-annotated videos and demonstrates good performance for the task of video-based face identification problem.
Ching-Hui Chen, Jun-Cheng Chen, Carlos Domingo Castillo, Rama Chellappa
FG4
2017 FaceNet2ExpNet: Regularizing a Deep Face Recognition Net for Expression Recognition
abstract
Relatively small data sets available for expression recognition research make the training of deep networks very challenging. Although fine-tuning can partially alleviate the issue, the performance is still below acceptable levels as the deep features probably contain redundant information from the pretrained domain. In this paper, we present FaceNet2ExpNet, a novel idea to train an expression recognition network based on static images. We first propose a new distribution function to model the high-level neurons of the expression network. Based on this, a two-stage training algorithm is carefully designed. In the pre-training stage, we train the convolutional layers of the expression net, regularized by the face net; In the refining stage, we append fully-connected layers to the pre-trained convolutional layers and train the whole network jointly. Visualization results show that the model trained with our method captures improved high-level expression semantics. Evaluations on four public expression databases, CK+, Oulu- CASIA, TFD, and SFEW demonstrate that our method achieves better results than state-of-the-art.
Hui Ding 0002, Shaohua Kevin Zhou, Rama Chellappa
FG3
2017 KEPLER: Keypoint and Pose Estimation of Unconstrained Faces by Learning Efficient H-CNN Regressors
abstract
Keypoint detection is one of the most important pre-processing steps in tasks such as face modeling, recognition and verification. In this paper, we present an iterative method for Keypoint Estimation and Pose prediction of unconstrained faces by Learning Efficient H-CNN Regressors (KEPLER) for addressing the face alignment problem. Recent state of the art methods have shown improvements in face keypoint detection by employing Convolution Neural Networks (CNNs). Although a simple feed forward neural network can learn the mapping between input and output spaces, it cannot learn the inherent structural dependencies. We present a novel architecture called H-CNN (Heatmap-CNN) which captures structured global and local features and thus favors accurate keypoint detection. H-CNN is jointly trained on the visibility, fiducials and 3D-pose of the face. As the iterations proceed, the error decreases making the gradients small and thus requiring efficient training of DCNNs to mitigate this. KEPLER performs global corrections in pose and fiducials for the first four iterations followed by local corrections in a subsequent stage. As a by-product, KEPLER also provides 3D pose (pitch, yaw and roll) of the face accurately. In this paper, we show that without using any 3D information, KEPLER outperforms state of the art methods for alignment on challenging datasets such as AFW [38] and AFLW [17].
Amit Kumar 0013, Azadeh Alavi, Rama Chellappa
FG3
2017 A Proximity-Aware Hierarchical Clustering of Faces
abstract
In this paper, we propose an unsupervised face clustering algorithm called “Proximity-Aware Hierarchical Clustering” (PAHC) that exploits the local structure of deep representations. In the proposed method, a similarity measure between deep features is computed by evaluating linear SVM margins. SVMs are trained using nearest neighbors of sample data, and thus do not require any external training data. Clus- ters are then formed by thresholding the similarity scores. We evaluate the clustering performance using three challenging un- constrained face datasets, including Celebrity in Frontal-Profile (CFP), IARPA JANUS Benchmark A (IJB-A), and JANUS Challenge Set 3 (JANUS CS3) datasets. Experimental results demonstrate that the proposed approach can achieve significant improvements over state-of-the-art methods. Moreover, we also show that the proposed clustering algorithm can be applied to curate a set of large-scale and noisy training dataset while maintaining sufficient amount of images and their variations due to nuisance factors. The face verification performance on JANUS CS3 improves significantly by finetuning a DCNN model with the curated MS-Celeb-1M dataset which contains over three million face images.
Wei-An Lin, Jun-Cheng Chen, Rama Chellappa
FG3
2017 Pooling Facial Segments to Face: The Shallow and Deep Ends
abstract
Generic face detection algorithms do not perform very well in the mobile domain due to significant presence of occluded and partially visible faces. One promising technique to handle the challenge of partial faces is to design face detectors based on facial segments. In this paper two such face detectors namely, SegFace and DeepSegFace, are proposed that detect the presence of a face given arbitrary combinations of certain face segments. Both methods use proposals from facial segments as input that are found using weak boosted classifiers. SegFace is a shallow and fast algorithm using traditional features, tailored for situations where real time constraints must be satisfied. On the other hand, DeepSegFace is a more powerful algorithm based on a deep convolutional neutral network (DCNN) architecture. DeepSegFace offers certain advantages over other DCNN-based face detectors as it requires relatively small amount of data to train by utilizing a novel data augmentation scheme and is very robust to occlusion by design. Extensive experiments show the superiority of the proposed methods, specially DeepSegFace, over other state-of-the-art face detectors in terms of precision-recall and ROC curve on two mobile face datasets.
Upal Mahbub, Sayantan Sarkar, Rama Chellappa
FG3
2017 An All-In-One Convolutional Neural Network for Face Analysis
abstract
We present a multi-purpose algorithm for simultaneous face detection, face alignment, pose estimation, gender recognition, smile detection, age estimation and face recognition using a single deep convolutional neural network (CNN). Theproposed method employs a multi-task learning framework that regularizes the shared parameters of CNN and builds a synergy among different domains and tasks. Extensive experiments show that the network has a better understanding of face and achieves state-of-the-art result for most of these tasks.
Rajeev Ranjan 0003, Swami Sankaranarayanan, Carlos Domingo Castillo, Rama Chellappa
FG4
2017 Deep Network Shrinkage Applied to Cross-Spectrum Face Recognition
abstract
In recent years, deep learning has emerged as a dominant methodology in virtually all machine learning problems. While it has been shown to produce state-of-the-art results for a variety of applicatons (including face recognition and heterogeneous face recognition), one aspect of deep networks that has not been extensively researched is how to determine the optimal network structure. This problem is generally solved by ad hoc methods. In this work we address a subproblem of this task: determining the breadth (number of nodes) of each layer. We show how to use group-sparsity-inducing regularization to effectively replace these hyper-parameters with a single hyperparameter which can be determined by cross-validation. We demonstrate our method by using it to reduce the size of networks on two commonly used NIR face datasets.
Christopher Reale, Hyungtae Lee, Heesung Kwon, Rama Chellappa
FG4
2017 UMDFaces: An annotated face dataset for training deep networks
abstract
Recent progress in face detection (including keypoint detection), and recognition is mainly being driven by (i) deeper convolutional neural network architectures, and (ii) larger datasets. However, most of the large datasets are maintained by private companies and are not publicly available. The academic computer vision community needs larger and more varied datasets to make further progress. In this paper, we introduce a new face dataset, called UMDFaces, which has 367,888 annotated faces of 8,277 subjects. We also introduce a new face recognition evaluation protocol which will help advance the state-of-the-art in this area. We discuss how a large dataset can be collected and annotated using human annotators and deep networks. We provide human curated bounding boxes for faces. We also provide estimated pose (roll, pitch and yaw), locations of twenty-one key-points and gender information generated by a pre-trained neural network. In addition, the quality of keypoint annotations has been verified by humans for about 115,000 images. Finally, we compare the quality of the dataset with other publicly available face datasets at similar scales.
Ankan Bansal, Anirudh Nanduri, Carlos Domingo Castillo, Rajeev Ranjan 0003, Rama Chellappa
IJCB5
2017 Soft-NMS - Improving Object Detection with One Line of Code
abstract
Non-maximum suppression is an integral part of the object detection pipeline. First, it sorts all detection boxes on the basis of their scores. The detection box M with the maximum score is selected and all other detection boxes with a significant overlap (using a pre-defined threshold) with M are suppressed. This process is recursively applied on the remaining boxes. As per the design of the algorithm, if an object lies within the predefined overlap threshold, it leads to a miss. To this end, we propose Soft-NMS, an algorithm which decays the detection scores of all other objects as a continuous function of their overlap with M. Hence, no object is eliminated in this process. Soft-NMS obtains consistent improvements for the coco-style mAP metric on standard datasets like PASCAL VOC2007 (1.7% for both R-FCN and Faster-RCNN) and MS-COCO (1.3% for R-FCN and 1.1% for Faster-RCNN) by just changing the NMS algorithm without any additional hyper-parameters. Using Deformable-RFCN, Soft-NMS improves state-of-the-art in object detection from 39.8% to 40.9% with a single model. Further, the computational complexity of Soft-NMS is the same as traditional NMS and hence it can be efficiently implemented. Since Soft-NMS does not require any extra training and is simple to implement, it can be easily integrated into any object detection pipeline. Code for Soft-NMS is publicly available on GitHub http://bit.ly/2nJLNMu.
Navaneeth Bodla, Rama Chellappa, Larry Davis 0001
ICCV3
2017 SSH: Single Stage Headless Face Detector
abstract
We introduce the Single Stage Headless (SSH) face detector. Unlike two stage proposal-classification detectors, SSH detects faces in a single stage directly from the early convolutional layers in a classification network. SSH is headless. That is, it is able to achieve state-of-the-art results while removing the “head” of its underlying classification network - i.e. all fully connected layers in the VGG-16 which contains a large number of parameters. Additionally, instead of relying on an image pyramid to detect faces with various scales, SSH is scale-invariant by design. We simultaneously detect faces with different scales in a single forward pass of the network, but from different layers. These properties make SSH fast and light-weight. Surprisingly, with a headless VGG-16, SSH beats the ResNet-101-based state-of-the-art on the WIDER dataset. Even though, unlike the current state-of-the-art, SSH does not use an image pyramid and is 5X faster. Moreover, if an image pyramid is deployed, our light-weight network achieves state-of-the-art on all subsets of the WIDER dataset, improving the AP by 2.5%. SSH also reaches state-of-the-art results on the FDDB and Pascal-Faces datasets while using a small input size, leading to a speed of 50 frames/second on a GPU.
Mahyar Najibi, Pouya Samangouei, Rama Chellappa, Larry Davis 0001
ICCV3
2017 Deep Heterogeneous Feature Fusion for Template-Based Face Recognition
abstract
Although deep learning has yielded impressive performance for face recognition, many studies have shown that different networks learn different feature maps: while some networks are more receptive to pose and illumination others appear to capture more local information. Thus, in this work, we propose a deep heterogeneous feature fusion network to exploit the complementary information present in features generated by different deep convolutional neural networks (DCNNs) for template-based face recognition, where a template refers to a set of still face images or video frames from different sources which introduces more blur, pose, illumination and other variations than traditional face datasets. The proposed approach efficiently fuses the discriminative information of different deep features by 1) jointly learning the non-linear high-dimensional projection of the deep features and 2) generating a more discriminative template representation which preserves the inherent geometry of the deep features in the feature space. Experimental results on the IARPA Janus Challenge Set 3 (Janus CS3) dataset demonstrate that the proposed method can effectively improve the recognition performance. In addition, we also present a series of covariate experiments on the face verification task for in-depth qualitative evaluations for the proposed approach.
Navaneeth Bodla, Jingxiao Zheng, Hongyu Xu, Jun-Cheng Chen, Carlos Domingo Castillo, Rama Chellappa
WACV6
2017 Image Set Classification Using Sparse Bayesian Regression
abstract
This paper presents Bayesian Representation-based Classification (BRC), an approach based on sparse Bayesian regression and subspace clustering for image set classification. Similar to existing representation-based approaches such as Sparse RC (SRC) and Collaborative RC (CRC), BRC assumes that a test image is approximated by a linear combination of the gallery images of the true class. However, we show through a Bayesian statistical framework that BRC employs precision hyperpriors that are more non-informative than those of CRC/SRC, while also showing that CRC and SRC are identical up to an implicit choice of different precision hyperpriors. Furthermore, we analyze the assumptions of existing strategies for selecting the images to classify from a probe set (e.g. sequence mean) and we show that these strategies can still work under milder assumptions. Finally, we present a more robust probe set handling strategy that balances efficiency and accuracy. Experiments on three datasets illustrate the effectiveness of our algorithm compared to state-of-the-art set-based methods.
Mohammed E. Fathy 0001, Rama Chellappa
WACV2
2017 Pose-Robust Face Verification by Exploiting Competing Tasks
abstract
In this paper, we propose a pose-robust metric learning framework for unconstrained face verification by jointly optimizing face and pose verification tasks. We learn a joint model for these two tasks and explicitly discourage the information sharing between pose and identity verification metrics so as to mitigate the information contained in the pose verification task leading to making the identity metrics for face verification more pose-robust. Specifically, we use the joint Bayesian metric learning framework to learn the metrics for both tasks and enforce an orthogonal regularization constraint on the learned projection matrices for the two tasks. The pose labels used for training the joint model are automatically estimated and do not require extra annotations. An efficient stochastic gradient descent (SGD) algorithm is used to solve the optimization problem. We conduct extensive experiments on three challenging unconstrained face datasets and show promising results compared to state-of-the-art methods.
Boyu Lu, Jingxiao Zheng, Jun-Cheng Chen, Rama Chellappa
WACV4
2017 Growing Regression Tree Forests by Classification for Continuous Object Pose Estimation
Kota Hara, Rama Chellappa
Int. J. Comput. Vis.2
2017 Robust local features for remote face recognition
Jie Chen 0001, Vishal M. Patel, Li Liu 0002, Vili Kellokumpu, Guoying Zhao 0001, Matti Pietikäinen, Rama Chellappa
Image Vis. Comput.7
2017 Facial attributes for active authentication on mobile devices
Pouya Samangouei, Vishal M. Patel, Rama Chellappa
Image Vis. Comput.3
2017 Submodular Attribute Selection for Visual Recognition
abstract
In real-world visual recognition problems, low-level features cannot adequately characterize the semantic content in images, or the spatio-temporal structure in videos. In this work, we encode objects or actions based on attributes that describe them as high-level concepts. We consider two types of attributes. One type of attributes is generated by humans, while the second type is data-driven attributes extracted from data using dictionary learning methods. Attribute-based representation may exhibit variations due to noisy and redundant attributes. We propose a discriminative and compact attribute-based representation by selecting a subset of discriminative attributes from a large attribute set. Three attribute selection criteria are proposed and formulated as a submodular optimization problem. A greedy optimization algorithm is presented and its solution is guaranteed to be at least (1-1/e)-approximation to the optimum. Experimental results on four public datasets demonstrate that the proposed attribute-based representation significantly boosts the performance of visual recognition and outperforms most recently proposed recognition approaches.
Zhuolin Jiang, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 Editorial: Special issue on ubiquitous biometrics
Ran He 0001, Brian C. Lovell, Rama Chellappa, Anil K. Jain 0001, Zhenan Sun
Pattern Recognit.3
2017 Low-Rank and Joint Sparse Representations for Multi-Modal Recognition
abstract
We propose multi-task and multivariate methods for multi-modal recognition based on low-rank and joint sparse representations. Our formulations can be viewed as generalized versions of multivariate low-rank and sparse regression, where sparse and low-rank representations across all modalities are imposed. One of our methods simultaneously couples information within different modalities by enforcing the common low-rank and joint sparse constraints among multi-modal observations. We also modify our formulations by including an occlusion term that is assumed to be sparse. The alternating direction method of multipliers is proposed to efficiently solve the resulting optimization problems. Extensive experiments on three publicly available multi-modal biometrics and object recognition data sets show that our methods compare favorably with other feature-level fusion methods.
Heng Zhang 0003, Vishal M. Patel, Rama Chellappa
IEEE Trans. Image Process.3
2017 Deep Multitask Learning for Railway Track Inspection
abstract
Railroad tracks need to be periodically inspected and monitored to ensure safe transportation. Automated track inspection using computer vision and pattern recognition methods has recently shown the potential to improve safety by allowing for more frequent inspections while reducing human errors. Achieving full automation is still very challenging due to the number of different possible failure modes, as well as the broad range of image variations that can potentially trigger false alarms. In addition, the number of defective components is very small, so not many training examples are available for the machine to learn a robust anomaly detector. In this paper, we show that detection performance can be improved by combining multiple detectors within a multitask learning framework. We show that this approach results in improved accuracy for detecting defects on railway ties and fasteners.
Xavier Gibert, Vishal M. Patel, Rama Chellappa
IEEE Trans. Intell. Transp. Syst.3
2017 Guest Editorial Introduction to the Special Issue on Large-Scale Video Analytics for Enhanced Security: Algorithms and Systems
abstract
Due to the rapid increase of the number of cameras used in the video surveillance and the huge needs of the smart city and public security, video surveillance by human beings is no longer suitable. Hence, since the end of the last century, video analytics for security or visual surveillance has become one of the hottest research topics. Wide-area video surveillance systems can have extremely high data rates and high data volumes. Therefore, the challenge of video analytics is to extract meaningful information efficiently from the huge flow of video data in order to produce high-level semantic descriptions of the activities occurring in the area under surveillance.
Kaiqi Huang, Tieniu Tan, Stephen J. Maybank, Rama Chellappa, Jake Aggarval
IEEE Trans. Syst. Man Cybern. Syst.4
2016 Rolling Rotations for Recognizing Human Actions from 3D Skeletal Data
abstract
Recently, skeleton-based human action recognition has been receiving significant attention from various research communities due to the availability of depth sensors and real-time depth-based 3D skeleton estimation algorithms. In this work, we use rolling maps for recognizing human actions from 3D skeletal data. The rolling map is a well-defined mathematical concept that has not been explored much by the vision community. First, we represent each skeleton using the relative 3D rotations between various body parts. Since 3D rotations are members of the special orthogonal group SO3, our skeletal representation becomes a point in the Lie group SO3× ... × SO3, which is also a Riemannian manifold. Then, using this representation, we model human actions as curves in this Lie group. Since classification of curves in this non-Euclidean space is a difficult task, we unwrap the action curves onto the Lie algebra so3× ... × so3(which is a vector space) by combining the logarithm map with rolling maps, and perform classification in the Lie algebra. Experimental results on three action datasets show that the proposed approach performs equally well or better when compared to state-of-the-art.
Raviteja Vemulapalli, Rama Chellappa
CVPR2
2016 Gaussian Conditional Random Field Network for Semantic Segmentation
abstract
In contrast to the existing approaches that use discrete Conditional Random Field (CRF) models, we propose to use a Gaussian CRF model for the task of semantic segmentation. We propose a novel deep network, which we refer to as Gaussian Mean Field (GMF) network, whose layers perform mean field inference over a Gaussian CRF. The proposed GMF network has the desired property that each of its layers produces an output that is closer to the maximum a posteriori solution of the Gaussian CRF compared to its input. By combining the proposed GMF network with deep Convolutional Neural Networks (CNNs), we propose a new end-to-end trainable Gaussian conditional random field network. The proposed Gaussian CRF network is composed of three sub-networks: (i) a CNN-based unary network for generating unary potentials, (ii) a CNN-based pairwise network for generating pairwise potentials, and (iii) a GMF network for performing Gaussian CRF inference. When trained end-to-end in a discriminative fashion, and evaluated on the challenging PASCALVOC 2012 segmentation dataset, the proposed Gaussian CRF network outperforms various recent semantic segmentation approaches that combine CNNs with discrete CRF models.
Raviteja Vemulapalli, Oncel Tuzel, Ming-Yu Liu 0001, Rama Chellappa
CVPR4
2016 Fisher vector encoded deep convolutional features for unconstrained face verification
abstract
We present a method to combine the Fisher vector representation and the Deep Convolutional Neural Network (DCNN) features to generate a rerpesentation, called the Fisher vector encoded DCNN (FV-DCNN) features, for unconstrained face verification. One of the key features of our method is that spatial and appearance information are simultaneously processed when learning the Gaussian mixture model to encode the DCNN features. Evaluations on two challenging verification datasets show that the proposed FV-DCNN method is able to capture the salient local features and also performs well when compared to many state-of-the-art face verification methods.
Jun-Cheng Chen, Jingxiao Zheng, Vishal M. Patel, Rama Chellappa
ICIP4
2016 Partial face detection for continuous authentication
abstract
In this paper, a part-based technique for real time detection of users' faces on mobile devices is proposed. This method is specifically designed for detecting partially cropped and occluded faces captured using a smartphone's front-facing camera for continuous authentication. The key idea is to detect facial segments in the frame and cluster the results to obtain the region which is most likely to contain a face. Extensive experimentation on a mobile dataset of 50 users shows that our method performs better than many state-of-the-art face detection methods in terms of accuracy and processing speed.
Upal Mahbub, Vishal M. Patel, Deepak Chandra, Brandon Barbello, Rama Chellappa
ICIP5
2016 Deep feature extraction in the DCT domain
abstract
We explore the effectiveness of deep features extracted by Convolutional Neural Networks(CNNs) in the Discrete Cosine Transform(DCT) domain for various image classification tasks such as pedestrian and face detection, material identification and object recognition. We perform the DCT operation on the feature maps generated by convolutional layers in CNNs. We compare the performance of the same network on the same datasets, with the same hyper-parameters with or without the DCT step. Our results indicate that a DCT operation incorporated into the network after convolution+thresholding and before pooling can have certain advantages such as convergence over fewer training epochs and sparser weight matrices that are more conducive to pruning and hashing techniques.
Arthita Ghosh, Rama Chellappa
ICPR2
2016 On the size of Convolutional Neural Networks and generalization performance
abstract
While Convolutional Neural Networks (CNNs) have recently achieved impressive results on many classification tasks, it is still unclear why they perform so well and how to properly design them. In this work, we investigate the effect of the convolutional depth of a CNN on its generalization performance for binary classification problems. We prove a sufficient condition -polynomial in the depth of the CNN- on the training database size to guarantee such performance. We empirically test our theory on the problem of gender classification and explore the effect of varying the CNN depth, as well as the training distribution and set size.
Maya Kabkab, Emily Morgan Hand, Rama Chellappa
ICPR3
2016 Regularized metric adaptation for unconstrained face verification
abstract
In this work, we propose a metric adaptation method for set-based face verification and evaluate it on the newly released IARPA Janus Benchmark A (IJB-A) dataset and its extended version, the Janus Challenging Set 2 (CS2). A template-specific metric is trained to adaptively learn the discriminative information in test templates and the negative training set, which contains subjects that are mutually exclusive to subjects in test templates. The proposed regularized joint Bayesian metric learning framework not only alleviates the over-fitting problem but also provides a way to efficiently reduce the model size. We also analyze the selection of the compact and representative negative set to speed up the training time and to reduce storage space. Experiments on the IJB-A and CS2 datasets yield promising results.
Boyu Lu, Jun-Cheng Chen, Rama Chellappa
ICPR3
2016 Localization of skin features on the hand and wrist from small image patches
abstract
Skin-based biometrics rely on the distinctiveness of skin patterns across individuals for identification. In this paper, we investigate whether small image patches of the skin can be localized on a user's body, determining not “who?” instead “where?” Applying techniques from biometrics and computer vision, we introduce a hierarchical classifier that estimates a location from the image texture and refines the estimate with keypoint matching and geometric verification. To evaluate our approach, we collected 10,198 close-up images of 17 hand and wrist locations across 30 participants. Within-person algorithmic experiments demonstrate that an individual's own skin features can be used to localize their skin surface image patches with an F1score of 96.5%. As secondary analyses, we assess the effects of training set size and between-person classification. We close with a discussion of the strengths and limitations of our approach and evaluation methods as well as implications for future applications using a wearable camera to support touch-based, location-specific taps and gestures on the surface of the skin.
Lee Stephan Stearns, Uran Oh, Bridget J. Cheng, Leah Findlater, Rama Chellappa, Jon Froehlich
ICPR6
2016 Template regularized sparse coding for face verification
abstract
In this paper, we propose a novel regularized sparse coding approach for template-based unconstrained face verification. Unlike traditional verification tasks, which require the evaluation on image-to-image or video-to-video pairs, template-based face verification/recognition methods can exploit training and/or gallery data containing a mixture of both images or videos from the person of interest. The proposed regularized sparse coding approach addresses the adaptation to training and gallery data using three steps. First, we construct a reference dictionary, which represents the training set. Then we learn the discriminative sparse codes of the templates for verification through the proposed template regularized sparse coding approach. Finally, we measure the similarity between templates. An efficient algorithm is employed to learn the template regularized sparse codes. Extensive experiments on the template-based verification benchmark dataset show that the proposed approach outperforms several state-of-the-art methods.
Hongyu Xu, Azadeh Alavi, Rama Chellappa
ICPR4
2016 VLAD encoded Deep Convolutional features for unconstrained face verification
abstract
We present a method for combining the Vector of Locally Aggregated Descriptor (VLAD) feature encoding with Deep Convolutional Neural Network (DCNN) features for unconstrained face verification. One of the key features of our method, called the VLAD-encoded DCNN (VLAD-DCNN) features, is that spatial and appearance information are simultaneously processed to learn an improved discriminative representation. Evaluations on the challenging IARPA Janus Benchmark A (IJB-A) face dataset show that the proposed VLAD-DCNN method is able to capture the salient local features and yield promising results for face verification. Furthermore, we show that additional performance gains can be achieved by simply fusing the VLAD-DCNN features that capture the local variations with the traditional DCNN features which characterize more global features.
Jingxiao Zheng, Jun-Cheng Chen, Navaneeth Bodla, Vishal M. Patel, Rama Chellappa
ICPR5
2016 Discriminative Log-Euclidean Feature Learning for Sparse Representation-Based Recognition of Faces from Videos
Mohammed E. Fathy 0001, Azadeh Alavi, Rama Chellappa
IJCAI3
2016 Unconstrained face verification using deep CNN features
abstract
In this paper, we present an algorithm for unconstrained face verification based on deep convolutional features and evaluate it on the newly released IARPA Janus Benchmark A (IJB-A) dataset as well as on the traditional Labeled Face in the Wild (LFW) dataset. The IJB-A dataset includes real-world unconstrained faces from 500 subjects with full pose and illumination variations which are much harder than the LFW and Youtube Face (YTF) datasets. The deep convolutional neural network (DCNN) is trained using the CASIA-WebFace dataset. Results of experimental evaluations on the IJB-A and the LFW datasets are provided.
Jun-Cheng Chen, Vishal M. Patel, Rama Chellappa
WACV3
2016 Frontal to profile face verification in the wild
abstract
We have collected a new face data set that will facilitate research in the problem of frontal to profile face verification `in the wild'. The aim of this data set is to isolate the factor of pose variation in terms of extreme poses like profile, where many features are occluded, along with other `in the wild' variations. We call this data set the Celebrities in Frontal-Profile (CFP) data set. We find that human performance on Frontal-Profile verification in this data set is only slightly worse (94.57% accuracy) than that on Frontal-Frontal verification (96.24% accuracy). However we evaluated many state-of-the-art algorithms, including Fisher Vector, Sub-SML and a Deep learning algorithm. We observe that all of them degrade more than 10% from Frontal-Frontal to Frontal-Profile verification. The Deep learning implementation, which performs comparable to humans on Frontal-Frontal, performs significantly worse (84.91% accuracy) on Frontal-Profile. This suggests that there is a gap between human performance and automatic face recognition methods for large pose variation in unconstrained images.
Roni Sengupta, Jun-Cheng Chen, Carlos Domingo Castillo, Vishal M. Patel, Rama Chellappa, David Jacobs 0001
WACV5
2016 Learning a structured dictionary for video-based face recognition
abstract
In this paper, we propose a structured dictionary learning framework for video-based face recognition. We discover the invariant structural information from different videos of each subject. Specifically, we employ dictionary learning and low-rank approximation to preserve the invariant structure of face images in videos. The learned dictionary is both discriminative and reconstructive. Thus, we not only minimize the reconstruction error of all the face images but also encourage a sub-dictionary to represent the corresponding subject from different videos. Moreover, by introducing the low-rank approximation, the proposed method is able to discover invariant structured information from different videos of the same subject. To this end, an efficient alternating algorithm is employed to learn our structured dictionary. Extensive experiments on three video-based face recognition databases show that our approach outperforms several state-of-the-art methods.
Hongyu Xu, Azadeh Alavi, Rama Chellappa
WACV4
2016 R3DG features: Relative 3D geometry-based skeletal representations for human action recognition
Raviteja Vemulapalli, Felipe Arrate, Rama Chellappa
Comput. Vis. Image Underst.3
2016 Editorial of special issue on spontaneous facial behaviour analysis
Stefanos Zafeiriou, Guoying Zhao 0001, Matti Pietikäinen, Rama Chellappa, Irene Kotsia, Jeffrey F. Cohn
Comput. Vis. Image Underst.4
2016 Ray Saliency: Bottom-Up Visual Saliency for a Rotating and Zooming Camera
Garrett Warnell, Philip David, Rama Chellappa
Int. J. Comput. Vis.3
2016 The changing fortunes of pattern recognition and computer vision
Rama Chellappa
Image Vis. Comput.1
2016 Handcrafted vs. learned representations for human action recognition
Xiantong Zhen, Ling Shao 0001, Stephen J. Maybank, Rama Chellappa
Image Vis. Comput.4
2016 Face Association for Videos Using Conditional Random Fields and Max-Margin Markov Networks
abstract
We address the video-based face association problem, in which one attempts to extract the face tracks of multiple subjects while maintaining label consistency. Traditional tracking algorithms have difficulty in handling this task, especially when challenging nuisance factors like motion blur, low resolution or significant camera motions are present. We demonstrate that contextual features, in addition to face appearance itself, play an important role in this case. We propose principled methods to combine multiple features using Conditional Random Fields and Max-Margin Markov networks to infer labels for the detected faces. Different from many existing approaches, our algorithms work in online mode and hence have a wider range of applications. We address issues such as parameter learning, inference and handling false positves/negatives that arise in the proposed approach. Finally, we evaluate our approach on several public databases.
Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 Special issue on ICPR 2014 awarded papers
Rama Chellappa, Anders Heyden, Denis Laurendeau, Michael Felsberg, Magnus Borga
Pattern Recognit. Lett.1
2016 Cross-View Action Recognition via Transferable Dictionary Learning
abstract
Discriminative appearance features are effective for recognizing actions in a fixed view, but may not generalize well to a new view. In this paper, we present two effective approaches to learn dictionaries for robust action recognition across views. In the first approach, we learn a set of view-specific dictionaries where each dictionary corresponds to one camera view. These dictionaries are learned simultaneously from the sets of correspondence videos taken at different views with the aim of encouraging each video in the set to have the same sparse representation. In the second approach, we additionally learn a common dictionary shared by different views to model view-shared features. This approach represents the videos in each view using a view-specific dictionary and the common dictionary. More importantly, it encourages the set of videos taken from the different views of the same action to have the similar sparse representations. The learned common dictionary not only has the capability to represent actions from unseen views, but also makes our approach effective in a semi-supervised setting where no correspondence videos exist and only a few labeled videos exist in the target view. The extensive experiments using three public datasets demonstrate that the proposed approach outperforms recently developed approaches for cross-view action recognition.
Zhuolin Jiang, Rama Chellappa
IEEE Trans. Image Process.3
2015 Supporting Everyday Activities for Persons with Visual Impairments Through Computer Vision-Augmented Touch
abstract
The HandSight project investigates how wearable micro-cameras can be used to augment a blind or visually impaired user--s sense of touch with computer vision. Our goal is to support an array of activities of daily living by sensing and feeding back non-tactile information (e.g., color, printed text, patterns) about an object as it is touched. In this poster paper, we provide an overview of the project, our current proof-of-concept prototype, and a summary of findings from finger-based text reading studies. As this is an early-stage project, we also enumerate current open questions.
Leah Findlater, Lee Stephan Stearns, Ruofei Du, Uran Oh, Rama Chellappa, Jon Froehlich
ASSETS6
2015 Character Identification in TV-series via Non-local Cost Aggregation
abstract
We propose a non-local cost aggregation algorithm to recognize the identity of face and person tracks in a TV-series. In our approach, the fundamental element for identification is a track node, which is built on top of face and person tracks. Track nodes with temporal dependency are grouped into a knot. These knots then serve as the basic units in the construction of a k-knot graph for exploring the video structure. We build the minimum-distance spanning tree (MST) from the k-knot graph such that track nodes of similar appearance are adjacent to each other in MST. Non-local cost aggregation is performed on MST, which ensures information from face and person tracks is utilized as a whole to improve the identification performance. The identification task is performed by minimizing the cost of each knot, which takes into account the unique presence of a subject in a venue. Experimental results demonstrate the effectiveness of our method.
Ching-Hui Chen, Rama Chellappa
BMVC2
2015 Incremental Dictionary Learning for Unsupervised Domain Adaptation
abstract
Domain adaptation (DA) methods attempt to solve the domain mismatch problem between source and target data. In this paper, we propose an incremental dictionary learning method where some target data called supportive samples are selected to assist adaptation. Supportive samples are close to the source domain and have two properties: first, their predicted class labels are reliable and can be used for building more discriminative classification models; second, they act as a bridge to connect the two domains and reduce the domain mismatch. Theoretical analysis shows that both properties are important for adaptation, enabling the idea of adding supportive samples to the source domain. A stopping criterion is designed to guarantee that the domain mismatch decreases monotonically during adaptation. Experimental results on several widely used visual datasets show that the proposed approach performs better than many state-of-the-art methods.
Boyu Lu, Rama Chellappa, Nasser M. Nasrabadi
BMVC2
2015 Bridging the Domain Shift by Domain Adaptive Dictionary Learning
abstract
Domain adaptation (DA) tackles the problem where data from the training set (source domain) and test set (target domain) have different underlying distributions. For instance, training and testing images may be acquired under different environments, viewpoints and illumination conditions. In this paper, we focus on the more challenging unsupervised DA problem where the samples in the target domain are unlabeled. It is noticed that dictionary learning has gained a lot of popularity due to the fact that images of interest could be reconstructed sparsely in an appropriately learned dictionary [1]. Specifically, we propose a novel domain-adaptive dictionary learning approach to generate a set of intermediate domains which bridge the gap between source and target domains. Our approach defines two types of dictionaries: a common dictionary and a domain-specific dictionary. The overall learning process illustrated in Figure 1 consists of three steps: (1) At the beginning, we first learn the common dictionary DC, domainspecific dictionaries D0 and Dt for source and target domains. (2) At the k-th step, we enforce the recovered feature representations of target data in all available domains to have the same sparse codes, while adapting the most recently obtained dictionary Dk to better represent the target domain. Then we multiply dictionaries in the k-th domain with the corresponding sparse codes to recover feature representations of target data Xk t in this domain. (3) We update Dk to find the next domain-specific dictionary Dk+1 by further minimizing the reconstruction error in representing the target data. Then we alternate between the steps of sparse coding and dictionary updating until the stopping criteria is satisfied. Notations: Let X s ∈ Rd×Ns , X t ∈ Rd×Nt be the feature representations of source and target data respectively, where d is the feature dimension, Ns and Nt are the number of samples in the two domains. The feature representations of recovered source and target data in the k-th intermediate domain are denoted as Xs ∈ Rd×Ns and Xk t ∈ Rd×Nt respectively. The common dictionary is denoted as DC, whereas source-specific and targetspecific dictionaries are denoted as D0, Dt respectively. Similarly, we use Dk,k = 1...N to denote the domain-specific dictionary for the k-th domain, where N is the number of intermediate domains. We set all the dictionaries to be of the same size ∈Rd×n. At the beginning, we learn the common dictionary DC by minimizing the reconstruction error of both source and target data as follows:
Hongyu Xu, Rama Chellappa
BMVC3
2015 Matrix completion for resolving label ambiguity
abstract
In real applications, data is not always explicitly-labeled. For instance, label ambiguity exists when we associate two persons appearing in a news photo with two names provided in the caption. We propose a matrix completion-based method for predicting the actual labels from the ambiguously labeled instances, and a standard supervised classifier can learn from the disambiguated labels to classify new data. We further generalize the method to handle the labeling constraints between instances when such prior knowledge is available. Compared to existing methods, our approach achieves 2.9% improvement on the labeling accuracy of the Lost dataset and comparable performance on the Labeled Yahoo! News dataset.
Ching-Hui Chen, Vishal M. Patel, Rama Chellappa
CVPR3
2015 Class consistent multi-modal fusion with binary features
abstract
Many existing recognition algorithms combine different modalities based on training accuracy but do not consider the possibility of noise at test time. We describe an algorithm that perturbs test features so that all modalities predict the same class. We enforce this perturbation to be as small as possible via a quadratic program (QP) for continuous features, and a mixed integer program (MIP) for binary features. To efficiently solve the MIP, we provide a greedy algorithm and empirically show that its solution is very close to that of a state-of-the-art MIP solver. We evaluate our algorithm on several datasets and show that the method outperforms existing approaches.
Ashish Shrivastava 0001, Mohammad Rastegari, Rama Chellappa, Larry Davis 0001
CVPR4
2015 Face-based Active Authentication on mobile devices
abstract
As mobile devices are becoming more ubiquitous, it becomes important to continuously verify the identity of the user during all interactions rather than just at login time. This paper investigates the effectiveness of methods for fully-automatic face recognition in solving the Active Authentication (AA) problem for smartphones. We report the results of face authentication using videos recorded by the front camera. The videos were acquired while the users were performing a number of tasks under three different ambient conditions to capture the type of variations caused by the 'mobility' of the devices. An inspection of these videos reveal a combination of favorable and challenging properties unique to smartphone face videos. In addition to variations caused by the mobility of the device, other challenges in the dataset include occlusion, occasional pose changes, blur and face/fiducial points localization errors. We evaluate still image and image set-based authentication algorithms using intensity features extracted around fiducial points. The recognition rates drop dramatically when enrollment and test videos come from different sessions. We will make the dataset and the computed features publicly available1to help the design of algorithms that are more robust to variations due to factors mentioned above.
Mohammed E. Fathy 0001, Vishal M. Patel, Rama Chellappa
ICASSP3
2015 Landmark-based fisher vector representation for video-based face verification
abstract
Unconstrained video-based face verification is a challenging problem because of dramatic variations in pose, illumination, and image quality of each face in a video. In this paper, we propose a landmark-based Fisher vector representation for video-to-video face verification. The proposed representation encodes dense multi-scale SIFT features extracted from patches centered at detected facial landmarks, and face similarity is computed with the distance measure learned from joint Bayesian metric learning. Experimental results demonstrate that our approach achieves significantly better performance than other competitive video-based face verification algorithms on two challenging unconstrained video face dataseis, Multiple Biometric Grand Challenge (MBGC) and Face and Ocular Challenge Series (FOCS).
Jun-Cheng Chen, Vishal M. Patel, Rama Chellappa
ICIP3
2015 Material classification and semantic segmentation of railway track images with deep convolutional neural networks
abstract
The condition of railway tracks needs to be periodically monitored to ensure passenger safety. Cameras mounted on a moving vehicle such as a hi-rail vehicle or a geometry inspection car can generate large volumes of high resolution images. Extracting accurate information from those images has been challenging due to background clutter in railroad environments. In this paper, we describe a novel approach to visual track inspection using material classification and semantic segmentation with Deep Convolutional Neural Networks (DCNN). We show that DCNNs trained end-to-end for material classification are more accurate than shallow learning machines with hand-engineered features and are more robust to noise. Our approach results in a material classification accuracy of 93.35% using 10 classes of materials. This allows for the detection of crumbling and chipped tie conditions at detection rates of 86.06% and 92.11%, respectively, at a false positive rate of 10 FP/mile on the 85-mile Northeast Corridor (NEC) 2012–2013 concrete tie dataset.
Xavier Giben, Vishal M. Patel, Rama Chellappa
ICIP3
2015 3D facial model synthesis using coupled dictionaries
abstract
In this work, we propose a generative way of modeling faces, where the 3D shape of a face is generated by a supervised learning procedure involving coupled sparse feature learning. To learn dictionaries using the proposed method, we use the USF-HUMAN ID database [1]. We provide as input to our training system, paired correspondences of 2D and 3D images of individuals and aim to learn the low-level patches both in 2D and 3D domains that describe the corresponding subspaces in a sparse manner. We demonstrate the efficacy of our method by quantitative results on the 3D database and qualitative results on images drawn from the internet.
Swami Sankaranarayanan, Vishal M. Patel, Rama Chellappa
ICIP3
2015 Integrability-regularized phase unwrapping via sparse error correction
abstract
We propose a new formulation of the classical two-dimensional phase unwrapping problem. Using a sparse-error, gradient-domain measurement model, we simultaneously seek the absolute phase and sparse gradient errors that minimize a novel energy functional that strongly encourages the integrability of the corrected gradient field. Our approach can be cast as a generalized lasso problem, and we compute the solution using the alternating direction method of multipliers (ADMM) algorithm. Adopting a commonly-used inter-ferometric synthetic aperture radar noise model, we evaluate our technique for several synthetic surfaces.
Garrett Warnell, Vishal M. Patel, Rama Chellappa
ICIP3
2015 Multitask multivariate common sparse representations for robust multimodal biometrics recognition
abstract
In this paper, we propose multitask multivairate common sparse representations for robust multimodal biometrics recognition. The proposed approach can be viewed as an extension of previous work on joint sparse representations for robust multimodal biometrics recognition. The proposed algorithm utilizes the discriminative information among different modalities simultaneously by enforcing the common sparse representation across all the modalities and achieves more robust multimodal recognition especially when all modalities are noisy and “weak”. Alternating direction method of multipliers is proposed to solve the resulting optimization problem. Experiments on two biometric datasets show that our method performs better than the state-of-the-art fusion methods.
Heng Zhang 0003, Vishal M. Patel, Rama Chellappa
ICIP3
2015 Robust Fastener Detection for Autonomous Visual Railway Track Inspection
abstract
Fasteners are critical railway components that maintain the rails in a fixed position. Their failure can lead to train derailments due to gage widening or wheel climb, so their condition needs to be periodically monitored. Several computer vision methods have been proposed in the literature for track inspection applications. However, these methods are not robust to clutter and background noise present in the railroad environment. This paper proposes a new method for fastener detection by 1) carefully aligning the training data, 2) reducing intra-class variation, and 3) bootstrapping difficult samples to improve the classification margin. Using the histogram of oriented gradients features and a combination of linear SVM classifiers, the system described in this paper can inspect ties for missing or defective rail fastener problems with a probability of detection of 98% and a false alarm rate of 1.23% on a new dataset of 85 miles of concrete tie images collected in the US Northeast Corridor (NEC) between 2012 and 2013. To the best of our knowledge, this dataset of 203,287 crossties is the largest ever reported in the literature.
Xavier Gibert, Vishal M. Patel, Rama Chellappa
WACV3
2015 Touch Gesture-Based Active User Authentication Using Dictionaries
abstract
Screen touch gesture has been shown to be a promising modality for touch-based active authentication of users of mobile devices. In this paper, we present an approach for active user authentication using screen touch gestures by building linear and kernelized dictionaries based on sparse representations and associated classifiers. Experiments using a new dataset collected by us as well as two other publicly available screen touch datasets show that the dictionary-based classification method compares favorably to those published in the literature. Experiments done using data collected in three different sessions corresponding to different environmental conditions show a drop in performance when the training and test data come from different sessions. This suggests a need for applying domain adaptation methods to further improve the performance of the classifiers.
Heng Zhang 0003, Vishal M. Patel, Mohammed E. Fathy 0001, Rama Chellappa
WACV4
2015 Generalized Dictionaries for Multiple Instance Learning
Ashish Shrivastava 0001, Vishal M. Patel, Jaishanker K. Pillai, Rama Chellappa
Int. J. Comput. Vis.4
2015 Fast detection of facial wrinkles based on Gabor features using image morphology and geometric constraints
Nazre Batool, Rama Chellappa
Pattern Recognit.2
2015 Salient views and view-dependent dictionaries for object recognition
Vishal M. Patel, Rama Chellappa, P. Jonathon Phillips
Pattern Recognit.3
2015 Non-linear dictionary learning with partially labeled data
Ashish Shrivastava 0001, Vishal M. Patel, Rama Chellappa
Pattern Recognit.3
2015 Joint Sparse Representation and Robust Feature-Level Fusion for Multi-Cue Visual Tracking
abstract
Visual tracking using multiple features has been proved as a robust approach because features could complement each other. Since different types of variations such as illumination, occlusion, and pose may occur in a video sequence, especially long sequence videos, how to properly select and fuse appropriate features has become one of the key problems in this approach. To address this issue, this paper proposes a new joint sparse representation model for robust feature-level fusion. The proposed method dynamically removes unreliable features to be fused for tracking by using the advantages of sparse representation. In order to capture the non-linear similarity of features, we extend the proposed method into a general kernelized framework, which is able to perform feature fusion on various kernel spaces. As a result, robust tracking performance is obtained. Both the qualitative and quantitative experimental results on publicly available videos show that the proposed method outperforms both sparse representation-based and fusion based-trackers.
Xiangyuan Lan, Andy Jinhua Ma, Pong C. Yuen, Rama Chellappa
IEEE Trans. Image Process.4
2015 DASH-N: Joint Hierarchical Domain Adaptation and Feature Learning
abstract
Complex visual data contain discriminative structures that are difficult to be fully captured by any single feature descriptor. While recent work on domain adaptation focuses on adapting a single hand-crafted feature, it is important to perform adaptation of a hierarchy of features to exploit the richness of visual data. We propose a novel framework for domain adaptation using a sparse and hierarchical network (DASH-N). Our method jointly learns a hierarchy of features together with transformations that rectify the mismatch between different domains. The building block of DASH-N is the latent sparse representation. It employs a dimensionality reduction step that can prevent the data dimension from increasing too fast as one traverses deeper into the hierarchy. The experimental results show that our method compares favorably with the competing state-of-the-art methods. In addition, it is shown that a multi-layer DASH-N performs better than a single-layer DASH-N.
Hien Van Nguyen, Huy Tho Ho, Vishal M. Patel, Rama Chellappa
IEEE Trans. Image Process.4
2015 Compositional Dictionaries for Domain Adaptive Face Recognition
abstract
We present a dictionary learning approach to compensate for the transformation of faces due to the changes in view point, illumination, resolution, and so on. The key idea of our approach is to force domain-invariant sparse coding, i.e., designing a consistent sparse representation of the same face in different domains. In this way, the classifiers trained on the sparse codes in the source domain consisting of frontal faces can be applied to the target domain (consisting of faces in different poses, illumination conditions, and so on) without much loss in recognition accuracy. The approach is to first learn a domain base dictionary, and then describe each domain shift (identity, pose, and illumination) using a sparse representation over the base dictionary. The dictionary adapted to each domain is expressed as the sparse linear combinations of the base dictionary. In the context of face recognition, with the proposed compositional dictionary approach, a face image can be decomposed into sparse representations for a given subject, pose, and illumination. This approach has three advantages. First, the extracted sparse representation for a subject is consistent across domains, and enables pose and illumination insensitive face recognition. Second, sparse representations for pose and illumination can be subsequently used to estimate the pose and illumination condition of a face image. Last, by composing sparse representations for the subject and the different domains, we can also perform pose alignment and illumination normalization. Extensive experiments using two public face data sets are presented to demonstrate the effectiveness of the proposed approach for face recognition.
Qiang Qiu 0002, Rama Chellappa
IEEE Trans. Image Process.2
2015 Coupled Projections for Adaptation of Dictionaries
abstract
Data-driven dictionaries have produced the state-of-the-art results in various classification tasks. However, when the target data has a different distribution than the source data, the learned sparse representation may not be optimal. In this paper, we investigate if it is possible to optimally represent both source and target by a common dictionary. In particular, we describe a technique which jointly learns projections of data in the two domains, and a latent dictionary which can succinctly represent both the domains in the projected low-dimensional space. The algorithm is modified to learn a common discriminative dictionary, which can further improve the classification performance. The algorithm is also effective for adaptation across multiple domains and is extensible to nonlinear feature spaces. The proposed approach does not require any explicit correspondences between the source and target domains, and yields good results even when there are only a few labels available in the target domain. We also extend it to unsupervised adaptation in cases where the same feature is extracted across all domains. Further, it can also be used for heterogeneous domain adaptation, where different features are extracted for different domains. Various recognition experiments show that the proposed method performs on par or better than competitive state-of-the-art methods.
Vishal M. Patel, Hien Van Nguyen, Rama Chellappa
IEEE Trans. Image Process.4
2015 Adaptive-Rate Compressive Sensing Using Side Information
abstract
We provide two novel adaptive-rate compressive sensing (CS) strategies for sparse, time-varying signals using side information. The first method uses extra cross-validation measurements, and the second one exploits extra low-resolution measurements. Unlike the majority of current CS techniques, we do not assume that we know an upper bound on the number of significant coefficients that comprises the images in the video sequence. Instead, we use the side information to predict the number of significant coefficients in the signal at the next time instant. We develop our techniques in the specific context of background subtraction using a spatially multiplexing CS camera such as the single-pixel camera. For each image in the video sequence, the proposed techniques specify a fixed number of CS measurements to acquire and adjust this quantity from image to image. We experimentally validate the proposed methods on real surveillance video sequences.
Garrett Warnell, Sourabh Bhattacharya, Rama Chellappa, Tamer Basar
IEEE Trans. Image Process.3
2014 Video-Based Face Recognition Using the Intra-Personal/Extra-Personal Difference Dictionary
Rama Chellappa
BMVC2
2014 Human Action Recognition by Representing 3D Skeletons as Points in a Lie Group
abstract
Recently introduced cost-effective depth sensors coupled with the real-time skeleton estimation algorithm of Shotton et al. [16] have generated a renewed interest in skeleton-based human action recognition. Most of the existing skeleton-based approaches use either the joint locations or the joint angles to represent a human skeleton. In this paper, we propose a new skeletal representation that explicitly models the 3D geometric relationships between various body parts using rotations and translations in 3D space. Since 3D rigid body motions are members of the special Euclidean group SE(3), the proposed skeletal representation lies in the Lie group SE(3)×.. .×SE(3), which is a curved manifold. Using the proposed representation, human actions can be modeled as curves in this Lie group. Since classification of curves in this Lie group is not an easy task, we map the action curves from the Lie group to its Lie algebra, which is a vector space. We then perform classification using a combination of dynamic time warping, Fourier temporal pyramid representation and linear SVM. Experimental results on three action datasets show that the proposed representation performs better than many existing skeletal representations. The proposed approach also outperforms various state-of-the-art skeleton-based human action recognition approaches.
Raviteja Vemulapalli, Felipe Arrate, Rama Chellappa
CVPR3
2014 Growing Regression Forests by Classification: Applications to Object Pose Estimation
Kota Hara, Rama Chellappa
ECCV (2)2
2014 Sparse localized facial motion dictionary learning for facial expression recognition
abstract
This paper presents a new framework for facial motion modeling with applications to facial expression recognition. First, we design sparse localized facial motion dictionaries from dense motion flow data of facial expression image sequences. Regularization based on spatial localized support map in addition to the sparsity constraints enables spatially localized dictionary learning. Proposed localized dictionaries are effective for local facial motion description as well as global facial motion analysis. Experimental results using CK+ database shows promising results for automatic facial expression recognition from motion flow data.
Chan-Su Lee, Rama Chellappa
ICASSP2
2014 Dictionary-based video face recognition using dense multi-scale facial landmark features
abstract
In video-based face recognition, different video sequences of the same subject contain variations in pose, illumination, and expression which contribute to the challenges in designing an effective video-based face-recognition system. In this paper, we propose a dictionary-based approach using dense and high-dimensional features extracted from multi-scale patches centered at detected facial landmarks for video-to-video face identification and verification. Experiments using unconstrained video sequences from Multiple Biometric Grand Challenge (MBGC) and Face and Ocular Challenge Series (FOCS) datasets show that our method performs significantly better than many state-of-the-art video-based face recognition algorithms.
Jun-Cheng Chen, Vishal M. Patel, Huy Tho Ho, Rama Chellappa
ICIP4
2014 Coupled dictionaries for thermal to visible face recognition
abstract
Thermal to visible face recognition is the problem of identifying a thermal infrared (IR) face image given a gallery of visible light face images. We attempt to solve this problem by learning coupled dictionaries to represent the two domains. The dictionaries provide a sparse representation which transforms the data into a single, domain-independent, latent space. We formulate the dictionary learning problem as a bi-level optimization problem and perform a stochastic gradient descent on the dictionaries to solve it. We present experimental results demonstrating the effectiveness of our approach.
Christopher Reale, Nasser M. Nasrabadi, Rama Chellappa
ICIP3
2014 Analysis sparse coding models for image-based classification
abstract
Data-driven sparse models have been shown to give superior performance for image classification tasks. Most of these works depend on learning a synthesis dictionary and the corresponding sparse code for recognition. However in recent years, an alternate analysis coding based framework (also known as co-sparse model) has been proposed for learning sparse models. In this paper, we study this framework for image classification. We demonstrate that the proposed approach is robust and efficient, while giving a comparable or better recognition performance than the traditional synthesis-based models.
Vishal M. Patel, Rama Chellappa
ICIP3
2014 Dictionary-based multiple instance learning
abstract
We present a multi-class, multiple instance learning (MIL) algorithm using the dictionary learning framework where the data is given in the form of bags. Each bag contains multiple samples, called instances, out of which at least one belongs to the class of the bag. We propose a noisy-OR model-based optimization framework for learning the dictionaries. Our method can be viewed as a generalized dictionary learning algorithm since it reduces to a novel discriminative dictionary learning framework when there is only one instance in each bag. Various experiments using the popular MIL datasets show that the proposed method performs better than existing methods.
Ashish Shrivastava 0001, Jaishanker K. Pillai, Vishal M. Patel, Rama Chellappa
ICIP4
2014 Toward a non-intrusive, physio- behavioral biometric for smartphones
abstract
Biometric authentication relies on an individual's inner characteristics and traits. We propose an active authentication system on a mobile device that relies on two biometric modalities: 3D gestures and face recognition. The novelty of our approach is to combine 3D gesture and face recognition in a nonintrusive and unconstrained environment; the active authentication system is running in the background while the user is performing his/her main task.
Esther Vasiete, Yan Chen 0033, Ian Char, Tom Yeh, Vishal M. Patel, Larry Davis 0001, Rama Chellappa
Mobile HCI7
2014 Submodular Attribute Selection for Action Recognition in Video
Zhuolin Jiang, Rama Chellappa, P. Jonathon Phillips
NIPS3
2014 Adaptive representations for video-based face recognition across pose
abstract
In this paper, we address the problem of matching faces across changes in pose in unconstrained videos. We propose two methods based on 3D rotation and sparse representation that compensate for changes in pose. The first is Sparse Representation-based Alignment (SRA) that generates pose aligned features under a sparsity constraint. The mapping for the pose aligned features are learned from a reference set of face images which is independent of the videos used in the experiment. Thus, they generalize across data sets. The second is a Dictionary Rotation (DR) method that directly rotates video dictionary atoms in both their harmonic basis and 3D geometry to match the poses of the probe videos. We demonstrate the effectiveness of our approach over several state-of-the-art algorithms through extensive experiments on three challenging unconstrained video datasets: the video challenge of the Face and Ocular Challenge Series (FOCS), the Multiple Biometrics Grand Challenge (MBGC), and the Human ID datasets.
Vishal M. Patel, Rama Chellappa, P. Jonathon Phillips
WACV3
2014 Guest Editor's Introduction to the Special Issue on Domain Adaptation for Vision Applications
Dong Xu 0001, Rama Chellappa, Trevor Darrell, Hal Daumé III
Int. J. Comput. Vis.2
2014 Entropy-Rate Clustering: Cluster Analysis via Maximizing a Submodular Function Subject to a Matroid Constraint
abstract
We propose a new objective function for clustering. This objective function consists of two components: the entropy rate of a random walk on a graph and a balancing term. The entropy rate favors formation of compact and homogeneous clusters, while the balancing function encourages clusters with similar sizes and penalizes larger clusters that aggressively group samples. We present a novel graph construction for the graph associated with the data and show that this construction induces a matroid--a combinatorial structure that generalizes the concept of linear independence in vector spaces. The clustering result is given by the graph topology that maximizes the objective function under the matroid constraint. By exploiting the submodular and monotonic properties of the objective function, we develop an efficient greedy algorithm. Furthermore, we prove an approximation bound of (1/2) for the optimality of the greedy solution. We validate the proposed algorithm on various benchmarks and show its competitive performances with respect to popular clustering algorithms. We further apply it for the task of superpixel segmentation. Experiments on the Berkeley segmentation data set reveal its superior performances over the state-of-the-art superpixel segmentation algorithms in all the standard evaluation metrics.
Ming-Yu Liu 0001, Oncel Tuzel, Srikumar Ramalingam, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.4
2014 Unsupervised Adaptation Across Domain Shifts by Generating Intermediate Data Representations
abstract
With unconstrained data acquisition scenarios widely prevalent, the ability to handle changes in data distribution across training and testing data sets becomes important. One way to approach this problem is through domain adaptation, and in this paper we primarily focus on the unsupervised scenario where the labeled source domain training data is accompanied by unlabeled target domain test data. We present a two-stage data-driven approach by generating intermediate data representations that could provide relevant information on the domain shift. Starting with a linear representation of domains in the form of generative subspaces of same dimensions for the source and target domains, we first utilize the underlying geometry of the space of these subspaces, the Grassmann manifold, to obtain a `shortest' geodesic path between the two domains. We then sample points along the geodesic to obtain intermediate cross-domain data representations, using which a discriminative classifier is learnt to estimate the labels of the target data. We subsequently incorporate non-linear representation of domains by considering a Reproducing Kernel Hilbert Space representation, and a low-dimensional manifold representation using Laplacian Eigenmaps, and also examine other domain adaptation settings such as (i) semi-supervised adaptation where the target domain is partially labeled, and (ii) multi-domain adaptation where there could be more than one domain in source and/or target data sets. Finally, we supplement our adaptation technique with (i) fine-grained reference domains that are created by blending samples from source and target data sets to provide some evidence on the actual domain shift, and (ii) a multi-class boosting analysis to obtain robustness to the choice of algorithm parameters. We evaluate our approach for object recognition problems and report competitive results on two widely used Office and Bing adaptation data sets.
Raghuraman Gopalan, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Cross-Sensor Iris Recognition through Kernel Learning
abstract
Due to the increasing popularity of iris biometrics, new sensors are being developed for acquiring iris images and existing ones are being continuously upgraded. Re-enrolling users every time a new sensor is deployed is expensive and time-consuming, especially in applications with a large number of enrolled users. However, recent studies show that cross-sensor matching, where the test samples are verified using data enrolled with a different sensor, often lead to reduced performance. In this paper, we propose a machine learning technique to mitigate the cross-sensor performance degradation by adapting the iris samples from one sensor to another. We first present a novel optimization framework for learning transformations on iris biometrics. We then utilize this framework for sensor adaptation, by reducing the distance between samples of the same class, and increasing it between samples of different classes, irrespective of the sensors acquiring them. Extensive evaluations on iris data from multiple sensors demonstrate that the proposed method leads to improvement in cross-sensor recognition accuracy. Furthermore, since the proposed technique requires minimal changes to the iris recognition pipeline, it can easily be incorporated into existing iris recognition systems.
Jaishanker K. Pillai, Maria Puertas, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Information-Theoretic Dictionary Learning for Image Classification
abstract
We present a two-stage approach for learning dictionaries for object classification tasks based on the principle of information maximization. The proposed method seeks a dictionary that is compact, discriminative, and generative. In the first stage, dictionary atoms are selected from an initial dictionary by maximizing the mutual information measure on dictionary compactness, discrimination and reconstruction. In the second stage, the selected dictionary atoms are updated for improved reconstructive and discriminative power using a simple gradient ascent algorithm on mutual information. Experiments using real data sets demonstrate the effectiveness of our approach for image classification tasks.
Qiang Qiu 0002, Vishal M. Patel, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Joint Sparse Representation for Robust Multimodal Biometrics Recognition
abstract
Traditional biometric recognition systems rely on a single biometric signature for authentication. While the advantage of using multiple sources of information for establishing the identity has been widely recognized, computational models for multimodal biometrics recognition have only recently received attention. We propose a multimodal sparse representation method, which represents the test data by a sparse linear combination of training data, while constraining the observations from different modalities of the test subject to share their sparse representations. Thus, we simultaneously take into account correlations as well as coupling information among biometric modalities. A multimodal quality measure is also proposed to weigh each modality as it gets fused. Furthermore, we also kernelize the algorithm to handle nonlinearity in data. The optimization problem is solved using an efficient alternative direction method. Various experiments show that the proposed method compares favorably with competing fusion-based methods.
Vishal M. Patel, Nasser M. Nasrabadi, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.4
2014 Screen-based active user authentication
Mohammed E. Fathy 0001, Vishal M. Patel, Tom Yeh, Yangmuzi Zhang, Rama Chellappa, Larry Davis 0001
Pattern Recognit. Lett.5
2014 Differential geometric representations and algorithms for some pattern recognition and computer vision problems
Pavan Turaga, Anuj Srivastava, Rama Chellappa
Pattern Recognit. Lett.4
2014 Separated Component-Based Restoration of Speckled SAR Images
abstract
Many coherent imaging modalities such as synthetic aperture radar suffer from a multiplicative noise, commonly referred to as speckle, which often makes the interpretation of data difficult. An effective strategy for speckle reduction is to use a dictionary that can sparsely represent the features in the speckled image. However, such approaches fail to capture important salient features such as texture. In this paper, we present a speckle reduction algorithm that handles this issue by formulating the restoration problem so that the structure and texture components can be separately estimated with different dictionaries. To solve this formulation, an iterative algorithm based on surrogate functionals is proposed. Experiments indicate the proposed method performs favorably compared to state-of-the-art speckle reduction methods.
Vishal M. Patel, Glenn R. Easley, Rama Chellappa, Nasser M. Nasrabadi
IEEE Trans. Geosci. Remote. Sens.3
2014 Ambiguously Labeled Learning Using Dictionaries
abstract
We propose a dictionary-based learning method for ambiguously labeled multiclass classification, where each training sample has multiple labels and only one of them is the correct label. The dictionary learning problem is solved using an iterative alternating algorithm. At each iteration of the algorithm, two alternating steps are performed: 1) a confidence update and 2) a dictionary update. The confidence of each sample is defined as the probability distribution on its ambiguous labels. The dictionaries are updated using either soft or hard decision rules. Furthermore, using the kernel methods, we make the dictionary learning framework nonlinear based on the soft decision rule. Extensive evaluations on four unconstrained face recognition datasets demonstrate that the proposed method performs significantly better than state-of-the-art ambiguously labeled learning approaches.
Vishal M. Patel, Rama Chellappa, P. Jonathon Phillips
IEEE Trans. Inf. Forensics Secur.3
2014 Detection and Inpainting of Facial Wrinkles Using Texture Orientation Fields and Markov Random Field Modeling
abstract
Facial retouching is widely used in media and entertainment industry. Professional software usually require a minimum level of user expertise to achieve the desirable results. In this paper, we present an algorithm to detect facial wrinkles/imperfection. We believe that any such algorithm would be amenable to facial retouching applications. The detection of wrinkles/imperfections can allow these skin features to be processed differently than the surrounding skin without much user interaction. For detection, Gabor filter responses along with texture orientation field are used as image features. A bimodal Gaussian mixture model (GMM) represents distributions of Gabor features of normal skin versus skin imperfections. Then, a Markov random field model is used to incorporate the spatial relationships among neighboring pixels for their GMM distributions and texture orientations. An expectation-maximization algorithm then classifies skin versus skin wrinkles/imperfections. Once detected automatically, wrinkles/imperfections are removed completely instead of being blended or blurred. We propose an exemplar-based constrained texture synthesis algorithm to inpaint irregularly shaped gaps left by the removal of detected wrinkles/imperfections. We present results conducted on images downloaded from the Internet to show the efficacy of our algorithms.
Nazre Batool, Rama Chellappa
IEEE Trans. Image Process.2
2014 Robust Face Recognition From Multi-View Videos
abstract
Multiview face recognition has become an active research area in the last few years. In this paper, we present an approach for video-based face recognition in camera networks. Our goal is to handle pose variations by exploiting the redundancy in the multiview video data. However, unlike traditional approaches that explicitly estimate the pose of the face, we propose a novel feature for robust face recognition in the presence of diffuse lighting and pose variations. The proposed feature is developed using the spherical harmonic representation of the face texture-mapped onto a sphere; the texture map itself is generated by back-projecting the multiview video data. Video plays an important role in this scenario. First, it provides an automatic and efficient way for feature extraction. Second, the data redundancy renders the recognition algorithm more robust. We measure the similarity between feature sets from different videos using the reproducing kernel Hilbert space. We demonstrate that the proposed approach outperforms traditional algorithms on a multiview video database.
Aswin C. Sankaranarayanan, Rama Chellappa
IEEE Trans. Image Process.3
2014 Multiple Kernel Learning for Sparse Representation-Based Classification
abstract
In this paper, we propose a multiple kernel learning (MKL) algorithm that is based on the sparse representation-based classification (SRC) method. Taking advantage of the nonlinear kernel SRC in efficiently representing the nonlinearities in the high-dimensional feature space, we propose an MKL method based on the kernel alignment criteria. Our method uses a two step training method to learn the kernel weights and sparse codes. At each iteration, the sparse codes are updated first while fixing the kernel mixing coefficients, and then the kernel mixing coefficients are updated while fixing the sparse codes. These two steps are repeated until a stopping criteria is met. The effectiveness of the proposed method is demonstrated using several publicly available image classification databases and it is shown that this method can perform significantly better than many competitive image classification algorithms.
Ashish Shrivastava 0001, Vishal M. Patel, Rama Chellappa
IEEE Trans. Image Process.3
2014 Structure-Preserving Sparse Decomposition for Facial Expression Analysis
abstract
Although facial expressions can be decomposed in terms of action units (AUs) as suggested by the facial action coding system, there have been only a few attempts that recognize expression using AUs and their composition rules. In this paper, we propose a dictionary-based approach for facial expression analysis by decomposing expressions in terms of AUs. First, we construct an AU-dictionary using domain experts' knowledge of AUs. To incorporate the high-level knowledge regarding expression decomposition and AUs, we then perform structure-preserving sparse coding by imposing two layers of grouping over AU-dictionary atoms as well as over the test image matrix columns. We use the computed sparse code matrix for each expressive face to perform expression decomposition and recognition. Since domain experts' knowledge may not always be available for constructing an AU-dictionary, we also propose a structure-preserving dictionary learning algorithm, which we use to learn a structured dictionary as well as divide expressive faces into several semantic regions. Experimental results on publicly available expression data sets demonstrate the effectiveness of the proposed approach for facial expression analysis.
Sima Taheri, Qiang Qiu 0002, Rama Chellappa
IEEE Trans. Image Process.3
2013 Dictionary Learning from Ambiguously Labeled Data
abstract
We propose a novel dictionary-based learning method for ambiguously labeled multiclass classification, where each training sample has multiple labels and only one of them is the correct label. The dictionary learning problem is solved using an iterative alternating algorithm. At each iteration of the algorithm, two alternating steps are performed: a confidence update and a dictionary update. The confidence of each sample is defined as the probability distribution on its ambiguous labels. The dictionaries are updated using either soft (EM-based) or hard decision rules. Extensive evaluations on existing datasets demonstrate that the proposed method performs significantly better than state-of-the-art ambiguously labeled learning approaches.
Vishal M. Patel, Jaishanker K. Pillai, Rama Chellappa, P. Jonathon Phillips
CVPR4
2013 Computationally Efficient Regression on a Dependency Graph for Human Pose Estimation
abstract
We present a hierarchical method for human pose estimation from a single still image. In our approach, a dependency graph representing relationships between reference points such as body joints is constructed and the positions of these reference points are sequentially estimated by a successive application of multidimensional output regressions along the dependency paths, starting from the root node. Each regressor takes image features computed from an image patch centered on the current node's position estimated by the previous regressor and is specialized for estimating its child nodes' positions. The use of the dependency graph allows us to decompose a complex pose estimation problem into a set of local pose estimation problems that are less complex. We design a dependency graph for two commonly used human pose estimation datasets, the Buffy Stickmen dataset and the ETHZ PASCAL Stickmen dataset, and demonstrate that our method achieves comparable accuracy to state-of-the-art results on both datasets with significantly lower computation time than existing methods. Furthermore, we propose an importance weighted boosted regression trees method for transductive learning settings and demonstrate the resulting improved performance for pose estimation tasks.
Kota Hara, Rama Chellappa
CVPR2
2013 Subspace Interpolation via Dictionary Learning for Unsupervised Domain Adaptation
abstract
Domain adaptation addresses the problem where data instances of a source domain have different distributions from that of a target domain, which occurs frequently in many real life scenarios. This work focuses on unsupervised domain adaptation, where labeled data are only available in the source domain. We propose to interpolate subspaces through dictionary learning to link the source and target domains. These subspaces are able to capture the intrinsic domain shift and form a shared feature representation for cross domain recognition. Further, we introduce a quantitative measure to characterize the shift between two domains, which enables us to select the optimal domain to adapt to the given multiple source domains. We present experiments on face recognition across pose, illumination and blur variations, cross dataset object recognition, and report improved performance over the state of the art.
Jie Ni, Qiang Qiu 0002, Rama Chellappa
CVPR3
2013 Generalized Domain-Adaptive Dictionaries
abstract
Data-driven dictionaries have produced state-of-the-art results in various classification tasks. However, when the target data has a different distribution than the source data, the learned sparse representation may not be optimal. In this paper, we investigate if it is possible to optimally represent both source and target by a common dictionary. Specifically, we describe a technique which jointly learns projections of data in the two domains, and a latent dictionary which can succinctly represent both the domains in the projected low-dimensional space. An efficient optimization technique is presented, which can be easily kernelized and extended to multiple domains. The algorithm is modified to learn a common discriminative dictionary, which can be further used for classification. The proposed approach does not require any explicit correspondence between the source and target domains, and shows good results even when there are only a few labels available in the target domain. Various recognition experiments show that the method performs on par or better than competitive state-of-the-art methods.
Vishal M. Patel, Hien Van Nguyen, Rama Chellappa
CVPR4
2013 Kernel Learning for Extrinsic Classification of Manifold Features
abstract
In computer vision applications, features often lie on Riemannian manifolds with known geometry. Popular learning algorithms such as discriminant analysis, partial least squares, support vector machines, etc., are not directly applicable to such features due to the non-Euclidean nature of the underlying spaces. Hence, classification is often performed in an extrinsic manner by mapping the manifolds to Euclidean spaces using kernels. However, for kernel based approaches, poor choice of kernel often results in reduced performance. In this paper, we address the issue of kernel selection for the classification of features that lie on Riemannian manifolds using the kernel learning approach. We propose two criteria for jointly learning the kernel and the classifier using a single optimization problem. Specifically, for the SVM classifier, we formulate the problem of learning a good kernel-classifier combination as a convex optimization problem and solve it efficiently following the multiple kernel learning approach. Experimental results on image set-based classification and activity recognition clearly demonstrate the superiority of the proposed approach over existing methods for classification of manifold features.
Raviteja Vemulapalli, Jaishanker K. Pillai, Rama Chellappa
CVPR3
2013 Special Section in Celebration of Professor J.K. Aggarwal
Rama Chellappa, Baba C. Vemuri
Comput. Vis. Image Underst.1
2013 Recognizing Interactive Group Activities Using Temporal Interaction Matrices and Their Riemannian Statistics
Rama Chellappa, Shaohua Kevin Zhou
Int. J. Comput. Vis.2
2013 Spatiotemporal Alignment of Visual Signals on a Special Manifold
abstract
We investigate the problem of spatiotemporal alignment of videos, signals, or feature sequences extracted from them. Specifically, we consider the scenario where the spatiotemporal misalignments can be characterized by parametric transformations. Using a nonlinear analytical structure referred to as an alignment manifold, we formulate the alignment problem as an optimization problem on this nonlinear space. We focus our attention on semantically meaningful videos or signals, e.g., those describing or capturing human motion or activities, and propose a new formalism for temporal alignment accounting for executing rate variations among instances of the same video event. The strategy taken in this effort bridges the family of geometric optimization and the family of stochastic algorithms: We regard the search for optimal alignment parameters as a recursive state estimation problem for a particular dynamic system evolving on the alignment manifold. Subsequently, a Sequential Importance Sampling procedure on the alignment manifold is designed for effective alignment. We further extend the basic Sequential Importance Sampling algorithm into a new version called Stochastic Gradient Sequential Importance Sampling, in which we incorporate a steepest descent structure on the alignment manifold and provide a more efficient particle propagation mechanism. We demonstrate the performance of alignment using manifolds on several types of input data that arise in vision problems.
Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Joint Albedo Estimation and Pose Tracking from Video
abstract
The albedo of a Lambertian object is a surface property that contributes to an object's appearance under changing illumination. As a signature independent of illumination, the albedo is useful for object recognition. Single image-based albedo estimation algorithms suffer due to shadows and non-Lambertian effects of the image. In this paper, we propose a sequential algorithm to estimate the albedo from a sequence of images of a known 3D object in varying poses and illumination conditions. We first show that by knowing/estimating the pose of the object at each frame of a sequence, the object's albedo can be efficiently estimated using a Kalman filter. We then extend this for the case of unknown pose by simultaneously tracking the pose as well as updating the albedo through a Rao-Blackwellized particle filter (RBPF). More specifically, the albedo is marginalized from the posterior distribution and estimated analytically using the Kalman filter, while the pose parameters are estimated using importance sampling and by minimizing the projection error of the face onto its spherical harmonic subspace, which results in an illumination-insensitive pose tracking algorithm. Illustrations and experiments are provided to validate the effectiveness of the approach using various synthetic and real sequences followed by applications to unconstrained, video-based face recognition.
Sima Taheri, Aswin C. Sankaranarayanan, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Compressive Acquisition of Linear Dynamical Systems
abstract
Compressive sensing (CS) enables the acquisition and recovery of sparse signals and images at sampling rates significantly below the classical Nyquist rate. Despite significant progress in the theory and methods of CS, little headway has been made in compressive video acquisition and recovery. Video CS is complicated by the ephemeral nature of dynamic events, which makes direct extensions of standard CS imaging architectures and signal models difficult. In this paper, we develop a new framework for video CS for dynamic textured scenes that models the evolution of the scene as a linear dynamical system (LDS). This reduces the video recovery problem to first estimating the model parameters of the LDS from compressive measurements and then reconstructing the image frames. We exploit the low-dimensional dynamic parameters (the state sequence) and high-dimensional static parameters (the observation matrix) of the LDS to devise a novel compressive measurement strategy that measures only the time-varying parameters at each instant and accumulates measurements over time to estimate the time-invariant parameters. This enables us to lower the compressive measurement rate considerably. We validate our approach and demonstrate its effectiveness with a range of experiments involving video recovery and scene classification.
Aswin C. Sankaranarayanan, Pavan Turaga, Rama Chellappa, Richard G. Baraniuk
SIAM J. Imaging Sci.3
2013 Component-Based Recognition of Facesand Facial Expressions
abstract
Most of the existing methods for the recognition of faces and expressions consider either the expression-invariant face recognition problem or the identity-independent facial expression recognition problem. In this paper, we propose joint face and facial expression recognition using a dictionary-based component separation algorithm (DCS). In this approach, the given expressive face is viewed as a superposition of a neutral face component with a facial expression component which is sparse with respect to the whole image. This assumption leads to a dictionary-based component separation algorithm which benefits from the idea of sparsity and morphological diversity. This entails building data-driven dictionaries for neutral and expressive components. The DCS algorithm then uses these dictionaries to decompose an expressive test face into its constituent components. The sparse codes we obtain as a result of this decomposition are then used for joint face and expression recognition. Experiments on publicly available expression and face data sets show the effectiveness of our method.
Sima Taheri, Vishal M. Patel, Rama Chellappa
IEEE Trans. Affect. Comput.3
2013 Guest Editorial: Special issue on intelligent video surveillance for public security and personal privacy
abstract
This Special Issue offers an overview of ongoing research on intelligent video surveillance (IVS) techniques, and brings together cutting-edge research work on security and privacy problems with respect to technological, behavioral, legal, and cultural aspects. We received 34 submissions and each submission was rigorously reviewed by at least two experts in the related fields based on the criteria of originality, significance, quality, and clarity. Eventually, 12 papers were accepted for the Special Issue, spanning a variety of topics including privacy protection, background modeling, tracking, action/activity analysis, and crowd behavior perception. The papers constituting this issue are then briefly summarized.
Noboru Babaguchi, Andrea Cavallaro, Rama Chellappa, Frédéric Dufaux, Liang Wang 0001
IEEE Trans. Inf. Forensics Secur.3
2013 In-Plane Rotation and Scale Invariant Clustering Using Dictionaries
abstract
In this paper, we present an approach that simultaneously clusters images and learns dictionaries from the clusters. The method learns dictionaries and clusters images in the radon transform domain. The main feature of the proposed approach is that it provides both in-plane rotation and scale invariant clustering, which is useful in numerous applications, including content-based image retrieval (CBIR). We demonstrate the effectiveness of our rotation and scale invariant clustering method on a series of CBIR experiments. Experiments are performed on the Smithsonian isolated leaf, Kimia shape, and Brodatz texture datasets. Our method provides both good retrieval performance and greater robustness compared to standard Gabor-based and three state-of-the-art shape-based methods that have similar objectives.
C. S. Sastry 0001, Vishal M. Patel, P. Jonathon Phillips, Rama Chellappa
IEEE Trans. Image Process.5
2013 Pose-Invariant Face Recognition Using Markov Random Fields
abstract
One of the key challenges for current face recognition techniques is how to handle pose variations between the probe and gallery face images. In this paper, we present a method for reconstructing the virtual frontal view from a given nonfrontal face image using Markov random fields (MRFs) and an efficient variant of the belief propagation algorithm. In the proposed approach, the input face image is divided into a grid of overlapping patches, and a globally optimal set of local warps is estimated to synthesize the patches at the frontal view. A set of possible warps for each patch is obtained by aligning it with images from a training database of frontal faces. The alignments are performed efficiently in the Fourier domain using an extension of the Lucas-Kanade algorithm that can handle illumination variations. The problem of finding the optimal warps is then formulated as a discrete labeling problem using an MRF. The reconstructed frontal face image can then be used with any face recognition technique. The two main advantages of our method are that it does not require manually selected facial landmarks or head pose estimation. In order to improve the performance of our pose normalization method in face recognition, we also present an algorithm for classifying whether a given face image is at a frontal or nonfrontal pose. Experimental results on different datasets are presented to demonstrate the effectiveness of the proposed approach.
Huy Tho Ho, Rama Chellappa
IEEE Trans. Image Process.2
2013 Design of Non-Linear Kernel Dictionaries for Object Recognition
abstract
In this paper, we present dictionary learning methods for sparse signal representations in a high dimensional feature space. Using the kernel method, we describe how the well known dictionary learning approaches, such as the method of optimal directions and KSVD, can be made nonlinear. We analyze their kernel constructions and demonstrate their effectiveness through several experiments on classification problems. It is shown that nonlinear dictionary learning approaches can provide significantly better performance compared with their linear counterparts and kernel principal component analysis, especially when the data is corrupted by different types of degradations.
Hien Van Nguyen, Vishal M. Patel, Nasser M. Nasrabadi, Rama Chellappa
IEEE Trans. Image Process.4
2013 Blur and Illumination Robust Face Recognition via Set-Theoretic Characterization
abstract
We address the problem of unconstrained face recognition from remotely acquired images. The main factors that make this problem challenging are image degradation due to blur, and appearance variations due to illumination and pose. In this paper, we address the problems of blur and illumination. We show that the set of all images obtained by blurring a given image forms a convex set. Based on this set-theoretic characterization, we propose a blur-robust algorithm whose main step involves solving simple convex optimization problems. We do not assume any parametric form for the blur kernels, however, if this information is available it can be easily incorporated into our algorithm. Furthermore, using the low-dimensional model for illumination variations, we show that the set of all images obtained from a face image by blurring it and by changing the illumination conditions forms a bi-convex set. Based on this characterization, we propose a blur and illumination-robust algorithm. Our experiments on a challenging real dataset obtained in uncontrolled settings illustrate the importance of jointly modeling blur and illumination.
Priyanka Vageeswaran, Kaushik Mitra, Rama Chellappa
IEEE Trans. Image Process.3
2013 Non-Uniform Deblurring in HDR Image Reconstruction
abstract
Hand-held cameras inevitably result in blurred images caused by camera-shake, and even more so in high dynamic range imaging applications where multiple images are captured over a wide range of exposure settings. The degree of blurring depends on many factors such as exposure time, stability of the platform, and user experience. Camera shake involves not only translations but also rotations resulting in nonuniform blurring. In this paper, we develop a method that takes input non-uniformly blurred and differently exposed images to extract the deblurred, latent irradiance image. We use transformation spread function (TSF) to effectively model the blur caused by camera motion. We first estimate the TSFs of the blurred images from locally derived point spread functions by exploiting their linear relationship. The scene irradiance is then estimated by minimizing a suitably derived cost functional. Two important cases are investigated wherein 1) only the higher exposures are blurred and 2) all the captured frames are blurred.
Channarayapatna Shivaram Vijay, Paramanand Chandramouli, A. N. Rajagopalan 0001, Rama Chellappa
IEEE Trans. Image Process.4
2013 Low-Resolution Face Tracker Robust to Illumination Variations
abstract
In many practical video surveillance applications, the faces acquired by outdoor cameras are of low resolution and are affected by uncontrolled illumination. Although significant efforts have been made to facilitate face tracking or illumination normalization in unconstrained videos, the approaches developed may not be effective in video surveillance applications. This is because: 1) a low-resolution face contains limited information, and 2) major changes in illumination on a small region of the face make the tracking ineffective. To overcome this problem, this paper proposes to perform tracking in an illumination-insensitive feature space, called the gradient logarithm field (GLF) feature space. The GLF feature mainly depends on the intrinsic characteristics of a face and is only marginally affected by the lighting source. In addition, the GLF feature is a global feature and does not depend on a specific face model, and thus is effective in tracking low-resolution faces. Experimental results show that the proposed GLF-based tracker works well under significant illumination changes and outperforms many state-of-the-art tracking algorithms.
Wilman W. W. Zou, Pong C. Yuen, Rama Chellappa
IEEE Trans. Image Process.3
2012 Design of Non-Linear Discriminative Dictionaries for Image Classification
Ashish Shrivastava 0001, Hien Van Nguyen, Vishal M. Patel, Rama Chellappa
ACCV (1)4
2012 Cross-View Action Recognition via a Transferable Dictionary Pair
abstract
Discriminative appearance features are effective for recognizing actions in a fixed view, but generalize poorly to changes in viewpoint. We present a method for viewinvariant action recognition based on sparse representations using a transferable dictionary pair. A transferable dictionary pair consists of two dictionaries that correspond to the source and target views respectively. The two dictionaries are learned simultaneously from pairs of videos taken at different views and aim to encourage each video in the pair to have the same sparse representation. Thus, the transferable dictionary pair links features between the two views that are useful for action recognition. Both unsupervised and supervised algorithms are presented for learning transferable dictionary pairs. Using the sparse representation as features, a classifier built in the source view can be directly transferred to the target view. We extend our approach to transferring an action model learned from multiple source views to one target view. We demonstrate the effectiveness of our approach on the multi-view IXMAS data set. Our results compare favorably to the the state of the art.
Zhuolin Jiang, P. Jonathon Phillips, Rama Chellappa
BMVC4
2012 Dictionary-Based Face Recognition from Video
Vishal M. Patel, P. Jonathon Phillips, Rama Chellappa
ECCV (6)4
2012 Face Association across Unconstrained Video Frames Using Conditional Random Fields
Rama Chellappa
ECCV (7)2
2012 Sparse Embedding: A Framework for Sparsity Promoting Dimensionality Reduction
Hien Van Nguyen, Vishal M. Patel, Nasser M. Nasrabadi, Rama Chellappa
ECCV (6)4
2012 Domain Adaptive Dictionary Learning
Qiang Qiu 0002, Vishal M. Patel, Pavan Turaga, Rama Chellappa
ECCV (4)4
2012 Age Invariant Face Verification with Relative Craniofacial Growth Model
Tao Wu 0009, Rama Chellappa
ECCV (6)2
2012 Rotation invariant simultaneous clustering and dictionary learning
abstract
In this paper, we present an approach that simultaneously clusters database members and learns dictionaries from the clusters. The method learns dictionaries in the Radon transform domain, while clustering in the image domain. Themain feature of the proposed approach is that it provides rotation invariant clustering which is useful in Content Based Image Retrieval (CBIR). We demonstrate through experimental results that the proposed rotation invariant clustering provides better retrieval performance than the standard Gabor-based method that has similar objectives.
C. S. Sastry 0001, Vishal M. Patel, P. Jonathon Phillips, Rama Chellappa
ICASSP5
2012 Kernel dictionary learning
abstract
In this paper, we present dictionary learning methods for sparse and redundant signal representations in high dimensional feature space. Using the kernel method, we describe how the well-known dictionary learning approaches such as the method of optimal directions and K-SVD can be made nonlinear. We analyze these constructions and demonstrate their improved performance through several experiments on classification problems. It is shown that nonlinear dictionary learning approaches can provide better discrimination compared to their linear counterparts and kernel PCA, especially when the data is corrupted by noise.
Hien Van Nguyen, Vishal M. Patel, Nasser M. Nasrabadi, Rama Chellappa
ICASSP4
2012 A hierarchical approach for human age estimation
abstract
We consider the problem of automatic age estimation from face images. Age estimation is usually formulated as a regression problem relating the facial features and the age variable, and a single regression model is learnt for all ages. We propose a hierarchical approach, where we first divide the face images into various age groups and then learn a separate regression model for each group. Given a test image, we first classify the image into one of the age groups and then use the regression model for that particular group. To improve our classification result, we use many different classifiers and fuse them using the majority rule. Experiments show that our approach outperforms many state of the art regression methods for age estimation.
Pavleen Thukral, Kaushik Mitra, Rama Chellappa
ICASSP3
2012 Adaptive rate compressive sensing for background subtraction
abstract
We study the problem of adaptive compressive sensing (CS) of a time-varying signal with slowly changing sparsity and rapidly varying support. We are specifically interested in visual surveillance applications such as background subtraction and tracking. Classical CS theory assumes prior knowledge of signal sparsity in order to determine the number of sensor measurements needed to ensure adequate signal reconstruction. However, when dealing with time-varying signals such as video, prior information regarding the exact sparsity may be difficult to obtain. Assuming a sensor that is able to take an adaptive number of compressive measurements, we present an algorithm based on cross validation that quantitatively evaluates the current measurement rate and adjusts it as needed.
Garrett Warnell, Dikpal Reddy, Rama Chellappa
ICASSP3
2012 Variable focus video: Reconstructing depth and video for dynamic scenes
abstract
Traditional depth from defocus (DFD) algorithms assume that the camera and the scene are static during acquisition time. In this paper, we examine the effects of camera and scene motion on DFD algorithms. We show that, given accurate estimates of optical flow (OF), one can robustly warp the focal stack (FS) images to obtain a virtual static FS and apply traditional DFD algorithms on the static FS. Acquiring accurate OF in the presence of varying focal blur is a challenging task. We show how defocus blur variations cause inherent biases in the estimates of optical flow. We then show how to robustly handle these biases and compute accurate OF estimates in the presence of varying focal blur. This leads to an architecture and an algorithm that converts a traditional 30 fps video camera into a co-located 30 fps image and a range sensor. Further, the ability to extract image and range information allows us to render images with artistic depth-of field effects, both extending and reducing the depth of field of the captured images. We demonstrate experimental results on challenging scenes captured using a camera prototype.
Nitesh Shroff, Ashok Veeraraghavan, Yuichi Taguchi, Oncel Tuzel, Amit K. Agrawal, Rama Chellappa
ICCP6
2012 A Markov Point Process model for wrinkles in human faces
abstract
In this paper, we present a new generative model for wrinkles on aging human faces based on Markov Point Processes (MPP) where wrinkles are considered as stochastic spatial arrangements of sequences of line segments. The model is then used in a Bayesian framework to localize the wrinkles in images. In aging human faces, wrinkles mostly appear as discontinuities in surrounding grayscale texture. The intensity gradients due to wrinkles are enhanced using filters and used as data to detect more probable locations and directions of line segments. Wrinkles are localized by sampling MPP using the Reversible Jump Markov Chain Monte Carlo (RJMCMC) algorithm. Experiments on images obtained from uncontrolled acquisition conditions are presented.
Nazre Batool, Rama Chellappa
ICIP2
2012 Salient view selection based on sparse representation
abstract
A sparse representation-based approach is proposed to find the salient views of 3D objects. Under the assumption that a meaningful object can appear in several perceptible views, we build the object's approximate convex shape that exhibits these apparent views. The salient views are categorized into two groups. The first are boundary representative views that have several visible sides and object surfaces attractive to human perceivers. The second are side representative views that best represent views from sides of the approximating convex shape. The side representative views are class-specific views that possess the most representative power compared to other within-class views. Using the concept of characteristic view class, we first present a sparse representation-based approach for estimating the boundary representative views. With the estimated boundaries, we determine the side representative view(s) based on a minimum reconstruction error.
Vishal M. Patel, Rama Chellappa, P. Jonathon Phillips
ICIP3
2012 Automatic head pose estimation using randomly projected dense SIFT descriptors
abstract
In this paper, we propose an automatic method for determining the head pose from a given face image. The face image is divided into a regular grid and a representation of the image is obtained by extracting dense SIFT descriptors from its grid points. Random Projection (RP) is then applied to reduce the dimension of the concatenated SIFT descriptor vector. Classification and regression using Support Vector Machine (SVM) are combined in order to obtain an accurate estimate of the head pose. The advantage of the proposed approach is that it does not require facial feature points such as eye corners, mouth corners and the nose tip to be extracted from the input face image as in many other methods. Experimental results are presented to demonstrate the effectiveness of the approach.
Huy Tho Ho, Rama Chellappa
ICIP2
2012 Learning discriminative dictionaries with partially labeled data
abstract
While recent techniques for discriminative dictionary learning have demonstrated tremendous success in image analysis applications, their performance is often limited by the amount of labeled data available for training. Even though labeling images is difficult, it is relatively easy to collect unlabeled images either by querying the web or from public datasets. In this paper, we propose a discriminative dictionary learning technique which utilizes both labeled and unlabeled data for learning dictionaries. Extensive evaluation on existing datasets demonstrate that the proposed method performs significantly better than state of the art dictionary learning approaches when unlabeled images are available for training.
Ashish Shrivastava 0001, Jaishanker K. Pillai, Vishal M. Patel, Rama Chellappa
ICIP4
2012 A Grassmann manifold-based domain adaptation approach
Ming-Yu Liu 0001, Rama Chellappa, P. Jonathon Phillips
ICPR3
2012 Introduction
Rama Chellappa
Image Vis. Comput.1
2012 Mathematical statistics and computer vision
abstract
In this discussion paper, I present my views on the role on mathematical statistics for solving computer vision problems.
Rama Chellappa
Image Vis. Comput.1
2012 A Blur-Robust Descriptor with Applications to Face Recognition
abstract
Understanding the effect of blur is an important problem in unconstrained visual analysis. We address this problem in the context of image-based recognition by a fusion of image-formation models and differential geometric tools. First, we discuss the space spanned by blurred versions of an image and then, under certain assumptions, provide a differential geometric analysis of that space. More specifically, we create a subspace resulting from convolution of an image with a complete set of orthonormal basis functions of a prespecified maximum size (that can represent an arbitrary blur kernel within that size), and show that the corresponding subspaces created from a clean image and its blurred versions are equal under the ideal case of zero noise and some assumptions on the properties of blur kernels. We then study the practical utility of this subspace representation for the problem of direct recognition of blurred faces by viewing the subspaces as points on the Grassmann manifold and present methods to perform recognition for cases where the blur is both homogenous and spatially varying. We empirically analyze the effect of noise, as well as the presence of other facial variations between the gallery and probe images, and provide comparisons with existing approaches on standard data sets.
Raghuraman Gopalan, Sima Taheri, Pavan Turaga, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.4
2012 Remote identification of faces: Problems, prospects, and progress
Rama Chellappa, Jie Ni, Vishal M. Patel
Pattern Recognit. Lett.1
2012 Dictionary-Based Face Recognition Under Variable Lighting and Pose
abstract
We present a face recognition algorithm based on simultaneous sparse approximations under varying illumination and pose. A dictionary is learned for each class based on given training examples which minimizes the representation error with a sparseness constraint. A novel test image is projected onto the span of the atoms in each learned dictionary. The resulting residual vectors are then used for classification. To handle variations in lighting conditions and pose, an image relighting technique based on pose-robust albedo estimation is used to generate multiple frontal images of the same person with variable lighting. As a result, the proposed algorithm has the ability to recognize human faces with high accuracy even when only a single or a very few images per person are provided for training. The efficiency of the proposed method is demonstrated using publicly available databases available databases and it is shown that this method is efficient and can perform significantly better than many competitive face recognition algorithms.
Vishal M. Patel, Tao Wu 0009, Soma Biswas, P. Jonathon Phillips, Rama Chellappa
IEEE Trans. Inf. Forensics Secur.5
2012 Age Estimation and Face Verification Across Aging Using Landmarks
abstract
Age estimation and face verification across aging are important problems with a wide range of applications. It is well known that age and identity information are encoded in both texture and shape of the face. Building on recent advances in landmark extraction and statistical techniques for landmark-based shape analysis, we consider these problems using facial shapes. We show that by using well-defined shape spaces and their associated geometry, one can obtain significant performance improvements in both age estimation and face verification. Toward this end, we propose to model the facial shapes as points on a Grassmann manifold. Age estimation and face verification are then considered as regression and classification problems on this manifold. Algorithms for regression and classification are designed to take into account the geometry of the underlying space. The proposed method is flexible and can be used as a standalone age estimator or classifier, and we also present methods for fusion with texture-based algorithms.
Tao Wu 0009, Pavan Turaga, Rama Chellappa
IEEE Trans. Inf. Forensics Secur.3
2012 Gradient-Based Image Recovery Methods From Incomplete Fourier Measurements
abstract
A major problem in imaging applications such as magnetic resonance imaging and synthetic aperture radar is the task of trying to reconstruct an image with the smallest possible set of Fourier samples, every single one of which has a potential time and/or power cost. The theory of compressive sensing (CS) points to ways of exploiting inherent sparsity in such images in order to achieve accurate recovery using sub-Nyquist sampling schemes. Traditional CS approaches to this problem consist of solving total-variation (TV) minimization programs with Fourier measurement constraints or other variations thereof. This paper takes a different approach. Since the horizontal and vertical differences of a medical image are each more sparse or compressible than the corresponding TV image, CS methods will be more successful in recovering these differences individually. We develop an algorithm called GradientRec that uses a CS algorithm to recover the horizontal and vertical gradients and then estimates the original image from these gradients. We present two methods of solving the latter inverse problem, i.e., one based on least-square optimization and the other based on a generalized Poisson solver. After a thorough derivation of our complete algorithm, we present the results of various experiments that compare the effectiveness of the proposed method against other leading methods.
Vishal M. Patel, Ray Maleh, Anna Gilbert 0001, Rama Chellappa
IEEE Trans. Image Process.4
2012 A Learning Approach Towards Detection and Tracking of Lane Markings
abstract
Road scene analysis is a challenging problem that has applications in autonomous navigation of vehicles. An integral component of this system is the robust detection and tracking of lane markings. It is a hard problem primarily due to large appearance variations in lane markings caused by factors such as occlusion (traffic on the road), shadows (from objects like trees), and changing lighting conditions of the scene (transition from day to night). In this paper, we address these issues through a learning-based approach using visual inputs from a camera mounted in front of a vehicle. We propose the following: 1) a pixel-hierarchy feature descriptor to model the contextual information shared by lane markings with the surrounding road region; 2) a robust boosting algorithm to select relevant contextual features for detecting lane markings; and 3) particle filters to track the lane markings, without knowledge of vehicle speed, by assuming the lane markings to be static through the video sequence and then learning the possible road scene variations from the statistics of tracked model parameters. We investigate the effectiveness of our algorithm on challenging daylight and night-time road video sequences.
Raghuraman Gopalan, Tsai Hong, Michael Shneier, Rama Chellappa
IEEE Trans. Intell. Transp. Syst.4
2011 Entropy rate superpixel segmentation
abstract
We propose a new objective function for superpixel segmentation. This objective function consists of two components: entropy rate of a random walk on a graph and a balancing term. The entropy rate favors formation of compact and homogeneous clusters, while the balancing function encourages clusters with similar sizes. We present a novel graph construction for images and show that this construction induces a matroid - a combinatorial structure that generalizes the concept of linear independence in vector spaces. The segmentation is then given by the graph topology that maximizes the objective function under the matroid constraint. By exploiting submodular and mono-tonic properties of the objective function, we develop an efficient greedy algorithm. Furthermore, we prove an approximation bound of ½ for the optimality of the solution. Extensive experiments on the Berkeley segmentation benchmark show that the proposed algorithm outperforms the state of the art in all the standard evaluation metrics.
Ming-Yu Liu 0001, Oncel Tuzel, Srikumar Ramalingam, Rama Chellappa
CVPR4
2011 P2C2: Programmable pixel compressive camera for high speed imaging
abstract
We describe an imaging architecture for compressive video sensing termed programmable pixel compressive camera (P2C2). P2C2 allows us to capture fast phenomena at frame rates higher than the camera sensor. In P2C2, each pixel has an independent shutter that is modulated at a rate higher than the camera frame-rate. The observed intensity at a pixel is an integration of the incoming light modulated by its specific shutter. We propose a reconstruction algorithm that uses the data from P2C2 along with additional priors about videos to perform temporal super-resolution. We model the spatial redundancy of videos using sparse representations and the temporal redundancy using brightness constancy constraints inferred via optical flow. We show that by modeling such spatio-temporal redundancies in a video volume, one can faithfully recover the underlying high-speed video frames from the observed low speed coded video. The imaging architecture and the reconstruction algorithm allows us to achieve temporal super-resolution without loss in spatial resolution. We implement a prototype of P2C2 using an LCOS modulator and recover several videos at 200 fps using a 25 fps camera.
Dikpal Reddy, Ashok Veeraraghavan, Rama Chellappa
CVPR3
2011 Recent advances in age and height estimation from still images and video
abstract
Soft-biometrics such as gender, age, race, etc have been found to be useful characterizations that enable fast pre-filtering and organization of data for biometric applications. In this paper, we focus on two useful soft-biometrics - age and height. We discuss their utility and the factors involved in their estimation from images and videos. In this context, we highlight the role that geometric constraints such as multiview-geometry, and shape-space geometry play. Then, we present methods based on these geometric constraints for age and height-estimation. These methods provide a principled means by fusing image-formation models, multi-view geometric constraints, and robust statistical methods for inference.
Rama Chellappa, Pavan Turaga
FG1
2011 Towards view-invariant expression analysis using analytic shape manifolds
abstract
Facial expression analysis is one of the important components for effective human-computer interaction. However, to develop robust and generalizable models for expression analysis one needs to break the dependence of the models on the choice of the coordinate frame of the camera i.e. expression models should generalize across facial poses. To perform this systematically, one needs to understand the space of observed images subject to projective transformations. However, since the projective shape-space is cumbersome to work with, we address this problem by deriving models for expressions on the affine shape-space as an approximation to the projective shape-space by using a Riemannian interpretation of deformations that facial expressions cause on different parts of the face. We use landmark configurations to represent facial deformations and exploit the fact that the affine shape-space can be studied using the Grassmann manifold. This representation enables us to perform various expression analysis and recognition algorithms without the need for the normalization as a preprocessing step. We extend some of the available approaches for expression analysis to the Grassmann manifold and experimentally show promising results, paving the way for a more general theory of view-invariant expression analysis.
Sima Taheri, Pavan Turaga, Rama Chellappa
FG3
2011 Synthesis-based recognition of low resolution faces
abstract
Recognition of low resolution face images is a challenging problem in many practical face recognition systems. Methods have been proposed in the face recognition literature for the problem when the probe is of low resolution, and a high resolution gallery is available for recognition. These methods modify the probe image such that the resultant image provides better discrimination. We formulate the problem differently by leveraging the information available in the high resolution gallery image and propose a generative approach for classifying the probe image. An important feature of our algorithm is that it can handle resolution changes along with illumination variations. The effective- ness of the proposed method is demonstrated using standard datasets and a challenging outdoor face dataset. It is shown that our method is efficient and can perform significantly better than many competitive low resolution face recognition algorithms.
Vishal M. Patel, Rama Chellappa
IJCB3
2011 Domain adaptation for object recognition: An unsupervised approach
abstract
Adapting the classifier trained on a source domain to recognize instances from a new target domain is an important problem that is receiving recent attention. In this paper, we present one of the first studies on unsupervised domain adaptation in the context of object recognition, where we have labeled data only from the source domain (and therefore do not have correspondences between object categories across domains). Motivated by incremental learning, we create intermediate representations of data between the two domains by viewing the generative subspaces (of same dimension) created from these domains as points on the Grassmann manifold, and sampling points along the geodesic between them to obtain subspaces that provide a meaningful description of the underlying domain shift. We then obtain the projections of labeled source domain data onto these subspaces, from which a discriminative classifier is learnt to classify projected data from the target domain. We discuss extensions of our approach for semi-supervised adaptation, and for cases with multiple source and target domains, and report competitive results on standard datasets.
Raghuraman Gopalan, Rama Chellappa
ICCV3
2011 Sparse dictionary-based representation and recognition of action attributes
abstract
We present an approach for dictionary learning of action attributes via information maximization. We unify the class distribution and appearance information into an objective function for learning a sparse dictionary of action attributes. The objective function maximizes the mutual information between what has been learned and what remains to be learned in terms of appearance information and class distribution for each dictionary item. We propose a Gaussian Process (GP) model for sparse representation to optimize the dictionary objective function. The sparse coding property allows a kernel with a compact support in GP to realize a very efficient dictionary learning process. Hence we can describe an action video by a set of compact and discriminative action attributes. More importantly, we can recognize modeled action categories in a sparse feature space, which can be generalized to unseen and unmodeled action categories. Experimental results demonstrate the effectiveness of our approach in action recognition applications.
Qiang Qiu 0002, Zhuolin Jiang, Rama Chellappa
ICCV3
2011 Blurring-invariant Riemannian metrics for comparing signals and images
abstract
We propose a novel Riemannian framework for comparing signals and images in a manner that is invariant to their levels of blur. This framework uses a log-Fourier representation of signals/images in which the set of all possible Gaussian blurs of a signal, i.e. its orbits under semigroup action of Gaussian blur functions, is a straight line. Using a set of Riemannian metrics under which the group actions are by isometries, the orbits are compared via distances between orbits. We demonstrate this framework using a number of experimental results involving 1D signals and 2D images.
Zhengwu Zhang, Eric Klassen, Anuj Srivastava, Pavan Turaga, Rama Chellappa
ICCV5
2011 Component-based restoration of speckled images
abstract
Many coherent imaging modalities are often characterized by a multiplicative noise, known as speckle which often makes the interpretation of data difficult. In this paper, we present a speckle reduction algorithm based on separating the structure and texture components of SAR images. An iterative algorithm based on surrogate functionals is presented that solves the component optimization formulation. Experiments indicate this proposed method performs favorably compared to state-of-the-art speckle reduction methods.
Vishal M. Patel, Glenn R. Easley, Rama Chellappa
ICIP3
2011 Illumination robust dictionary-based face recognition
abstract
In this paper, we present a face recognition method based on simultaneous sparse approximations under varying illumination. Our method consists of two main stages. In the first stage, a dictionary is learned for each face class based on given training examples which minimizes the representation error with a sparseness constraint. In the second stage, a test image is projected onto the span of the atoms in each learned dictionary. The resulting residual vectors are then used for classification. Furthermore, to handle changes in lighting conditions, we use a relighting approach based on a non-stationary stochastic filter to generate multiple images of the same person with different lighting. As a result, our algorithm has the ability to recognize human faces with good accuracy even when only a single or a very few images are provided for training. The effectiveness of the proposed method is demonstrated on publicly available databases and it is shown that this method is efficient and can perform significantly better than many competitive face recognition algorithms.
Vishal M. Patel, Tao Wu 0009, Soma Biswas, P. Jonathon Phillips, Rama Chellappa
ICIP5
2011 Variable remapping of images from very different sources
abstract
We present a system which registers image sequences acquired by very different sources, so that multiple views could be transformed to the same coordinates system. This enables the functionality of automatic object identification and confirmation across views and platforms. The capability of the system comes from three ingredients: 1) image context enlargement through temporal integration; 2) robust motion estimation using the G-RANSAC framework with a relaxed correspondence criteria; 3) constrained motion estimation within the G-RANSAC framework. The proposed system has worked successfully on thousands of frames from multiple collections with significant variations in scale and resolution.
Yanlin Guo, Reuven Meth, Harvey Sokoloff, Art Pope, Thomas M. Strat, Rama Chellappa
ICIP7
2011 Face tracking in low resolution videos under illumination variations
abstract
In practical face tracking applications, the face region is often small and affected by illumination variations. We address this problem by using a new feature, namely the Gradient-Logarithmic Field (GLF) feature, in the particle filter framework. The GLF feature is robust under illumination variations and the GLF-based tracker does not assume any model for the face being tracked and is effective in low-resolution video. Experimental results show that the proposed GFL-based tracker works well under significant illumination changes and outperforms some of the state-of-the-art algorithms.
Wilman W. W. Zou, Rama Chellappa, Pong C. Yuen
ICIP2
2011 Manifold Precis: An Annealing Technique for Diverse Sampling of Manifolds
abstract
In this paper, we consider the 'Precis' problem of sampling K representative yet diverse data points from a large dataset. This problem arises frequently in applications such as video and document summarization, exploratory data analysis, and pre-filtering. We formulate a general theory which encompasses not just traditional techniques devised for vector spaces, but also non-Euclidean manifolds, thereby enabling these techniques to shapes, human activities, textures and many other image and video based datasets. We propose intrinsic manifold measures for measuring the quality of a selection of points with respect to their representative power, and their diversity. We then propose efficient algorithms to optimize the cost function using a novel annealing-based iterative alternation algorithm. The proposed formulation is applicable to manifolds of known geometry as well as to manifolds whose geometry needs to be estimated from samples. Experimental results show the strength and generality of the proposed approach.
Nitesh Shroff, Pavan Turaga, Rama Chellappa
NIPS3
2011 Silhouette-based gesture and action recognition via modeling trajectories on Riemannian shape manifolds
Mohamed F. Abdelkader, Wael Abd-Almageed, Anuj Srivastava, Rama Chellappa
Comput. Vis. Image Underst.4
2011 Editorial
Sinisa Todorovic, Rama Chellappa
Int. J. Comput. Vis.2
2011 Secure and Robust Iris Recognition Using Random Projections and Sparse Representations
abstract
Noncontact biometrics such as face and iris have additional benefits over contact-based biometrics such as fingerprint and hand geometry. However, three important challenges need to be addressed in a noncontact biometrics-based authentication system: ability to handle unconstrained acquisition, robust and accurate matching, and privacy enhancement without compromising security. In this paper, we propose a unified framework based on random projections and sparse representations, that can simultaneously address all three issues mentioned above in relation to iris biometrics. Our proposed quality measure can handle segmentation errors and a wide variety of possible artifacts during iris acquisition. We demonstrate how the proposed approach can be easily extended to handle alignment variations and recognition from iris videos, resulting in a robust and accurate system. The proposed approach includes enhancements to privacy and security by providing ways to create cancelable iris templates. Results on public data sets show significant benefits of the proposed approach.
Jaishanker K. Pillai, Vishal M. Patel, Rama Chellappa, Nalini K. Ratha
IEEE Trans. Pattern Anal. Mach. Intell.3
2011 A Fast Bilinear Structure from Motion Algorithm Using a Video Sequence and Inertial Sensors
abstract
In this paper, we study the benefits of the availability of a specific form of additional information—the vertical direction (gravity) and the height of the camera, both of which can be conveniently measured using inertial sensors and a monocular video sequence for 3D urban modeling. We show that in the presence of this information, the SfM equations can be rewritten in a bilinear form. This allows us to derive a fast, robust, and scalable SfM algorithm for large scale applications. The SfM algorithm developed in this paper is experimentally demonstrated to have favorable properties compared to the sparse bundle adjustment algorithm. We provide experimental evidence indicating that the proposed algorithm converges in many cases to solutions with lower error than state-of-art implementations of bundle adjustment. We also demonstrate that for the case of large reconstruction problems, the proposed algorithm takes lesser time to reach its solution compared to bundle adjustment. We also present SfM results using our algorithm on the Google StreetView research data set.
Mahesh Ramachandran, Ashok Veeraraghavan, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.3
2011 Statistical Computations on Grassmann and Stiefel Manifolds for Image and Video-Based Recognition
abstract
In this paper, we examine image and video-based recognition applications where the underlying models have a special structure—the linear subspace structure. We discuss how commonly used parametric models for videos and image sets can be described using the unified framework of Grassmann and Stiefel manifolds. We first show that the parameters of linear dynamic models are finite-dimensional linear subspaces of appropriate dimensions. Unordered image sets as samples from a finite-dimensional linear subspace naturally fall under this framework. We show that an inference over subspaces can be naturally cast as an inference problem on the Grassmann manifold. To perform recognition using subspace-based models, we need tools from the Riemannian geometry of the Grassmann manifold. This involves a study of the geometric properties of the space, appropriate definitions of Riemannian metrics, and definition of geodesics. Further, we derive statistical modeling of inter and intraclass variations that respect the geometry of the space. We apply techniques such as intrinsic and extrinsic statistics to enable maximum-likelihood classification. We also provide algorithms for unsupervised clustering derived from the geometry of the manifold. Finally, we demonstrate the improved performance of these methods in a wide variety of vision applications such as activity recognition, video-based face recognition, object recognition from image sets, and activity-based video clustering.
Pavan Turaga, Ashok Veeraraghavan, Anuj Srivastava, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.4
2011 Special Issue on Video Analysis on Resource-Limited Systems
abstract
The 17 papers in this special issue focus on resource-limited systems.
Rama Chellappa, Andrea Cavallaro, Ying Wu 0001, Caifeng Shan, Yun Fu 0001, Kari Pulli
IEEE Trans. Circuits Syst. Video Technol.1
2011 Example-Driven Manifold Priors for Image Deconvolution
abstract
Image restoration methods that exploit prior information about images to be estimated have been extensively studied, typically using the Bayesian framework. In this paper, we consider the role of prior knowledge of the object class in the form of a patch manifold to address the deconvolution problem. Specifically, we incorporate unlabeled image data of the object class, say natural images, in the form of a patch-manifold prior for the object class. The manifold prior is implicitly estimated from the given unlabeled data. We show how the patch-manifold prior effectively exploits the available sample class data for regularizing the deblurring problem. Furthermore, we derive a generalized cross-validation (GCV) function to automatically determine the regularization parameter at each iteration without explicitly knowing the noise variance. Extensive experiments show that this method performs better than many competitive image deconvolution methods.
Jie Ni, Pavan Turaga, Vishal M. Patel, Rama Chellappa
IEEE Trans. Image Process.4
2010 Pose-robust albedo estimation from a single image
abstract
We present a stochastic filtering approach to perform albedo estimation from a single non-frontal face image. Albedo estimation has far reaching applications in various computer vision tasks like illumination-insensitive matching, shape recovery, etc. We extend the formulation proposed in that assumes face in known pose and present an algorithm that can perform albedo estimation from a single image even when pose information is inaccurate. 3D pose of the input face image is obtained as a byproduct of the algorithm. The proposed approach utilizes class-specific statistics of faces to iteratively improve albedo and pose estimates. Illustrations and experimental results are provided to show the effectiveness of the approach. We highlight the usefulness of the method for the task of matching faces across variations in pose and illumination. The facial pose estimates obtained are also compared against ground truth.
Soma Biswas, Rama Chellappa
CVPR2
2010 Group motion segmentation using a Spatio-Temporal Driving Force Model
abstract
We consider the `group motion segmentation' problem and provide a solution for it. The group motion segmentation problem aims at analyzing motion trajectories of multiple objects in video and finding among them the ones involved in a `group motion pattern'. This problem is motivated by and serves as the basis for the `multi-object activity recognition' problem, which is currently an active research topic in event analysis and activity recognition. Specifically, we learn a Spatio-Temporal Driving Force Model to characterize a group motion pattern and design an approach for segmenting the group motion. We illustrate the approach using videos of American football plays, where we identify the offensive players, who follow an offensive motion pattern, from motions of all players in the field. Experiments using GaTech Football Play Dataset validate the effectiveness of the segmentation algorithm.
Rama Chellappa
CVPR2
2010 Fast directional chamfer matching
abstract
We study the object localization problem in images given a single hand-drawn example or a gallery of shapes as the object model. Although many shape matching algorithms have been proposed for the problem over the decades, chamfer matching remains to be the preferred method when speed and robustness are considered. In this paper, we significantly improve the accuracy of chamfer matching while reducing the computational time from linear to sublinear (shown empirically). Specifically, we incorporate edge orientation information in the matching algorithm such that the resulting cost function is piecewise smooth and the cost variation is tightly bounded. Moreover, we present a sublinear time algorithm for exact computation of the directional chamfer matching score using techniques from 3D distance transforms and directional integral images. In addition, the smooth cost function allows to bound the cost distribution of large neighborhoods and skip the bad hypotheses within. Experiments show that the proposed approach improves the speed of the original chamfer matching upto an order of 45×, and it is much faster than many state of art techniques while the accuracy is comparable.
Ming-Yu Liu 0001, Oncel Tuzel, Ashok Veeraraghavan, Rama Chellappa
CVPR4
2010 Robust RVM regression using sparse outlier model
abstract
Kernel regression techniques such as Relevance Vector Machine (RVM) regression, Support Vector Regression and Gaussian processes are widely used for solving many computer vision problems such as age, head pose, 3D human pose and lighting estimation. However, the presence of outliers in the training dataset makes the estimates from these regression techniques unreliable. In this paper, we propose robust versions of the RVM regression that can handle outliers in the training dataset. We decompose the noise term in the RVM formulation into a (sparse) outlier noise term and a Gaussian noise term. We then estimate the outlier noise along with the model parameters. We present two approaches for solving this estimation problem: (1) a Bayesian approach, which essentially follows the RVM framework and (2) an optimization approach based on Basis Pursuit Denoising. In the Bayesian approach, the robust RVM problem essentially becomes a bigger RVM problem with the advantage that it can be solved efficiently by a fast algorithm. Empirical evaluations, and real experiments on image de-noising and age estimation demonstrate the better performance of the robust RVM algorithms over that of the RVM reg ression.
Kaushik Mitra, Ashok Veeraraghavan, Rama Chellappa
CVPR3
2010 Moving vistas: Exploiting motion for describing scenes
abstract
Scene recognition in an unconstrained setting is an open and challenging problem with wide applications. In this paper, we study the role of scene dynamics for improved representation of scenes. We subsequently propose dynamic attributes which can be augmented with spatial attributes of a scene for semantically meaningful categorization of dynamic scenes. We further explore accurate and generalizable computational models for characterizing the dynamics of unconstrained scenes. The large intra-class variation due to unconstrained settings and the complex underlying physics present challenging problems in modeling scene dynamics. Motivated by these factors, we propose using the theory of chaotic systems to capture dynamics. Due to the lack of a suitable dataset, we compiled a dataset of `in-the-wild' dynamic scenes. Experimental results show that the proposed framework leads to the best classification rate among other well-known dynamic modeling techniques. We also show how these dynamic features provide a means to describe dynamic scenes with motion-attributes, which then leads to meaningful organization of the video data.
Nitesh Shroff, Pavan Turaga, Rama Chellappa
CVPR3
2010 Articulation-Invariant Representation of Non-planar Shapes
Raghuraman Gopalan, Pavan Turaga, Rama Chellappa
ECCV (3)3
2010 Aligning Spatio-Temporal Signals on a Special Manifold
Rama Chellappa
ECCV (5)2
2010 Compressive Acquisition of Dynamic Scenes
Aswin C. Sankaranarayanan, Pavan Turaga, Richard G. Baraniuk, Rama Chellappa
ECCV (1)4
2010 Sparse representations and Random Projections for robust and cancelable biometrics
abstract
In recent years, the theories of Sparse Representation (SR) and Compressed Sensing (CS) have emerged as powerful tools for efficiently processing data in non-traditional ways. An area of promise for these theories is biométrie identification. In this paper, we review the role of sparse representation and CS for efficient biométrie identification. Algorithms to perform identification from face and iris data are reviewed. By applying Random Projections it is possible to purposively hide the biométrie data within a template. This procedure can be effectively employed for securing and protecting personal biométrie data against theft. Some of the most compelling challenges and issues that confront research in biometrics using sparse representations and CS are also addressed.
Vishal M. Patel, Rama Chellappa, Massimo Tistarelli
ICARCV2
2010 Robust regression using sparse learning for high dimensional parameter estimation problems
abstract
Algorithms such as Least Median of Squares (LMedS) and Random Sample Consensus (RANSAC) have been very successful for low-dimensional robust regression problems. However, the combinatorial nature of these algorithms makes them practically unusable for high-dimensional applications. In this paper, we introduce algorithms that have cubic time complexity in the dimension of the problem, which make them computationally efficient for high-dimensional problems. We formulate the robust regression problem by projecting the dependent variable onto the null space of the independent variables which receives significant contributions only from the outliers. We then identify the outliers using sparse representation/learning based algorithms. Under certain conditions, that follow from the theory of sparse representation, these polynomial algorithms can accurately solve the robust regression problem which is, in general, a combinatorial problem. We present experimental results that demonstrate the efficacy of the proposed algorithms. We also analyze the intrinsic parameter space of robust regression and identify an efficient and accurate class of algorithms for different operating conditions. An application to facial age estimation is presented.
Kaushik Mitra, Ashok Veeraraghavan, Rama Chellappa
ICASSP3
2010 Sectored Random Projections for Cancelable Iris Biometrics
abstract
Privacy and security are essential requirements in practical biometric systems. In order to prevent the theft of biometric patterns, it is desired to modify them through revocable and non invertible transformations called Cancelable Biometrics. In this paper, we propose an efficient algorithm for generating a Cancelable Iris Biometric based on Sectored Random Projections. Our algorithm can generate a new pattern if the existing one is stolen, retain the original recognition performance and prevent extraction of useful information from the transformed patterns. Our method also addresses some of the drawbacks of existing techniques and is robust to degradations due to eyelids and eyelashes.
Jaishanker K. Pillai, Vishal M. Patel, Rama Chellappa, Nalini K. Ratha
ICASSP3
2010 The role of geometry in age estimation
abstract
Understanding and modeling of aging in human faces is an important problem in many real-world applications such as biometrics, authentication, and synthesis. In this paper, we consider the role of geometric attributes of faces, as described by a set of landmark points on the face, in age perception. Towards this end, we show that the space of landmarks can be interpreted as a Grassmann manifold. Then the problem of age estimation is posed as a problem of function estimation on the manifold. The warping of an average face to a given face is quantified as a velocity vector that transforms the average to a given face along a smooth geodesic in unit-time. This deformation is then shown to contain important information about the age of the face. We show in experiments that exploiting geometric cues in a principled manner provides comparable performance to several systems that utilize both geometric and textural cues. We show results on age estimation using the standard FG-Net dataset and a passport dataset which illustrate the effectiveness of the approach.
Pavan Turaga, Soma Biswas, Rama Chellappa
ICASSP3
2010 Recognizing offensive strategies from football videos
abstract
We address the problem of recognizing offensive play strategies from American football play videos. Specifically, we propose a probabilistic model which describes the generative process of an observed football play and takes into account practical issues in real football videos, such as difficulty in identifying offensive players, view changes, and tracking errors. In particular, we exploit the geometric properties of nonlinear spaces of involved variables and design statistical models on these manifolds. Then recognition is performed via 'analysis-by-synthesis' technique. Experiments on a newly established dataset of American football videos demonstrate the effectiveness of the approach.
Rama Chellappa
ICIP2
2010 Evaluation of state-of-the-art algorithms for remote face recognition
abstract
In this paper, we describe a remote face database which has been acquired in an unconstrained outdoor environment. The face images in this database suffer from variations due to blur, poor illumination, pose, and occlusion. It is well known that many state-of-the-art still image-based face recognition algorithms work well, when constrained (frontal, well illuminated, high-resolution, sharp, and complete) face images are presented. In this paper, we evaluate the effectiveness of a subset of existing still image-based face recognition algorithms for the remote face data set. We demonstrate that in addition to applying a good classification algorithm, consistent detection of faces with fewer false alarms and finding features that are robust to variations mentioned above are very important for remote face recognition. Also setting up a comprehensive metric to evaluate the quality of face images is necessary in order to reject images that are of low quality.
Jie Ni, Rama Chellappa
ICIP2
2010 Automatic target recognition based on simultaneous sparse representation
abstract
In this paper, an automatic target recognition algorithm is presented based on a framework for learning dictionaries for simultaneous sparse signal representation and feature extraction. The dictionary learning algorithm is based on class supervised simultaneous orthogonal matching pursuit while a matching pursuit-based similarity measure is used for classification. We show how the proposed framework can be helpful for efficient utilization of data, with the possibility of developing real-time, robust target classification. We verify the efficacy of the proposed algorithm using confusion matrices on the well known Comanche forward-looking infrared data set consisting of ten different military targets at different orientations.
Vishal M. Patel, Nasser M. Nasrabadi, Rama Chellappa
ICIP3
2010 Pose estimation in heavy clutter using a multi-flash camera
abstract
We propose a novel solution to object detection, localization and pose estimation with applications in robot vision. The proposed method is especially applicable when the objects of interest may not be richly textured and are immersed in heavy clutter. We show that a multi-flash camera (MFC) provides accurate separation of depth edges and texture edges in such scenes. Then, we reformulate the problem, as one of finding matches between the depth edges obtained in one or more MFC images to the rendered depth edges that are computed offline using 3D CAD model of the objects. In order to facilitate accurate matching of these binary depth edge maps, we introduce a novel cost function that respects both the position and the local orientation of each edge pixel. This cost function is significantly superior to traditional Chamfer cost and leads to accurate matching even in heavily cluttered scenes where traditional methods are unreliable. We present a sub-linear time algorithm to compute the cost function using techniques from 3D distance transforms and integral images. Finally, we also propose a multi-view based pose-refinement algorithm to improve the estimated pose. We implemented the algorithm on an industrial robot arm and obtained location and angular estimation accuracy of the order of 1 mm and 2° respectively for a variety of parts with minimal texture.
Ming-Yu Liu 0001, Oncel Tuzel, Ashok Veeraraghavan, Rama Chellappa, Amit K. Agrawal, Haruhisa Okuda
ICRA4
2010 Large-Scale Matrix Factorization with Missing Data under Additional Constraints
abstract
Matrix factorization in the presence of missing data is at the core of many computer vision problems such as structure from motion (SfM), non-rigid SfM and photometric stereo. We formulate the problem of matrix factorization with missing data as a low-rank semidefinite program (LRSDP) with the advantage that: $1)$ an efficient quasi-Newton implementation of the LRSDP enables us to solve large-scale factorization problems, and $2)$ additional constraints such as ortho-normality, required in orthographic SfM, can be directly incorporated in the new formulation. Our empirical evaluations suggest that, under the conditions of matrix completion theory, the proposed algorithm finds the optimal solution, and also requires fewer observations compared to the current state-of-the-art algorithms. We further demonstrate the effectiveness of the proposed algorithm in solving the affine SfM problem, non-rigid SfM and photometric stereo problems.
Kaushik Mitra, Sameer Sheorey, Rama Chellappa
NIPS3
2010 PADS: A Probabilistic Activity Detection Framework for Video Data
abstract
There is now a growing need to identify various kinds of activities that occur in videos. In this paper, we first present a logical language called Probabilistic Activity Description Language (PADL) in which users can specify activities of interest. We then develop a probabilistic framework which assigns to any subvideo of a given video sequence a probability that the subvideo contains the given activity, and we finally develop two fast algorithms to detect activities within this framework. OffPad finds all minimal segments of a video that contain a given activity with a probability exceeding a given threshold. In contrast, the OnPad algorithm examines a video during playout (rather than afterwards as OffPad does) and computes the probability that a given activity is occurring (even if the activity is only partially complete). Our prototype Probabilistic Activity Detection System (PADS) implements the framework and the two algorithms, building on top of existing image processing algorithms. We have conducted detailed experiments and compared our approach to four different approaches presented in the literature. We show that-for complex activity definitions-our approach outperforms all the other approaches.
Massimiliano Albanese, Rama Chellappa, Naresh P. Cuntoor, Vincenzo Moscato, Antonio Picariello, V. S. Subrahmanian, Octavian Udrea
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Video Metrology Using a Single Camera
abstract
This paper presents a video metrology approach using an uncalibrated single camera that is either stationary or in planar motion. Although theoretically simple, measuring the length of even a line segment in a given video is often a difficult problem. Most existing techniques for this task are extensions of single image-based techniques and do not achieve the desired accuracy especially in noisy environments. In contrast, the proposed algorithm moves line segments on the reference plane to share a common endpoint using the vanishing line information followed by fitting multiple concentric circles on the image plane. A fully automated real-time system based on this algorithm has been developed to measure vehicle wheelbases using an uncalibrated stationary camera. The system estimates the vanishing line using invariant lengths on the reference plane from multiple frames rather than the given parallel lines, which may not exist in videos. It is further extended to a camera undergoing a planar motion by automatically selecting frames with similar vanishing lines from the video. Experimental results show that the measurement results are accurate enough to classify moving vehicles based on their size.
Feng Guo 0006, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Online Empirical Evaluation of Tracking Algorithms
abstract
Evaluation of tracking algorithms in the absence of ground truth is a challenging problem. There exist a variety of approaches for this problem, ranging from formal model validation techniques to heuristics that look for mismatches between track properties and the observed data. However, few of these methods scale up to the task of visual tracking, where the models are usually nonlinear and complex and typically lie in a high-dimensional space. Further, scenarios that cause track failures and/or poor tracking performance are also quite diverse for the visual tracking problem. In this paper, we propose an online performance evaluation strategy for tracking systems based on particle filters using a time-reversed Markov chain. The key intuition of our proposed methodology relies on the time-reversible nature of physical motion exhibited by most objects, which in turn should be possessed by a good tracker. In the presence of tracking failures due to occlusion, low SNR, or modeling errors, this reversible nature of the tracker is violated. We use this property for detection of track failures. To evaluate the performance of the tracker at time instant t, we use the posterior of the tracking algorithm to initialize a time-reversed Markov chain. We compute the posterior density of track parameters at the starting time t=0 by filtering back in time to the initial time instant. The distance between the posterior density of the time-reversed chain (at t=0) and the prior density used to initialize the tracking algorithm forms the decision statistic for evaluation. It is observed that when the data are generated by the underlying models, the decision statistic takes a low value. We provide a thorough experimental analysis of the evaluation methodology. Specifically, we demonstrate the effectiveness of our approach for tackling common challenges such as occlusion, pose, and illumination changes and provide the Receiver Operating Characteristic (ROC) curves. Finally, we also show the applicability of the core ideas of the paper to other tracking algorithms such as the Kanade-Lucas-Tomasi (KLT) feature tracker and the mean-shift tracker.
Hao Wu 0014, Aswin C. Sankaranarayanan, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.3
2010 Special Section on Distributed Camera Networks: Sensing, Processing, Communication, and Implementation
abstract
The eight papers in this special section span across theoretical and practical considerations of various aspects of distributed camera networks, including adaptive sensing, distributed processing, efficient communications, and versatile implementations.
Rama Chellappa, Wendi B. Heinzelman, Janusz Konrad, Dan Schonfeld, Marilyn Wolf
IEEE Trans. Image Process.1
2010 Robust Height Estimation of Moving Objects From Uncalibrated Videos
abstract
This paper presents an approach for video metrology. From videos acquired by an uncalibrated stationary camera, we first recover the vanishing line and the vertical point of the scene based upon tracking moving objects that primarily lie on a ground plane. Using geometric properties of moving objects, a probabilistic model is constructed for simultaneously grouping trajectories and estimating vanishing points. Then we apply a single view mensuration algorithm to each of the frames to obtain height measurements. We finally fuse the multiframe measurements using the least median of squares (LMedS) as a robust cost function and the Robbins-Monro stochastic approximation (RMSA) technique. This method enables less human supervision, more flexibility and improved robustness. From the uncertainty analysis, we conclude that the method with auto-calibration is robust in practice. Results are shown based upon realistic tracking data from a variety of scenes.
Jie Shao 0007, Shaohua Kevin Zhou, Rama Chellappa
IEEE Trans. Image Process.3
2010 An Efficient and Robust Algorithm for Shape Indexing and Retrieval
abstract
Many shape matching methods are either fast but too simplistic to give the desired performance or promising as far as performance is concerned but computationally demanding. In this paper, we present a very simple and efficient approach that not only performs almost as good as many state-of-the-art techniques but also scales up to large databases. In the proposed approach, each shape is indexed based on a variety of simple and easily computable features which are invariant to articulations, rigid transformations, etc. The features characterize pairwise geometric relationships between interest points on the shape. The fact that each shape is represented using a number of distributed features instead of a single global feature that captures the shape in its entirety provides robustness to the approach. Shapes in the database are ordered according to their similarity with the query shape and similar shapes are retrieved using an efficient scheme which does not involve costly operations like shape-wise alignment or establishing correspondences. Depending on the application, the approach can be used directly for matching or as a first step for obtaining a short list of candidate shapes for more rigorous matching. We show that the features proposed to perform shape indexing can be used to perform the rigorous matching as well, to further improve the retrieval performance.
Soma Biswas, Gaurav Aggarwal, Rama Chellappa
IEEE Trans. Multim.3
2010 Video Précis: Highlighting Diverse Aspects of Videos
abstract
Summarizing long unconstrained videos is gaining importance in surveillance, web-based video browsing, and video-archival applications. Summarizing a video requires one to identify key aspects that contain the essence of the video. In this paper, we propose an approach that optimizes two criteria that a video summary should embody. The first criterion, “coverage,” requires that the summary be able to represent the original video well. The second criterion, “diversity,” requires that the elements of the summary be as distinct from each other as possible. Given a user-specified summary length, we propose a cost function to measure the quality of a summary. The problem of generating a précis is then reduced to a combinatorial optimization problem of minimizing the proposed cost function. We propose an efficient method to solve the optimization problem. We demonstrate through experiments (on KTH data, unconstrained skating video, a surveillance video, and a YouTube home video) that optimizing the proposed criterion results in meaningful video summaries over a wide range of scenarios. Summaries thus generated are then evaluated using both quantitative measures and user studies.
Nitesh Shroff, Pavan Turaga, Rama Chellappa
IEEE Trans. Multim.3
2010 Applications of a Simple Characterization of Human Gait in Surveillance
abstract
Applications of a simple spatiotemporal characterization of human gait in the surveillance domain are presented. The approach is based on decomposing a video sequence into x-t slices, which generate periodic patterns referred to as double helical signatures (DHSs). The features of DHS are given as follows: 1) they naturally encode the appearance and kinematics of human motion and reveal geometric symmetries and 2) they are effective and efficient for recovering gait parameters and detecting simple events. We present an iterative local curve embedding algorithm to extract the DHS from video sequences. Two applications are then considered. First, the DHS is used for simultaneous segmentation and labeling of body parts in cluttered scenes. Experimental results showed that the algorithm is robust to size, viewing angles, camera motion, and severe occlusion. Then, the DHS is used to classify load-carrying conditions. By examining various symmetries in DHS, activities such as carrying, holding, and walking with objects that are attached to legs are detected. Our approach possesses several advantages: a compact representation that can be computed in real time is used; furthermore, it does not depend on silhouettes or landmark tracking, which are sensitive to errors in background subtraction stage.
Yang Ran, Qinfen Zheng, Rama Chellappa, Thomas M. Strat
IEEE Trans. Syst. Man Cybern. Part B3
2009 Learning multi-modal densities on Discriminative Temporal Interaction Manifold for group activity recognition
abstract
While video-based activity analysis and recognition has received much attention, existing body of work mostly deals with single object/person case. Coordinated multi-object activities, or group activities, present in a variety of applications such as surveillance, sports, and biological monitoring records, etc., are the main focus of this paper. Unlike earlier attempts which model the complex spatial temporal constraints among multiple objects with a parametric Bayesian network, we propose a Discriminative Temporal Interaction Manifold (DTIM) framework as a data-driven strategy to characterize the group motion pattern without employing specific domain knowledge. In particular, we establish probability densities on the DTIM, whose element, the discriminative temporal interaction matrix, compactly describes the coordination and interaction among multiple objects in a group activity. For each class of group activity we learn a multi-modal density function on the DTIM. A Maximum a Posteriori (MAP) classifier on the manifold is then designed for recognizing new activities. Experiments on football play recognition demonstrate the effectiveness of the approach.
Rama Chellappa, Shaohua Kevin Zhou
CVPR2
2009 Enforcing integrability by error correction using l1-minimization
abstract
Surface reconstruction from gradient fields is an important final step in several applications involving gradient manipulations and estimation. Typically, the resulting gradient field is non-integrable due to linear/non-linear gradient manipulations, or due to presence of noise/outliers in gradient estimation. In this paper, we analyze integrability as error correction, inspired from recent work in compressed sensing, particulary ℓ0- ℓ1equivalence. We propose to obtain the surface by finding the gradient field which best fits the corrupted gradient field in ℓ1sense. We present an exhaustive analysis of the properties of ℓ1solution for gradient field integration using linear algebra and graph analogy. We consider three cases: (a) noise, but no outliers (b) no-noise but outliers and (c) presence of both noise and outliers in the given gradient field. We show that ℓ1solution performs as well as least squares in the absence of outliers. While previous ℓ0- ℓ1equivalence work has focused on the number of errors (outliers), we show that the location of errors is equally important for gradient field integration. We characterize the ℓ1solution both in terms of location and number of outliers, and outline scenarios where ℓ1solution is equivalent to ℓ0solution. We also show that when ℓ1solution is not able to remove outliers, the property of local error confinement holds: i.e., the errors do not propagate to the entire surface as in least squares. We compare with previous techniques and show that ℓ1solution performs well across all scenarios without the need for any tunable parameter adjustments.
Dikpal Reddy, Amit K. Agrawal, Rama Chellappa
CVPR3
2009 Locally time-invariant models of human activities using trajectories on the grassmannian
abstract
Human activity analysis is an important problem in computer vision with applications in surveillance and summarization and indexing of consumer content. Complex human activities are characterized by non-linear dynamics that make learning, inference and recognition hard. In this paper, we consider the problem of modeling and recognizing complex activities which exhibit time-varying dynamics. To this end, we describe activities as outputs of linear dynamic systems (LDS) whose parameters vary with time, or a time-varying linear dynamic system (TV-LDS). We discuss parameter estimation methods for this class of models by assuming that the parameters are locally time-invariant. Then, we represent the space of LDS models as a Grassmann manifold. Then, the TV-LDS model is defined as a trajectory on the Grassmann manifold. We show how trajectories on the Grassmannian can be characterized using appropriate distance metrics and statistical methods that reflect the underlying geometry of the manifold. This results in more expressive and powerful models for complex human activities. We demonstrate the strength of the framework for activity-based summarization of long videos and recognition of complex human actions on two datasets.
Pavan Turaga, Rama Chellappa
CVPR2
2009 Recognizing coordinated multi-object activities using a dynamic event ensemble model
abstract
While video-based activity analysis and recognition has received broad attention, existing body of work mostly deals with single object/person case. Modeling involving multiple objects and recognition of coordinated group activities, present in a variety of applications such as surveillance, sports, biological records, and so on, is the main focus of this paper. Unlike earlier attempts which model the complex spatial temporal constraints among different activities of multiple objects with a parametric Bayesian network, we propose a dynamic ‘event ensemble’ framework as a data-driven strategy to characterize the group motion pattern without employing any specific domain knowledge. In particular, we exploit the Riemannian geometric property of the set of ensemble description functions and develop a compact representation for group activities on the ensemble manifold. An appropriate classifier on the manifold is then designed for recognizing new activities. Experiments on football play recognition demonstrate the effectiveness of the framework.
Rama Chellappa
ICASSP2
2009 Enhancing sparsity using gradients for compressive sensing
abstract
In this paper, we propose a reconstruction method that recovers images assumed to have a sparse representation in a gradient domain by using partial measurement samples that are collected in the Fourier domain. A key improvement of this technique is that it makes use of a robust generalized Poisson solver that greatly aids in achieving a significantly improved performance over similar proposed methods. Experiments provided also demonstrate that this new technique is more flexible to work with either random or restricted sampling scenarios better than its competitors.
Vishal M. Patel, Glenn R. Easley, Rama Chellappa, Dennis M. Healy Jr.
ICIP3
2009 Compressed sensing for Synthetic Aperture Radar imaging
abstract
In this paper, we introduce a new Synthetic Aperture Radar (SAR) imaging modality that provides a high resolution map of the spatial distribution of targets and terrain based on a significant reduction in the number of transmitted and/or received electromagnetic waveforms. This new imaging scheme, which requires no new hardware components, allows the aperture to be compressed and presents many important applications and advantages among which include resolving ambiguities, strong resistance to countermeasures and interception, and reduced on-board storage constraints.
Vishal M. Patel, Glenn R. Easley, Dennis M. Healy Jr., Rama Chellappa
ICIP4
2009 How would you look as you age?
abstract
Facial appearances change with increase in age. While generic growth patterns that are characteristic of different age groups can be identified, facial growth is also observed to be influenced by individual-specific attributes such as one's gender, ethnicity, life-style etc. In this paper, we propose a facial growth model that comprises of transformation models for facial shape and texture. We collected empirical data pertaining to facial growth from a database of age-separated face images of adults and used the same in developing the aforementioned transformation models. The proposed model finds applications in predicting one's appearance across ages and in performing face verification across ages.
Narayanan Ramanathan, Rama Chellappa
ICIP2
2009 Unsupervised view and rate invariant clustering of video sequences
Pavan Turaga, Ashok Veeraraghavan, Rama Chellappa
Comput. Vis. Image Underst.3
2009 Robust Estimation of Albedo for Illumination-Invariant Matching and Shape Recovery
abstract
We present a nonstationary stochastic filtering framework for the task of albedo estimation from a single image. There are several approaches in the literature for albedo estimation, but few include the errors in estimates of surface normals and light source direction to improve the albedo estimate. The proposed approach effectively utilizes the error statistics of surface normals and illumination direction for robust estimation of albedo, for images illuminated by single and multiple light sources. The albedo estimate obtained is subsequently used to generate albedo-free normalized images for recovering the shape of an object. Traditional Shape-from-Shading (SFS) approaches often assume constant/piecewise constant albedo and known light source direction to recover the underlying shape. Using the estimated albedo, the general problem of estimating the shape of an object with varying albedo map and unknown illumination source is reduced to one that can be handled by traditional SFS approaches. Experimental results are provided to show the effectiveness of the approach and its application to illumination-invariant matching and shape recovery. The estimated albedo maps are compared with the ground truth. The maps are used as illumination-invariant signatures for the task of face recognition across illumination variations. The recognition results obtained compare well with the current state-of-the-art approaches. Impressive shape recovery results are obtained using images downloaded from the Web with little control over imaging conditions. The recovered shapes are also used to synthesize novel views under novel illumination conditions.
Soma Biswas, Gaurav Aggarwal, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.3
2009 Moving Object Verification in Airborne Video Sequences
abstract
This paper presents an end-to-end system for moving object verification in airborne video sequences. Using a sample selection module, the system first selects frames from a short sequence and stores them in an exemplar database. To handle appearance change due to potentially large aspect angle variations, a homography-based view synthesis method is then used to generate a novel view of each image in the exemplar database at the same pose as the testing object in each frame of a testing video segment. A rotationally invariant color matcher and a spatial-feature matcher based on distance transforms are combined using a weighted average rule to compare the novel view and the testing object. After looping over all testing frames, the set of match scores is passed to a temporal analysis module to examine the behavior of the testing object, and calculate a final likelihood. Very good verification performance is achieved over thousands of trials for both color and infrared video sequences using the proposed system.
Zhanfeng Yue, David Guarino, Rama Chellappa
IEEE Trans. Circuits Syst. Video Technol.3
2009 Appearance Modeling Using a Geometric Transform
abstract
A general transform, called the geometric transform (GeT), that models the appearance inside a closed contour is proposed. The proposed GeT is a functional of an image intensity function and a region indicator function derived from a closed contour. It can be designed to combine the shape and appearance information at different resolutions and to generate models invariant to deformation, articulation, or occlusion. By choosing appropriate functionals and region indicator functions, the GeT unifies Radon transform, trace transform, and a class of image warpings. By varying the region indicator and the types of features used for appearance modeling, five novel types of GeTs are introduced and applied to fingerprinting the appearance inside a contour. They include the GeTs based on a level set, shape matching, feature curves, and the GeT invariant to occlusion, and a multiresolution GeT (MRGeT). Applications of GeT to pedestrian identity recognition, human body part segmentation, and image synthesis are illustrated. The proposed approach produces promising results when applied to fingerprinting the appearance of a human and body parts despite the presence of nonrigid deformations and articulated motion.
Jian Li 0022, Shaohua Kevin Zhou, Rama Chellappa
IEEE Trans. Image Process.3
2009 Multicamera Tracking of Articulated Human Motion Using Shape and Motion Cues
abstract
We present a completely automatic algorithm for initializing and tracking the articulated motion of humans using image sequences obtained from multiple cameras. A detailed articulated human body model composed of sixteen rigid segments that allows both translation and rotation at joints is used. Voxel data of the subject obtained from the images is segmented into the different articulated chains using Laplacian Eigenmaps. The segmented chains are registered in a subset of the frames using a single-frame registration technique and subsequently used to initialize the pose in the sequence. A temporal registration method is proposed to identify the partially segmented or unregistered articulated chains in the remaining frames in the sequence. The proposed tracker uses motion cues such as pixel displacement as well as 2-D and 3-D shape cues such as silhouettes, motion residue, and skeleton curves. The tracking algorithm consists of a predictor that uses motion cues and a corrector that uses shape cues. The use of complementary cues in the tracking alleviates the twin problems of drift and convergence to local minima. The use of multiple cameras also allows us to deal with the problems due to self-occlusion and kinematic singularity. We present tracking results on sequences with different kinds of motion to illustrate the effectiveness of our approach. The pose of the subject is correctly tracked for the duration of the sequence as can be verified by inspection.
Aravind Sundaresan, Rama Chellappa
IEEE Trans. Image Process.2
2009 Rate-Invariant Recognition of Humans and Their Activities
abstract
Pattern recognition in video is a challenging task because of the multitude of spatio-temporal variations that occur in different videos capturing the exact same event. While traditional pattern-theoretic approaches account for the spatial changes that occur due to lighting and pose, very little has been done to address the effect of temporal rate changes in the executions of an event. In this paper, we provide a systematic model-based approach to learn the nature of such temporal variations (time warps) while simultaneously allowing for the spatial variations in the descriptors. We illustrate our approach for the problem of action recognition and provide experimental justification for the importance of accounting for rate variations in action recognition. The model is composed of a nominal activity trajectory and a function space capturing the probability distribution of activity-specific time warping transformations. We use the square-root parameterization of time warps to derive geodesics, distance measures, and probability distributions on the space of time warping functions. We then design a Bayesian algorithm which treats the execution rate function as a nuisance variable and integrates it out using Monte Carlo sampling, to generate estimates of class posteriors. This approach allows us to learn the space of time warps for each activity while simultaneously capturing other intra- and interclass variations. Next, we discuss a special case of this approach which assumes a uniform distribution on the space of time warping functions and show how computationally efficient inference algorithms may be derived for this special case. We discuss the relative advantages and disadvantages of both approaches and show their efficacy using experiments on gait-based person identification and activity recognition.
Ashok Veeraraghavan, Anuj Srivastava, Amit K. Roy-Chowdhury, Rama Chellappa
IEEE Trans. Image Process.4
2008 Statistical analysis on Stiefel and Grassmann manifolds with applications in computer vision
abstract
Many applications in computer vision and pattern recognition involve drawing inferences on certain manifold-valued parameters. In order to develop accurate inference algorithms on these manifolds we need to a) understand the geometric structure of these manifolds b) derive appropriate distance measures and c) develop probability distribution functions (pdf) and estimation techniques that are consistent with the geometric structure of these manifolds. In this paper, we consider two related manifolds - the Stiefel manifold and the Grassmann manifold, which arise naturally in several vision applications such as spatio-temporal modeling, affine invariant shape analysis, image matching and learning theory. We show how accurate statistical characterization that reflects the geometry of these manifolds allows us to design efficient algorithms that compare favorably to the state of the art in these very different applications. In particular, we describe appropriate distance measures and parametric and non-parametric density estimators on these manifolds. These methods are then used to learn class conditional densities for applications such as activity recognition, video based face recognition and shape classification.
Pavan Turaga, Ashok Veeraraghavan, Rama Chellappa
CVPR3
2008 Compressive Sensing for Background Subtraction
Volkan Cevher, Aswin C. Sankaranarayanan, Marco F. Duarte, Dikpal Reddy, Richard G. Baraniuk, Rama Chellappa
ECCV (2)6
2008 Modeling shape and textural variations in aging faces
abstract
We propose a two fold approach towards modeling facial aging in adults. Firstly, we develop a shape transformation model that is formulated as a physically-based parametric muscle model that captures the subtle deformations facial features undergo with age. The model implicitly accounts for the physical properties and geometric orientations of the individual facial muscles. Next, we develop an image gradient based texture transformation function that characterizes facial wrinkles and other skin artifacts often observed during different ages. Facial growth statistics (both in terms of shape and texture) play a crucial role in developing the aforementioned transformation models. From a database that comprises of pairs of age separated face images of many individuals, we extract age-based facial measurements across key fiducial features and further, study textural variations across ages. We present experimental results that illustrate the applications of the proposed facial aging model in tasks such as face recognition and facial appearance prediction across aging.
Narayanan Ramanathan, Rama Chellappa
FG2
2008 Multi-biometric cohort analysis for biometric fusion
abstract
Biometric matching decisions have traditionally been made based solely on a score that represents the similarity of the query biometric to the enrolled biometric(s) of the claimed identity. Fusion schemes have been proposed to benefit from the availability of multiple biometric samples (e.g., multiple samples of the same fingerprint) or multiple different biometrics (e.g., face and fingerprint). These commonly adopted fusion approaches rarely make use of the large number of non-matching biometric samples available in the database in the form of other enrolled identities or training data. In this paper, we study the impact of combining this information with the existing fusion methodologies in a cohort analysis framework. Experimental results are provided to show the usefulness of such a cohort-based fusion of face and fingerprint biometrics.
Gaurav Aggarwal, Nalini K. Ratha, Ruud M. Bolle, Rama Chellappa
ICASSP4
2008 Compressive wireless arrays for bearing estimation
abstract
Joint processing of sensor array outputs improves the performance of parameter estimation and hypothesis testing problems beyond the sum of the individual sensor processing results. When the sensors have high data sampling rates, arrays are tethered, creating a disadvantage for their deployment and also limiting their aperture size. In this paper, we develop the signal processing algorithms for randomly deployable wireless sensor arrays that are severely constrained in communication bandwidth. We focus on the acoustic bearing estimation problem and show that when the target bearings are modeled as a sparse vector in the angle space, low dimensional random projections of the microphone signals can be used to determine multiple source bearings by solving an ℓ1-norm minimization problem. Field data results are shown where only 10 bits of information is passed from each microphone to estimate multiple target bearings.
Volkan Cevher, Ali Cafer Gürbüz, James H. McClellan, Rama Chellappa
ICASSP4
2008 Factorized variational approximations for acoustic multi source localization
abstract
Estimation based on received signal strength (RSS) is crucial in sensor networks for sensor localization, target tracking, etc. In this paper, we present a Gaussian approximation of the Chi distribution that is applicable to general RSS source localization problems in sensor networks. Using our Gaussian approximation, we provide a factorized variational Bayes (VB) approximation to the location and power posterior of multiple sources using a sensor network. When the source signal and the sensor noise have uncorrelated Gaussian distributions, we demonstrate that the envelope of the sensor output can be accurately modeled with a multiplicative Gaussian noise model. In turn, our factorized VB approximations decrease the computational complexity and provide computational robustness as the number of targets increases. Simulations are provided to demonstrate the effectiveness of the proposed approximations.
Volkan Cevher, Aswin C. Sankaranarayanan, Rama Chellappa
ICASSP3
2008 Compressed sensing for multi-view tracking and 3-D voxel reconstruction
abstract
Compressed sensing (CS) suggests that a signal, sparse in some basis, can be recovered from a small number of random projections. In this paper, we apply the CS theory on sparse background-subtracted silhouettes and show the usefulness of such an approach in various multi-view estimation problems. The sparsity of the silhouette images corresponds to sparsity of object parameters (location, volume etc.) in the scene. We use random projections (compressed measurements) of the silhouette images for directly recovering object parameters in the scene coordinates. To keep the computational requirements of this recovery procedure reasonable, we tessellate the scene into a bunch of non-overlapping lines and perform estimation on each of these lines. Our method is scalable in the number of cameras and utilizes very few measurements for transmission among cameras. We illustrate the usefulness of our approach for multi-view tracking and 3-D voxel reconstruction problems.
Dikpal Reddy, Aswin C. Sankaranarayanan, Volkan Cevher, Rama Chellappa
ICIP4
2008 Stochastic fusion of multi-view gradient fields
abstract
Image gradients form powerful cues in a host of vision and graphics applications. In this paper, we consider multiple views of a textured planar scene and consider the problem of estimating the scene texture map using these multi-view inputs. Modeling each camera view as a projective transformation of the scene, we show that the problem is equivalent to that of studying the effect of noise (and the projective imaging) on the gradient fields induced by this texture map. We show that these noisy gradient fields can be modeled as complete observers of the scene radiance. Further, the corrupting noise can be shown to be additive and linear, although spatially varying. However, the specific form of the noise term can be exploited to design linear estimators that fuse the gradient fields obtained from each of the individual views. The fused gradient field forms a robust estimate of the scene gradients and can be used for scene reconstruction.
Aswin C. Sankaranarayanan, Rama Chellappa
ICIP2
2008 Learning action dictionaries from video
abstract
Summarizing the contents of a video containing human activities is an important problem in computer vision and has important applications in automated surveillance systems. Summarizing a video requires one to identify and learn a 'vocabulary' of action-phrases corresponding to specific events and actions occurring in the video. We propose a generative model for dynamic scenes containing human activities as a composition of independent action-phrases - each of which is derived from an underlying vocabulary. Given a long video sequence, we propose a completely unsupervised approach to learn the vocabulary. Once the vocabulary is learnt, a video segment can be decomposed into a collection of phrases for summarization. We then describe methods to learn the correlations between activities and sequentiality of events. We also propose a novel method for building invariances to spatial transforms in the summarization scheme.
Pavan Turaga, Rama Chellappa
ICIP2
2008 An ontology based approach for activity recognition from video
abstract
Representation and recognition of human activities is an important problem for video surveillance and security applications. Considering the wide variety of settings in which surveillance systems are being deployed, it is necessary to create a common knowledge-base or ontology of human activities. Most current attempts at ontology design in computer vision for human activities have been empirical in nature. In this paper, we present a more systematic approach to address the problem of designing ontologies for visual activity recognition. We draw on general ontology design principles and adapt them to the specific domain of human activity ontologies. Then, we discuss qualitative evaluation principles and provide several examples from existing ontologies and how they can be improved upon. Finally, we demonstrate quantitatively in terms of recognition performance, the efficacy and validity of our approach for bank and airport tarmac surveillance domains.
Umut Akdemir, Pavan Turaga, Rama Chellappa
ACM Multimedia3
2008 Model Driven Segmentation of Articulating Humans in Laplacian Eigenspace
abstract
We propose a general approach using Laplacian Eigenmaps and a graphical model of the human body to segment 3D voxel data of humans into different articulated chains. In the bottom-up stage, the voxels are transformed into a high-dimensional (6D or less) Laplacian Eigenspace (LE) of the voxel neighborhood graph. We show that LE is effective at mapping voxels on long articulated chains to nodes on smooth 1D curves that can be easily discriminated, and prove these properties using representative graphs. We fit 1D splines to voxels belonging to different articulated chains such as the limbs, head and trunk, and we determine the boundary between splines by thresholding the spline fit error, which is high at junctions. A top-down probabilistic approach is then used to register the segmented chains, utilizing both their mutual connectivity and their individual properties such as length and thickness. Our approach enables us to deal with complex poses such as those where the limbs form loops. We use the segmentation results to automatically estimate the human body models. Although we use human subjects in our experiments, the method is fairly general and can be applied to voxel-based registration of any articulated object, which is composed of long chains. We present results on real and synthetic data that illustrate the usefulness of this approach.
Aravind Sundaresan, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Shape-and-Behavior Encoded Tracking of Bee Dances
abstract
Behavior analysis of social insects has garnered impetus in recent years and has led to some advances in fields like control systems, flight navigation etc. Manual labeling of insect motions required for analyzing the behaviors of insects requires significant investment of time and effort. In this paper, we propose certain general principles that help in simultaneous automatic tracking and behavior analysis with applications in tracking bees and recognizing specific behaviors exhibited by them. The state space for tracking is defined using position, orientation and the current behavior of the insect being tracked. The position and orientation are parametrized using a shape model while the behavior is explicitly modeled using a three-tier hierarchical motion model. The first tier (dynamics) models the local motions exhibited and the models built in this tier act as a vocabulary for behavior modeling. The second tier is a Markov motion model built on top of the local motion vocabulary which serves as the behavior model. The third tier of the hierarchy models the switching between behaviors and this is also modeled as a Markov model. We address issues in learning the three-tier behavioral model, in discriminating between models, detecting and in modeling abnormal behaviors. Another important aspect of this work is that it leads to joint tracking and behavior analysis instead of the traditional track and then recognize approach. We apply these principles for tracking bees in a hive while they are executing the waggle dance and the round dance.
Ashok Veeraraghavan, Rama Chellappa, Mandyam V. Srinivasan
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Object Detection, Tracking and Recognition for Multiple Smart Cameras
abstract
Video cameras are among the most commonly used sensors in a large number of applications, ranging from surveillance to smart rooms for videoconferencing. There is a need to develop algorithms for tasks such as detection, tracking, and recognition of objects, specifically using distributed networks of cameras. The projective nature of imaging sensors provides ample challenges for data association across cameras. We first discuss the nature of these challenges in the context of visual sensor networks. Then, we show how real-world constraints can be favorably exploited in order to tackle these challenges. Examples of real-world constraints are (a) the presence of a world plane, (b) the presence of a three-dimiensional scene model, (c) consistency of motion across cameras, and (d) color and texture properties. In this regard, the main focus of this paper is towards highlighting the efficient use of the geometric constraints induced by the imaging devices to derive distributed algorithms for target detection, tracking, and recognition. Our discussions are supported by several examples drawn from real applications. Lastly, we also describe several potential research problems that remain to be addressed.
Aswin C. Sankaranarayanan, Ashok Veeraraghavan, Rama Chellappa
Proc. IEEE3
2008 Machine Recognition of Human Activities: A Survey
abstract
The past decade has witnessed a rapid proliferation of video cameras in all walks of life and has resulted in a tremendous explosion of video content. Several applications such as content-based video annotation and retrieval, highlight extraction and video summarization require recognition of the activities occurring in the video. The analysis of human activities in videos is an area with increasingly important consequences from security and surveillance to entertainment and personal archiving. Several challenges at various levels of processing-robustness against errors in low-level processing, view and rate-invariant representations at midlevel processing and semantic representation of human activities at higher level processing-make this problem hard to solve. In this review paper, we present a comprehensive survey of efforts in the past couple of decades to address the problems of representation, recognition, and learning of human activities from video and related applications. We discuss the problem at two major levels of complexity: 1) "actions" and 2) "activities." "Actions" are characterized by simple motion patterns typically executed by a single human. "Activities" are more complex and involve coordinated actions among a small number of humans. We will discuss several approaches and classify them according to their ability to handle varying degrees of complexity as interpreted above. We begin with a discussion of approaches to model the simplest of action classes known as atomic or primitive actions that do not require sophisticated dynamical modeling. Then, methods to model actions with more complex dynamics are discussed. The discussion then leads naturally to methods for higher level representation of complex activities.
Pavan Turaga, Rama Chellappa, V. S. Subrahmanian, Octavian Udrea
IEEE Trans. Circuits Syst. Video Technol.2
2008 Activity Modeling Using Event Probability Sequences
abstract
Changes in motion properties of trajectories provide useful cues for modeling and recognizing human activities. We associate an event with significant changes that are localized in time and space, and represent activities as a sequence of such events. The localized nature of events allows for detection of subtle changes or anomalies in activities. In this paper, we present a probabilistic approach for representing events using the hidden Markov model (HMM) framework. Using trained HMMs for activities, an event probability sequence is computed for every motion trajectory in the training set. It reflects the probability of an event occurring at every time instant. Though the parameters of the trained HMMs depend on viewing direction, the event probability sequences are robust to changes in viewing direction. We describe sufficient conditions for the existence of view invariance. The usefulness of the proposed event representation is illustrated using activity recognition and anomaly detection. Experiments using the indoor University of Central Florida human action dataset, the Carnegie Mellon University Credo Intelligence, Inc., Motion Capture dataset, and the outdoor Transportation Security Administration airport tarmac surveillance dataset show encouraging results.
Naresh P. Cuntoor, Bayya Yegnanarayana, Rama Chellappa
IEEE Trans. Image Process.3
2008 Algorithmic and Architectural Optimizations for Computationally Efficient Particle Filtering
abstract
In this paper, we analyze the computational challenges in implementing particle filtering, especially to video sequences. Particle filtering is a technique used for filtering nonlinear dynamical systems driven by non-Gaussian noise processes. It has found widespread applications in detection, navigation, and tracking problems. Although, in general, particle filtering methods yield improved results, it is difficult to achieve real time performance. In this paper, we analyze the computational drawbacks of traditional particle filtering algorithms, and present a method for implementing the particle filter using the Independent Metropolis Hastings sampler, that is highly amenable to pipelined implementations and parallelization. We analyze the implementations of the proposed algorithm, and, in particular, concentrate on implementations that have minimum processing times. It is shown that the design parameters for the fastest implementation can be chosen by solving a set of convex programs. The proposed computational methodology was verified using a cluster of PCs for the application of visual tracking. We demonstrate a linear speed-up of the algorithm using the methodology proposed in the paper.
Aswin C. Sankaranarayanan, Ankur Srivastava 0001, Rama Chellappa
IEEE Trans. Image Process.3
2008 A Constrained Probabilistic Petri Net Framework for Human Activity Detection in Video
abstract
Recognition of human activities in restricted settings such as airports, parking lots and banks is of significant interest in security and automated surveillance systems. In such settings, data is usually in the form of surveillance videos with wide variation in quality and granularity. Interpretation and identification of human activities requires an activity model that a) is rich enough to handle complex multi-agent interactions, b) is robust to uncertainty in low-level processing and c) can handle ambiguities in the unfolding of activities. We present a computational framework for human activity representation based on Petri nets. We propose an extension-Probabilistic Petri Nets (PPN)-and show how this model is well suited to address each of the above requirements in a wide variety of settings. We then focus on answering two types of questions: (i) what are the minimal sub-videos in which a given activity is identified with a probability above a certain threshold and (ii) for a given video, which activity from a given set occurred with the highest probability? We provide the PPN-MPS algorithm for the first problem, as well as two different algorithms (naive PPN-MPA and PPN-MPA) to solve the second. Our experimental results on a dataset consisting of bank surveillance videos and an unconstrained TSA tarmac surveillance dataset show that our algorithms are both fast and provide high quality results.
Massimiliano Albanese, Rama Chellappa, Naresh P. Cuntoor, Vincenzo Moscato, Antonio Picariello, V. S. Subrahmanian, Octavian Udrea
IEEE Trans. Multim.2
2008 A Constrained Probabilistic Petri Net Framework for Human Activity Detection in Video
abstract
Recognition of human activities in restricted settings such as airports, parking lots and banks is of significant interest in security and automated surveillance systems. In such settings, data is usually in the form of surveillance videos with wide variation in quality and granularity. Interpretation and identification of human activities requires an activity model that a) is rich enough to handle complex multi-agent interactions, b) is robust to uncertainty in low-level processing and c) can handle ambiguities in the unfolding of activities. We present a computational framework for human activity representation based on Petri nets. We propose an extension—Probabilistic Petri Nets (PPN)—and show how this model is well suited to address each of the above requirements in a wide variety of settings. We then focus on answering two types of questions: (i) what are the minimal sub-videos in which a given activity is identified with a probability above a certain threshold and (ii) for a given video, which activity from a given set occurred with the highest probability? We provide the PPN-MPS algorithm for the first problem, as well as two different algorithms (naive PPN-MPA and PPN-MPA) to solve the second. Our experimental results on a dataset consisting of bank surveillance videos and an unconstrained TSA tarmac surveillance dataset show that our algorithms are both fast and provide high quality results.
Massimiliano Albanese, Rama Chellappa, Naresh P. Cuntoor, Vincenzo Moscato, Antonio Picariello, V. S. Subrahmanian, Octavian Udrea
IEEE Trans. Multim.2
2008 Synthesis of Silhouettes and Visual Hull Reconstruction for Articulated Humans
abstract
In this paper, we propose a complete framework for improved synthesis and understanding of the human pose from a limited number of silhouette images. It combines the active image-based visual hull (IBVH) algorithm and a contour-based body part segmentation technique. We derive a simple, approximate algorithm to decide the extrinsic parameters of a virtual camera, and synthesize the turntable image collection of the person using the IBVH algorithm by actively moving the virtual camera on a properly computed circular trajectory around the person. Using the turning function distance as the silhouette similarity measurement, this approach can be used to generate the desired pose-normalized images for recognition applications. In order to overcome the inability of the visual hull (VH) method to reconstruct concave regions, we propose a contour-based human body part localization algorithm to segment the silhouette images into convex body parts. The body parts observed from the virtual view are generated separately from the corresponding body parts observed from the input views and then assembled together for a more accurate VH reconstruction. Furthermore, the obtained turntable image collection helps to improve the body part segmentation and identification process. By using the inner distance shape context (IDSC) measurement, we are able to estimate the body part locations more accurately from a synthesized view where we can localize the body part more precisely. Experiments show that the proposed algorithm can greatly improve body part segmentation and hence shape reconstruction results.
Zhanfeng Yue, Rama Chellappa
IEEE Trans. Multim.2
2007 Symmetric Objects are Hardly Ambiguous
abstract
Given any two images taken under different illumination conditions, there always exist a physically realizable object which is consistent with both the images even if the lighting in each scene is constrained to be a known point light source at infinity. In this paper, we show that images are much less ambiguous for the class of bilaterally symmetric Lambertian objects. In fact, the set of such objects can be partitioned into equivalence classes such that it is always possible to distinguish between two objects belonging to different equivalence classes using just one image per object. The conditions required for two objects to belong to the same equivalence class are very restrictive, thereby leading to the conclusion that images of symmetric objects are hardly ambiguous. The observation leads to an illumination-invariant matching algorithm to compare images of bilaterally symmetric Lambertian objects. Experiments on real data are performed to show the implications of the theoretical result even when the symmetry and Lambertian assumptions are not strictly satisfied.
Gaurav Aggarwal, Soma Biswas, Rama Chellappa
CVPR3
2007 Efficient Indexing For Articulation Invariant Shape Matching And Retrieval
abstract
Most shape matching methods are either fast but too simplistic to give the desired performance or promising as far as performance is concerned but computationally demanding. In this paper, we present a very simple and efficient approach that not only performs almost as good as many state-of-the-art techniques but also scales up to large databases. In the proposed approach, each shape is indexed based on a variety of simple and easily computable features which are invariant to articulations and rigid transformations. The features characterize pairwise geometric relationships between interest points on the shape, thereby providing robustness to the approach. Shapes are retrieved using an efficient scheme which does not involve costly operations like shape-wise alignment or establishing correspondences. Even for a moderate size database of 1000 shapes, the retrieval process is several times faster than most techniques with similar performance. Extensive experimental results are presented to illustrate the advantages of our approach as compared to the best in the field.
Soma Biswas, Gaurav Aggarwal, Rama Chellappa
CVPR3
2007 Human Identification using Gait and Face
abstract
In general the visual-hull approach for performing integrated face and gait recognition requires at least two cameras. In this paper we present experimental results for fusion of face and gait for the single camera case. We considered the NIST database which contains outdoor face and gait data for 30 subjects. In the NIST database, subjects walk along an inverted Sigma pattern. In (A. Kale, et al., 2003), we presented a view-invariant gait recognition algorithm for the single camera case along with some experimental evaluations. In this chapter we present the results of our view-invariant gait recognition algorithm in (A. Kale, et al., 2003) on the NIST database. The algorithm is based on the planar approximation of the person which is valid when the person walks far away from the camera. In (S. Zhou et al., 2003), an algorithm for probabilistic recognition of human faces from video was proposed and the results were demonstrated on the NIST database. Details of these methods can be found in the respective papers. We give an outline of the fusion strategy here.
Rama Chellappa, Amit K. Roy-Chowdhury, Amit A. Kale
CVPR1
2007 Epitomic Representation of Human Activities
abstract
We introduce an epitomic representation for modeling human activities in video sequences. A video sequence is divided into segments within which the dynamics of objects is assumed to be linear and modeled using linear dynamical systems. The tuple consisting of the estimated system matrix, statistics of the input signal and the initial state value is said to form an epitome. The system matrices are decomposed using the Iwasawa matrix decomposition to isolate the effect of rotation, scaling and projective action on the state vector. "We demonstrate the usefulness of the proposed representation and decomposition for activity recognition using the TSA airport surveillance dataset and the UCF indoor human action dataset.
Naresh P. Cuntoor, Rama Chellappa
CVPR2
2007 From Videos to Verbs: Mining Videos for Activities using a Cascade of Dynamical Systems
abstract
Clustering video sequences in order to infer and extract activities from a single video stream is an extremely important problem and has significant potential in video indexing, surveillance, activity discovery and event recognition. Clustering a video sequence into activities requires one to simultaneously recognize activity boundaries (activity consistent subsequences) and cluster these activity subsequences. In order to do this, we build a generative model for activities (in video) using a cascade of dynamical systems and show that this model is able to capture and represent a diverse class of activities. We then derive algorithms to learn the model parameters from a video stream and also show how a single video sequence may be clustered into different clusters where each cluster represents an activity. We also propose a novel technique to build affine, view, rate invariance of the activity into the distance metric for clustering. Experiments show that the clusters found by the algorithm correspond to semantically meaningful activities.
Pavan Turaga, Ashok Veeraraghavan, Rama Chellappa
CVPR3
2007 In Situ Evaluation of Tracking Algorithms Using Time Reversed Chains
abstract
Automatic evaluation of visual tracking algorithms in the absence of ground truth is a very challenging and important problem. In the context of online appearance modeling, there is an additional ambiguity involving the correctness of the appearance model. In this paper, we propose a novel performance evaluation strategy for tracking systems based on particle filter using a time reversed Markov chain. Starting from the latest observation, the time reversed chain is propagated back till the starting time t = 0 of the tracking algorithm. The posterior density of the time reversed chain is also computed. The distance between the posterior density of the time reversed chain (at t = 0) and the prior density used to initialize the tracking algorithm forms the decision statistic for evaluation. It is postulated that when the data is generated true to the underlying models, the decision statistic takes a low value. We empirically demonstrate the performance of the algorithm against various common failure modes in the generic visual tracking problem. Finally, we derive a small frame approximation that allows for very efficient computation of the decision statistic.
Hao Wu 0014, Aswin C. Sankaranarayanan, Rama Chellappa
CVPR3
2007 Joint Acoustic-Video Fingerprinting of Vehicles, Part I
abstract
We address vehicle classification and measurement problems using acoustic and video sensors. In this paper, we show how to estimate a vehicle's speed, width, and length by jointly estimating its acoustic wave-pattern using a single passive sensor that records the vehicle's drive-by noise. The acoustic wave-pattern is approximated using three envelope shape (ES) components, which approximate the shape of the received signal's power envelope. We incorporate the parameters of the ES components along with the estimates of the vehicle engine RPM and number of cylinders to create a vehicle profile vector that forms an intuitive discriminatory feature space. In the companion paper, we discuss vehicle classification and mensuration based on silhouette extraction and wheel detection, using a video sensor. Vehicle speed estimation and classification results are provided using field data.
Volkan Cevher, Rama Chellappa, James H. McClellan
ICASSP (2)2
2007 Joint Acoustic-Video Fingerprinting of Vehicles, Part II
abstract
In this second paper, we first show how to estimate the wheelbase length of a vehicle using line metrology in video. We then address the vehicle fingerprinting problem using vehicle silhouettes and color invariants. We combine the acoustic metrology and classification results discussed in Part I with the video results to improve estimation performance and robustness. The acoustic video fusion is achieved in a Bayesian framework by assuming conditional independence of the observations of each modality. For the metrology density functions, Laplacian approximations are used for computational efficiency. Experimental results are given using field data.
Volkan Cevher, Feng Guo 0006, Aswin C. Sankaranarayanan, Rama Chellappa
ICASSP (2)4
2007 Coarse-to-Fine Event Model for Human Activities
abstract
We analyze coarse-to-fine hierarchical representation of human activities in video sequences. It can be used for efficient video browsing and activity recognition. Activities are modeled using a sequence of instantaneous events. Events in activities can be represented in a coarse-to-fine hierarchy in several ways, i.e., there may not be a unique hierarchical structure. We present five criteria and quantitative measures for evaluating their effectiveness. The criteria are minimalism, stability, consistency, accessibility and applicability. It is desirable to develop activity models that rank highly on these criteria at all levels of hierarchy. In this paper, activities are represented as sequence of event probabilities computed using the hidden Markov model framework. Two aspects of hierarchies are analyzed: the effect of reduced frame rate on the accuracy of events detected at a finer scale; and the effect of reduced spatial resolution on activity recognition. Experiments using the UCF indoor human action dataset and the TSA airport tarmac surveillance dataset show encouraging results.
Naresh P. Cuntoor, Rama Chellappa
ICASSP (1)2
2007 Markerless Monocular Tracking of Articulated Human Motion
abstract
This paper presents a method for tracking general 3D general articulated human motion using a single camera with unknown calibration data. No markers, special clothes, or devices are assumed to be attached to the subject. In addition, both the camera and the subject are allowed to move freely, so that long-term view-independent human motion tracking and recognition are possible. We exploit the fact that the anatomical structure of the human body can be approximated by an articulated blob model. The optical flow under scaled orthographic projection is used to relate the spatial-temporal intensity change of the image sequence to the human motion parameters. These motion parameters are obtained by solving a set of linear equations to achieve global optimization. The correctness and robustness of the proposed method are demonstrated using Tai Chi sequences.
Rama Chellappa
ICASSP (1)2
2007 Simulation and Analysis of Human Walking Motion
abstract
Simulation and analysis of human walking motion has applications in surveillance and healthcare. In this paper we discuss an approach for modeling human walking motion using a mechanical model in the form of a kinematic chain consisting of rigid links and revolute joints. Our goal is to discriminate different types of walking motions using information such as joint torque and angle sequences extracted from the model. The angle sequences are initially extracted using 3D geometry. From these angle sequences we extract the torque sequences using a recursive Newton Euler inverse dynamics algorithm. Time series models and dynamic time warping of the torque and angle sequences are used to characterize and discriminate different walking patterns. A forward dynamics algorithm is also presented for synthesizing different walking sequences like limping from a normal walking torque sequence.
Kaustav Nandy, Rama Chellappa
ICASSP (1)2
2007 Robust Estimation of Albedo for Illumination-invariant Matching and Shape Recovery
abstract
In this paper, we propose a non-stationary stochastic filtering framework for the task of albedo estimation from a single image. There are several approaches in literature for albedo estimation, but few include the errors in estimates of surface normals and light source directions to improve the albedo estimate. The proposed approach effectively utilizes the error statistics of surface normals and illumination direction for robust estimation of albedo. The albedo estimate obtained is further used to generate albedo-free normalized images for recovering the shape of an object. Illustrations and experiments are provided to show the efficacy of the approach and its application to illumination-invariant matching and shape recovery.
Soma Biswas, Gaurav Aggarwal, Rama Chellappa
ICCV3
2007 Fast Bilinear SfM with Side Information
abstract
We study the beneficial effect of side information on the Structure from Motion (SfM) estimation problem. The side information that we consider is measurement of a 'reference vector' and distance from fixed plane perpendicular to that reference vector. Firstly, we show that in the presence of this information, the SfM equations can be rewritten similar to a bilinear form in its unknowns. Secondly, we describe a fast iterative estimation procedure to recover the structure of both stationary scenes and moving objects that capitalizes on this information. We also provide a refinement procedure in order to tackle incomplete or noisy side information. We characterize the algorithm with respect to its reconstruction accuracy, memory requirements and stability. Finally, we describe two classes of commonly occurring real-world scenarios in which this algorithm will be effective: (a) presence of a dominant ground plane in the scene and (b) presence of an inertial measurement unit on board. Experiments using both real data and rigorous simulations show the efficacy of the algorithm.
Mahesh Ramachandran, Ashok Veeraraghavan, Rama Chellappa
ICCV3
2007 Robust Visual Tracking Using the Time-Reversibility Constraint
abstract
Visual tracking is a very important front-end to many vision applications. We present a new framework for robust visual tracking in this paper. Instead of just looking forward in the time domain, we incorporate both forward and backward processing of video frames using a novel time-reversibility constraint. This leads to a new minimization criterion that combines the forward and backward similarity functions and the distances of the state vectors between the forward and backward states of the tracker. The new framework reduces the possibility of the tracker getting stuck in local minima and significantly improves the tracking robustness and accuracy. Our approach is general enough to be incorporated into most of the current tracking algorithms. We illustrate the improvements due to the proposed approach for the popular KLT tracker and a search based tracker. The experimental results show that the improved KLT tracker significantly outperforms the original KLT tracker. The time-reversibility constraint used for tracking can be incorporated to improve the performance of optical flow, mean shift tracking and other algorithms.
Hao Wu 0014, Rama Chellappa, Aswin C. Sankaranarayanan, Shaohua Kevin Zhou
ICCV2
2007 Kernel fully constrained least squares abundance estimates
abstract
A critical step for fitting a linear mixing model to hyperspectral imagery is the estimation of the abundances. The abundances are the percentage of each end member within a given pixel; therefore, they should be non-negative and sum to one. With the advent of kernel based algorithms for hyperspectral imagery, kernel based abundance estimates have become necessary. This paper presents such an algorithm that estimates the abundances in the kernel feature space while maintaining the non-negativity and sum-to-one constraints. The usefulness of the algorithm is shown using the AVIRIS Cuprite, Nevada image.
Joshua B. Broadwater, Rama Chellappa, Amit Banerjee, Philippe Burlina
IGARSS2
2007 Hybrid Detectors for Subpixel Targets
abstract
Subpixel detection is a challenging problem in hyperspectral imagery analysis. Since the target size is smaller than the size of a pixel, detection algorithms must rely solely on spectral information. A number of different algorithms have been developed over the years to accomplish this task, but most detectors have taken either a purely statistical or a physics-based approach to the problem. We present two new hybrid detectors that take advantage of these approaches by modeling the background using both physics and statistics. Results demonstrate improved performance over the well known AMSD and ACE subpixel algorithms in experiments that include multiple targets, images, and area types--especially when dealing with weak targets in complex backgrounds.
Joshua B. Broadwater, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Appearance Characterization of Linear Lambertian Objects, Generalized Photometric Stereo, and Illumination-Invariant Face Recognition
abstract
Traditional photometric stereo algorithms employ a Lambertian reflectance model with a varying albedo field and involve the appearance of only one object. In this paper, we generalize photometric stereo algorithms to handle all appearances of all objects in a class, in particular the human face class, by making use of the linear Lambertian property. A linear Lambertian object is one which is linearly spanned by a set of basis objects and has a Lambertian surface. The linear property leads to a rank constraint and, consequently, a factorization of an observation matrix that consists of exemplar images of different objects (e.g., faces of different subjects) under different, unknown illuminations. Integrability and symmetry constraints are used to fully recover the subspace bases using a novel linearized algorithm that takes the varying albedo field into account. The effectiveness of the linear Lambertian property is further investigated by using it for the problem of illumination-invariant face recognition using just one image. Attached shadows are incorporated in the model by a careful treatment of the inherent nonlinearity in Lambert's law. This enables us to extend our algorithm to perform face recognition in the presence of multiple illumination sources. Experimental results using standard data sets are presented.
Shaohua Kevin Zhou, Gaurav Aggarwal, Rama Chellappa, David Jacobs 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2007 Guest Editorial: Special Issue on Human Detection and Recognition
abstract
The 12 regular papers and three correspondences in this special issue focus on human detection and recognition. The papers represent gait, face (3-D, 2-D, video), iris, palmprint, cardiac sounds, and vulnerability of biometrics and protection against the spoof attacks.
Bir Bhanu, Nalini K. Ratha, B. V. K. Vijaya Kumar, Rama Chellappa, Josef Bigün
IEEE Trans. Inf. Forensics Secur.4
2007 A Multiple-Hypothesis Approach for Multiobject Visual Tracking
abstract
In multiple-object tracking applications, it is essential to address the problem of associating targets and observation data. For visual tracking of multiple targets which involves objects that split and merge, a target may be associated with multiple measurements and many targets may be associated with a single measurement. The space of such data association is exponential in the number of targets and exhaustive enumeration is impractical. We pose the association problem as a bipartite graph edge covering problem given the targets and the object detection information. We propose an efficient method of maintaining multiple association hypotheses with the highest probabilities over all possible histories of associations. Our approach handles objects entering and exiting the field of view, merging and splitting objects, as well as objects that are detected as fragmented parts. Experimental results are given for tracking multiple players in a soccer game and for tracking people with complex interaction in a surveillance setting. It is shown through quantitative evaluation that our method tracks through varying degrees of interactions among the targets with high success rate.
Seong-Wook Joo, Rama Chellappa
IEEE Trans. Image Process.2
2007 Target Tracking Using a Joint Acoustic Video System
abstract
In this paper, a multitarget tracking system for collocated video and acoustic sensors is presented. We formulate the tracking problem using a particle filter based on a state-space approach. We first discuss the acoustic state-space formulation whose observations use a sliding window of direction-of-arrival estimates. We then present the video state space that tracks a target's position on the image plane based on online adaptive appearance models. For the joint operation of the filter, we combine the state vectors of the individual modalities and also introduce a time-delay variable to handle the acoustic-video data synchronization issue, caused by acoustic propagation delays. A novel particle filter proposal strategy for joint state-space tracking is introduced, which places the random support of the joint filter where the final posterior is likely to lie. By using the Kullback-Leibler divergence measure, it is shown that the joint operation of the filter decreases the worst case divergence of the individual modalities. The resulting joint tracking filter is quite robust against video and acoustic occlusions due to our proposal strategy. Computer simulations are presented with synthetic and field data to demonstrate the filter's performance
Volkan Cevher, Aswin C. Sankaranarayanan, James H. McClellan, Rama Chellappa
IEEE Trans. Multim.4
2007 Super-Resolution of Face Images Using Kernel PCA-Based Prior
abstract
We present a learning-based method to super-resolve face images using a kernel principal component analysis-based prior model. A prior probability is formulated based on the energy lying outside the span of principal components identified in a higher-dimensional feature space. This is used to regularize the reconstruction of the high-resolution image. We demonstrate with experiments that including higher-order correlations results in significant improvements
Ayan Chakrabarti, A. N. Rajagopalan 0001, Rama Chellappa
IEEE Trans. Multim.3
2006 Key Frame-Based Activity Representation Using Antieigenvalues
Naresh P. Cuntoor, Rama Chellappa
ACCV (2)2
2006 Multi-camera Tracking of Articulated Human Motion Using Motion and Shape Cues
Aravind Sundaresan, Rama Chellappa
ACCV (2)2
2006 Edge Suppression by Gradient Field Transformation Using Cross-Projection Tensors
abstract
We propose a new technique for edge-suppressing operations on images. We introduce cross projection tensors to achieve affine transformations of gradient fields. We use these tensors, for example, to remove edges in one image based on the edge-information in a second image. Traditionally, edge suppression is achieved by setting image gradients to zero based on thresholds. A common application is in the Retinex problem, where the illumination map is recovered by suppressing the reflectance edges, assuming it is slowly varying. We present a class of problems where edge-suppression can be a useful tool. These problems involve analyzing images of the same scene under variable illumination. Instead of resetting gradients, the key idea in our approach is to derive local tensors using one image and to transform the gradient field of another image using them. Reconstructed image from the modified gradient field shows suppressed edges or textures at the corresponding locations. All operations are local and our approach does not require any global analysis. We demonstrate the algorithm in the context of several applications such as (a) recovering the foreground layer under varying illumination, (b) estimating intrinsic images in non-Lambertian scenes, (c) removing shadows from color images and obtaining the illumination map, and (d) removing glass reflections.
Amit K. Agrawal, Ramesh Raskar, Rama Chellappa
CVPR (2)3
2006 Modeling Age Progression in Young Faces
abstract
We propose a craniofacial growth model that characterizes growth related shape variations observed in human faces during formative years. The model draws inspiration from the ‘revised’ cardioidal strain transformation model proposed in psychophysical studies related to craniofacial growth. The model takes into account anthropometric evidences collected on facial growth and hence is in accordance with the observed growth patterns in human faces across years. We characterize facial growth by means of growth parameters defined over facial landmarks often used in anthropometric studies. We illustrate how the age-based anthropometric constraints on facial proportions translate into linear and non-linear constraints on facial growth parameters and propose methods to compute the optimal growth parameters. The proposed craniofacial growth model can be used to predict one’s appearance across years and to perform face recognition across age progression. This is demonstrated on a database of age separated face images of individuals under 18 years of age.
Narayanan Ramanathan, Rama Chellappa
CVPR (1)2
2006 What Is the Range of Surface Reconstructions from a Gradient Field?
Amit K. Agrawal, Ramesh Raskar, Rama Chellappa
ECCV (1)3
2006 Video Mensuration Using a Stationary Camera
Feng Guo 0006, Rama Chellappa
ECCV (3)2
2006 An Adaptive Threshold Method for Hyperspectral Target Detection
abstract
In this paper, we present a new approach to automatically determine a detector threshold. This research problem is especially important in hyperspectral target detection as targets are typically very similar to the background. While a number of methods exist to determine the threshold, these methods require either large amounts of data or make simplifying assumptions about the background distribution. We use a method called inverse blind importance sampling which requires few samples and makes no a-priori assumptions about the background statistics. Results show the promise of this algorithm to determine thresholds for fixed false alarm densities in hyperspectral detectors
Joshua B. Broadwater, Rama Chellappa
ICASSP (5)2
2006 Motion Based Correspondence for 3D Tracking of Multiple Dim Objects
abstract
Tracking multiple objects in a video is a demanding task that is frequently encountered in several systems such as surveillance and motion analysis. Ability to track objects in 3D requires the use of multiple cameras. While tracking multiple objects using multiples video cameras, establishing correspondence between objects in the various cameras is a non-trivial task. Specifically, when the targets are dim or are very far away from the camera, appearance cannot be used in order to establish this correspondence. Here, we propose a technique to establish correspondence across cameras using the motion features extracted from the targets, even when the relative position of the cameras is unknown. Experimental results are provided for the problem of tracking multiple bees in natural flight using two cameras. The reconstructed 3D flight paths of the bees show some interesting flight patterns.
Ashok Veeraraghavan, Mandyam V. Srinivasan, Rama Chellappa, Emily Baird, Richard Lamont
ICASSP (2)3
2006 Invariant Geometric Representation of 3D Point Clouds for Registration and Matching
abstract
Though implicit representations of surfaces have often been used for various computer graphics tasks like modeling and morphing of objects, it has rarely been used for registration and matching of 3D point clouds. Unlike in graphics, where the goal is precise reconstruction, we use isosurfaces to derive a smooth and approximate representation of the underlying point cloud which helps in generalization. Implicit surfaces are generated using a variational interpolation technique. Implicit function values on a set of concentric spheres around the 3D point cloud of object are used as features for matching. Geometric-invariance is achieved by decomposing implicit values based feature set into various spherical harmonics. The decomposition provides a compact representation of 3D point clouds while achieving rotation invariance.
Soma Biswas, Gaurav Aggarwal, Rama Chellappa
ICIP3
2006 Recognition of Multi-Object Events Using Attribute Grammars
abstract
We present a method for representing and recognizing visual events using attribute grammars. In contrast to conventional grammars, attribute grammars are capable of describing features that are not easily represented by finite symbols. Our approach handles multiple concurrent events involving multiple entities by associating unique object identification labels with multiple event threads. Probabilistic parsing and probabilistic conditions on the attributes are used to achieve a robust recognition system. We demonstrate the effectiveness of our method for the task of recognizing vehicle casing in parking lots and events occurring in an airport tarmac.
Seong-Wook Joo, Rama Chellappa
ICIP2
2006 Stabilization and Mosaicing of Airborne Videos
abstract
We present an algorithm for stabilizing low-quality and low-resolution video sequences obtained from UAVs and MAVs flying over a predominantly planar terrain. The problem is important for the processing of videos captured from airborne platforms where most of the image region gives little information about the image-motion. The algorithm consists of approximately aligning the images using phase correlation, then refining the transformation parameters using available optical flow measurements and finally performing a minimization of the image difference by carefully selecting the image regions to make the iterations converge. In the second part of the paper, we describe an algorithm for camera ego-motion estimation by incorporating available metadata along with the video sequence. The availability of metadata such as IMU measurements provides a good initial solution for the camera orientation and helps the iterations converge faster by reducing the search space.
Mahesh Ramachandran, Rama Chellappa
ICIP2
2006 Shape-Regulated Particle Filtering for Tracking Non-Rigid Objects
abstract
This paper presents an active contour based algorithm for tracking non-rigid objects in heavily cluttered scenes. We decompose the non-rigid contour tracking problem into three subproblems: 2D motion estimation, deformation detection, and shape regulation. First, we employ a particle filter to estimate the affine transform parameters between successive frames. Second, by using a dynamic object model, we generate a probabilistic map of deformation to reshape its contour. Finally, we project the updated model onto a trained shape subspace to constrain deformations to be within possible object appearances. Our experiments show that the proposed algorithm significantly improves the performance of the tracker.
Jie Shao 0007, Rama Chellappa, Fatih Porikli
ICIP2
2006 Integrated Motion Detection and Tracking for Visual Surveillance
abstract
Visual surveillance systems have gained a lot of interest in the last few years. In this paper, we present a visual surveillance system that is based on the integration of motion detection and visual tracking to achieve better performance. Motion detection is achieved using an algorithm that combines temporal variance with background modeling methods. The tracking algorithm combines motion and appearance information into an appearance model and uses a particle filter framework for tracking the object in subsequent frames. The systems was tested on a large ground-truthed data set containing hundreds of color and FLIR image sequences. A performance evaluation for the system was performed and the average evaluation results are reported in this paper.
Mohamed F. Abdelkader, Rama Chellappa, Qinfen Zheng, LipChen Alex Chan
ICVS2
2006 Moving Object Verification from Airborne Video
abstract
This paper presents an end-to-end verification system for moving objects in airborne video. Lacking prior training data, the object information is collected on the fly from a short real-time learning sequence. Using a sample selection module, the system selects samples from the learning sequence and stores them in an exemplar database. To handle appearance change due to potentially large aspect angle variations, a homography-based view synthesis method is used to generate a novel view of each image in the exemplar database at the same pose as the query object in each frame of a query sequence. A spatial match score is obtained using a Distance Transform to compare the novel view and query object. After looping over all query frames, the set of match scores is passed to a temporal analysis module to examine the behavior of the query object, and calculate a final likelihood. Very good verification performance is achieved over thousands of trials for both color and infrared video sequences using the proposed system.
Zhanfeng Yue, Rama Chellappa, David Guarino
ICVS2
2006 Editorial
Aaron F. Bobick, Rama Chellappa, Larry Davis 0001
Int. J. Comput. Vis.2
2006 View Invariance for Human Action Recognition
Vasu Parameswaran, Rama Chellappa
Int. J. Comput. Vis.2
2006 From Sample Similarity to Ensemble Similarity: Probabilistic Distance Measures in Reproducing Kernel Hilbert Space
abstract
This paper addresses the problem of characterizing ensemble similarity from sample similarity in a principled manner. Using reproducing kernel as a characterization of sample similarity, we suggest a probabilistic distance measure in the reproducing kernel Hilbert space (RKHS) as the ensemble similarity. Assuming normality in the RKHS, we derive analytic expressions for probabilistic distance measures that are commonly used in many applications, such as Chernoff distance (or the Bhattacharyya distance as its special case), Kullback-Leibler divergence, etc. Since the reproducing kernel implicitly embeds a nonlinear mapping, our approach presents a new way to study these distances whose feasibility and efficiency is demonstrated using experiments with synthetic and real examples. Further, we extend the ensemble similarity to the reproducing kernel for ensemble and study the ensemble similarity for more general data representations.
Shaohua Kevin Zhou, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 Robust ego-motion estimation and 3-D model refinement using surface parallax
abstract
We present an iterative algorithm for robustly estimating the ego-motion and refining and updating a coarse depth map using parametric surface parallax models and brightness derivatives extracted from an image pair. Given a coarse depth map acquired by a range-finder or extracted from a digital elevation map (DEM), ego-motion is estimated by combining a global ego-motion constraint and a local brightness constancy constraint. Using the estimated camera motion and the available depth estimate, motion of the three-dimensional (3-D) points is compensated. We utilize the fact that the resulting surface parallax field is an epipolar field, and knowing its direction from the previous motion estimates, estimate its magnitude and use it to refine the depth map estimate. The parallax magnitude is estimated using a constant parallax model (CPM) which assumes a smooth parallax field and a depth based parallax model (DBPM), which models the parallax magnitude using the given depth map. We obtain confidence measures for determining the accuracy of the estimated depth values which are used to remove regions with potentially incorrect depth estimates for robustly estimating ego-motion in subsequent iterations. Experimental results using both synthetic and real data (both indoor and outdoor sequences) illustrate the effectiveness of the proposed algorithm.
Amit K. Agrawal, Rama Chellappa
IEEE Trans. Image Process.2
2006 Structure From Planar Motion
abstract
Planar motion is arguably the most dominant type of motion in surveillance videos. The constraints on motion lead to a simplified factorization method for structure from planar motion when using a stationary perspective camera. Compared with methods for general motion, our approach has two major advantages: a measurement matrix that fully exploits the motion constraints is formed such that the new measurement matrix has a rank of at most 3, instead of 4; the measurement matrix needs similar scalings, but the estimation of fundamental matrices or epipoles is not needed. Experimental results show that the algorithm is accurate and fairly robust to noise and inaccurate calibration. As the new measurement matrix is a nonlinear function of the observed variables, a different method is introduced to deal with the directional uncertainty in the observed variables. Differences and the dual relationship between planar motion and planar object are also clarified. Based on our method, a fully automated vehicle reconstruction system has been designed.
Jian Li 0022, Rama Chellappa
IEEE Trans. Image Process.2
2006 Face Verification Across Age Progression
abstract
Human faces undergo considerable amounts of varialions with aging. While face recognition systems have been proven to be sensitive to factors such as illumination and pose, their sensitivity to facial aging effects is yet to be studied. How does age progression affect the similarity between a pair of face images of an individual? What is the confidence associated with establishing the identity between a pair of age separated face images? In this paper, we develop a Bayesian age difference classifier that classifies face images of individuals based on age differences and performs face verification across age progression. Further, we study the similarity of faces across age progression. Since age separated face images invariably differ in illumination and pose, we propose preprocessing methods for minimizing such variations. Experimental results using a database comprising of pairs of face images that were retrieved from the passports of 465 individuals are presented. The verification system for faces separated by as many as nine years, attains an equal error rate of 8.5%.
Narayanan Ramanathan, Rama Chellappa
IEEE Trans. Image Process.2
2006 Principal Components Null Space Analysis for Image and Video Classification
abstract
We present a new classification algorithm, principal component null space analysis (PCNSA), which is designed for classification problems like object recognition where different classes have unequal and nonwhite noise covariance matrices. PCNSA first obtains a principal components subspace (PCA space) for the entire data. In this PCA space, it finds for each class "i," an Mi-dimensional subspace along which the class' intraclass variance is the smallest. We call this subspace an approximate null space (ANS) since the lowest variance is usually "much smaller" than the highest. A query is classified into class "i" if its distance from the class' mean in the class' ANS is a minimum. We derive upper bounds on classification error probability of PCNSA and use these expressions to compare classification performance of PCNSA with that of subspace linear discriminant analysis (SLDA). We propose a practical modification of PCNSA called progressive-PCNSA that also detects "new" (untrained classes). Finally, we provide an experimental comparison of PCNSA and progressive PCNSA with SLDA and PCA and also with other classification algorithms-linear SVMs, kernel PCA, kernel discriminant analysis, and kernel SLDA, for object recognition and face recognition under large pose/expression variation. We also show applications of PCNSA to two classification problems in video--an action retrieval problem and abnormal activity detection.
Namrata Vaswani, Rama Chellappa
IEEE Trans. Image Process.2
2005 Synthesis of Novel Views of Moving Objects in Airborne Video
abstract
This paper presents a method for synthesizing novel views of moving objects in airborne video. The object of interest is tracked using an appearance based visual tracking method. The on-object point correspondence is then built and used to estimate the homography induced by the ground plane. With known camera focal length, the surface normal to the ground plane and the camera motion between two views are factored out from the homography. In order to assure robustness of surface normal estimation, a rank one constraint is applied to decompose a matrix which contains the homographies from multiple frame pairs. Given a desired viewing direction, the novel image of the object is generated by warping the reference frame using the new homography between the desired viewpoint and the reference frame. Experimental results show that the method is robust and errors due to small depth variations of the object is negligible. 1
Zhanfeng Yue, Rama Chellappa
BMVC2
2005 Face Verification across Age Progression
abstract
Human faces undergo considerable amount of variations with aging. While studies have revealed the extent to which factors such as illumination variations, pose variations, facial expression and occlusions affect face recognition, the role of natural factors such as aging effects in affecting the same are yet to be studied. How does age progression affect the similarity between two images of an individual? What is the confidence associated with establishing the identity between two age separated face images of an individual? On a database of pairs of passport images, we study similarity of faces as a function time. We propose a Bayesian age-difference classifier that is built on a probabilistic eigenspaces framework. Since age separated face images invariably differ in illumination and have facial variations due to aging, we propose a method to overcome non uniform illumination across face images. The problem discussed in this paper has direct applications in passport renewal and homeland security.
Narayanan Ramanathan, Rama Chellappa
CVPR (2)2
2005 Moving Object Segmentation and Dynamic Scene Reconstruction Using Two Frames
abstract
In this paper, a two-frame approach is presented for segmentation of independent moving objects in video along with estimation of ego-motion, independent object motion and reconstruction of the dynamic scene using intensity images. The proposed method utilizes the least median of squares in estimating ego-motion and parallax constraints for segmenting independently moving objects. A 3D structure for the static scene is also estimated using surface parallax. The motion of moving objects is estimated by first fitting a parametric flow model followed by subspace analysis. The algorithm works well for unconstrained translational motion of moving objects.
Amit K. Agrawal, Rama Chellappa
ICASSP (2)2
2005 Interpretation of State Sequences in HMM for Activity Representation
abstract
We propose a method for activity representation based on semantic events, using the HMM framework. For every time instant, the probability of event occurrence is computed by exploring a subset of state sequences. The idea is that while activity trajectories may have large variations at the data or the state levels, they may exhibit similarities at the event level. Our experiments show the application of these events to activity recognition in an office environment and to anomalous trajectory detection using surveillance video data.
Naresh P. Cuntoor, Bayya Yegnanarayana, Rama Chellappa
ICASSP (2)3
2005 A method for converting a smiling face to a neutral face with applications to face recognition
abstract
The human face displays a variety of expressions, like smile, sorrow, surprise, etc. All these expressions constitute nonrigid motions of various features of the face. These expressions lead to a significant change in the appearance of a facial image which leads to a drop in the recognition accuracy of a face-recognition system trained with neutral faces. There are other factors like pose and illumination which also lead to performance drops. Researchers have proposed methods to tackle the effects of pose and illumination; however, there has been little work on how to tackle expressions. We attempt to address the issue of expression invariant face-recognition. We present preprocessing steps for converting a smiling face to a neutral face. We expect that this would in turn make the vector in the feature space to be closer to the correct vector in the gallery, in an appearance-based face recognition. This conjecture is supported by our recognition results which demonstrate that the accuracy goes up if we include the expression-normalization block.
Mahesh Ramachandran, Shaohua Kevin Zhou, Divya Jhalani, Rama Chellappa
ICASSP (2)4