VLDB 2026 Research / reviewers in the wild / expert
Amit K. Roy-Chowdhury
dblp:c/AmitKRoyChowdhury · also Amit Roy-Chowdhury 0001, Amit Roychowdhury 0001
· DBLP profile ↗
224ranked-venue papers
12as first author
60since 2021 · last 2026
0000-0001-6690-9725ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 170 · 9 first-author · 40 since 2021Artificial intelligence and machine learning · 107 · 4 first-author · 37 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 since 2021Computer networks · 5 · 1 since 2021Systems, architecture and hardware · 3Security and privacy · 2Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Visibility guided Self-Supervised Occlusion-Resilient Human Pose EstimationabstractOcclusion remains a significant challenge for existing human pose estimation algorithms, often resulting in inaccurate and anatomically implausible predictions. Although recent occlusion-robust methods report strong performance, they typically rely heavily on supervised learning and privileged information, such as multiview data or temporal sequences. Furthermore, these models often fail under domain changes. Domain-adaptive human pose estimation seeks to mitigate this issue; however, when occlusions are present in the target domain, a common occurrence in real-world applications, performance of these algorithms deteriorates significantly. To address these challenges, we propose VisOR, a novel Visibility guided Self-Supervised algorithm for Occlusion-Resilient Human Pose Estimation. VisOR achieves robustness to both domain shifts and occlusions by integrating contextual reasoning with iterative pseudo-label refinement. It mitigates the overfitting to noisy labels from occluded regions via a visibility-driven curriculum learning strategy, which progressively introduces the model to increasingly occluded training samples. Additionally, VisOR is regularized by a learned human pose prior that maintains anatomical plausibility throughout the adaptation process. Recognizing the scarcity of human pose datasets with realistic occlusions, we introduce BOW Blended Occlusions in-the-Wild, a rigorously constructed context-aware synthetic benchmark designed to evaluate the occlusion resilience of human pose estimation algorithms. BOW offers a diverse range of context-aware occlusions across both indoor and outdoor environments, simulating real-world conditions. Through extensive experiments, we demonstrate that VisOR outperforms current state-of-the-art methods by ∼ 7% in challenging occluded human pose estimation benchmarks and provides a baseline performance on BOW, against existing algorithms. Arindam Dutta, Sarosij Bose, Rohit Kundu, Calvin-Khang Ta, Saketh Bachu, Konstantinos Karydis, Amit K. Roy-Chowdhury |
WACV | 7 |
| 2026 | Pose Guided Unsupervised Domain Adaptation for Human Body Part SegmentationabstractExisting algorithms for human body part segmentation have shown promising results on challenging datasets, primarily relying on end-to-end supervision. However, these algorithms exhibit severe performance drops in the face of domain shifts, leading to inaccurate segmentation masks. To tackle this issue, we introduce POSTURE: Pose Guided Unsupervised Domain Adaptation for Human Body Part Segmentation - an innovative pseudo-labelling approach 0designed to improve segmentation performance on the unlabeled target data. Distinct from conventional domain adaptive methods for general semantic segmentation, POSTURE stands out by considering the underlying structure of the human body and uses anatomical guidance from pose keypoints to drive the adaptation process. This strong inductive prior translates to impressive performance improvements, averaging 8% over existing state-of-the-art domain adaptive semantic segmentation methods across three benchmark datasets. Furthermore, the inherent flexibility of our proposed approach facilitates seamless extension to source-free settings (SF-POSTURE), effectively mitigating potential privacy and computational concerns, with negligible drop in performance. Arindam Dutta, Rohit Lal, Yash Garg, Calvin-Khang Ta, Dripta S. Raychaudhuri, Amit K. Roy-Chowdhury |
IEEE Trans. Image Process. | 6 |
| 2025 | Provable Benefits of Task-Specific Prompts for In-context LearningabstractThe in-context learning capabilities of modern language models have motivated a deeper mathematical understanding of sequence models. A line of recent work has shown that linear attention models can emulate projected gradient descent iterations to implicitly learn the task vector from the data provided in the context window. In this work, we consider a novel setting where the global task distribution can be partitioned into a union of conditional task distributions. We then examine the use of task-specific prompts and prediction heads for learning the prior information associated with the conditional task distribution using a one-layer attention model. Our results on loss landscape show that task-specific prompts facilitate a covariance-mean decoupling where prompt-tuning explains the conditional mean of the distribution whereas the variance is learned/explained through in-context learning. Incorporating task-specific head further aids this process by entirely decoupling estimation of mean and variance components. This covariance-mean perspective similarly explains how jointly training prompt and attention weights can provably help over fine-tuning after pretraining. Xiangyu Chang, Yingcong Li, Muti Kara, Samet Oymak, Amit K. Roy-Chowdhury |
AISTATS | 5 |
| 2025 | Towards Source-Free Machine UnlearningabstractAs machine learning becomes more pervasive and data privacy regulations evolve, the ability to remove private or copyrighted information from trained models is becoming an increasingly critical requirement. Existing unlearning methods often rely on the assumption of having access to the entire training dataset during the forgetting process. However, this assumption may not hold true in practical scenarios where the original training data may not be accessible, i.e., the source-free setting. To address this challenge, we focus on the source-free unlearning scenario, where an unlearning algorithm must be capable of removing specific data from a trained model without requiring access to the original training dataset. Building on recent work, we present a method that can estimate the Hessian of the unknown remaining training data, a crucial component required for efficient unlearning. Leveraging this estimation technique, our method enables efficient zero-shot unlearning while providing robust theoretical guarantees on the unlearning performance, while maintaining performance on the remaining data. Extensive experiments over a wide range of datasets verify the efficacy of our method. Sk Miraj Ahmed, Umit Yigit Basaran, Dripta S. Raychaudhuri, Arindam Dutta, Rohit Kundu, Fahim Faisal Niloy, Basak Guler, Amit K. Roy-Chowdhury |
CVPR | 8 |
| 2025 | AdMiT: Adaptive Multi-Source Tuning in Dynamic EnvironmentsabstractIncorporating transformer models into edge devices poses a significant challenge due to the computational demands of adapting these large models across diverse applications. Parameter-efficient tuning (PET) methods (e.g. LoRA, Adapter, Visual Prompt Tuning, etc.) allow for targeted adaptation by modifying only small parts of the transformer model. However, adapting to dynamic unlabeled target distributions at the test time remains complex. To address this, we introduce AdMiT: Adaptive Multi-Source Tuning in Dynamic Environments. AdMiT innovates by pre-training a set of PET modules, each optimized for different source distributions or tasks, and dynamically selecting and integrating a sparse subset of relevant modules when encountering a new, few-shot, unlabeled target distribution. This integration leverages Kernel Mean Embedding (KME)-based matching to align the target distribution with relevant source knowledge efficiently, without requiring additional routing networks or hyperparameter tuning. AdMiT achieves adaptation with a single inference step, making it particularly suitable for resource-constrained edge deployments. Furthermore, AdMiT preserves privacy by performing an adaptation locally on each edge device, without the need for data exchange. Our theoretical analysis establishes guarantees for AdMiT’s generalization, while extensive benchmarks demonstrate that AdMiT consistently outperforms other PET methods across a range of tasks, achieving robust and efficient adaptation. Xiangyu Chang, Fahim Faisal Niloy, Sk Miraj Ahmed, Srikanth V. Krishnamurthy, Basak Guler, Ananthram Swami, Samet Oymak, Amit K. Roy-Chowdhury |
CVPR | 8 |
| 2025 | Towards a Universal Synthetic Video Detector: From Face or Background Manipulations to Fully AI-Generated ContentabstractExisting DeepFake detection techniques primarily focus on facial manipulations, such as face-swapping or lip-syncing. However, advancements in text-to-video (T2V) and image-to-video (I2V) generative models now allow fully AI-generated synthetic content and seamless background alterations, challenging face-centric detection methods and demanding more versatile approaches.To address this, we introduce the Universal Network for Identifying Tampered and synthEtic videos (UNITE) model, which, unlike traditional detectors, captures full-frame manipulations. UNITE extends detection capabilities to scenarios without faces, non-human subjects, and complex background modifications. It leverages a transformer-based architecture that processes domain-agnostic features extracted from videos via the SigLIP-So400M foundation model. Given limited datasets encompassing both facial/background alterations and T2V/I2V content, we integrate task-irrelevant data alongside standard DeepFake datasets in training. We further mitigate the model’s tendency to over-focus on faces by incorporating an attention-diversity (AD) loss, which promotes diverse spatial attention across video frames. Combining AD loss with cross-entropy improves detection performance across varied contexts. Comparative evaluations demonstrate that UNITE outperforms state-of-the-art detectors on datasets featuring face/background manipulations and fully synthetic T2V/I2V videos, showcasing its adaptability and generalizable detection capabilities. Rohit Kundu, Vishal Mohanty, Athula Balachandran, Amit K. Roy-Chowdhury |
CVPR | 5 |
| 2025 | Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph GenerationabstractScene Graph Generation (SGG) aims to represent visual scenes by identifying objects and their pairwise relationships, providing a structured understanding of image content. However, inherent challenges like long-tailed class distributions and prediction variability necessitate uncertainty quantification in SGG for its practical viability. In this paper, we introduce a novel Conformal Prediction based framework, adaptive to any existing SGG method, for quantifying their predictive uncertainty by constructing well-calibrated prediction sets over their generated scene graphs. These scene graph prediction sets are designed to achieve statistically rigorous coverage guarantees under exchangeability assumptions. Additionally, to ensure the prediction sets contain the most practically interpretable scene graphs, we propose an effective MLLM-based post-processing strategy for selecting the most visually and semantically plausible scene graphs within each set. We show that our proposed approach can produce diverse possible scene graphs from an image, assess the reliability of SGG methods, and improve overall SGG performance. Sayak Nag, Udita Ghosh, Calvin-Khang Ta, Sarosij Bose, Amit K. Roy-Chowdhury |
CVPR | 6 |
| 2025 | Gradient Inversion Attacks on Parameter-Efficient Fine-TuningabstractFederated learning (FL) allows multiple data-owners to collaboratively train machine learning models by exchanging local gradients, while keeping their private data on-device. To simultaneously enhance privacy and training efficiency, recently parameter-efficient fine-tuning (PEFT) of large-scale pretrained models has gained substantial attention in FL. While keeping a pretrained (backbone) model frozen, each user fine-tunes only a few lightweight modules to be used in conjunction, to fit specific downstream applications. Accordingly, only the gradients with respect to these lightweight modules are shared with the server. In this work, we investigate how the privacy of the fine-tuning data of the users can be compromised via a malicious design of the pretrained model and trainable adapter modules. We demonstrate gradient inversion attacks on a popular PEFT mechanism, the adapter, which allow an attacker to reconstruct local data samples of a target user, using only the accessible adapter gradients. Via extensive experiments, we demonstrate that a large batch of fine-tuning images can be retrieved with high fidelity. Our attack highlights the need for privacy-preserving mechanisms for PEFT, while opening up several future directions. Our code is available at https://github.com/info-ucr/PEFTLeak. Hasin Us Sami, Swapneel Sen, Amit K. Roy-Chowdhury, Srikanth V. Krishnamurthy, Basak Guler |
CVPR | 3 |
| 2025 | FedBand: Adaptive Federated Learning Under Strict Bandwidth ConstraintsabstractFederated Learning (FL) enables model training across decentralized clients while preserving data privacy. However, bandwidth constraints limit the volume of information exchanged, making communication efficiency a critical challenge. In addition, non-IID data distributions require fairness-aware mechanisms to prevent performance degradation for certain clients. Existing sparsification techniques often apply fixed compression ratios uniformly, ignoring variations in client importance and bandwidth. We propose Fed-Band, a dynamic bandwidth allocation framework that prioritizes clients based on their contribution to the global model. Unlike conventional approaches, FedBand does not enforce uniform client participation in every communication round. Instead, it allocates more bandwidth to clients whose local updates deviate significantly from the global model, enabling them to transmit a greater number of parameters. Clients with less impactful updates contribute proportionally less or may defer transmission, reducing unnecessary overhead while maintaining generalizability. By optimizing the trade-off between communication efficiency and learning performance, FedBand substantially reduces transmission costs while preserving model accuracy. Experiments on non-IID CIFAR-10 and UTMobileNet2021 datasets, demonstrate that FedBand achieves up to 99.81% bandwidth savings per round while maintaining accuracies close to that of an unsparsified model (80% on CIFAR-10, 95% on UTMobileNet), despite transmitting less than 1% of the model parameters in each round. Moreover, FedBand accelerates convergence by 37.4%, further improving learning efficiency under bandwidth constraints. Mininet emulations further show a 42.6% reduction in communication costs and a 65.57% acceleration in convergence compared to baseline methods, validating its real-world efficiency. These results demonstrate that adaptive bandwidth allocation can significantly enhance the scalability and communication efficiency of federated learning, making it more viable for real-world, bandwidth-constrained networking environments. Taghreed Alanazi, Abdulrahman Fahim, Muntaka Ibnath, Basak Guler, Amit K. Roy-Chowdhury, Ananthram Swami, Evangelos E. Papalexakis, Srikanth V. Krishnamurthy |
ICCCN | 5 |
| 2025 | Uncertainty-Aware Diffusion-Guided Refinement of 3D ScenesabstractReconstructing 3D scenes from a single image is a fundamentally ill-posed task due to the severely under-constrained nature of the problem. Consequently, when the scene is rendered from novel camera views, existing single image to 3D reconstruction methods render incoherent and blurry views. This problem is exacerbated when the unseen regions are far away from the input camera. In this work, we address these inherent limitations in existing single image-to-3D scene feedforward networks. To alleviate the poor performance due to insufficient information beyond the input image's view, we leverage a strong generative prior in the form of a pre-trained latent video diffusion model, for iterative refinement of a coarse scene represented by optimizable Gaussian parameters. To ensure that the style and texture of the generated images align with that of the input image, we incorporate on-the-fly Fourier-style transfer between the generated images and the input image. Additionally, we design a semantic uncertainty quantification module that calculates the per-pixel entropy and yields uncertainty maps used to guide the refinement process from the most confident pixels while discarding the remaining highly uncertain ones. We conduct extensive experiments on real-world scene datasets, including in-domain RealEstate-10K and out-of-domain KITTI-v2, showing that our approach can provide more realistic and high-fidelity novel view synthesis results compared to existing state-of-the-art methods. Sarosij Bose, Arindam Dutta, Sayak Nag, Junge Zhang, Konstantinos Karydis, Amit K. Roy-Chowdhury |
ICCV | 7 |
| 2025 | CHROME: Clothed Human Reconstruction with Occlusion-Resilience and Multiview-Consistency from a Single ImageabstractReconstructing clothed humans from a single image is a fundamental task in computer vision with wide-ranging applications. Although existing monocular clothed human reconstruction solutions have shown promising results, they often rely on the assumption that the human subject is in an occlusion-free environment. Thus, when encountering in-the-wild occluded images, these algorithms produce multiview inconsistent and fragmented reconstructions. Additionally, most algorithms for monocular 3D human reconstruction leverage geometric priors such as SMPL annotations for training and inference, which are extremely challenging to acquire in real-world applications. To address these limitations, we propose CHROME: Clothed Human Reconstruction with Occlusion-Resilience and Multiview-ConsistEncy from a Single Image, a novel pipeline designed to reconstruct occlusion-resilient 3D humans with multiview consistency from a single occluded image, without requiring either ground-truth geometric prior annotations or 3D supervision. Specifically, CHROME leverages a multiview diffusion model to first synthesize occlusion-free human images from the occluded input, compatible with off-the-shelf pose control to explicitly enforce cross-view consistency during synthesis. A 3D reconstruction model is then trained to predict a set of 3D Gaussians conditioned on both the occluded input and synthesized views, aligning cross-view details to produce a cohesive and accurate 3D representation. CHROME achieves significant improvements in terms of both novel view synthesis (upto 3 db PSNR) and geometric reconstruction under challenging conditions. Arindam Dutta, Meng Zheng 0002, Zhongpai Gao, Benjamin Planche, Anwesa Choudhuri, Terrence Chen, Amit K. Roy-Chowdhury, Ziyan Wu 0001 |
ICCV | 7 |
| 2025 | VOccl3D: A Video Benchmark Dataset for 3D Human Pose and Shape Estimation Under Real Occlusions
Yash Garg, Saketh Bachu, Arindam Dutta, Rohit Lal, Sarosij Bose, Calvin-Khang Ta, Muhammad Salman Asif, Amit K. Roy-Chowdhury |
ICCV | 8 |
| 2025 | Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language ModelsabstractVision-language models (VLMs) have improved significantly in their capabilities, but their complex architecture makes their safety alignment challenging. In this paper, we reveal an uneven distribution of harmful information across the intermediate layers of the image encoder and show that skipping a certain set of layers and exiting early can increase the chance of the VLM generating harmful responses. We call it as “Image enCoder Early-exiT” based vulnerability (ICET). Our experiments across three VLMs: LLaVA-1.5, LLaVA-NeXT, and Llama 3.2 show that performing early exits from the image encoder significantly increases the likelihood of generating harmful outputs. To tackle this, we propose a simple yet effective modification of the Clipped-Proximal Policy Optimization (Clip-PPO) algorithm for performing layer-wise multi-modal RLHF for VLMs. We term this as Layer-Wise PPO (L-PPO). We evaluate our L-PPO algorithm across three multi-modal datasets and show that it consistently reduces the harmfulness caused by early exits. Saketh Bachu, Erfan Shayegani, Rohit Lal, Trishna Chakraborty, Arindam Dutta, Chengyu Song, Yue Dong 0002, Nael B. Abu-Ghazaleh, Amit K. Roy-Chowdhury |
ICML | 9 |
| 2025 | A Certified Unlearning Approach without Access to Source DataabstractWith the growing adoption of data privacy regulations, the ability to erase private or copyrighted information from trained models has become a crucial requirement. Traditional unlearning methods often assume access to the complete training dataset, which is unrealistic in scenarios where the source data is no longer available. To address this challenge, we propose a certified unlearning framework that enables effective data removal without access to the original training data samples. Our approach utilizes a surrogate dataset that approximates the statistical properties of the source data, allowing for controlled noise scaling based on the statistical distance between the two. While our theoretical guarantees assume knowledge of the exact statistical distance, practical implementations typically approximate this distance, resulting in potentially weaker but still meaningful privacy guarantees. This ensures strong guarantees on the model’s behavior post-unlearning while maintaining its overall utility. We establish theoretical bounds, introduce practical noise calibration techniques, and validate our method through extensive experiments on both synthetic and real-world datasets. The results demonstrate the effectiveness and reliability of our approach in privacy-sensitive settings. Umit Yigit Basaran, Sk Miraj Ahmed, Amit K. Roy-Chowdhury, Basak Guler |
ICML | 3 |
| 2025 | ODES: Online Domain Adaptation with Expert Guidance for Medical Image Segmentation
Md Shazid Islam, Sayak Nag, Arindam Dutta, Sk Miraj Ahmed, Fahim Faisal Niloy, Shreyangshu Bera, Amit K. Roy-Chowdhury |
MICCAI (4) | 7 |
| 2025 | When and How Unlabeled Data Provably Improve In-Context LearningabstractRecent research shows that in-context learning (ICL) can be effective even when demonstrations have missing or incorrect labels. To shed light on this capability, we examine a canonical setting where the demonstrations are drawn according to a binary Gaussian mixture model (GMM) and a certain fraction of the demonstrations have missing labels. We provide a comprehensive theoretical study to show that: (1) The loss landscape of one-layer linear attention models recover the optimal fully-supervised estimator but completely fail to exploit unlabeled data; (2) In contrast, multilayer or looped transformers can effectively leverage unlabeled data by implicitly constructing estimators of the form $\sum_{i\ge 0} a_i (X^\top X)^iX^\top y$ with $X$ and $y$ denoting features and partially-observed labels (with missing entries set to zero). We characterize the class of polynomials that can be expressed as a function of depth and draw connections to Expectation Maximization, an iterative pseudo-labeling algorithm commonly used in semi-supervised learning. Importantly, the leading polynomial power is exponential in depth, so mild amount of depth/looping suffices. As an application of theory, we propose looping off-the-shelf tabular foundation models to enhance their semi-supervision capabilities. Extensive evaluations on real-world datasets show that our method significantly improves the semisupervised tabular learning performance over the standard single pass inference. Yingcong Li, Xiangyu Chang, Muti Kara, Amit K. Roy-Chowdhury, Samet Oymak |
NeurIPS | 5 |
| 2025 | iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video ReasoningabstractGrounding large language models (LLMs) in domain-specific tasks like post-hoc dash-cam driving video analysis is challenging due to their general-purpose training and lack of structured inductive biases. As vision is often the sole modality available for such analysis (i.e., no LiDAR, GPS, etc.), existing video-based vision-language models (V-VLMs) struggle with spatial reasoning, causal inference, and explainability of events in the input video. To this end, we introduce iFinder, a structured semantic grounding framework that decouples perception from reasoning by translating dash-cam videos into a hierarchical, interpretable data structure for LLMs. iFinder operates as a modular, training-free pipeline that employs pretrained vision models to extract critical cues—object pose, lane positions, and object trajectories—which are hierarchically organized into frame- and video-level structures. Combined with a three-block prompting strategy, it enables step-wise, grounded reasoning for the LLM to refine a peer V-VLM's outputs and provide accurate reasoning.
Evaluations on four public dash-cam video benchmarks show that iFinder's proposed grounding with domain-specific cues—especially object orientation and global context—significantly outperforms end-to-end V-VLMs on four zero-shot driving benchmarks, with up to 39% gains in accident reasoning accuracy. By grounding LLMs with driving domain-specific representations, iFinder offers a zero-shot, interpretable, and reliable alternative to end-to-end V-VLMs for post-hoc driving video understanding. Manyi Yao, Bingbing Zhuang, Sparsh Garg, Amit K. Roy-Chowdhury, Christian R. Shelton, Manmohan Krishna Chandraker, Abhishek Aich |
NeurIPS | 4 |
| 2025 | STRIDE: Single-Video Based Temporally Continuous Occlusion-Robust 3D Pose EstimationabstractAccurately estimating 3D human poses is crucial for fields like action recognition, gait recognition, and virtual/augmented reality. However, predicting human poses under severe occlusion remains a persistent and significant challenge. Existing image-based estimators struggle with heavy occlusions due to a lack of temporal context, resulting in inconsistent predictions, while video-based models, despite benefiting from temporal data, face limitations with prolonged occlusions over multiple frames. Additionally, existing algorithms often struggle to generalize unseen videos. Addressing these challenges, we propose STRIDE (Single-video based TempoRally contInuous Occlusion-Robust 3D Pose Estimation), a novel Test-Time Training (TTT) approach to fit a human motion prior for estimating 3D human poses for each video. Our proposed approach handles occlusions not encountered during the model's training by refining a sequence of noisy initial pose estimates into accurate, temporally coherent poses at test time, effectively overcoming the limitations of existing methods. Our flexible, model-agnostic framework allows us to use any off-the-shelf 3D pose estimation method to improve robustness and temporal consistency. We validate STRIDE's efficacy through comprehensive experiments on multiple challenging datasets where it not only outperforms existing single-image and video-based pose estimation models but also showcases superior handling of substantial occlusions, achieving fast, robust, accurate, and temporally consistent 3D pose estimates. Code is made publicly available at https://github.com/take2rohit/stride Rohit Lal, Saketh Bachu, Yash Garg, Arindam Dutta, Calvin-Khang Ta, Hannah Dela Cruz, Dripta S. Raychaudhuri, Muhammad Salman Asif, Amit K. Roy-Chowdhury |
WACV | 9 |
| 2025 | Egocentric and exocentric methods: A short survey
Anirudh Thatipelli, Shao-Yuan Lo, Amit K. Roy-Chowdhury |
Comput. Vis. Image Underst. | 3 |
| 2025 | DEGAST3D: Learning Deformable 3D Graph Similarity to Track Plant Cells in Unregistered Time Lapse ImagesabstractTracking plant cells in three-dimensional (3D) tissue captured through light microscopy presents significant challenges due to the large number of densely packed cells, non-uniform growth patterns, and variations in cell division planes across different cell layers. In addition, images of deeper tissue layers are often noisy, and systemic imaging errors further exacerbate the complexity of the task. In this paper, we propose a novel learning-based method DEGAST3D: Learning Deformable 3D GrAph Similarity to Track Plant Cells in Unregistered Time Lapse Images exploits the tightly packed 3D cell structure of plant cells to create a three-dimensional graph for accurate cell tracking. We also propose a novel algorithm for cell division detection and an effective three-dimensional registration, improving state-of-the-art algorithms. On a public dataset, our novel cell pair matching method outperforms the baseline by $6.83 \%$, $5.96 \%$, $6.40 \%$ in precision, recall, and F-1 score, respectively. On the same dataset, our proposed novel cell division technique improves the results of the baseline method by $15.38 \%$ and $14.78 \%$ in terms of recall and F1-score, respectively. Md Shazid Islam, Arindam Dutta, Calvin-Khang Ta, Kevin Rodriguez, Christian Michael, Mark S. Alber, G. Venugopala Reddy, Amit K. Roy-Chowdhury |
IEEE Trans. Comput. Biol. Bioinform. | 8 |
| 2024 | Selective Attention: Enhancing Transformer through Principled Context ControlabstractThe attention mechanism within the transformer architecture enables the model to weigh and combine tokens based on their relevance to the query. While self-attention has enjoyed major success, it notably treats all queries $q$ in the same way by applying the mapping $V^\top\text{softmax}(Kq)$, where $V,K$ are the value and key embeddings respectively. In this work, we argue that this uniform treatment hinders the ability to control contextual sparsity and relevance. As a solution, we introduce the Selective Self-Attention (SSA) layer that augments the softmax nonlinearity with a principled temperature scaling strategy. By controlling temperature, SSA adapts the contextual sparsity of the attention map to the query embedding and its position in the context window. Through theory and experiments, we demonstrate that this alleviates attention dilution, aids the optimization process, and enhances the model's ability to control softmax spikiness of individual queries. We also incorporate temperature scaling for value embeddings and show that it boosts the model's ability to suppress irrelevant/noisy tokens. Notably, SSA is a lightweight method which introduces less than 0.5\% new parameters through a weight-sharing strategy and can be fine-tuned on existing LLMs. Extensive empirical evaluations demonstrate that SSA-equipped models achieve a noticeable and consistent accuracy improvement on language modeling benchmarks. Xuechen Zhang 0002, Xiangyu Chang, Amit K. Roy-Chowdhury, Jiasi Chen, Samet Oymak |
NeurIPS | 4 |
| 2024 | CONTRAST: Continual Multi-source Adaptation to Dynamic DistributionsabstractAdapting to dynamic data distributions is a practical yet challenging task. One effective strategy is to use a model ensemble, which leverages the diverse expertise of different models to transfer knowledge to evolving data distributions. However, this approach faces difficulties when the dynamic test distribution is available only in small batches and without access to the original source data. To address the challenge of adapting to dynamic distributions in such practical settings, we propose continual multi-source adaptation to dynamic distributions (CONTRAST), a novel method that optimally combines multiple source models to adapt to the dynamic test data. CONTRAST has two distinguishing features. First, it efficiently computes the optimal combination weights to combine the source models to adapt to the test data distribution continuously as a function of time. Second, it identifies which of the source model parameters to update so that only the model which is most correlated to the target data is adapted, leaving the less correlated ones untouched; this mitigates the issue of ``forgetting" the source model parameters by focusing only on the source model that exhibits the strongest correlation with the test batch distribution. Through theoretical analysis we show that the proposed method is able to optimally combine the source models and prioritize updates to the model least prone to forgetting. Experimental analysis on diverse datasets demonstrates that the combination of multiple source models does at least as well as the best source (with hindsight knowledge), and performance does not degrade as the test data distribution changes over time (robust to forgetting). Sk Miraj Ahmed, Fahim Faisal Niloy, Xiangyu Chang, Dripta S. Raychaudhuri, Samet Oymak, Amit K. Roy-Chowdhury |
NeurIPS | 6 |
| 2024 | POISE: Pose Guided Human Silhouette Extraction under OcclusionsabstractHuman silhouette extraction is a fundamental task in computer vision with applications in various downstream tasks. However, occlusions pose a significant challenge, leading to incomplete and distorted silhouettes. To address this challenge, we introduce POISE: Pose Guided Human Silhouette Extraction under Occlusions, a novel self-supervised fusion framework that enhances accuracy and robustness in human silhouette prediction. By combining initial silhouette estimates from a segmentation model with human joint predictions from a 2D pose estimation model, POISE leverages the complementary strengths of both approaches, effectively integrating precise body shape information and spatial information to tackle occlusions. Furthermore, the self-supervised nature of POISE eliminates the need for costly annotations, making it scalable and practical. Extensive experimental results demonstrate its superiority in improving silhouette extraction under occlusions, with promising results in downstream tasks such as gait recognition. The code for our method is available https://github.com/take2rohit/poise. Arindam Dutta, Rohit Lal, Dripta S. Raychaudhuri, Calvin-Khang Ta, Amit K. Roy-Chowdhury |
WACV | 5 |
| 2024 | Effective Restoration of Source Knowledge in Continual Test Time AdaptationabstractTraditional test-time adaptation (TTA) methods face significant challenges in adapting to dynamic environments characterized by continuously changing long-term target distributions. These challenges primarily stem from two factors: catastrophic forgetting of previously learned valuable source knowledge and gradual error accumulation caused by miscalibrated pseudo labels. To address these issues, this paper introduces an unsupervised domain change detection method that is capable of identifying domain shifts in dynamic environments and subsequently resets the model parameters to the original source pre-trained values. By restoring the knowledge from the source, it effectively corrects the negative consequences arising from the gradual deterioration of model parameters caused by ongoing shifts in the domain. Our method involves progressive estimation of global batch-norm statistics specific to each domain, while keeping track of changes in the statistics triggered by domain shifts. Importantly, our method is agnostic to the specific adaptation technique employed and thus, can be incorporated to existing TTA methods to enhance their performance in dynamic environments. We perform extensive experiments on benchmark datasets to demonstrate the superior performance of our method compared to state-of-the-art adaptation methods. Fahim Faisal Niloy, Sk Miraj Ahmed, Dripta S. Raychaudhuri, Samet Oymak, Amit K. Roy-Chowdhury |
WACV | 5 |
| 2024 | A Two-Stage Noise-Tolerant Paradigm for Label Corrupted Person Re-IdentificationabstractSupervised person re-identification (Re-ID) approaches are sensitive to label corrupted data, which is inevitable and generally ignored in the field of person Re-ID. In this paper, we propose a two-stage noise-tolerant paradigm (TSNT) for labeling corrupted person Re-ID. Specifically, at stage one, we present a self-refining strategy to separately train each network in TSNT by concentrating more on pure samples. These pure samples are progressively refurbished via mining the consistency between annotations and predictions. To enhance the tolerance of TSNT to noisy labels, at stage two, we employ a co-training strategy to collaboratively supervise the learning of the two networks. Concretely, a rectified cross-entropy loss is proposed to learn the mutual information from the peer network by assigning large weights to the refurbished reliable samples. Moreover, a noise-robust triplet loss is formulated for further improving the robustness of TSNT by increasing inter-class distances and reducing intra-class distances in the label-corrupted dataset, where a constraint condition for reliability discrimination is carefully designed to select reliable triplets. Extensive experiments demonstrate the superiority of TSNT, for instance, on the Market1501 dataset, our paradigm achieves 90.3% rank-1 accuracy (6.2% improvement over the state-of-the-art method) under noise ratio 20%. Min Liu 0008, Fei Wang 0124, Yaonan Wang 0001, Amit K. Roy-Chowdhury |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Collaborative Multi-Agent Video Fast-ForwardingabstractMulti-agent applications have recently gained significant popularity. In many computer vision tasks, a network of agents, such as a team of robots with cameras, could work collaboratively to perceive the environment for efficient and accurate situation awareness. However, these agents often have limited computation, communication, and storage resources. Thus, reducing resource consumption while still providing an accurate perception of the environment becomes an important goal when deploying multi-agent systems. To achieve this goal, we identify and leverage the overlap among different camera views in multi-agent systems for reducing the processing, transmission and storage of redundant/unimportant video frames. Specifically, we have developed two collaborative multi-agent video fast-forwarding frameworks in distributed and centralized settings, respectively. In these frameworks, each individual agent can selectively process or skip video frames at adjustable paces based on multiple strategies via reinforcement learning. Multiple agents then collaboratively sense the environment via either 1) a consensus-based distributed framework calledDMVFthat periodically updates the fast-forwarding strategies of agents by establishing communication and consensus among connected neighbors, or 2) a centralized framework calledMFFNetthat utilizes a central controller to decide the fast-forwarding strategies for agents based on collected data. We demonstrate the efficacy and efficiency of our proposed frameworks on a real-world surveillance video dataset VideoWeb and a new simulated driving dataset CarlaSim, through extensive simulations and deployment on an embedded platform with TCP communication. We show that compared with other approaches in the literature, our frameworks achieve better coverage of important frames, while significantly reducing the number of frames processed at each agent. Shuyue Lan, Zhilu Wang, Ermin Wei, Amit K. Roy-Chowdhury, Qi Zhu 0002 |
IEEE Trans. Multim. | 4 |
| 2023 | Unbiased Scene Graph Generation in VideosabstractThe task of dynamic scene graph generation (SGG) from videos is complicated and challenging due to the inherent dynamics of a scene, temporal fluctuation of model predictions, and the long-tailed distribution of the visual relationships in addition to the already existing challenges in image-based SGG. Existing methods for dynamic SGG have primarily focused on capturing spatio-temporal context using complex architectures without addressing the challenges mentioned above, especially the long-tailed distribution of relationships. This often leads to the generation of biased scene graphs. To address these challenges, we introduce a new framework called TEMPURA: TEmporal consistency and Memory Prototype guided UnceR-tainty Attenuation for unbiased dynamic SGG. TEMPURA employs object-level temporal consistencies via transformer-based sequence modeling, learns to synthesize unbiased relationship representations using memory-guided training, and attenuates the predictive uncertainty of visual relations using a Gaussian Mixture Model (GMM). Extensive experiments demonstrate that our method achieves significant (up to 10% in some cases) performance gain over existing methods highlighting its superiority in generating more unbiased scene graphs. Code: https://github.com/sayaknag/unbiasedSGG.git Sayak Nag, Kyle Min 0001, Subarna Tripathi, Amit K. Roy-Chowdhury |
CVPR | 4 |
| 2023 | Efficient Controllable Multi-Task ArchitecturesabstractWe aim to train a multi-task model such that users can adjust the desired compute budget and relative importance of task performances after deployment, without retraining. This enables optimizing performance for dynamically varying user needs, without heavy computational overhead to train and save models for various scenarios. To this end, we propose a multi-task model consisting of a shared encoder and task-specific decoders where both encoder and decoder channel widths are slimmable. Our key idea is to control the task importance by varying the capacities of task-specific decoders, while controlling the total computational cost by jointly adjusting the encoder capacity. This improves overall accuracy by allowing a stronger encoder for a given budget, increases control over computational cost, and delivers high-quality slimmed sub-architectures based on user’s constraints. Our training strategy involves a novel ‘Configuration-Invariant Knowledge Distillation’ loss that enforces backbone representations to be invariant under different runtime width configurations to enhance accuracy. Further, we present a simple but effective search algorithm that translates user constraints to runtime width configurations of both the shared encoder and task decoders, for sampling the sub-architectures. The key rule for the search algorithm is to provide a larger computational budget to the higher preferred task decoder, while searching a shared encoder configuration that enhances the overall MTL performance. Various experiments on three multi-task benchmarks (PASCALContext, NYUDv2, and CIFAR100-MTL) with diverse backbone architectures demonstrate the advantage of our approach. For example, our method shows a higher controllability by ∼ 33.5% in the NYUD-v2 dataset over prior methods, while incurring much less compute cost. Abhishek Aich, Samuel Schulter, Amit K. Roy-Chowdhury, Manmohan Krishna Chandraker, Yumin Suh |
ICCV | 3 |
| 2023 | Prior-guided Source-free Domain Adaptation for Human Pose EstimationabstractDomain adaptation methods for 2D human pose estimation typically require continuous access to the source data during adaptation, which can be challenging due to privacy, memory, or computational constraints. To address this limitation, we focus on the task of source-free domain adaptation for pose estimation, where a source model must adapt to a new target domain using only unlabeled target data. Although recent advances have introduced source-free methods for classification tasks, extending them to the regression task of pose estimation is non-trivial. In this paper, we present Prior-guided Self-training (POST), a pseudo-labeling approach that builds on the popular Mean Teacher framework to compensate for the distribution shift. POST leverages prediction-level and feature-level consistency between a student and teacher model against certain image transformations. In the absence of source data, POST utilizes a human pose prior that regularizes the adaptation process by directing the model to generate more accurate and anatomically plausible pose pseudo-labels. Despite being simple and intuitive, our framework can deliver significant performance gains compared to applying the source model directly to the target data, as demonstrated in our extensive experiments and ablation studies. In fact, our approach achieves comparable performance to recent state-of-the-art methods that use source data for adaptation. Dripta S. Raychaudhuri, Calvin-Khang Ta, Arindam Dutta, Rohit Lal, Amit K. Roy-Chowdhury |
ICCV | 5 |
| 2023 | SUMMIT: Source-Free Adaptation of Uni-Modal Models to Multi-Modal TargetsabstractScene understanding using multi-modal data is necessary in many applications, e.g., autonomous navigation. To achieve this in a variety of situations, existing models must be able to adapt to shifting data distributions without arduous data annotation. Current approaches assume that the source data is available during adaptation and that the source consists of paired multi-modal data. Both these assumptions may be problematic for many applications. Source data may not be available due to privacy, security, or economic concerns. Assuming the existence of paired multi-modal data for training also entails significant data collection costs and fails to take advantage of widely available freely distributed pre-trained uni-modal models. In this work, we relax both of these assumptions by addressing the problem of adapting a set of models trained independently on uni-modal data to a target domain consisting of unlabeled multi-modal data, without having access to the original source dataset. Our proposed approach solves this problem through a switching framework which automatically chooses between two complementary methods of cross-modal pseudo-label fusion – agreement filtering and entropy weighting – based on the estimated domain gap. We demonstrate our work on the semantic segmentation problem. Experiments across seven challenging adaptation scenarios verify the efficacy of our approach, achieving results comparable to, and in some cases outperforming, methods which assume access to source data. Our method achieves an improvement in mIoU of up to 12% over competing baselines. Our code is publicly available at https://github.com/csimo005/SUMMIT. Cody Simons, Dripta S. Raychaudhuri, Sk Miraj Ahmed, Suya You, Konstantinos Karydis, Amit K. Roy-Chowdhury |
ICCV | 6 |
| 2023 | Leveraging Local Patch Differences in Multi-Object Scenes for Generative Adversarial AttacksabstractState-of-the-art generative model-based attacks against image classifiers overwhelmingly focus on single-object(i.e., single dominant object) images. Different from such settings, we tackle a more practical problem of generating adversarial perturbations using multi-object (i.e., multiple dominant objects) images as they are representative of most real-world scenes. Our goal is to design an attack strategy that can learn from such natural scenes by leveraging the local patch differences that occur inherently in such images (e.g. difference between the local patch on the object ‘person’ and the object ‘bike’ in a traffic scene). Our key idea is to misclassify an adversarial multi-object image by confusing the victim classifier for each local patch in the image. Based on this, we propose a novel generative attack (called Local Patch Difference or LPD-Attack) where a novel contrastive loss function uses the aforesaid local differences in feature space of multi-object scenes to optimize the perturbation generator. Through various experiments across diverse victim convolutional neural networks, we show that our approach outperforms baseline generative attacks with highly transferable perturbations when evaluated under different white-box and black-box settings. Abhishek Aich, Shasha Li 0002, Chengyu Song, Muhammad Salman Asif, Srikanth V. Krishnamurthy, Amit K. Roy-Chowdhury |
WACV | 6 |
| 2023 | Cross-Domain Video Anomaly Detection without Target Domain AdaptationabstractMost cross-domain unsupervised Video Anomaly Detection (VAD) works assume that at least few task-relevant target domain training data are available for adaptation from the source to the target domain. However, this requires laborious model-tuning by the end-user who may prefer to have a system that works "out-of-the-box." To address such practical scenarios, we identify a novel target domain (inference-time) VAD task where no target domain training data are available. To this end, we propose a new ‘Zero-shot Cross-domain Video Anomaly Detection (zxVAD)’ framework that includes a future-frame prediction generative model setup. Different from prior future-frame prediction models, our model uses a novel Normalcy Classifier module to learn the features of normal event videos by learning how such features are different "relatively" to features in pseudo-abnormal examples. A novel Untrained Convolutional Neural Network based Anomaly Synthesis module crafts these pseudo-abnormal examples by adding foreign objects in normal video frames with no extra training cost. With our novel relative normalcy feature learning strategy, zxVAD generalizes and learns to distinguish between normal and abnormal frames in a new target domain without adaptation during inference. Through evaluations on common datasets, we show that zxVAD outperforms the state-of-the-art (SOTA), regardless of whether task-relevant (i.e., VAD) source training data are available or not. Lastly, zxVAD also beats the SOTA methods in inference-time efficiency metrics including the model size, total parameters, GPU energy consumption, and GMACs. Abhishek Aich, Kuan-Chuan Peng, Amit K. Roy-Chowdhury |
WACV | 3 |
| 2023 | Joint Video Rolling Shutter Correction and Super-ResolutionabstractWith the prevalence of CMOS cameras in many computer vision applications, there is an increase in the appearance of rolling shutter (RS) artifacts in captured videos. However, existing video super-resolution algorithms assume that the motion is globally consistent in each video frame and no rolling shutter effect is present. The problem of video super-resolution for video captured using RS cameras is challenging as the model needs to learn the row-wise local pixel displacements and the global structure of the frame for RS correction and super-resolution, simultaneously. Different from existing works, we address a more realistic problem of joint rolling shutter correction and super-resolution (RS-SR). We introduce a novel architecture, deformable Patch Attention Network (PatchNet), that utilizes patch-recurrence property along with deformable receptive fields to learn the global and local structure of the video. Specifically, PatchNet leverages bi-directional motion field in the feature space to extract relevant information from neighboring patches using attention mechanism, and deformable fields using deformable convolutions to extract local pixel-level information for joint rolling shutter correction and super-resolution. Our work is the first to tackle the task of RS correction and super-resolution on the recently released BS-RSCD dataset. Experiments on the BS-RSCD and FastecRS datasets demonstrate that our model performs favorably against various state-of-the-art approaches. Project details are available at https://akashagupta.com/publication/wacv23_patchnet/project.html Akash Gupta 0001, Sudhir Kumar Singh, Amit K. Roy-Chowdhury |
WACV | 3 |
| 2023 | Semantics Guided Contrastive Learning of Transformers for Zero-shot Temporal Activity DetectionabstractZero-shot temporal activity detection (ZSTAD) is the problem of simultaneous temporal localization and classification of activity segments that are previously unseen during training. This is achieved by transferring the knowledge learned from semantically-related seen activities. This ability to reason about unseen concepts without supervision makes ZSTAD very promising for applications where the acquisition of annotated training videos is difficult. In this paper, we design a transformer-based framework titled TranZAD, which streamlines the detection of unseen activities by casting ZSTAD as a direct set-prediction problem, removing the need for hand-crafted designs and manual post-processing. We show how a semantic information-guided contrastive learning strategy can effectively train TranZAD for the zero-shot setting, enabling the efficient transfer of knowledge from the seen to the unseen activities. To reduce confusion between unseen activities and unrelated background information in videos, we introduce a more efficient method of computing the background class embedding by dynamically adapting it as part of the end-to-end learning. Additionally, unlike existing work on ZSTAD, we do not assume the knowledge of which classes are unseen during training and use the visual and semantic information of only the seen classes for the knowledge transfer. This makes TranZAD more viable for practical scenarios, which we evaluate by conducting extensive experiments on Thumos’14 and Charades. Sayak Nag, Orpaz Goldstein, Amit K. Roy-Chowdhury |
WACV | 3 |
| 2023 | Centroid Distance Keypoint Detector for Colored Point CloudsabstractKeypoint detection serves as the basis for many computer vision and robotics applications. Despite the fact that colored point clouds can be readily obtained, most existing keypoint detectors extract only geometry-salient keypoints, which can impede the overall performance of systems that intend to (or have the potential to) leverage color information. To promote advances in such systems, we propose an efficient multi-modal keypoint detector that can extract both geometry-salient and color-salient keypoints in colored point clouds. The proposed CEntroid Distance (CED) key- point detector comprises an intuitive and effective saliency measure, the centroid distance, that can be used in both 3D space and color space, and a multi-modal non-maximum suppression algorithm that can select keypoints with high saliency in two or more modalities. The proposed saliency measure leverages directly the distribution of points in a local neighborhood and does not require normal estimation or eigenvalue decomposition. We evaluate the proposed method in terms of repeatability and computational efficiency (i.e. running time) against state-of-the-art key- point detectors on both synthetic and real-world datasets. Results demonstrate that our proposed CED keypoint detector requires minimal computational time while attaining high repeatability. To showcase one of the potential applications of the proposed method, we further investigate the task of colored point cloud registration. Results suggest that our proposed CED detector outperforms state-of- the-art handcrafted and learning-based keypoint detectors in the evaluated scenes. The C++ implementation of the proposed method is made publicly available at https://github.com/UCR-Robotics/CED_Detector. Hanzhe Teng, Dimitrios Chatziparaschis, Xinyue Kan, Amit K. Roy-Chowdhury, Konstantinos Karydis |
WACV | 4 |
| 2023 | Reconstruction Guided Meta-Learning for Few Shot Open Set RecognitionabstractIn many applications, we are constrained to learn classifiers from very limited data (few-shot classification). The task becomes even more challenging if it is also required to identify samples from unknown categories (open-set classification). Learning a good abstraction for a class with very few samples is extremely difficult, especially under open-set settings. As a result, open-set recognition has received limited attention in the few-shot setting. However, it is a critical task in many applications like environmental monitoring, where the number of labeled examples for each class is limited. Existing few-shot open-set recognition (FSOSR) methods rely on thresholding schemes, with some considering uniform probability for open-class samples. However, this approach is often inaccurate, especially for fine-grained categorization, and makes them highly sensitive to the choice of a threshold. To address these concerns, we propose Reconstructing Exemplar-based Few-shot Open-set ClaSsifier (ReFOCS). By using a novel exemplar reconstruction-based meta-learning strategy ReFOCS streamlines FSOSR eliminating the need for a carefully tuned threshold by learning to be self-aware of the openness of a sample. The exemplars, act as class representatives and can be either provided in the training dataset or estimated in the feature domain. By testing on a wide variety of datasets, we show ReFOCS to outperform multiple state-of-the-art methods. Sayak Nag, Dripta S. Raychaudhuri, Sujoy Paul, Amit K. Roy-Chowdhury |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | AcTrak: Controlling a Steerable Surveillance Camera using Reinforcement LearningabstractSteerable cameras that can be controlled via a network, to retrieve telemetries of interest have become popular. In this paper, we develop a framework called AcTrak , to automate a camera’s motion to appropriately switch between (a) zoom ins on existing targets in a scene to track their activities, and (b) zoom out to search for new targets arriving to the area of interest. Specifically, we seek to achieve a good trade-off between the two tasks, i.e., we want to ensure that new targets are observed by the camera before they leave the scene, while also zooming in on existing targets frequently enough to monitor their activities. There exist prior control algorithms for steering cameras to optimize certain objectives; however, to the best of our knowledge, none have considered this problem, and do not perform well when target activity tracking is required. AcTrak automatically controls the camera’s PTZ configurations using reinforcement learning (RL ), to select the best camera position given the current state. Via simulations using real datasets, we show that AcTrak detects newly arriving targets 30% faster than a non-adaptive baseline and rarely misses targets, unlike the baseline which can miss up to 5% of the targets. We also implement AcTrak to control a real camera and demonstrate that in comparison with the baseline, it acquires about 2× more high resolution images of targets. Abdulrahman Fahim, Evangelos E. Papalexakis, Srikanth V. Krishnamurthy, Amit K. Roy-Chowdhury, Lance M. Kaplan, Tarek F. Abdelzaher |
ACM Trans. Cyber Phys. Syst. | 4 |
| 2022 | Context-Aware Transfer Attacks for Object DetectionabstractBlackbox transfer attacks for image classifiers have been extensively studied in recent years. In contrast, little progress has been made on transfer attacks for object detectors. Object detectors take a holistic view of the image and the detection of one object (or lack thereof) often depends on other objects in the scene. This makes such detectors inherently context-aware and adversarial attacks in this space are more challenging than those targeting image classifiers. In this paper, we present a new approach to generate context-aware attacks for object detectors. We show that by using co-occurrence of objects and their relative locations and sizes as context information, we can successfully generate targeted mis-categorization attacks that achieve higher transfer success rates on blackbox object detectors than the state-of-the-art. We test our approach on a variety of object detectors with images from PASCAL VOC and MS COCO datasets and demonstrate up to 20 percentage points improvement in performance compared to the other state-of-the-art methods. Zikui Cai, Xinxin Xie, Shasha Li 0001, Mingjun Yin, Chengyu Song, Srikanth V. Krishnamurthy, Amit K. Roy-Chowdhury, Muhammad Salman Asif |
AAAI | 7 |
| 2022 | Zero-Query Transfer Attacks on Context-Aware Object DetectorsabstractAdversarial attacks perturb images such that a deep neural network produces incorrect classification results. A promising approach to defend against adversarial attacks on natural multi-object scenes is to impose a context-consistency check, wherein, if the detected objects are not consistent with an appropriately defined context, then an attack is suspected. Stronger attacks are needed to fool such context-aware detectors. We present the first approach for generating context-consistent adversarial attacks that can evade the context-consistency check of black-box object detectors operating on complex, natural scenes. Unlike many black-box attacks that perform repeated attempts and open themselves to detection, we assume a “zero-query” setting, where the attacker has no knowledge of the classification decisions of the victim system. First, we derive multiple attack plans that assign incorrect labels to victim objects in a context-consistent manner. Then we design and use a novel data structure that we call the perturbation success probability matrix, which enables us to filter the attack plans and choose the one most likely to succeed. This final attack plan is implemented using a perturbation-bounded adversarial attack algorithm. We compare our zero-query attack against a few-query scheme that repeatedly checks if the victim system is fooled. We also compare against state-of-the-art context-agnostic attacks. Against a context-aware defense, the fooling rate of our zero-query approach is significantly higher than context-agnostic approaches and higher than that achievable with up to three rounds of the fewquery scheme. Zikui Cai, Shantanu Rane, Alejandro E. Brito, Chengyu Song, Srikanth V. Krishnamurthy, Amit K. Roy-Chowdhury, Muhammad Salman Asif |
CVPR | 6 |
| 2022 | Controllable Dynamic Multi-Task ArchitecturesabstractMulti-task learning commonly encounters competition for resources among tasks, specifically when model capac-ity is limited. This challenge motivates models which al-low control over the relative importance of tasks and total compute cost during inference time. In this work, we pro-pose such a controllable multi-task network that dynami-cally adjusts its architecture and weights to match the de-sired task preference as well as the resource constraints. In contrast to the existing dynamic multi-task approaches that adjust only the weights within a fixed architecture, our approach affords the flexibility to dynamically control the total computational cost and match the user-preferred task importance better. We propose a disentangled training of two hype rnetwo rks, by exploiting task affinity and a novel branching regularized loss, to take input prefer-ences and accordingly predict tree-structured models with adapted weights. Experiments on three multi-task bench-marks, namely PASCAL-Context, NYU-v2, and CIFAR-100, show the efficacy of our approach. Project page is available at https://www.nec-labs.com/-mas/DYMU. Dripta S. Raychaudhuri, Yumin Suh, Samuel Schulter, Xiang Yu 0002, Masoud Faraki, Amit K. Roy-Chowdhury, Manmohan Krishna Chandraker |
CVPR | 6 |
| 2022 | Cross-Modal Knowledge Transfer Without Task-Relevant Source Data
Sk Miraj Ahmed, Suhas Lohit, Kuan-Chuan Peng, Michael J. Jones 0001, Amit K. Roy-Chowdhury |
ECCV (34) | 5 |
| 2022 | Text-Based Temporal Localization of Novel Events
Sudipta Paul 0007, Niluthpol Chowdhury Mithun, Amit K. Roy-Chowdhury |
ECCV (14) | 3 |
| 2022 | Poisson2Sparse: Self-supervised Poisson Denoising from a Single Image
Calvin-Khang Ta, Abhishek Aich, Akash Gupta 0001, Amit K. Roy-Chowdhury |
MICCAI (8) | 4 |
| 2022 | AVLEN: Audio-Visual-Language Embodied Navigation in 3D EnvironmentsabstractRecent years have seen embodied visual navigation advance in two distinct directions: (i) in equipping the AI agent to follow natural language instructions, and (ii) in making the navigable world multimodal, e.g., audio-visual navigation. However, the real world is not only multimodal, but also often complex, and thus in spite of these advances, agents still need to understand the uncertainty in their actions and seek instructions to navigate. To this end, we present AVLEN -- an interactive agent for Audio-Visual-Language Embodied Navigation. Similar to audio-visual navigation tasks, the goal of our embodied agent is to localize an audio event via navigating the 3D visual world; however, the agent may also seek help from a human (oracle), where the assistance is provided in free-form natural language. To realize these abilities, AVLEN uses a multimodal hierarchical reinforcement learning backbone that learns: (a) high-level policies to choose either audio-cues for navigation or to query the oracle, and (b) lower-level policies to select navigation actions based on its audio-visual and language inputs. The policies are trained via rewarding for the success on the navigation task while minimizing the number of queries to the oracle. To empirically evaluate AVLEN, we present experiments on the SoundSpaces framework for semantic audio-visual navigation tasks. Our results show that equipping the agent to ask for help leads to a clear improvement in performances, especially in challenging cases, e.g., when the sound is unheard during training or in the presence of distractor sounds. Sudipta Paul 0007, Amit K. Roy-Chowdhury, Anoop Cherian |
NeurIPS | 2 |
| 2022 | GAMA: Generative Adversarial Multi-Object Scene AttacksabstractThe majority of methods for crafting adversarial attacks have focused on scenes with a single dominant object (e.g., images from ImageNet). On the other hand, natural scenes include multiple dominant objects that are semantically related. Thus, it is crucial to explore designing attack strategies that look beyond learning on single-object scenes or attack single-object victim classifiers. Due to their inherent property of strong transferability of perturbations to unknown models, this paper presents the first approach of using generative models for adversarial attacks on multi-object scenes. In order to represent the relationships between different objects in the input scene, we leverage upon the open-sourced pre-trained vision-language model CLIP (Contrastive Language-Image Pre-training), with the motivation to exploit the encoded semantics in the language space along with the visual space. We call this attack approach Generative Adversarial Multi-object Attacks (GAMA). GAMA demonstrates the utility of the CLIP model as an attacker's tool to train formidable perturbation generators for multi-object scenes. Using the joint image-text features to train the generator, we show that GAMA can craft potent transferable perturbations in order to fool victim classifiers in various attack settings. For example, GAMA triggers ~16% more misclassification than state-of-the-art generative approaches in black-box settings where both the classifier architecture and data distribution of the attacker are different from the victim. Our code is available here: https://abhishekaich27.github.io/gama.html Abhishek Aich, Calvin-Khang Ta, Akash Gupta 0001, Chengyu Song, Srikanth V. Krishnamurthy, Muhammad Salman Asif, Amit K. Roy-Chowdhury |
NeurIPS | 7 |
| 2022 | Blackbox Attacks via Surrogate Ensemble SearchabstractBlackbox adversarial attacks can be categorized into transfer- and query-based attacks. Transfer methods do not require any feedback from the victim model, but provide lower success rates compared to query-based methods. Query attacks often require a large number of queries for success. To achieve the best of both approaches, recent efforts have tried to combine them, but still require hundreds of queries to achieve high success rates (especially for targeted attacks). In this paper, we propose a novel method for Blackbox Attacks via Surrogate Ensemble Search (BASES) that can generate highly successful blackbox attacks using an extremely small number of queries. We first define a perturbation machine that generates a perturbed image by minimizing a weighted loss function over a fixed set of surrogate models. To generate an attack for a given victim model, we search over the weights in the loss function using queries generated by the perturbation machine. Since the dimension of the search space is small (same as the number of surrogate models), the search requires a small number of queries. We demonstrate that our proposed method achieves better success rate with at least $30\times$ fewer queries compared to state-of-the-art methods on different image classifiers trained with ImageNet (including VGG-19, DenseNet-121, and ResNext-50). In particular, our method requires as few as 3 queries per image (on average) to achieve more than a $90\%$ success rate for targeted attacks and 1--2 queries per image for over a $99\%$ success rate for untargeted attacks. Our method is also effective on Google Cloud Vision API and achieved a $91\%$ untargeted attack success rate with 2.9 queries per image. We also show that the perturbations generated by our proposed method are highly transferable and can be adopted for hard-label blackbox attacks. Furthermore, we argue that BASES can be used to create attacks for a variety of tasks and show its effectiveness for attacks on object detection models. Our code is available at https://github.com/CSIPlab/BASES. Zikui Cai, Chengyu Song, Srikanth V. Krishnamurthy, Amit K. Roy-Chowdhury, Salman Asif |
NeurIPS | 4 |
| 2022 | Detection and Localization of Facial Expression ManipulationsabstractConcerns regarding the wide-spread use of forged images and videos in social media necessitate precise detection of such fraud. Facial manipulations can be created by Identity swap (DeepFake) or Expression swap. Contrary to the identity swap, which can easily be detected with novel deepfake detection methods, expression swap detection has not yet been addressed extensively. The importance of facial expressions in inter-person communication is known. Consequently, it is important to develop methods that can detect and localize manipulations in facial expressions.To this end, we present a novel framework to exploit the underlying feature representations of facial expressions learned from expression recognition models to identify the manipulated features. Using discriminative feature maps extracted from a facial expression recognition framework, our manipulation detector is able to localize the manipulated regions of input images and videos. On the Face2Face dataset, (abundant expression manipulation), and Neural-Textures dataset (facial expressions manipulation corresponding to the mouth regions), our method achieves higher accuracy for both classification and localization of manipulations compared to state-of-the-art methods. Furthermore, we demonstrate that our method performs at-par with the state-of-the-art methods in cases where the expression is not manipulated, but rather the identity is changed, leading to a generalized approach for facial manipulation detection. Ghazal Mazaheri, Amit K. Roy-Chowdhury |
WACV | 2 |
| 2022 | ADC: Adversarial attacks against object Detection that evade Context consistency checksabstractDeep Neural Networks (DNNs) have been shown to be vulnerable to adversarial examples, which are slightly perturbed input images which lead DNNs to make wrong predictions. To protect from such examples, various defense strategies have been proposed. A very recent defense strategy for detecting adversarial examples, that has been shown to be robust to current attacks, is to check for intrinsic context consistencies in the input data, where context refers to various relationships (e.g., object-to-object co-occurrence relationships) in images. In this paper, we show that even context consistency checks can be brittle to properly crafted adversarial examples and to the best of our knowledge, we are the first to do so. Specifically, we propose an adaptive framework to generate examples that subvert such defenses, namely, Adversarial attacks against object Detection that evade Context consistency checks (ADC). In ADC, we formulate a joint optimization problem which has two attack goals, viz., (i) fooling the object detector and (ii) evading the context consistency check system, at the same time. Experiments on both PASCAL VOC and MS COCO datasets show that examples generated with ADC fool the object detector with a success rate of over 85% in most cases, and at the same time evade the recently proposed context consistency checks, with a "bypassing" rate of over 80% in most cases. Our results suggest that "how to robustly model con- text and check its consistency," is still an open problem. Mingjun Yin, Shasha Li 0001, Chengyu Song, Muhammad Salman Asif, Amit K. Roy-Chowdhury, Srikanth V. Krishnamurthy |
WACV | 5 |
| 2022 | Combined computational modeling and experimental analysis integrating chemical and mechanical signals suggests possible mechanism of shoot meristem maintenanceabstractStem cell maintenance in multilayered shoot apical meristems (SAMs) of plants requires strict regulation of cell growth and division. Exactly how the complex milieu of chemical and mechanical signals interact in the central region of the SAM to regulate cell division plane orientation is not well understood. In this paper, simulations using a newly developed multiscale computational model are combined with experimental studies to suggest and test three hypothesized mechanisms for the regulation of cell division plane orientation and the direction of anisotropic cell expansion in the corpus. Simulations predict that in the Apical corpus, WUSCHEL and cytokinin regulate the direction of anisotropic cell expansion, and cells divide according to tensile stress on the cell wall. In the Basal corpus, model simulations suggest dual roles for WUSCHEL and cytokinin in regulating both the direction of anisotropic cell expansion and cell division plane orientation. Simulation results are followed by a detailed analysis of changes in cell characteristics upon manipulation of WUSCHEL and cytokinin in experiments that support model predictions. Moreover, simulations predict that this layer-specific mechanism maintains both the experimentally observed shape and structure of the SAM as well as the distribution of WUSCHEL in the tissue. This provides an additional link between the roles of WUSCHEL, cytokinin, and mechanical stress in regulating SAM growth and proper stem cell maintenance in the SAM. Mikahl Banwarth-Kuhn, Kevin Rodriguez, Christian Michael, Calvin-Khang Ta, Alexander Plong, Eric Bourgain-Chang, Ali Nematbakhsh, Amit K. Roy-Chowdhury, G. Venugopala Reddy, Mark S. Alber |
PLoS Comput. Biol. | 9 |
| 2021 | Unsupervised Multi-Source Domain Adaptation Without Access to Source DataabstractUnsupervised Domain Adaptation (UDA) aims to learn a predictor model for an unlabeled domain by transferring knowledge from a separate labeled source domain. However, most of these conventional UDA approaches make the strong assumption of having access to the source data during training, which may not be very practical due to privacy, security and storage concerns. A recent line of work addressed this problem and proposed an algorithm that transfers knowledge to the unlabeled target domain from a single source model without requiring access to the source data. However, for adaptation purposes, if there are multiple trained source models available to choose from, this method has to go through adapting each and every model individually, to check for the best source. Thus, we ask the question: can we find the optimal combination of source models, with no source data and without target labels, whose performance is no worse than the single best source? To answer this, we propose a novel and efficient algorithm which automatically combines the source models with suitable weights in such a way that it performs at least as good as the best source model. We provide intuitive theoretical insights to justify our claim. Furthermore, extensive experiments are conducted on several benchmark datasets to show the effectiveness of our algorithm, where in most cases, our method not only reaches best source accuracy but also outperforms it. Sk Miraj Ahmed, Dripta S. Raychaudhuri, Sujoy Paul, Samet Oymak, Amit K. Roy-Chowdhury |
CVPR | 5 |
| 2021 | Spatio-Temporal Representation Factorization for Video-based Person Re-IdentificationabstractDespite much recent progress in video-based person re-identification (re-ID), the current state-of-the-art still suffers from common real-world challenges such as appearance similarity among various people, occlusions, and frame misalignment. To alleviate these problems, we propose Spatio-Temporal Representation Factorization (STRF), a flexible new computational unit that can be used in conjunction with most existing 3D convolutional neural network architectures for re-ID. The key innovations of STRF over prior work include explicit pathways for learning discriminative temporal and spatial features, with each component further factorized to capture complementary person-specific appearance and motion information. Specifically, temporal factorization comprises two branches, one each for static features (e.g., the color of clothes) that do not change much over time, and dynamic features (e.g., walking patterns) that change over time. Further, spatial factorization also comprises two branches to learn both global (coarse segments) as well as local (finer segments) appearance features, with the local features particularly useful in cases of occlusion or spatial misalignment. These two factorization operations taken together result in a modular architecture for our parameter-wise light STRF unit that can be plugged in between any two 3D convolutional layers, resulting in an end-to-end learning framework. We empirically show that STRF improves performance of various existing baseline architectures while demonstrating new state-of-the-art results using standard person re-ID evaluation protocols on three benchmarks. Abhishek Aich, Meng Zheng 0002, Srikrishna Karanam, Terrence Chen, Amit K. Roy-Chowdhury, Ziyan Wu 0001 |
ICCV | 5 |
| 2021 | Multi-Expert Adversarial Attack Detection in Person Re-identification Using Context InconsistencyabstractThe success of deep neural networks (DNNs) has promoted the widespread applications of person re-identification (ReID). However, ReID systems inherit the vulnerability of DNNs to malicious attacks of visually in-conspicuous adversarial perturbations. Detection of adversarial attacks is, therefore, a fundamental requirement for robust ReID systems. In this work, we propose a Multi-Expert Adversarial Attack Detection (MEAAD) approach to achieve this goal by checking context inconsistency, which is suitable for any DNN-based ReID systems. Specifically, three kinds of context inconsistencies caused by adversarial attacks are employed to learn a detector for distinguishing the perturbed examples, i.e., a) the embedding distances between a perturbed query person image and its top-K retrievals are generally larger than those between a benign query image and its top-K retrievals, b) the embedding distances among the top-K retrievals of a perturbed query image are larger than those of a benign query image, c) the top-K retrievals of a benign query image obtained with multiple expert ReID models tend to be consistent, which is not preserved when attacks are present. Extensive experiments on the Market1501 and DukeMTMC-ReID datasets show that, as the first adversarial attack detection approach for ReID, MEAAD effectively detects various adversarial attacks and achieves high ROC-AUC (over 97.5%). Shasha Li 0001, Min Liu 0008, Yaonan Wang 0001, Amit K. Roy-Chowdhury |
ICCV | 5 |
| 2021 | Exploiting Multi-Object Relationships for Detecting Adversarial Attacks in Complex ScenesabstractVision systems that deploy Deep Neural Networks (DNNs) are known to be vulnerable to adversarial examples. Recent research has shown that checking the intrinsic consistencies in the input data is a promising way to detect adversarial attacks (e.g., by checking the object co-occurrence relationships in complex scenes). However, existing approaches are tied to specific models and do not offer generalizability. Motivated by the observation that language descriptions of natural scene images have already captured the object co-occurrence relationships that can be learned by a language model, we develop a novel approach to perform context consistency checks using such language models. The distinguishing aspect of our approach is that it is independent of the deployed object detector and yet offers very high accuracy in terms of detecting adversarial examples in practical scenes with multiple objects. Experiments on the PASCAL VOC and MS COCO datasets show that our method can outperform state-of-the-art methods in detecting adversarial attacks. Mingjun Yin, Shasha Li 0001, Zikui Cai, Chengyu Song, Muhammad Salman Asif, Amit K. Roy-Chowdhury, Srikanth V. Krishnamurthy |
ICCV | 6 |
| 2021 | Cross-domain Imitation from ObservationsabstractImitation learning seeks to circumvent the difficulty in designing proper reward functions for training agents by utilizing expert behavior. With environments modeled as Markov Decision Processes (MDP), most of the existing imitation algorithms are contingent on the availability of expert demonstrations in the same MDP as the one in which a new imitation policy is to be learned. In this paper, we study the problem of how to imitate tasks when discrepancies exist between the expert and agent MDP. These discrepancies across domains could include differing dynamics, viewpoint, or morphology; we present a novel framework to learn correspondences across such domains. Importantly, in contrast to prior works, we use unpaired and unaligned trajectories containing only states in the expert domain, to learn this correspondence. We utilize a cycle-consistency constraint on both the state space and a domain agnostic latent space to do this. In addition, we enforce consistency on the temporal position of states via a normalized position estimator function, to align the trajectories across the two domains. Once this correspondence is found, we can directly transfer the demonstrations on one domain to the other and use it for imitation. Experiments across a wide variety of challenging domains demonstrate the efficacy of our approach. Dripta S. Raychaudhuri, Sujoy Paul, Jeroen van Baar, Amit K. Roy-Chowdhury |
ICML | 4 |
| 2021 | Ada-VSR: Adaptive Video Super-Resolution with Meta-LearningabstractMost of the existing works in supervised spatio-temporal video super-resolution (STVSR) heavily rely on a large-scale external dataset consisting of paired low-resolution low-frame rate (LR-LFR) and high-resolution high-frame-rate (HR-HFR) videos. Despite their remarkable performance, these methods make a prior assumption that the low-resolution video is obtained by down-scaling the high-resolution video using a known degradation kernel, which does not hold in practical settings. Another problem with these methods is that they cannot exploit instance-specific internal information of a video at testing time. Recently, deep internal learning approaches have gained attention due to their ability to utilize the instance-specific statistics of a video. However, these methods have a large inference time as they require thousands of gradient updates to learn the intrinsic structure of the data. In this work, we present Adaptive VideoSuper-Resolution (Ada-VSR) which leverages external, as well as internal, information through meta-transfer learning and internal learning, respectively. Specifically, meta-learning is employed to obtain adaptive parameters, using a large-scale external dataset, that can adapt quickly to the novel condition (degradation model) of the given test video during the internal learning task, thereby exploiting external and internal information of a video for super-resolution. The model trained using our approach can quickly adapt to a specific video condition with only a few gradient updates, which reduces the inference time significantly. Extensive experiments on standard datasets demonstrate that our method performs favorably against various state-of-the-art approaches. Akash Gupta 0001, Padmaja Jonnalagedda, Bir Bhanu, Amit K. Roy-Chowdhury |
ACM Multimedia | 4 |
| 2021 | Adversarial Attacks on Black Box Video Classifiers: Leveraging the Power of Geometric TransformationsabstractWhen compared to the image classification models, black-box adversarial attacks against video classification models have been largely understudied. This could be possible because, with video, the temporal dimension poses significant additional challenges in gradient estimation. Query-efficient black-box attacks rely on effectively estimated gradients towards maximizing the probability of misclassifying the target video. In this work, we demonstrate that such effective gradients can be searched for by parameterizing the temporal structure of the search space with geometric transformations. Specifically, we design a novel iterative algorithm GEOmetric TRAnsformed Perturbations (GEO-TRAP), for attacking video classification models. GEO-TRAP employs standard geometric transformation operations to reduce the search space for effective gradients into searching for a small group of parameters that define these operations. This group of parameters describes the geometric progression of gradients, resulting in a reduced and structured search space. Our algorithm inherently leads to successful perturbations with surprisingly few queries. For example, adversarial examples generated from GEO-TRAP have better attack success rates with ~73.55% fewer queries compared to the state-of-the-art method for video adversarial attacks on the widely used Jester dataset. Overall, our algorithm exposes vulnerabilities of diverse video classification models and achieves new state-of-the-art results under black-box settings on two large datasets. Shasha Li 0001, Abhishek Aich, Shitong Zhu, Muhammad Salman Asif, Chengyu Song, Amit K. Roy-Chowdhury, Srikanth V. Krishnamurthy |
NeurIPS | 6 |
| 2021 | Prediction and Description of Near-Future Activities in VideoabstractMost of the existing works on human activity analysis focus on recognition or early recognition of the activity labels from complete or partial observations. Similarly, almost all of the existing video captioning approaches focus on the observed events in videos. Predicting the labels and the captions of future activities where no frames of the predicted activities have been observed is a challenging problem, with important applications that require anticipatory response. In this work, we propose a system that can infer the labels and the captions of a sequence of future activities. Our proposed network for label prediction of a future activity sequence has three branches where the first branch takes visual features from the objects present in the scene, the second branch takes observed sequential activity features, and the third branch captures the last observed activity features. The predicted labels and the observed scene context are then mapped to meaningful captions using a sequence-to-sequence learning-based method. Experiments on four challenging activity analysis datasets and a video description dataset demonstrate that our label prediction approach achieves comparable performance with the state-of-the-arts and our captioning framework outperform the state-of-the-arts. Tahmida Mahmud, Mohammad Billah 0001, Mahmudul Hasan 0003, Amit K. Roy-Chowdhury |
Comput. Vis. Image Underst. | 4 |
| 2021 | Exploiting Global Camera Network Constraints for Unsupervised Video Person Re-IdentificationabstractMany unsupervised approaches have been proposed recently for the video-based re-identification problem since annotations of samples across cameras are time-consuming. However, higher-order relationships across the entire camera network are ignored by these methods, leading to contradictory outputs when matching results from different camera pairs are combined. In this paper, we address the problem of unsupervised video-based re-identification by proposing a consistent cross-view matching (CCM) framework, in which global camera network constraints are exploited to guarantee the matched pairs are with consistency. Specifically, we first propose to utilize the first neighbor of each sample to discover relations among samples and find the groups in each camera. Additionally, a cross-view matching strategy followed by global camera network constraints is proposed to explore the matching relationships across the entire camera network. Finally, we learn metric models for camera pairs progressively by alternatively mining consistent cross-view matching pairs and updating metric models using these obtained matches. Rigorous experiments on two widely-used benchmarks for video re-identification demonstrate the superiority of the proposed method over current state-of-the-art unsupervised methods; for example, on the MARS dataset, our method achieves an improvement of 4.2% over unsupervised methods, and even 2.5% over one-shot supervision-based methods for rank-1 accuracy. Rameswar Panda, Min Liu 0008, Yaonan Wang 0001, Amit K. Roy-Chowdhury |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Text-Based Localization of Moments in a Video CorpusabstractPrior works on text-based video moment localization focus on temporally grounding the textual query in an untrimmed video. These works assume that the relevant video is already known and attempt to localize the moment on that relevant video only. Different from such works, we relax this assumption and address the task of localizing moments in a corpus of videos for a given sentence query. This task poses a unique challenge as the system is required to perform: 2) retrieval of the relevant video where only a segment of the video corresponds with the queried sentence, 2) temporal localization of moment in the relevant video based on sentence query. Towards overcoming this challenge, we propose Hierarchical Moment Alignment Network (HMAN) which learns an effective joint embedding space for moments and sentences. In addition to learning subtle differences between intra-video moments, HMAN focuses on distinguishing inter-video global semantic concepts based on sentence queries. Qualitative and quantitative results on three benchmark text-based video moment retrieval datasets - Charades-STA, DiDeMo, and ActivityNet Captions - demonstrate that our method achieves promising performance on the proposed task of temporal localization of moments in a corpus of videos. Sudipta Paul 0007, Niluthpol Chowdhury Mithun, Amit K. Roy-Chowdhury |
IEEE Trans. Image Process. | 3 |
| 2021 | Learning Person Re-Identification Models From Videos With Weak SupervisionabstractMost person re-identification methods, being supervised techniques, suffer from the burden of massive annotation requirement. Unsupervised methods overcome this need for labeled data, but perform poorly compared to the supervised alternatives. In order to cope with this issue, we introduce the problem of learning person re-identification models from videos with weak supervision. The weak nature of the supervision arises from the requirement of video-level labels, i.e. person identities who appear in the video, in contrast to the more precise frame-level annotations. Towards this goal, we propose a multiple instance attention learning framework for person re-identification using such video-level labels. Specifically, we first cast the video person re-identification task into a multiple instance learning setting, in which person images in a video are collected into a bag. The relations between videos with similar labels can be utilized to identify persons, on top of that, we introduce a co-person attention mechanism which mines the similarity correlations between videos with person identities in common. The attention weights are obtained based on all person images instead of person tracklets in a video, making our learned model less affected by noisy annotations. Extensive experiments demonstrate the superiority of the proposed method over the related methods on two weakly labeled person re-identification datasets. Min Liu 0008, Dripta S. Raychaudhuri, Sujoy Paul, Yaonan Wang 0001, Amit K. Roy-Chowdhury |
IEEE Trans. Image Process. | 6 |
| 2020 | Camera On-Boarding for Person Re-Identification Using Hypothesis Transfer LearningabstractMost of the existing approaches for person re-identification consider a static setting where the number of cameras in the network is fixed. An interesting direction, which has received little attention, is to explore the dynamic nature of a camera network, where one tries to adapt the existing re-identification models after on-boarding new cameras, with little additional effort. There have been a few recent methods proposed in person re-identification that attempt to address this problem by assuming the labeled data in the existing network is still available while adding new cameras. This is a strong assumption since there may exist some privacy issues for which one may not have access to those data. Rather, based on the fact that it is easy to store the learned re-identifications models, which mitigates any data privacy concern, we develop an efficient model adaptation approach using hypothesis transfer learning that aims to transfer the knowledge using only source models and limited labeled data, but without using any source camera data from the existing network. Our approach minimizes the effect of negative transfer by finding an optimal weighted combination of multiple source models for transferring the knowledge. Extensive experiments on four challenging benchmark datasets with variable number of cameras well demonstrate the efficacy of our proposed approach over state-of-the-art methods. Sk Miraj Ahmed, Aske R. Lejbølle, Rameswar Panda, Amit K. Roy-Chowdhury |
CVPR | 4 |
| 2020 | Non-Adversarial Video Synthesis with Learned PriorsabstractMost of the existing works in video synthesis focus on generating videos using adversarial learning. Despite their success, these methods often require input reference frame or fail to generate diverse videos from the given data distribution, with little to no uniformity in the quality of videos that can be generated. Different from these methods, we focus on the problem of generating videos from latent noise vectors, without any reference input frames. To this end, we develop a novel approach that jointly optimizes the input latent space, the weights of a recurrent neural network and a generator through non-adversarial learning. Optimizing for the input latent space along with the network weights allows us to generate videos in a controlled environment, i.e., we can faithfully generate all videos the model has seen during the learning process as well as new unseen videos. Extensive experiments on three challenging and diverse datasets well demonstrate that our proposed approach generates superior quality videos compared to the existing state-of-the-art methods. Abhishek Aich, Akash Gupta 0001, Rameswar Panda, Rakib Hyder, Muhammad Salman Asif, Amit K. Roy-Chowdhury |
CVPR | 6 |
| 2020 | Connecting the Dots: Detecting Adversarial Perturbations Using Context Inconsistency
Shasha Li 0001, Shitong Zhu, Sudipta Paul 0007, Amit K. Roy-Chowdhury, Chengyu Song, Srikanth V. Krishnamurthy, Ananthram Swami, Kevin S. Chan |
ECCV (23) | 4 |
| 2020 | Domain Adaptive Semantic Segmentation Using Weak Labels
Sujoy Paul, Yi-Hsuan Tsai, Samuel Schulter, Amit K. Roy-Chowdhury, Manmohan Krishna Chandraker |
ECCV (9) | 4 |
| 2020 | Exploiting Temporal Coherence for Self-Supervised One-Shot Video Re-identification
Dripta S. Raychaudhuri, Amit K. Roy-Chowdhury |
ECCV (27) | 2 |
| 2020 | Complex Pairwise Activity Analysis Via Instance Level Evolution ReasoningabstractVideo activity analysis systems are often trained on large datasets. Activities and events in the real world do not occur in isolation, instead, they occur as interactions between related objects. This work introduces a novel method that jointly exploits relational information between pairs of objects and temporal dynamics of each object. The proposed method effectively leverages a new simple architecture that is flexible and easily trained to detect relational activities and events using small datasets (hundreds of samples). The solution is constructed and tested using synthetic videos of car-collision events. The annotated datasets in this work will be made available online to the research community. Experimental results demonstrate the efficacy of the network to perform complex activity analysis. Sudipta Paul 0007, Shivkumar Chandrasekaran, Amit K. Roy-Chowdhury |
ICASSP | 4 |
| 2020 | Webly Supervised Image-Text Embedding with Noisy Tag RefinementabstractIn this paper, we address the problem of utilizing web images in training robust joint embedding models for the image-text retrieval task. Prior webly supervised approaches directly leverage weakly annotated web images in the joint embedding learning framework. The objective of these approaches would suffer significantly when the ratio of noisy and missing tags associated with the web images is very high. In this regard, we propose a CP decomposition based tensor completion framework to refine the tags of web images by modeling observed ternary inter-relations between the sets of labeled images, tags, and web images as a tensor. To effectively deal with the high ratio of missing entries likely in our case, we incorporate intra-modal correlation as side information in the proposed framework. Our tag refinement approach combined with existing web supervised image-text embedding approaches provide a more principled way for learning the joint embedding models in the presence of significant noise from web data and limited clean labeled data. Experiments on benchmark datasets demonstrate that the proposed approach helps to achieve a significant performance gain in image-text retrieval. Niluthpol Chowdhury Mithun, Ravdeep Pasricha, Evangelos E. Papalexakis, Amit K. Roy-Chowdhury |
ICPR | 4 |
| 2020 | ALANET: Adaptive Latent Attention Network for Joint Video Deblurring and InterpolationabstractExisting works address the problem of generating high frame-rate sharp videos by separately learning the frame deblurring and frame interpolation modules. Most of these approaches have a strong prior assumption that all the input frames are blurry whereas in a real-world setting, the quality of frames varies. Moreover, such approaches are trained to perform either of the two tasks - deblurring or interpolation - in isolation, while many practical situations call for both. Different from these works, we address a more realistic problem of high frame-rate sharp video synthesis with no prior assumption that input is always blurry. We introduce a novel architecture, Adaptive Latent Attention Network (ALANET), which synthesizes sharp high frame-rate videos with no prior knowledge of input frames being blurry or not, thereby performing the task of both deblurring and interpolation. We hypothesize that information from the latent representation of the consecutive frames can be utilized to generate optimized representations for both frame deblurring and frame interpolation. Specifically, we employ combination of self-attention and cross-attention module between consecutive frames in the latent space to generate optimized representation for each frame. The optimized representation learnt using these attention modules help the model to generate and interpolate sharp frames. Extensive experiments on standard datasets demonstrate that our method performs favorably against various state-of-the-art approaches, even though we tackle a much more difficult problem. The project page is available at https://agupt013.github.io/ALANET.html. Akash Gupta 0001, Abhishek Aich, Amit K. Roy-Chowdhury |
ACM Multimedia | 3 |
| 2020 | Adversarial Knowledge Transfer from Unlabeled DataabstractWhile machine learning approaches to visual recognition offer great promise, most of the existing methods rely heavily on the availability of large quantities of labeled training data. However, in the vast majority of real-world settings, manually collecting such large labeled datasets is infeasible due to the cost of labeling data or the paucity of data in a given domain. In this paper, we present a novel Adversarial Knowledge Transfer (AKT) framework for transferring knowledge from internet-scale unlabeled data to improve the performance of a classifier on a given visual recognition task. The proposed adversarial learning framework aligns the feature space of the unlabeled source data with the labeled target data such that the target classifier can be used to predict pseudo labels on the source data. An important novel aspect of our method is that the unlabeled source data can be of different classes from those of the labeled target data, and there is no need to define a separate pretext task, unlike some existing approaches. Extensive experiments well demonstrate that models learned using our approach hold a lot of promise across a variety of visual recognition tasks on multiple standard datasets. Project page is at \texttthttps://agupt013.github.io/akt.html. Akash Gupta 0001, Rameswar Panda, Sujoy Paul, Jianming Zhang 0001, Amit K. Roy-Chowdhury |
ACM Multimedia | 5 |
| 2020 | Distributed Multi-agent Video Fast-forwardingabstractIn many intelligent systems, a network of agents collaboratively perceives the environment for better and more efficient situation awareness. As these agents often have limited resources, it could be greatly beneficial to identify the content overlapping among camera views from different agents and leverage it for reducing the processing, transmission and storage of redundant/unimportant video frames. This paper presents a consensus-based distributed multi-agent video fast-forwarding framework, named DMVF, that fast-forwards multi-view video streams collaboratively and adaptively. In our framework, each camera view is addressed by a reinforcement learning based fast-forwarding agent, which periodically chooses from multiple strategies to selectively process video frames and transmits the selected frames at adjustable paces. During every adaptation period, each agent communicates with a number of neighboring agents, evaluates the importance of the selected frames from itself and those from its neighbors, refines such evaluation together with other agents via a system-wide consensus algorithm, and uses such evaluation to decide their strategy for the next period. Compared with approaches in the literature on a real-world surveillance video dataset VideoWeb, our method significantly improves the coverage of important frames and also reduces the number of frames processed in the system. Shuyue Lan, Zhilu Wang, Amit K. Roy-Chowdhury, Ermin Wei, Qi Zhu 0002 |
ACM Multimedia | 3 |
| 2020 | Context-Aware Query Selection for Active Learning in Event RecognitionabstractActivity recognition is a challenging problem with many practical applications. In addition to the visual features, recent approaches have benefited from the use of context, e.g., inter-relationships among the activities and objects. However, these approaches require data to be labeled, entirely available beforehand, and not designed to be updated continuously, which make them unsuitable for surveillance applications. In contrast, we propose a continuous-learning framework for context-aware activity recognition from unlabeled video, which has two distinct advantages over existing methods. First, it employs a novel active-learning technique that not only exploits the informativeness of the individual activities but also utilizes their contextual information during query selection; this leads to significant reduction in expensive manual annotation effort. Second, the learned models can be adapted online as more data is available. We formulate a conditional random field model that encodes the context and devise an information-theoretic approach that utilizes entropy and mutual information of the nodes to compute the set of most informative queries, which are labeled by a human. These labels are combined with graphical inference techniques for incremental updates. We provide a theoretical formulation of the active learning framework with an analytic solution. Experiments on six challenging datasets demonstrate that our framework achieves superior performance with significantly less manual labeling. Mahmudul Hasan 0003, Sujoy Paul, Anastasios I. Mourikis, Amit K. Roy-Chowdhury |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Heterogeneous Acceleration of HAR ApplicationsabstractHuman action recognition (HAR) is an important field of research that intercepts with areas such as image processing, computer vision, and the design of fast algorithms, among others. HAR has several important applications including healthcare monitoring, security and surveillance, assisted living, smart homes, and video search and indexing. Despite recent developments in the field, major challenges remain. For instance, HAR is computationally expensive. Tasks such as video preprocessing, feature extraction, feature quantization, and feature classification require the execution of millions of arithmetic operations for a video sequence lasting a few seconds. To address these problems, we propose a heterogeneous approach that is based on an extensive algorithmic and experimental analysis of the histogram of gradients application. We divide the application into four stages and evaluate each on the CPU, GPU, and FPGA platforms. Our heterogeneous design combines the strengths of both the FPGA and GPU platforms, and achieves a 1.3X speedup compared with a state-of-the-art GPU while being 1.5X more energy efficient than other homogeneous solutions, including FPGA-based designs. Moreover, our heterogeneous HAR design using fixed-point arithmetic has comparable accuracy to those of HAR algorithms using single precision floating point arithmetic. Jose M. Rodriguez Borbon, Xiaoyin Ma, Amit K. Roy-Chowdhury, Walid A. Najjar |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Construction of Diverse Image Datasets From Web Collections With Limited LabelingabstractImage datasets play a pivotal role in advancing computer vision and multimedia research. However, most of the datasets are created by extensive human effort and are extremely expensive to scale-up. To address these issues, several automatic and semi-automatic approaches have been proposed for creating datasets by refining web images. However, these approaches either include significant redundant images in the dataset or fail to provide a diverse enough set to train a robust classifier. Ideally, a representative subset should be both semantically and visually diverse so as to provide the maximum amount of information under the current budget. Most current approaches are entirely based on the analysis of visual features, which may not correlate well with image semantics, and hence, collected images may not be sufficient to give a detailed understanding of a category. In this paper, we propose a system for creating diverse image dataset collections from the web with limited manual labeling effort. It is based upon a semi-supervised sparse coding framework that employs a joint visual-semantic space to simultaneously utilize both the images and associated textual information from the web for dataset construction. In addition, the proposed system is online and is capable of collecting more discriminative images continuously as new data becomes available, which is also suitable for enriching the existing datasets. The experiments demonstrate that our system can create and enrich datasets with limited manual labeling, with better cross-dataset generalization capability and diversity compared to the state-of-the-art datasets. Niluthpol Chowdhury Mithun, Rameswar Panda, Amit K. Roy-Chowdhury |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Weakly Supervised Video Moment Retrieval From Text QueriesabstractThere have been a few recent methods proposed in text to video moment retrieval using natural language queries, but requiring full supervision during training. However, acquiring a large number of training videos with temporal boundary annotations for each text description is extremely time-consuming and often not scalable. In order to cope with this issue, in this work, we introduce the problem of learning from weak labels for the task of text to video moment retrieval. The weak nature of the supervision is because, during training, we only have access to the video-text pairs rather than the temporal extent of the video to which different text descriptions relate. We propose a joint visual-semantic embedding based framework that learns the notion of relevant segments from video using only video-level sentence descriptions. Specifically, our main idea is to utilize latent alignment between video frames and sentence descriptions using Text-Guided Attention (TGA). TGA is then used during the test phase to retrieve relevant moments. Experiments on two benchmark datasets demonstrate that our method achieves comparable performance to state-of-the-art fully supervised approaches. Niluthpol Chowdhury Mithun, Sujoy Paul, Amit K. Roy-Chowdhury |
CVPR | 3 |
| 2019 | Stealthy Adversarial Perturbations Against Real-Time Video Classification Systems
Shasha Li 0001, Ajaya Neupane, Sujoy Paul, Chengyu Song, Srikanth V. Krishnamurthy, Amit K. Roy-Chowdhury, Ananthram Swami |
NDSS | 6 |
| 2019 | Learning from Trajectories via Subgoal DiscoveryabstractLearning to solve complex goal-oriented tasks with sparse terminal-only rewards often requires an enormous number of samples. In such cases, using a set of expert trajectories could help to learn faster. However, Imitation Learning (IL) via supervised pre-training with these trajectories may not perform as well and generally requires additional finetuning with expert-in-the-loop. In this paper, we propose an approach which uses the expert trajectories and learns to decompose the complex main task into smaller sub-goals. We learn a function which partitions the state-space into sub-goals, which can then be used to design an extrinsic reward function. We follow a strategy where the agent first learns from the trajectories using IL and then switches to Reinforcement Learning (RL) using the identified sub-goals, to alleviate the errors in the IL step. To deal with states which are under-represented by the trajectory set, we also learn a function to modulate the sub-goal predictions. We show that our method is able to solve complex goal-oriented tasks, which other RL, IL or their combinations in literature are not able to solve. Sujoy Paul, Jeroen van Baar, Amit K. Roy-Chowdhury |
NeurIPS | 3 |
| 2019 | Frugal following: power thrifty object detection and tracking for mobile augmented realityabstractAccurate tracking of objects in the real world is highly desirable in Augmented Reality (AR) to aid proper placement of virtual objects in a user's view. Deep neural networks (DNNs) yield high precision in detecting and tracking objects, but they are energy-heavy and can thus be prohibitive for deployment on mobile devices. Towards reducing energy drain while maintaining good object tracking precision, we develop a novel software framework called MARLIN. MARLIN only uses a DNN as needed, to detect new objects or recapture objects that significantly change in appearance. It employs lightweight methods in between DNN executions to track the detected objects with high fidelity. We experiment with several baseline DNN models optimized for mobile devices, and via both offline and live object tracking experiments on two different Android phones (one utilizing a mobile GPU), we show that MARLIN compares favorably in terms of accuracy while saving energy significantly. Specifically, we show that MARLIN reduces the energy consumption by up to 73.3% (compared to an approach that executes the best baseline DNN continuously), and improves accuracy by up to 19× (compared to an approach that infrequently executes the same best baseline DNN). Moreover, while in 75% or more cases, MARLIN incurs at most a 7.36% reduction in location accuracy (using the common IOU metric), in more than 46% of the cases, MARLIN even improves the IOU compared to the continuous, best DNN approach. Kittipat Apicharttrisorn, Xukan Ran, Jiasi Chen, Srikanth V. Krishnamurthy, Amit K. Roy-Chowdhury |
SenSys | 5 |
| 2019 | Adaptation of person re-identification models for on-boarding new camera(s)
Rameswar Panda, Amran Bhuiyan, Vittorio Murino, Amit K. Roy-Chowdhury |
Pattern Recognit. | 4 |
| 2019 | Exploiting Typicality for Selecting Informative and Anomalous Samples in VideosabstractIn this paper, we present a novel approach to find informative and anomalous samples in videos exploiting the concept of typicality from information theory. In most video analysis tasks, selection of the most informative samples from a huge pool of training data in order to learn a good recognition model is an important problem. Furthermore, it is also useful to reduce the annotation cost as it is time-consuming to annotate unlabeled samples. Typicality is a simple and powerful technique which can be applied to compress the training data to learn a good classification model. In a continuous video clip, an activity shares a strong correlation with its previous activities. We assume that the activity samples that appear in a video form a Markov chain. We explicitly show how typicality can be utilized in this scenario. We compute an atypical score for a sample using typicality and the Markovian property, which can be applied to two challenging vision problems-(a) sample selection for learning activity recognition models, and (b) anomaly detection. In the first case, our approach leads to a significant reduction of manual labeling cost while achieving similar or better recognition performance compared to a model trained with the entire training set. For the latter case, the atypical score has been exploited in identifying anomalous activities in videos where our results demonstrate the effectiveness of the proposed framework over other recent strategies. Jawadul H. Bappy, Sujoy Paul, Ertem Tuncel, Amit K. Roy-Chowdhury |
IEEE Trans. Image Process. | 4 |
| 2019 | Hybrid LSTM and Encoder-Decoder Architecture for Detection of Image ForgeriesabstractWith advanced image journaling tools, one can easily alter the semantic meaning of an image by exploiting certain manipulation techniques such as copy clone, object splicing, and removal, which mislead the viewers. In contrast, the identification of these manipulations becomes a very challenging task as manipulated regions are not visually apparent. This paper proposes a high-confidence manipulation localization architecture that utilizes resampling features, long short-term memory (LSTM) cells, and an encoder-decoder network to segment out manipulated regions from non-manipulated ones. Resampling features are used to capture artifacts, such as JPEG quality loss, upsampling, downsampling, rotation, and shearing. The proposed network exploits larger receptive fields (spatial maps) and frequency-domain correlation to analyze the discriminative characteristics between the manipulated and non-manipulated regions by incorporating the encoder and LSTM network. Finally, the decoder network learns the mapping from low-resolution feature maps to pixel-wise predictions for image tamper localization. With the predicted mask provided by the final layer (softmax) of the proposed architecture, end-to-end training is performed to learn the network parameters through back-propagation using the ground-truth masks. Furthermore, a large image splicing dataset is introduced to guide the training process. The proposed method is capable of localizing image manipulations at the pixel level with high precision, which is demonstrated through rigorous experimentation on three diverse datasets. Jawadul H. Bappy, Cody Simons, Lakshmanan Nataraj, B. S. Manjunath, Amit K. Roy-Chowdhury |
IEEE Trans. Image Process. | 5 |
| 2018 | FFNet: Video Fast-Forwarding via Reinforcement LearningabstractFor many applications with limited computation, communication, storage and energy resources, there is an imperative need of computer vision methods that could select an informative subset of the input video for efficient processing at or near real time. In the literature, there are two relevant groups of approaches: generating a "trailer" for a video or fast-forwarding while watching/processing the video. The first group is supported by video summarization techniques, which require processing of the entire video to select an important subset for showing to users. In the second group, current fast-forwarding methods depend on either manual control or automatic adaptation of playback speed, which often do not present an accurate representation and may still require processing of every frame. In this paper, we introduce FastForwardNet (FFNet), a reinforcement learning agent that gets inspiration from video summarization and does fast-forwarding differently. It is an online framework that automatically fast-forwards a video and presents a representative subset of frames to users on the fly. It does not require processing the entire video, but just the portion that is selected by the fast-forward agent, which makes the process very computationally efficient. The online nature of our proposed method also enables the users to begin fast-forwarding at any point of the video. Experiments on two real-world datasets demonstrate that our method can provide better representation of the input video (about 6%-20% improvement on coverage of important frames) with much less processing requirement (more than 80% reduction in the number of frames processed). Shuyue Lan, Rameswar Panda, Qi Zhu 0002, Amit K. Roy-Chowdhury |
CVPR | 4 |
| 2018 | Exploiting Transitivity for Learning Person Re-Identification Models on a BudgetabstractMinimization of labeling effort for person re-identification in camera networks is an important problem as most of the existing popular methods are supervised and they require large amount of manual annotations, acquiring which is a tedious job. In this work, we focus on this labeling effort minimization problem and approach it as a subset selection task where the objective is to select an optimal subset of image-pairs for labeling without compromising performance. Towards this goal, our proposed scheme first represents any camera network (with k number of cameras) as an edge weighted complete k-partite graph where each vertex denotes a person and similarity scores between persons are used as edge-weights. Then in the second stage, our algorithm selects an optimal subset of pairs by solving a triangle free subgraph maximization problem on the k-partite graph. This sub-graph weight maximization problem is NP-hard (at least for k = 4) which means for large datasets the optimization problem becomes intractable. In order to make our framework scalable, we propose two polynomial time approximately-optimal algorithms. The first algorithm is a 1/2-approximation algorithm which runs in linear time in the number of edges. The second algorithm is a greedy algorithm with sub-quadratic (in number of edges) time-complexity. Experiments on three state-of-the-art datasets depict that the proposed approach requires on an average only 8-15% manually labeled pairs in order to achieve the performance when all the pairs are manually annotated. Sourya Roy, Sujoy Paul, Neal E. Young, Amit K. Roy-Chowdhury |
CVPR | 4 |
| 2018 | Contemplating Visual Emotions: Understanding and Overcoming Dataset Bias
Rameswar Panda, Jianming Zhang 0001, Joon-Young Lee, Xin Lu 0006, Amit K. Roy-Chowdhury |
ECCV (2) | 6 |
| 2018 | W-TALC: Weakly-Supervised Temporal Activity Localization and Classification
Sujoy Paul, Sourya Roy, Amit K. Roy-Chowdhury |
ECCV (4) | 3 |
| 2018 | Incorporating Scalability in Unsupervised Spatio- Temporal Feature LearningabstractDeep neural networks are efficient learning machines which leverage upon a large amount of manually labeled data for learning discriminative features. However, acquiring substantial amount of supervised data, especially for videos can be a tedious job across various computer vision tasks. This necessitates learning of visual features from videos in an unsupervised setting. In this paper, we propose a computationally simple, yet effective, framework to learn spatio-temporal feature embedding from unlabeled videos. We train a Convolutional 3D Siamese network using positive and negative pairs mined from videos under certain probabilistic assumptions. Experimental results on three datasets demonstrate that our proposed framework is able to learn weights which can be used for same as well as cross dataset and tasks. Sujoy Paul, Sourya Roy, Amit K. Roy-Chowdhury |
ICASSP | 3 |
| 2018 | Deep Learning Based Identity Verification in Renaissance PortraitsabstractThe identity of subjects in many portraits has been a matter of debate for art historians that relied upon subjective analysis of facial features to resolve ambiguity in sitter identity. Developing automated face verification technique has thus garnered interest to provide a quantitative way to reinforce the decision arrived at by the art historians. However, most existing works often fail to resolve ambiguities concerning the identity of the subjects due to significant variation in artistic styles and the limited availability and authenticity of art images. To these ends, we explore the use of deep Siamese Convolutional Neural Networks (CNN) to provide a measure of similarity between a pair of portraits. To mitigate limited training data issue, we employ CNN based style-transfer technique that creates several new images by recasting an art style to other images, keeping original image content unchanged. The resulting system thereby learns features which are discriminative and invariant to changes in artistic styles. Our approach shows significant improvement over baselines and state-of-the-art methods on several examples which are identified by art historians as being very challenging and controversial. Akash Gupta 0001, Niluthpol Chowdhury Mithun, Conrad Rudolph, Amit K. Roy-Chowdhury |
ICME | 4 |
| 2018 | Learning Joint Embedding with Multimodal Cues for Cross-Modal Video-Text RetrievalabstractConstructing a joint representation invariant across different modalities (e.g., video, language) is of significant importance in many multimedia applications. While there are a number of recent successes in developing effective image-text retrieval methods by learning joint representations, the video-text retrieval task, however, has not been explored to its fullest extent. In this paper, we study how to effectively utilize available multimodal cues from videos for the cross-modal video-text retrieval task. Based on our analysis, we propose a novel framework that simultaneously utilizes multi-modal features (different visual characteristics, audio inputs, and text) by a fusion strategy for efficient retrieval. Furthermore, we explore several loss functions in training the embedding and propose a modified pairwise ranking loss for the task. Experiments on MSVD and MSR-VTT datasets demonstrate that our method achieves significant performance gain compared to the state-of-the-art approaches. Niluthpol Chowdhury Mithun, Juncheng Li 0001, Florian Metze, Amit K. Roy-Chowdhury |
ICMR | 4 |
| 2018 | Webly Supervised Joint Embedding for Cross-Modal Image-Text RetrievalabstractCross-modal retrieval between visual data and natural language description remains a long-standing challenge in multimedia. While recent image-text retrieval methods offer great promise by learning deep representations aligned across modalities, most of these methods are plagued by the issue of training with small-scale datasets covering a limited number of images with ground-truth sentences. Moreover, it is extremely expensive to create a larger dataset by annotating millions of images with sentences and may lead to a biased model. Inspired by the recent success of webly supervised learning in deep neural networks, we capitalize on readily-available web images with noisy annotations to learn robust image-text joint representation. Specifically, our main idea is to leverage web images and corresponding tags, along with fully annotated datasets, in training for learning the visual-semantic joint embedding. We propose a two-stage approach for the task that can augment a typical supervised pair-wise ranking loss based formulation with weakly-annotated web images to learn a more robust visual-semantic embedding. Experiments on two standard benchmark datasets demonstrate that our method achieves a significant performance gain in image-text retrieval compared to state-of-the-art approaches. Niluthpol Chowdhury Mithun, Rameswar Panda, Evangelos E. Papalexakis, Amit K. Roy-Chowdhury |
ACM Multimedia | 4 |
| 2018 | Learning Long-Term Invariant Features for Vision-Based LocalizationabstractConstructing a feature representation invariant to certain types of geometric and photometric transformations is of significant importance in many computer vision applications. In spite of significant effort, developing invariant feature representations remains a challenging problem. Most of the existing representations often fail to satisfy the longterm repeatability requirements of specific applications like vision-based localization, applications whose domain includes significant, non-uniform illumination and environmental changes. To these ends, we explore the use of natural image pairs (i.e. images captured of the same location but at different times) as an additional source of supervision to generate an improved feature representation for the task of vision-based localization. Specifically, we resort to training deep denoising autoencoder, with CNN feature representation of one image in the pair being treated as a noisy version of the other. The resulting system thereby learns localization features which are both discriminative and invariant to illumination and environmental changes. In experiments tailored towards vision-based localization, features generated using the proposed method produced higher matching rates than state-of-the-art image features. Niluthpol Chowdhury Mithun, Cody Simons, Robert Casey, Stefan Hilligardt, Amit K. Roy-Chowdhury |
WACV | 5 |
| 2017 | The Impact of Typicality for Informative Representative SelectionabstractIn computer vision, selection of the most informative samples from a huge pool of training data in order to learn a good recognition model is an active research problem. Furthermore, it is also useful to reduce the annotation cost, as it is time consuming to annotate unlabeled samples. In this paper, motivated by the theories in data compression, we propose a novel sample selection strategy which exploits the concept of typicality from the domain of information theory. Typicality is a simple and powerful technique which can be applied to compress the training data to learn a good classification model. In this work, typicality is used to identify a subset of the most informative samples for labeling, which is then used to update the model using active learning. The proposed model can take advantage of the inter-relationships between data samples. Our approach leads to a significant reduction of manual labeling cost while achieving similar or better recognition performance compared to a model trained with entire training set. This is demonstrated through rigorous experimentation on five datasets. Jawadul H. Bappy, Sujoy Paul, Ertem Tuncel, Amit K. Roy-Chowdhury |
CVPR | 4 |
| 2017 | Unsupervised Adaptive Re-identification in Open World Dynamic Camera NetworksabstractPerson re-identification is an open and challenging problem in computer vision. Existing approaches have concentrated on either designing the best feature representation or learning optimal matching metrics in a static setting where the number of cameras are fixed in a network. Most approaches have neglected the dynamic and open world nature of the re-identification problem, where a new camera may be temporarily inserted into an existing system to get additional information. To address such a novel and very practical problem, we propose an unsupervised adaptation scheme for re-identification models in a dynamic camera network. First, we formulate a domain perceptive re-identification method based on geodesic flow kernel that can effectively find the best source camera (already installed) to adapt with a newly introduced target camera, without requiring a very expensive training phase. Second, we introduce a transitive inference algorithm for re-identification that can exploit the information from best source camera to improve the accuracy across other camera pairs in a network of multiple cameras. Extensive experiments on four benchmark datasets demonstrate that the proposed approach significantly outperforms the state-of-the-art unsupervised learning based alternatives whilst being extremely efficient to compute. Rameswar Panda, Amran Bhuiyan, Vittorio Murino, Amit K. Roy-Chowdhury |
CVPR | 4 |
| 2017 | Collaborative Summarization of Topic-Related Videos
Rameswar Panda, Amit K. Roy-Chowdhury |
CVPR | 2 |
| 2017 | Non-uniform Subset Selection for Active Learning in Structured DataabstractSeveral works have shown that relationships between data points (i.e., context) in structured data can be exploited to obtain better recognition performance. In this paper, we explore a different, but related, problem: how can these inter-relationships be used to efficiently learn and continuously update a recognition model, with minimal human labeling effort. Towards this goal, we propose an active learning framework to select an optimal subset of data points for manual labeling by exploiting the relationships between them. We construct a graph from the unlabeled data to represent the underlying structure, such that each node represents a data point, and edges represent the inter-relationships between them. Thereafter, considering the flow of beliefs in this graph, we choose those samples for labeling which minimize the joint entropy of the nodes of the graph. This results in significant reduction in manual labeling effort without compromising recognition performance. Our method chooses non-uniform number of samples from each batch of streaming data depending on its information content. Also, the submodular property of our objective function makes it computationally efficient to optimize. The proposed framework is demonstrated in various applications, including document analysis, scene-object recognition, and activity recognition. Sujoy Paul, Jawadul H. Bappy, Amit K. Roy-Chowdhury |
CVPR | 3 |
| 2017 | Sparse modeling for topic-oriented video summarizationabstractWhile most existing video summarization approaches aim to extract an informative summary of a single video, we propose an unsupervised framework for summarizing topic-related videos by exploring complementarity within videos. We develop a novel sparse optimization method to extract a diverse summary that is both interesting and representative in describing the video collection. To efficiently solve our optimization problem, we develop an alternating minimization algorithm that minimizes the overall objective function with respect to one video at a time while fixing the other videos. Experimental results demonstrate that our approach clearly outperforms the state-of-the-art methods. Rameswar Panda, Amit K. Roy-Chowdhury |
ICASSP | 2 |
| 2017 | Exploiting Spatial Structure for Localizing Manipulated Image RegionsabstractThe advent of high-tech journaling tools facilitates an image to be manipulated in a way that can easily evade state-of-the-art image tampering detection approaches. The recent success of the deep learning approaches in different recognition tasks inspires us to develop a high confidence detection framework which can localize manipulated regions in an image. Unlike semantic object segmentation where all meaningful regions (objects) are segmented, the localization of image manipulation focuses only the possible tampered region which makes the problem even more challenging. In order to formulate the framework, we employ a hybrid CNN-LSTM model to capture discriminative features between manipulated and non-manipulated regions. One of the key properties of manipulated regions is that they exhibit discriminative features in boundaries shared with neighboring non-manipulated pixels. Our motivation is to learn the boundary discrepancy, i.e., the spatial structure, between manipulated and non-manipulated regions with the combination of LSTM and convolution layers. We perform end-to-end training of the network to learn the parameters through back-propagation given ground-truth mask information. The overall framework is capable of detecting different types of image manipulations, including copy-move, removal and splicing. Our model shows promising results in localizing manipulated regions, which is demonstrated through rigorous experimentation on three diverse datasets. Jawadul H. Bappy, Amit K. Roy-Chowdhury, Jason Bunk, Lakshmanan Nataraj, B. S. Manjunath |
ICCV | 2 |
| 2017 | Joint Prediction of Activity Labels and Starting Times in Untrimmed VideosabstractMost of the existing works on human activity analysis focus on recognition or early recognition of the activity labels from complete or partial observations. Predicting the labels of future unobserved activities where no frames of the predicted activities have been observed is a challenging problem, with important applications, which has not been explored much. Associated with the future label prediction problem is the problem of predicting the starting time of the next activity. In this work, we propose a system that is able to infer about the labels and the starting times of future activities. Activities are characterized by the previous activity sequence (which is observed), as well as the objects present in the scene during their occurrence. We propose a network similar to a hybrid Siamese network with three branches to jointly learn both the future label and the starting time. The first branch takes visual features from the objects present in the scene using a fully connected network, the second branch takes previous activity features using a LSTM network to model long-term sequential relationships and the third branch captures the last observed activity features to model the context of inter-activity time using another fully connected network. These concatenated features are used for both label and time prediction. Experiments on two challenging datasets demonstrate that our framework for joint prediction of activity label and starting time improves the performance of both, and outperforms the state-of-the-arts. Tahmida Mahmud, Mahmudul Hasan 0003, Amit K. Roy-Chowdhury |
ICCV | 3 |
| 2017 | Weakly Supervised Summarization of Web VideosabstractMost of the prior works summarize videos by either exploring different heuristically designed criteria in an unsupervised way or developing fully supervised algorithms by leveraging human-crafted training data in form of video-summary pairs or importance annotations. However, unsupervised methods are blind to the video category and often fail to produce semantically meaningful video summaries. On the other hand, acquisition of large amount of training data in supervised approaches is non-trivial and may lead to a biased model. Different from existing methods, we introduce a weakly supervised approach that requires only video-level annotation for summarizing web videos. Casting the problem as a weakly supervised learning problem, we propose a flexible deep 3D CNN architecture to learn the notion of importance using only video-level annotation, and without any human-crafted training data. Specifically, our main idea is to leverage multiple videos of a category to automatically learn a parametric model for categorizing videos and then adopt the model to find important segments from a given video as the ones which have maximum influence to the model output. Furthermore, to unleash the full potential of our 3D CNN architecture, we also explored a series of good practices to reduce the influence of limited training data while summarizing videos. Experiments on two challenging and diverse datasets well demonstrate that our approach produces superior quality video summaries compared to several recently proposed approaches. Rameswar Panda, Abir Das, Ziyan Wu 0001, Jan Ernst, Amit K. Roy-Chowdhury |
ICCV | 5 |
| 2017 | Energy Efficient Object Detection in Camera Sensor NetworksabstractA wireless camera network can provide situation awareness information (e.g., humans in distress) in scenarios such as disaster recovery. If such camera sensors are battery operated, sending raw video feeds back to a central controller can be expensive in terms of energy consumption. Further, if all cameras were to use the optimal processing algorithm for object decision, they may also expend unnecessary energy. Stated otherwise, cameras that capture the same objects may not all have to use the optimal algorithm to achieve a desired accuracy, and this can save processing energy costs. In this paper, our objective is to design and implement a framework that can support coordination among cameras to deliver highly accurate detection of objects in an energy efficient way. The framework, which we call EECS (for energy efficient camera sensors), estimates the detection accuracy and energy costs incurred (both the processing and communication costs are taken into account) with each detection algorithm for each camera, and comes up with a choice of cameras for sending information pertaining to the object of interest. This set of cameras and the video processing algorithms that they must use, are chosen so as to minimize the energy expenditures, given a desired detection accuracy. We implement EECS on a camera network built with smartphones, and demonstrate that it reduces the energy consumption by up to 40% while ensuring a object detection accuracy of over 86%. Tuan Dao, Karim Khalil, Amit K. Roy-Chowdhury, Srikanth V. Krishnamurthy, Lance M. Kaplan |
ICDCS | 3 |
| 2017 | Accurate and Timely Situation Awareness Retrieval from a Bandwidth Constrained Camera NetworkabstractWireless cameras can be used to gather situation awareness information (e.g., humans in distress) in disaster recovery scenarios. However, blindly sending raw video streams from such cameras, to an operations center or controller can be prohibitive in terms of bandwidth. Further, these raw streams could contain either redundant or irrelevant information. Thus, we ask "how do we extract accurate situation awareness information from such camera nodes and send it in a timely manner, back to the operations center?" Towards this, we design ACTION, a framework that (a) detects objects of interest (e.g., humans) from the video streams, (b) combines these streams intelligently to eliminate redundancies and (c) transmits only parts of the feeds that are sufficient in achieving a desired detection accuracy to the controller. ACTION uses small amounts of metadata to determine if the objects from different camera feeds are the same. A resource-aware greedy algorithm is used to select a subset of video feeds that are associated with the same object, so as to provide a desired accuracy, for being sent to the operations center. Our evaluations show that ACTION helps reduce the network usage up to threefold, and yet achieves a high detection accuracy of ≈ 90%. Tuan Dao, Amit K. Roy-Chowdhury, Nasser M. Nasrabadi, Srikanth V. Krishnamurthy, Prasant Mohapatra, Lance M. Kaplan |
MASS | 2 |
| 2017 | Real Estate Image ClassificationabstractPosting pictures is a necessary part of advertising a home for sale. Agents typically sort through dozens of images from which to pick the most complimentary ones. This is a manual effort involving annotating images accompanied by descriptions (bedroom, bathroom, attic, etc.). When volumes are small, manual annotation is not a problem, but there is a point where this becomes too burdensome and ultimately infeasible. Here, we propose an approach based on computer vision methodology to radically increase the efficiency of such tasks. We present a high-confidence image classification framework, whose inputs are images and outputs are labels. The core of the classification algorithm is long short term memory (LSTM), and fully connected neural networks, along with a substantial preprocessing using 'contrast-limited adaptive histogram equalization (CLAHE) for image enhancement. Since, there is no standard benchmark containing a comprehensive dataset of well-annotated real estate images, we introduce Real Estate Image (REI) database for evaluating the image classification algorithms. Therein we demonstrate empirics based on our proposed framework on the new REI dataset, as well as on the SUN dataset. Jawadul H. Bappy, Joseph R. Barr, Narayanan Srinivasan 0002, Amit K. Roy-Chowdhury |
WACV | 4 |
| 2017 | Continuous adaptation of multi-camera person identification models through sparse non-redundant representative selection
Abir Das, Rameswar Panda, Amit K. Roy-Chowdhury |
Comput. Vis. Image Underst. | 3 |
| 2017 | Optimal Landmark Selection for Registration of 4D Confocal Image Stacks in ArabidopsisabstractTechnologically advanced imaging techniques have allowed us to generate and study the internal part of a tissue over time by capturing serial optical images that contain spatio-temporal slices of hundreds of tightly packed cells. Image registration of such live-imaging datasets of developing multicelluar tissues is one of the essential components of all image analysis pipelines. In this paper, we present a fully automated 4D(X-Y-Z-T) registration method of live imaging stacks that takes care of both temporal and spatial misalignments. We present a novel landmark selection methodology where the shape features of individual cells are not of high quality and highly distinguishable. The proposed registration method finds the best image slice correspondence from consecutive image stacks to account for vertical growth in the tissue and the discrepancy in the choice of the starting focal point. Then, it uses local graph-based approach to automatically find corresponding landmark pairs, and finally the registration parameters are used to register the entire image stack. The proposed registration algorithm combined with an existing tracking method is tested on multiple image stacks of tightly packed cells of Arabidopsis shoot apical meristem and the results show that it significantly improves the accuracy of cell lineages and division statistics. Katya Mkrtchyan, Anirban Chakraborty 0001, Amit K. Roy-Chowdhury |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2017 | Opportunistic Image Acquisition of Individual and Group Activities in a Distributed Camera NetworkabstractThe decreasing cost and size of video sensors has led to camera networks becoming pervasive in our lives. However, the ability to analyze these images effectively is very much a function of the quality of the acquired images. In this paper, we consider the problem of automatically controlling the fields of view of individual pan-tilt-zoom (PTZ) cameras in a camera network leading to improved situation awareness (e.g., where and what are the critical targets and events) in a region of interest. The network of cameras attempts to observe the entire region of interest at some minimum resolution while opportunistically acquiring high resolution images of critical events in real time. Since many activities involve groups of people interacting, an important decision that the network needs to make is whether to focus on individuals or groups of them. This is achieved by understanding the performance of video analysis tasks and designing camera control strategies to improve a metric that quantifies the quality of the source imagery. Optimization strategies, along with a distributed implementation, are proposed, and their theoretical properties analyzed. The proposed methods bring together computer vision and network control ideas. The performance of the proposed methodologies discussed herein has been evaluated on a real-life wireless network of PTZ capable cameras. Chong Ding, Jawadul H. Bappy, Jay A. Farrell, Amit K. Roy-Chowdhury |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Diversity-Aware Multi-Video SummarizationabstractMost video summarization approaches have focused on extracting a summary from a single video; we propose an unsupervised framework for summarizing a collection of videos. We observe that each video in the collection may contain some information that other videos do not have, and thus exploring the underlying complementarity could be beneficial in creating a diverse informative summary. We develop a novel diversity-aware sparse optimization method for multi-video summarization by exploring the complementarity within the videos. Our approach extracts a multi-video summary, which is both interesting and representative in describing the whole video collection. To efficiently solve our optimization problem, we develop an alternating minimization algorithm that minimizes the overall objective function with respect to one video at a time while fixing the other videos. Moreover, we introduce a new benchmark data set, Tour20, that contains 140 videos with multiple manually created summaries, which were acquired in a controlled experiment. Finally, by extensive experiments on the new Tour20 data set and several other multi-view data sets, we show that the proposed approach clearly outperforms the state-of-the-art methods on the two problems-topic-oriented video summarization and multi-view video summarization in a camera network. Rameswar Panda, Niluthpol Chowdhury Mithun, Amit K. Roy-Chowdhury |
IEEE Trans. Image Process. | 3 |
| 2017 | Multi-View Surveillance Video Summarization via Joint Embedding and Sparse OptimizationabstractMost traditional video summarization methods are designed to generate effective summaries for single-view videos, and thus, they cannot fully exploit the complicated intra- and inter-view correlations in summarizing multi-view videos in a camera network. In this paper, with the aim of summarizing multi-view videos, we introduce a novel unsupervised framework via joint embedding and sparse representative selection. The objective function is twofold. The first is to capture the multiview correlations via an embedding, which helps in extracting a diverse set of representatives. The second is to use a ℓ2,1-norm to model the sparsity while selecting representative shots for the summary. We propose to jointly optimize both of the objectives, such that embedding cannot only characterize the correlations, but also indicate the requirements of sparse representative selection. We present an efficient alternating algorithm based on half-quadratic minimization to solve the proposed non-smooth and non-convex objective with convergence analysis. A key advantage of the proposed approach with respect to the state-of-the-art is that it can summarize multi-view videos without assuming any prior correspondences/alignment between them, e.g., uncalibrated camera networks. Rigorous experiments on several multi-view datasets demonstrate that our approach clearly outperforms the state-of-the-art methods. Rameswar Panda, Amit K. Roy-Chowdhury |
IEEE Trans. Multim. | 2 |
| 2017 | Managing Redundant Content in Bandwidth Constrained Wireless NetworksabstractImages/videos are often uploaded in situations like disasters. This can tax the network in terms of increased load and thereby upload latency, and this can be critical for response activities. In such scenarios, prior work has shown that there is significant redundancy in the content (e.g., similar photos taken by users) transferred. By intelligently suppressing/deferring transfers of redundant content, the load can be significantly reduced, thereby facilitating the timely delivery of unique, possibly critical information. A key challenge here however, is detecting “what content is similar,” given that the content is generated by uncoordinated user devices. Toward addressing this challenge, we propose a framework, wherein a service to which the content is to be uploaded first solicits metadata (e.g., image features) from any device uploading content. By intelligently comparing this metadata with that associated with previously uploaded content, the service effectively identifies (and thus enables the suppression of) redundant content. Our evaluations on a testbed of 20 Android smartphones and via ns3 simulations show that we can identify similar content with a 70% true positive rate and a 1% false positive rate. The resulting reduction in redundant content transfers translates to a latency reduction of 44 % for unique content. Tuan Dao, Amit K. Roy-Chowdhury, Harsha V. Madhyastha, Srikanth V. Krishnamurthy, Thomas La Porta |
IEEE/ACM Trans. Netw. | 2 |
| 2016 | Learning Temporal Regularity in Video SequencesabstractPerceiving meaningful activities in a long video sequence is a challenging problem due to ambiguous definition of 'meaningfulness' as well as clutters in the scene. We approach this problem by learning a generative model for regular motion patterns (termed as regularity) using multiple sources with very limited supervision. Specifically, we propose two methods that are built upon the autoencoders for their ability to work with little to no supervision. We first leverage the conventional handcrafted spatio-temporal local features and learn a fully connected autoencoder on them. Second, we build a fully convolutional feed-forward autoencoder to learn both the local features and the classifiers as an end-to-end learning framework. Our model can capture the regularities from multiple datasets. We evaluate our methods in both qualitative and quantitative ways - showing the learned regularity of videos in various aspects and demonstrating competitive performance on anomaly detection datasets as an application. Mahmudul Hasan 0003, Jan Neumann, Amit K. Roy-Chowdhury, Larry Davis 0001 |
CVPR | 4 |
| 2016 | Online Adaptation for Joint Scene and Object Classification
Jawadul H. Bappy, Sujoy Paul, Amit K. Roy-Chowdhury |
ECCV (8) | 3 |
| 2016 | Temporal Model Adaptation for Person Re-identification
Niki Martinel, Abir Das, Christian Micheloni, Amit K. Roy-Chowdhury |
ECCV (4) | 4 |
| 2016 | OSNI: Searching for Needles in a Haystack of Social Network DataabstractThis paper presents the Online Social Network Investigator (OSNI), a scalable distributed system to search social net- work data, based on a spatiotemporal window and a list of keywords. Given that only 2% of tweets are geolocated, we have implemented and compared various state-of-art loca- tion estimation techniques. Further, to enrich the context of posts, associations of images to terms are estimated through various classication techniques. The accuracies of these es- timations are evaluated on large real datasets. OSNI's query interface is available on the Web. Shiwen Cheng, James Fang, Vagelis Hristidis, Harsha V. Madhyastha, Niluthpol Chowdhury Mithun, Dorian Jean Perkins, Amit K. Roy-Chowdhury, Moloud Shahbazi, Vassilis J. Tsotras |
EDBT | 7 |
| 2016 | Optimizing hardware design for Human Action RecognitionabstractHuman action recognition (HAR) is an important topic in computer vision having a wide range of applications: health care, assisted living, surveillance, security, gaming, etc. Despite significant amount of work having been conducted in this area in recent years, the execution speed still limits real-time applications. Moreover, it is highly desirable to have the compute-intensive feature extraction stage done right at the output of the camera to extract and transfer only action feature in multi-camera network setting and hence reduce network bandwidth requirement. In this work, we first evaluate the possibility to perform feature extraction under reduced precision fixed-point arithmetic to ease hardware resource requirements. We compared the Histogram of Oriented Gradient in 3D (HOG3D) feature extraction with state-of-the-art Convolutional Neural Networks (CNNs) methods and shown the later to be 75× slower than the former. Our experiment shows that by re-training the classifier with reduced data precision, the classification performs as well as the original double-precision floating-point. Based on this result, we implement an FPGA-based HAR feature extraction for near camera processing using fixed-point data representation and arithmetic. This implementation, using a single Xilinx Virtex 6 FPGA, achieves about 70× speedup over multicore CPU. Furthermore, a GPU implementation of HAR is introduced with 80× speedup over CPU (on an Nvidia Tesla K20). Last but not least, a power comparison is presented for the three platforms. Xiaoyin Ma, Jose M. Rodriguez Borbon, Walid A. Najjar, Amit K. Roy-Chowdhury |
FPL | 4 |
| 2016 | CNN based region proposals for efficient object detectionabstractIn computer vision, object detection is addressed as one of the most challenging problems as it is prone to localization and classification error. The current best-performing detectors are based on the technique of finding region proposals in order to localize objects. Despite having very good performance, these techniques are computationally expensive due to having large number of proposed regions. In this paper, we develop a high-confidence region-based object detection framework that boosts up the classification performance with less computational burden. In order to formulate our framework, we consider a deep network that activates the semantically meaningful regions in order to localize objects. These activated regions are used as input to a convolutional neural network (CNN) to extract deep features. With these features, we train a set of class-specific binary classifiers to predict the object labels. Our new region-based detection technique significantly reduces the computational complexity and improves the performance in object detection. We perform rigorous experiments on PASCAL, SUN, MIT-67 Indoor and MSRC datasets to demonstrate that our proposed framework outperforms other state-of-the-art methods in recognizing objects. Jawadul H. Bappy, Amit K. Roy-Chowdhury |
ICIP | 2 |
| 2016 | A poisson process model for activity forecastingabstractActivity forecasting has recently become an active research area for its importance in critical applications like automated navigation and human-computer interaction. However, for a video observed upto a certain time, all of the existing forecasting works focus on predicting the activity label, i.e., predicting what the next unobserved activity is. To the best of our knowledge, no work has answered the crucial question yet: when the next unobserved activity will occur. In this paper, we propose an approach for predicting the starting time of the next unobserved activity without assuming that we know its label. We model activities occurring at a variable rate using a Log-Gaussian Cox Process (LGCP) and learn the rate function from the training data. Then the starting time is predicted using importance sampling algorithm. In our experiments on the challenging MPII-Cooking dataset, we find that both the label of the last observed activity and the label of the activity being predicted affect the time prediction accuracy. Tahmida Mahmud, Mahmudul Hasan 0003, Anirban Chakraborty 0001, Amit K. Roy-Chowdhury |
ICIP | 4 |
| 2016 | Embedded sparse coding for summarizing multi-view videosabstractMost traditional video summarization methods are designed to generate effective summaries for single-view videos, and thus they cannot fully exploit the complicated intra- and inter-view correlations in summarizing multi-view videos. In this paper, we introduce a novel framework for summarizing multi-view videos in a way that takes into consideration both intra- and inter-view correlations in a joint embedding space. We learn the embedding by minimizing an objective function that has two terms: one due to intra-view correlations and another due to inter-view correlations across the multiple views. The solution is obtained by using a Majorization-Minimization algorithm that monotonically decreases the cost function in each iteration. We then employ a sparse representative selection approach over the learned embedding space to summarize the multi-view videos. Experiments on several multi-view datasets demonstrate that the proposed approach clearly outperforms the state-of-the-art methods. Rameswar Panda, Abir Das, Amit K. Roy-Chowdhury |
ICIP | 3 |
| 2016 | Efficient selection of informative and diverse training samples with applications in scene classificationabstractThe huge amount of time required to construct a set of labeled images to train a classifier has led researchers to develop algorithms which can identify the most informative training images, such that labelling those will be sufficient to achieve a considerable classification accuracy. In this paper we focus on choosing a subset of the most informative and diverse images based on which the classification model can be learned efficiently. The size of the subset to be chosen is determined by the available budget for manual labeling. Although the problem of identifying the informative images can be solved by active learning algorithms, it will require a set of labeled images for initial model construction, which is not required in our method as we identify the best samples at one shot. We incorporate the concepts of strong and weak teacher to help the learner to learn the model efficiently with limited budget for manual labeling. We perform rigorous experiments on two challenging scene classification datasets to demonstrate the effectiveness of our algorithm. Sujoy Paul, Jawadul H. Bappy, Amit K. Roy-Chowdhury |
ICIP | 3 |
| 2016 | Adaptive algorithm selection, with applications in pedestrian detectionabstractComputer vision algorithms are known to be extremely sensitive to the environmental conditions in which the data is captured, e.g., lighting conditions and target density. Tuning of parameters or choosing a completely new algorithm is often needed to achieve a certain performance level. In this paper, we focus on this problem and propose a framework to automatically choose the “best” algorithm-parameter combination (often referred to as the best algorithm for simplicity in this paper) for a certain input data. This necessitates developing a mechanism to switch among different algorithms and parameters as the nature of the input video changes. Specifically, our proposed algorithm calculates a similarity function between a test video segment and a training video segment. Similarity between training and test dataset indicates the same algorithm can be applied to both of them. We design a cost function with this similarity measure and a constraint on the number of switches. In the experiments, we apply our algorithm to the problem of pedestrian detection. We show how to adaptively select among 7 algorithm-parameter combinations and provide promising results on 3 publicly available datasets. Shu Zhang 0007, Qi Zhu 0002, Amit K. Roy-Chowdhury |
ICIP | 3 |
| 2016 | Inter-dependent CNNs for joint scene and object recognitionabstractIn this paper, we consider two inter-dependent deep networks, where one network taps into the other, to perform two challenging cognitive vision tasks - scene classification and object recognition jointly. Recently, convolutional neura networks have shown promising results in each of these tasks. However, as scene and objects are interrelated, the performance of both of these recognition tasks can be further improved by exploiting dependencies between scene and object deep networks. The advantages of considering the inter-dependency between these networks are the following: 1. improvement of accuracy in both scene and object classification, and 2. significant reduction of computational cost in object detection. In order to formulate our framework, we employ two convolutional neural networks (CNNs), scene-CNN and object-CNN. We utilize scene-CNN to generate object proposals which indicate the probable object locations in an image. Object proposals found in the process are semantically relevant to the object. More importantly, the number of object proposals is fewer in amount when compared to other existing methods which reduces the computational cost significantly. Thereafter, in scene classification, we train three hidden layers in order to combine the global (image as a whole) and local features (object information in an image). Features extracted from CNN architecture along with the features processed from object-CNN are combined to perform efficient classification. We perform rigorous experiments on five datasets to demonstrate that our proposed framework outperforms other state-of-the-art methods in classifying scenes as well as recognizing objects. Jawadul H. Bappy, Amit K. Roy-Chowdhury |
ICPR | 2 |
| 2016 | Video summarization in a multi-view camera networkabstractWhile most existing video summarization approaches aim to extract an informative summary of a single video, we propose a novel framework for summarizing multi-view videos by exploiting both intra- and inter-view content correlations in a joint embedding space. We learn the embedding by minimizing an objective function that has two terms: one due to intra-view correlations and another due to inter-view correlations across the multiple views. The solution can be obtained directly by solving one Eigen-value problem that is linear in the number of multi-view videos. We then employ a sparse representative selection approach over the learned embedding space to summarize the multi-view videos. Experimental results on several benchmark datasets demonstrate that our proposed approach clearly out-performs the state-of-the-art. Rameswar Panda, Abir Das, Amit K. Roy-Chowdhury |
ICPR | 3 |
| 2016 | Generating Diverse Image Datasets with Limited LabelingabstractImage datasets play a pivotal role in advancing multimedia and image analysis research. However, most of these datasets are created by extensive human effort and extremely expensive to scale up. There is high chance that we may have no instances for some required concepts in these data-sets or the available instances do not cover the diversity of real-world scenarios. In this regard, several approaches for learning from web images and refining them have been proposed, but these approaches either include significant redundant instances in the dataset or fail to guarantee a diverse enough set to train a robust classifier. In this work, we propose a semi-supervised sparse coding framework to collect a diverse set of images with minimal human effort, which can be used to both create a dataset from scratch or enrich an existing dataset with diverse examples. To evaluate our method, we constructed an image dataset with our framework, which is named as DivNet. Experiments on this dataset demonstrate that our method not only reduces manual effort, but also the created dataset has excellent accuracy, diversity and cross-dataset generalization ability. Niluthpol Chowdhury Mithun, Rameswar Panda, Amit K. Roy-Chowdhury |
ACM Multimedia | 3 |
| 2016 | Incremental learning of human activity models from videos
Mahmudul Hasan 0003, Amit K. Roy-Chowdhury |
Comput. Vis. Image Underst. | 2 |
| 2016 | Network Consistent Data AssociationabstractExisting data association techniques mostly focus on matching pairs of data-point sets and then repeating this process along space-time to achieve long term correspondences. However, in many problems such as person re-identification, a set of data-points may be observed at multiple spatio-temporal locations and/or by multiple agents in a network and simply combining the local pairwise association results between sets of data-points often leads to inconsistencies over the global space-time horizons. In this paper, we propose a Novel Network Consistent Data Association (NCDA) framework formulated as an optimization problem that not only maintains consistency in association results across the network, but also improves the pairwise data association accuracies. The proposed NCDA can be solved as a binary integer program leading to a globally optimal solution and is capable of handling the challenging data-association scenario where the number of data-points varies across different sets of instances in the network. We also present an online implementation of NCDA method that can dynamically associate new observations to already observed data-points in an iterative fashion, while maintaining network consistency. We have tested both the batch and the online NCDA in two application areas-person re-identification and spatio-temporal cell tracking and observed consistent and highly accurate data association results in all the cases. Anirban Chakraborty 0001, Abir Das, Amit K. Roy-Chowdhury |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | Distributed Multi-Target Tracking and Data Association in Vision NetworksabstractDistributed algorithms have recently gained immense popularity. With regards to computer vision applications, distributed multi-target tracking in a camera network is a fundamental problem. The goal is for all cameras to have accurate state estimates for all targets. Distributed estimation algorithms work by exchanging information between sensors that are communication neighbors. Vision-based distributed multi-target state estimation has at least two characteristics that distinguishes it from other applications. First, cameras are directional sensors and often neighboring sensors may not be sensing the same targets, i.e., they are naive with respect to that target. Second, in the presence of clutter and multiple targets, each camera must solve a data association problem. This paper presents an information-weighted, consensus-based, distributed multi-target tracking algorithm referred to as the Multi-target Information Consensus (MTIC) algorithm that is designed to address both the naivety and the data association problems. It converges to the centralized minimum mean square error estimate. The proposed MTIC algorithm and its extensions to non-linear camera models, termed as the Extended MTIC (EMTIC), are robust to false measurements and limited resources like power, bandwidth and the real-time operational requirements. Simulation and experimental analysis are provided to support the theoretical results. Ahmed Tashrif Kamal, Jawadul H. Bappy, Jay A. Farrell, Amit K. Roy-Chowdhury |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Introduction of New Associate EditorsabstractPresents a listing of the new Associate Editors for this issue of the publication. Nikolaos V. Boulgouris, David Bull 0001, Marco Cagnazzo, Andrea Cavallaro, Gene Cheung, Amit K. Roy-Chowdhury, Pedro Comesaña Alfaro, Sarp Ertürk, Markus Flierl, Gian Luca Foresti, Gang Hua 0001, Zhu Li 0001, Weisi Lin, Siwei Ma 0001, Pramod Kumar Meher, Debargha Mukherjee, Aleksandra Pizurica, Andrea Prati 0001, Paolo Remagnino, Arun Ross, Shin'ichi Satoh 0001, Andreas E. Savakis, Heiko Schwarz, Ling Shao 0001, Shervin Shirmohammadi, Giuseppe Valenzise, Meng Wang 0001, Zhou Wang 0001, Yonggang Wen 0001, Dong Xu 0001, Junsong Yuan 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2016 | Context-Aware Surveillance Video SummarizationabstractWe present a method that is able to find the most informative video portions, leading to a summarization of video sequences. In contrast to the existing works, our method is able to capture the important video portions through information about individual local motion regions, as well as the interactions between these motion regions. Specifically, our proposed Context-Aware Video Summarization (CAVS) framework adopts the methodology of sparse coding with generalized sparse group lasso to learn a dictionary of video features and a dictionary of spatio-temporal feature correlation graphs. Sparsity ensures that the most informative features and relationships are retained. The feature correlations, represented by a dictionary of graphs, indicate how motion regions correlate to each other globally. When a new video segment is processed by CAVS, both dictionaries are updated in an online fashion. Specifically, CAVS scans through every video segment to determine if the new features along with the feature correlations, can be sparsely represented by the learned dictionaries. If not, the dictionaries are updated, and the corresponding video segments are incorporated into the summarized video. The results on four public datasets, mostly composed of surveillance videos and a small amount of other online videos, show the effectiveness of our proposed method. Shu Zhang 0007, Yingying Zhu 0002, Amit K. Roy-Chowdhury |
IEEE Trans. Image Process. | 3 |
| 2015 | An end-to-end system for content-based video retrieval using behavior, actions, and appearance with interactive query refinementabstractWe describe a system for content-based retrieval from large surveillance video archives, using behavior, action and appearance of objects. Objects are detected, tracked, and classified into broad categories. Their behavior and appearance are characterized by action detectors and descriptors, which are indexed in an archive. Queries can be posed as video exemplars, and the results can be refined through relevance feedback. The contributions of our system include the fusion of behavior and action detectors with appearance for matching; the improvement of query results through interactive query refinement (IQR), which learns a discriminative classifier online based on user feedback; and reasonable performance on low resolution, poor quality video. The system operates on video from ground cameras and aerial platforms, both RGB and IR. Performance is evaluated on publicly-available surveillance datasets, showing that subtle actions can be detected under difficult conditions, with reasonable improvement from IQR. Anthony Hoogs, A. G. Amitha Perera, Roderic Collins, Arslan Basharat, Keith Fieldhouse, Chuck Atkins, Linus Sherrill, Benjamin Boeckel, Russell Blue, Matthew Woehlke, C. Greco, Zhaohui Sun, Eran Swears, Naresh P. Cuntoor, J. Luck, B. Drew, D. Hanson, D. Rowley, J. Kopaz, T. Rude, D. Keefe, Amit Srivastava, Saurabh Khanwalkar, Chia-Chih Chen, Jake K. Aggarwal, Larry Davis 0001, Yaser Yacoob, Dong Liu 0001, Shih-Fu Chang, Bi Song, Amit K. Roy-Chowdhury, Kenneth Sullivan, Jelena Tesic, Shivkumar Chandrasekaran, B. S. Manjunath, K. Reddy, Mubarak Shah, K. Chang, Tsuhan Chen, Mita Desai |
AVSS | 33 |
| 2015 | Context Aware Active Learning of Activity Recognition ModelsabstractActivity recognition in video has recently benefited from the use of the context e.g., inter-relationships among the activities and objects. However, these approaches require data to be labeled and entirely available at the outset. In contrast, we formulate a continuous learning framework for context aware activity recognition from unlabeled video data which has two distinct advantages over most existing methods. First, we propose a novel active learning technique which not only exploits the informativeness of the individual activity instances but also utilizes their contextual information during the query selection process, this leads to significant reduction in expensive manual annotation effort. Second, the learned models can be adapted online as more data is available. We formulate a conditional random field (CRF) model that encodes the context and devise an information theoretic approach that utilizes entropy and mutual information of the nodes to compute the set of most informative query instances, which need to be labeled by a human. These labels are combined with graphical inference techniques for incrementally updating the model as new videos come in. Experiments on four challenging datasets demonstrate that our framework achieves superior performance with significantly less amount of manual labeling. Mahmudul Hasan 0003, Amit K. Roy-Chowdhury |
ICCV | 2 |
| 2015 | Active image pair selection for continuous person re-identificationabstractMost traditional multi-camera person re-identification systems rely on learning a static model on tediously labeled training data. Such a framework may not be suitable for situations when new data arrives continuously or all the data is not available for labeling beforehand. Inspired by the `value of information' active learning framework, we propose a continuous learning person re-identification system with a human in the loop. In brief, we term this `continuous person re-identification'. The human in the loop not only provides labels to the incoming images but also improves the learned model by providing most appropriate attribute based explanations. These attribute based explanations are used to learn attribute predictors along the way. The overall effect of such a stratgey is that starting with a few annotated images, the system begins to improve via a symbiotic relationship between the man and the machine. The machine assists the human to speed the annotation and the human assists the machine to update itself with more annotation so that more and more distinct persons are re-identified as more and more images come in. Using a benchmark dataset, we validate our approach and compare with state-of-the-art methods. Abir Das, Rameswar Panda, Amit K. Roy-Chowdhury |
ICIP | 3 |
| 2015 | Video summarization through change detection in a non-overlapping camera networkabstractWe present a method that is able to find the most informative video portions in a non-overlapping camera network, leading to a summarization of the multiple video sequences. This is posed as a problem of detecting changes in the interactions between the targets in the network of cameras. Examples include formation and dispersal of groups within the view of a single camera, as well as identifying changes between cameras. The latter includes prediction of events that may have occurred in the gaps between the cameras. The solution strategy is built upon a social group identification method and a track association strategy, which together are used to indicate conflicts in the interactions between the targets, leading to identification of the most informative video portions in a non-overlapping camera network. We apply our algorithm on a public dataset with multiple non-overlapping cameras on a university campus. We show examples of informative video segments, as well as perform a statistical analysis of the results. Shu Zhang 0007, Amit K. Roy-Chowdhury |
ICIP | 2 |
| 2015 | A Camera Network Tracking (CamNeT) Dataset and Performance BaselineabstractIn this paper, we propose a novel Non-Overlapping Camera Network Tracking Dataset (CamNeT) for evaluating multi-target tracking algorithms. The dataset is composed of five to eight cameras covering both indoor and outdoor scenes at a university. This dataset consists of six scenarios. Within each scenario are challenges relevant to lighting changes, complex topographies, crowded scenes, and changing grouping dynamics. Persons with predefined trajectories are combined with persons with random trajectories. Ground truth data for predefined trajectories is provided for each camera. Also, a baseline multi-target tracking system is presented. The tracking results using the baseline system are provided, which can be compared with future works. The work provides a comprehensive multicamera dataset for performance evaluation in this challenging application domain, as well as an initial set of results. Shu Zhang 0007, Elliot Staudt, Tim Faltemier, Amit K. Roy-Chowdhury |
WACV | 4 |
| 2015 | Tracking multiple interacting targets in a camera network
Shu Zhang 0007, Yingying Zhu 0002, Amit K. Roy-Chowdhury |
Comput. Vis. Image Underst. | 3 |
| 2015 | Context aware spatio-temporal cell tracking in densely packed multilayer tissues
Anirban Chakraborty 0001, Amit K. Roy-Chowdhury |
Medical Image Anal. | 2 |
| 2015 | Re-Identification in the Function Space of Feature WarpsabstractPerson re-identification in a non-overlapping multicamera scenario is an open challenge in computer vision because of the large changes in appearances caused by variations in viewing angle, lighting, background clutter, and occlusion over multiple cameras. As a result of these variations, features describing the same person get transformed between cameras. To model the transformation of features, the feature space is nonlinearly warped to get the "warp functions". The warp functions between two instances of the same target form the set of feasible warp functions while those between instances of different targets form the set of infeasible warp functions. In this work, we build upon the observation that feature transformations between cameras lie in a nonlinear function space of all possible feature transformations. The space consisting of all the feasible and infeasible warp functions is the warp function space (WFS). We propose to learn a discriminating surface separating these two sets of warp functions in the WFS and to re-identify persons by classifying a test warp function as feasible or infeasible. Towards this objective, a Random Forest (RF) classifier is employed which effectively chooses the warp function components according to their importance in separating the feasible and the infeasible warp functions in the WFS. Extensive experiments on five datasets are carried out to show the superior performance of the proposed approach over state-of-the-art person re-identification methods. We show that our approach outperforms all other methods when large illumination variations are considered. At the same time it has been shown that our method reaches the best average performance over multiple combinations of the datasets, thus, showing that our method is not designed only to address a specific challenge posed by a particular dataset. Niki Martinel, Abir Das, Christian Micheloni, Amit K. Roy-Chowdhury |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Context-Aware Activity Modeling Using Hierarchical Conditional Random FieldsabstractIn this paper, rather than modeling activities in videos individually, we jointly model and recognize related activities in a scene using both motion and context features. This is motivated from the observations that activities related in space and time rarely occur independently and can serve as the context for each other. We propose a two-layer conditional random field model, that represents the action segments and activities in a hierarchical manner. The model allows the integration of both motion and various context features at different levels and automatically learns the statistics that capture the patterns of the features. With weakly labeled training data, the learning problem is formulated as a max-margin problem and is solved by an iterative algorithm. Rather than generating activity labels for individual activities, our model simultaneously predicts an optimum structural label for the related activities in the scene. We show promising results on the UCLA Office Dataset and VIRAT Ground Dataset that demonstrate the benefit of hierarchical modeling of related activities using both motion and context features. Yingying Zhu 0002, Nandita M. Nayak, Amit K. Roy-Chowdhury |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Evaluation and Acceleration of High-Throughput Fixed-Point Object Detection on FPGAsabstractReliance on object or people detection is rapidly growing beyond surveillance to industrial and social applications. The histogram of oriented gradients (HOG), one of the most popular object detection algorithms, achieves high detection accuracy but delivers just under 1 frame/s on a high-end CPU. Field-programmable gate array (FPGA) accelerations of this algorithm are limited by the intensive floating-point computations. All current fixed-point HOG implementations use large bit width to maintain detection accuracy, or perform poorly at reduced data precision. In this paper, we introduce the full-image evaluation methodology to explore the FPGA implementation of HOG using reduced bit width. This approach lessens the required area resources on the FPGA, and increases the clock frequency and hence the throughput per device through increased parallelism. We evaluate the detection accuracy of the fixed-point HOG by applying state-of-the-art computer vision pedestrian detection evaluation metrics and show it performs as well as the original floating-point code from OpenCV. We then show our single FPGA implementation achieves a 68.7 × higher throughput than a highend CPU, 5.1 × higher than a high-end graphics processing unit (GPU), and 7.8 × higher than the same implementation using floating-point on the same FPGA. A power consumption comparison for different platforms shows our fixed-point FPGA implementation uses 130 × less power than CPU, and 31 × less energy than GPU to process one image. Xiaoyin Ma, Walid A. Najjar, Amit K. Roy-Chowdhury |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | A Continuous Learning Framework for Activity Recognition Using Deep Hybrid Feature ModelsabstractMost of the research on human activity recognition has focused on learning a static model, considering that all the training instances are labeled and present in advance, while in streaming videos new instances continuously arrive and are not labeled. Moreover, these methods generally use application- specific hand-engineered and static feature models, which are not suitable for continuous learning. Some recent approaches on activity recognition use deep-learning-based hierarchical feature models, but the large size of these networks constrain them from being used in continuous learning scenarios. In this work, we propose a continuous activity learning framework for streaming videos by intricately tying together deep hybrid feature models and active learning. This allows us to automatically select the most suitable features and take the advantage of incoming unlabeled instances to improve the existing model incrementally. Given the segmented activities from streaming videos, we learn features in an unsupervised manner using deep hybrid networks, which have the ability to take the advantage of both the local hand-engineered features and the deep model in an efficient way. Additionally, we use active learning to train the activity classifier using a reduced amount of manually labeled instances. Retraining the models with a huge amount of accumulated examples is computationally expensive and not suitable for continuous learning. Hence, we propose a method to select the best subset of these examples to update the models incrementally. We conduct rigorous experiments on four challenging human activity datasets to demonstrate the effectiveness of our framework. Mahmudul Hasan 0003, Amit K. Roy-Chowdhury |
IEEE Trans. Multim. | 2 |
| 2014 | Context-Aware Activity Forecasting
Anirban Chakraborty 0001, Amit K. Roy-Chowdhury |
ACCV (5) | 2 |
| 2014 | Managing Redundant Content in Bandwidth Constrained Wireless NetworksabstractImages/videos are often uploaded in situations like disasters. This can tax the network in terms of increased load and thereby upload latency, and this can be critical for response activities. In such scenarios, prior work has shown that there is significant redundancy in the content (e.g., similar photos taken by users) transferred. By intelligently suppressing/deferring transfers of redundant content, the load can be significantly reduced, thereby facilitating the timely delivery of unique, possibly critical information. A key challenge here however, is detecting 'what content is similar,' given that the content is generated by uncoordinated user devices. Towards addressing this challenge, we propose a framework, wherein a service to which the content is to be uploaded first solicits metadata (e.g, image features) from any device uploading content. By intelligently comparing this metadata with that associated with previously uploaded content, the service effectively identifies (and thus enables the suppression of) redundant content. Our evaluations on a testbed of 20 Android smartphones and via ns3 simulations show that we can identify similar content with a 70% true positive rate and a 1% false positive rate. The resulting reduction in redundant content transfers translates to a latency reduction of 44 % for unique content. Tuan Dao, Amit K. Roy-Chowdhury, Harsha V. Madhyastha, Srikanth V. Krishnamurthy, Thomas La Porta |
CoNEXT | 2 |
| 2014 | Incremental Activity Modeling and Recognition in Streaming VideosabstractMost of the state-of-the-art approaches to human activity recognition in video need an intensive training stage and assume that all of the training examples are labeled and available beforehand. But these assumptions are unrealistic for many applications where we have to deal with streaming videos. In these videos, as new activities are seen, they can be leveraged upon to improve the current activity recognition models. In this work, we develop an incremental activity learning framework that is able to continuously update the activity models and learn new ones as more videos are seen. Our proposed approach leverages upon state-of-the-art machine learning tools, most notably active learning systems. It does not require tedious manual labeling of every incoming example of each activity class. We perform rigorous experiments on challenging human activity datasets, which demonstrate that the incremental activity modeling framework can achieve performance very close to the cases when all examples are available a priori. Mahmudul Hasan 0003, Amit K. Roy-Chowdhury |
CVPR | 2 |
| 2014 | Consistent Re-identification in a Camera Network
Abir Das, Anirban Chakraborty 0001, Amit K. Roy-Chowdhury |
ECCV (2) | 3 |
| 2014 | Continuous Learning of Human Activity Models Using Deep Nets
Mahmudul Hasan 0003, Amit K. Roy-Chowdhury |
ECCV (3) | 2 |
| 2014 | High-Throughput Fixed-Point Object Detection on FPGAsabstractComputer vision applications make extensive use of floating-point number representation, both single and double precision. The major advantage of floating-point representation is the very large range of values that can be represented with a limited number of bits. Most CPU, and all GPU designs have been extensively optimized for short latency and high-throughput processing of floating-point operations. On an FPGA, the bit-width of operands is a major determinant of its resource utilization, the achievable clock frequency and hence its throughput. By using a fixed-point representation with fewer bits, an application developer could implement more processing units and a higher-clock frequency and a dramatically larger throughput. However, smaller bit-widths may lead to inaccurate or incorrect results. Object and human detection are fundamental problems in computer vision and a very active research area. In these applications a high throughput and an economy of resources are highly desirable features allowing the applications to be embedded in mobile or fielddeployable equipment. The Histogram of Oriented Gradients (HOG) algorithm [1], developed for human detection and expanded to object detection, is one of the most successful and popular algorithm in its class. In this algorithm, object descriptors are extracted from detection window with grids of overlapping blocks. Each block is divided into cells in which histograms of intensity gradients are collected as HOG features. Vectors of histograms are normalized and passed to a Support Vector Machine (SVM) classifier to recognize a person or an object. Xiaoyin Ma, Walid A. Najjar, Amit K. Roy-Chowdhury |
FCCM | 3 |
| 2014 | Face recognition based on SIGMA sets of image featuresabstractAutomatic face recognition is prevalent in a wide range of systems these days and it is critical to explore new techniques in order to enhance the state of the art. In this paper, we analyze the Region Covariance Matrix (RCM) and its enhancement based on Sigma sets as a feature extraction procedure for face images. The RCM features encode the covariance of various low level features, e.g., pixel intensities and gradients. Sigma sets, on the other hand, reduce the computational complexity of comparing two RCMs. Based on our experiments on the Labeled Faces in the Wild (LFW) dataset, we show that the proposed technique outperforms the popular Local Binary Patterns (LBP) technique and is on par with other better performing techniques that use complex classifiers. Ramya Srinivasan 0002, Abhishek Nagar, Anshuman Tewari, Donato Mitrani, Amit K. Roy-Chowdhury |
ICASSP | 5 |
| 2014 | A conditional random field model for tracking in densely packed cell structuresabstractAutomated tracking of plant and animal cells in time lapse live-imaging datasets of developing multicellular tissues is required for quantitative, high throughput analysis of cell division, migration and cell growth. In this paper, we present a novel cell tracking method that exploits the tight spatial topology of neighboring cells in a multicellular field as contextual information and combines it with physical features of individual cells for generating reliable cell lineages. The 2D image slices of multicellular tissues are modeled as CRFs and spatio-temporal cell to cell correspondences are obtained by performing inference on this CRF using loopy belief propagation. We present results on a (3D+t) confocal image stack of Arabidopsis shoot meristem and show that the method can handle many visual analysis challenges associated with such cell tracking problems, viz. poor feature quality of individual cells, low SNR in parts of images, variable number of cells across slices and cell division detection. Anirban Chakraborty 0001, Amit K. Roy-Chowdhury |
ICIP | 2 |
| 2014 | Learning a sparse dictionary of video structure for activity modelingabstractWe present an approach which incorporates spatiotemporal features as well as the relationships between them, into a sparse dictionary learning framework for activity recognition. We propose that the dictionary learning framework can be adapted to learning complex relationships between features in an unsupervised manner. From a set of training videos, a dictionary is learned for individual features, as well as the relationships between them using a stacked predictive sparse decomposition framework. This combined dictionary provides a representation of the structure of the video and is spatio-temporally pooled in a local manner to obtain descriptors. The descriptors are then combined using a multiple kernel learning framework to design classifiers. Experiments have been conducted on two popular activity recognition datasets to demonstrate the superior performance of our approach on single person as well as multi-person activities. Nandita M. Nayak, Amit K. Roy-Chowdhury |
ICIP | 2 |
| 2013 | Information Consensus for Distributed Multi-target TrackingabstractDue to their high fault-tolerance, ease of installation and scalability to large networks, distributed algorithms have recently gained immense popularity in the sensor networks community, especially in computer vision. Multi-target tracking in a camera network is one of the fundamental problems in this domain. Distributed estimation algorithms work by exchanging information between sensors that are communication neighbors. Since most cameras are directional sensors, it is often the case that neighboring sensors may not be sensing the same target. Such sensors that do not have information about a target are termed as ``naive'' with respect to that target. In this paper, we propose consensus-based distributed multi-target tracking algorithms in a camera network that are designed to address this issue of naivety. The estimation errors in tracking and data association, as well as the effect of naivety, are jointly addressed leading to the development of an information-weighted consensus algorithm, which we term as the Multi-target Information Consensus (MTIC) algorithm. The incorporation of the probabilistic data association mechanism makes the MTIC algorithm very robust to false measurements/clutter. Experimental analysis is provided to support the theoretical results. Ahmed Tashrif Kamal, Jay A. Farrell, Amit K. Roy-Chowdhury |
CVPR | 3 |
| 2013 | Context-Aware Modeling and Recognition of Activities in VideoabstractIn this paper, rather than modeling activities in videos individually, we propose a hierarchical framework that jointly models and recognizes related activities using motion and various context features. This is motivated from the observations that the activities related in space and time rarely occur independently and can serve as the context for each other. Given a video, action segments are automatically detected using motion segmentation based on a nonlinear dynamical model. We aim to merge these segments into activities of interest and generate optimum labels for the activities. Towards this goal, we utilize a structural model in a max-margin framework that jointly models the underlying activities which are related in space and time. The model explicitly learns the duration, motion and context patterns for each activity class, as well as the spatio-temporal relationships for groups of them. The learned model is then used to optimally label the activities in the testing videos using a greedy search method. We show promising results on the VIRAT Ground Dataset demonstrating the benefit of joint modeling and recognizing activities in a wide-area scene. Yingying Zhu 0002, Nandita M. Nayak, Amit K. Roy-Chowdhury |
CVPR | 3 |
| 2013 | A generalized data-driven Hamiltonian Monte Carlo for hierarchical activity searchabstractMotion and image analysis are both important for robust solutions to video search of activities; the physics-based, data-driven Hamiltonian Monte Carlo (HMC), a Markov chain Monte Carlo variant that is efficient in searching large dimensional spaces, simultaneously examines the combined motion and image space. In this paper, we generalize the data-driven HMC to no longer depend upon ad hoc Guide Hamiltonians and to no longer require physics-based features from tracks as pre-requisites. Our generalization thus allows it to be used with or without a tracker, overcoming a significant limitation of the physics-based approach, as well as being extensible to utilizing any pre-existing image- or motion-based method. We demonstrate the generalizability of our framework by considering situations when tracking is available and when it is not available. When tracking is available, we utilize Histogram of Oriented Gradients, shapes of trajectories, and Hamiltonian Energy Signatures; when tracking is not available, we use Space-time Interest Points and GIST features. In addition, we show our generalized framework performs better than the physics-based, data-driven HMC, as well as state-of-the-art, by demonstrating the efficacy of our system on real-life video sequences using the well-known Weizmann and YouTube Action datasets. Ricky J. Sethi, Hyunjoon Jo, Amit K. Roy-Chowdhury |
ICIP | 3 |
| 2013 | Recognizing the royals: leveraging computerized face recognition for identifying subjects in ancient artworksabstractWe present a work that explores the feasibility of automated face recognition technologies for analyzing identities in works of portraiture, and in the process provide additional evidence to settle some long-standing questions in art history. Works of portrait art bear the mark of visual interpretation of the artist. Moreover, the number of samples available to model these effects is often limited. From a set of portraiture of the Renaissance and Baroque periods, where the identities of subjects are known, we derive appropriate features that are based on domain knowledge of artistic renderings, and learn and validate statistical models for the distribution of the match and non-match scores, which we refer to as portrait feature space (PFS). Thereafter, we use this PFS on a number of cases that have been "open questions" to art historians. They are usually in the form of validating two portraits as belonging to the same person. Using statistical hypothesis tests on the PFS, we provide quantitative measures of similarity for each of these questions. It is, to the best of our knowledge, the first study that applies automated face recognition technologies to the analysis of portraits of multiple subjects in various forms - paintings, death masks, sculptures. Ramya Srinivasan 0002, Amit K. Roy-Chowdhury, Conrad Rudolph, Jeanette Kohl |
ACM Multimedia | 2 |
| 2013 | Modeling multi-object interactions using "string of feature graphs"
Yingying Zhu 0002, Nandita M. Nayak, Utkarsh Gaur, Bi Song, Amit K. Roy-Chowdhury |
Comput. Vis. Image Underst. | 5 |
| 2013 | Vector field analysis for multi-object behavior modeling
Nandita M. Nayak, Yingying Zhu 0002, Amit K. Roy-Chowdhury |
Image Vis. Comput. | 3 |
| 2013 | Quantitative Analysis of Live-Cell Growth at the Shoot Apex of Arabidopsis thaliana: Algorithms for Feature Measurement and Temporal AlignmentabstractStudy of the molecular control of organ growth requires establishment of the causal relationship between gene expression and cell behaviors. We seek to understand this relationship at the shoot apical meristem (SAM) of model plant Arabidopsis thaliana. This requires the spatial mapping and temporal alignment of different functional domains into a single template. Live-cell imaging techniques allow us to observe real-time organ primordia growth and gene expression dynamics at cellular resolution. In this paper, we propose a framework for the measurement of growth features at the 3D reconstructed surface of organ primordia, as well as algorithms for robust time alignment of primordia. We computed areas and deformation values from reconstructed 3D surfaces of individual primordia from live-cell imaging data. Based on these growth measurements, we applied a multiple feature landscape matching (LAM-M) algorithm to ensure a reliable temporal alignment of multiple primordia. Although the original landscape matching (LAM) algorithm motivated our alignment approach, it sometimes fails to properly align growth curves in the presence of high noise/distortion. To overcome this shortcoming, we modified the cost function to consider the landscape of the corresponding growth features. We also present an alternate parameter-free growth alignment algorithm which performs as well as LAM-M for high-quality data, but is more robust to the presence of outliers or noise. Results on primordia and guppy evolutionary growth data show that the proposed alignment framework performs at least as well as the LAM algorithm in the general case, and significantly better in the case of increased noise. Oben M. Tataw, G. Venugopala Reddy, Eamonn J. Keogh, Amit K. Roy-Chowdhury |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2013 | Exploiting Spatio-Temporal Scene Structure for Wide-Area Activity Analysis in Unconstrained EnvironmentsabstractSurveillance videos in unconstrained environments typically consist of long duration sequences of activities which occur at different spatio-temporal locations and can involve multiple people acting simultaneously. Often, the activities have contextual relationships with one another. Although context has been studied in the past for the purpose of activity recognition, the use of context in recognition of activities in such challenging environments is relatively unexplored. In this paper, we propose a novel method for capturing the spatio-temporal context between activities in a Markov random field. The structure of the MRF is improvised upon during test time and not predefined, unlike many approaches that model the contextual relationships between activities. Given a collection of videos and a set of weak classifiers for individual activities, the spatio-temporal relationships between activities are represented as probabilistic edge weights in the MRF. This model provides a generic representation for an activity sequence that can extend to any number of objects and interactions in a video. We show that the recognition of activities in a video can be posed as an inference problem on the graph. We conduct experiments on the publicly available UCLA office dataset and the VIRAT dataset, to demonstrate the improvement in recognition accuracy using our proposed model as opposed to recognition using state-of-the-art features on individual activity regions. Nandita M. Nayak, Yingying Zhu 0002, Amit K. Roy-Chowdhury |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2012 | Features with Feelings - Incorporating User Preferences in Video Categorization
Ramya Srinivasan 0002, Amit K. Roy-Chowdhury |
ACCV (3) | 2 |
| 2012 | Consensus-based distributed estimation in camera networksabstractDistributed algorithms in the sensors networks community usually require each sensor to have its own measurement. In practice, this constraint can not always be met. For example, in a camera network, all cameras might not observe a particular target as cameras are directional sensors and have a limited field-of-view (FOV). Moreover, different sensors might provide different quality measures related to different elements of the measurement vector depending on various factors as directionality, occlusion etc. This requires the designing of a new type of distributed algorithm that considers the quality and/or absence of measurements. In this paper, we present a distributed algorithm to compute the maximum likelihood estimate of the state of a target viewed by the network of cameras, taking into account the above-mentioned factors. We provide step-by-step derivation along with theoretical guarantee of optimality and convergence of the method. Experimental results are provided to show the performance of the proposed algorithm. Ahmed Tashrif Kamal, Jay A. Farrell, Amit K. Roy-Chowdhury |
ICIP | 3 |
| 2012 | Video classification based on social attitudesabstractOrganizing large video databases is a pressing need and a challenging problem. Social attitudes in the form of users' beliefs and evaluations can benefit classification. For instance, news videos do not gather as much user attention as music videos while sports videos trigger interest mainly during the time of event. In this paper, we provide an extensive analysis of the role of usage statistics in aiding classification. Towards this, we propose a novel framework motivated by evolutionary biology to characterize growth, persistence and decline of contents in online environments. We then incorporate this information in a nearest neighbor classifier to establish categories. The effectiveness of the approach is demonstrated by comparing against results obtained using principal component analysis followed by nearest neighbor based classification. Ramya Srinivasan 0002, Amit K. Roy-Chowdhury |
ICIP | 2 |
| 2012 | Collaborative Sensing in a Distributed PTZ Camera NetworkabstractThe performance of dynamic scene algorithms often suffers because of the inability to effectively acquire features on the targets, particularly when they are distributed over a wide field of view. In this paper, we propose an integrated analysis and control framework for a pan, tilt, zoom (PTZ) camera network in order to maximize various scene understanding performance criteria (e.g., tracking accuracy, best shot, and image resolution) through dynamic camera-to-target assignment and efficient feature acquisition. Moreover, we consider the situation where processing is distributed across the network since it is often unrealistic to have all the image data at a central location. In such situations, the cameras, although autonomous, must collaborate among themselves because each camera's PTZ parameter entails constraints on the others. Motivated by recent work in cooperative control of sensor networks, we propose a distributed optimization strategy, which can be modeled as a game involving the cameras and targets. The cameras gain by reducing the error covariance of the tracked targets or through higher resolution feature acquisition, which, however, comes at the risk of losing the dynamic target. Through the optimization of this reward-versus-risk tradeoff, we are able to control the PTZ parameters of the cameras and assign them to targets dynamically. The tracks, upon which the control algorithm is dependent, are obtained through a consensus estimation algorithm whereby cameras can arrive at a consensus on the state of each target through a negotiation strategy. We analyze the performance of this collaborative sensing strategy in active camera networks in a simulation environment, as well as a real-life camera network. Chong Ding, Bi Song, Akshay A. Morye, Jay A. Farrell, Amit K. Roy-Chowdhury |
IEEE Trans. Image Process. | 5 |
| 2011 | AVSS 2011 demo session: A large-scale benchmark dataset for event recognition in surveillance videoabstractSummary form only given. We present a concept for automatic construction site monitoring by taking into account 4D information (3D over time), that is acquired from highly-overlapping digital aerial images. On the one hand today's maturity of flying micro aerial vehicles (MAVs) enables a low-cost and an efficient image acquisition of high-quality data that maps construction sites entirely from many varying viewpoints. On the other hand, due to low-noise sensors and high redundancy in the image data, recent developments in 3D reconstruction workflows have benefited the automatic computation of accurate and dense 3D scene information. Having both an inexpensive high-quality image acquisition and an efficient 3D analysis workflow enables monitoring, documentation and visualization of observed sites over time with short intervals. Relating acquired 4D site observations, composed of color, texture, geometry over time, largely supports automated methods toward full scene understanding, the acquisition of both the change and the construction site's progress. Sangmin Oh, Anthony Hoogs, A. G. Amitha Perera, Naresh P. Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, Jake K. Aggarwal, Hyungtae Lee, Larry Davis 0001, Eran Swears, Xiaoyang Wang 0001, Kishore K. Reddy, Mubarak Shah, Carl Vondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba 0001, Bi Song, Anesco Fong, Amit K. Roy-Chowdhury, Mita Desai |
AVSS | 23 |
| 2011 | Cell Resolution 3D Reconstruction of Developing Multilayer Tissues from Sparsely Sampled Volumetric Microscopy ImagesabstractUnderstanding of the growth dynamics in developmental biology is often pursued through the analysis of cell sizes and shapes obtained from CLSM based imaging at cell resolution of multi-layer tissues. This necessitates the development of robust 3D reconstruction methods using such images. However, all of the current methods of 3D reconstruction using CLSM imaging require large number of cell slices. But in the case of live cell imaging, i.e., imaging a growing tissue, such high depth resolution is not feasible in order to avoid photodynamic damage to the growing cells from prolonged exposure to laser radiation. In this work, we have addressed the problem of 3D reconstruction at cell resolution of a developing multi-layer tissue in the plant meristem when the amount of data is as limited as two to four slices per cell. This introduces significant image analysis challenges in terms of sparsity of the data, low signal-to-noise ratio, and a wide range of shapes and sizes. Motivated by the physical structure of the cells, we propose to reconstruct a cell cluster as a packing of truncated ellipsoids representing the individual cells. We test the proposed computational method on time-lapse CLSM images of Shoot Apical Meristem (SAM) cells of model plant Arabidopsis Thaliana. We show that the 3D reconstruction can lead to 3D shape models of complete cell clusters, which is an essential first step towards obtaining growth statistics for individual cells. Anirban Chakraborty 0001, Ram Kishor Yadav, G. Venugopala Reddy, Amit K. Roy-Chowdhury |
BIBM | 4 |
| 2011 | 3D Neuron Tip Detection in Volumetric Microscopy ImagesabstractThis paper addresses the problem of 3D neuron tips detection in volumetric microscopy image stacks. We focus particularly on neuron tracing applications, where the detected 3D tips could be used as the seeding points. Most of the existing neuron tracing methods require a good choice of seeding points. In this paper, we propose an automated neuron tips detection method for volumetric microscopy image stacks. Our method is based on first detecting 2D tips using curvature information and a ray-shooting intensity distribution model, and then extending it to the 3D stack by rejecting false positives. We tested this method based on the V3D platform, which can reconstruct a neuron based on automated searching of the optimal 'paths' connecting those detected 3D tips. The experiments demonstrate the effectiveness of the proposed method in building a fully automatic neuron tracing system. Min Liu 0008, Hanchuan Peng, Amit K. Roy-Chowdhury, Eugene W. Myers |
BIBM | 3 |
| 2011 | A large-scale benchmark dataset for event recognition in surveillance videoabstractWe introduce a new large-scale video dataset designed to assess the performance of diverse visual event recognition algorithms with a focus on continuous visual event recognition (CVER) in outdoor areas with wide coverage. Previous datasets for action recognition are unrealistic for real-world surveillance because they consist of short clips showing one action by one individual [15, 8]. Datasets have been developed for movies [11] and sports [12], but, these actions and scene conditions do not apply effectively to surveillance videos. Our dataset consists of many outdoor scenes with actions occurring naturally by non-actors in continuously captured videos of the real world. The dataset includes large numbers of instances for 23 event types distributed throughout 29 hours of video. This data is accompanied by detailed annotations which include both moving object tracks and event examples, which will provide solid basis for large-scale evaluation. Additionally, we propose different types of evaluation modes for visual recognition tasks and evaluation metrics along with our preliminary experimental results. We believe that this dataset will stimulate diverse aspects of computer vision research and help us to advance the CVER tasks in the years ahead. Sangmin Oh, Anthony Hoogs, A. G. Amitha Perera, Naresh P. Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, Jake K. Aggarwal, Hyungtae Lee, Larry Davis 0001, Eran Swears, Xiaoyang Wang 0001, Kishore K. Reddy, Mubarak Shah, Carl Vondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba 0001, Bi Song, Anesco Fong, Amit K. Roy-Chowdhury, Mita Desai |
CVPR | 23 |
| 2011 | A "string of feature graphs" model for recognition of complex activities in natural videosabstractVideos usually consist of activities involving interactions between multiple actors, sometimes referred to as complex activities. Recognition of such activities requires modeling the spatio-temporal relationships between the actors and their individual variabilities. In this paper, we consider the problem of recognition of complex activities in a video given a query example. We propose a new feature model based on a string representation of the video which respects the spatio-temporal ordering. This ordered arrangement of local collections of features (e.g., cuboids, STIP), which are the characters in the string, are initially matched using graph-based spectral techniques. Final recognition is obtained by matching the string representations of the query and the test videos in a dynamic programming framework which allows for variability in sampling rates and speed of activity execution. The method does not require tracking or recognition of body parts, is able to identify the region of interest in a cluttered scene, and gives reasonable performance with even a single query example. We test our approach in an example-based video retrieval framework with two publicly available complex activity datasets and provide comparisons against other methods that have studied this problem. Utkarsh Gaur, Yingying Zhu 0002, Bi Song, Amit K. Roy-Chowdhury |
ICCV | 4 |
| 2011 | Belief consensus for distributed action recognitionabstractIn this work, we consider a camera network where processing is distributed across the cameras. Our goal is to recognize actions of multiple targets consistently observed over the entire network. To obtain consistent and better results we need to properly fuse the action scores from multiple cameras. There have been multiple works on distributed tracking and distributed data association for multiple targets in a camera network. We can use the data association results and tracking confidence scores to improve the action recognition results. We propose a consensus based framework for solving this problem in an integrated manner and with a completely distributed camera network architecture. We propose a novel method for weighting the action scores based on tracking confidences and show how the cameras can reach a consensus about the action of a target using belief consensus. We show real life experiments and performance metrics with multiple cameras and targets. Ahmed Tashrif Kamal, Bi Song, Amit K. Roy-Chowdhury |
ICIP | 3 |
| 2011 | Efficient cell segmentation and tracking of developing plant meristemabstractAnalysis of Confocal Laser Scanning Microscopy (CLSM) images is gaining popularity in developmental biology for understanding growth dynamics. The automated analysis of such images is highly desirable for efficiency and accuracy. The first step in this process is segmentation and tracking leading to computation of cell lineages. In this paper, we present efficient, accurate, and robust segmentation and tracking algorithms for cells and detection of cell divisions in a 4D spatio-temporal image stack of a growing plant meristem. We show how to optimally choose the parameters in the watershed algorithm for high quality segmentation results. This yields high quality tracking results using cell correspondence evaluation functions. We show segmentation and tracking results on Confocal laser scanning microscopy data captured for 72 hours at every 3 hour intervals. Compared to recent results in this area, the proposed algorithms provide significantly longer cell lineages and more comprehensive identification of cell divisions. Katya Mkrtchyan, Damanpreet Singh, Min Liu 0008, G. Venugopala Reddy, Amit K. Roy-Chowdhury, Meenakshisundaram Gopi |
ICIP | 5 |
| 2011 | Vector field analysis for motion pattern identification in videoabstractIdentification of motion patterns in video is an important problem because it is the first step towards analysis of complex multi-person behaviors to obtain long-term interaction models. In this paper, we will present a flow based technique to identify spatio-temporal motion patterns in a multi-object video. We use the Helmholtz decomposition of optical flow and compute singular points corresponding to component fields. We will show that the optical flow can be used to identify regions which correspond to different moving entities in the video. The singular points in these regions capture the characteristics of the field around them and can be used to identify these regions. This representation would provide us with a framework to analyze activities of individual entities in the scene as well as the global interactions between them. We demonstrate our algorithm on a dataset composed of multi-object videos recorded in a realistic environment. Nandita M. Nayak, Ahmed Tashrif Kamal, Amit K. Roy-Chowdhury |
ICIP | 3 |
| 2011 | A Physics-Based Analysis of Image Appearance ModelsabstractLinear and multilinear models (PCA, 3DMM, AAM/ASM, and multilinear tensors) of object shape/appearance have been very popular in computer vision. In this paper, we analyze the applicability of these heuristic models from the fundamental physical laws of object motion and image formation. We prove that under suitable conditions, the image appearance space can be closely approximated to be multilinear, with the illumination and texture subspaces being trilinearly combined with the direct sum of the motion and deformation subspaces. This result provides a physics-based understanding of many of the successes and limitations of the linear and multilinear approaches existing in the computer vision literature, and also identifies some of the conditions under which they are valid. It provides an analytical representation of the image space in terms of different physical factors that affect the image formation process. Numerical analysis of the accuracy of the physics-based models is performed, and tracking results on real data are presented. Yilei Xu, Amit K. Roy-Chowdhury |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Interactive Event Search through Transfer Learning
Antony Lam, Amit K. Roy-Chowdhury, Christian R. Shelton |
ACCV (3) | 2 |
| 2010 | Multilinear feature extraction and classification of multi-focal images, with applications in nematode taxonomyabstractIn this paper, we present a 3D X-Ray Transform based multilinear feature extraction and classification method for Digital Multi-focal Images (DMI). In such images, morphological information for a transparent specimen can be captured in the form of a stack of high-quality images, representing individual focal planes through the specimen's body. We present a method that can effectively exploit the entire information in the stack using the 3D X-Ray projections at different viewing angles. These DMI stacks represent the effect of different factors - shape, texture, viewpoint, different instances within the same class and different classes of specimens. For this purpose, we embed the 3D X-Ray Transform within a multilinear framework and propose a Multilinear X-Ray Transform (MXRT) feature representation. By combining the tensor texture and shape information we can get better recognition rates than just relying on the original or key frames of DMI stacks. The experimental results on the nematode DMI data show that the 3D X-Ray Transform based multilinear analysis method can effectively give 100% recognition rate on a real-life database. Min Liu 0008, Amit K. Roy-Chowdhury |
CVPR | 2 |
| 2010 | A Stochastic Graph Evolution Framework for Robust Multi-target Tracking
Bi Song, Ting-Yueh Jeng, Elliot Staudt, Amit K. Roy-Chowdhury |
ECCV (1) | 4 |
| 2010 | Multi-target tracking using long-term stochastic associationsabstractMaintaining the stability of tracks on multiple targets in video over extended time periods remains a challenging problem. A few methods which have recently shown encouraging results in this direction rely on learning context models or the availability of training data. However, this may not be feasible in many application scenarios. Moreover, tracking methods should be able to work across multiple resolutions of the video. In this paper, we consider the problem of long-term tracking in video in application domains where context information is not available a priori, nor can it be learned online. We build our solution on the hypothesis that most existing trackers can obtain reasonable short-term tracks (tracklets). By analyzing the statistical properties of these tracklets, we develop associations between them so as to come up with longer tracks. On multiple real-life video sequences spanning low and high resolution data, we show the ability to accurately track over extended time periods. Ting-Yueh Jeng, Bi Song, Elliot Staudt, Min Liu 0008, Amit K. Roy-Chowdhury, Ashis SenGupta |
ICIP | 5 |
| 2010 | Multi-focal nematode image classification using the 3D X-Ray TransformabstractIn this paper, we present a 3D X-Ray Transform based feature extraction and classification method for Digital Multi-focal Images (DMI). In such images, morphological information for a transparent specimen can be captured in the form of a stack of high-quality images, representing individual focal planes through the specimen's body. We present a method that can effectively exploit the entire information in the stack using the 3D X-Ray Transform at different angle views. By combining the texture and shape information from different angles, we can get better recognition rates than just relying on the original or key frames of DMI stacks. The experimental results on the nematode DMI data show that the 3D X-Ray Transform based classification method can effectively improve the recognition rate from 60% (PCA) to 96.8%. Min Liu 0008, Amit K. Roy-Chowdhury, Melissa Yoder, Paul De Ley |
ICIP | 2 |
| 2010 | Pattern analysis of stem cell growth dynamics in the shoot apex of arabidopsisabstractThe Shoot Apical Meristem (SAM) is made of stem cells that are responsible for all above ground plant structures. Differentiating cells in the development of SAM form primordia. Primordia develop to become various plant organs. Understanding the growth dynamics of primordia is critical to understanding the developmental dynamics of the entire SAM. We present a method for performing quantitative analysis of primordia development in model plant Arabidopsis thaliana. A contour based approach is used to detect and isolate individual primordia from 3D live imaging data. Regions of growth are detected by analyzing eigenvalues of curvature covariance matrices. After primordia detection and isolation, a Dynamic Time Warping (DTW) Algorithm is applied to compute the rate of growth. Results show the successful use of our method to quantitatively analyze primordial growth. Oben M. Tataw, Min Liu 0008, Amit K. Roy-Chowdhury, Ram Kishor Yadav, G. Venugopala Reddy |
ICIP | 3 |
| 2010 | A Neurobiologically Motivated Stochastic Method for Analysis of Human Activities in VideoabstractIn this paper, we develop a neurobiologically-motivated statistical method for video analysis that simultaneously searches the combined motion and form space in a concerted and efficient manner using well-known Markov chain Monte Carlo (MCMC) techniques. Specifically, we leverage upon an MCMC variant called the Hamiltonian Monte Carlo (HMC), which we extend to utilize data-based proposals rather than the blind proposals in a traditional HMC, thus creating the Data-Driven HMC (DDHMC). We demonstrate the efficacy of our system on real-life video sequences. Ricky J. Sethi, Amit K. Roy-Chowdhury |
ICPR | 2 |
| 2010 | The Human Action ImageabstractRecognizing a person's motion is intuitive for humans but represents a challenging problem in machine vision. In this paper, we present a multi-disciplinary framework for recognizing human actions. We develop a novel descriptor, the Human Action Image (HAI): a physically-significant, compact representation for the motion of a person, which we derive from first principles in physics using Hamilton's Action. We embed the HAI as the Motion Energy Pathway of the latest Neurobiological model of motion recognition. The Form Pathway is modelled using existing low-level feature descriptors based on shape and appearance. Experimental validation of the theory is provided on the well-known Weizmann and USF Gait datasets. Ricky J. Sethi, Amit K. Roy-Chowdhury |
ICPR | 2 |
| 2010 | Tracking and Activity Recognition Through Consensus in Distributed Camera NetworksabstractCamera networks are being deployed for various applications like security and surveillance, disaster response and environmental modeling. However, there is little automated processing of the data. Moreover, most methods for multicamera analysis are centralized schemes that require the data to be present at a central server. In many applications, this is prohibitively expensive, both technically and economically. In this paper, we investigate distributed scene analysis algorithms by leveraging upon concepts of consensus that have been studied in the context of multiagent systems, but have had little applications in video analysis. Each camera estimates certain parameters based upon its own sensed data which is then shared locally with the neighboring cameras in an iterative fashion, and a final estimate is arrived at in the network using consensus algorithms. We specifically focus on two basic problems-tracking and activity recognition. For multitarget tracking in a distributed camera network, we show how the Kalman-Consensus algorithm can be adapted to take into account the directional nature of video sensors and the network topology. For the activity recognition problem, we derive a probabilistic consensus scheme that combines the similarity scores of neighboring cameras to come up with a probability for each action at the network level. Thorough experimental results are shown on real data along with a quantitative analysis. Bi Song, Ahmed Tashrif Kamal, Cristian Soto, Chong Ding, Jay A. Farrell, Amit K. Roy-Chowdhury |
IEEE Trans. Image Process. | 6 |
| 2009 | Distributed multi-target tracking in a self-configuring camera networkabstractThis paper deals with the problem of tracking multiple targets in a distributed network of self-configuring pan-tilt-zoom cameras. We focus on applications where events unfold over a large geographic area and need to be analyzed by multiple overlapping and non-overlapping active cameras without a central unit accumulating and analyzing all the data. The overall goal is to keep track of all targets in the region of deployment of the cameras, while selectively focusing at a high resolution on some particular target features. To acquire all the targets at the desired resolutions while keeping the entire scene in view, we use cooperative network control ideas based on multi-player learning in games. For tracking the targets as they move through the area covered by the cameras, we propose a special application of the distributed estimation algorithm known as Kalman-Consensus filter through which each camera comes to a consensus with its neighboring cameras about the actual state of the target. This leads to a camera network topology that changes with time. Combining these ideas with single-view analysis, we have a completely distributed approach for multi-target tracking and camera network self-configuration. We show performance analysis results with real-life experiments on a network of 10 cameras. Cristian Soto, Bi Song, Amit K. Roy-Chowdhury |
CVPR | 3 |
| 2009 | Exploiting local structure for tracking plant cells in noisy imagesabstractIn this paper, we present a local graph matching based method for tracking cells and cell divisions in noisy images. We work with plant cells, where the cells are tightly clustered in space and computing correspondences across time can be very challenging. The local graph matching method is able to track the cells and cell divisions even when significant portions of the images are corrupted due to sensor noise in the imaging process. The geometric structure and topology of the cells' relative positions are efficiently exploited to solve the tracking problem using the local graph matching technique. Using this method we can track almost all of the properly segmented cells, even when some of those images are highly noisy. Min Liu 0008, Amit K. Roy-Chowdhury, G. Venugopala Reddy |
ICIP | 2 |
| 2009 | Rate-Invariant Recognition of Humans and Their ActivitiesabstractPattern recognition in video is a challenging task because of the multitude of spatio-temporal variations that occur in different videos capturing the exact same event. While traditional pattern-theoretic approaches account for the spatial changes that occur due to lighting and pose, very little has been done to address the effect of temporal rate changes in the executions of an event. In this paper, we provide a systematic model-based approach to learn the nature of such temporal variations (time warps) while simultaneously allowing for the spatial variations in the descriptors. We illustrate our approach for the problem of action recognition and provide experimental justification for the importance of accounting for rate variations in action recognition. The model is composed of a nominal activity trajectory and a function space capturing the probability distribution of activity-specific time warping transformations. We use the square-root parameterization of time warps to derive geodesics, distance measures, and probability distributions on the space of time warping functions. We then design a Bayesian algorithm which treats the execution rate function as a nuisance variable and integrates it out using Monte Carlo sampling, to generate estimates of class posteriors. This approach allows us to learn the space of time warps for each activity while simultaneously capturing other intra- and interclass variations. Next, we discuss a special case of this approach which assumes a uniform distribution on the space of time warping functions and show how computationally efficient inference algorithms may be derived for this special case. We discuss the relative advantages and disadvantages of both approaches and show their efficacy using experiments on gait-based person identification and activity recognition. Ashok Veeraraghavan, Anuj Srivastava, Amit K. Roy-Chowdhury, Rama Chellappa |
IEEE Trans. Image Process. | 3 |
| 2008 | Learning a geometry integrated image appearance manifold from a small training setabstractWhile low-dimensional image representations have been very popular in computer vision, they suffer from two limitations: (i) they require collecting a large and varied training set to learn a low-dimensional set of basis functions, and (ii) they do not retain information about the 3D geometry of the object being imaged. In this paper, we show that it is possible to estimate low-dimensional manifolds that describe object appearance while retaining the geometrical information about the 3D structure of the object. By using a combination of analytically derived geometrical models and statistical learning methods, this can be achieved using a much smaller training set than most of the existing approaches. Specifically, we derive a quadrilinearmanifold of object appearance that can represent the effects of illumination, pose, identity and deformation, and the basis functions of the tangent space to this manifold depend on the 3D surface normals of the objects. We show experimental results on constructing this manifold and how to efficiently track on it using an inverse compositional algorithm. Yilei Xu, Amit K. Roy-Chowdhury |
CVPR | 2 |
| 2008 | A theoretical analysis of linear and multi-linear models of image appearanceabstractLinear and multi-linear models of object shape/appearance (PCA, 3 DMM, AAM/ASM, multilinear tensors) have been very popular in computer vision. In this paper, we analyze the validity of these models from the fundamental physical laws of object motion and image formation. We rigorously prove that the image appearance space can be closely approximated to be locally multilinear, with the illumination subspace being bilinearly combined with the direct sum of the motion, deformation and texture subspaces. This result allows us to understand theoretically many of the successes and limitations of the linear and multi-linear approaches existing in the computer vision literature, and also identifies some of the conditions under which they are valid. It provides an analytical representation of the image space in terms of different physical factors that affect the image formation process. Experimental analysis of the accuracy of the theoretical models is performed as well as tracking on real data using the analytically derived basis functions of this space. Yilei Xu, Amit K. Roy-Chowdhury |
CVPR | 2 |
| 2008 | Efficient motion estimation under varying illuminationabstractIn this paper, we show how to estimate, accurately and efficiently, the 3D motion of a rigid or non-rigid object, and time-varying lighting in a dynamic scene. This is achieved in an inverse compositional tracking framework with a novel warping function that involves a 2D rarr3Drarr2D transformation. The method is guaranteed to converge, is able to work with rigid and non-rigid objects, and estimates the lighting and motion from a video sequence. Experimental analysis on multiple face video sequences shows impressive speed-up over existing methods while retaining a high level of accuracy. Yilei Xu, Amit K. Roy-Chowdhury |
ICIP | 2 |
| 2008 | Inverse Compositional Estimation of 3D Pose And Lighting in Dynamic ScenesabstractIn this paper, we show how to estimate, accurately and efficiently, the 3D motion of a rigid object and time-varying lighting in a dynamic scene. This is achieved in an inverse compositional tracking framework with a novel warping function that involves a 2D --> 3D --> 2D transformation. This also allows us to extend traditional two frame inverse compositional tracking to a sequence of frames, leading to even higher computational savings. We prove the theoretical convergence of this method and show that it leads to significant reduction in computational burden. Experimental analysis on multiple video sequences shows impressive speed-up over existing methods while retaining a high level of accuracy. Yilei Xu, Amit K. Roy-Chowdhury |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Human Identification using Gait and FaceabstractIn general the visual-hull approach for performing integrated face and gait recognition requires at least two cameras. In this paper we present experimental results for fusion of face and gait for the single camera case. We considered the NIST database which contains outdoor face and gait data for 30 subjects. In the NIST database, subjects walk along an inverted Sigma pattern. In (A. Kale, et al., 2003), we presented a view-invariant gait recognition algorithm for the single camera case along with some experimental evaluations. In this chapter we present the results of our view-invariant gait recognition algorithm in (A. Kale, et al., 2003) on the NIST database. The algorithm is based on the planar approximation of the person which is valid when the person walks far away from the camera. In (S. Zhou et al., 2003), an algorithm for probabilistic recognition of human faces from video was proposed and the results were demonstrated on the NIST database. Details of these methods can be found in the respective papers. We give an outline of the fusion strategy here. Rama Chellappa, Amit K. Roy-Chowdhury, Amit A. Kale |
CVPR | 2 |
| 2007 | Closed-Loop Tracking and Change Detection in Multi-Activity SequencesabstractWe present a novel framework for tracking of a long sequence of human activities, including the time instances of change from one activity to the next, using a closed-loop, non-linear dynamical feedback system. A composite feature vector describing the shape, color and motion of the objects, and a non-linear, piecewise stationary, stochastic dynamical model describing its spatio-temporal evolution, are used for tracking. The tracking error or expected log likelihood, which serves as a feedback signal, is used to automatically detect changes and switch between activities happening one after another in a long video sequence. Whenever a change is detected, the tracker is re initialized automatically by comparing the input image with learned models of the activities. Unlike some other approaches that can track a sequence of activities, we do not need to know the transition probabilities between the activities, which can be difficult to estimate in many application scenarios. We demonstrate the effectiveness of the method on multiple indoor and outdoor real-life videos and analyze its performance. Bi Song, Namrata Vaswani, Amit K. Roy-Chowdhury |
CVPR | 3 |
| 2007 | Pose and Illumination Invariant Face Recognition in VideoabstractThe use of video sequences for face recognition has been relatively less studied than image-based approaches. In this paper, we present a framework for face recognition from video sequences that is robust to large changes in facial pose and lighting conditions. Our method is based on a recently obtained theoretical result that can integrate the effects of motion, lighting and shape in generating an image using a perspective camera. This result can be used to estimate the pose and illumination conditions for each frame of the probe sequence. Then, using a 3D face model, we synthesize images corresponding to the pose and illumination conditions estimated in the probe sequences. Similarity between the synthesized images and the probe video is computed by integrating over the entire sequence. The method can handle situations where the pose and lighting conditions in the training and testing data are completely disjoint. Yilei Xu, Amit K. Roy-Chowdhury, Keyur Patel |
CVPR | 2 |
| 2007 | Stochastic Adaptive Tracking In A Camera NetworkabstractWe present a novel stochastic, adaptive strategy for tracking multiple people in a large network of video cameras. Similarities between features (appearance and biometrics) observed at different cameras are continuously adapted and the stochastically optimal path for each person computed. The following are the major contributions of the proposed approach. First, we consider situations where the feature similarities are uncertain and treat them as random variables. We show how the distributions of these random variables can be learned and how to compute the tracks in a stochastically optimal manner. Second, we consider the possibility of long-term interdependence of the features over space and time. This allows us to adoptively evolve the feature correspondences by observing the system performance over a time window, and correct for errors in the similarity computations. Third, we show that the above two conditions can be addressed by treating the issue of tracking in a camera network as an optimization problem in a stochastic adaptive system. We show results on data collected by a large camera network. The proposed approach is particularly suitable for distributed processing over the entire network. Bi Song, Amit K. Roy-Chowdhury |
ICCV | 2 |
| 2007 | Modeling Time-Varying Illumination Patterns in VideoabstractRecreating the temporal illumination variations of natural scenes has great potential for realistic synthesis of video sequences. In this paper, we present a 3D (model-based) approach that achieves this goal. The approach requires a training sequence to learn the time-varying illumination models, which can then be used for synthesis in another sequence. The motion and illumination parameters in the training sequence are estimated alternately by projecting onto appropriate basis functions of a bilinear space defined in terms of the 3D surface normals of the objects. The motion is represented in terms of 3D translation and rotation of the object centroid in the camera frame, and the illumination is represented using a spherical harmonics linear basis. We show video synthesis results using the proposed approach. Yilei Xu, Amit K. Roy-Chowdhury |
ICIP (2) | 2 |
| 2007 | Super-Resolved Facial Texture Under Changing Pose and IlluminationabstractIn this paper, we propose a method to incrementally super-resolve 3D facial texture by integrating information frame by frame from a video captured under changing poses and illuminations. First, we recover illumination, 3D motion and shape parameters from our tracking algorithm. This information is then used to super-resolve 3D texture using iterative back-projection (IBP) method. Finally, the super-resolved texture is fed back to the tracking part to improve the estimation of illumination and motion parameters. This closed-loop process continues to refine the texture as new frames come in. We also propose a local-region based scheme to handle non-rigidity of the human face. Experiments demonstrate that our framework not only incrementally super-resolves facial images, but recovers the detailed expression changes in high quality. Jiangang Yu, Bir Bhanu, Yilei Xu, Amit K. Roy-Chowdhury |
ICIP (3) | 4 |
| 2007 | Determining Topology in a Distributed Camera NetworkabstractRecently, 'entry/exit' events of objects in the field-of-views of cameras were used to learn the topology of the camera network. The integration of object appearance was also proposed to employ the visual information provided by the imaging sensors. A problem with these methods is the lack of robustness to appearance changes. This paper integrates face recognition in the statistical model to better estimate the correspondence in the time-varying network. The statistical dependence between the entry and exit nodes indicates the connectivity and traffic patterns of the camera network, which are represented by a weighted directed graph and transition time distributions. A nine-camera network with 25 nodes is analyzed both in simulation and in real-life experiments, and compared with the previous approaches. Xiaotao Zou, Bir Bhanu, Bi Song, Amit K. Roy-Chowdhury |
ICIP (5) | 4 |
| 2007 | Integrating Motion, Illumination, and Structure in Video Sequences with Applications in Illumination-Invariant TrackingabstractIn this paper, we present a theory for combining the effects of motion, illumination, 3D structure, albedo, and camera parameters in a sequence of images obtained by a perspective camera. We show that the set of all Lambertian reflectance functions of a moving object, at any position, illuminated by arbitrarily distant light sources, lies "close" to a bilinear subspace consisting of nine illumination variables and six motion variables. This result implies that, given an arbitrary video sequence, it is possible to recover the 3D structure, motion, and illumination conditions simultaneously using the bilinear subspace formulation. The derivation builds upon existing work on linear subspace representations of reflectance by generalizing it to moving objects. Lighting can change slowly or suddenly, locally or globally, and can originate from a combination of point and extended sources. We experimentally compare the results of our theory with ground truth data and also provide results on real data by using video sequences of a 3D face and the entire human body with various combinations of motion and illumination directions. We also show results of our theory in estimating 3D motion and illumination model parameters from a video sequence. Yilei Xu, Amit K. Roy-Chowdhury |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Towards a measure of deformability of shape sequences
Amit K. Roy-Chowdhury |
Pattern Recognit. Lett. | 1 |
| 2006 | The Function Space of an ActivityabstractAn activity consists of an actor performing a series of actions in a pre-defined temporal order. An action is an individual atomic unit of an activity. Different instances of the same activity may consist of varying relative speeds at which the various actions are executed, in addition to other intra- and inter- person variabilities. Most existing algorithms for activity recognition are not very robust to intra- and inter-personal changes of the same activity, and are extremely sensitive to warping of the temporal axis due to variations in speed profile. In this paper, we provide a systematic approach to learn the nature of such time warps while simultaneously allowing for the variations in descriptors for actions. For each activity we learn an ‘average’ sequence that we denote as the nominal activity trajectory. We also learn a function space of time warpings for each activity separately. The model can be used to learn individualspecific warping patterns so that it may also be used for activity based person identification. The proposed model leads us to algorithms for learning a model for each activity, clustering activity sequences and activity recognition that are robust to temporal, intra- and inter-person variations. We provide experimental results using two datasets. Ashok Veeraraghavan, Amit K. Roy-Chowdhury |
CVPR (1) | 2 |
| 2006 | Towards a Multi-Terminal Video Compression Algorithm Using Epipolar GeometryabstractWe present a novel distributed video coding algorithm based on transform coding of distributed sources and exploiting the geometrical relationships between the location of the sensors. The geometry is used to align the video sequences and distributed quantization of transform coefficients is used to eliminate spatial and inter-sensor redundancy. In contrast with most of the current video compression standards which only exploit spatial and temporal dundancy within each video sequence, we also consider the significant redundancy between the sequences. Results demonstrate that our algorithm yields a significant saving in bit rate on the overlapping portion of multiple views. Bi Song, Ozgun Y. Bursalioglu, Amit K. Roy-Chowdhury, Ertem Tuncel |
ICASSP (2) | 3 |
| 2006 | An Illumination Invariant 3D Model Based Tracking Algorithm, with Application in Video CompressionabstractWe present an algorithm for illumination invariant 3D model based tracking and video compression. While model-based coding schemes are well developed, the compression rate reduces if the illumination conditions within the sequence change drastically. Our proposed scheme can accurately track under both slow and drastic changes of lighting and achieves a significant reduction in the bit rate compared to a scheme that does not consider any illumination models. The tracking scheme is based on a recently developed framework showing that the joint motion and illumination space of video sequences is approximately bilinear. We show results of the reconstructed video sequences and the reduction in the distortion (for a constant bit rate) for a compression scheme that uses the illumination invariant tracking algorithm. Long Nyugen, Yilei Xu, Amit K. Roy-Chowdhury |
ICIP | 3 |
| 2006 | A Multi-Terminal Model-Based Video Compression AlgorithmabstractWe present a novel 3D model-based distributed video coding algorithm. It is based on independent, model-based tracking of multiple sources and distributed coding of the tracked feature points. The model-based tracking scheme provides correspondence between the overlapping set of features that are visible in the different views. While the motion estimates obtained from the tracking algorithm remove temporal redundancy and the 3D model accounts for removing spatial redundancy, distributed coding is used to eliminate inter-sensor redundancy. Thus, in contrast to most of the current video compression standards which only exploit spatial and temporal redundancy within each video sequence, we also consider the significant redundancy between the sequences. Results demonstrate that our algorithm yields a significant saving in bit rate on the overlapping portion of multiple views. Bi Song, Amit K. Roy-Chowdhury, Ertem Tuncel |
ICIP | 2 |
| 2006 | Summarization and Indexing of Human Activity SequencesabstractIn order to summarize a video consisting of a sequence of different activities, there are three fundamental problems: tracking the objects of interest, detecting the activity change times and recognizing the new activity. This paper presents an algorithm for achieving all these three tasks simultaneously and presents results on how it can used for indexing and summarizing a real-life video sequence. Human activities are represented by a model for the dynamics of the shape of the human body contour. Measures are designed for detecting both gradual transitions and sudden changes between activity models. Bi Song, Namrata Vaswani, Amit K. Roy-Chowdhury |
ICIP | 3 |
| 2005 | A Measure of Deformability of Shapes, with Applications to Human Motion AnalysisabstractIn this paper we develop a theory for characterizing how deformable a shape is. We define a term called "deformability index" for shapes. The deformability index is computed from the tracked positions of a sequence of deformable shapes, using a scaled orthographic camera projection model. Our method assumes that a deformable shape sequence can be represented by a linear combination of basis shapes, where the weights assigned to each basis shape changes with time. The tracked points obtained from the shape sequence is transformed to a 3D shape space. Using statistical models to separate out the "true" deformations from those induced by noise in the trajectories, the dimension of this shape space is estimated using spectral analysis methods. The dimension of this shape space determines the number of basis shapes needed to represent the shape sequence, which, in turn, determines the deformability index. Rigid 3D transformations of the shape are taken into account in estimating the deformability index; however, the method does not require estimation of 3D structure or motion. Experimental results are shown using motion capture data as well as real imagery of different human activities. The results show that the deformability index is in accordance with our intuitive judgement and corroborates certain hypotheses in human movement analysis studies. Amit K. Roy-Chowdhury |
CVPR (1) | 1 |
| 2005 | An algorithm for 3D reconstruction of deformable shape sequencesabstractIn this paper, we present an algorithm for estimating the 3D model of a deformable shape from a video sequence. Our method assumes that a deformable shape sequence can be represented by a linear combination of basis shapes, where the weights assigned to each basis shape change with time. While there is existing work on estimating the basis shapes and their combination coefficients, they lack the crucial information about the number of basis shapes that are required for the model. This is usually determined through heuristics about the physics of the underlying structure. We show that it is possible to estimate the number of basis shapes from the tracked points obtained from the video sequence, using a scaled orthographic camera projection model. This estimate is then used to compute the 3D structure of each of the basis shapes. We present experimental results in recreating the structure of the human body during various activities from a video sequence. Amit K. Roy-Chowdhury |
ICASSP (2) | 1 |
| 2005 | Integrating the Effects of Motion, Illumination and Structure in Video SequencesabstractMost work in computer vision has concentrated on studying the individual effect of motion and illumination on a 3D object. In this paper, we present a theory for combining the effects of motion, illumination, 3D structure, albedo, and camera parameters in a sequence of images obtained by a perspective camera. We show that the set of all Lambertian reflectance functions of a moving object, illuminated by arbitrarily distant light sources, lies close to a bilinear subspace consisting of nine illumination variables and six motion variables. This result implies that, given an arbitrary video sequence, it is possible to recover the 3D structure, motion and illumination conditions simultaneously using the bilinear subspace formulation. The derivation is based on the intuitive notion that, given an illumination direction, the images of a moving surface cannot change suddenly over a short time period. We experimentally compare the images obtained using our theory with ground truth data and show that the difference is small and acceptable. We also provide experimental results on real data by synthesizing video sequences of a 3D face with various combinations of motion and illumination directions. Yilei Xu, Amit K. Roy-Chowdhury |
ICCV | 2 |
| 2005 | A study on view-insensitive gait recognitionabstractMost gait recognition approaches only study human walking frontoparallel to the image plane which is not realistic in video surveillance applications. Human gait appearance depends on various factors including locations of the camera and the person, the camera axis and the walking direction. By analyzing these factors, we propose a statistical approach for view-insensitive gait recognition. The proposed approach recognizes human using a single camera, and avoids the difficulties of recovering the human body structure and camera calibration. Experimental results show that the proposed approach achieves good performance in recognizing individuals walking along different directions. Ju Han, Bir Bhanu, Amit K. Roy-Chowdhury |
ICIP (3) | 3 |
| 2005 | The joint illumination and motion space of video sequencesabstractIt has been proved that the set of all Lambertian reflectance functions obtained with arbitrarily distant light sources lies close to a 9D linear subspace. We extend this result from still images to video sequences. We show that the set of all Lambertian reflectance functions of a moving object at any position, illuminated by arbitrarily distant light sources, lies close to a bilinear subspace consisting of nine illumination variables and six motion variables. This result implies that, when the position and 3D model of an object at one instance of time is known, the reflectance images at future time instances can be estimated using the bilinear subspace. This is based on the fact that, given the illumination direction, the image of a moving surface cannot change suddenly over a short time period. We apply our theory to synthesize video sequences of a 3D face with various combinations of motion and illumination directions. Yilei Xu, Amit K. Roy-Chowdhury |
ICIP (2) | 2 |
| 2005 | Matching Shape Sequences in Video with Applications in Human Movement AnalysisabstractWe present an approach for comparing two sequences of deforming shapes using both parametric models and nonparametric methods. In our approach, Kendall's definition of shape is used for feature extraction. Since the shape feature rests on a non-Euclidean manifold, we propose parametric models like the autoregressive model and autoregressive moving average model on the tangent space and demonstrate the ability of these models to capture the nature of shape deformations using experiments on gait-based human recognition. The nonparametric model is based on Dynamic Time-Warping. We suggest a modification of the Dynamic time-warping algorithm to include the nature of the non-Euclidean space in which the shape deformations take place. We also show the efficacy of this algorithm by its application to gait-based human recognition. We exploit the shape deformations of a person's silhouette as a discriminating feature and provide recognition results using the nonparametric model. Our analysis leads to some interesting observations on the role of shape and kinematics in automated gait-based person authentication. Ashok Veeraraghavan, Amit K. Roy-Chowdhury, Rama Chellappa |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Statistical bias in 3-D reconstruction from a monocular videoabstractThe present state-of-the-art in computing the error statistics in three-dimensional (3-D) reconstruction from video concentrates on estimating the error covariance. A different source of error which has not received much attention is the fact that the reconstruction estimates are often significantly statistically biased. In this paper, we derive a precise expression for the bias in the depth estimate, based on the continuous (differentiable) version of structure from motion (SfM). Many SfM algorithms, or certain portions of them, can be posed in a linear least-squares (LS) framework Ax = b. Examples include initialization procedures for bundle adjustment or algorithms that alternately estimate depth and camera motion. It is a well-known fact that the LS estimate is biased if the system matrix A is noisy. In SfM, the matrix A contains point correspondences, which are always difficult to obtain precisely; thus, it is expected that the structure and motion estimates in such a formulation of the problem would be biased. Existing results on the minimum achievable variance of the SfM estimator are extended by deriving a generalized Cramer-Rao lower bound. A detailed analysis of the effect of various camera motion parameters on the bias is presented. We conclude by presenting the effect of bias compensation on reconstructing 3-D face models from rendered images. Amit K. Roy-Chowdhury, Rama Chellappa |
IEEE Trans. Image Process. | 1 |
| 2005 | "Shape Activity": A Continuous-State HMM for Moving/Deforming Shapes With Application to Abnormal Activity DetectionabstractThe aim is to model "activity" performed by a group of moving and interacting objects (which can be people, cars, or different rigid components of the human body) and use the models for abnormal activity detection. Previous approaches to modeling group activity include co-occurrence statistics (individual and joint histograms) and dynamic Bayesian networks, neither of which is applicable when the number of interacting objects is large. We treat the objects as point objects (referred to as "landmarks") and propose to model their changing configuration as a moving and deforming "shape" (using Kendall's shape theory for discrete landmarks). A continuous-state hidden Markov model is defined for landmark shape dynamics in an activity. The configuration of landmarks at a given time forms the observation vector, and the corresponding shape and the scaled Euclidean motion parameters form the hidden-state vector. An abnormal activity is then defined as a change in the shape activity model, which could be slow or drastic and whose parameters are unknown. Results are shown on a real abnormal activity-detection problem involving multiple moving objects. Namrata Vaswani, Amit K. Roy-Chowdhury, Rama Chellappa |
IEEE Trans. Image Process. | 2 |
| 2004 | Contour-based 3D Face Modeling from a Monocular VideoabstractConstructing 3D models from video is one of the most important problems in computer vision. We propose a novel 3D face modeling approach from monocular video captured by a conventional camera. An algorithm is proposed to estimate the head pose by comparing the edges of video frame, and the contours extracted from a generic face model. A generic 3D face model is assumed to be the initial estimate of the true 3D model. The generic face model is adapted to the actual 3D face model by global and local deformations. An affine model is used for global deformation. The 3D model is locally deformed by computing the optimal perturbations of a sparse set of control points using a stochastic search optimization method. The deformations are integrated over a set of poses in the video sequence, leading to an accurate 3D model. Himaanshu Gupta, Amit K. Roy-Chowdhury, Rama Chellappa |
BMVC | 2 |
| 2004 | Role of Shape and Kinematics in Human Movement Analysis
Ashok Veeraraghavan, Amit K. Roy-Chowdhury, Rama Chellappa |
CVPR (1) | 2 |
| 2004 | Fusion of gait and face for human identificationabstractIdentification of humans from arbitrary view points is an important requirement for different tasks including perceptual interfaces for intelligent environments, covert security and access control etc. For optimal performance, the system must use as many cues as possible and combine them in meaningful ways. In this paper, we discuss fusion of face and gait cues for the single camera case. We present a view invariant gait recognition algorithm for gait recognition. We employ decision fusion to combine the results of our gait recognition algorithm and a face recognition algorithm based on sequential importance sampling. We consider two fusion scenarios: hierarchical and holistic. The first involves using the gait recognition algorithm as a filter to pass on a smaller set of candidates to the face recognition algorithm. The second involves combining the similarity scores obtained individually from the face and gait recognition algorithms. Simple rules like the SUM, MIN and PRODUCT are used for combining the scores. The results of fusion experiments are demonstrated on the NIST database which has outdoor gait and face data of 30 subjects. Amit A. Kale, Amit K. Roy-Chowdhury, Rama Chellappa |
ICASSP (5) | 2 |
| 2004 | Facial similarity across age disguise illumination and pose
Narayanan Ramanathan, Rama Chellappa, Amit K. Roy-Chowdhury |
ICIP | 3 |
| 2004 | Multiple view tracking of humans modelled by kinematic chainsabstractWe use a kinematic chain to model human body motion. We estimate the kinematic chain motion parameters using pixel displacements calculated from video sequences obtained from multiple calibrated cameras to perform tracking. We derive a linear relation between the 2D motion of pixels in terms of the 3D motion parameters of various body parts using a perspective projection model for the cameras, a rigid body motion model for the base body and the kinematic chain model for the body parts. An error analysis of the estimator is provided, leading to an iterative algorithm for calculating the motion parameters from the pixel displacements. We provide experimental results to demonstrate the accuracy of our formulation. We also compare our iterative algorithm to the noniterative algorithm and discuss its robustness in the presence of noise. Aravind Sundaresan, Rama Chellappa, Amit K. Roy-Chowdhury |
ICIP | 3 |
| 2004 | An information theoretic criterion for evaluating the quality of 3-D reconstructions from videoabstractEven though numerous algorithms exist for estimating the three-dimensional (3-D) structure of a scene from its video, the solutions obtained are often of unacceptable quality. To overcome some of the deficiencies, many application systems rely on processing more data than necessary, thus raising the question: how is the accuracy of the solution related to the amount of data processed by the algorithm? Can we automatically recognize situations where the quality of the data is so bad that even a large number of additional observations will not yield the desired solution? Previous efforts to answer this question have used statistical measures like second order moments. They are useful if the estimate of the structure is unbiased and the higher order statistical effects are negligible, which is often not the case. This paper introduces an alternative information-theoretic criterion for evaluating the quality of a 3-D reconstruction. The accuracy of the reconstruction is judged by considering the change in mutual information (MI) (termed as the incremental MI) between a scene and its reconstructions. An example of 3-D reconstruction from a video sequence using optical flow equations and known noise distribution is considered and it is shown how the MI can be computed from first principles. We present simulations on both synthetic and real data to demonstrate the effectiveness of the proposed criterion. Amit K. Roy-Chowdhury, Rama Chellappa |
IEEE Trans. Image Process. | 1 |
| 2004 | Identification of humans using gaitabstractWe propose a view-based approach to recognize humans from their gait. Two different image features have been considered: the width of the outer contour of the binarized silhouette of the walking person and the entire binary silhouette itself. To obtain the observation vector from the image features, we employ two different methods. In the first method, referred to as the indirect approach, the high-dimensional image feature is transformed to a lower dimensional space by generating what we call the frame to exemplar (FED) distance. The FED vector captures both structural and dynamic traits of each individual. For compact and effective gait representation and recognition, the gait information in the FED vector sequences is captured in a hidden Markov model (HMM). In the second method, referred to as the direct approach, we work with the feature vector directly (as opposed to computing the FED) and train an HMM. We estimate the HMM parameters (specifically the observation probability B) based on the distance between the exemplars and the image features. In this way, we avoid learning high-dimensional probability density functions. The statistical nature of the HMM lends overall robustness to representation and recognition. The performance of the methods is illustrated using several databases. Amit A. Kale, Aravind Sundaresan, A. N. Rajagopalan 0001, Naresh P. Cuntoor, Amit K. Roy-Chowdhury, Volker Krüger, Rama Chellappa |
IEEE Trans. Image Process. | 5 |
| 2004 | Wide baseline image registration with application to 3-D face modelingabstractEstablishing correspondence between features in two images of the same scene taken from different viewing angles is a challenging problem in image processing and computer vision. However, its solution is an important step in many applications like wide baseline stereo, three-dimensional (3-D) model alignment, creation of panoramic views, etc. In this paper, we propose a technique for registration of two images of a face obtained from different viewing angles. We show that prior information about the general characteristics of a face obtained from video sequences of different faces can be used to design a robust correspondence algorithm. The method works by matching two-dimensional (2-D) shapes of the different features of the face (e.g., eyes, nose etc.). A doubly stochastic matrix, representing the probability of match between the features, is derived using the Sinkhorn normalization procedure. The final correspondence is obtained by minimizing the probability of error of a match between the entire constellation of features in the two sets, thus taking into account the global spatial configuration of the features. The method is applied for creating holistic 3-D models of a face from partial representations. Although this paper focuses primarily on faces, the algorithm can also be used for other objects with small modifications. Amit K. Roy-Chowdhury, Rama Chellappa, Trish Keaton |
IEEE Trans. Multim. | 1 |
| 2003 | Towards a View Invariant Gait Recognition AlgorithmabstractHuman gait is a spatio-temporal phenomenon and typifies the motion characteristics of an individual. The gait of a person is easily recognizable when extracted from a side-view of the person. Accordingly, gait-recognition algorithms work best when presented with images where the person walks parallel to the camera image plane. However, it is not realistic to expect this assumption to be valid in most real-life scenarios. Hence, it is important to develop methods whereby the side-view can be generated from any other arbitrary view in a simple, yet accurate, manner. This is the main theme of the paper. We show that if the person is far enough from the camera, it is possible to synthesize a side view (referred to as canonical view) from any other arbitrary view using a single camera. Two methods are proposed for doing this: (i) using the perspective projection model; (ii) using the optical flow based structure from motion equations. A simple camera calibration scheme for this method is also proposed. Examples of synthesized views are presented. Preliminary testing with gait recognition algorithms gives encouraging results. A by-product of this method is a simple algorithm for synthesizing novel views of a planar scene. Amit A. Kale, Amit K. Roy-Chowdhury, Rama Chellappa |
AVSS | 2 |
| 2003 | Activity Recognition Using the Dynamics of the Configuration of Interacting ObjectsabstractMonitoring activities using video data is an important surveillance problem. A special scenario is to learn the pattern of normal activities and detect abnormal events from a very low resolution video where the moving objects are small enough to be modeled as point objects in a 2D plane. Instead of tracking each point separately, we propose to model an activity by the polygonal 'shape' of the configuration of these point masses at any time t, and its deformation over time. We learn the mean shape and the dynamics of the shape change using hand-picked location data (no observation noise) and define an abnormality detection statistic for the simple case of a test sequence with negligible observation noise. For the more practical case where observation (point locations) noise is large and cannot be ignored, we use a particle filter to estimate the probability distribution of the shape given the noisy observations up to the current time. Abnormality detection in this case is formulated as a change detection problem. We propose a detection strategy that can detect both 'drastic' and 'slow' abnormalities. Our framework can be directly applied for object location data obtained using any type of sensors - visible, radar, infrared or acoustic. Namrata Vaswani, Amit K. Roy-Chowdhury, Rama Chellappa |
CVPR (2) | 2 |
| 2003 | Video synthesis of arbitrary views for approximately planar scenesabstractIn this paper, we propose a method to synthesize arbitrary views of a planar scene, given a monocular video sequence. The method is based on the availability of knowledge of the angle between the original and synthesized views. Such a method has many important applications, one of them being gait recognition. Gait recognition algorithms rely on the availability of an approximate side-view of the person. From a realistic viewpoint, such an assumption is impractical in surveillance applications and it is of interest to develop methods to synthesize a side view of the person, given an arbitrary view. For large distances from the camera, a planar approximation for the individual can be assumed. In this paper, we propose a perspective projection approach for recovering the direction of motion of the person purely from the video data, followed by synthesis of a new video sequence at a different angle. The algorithm works purely in the image and video domain, though 3D structure plays an implicit role in its theoretical justification. Examples of synthesized views using our method and performance evaluation are presented. Amit K. Roy-Chowdhury, Amit A. Kale, Rama Chellappa |
ICASSP (3) | 1 |
| 2003 | Statistical shape theory for activity modelingabstractMonitoring activities in a certain region from video data is an important surveillance problem. The goal is to learn the pattern of normal activities and detect unusual ones by identifying activities that deviate appreciably from the typical ones. We propose an approach using statistical shape theory based on the shape model of D.G. Kendall et al. (see "Shape and Shape Theory", John Wiley and Sons, 1999). In a low resolution video, each moving object is best represented as a moving point mass or particle. In this case, an activity can be defined by the interactions of all or some of these moving particles over time. We model this configuration of the particles by a polygonal shape formed from the locations of the points in a frame and the activity by the deformation of the polygons in time. These parameters are learned for each typical activity. Given a test video sequence, an activity is classified as abnormal if the probability for the sequence (represented by the mean shape and the dynamics of the deviations), given the model, is below a certain threshold The approach gives very encouraging results in surveillance applications using a single camera and is able to identify various kinds of abnormal behavior. Namrata Vaswani, Amit K. Roy-Chowdhury, Rama Chellappa |
ICASSP (3) | 2 |
| 2003 | A hidden Markov model based framework for recognition of humans from gait sequencesabstractIn this paper we propose a generic framework based on hidden Markov models (HMMs) for recognition of individuals from their gait. The HMM framework is suitable, because the gait of an individual can be visualized as his adopting postures from a set, in a sequence which has an underlying structured probabilistic nature. The postures that the individual adopts can be regarded as the states of the HMM and are typical to that individual and provide a means of discrimination. The framework assumes that, during gait, the individual transitions between N discrete postures or states but it is not dependent on the particular feature vector used to represent the gait information contained in the postures. The framework, thus, provides flexibility in the selection of the feature vector. The statistical nature of the HMM lends robustness to the model. In this paper we use the binarized background-subtracted image as the feature vector and use different distance metrics, such as those based on the L/sub 1/ and L/sub 2/ norms of the vector difference, and the normalized inner product of the vectors, to measure the similarity between feature vectors. The results we obtain are better than the baseline recognition rates reported before. Aravind Sundaresan, Amit K. Roy-Chowdhury, Rama Chellappa |
ICIP (2) | 2 |
| 2003 | Video based rendering of planar dynamic scenesabstractIn this paper, we propose a method to synthesize arbitrary views of a planar scene from a monocular video sequence of it. The 3-D direction of motion of the object is robustly estimated from the video sequence. Given this direction any other view of the object can be synthesized through a perspective projection approach, under assumptions of planarity. If the distance of the object from the camera is large, a planar approximation is reasonable even for non-planar scenes. Such a method has many important applications, one of them being gait recognition where a side view of the person is required. Our method can be used to synthesize the side-view of the person in case he/she does not present a side view to the camera. Since the planarity assumption is often an approximation, the effects of non-planarity can lead to inaccuracies in rendering and needs to be corrected for. Regions where this happens are examined and a simple technique based on weak perspective approximation is proposed to offset rendering inaccuracies. Examples of synthesized views using our method and performance evaluation are presented. Amit A. Kale, Amit K. Roy-Chowdhury, Rama Chellappa |
ICME | 2 |
| 2003 | Statistical shape theory for activity modelingabstractMonitoring activities in a certain region from video data is an important surveillance problem today. The goal is to learn the pattern of normal activities and detect unusual ones by identifying activities that deviate appreciably from the typical ones. In this paper we propose an approach using statistical shape theory (based on Kendall's shape model) [D.G. Kendall et al., 1999]. In a low resolution video each moving object is best represented as a moving point mass or particle. In this case, an activity can be defined by the interactions of all or some of these moving particles over time. We model this configuration of the particles by a polygonal shape formed from the locations of the points in a frame and the activity by the deformation of the polygons in time. These parameters are learnt for each typical activity. Given a test video sequence, an activity is classified as abnormal if the probability for the sequence (represented by the mean shape and the dynamics of the deviations), given the model is below a certain threshold. The approach gives very encouraging results in surveillance applications using a single camera and is able to identify various kinds of abnormal behaviors. Namrata Vaswani, Amit K. Roy-Chowdhury, Rama Chellappa |
ICME | 2 |
| 2003 | Face reconstruction from monocular video using uncertainty analysis and a generic model
Amit K. Roy-Chowdhury, Rama Chellappa |
Comput. Vis. Image Underst. | 1 |
| 2003 | Stochastic Approximation and Rate-Distortion Analysis for Robust Structure and Motion Estimation
Amit K. Roy-Chowdhury, Rama Chellappa |
Int. J. Comput. Vis. | 1 |
| 2002 | Towards a criterion for evaluating the quality of 3D reconstructionsabstractEven though numerous algorithms exist for estimating the structure of a scene from its video, the solutions obtained are often of unacceptable quality. To overcome some of the deficiencies, many application systems rely on processing more information than necessary with the hope that the redundancy will help improve the quality. This raises the question about how the accuracy of the solution is related to the amount of information processed by the algorithm. Can we define the accuracy of the solution precisely enough that we automatically recognize situations where the quality of the data is so bad that even a large number of additional observations will not yield the desired solution? This paper proposes an information theoretic criterion for evaluating the quality of a 3D reconstruction in terms of the statistics of the observed parameters (i.e. the image correspondences). The accuracy of the reconstruction is judged by considering the change in mutual information (or equivalently the conditional differential entropy) between a scene and its reconstructions and its effectiveness is shown through simulations. Amit K. Roy-Chowdhury, Rama Chellappa |
ICASSP | 1 |
| 2002 | 3D face reconstruction from video using a generic modelabstractReconstructing a 3D model of a human face from a video sequence is an important problem in computer vision, with applications to recognition, surveillance, multimedia etc. However, the quality of 3D reconstructions using structure from motion (SfM) algorithms is often not satisfactory. One common method of overcoming this problem is to use a generic model of a face. Existing work using this approach initializes the reconstruction algorithm with this generic model. The problem with this approach is that the algorithm can converge to a solution very close to this initial value, resulting in a reconstruction which resembles the generic model rather than the particular face in the video which needs to be modeled. We propose a method of 3D reconstruction of a human face from video in which the 3D reconstruction algorithm and the generic model are handled separately. A 3D estimate is obtained purely from the video sequence using SfM algorithms without use of the generic model. The final 3D model is obtained after combining the SfM estimate and the generic model using an energy function that corrects for the errors in the estimate by comparing local regions in the two models. The optimization is done using a Markov chain Monte Carlo (MCMC) sampling strategy. The main advantage of our algorithm over others is that it is able to retain the specific features of the face in the video sequence even when these features are different from those of the generic model. The evolution of the 3D model through the various stages of the algorithm is presented. Amit K. Roy-Chowdhury, Rama Chellappa, Sandeep Krishnamurthy, Tai Vo |
ICME (1) | 1 |
| 2001 | A robust algorithm for fusing noisy depth estimates using stochastic approximationabstractThe problem of structure from motion (SFM) is to extract the three-dimensional model of a moving scene from a sequence of images. Most of the algorithms which work by fusing the two-frame depth estimates (observations) assume an underlying statistical model for the observations and do not evaluate the quality of the individual observations. However, in real scenarios, it is often difficult to justify the statistical assumptions. Also, outliers are present in any observation sequence and need to be identified and removed from the fusion algorithm. We present a recursive fusion algorithm using the Robbins-Monro stochastic approximation (RMSA) which takes care of both these problems to provide an estimate of the real depth of the scene point. The estimate converges to the true value asymptotically. We also propose a method to evaluate the importance of the successive observations by computing the Fisher information (FI) recursively. Though we apply our algorithm in the SFM problem by modeling of a human face, it can be easily adopted to other data fusion applications. Amit K. Roy-Chowdhury, Rama Chellappa |
ICASSP | 1 |
| 2001 | Robust estimation of depth and motion using stochastic approximationabstractThe problem of structure from motion (SfM) is to extract the three-dimensional model of a moving scene from a sequence of images. Though two images are sufficient to produce a 3D reconstruction, they usually perform poorly because of errors in the estimation of the camera motion and image correspondences, thus motivating the need for multiple frame algorithms. One common approach to this problem is to determine the estimate from pairs of images and then fuse them together. Data fusion techniques, like the Kalman filter, require estimates of the error in modeling and observations. The complexity of the SfM problem makes it difficult to reliably estimate these errors. This paper describes a new recursive algorithm to estimate the camera motion and scene structure by fusing the two-frame estimates, using stochastic approximation techniques. The method does not require estimates of the error in the two-frame case and can reconstruct the scene to arbitrary accuracy given a sufficient number of frames. Experimental results are reported to support these claims. Rama Chellappa, Amit K. Roy-Chowdhury |
ICIP (1) | 2 |