Ramakant Nevatia

dblp:n/RamakantNevatia · also Ram Nevatia · DBLP profile ↗
← Back
229ranked-venue papers
8as first author
22since 2021 · last 2024
0009-0003-8079-4209ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 174 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 173 · 4 first-author · 20 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Large Language Models are Good Prompt Learners for Low-Shot Image Classification
abstract
Low-shot image classification, where training images are limited or inaccessible, has benefited from recent progress on pretrained vision-language (VL) models with strong generalizability. e.g. CLIP. Prompt learning methods built with VL models generate text features from the class names that only have confined class-specific information. Large Language Models (LLMs), with their vast en-cyclopedic knowledge, emerge as the complement. Thus, in this paper, we discuss the integration of LLMs to enhance pretrained VL models, specifically on low-shot classification. However, the domain gap between language and vision blocks the direct application of LLMs. Thus, we propose LLaMp, Large Language Models as Prompt learners, that produces adaptive prompts for the CLIP text encoder, establishing it as the connecting bridge. Experiments show that, compared with other state-of-the-art prompt learning methods, LLaMP yields better performance on both zero-shot generalization and few-shot image classification, over a spectrum of 11 datasets. Code will be made available at: https://github.com/zhaohengz/LLaMP.
Zhaoheng Zheng, Jingmin Wei, Xuefeng Hu, Haidong Zhu, Ramakant Nevatia
CVPR5
2024 SEAS: ShapE-Aligned Supervision for Person Re-Identification
abstract
We introduce SEAS, using ShapE-Aligned Supervision, to enhance appearance-based person re-identification. When recognizing an individual's identity, existing methods primarily rely on appearance, which can be influenced by the background environment due to a lack of body shape awareness. Although some methods attempt to incorporate other modalities, such as gait or body shape, they encode the additional modality separately, resulting in extra computational costs and lacking an inherent connection with appearance. In this paper, we explore the use of implicit 3-D body shape representations as pixel-level guidance to augment the extraction of identity features with body shape knowledge, in addition to appearance. Using body shape as supervision, rather than as input, provides shapeaware enhancements without any increase in computational cost and delivers coherent integration with pixel-wise appearance features. Moreover, for video-based person reidentification, we align pixel-level features across frames with shape awareness to ensure temporal consistency. Our results demonstrate that incorporating body shape as pixel-level supervision reduces rank-1 errors by 1.4% for framebased and by 2.5% for video-based re-identification tasks, respectively, and can also be generalized to other existing appearance-based person re-identification methods.
Haidong Zhu, Pranav Budhwant, Zhaoheng Zheng, Ramakant Nevatia
CVPR4
2024 CaesarNeRF: Calibrated Semantic Representation for Few-Shot Generalizable Neural Rendering
Haidong Zhu, Tianyu Ding, Ilya Zharkov, Ramakant Nevatia, Luming Liang
ECCV (6)5
2024 ReCLIP: Refine Contrastive Language Image Pre-Training with Source Free Domain Adaptation
abstract
Large-scale pre-trained vision-language models (VLM) such as CLIP [32] have demonstrated noteworthy zero-shot classification capability, achieving 76.3% top-1 accuracy on ImageNet without seeing any examples. However, while applying CLIP to a downstream target domain, the presence of visual and text domain gaps and cross-modality misalignment can greatly impact the model performance. To address such challenges, we propose ReCLIP, a novel source-free domain adaptation method for VLMs, which does not require any source data or target labeled data. ReCLIP first learns a projection space to mitigate the misaligned visual-text embeddings and learns pseudo labels. Then, it deploys cross-modality self-training with the pseudo labels to update visual and text encoders, refine labels and reduce domain gaps and misalignment iteratively. With extensive experiments, we show that ReCLIP outperforms all the baselines significantly and improves the average accuracy of CLIP from 69.83% to 74.94% on 22 image classification benchmarks.
Xuefeng Hu, Ke Zhang 0028, Albert Chen 0001, Jiajia Luo, Yuyin Sun, Ken Wang, Nan Qiao 0009, Min Sun 0001, Cheng-Hao Kuo, Ramakant Nevatia
WACV12
2024 Efficient Feature Distillation for Zero-shot Annotation Object Detection
abstract
We propose a new setting for detecting unseen objects called Zero-shot Annotation object Detection (ZAD). It expands the zero-shot object detection setting by allowing the novel objects to exist in the training images and restricts the additional information the detector uses to novel category names. Recently, to detect unseen objects, largescale vision-language models (e.g., CLIP) are leveraged by different methods. The distillation-based methods have good overall performance but suffer from a long training schedule caused by two factors. First, existing work creates distillation regions biased to the base categories, which limits the distillation of novel category information. Second, directly using the raw feature from CLIP for distillation neglects the domain gap between the training data of CLIP and the detection datasets, which makes it difficult to learn the mapping from the image region to the vision-language feature space. To solve these problems, we propose Efficient feature distillation for Zero-shot Annotation object Detection (EZAD). Firstly, EZAD adapts the CLIP’s feature space to the target detection domain by re-normalizing CLIP; Secondly, EZAD uses CLIP to generate distillation proposals with potential novel category names to avoid the distillation being overly biased toward the base categories. Finally, EZAD takes advantage of semantic meaning for regression to further improve the model performance. As a result, EZAD outperforms the previous distillation-based methods in COCO by 4% with a much shorter training schedule and achieves a 3% improvement on the LVIS dataset. Our code is available at https://github.com/dragonlzm/EZAD
Zhuoming Liu 0001, Xuefeng Hu, Ramakant Nevatia
WACV3
2024 Leveraging Task-Specific Pre-Training to Reason across Images and Videos
abstract
We explore the task of Reasoning Across Images and Video (RAIV), which requires models to reason on a pair of visual inputs comprising various combinations of images and/or videos. Previous work in this area has been limited to image pairs focusing primarily on the existence and/or cardinality of objects. To address this, we leverage existing datasets with rich annotations to generate semantically meaningful queries about actions, objects, and their relationships. We introduce new datasets that encompass visually similar inputs, reasoning over images, across images and videos, or across videos. Recognizing the distinct nature of RAIV compared to existing pre-training objectives which work on single image-text pairs, we explore task-specific pre-training, wherein a pre-trained model is trained on an objective similar to downstream tasks without utilizing fine-tuning datasets. Experiments with several state-of-the-art pre-trained image-language models reveal that task-specific pre-training significantly enhances performance on downstream datasets, even in the absence of additional pre-training data. We provide further ablative studies to guide future work.
Arka Sadhu, Ramakant Nevatia
WACV2
2024 CAILA: Concept-Aware Intra-Layer Adapters for Compositional Zero-Shot Learning
abstract
In this paper, we study the problem of Compositional Zero-Shot Learning (CZSL), which is to recognize novel attribute-object combinations with pre-existing concepts. Recent researchers focus on applying large-scale Vision-Language Pre-trained (VLP) models like CLIP with strong generalization ability. However, these methods treat the pre-trained model as a black box and focus on pre- and post-CLIP operations, which do not inherently mine the semantic concept between the layers inside CLIP. We propose to dive deep into the architecture and insert adapters, a parameter-efficient technique proven to be effective among large language models, into each CLIP encoder layer. We further equip adapters with concept awareness so that concept-specific features of "object", "attribute", and "composition" can be extracted. We assess our method on four popular CZSL datasets, MIT-States, C-GQA, UT-Zappos, and VAW-CZSL, which shows state-of-the-art performance compared to existing methods on all of them.
Zhaoheng Zheng, Haidong Zhu, Ramakant Nevatia
WACV3
2024 ShARc: Shape and Appearance Recognition for Person Identification In-the-wild
abstract
Identifying individuals in unconstrained video settings is a valuable yet challenging task in biometric analysis due to variations in appearances, environments, degradations, and occlusions. In this paper, we present ShARc, a multimodal approach for video-based person identification in uncontrolled environments that emphasizes 3-D body shape, pose, and appearance. We introduce two encoders: a Pose and Shape Encoder (PSE) and an Aggregated Appearance Encoder (AAE). PSE encodes the body shape via binarized silhouettes, skeleton motions, and 3-D body shape, while AAE provides two levels of temporal appearance feature aggregation: attention-based feature aggregation and averaging aggregation. For attention-based feature aggregation, we employ spatial and temporal attention to focus on key areas for person distinction. For averaging aggregation, we introduce a novel flattening layer after averaging to extract more distinguishable information and reduce overfitting of attention. We utilize centroid feature averaging for gallery registration. We demonstrate significant improvements over existing state-of-the-art methods on public datasets, including CCVID, MEVID, and BRIAR.
Haidong Zhu, Wanrong Zheng, Zhaoheng Zheng, Ramakant Nevatia
WACV4
2023 AG-ReID 2023: Aerial-Ground Person Re-identification Challenge Results
abstract
Person re-identification (Re-ID) on aerial-ground platforms has emerged as an intriguing topic within computer vision, presenting a plethora of unique challenges. Highflying altitudes of aerial cameras make persons appear differently in terms of viewpoints, poses, and resolution compared to the images of the same person viewed from ground cameras. Despite its potential, few algorithms have been developed for person re-identification on aerial-ground data, mainly due to the absence of comprehensive datasets. In response, we have collected a large-scale dataset and organized the Aerial-Ground person Re-IDentification Challenge (AG-ReID2023) to foster advancements in the field. The dataset comprises 100,502 images with 1,615 unique identities, including 51,530 training images featuring 807 identities. The test set is divided into two subsets: Aerial to Ground (808 ids, 4,348 query images, 19,259 gallery images) and Ground to Aerial (808 ids, 4,151 query images, 21,214 gallery images). In addition, we manually annotate individuals with their matching IDs across cameras and provide 15 soft attribute labels. The AG-ReID2023 Challenge in conjunction with the 7thIEEE International Joint Conference on Biometrics (IJCB) has garnered interest from numerous institutes, resulting in the submission of five distinct algorithms. We provide an in-depth examination of the evaluation outcomes and present our findings from the contest. For additional details, kindly refer to the official website1.1https://agreid23.github.io.
Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Dana Michalski, Debayan Deb, Mahak Kothari, Manisha Saini, Dawei Du, Scott McCloskey, Gabriel Bertocco, Fernanda A. Andaló, Terrance E. Boult, Anderson Rocha 0001, Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia, Zaigham A. Randhawa, Sinan Sabri, Gianfranco Doretto
IJCB20
2023 GaitRef: Gait Recognition with Refined Sequential Skeletons
abstract
Identifying humans with their walking sequences, known as gait recognition, is a useful biometric understanding task as it can be observed from a long distance and does not require cooperation from the subject. Two common modalities used for representing the walking sequence of a person are silhouettes and joint skeletons. Silhouette sequences, which record the boundary of the walking person in each frame, may suffer from the variant appearances from carried-on objects and clothes of the person. Framewise joint detections are noisy and introduce some jitters that are not consistent with sequential detections. In this paper, we combine the silhouettes and skeletons and refine the framewise joint predictions for gait recognition. With temporal information from the silhouette sequences. We show that the refined skeletons can improve gait recognition performance without extra annotations. We compare our methods on four public datasets, CASIA-B, OUMVLP, Gait3D and GREW, and show state-of-the-art performance.
Haidong Zhu, Wanrong Zheng, Zhaoheng Zheng, Ramakant Nevatia
IJCB4
2023 Multimodal Neural Radiance Field
abstract
This paper addresses the challenge of reconstructing a scene with a neural radiance field (NeRF) for robot vision and scene understanding using multiple modalities. Researchers have introduced the use of NeRF to represent an object for synthesizing and rendering novel views of complex scenes by optimizing a 3-D radiance field for ray casting and rendering for 2-D RGB images. However, using RGB images alone introduces additional geometry ambiguities with transparent objects or complex scenes and cannot accurately depict the 3-D shapes. We discuss and solve this problem and use multiple modalities as input for the same NeRF model to build a multimodal NeRF by incorporating point clouds and infrared image supervision to prevent such bias. In contrast to RGB images, infrared images and point clouds are typically taken by separate cameras that cannot be aligned with the RGB camera. We further introduce the alignment of different modalities based on point cloud registration to estimate the relative transformation matrices between them before training a NeRF model with multiple modalities. We evaluate our model on chosen scenes from the ScanNet and M2DGR datasets and demonstrate that it outperforms existing state-of-the-art methods.
Haidong Zhu, Yuyin Sun, Jiajia Luo, Nan Qiao 0009, Ramakant Nevatia, Cheng-Hao Kuo
ICRA7
2023 PatchZero: Defending against Adversarial Patch Attacks by Detecting and Zeroing the Patch
abstract
Adversarial patch attacks mislead neural networks by injecting adversarial pixels within a local region. Patch attacks can be highly effective in a variety of tasks and physically realizable via attachment (e.g. a sticker) to the real-world objects. Despite the diversity in attack patterns, adversarial patches tend to be highly textured and different in appearance from natural images. We exploit this property and present PatchZero, a general defense pipeline against white-box adversarial patches without retraining the downstream classifier or detector. Specifically, our defense detects adversaries at the pixel-level and "zeros out" the patch region by repainting with mean pixel values. We further design a two-stage adversarial training scheme to defend against the stronger adaptive attacks. PatchZero achieves SOTA defense performance on the image classification (ImageNet, RESISC45), object detection (PASCAL VOC), and video classification (UCF101) tasks with little degradation in benign performance. In addition, PatchZero transfers to different patch shapes and attack types.
Zhaoheng Zheng, Kaijie Cai, Ramakant Nevatia
WACV5
2023 Gait Recognition Using 3-D Human Body Shape Inference
abstract
Gait recognition, which identifies individuals based on their walking patterns, is an important biometric technique since it can be observed from a distance and does not require the subject’s cooperation. Recognizing a person’s gait is difficult because of the appearance variants in human silhouette sequences produced by varying viewing angles, carrying objects, and clothing. Recent research has produced a number of ways for coping with these variants. In this paper, we present the usage of inferring 3-D body shapes distilled from limited images, which are, in principle, invariant to the specified variants. Inference of 3-D shape is a difficult task, especially when only silhouettes are provided in a dataset. We provide a method for learning 3-D body inference from silhouettes by transferring knowledge from 3-D shape prior from RGB photos. We use our method on multiple existing state-of-the-art gait baselines and obtain consistent improvements for gait identification on two public datasets, CASIA-B and OUMVLP, on several variants and settings, including a new setting of novel views not seen during training.
Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia
WACV3
2022 Self-Supervised Learning for Sentiment Analysis via Image-Text Matching
abstract
There is often a resemblance in the sentiment expressed in social media posts (text) and their accompanying images. In this paper, We leverage this sentiment congruence for self-supervised representation learning for sentiment analysis. By teaching the model to pair an image with its corresponding social media post, the model can learn a representation capturing sentiment features from the image and text without supervision. We then use the pre-trained encoder for feature extraction for sentiment analysis in downstream tasks. We show significant improvement and good transferability for sentiment classification in addition to robustness in performance when available data decreases on public datasets (B-T4SA and IMDb Movie Review). With this work, we demonstrate the effectiveness of self-supervised learning through cross-modal matching for sentiment analysis.
Haidong Zhu, Zhaoheng Zheng, Mohammad Soleymani 0001, Ramakant Nevatia
ICASSP4
2022 Improving Weakly Supervised Scene Graph Parsing through Object Grounding
abstract
Weakly supervised scene graph parsing, which learns structured image representations without annotated correspondences between graph nodes and visual objects, has been prevalent in recent computer vision research. Existing methods mainly focus on designing task-specific loss functions, model architectures, or optimization algorithms. We argue that correspondences between objects and graph nodes are crucial for the weakly supervised scene graph parsing task and are worth learning explicitly. Thus we propose GroParser, a framework that improves weakly supervised scene graph parsing models by grounding visual objects. The proposed weakly supervised grounding method learns a metric among visual objects and scene graph nodes by incorporating information from both object features and relational features. Specifically, we apply multi-instance learning to learn the object category information and exploit a two-stream graph neural network to model the relational similarity metric. Extensive experiments on the scene graph parsing task verify the grounding found by our model can reinforce the performance of the existing weakly supervised scene graph parsing methods, including the current state-of-the-art. Further experiments on Visual Genome (VG) and Visual Relation Detection (VRD) datasets verify that our model brings an improvement on scene graph grounding task over existing approaches.
Zhaoheng Zheng, Ramakant Nevatia, Yan Liu 0002
ICPR3
2022 OPEN: Order-preserving Pointcloud Encoder Decoder Network for Body Shape Refinement
abstract
Image-based 3-D human body shape estimation and reconstruction have shown significant improvement by using deep neural networks. Compared with reconstructing from a single image, reconstructing 3-D human body shapes from video or image sequences requires high precision and dense correspondences between the keypoints of the reconstructed shape sequence. Existing methods cannot achieve both high accuracy and keep the dense correspondence between different shapes after reconstruction. In this paper, we propose a method named Order-preserving Point cloud Encoder-decoder Network to refine the reconstructed human body shape from SMPL with the assistance of RGB images while preserving its original dense correspondence. We further introduce using 2-D RGB images as weak supervision when 3-D labels are not available. We assess our methods on the public dataset and show improved results compared with the baseline methods.
Haidong Zhu, Ramakant Nevatia
ICPR5
2022 Temporal Shift and Attention Modules for Graphical Skeleton Action Recognition
abstract
Skeletons, consisting of joint positions and connections between them, are an important representation for modeling human bodies in image frames. Compared with understanding RGB videos, recognizing actions from the skeletons removes the biases of background and body shapes. Researchers use spatial-temporal graphs to model the skeleton sequences. These methods weigh all frames in the sequence equally even though many of the frames may not be useful for action and prediction and dilute the influence of important frames. Also, the temporal graph focuses on understanding only the low-level feature of the joints for the motion of the skeleton. In this paper, we introduce two modules, temporal shift module and temporal attention module that can be added to graph convolution networks for skeleton action recognition. Temporal attention module focuses on keyframes for making predictions, and temporal shift module helps to exchange the high-level features between different frames along the temporal dimension besides the local patterns. We evaluate the two modules with two existing skeleton action recognition networks, ST-GCN and MS-G3D, on three public datasets and show better results than the original methods.
Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia
ICPR3
2021 SimPLE: Similar Pseudo Label Exploitation for Semi-Supervised Classification
abstract
A common classification task situation is where one has a large amount of data available for training, but only a small portion is annotated with class labels. The goal of semi-supervised training, in this context, is to improve classification accuracy by leverage information not only from labeled data but also from a large amount of unlabeled data. Recent works have developed significant improvements by exploring the consistency constrain between differently augmented labeled and unlabeled data. Following this path, we propose a novel unsupervised objective that focuses on the less studied relationship between the high confidence unlabeled data that are similar to each other. The new proposed Pair Loss minimizes the statistical distance between high confidence pseudo labels with similarity above a certain threshold. Combining the Pair Loss with the techniques developed by the MixMatch family, our proposed SimPLE algorithm shows significant performance gains over previous algorithms on CIFAR-100 and Mini-ImageNet, and is on par with the state-of-the-art methods on CIFAR-10 and SVHN. Furthermore, SimPLE also outperforms the state-of-the-art methods in the transfer learning setting, where models are initialized by the weights pre-trained on ImageNet or DomainNet-Real. The code is available at github.com/zijian-hu/SimPLE.
Zijian Hu 0001, Zhengyu Yang 0003, Xuefeng Hu, Ramakant Nevatia
CVPR4
2021 Visual Semantic Role Labeling for Video Understanding
abstract
We propose a new framework for understanding and representing related salient events in a video using visual semantic role labeling. We represent videos as a set of related events, wherein each event consists of a verb and multiple entities that fulfill various roles relevant to that event. To study the challenging task of semantic role labeling in videos or VidSRL, we introduce the VidSitu benchmark, a large scale video understanding data source with 29K 10-second movie clips richly annotated with a verb and semantic-roles every 2 seconds. Entities are co-referenced across events within a movie clip and events are connected to each other via event-event relations. Clips in VidSitu are drawn from a large collection of movies (∼3K) and have been chosen to be both complex (∼4.2 unique verbs within a video) as well as diverse (∼200 verbs have more than 100 annotations each). We provide a comprehensive analysis of the dataset in comparison to other publicly available video understanding benchmarks, several illustrative baselines and evaluate a range of standard video recognition models. Our code and dataset is available at vidsitu.org.
Arka Sadhu, Tanmay Gupta, Mark Yatskar, Ramakant Nevatia, Aniruddha Kembhavi
CVPR4
2021 Improving Object Detection And Attribute Recognition By Feature Entanglement Reduction
abstract
We explore object detection with two attributes: color and material. The task aims to simultaneously detect objects and infer their color and material. A straight-forward approach is to add attribute heads at the very end of a usual object detection pipeline. However, we observe that the two goals are in conflict: Object detection should be attribute-independent and attributes be largely object-independent. Features computed by a standard detection network entangle the category and attribute features; we disentangle them by the use of a two-stream model where the category and attribute features are computed independently but the classification heads share Regions of Interest (RoIs). Compared with a traditional single-stream model, our model shows significant improvements over VG-20, a subset of Visual Genome, on both supervised and attribute transfer tasks.
Zhaoheng Zheng, Arka Sadhu, Ramakant Nevatia
ICIP3
2021 Video Question Answering with Phrases via Semantic Roles
abstract
Video Question Answering (VidQA) evaluation metrics have been limited to a single-word answer or selecting a phrase from a fixed set of phrases.These metrics limit the VidQA models' application scenario.In this work, we leverage semantic roles derived from video descriptions to mask out certain phrases, to introduce VidQAP which poses VidQA as a fillin-the-phrase task.To enable evaluation of answer phrases, we compute the relative improvement of the predicted answer compared to an empty string.To reduce the influence of language-bias in VidQA datasets, we retrieve a video having a different answer for the same question.To facilitate research, we construct ActivityNet-SRL-QA and Charades-SRL-QA and benchmark them by extending three vision-language models.We perform extensive analysis and ablative studies to guide future work.Code and data are public.Video description: A man on top of a building throws a bowling ball towards the pins Q4: throws a bowling ball towards the pins.Model's generated answer: A man standing on a house Correct answer: A man on top of a building Q5: A man on top of a building a bowling ball towards the pins.Model's generated answer: throws Correct answer: throws Q6: A man on top of a building throws towards the pins.Model's generated answer: a ball Correct answer: a bowling ball Q7: A man on top of a building throws a bowling ball Model's generated answer: towards some bottles Correct answer: towards the pins (b) Free-form Answer Generation ARG0
Arka Sadhu, Ramakant Nevatia
NAACL-HLT3
2021 Utilizing Every Image Object for Semi-supervised Phrase Grounding
abstract
Phrase grounding models localize an object in the image given a referring expression. The annotated language queries available during training are limited, which also limits the variations of language combinations that a model can see during training. In this paper, we study the case applying objects without labeled queries for training the semi-supervised phrase grounding. We propose to use learned location and subject embedding predictors (LSEP) to generate the corresponding language embeddings for objects lacking annotated queries in the training set. With the assistance of the detector, we also apply LSEP to train a grounding model on images without any annotation. We evaluate our method based on MAttNet on three public datasets: RefCOCO, RefCOCO+, and RefCOCOg. We show that our predictors allow the grounding system to learn from the objects without labeled queries and improve accuracy by 34.9% relatively with the detection results.
Haidong Zhu, Arka Sadhu, Zhaoheng Zheng, Ramakant Nevatia
WACV4
2020 Video Object Grounding Using Semantic Roles in Language Description
abstract
We explore the task of Video Object Grounding (VOG), which grounds objects in videos referred to in natural language descriptions. Previous methods apply image grounding based algorithms to address VOG, fail to explore the object relation information and suffer from limited generalization. Here, we investigate the role of object relations in VOG and propose a novel framework VOGNet to encode multi-modal object relations via self-attention with relative position encoding. To evaluate VOGNet, we propose novel contrasting sampling methods to generate more challenging grounding input samples, and construct a new dataset called ActivityNet-SRL (ASRL) based on existing caption and grounding datasets. Experiments on ASRL validate the need of encoding object relations in VOG, and our VOGNet outperforms competitive baselines by a significant margin.
Arka Sadhu, Ramakant Nevatia
CVPR3
2020 Curriculum DeepSDF
Yueqi Duan, Haidong Zhu, He Wang 0010, Li Yi 0001, Ramakant Nevatia, Leonidas J. Guibas
ECCV (8)5
2020 SPAN: Spatial Pyramid Attention Network for Image Manipulation Localization
Xuefeng Hu, Zhenye Jiang, Syomantak Chaudhuri, Zhenheng Yang, Ramakant Nevatia
ECCV (21)6
2020 Visually Grounded Continual Learning of Compositional Phrases
abstract
Humans acquire language continually with much more limited access to data samples at a time, as compared to contemporary NLP systems.To study this human-like language acquisition ability, we present VisCOLL, a visually grounded language learning task, which simulates the continual acquisition of compositional phrases from streaming visual scenes.In the task, models are trained on a paired image-caption stream which has shifting object distribution; while being constantly evaluated by a visually-grounded masked language prediction task on held-out test sets.VisCOLL compounds the challenges of continual learning (i.e., learning from continuously shifting data distribution) and compositional generalization (i.e., generalizing to novel compositions).To facilitate research on VisCOLL, we construct two datasets, COCO-shift and Flickrshift, and benchmark them using different continual learning methods.Results reveal that SoTA continual learning approaches provide little to no improvements on VisCOLL, since storing examples of all possible compositions is infeasible.We conduct further ablations and analysis to guide future work 1 .
Xisen Jin, Junyi Du, Arka Sadhu, Ramakant Nevatia, Xiang Ren 0001
EMNLP (1)4
2020 Every Pixel Counts ++: Joint Learning of Geometry and Motion with 3D Holistic Understanding
abstract
Learning to estimate 3D geometry in a single frame and optical flow from consecutive frames by watching unlabeled videos via deep convolutional network has made significant progress recently. Current state-of-the-art (SoTA) methods treat the two tasks independently. One typical assumption of the existing depth estimation methods is that the scenes contain no independent moving objects. while object moving could be easily modeled using optical flow. In this paper, we propose to address the two tasks as a whole, i.e., to jointly understand per-pixel 3D geometry and motion. This eliminates the need of static scene assumption and enforces the inherent geometrical consistency during the learning process, yielding significantly improved results for both tasks. We call our method as “Every Pixel Counts++” or “EPC++”. Specifically, during training, given two consecutive frames from a video, we adopt three parallel networks to predict the camera motion (MotionNet), dense depth map (DepthNet), and per-pixel optical flow between two frames (OptFlowNet) respectively. The three types of information, are fed into a holistic 3D motion parser (HMP), and per-pixel 3D motion of both rigid background and moving objects are disentangled and recovered. Various loss terms are formulated to jointly supervise the three networks. An effective adaptive training strategy is proposed to achieve better performance and more efficient convergence. Comprehensive experiments were conducted on datasets with different scenes, including driving scenario (KITTI 2012 and KITTI 2015 datasets), mixed outdoor/indoor scenes (Make3D) and synthetic animation (MPI Sintel dataset). Performance on the five tasks of depth estimation, optical flow estimation, odometry, moving object segmentation and scene flow estimation shows that our approach outperforms other SoTA methods, demonstrating the effectiveness of each module of our proposed method. Code will be available at: https://github.com/chenxuluo/EPC.
Chenxu Luo, Zhenheng Yang, Peng Wang 0001, Yang Wang 0046, Wei Xu 0017, Ramakant Nevatia, Alan L. Yuille
IEEE Trans. Pattern Anal. Mach. Intell.6
2019 Activity Driven Weakly Supervised Object Detection
abstract
Weakly supervised object detection aims at reducing the amount of supervision required to train detection models. Such models are traditionally learned from images/videos labelled only with the object class and not the object bounding box. In our work, we try to leverage not only the object class labels but also the action labels associated with the data. We show that the action depicted in the image/video can provide strong cues about the location of the associated object. We learn a spatial prior for the object dependent on the action (e.g. "ball" is closer to "leg of the person" in "kicking ball"), and incorporate this prior to simultaneously train a joint object detection and action classification model. We conducted experiments on both video datasets and image datasets to evaluate the performance of our weakly supervised object detection model. Our approach outperformed the current state-of-the-art (SOTA) method by more than 6% in mAP on the Charades video dataset.
Zhenheng Yang, Dhruv Mahajan 0001, Deepti Ghadiyaram, Ramakant Nevatia, Vignesh Ramanathan
CVPR4
2019 NOTE-RCNN: NOise Tolerant Ensemble RCNN for Semi-Supervised Object Detection
abstract
The labeling cost of large number of bounding boxes is one of the main challenges for training modern object detectors. To reduce the dependence on expensive bounding box annotations, we propose a new semi-supervised object detection formulation, in which a few seed box level annotations and a large scale of image level annotations are used to train the detector. We adopt a training-mining framework, which is widely used in weakly supervised object detection tasks. However, the mining process inherently introduces various kinds of labelling noises: false negatives, false positives and inaccurate boundaries, which can be harmful for training the standard object detectors (e.g. Faster RCNN). We propose a novel NOise Tolerant Ensemble RCNN (NOTE-RCNN) object detector to handle such noisy labels. Comparing to standard Faster RCNN, it contains three highlights: an ensemble of two classification heads and a distillation head to avoid overfitting on noisy labels and improve the mining precision, masking the negative sample loss in box predictor to avoid the harm of false negative labels, and training box regression head only on seed annotations to eliminate the harm from inaccurate boundaries of mined bounding boxes. We evaluate the methods on ILSVRC 2013 and MSCOCO 2017 dataset; we observe that the detection accuracy consistently improves as we iterate between mining and training steps, and state-of-the-art performance is achieved.
Jiyang Gao, Jiang Wang 0001, Shengyang Dai, Li-Jia Li 0001, Ramakant Nevatia
ICCV5
2019 Zero-Shot Grounding of Objects From Natural Language Queries
abstract
A phrase grounding system localizes a particular object in an image referred to by a natural language query. In previous work, the phrases were restricted to have nouns that were encountered in training, we extend the task to Zero-Shot Grounding(ZSG) which can include novel, “unseen” nouns. Current phrase grounding systems use an explicit object detection network in a 2-stage framework where one stage generates sparse proposals and the other stage evaluates them. In the ZSG setting, generating appropriate proposals itself becomes an obstacle as the proposal generator is trained on the entities common in the detection and grounding datasets. We propose a new single-stage model called ZSGNet which combines the detector network and the grounding system and predicts classification scores and regression parameters. Evaluation of ZSG system brings additional subtleties due to the influence of the relationship between the query and learned categories; we define four distinct conditions that incorporate different levels of difficulty. We also introduce new datasets, sub-sampled from Flickr30k Entities and Visual Genome, that enable evaluations for the four conditions. Our experiments show that ZSGNet achieves state-of-the-art performance on Flickr30k and ReferIt under the usual “seen” settings and performs significantly better than baseline in the zero-shot setting.
Arka Sadhu, Ramakant Nevatia
ICCV3
2019 MAC: Mining Activity Concepts for Language-Based Temporal Localization
abstract
We address the problem of language-based temporal localization in untrimmed videos. Compared to temporal localization with fixed categories, this problem is more challenging as the language-based queries not only have no pre-defined activity list but also may contain complex descriptions. Previous methods address the problem by considering features from video sliding windows and language queries and learning a subspace to encode their correlation, which ignore rich semantic cues about activities in videos and queries. We propose to mine activity concepts from both video and language modalities by applying the actionness score enhanced Activity Concepts based Localizer (ACL). Specifically, the novel ACL encodes the semantic concepts from verb-obj pairs in language queries and leverages activity classifiers' prediction scores to encode visual concepts. Besides, ACL also has the capability to regress sliding windows as localization results. Experiments show that ACL significantly outperforms state-of-the-arts under the widely used metric, with more than 5% increase on both Charades-STA and TACoS datasets.
Runzhou Ge, Jiyang Gao, Ramakant Nevatia
WACV4
2019 Deep, Landmark-Free FAME: Face Alignment, Modeling, and Expression Estimation
Feng-Ju Chang, Anh Tuan Tran 0001, Tal Hassner, Iacopo Masi, Ramakant Nevatia, Gérard G. Medioni
Int. J. Comput. Vis.5
2019 Learning Pose-Aware Models for Pose-Invariant Face Recognition in the Wild
abstract
We propose a method designed to push the frontiers of unconstrained face recognition in the wild with an emphasis on extreme out-of-plane pose variations. Existing methods either expect a single model to learn pose invariance by training on massive amounts of data or else normalize images by aligning faces to a single frontal pose. Contrary to these, our method is designed to explicitly tackle pose variations. Our proposed Pose-Aware Models (PAM) process a face image using several pose-specific, deep convolutional neural networks (CNN). 3D rendering is used to synthesize multiple face poses from input images to both train these models and to provide additional robustness to pose variations at test time. Our paper presents an extensive analysis of the IARPA Janus Benchmark A (IJB-A), evaluating the effects that landmark detection accuracy, CNN layer selection, and pose model selection all have on the performance of the recognition pipeline. It further provides comparative evaluations on IJB-A and the PIPA dataset. These tests show that our approach outperforms existing methods, even surprisingly matching the accuracy of methods that were specifically fine-tuned to the target dataset. Parts of this work previously appeared in [1] and [2].
Iacopo Masi, Feng-Ju Chang, Jongmoo Choi, Shai Harel, Jungyeon Kim, KangGeon Kim, Jatuporn Toy Leksut, Stephen Rawls, Yue Wu 0001, Tal Hassner, Wael Abd-Almageed, Gérard G. Medioni, Louis-Philippe Morency, Premkumar Natarajan, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.15
2018 SPOT Poachers in Action: Augmenting Conservation Drones With Automatic Detection in Near Real Time
abstract
The unrelenting threat of poaching has led to increased development of new technologies to combat it. One such example is the use of long wave thermal infrared cameras mounted on unmanned aerial vehicles (UAVs or drones) to spot poachers at night and report them to park rangers before they are able to harm animals. However, monitoring the live video stream from these conservation UAVs all night is an arduous task. Therefore, we build SPOT (Systematic POacher deTector), a novel application that augments conservation drones with the ability to automatically detect poachers and animals in near real time. SPOT illustrates the feasibility of building upon state-of-the-art AI techniques, such as Faster RCNN, to address the challenges of automatically detecting animals and poachers in infrared images. This paper reports (i) the design and architecture of SPOT, (ii) a series of efforts towards more robust and faster processing to make SPOT usable in the field and provide detections in near real time, and (iii) evaluation of SPOT based on both historical videos and a real-world test run by the end users in the field. The promising results from the test in the field have led to a plan for larger-scale deployment in a national park in Botswana. While SPOT is developed for conservation drones, its design and novel techniques have wider application for automated detection from UAV videos.
Elizabeth Bondi-Kelly, Fei Fang 0001, Mark Hamilton, Debarun Kar, Donnabell Dmello, Jongmoo Choi, Robert Hannaford, Arvind Iyer, Lucas Joppa, Milind Tambe, Ramakant Nevatia
AAAI11
2018 Unsupervised Learning of Geometry From Videos With Edge-Aware Depth-Normal Consistency
abstract
Learning to reconstruct depths from a single image by watching unlabeled videos via deep convolutional network (DCN) is attracting significant attention in recent years, e.g. (Zhou et al. 2017). In this paper, we propose to use surface normal representation for unsupervised depth estimation framework. Our estimated depths are constrained to be compatible with predicted normals, yielding more robust geometry results. Specifically, we formulate an edge-aware depth-normal consistency term, and solve it by constructing a depth-to-normal layer and a normal-to-depth layer inside of the DCN. The depth-to-normal layer takes estimated depths as input, and computes normal directions using cross production based on neighboring pixels. Then given the estimated normals, the normal-to-depth layer outputs a regularized depth map through local planar smoothness. Both layers are computed with awareness of edges inside the image to help address the issue of depth/normal discontinuity and preserve sharp edges. Finally, to train the network, we apply the photometric error and gradient smoothness to supervise both depth and normal predictions. We conducted experiments on both outdoor (KITTI) and indoor (NYUv2) datasets, and showed that our algorithm vastly outperforms state-of-the-art, which demonstrates the benefits of our approach.
Zhenheng Yang, Peng Wang 0001, Wei Xu 0017, Liang Zhao 0006, Ramakant Nevatia
AAAI5
2018 PIRC Net: Using Proposal Indexing, Relationships and Context for Phrase Grounding
Rama Kovvuri, Ramakant Nevatia
ACCV (4)2
2018 Knowledge Aided Consistency for Weakly Supervised Phrase Grounding
abstract
Given a natural language query, a phrase grounding system aims to localize mentioned objects in an image. In weakly supevised scenario, mapping between image regions (i.e., proposals) and language is not available in the training set. Previous methods address this deficiency by training a grounding system via learning to reconstruct language information contained in input queries from predicted proposals. However, the optimization is solely guided by the reconstruction loss from the language modality, and ignores rich visual information contained in proposals and useful cues from external knowledge. In this paper, we explore the consistency contained in both visual and language modalities, and leverage complementary external knowledge to facilitate weakly supervised grounding. We propose a novel Knowledge Aided Consistency Network (KAC Net) which is optimized by reconstructing input query and proposal's information. To leverage complementary knowledge contained in the visual features, we introduce a Knowledge Based Pooling (KBP) gate to focus on query-related proposals. Experiments show that KAC Net provides a significant improvement on two popular datasets.
Jiyang Gao, Ramakant Nevatia
CVPR3
2018 Motion-Appearance Co-Memory Networks for Video Question Answering
abstract
Video Question Answering (QA) is an important task in understanding video temporal structure. We observe that there are three unique attributes of video QA compared with image QA: (1) it deals with long sequences of images containing richer information not only in quantity but also in variety; (2) motion and appearance information are usually correlated with each other and able to provide useful attention cues to the other; (3) different questions require different number of frames to infer the answer. Based on these observations, we propose a motion-appearance co-memory network for video QA. Our networks are built on concepts from Dynamic Memory Network (DMN) and introduces new mechanisms for video QA. Specifically, there are three salient aspects: (1) a co-memory attention mechanism that utilizes cues from both motion and appearance to generate attention; (2) a temporal conv-deconv network to generate multi-level contextual facts; (3) a dynamic fact ensemble method to construct temporal representation dynamically for different questions. We evaluate our method on TGIF-QA dataset, and the results outperform state-of-the-art significantly on all four tasks of TGIF-QA.
Jiyang Gao, Runzhou Ge, Ramakant Nevatia
CVPR4
2018 LEGO: Learning Edge With Geometry All at Once by Watching Videos
abstract
Learning to estimate 3D geometry in a single image by watching unlabeled videos via deep convolutional network is attracting significant attention. In this paper, we introduce a "3D as-smooth-as-possible (3D-ASAP)" prior inside the pipeline, which enables joint estimation of edges and 3D scene, yielding results with significant improvement in accuracy for fine detailed structures. Specifically, we define the 3D-ASAP prior by requiring that any two points recovered in 3D from an image should lie on an existing planar surface if no other cues provided. We design an unsupervised framework that Learns Edges and Geometry (depth, normal) all at Once (LEGO). The predicted edges are embedded into depth and surface normal smoothness terms, where pixels without edges in-between are constrained to satisfy the prior. In our framework, the predicted depths, normals and edges are forced to be consistent all the time. We conduct experiments on KITTI to evaluate our estimated geometry and CityScapes to perform edge evaluation. We show that in all of the tasks, i.e. depth, normal and edge, our algorithm vastly outperforms other state-of-the-art (SOTA) algorithms, demonstrating the benefits of our approach.
Zhenheng Yang, Peng Wang 0001, Yang Wang 0046, Wei Xu 0017, Ramakant Nevatia
CVPR5
2018 CTAP: Complementary Temporal Action Proposal Generation
Jiyang Gao, Ramakant Nevatia
ECCV (2)3
2018 ExpNet: Landmark-Free, Deep, 3D Facial Expressions
abstract
We describe a deep learning based method for estimating 3D facial expression coefficients. Unlike previous work, our process does not relay on facial landmark detection methods as a proxy step. Recent methods have shown that a CNN can be trained to regress accurate and discriminative 3D morphable model (3DMM) representations, directly from image intensities. By foregoing landmark detection, these methods were able to estimate shapes for occluded faces appearing in unprecedented viewing conditions. We build on those methods by showing that facial expressions can also be estimated by a robust, deep, landmark-free approach. Our ExpNet CNN is applied directly to the intensities of a face image and regresses a 29D vector of 3D expression coefficients. We propose a unique method for collecting data to train our network, leveraging on the robustness of deep networks to training label noise. We further offer a novel means of evaluating the accuracy of estimated expression coefficients: by measuring how well they capture facial emotions on the CK+ and EmotiW-17 emotion recognition benchmarks. We show that our ExpNet produces expression coefficients which better discriminate between facial emotions than those obtained using state of the art, facial landmark detectors. Moreover, this advantage grows as image scales drop, demonstrating that our ExpNet is more robust to scale changes than landmark detectors. Finally, our ExpNet is orders of magnitude faster than its alternatives.
Feng-Ju Chang, Anh Tuan Tran 0001, Tal Hassner, Iacopo Masi, Ramakant Nevatia, Gérard G. Medioni
FG5
2018 Face and Body Association for Video-Based Face Recognition
abstract
In recent years face recognition has made extraordinary leaps, yet unconstrained video-based face identification in the wild remains an open and interesting problem. Videos, unlike still-images, offer a myriad of data for face modeling, sampling, and recognition, but, on the other hand, contain low-quality frames and motion blur. A key component in video-based face recognition is the way in which faces are associated through the video sequence before being used for recognition. In this paper, we present a video-based face recognition method taking advantage of face and body association (FBA). To track and associate subjects that appear across frames in multiple shots, we solve a data association problem using both face and body appearance. The final recovered track is then used to build a face representation for recognition. We evaluate our FBA method for video-based face recognition on a challenging dataset. Our experiments show up to 5% improvement in the identification rate over the state-of-the-art.
KangGeon Kim, Zhenheng Yang, Iacopo Masi, Ramakant Nevatia, Gérard G. Medioni
WACV4
2017 DECK: Discovering Event Composition Knowledge from Web Images for Zero-Shot Event Detection and Recounting in Videos
abstract
We address the problem of zero-shot event recognition in consumer videos. An event usually consists of multiple human-human and human-object interactions over a relative long period of time. A common approach proceeds by representing videos with banks of object and action concepts, but requires additional user inputs to specify the desired concepts per event. In this paper, we provide a fully automatic algorithm to select representative and reliable concepts for event queries. This is achieved by discovering event composition knowledge (DECK) from web images. To evaluate our proposed method, we use the standard zero-shot event detection protocol (ZeroMED), but also introduce a novel zero-shot event recounting (ZeroMER) problem to select supporting evidence of the events. Our ZeroMER formulation aims to select video snippets that are relevant and diverse. Evaluation on the challenging TRECVID MED dataset show that our proposed method achieves promising results on both tasks.
Chuang Gan 0001, Chen Sun 0002, Ramakant Nevatia
AAAI3
2017 Cascaded Boundary Regression for Temporal Action Detection
Jiyang Gao, Zhenheng Yang, Ramakant Nevatia
BMVC3
2017 RED: Reinforced Encoder-Decoder Networks for Action Anticipation
Jiyang Gao, Zhenheng Yang, Ramakant Nevatia
BMVC3
2017 Spatio-Temporal Action Detection with Cascade Proposal and Location Anticipation
Zhenheng Yang, Jiyang Gao, Ramakant Nevatia
BMVC3
2017 AMC: Attention Guided Multi-modal Correlation Learning for Image Search
abstract
Given a users query, traditional image search systems rank images according to its relevance to a single modality (e.g., image content or surrounding text). Nowadays, an increasing number of images on the Internet are available with associated meta data in rich modalities (e.g., titles, keywords, tags, etc.), which can be exploited for better similarity measure with queries. In this paper, we leverage visual and textual modalities for image search by learning their correlation with input query. According to the intent of query, attention mechanism can be introduced to adaptively balance the importance of different modalities. We propose a novel Attention guided Multi-modal Correlation (AMC) learning method which consists of a jointly learned hierarchy of intra and inter-attention networks. Conditioned on querys intent, intra-attention networks (i.e., visual intra-attention network and language intra-attention network) attend on informative parts within each modality, a multi-modal inter-attention network promotes the importance of the most query-relevant modalities. In experiments, we evaluate AMC models on the search logs from two real world image search engines and show a significant boost on the ranking of user-clicked images in search results. Additionally, we extend AMC models to caption ranking task on COCO dataset and achieve competitive results compared with recent state-of-the-arts.
Trung Bui, Ramakant Nevatia
CVPR5
2017 Local-Global Landmark Confidences for Face Recognition
abstract
A key to successful face recognition is accurate and reliable face alignment using automatically-detected facial landmarks. Given this strong dependency between face recognition and facial landmark detection, robust face recognition requires knowledge of when the facial landmark detection algorithm succeeds and when it fails. Facial landmark confidence represents this measure of success. In this paper, we propose two methods to measure landmark detection confidence: local confidence based on local predictors of each facial landmark, and global confidence based on a 3D rendered face model. A score fusion approach is also introduced to integrate these two confidences effectively. We evaluate both confidence metrics on two datasets for face recognition: JANUS CS2 and IJB-A datasets. Our experiments show up to 9% improvements when face recognition algorithm integrates the local-global confidence metrics.
KangGeon Kim, Feng-Ju Chang, Jongmoo Choi, Louis-Philippe Morency, Ramakant Nevatia, Gérard G. Medioni
FG5
2017 Query-Guided Regression Network with Context Policy for Phrase Grounding
abstract
Given a textual description of an image, phrase grounding localizes objects in the image referred by query phrases in the description. State-of-the-art methods address the problem by ranking a set of proposals based on the relevance to each query, which are limited by the performance of independent proposal generation systems and ignore useful cues from context in the description. In this paper, we adopt a spatial regression method to break the performance limit, and introduce reinforcement learning techniques to further leverage semantic context information. We propose a novel Query-guided Regression network with Context policy (QRC Net) which jointly learns a Proposal Generation Network (PGN), a Query-guided Regression Network (QRN) and a Context Policy Network (CPN). Experiments show QRC Net provides a significant improvement in accuracy on two popular datasets: Flickr30K Entities and Referit Game, with 14.25% and 17.14% increase over the state-of-the-arts respectively.
Rama Kovvuri, Ramakant Nevatia
ICCV3
2017 TALL: Temporal Activity Localization via Language Query
abstract
This paper focuses on temporal localization of actions in untrimmed videos. Existing methods typically train classifiers for a pre-defined list of actions and apply them in a sliding window fashion. However, activities in the wild consist of a wide combination of actors, actions and objects; it is difficult to design a proper activity list that meets users' needs. We propose to localize activities by natural language queries. Temporal Activity Localization via Language (TALL) is challenging as it requires: (1) suitable design of text and video representations to allow cross-modal matching of actions and language queries; (2) ability to locate actions accurately given features from sliding windows of limited granularity. We propose a novel Cross-modal Temporal Regression Localizer (CTRL) to jointly model text query and video clips, output alignment scores and action boundary regression results for candidate clips. Lor evaluation, we adopt TaCoS dataset, and build a new dataset for this task on top of Charades by adding sentence temporal annotations, called Charades-STA. We also build complex sentence queries in Charades-STA for test. Experimental results show that CTRL outperforms previous methods significantly on both datasets.
Jiyang Gao, Chen Sun 0002, Zhenheng Yang, Ramakant Nevatia
ICCV4
2017 TURN TAP: Temporal Unit Regression Network for Temporal Action Proposals
abstract
Temporal Action Proposal (TAP) generation is an important problem, as fast and accurate extraction of semantically important (e.g. human actions) segments from untrimmed videos is an important step for large-scale video analysis. We propose a novel Temporal Unit Regression Network (TURN) model. There are two salient aspects of TURN: (1) TURN jointly predicts action proposals and refines the temporal boundaries by temporal coordinate regression: (2) Fast computation is enabled by unit feature reuse: a long untrimmed video is decomposed into video units, which are reused as basic building blocks of temporal proposals. TURN outperforms the previous state-of-the-art methods under average recall (AR) by a large margin on THUMOS-14 and ActivityNet datasets, and runs at over 880 frames per second (FPS) on a TITAN X GPU. We further apply TURN as a proposal generation stage for existing temporal action localization pipelines, it outperforms state-of-the-art performance on THUMOS-14 and ActivityNet.
Jiyang Gao, Zhenheng Yang, Chen Sun 0002, Ramakant Nevatia
ICCV5
2017 MSRC: Multimodal Spatial Regression with Semantic Context for Phrase Grounding
abstract
Given an image and a natural language query phrase, a grounding system localizes the mentioned objects in the image according to the query's specifications. State-of-the-art methods address the problem by ranking a set of proposal bounding boxes according to the query's semantics, which makes them dependent on the performance of proposal generation systems. Besides, query phrases in one sentence may be semantically related in one sentence and can provide useful cues to ground objects. We propose a novel Multimodal Spatial Regression with semantic Context (MSRC) system which not only predicts the location of ground truth based on proposal bounding boxes, but also refines prediction results by penalizing similarities of different queries coming from same sentences. The advantages of MSRC are twofold: first, it removes the limitation of performance from proposal generation algorithms by using a spatial regression network. Second, MSRC not only encodes the semantics of a query phrase, but also deals with its relation with other queries in the same sentence (i.e., context) by a context refinement network. Experiments show MSRC system provides a significant improvement in accuracy on two popular datasets: Flickr30K Entities and Refer-it Game, with 6.64% and 5.28% increase over the state-of-the-arts respectively.
Rama Kovvuri, Jiyang Gao, Ramakant Nevatia
ICMR4
2016 Image Set Classification via Template Triplets and Context-Aware Similarity Embedding
Feng-Ju Chang, Ramakant Nevatia
ACCV (5)2
2016 Learning Action Concept Trees and Semantic Alignment Networks from Image-Description Data
Jiyang Gao, Ramakant Nevatia
ACCV (2)2
2016 ProNet: Learning to Propose Object-Specific Boxes for Cascaded Neural Networks
abstract
This paper aims to classify and locate objects accurately and efficiently, without using bounding box annotations. It is challenging as objects in the wild could appear at arbitrary locations and in different scales. In this paper, we propose a novel classification architecture ProNet based on convolutional neural networks. It uses computationally efficient neural networks to propose image regions that are likely to contain objects, and applies more powerful but slower networks on the proposed regions. The basic building block is a multi-scale fully-convolutional network which assigns object confidence scores to boxes at different locations and scales. We show that such networks can be trained effectively using image-level annotations, and can be connected into cascades or trees for efficient object classification. ProNet outperforms previous state-of-the-art significantly on PASCAL VOC 2012 and MS COCO datasets for object classification and point-based localization.
Chen Sun 0002, Manohar Paluri, Ronan Collobert, Ramakant Nevatia, Lubomir D. Bourdev
CVPR4
2016 Exploring deep learning based solutions in fine grained activity recognition in the wild
abstract
In this paper, we explore the usage of deep learning based solutions in fine grained activity recognition in the wild. As a powerful tool, deep learning has been widely used in image classification, object detection and activity recognition. We focus on implementing deep learning methods into the more complicated fine grained activity recognition problems. We test our solutions on MPII activity dataset with 410 activities. We find that due to the challenges of large intra class variances, small inter class variances, and limited training samples per activity, the classical two stream deep ConvNets method does not perform that well for fine grained activity recognition. Observing these issues, we propose a solution to directly use deep features learned from ImageNet in an SVM. In experiments, we achieve a 20 percent improvement compared to the classical two stream deep ConvNets solutions, on MPII fine grained activity challenge videos.
Ramakant Nevatia
ICPR2
2016 Segment-based models for event detection and recounting
abstract
We present a novel approach towards web video classification and recounting that uses video segments to model an event. This approach overcomes the limitations faced by the classical video-level models such as modeling semantics, identifying informative segments in a video and background segment suppression. We posit that segment-based models are able to identify both the frequently-occurring and rarer patterns in an event effectively, despite being trained on only a fraction of the training data. Our framework employs a discriminative approach to optimize our models in distributed and data-driven fashion while maintaining semantic interpretability. We evaluate the effectiveness of our approach on the challenging TRECVID MEDTest 2014 dataset. We demonstrate improvements in recounting and classification, particularly in events characterized by inherent intra-class variations.
Rama Kovvuri, Ramakant Nevatia, Cees Snoek
ICPR2
2016 A multi-scale cascade fully convolutional network face detector
abstract
Face detection is challenging as faces in images could be present at arbitrary locations and in different scales. We propose a three-stage cascade structure based on fully convolutional neural networks (FCNs). It first proposes the approximate locations where the faces may be, then aims to find the accurate location by zooming on to the faces. Each level of the FCN cascade is a multi-scale fully-convolutional network, which generates scores at different locations and in different scales. A score map is generated after each FCN stage. Probable regions of face are selected and fed to the next stage. The number of proposals is decreased after each level, and the areas of regions are decreased to more precisely fit the face. Compared to passing proposals directly between stages, passing probable regions can decrease the number of proposals and reduce the cases where first stage doesn't propose good bounding boxes. We show that by using FCN and score map, the FCN cascade face detector can achieve strong performance on public datasets.
Zhenheng Yang, Ramakant Nevatia
ICPR2
2016 ACD: Action Concept Discovery from Image-Sentence Corpora
abstract
Action classification in still images is an important task in computer vision. It is challenging as the appearances of actions may vary depending on their context (e.g. associated objects). Manually labeling of context information would be time consuming and difficult to scale up. To address this challenge, we propose a method to automatically discover and cluster action concepts, and learn their classifiers from weakly supervised image-sentence corpora. It obtains candidate action concepts by extracting verb-object pairs from sentences and verifies their visualness with the associated images. Candidate action concepts are then clustered by using a multi-modal representation with image embeddings from deep convolutional networks and text embeddings from word2vec. More than one hundred human action concept classifiers are learned from the Flickr 30k dataset with no additional human effort and promising classification results are obtained. We further apply the AdaBoost algorithm to automatically select and combine relevant action concepts given an action query. Promising results have been shown on the PASCAL VOC 2012 action classification benchmark, which has zero overlap with Flickr30k.
Jiyang Gao, Chen Sun 0002, Ramakant Nevatia
ICMR3
2016 Face recognition using deep multi-pose representations
abstract
We introduce our method and system for face recognition using multiple pose-aware deep learning models. In our representation, a face image is processed by several pose-specific deep convolutional neural network (CNN) models to generate multiple pose-specific features. 3D rendering is used to generate multiple face poses from the input image. Sensitivity of the recognition system to pose variations is reduced since we use an ensemble of pose-specific CNN features. The paper presents extensive experimental results on the effect of landmark detection, CNN layer selection and pose model selection on the performance of the recognition pipeline. Our novel representation achieves better results than the state-of-the-art on IARPA's CS2 and NIST's IJB-A in both verification and identification (i.e. search) tasks.
Wael Abd-Almageed, Yue Wu 0001, Stephen Rawls, Shai Harel, Tal Hassner, Iacopo Masi, Jongmoo Choi, Jatuporn Toy Leksut, Jungyeon Kim, Premkumar Natarajan, Ramakant Nevatia, Gérard G. Medioni
WACV11
2016 Tag-based video retrieval by embedding semantic content in a continuous word space
abstract
Content-based event retrieval in unconstrained web videos, based on query tags, is a hard problem due to large intra-class variances, and limited vocabulary and accuracy of the video concept detectors, creating a "semantic query gap". We present a technique to overcome this gap by using continuous word space representations to explicitly compute query and detector concept similarity. This not only allows for fast query-video similarity computation with implicit query expansion, but leads to a compact video representation, which allows implementation of a real-time retrieval system that can fit several thousand videos in a few hundred megabytes of memory. We evaluate the effectiveness of our representation on the challenging NIST MEDTest 2014 dataset.
Arnav Agharwal, Rama Kovvuri, Ramakant Nevatia, Cees Snoek
WACV3
2016 Abstraction hierarchy and self annotation update for fine grained activity recognition
abstract
Fine-grained activity recognition focuses recognition on sub-ordinate levels. This task is made difficult due to low inter-class variability and high intra-class variability caused by human motion and objects. We propose that recognition of such activities can be significantly improved by grouping and decomposing them into a hierarchy of multiple abstraction layers; we introduce a Hierarchical Activity Network (HAN). Recognition in HAN is guided by classifiers operating at multiple levels; furthermore, descriptions of different levels of abstraction are also generated from HAN, which may be useful for different tasks. We show significant improvements in accuracy of recognition compared to earlier methods. Besides, annotation for fine grained activity is challenging, and inaccurate annotation influences classification performance. We explore an automatic solution for improving the classification results while auto enhancing the annotation quality.
Ramakant Nevatia
WACV3
2016 Activity recognition and prediction with pose based discriminative patch model
abstract
We describe an image based activity recognition solution which can be applied to both off-line video classification and activity prediction in frames. We propose a Pose based Discriminative Patch Model to make activity recognition and prediction on image level (only observing several frames). This model enables a general and flexible framework to add in discriminative patches and consider their mutual relations to an efficient tree structure. PDP makes contribution in two aspects: (1) PDP provides a novel solution to improve activity recognition and prediction, by utilizing pose based discriminative patches instead of pose configuration feature, and modeling the patches' mutual relations. (2) PDP is an image-based algorithm, so it can make predictions using limited frames, even a single image. PDP focuses on challenging data captured from Internet and movies, where we achieve a 6% improvement compared with state-of-the-art method on video level recognition dataset - Sub-JHMDB, and image level action recognition dataset. We also obtain good improvement on activity prediction task.
Ramakant Nevatia
WACV3
2015 Video event classification with temporal partitioning
abstract
This paper addresses the problem of temporal pruning of noisy parts to improve event recognition performance. We present a new technique based on the temporal partitioning of the processed videos according to their motion patterns and the subsequent analysis of the yielded time segments. For each event type, we automatically learn the types of segments that are discriminative and those that perturb the classification. This process does not require detailed annotation of actions within an event type. A video is described with a set of quantized features and the final classification is performed according to the features that fall within the discriminative segments only. Experimental results show increased classification performance on the NIST MED11 dataset using two types of local features.
Rémi Trichet, Ramakant Nevatia, J. Brian Burns
AVSS2
2015 Automatic Concept Discovery from Parallel Text and Visual Corpora
abstract
Humans connect language and vision to perceive the world. How to build a similar connection for computers? One possible way is via visual concepts, which are text terms that relate to visually discriminative entities. We propose an automatic visual concept discovery algorithm using parallel text and visual corpora, it filters text terms based on the visual discriminative power of the associated images, and groups them into concepts using visual and semantic similarities. We illustrate the applications of the discovered concepts using bidirectional image and sentence retrieval task and image tagging task, and show that the discovered concepts not only outperform several large sets of manually selected concepts significantly, but also achieves the state-of-the-art performance in the retrieval task.
Chen Sun 0002, Chuang Gan 0001, Ramakant Nevatia
ICCV3
2015 Temporal Localization of Fine-Grained Actions in Videos by Domain Transfer from Web Images
abstract
We address the problem of fine-grained action localization from temporally untrimmed web videos. We assume that only weak video-level annotations are available for training. The goal is to use these weak labels to identify temporal segments corresponding to the actions, and learn models that generalize to unconstrained web videos. We find that web images queried by action names serve as well-localized highlights for many actions, but are noisily labeled. To solve this problem, we propose a simple yet effective method that takes weak video labels and noisy image labels as input, and generates localized action frames as output. This is achieved by cross-domain transfer between video frames and web images, using pre-trained deep convolutional neural networks. We then use the localized action frames to train action recognition models with long short-term memory networks. We collect a fine-grained sports action data set FGA-240 of more than 130,000 YouTube videos. It has 240 fine-grained actions under 85 sports activities. Convincing results are shown on the FGA-240 data set, as well as the THUMOS 2014 localization data set with untrimmed training videos.
Chen Sun 0002, Sanketh Shetty, Rahul Sukthankar, Ramakant Nevatia
ACM Multimedia4
2015 Forecasting Human Pose and Motion with Multibody Dynamic Model
abstract
Understanding human motion with dynamics is in its infancy, but it is a highly promising approach in computer vision, robotics and computer graphics. We propose a Multibody Dynamic Model (MDM) which estimates poses and motions through analyzing forces-the intrinsic motivation for motion. With the 23 degrees of freedom Multibody Dynamic Model, we analyze human motion dynamics in the whole body, and then forecast human motion or pose in occluded or non-captured circumstances. Our two main contributions are essential for understanding human motion with dynamics. The first one is to provide effective representations and computational models for dynamic analysis of human motion in the whole body, via the intrinsic connection between force and motion in the biomechanical system. The second contribution is to offer a more natural method to forecast pose and motion with the estimated forces. In our experiments, MDM has been successfully applied to running, jumping and other challenging sports activities.
Ramakant Nevatia
WACV2
2015 A Robust Adaptive Classifier for Detector Adaptation in a Video
abstract
We propose a novel method for improved object detection in a video. Our approach adapts a generic offline trained detector (OTD) to a specific test video by collecting online samples in an unsupervised manner. Most of the existing adaptation methods focus on collecting confident online samples and do not address how to deal with ambiguous and noisy online samples. We address the importance of collecting online samples which are true representative of the actual objects present in the video and propose a Boosted Multiple Instance Random Fern (B-MIRF) classifier as the adaptive classifier. Multiple Instance Learning (MIL) provides reliability for training with noisy online samples and boosting process enables in obtaining more discriminative random ferns. We apply B-MIRF classifier on the detection responses obtained from OTD, hence our method improves the performance by improving the precision of OTD. We evaluate performance of our method on two challenging public datasets and show better performance than other state of the art methods.
Pramod Sharma, Ramakant Nevatia
WACV2
2015 Beyond Pedestrians: A Hybrid Approach of Tracking Multiple Articulating Humans
abstract
We propose a hybrid framework to address the problem of tracking multiple articulated humans from a single camera. Our method incorporates offline learned category-level detector with online learned instance-specific detector as a hybrid system. To deal with humans in large pose articulation, which can not be reliably detected by off-line trained detectors, we propose an online learned instance specific patch-based detector, consisting of layered patch classifiers. With extrapolated track lets by online learned detectors, we use the discriminative color filters learned online to compute the appearance affinity score for further global association. Experimental evaluation on both standard pedestrian datasets and articulated human datasets shows significant improvement compared to state-of-the-art multi-human tracking methods.
Ramakant Nevatia, Bo Yang 0008
WACV2
2014 Multi-state Discriminative Video Segment Selection for Complex Event Classification
Prithviraj Banerjee, Ramakant Nevatia
ACCV (5)2
2014 DISCOVER: Discovering Important Segments for Classification of Video Events and Recounting
abstract
We propose a unified framework DISCOVER to simultaneously discover important segments, classify high-level events and generate recounting for large amounts of unconstrained web videos. The motivation is our observation that many video events are characterized by certain important segments. Our goal is to find the important segments and capture their information for event classification and recounting. We introduce an evidence localization model where evidence locations are modeled as latent variables. We impose constraints on global video appearance, local evidence appearance and the temporal structure of the evidence. The model is learned via a max-margin framework and allows efficient inference. Our method does not require annotating sources of evidence, and is jointly optimized for event classification and recounting. Experimental results are shown on the challenging TRECVID 2013 MEDTest dataset.
Chen Sun 0002, Ramakant Nevatia
CVPR2
2014 Pose Filter Based Hidden-CRF Models for Activity Detection
Prithviraj Banerjee, Ramakant Nevatia
ECCV (2)2
2014 Semantic Aware Video Transcription Using Random Forest Classifiers
Chen Sun 0002, Ramakant Nevatia
ECCV (1)2
2014 Late fusion and calibration for multimedia event detection using few examples
abstract
The state-of-the-art in example-based multimedia event detection (MED) rests on heterogeneous classifiers whose scores are typically combined in a late-fusion scheme. Recent studies on this topic have failed to reach a clear consensus as to whether machine learning techniques can outperform rule-based fusion schemes with varying amount of training data. In this paper, we present two parametric approaches to late fusion: a normalization scheme for arithmetic mean fusion (logistic averaging) and a fusion scheme based on logistic regression, and compare them to widely used rule-based fusion schemes. We also describe how logistic regression can be used to calibrate the fused detection scores to predict an optimal threshold given a detection prior and costs on errors. We discuss the advantages and shortcomings of each approach when the amount of positives available for training varies from 10 positives (10Ex) to 100 positives (100Ex). Experiments were run using video data from the NIST TRECVID MED 2013 evaluation and results were reported in terms of a ranking metric: the mean average precision (mAP) and R0, a cost-based metric introduced in TRECVID MED 2013.
Julien van Hout, Eric Yeh, Dennis C. Koelma, Cees Snoek, Chen Sun 0002, Ramakant Nevatia, Julie Wong, Gregory K. Myers
ICASSP6
2014 Video Segmentation Descriptors for Event Recognition
abstract
This paper presents a new video motion descriptor based on a multi-scale video segmentation to provide a multi-layered output as well as connections with the rich interactions that occur between objects at the semantic level. We also put the emphasis on relationships between motion clusters by providing a new relative motion descriptor encapsulating relative motion patterns within a local spatio-temporal neighborhood. Experimental results on the challenging TRECVID MED11 event recognition dataset validate the approach.
Rémi Trichet, Ramakant Nevatia
ICPR2
2014 ISOMER: Informative Segment Observations for Multimedia Event Recounting
abstract
This paper describes a system for multimedia event detection and recounting. The goal is to detect a high level event class in unconstrained web videos and generate event oriented summarization for display to users. For this purpose, we detect informative segments and collect observations for them, leading to our ISOMER system. We combine a large collection of both low level and semantic level visual and audio features for event detection. For event recounting, we propose a novel approach to identify event oriented discriminative video segments and their descriptions with a linear SVM event classifier. User friendly concepts including objects, actions, scenes, speech and optical character recognition are used in generating descriptions. We also develop several mapping and filtering strategies to cope with noisy concept detectors. Our system performed competitively in the TRECVID 2013 Multimedia Event Detection task with near 100,000 videos and was the highest performer in TRECVID 2013 Multimedia Event Recounting task.
Chen Sun 0002, J. Brian Burns, Ramakant Nevatia, Cees Snoek, Robert C. Bolles, Gregory K. Myers, Wen Wang 0001, Eric Yeh
ICMR3
2014 Multi class boosted random ferns for adapting a generic object detector to a specific video
abstract
Detector adaptation is a challenging problem and several methods have been proposed in recent years. We propose multi class boosted random ferns for detector adaptation. First we collect online samples in an unsupervised manner and collected positive online samples are divided into different categories for different poses of the object. Then we train a multi-class boosted random fern adaptive classifier. Our adaptive classifier training focuses on two aspects: discriminability and efficiency. Boosting provides discriminative random ferns. For efficiency, our boosting procedure focuses on sharing the same feature among different classes and multiple strong classifiers are trained in a single boosting framework. Experiments on challenging public datasets demonstrate effectiveness of our approach.
Pramod Sharma, Ramakant Nevatia
WACV2
2014 Video segmentation and feature co-occurrences for activity classification
abstract
Bag-of-Word scheme has almost become de rigueur for event recognition tasks due to its robustness and simplicity. Despite its effectiveness, this technique discards spatial and temporal relationships between codewords. This paper tackles the problem of building a video codeword representation that captures such relationships. We developed a new method that harnesses spatio-temporal boundaries and discriminative codeword co-occurrences. Given a set of videos and their corresponding quantized features, the video is first decomposed in spatio-temporal volumes according to a multi-scale video segmentation algorithm. Meaningful codeword co-occurrences are then extracted within each volume and videos are then represented with histograms of co-occurring features. The set of histograms is finally fed to an SVM for classification. Evaluation under the realistic TRECVID MED11 challenge database validates the approach.
Rémi Trichet, Ramakant Nevatia
WACV2
2014 Multi-Target Tracking by Online Learning a CRF Model of Appearance and Motion Patterns
Bo Yang 0008, Ramakant Nevatia
Int. J. Comput. Vis.2
2014 Hierarchical abnormal event detection by real time and semi-real time multi-tasking video surveillance system
Sung Chun Lee, Ramakant Nevatia
Mach. Vis. Appl.2
2014 Evaluating multimedia features and fusion for example-based event detection
abstract
Multimedia event detection (MED) is a challenging problem because of the heterogeneous content and variable quality found in large collections of Internet videos. To study the value of multimedia features and fusion for representing and learning events from a set of example video clips, we created SESAME, a system for video SEarch with Speed and Accuracy for Multimedia Events. SESAME includes multiple bag-of-words event classifiers based on single data types: low-level visual, motion, and audio features; high-level semantic visual concepts; and automatic speech recognition. Event detection performance was evaluated for each event classifier. The performance of low-level visual and motion features was improved by the use of difference coding. The accuracy of the visual concepts was nearly as strong as that of the low-level visual features. Experiments with a number of fusion methods for combining the event detection scores from these classifiers revealed that simple fusion methods, such as arithmetic mean, perform as well as or better than other, more complex fusion methods. SESAME’s performance in the 2012 TRECVID MED evaluation was one of the best reported.
Gregory K. Myers, Ramesh Nallapati, Julien van Hout, Stephanie Pancoast, Ramakant Nevatia, Chen Sun 0002, AmirHossein Habibian, Dennis C. Koelma, Koen E. A. van de Sande, Arnold W. M. Smeulders, Cees Snoek
Mach. Vis. Appl.5
2013 Conditional Bayesian networks for action detection
abstract
The task of understanding video content has seen great interest from computer vision community with the increase in camera based surveillance at grocery stores, airports, train stations, etc. What makes up a scene (objects) and what happens in the scene (actions) are two important dimensions of video understanding. In this work, we aim to identify both actions and objects in the video, however, we focus only on the objects with which human interacts. We use videos which may have multiple actions taking place during possibly overlapping intervals. Our system can recognize actions having high intra-class variance performed in complex environments using objects of different types, sizes and shapes. We produce structured descriptions for the videos as output. The descriptions identify the subject, the object, the verb and the interval of each activity recognized.
Furqan Muhammad Khan, Sung Chun Lee, Ramakant Nevatia
AVSS3
2013 Video segmentation with spatio-temporal tubes
abstract
Long-term temporal interactions among objects are an important cue for video understanding. To capture such object relations, we propose a novel method for spatiotemporal video segmentation based on dense trajectory clustering that is also effective when objects articulate. We use superpixels of homogeneous size jointly with optical flow information to ease the matching of regions from one frame to another. Our second main contribution is a hierarchical fusion algorithm that yields segmentation information available at multiple linked scales. We test the algorithm on several videos from the web showing a large variety of difficulties.
Rémi Trichet, Ramakant Nevatia
AVSS2
2013 Efficient Detector Adaptation for Object Detection in a Video
abstract
In this work, we present a novel and efficient detector adaptation method which improves the performance of an offline trained classifier (baseline classifier) by adapting it to new test datasets. We address two critical aspects of adaptation methods: generalizability and computational efficiency. We propose an adaptation method, which can be applied to various baseline classifiers and is computationally efficient also. For a given test video, we collect online samples in an unsupervised manner and train a random fern adaptive classifier. The adaptive classifier improves precision of the baseline classifier by validating the obtained detection responses from baseline classifier as correct detections or false alarms. Experiments demonstrate generalizability, computational efficiency and effectiveness of our method, as we compare our method with state of the art approaches for the problem of human detection and show good performance with high computational efficiency on two different baseline classifiers.
Pramod Sharma, Ramakant Nevatia
CVPR2
2013 ACTIVE: Activity Concept Transitions in Video Event Classification
abstract
The goal of high level event classification from videos is to assign a single, high level event label to each query video. Traditional approaches represent each video as a set of low level features and encode it into a fixed length feature vector (e.g. Bag-of-Words), which leave a big gap between low level visual features and high level events. Our paper tries to address this problem by exploiting activity concept transitions in video events (ACTIVE). A video is treated as a sequence of short clips, all of which are observations corresponding to latent activity concept variables in a Hidden Markov Model (HMM). We propose to apply Fisher Kernel techniques so that the concept transitions over time can be encoded into a compact and fixed length feature vector very efficiently. Our approach can utilize concept annotations from independent datasets, and works well even with a very small number of training samples. Experiments on the challenging NIST TRECVID Multimedia Event Detection (MED) dataset shows our approach performs favorably over the state-of-the-art.
Chen Sun 0002, Ramakant Nevatia
ICCV2
2013 Large-scale web video event classification by use of Fisher Vectors
abstract
Event recognition has been an important topic in computer vision research due to its many applications. However, most of the work has focused on videos taken from a fixed camera, known environments and basic events. Here, we focus on classification of unconstrained, web videos into much higher level activities. We follow the approach of constructing fixed length feature vectors from local feature descriptors for classification using an SVM. Our key contribution is the study of the utility of Fisher Vector representation in improving results compared to the conventional Bag-of-Words (BoW) approach. Such coding has shown to be useful for static image classification in the past but not applied to video categorization. We perform tests on the challenging NIST TRECVID Multimedia Event Detection (MED) dataset, which has thousand hours of unconstrained user generated videos; our approach achieves as much as 35% improvement over the BoW baseline. We also offer an analysis of possible causes of such improvements.
Chen Sun 0002, Ramakant Nevatia
WACV2
2013 Hierarchical multi-channel hidden semi Markov graphical models for activity recognition
Pradeep Natarajan, Ramakant Nevatia
Comput. Vis. Image Underst.2
2013 Multiple Target Tracking by Learning-Based Hierarchical Association of Detection Responses
abstract
We propose a hierarchical association approach to multiple target tracking from a single camera by progressively linking detection responses into longer track fragments (i.e., tracklets). Given frame-by-frame detection results, a conservative dual-threshold method that only links very similar detection responses between consecutive frames is adopted to generate initial tracklets with minimum identity switches. Further association of these highly fragmented tracklets at each level of the hierarchy is formulated as a Maximum A Posteriori (MAP) problem that considers initialization, termination, and transition of tracklets as well as the possibility of them being false alarms, which can be efficiently computed by the Hungarian algorithm. The tracklet affinity model, which measures the likelihood of two tracklets belonging to the same target, is a linear combination of automatically learned weak nonparametric models upon various features, which is distinct from most of previous work that relies on heuristic selection of parametric models and manual tuning of their parameters. For this purpose, we develop a novel bag ranking method and train the crucial tracklet affinity models by the boosting algorithm. This bag ranking method utilizes the soft max function to relax the oversufficient objective function used by the conventional instance ranking method. It provides a tighter upper bound of empirical errors in distinguishing correct associations from the incorrect ones, and thus yields more accurate tracklet affinity models for the tracklet association problem. We apply this approach to the challenging multiple pedestrian tracking task. Systematic experiments conducted on two real-life datasets show that the proposed approach outperforms previous state-of-the-art algorithms in terms of tracking accuracy, in particular, considerably reducing fragmentations and identity switches.
Chang Huang, Yuan Li 0022, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 Robust Object Tracking Using Constellation Model with Superpixel
Ramakant Nevatia
ACCV (3)2
2012 Unsupervised incremental learning for improved object detection in a video
abstract
Most common approaches for object detection collect thousands of training examples and train a detector in an offline setting, using supervised learning methods, with the objective of obtaining a generalized detector that would give good performance on various test datasets. However, when an offline trained detector is applied on challenging test datasets, it may fail in some cases by not being able to detect some objects or by producing false alarms. We propose an unsupervised multiple instance learning (MIL) based incremental solution to deal with this issue. We introduce an MIL loss function for Real Adaboost and present a tracking based effective unsupervised online sample collection mechanism to collect the online samples for incremental learning. Experiments demonstrate the effectiveness of our approach by improving the performance of a state of the art offline trained detector on the challenging datasets for pedestrian category.
Pramod Sharma, Chang Huang, Ramakant Nevatia
CVPR3
2012 Multi-target tracking by online learning of non-linear motion patterns and robust appearance models
abstract
We describe an online approach to learn non-linear motion patterns and robust appearance models for multi-target tracking in a tracklet association framework. Unlike most previous approaches that use linear motion methods only, we online build a non-linear motion map to better explain direction changes and produce more robust motion affinities between tracklets. Moreover, based on the incremental learned entry/exit map, a multiple instance learning method is devised to produce strong appearance models for tracking; positive sample pairs are collected from different track-lets so that training samples have high diversity. Finally, using online learned moving groups, a tracklet completion process is introduced to deal with tracklets not reaching entry/exit points. We evaluate our approach on three public data sets, and show significant improvements compared with state-of-art methods.
Bo Yang 0008, Ramakant Nevatia
CVPR2
2012 An online learned CRF model for multi-target tracking
abstract
We introduce an online learning approach for multitarget tracking. Detection responses are gradually associated into tracklets in multiple levels to produce final tracks. Unlike most previous approaches which only focus on producing discriminative motion and appearance models for all targets, we further consider discriminative features for distinguishing difficult pairs of targets. The tracking problem is formulated using an online learned CRF model, and is transformed into an energy minimization problem. The energy functions include a set of unary functions that are based on motion and appearance models for discriminating all targets, as well as a set of pairwise functions that are based on models for differentiating corresponding pairs of tracklets. The online CRF approach is more powerful at distinguishing spatially close targets with similar appearances, as well as in dealing with camera motions. An efficient algorithm is introduced for finding an association with low energy cost. We evaluate our approach on three public data sets, and show significant improvements compared with several state-of-art methods.
Bo Yang 0008, Ramakant Nevatia
CVPR2
2012 Online Learned Discriminative Part-Based Appearance Models for Multi-human Tracking
Bo Yang 0008, Ramakant Nevatia
ECCV (1)2
2012 Pose based activity recognition using Multiple Kernel learning
Prithviraj Banerjee, Ramakant Nevatia
ICPR2
2012 Robust multi-pose face tracking by multi-stage tracklet association
Markus Roth, Martin Bäuml, Ramakant Nevatia, Rainer Stiefelhagen
ICPR3
2012 Efficient incremental learning of boosted classifiers for object detection
Pramod Sharma, Chang Huang, Ramakant Nevatia
ICPR3
2012 Simultaneous inference of activity, pose and object
abstract
Human movements are important cues for recognizing human actions, which can be captured by explicit modeling and tracking of actor or through space-time low-level features. However, relying solely on human dynamics is not enough to discriminate between actions which have similar human dynamics, such as smoking and drinking, irrespective of the modeling method. Object perception plays an important role in such cases. Conversely, human movements are indicative of type of object used for the action. These two processes of object perception and action understanding are thus not independent. Consequently, action recognition improves when human movements and object perception are used in conjunction. Therefore, we propose a probabilistic approach to simultaneously infer what action was performed, what object was used and what poses the actor went through. This joint inference framework can better discriminate between actions and objects which are too similar and lack discriminative features.
Furqan Muhammad Khan, Vivek K. Singh 0002, Ramakant Nevatia
WACV3
2012 A systems level approach to perimeter protection
abstract
Effective perimeter protection mechanisms for industrial sites and critical infrastructure must contend with a large variety of potential threats as well as with the fact that normal site activity can be both complex and diverse. This paper documents the development of a system level approach capable of functioning under such challenging conditions. A multi-view tracking system is used to provide real-time site wide trajectories of all observed individuals. A Radar-based system is also used for tracking if and when camera coverage of various regions is not available. Track information is then analyzed with respect to articulated motion analysis, complex event analysis and normalcy analysis. In addition, object recognition is used to classify left behind objects using high resolution PTZ imagery. A real-time integrated version of this comprehensive approach to perimeter protection was deployed using a single standard off-the-shelf desktop computer.
Peter H. Tu, Ting Yu 0003, Ramakant Nevatia, Sung Chun Lee, Hale Kim, Phill-Kyu Rhee, Joong-Hwan Baek
WACV4
2011 Learning neighborhood cooccurrence statistics of sparse features for human activity recognition
abstract
A common approach to activity recognition has been the use of histogram of codewords computed from Spatio Temporal Interest Points (STIPs). Recent methods have focused on leveraging the spatio-temporal neighborhood structure of the features, but they are generally restricted to aggregate statistics over the entire video volume, and ignore local pairwise relationships. Our goal is to capture these relations in terms of pairwise cooccurrence statistics of codewords. We show a reduction of such cooccurrence relations to the edges connecting the latent variables of a Conditional Random Field (CRF) classifier. As a consequence, we also learn the codeword dictionary as a part of the maximum likelihood learning process, with each interest point assigned a probability distribution over the codewords. We show results on two widely used activity recognition datasets.
Prithviraj Banerjee, Ramakant Nevatia
AVSS2
2011 AVSS 2011 demo session: A systems level approach to perimeter protection
abstract
Summary form only given. The rapid evolution of tools and software systems to design experiments, automatically monitor, collect and warehouse large amounts of data, from applications such as life sciences and industrial processes has resulted in a new paradigm shift. This change of paradigm is so fast that some of the practices for optimization and management of these processes that were valid only 5–10 years ago may no longer be fully acceptable or sufficient for today's business optimization and management. This has a direct influence on the best practices for knowledge discovery and management of the discovered knowledge in real-world data mining applications. Establishing and managing a real-world data mining project in any domain, in particular in today's life science industry, is not a trivial task. A few approaches have been proposed in the literature. However, initiation and successful management of such efforts may depend on where a given case study fits in the overall classification of data mining approaches. Today's knowledge discovery from data can be classified in several ways: (i) data mining on engineered systems (e.g. complex equipment) or systems designed by nature (e.g. life sciences), (ii) explanatory or predictive data mining, (iii) data mining from static data (e.g. data warehouse) or dynamic data (e.g. data streams), (iv) user operated or automated data mining. There could still be other ways to classify data mining applications. This talk provides an overview of the above listed knowledge discovery applications. We provide examples where we demonstrate how small or large amounts of data, when understood from a real-world data mining point of view and the required data is properly integrated, can result in novel knowledge discovery case studies. We explain motivations and challenges of establishing real-world dat
Peter H. Tu, Ting Yu 0003, Ramakant Nevatia, Sung Chun Lee, Hale Kim, Phill-Kyu Rhee, Joong-Hwan Baek
AVSS4
2011 How does person identity recognition help multi-person tracking?
abstract
We address the problem of multi-person tracking in a complex scene from a single camera. Although tracklet-association methods have shown impressive results in several challenging datasets, discriminability of the appearance model remains a limitation. Inspired by the work of person identity recognition, we obtain discriminative appearance-based affinity models by a novel framework to incorporate the merits of person identity recognition, which help multi-person tracking performance. During off-line learning, a small set of local image descriptors is selected to be used in on-line learned appearances-based affinity models effectively and efficiently. Given short but reliable track-lets generated by frame-to-frame association of detection responses, we identify them as query tracklets and gallery tracklets. For each gallery tracklet, a target-specific appearance model is learned from the on-line training samples collected by spatio-temporal constraints. Both gallery tracklets and query tracklets are fed into hierarchical association framework to obtain final tracking results. We evaluate our proposed system on several public datasets and show significant improvements in terms of tracking evaluation metrics.
Cheng-Hao Kuo, Ramakant Nevatia
CVPR2
2011 Learning affinities and dependencies for multi-target tracking using a CRF model
abstract
We propose a learning-based Conditional Random Field (CRF) model for tracking multiple targets by progressively associating detection responses into long tracks. Tracking task is transformed into a data association problem, and most previous approaches developed heuristical parametric models or learning approaches for evaluating independent affinities between track fragments (tracklets). We argue that the independent assumption is not valid in many cases, and adopt a CRF model to consider both tracklet affinities and dependencies among them, which are represented by unary term costs and pairwise term costs respectively. Unlike previous methods, we learn the best global associations instead of the best local affinities between tracklets, and transform the task of finding the best association into an energy minimization problem. A RankBoost algorithm is proposed to select effective features for estimation of term costs in the CRF model, so that better associations have lower costs. Our approach is evaluated on challenging pedestrian data sets, and are compared with state-of-art methods. Experiments show effectiveness of our algorithm as well as improvement in tracking performance.
Bo Yang 0008, Chang Huang, Ramakant Nevatia
CVPR3
2011 Action recognition in cluttered dynamic scenes using Pose-Specific Part Models
abstract
We present an approach to recognizing single actor human actions in complex backgrounds. We adopt a Joint Tracking and Recognition approach, which track the actor pose by sampling from 3D action models. Most existing such approaches require large training data or MoCAP to handle multiple viewpoints, and often rely on clean actor silhouettes. The action models in our approach are obtained by annotating keyposes in 2D, lifting them to 3D stick figures and then computing the transformation matrices between the 3D keypose figures. Poses sampled from coarse action models may not fit the observations well; to overcome this difficulty, we propose an approach for efficiently localizing a pose by generating a Pose-Specific Part Model (PSPM) which captures appropriate kinematic and occlusion constraints in a tree-structure. In addition, our approach also does not require pose silhouettes. We show improvements to previous results on two publicly available datasets as well as on a novel, augmented dataset with dynamic backgrounds.
Vivek K. Singh 0002, Ramakant Nevatia
ICCV2
2011 Vehicle detection from low quality aerial LIDAR data
abstract
In this paper we propose a vehicle detection framework on low resolution aerial range data. Our system consists of three steps: data mapping, 2D vehicle detection and postprocessing. First, we map the range data into 2D grayscale images by using the depth information only. For this purpose we propose a novel local ground plane estimation method, and the estimated ground plane is further refined by a global refinement process. Then we compute the depth value of missing points (points for which no depth information is available) by an effective interpolation method. In the second step, to train a classifier for the vehicles, we describe a method to generate more training examples from very few training annotations and adopt the fast cascade Adaboost approach for detecting vehicles in 2D grayscale images. Finally, in post-processing step we design a novel method to detect some vehicles which are comprised of clusters of missing points. We evaluate our method on real aerial data and the experiments demonstrate the effectiveness of our approach.
Bo Yang 0008, Pramod Sharma, Ramakant Nevatia
WACV3
2011 Segmentation of objects in a detection window by Nonparametric Inhomogeneous CRFs
Bo Yang 0008, Chang Huang, Ramakant Nevatia
Comput. Vis. Image Underst.3
2011 Simultaneous tracking and action recognition for single actor human actions
Vivek K. Singh 0002, Ramakant Nevatia
Vis. Comput.2
2010 Dynamics Based Trajectory Segmentation for UAV videos
abstract
A novel representation of vehicle trajectories is proposed for applications in trajectory analysis and activity detection. Specifically, a piecewise arc fitting based smoothing algorithm is proposed for denoising the trajectories. A dynamic program is used to find the optimal arc fit to a given trajectory. We motivate the usage of dynamic primitives to parametrize common vehicular activities, and propose a dynamics based trajectory segmentation algorithm. Each primitive is modeled using a second order Auto-Regressive model, and form useful descriptors for a given vehicular trajectory. We evaluate both our trajectory smoothing and dynamic trajectory segmentation algorithm on a real UAV video dataset, and show performance improvements which clearly motivate its wide applicability in a general trajectory analysis system.
Prithviraj Banerjee, Ramakant Nevatia
AVSS2
2010 High performance object detection by collaborative learning of Joint Ranking of Granules features
abstract
Object detection remains an important but challenging task in computer vision. We present a method that combines high accuracy with high efficiency. We adopt simplified forms of APCF features [3], which we term Joint Ranking of Granules (JRoG) features; the features consists of discrete values by uniting binary ranking results of pair-wise granules in the image. We propose a novel collaborative learning method for JRoG features, which consists of a Simulated Annealing (SA) module and an incremental feature selection module. The two complementary modules collaborate to efficiently search the formidably large JRoG feature space for discriminative features, which are fed into a boosted cascade for object detection. To cope with occlusions in crowded environments, we employ the strategy of part based detection, as in [19] but propose a new dynamic search method to improve the Bayesian combination of the part detection results. Experiments on several challenging data sets show that our approach achieves not only considerable improvement in detection accuracy but also major improvements in computational efficiency; on a Xeon 3GHz computer, with only a single thread, it can process a million scanning windows per second, sufficing for many practical real-time detection tasks.
Chang Huang, Ramakant Nevatia
CVPR2
2010 Multi-target tracking by on-line learned discriminative appearance models
abstract
We present an approach for online learning of discriminative appearance models for robust multi-target tracking in a crowded scene from a single camera. Although much progress has been made in developing methods for optimal data association, there has been comparatively less work on the appearance models, which are key elements for good performance. Many previous methods either use simple features such as color histograms, or focus on the discriminability between a target and the background which does not resolve ambiguities between the different targets. We propose an algorithm for learning a discriminative appearance model for different targets. Training samples are collected online from tracklets within a time sliding window based on some spatial-temporal constraints; this allows the models to adapt to target instances. Learning uses an Ad-aBoost algorithm that combines effective image descriptors and their corresponding similarity measurements. We term the learned models as OLDAMs. Our evaluations indicate that OLDAMs have significantly higher discrimination between different targets than conventional holistic color histograms, and when integrated into a hierarchical association framework, they help improve the tracking accuracy, particularly reducing the false alarms and identity switches.
Cheng-Hao Kuo, Chang Huang, Ramakant Nevatia
CVPR3
2010 Learning 3D action models from a few 2D videos for view invariant action recognition
abstract
Most existing approaches for learning action models work by extracting suitable low-level features and then training appropriate classifiers. Such approaches require large amounts of training data and do not generalize well to variations in viewpoint, scale and across datasets. Some work has been done recently to learn multi-view action models from Mocap data, but obtaining such data is time consuming and requires costly infrastructure. We present a method that addresses both these issues by learning action models from just a few video training samples. We model each action as a sequence of primitive actions, represented as functions which transform the actor's state. We formulate model learning as a curve-fitting problem, and present a novel algorithm for learning human actions by lifting 2D annotations of a few keyposes to 3D and interpolating between them. Actions are inferred by sampling the models and accumulating the feature weights learned discriminatively using a latent state Perceptron algorithm. We show results comparable to state-of-art on the standard Weizmann dataset, with a much smaller train:test ratio, and also in datasets for visual gesture recognition and cluttered grocery store environments.
Pradeep Natarajan, Vivek K. Singh 0002, Ramakant Nevatia
CVPR3
2010 Inter-camera Association of Multi-target Tracks by On-Line Learned Appearance Affinity Models
Cheng-Hao Kuo, Chang Huang, Ramakant Nevatia
ECCV (1)3
2010 Efficient Inference with Multiple Heterogeneous Part Detectors for Human Pose Estimation
Vivek K. Singh 0002, Ramakant Nevatia, Chang Huang
ECCV (3)2
2009 Learning to associate: HybridBoosted multi-target tracker for crowded scene
abstract
We propose a learning-based hierarchical approach of multi-target tracking from a single camera by progressively associating detection responses into longer and longer track fragments (tracklets) and finally the desired target trajectories. To define tracklet affinity for association, most previous work relies on heuristically selected parametric models; while our approach is able to automatically select among various features and corresponding non-parametric models, and combine them to maximize the discriminative power on training data by virtue of a HybridBoost algorithm. A hybrid loss function is used in this algorithm because the association of tracklet is formulated as a joint problem of ranking and classification: the ranking part aims to rank correct tracklet associations higher than other alternatives; the classification part is responsible to reject wrong associations when no further association should be done. Experiments are carried out by tracking pedestrians in challenging datasets. We compare our approach with state-of-the-art algorithms to show its improvement in terms of tracking accuracy.
Yuan Li 0022, Chang Huang, Ramakant Nevatia
CVPR3
2009 Robust multi-view car detection using unsupervised sub-categorization
abstract
This paper presents a novel approach for multi-view car detection using unsupervised sub-categorization instead of manual labeling. Cars have large variability of models and the view-point makes the appearance change dramatically. For object classes with a large intra-class variation like cars, a divide-and-conquer strategy may be applied. Instead of using manually predefined intra-class sub-categorization, we examine several non-linear dimension reduction methods and group samples in the low-dimension embedding in an unsupervised way. The clustered samples have strong view-point similarities internally. A boosting-based cascade tree classifier is trained based on these sub-categorizations. To demonstrate the capability of our multi-view car detector, we create a more challenging test set with annotations. Compared to the UIUC side-view car data set, our test set contains a large range of car models, view points, and complex backgrounds. We compare our approach with previous methods and the result shows that ours outperforms the state-of-the-art methods.
Cheng-Hao Kuo, Ramakant Nevatia
WACV2
2009 Extensive articulated human detection by voting Cluster Boosted Tree
abstract
Our goal is to detect people in highly articulated poses, including bending, crouching, etc. Such formidable diversity in human poses makes detection much more difficult than for pedestrian poses. ¿Divide-and-conquer¿ is a favorable strategy for detecting objects with large intra class variations, which splits object instances into several subcategories and trains relatively simple classifiers for each sub-category. We propose a novel sample split method, which benefits the learning results of articulated humans. We adopt the cluster boosted tree (CBT) structure to automatically decide when a split should be triggered. Unlike the simple k-means used in CBT for sample split, our approach aims at minimizing the training loss after the split. Since this minimization is an NP-hard problem, we design a heuristic algorithm, in which we find optimal sample divisions according to each single feature, and then make compromises to get a final division by a voting-like process. We name our training method as voting cluster boosted tree (VCBT). Furthermore, to avoid large background area in training samples, we first cluster samples according to their width/height ratios, and then train a VCBT for each subset. We conduct an experiment on 17 infrared surveillance video clips, report superior performance compared with previous human detection methods, and show how our approach benefits the learning results by reducing training loss.
Bo Yang 0008, Chang Huang, Ramakant Nevatia
WACV3
2009 Detection and Segmentation of Multiple, Partially Occluded Objects by Grouping, Merging, Assigning Part Detection Responses
Bo Wu 0001, Ramakant Nevatia
Int. J. Comput. Vis.2
2009 Human Pose Tracking in Monocular Sequence Using Multilevel Structured Models
abstract
Tracking human body poses in monocular video has many important applications. The problem is challenging in realistic scenes due to background clutter, variation in human appearance and self-occlusion. The complexity of pose tracking is further increased when there are multiple people whose bodies may inter-occlude. We proposed a three-stage approach with multi-level state representation that enables a hierarchical estimation of 3D body poses. Our method addresses various issues including automatic initialization, data association, self and inter-occlusion. At the first stage, humans are tracked as foreground blobs and their positions and sizes are coarsely estimated. In the second stage, parts such as face, shoulders and limbs are detected using various cues and the results are combined by a grid-based belief propagation algorithm to infer 2D joint positions. The derived belief maps are used as proposal functions in the third stage to infer the 3D pose using data-driven Markov chain Monte Carlo. Experimental results on several realistic indoor video sequences show that the method is able to track multiple persons during complex movement including sitting and turning movements with self and inter-occlusion.
Mun Wai Lee, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 View and scale invariant action recognition using multiview shape-flow models
abstract
Actions in real world applications typically take place in cluttered environments with large variations in the orientation and scale of the actor. We present an approach to simultaneously track and recognize known actions that is robust to such variations, starting from a person detection in the standing pose. In our approach we first render synthetic poses from multiple viewpoints using Mocap data for known actions and represent them in a conditional random field (CRF) whose observation potentials are computed using shape similarity and the transition potentials are computed using optical flow. We enhance these basic potentials with terms to represent spatial and temporal constraints and call our enhanced model the shape, flow, duration-conditional random field (SFD-CRF). We find the best sequence of actions using Viterbi search in the SFD-CRF. We demonstrate our approach on videos from multiple viewpoints and in the presence of background clutter.
Pradeep Natarajan, Ramakant Nevatia
CVPR2
2008 Optimizing discrimination-efficiency tradeoff in integrating heterogeneous local features for object detection
abstract
A large variety of image features has been invented for detection of objects of a known class. We propose a framework to optimize the discrimination-efficiency tradeoff in integrating multiple, heterogeneous features for object detection. Cascade structured detectors are learned by boosting local feature based weak classifiers. Each weak classifier corresponds to a local image region, from which several different types of features are extracted. The weak classifier makes predictions by examining the features one by one; this classifier goes to the next feature only when the prediction from the already examined features is not confident enough. The order in which the features are evaluated is determined based on their computational cost normalized classification powers. We apply our approach to two object classes, pedestrians and cars. The experimental results show that our approach outperforms the state-of-the-art methods.
Bo Wu 0001, Ramakant Nevatia
CVPR2
2008 Segmentation of multiple, partially occluded objects by grouping, merging, assigning part detection responses
abstract
We propose a method that detects and segments multiple, partially occluded objects in images. A part hierarchy is defined for the object class. Whole-object segmentor and part detectors are learned by boosting shape oriented local image features. During detection, the part detectors are applied to the input image. All the edge pixels in the image that positively contribute to part detection responses are extracted. A joint likelihood of multiple objects is defined based on the part detection responses and the object edges. Computing the joint likelihood includes an inter-object occlusion reasoning that is based on the object silhouettes extracted with the whole-object segmentor. By maximizing the joint likelihood, part detection responses are grouped, merged, and assigned to multiple object hypotheses. The proposed approach is applied to the pedestrian class, and evaluated on two public test sets. The experimental results show that our method outperforms the previous ones.
Bo Wu 0001, Ramakant Nevatia, Yuan Li 0022
CVPR2
2008 Global data association for multi-object tracking using network flows
abstract
We propose a network flow based optimization method for data association needed for multiple object tracking. The maximum-a-posteriori (MAP) data association problem is mapped into a cost-flow network with a non-overlap constraint on trajectories. The optimal data association is found by a min-cost flow algorithm in the network. The network is augmented to include an Explicit Occlusion Model(EOM) to track with long-term inter-object occlusions. A solution to the EOM-based network is found by an iterative approach built upon the original algorithm. Initialization and termination of trajectories and potential false observations are modeled by the formulation intrinsically. The method is efficient and does not require hypotheses pruning. Performance is compared with previous results on two public pedestrian datasets to show its improvement.
Yuan Li 0022, Ramakant Nevatia
CVPR3
2008 Robust Object Tracking by Hierarchical Association of Detection Responses
Chang Huang, Bo Wu 0001, Ramakant Nevatia
ECCV (2)3
2008 Key Object Driven Multi-category Object Recognition, Localization and Tracking Using Spatio-temporal Context
Yuan Li 0022, Ramakant Nevatia
ECCV (4)2
2008 Human detection by searching in 3d space using camera and scene knowledge
abstract
Many existing human detection systems are based on sub-window classification, namely detection is done by enumerating rectangular sub-images in the 2D image space. Detection rate of such approaches may be affected by perspective distortion and tilted orientation of the human in images. To overcome this problem without re-training the classifier, we develop a 3D search method. A search grid is defined in the 3D scene. At each grid point a rectified sub-image is generated to approximate the orthogonal projection of the target, so that the distortion due to camera setting is reduced. In addition, 3D target position can be estimated from single camera data. Experiments on challenging data from the PETS2007 and CAVIAR INRIA datasets show significantly improved detection performance of our approach compared with the 2D search-based methods.
Yuan Li 0022, Bo Wu 0001, Ramakant Nevatia
ICPR3
2008 Segmentation and Tracking of Multiple Humans in Crowded Environments
abstract
Segmentation and tracking of multiple humans in crowded situations is made difficult by interobject occlusion. We propose a model based approach to interpret the image observations by multiple, partially occluded human hypotheses in a Bayesian framework. We define a joint image likelihood for multiple humans based on the appearance of the humans, the visibility of body obtained by occlusion reasoning, and foreground/background separation. The optimal solution is obtained by using an efficient sampling method, data-driven Markov chain Monte Carlo (DDMCMC), which uses image observations for proposal probabilities. Knowledge of various aspects including human shape, camera model, and image cues are integrated in one theoretically sound framework. We present experimental results and quantitative evaluation, demonstrating that the resulting approach is effective for very challenging data.
Tao Zhao 0001, Ramakant Nevatia, Bo Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Single View Human Action Recognition using Key Pose Matching and Viterbi Path Searching
abstract
3D human pose recovery is considered as a fundamental step in view-invariant human action recognition. However, inferring 3D poses from a single view usually is slow due to the large number of parameters that need to be estimated and recovered poses are often ambiguous due to the perspective projection. We present an approach that does not explicitly infer 3D pose at each frame. Instead, from existing action models we search for a series of actions that best match the input sequence. In our approach, each action is modeled as a series of synthetic 2D human poses rendered from a wide range of viewpoints. The constraints on transition of the synthetic poses is represented by a graph model called Action Net. Given the input, silhouette matching between the input frames and the key poses is performed first using an enhanced Pyramid Match Kernel algorithm. The best matched sequence of actions is then tracked using the Viterbi algorithm. We demonstrate this approach on a challenging video sets consisting of 15 complex action classes.
Fengjun Lv, Ramakant Nevatia
CVPR2
2007 Simultaneous Object Detection and Segmentation by Boosting Local Shape Feature based Classifier
abstract
This paper proposes an approach to simultaneously detect and segment objects of a known category. Edgelet features are used to capture the local shape of the objects. For each feature a pair of base classifiers for detection and segmentation is built. The base segmentor is designed to predict the per-pixel figure-ground assignment around a neighborhood of the edgelet based on the feature response. The neighborhood is represented as an effective field which is determined by the shape of the edgelet. A boosting algorithm is used to learn the ensemble classifier with cascade decision strategy from the base classifier pool. The simultaneousness is achieved for both training and testing. The system is evaluated on a number of public image sets and compared with several previous methods.
Bo Wu 0001, Ramakant Nevatia
CVPR2
2007 Improving Part based Object Detection by Unsupervised, Online Boosting
abstract
Detection of objects of a given class is important for many applications. However it is difficult to learn a general detector with high detection rate as well as low false alarm rate. Especially, the labor needed for manually labeling a huge training sample set is usually not affordable. We propose an unsupervised, incremental learning approach based on online boosting to improve the performance on special applications of a set of general part detectors, which are learned from a small amount of labeled data and have moderate accuracy. Our oracle for unsupervised learning, which has high precision, is based on a combination of a set of shape based part detectors learned by off-line boosting. Our online boosting algorithm, which is designed for cascade structure detector, is able to adapt the simple features, the base classifiers, the cascade decision strategy, and the complexity of the cascade automatically to the special application. We integrate two noise restraining strategies in both the oracle and the online learner. The system is evaluated on two public video corpora.
Bo Wu 0001, Ramakant Nevatia
CVPR2
2007 Pedestrian Detection in Infrared Images based on Local Shape Features
abstract
Use of IR images is advantageous for many surveillance applications where the systems must operate around the clock and external illumination is not always available. We investigate the methods derived from visible spectrum analysis for the task of human detection. Two feature classes (edgelets and HOG features) and two classification models(AdaBoost and SVM cascade) are extended to IR images. We find out that it is possible to get detection performance in IR images that is comparable to state-of-the-art results for visible spectrum images. It is also shown that the two domains share many features, likely originating from the silhouettes, in spite of the starkly different appearances of the two modalities.
Bo Wu 0001, Ramakant Nevatia
CVPR3
2007 Cluster Boosted Tree Classifier for Multi-View, Multi-Pose Object Detection
abstract
Detection of object of a known class is a fundamental problem of computer vision. The appearance of objects can change greatly due to illumination, view point, and articulation. For object classes with large intra-class variation, some divide-and-conquer strategy is necessary. Tree structured classifier models have been used for multi-view multi- pose object detection in previous work. This paper proposes a boosting based learning method, called Cluster Boosted Tree (CBT), to automatically construct tree structured object detectors. Instead of using predefined intra-class sub- categorization based on domain knowledge, we divide the sample space by unsupervised clustering based on discriminative image features selected by boosting algorithm. The sub-categorization information of the leaf nodes is sent back to refine their ancestors' classification functions. We compare our approach with previous related methods on several public data sets. The results show that our approach outperforms the state-of-the-art methods.
Bo Wu 0001, Ramakant Nevatia
ICCV2
2007 Detection and Tracking of Multiple Humans with Extensive Pose Articulation
abstract
We describe a method for detecting and tracking humans. Different from most of the previous work, we focus on humans with extensive pose articulations, under situations where there is typically only a single camera, multiple humans are present and the image resolution is low. In our method pose clusters are learned from an embedded silhouette manifold. A set of object detectors, each of which corresponds to one pose cluster, are trained based on a novel Object-Weighted Appearance Model. A probabilistic pose-based transition model is used to track multiple objects within a sliding window buffer, making use of the detection responses. The track segments in the sliding windows are connected sequentially into full trajectories. Experiments on a set of challenging surveillance videos are presented; these show good performance of our approach compared to standard pedestrian detectors, under difficult conditions.
Bo Wu 0001, Ramakant Nevatia
ICCV3
2007 Hierarchical Multi-channel Hidden Semi Markov Models
Pradeep Natarajan, Ramakant Nevatia
IJCAI2
2007 Detection and Tracking of Multiple, Partially Occluded Humans by Bayesian Combination of Edgelet based Part Detectors
Bo Wu 0001, Ramakant Nevatia
Int. J. Comput. Vis.2
2006 Tracking of Multiple, Partially Occluded Humans based on Static Body Part Detection
abstract
Tracking of humans in videos is important for many applications. A major source of difficulty in performing this task is due to inter-human or scene occlusion. We present an approach based on representing humans as an assembly of four body parts and detection of the body parts in single frames which makes the method insensitive to camera motions. The responses of the body part detectors and a combined human detector provide the "observations" used for tracking. Trajectory initialization and termination are both fully automatic and rely on the confidences computed from the detection responses. An object is tracked by data association if its corresponding detection response can be found; otherwise it is tracked by a meanshift style tracker. Our method can track humans with both inter-object and scene occlusions. The system is evaluated on three sets of videos and compared with previous method.
Bo Wu 0001, Ramakant Nevatia
CVPR (1)2
2006 Human Pose Tracking Using Multi-level Structured Models
Mun Wai Lee, Ramakant Nevatia
ECCV (3)2
2006 Recognition and Segmentation of 3-D Human Action Using HMM and Multi-class AdaBoost
Fengjun Lv, Ramakant Nevatia
ECCV (4)2
2006 Geodec: Enabling Geospatial Decision Making
abstract
The rapid increase in the availability of geospatial data has motivated the effort to seamlessly integrate this information into an information-rich and realistic 3D environment. However, heterogeneous data sources with varying degrees of consistency and accuracy pose a challenge to such efforts. We describe the geospatial decision making (GeoDec) system, which accurately integrates satellite imagery, three-dimensional models, textures and video streams, road data, maps, point data and temporal data. The system also includes a glove-based user interface
Cyrus Shahabi, Yao-Yi Chiang, Kelvin Chung, Kai-Chen Huang, Ali Khoshgozaran, Craig A. Knoblock, Sung Lee, Ulrich Neumann, Ramakant Nevatia, Arjun Rihan, Snehal Thakkar, Suya You
ICME9
2006 Camera Calibration from Video of a Walking Human
abstract
A self-calibration method to estimate a camera's intrinsic and extrinsic parameters from vertical line segments of the same height is presented. An algorithm to obtain the needed line segments by detecting the head and feet positions of a walking human in his leg-crossing phases is described. Experimental results show that the method is accurate and robust with respect to various viewing angles and subjects.
Fengjun Lv, Tao Zhao 0001, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.3
2005 A Model-Based Vehicle Segmentation Method for Tracking
abstract
Our goal is to detect and track moving vehicles on a road observed from cameras placed on poles or buildings. Inter-vehicle occlusion is significant under these conditions and traditional blob tracking methods is unable to separate the vehicles in the merged blobs. We use vehicle shape models, in addition to camera calibration and ground plane knowledge, to detect, track and classify moving vehicles in presence of occlusion. We use a 2-stage approach. In the first stage, hypothesis for vehicle types, positions and orientations are formed by a coarse search, which is then refined by a data driven Markov chain Monte Carlo (DDMCMC) process. We show results and evaluations on some real urban traffic video sequence using three types of vehicle models
Xuefeng Song, Ramakant Nevatia
ICCV2
2005 Detection of Multiple, Partially Occluded Humans in a Single Image by Bayesian Combination of Edgelet Part Detectors
abstract
This paper proposes a method for human detection in crowded scene from static images. An individual human is modeled as an assembly of natural body parts. We introduce edgelet features, which are a new type of silhouette oriented features. Part detectors, based on these features, are learned by a boosting method. Responses of part detectors are combined to form a joint likelihood model that includes cases of multiple, possibly inter-occluded humans. The human detection problem is formulated as maximum a posteriori (MAP) estimation. We show results on a commonly used previous dataset as well as new data sets that could not be processed by earlier methods.
Bo Wu 0001, Ramakant Nevatia
ICCV2
2004 Extraction and Integration of Window in a 3D Building Model from Ground View Image
Sung Chun Lee, Ramakant Nevatia
CVPR (2)2
2004 Tracking Multiple Humans in Crowded Environment
Tao Zhao 0001, Ramakant Nevatia
CVPR (2)2
2004 Video-based event recognition: activity representation and probabilistic recognition methods
Somboon Hongeng, Ramakant Nevatia, François Brémond
Comput. Vis. Image Underst.2
2004 Automatic description of complex buildings from multiple images
Zu Whan Kim, Ramakant Nevatia
Comput. Vis. Image Underst.2
2004 Tracking Multiple Humans in Complex Situations
abstract
Tracking multiple humans in complex situations is challenging. The difficulties are tackled with appropriate knowledge in the form of various models in our approach. Human motion is decomposed into its global motion and limb motion. In the first part, we show how multiple human objects are segmented and their global motions are tracked in 3D using ellipsoid human shape models. Experiments show that it successfully applies to the cases where a small number of people move together, have occlusion, and cast shadow or reflection. In the second part, we estimate the modes (e.g., walking, running, standing) of the locomotion and 3D body postures by making inference in a prior locomotion model. Camera model and ground plane assumptions provide geometric constraints in both parts. Robust results are shown on some difficult sequences.
Tao Zhao 0001, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Bayesian Human Segmentation in Crowded Situations
abstract
The problem of segmenting individual humans in crowded situations from stationary video camera sequences is exacerbated by object inter-occlusion. We pose this problem as a "model-based segmentation" problem in which human shape models are used to interpret the foreground in a Bayesian framework. The solution is obtained by using an efficient Markov chain Monte Carlo (MCMC) method that uses domain knowledge as proposal probabilities. Knowledge of various aspects including human shape, human height, camera model, and image cues including human head candidates, foreground/background separation are integrated in one theoretically sound framework. We show promising results and evaluations on some challenging data.
Tao Zhao 0001, Ramakant Nevatia
CVPR (2)2
2003 Large-Scale Event Detection Using Semi-Hidden Markov Models
abstract
We present a new approach to recognizing events in videos. We first detect and track moving objects in the scene. Based on the shape and motion properties of these objects, we infer probabilities of primitive events frame-by-frame by using Bayesian networks. Composite events, consisting of multiple primitive events, over extended periods of time are analyzed by using a hidden, semi-Markov finite state model. This results in more reliable event segmentation compared to the use of standard HMMs in noisy video sequences at the cost of some increase in computational complexity. We describe our approach to reducing this complexity. We demonstrate the effectiveness of our algorithm using both real-world and perturbed data.
Somboon Hongeng, Ramakant Nevatia
ICCV2
2003 Car detection in low resolution aerial images
Tao Zhao 0001, Ramakant Nevatia
Image Vis. Comput.2
2003 Improved Rooftop Detection in Aerial Images with Machine Learning
Marcus A. Maloof, Pat Langley, Thomas O. Binford, Ramakant Nevatia, Stephanie Sage
Mach. Learn.4
2003 Expandable Bayesian Networks for 3D Object Description from Multiple Views and Multiple Mode Inputs
abstract
Computing 3D object descriptions from images is an important goal of computer vision. A key problem here is the evaluation of a hypothesis based on evidence that is uncertain. There have been few efforts on applying formal reasoning methods to this problem. In multiview and multimode object description problems, reasoning is required on evidence features extracted from multiple images and nonintensity data. One challenge here is that the number of the evidence features varies at runtime because the number of images being used is not fixed and some modalities may not always be available. We introduce an augmented Bayesian network, the expandable Bayesian network (EBN), which instantiates its structure at runtime according to the structure of input. We introduce the use of hidden variables to handle correlation of evidence features across images. We show an application of an EBN to a multiview building description system. Experimental results show that the proposed method gives significant and consistent performance improvement to others.
Zu Whan Kim, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 Automatic and interactive modeling of buildings in urban environments from aerial images
abstract
Automatically extracting object models from images is a complex task. We describe research in extracting 3D models of buildings from aerial images. This work has resulted in several related systems including assisted extraction (minimal manual interaction to guide automatic processing), automatic extraction with limited imagery and limited building models, and automatic extraction with very good imagery and digital elevation models and more complex building models. Some results are provided for the assisted system and one of the automatic systems.
Ramakant Nevatia, Keith E. Price
ICIP (3)1
2002 Automatic Pose Estimation of Complex 3D Building Models
abstract
3D models of urban sites with geometry and facade textures are needed for many planning and visualization applications. Approximate 3D wireframe model can be derived from aerial images but detailed textures must be obtained from ground level images. Integrating such views with the 3D models is difficult as only small parts of buildings may be visible in a single view. We describe a method that uses two or three vanishing points, and three 3D to 2D line correspondences to estimate the rotational and translational parameters of the ground level cameras. The valid set of multiple combinations of 3D to 2D line pairs is chosen by a hypotheses generation and evaluation Some experimental results are presented.
Sung Chun Lee, Soon Ki Jung, Ramakant Nevatia
WACV3
2002 Automatic Integration of Facade Textures into 3D Building Models with a Projective Geometry Based Line Clustering
abstract
Visualization of city scenes is important for many applications including entertainment and urban mission planning. Models covering wide areas can be efficiently constructed from aerial images. However, only roof details are visible from aerial views; ground views are needed to provide details of the building facades for high quality 'fly-through' visualization or simulation applications. We present an automatic method of integrating facade textures from ground view images into 3D building models for urban site modeling. We first segment the input image into building facade regions using a hybrid feature extraction method, which combines global feature extraction with Hough transform on an adaptively tessellated Gaussian Sphere and local region segmentation. We estimate the external camera parameters by using the corner points of the extracted facade regions to integrate the facade textures into the 3D building models. We validate our approach with a set of experiments on some urban sites. Categories and Subject Descriptors (according to ACM CCS): I.3.3 [Computer Graphics]: Modeling packages
Sung Chun Lee, Soon Ki Jung, Ramakant Nevatia
Comput. Graph. Forum3
2001 Automatic Description of Buildings with Complex Rooftops from Multiple Images
abstract
We present a model-based approach to detecting and describing compositions of buildings with complex rooftops. Previous approaches have dealt with either simpler models or models which lack geometric information. In spite of increasing model complexity, we maintain the computation affordable by effectively using multiple overlapping images. We obtain rooftop hypotheses in 3-D by using 3-D lines and junctions generated from multiple images. Image-derived unedited elevation data is used to assist feature matching, and to generate rough cues of the presence of 3-D structures. Experimental results are shown on complex buildings.
Zu Whan Kim, Andres Huertas, Ramakant Nevatia
CVPR (2)3
2001 Segmentation and Tracking of Multiple Humans in Complex Situations
abstract
Segmenting and tracking multiple humans is a challenging problem in complex situations in which extended occlusion, shadow and/or reflection exists. We tackle this problem with a 3D model-based approach. Our method includes two stages, segmentation (detection) and tracking. Human hypotheses are generated by shape analysis of the foreground blobs using a human shape model. The segmented human hypotheses are tracked with a Kalman filter with explicit handling of occlusion. Hypotheses are verified while being tracked for the first second or so. The verification is done by walking recognition using an articulated human walking model. We propose a new method to recognize walking using a motion template and temporal integration. Experiments show that our approach works robustly in very challenging sequences.
Tao Zhao 0001, Ramakant Nevatia, Fengjun Lv
CVPR (2)2
2001 Multi-Agent Event Recognition
abstract
This paper presents a new approach to recognizing multiagent events observed by a static camera. To track objects robustly, knowledge about the ground plane and the events is used. An event is considered as composed of action threads, each thread being executed by a single actor. A single thread of action is recognized from the characteristics of the trajectory and moving blob of the actor using Bayesian methods. A multi-agent event is represented by a number of action threads related by temporal constraints. Multi-agent events are recognized by propagating the constraints and likelihoods of event threads in a temporal logic network.
Somboon Hongeng, Ramakant Nevatia
ICCV2
2001 Car Detection in Low Resolution Aerial Image
Tao Zhao 0001, Ramakant Nevatia
ICCV2
2001 Event Detection and Analysis from Video Streams
abstract
We present a system which takes as input a video stream obtained from an airborne moving platform and produces an analysis of the behavior of the moving objects in the scene. To achieve this functionality, our system relies on two modular blocks. The first one detects and tracks moving regions in the sequence. It uses a set of features at multiple scales to stabilize the image sequence, that is, to compensate for the motion of the observer, then extracts regions with residual motion and uses an attribute graph representation to infer their trajectories. The second module takes as input these trajectories, together with user-provided information in the form of geospatial context and goal context to instantiate likely scenarios. We present details of the system, together with results on a number of real video sequences and also provide a quantitative analysis of the results.
Gérard G. Medioni, Isaac Cohen, François Brémond, Somboon Hongeng, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.5
2001 Detection and Modeling of Buildings from Multiple Aerial Images
abstract
Automatic detection and description of cultural features, such as buildings, from aerial images is becoming increasingly important for a number of applications. This task also offers an excellent domain for studying the general problems of scene segmentation, 3D inference, and shape description under highly challenging conditions. We describe a system that detects and constructs 3D models for rectilinear buildings with either flat or symmetric gable roofs from multiple aerial images; the multiple images, however, need not be stereo pairs (i.e., they may be acquired at different times). Hypotheses for rectangular roof components are generated by grouping lines in the images hierarchically; the hypotheses are verified by searching for presence of predicted walls and shadows. The hypothesis generation process combines the tasks of hierarchical grouping with matching at successive stages. Overlap and containment relations between 3D structures are analyzed to resolve conflicts. This system has been tested on a large number of real examples with good results, some of which are included in the paper along with their evaluations.
Sanjay Noronha, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 Representation and Optimal Recognition of Human Activities
abstract
Towards the goal of realizing a generic automatic human activity recognition system, a new formalism is proposed. Activities are described by a chained hierarchical representation using three type of entities: image features, mobile object properties and scenarios. Taking image features of tracked moving regions from an image sequence as input, mobile object properties are first computed by specific methods while noise is suppressed by statistical methods. Scenarios are recognized from mobile object properties based on Bayesian analysis. Several scenarios are recognized by an algorithm using a probabilistic finite-state automaton (a variant of structured HMM). A demonstration of the optimality of this recognition method is discussed. Finally, the validity and the effectiveness of our approach is demonstrated on both real-world and perturbed data.
Somboon Hongeng, François Brémond, Ramakant Nevatia
CVPR3
2000 Multisensor Integration for Building Modeling
abstract
Machine perception can benefit from the use of features extracted from data provided by a variety of sensor modalities. Recent advances in sensor design makes it possible to incorporate multiple sensors into vision systems for increased capability. Two important issues must be considered for the integration task. The sensors must be spatially coregistered and the phenomenologies must be compatible. In this paper we address these issues as they apply to the problem of automatic modeling of building structures from aerial views. We present a methodology to incorporate cues extracted from IFSAR (Interferometric Synthetic Aperture Radar) to significantly improve the performance and the quality of the results of an existing system that relies on electro-optical panchromatic images, while reducing processing time. Quantitative evaluations are given.
Andres Huertas, Zu Whan Kim, Ramakant Nevatia
CVPR3
2000 Learning Bayesian Networks for Diverse and Varying numbers of Evidence Sets
Zu Whan Kim, Ramakant Nevatia
ICML2
2000 Bayesian Framework for Video Surveillance Application
abstract
The goal of this paper is to describe and demonstrate the application of Bayesian networks in a generic automatic video surveillance system. Taking image features of tracked moving regions from an image sequence as input, mobile object properties are first computed and noise is suppressed by statistical methods. The probability that a scenario occurs is then computed from these mobile object properties through several layers of naive Bayesian classifiers (or a Bayesian network). Several issues and solutions regarding the efficiency of the Bayesian network are discussed. For example, the parameters of the networks, which represent rare activities (typical of video surveillance applications), can be learned from image sequences of similar scenarios which are more common. We demonstrate the effectiveness of our approach by training the networks with 600 image frames belonging to one domain of interest and applying them to image sequences in a different domain.
Somboon Hongeng, François Brémond, Ramakant Nevatia
ICPR3
2000 Automatic description of complex buildings with multiple images
abstract
3-D building detection and description is a practical application of 3-D object description, a key task of computer vision. We present an approach to detecting and describing buildings of polygonal rooftops by using multiple, overlapping images of the scene. First, 3-D features are generated by using multiple images, and rooftop hypotheses are generated by neighborhood searches on those features. For robust generation of 3-D features, we present a probabilistic approach to address the epipolar alignment problem in line matching. Image-derived unedited elevation data is used to assist feature matching, and to generate rough cues of the presence of 3-D structures. These cues help reduce the search space significantly. Experimental results are shown on some complex buildings.
Zu Whan Kim, Andres Huertas, Ramakant Nevatia
WACV3
2000 Modeling 3-D complex buildings with user assistance
abstract
An effective 3D method incorporating user assistance for modeling complex buildings is proposed. This method utilizes the connectivity and similar structure information among unit blocks in a multi-component building structure, to enable the user to incrementally construct models of many types of buildings. The system attempts to minimize the time and the number of user interactions needed to assist an existing automatic system in this task. Several examples are presented that demonstrate significant improvement and efficiency compared with other approaches and with purely manual systems.
Sung Chun Lee, Andres Huertas, Ramakant Nevatia
WACV3
2000 Detecting changes in aerial views of man-made structures
Andres Huertas, Ramakant Nevatia
Image Vis. Comput.2
1999 User Assisted Modeling of Buildings from Aerial Images
abstract
An approach that allows a user to assist an automatic system in modeling buildings is described. The approach is designed to be efficient in user time and effort while preserving the quality of the models created. Currently our system is able to handle the rectangular buildings with flat roof or symmetric gabled roof. Models can be created by only one or two clicks in many cases. Efficient editing of automatically derived models is also possible.
Ramakant Nevatia, Sanjay Noronha
CVPR2
1999 Uncertain Reasoning and Learning for Feature Grouping
Zu Whan Kim, Ramakant Nevatia
Comput. Vis. Image Underst.2
1999 Part-Based 3D Descriptions of Complex Objects from a Single Image
abstract
Volumetric, 3D, part-based descriptions of complex objects in a scene can be highly beneficial for many tasks such as generic object recognition, navigation, and manipulation. However, it has been difficult to derive such descriptions from image data. There has been some progress in getting such descriptions from range data or from perfect contours, but analysis of a real intensity image presents many difficulties. The object and part boundaries do not completely correspond to image boundaries. The detected boundaries are often fragmented and many boundaries due to surface markings, shadows, and noise are present. In addition, inference of 3D from a 2D image is difficult. The paper describes a method to compute the desired descriptions from a single image by exploiting projective properties of a class of generalized cylinders and of possible joints between them. Experimental results on some examples are given.
Mourad Zerroug, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.2
1998 Recent Advances in Detection and Description of Buildings from Multiple Aerial Images
Sanjay Noronha, Ramakant Nevatia
ACCV (2)2
1998 Detecting Changes in Aerial Views of Man-Made Structures
abstract
Many applications require detecting structural changes in a scene over a period of time. Comparing intensity values of successive images is not effective as such changes don't necessarily reflect actual changes at a site but might be caused by changes in the view point, illumination and seasons. We take the approach of comparing a 3-D model of the site, prepared from previous images, with new images to infer significant changes. This task is difficult as the images and the models have very different levels of abstract representations. Our approach consists of several steps: registering a site model to a new image, model validation to confirm the presence of model objects in the image; structural change detection seeks to resolve matching problems and indicate possibly changed structures; and finally updating models to reflect the changes. Our system is able to detect missing (or mis-modeled) buildings, changes in model dimensions, and new buildings under some conditions.
Andres Huertas, Ramakant Nevatia
ICCV2
1998 Generalizing over aspect and location for rooftop detection
abstract
We present the results of an empirical study in which we evaluated cost-sensitive learning algorithms on a rooftop detection task, which is one level of processing in a building detection system. Specifically, we investigated how well machine learning methods generalized to unseen images that differed in location and in aspect. For the purpose of comparison, we included in our evaluation a handcrafted linear classifier, which is the selection heuristic currently used in the building detection system. ROC analysis showed that, when generalizing to unseen images that differed in location and aspect, a naive Bayesian classifier outperformed nearest neighbor and the handcrafted solution.
Marcus A. Maloof, Pat Langley, Thomas O. Binford, Ramakant Nevatia
WACV4
1998 Automatic Building Extraction from Aerial Images
Armin Grün, Ramakant Nevatia
Comput. Vis. Image Underst.2
1998 Building Detection and Description from a Single Intensity Image
Chungan Lin, Ramakant Nevatia
Comput. Vis. Image Underst.2
1998 Recognition and localization of generic objects for indoor navigation using functionality
Ramakant Nevatia
Image Vis. Comput.2
1997 Detection and Description of Buildings from Multiple Aerial Images
abstract
A method for detection and description of rectangular buildings from two or more registered aerial intensity images is proposed. The output is a 3D description of the buildings, with an associated confidence measure for each building. Hierarchical perceptual grouping and matching across views is employed to increase the robustness of the system. Verification of selected building hypotheses is done using shadow and wall evidence of the buildings. The system is largely feature-based. Grouping and matching are performed in a hierarchical manner utilizing primitives of increasing complexity, starting with line segments and junctions, and proceeding to higher level features. Binocular and trinocular epipolar constraints are used to reduce the search space for matching features.
Sanjay Noronha, Ramakant Nevatia
CVPR2
1996 Load balancing strategies for symbolic vision computations
abstract
Most intermediate and high-level vision algorithms manipulate symbolic features. A key operation in these vision algorithms is to search symbolic features satisfying certain geometric constraints. Parallelizing this symbolic search needs a non-trivial algorithmic technique due to the unpredictable workload. In this paper, we propose load balancing strategies for parallelizing symbolic search operations on distributed memory machines. By using an initial workload estimate, we first partition the computations such that the workload is distributed evenly across the processors. In addition, we perform fast migrations dynamically to adapt to the evolving workload. To demonstrate the usefulness of our load balancing strategies, experiments were conducted on an IBM SP2 and a Cray T3D. Our results show that our task migration strategy can balance the unpredictable workload with little overhead. Our code using C and MPI is portable onto other high performance computing platforms.
Yongwha Chung, Jongwook Woo, Ramakant Nevatia, Viktor Prasanna 0001
HiPC3
1996 Recovering LSHGCs and SHGCs from stereo
Ronald Chung, Ramakant Nevatia
Int. J. Comput. Vis.2
1996 Computer vision research at the University of Southern California
Ramakant Nevatia, Gérard G. Medioni
Int. J. Comput. Vis.1
1996 Volumetric descriptions from a single intensity image
Mourad Zerroug, Ramakant Nevatia
Int. J. Comput. Vis.2
1996 Three-Dimensional Descriptions Based on the Analysis of the Invariant and Quasi-Invariant Properties of Some Curved-Axis Generalized Cylinders
abstract
We address the recovery of object-level 3D descriptions of some classes of curved-axis generalized cylinders. For this, the first part of the paper analyzes the projective properties of two common generic shapes, planar right constant generalized cylinders (PRCGCs) and circular planar right generalized cylinders (circular PRGCs). The properties we analyze include new geometric invariant and quasi-invariant properties of the orthographic projection of the above shapes and a useful classification of their structural properties as functions of their pose. The second part of the paper describes an implemented system which detects and recovers PRCGCs and circular PRGCs from an intensity image in the presence of noise, surface markings, shadows, and partial occlusion. The methods exploit the projective properties to hypothesize and verify relevant curved-axis objects, thus explicitly using the three-dimensionality of the objects and of the desired descriptions. This work extends past work on the recovery of volumetric shapes from an intensity image by addressing new primitives, deriving new properties and by developing a system that recovers them from an intensity image. We demonstrate our method on several real intensity images.
Mourad Zerroug, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.2
1995 Use of Monocular Groupings and Occlusion Analysis in a Hierarchical Stereo System
Ronald Chung, Ramakant Nevatia
Comput. Vis. Image Underst.2
1995 Shape from Contour: Straight Homogeneous Generalized Cylinders and Constant Cross Section Generalized Cylinders
abstract
We analyze the properties of straight homogeneous generalized cylinders (SHGCs) and constant cross section generalized cylinders (CGCs), and derive the types of symmetries that the limb boundaries and cross sections of these objects produce on the image plane. The constraints on the 3D shape of the objects are formulated based on the symmetries and from the geometry of the projection models. Finally, the methods that recover the 3D shape from the image of their contours are discussed and recovered surfaces are shown for sample objects.>
Fatih Ulupinar, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.2
1994 Representation and computation of the spatial environment for indoor navigation
abstract
We introduce a spatial representation, s-map, for an indoor navigation robot. The s-map represents the locations of obstacles in a planar domain, where obstacles are defined as any objects that can block movement of the robot. In building the s-map, the viewing triangle constraint and the stability constraint are introduced for efficient verification of vertical surfaces. These verified vertical surfaces and 3-D segments of obstacles smaller than a robot, are mapped to the s-map by simply dropping height information. Thus, the s-map is made directly from 3-D segments with simple verification, and represents obstacles in a planar domain so that it becomes a navigable map for the robot without further processing. In addition to efficient map building, the s-map represents the environment more realistically and completely. Furthermore, the s-map converts many navigation problems in 3-D, such as map fusion and path planning, into 2-D ones. We present the analysis of the s-map in terms of complexity and reliability, and discuss its pros and cons. Moreover, we show the results of the s-maps for indoor environments.>
Ramakant Nevatia
CVPR2
1994 Detection of buildings using perceptual grouping and shadows
abstract
We describe a system for detection and description of buildings in aerial scenes. This is a difficult task as the aerial images contain a variety of objects. Low-level segmentation processes give highly fragmented segments due to a number of reasons. We use a perceptual grouping approach to collect these fragments and discard those that come from other sources. We use shape properties of the buildings for this. We use shadows to help form and verify the hypotheses generated by the grouping process. This latter step also provides 3-D descriptions of the buildings. Our system has been tested on a number of examples and is able to work with overhead or oblique views.>
Chungan Lin, Andres Huertas, Ramakant Nevatia
CVPR3
1994 Segmentation and Recovery of SHGCs from a Real Intensity Image
Mourad Zerroug, Ramakant Nevatia
ECCV (1)2
1994 Parallel processing for spatial grouping and matching
abstract
In this paper, we identify the computational requirements for structural pattern analysis, particularly for the operations of spatial grouping and matching. We describe two such algorithms that are in wide use here at USC and discuss approaches to reducing their execution times via parallel implementation. We provide brief descriptions and results of two research projects geared generally, toward the parallel implementation of computer vision systems and specifically, towards these algorithms.
Ramakant Nevatia, Craig C. Reinhart
ICPR (3)1
1994 From an intensity image to 3-D segmented descriptions
abstract
Addresses the inference of 3-D segmented descriptions of complex objects from a single intensity image. The authors' approach is based on the analysis of the projective properties of a small number of generalized cylinder primitives and their relationships in the image which make up common man-made objects. Past work on this problem has either assumed perfect contours as input or used 2-dimensional shape primitives without relating them to 3-D shape. The method the authors present explicitly uses the 3-dimensionality of the desired descriptions and directly addresses the segmentation problem in the presence of contour breaks, markings shadows and occlusion. This work has many significant applications including recognition of complex curved objects from a single real intensity image. The authors demonstrate their method on real images.
Mourad Zerroug, Ramakant Nevatia
ICPR (1)2
1994 Segmentation and 3-D recovery of curved-axis generalized cylinders from an intensity image
abstract
Addresses the problem of segmentation and recovery of 3-D object-centered descriptions of two large sub-classes of curved axis generalized cylinders, PRCGCs and circular PRGCs, from a single real intensity image. The purpose of this work is to augment the set of 3-D primitives which can be recovered, beyond previous work which has addressed mainly straight axis ones, so that more complex objects can be handled. The authors' approach is based on the exploitation of geometric projective properties as well as structural properties of the contours of circular PRGCs. The implemented method works in the presence of noise, contour breaks, markings, shadows and occlusion. The authors demonstrate their method on real images.
Mourad Zerroug, Ramakant Nevatia
ICPR (1)2
1994 Model validation for change detection [machine vision]
abstract
An important application of machine vision is to provide a means to monitor a scene over a period of time and report changes in the content of the scene. We have developed a validation mechanism that implements the first step towards a system for detecting changes in images of aerial scenes. By validation we mean the confirmation of the presence of model objects in the image. Our system uses a 3-D site model of the scene as a basis for model validation, and eventually for detecting changes and to update the site model. The scenario for our present validation system consists of adding a new image to a database associated with the site. The validation process is implemented in three steps: registration of the image to the model, or equivalently, determination of the position and orientation of the camera; matching of model features to image features; and validation of the objects in the model. Our system processes the new image monocularly and uses shadows as 3-D clues to help validate the model. The system has been tested using a hand-generated site model and several images of a 500:1 scale model of the site, acquired form several viewpoints.>
Mathias Bejanin, Andres Huertas, Gérard G. Medioni, Ramakant Nevatia
WACV4
1994 A method for recognition and localization of generic objects for indoor navigation
abstract
We introduce an efficient method for recognition and localization of generic objects for robot navigation, which works on real scenes. The generic objects used in our experiments are desks and doors as they are suitable landmarks for navigation. The recognition method uses significant surfaces and accompanying functional evidence for recognition of such objects. Currently, our system works with planar surfaces only and assumes that the objects are in a "standard" pose. The localization and orientation of an object are represented with the most significant surface in an "s-map". Some results for laboratory scenes are given.>
Ramakant Nevatia
WACV2
1994 Recovery of 3-D Objects with Multiple Curved Surfaces from 2-D Contours
Fatih Ulupinar, Ramakant Nevatia
Artif. Intell.2
1993 Quasi-invariant properties and 3-D shape recovery of non-straight, non-constant generalized cylinders
abstract
The geometric protective properties of the contours of right generalized cylinders with a planar, but not necessarily straight, axis and circular, but possibly varying in size, cross-sections (called circular PRGCs) are addressed. Important rigourous quasi-invariant properties of circular PRGCs and invariant properties for their subclasses are derived. Their application for 2-D descriptions and for recovery of complete 3-D object-centered descriptions from the 2-D contours is shown.>
Mourad Zerroug, Ramakant Nevatia
CVPR2
1993 Perception of 3-D Surfaces from 2-D Contours
abstract
Inference of 3-D shape from 2-D contours in a single image is an important problem in machine vision. The authors survey classes of techniques proposed in the past and provide a critical analysis. They show that two kinds of symmetries in figures, which are known as parallel and skew symmetries, give significant information about surface shape for a variety of objects. They derive the constraints imposed by these symmetries and show how to use them to infer 3-D shape. They also discuss the zero Gaussian curvature (ZGC) surfaces in depth and show results on the recovery of surface orientation for various ZGC surfaces.>
Fatih Ulupinar, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.2
1992 Recovering LSHGCs and SHGCs from stereo
abstract
The problem of computing volumetric shape from stereo is examined. It is argued that intermediate two-and-one-half-dimensional dense or wire-frame descriptions may not be always possible from stereo, especially when there are curved surfaces in the scene, and that 3D volumetric descriptions of objects may have to be derived directly from stereo correspondences. Methods are then presented for recovering volumetric shape with linear straight homogeneous generalized cones (LSHGCs) and straight homogeneous generalized cones (SHGCs) as the shape models, using some invariant properties in their monocular and stereo projections. Experimental results on images of objects with curved surfaces are given.>
Ronald C.-K. Chung, Ramakant Nevatia
CVPR2
1992 Recovery of 3-D objects with multiple curved surfaces from 2-D contours
abstract
The authors describe a technique for inference of 3-D shape from 2-D contours that utilizes not only the shapes of individual surfaces but also the interactions between them. The analysis applies to objects made of zero-Gaussian curvature surfaces viewed under orthographic projection.>
Fatih Ulupinar, Ramakant Nevatia
CVPR2
1992 Description and tracking of moving articulated objects
abstract
Proposes a method to obtain reliable shape description of articulated objects by integrating initial descriptions computed from different view images. Ribbon, which is a 2-D analog of a generalized cone, is used as the basic shape representation scheme. An initial description for each frame is the collection of composed ribbons, which is obtained after filtering out most inadequate ribbons and grouping the remaining ribbons by using geometric constraints. Ribbon matching is then conducted between different frames and ribbons which match are retained. From the retained ribbons, the geometric constraints make integrated descriptions and the tracking of parts is established from the ribbon matching results. Since the ribbon matching allows one ribbon to be matched with two ribbons, an articulation which is not detected in one frame but is detected in another frame can be recovered. Experimental results are also shown.>
Shoji Kurakake, Ramakant Nevatia
ICPR (1)2
1992 Issues in parallel tree search for object recognition
abstract
The authors describe the parallel implementation of a 3D object recognition algorithm. The algorithm is representative of methods utilized by various computer vision researchers and presents some interesting problems that are generally overlooked by parallel processing researchers that have studied tree search problems. They describe their objectives in developing the parallel implementation and discuss its performance. They also (briefly) discuss the affects that the parallel implementation has on the runtime characteristics of the algorithm.>
Craig C. Reinhart, Ramakant Nevatia
ICPR (4)2
1992 Recovering building structures from stereo
abstract
Addresses the problem of extracting polyhedral building structures from a stereo pair of aerial intensity images. The authors describe a system that computes a hierarchy of descriptions such as segments, junctions, and links between junctions from each view, and matches these features at the different levels. Such high level features not only help reduce correspondence ambiguity during stereo matching, but also allow us to infer surface boundaries even though the boundaries may be broken because of noise and weak contrast. The authors hypothesize surface boundaries by examining global information such as continuity and coplanarity of linked edges in 3-D, rather than merely by looking at local depth information. When the walls of the buildings are visible, they also exploit the relationship among adjacent surfaces in a polyhedral object to help confirm the different levels of descriptions. The authors give some experimental results for aerial images taken from overhead views and oblique views.>
Ronald Chung, Ramakant Nevatia
WACV2
1992 Perceptual Organization for Scene Segmentation and Description
abstract
A data-driven system for segmenting scenes into objects and their components is presented. This segmentation system generates hierarchies of features that correspond to structural elements such as boundaries and surfaces of objects. The technique is based on perceptual organization, implemented as a mechanism for exploiting geometrical regularities in the shapes of objects as projected on images. Edges are recursively grouped on geometrical relationships into a description hierarchy ranging from edges to the visible surfaces of objects. These edge groupings, which are termed collated features, are abstract descriptors encoding structural information. The geometrical relationships employed are quasi-invariant over 2-D projections and are common to structures of most objects. Thus, collations have a high likelihood of corresponding to parts of objects. Collations serve as intermediate and high-level features for various visual processes. Applications of collations to stereo correspondence, object-level segmentation, and shape description are illustrated.>
Rakesh Mohan, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.2
1991 Use of monocular groupings and occlusion analysis in a hierarchical stereo system
abstract
A hierarchical stereo system is described that uses structural descriptions up to the surface level. Surface descriptions are computed from monocular images, by using a perceptual grouping technique. Occlusion can be a major problem in stereo analysis and is often not treated explicitly. An analysis is presented of occlusion effects in stereo, and it is shown how structural descriptions can be used to deal with them. Experimental results are given for scenes with curved objects and significant occlusions.>
Ronald Chung, Ramakant Nevatia
CVPR2
1991 Recovering shape from contour for constant cross section generalized cylinders
abstract
The authors analyze the properties of constant cross section generalized cylinders (CGCs), and derive the types of symmetries that the limb boundaries and cross sections of these objects produce on the image plane. The constraints on the 3-D shape of the objects are formulated based on the symmetries and from the geometry of the projection models. The methods that recover the 3-D shape from the image of their contours are discussed and recovered surfaces are shown for sample objects.>
Fatih Ulupinar, Ramakant Nevatia
CVPR2
1991 Constraints for interpretation of line drawings under perspective projection
Fatih Ulupinar, Ramakant Nevatia
CVGIP Image Underst.2
1990 Shape from contour: straight homogeneous generalized cones
abstract
The authors analyze the properties of straight homogeneous generalized cones (SHGCs) and derive the types of symmetries, that the limb boundaries and cross sections of these objects produce on the image plane. The constraints on the 3-D shape of the objects are formulated based on the symmetries and from the geometry of the projection models. Finally the methods that recover the 3-D shape from the image of their contours are discussed and recovered surfaces are shown for sample objects.>
Fatih Ulupinar, Ramakant Nevatia
ICCV2
1990 Shape description from imperfect and incomplete data
abstract
Usually, shape description systems assume that a scene has been segmented into objects and that object boundaries are given. This, however, is not realistic when working with intensity images; the resulting boundaries are fragmented and contain surface markings, and shadow and noise boundaries. A system is described which works with such input and computes shape descriptions of complex objects. Scene segmentation takes place through shape description. Generalized cones or, more precisely, their 2D analogs of ribbons are used as the basic shape representation scheme. Results for synthetic and real examples are shown. The output of the system is useful for object recognition, learning, further inference of 3D shape, grasping, and navigation.>
Kashipati Rao, Ramakant Nevatia
ICPR (1)2
1990 Inferring shape from contour for curved surfaces
abstract
A technique based on analysis of symmetries in an image is proposed for inferring the 3-D shapes of surfaces of objects in it. This technique is analyzed and applied to zero-Gaussian-curvature surfaces. The method consists of deriving a number of constraints based on a few simple assumptions. The combination of constraints to give unique (or few) solutions is discussed. Experimental results on selected scenes are given and are shown to conform well with human perception.>
Fatih Ulupinar, Ramakant Nevatia
ICPR (1)2
1990 Detecting runways in complex airport scenes
Andres Huertas, William Cole, Ramakant Nevatia
Comput. Vis. Graph. Image Process.3
1989 Segmentation and description based on perceptual organization
abstract
The authors present a description framework, motivated by perceptual organization, which consists of representations of the geometrical organizations of intensity discontinuities. The descriptors in this framework are called collated features, and are groupings identified by perceptual organization. The processes that operate on the image to obtain these descriptors and the visual processes that utilize them are discussed. The detection of collated features is robust to local problems. The structural information encoded in them aids various visual tasks such as object segmentation, correspondence processes (stereo, motion, and model matching), and shape inferences. Two primary grouping processes, cocurvilinearity and symmetry are applied to intensity edge contours to generate the collated features, including curves, symmetries, and ribbons. These collations can be used to segment into visible surfaces of objects and to describe the 2D shapes of those surfaces.>
Rakesh Mohan, Ramakant Nevatia
CVPR2
1989 Using Generic Knowledge in Analysis of Aerial Scenes: A Case Study
Andres Huertas, William Cole, Ramakant Nevatia
IJCAI3
1989 Dissertation abstracts
Steven L. Tanimoto, Sargur N. Srihari, Martin D. Levine, Warren P. Seering, Ramakant Nevatia
Mach. Vis. Appl.5
1989 Recognizing 3-D Objects Using Surface Descriptions
abstract
The authors provide a complete method for describing and recognizing 3-D objects, using surface information. Their system takes as input dense range date and automatically produces a symbolic description of the objects in the scene in terms of their visible surface patches. This segmented representation may be viewed as a graph whose nodes capture information about the individual surface patches and whose links represent the relationships between them, such as occlusion and connectivity. On the basis of these relations, a graph for a given scene is decomposed into subgraphs corresponding to different objects. A model is represented by a set of such descriptions from multiple viewing angles, typically four to six. Models can therefore be acquired and represented automatically. Matching between the objects in a scene and the models is performed by three modules: the screener, in which the most likely candidate views for each object are found; the graph matcher, which compares the potential matching graphs and computes the 3-D transformation between them; and the analyzer, which takes a critical look at the results and proposes to split and merge object graphs.>
Ting-Jun Fan, Gérard G. Medioni, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.3
1989 Stereo Error Detection, Correction, and Evaluation
abstract
An algorithm is presented for error detection and correction of disparity, as a process separate from stereo matching, with the contention that matching is not necessarily the best way to utilize all the physical constraints characteristic to stereopsis. As a result of the bias in stereo research towards matching, vision tasks like surface interpolation and object modeling have to accept erroneous data from the stereo matchers without the benefits of any intervening stage of error correction. An algorithm which identifies all errors in disparity data that can be detected on the basis of figural continuity and corrects them is presented. The algorithm can be used as a postprocessor to any edged-based stereo matching algorithm, and can additionally be used to automatically provide quantitative evaluations on the performance of matching algorithms of this class.>
Rakesh Mohan, Gérard G. Medioni, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.3
1989 Using Perceptual Organization to Extract 3-D Structures
abstract
The authors describe an approach to perceptual grouping for detecting and describing 3-D objects in complex images and apply it to the task of detecting and describing complex buildings in aerial images. They argue that representations of structural relationships in the arrangements of primitive image features, as detected by the perceptual organization process, are essential for analyzing complex imagery. They term these representations collated features. The choice of collated features is determined by the generic shape of the desired objects in the scene. The detection process for collated features is more robust than the local operations for region segmentation and contour tracing. The important structural information encoded in collated features aids various visual tasks such as object segmentation, correspondence processes, and shape description. The proposed method initially detects all reasonable feature groupings. A constraint satisfaction network is then used to model the complex interactions between the collations and select the promising ones. Stereo matching is performed on the collations to obtain height information. This aids in further reasoning on the collated features and results in the 3-D description of the desired objects.>
Rakesh Mohan, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.2
1988 Using Symmetries For Analysis Of Shape From Contour
abstract
Inference of 3-D shape from 2-D contours in a single image is an important problem in machine vision. We survey classes of techniques proposed in the past and provide a critical analysis. We propose two kinds of symmetries in figures, which we call parallel and mirror symmetries, give significant information about surface shape for a variety of objects. We show the constraints imposed by these symmetries and how to use them to infer 3-D shape. Our method is applicable to any zero-gaussian curvature surface, and also to a variety of doubly curved surfaces. One of our mathematical results is that for a cone, the surface shape can be constructed uniquely under very simple assumptions. We also show some preliminary results on extraction of symmetries from real images.
Fatih Ulupinar, Ramakant Nevatia
ICCV2
1988 Matching 3-D objects using surface descriptions
abstract
A method is developed to extract important curves, corresponding to physical boundaries of objects, from a range image. It is shown how to infer, from these curves, a segmentation of the scene into surface patches, and how to use these descriptions to establish correspondences between two scenes. In a first step, labeled curves corresponding to jump boundaries, creases, and limbs of objects are grouped into boundaries of regions. Each region is therefore described by its boundaries and by a polynomial approximation, which allows each path individually and also the complete scene to be reconstructed. In a second step, objects (or partial objects) are inferred from surface patches, and then two range images are at this partial object level. Graphs of objects are matched using a best-first search under three types of constraints: unary constraints between corresponding nodes, binary constraints between corresponding linked pairs of nodes, and constraints imposed by the computed geometric transformation. Substantial partial occlusion is allowed. The generality and robustness of this approach is illustrated by several examples.>
Ting-Jun Fan, Gérard G. Medioni, Ramakant Nevatia
ICRA3
1988 Detecting buildings in aerial images
Andres Huertas, Ramakant Nevatia
Comput. Vis. Graph. Image Process.2
1988 Computing volume descriptions from sparse 3-D data
Kashipati Rao, Ramakant Nevatia
Int. J. Comput. Vis.2
1987 Detecting Runways in Aerial Images
Andres Huertas, William Cole, Ramakant Nevatia
AAAI3
1987 Segmented descriptions of 3-D surfaces
abstract
A method to segment and describe visible surfaces of three-dimensional (3-D) objects is presented by first segmenting the surfaces into simple surface patches and then using these patches and their boundaries to describe the 3-D surfaces. First, distinguished points are extracted which will comprise the edges of segmented surface patches, using the zero-crossings and extrema of curvature along a given direction. Two different methods are used: if the sensor provides relatively noise-free range images, the principal curvatures are computed at only one resolution, otherwise, a multiple scale approach is used and curvature is computed in four directions 45° apart to facilitate interscale tracking. These points are then grouped into curves and these curves are classified into different classes which correspond to significant physical properties such as jump boundaries, folds, and ridge lines (or smooth extrema). Then jump boundaries and folds are used to segment the surfaces into surface patches, and a simple surface is fitted to each patch to reconstruct the original objects. These descriptions not only make explicit most of the salient properties present in the original input, but are more suited to further processing, such as matching with a given model. The generality and robustness of this approach is illustrated on scene images with different available range sensors.
Ting-Jun Fan, Gérard G. Medioni, Ramakant Nevatia
IEEE J. Robotics Autom.3
1986 Structural Analysis of Natural Textures
abstract
Many textures can be described structurally, in terms of the individual textural elements and their spatial relationships. This paper describes a system to generate useful descriptions of natural textures in these terms. The basic approach is to determine an initial, partial description of the elements using edge features. This description controls the extraction of the texture elements. The elements are grouped by type, and spatial relationships between elements are computed. The descriptions are shown to be useful for recognition of the textures, and for reconstruction of periodic textures.
Felicia M. Vilnrotter, Ramakant Nevatia, Keith E. Price
IEEE Trans. Pattern Anal. Mach. Intell.2
1985 Segment-based stereo matching
Gérard G. Medioni, Ramakant Nevatia
Comput. Vis. Graph. Image Process.2
1984 Matching Images Using Linear Features
abstract
We describe techniques for matching two images or an image and a map. This operation is basic for machine vision and is needed for the tasks of object recognition, change detection, map up-dating, passive navigation, and other tasks. Our system uses line-based descriptions, and matching is accomplished by a relaxation operation which computes most similar geometrical structures. A more efficient variation, called the ``kernel'' method, is also described. We give results on complex aerial images which contain many image differences, caused by varying sun position, different seasons, and imaging environments, and also structural changes caused by man-made alterations such as new construction.
Gérard G. Medioni, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.2
1984 Visual inspection using linear features
Gérard G. Medioni, Ramakant Nevatia
Pattern Recognit.3
1983 Detection of Buildings in Aerial Images Using Shape and Shadows
Andres Huertas, Ramakant Nevatia
IJCAI2
1982 Locating Structures in Aerial Images
abstract
A technique for locating desired structures utilizing user specified information about properties of these structures and their relationships with other more easily extracted objects is described. An edge-based and region-based technique is used for scene segmentation. Experimental results of the processing of aerial pictures are presented.
Ramakant Nevatia, Keith E. Price
IEEE Trans. Pattern Anal. Mach. Intell.1
1979 Describing Natural Textures
Ramakant Nevatia, Keith E. Price, Felicia M. Vilnrotter
IJCAI1
1977 Description and Recognition of Curved Objects
Ramakant Nevatia, Thomas O. Binford
Artif. Intell.1
1976 Locating Object Boundaries in Textured Environments
abstract
Detection of object boundaries is an important step in the analysis of an image. In the presence of a textured background, local edge operators generate many edges that do not correspond to the object boundaries. However, edges along the object boundaries link in elongated segments. An efficient algorithm to perform such linking is described and experimental results of a working program are presented.
Ramakant Nevatia
IEEE Trans. Computers1
1973 Structured Descriptions of Complex Objects
Ramakant Nevatia, Thomas O. Binford
IJCAI1