EDBT 2026 Demo / reviewers in the wild / expert
Ramakant Nevatia
dblp:n/RamakantNevatia · also Ram Nevatia
· DBLP profile ↗
229ranked-venue papers
8as first author
22since 2021 · last 2024
0009-0003-8079-4209ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 174 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 173 · 4 first-author · 20 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Large Language Models are Good Prompt Learners for Low-Shot Image ClassificationabstractLow-shot image classification, where training images are limited or inaccessible, has benefited from recent progress on pretrained vision-language (VL) models with strong generalizability. e.g. CLIP. Prompt learning methods built with VL models generate text features from the class names that only have confined class-specific information. Large Language Models (LLMs), with their vast en-cyclopedic knowledge, emerge as the complement. Thus, in this paper, we discuss the integration of LLMs to enhance pretrained VL models, specifically on low-shot classification. However, the domain gap between language and vision blocks the direct application of LLMs. Thus, we propose LLaMp, Large Language Models as Prompt learners, that produces adaptive prompts for the CLIP text encoder, establishing it as the connecting bridge. Experiments show that, compared with other state-of-the-art prompt learning methods, LLaMP yields better performance on both zero-shot generalization and few-shot image classification, over a spectrum of 11 datasets. Code will be made available at: https://github.com/zhaohengz/LLaMP. Zhaoheng Zheng, Jingmin Wei, Xuefeng Hu, Haidong Zhu, Ramakant Nevatia |
CVPR | 5 |
| 2024 | SEAS: ShapE-Aligned Supervision for Person Re-IdentificationabstractWe introduce SEAS, using ShapE-Aligned Supervision, to enhance appearance-based person re-identification. When recognizing an individual's identity, existing methods primarily rely on appearance, which can be influenced by the background environment due to a lack of body shape awareness. Although some methods attempt to incorporate other modalities, such as gait or body shape, they encode the additional modality separately, resulting in extra computational costs and lacking an inherent connection with appearance. In this paper, we explore the use of implicit 3-D body shape representations as pixel-level guidance to augment the extraction of identity features with body shape knowledge, in addition to appearance. Using body shape as supervision, rather than as input, provides shapeaware enhancements without any increase in computational cost and delivers coherent integration with pixel-wise appearance features. Moreover, for video-based person reidentification, we align pixel-level features across frames with shape awareness to ensure temporal consistency. Our results demonstrate that incorporating body shape as pixel-level supervision reduces rank-1 errors by 1.4% for framebased and by 2.5% for video-based re-identification tasks, respectively, and can also be generalized to other existing appearance-based person re-identification methods. Haidong Zhu, Pranav Budhwant, Zhaoheng Zheng, Ramakant Nevatia |
CVPR | 4 |
| 2024 | CaesarNeRF: Calibrated Semantic Representation for Few-Shot Generalizable Neural Rendering
Haidong Zhu, Tianyu Ding, Ilya Zharkov, Ramakant Nevatia, Luming Liang |
ECCV (6) | 5 |
| 2024 | ReCLIP: Refine Contrastive Language Image Pre-Training with Source Free Domain AdaptationabstractLarge-scale pre-trained vision-language models (VLM) such as CLIP [32] have demonstrated noteworthy zero-shot classification capability, achieving 76.3% top-1 accuracy on ImageNet without seeing any examples. However, while applying CLIP to a downstream target domain, the presence of visual and text domain gaps and cross-modality misalignment can greatly impact the model performance. To address such challenges, we propose ReCLIP, a novel source-free domain adaptation method for VLMs, which does not require any source data or target labeled data. ReCLIP first learns a projection space to mitigate the misaligned visual-text embeddings and learns pseudo labels. Then, it deploys cross-modality self-training with the pseudo labels to update visual and text encoders, refine labels and reduce domain gaps and misalignment iteratively. With extensive experiments, we show that ReCLIP outperforms all the baselines significantly and improves the average accuracy of CLIP from 69.83% to 74.94% on 22 image classification benchmarks. Xuefeng Hu, Ke Zhang 0028, Albert Chen 0001, Jiajia Luo, Yuyin Sun, Ken Wang, Nan Qiao 0009, Min Sun 0001, Cheng-Hao Kuo, Ramakant Nevatia |
WACV | 12 |
| 2024 | Efficient Feature Distillation for Zero-shot Annotation Object DetectionabstractWe propose a new setting for detecting unseen objects called Zero-shot Annotation object Detection (ZAD). It expands the zero-shot object detection setting by allowing the novel objects to exist in the training images and restricts the additional information the detector uses to novel category names. Recently, to detect unseen objects, largescale vision-language models (e.g., CLIP) are leveraged by different methods. The distillation-based methods have good overall performance but suffer from a long training schedule caused by two factors. First, existing work creates distillation regions biased to the base categories, which limits the distillation of novel category information. Second, directly using the raw feature from CLIP for distillation neglects the domain gap between the training data of CLIP and the detection datasets, which makes it difficult to learn the mapping from the image region to the vision-language feature space. To solve these problems, we propose Efficient feature distillation for Zero-shot Annotation object Detection (EZAD). Firstly, EZAD adapts the CLIP’s feature space to the target detection domain by re-normalizing CLIP; Secondly, EZAD uses CLIP to generate distillation proposals with potential novel category names to avoid the distillation being overly biased toward the base categories. Finally, EZAD takes advantage of semantic meaning for regression to further improve the model performance. As a result, EZAD outperforms the previous distillation-based methods in COCO by 4% with a much shorter training schedule and achieves a 3% improvement on the LVIS dataset. Our code is available at https://github.com/dragonlzm/EZAD Zhuoming Liu 0001, Xuefeng Hu, Ramakant Nevatia |
WACV | 3 |
| 2024 | Leveraging Task-Specific Pre-Training to Reason across Images and VideosabstractWe explore the task of Reasoning Across Images and Video (RAIV), which requires models to reason on a pair of visual inputs comprising various combinations of images and/or videos. Previous work in this area has been limited to image pairs focusing primarily on the existence and/or cardinality of objects. To address this, we leverage existing datasets with rich annotations to generate semantically meaningful queries about actions, objects, and their relationships. We introduce new datasets that encompass visually similar inputs, reasoning over images, across images and videos, or across videos. Recognizing the distinct nature of RAIV compared to existing pre-training objectives which work on single image-text pairs, we explore task-specific pre-training, wherein a pre-trained model is trained on an objective similar to downstream tasks without utilizing fine-tuning datasets. Experiments with several state-of-the-art pre-trained image-language models reveal that task-specific pre-training significantly enhances performance on downstream datasets, even in the absence of additional pre-training data. We provide further ablative studies to guide future work. Arka Sadhu, Ramakant Nevatia |
WACV | 2 |
| 2024 | CAILA: Concept-Aware Intra-Layer Adapters for Compositional Zero-Shot LearningabstractIn this paper, we study the problem of Compositional Zero-Shot Learning (CZSL), which is to recognize novel attribute-object combinations with pre-existing concepts. Recent researchers focus on applying large-scale Vision-Language Pre-trained (VLP) models like CLIP with strong generalization ability. However, these methods treat the pre-trained model as a black box and focus on pre- and post-CLIP operations, which do not inherently mine the semantic concept between the layers inside CLIP. We propose to dive deep into the architecture and insert adapters, a parameter-efficient technique proven to be effective among large language models, into each CLIP encoder layer. We further equip adapters with concept awareness so that concept-specific features of "object", "attribute", and "composition" can be extracted. We assess our method on four popular CZSL datasets, MIT-States, C-GQA, UT-Zappos, and VAW-CZSL, which shows state-of-the-art performance compared to existing methods on all of them. Zhaoheng Zheng, Haidong Zhu, Ramakant Nevatia |
WACV | 3 |
| 2024 | ShARc: Shape and Appearance Recognition for Person Identification In-the-wildabstractIdentifying individuals in unconstrained video settings is a valuable yet challenging task in biometric analysis due to variations in appearances, environments, degradations, and occlusions. In this paper, we present ShARc, a multimodal approach for video-based person identification in uncontrolled environments that emphasizes 3-D body shape, pose, and appearance. We introduce two encoders: a Pose and Shape Encoder (PSE) and an Aggregated Appearance Encoder (AAE). PSE encodes the body shape via binarized silhouettes, skeleton motions, and 3-D body shape, while AAE provides two levels of temporal appearance feature aggregation: attention-based feature aggregation and averaging aggregation. For attention-based feature aggregation, we employ spatial and temporal attention to focus on key areas for person distinction. For averaging aggregation, we introduce a novel flattening layer after averaging to extract more distinguishable information and reduce overfitting of attention. We utilize centroid feature averaging for gallery registration. We demonstrate significant improvements over existing state-of-the-art methods on public datasets, including CCVID, MEVID, and BRIAR. Haidong Zhu, Wanrong Zheng, Zhaoheng Zheng, Ramakant Nevatia |
WACV | 4 |
| 2023 | AG-ReID 2023: Aerial-Ground Person Re-identification Challenge ResultsabstractPerson re-identification (Re-ID) on aerial-ground platforms has emerged as an intriguing topic within computer vision, presenting a plethora of unique challenges. Highflying altitudes of aerial cameras make persons appear differently in terms of viewpoints, poses, and resolution compared to the images of the same person viewed from ground cameras. Despite its potential, few algorithms have been developed for person re-identification on aerial-ground data, mainly due to the absence of comprehensive datasets. In response, we have collected a large-scale dataset and organized the Aerial-Ground person Re-IDentification Challenge (AG-ReID2023) to foster advancements in the field. The dataset comprises 100,502 images with 1,615 unique identities, including 51,530 training images featuring 807 identities. The test set is divided into two subsets: Aerial to Ground (808 ids, 4,348 query images, 19,259 gallery images) and Ground to Aerial (808 ids, 4,151 query images, 21,214 gallery images). In addition, we manually annotate individuals with their matching IDs across cameras and provide 15 soft attribute labels. The AG-ReID2023 Challenge in conjunction with the 7thIEEE International Joint Conference on Biometrics (IJCB) has garnered interest from numerous institutes, resulting in the submission of five distinct algorithms. We provide an in-depth examination of the evaluation outcomes and present our findings from the contest. For additional details, kindly refer to the official website1.1https://agreid23.github.io. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Dana Michalski, Debayan Deb, Mahak Kothari, Manisha Saini, Dawei Du, Scott McCloskey, Gabriel Bertocco, Fernanda A. Andaló, Terrance E. Boult, Anderson Rocha 0001, Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia, Zaigham A. Randhawa, Sinan Sabri, Gianfranco Doretto |
IJCB | 20 |
| 2023 | GaitRef: Gait Recognition with Refined Sequential SkeletonsabstractIdentifying humans with their walking sequences, known as gait recognition, is a useful biometric understanding task as it can be observed from a long distance and does not require cooperation from the subject. Two common modalities used for representing the walking sequence of a person are silhouettes and joint skeletons. Silhouette sequences, which record the boundary of the walking person in each frame, may suffer from the variant appearances from carried-on objects and clothes of the person. Framewise joint detections are noisy and introduce some jitters that are not consistent with sequential detections. In this paper, we combine the silhouettes and skeletons and refine the framewise joint predictions for gait recognition. With temporal information from the silhouette sequences. We show that the refined skeletons can improve gait recognition performance without extra annotations. We compare our methods on four public datasets, CASIA-B, OUMVLP, Gait3D and GREW, and show state-of-the-art performance. Haidong Zhu, Wanrong Zheng, Zhaoheng Zheng, Ramakant Nevatia |
IJCB | 4 |
| 2023 | Multimodal Neural Radiance FieldabstractThis paper addresses the challenge of reconstructing a scene with a neural radiance field (NeRF) for robot vision and scene understanding using multiple modalities. Researchers have introduced the use of NeRF to represent an object for synthesizing and rendering novel views of complex scenes by optimizing a 3-D radiance field for ray casting and rendering for 2-D RGB images. However, using RGB images alone introduces additional geometry ambiguities with transparent objects or complex scenes and cannot accurately depict the 3-D shapes. We discuss and solve this problem and use multiple modalities as input for the same NeRF model to build a multimodal NeRF by incorporating point clouds and infrared image supervision to prevent such bias. In contrast to RGB images, infrared images and point clouds are typically taken by separate cameras that cannot be aligned with the RGB camera. We further introduce the alignment of different modalities based on point cloud registration to estimate the relative transformation matrices between them before training a NeRF model with multiple modalities. We evaluate our model on chosen scenes from the ScanNet and M2DGR datasets and demonstrate that it outperforms existing state-of-the-art methods. Haidong Zhu, Yuyin Sun, Jiajia Luo, Nan Qiao 0009, Ramakant Nevatia, Cheng-Hao Kuo |
ICRA | 7 |
| 2023 | PatchZero: Defending against Adversarial Patch Attacks by Detecting and Zeroing the PatchabstractAdversarial patch attacks mislead neural networks by injecting adversarial pixels within a local region. Patch attacks can be highly effective in a variety of tasks and physically realizable via attachment (e.g. a sticker) to the real-world objects. Despite the diversity in attack patterns, adversarial patches tend to be highly textured and different in appearance from natural images. We exploit this property and present PatchZero, a general defense pipeline against white-box adversarial patches without retraining the downstream classifier or detector. Specifically, our defense detects adversaries at the pixel-level and "zeros out" the patch region by repainting with mean pixel values. We further design a two-stage adversarial training scheme to defend against the stronger adaptive attacks. PatchZero achieves SOTA defense performance on the image classification (ImageNet, RESISC45), object detection (PASCAL VOC), and video classification (UCF101) tasks with little degradation in benign performance. In addition, PatchZero transfers to different patch shapes and attack types. Zhaoheng Zheng, Kaijie Cai, Ramakant Nevatia |
WACV | 5 |
| 2023 | Gait Recognition Using 3-D Human Body Shape InferenceabstractGait recognition, which identifies individuals based on their walking patterns, is an important biometric technique since it can be observed from a distance and does not require the subject’s cooperation. Recognizing a person’s gait is difficult because of the appearance variants in human silhouette sequences produced by varying viewing angles, carrying objects, and clothing. Recent research has produced a number of ways for coping with these variants. In this paper, we present the usage of inferring 3-D body shapes distilled from limited images, which are, in principle, invariant to the specified variants. Inference of 3-D shape is a difficult task, especially when only silhouettes are provided in a dataset. We provide a method for learning 3-D body inference from silhouettes by transferring knowledge from 3-D shape prior from RGB photos. We use our method on multiple existing state-of-the-art gait baselines and obtain consistent improvements for gait identification on two public datasets, CASIA-B and OUMVLP, on several variants and settings, including a new setting of novel views not seen during training. Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia |
WACV | 3 |
| 2022 | Self-Supervised Learning for Sentiment Analysis via Image-Text MatchingabstractThere is often a resemblance in the sentiment expressed in social media posts (text) and their accompanying images. In this paper, We leverage this sentiment congruence for self-supervised representation learning for sentiment analysis. By teaching the model to pair an image with its corresponding social media post, the model can learn a representation capturing sentiment features from the image and text without supervision. We then use the pre-trained encoder for feature extraction for sentiment analysis in downstream tasks. We show significant improvement and good transferability for sentiment classification in addition to robustness in performance when available data decreases on public datasets (B-T4SA and IMDb Movie Review). With this work, we demonstrate the effectiveness of self-supervised learning through cross-modal matching for sentiment analysis. Haidong Zhu, Zhaoheng Zheng, Mohammad Soleymani 0001, Ramakant Nevatia |
ICASSP | 4 |
| 2022 | Improving Weakly Supervised Scene Graph Parsing through Object GroundingabstractWeakly supervised scene graph parsing, which learns structured image representations without annotated correspondences between graph nodes and visual objects, has been prevalent in recent computer vision research. Existing methods mainly focus on designing task-specific loss functions, model architectures, or optimization algorithms. We argue that correspondences between objects and graph nodes are crucial for the weakly supervised scene graph parsing task and are worth learning explicitly. Thus we propose GroParser, a framework that improves weakly supervised scene graph parsing models by grounding visual objects. The proposed weakly supervised grounding method learns a metric among visual objects and scene graph nodes by incorporating information from both object features and relational features. Specifically, we apply multi-instance learning to learn the object category information and exploit a two-stream graph neural network to model the relational similarity metric. Extensive experiments on the scene graph parsing task verify the grounding found by our model can reinforce the performance of the existing weakly supervised scene graph parsing methods, including the current state-of-the-art. Further experiments on Visual Genome (VG) and Visual Relation Detection (VRD) datasets verify that our model brings an improvement on scene graph grounding task over existing approaches. Zhaoheng Zheng, Ramakant Nevatia, Yan Liu 0002 |
ICPR | 3 |
| 2022 | OPEN: Order-preserving Pointcloud Encoder Decoder Network for Body Shape RefinementabstractImage-based 3-D human body shape estimation and reconstruction have shown significant improvement by using deep neural networks. Compared with reconstructing from a single image, reconstructing 3-D human body shapes from video or image sequences requires high precision and dense correspondences between the keypoints of the reconstructed shape sequence. Existing methods cannot achieve both high accuracy and keep the dense correspondence between different shapes after reconstruction. In this paper, we propose a method named Order-preserving Point cloud Encoder-decoder Network to refine the reconstructed human body shape from SMPL with the assistance of RGB images while preserving its original dense correspondence. We further introduce using 2-D RGB images as weak supervision when 3-D labels are not available. We assess our methods on the public dataset and show improved results compared with the baseline methods. Haidong Zhu, Ramakant Nevatia |
ICPR | 5 |
| 2022 | Temporal Shift and Attention Modules for Graphical Skeleton Action RecognitionabstractSkeletons, consisting of joint positions and connections between them, are an important representation for modeling human bodies in image frames. Compared with understanding RGB videos, recognizing actions from the skeletons removes the biases of background and body shapes. Researchers use spatial-temporal graphs to model the skeleton sequences. These methods weigh all frames in the sequence equally even though many of the frames may not be useful for action and prediction and dilute the influence of important frames. Also, the temporal graph focuses on understanding only the low-level feature of the joints for the motion of the skeleton. In this paper, we introduce two modules, temporal shift module and temporal attention module that can be added to graph convolution networks for skeleton action recognition. Temporal attention module focuses on keyframes for making predictions, and temporal shift module helps to exchange the high-level features between different frames along the temporal dimension besides the local patterns. We evaluate the two modules with two existing skeleton action recognition networks, ST-GCN and MS-G3D, on three public datasets and show better results than the original methods. Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia |
ICPR | 3 |
| 2021 | SimPLE: Similar Pseudo Label Exploitation for Semi-Supervised ClassificationabstractA common classification task situation is where one has a large amount of data available for training, but only a small portion is annotated with class labels. The goal of semi-supervised training, in this context, is to improve classification accuracy by leverage information not only from labeled data but also from a large amount of unlabeled data. Recent works have developed significant improvements by exploring the consistency constrain between differently augmented labeled and unlabeled data. Following this path, we propose a novel unsupervised objective that focuses on the less studied relationship between the high confidence unlabeled data that are similar to each other. The new proposed Pair Loss minimizes the statistical distance between high confidence pseudo labels with similarity above a certain threshold. Combining the Pair Loss with the techniques developed by the MixMatch family, our proposed SimPLE algorithm shows significant performance gains over previous algorithms on CIFAR-100 and Mini-ImageNet, and is on par with the state-of-the-art methods on CIFAR-10 and SVHN. Furthermore, SimPLE also outperforms the state-of-the-art methods in the transfer learning setting, where models are initialized by the weights pre-trained on ImageNet or DomainNet-Real. The code is available at github.com/zijian-hu/SimPLE. Zijian Hu 0001, Zhengyu Yang 0003, Xuefeng Hu, Ramakant Nevatia |
CVPR | 4 |
| 2021 | Visual Semantic Role Labeling for Video UnderstandingabstractWe propose a new framework for understanding and representing related salient events in a video using visual semantic role labeling. We represent videos as a set of related events, wherein each event consists of a verb and multiple entities that fulfill various roles relevant to that event. To study the challenging task of semantic role labeling in videos or VidSRL, we introduce the VidSitu benchmark, a large scale video understanding data source with 29K 10-second movie clips richly annotated with a verb and semantic-roles every 2 seconds. Entities are co-referenced across events within a movie clip and events are connected to each other via event-event relations. Clips in VidSitu are drawn from a large collection of movies (∼3K) and have been chosen to be both complex (∼4.2 unique verbs within a video) as well as diverse (∼200 verbs have more than 100 annotations each). We provide a comprehensive analysis of the dataset in comparison to other publicly available video understanding benchmarks, several illustrative baselines and evaluate a range of standard video recognition models. Our code and dataset is available at vidsitu.org. Arka Sadhu, Tanmay Gupta, Mark Yatskar, Ramakant Nevatia, Aniruddha Kembhavi |
CVPR | 4 |
| 2021 | Improving Object Detection And Attribute Recognition By Feature Entanglement ReductionabstractWe explore object detection with two attributes: color and material. The task aims to simultaneously detect objects and infer their color and material. A straight-forward approach is to add attribute heads at the very end of a usual object detection pipeline. However, we observe that the two goals are in conflict: Object detection should be attribute-independent and attributes be largely object-independent. Features computed by a standard detection network entangle the category and attribute features; we disentangle them by the use of a two-stream model where the category and attribute features are computed independently but the classification heads share Regions of Interest (RoIs). Compared with a traditional single-stream model, our model shows significant improvements over VG-20, a subset of Visual Genome, on both supervised and attribute transfer tasks. Zhaoheng Zheng, Arka Sadhu, Ramakant Nevatia |
ICIP | 3 |
| 2021 | Video Question Answering with Phrases via Semantic RolesabstractVideo Question Answering (VidQA) evaluation metrics have been limited to a single-word answer or selecting a phrase from a fixed set of phrases.These metrics limit the VidQA models' application scenario.In this work, we leverage semantic roles derived from video descriptions to mask out certain phrases, to introduce VidQAP which poses VidQA as a fillin-the-phrase task.To enable evaluation of answer phrases, we compute the relative improvement of the predicted answer compared to an empty string.To reduce the influence of language-bias in VidQA datasets, we retrieve a video having a different answer for the same question.To facilitate research, we construct ActivityNet-SRL-QA and Charades-SRL-QA and benchmark them by extending three vision-language models.We perform extensive analysis and ablative studies to guide future work.Code and data are public.Video description: A man on top of a building throws a bowling ball towards the pins Q4: throws a bowling ball towards the pins.Model's generated answer: A man standing on a house Correct answer: A man on top of a building Q5: A man on top of a building a bowling ball towards the pins.Model's generated answer: throws Correct answer: throws Q6: A man on top of a building throws towards the pins.Model's generated answer: a ball Correct answer: a bowling ball Q7: A man on top of a building throws a bowling ball Model's generated answer: towards some bottles Correct answer: towards the pins (b) Free-form Answer Generation ARG0 Arka Sadhu, Ramakant Nevatia |
NAACL-HLT | 3 |
| 2021 | Utilizing Every Image Object for Semi-supervised Phrase GroundingabstractPhrase grounding models localize an object in the image given a referring expression. The annotated language queries available during training are limited, which also limits the variations of language combinations that a model can see during training. In this paper, we study the case applying objects without labeled queries for training the semi-supervised phrase grounding. We propose to use learned location and subject embedding predictors (LSEP) to generate the corresponding language embeddings for objects lacking annotated queries in the training set. With the assistance of the detector, we also apply LSEP to train a grounding model on images without any annotation. We evaluate our method based on MAttNet on three public datasets: RefCOCO, RefCOCO+, and RefCOCOg. We show that our predictors allow the grounding system to learn from the objects without labeled queries and improve accuracy by 34.9% relatively with the detection results. Haidong Zhu, Arka Sadhu, Zhaoheng Zheng, Ramakant Nevatia |
WACV | 4 |
| 2020 | Video Object Grounding Using Semantic Roles in Language DescriptionabstractWe explore the task of Video Object Grounding (VOG), which grounds objects in videos referred to in natural language descriptions. Previous methods apply image grounding based algorithms to address VOG, fail to explore the object relation information and suffer from limited generalization. Here, we investigate the role of object relations in VOG and propose a novel framework VOGNet to encode multi-modal object relations via self-attention with relative position encoding. To evaluate VOGNet, we propose novel contrasting sampling methods to generate more challenging grounding input samples, and construct a new dataset called ActivityNet-SRL (ASRL) based on existing caption and grounding datasets. Experiments on ASRL validate the need of encoding object relations in VOG, and our VOGNet outperforms competitive baselines by a significant margin. Arka Sadhu, Ramakant Nevatia |
CVPR | 3 |
| 2020 | Curriculum DeepSDF
Yueqi Duan, Haidong Zhu, He Wang 0010, Li Yi 0001, Ramakant Nevatia, Leonidas J. Guibas |
ECCV (8) | 5 |
| 2020 | SPAN: Spatial Pyramid Attention Network for Image Manipulation Localization
Xuefeng Hu, Zhenye Jiang, Syomantak Chaudhuri, Zhenheng Yang, Ramakant Nevatia |
ECCV (21) | 6 |
| 2020 | Visually Grounded Continual Learning of Compositional PhrasesabstractHumans acquire language continually with much more limited access to data samples at a time, as compared to contemporary NLP systems.To study this human-like language acquisition ability, we present VisCOLL, a visually grounded language learning task, which simulates the continual acquisition of compositional phrases from streaming visual scenes.In the task, models are trained on a paired image-caption stream which has shifting object distribution; while being constantly evaluated by a visually-grounded masked language prediction task on held-out test sets.VisCOLL compounds the challenges of continual learning (i.e., learning from continuously shifting data distribution) and compositional generalization (i.e., generalizing to novel compositions).To facilitate research on VisCOLL, we construct two datasets, COCO-shift and Flickrshift, and benchmark them using different continual learning methods.Results reveal that SoTA continual learning approaches provide little to no improvements on VisCOLL, since storing examples of all possible compositions is infeasible.We conduct further ablations and analysis to guide future work 1 . Xisen Jin, Junyi Du, Arka Sadhu, Ramakant Nevatia, Xiang Ren 0001 |
EMNLP (1) | 4 |
| 2020 | Every Pixel Counts ++: Joint Learning of Geometry and Motion with 3D Holistic UnderstandingabstractLearning to estimate 3D geometry in a single frame and optical flow from consecutive frames by watching unlabeled videos via deep convolutional network has made significant progress recently. Current state-of-the-art (SoTA) methods treat the two tasks independently. One typical assumption of the existing depth estimation methods is that the scenes contain no independent moving objects. while object moving could be easily modeled using optical flow. In this paper, we propose to address the two tasks as a whole, i.e., to jointly understand per-pixel 3D geometry and motion. This eliminates the need of static scene assumption and enforces the inherent geometrical consistency during the learning process, yielding significantly improved results for both tasks. We call our method as “Every Pixel Counts++” or “EPC++”. Specifically, during training, given two consecutive frames from a video, we adopt three parallel networks to predict the camera motion (MotionNet), dense depth map (DepthNet), and per-pixel optical flow between two frames (OptFlowNet) respectively. The three types of information, are fed into a holistic 3D motion parser (HMP), and per-pixel 3D motion of both rigid background and moving objects are disentangled and recovered. Various loss terms are formulated to jointly supervise the three networks. An effective adaptive training strategy is proposed to achieve better performance and more efficient convergence. Comprehensive experiments were conducted on datasets with different scenes, including driving scenario (KITTI 2012 and KITTI 2015 datasets), mixed outdoor/indoor scenes (Make3D) and synthetic animation (MPI Sintel dataset). Performance on the five tasks of depth estimation, optical flow estimation, odometry, moving object segmentation and scene flow estimation shows that our approach outperforms other SoTA methods, demonstrating the effectiveness of each module of our proposed method. Code will be available at: https://github.com/chenxuluo/EPC. Chenxu Luo, Zhenheng Yang, Peng Wang 0001, Yang Wang 0046, Wei Xu 0017, Ramakant Nevatia, Alan L. Yuille |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2019 | Activity Driven Weakly Supervised Object DetectionabstractWeakly supervised object detection aims at reducing the amount of supervision required to train detection models. Such models are traditionally learned from images/videos labelled only with the object class and not the object bounding box. In our work, we try to leverage not only the object class labels but also the action labels associated with the data. We show that the action depicted in the image/video can provide strong cues about the location of the associated object. We learn a spatial prior for the object dependent on the action (e.g. "ball" is closer to "leg of the person" in "kicking ball"), and incorporate this prior to simultaneously train a joint object detection and action classification model. We conducted experiments on both video datasets and image datasets to evaluate the performance of our weakly supervised object detection model. Our approach outperformed the current state-of-the-art (SOTA) method by more than 6% in mAP on the Charades video dataset. Zhenheng Yang, Dhruv Mahajan 0001, Deepti Ghadiyaram, Ramakant Nevatia, Vignesh Ramanathan |
CVPR | 4 |
| 2019 | NOTE-RCNN: NOise Tolerant Ensemble RCNN for Semi-Supervised Object DetectionabstractThe labeling cost of large number of bounding boxes is one of the main challenges for training modern object detectors. To reduce the dependence on expensive bounding box annotations, we propose a new semi-supervised object detection formulation, in which a few seed box level annotations and a large scale of image level annotations are used to train the detector. We adopt a training-mining framework, which is widely used in weakly supervised object detection tasks. However, the mining process inherently introduces various kinds of labelling noises: false negatives, false positives and inaccurate boundaries, which can be harmful for training the standard object detectors (e.g. Faster RCNN). We propose a novel NOise Tolerant Ensemble RCNN (NOTE-RCNN) object detector to handle such noisy labels. Comparing to standard Faster RCNN, it contains three highlights: an ensemble of two classification heads and a distillation head to avoid overfitting on noisy labels and improve the mining precision, masking the negative sample loss in box predictor to avoid the harm of false negative labels, and training box regression head only on seed annotations to eliminate the harm from inaccurate boundaries of mined bounding boxes. We evaluate the methods on ILSVRC 2013 and MSCOCO 2017 dataset; we observe that the detection accuracy consistently improves as we iterate between mining and training steps, and state-of-the-art performance is achieved. Jiyang Gao, Jiang Wang 0001, Shengyang Dai, Li-Jia Li 0001, Ramakant Nevatia |
ICCV | 5 |
| 2019 | Zero-Shot Grounding of Objects From Natural Language QueriesabstractA phrase grounding system localizes a particular object in an image referred to by a natural language query. In previous work, the phrases were restricted to have nouns that were encountered in training, we extend the task to Zero-Shot Grounding(ZSG) which can include novel, “unseen” nouns. Current phrase grounding systems use an explicit object detection network in a 2-stage framework where one stage generates sparse proposals and the other stage evaluates them. In the ZSG setting, generating appropriate proposals itself becomes an obstacle as the proposal generator is trained on the entities common in the detection and grounding datasets. We propose a new single-stage model called ZSGNet which combines the detector network and the grounding system and predicts classification scores and regression parameters. Evaluation of ZSG system brings additional subtleties due to the influence of the relationship between the query and learned categories; we define four distinct conditions that incorporate different levels of difficulty. We also introduce new datasets, sub-sampled from Flickr30k Entities and Visual Genome, that enable evaluations for the four conditions. Our experiments show that ZSGNet achieves state-of-the-art performance on Flickr30k and ReferIt under the usual “seen” settings and performs significantly better than baseline in the zero-shot setting. Arka Sadhu, Ramakant Nevatia |
ICCV | 3 |
| 2019 | MAC: Mining Activity Concepts for Language-Based Temporal LocalizationabstractWe address the problem of language-based temporal localization in untrimmed videos. Compared to temporal localization with fixed categories, this problem is more challenging as the language-based queries not only have no pre-defined activity list but also may contain complex descriptions. Previous methods address the problem by considering features from video sliding windows and language queries and learning a subspace to encode their correlation, which ignore rich semantic cues about activities in videos and queries. We propose to mine activity concepts from both video and language modalities by applying the actionness score enhanced Activity Concepts based Localizer (ACL). Specifically, the novel ACL encodes the semantic concepts from verb-obj pairs in language queries and leverages activity classifiers' prediction scores to encode visual concepts. Besides, ACL also has the capability to regress sliding windows as localization results. Experiments show that ACL significantly outperforms state-of-the-arts under the widely used metric, with more than 5% increase on both Charades-STA and TACoS datasets. Runzhou Ge, Jiyang Gao, Ramakant Nevatia |
WACV | 4 |
| 2019 | Deep, Landmark-Free FAME: Face Alignment, Modeling, and Expression Estimation
Feng-Ju Chang, Anh Tuan Tran 0001, Tal Hassner, Iacopo Masi, Ramakant Nevatia, Gérard G. Medioni |
Int. J. Comput. Vis. | 5 |
| 2019 | Learning Pose-Aware Models for Pose-Invariant Face Recognition in the WildabstractWe propose a method designed to push the frontiers of unconstrained face recognition in the wild with an emphasis on extreme out-of-plane pose variations. Existing methods either expect a single model to learn pose invariance by training on massive amounts of data or else normalize images by aligning faces to a single frontal pose. Contrary to these, our method is designed to explicitly tackle pose variations. Our proposed Pose-Aware Models (PAM) process a face image using several pose-specific, deep convolutional neural networks (CNN). 3D rendering is used to synthesize multiple face poses from input images to both train these models and to provide additional robustness to pose variations at test time. Our paper presents an extensive analysis of the IARPA Janus Benchmark A (IJB-A), evaluating the effects that landmark detection accuracy, CNN layer selection, and pose model selection all have on the performance of the recognition pipeline. It further provides comparative evaluations on IJB-A and the PIPA dataset. These tests show that our approach outperforms existing methods, even surprisingly matching the accuracy of methods that were specifically fine-tuned to the target dataset. Parts of this work previously appeared in [1] and [2]. Iacopo Masi, Feng-Ju Chang, Jongmoo Choi, Shai Harel, Jungyeon Kim, KangGeon Kim, Jatuporn Toy Leksut, Stephen Rawls, Yue Wu 0001, Tal Hassner, Wael Abd-Almageed, Gérard G. Medioni, Louis-Philippe Morency, Premkumar Natarajan, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 15 |
| 2018 | SPOT Poachers in Action: Augmenting Conservation Drones With Automatic Detection in Near Real TimeabstractThe unrelenting threat of poaching has led to increased development of new technologies to combat it. One such example is the use of long wave thermal infrared cameras mounted on unmanned aerial vehicles (UAVs or drones) to spot poachers at night and report them to park rangers before they are able to harm animals. However, monitoring the live video stream from these conservation UAVs all night is an arduous task. Therefore, we build SPOT (Systematic POacher deTector), a novel application that augments conservation drones with the ability to automatically detect poachers and animals in near real time. SPOT illustrates the feasibility of building upon state-of-the-art AI techniques, such as Faster RCNN, to address the challenges of automatically detecting animals and poachers in infrared images. This paper reports (i) the design and architecture of SPOT, (ii) a series of efforts towards more robust and faster processing to make SPOT usable in the field and provide detections in near real time, and (iii) evaluation of SPOT based on both historical videos and a real-world test run by the end users in the field. The promising results from the test in the field have led to a plan for larger-scale deployment in a national park in Botswana. While SPOT is developed for conservation drones, its design and novel techniques have wider application for automated detection from UAV videos. Elizabeth Bondi-Kelly, Fei Fang 0001, Mark Hamilton, Debarun Kar, Donnabell Dmello, Jongmoo Choi, Robert Hannaford, Arvind Iyer, Lucas Joppa, Milind Tambe, Ramakant Nevatia |
AAAI | 11 |
| 2018 | Unsupervised Learning of Geometry From Videos With Edge-Aware Depth-Normal ConsistencyabstractLearning to reconstruct depths from a single image by watching unlabeled videos via deep convolutional network (DCN) is attracting significant attention in recent years, e.g. (Zhou et al. 2017). In this paper, we propose to use surface normal representation for unsupervised depth estimation framework. Our estimated depths are constrained to be compatible with predicted normals, yielding more robust geometry results. Specifically, we formulate an edge-aware depth-normal consistency term, and solve it by constructing a depth-to-normal layer and a normal-to-depth layer inside of the DCN. The depth-to-normal layer takes estimated depths as input, and computes normal directions using cross production based on neighboring pixels. Then given the estimated normals, the normal-to-depth layer outputs a regularized depth map through local planar smoothness. Both layers are computed with awareness of edges inside the image to help address the issue of depth/normal discontinuity and preserve sharp edges. Finally, to train the network, we apply the photometric error and gradient smoothness to supervise both depth and normal predictions. We conducted experiments on both outdoor (KITTI) and indoor (NYUv2) datasets, and showed that our algorithm vastly outperforms state-of-the-art, which demonstrates the benefits of our approach. Zhenheng Yang, Peng Wang 0001, Wei Xu 0017, Liang Zhao 0006, Ramakant Nevatia |
AAAI | 5 |
| 2018 | PIRC Net: Using Proposal Indexing, Relationships and Context for Phrase Grounding
Rama Kovvuri, Ramakant Nevatia |
ACCV (4) | 2 |
| 2018 | Knowledge Aided Consistency for Weakly Supervised Phrase GroundingabstractGiven a natural language query, a phrase grounding system aims to localize mentioned objects in an image. In weakly supevised scenario, mapping between image regions (i.e., proposals) and language is not available in the training set. Previous methods address this deficiency by training a grounding system via learning to reconstruct language information contained in input queries from predicted proposals. However, the optimization is solely guided by the reconstruction loss from the language modality, and ignores rich visual information contained in proposals and useful cues from external knowledge. In this paper, we explore the consistency contained in both visual and language modalities, and leverage complementary external knowledge to facilitate weakly supervised grounding. We propose a novel Knowledge Aided Consistency Network (KAC Net) which is optimized by reconstructing input query and proposal's information. To leverage complementary knowledge contained in the visual features, we introduce a Knowledge Based Pooling (KBP) gate to focus on query-related proposals. Experiments show that KAC Net provides a significant improvement on two popular datasets. Jiyang Gao, Ramakant Nevatia |
CVPR | 3 |
| 2018 | Motion-Appearance Co-Memory Networks for Video Question AnsweringabstractVideo Question Answering (QA) is an important task in understanding video temporal structure. We observe that there are three unique attributes of video QA compared with image QA: (1) it deals with long sequences of images containing richer information not only in quantity but also in variety; (2) motion and appearance information are usually correlated with each other and able to provide useful attention cues to the other; (3) different questions require different number of frames to infer the answer. Based on these observations, we propose a motion-appearance co-memory network for video QA. Our networks are built on concepts from Dynamic Memory Network (DMN) and introduces new mechanisms for video QA. Specifically, there are three salient aspects: (1) a co-memory attention mechanism that utilizes cues from both motion and appearance to generate attention; (2) a temporal conv-deconv network to generate multi-level contextual facts; (3) a dynamic fact ensemble method to construct temporal representation dynamically for different questions. We evaluate our method on TGIF-QA dataset, and the results outperform state-of-the-art significantly on all four tasks of TGIF-QA. Jiyang Gao, Runzhou Ge, Ramakant Nevatia |
CVPR | 4 |
| 2018 | LEGO: Learning Edge With Geometry All at Once by Watching VideosabstractLearning to estimate 3D geometry in a single image by watching unlabeled videos via deep convolutional network is attracting significant attention. In this paper, we introduce a "3D as-smooth-as-possible (3D-ASAP)" prior inside the pipeline, which enables joint estimation of edges and 3D scene, yielding results with significant improvement in accuracy for fine detailed structures. Specifically, we define the 3D-ASAP prior by requiring that any two points recovered in 3D from an image should lie on an existing planar surface if no other cues provided. We design an unsupervised framework that Learns Edges and Geometry (depth, normal) all at Once (LEGO). The predicted edges are embedded into depth and surface normal smoothness terms, where pixels without edges in-between are constrained to satisfy the prior. In our framework, the predicted depths, normals and edges are forced to be consistent all the time. We conduct experiments on KITTI to evaluate our estimated geometry and CityScapes to perform edge evaluation. We show that in all of the tasks, i.e. depth, normal and edge, our algorithm vastly outperforms other state-of-the-art (SOTA) algorithms, demonstrating the benefits of our approach. Zhenheng Yang, Peng Wang 0001, Yang Wang 0046, Wei Xu 0017, Ramakant Nevatia |
CVPR | 5 |
| 2018 | CTAP: Complementary Temporal Action Proposal Generation
Jiyang Gao, Ramakant Nevatia |
ECCV (2) | 3 |
| 2018 | ExpNet: Landmark-Free, Deep, 3D Facial ExpressionsabstractWe describe a deep learning based method for estimating 3D facial expression coefficients. Unlike previous work, our process does not relay on facial landmark detection methods as a proxy step. Recent methods have shown that a CNN can be trained to regress accurate and discriminative 3D morphable model (3DMM) representations, directly from image intensities. By foregoing landmark detection, these methods were able to estimate shapes for occluded faces appearing in unprecedented viewing conditions. We build on those methods by showing that facial expressions can also be estimated by a robust, deep, landmark-free approach. Our ExpNet CNN is applied directly to the intensities of a face image and regresses a 29D vector of 3D expression coefficients. We propose a unique method for collecting data to train our network, leveraging on the robustness of deep networks to training label noise. We further offer a novel means of evaluating the accuracy of estimated expression coefficients: by measuring how well they capture facial emotions on the CK+ and EmotiW-17 emotion recognition benchmarks. We show that our ExpNet produces expression coefficients which better discriminate between facial emotions than those obtained using state of the art, facial landmark detectors. Moreover, this advantage grows as image scales drop, demonstrating that our ExpNet is more robust to scale changes than landmark detectors. Finally, our ExpNet is orders of magnitude faster than its alternatives. Feng-Ju Chang, Anh Tuan Tran 0001, Tal Hassner, Iacopo Masi, Ramakant Nevatia, Gérard G. Medioni |
FG | 5 |
| 2018 | Face and Body Association for Video-Based Face RecognitionabstractIn recent years face recognition has made extraordinary leaps, yet unconstrained video-based face identification in the wild remains an open and interesting problem. Videos, unlike still-images, offer a myriad of data for face modeling, sampling, and recognition, but, on the other hand, contain low-quality frames and motion blur. A key component in video-based face recognition is the way in which faces are associated through the video sequence before being used for recognition. In this paper, we present a video-based face recognition method taking advantage of face and body association (FBA). To track and associate subjects that appear across frames in multiple shots, we solve a data association problem using both face and body appearance. The final recovered track is then used to build a face representation for recognition. We evaluate our FBA method for video-based face recognition on a challenging dataset. Our experiments show up to 5% improvement in the identification rate over the state-of-the-art. KangGeon Kim, Zhenheng Yang, Iacopo Masi, Ramakant Nevatia, Gérard G. Medioni |
WACV | 4 |
| 2017 | DECK: Discovering Event Composition Knowledge from Web Images for Zero-Shot Event Detection and Recounting in VideosabstractWe address the problem of zero-shot event recognition in consumer videos. An event usually consists of multiple human-human and human-object interactions over a relative long period of time. A common approach proceeds by representing videos with banks of object and action concepts, but requires additional user inputs to specify the desired concepts per event. In this paper, we provide a fully automatic algorithm to select representative and reliable concepts for event queries. This is achieved by discovering event composition knowledge (DECK) from web images. To evaluate our proposed method, we use the standard zero-shot event detection protocol (ZeroMED), but also introduce a novel zero-shot event recounting (ZeroMER) problem to select supporting evidence of the events. Our ZeroMER formulation aims to select video snippets that are relevant and diverse. Evaluation on the challenging TRECVID MED dataset show that our proposed method achieves promising results on both tasks. Chuang Gan 0001, Chen Sun 0002, Ramakant Nevatia |
AAAI | 3 |
| 2017 | Cascaded Boundary Regression for Temporal Action Detection
Jiyang Gao, Zhenheng Yang, Ramakant Nevatia |
BMVC | 3 |
| 2017 | RED: Reinforced Encoder-Decoder Networks for Action Anticipation
Jiyang Gao, Zhenheng Yang, Ramakant Nevatia |
BMVC | 3 |
| 2017 | Spatio-Temporal Action Detection with Cascade Proposal and Location Anticipation
Zhenheng Yang, Jiyang Gao, Ramakant Nevatia |
BMVC | 3 |
| 2017 | AMC: Attention Guided Multi-modal Correlation Learning for Image SearchabstractGiven a users query, traditional image search systems rank images according to its relevance to a single modality (e.g., image content or surrounding text). Nowadays, an increasing number of images on the Internet are available with associated meta data in rich modalities (e.g., titles, keywords, tags, etc.), which can be exploited for better similarity measure with queries. In this paper, we leverage visual and textual modalities for image search by learning their correlation with input query. According to the intent of query, attention mechanism can be introduced to adaptively balance the importance of different modalities. We propose a novel Attention guided Multi-modal Correlation (AMC) learning method which consists of a jointly learned hierarchy of intra and inter-attention networks. Conditioned on querys intent, intra-attention networks (i.e., visual intra-attention network and language intra-attention network) attend on informative parts within each modality, a multi-modal inter-attention network promotes the importance of the most query-relevant modalities. In experiments, we evaluate AMC models on the search logs from two real world image search engines and show a significant boost on the ranking of user-clicked images in search results. Additionally, we extend AMC models to caption ranking task on COCO dataset and achieve competitive results compared with recent state-of-the-arts. Trung Bui, Ramakant Nevatia |
CVPR | 5 |
| 2017 | Local-Global Landmark Confidences for Face RecognitionabstractA key to successful face recognition is accurate and reliable face alignment using automatically-detected facial landmarks. Given this strong dependency between face recognition and facial landmark detection, robust face recognition requires knowledge of when the facial landmark detection algorithm succeeds and when it fails. Facial landmark confidence represents this measure of success. In this paper, we propose two methods to measure landmark detection confidence: local confidence based on local predictors of each facial landmark, and global confidence based on a 3D rendered face model. A score fusion approach is also introduced to integrate these two confidences effectively. We evaluate both confidence metrics on two datasets for face recognition: JANUS CS2 and IJB-A datasets. Our experiments show up to 9% improvements when face recognition algorithm integrates the local-global confidence metrics. KangGeon Kim, Feng-Ju Chang, Jongmoo Choi, Louis-Philippe Morency, Ramakant Nevatia, Gérard G. Medioni |
FG | 5 |
| 2017 | Query-Guided Regression Network with Context Policy for Phrase GroundingabstractGiven a textual description of an image, phrase grounding localizes objects in the image referred by query phrases in the description. State-of-the-art methods address the problem by ranking a set of proposals based on the relevance to each query, which are limited by the performance of independent proposal generation systems and ignore useful cues from context in the description. In this paper, we adopt a spatial regression method to break the performance limit, and introduce reinforcement learning techniques to further leverage semantic context information. We propose a novel Query-guided Regression network with Context policy (QRC Net) which jointly learns a Proposal Generation Network (PGN), a Query-guided Regression Network (QRN) and a Context Policy Network (CPN). Experiments show QRC Net provides a significant improvement in accuracy on two popular datasets: Flickr30K Entities and Referit Game, with 14.25% and 17.14% increase over the state-of-the-arts respectively. Rama Kovvuri, Ramakant Nevatia |
ICCV | 3 |
| 2017 | TALL: Temporal Activity Localization via Language QueryabstractThis paper focuses on temporal localization of actions in untrimmed videos. Existing methods typically train classifiers for a pre-defined list of actions and apply them in a sliding window fashion. However, activities in the wild consist of a wide combination of actors, actions and objects; it is difficult to design a proper activity list that meets users' needs. We propose to localize activities by natural language queries. Temporal Activity Localization via Language (TALL) is challenging as it requires: (1) suitable design of text and video representations to allow cross-modal matching of actions and language queries; (2) ability to locate actions accurately given features from sliding windows of limited granularity. We propose a novel Cross-modal Temporal Regression Localizer (CTRL) to jointly model text query and video clips, output alignment scores and action boundary regression results for candidate clips. Lor evaluation, we adopt TaCoS dataset, and build a new dataset for this task on top of Charades by adding sentence temporal annotations, called Charades-STA. We also build complex sentence queries in Charades-STA for test. Experimental results show that CTRL outperforms previous methods significantly on both datasets. Jiyang Gao, Chen Sun 0002, Zhenheng Yang, Ramakant Nevatia |
ICCV | 4 |
| 2017 | TURN TAP: Temporal Unit Regression Network for Temporal Action ProposalsabstractTemporal Action Proposal (TAP) generation is an important problem, as fast and accurate extraction of semantically important (e.g. human actions) segments from untrimmed videos is an important step for large-scale video analysis. We propose a novel Temporal Unit Regression Network (TURN) model. There are two salient aspects of TURN: (1) TURN jointly predicts action proposals and refines the temporal boundaries by temporal coordinate regression: (2) Fast computation is enabled by unit feature reuse: a long untrimmed video is decomposed into video units, which are reused as basic building blocks of temporal proposals. TURN outperforms the previous state-of-the-art methods under average recall (AR) by a large margin on THUMOS-14 and ActivityNet datasets, and runs at over 880 frames per second (FPS) on a TITAN X GPU. We further apply TURN as a proposal generation stage for existing temporal action localization pipelines, it outperforms state-of-the-art performance on THUMOS-14 and ActivityNet. Jiyang Gao, Zhenheng Yang, Chen Sun 0002, Ramakant Nevatia |
ICCV | 5 |
| 2017 | MSRC: Multimodal Spatial Regression with Semantic Context for Phrase GroundingabstractGiven an image and a natural language query phrase, a grounding system localizes the mentioned objects in the image according to the query's specifications. State-of-the-art methods address the problem by ranking a set of proposal bounding boxes according to the query's semantics, which makes them dependent on the performance of proposal generation systems. Besides, query phrases in one sentence may be semantically related in one sentence and can provide useful cues to ground objects. We propose a novel Multimodal Spatial Regression with semantic Context (MSRC) system which not only predicts the location of ground truth based on proposal bounding boxes, but also refines prediction results by penalizing similarities of different queries coming from same sentences. The advantages of MSRC are twofold: first, it removes the limitation of performance from proposal generation algorithms by using a spatial regression network. Second, MSRC not only encodes the semantics of a query phrase, but also deals with its relation with other queries in the same sentence (i.e., context) by a context refinement network. Experiments show MSRC system provides a significant improvement in accuracy on two popular datasets: Flickr30K Entities and Refer-it Game, with 6.64% and 5.28% increase over the state-of-the-arts respectively. Rama Kovvuri, Jiyang Gao, Ramakant Nevatia |
ICMR | 4 |
| 2016 | Image Set Classification via Template Triplets and Context-Aware Similarity Embedding
Feng-Ju Chang, Ramakant Nevatia |
ACCV (5) | 2 |
| 2016 | Learning Action Concept Trees and Semantic Alignment Networks from Image-Description Data
Jiyang Gao, Ramakant Nevatia |
ACCV (2) | 2 |
| 2016 | ProNet: Learning to Propose Object-Specific Boxes for Cascaded Neural NetworksabstractThis paper aims to classify and locate objects accurately and efficiently, without using bounding box annotations. It is challenging as objects in the wild could appear at arbitrary locations and in different scales. In this paper, we propose a novel classification architecture ProNet based on convolutional neural networks. It uses computationally efficient neural networks to propose image regions that are likely to contain objects, and applies more powerful but slower networks on the proposed regions. The basic building block is a multi-scale fully-convolutional network which assigns object confidence scores to boxes at different locations and scales. We show that such networks can be trained effectively using image-level annotations, and can be connected into cascades or trees for efficient object classification. ProNet outperforms previous state-of-the-art significantly on PASCAL VOC 2012 and MS COCO datasets for object classification and point-based localization. Chen Sun 0002, Manohar Paluri, Ronan Collobert, Ramakant Nevatia, Lubomir D. Bourdev |
CVPR | 4 |
| 2016 | Exploring deep learning based solutions in fine grained activity recognition in the wildabstractIn this paper, we explore the usage of deep learning based solutions in fine grained activity recognition in the wild. As a powerful tool, deep learning has been widely used in image classification, object detection and activity recognition. We focus on implementing deep learning methods into the more complicated fine grained activity recognition problems. We test our solutions on MPII activity dataset with 410 activities. We find that due to the challenges of large intra class variances, small inter class variances, and limited training samples per activity, the classical two stream deep ConvNets method does not perform that well for fine grained activity recognition. Observing these issues, we propose a solution to directly use deep features learned from ImageNet in an SVM. In experiments, we achieve a 20 percent improvement compared to the classical two stream deep ConvNets solutions, on MPII fine grained activity challenge videos. Ramakant Nevatia |
ICPR | 2 |
| 2016 | Segment-based models for event detection and recountingabstractWe present a novel approach towards web video classification and recounting that uses video segments to model an event. This approach overcomes the limitations faced by the classical video-level models such as modeling semantics, identifying informative segments in a video and background segment suppression. We posit that segment-based models are able to identify both the frequently-occurring and rarer patterns in an event effectively, despite being trained on only a fraction of the training data. Our framework employs a discriminative approach to optimize our models in distributed and data-driven fashion while maintaining semantic interpretability. We evaluate the effectiveness of our approach on the challenging TRECVID MEDTest 2014 dataset. We demonstrate improvements in recounting and classification, particularly in events characterized by inherent intra-class variations. Rama Kovvuri, Ramakant Nevatia, Cees Snoek |
ICPR | 2 |
| 2016 | A multi-scale cascade fully convolutional network face detectorabstractFace detection is challenging as faces in images could be present at arbitrary locations and in different scales. We propose a three-stage cascade structure based on fully convolutional neural networks (FCNs). It first proposes the approximate locations where the faces may be, then aims to find the accurate location by zooming on to the faces. Each level of the FCN cascade is a multi-scale fully-convolutional network, which generates scores at different locations and in different scales. A score map is generated after each FCN stage. Probable regions of face are selected and fed to the next stage. The number of proposals is decreased after each level, and the areas of regions are decreased to more precisely fit the face. Compared to passing proposals directly between stages, passing probable regions can decrease the number of proposals and reduce the cases where first stage doesn't propose good bounding boxes. We show that by using FCN and score map, the FCN cascade face detector can achieve strong performance on public datasets. Zhenheng Yang, Ramakant Nevatia |
ICPR | 2 |
| 2016 | ACD: Action Concept Discovery from Image-Sentence CorporaabstractAction classification in still images is an important task in computer vision. It is challenging as the appearances of actions may vary depending on their context (e.g. associated objects). Manually labeling of context information would be time consuming and difficult to scale up. To address this challenge, we propose a method to automatically discover and cluster action concepts, and learn their classifiers from weakly supervised image-sentence corpora. It obtains candidate action concepts by extracting verb-object pairs from sentences and verifies their visualness with the associated images. Candidate action concepts are then clustered by using a multi-modal representation with image embeddings from deep convolutional networks and text embeddings from word2vec. More than one hundred human action concept classifiers are learned from the Flickr 30k dataset with no additional human effort and promising classification results are obtained. We further apply the AdaBoost algorithm to automatically select and combine relevant action concepts given an action query. Promising results have been shown on the PASCAL VOC 2012 action classification benchmark, which has zero overlap with Flickr30k. Jiyang Gao, Chen Sun 0002, Ramakant Nevatia |
ICMR | 3 |
| 2016 | Face recognition using deep multi-pose representationsabstractWe introduce our method and system for face recognition using multiple pose-aware deep learning models. In our representation, a face image is processed by several pose-specific deep convolutional neural network (CNN) models to generate multiple pose-specific features. 3D rendering is used to generate multiple face poses from the input image. Sensitivity of the recognition system to pose variations is reduced since we use an ensemble of pose-specific CNN features. The paper presents extensive experimental results on the effect of landmark detection, CNN layer selection and pose model selection on the performance of the recognition pipeline. Our novel representation achieves better results than the state-of-the-art on IARPA's CS2 and NIST's IJB-A in both verification and identification (i.e. search) tasks. Wael Abd-Almageed, Yue Wu 0001, Stephen Rawls, Shai Harel, Tal Hassner, Iacopo Masi, Jongmoo Choi, Jatuporn Toy Leksut, Jungyeon Kim, Premkumar Natarajan, Ramakant Nevatia, Gérard G. Medioni |
WACV | 11 |
| 2016 | Tag-based video retrieval by embedding semantic content in a continuous word spaceabstractContent-based event retrieval in unconstrained web videos, based on query tags, is a hard problem due to large intra-class variances, and limited vocabulary and accuracy of the video concept detectors, creating a "semantic query gap". We present a technique to overcome this gap by using continuous word space representations to explicitly compute query and detector concept similarity. This not only allows for fast query-video similarity computation with implicit query expansion, but leads to a compact video representation, which allows implementation of a real-time retrieval system that can fit several thousand videos in a few hundred megabytes of memory. We evaluate the effectiveness of our representation on the challenging NIST MEDTest 2014 dataset. Arnav Agharwal, Rama Kovvuri, Ramakant Nevatia, Cees Snoek |
WACV | 3 |
| 2016 | Abstraction hierarchy and self annotation update for fine grained activity recognitionabstractFine-grained activity recognition focuses recognition on sub-ordinate levels. This task is made difficult due to low inter-class variability and high intra-class variability caused by human motion and objects. We propose that recognition of such activities can be significantly improved by grouping and decomposing them into a hierarchy of multiple abstraction layers; we introduce a Hierarchical Activity Network (HAN). Recognition in HAN is guided by classifiers operating at multiple levels; furthermore, descriptions of different levels of abstraction are also generated from HAN, which may be useful for different tasks. We show significant improvements in accuracy of recognition compared to earlier methods. Besides, annotation for fine grained activity is challenging, and inaccurate annotation influences classification performance. We explore an automatic solution for improving the classification results while auto enhancing the annotation quality. Ramakant Nevatia |
WACV | 3 |
| 2016 | Activity recognition and prediction with pose based discriminative patch modelabstractWe describe an image based activity recognition solution which can be applied to both off-line video classification and activity prediction in frames. We propose a Pose based Discriminative Patch Model to make activity recognition and prediction on image level (only observing several frames). This model enables a general and flexible framework to add in discriminative patches and consider their mutual relations to an efficient tree structure. PDP makes contribution in two aspects: (1) PDP provides a novel solution to improve activity recognition and prediction, by utilizing pose based discriminative patches instead of pose configuration feature, and modeling the patches' mutual relations. (2) PDP is an image-based algorithm, so it can make predictions using limited frames, even a single image. PDP focuses on challenging data captured from Internet and movies, where we achieve a 6% improvement compared with state-of-the-art method on video level recognition dataset - Sub-JHMDB, and image level action recognition dataset. We also obtain good improvement on activity prediction task. Ramakant Nevatia |
WACV | 3 |
| 2015 | Video event classification with temporal partitioningabstractThis paper addresses the problem of temporal pruning of noisy parts to improve event recognition performance. We present a new technique based on the temporal partitioning of the processed videos according to their motion patterns and the subsequent analysis of the yielded time segments. For each event type, we automatically learn the types of segments that are discriminative and those that perturb the classification. This process does not require detailed annotation of actions within an event type. A video is described with a set of quantized features and the final classification is performed according to the features that fall within the discriminative segments only. Experimental results show increased classification performance on the NIST MED11 dataset using two types of local features. Rémi Trichet, Ramakant Nevatia, J. Brian Burns |
AVSS | 2 |
| 2015 | Automatic Concept Discovery from Parallel Text and Visual CorporaabstractHumans connect language and vision to perceive the world. How to build a similar connection for computers? One possible way is via visual concepts, which are text terms that relate to visually discriminative entities. We propose an automatic visual concept discovery algorithm using parallel text and visual corpora, it filters text terms based on the visual discriminative power of the associated images, and groups them into concepts using visual and semantic similarities. We illustrate the applications of the discovered concepts using bidirectional image and sentence retrieval task and image tagging task, and show that the discovered concepts not only outperform several large sets of manually selected concepts significantly, but also achieves the state-of-the-art performance in the retrieval task. Chen Sun 0002, Chuang Gan 0001, Ramakant Nevatia |
ICCV | 3 |
| 2015 | Temporal Localization of Fine-Grained Actions in Videos by Domain Transfer from Web ImagesabstractWe address the problem of fine-grained action localization from temporally untrimmed web videos. We assume that only weak video-level annotations are available for training. The goal is to use these weak labels to identify temporal segments corresponding to the actions, and learn models that generalize to unconstrained web videos. We find that web images queried by action names serve as well-localized highlights for many actions, but are noisily labeled. To solve this problem, we propose a simple yet effective method that takes weak video labels and noisy image labels as input, and generates localized action frames as output. This is achieved by cross-domain transfer between video frames and web images, using pre-trained deep convolutional neural networks. We then use the localized action frames to train action recognition models with long short-term memory networks. We collect a fine-grained sports action data set FGA-240 of more than 130,000 YouTube videos. It has 240 fine-grained actions under 85 sports activities. Convincing results are shown on the FGA-240 data set, as well as the THUMOS 2014 localization data set with untrimmed training videos. Chen Sun 0002, Sanketh Shetty, Rahul Sukthankar, Ramakant Nevatia |
ACM Multimedia | 4 |
| 2015 | Forecasting Human Pose and Motion with Multibody Dynamic ModelabstractUnderstanding human motion with dynamics is in its infancy, but it is a highly promising approach in computer vision, robotics and computer graphics. We propose a Multibody Dynamic Model (MDM) which estimates poses and motions through analyzing forces-the intrinsic motivation for motion. With the 23 degrees of freedom Multibody Dynamic Model, we analyze human motion dynamics in the whole body, and then forecast human motion or pose in occluded or non-captured circumstances. Our two main contributions are essential for understanding human motion with dynamics. The first one is to provide effective representations and computational models for dynamic analysis of human motion in the whole body, via the intrinsic connection between force and motion in the biomechanical system. The second contribution is to offer a more natural method to forecast pose and motion with the estimated forces. In our experiments, MDM has been successfully applied to running, jumping and other challenging sports activities. Ramakant Nevatia |
WACV | 2 |
| 2015 | A Robust Adaptive Classifier for Detector Adaptation in a VideoabstractWe propose a novel method for improved object detection in a video. Our approach adapts a generic offline trained detector (OTD) to a specific test video by collecting online samples in an unsupervised manner. Most of the existing adaptation methods focus on collecting confident online samples and do not address how to deal with ambiguous and noisy online samples. We address the importance of collecting online samples which are true representative of the actual objects present in the video and propose a Boosted Multiple Instance Random Fern (B-MIRF) classifier as the adaptive classifier. Multiple Instance Learning (MIL) provides reliability for training with noisy online samples and boosting process enables in obtaining more discriminative random ferns. We apply B-MIRF classifier on the detection responses obtained from OTD, hence our method improves the performance by improving the precision of OTD. We evaluate performance of our method on two challenging public datasets and show better performance than other state of the art methods. Pramod Sharma, Ramakant Nevatia |
WACV | 2 |
| 2015 | Beyond Pedestrians: A Hybrid Approach of Tracking Multiple Articulating HumansabstractWe propose a hybrid framework to address the problem of tracking multiple articulated humans from a single camera. Our method incorporates offline learned category-level detector with online learned instance-specific detector as a hybrid system. To deal with humans in large pose articulation, which can not be reliably detected by off-line trained detectors, we propose an online learned instance specific patch-based detector, consisting of layered patch classifiers. With extrapolated track lets by online learned detectors, we use the discriminative color filters learned online to compute the appearance affinity score for further global association. Experimental evaluation on both standard pedestrian datasets and articulated human datasets shows significant improvement compared to state-of-the-art multi-human tracking methods. Ramakant Nevatia, Bo Yang 0008 |
WACV | 2 |
| 2014 | Multi-state Discriminative Video Segment Selection for Complex Event Classification
Prithviraj Banerjee, Ramakant Nevatia |
ACCV (5) | 2 |
| 2014 | DISCOVER: Discovering Important Segments for Classification of Video Events and RecountingabstractWe propose a unified framework DISCOVER to simultaneously discover important segments, classify high-level events and generate recounting for large amounts of unconstrained web videos. The motivation is our observation that many video events are characterized by certain important segments. Our goal is to find the important segments and capture their information for event classification and recounting. We introduce an evidence localization model where evidence locations are modeled as latent variables. We impose constraints on global video appearance, local evidence appearance and the temporal structure of the evidence. The model is learned via a max-margin framework and allows efficient inference. Our method does not require annotating sources of evidence, and is jointly optimized for event classification and recounting. Experimental results are shown on the challenging TRECVID 2013 MEDTest dataset. Chen Sun 0002, Ramakant Nevatia |
CVPR | 2 |
| 2014 | Pose Filter Based Hidden-CRF Models for Activity Detection
Prithviraj Banerjee, Ramakant Nevatia |
ECCV (2) | 2 |
| 2014 | Semantic Aware Video Transcription Using Random Forest Classifiers
Chen Sun 0002, Ramakant Nevatia |
ECCV (1) | 2 |
| 2014 | Late fusion and calibration for multimedia event detection using few examplesabstractThe state-of-the-art in example-based multimedia event detection (MED) rests on heterogeneous classifiers whose scores are typically combined in a late-fusion scheme. Recent studies on this topic have failed to reach a clear consensus as to whether machine learning techniques can outperform rule-based fusion schemes with varying amount of training data. In this paper, we present two parametric approaches to late fusion: a normalization scheme for arithmetic mean fusion (logistic averaging) and a fusion scheme based on logistic regression, and compare them to widely used rule-based fusion schemes. We also describe how logistic regression can be used to calibrate the fused detection scores to predict an optimal threshold given a detection prior and costs on errors. We discuss the advantages and shortcomings of each approach when the amount of positives available for training varies from 10 positives (10Ex) to 100 positives (100Ex). Experiments were run using video data from the NIST TRECVID MED 2013 evaluation and results were reported in terms of a ranking metric: the mean average precision (mAP) and R0, a cost-based metric introduced in TRECVID MED 2013. Julien van Hout, Eric Yeh, Dennis C. Koelma, Cees Snoek, Chen Sun 0002, Ramakant Nevatia, Julie Wong, Gregory K. Myers |
ICASSP | 6 |
| 2014 | Video Segmentation Descriptors for Event RecognitionabstractThis paper presents a new video motion descriptor based on a multi-scale video segmentation to provide a multi-layered output as well as connections with the rich interactions that occur between objects at the semantic level. We also put the emphasis on relationships between motion clusters by providing a new relative motion descriptor encapsulating relative motion patterns within a local spatio-temporal neighborhood. Experimental results on the challenging TRECVID MED11 event recognition dataset validate the approach. Rémi Trichet, Ramakant Nevatia |
ICPR | 2 |
| 2014 | ISOMER: Informative Segment Observations for Multimedia Event RecountingabstractThis paper describes a system for multimedia event detection and recounting. The goal is to detect a high level event class in unconstrained web videos and generate event oriented summarization for display to users. For this purpose, we detect informative segments and collect observations for them, leading to our ISOMER system. We combine a large collection of both low level and semantic level visual and audio features for event detection. For event recounting, we propose a novel approach to identify event oriented discriminative video segments and their descriptions with a linear SVM event classifier. User friendly concepts including objects, actions, scenes, speech and optical character recognition are used in generating descriptions. We also develop several mapping and filtering strategies to cope with noisy concept detectors. Our system performed competitively in the TRECVID 2013 Multimedia Event Detection task with near 100,000 videos and was the highest performer in TRECVID 2013 Multimedia Event Recounting task. Chen Sun 0002, J. Brian Burns, Ramakant Nevatia, Cees Snoek, Robert C. Bolles, Gregory K. Myers, Wen Wang 0001, Eric Yeh |
ICMR | 3 |
| 2014 | Multi class boosted random ferns for adapting a generic object detector to a specific videoabstractDetector adaptation is a challenging problem and several methods have been proposed in recent years. We propose multi class boosted random ferns for detector adaptation. First we collect online samples in an unsupervised manner and collected positive online samples are divided into different categories for different poses of the object. Then we train a multi-class boosted random fern adaptive classifier. Our adaptive classifier training focuses on two aspects: discriminability and efficiency. Boosting provides discriminative random ferns. For efficiency, our boosting procedure focuses on sharing the same feature among different classes and multiple strong classifiers are trained in a single boosting framework. Experiments on challenging public datasets demonstrate effectiveness of our approach. Pramod Sharma, Ramakant Nevatia |
WACV | 2 |
| 2014 | Video segmentation and feature co-occurrences for activity classificationabstractBag-of-Word scheme has almost become de rigueur for event recognition tasks due to its robustness and simplicity. Despite its effectiveness, this technique discards spatial and temporal relationships between codewords. This paper tackles the problem of building a video codeword representation that captures such relationships. We developed a new method that harnesses spatio-temporal boundaries and discriminative codeword co-occurrences. Given a set of videos and their corresponding quantized features, the video is first decomposed in spatio-temporal volumes according to a multi-scale video segmentation algorithm. Meaningful codeword co-occurrences are then extracted within each volume and videos are then represented with histograms of co-occurring features. The set of histograms is finally fed to an SVM for classification. Evaluation under the realistic TRECVID MED11 challenge database validates the approach. Rémi Trichet, Ramakant Nevatia |
WACV | 2 |
| 2014 | Multi-Target Tracking by Online Learning a CRF Model of Appearance and Motion Patterns
Bo Yang 0008, Ramakant Nevatia |
Int. J. Comput. Vis. | 2 |
| 2014 | Hierarchical abnormal event detection by real time and semi-real time multi-tasking video surveillance system
Sung Chun Lee, Ramakant Nevatia |
Mach. Vis. Appl. | 2 |
| 2014 | Evaluating multimedia features and fusion for example-based event detectionabstractMultimedia event detection (MED) is a challenging problem because of the heterogeneous content and variable quality found in large collections of Internet videos. To study the value of multimedia features and fusion for representing and learning events from a set of example video clips, we created SESAME, a system for video SEarch with Speed and Accuracy for Multimedia Events. SESAME includes multiple bag-of-words event classifiers based on single data types: low-level visual, motion, and audio features; high-level semantic visual concepts; and automatic speech recognition. Event detection performance was evaluated for each event classifier. The performance of low-level visual and motion features was improved by the use of difference coding. The accuracy of the visual concepts was nearly as strong as that of the low-level visual features. Experiments with a number of fusion methods for combining the event detection scores from these classifiers revealed that simple fusion methods, such as arithmetic mean, perform as well as or better than other, more complex fusion methods. SESAME’s performance in the 2012 TRECVID MED evaluation was one of the best reported. Gregory K. Myers, Ramesh Nallapati, Julien van Hout, Stephanie Pancoast, Ramakant Nevatia, Chen Sun 0002, AmirHossein Habibian, Dennis C. Koelma, Koen E. A. van de Sande, Arnold W. M. Smeulders, Cees Snoek |
Mach. Vis. Appl. | 5 |
| 2013 | Conditional Bayesian networks for action detectionabstractThe task of understanding video content has seen great interest from computer vision community with the increase in camera based surveillance at grocery stores, airports, train stations, etc. What makes up a scene (objects) and what happens in the scene (actions) are two important dimensions of video understanding. In this work, we aim to identify both actions and objects in the video, however, we focus only on the objects with which human interacts. We use videos which may have multiple actions taking place during possibly overlapping intervals. Our system can recognize actions having high intra-class variance performed in complex environments using objects of different types, sizes and shapes. We produce structured descriptions for the videos as output. The descriptions identify the subject, the object, the verb and the interval of each activity recognized. Furqan Muhammad Khan, Sung Chun Lee, Ramakant Nevatia |
AVSS | 3 |
| 2013 | Video segmentation with spatio-temporal tubesabstractLong-term temporal interactions among objects are an important cue for video understanding. To capture such object relations, we propose a novel method for spatiotemporal video segmentation based on dense trajectory clustering that is also effective when objects articulate. We use superpixels of homogeneous size jointly with optical flow information to ease the matching of regions from one frame to another. Our second main contribution is a hierarchical fusion algorithm that yields segmentation information available at multiple linked scales. We test the algorithm on several videos from the web showing a large variety of difficulties. Rémi Trichet, Ramakant Nevatia |
AVSS | 2 |
| 2013 | Efficient Detector Adaptation for Object Detection in a VideoabstractIn this work, we present a novel and efficient detector adaptation method which improves the performance of an offline trained classifier (baseline classifier) by adapting it to new test datasets. We address two critical aspects of adaptation methods: generalizability and computational efficiency. We propose an adaptation method, which can be applied to various baseline classifiers and is computationally efficient also. For a given test video, we collect online samples in an unsupervised manner and train a random fern adaptive classifier. The adaptive classifier improves precision of the baseline classifier by validating the obtained detection responses from baseline classifier as correct detections or false alarms. Experiments demonstrate generalizability, computational efficiency and effectiveness of our method, as we compare our method with state of the art approaches for the problem of human detection and show good performance with high computational efficiency on two different baseline classifiers. Pramod Sharma, Ramakant Nevatia |
CVPR | 2 |
| 2013 | ACTIVE: Activity Concept Transitions in Video Event ClassificationabstractThe goal of high level event classification from videos is to assign a single, high level event label to each query video. Traditional approaches represent each video as a set of low level features and encode it into a fixed length feature vector (e.g. Bag-of-Words), which leave a big gap between low level visual features and high level events. Our paper tries to address this problem by exploiting activity concept transitions in video events (ACTIVE). A video is treated as a sequence of short clips, all of which are observations corresponding to latent activity concept variables in a Hidden Markov Model (HMM). We propose to apply Fisher Kernel techniques so that the concept transitions over time can be encoded into a compact and fixed length feature vector very efficiently. Our approach can utilize concept annotations from independent datasets, and works well even with a very small number of training samples. Experiments on the challenging NIST TRECVID Multimedia Event Detection (MED) dataset shows our approach performs favorably over the state-of-the-art. Chen Sun 0002, Ramakant Nevatia |
ICCV | 2 |
| 2013 | Large-scale web video event classification by use of Fisher VectorsabstractEvent recognition has been an important topic in computer vision research due to its many applications. However, most of the work has focused on videos taken from a fixed camera, known environments and basic events. Here, we focus on classification of unconstrained, web videos into much higher level activities. We follow the approach of constructing fixed length feature vectors from local feature descriptors for classification using an SVM. Our key contribution is the study of the utility of Fisher Vector representation in improving results compared to the conventional Bag-of-Words (BoW) approach. Such coding has shown to be useful for static image classification in the past but not applied to video categorization. We perform tests on the challenging NIST TRECVID Multimedia Event Detection (MED) dataset, which has thousand hours of unconstrained user generated videos; our approach achieves as much as 35% improvement over the BoW baseline. We also offer an analysis of possible causes of such improvements. Chen Sun 0002, Ramakant Nevatia |
WACV | 2 |
| 2013 | Hierarchical multi-channel hidden semi Markov graphical models for activity recognition
Pradeep Natarajan, Ramakant Nevatia |
Comput. Vis. Image Underst. | 2 |
| 2013 | Multiple Target Tracking by Learning-Based Hierarchical Association of Detection ResponsesabstractWe propose a hierarchical association approach to multiple target tracking from a single camera by progressively linking detection responses into longer track fragments (i.e., tracklets). Given frame-by-frame detection results, a conservative dual-threshold method that only links very similar detection responses between consecutive frames is adopted to generate initial tracklets with minimum identity switches. Further association of these highly fragmented tracklets at each level of the hierarchy is formulated as a Maximum A Posteriori (MAP) problem that considers initialization, termination, and transition of tracklets as well as the possibility of them being false alarms, which can be efficiently computed by the Hungarian algorithm. The tracklet affinity model, which measures the likelihood of two tracklets belonging to the same target, is a linear combination of automatically learned weak nonparametric models upon various features, which is distinct from most of previous work that relies on heuristic selection of parametric models and manual tuning of their parameters. For this purpose, we develop a novel bag ranking method and train the crucial tracklet affinity models by the boosting algorithm. This bag ranking method utilizes the soft max function to relax the oversufficient objective function used by the conventional instance ranking method. It provides a tighter upper bound of empirical errors in distinguishing correct associations from the incorrect ones, and thus yields more accurate tracklet affinity models for the tracklet association problem. We apply this approach to the challenging multiple pedestrian tracking task. Systematic experiments conducted on two real-life datasets show that the proposed approach outperforms previous state-of-the-art algorithms in terms of tracking accuracy, in particular, considerably reducing fragmentations and identity switches. Chang Huang, Yuan Li 0022, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Robust Object Tracking Using Constellation Model with Superpixel
Ramakant Nevatia |
ACCV (3) | 2 |
| 2012 | Unsupervised incremental learning for improved object detection in a videoabstractMost common approaches for object detection collect thousands of training examples and train a detector in an offline setting, using supervised learning methods, with the objective of obtaining a generalized detector that would give good performance on various test datasets. However, when an offline trained detector is applied on challenging test datasets, it may fail in some cases by not being able to detect some objects or by producing false alarms. We propose an unsupervised multiple instance learning (MIL) based incremental solution to deal with this issue. We introduce an MIL loss function for Real Adaboost and present a tracking based effective unsupervised online sample collection mechanism to collect the online samples for incremental learning. Experiments demonstrate the effectiveness of our approach by improving the performance of a state of the art offline trained detector on the challenging datasets for pedestrian category. Pramod Sharma, Chang Huang, Ramakant Nevatia |
CVPR | 3 |
| 2012 | Multi-target tracking by online learning of non-linear motion patterns and robust appearance modelsabstractWe describe an online approach to learn non-linear motion patterns and robust appearance models for multi-target tracking in a tracklet association framework. Unlike most previous approaches that use linear motion methods only, we online build a non-linear motion map to better explain direction changes and produce more robust motion affinities between tracklets. Moreover, based on the incremental learned entry/exit map, a multiple instance learning method is devised to produce strong appearance models for tracking; positive sample pairs are collected from different track-lets so that training samples have high diversity. Finally, using online learned moving groups, a tracklet completion process is introduced to deal with tracklets not reaching entry/exit points. We evaluate our approach on three public data sets, and show significant improvements compared with state-of-art methods. Bo Yang 0008, Ramakant Nevatia |
CVPR | 2 |
| 2012 | An online learned CRF model for multi-target trackingabstractWe introduce an online learning approach for multitarget tracking. Detection responses are gradually associated into tracklets in multiple levels to produce final tracks. Unlike most previous approaches which only focus on producing discriminative motion and appearance models for all targets, we further consider discriminative features for distinguishing difficult pairs of targets. The tracking problem is formulated using an online learned CRF model, and is transformed into an energy minimization problem. The energy functions include a set of unary functions that are based on motion and appearance models for discriminating all targets, as well as a set of pairwise functions that are based on models for differentiating corresponding pairs of tracklets. The online CRF approach is more powerful at distinguishing spatially close targets with similar appearances, as well as in dealing with camera motions. An efficient algorithm is introduced for finding an association with low energy cost. We evaluate our approach on three public data sets, and show significant improvements compared with several state-of-art methods. Bo Yang 0008, Ramakant Nevatia |
CVPR | 2 |
| 2012 | Online Learned Discriminative Part-Based Appearance Models for Multi-human Tracking
Bo Yang 0008, Ramakant Nevatia |
ECCV (1) | 2 |
| 2012 | Pose based activity recognition using Multiple Kernel learning
Prithviraj Banerjee, Ramakant Nevatia |
ICPR | 2 |
| 2012 | Robust multi-pose face tracking by multi-stage tracklet association
Markus Roth, Martin Bäuml, Ramakant Nevatia, Rainer Stiefelhagen |
ICPR | 3 |
| 2012 | Efficient incremental learning of boosted classifiers for object detection
Pramod Sharma, Chang Huang, Ramakant Nevatia |
ICPR | 3 |
| 2012 | Simultaneous inference of activity, pose and objectabstractHuman movements are important cues for recognizing human actions, which can be captured by explicit modeling and tracking of actor or through space-time low-level features. However, relying solely on human dynamics is not enough to discriminate between actions which have similar human dynamics, such as smoking and drinking, irrespective of the modeling method. Object perception plays an important role in such cases. Conversely, human movements are indicative of type of object used for the action. These two processes of object perception and action understanding are thus not independent. Consequently, action recognition improves when human movements and object perception are used in conjunction. Therefore, we propose a probabilistic approach to simultaneously infer what action was performed, what object was used and what poses the actor went through. This joint inference framework can better discriminate between actions and objects which are too similar and lack discriminative features. Furqan Muhammad Khan, Vivek K. Singh 0002, Ramakant Nevatia |
WACV | 3 |
| 2012 | A systems level approach to perimeter protectionabstractEffective perimeter protection mechanisms for industrial sites and critical infrastructure must contend with a large variety of potential threats as well as with the fact that normal site activity can be both complex and diverse. This paper documents the development of a system level approach capable of functioning under such challenging conditions. A multi-view tracking system is used to provide real-time site wide trajectories of all observed individuals. A Radar-based system is also used for tracking if and when camera coverage of various regions is not available. Track information is then analyzed with respect to articulated motion analysis, complex event analysis and normalcy analysis. In addition, object recognition is used to classify left behind objects using high resolution PTZ imagery. A real-time integrated version of this comprehensive approach to perimeter protection was deployed using a single standard off-the-shelf desktop computer. Peter H. Tu, Ting Yu 0003, Ramakant Nevatia, Sung Chun Lee, Hale Kim, Phill-Kyu Rhee, Joong-Hwan Baek |
WACV | 4 |
| 2011 | Learning neighborhood cooccurrence statistics of sparse features for human activity recognitionabstractA common approach to activity recognition has been the use of histogram of codewords computed from Spatio Temporal Interest Points (STIPs). Recent methods have focused on leveraging the spatio-temporal neighborhood structure of the features, but they are generally restricted to aggregate statistics over the entire video volume, and ignore local pairwise relationships. Our goal is to capture these relations in terms of pairwise cooccurrence statistics of codewords. We show a reduction of such cooccurrence relations to the edges connecting the latent variables of a Conditional Random Field (CRF) classifier. As a consequence, we also learn the codeword dictionary as a part of the maximum likelihood learning process, with each interest point assigned a probability distribution over the codewords. We show results on two widely used activity recognition datasets. Prithviraj Banerjee, Ramakant Nevatia |
AVSS | 2 |
| 2011 | AVSS 2011 demo session: A systems level approach to perimeter protectionabstractSummary form only given. The rapid evolution of tools and software systems to design experiments, automatically monitor, collect and warehouse large amounts of data, from applications such as life sciences and industrial processes has resulted in a new paradigm shift. This change of paradigm is so fast that some of the practices for optimization and management of these processes that were valid only 5–10 years ago may no longer be fully acceptable or sufficient for today's business optimization and management. This has a direct influence on the best practices for knowledge discovery and management of the discovered knowledge in real-world data mining applications. Establishing and managing a real-world data mining project in any domain, in particular in today's life science industry, is not a trivial task. A few approaches have been proposed in the literature. However, initiation and successful management of such efforts may depend on where a given case study fits in the overall classification of data mining approaches. Today's knowledge discovery from data can be classified in several ways: (i) data mining on engineered systems (e.g. complex equipment) or systems designed by nature (e.g. life sciences), (ii) explanatory or predictive data mining, (iii) data mining from static data (e.g. data warehouse) or dynamic data (e.g. data streams), (iv) user operated or automated data mining. There could still be other ways to classify data mining applications. This talk provides an overview of the above listed knowledge discovery applications. We provide examples where we demonstrate how small or large amounts of data, when understood from a real-world data mining point of view and the required data is properly integrated, can result in novel knowledge discovery case studies. We explain motivations and challenges of establishing real-world dat Peter H. Tu, Ting Yu 0003, Ramakant Nevatia, Sung Chun Lee, Hale Kim, Phill-Kyu Rhee, Joong-Hwan Baek |
AVSS | 4 |
| 2011 | How does person identity recognition help multi-person tracking?abstractWe address the problem of multi-person tracking in a complex scene from a single camera. Although tracklet-association methods have shown impressive results in several challenging datasets, discriminability of the appearance model remains a limitation. Inspired by the work of person identity recognition, we obtain discriminative appearance-based affinity models by a novel framework to incorporate the merits of person identity recognition, which help multi-person tracking performance. During off-line learning, a small set of local image descriptors is selected to be used in on-line learned appearances-based affinity models effectively and efficiently. Given short but reliable track-lets generated by frame-to-frame association of detection responses, we identify them as query tracklets and gallery tracklets. For each gallery tracklet, a target-specific appearance model is learned from the on-line training samples collected by spatio-temporal constraints. Both gallery tracklets and query tracklets are fed into hierarchical association framework to obtain final tracking results. We evaluate our proposed system on several public datasets and show significant improvements in terms of tracking evaluation metrics. Cheng-Hao Kuo, Ramakant Nevatia |
CVPR | 2 |
| 2011 | Learning affinities and dependencies for multi-target tracking using a CRF modelabstractWe propose a learning-based Conditional Random Field (CRF) model for tracking multiple targets by progressively associating detection responses into long tracks. Tracking task is transformed into a data association problem, and most previous approaches developed heuristical parametric models or learning approaches for evaluating independent affinities between track fragments (tracklets). We argue that the independent assumption is not valid in many cases, and adopt a CRF model to consider both tracklet affinities and dependencies among them, which are represented by unary term costs and pairwise term costs respectively. Unlike previous methods, we learn the best global associations instead of the best local affinities between tracklets, and transform the task of finding the best association into an energy minimization problem. A RankBoost algorithm is proposed to select effective features for estimation of term costs in the CRF model, so that better associations have lower costs. Our approach is evaluated on challenging pedestrian data sets, and are compared with state-of-art methods. Experiments show effectiveness of our algorithm as well as improvement in tracking performance. Bo Yang 0008, Chang Huang, Ramakant Nevatia |
CVPR | 3 |
| 2011 | Action recognition in cluttered dynamic scenes using Pose-Specific Part ModelsabstractWe present an approach to recognizing single actor human actions in complex backgrounds. We adopt a Joint Tracking and Recognition approach, which track the actor pose by sampling from 3D action models. Most existing such approaches require large training data or MoCAP to handle multiple viewpoints, and often rely on clean actor silhouettes. The action models in our approach are obtained by annotating keyposes in 2D, lifting them to 3D stick figures and then computing the transformation matrices between the 3D keypose figures. Poses sampled from coarse action models may not fit the observations well; to overcome this difficulty, we propose an approach for efficiently localizing a pose by generating a Pose-Specific Part Model (PSPM) which captures appropriate kinematic and occlusion constraints in a tree-structure. In addition, our approach also does not require pose silhouettes. We show improvements to previous results on two publicly available datasets as well as on a novel, augmented dataset with dynamic backgrounds. Vivek K. Singh 0002, Ramakant Nevatia |
ICCV | 2 |
| 2011 | Vehicle detection from low quality aerial LIDAR dataabstractIn this paper we propose a vehicle detection framework on low resolution aerial range data. Our system consists of three steps: data mapping, 2D vehicle detection and postprocessing. First, we map the range data into 2D grayscale images by using the depth information only. For this purpose we propose a novel local ground plane estimation method, and the estimated ground plane is further refined by a global refinement process. Then we compute the depth value of missing points (points for which no depth information is available) by an effective interpolation method. In the second step, to train a classifier for the vehicles, we describe a method to generate more training examples from very few training annotations and adopt the fast cascade Adaboost approach for detecting vehicles in 2D grayscale images. Finally, in post-processing step we design a novel method to detect some vehicles which are comprised of clusters of missing points. We evaluate our method on real aerial data and the experiments demonstrate the effectiveness of our approach. Bo Yang 0008, Pramod Sharma, Ramakant Nevatia |
WACV | 3 |
| 2011 | Segmentation of objects in a detection window by Nonparametric Inhomogeneous CRFs
Bo Yang 0008, Chang Huang, Ramakant Nevatia |
Comput. Vis. Image Underst. | 3 |
| 2011 | Simultaneous tracking and action recognition for single actor human actions
Vivek K. Singh 0002, Ramakant Nevatia |
Vis. Comput. | 2 |
| 2010 | Dynamics Based Trajectory Segmentation for UAV videosabstractA novel representation of vehicle trajectories is proposed for applications in trajectory analysis and activity detection. Specifically, a piecewise arc fitting based smoothing algorithm is proposed for denoising the trajectories. A dynamic program is used to find the optimal arc fit to a given trajectory. We motivate the usage of dynamic primitives to parametrize common vehicular activities, and propose a dynamics based trajectory segmentation algorithm. Each primitive is modeled using a second order Auto-Regressive model, and form useful descriptors for a given vehicular trajectory. We evaluate both our trajectory smoothing and dynamic trajectory segmentation algorithm on a real UAV video dataset, and show performance improvements which clearly motivate its wide applicability in a general trajectory analysis system. Prithviraj Banerjee, Ramakant Nevatia |
AVSS | 2 |
| 2010 | High performance object detection by collaborative learning of Joint Ranking of Granules featuresabstractObject detection remains an important but challenging task in computer vision. We present a method that combines high accuracy with high efficiency. We adopt simplified forms of APCF features [3], which we term Joint Ranking of Granules (JRoG) features; the features consists of discrete values by uniting binary ranking results of pair-wise granules in the image. We propose a novel collaborative learning method for JRoG features, which consists of a Simulated Annealing (SA) module and an incremental feature selection module. The two complementary modules collaborate to efficiently search the formidably large JRoG feature space for discriminative features, which are fed into a boosted cascade for object detection. To cope with occlusions in crowded environments, we employ the strategy of part based detection, as in [19] but propose a new dynamic search method to improve the Bayesian combination of the part detection results. Experiments on several challenging data sets show that our approach achieves not only considerable improvement in detection accuracy but also major improvements in computational efficiency; on a Xeon 3GHz computer, with only a single thread, it can process a million scanning windows per second, sufficing for many practical real-time detection tasks. Chang Huang, Ramakant Nevatia |
CVPR | 2 |
| 2010 | Multi-target tracking by on-line learned discriminative appearance modelsabstractWe present an approach for online learning of discriminative appearance models for robust multi-target tracking in a crowded scene from a single camera. Although much progress has been made in developing methods for optimal data association, there has been comparatively less work on the appearance models, which are key elements for good performance. Many previous methods either use simple features such as color histograms, or focus on the discriminability between a target and the background which does not resolve ambiguities between the different targets. We propose an algorithm for learning a discriminative appearance model for different targets. Training samples are collected online from tracklets within a time sliding window based on some spatial-temporal constraints; this allows the models to adapt to target instances. Learning uses an Ad-aBoost algorithm that combines effective image descriptors and their corresponding similarity measurements. We term the learned models as OLDAMs. Our evaluations indicate that OLDAMs have significantly higher discrimination between different targets than conventional holistic color histograms, and when integrated into a hierarchical association framework, they help improve the tracking accuracy, particularly reducing the false alarms and identity switches. Cheng-Hao Kuo, Chang Huang, Ramakant Nevatia |
CVPR | 3 |
| 2010 | Learning 3D action models from a few 2D videos for view invariant action recognitionabstractMost existing approaches for learning action models work by extracting suitable low-level features and then training appropriate classifiers. Such approaches require large amounts of training data and do not generalize well to variations in viewpoint, scale and across datasets. Some work has been done recently to learn multi-view action models from Mocap data, but obtaining such data is time consuming and requires costly infrastructure. We present a method that addresses both these issues by learning action models from just a few video training samples. We model each action as a sequence of primitive actions, represented as functions which transform the actor's state. We formulate model learning as a curve-fitting problem, and present a novel algorithm for learning human actions by lifting 2D annotations of a few keyposes to 3D and interpolating between them. Actions are inferred by sampling the models and accumulating the feature weights learned discriminatively using a latent state Perceptron algorithm. We show results comparable to state-of-art on the standard Weizmann dataset, with a much smaller train:test ratio, and also in datasets for visual gesture recognition and cluttered grocery store environments. Pradeep Natarajan, Vivek K. Singh 0002, Ramakant Nevatia |
CVPR | 3 |
| 2010 | Inter-camera Association of Multi-target Tracks by On-Line Learned Appearance Affinity Models
Cheng-Hao Kuo, Chang Huang, Ramakant Nevatia |
ECCV (1) | 3 |
| 2010 | Efficient Inference with Multiple Heterogeneous Part Detectors for Human Pose Estimation
Vivek K. Singh 0002, Ramakant Nevatia, Chang Huang |
ECCV (3) | 2 |
| 2009 | Learning to associate: HybridBoosted multi-target tracker for crowded sceneabstractWe propose a learning-based hierarchical approach of multi-target tracking from a single camera by progressively associating detection responses into longer and longer track fragments (tracklets) and finally the desired target trajectories. To define tracklet affinity for association, most previous work relies on heuristically selected parametric models; while our approach is able to automatically select among various features and corresponding non-parametric models, and combine them to maximize the discriminative power on training data by virtue of a HybridBoost algorithm. A hybrid loss function is used in this algorithm because the association of tracklet is formulated as a joint problem of ranking and classification: the ranking part aims to rank correct tracklet associations higher than other alternatives; the classification part is responsible to reject wrong associations when no further association should be done. Experiments are carried out by tracking pedestrians in challenging datasets. We compare our approach with state-of-the-art algorithms to show its improvement in terms of tracking accuracy. Yuan Li 0022, Chang Huang, Ramakant Nevatia |
CVPR | 3 |
| 2009 | Robust multi-view car detection using unsupervised sub-categorizationabstractThis paper presents a novel approach for multi-view car detection using unsupervised sub-categorization instead of manual labeling. Cars have large variability of models and the view-point makes the appearance change dramatically. For object classes with a large intra-class variation like cars, a divide-and-conquer strategy may be applied. Instead of using manually predefined intra-class sub-categorization, we examine several non-linear dimension reduction methods and group samples in the low-dimension embedding in an unsupervised way. The clustered samples have strong view-point similarities internally. A boosting-based cascade tree classifier is trained based on these sub-categorizations. To demonstrate the capability of our multi-view car detector, we create a more challenging test set with annotations. Compared to the UIUC side-view car data set, our test set contains a large range of car models, view points, and complex backgrounds. We compare our approach with previous methods and the result shows that ours outperforms the state-of-the-art methods. Cheng-Hao Kuo, Ramakant Nevatia |
WACV | 2 |
| 2009 | Extensive articulated human detection by voting Cluster Boosted TreeabstractOur goal is to detect people in highly articulated poses, including bending, crouching, etc. Such formidable diversity in human poses makes detection much more difficult than for pedestrian poses. ¿Divide-and-conquer¿ is a favorable strategy for detecting objects with large intra class variations, which splits object instances into several subcategories and trains relatively simple classifiers for each sub-category. We propose a novel sample split method, which benefits the learning results of articulated humans. We adopt the cluster boosted tree (CBT) structure to automatically decide when a split should be triggered. Unlike the simple k-means used in CBT for sample split, our approach aims at minimizing the training loss after the split. Since this minimization is an NP-hard problem, we design a heuristic algorithm, in which we find optimal sample divisions according to each single feature, and then make compromises to get a final division by a voting-like process. We name our training method as voting cluster boosted tree (VCBT). Furthermore, to avoid large background area in training samples, we first cluster samples according to their width/height ratios, and then train a VCBT for each subset. We conduct an experiment on 17 infrared surveillance video clips, report superior performance compared with previous human detection methods, and show how our approach benefits the learning results by reducing training loss. Bo Yang 0008, Chang Huang, Ramakant Nevatia |
WACV | 3 |
| 2009 | Detection and Segmentation of Multiple, Partially Occluded Objects by Grouping, Merging, Assigning Part Detection Responses
Bo Wu 0001, Ramakant Nevatia |
Int. J. Comput. Vis. | 2 |
| 2009 | Human Pose Tracking in Monocular Sequence Using Multilevel Structured ModelsabstractTracking human body poses in monocular video has many important applications. The problem is challenging in realistic scenes due to background clutter, variation in human appearance and self-occlusion. The complexity of pose tracking is further increased when there are multiple people whose bodies may inter-occlude. We proposed a three-stage approach with multi-level state representation that enables a hierarchical estimation of 3D body poses. Our method addresses various issues including automatic initialization, data association, self and inter-occlusion. At the first stage, humans are tracked as foreground blobs and their positions and sizes are coarsely estimated. In the second stage, parts such as face, shoulders and limbs are detected using various cues and the results are combined by a grid-based belief propagation algorithm to infer 2D joint positions. The derived belief maps are used as proposal functions in the third stage to infer the 3D pose using data-driven Markov chain Monte Carlo. Experimental results on several realistic indoor video sequences show that the method is able to track multiple persons during complex movement including sitting and turning movements with self and inter-occlusion. Mun Wai Lee, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | View and scale invariant action recognition using multiview shape-flow modelsabstractActions in real world applications typically take place in cluttered environments with large variations in the orientation and scale of the actor. We present an approach to simultaneously track and recognize known actions that is robust to such variations, starting from a person detection in the standing pose. In our approach we first render synthetic poses from multiple viewpoints using Mocap data for known actions and represent them in a conditional random field (CRF) whose observation potentials are computed using shape similarity and the transition potentials are computed using optical flow. We enhance these basic potentials with terms to represent spatial and temporal constraints and call our enhanced model the shape, flow, duration-conditional random field (SFD-CRF). We find the best sequence of actions using Viterbi search in the SFD-CRF. We demonstrate our approach on videos from multiple viewpoints and in the presence of background clutter. Pradeep Natarajan, Ramakant Nevatia |
CVPR | 2 |
| 2008 | Optimizing discrimination-efficiency tradeoff in integrating heterogeneous local features for object detectionabstractA large variety of image features has been invented for detection of objects of a known class. We propose a framework to optimize the discrimination-efficiency tradeoff in integrating multiple, heterogeneous features for object detection. Cascade structured detectors are learned by boosting local feature based weak classifiers. Each weak classifier corresponds to a local image region, from which several different types of features are extracted. The weak classifier makes predictions by examining the features one by one; this classifier goes to the next feature only when the prediction from the already examined features is not confident enough. The order in which the features are evaluated is determined based on their computational cost normalized classification powers. We apply our approach to two object classes, pedestrians and cars. The experimental results show that our approach outperforms the state-of-the-art methods. Bo Wu 0001, Ramakant Nevatia |
CVPR | 2 |
| 2008 | Segmentation of multiple, partially occluded objects by grouping, merging, assigning part detection responsesabstractWe propose a method that detects and segments multiple, partially occluded objects in images. A part hierarchy is defined for the object class. Whole-object segmentor and part detectors are learned by boosting shape oriented local image features. During detection, the part detectors are applied to the input image. All the edge pixels in the image that positively contribute to part detection responses are extracted. A joint likelihood of multiple objects is defined based on the part detection responses and the object edges. Computing the joint likelihood includes an inter-object occlusion reasoning that is based on the object silhouettes extracted with the whole-object segmentor. By maximizing the joint likelihood, part detection responses are grouped, merged, and assigned to multiple object hypotheses. The proposed approach is applied to the pedestrian class, and evaluated on two public test sets. The experimental results show that our method outperforms the previous ones. Bo Wu 0001, Ramakant Nevatia, Yuan Li 0022 |
CVPR | 2 |
| 2008 | Global data association for multi-object tracking using network flowsabstractWe propose a network flow based optimization method for data association needed for multiple object tracking. The maximum-a-posteriori (MAP) data association problem is mapped into a cost-flow network with a non-overlap constraint on trajectories. The optimal data association is found by a min-cost flow algorithm in the network. The network is augmented to include an Explicit Occlusion Model(EOM) to track with long-term inter-object occlusions. A solution to the EOM-based network is found by an iterative approach built upon the original algorithm. Initialization and termination of trajectories and potential false observations are modeled by the formulation intrinsically. The method is efficient and does not require hypotheses pruning. Performance is compared with previous results on two public pedestrian datasets to show its improvement. Yuan Li 0022, Ramakant Nevatia |
CVPR | 3 |
| 2008 | Robust Object Tracking by Hierarchical Association of Detection Responses
Chang Huang, Bo Wu 0001, Ramakant Nevatia |
ECCV (2) | 3 |
| 2008 | Key Object Driven Multi-category Object Recognition, Localization and Tracking Using Spatio-temporal Context
Yuan Li 0022, Ramakant Nevatia |
ECCV (4) | 2 |
| 2008 | Human detection by searching in 3d space using camera and scene knowledgeabstractMany existing human detection systems are based on sub-window classification, namely detection is done by enumerating rectangular sub-images in the 2D image space. Detection rate of such approaches may be affected by perspective distortion and tilted orientation of the human in images. To overcome this problem without re-training the classifier, we develop a 3D search method. A search grid is defined in the 3D scene. At each grid point a rectified sub-image is generated to approximate the orthogonal projection of the target, so that the distortion due to camera setting is reduced. In addition, 3D target position can be estimated from single camera data. Experiments on challenging data from the PETS2007 and CAVIAR INRIA datasets show significantly improved detection performance of our approach compared with the 2D search-based methods. Yuan Li 0022, Bo Wu 0001, Ramakant Nevatia |
ICPR | 3 |
| 2008 | Segmentation and Tracking of Multiple Humans in Crowded EnvironmentsabstractSegmentation and tracking of multiple humans in crowded situations is made difficult by interobject occlusion. We propose a model based approach to interpret the image observations by multiple, partially occluded human hypotheses in a Bayesian framework. We define a joint image likelihood for multiple humans based on the appearance of the humans, the visibility of body obtained by occlusion reasoning, and foreground/background separation. The optimal solution is obtained by using an efficient sampling method, data-driven Markov chain Monte Carlo (DDMCMC), which uses image observations for proposal probabilities. Knowledge of various aspects including human shape, camera model, and image cues are integrated in one theoretically sound framework. We present experimental results and quantitative evaluation, demonstrating that the resulting approach is effective for very challenging data. Tao Zhao 0001, Ramakant Nevatia, Bo Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Single View Human Action Recognition using Key Pose Matching and Viterbi Path Searchingabstract3D human pose recovery is considered as a fundamental step in view-invariant human action recognition. However, inferring 3D poses from a single view usually is slow due to the large number of parameters that need to be estimated and recovered poses are often ambiguous due to the perspective projection. We present an approach that does not explicitly infer 3D pose at each frame. Instead, from existing action models we search for a series of actions that best match the input sequence. In our approach, each action is modeled as a series of synthetic 2D human poses rendered from a wide range of viewpoints. The constraints on transition of the synthetic poses is represented by a graph model called Action Net. Given the input, silhouette matching between the input frames and the key poses is performed first using an enhanced Pyramid Match Kernel algorithm. The best matched sequence of actions is then tracked using the Viterbi algorithm. We demonstrate this approach on a challenging video sets consisting of 15 complex action classes. Fengjun Lv, Ramakant Nevatia |
CVPR | 2 |
| 2007 | Simultaneous Object Detection and Segmentation by Boosting Local Shape Feature based ClassifierabstractThis paper proposes an approach to simultaneously detect and segment objects of a known category. Edgelet features are used to capture the local shape of the objects. For each feature a pair of base classifiers for detection and segmentation is built. The base segmentor is designed to predict the per-pixel figure-ground assignment around a neighborhood of the edgelet based on the feature response. The neighborhood is represented as an effective field which is determined by the shape of the edgelet. A boosting algorithm is used to learn the ensemble classifier with cascade decision strategy from the base classifier pool. The simultaneousness is achieved for both training and testing. The system is evaluated on a number of public image sets and compared with several previous methods. Bo Wu 0001, Ramakant Nevatia |
CVPR | 2 |
| 2007 | Improving Part based Object Detection by Unsupervised, Online BoostingabstractDetection of objects of a given class is important for many applications. However it is difficult to learn a general detector with high detection rate as well as low false alarm rate. Especially, the labor needed for manually labeling a huge training sample set is usually not affordable. We propose an unsupervised, incremental learning approach based on online boosting to improve the performance on special applications of a set of general part detectors, which are learned from a small amount of labeled data and have moderate accuracy. Our oracle for unsupervised learning, which has high precision, is based on a combination of a set of shape based part detectors learned by off-line boosting. Our online boosting algorithm, which is designed for cascade structure detector, is able to adapt the simple features, the base classifiers, the cascade decision strategy, and the complexity of the cascade automatically to the special application. We integrate two noise restraining strategies in both the oracle and the online learner. The system is evaluated on two public video corpora. Bo Wu 0001, Ramakant Nevatia |
CVPR | 2 |
| 2007 | Pedestrian Detection in Infrared Images based on Local Shape FeaturesabstractUse of IR images is advantageous for many surveillance applications where the systems must operate around the clock and external illumination is not always available. We investigate the methods derived from visible spectrum analysis for the task of human detection. Two feature classes (edgelets and HOG features) and two classification models(AdaBoost and SVM cascade) are extended to IR images. We find out that it is possible to get detection performance in IR images that is comparable to state-of-the-art results for visible spectrum images. It is also shown that the two domains share many features, likely originating from the silhouettes, in spite of the starkly different appearances of the two modalities. Bo Wu 0001, Ramakant Nevatia |
CVPR | 3 |
| 2007 | Cluster Boosted Tree Classifier for Multi-View, Multi-Pose Object DetectionabstractDetection of object of a known class is a fundamental problem of computer vision. The appearance of objects can change greatly due to illumination, view point, and articulation. For object classes with large intra-class variation, some divide-and-conquer strategy is necessary. Tree structured classifier models have been used for multi-view multi- pose object detection in previous work. This paper proposes a boosting based learning method, called Cluster Boosted Tree (CBT), to automatically construct tree structured object detectors. Instead of using predefined intra-class sub- categorization based on domain knowledge, we divide the sample space by unsupervised clustering based on discriminative image features selected by boosting algorithm. The sub-categorization information of the leaf nodes is sent back to refine their ancestors' classification functions. We compare our approach with previous related methods on several public data sets. The results show that our approach outperforms the state-of-the-art methods. Bo Wu 0001, Ramakant Nevatia |
ICCV | 2 |
| 2007 | Detection and Tracking of Multiple Humans with Extensive Pose ArticulationabstractWe describe a method for detecting and tracking humans. Different from most of the previous work, we focus on humans with extensive pose articulations, under situations where there is typically only a single camera, multiple humans are present and the image resolution is low. In our method pose clusters are learned from an embedded silhouette manifold. A set of object detectors, each of which corresponds to one pose cluster, are trained based on a novel Object-Weighted Appearance Model. A probabilistic pose-based transition model is used to track multiple objects within a sliding window buffer, making use of the detection responses. The track segments in the sliding windows are connected sequentially into full trajectories. Experiments on a set of challenging surveillance videos are presented; these show good performance of our approach compared to standard pedestrian detectors, under difficult conditions. Bo Wu 0001, Ramakant Nevatia |
ICCV | 3 |
| 2007 | Hierarchical Multi-channel Hidden Semi Markov Models
Pradeep Natarajan, Ramakant Nevatia |
IJCAI | 2 |
| 2007 | Detection and Tracking of Multiple, Partially Occluded Humans by Bayesian Combination of Edgelet based Part Detectors
Bo Wu 0001, Ramakant Nevatia |
Int. J. Comput. Vis. | 2 |
| 2006 | Tracking of Multiple, Partially Occluded Humans based on Static Body Part DetectionabstractTracking of humans in videos is important for many applications. A major source of difficulty in performing this task is due to inter-human or scene occlusion. We present an approach based on representing humans as an assembly of four body parts and detection of the body parts in single frames which makes the method insensitive to camera motions. The responses of the body part detectors and a combined human detector provide the "observations" used for tracking. Trajectory initialization and termination are both fully automatic and rely on the confidences computed from the detection responses. An object is tracked by data association if its corresponding detection response can be found; otherwise it is tracked by a meanshift style tracker. Our method can track humans with both inter-object and scene occlusions. The system is evaluated on three sets of videos and compared with previous method. Bo Wu 0001, Ramakant Nevatia |
CVPR (1) | 2 |
| 2006 | Human Pose Tracking Using Multi-level Structured Models
Mun Wai Lee, Ramakant Nevatia |
ECCV (3) | 2 |
| 2006 | Recognition and Segmentation of 3-D Human Action Using HMM and Multi-class AdaBoost
Fengjun Lv, Ramakant Nevatia |
ECCV (4) | 2 |
| 2006 | Geodec: Enabling Geospatial Decision MakingabstractThe rapid increase in the availability of geospatial data has motivated the effort to seamlessly integrate this information into an information-rich and realistic 3D environment. However, heterogeneous data sources with varying degrees of consistency and accuracy pose a challenge to such efforts. We describe the geospatial decision making (GeoDec) system, which accurately integrates satellite imagery, three-dimensional models, textures and video streams, road data, maps, point data and temporal data. The system also includes a glove-based user interface Cyrus Shahabi, Yao-Yi Chiang, Kelvin Chung, Kai-Chen Huang, Ali Khoshgozaran, Craig A. Knoblock, Sung Lee, Ulrich Neumann, Ramakant Nevatia, Arjun Rihan, Snehal Thakkar, Suya You |
ICME | 9 |
| 2006 | Camera Calibration from Video of a Walking HumanabstractA self-calibration method to estimate a camera's intrinsic and extrinsic parameters from vertical line segments of the same height is presented. An algorithm to obtain the needed line segments by detecting the head and feet positions of a walking human in his leg-crossing phases is described. Experimental results show that the method is accurate and robust with respect to various viewing angles and subjects. Fengjun Lv, Tao Zhao 0001, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2005 | A Model-Based Vehicle Segmentation Method for TrackingabstractOur goal is to detect and track moving vehicles on a road observed from cameras placed on poles or buildings. Inter-vehicle occlusion is significant under these conditions and traditional blob tracking methods is unable to separate the vehicles in the merged blobs. We use vehicle shape models, in addition to camera calibration and ground plane knowledge, to detect, track and classify moving vehicles in presence of occlusion. We use a 2-stage approach. In the first stage, hypothesis for vehicle types, positions and orientations are formed by a coarse search, which is then refined by a data driven Markov chain Monte Carlo (DDMCMC) process. We show results and evaluations on some real urban traffic video sequence using three types of vehicle models Xuefeng Song, Ramakant Nevatia |
ICCV | 2 |
| 2005 | Detection of Multiple, Partially Occluded Humans in a Single Image by Bayesian Combination of Edgelet Part DetectorsabstractThis paper proposes a method for human detection in crowded scene from static images. An individual human is modeled as an assembly of natural body parts. We introduce edgelet features, which are a new type of silhouette oriented features. Part detectors, based on these features, are learned by a boosting method. Responses of part detectors are combined to form a joint likelihood model that includes cases of multiple, possibly inter-occluded humans. The human detection problem is formulated as maximum a posteriori (MAP) estimation. We show results on a commonly used previous dataset as well as new data sets that could not be processed by earlier methods. Bo Wu 0001, Ramakant Nevatia |
ICCV | 2 |
| 2004 | Extraction and Integration of Window in a 3D Building Model from Ground View Image
Sung Chun Lee, Ramakant Nevatia |
CVPR (2) | 2 |
| 2004 | Tracking Multiple Humans in Crowded Environment
Tao Zhao 0001, Ramakant Nevatia |
CVPR (2) | 2 |
| 2004 | Video-based event recognition: activity representation and probabilistic recognition methods
Somboon Hongeng, Ramakant Nevatia, François Brémond |
Comput. Vis. Image Underst. | 2 |
| 2004 | Automatic description of complex buildings from multiple images
Zu Whan Kim, Ramakant Nevatia |
Comput. Vis. Image Underst. | 2 |
| 2004 | Tracking Multiple Humans in Complex SituationsabstractTracking multiple humans in complex situations is challenging. The difficulties are tackled with appropriate knowledge in the form of various models in our approach. Human motion is decomposed into its global motion and limb motion. In the first part, we show how multiple human objects are segmented and their global motions are tracked in 3D using ellipsoid human shape models. Experiments show that it successfully applies to the cases where a small number of people move together, have occlusion, and cast shadow or reflection. In the second part, we estimate the modes (e.g., walking, running, standing) of the locomotion and 3D body postures by making inference in a prior locomotion model. Camera model and ground plane assumptions provide geometric constraints in both parts. Robust results are shown on some difficult sequences. Tao Zhao 0001, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | Bayesian Human Segmentation in Crowded SituationsabstractThe problem of segmenting individual humans in crowded situations from stationary video camera sequences is exacerbated by object inter-occlusion. We pose this problem as a "model-based segmentation" problem in which human shape models are used to interpret the foreground in a Bayesian framework. The solution is obtained by using an efficient Markov chain Monte Carlo (MCMC) method that uses domain knowledge as proposal probabilities. Knowledge of various aspects including human shape, human height, camera model, and image cues including human head candidates, foreground/background separation are integrated in one theoretically sound framework. We show promising results and evaluations on some challenging data. Tao Zhao 0001, Ramakant Nevatia |
CVPR (2) | 2 |
| 2003 | Large-Scale Event Detection Using Semi-Hidden Markov ModelsabstractWe present a new approach to recognizing events in videos. We first detect and track moving objects in the scene. Based on the shape and motion properties of these objects, we infer probabilities of primitive events frame-by-frame by using Bayesian networks. Composite events, consisting of multiple primitive events, over extended periods of time are analyzed by using a hidden, semi-Markov finite state model. This results in more reliable event segmentation compared to the use of standard HMMs in noisy video sequences at the cost of some increase in computational complexity. We describe our approach to reducing this complexity. We demonstrate the effectiveness of our algorithm using both real-world and perturbed data. Somboon Hongeng, Ramakant Nevatia |
ICCV | 2 |
| 2003 | Car detection in low resolution aerial images
Tao Zhao 0001, Ramakant Nevatia |
Image Vis. Comput. | 2 |
| 2003 | Improved Rooftop Detection in Aerial Images with Machine Learning
Marcus A. Maloof, Pat Langley, Thomas O. Binford, Ramakant Nevatia, Stephanie Sage |
Mach. Learn. | 4 |
| 2003 | Expandable Bayesian Networks for 3D Object Description from Multiple Views and Multiple Mode InputsabstractComputing 3D object descriptions from images is an important goal of computer vision. A key problem here is the evaluation of a hypothesis based on evidence that is uncertain. There have been few efforts on applying formal reasoning methods to this problem. In multiview and multimode object description problems, reasoning is required on evidence features extracted from multiple images and nonintensity data. One challenge here is that the number of the evidence features varies at runtime because the number of images being used is not fixed and some modalities may not always be available. We introduce an augmented Bayesian network, the expandable Bayesian network (EBN), which instantiates its structure at runtime according to the structure of input. We introduce the use of hidden variables to handle correlation of evidence features across images. We show an application of an EBN to a multiview building description system. Experimental results show that the proposed method gives significant and consistent performance improvement to others. Zu Whan Kim, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Automatic and interactive modeling of buildings in urban environments from aerial imagesabstractAutomatically extracting object models from images is a complex task. We describe research in extracting 3D models of buildings from aerial images. This work has resulted in several related systems including assisted extraction (minimal manual interaction to guide automatic processing), automatic extraction with limited imagery and limited building models, and automatic extraction with very good imagery and digital elevation models and more complex building models. Some results are provided for the assisted system and one of the automatic systems. Ramakant Nevatia, Keith E. Price |
ICIP (3) | 1 |
| 2002 | Automatic Pose Estimation of Complex 3D Building Modelsabstract3D models of urban sites with geometry and facade textures are needed for many planning and visualization applications. Approximate 3D wireframe model can be derived from aerial images but detailed textures must be obtained from ground level images. Integrating such views with the 3D models is difficult as only small parts of buildings may be visible in a single view. We describe a method that uses two or three vanishing points, and three 3D to 2D line correspondences to estimate the rotational and translational parameters of the ground level cameras. The valid set of multiple combinations of 3D to 2D line pairs is chosen by a hypotheses generation and evaluation Some experimental results are presented. Sung Chun Lee, Soon Ki Jung, Ramakant Nevatia |
WACV | 3 |
| 2002 | Automatic Integration of Facade Textures into 3D Building Models with a Projective Geometry Based Line ClusteringabstractVisualization of city scenes is important for many applications including entertainment and urban mission planning. Models covering wide areas can be efficiently constructed from aerial images. However, only roof details are visible from aerial views; ground views are needed to provide details of the building facades for high quality 'fly-through' visualization or simulation applications. We present an automatic method of integrating facade textures from ground view images into 3D building models for urban site modeling. We first segment the input image into building facade regions using a hybrid feature extraction method, which combines global feature extraction with Hough transform on an adaptively tessellated Gaussian Sphere and local region segmentation. We estimate the external camera parameters by using the corner points of the extracted facade regions to integrate the facade textures into the 3D building models. We validate our approach with a set of experiments on some urban sites. Categories and Subject Descriptors (according to ACM CCS): I.3.3 [Computer Graphics]: Modeling packages Sung Chun Lee, Soon Ki Jung, Ramakant Nevatia |
Comput. Graph. Forum | 3 |
| 2001 | Automatic Description of Buildings with Complex Rooftops from Multiple ImagesabstractWe present a model-based approach to detecting and describing compositions of buildings with complex rooftops. Previous approaches have dealt with either simpler models or models which lack geometric information. In spite of increasing model complexity, we maintain the computation affordable by effectively using multiple overlapping images. We obtain rooftop hypotheses in 3-D by using 3-D lines and junctions generated from multiple images. Image-derived unedited elevation data is used to assist feature matching, and to generate rough cues of the presence of 3-D structures. Experimental results are shown on complex buildings. Zu Whan Kim, Andres Huertas, Ramakant Nevatia |
CVPR (2) | 3 |
| 2001 | Segmentation and Tracking of Multiple Humans in Complex SituationsabstractSegmenting and tracking multiple humans is a challenging problem in complex situations in which extended occlusion, shadow and/or reflection exists. We tackle this problem with a 3D model-based approach. Our method includes two stages, segmentation (detection) and tracking. Human hypotheses are generated by shape analysis of the foreground blobs using a human shape model. The segmented human hypotheses are tracked with a Kalman filter with explicit handling of occlusion. Hypotheses are verified while being tracked for the first second or so. The verification is done by walking recognition using an articulated human walking model. We propose a new method to recognize walking using a motion template and temporal integration. Experiments show that our approach works robustly in very challenging sequences. Tao Zhao 0001, Ramakant Nevatia, Fengjun Lv |
CVPR (2) | 2 |
| 2001 | Multi-Agent Event RecognitionabstractThis paper presents a new approach to recognizing multiagent events observed by a static camera. To track objects robustly, knowledge about the ground plane and the events is used. An event is considered as composed of action threads, each thread being executed by a single actor. A single thread of action is recognized from the characteristics of the trajectory and moving blob of the actor using Bayesian methods. A multi-agent event is represented by a number of action threads related by temporal constraints. Multi-agent events are recognized by propagating the constraints and likelihoods of event threads in a temporal logic network. Somboon Hongeng, Ramakant Nevatia |
ICCV | 2 |
| 2001 | Car Detection in Low Resolution Aerial Image
Tao Zhao 0001, Ramakant Nevatia |
ICCV | 2 |
| 2001 | Event Detection and Analysis from Video StreamsabstractWe present a system which takes as input a video stream obtained from an airborne moving platform and produces an analysis of the behavior of the moving objects in the scene. To achieve this functionality, our system relies on two modular blocks. The first one detects and tracks moving regions in the sequence. It uses a set of features at multiple scales to stabilize the image sequence, that is, to compensate for the motion of the observer, then extracts regions with residual motion and uses an attribute graph representation to infer their trajectories. The second module takes as input these trajectories, together with user-provided information in the form of geospatial context and goal context to instantiate likely scenarios. We present details of the system, together with results on a number of real video sequences and also provide a quantitative analysis of the results. Gérard G. Medioni, Isaac Cohen, François Brémond, Somboon Hongeng, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2001 | Detection and Modeling of Buildings from Multiple Aerial ImagesabstractAutomatic detection and description of cultural features, such as buildings, from aerial images is becoming increasingly important for a number of applications. This task also offers an excellent domain for studying the general problems of scene segmentation, 3D inference, and shape description under highly challenging conditions. We describe a system that detects and constructs 3D models for rectilinear buildings with either flat or symmetric gable roofs from multiple aerial images; the multiple images, however, need not be stereo pairs (i.e., they may be acquired at different times). Hypotheses for rectangular roof components are generated by grouping lines in the images hierarchically; the hypotheses are verified by searching for presence of predicted walls and shadows. The hypothesis generation process combines the tasks of hierarchical grouping with matching at successive stages. Overlap and containment relations between 3D structures are analyzed to resolve conflicts. This system has been tested on a large number of real examples with good results, some of which are included in the paper along with their evaluations. Sanjay Noronha, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2000 | Representation and Optimal Recognition of Human ActivitiesabstractTowards the goal of realizing a generic automatic human activity recognition system, a new formalism is proposed. Activities are described by a chained hierarchical representation using three type of entities: image features, mobile object properties and scenarios. Taking image features of tracked moving regions from an image sequence as input, mobile object properties are first computed by specific methods while noise is suppressed by statistical methods. Scenarios are recognized from mobile object properties based on Bayesian analysis. Several scenarios are recognized by an algorithm using a probabilistic finite-state automaton (a variant of structured HMM). A demonstration of the optimality of this recognition method is discussed. Finally, the validity and the effectiveness of our approach is demonstrated on both real-world and perturbed data. Somboon Hongeng, François Brémond, Ramakant Nevatia |
CVPR | 3 |
| 2000 | Multisensor Integration for Building ModelingabstractMachine perception can benefit from the use of features extracted from data provided by a variety of sensor modalities. Recent advances in sensor design makes it possible to incorporate multiple sensors into vision systems for increased capability. Two important issues must be considered for the integration task. The sensors must be spatially coregistered and the phenomenologies must be compatible. In this paper we address these issues as they apply to the problem of automatic modeling of building structures from aerial views. We present a methodology to incorporate cues extracted from IFSAR (Interferometric Synthetic Aperture Radar) to significantly improve the performance and the quality of the results of an existing system that relies on electro-optical panchromatic images, while reducing processing time. Quantitative evaluations are given. Andres Huertas, Zu Whan Kim, Ramakant Nevatia |
CVPR | 3 |
| 2000 | Learning Bayesian Networks for Diverse and Varying numbers of Evidence Sets
Zu Whan Kim, Ramakant Nevatia |
ICML | 2 |
| 2000 | Bayesian Framework for Video Surveillance ApplicationabstractThe goal of this paper is to describe and demonstrate the application of Bayesian networks in a generic automatic video surveillance system. Taking image features of tracked moving regions from an image sequence as input, mobile object properties are first computed and noise is suppressed by statistical methods. The probability that a scenario occurs is then computed from these mobile object properties through several layers of naive Bayesian classifiers (or a Bayesian network). Several issues and solutions regarding the efficiency of the Bayesian network are discussed. For example, the parameters of the networks, which represent rare activities (typical of video surveillance applications), can be learned from image sequences of similar scenarios which are more common. We demonstrate the effectiveness of our approach by training the networks with 600 image frames belonging to one domain of interest and applying them to image sequences in a different domain. Somboon Hongeng, François Brémond, Ramakant Nevatia |
ICPR | 3 |
| 2000 | Automatic description of complex buildings with multiple imagesabstract3-D building detection and description is a practical application of 3-D object description, a key task of computer vision. We present an approach to detecting and describing buildings of polygonal rooftops by using multiple, overlapping images of the scene. First, 3-D features are generated by using multiple images, and rooftop hypotheses are generated by neighborhood searches on those features. For robust generation of 3-D features, we present a probabilistic approach to address the epipolar alignment problem in line matching. Image-derived unedited elevation data is used to assist feature matching, and to generate rough cues of the presence of 3-D structures. These cues help reduce the search space significantly. Experimental results are shown on some complex buildings. Zu Whan Kim, Andres Huertas, Ramakant Nevatia |
WACV | 3 |
| 2000 | Modeling 3-D complex buildings with user assistanceabstractAn effective 3D method incorporating user assistance for modeling complex buildings is proposed. This method utilizes the connectivity and similar structure information among unit blocks in a multi-component building structure, to enable the user to incrementally construct models of many types of buildings. The system attempts to minimize the time and the number of user interactions needed to assist an existing automatic system in this task. Several examples are presented that demonstrate significant improvement and efficiency compared with other approaches and with purely manual systems. Sung Chun Lee, Andres Huertas, Ramakant Nevatia |
WACV | 3 |
| 2000 | Detecting changes in aerial views of man-made structures
Andres Huertas, Ramakant Nevatia |
Image Vis. Comput. | 2 |
| 1999 | User Assisted Modeling of Buildings from Aerial ImagesabstractAn approach that allows a user to assist an automatic system in modeling buildings is described. The approach is designed to be efficient in user time and effort while preserving the quality of the models created. Currently our system is able to handle the rectangular buildings with flat roof or symmetric gabled roof. Models can be created by only one or two clicks in many cases. Efficient editing of automatically derived models is also possible. Ramakant Nevatia, Sanjay Noronha |
CVPR | 2 |
| 1999 | Uncertain Reasoning and Learning for Feature Grouping
Zu Whan Kim, Ramakant Nevatia |
Comput. Vis. Image Underst. | 2 |
| 1999 | Part-Based 3D Descriptions of Complex Objects from a Single ImageabstractVolumetric, 3D, part-based descriptions of complex objects in a scene can be highly beneficial for many tasks such as generic object recognition, navigation, and manipulation. However, it has been difficult to derive such descriptions from image data. There has been some progress in getting such descriptions from range data or from perfect contours, but analysis of a real intensity image presents many difficulties. The object and part boundaries do not completely correspond to image boundaries. The detected boundaries are often fragmented and many boundaries due to surface markings, shadows, and noise are present. In addition, inference of 3D from a 2D image is difficult. The paper describes a method to compute the desired descriptions from a single image by exploiting projective properties of a class of generalized cylinders and of possible joints between them. Experimental results on some examples are given. Mourad Zerroug, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1998 | Recent Advances in Detection and Description of Buildings from Multiple Aerial Images
Sanjay Noronha, Ramakant Nevatia |
ACCV (2) | 2 |
| 1998 | Detecting Changes in Aerial Views of Man-Made StructuresabstractMany applications require detecting structural changes in a scene over a period of time. Comparing intensity values of successive images is not effective as such changes don't necessarily reflect actual changes at a site but might be caused by changes in the view point, illumination and seasons. We take the approach of comparing a 3-D model of the site, prepared from previous images, with new images to infer significant changes. This task is difficult as the images and the models have very different levels of abstract representations. Our approach consists of several steps: registering a site model to a new image, model validation to confirm the presence of model objects in the image; structural change detection seeks to resolve matching problems and indicate possibly changed structures; and finally updating models to reflect the changes. Our system is able to detect missing (or mis-modeled) buildings, changes in model dimensions, and new buildings under some conditions. Andres Huertas, Ramakant Nevatia |
ICCV | 2 |
| 1998 | Generalizing over aspect and location for rooftop detectionabstractWe present the results of an empirical study in which we evaluated cost-sensitive learning algorithms on a rooftop detection task, which is one level of processing in a building detection system. Specifically, we investigated how well machine learning methods generalized to unseen images that differed in location and in aspect. For the purpose of comparison, we included in our evaluation a handcrafted linear classifier, which is the selection heuristic currently used in the building detection system. ROC analysis showed that, when generalizing to unseen images that differed in location and aspect, a naive Bayesian classifier outperformed nearest neighbor and the handcrafted solution. Marcus A. Maloof, Pat Langley, Thomas O. Binford, Ramakant Nevatia |
WACV | 4 |
| 1998 | Automatic Building Extraction from Aerial Images
Armin Grün, Ramakant Nevatia |
Comput. Vis. Image Underst. | 2 |
| 1998 | Building Detection and Description from a Single Intensity Image
Chungan Lin, Ramakant Nevatia |
Comput. Vis. Image Underst. | 2 |
| 1998 | Recognition and localization of generic objects for indoor navigation using functionality
Ramakant Nevatia |
Image Vis. Comput. | 2 |
| 1997 | Detection and Description of Buildings from Multiple Aerial ImagesabstractA method for detection and description of rectangular buildings from two or more registered aerial intensity images is proposed. The output is a 3D description of the buildings, with an associated confidence measure for each building. Hierarchical perceptual grouping and matching across views is employed to increase the robustness of the system. Verification of selected building hypotheses is done using shadow and wall evidence of the buildings. The system is largely feature-based. Grouping and matching are performed in a hierarchical manner utilizing primitives of increasing complexity, starting with line segments and junctions, and proceeding to higher level features. Binocular and trinocular epipolar constraints are used to reduce the search space for matching features. Sanjay Noronha, Ramakant Nevatia |
CVPR | 2 |
| 1996 | Load balancing strategies for symbolic vision computationsabstractMost intermediate and high-level vision algorithms manipulate symbolic features. A key operation in these vision algorithms is to search symbolic features satisfying certain geometric constraints. Parallelizing this symbolic search needs a non-trivial algorithmic technique due to the unpredictable workload. In this paper, we propose load balancing strategies for parallelizing symbolic search operations on distributed memory machines. By using an initial workload estimate, we first partition the computations such that the workload is distributed evenly across the processors. In addition, we perform fast migrations dynamically to adapt to the evolving workload. To demonstrate the usefulness of our load balancing strategies, experiments were conducted on an IBM SP2 and a Cray T3D. Our results show that our task migration strategy can balance the unpredictable workload with little overhead. Our code using C and MPI is portable onto other high performance computing platforms. Yongwha Chung, Jongwook Woo, Ramakant Nevatia, Viktor Prasanna 0001 |
HiPC | 3 |
| 1996 | Recovering LSHGCs and SHGCs from stereo
Ronald Chung, Ramakant Nevatia |
Int. J. Comput. Vis. | 2 |
| 1996 | Computer vision research at the University of Southern California
Ramakant Nevatia, Gérard G. Medioni |
Int. J. Comput. Vis. | 1 |
| 1996 | Volumetric descriptions from a single intensity image
Mourad Zerroug, Ramakant Nevatia |
Int. J. Comput. Vis. | 2 |
| 1996 | Three-Dimensional Descriptions Based on the Analysis of the Invariant and Quasi-Invariant Properties of Some Curved-Axis Generalized CylindersabstractWe address the recovery of object-level 3D descriptions of some classes of curved-axis generalized cylinders. For this, the first part of the paper analyzes the projective properties of two common generic shapes, planar right constant generalized cylinders (PRCGCs) and circular planar right generalized cylinders (circular PRGCs). The properties we analyze include new geometric invariant and quasi-invariant properties of the orthographic projection of the above shapes and a useful classification of their structural properties as functions of their pose. The second part of the paper describes an implemented system which detects and recovers PRCGCs and circular PRGCs from an intensity image in the presence of noise, surface markings, shadows, and partial occlusion. The methods exploit the projective properties to hypothesize and verify relevant curved-axis objects, thus explicitly using the three-dimensionality of the objects and of the desired descriptions. This work extends past work on the recovery of volumetric shapes from an intensity image by addressing new primitives, deriving new properties and by developing a system that recovers them from an intensity image. We demonstrate our method on several real intensity images. Mourad Zerroug, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1995 | Use of Monocular Groupings and Occlusion Analysis in a Hierarchical Stereo System
Ronald Chung, Ramakant Nevatia |
Comput. Vis. Image Underst. | 2 |
| 1995 | Shape from Contour: Straight Homogeneous Generalized Cylinders and Constant Cross Section Generalized CylindersabstractWe analyze the properties of straight homogeneous generalized cylinders (SHGCs) and constant cross section generalized cylinders (CGCs), and derive the types of symmetries that the limb boundaries and cross sections of these objects produce on the image plane. The constraints on the 3D shape of the objects are formulated based on the symmetries and from the geometry of the projection models. Finally, the methods that recover the 3D shape from the image of their contours are discussed and recovered surfaces are shown for sample objects.> Fatih Ulupinar, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1994 | Representation and computation of the spatial environment for indoor navigationabstractWe introduce a spatial representation, s-map, for an indoor navigation robot. The s-map represents the locations of obstacles in a planar domain, where obstacles are defined as any objects that can block movement of the robot. In building the s-map, the viewing triangle constraint and the stability constraint are introduced for efficient verification of vertical surfaces. These verified vertical surfaces and 3-D segments of obstacles smaller than a robot, are mapped to the s-map by simply dropping height information. Thus, the s-map is made directly from 3-D segments with simple verification, and represents obstacles in a planar domain so that it becomes a navigable map for the robot without further processing. In addition to efficient map building, the s-map represents the environment more realistically and completely. Furthermore, the s-map converts many navigation problems in 3-D, such as map fusion and path planning, into 2-D ones. We present the analysis of the s-map in terms of complexity and reliability, and discuss its pros and cons. Moreover, we show the results of the s-maps for indoor environments.> Ramakant Nevatia |
CVPR | 2 |
| 1994 | Detection of buildings using perceptual grouping and shadowsabstractWe describe a system for detection and description of buildings in aerial scenes. This is a difficult task as the aerial images contain a variety of objects. Low-level segmentation processes give highly fragmented segments due to a number of reasons. We use a perceptual grouping approach to collect these fragments and discard those that come from other sources. We use shape properties of the buildings for this. We use shadows to help form and verify the hypotheses generated by the grouping process. This latter step also provides 3-D descriptions of the buildings. Our system has been tested on a number of examples and is able to work with overhead or oblique views.> Chungan Lin, Andres Huertas, Ramakant Nevatia |
CVPR | 3 |
| 1994 | Segmentation and Recovery of SHGCs from a Real Intensity Image
Mourad Zerroug, Ramakant Nevatia |
ECCV (1) | 2 |
| 1994 | Parallel processing for spatial grouping and matchingabstractIn this paper, we identify the computational requirements for structural pattern analysis, particularly for the operations of spatial grouping and matching. We describe two such algorithms that are in wide use here at USC and discuss approaches to reducing their execution times via parallel implementation. We provide brief descriptions and results of two research projects geared generally, toward the parallel implementation of computer vision systems and specifically, towards these algorithms. Ramakant Nevatia, Craig C. Reinhart |
ICPR (3) | 1 |
| 1994 | From an intensity image to 3-D segmented descriptionsabstractAddresses the inference of 3-D segmented descriptions of complex objects from a single intensity image. The authors' approach is based on the analysis of the projective properties of a small number of generalized cylinder primitives and their relationships in the image which make up common man-made objects. Past work on this problem has either assumed perfect contours as input or used 2-dimensional shape primitives without relating them to 3-D shape. The method the authors present explicitly uses the 3-dimensionality of the desired descriptions and directly addresses the segmentation problem in the presence of contour breaks, markings shadows and occlusion. This work has many significant applications including recognition of complex curved objects from a single real intensity image. The authors demonstrate their method on real images. Mourad Zerroug, Ramakant Nevatia |
ICPR (1) | 2 |
| 1994 | Segmentation and 3-D recovery of curved-axis generalized cylinders from an intensity imageabstractAddresses the problem of segmentation and recovery of 3-D object-centered descriptions of two large sub-classes of curved axis generalized cylinders, PRCGCs and circular PRGCs, from a single real intensity image. The purpose of this work is to augment the set of 3-D primitives which can be recovered, beyond previous work which has addressed mainly straight axis ones, so that more complex objects can be handled. The authors' approach is based on the exploitation of geometric projective properties as well as structural properties of the contours of circular PRGCs. The implemented method works in the presence of noise, contour breaks, markings, shadows and occlusion. The authors demonstrate their method on real images. Mourad Zerroug, Ramakant Nevatia |
ICPR (1) | 2 |
| 1994 | Model validation for change detection [machine vision]abstractAn important application of machine vision is to provide a means to monitor a scene over a period of time and report changes in the content of the scene. We have developed a validation mechanism that implements the first step towards a system for detecting changes in images of aerial scenes. By validation we mean the confirmation of the presence of model objects in the image. Our system uses a 3-D site model of the scene as a basis for model validation, and eventually for detecting changes and to update the site model. The scenario for our present validation system consists of adding a new image to a database associated with the site. The validation process is implemented in three steps: registration of the image to the model, or equivalently, determination of the position and orientation of the camera; matching of model features to image features; and validation of the objects in the model. Our system processes the new image monocularly and uses shadows as 3-D clues to help validate the model. The system has been tested using a hand-generated site model and several images of a 500:1 scale model of the site, acquired form several viewpoints.> Mathias Bejanin, Andres Huertas, Gérard G. Medioni, Ramakant Nevatia |
WACV | 4 |
| 1994 | A method for recognition and localization of generic objects for indoor navigationabstractWe introduce an efficient method for recognition and localization of generic objects for robot navigation, which works on real scenes. The generic objects used in our experiments are desks and doors as they are suitable landmarks for navigation. The recognition method uses significant surfaces and accompanying functional evidence for recognition of such objects. Currently, our system works with planar surfaces only and assumes that the objects are in a "standard" pose. The localization and orientation of an object are represented with the most significant surface in an "s-map". Some results for laboratory scenes are given.> Ramakant Nevatia |
WACV | 2 |
| 1994 | Recovery of 3-D Objects with Multiple Curved Surfaces from 2-D Contours
Fatih Ulupinar, Ramakant Nevatia |
Artif. Intell. | 2 |
| 1993 | Quasi-invariant properties and 3-D shape recovery of non-straight, non-constant generalized cylindersabstractThe geometric protective properties of the contours of right generalized cylinders with a planar, but not necessarily straight, axis and circular, but possibly varying in size, cross-sections (called circular PRGCs) are addressed. Important rigourous quasi-invariant properties of circular PRGCs and invariant properties for their subclasses are derived. Their application for 2-D descriptions and for recovery of complete 3-D object-centered descriptions from the 2-D contours is shown.> Mourad Zerroug, Ramakant Nevatia |
CVPR | 2 |
| 1993 | Perception of 3-D Surfaces from 2-D ContoursabstractInference of 3-D shape from 2-D contours in a single image is an important problem in machine vision. The authors survey classes of techniques proposed in the past and provide a critical analysis. They show that two kinds of symmetries in figures, which are known as parallel and skew symmetries, give significant information about surface shape for a variety of objects. They derive the constraints imposed by these symmetries and show how to use them to infer 3-D shape. They also discuss the zero Gaussian curvature (ZGC) surfaces in depth and show results on the recovery of surface orientation for various ZGC surfaces.> Fatih Ulupinar, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1992 | Recovering LSHGCs and SHGCs from stereoabstractThe problem of computing volumetric shape from stereo is examined. It is argued that intermediate two-and-one-half-dimensional dense or wire-frame descriptions may not be always possible from stereo, especially when there are curved surfaces in the scene, and that 3D volumetric descriptions of objects may have to be derived directly from stereo correspondences. Methods are then presented for recovering volumetric shape with linear straight homogeneous generalized cones (LSHGCs) and straight homogeneous generalized cones (SHGCs) as the shape models, using some invariant properties in their monocular and stereo projections. Experimental results on images of objects with curved surfaces are given.> Ronald C.-K. Chung, Ramakant Nevatia |
CVPR | 2 |
| 1992 | Recovery of 3-D objects with multiple curved surfaces from 2-D contoursabstractThe authors describe a technique for inference of 3-D shape from 2-D contours that utilizes not only the shapes of individual surfaces but also the interactions between them. The analysis applies to objects made of zero-Gaussian curvature surfaces viewed under orthographic projection.> Fatih Ulupinar, Ramakant Nevatia |
CVPR | 2 |
| 1992 | Description and tracking of moving articulated objectsabstractProposes a method to obtain reliable shape description of articulated objects by integrating initial descriptions computed from different view images. Ribbon, which is a 2-D analog of a generalized cone, is used as the basic shape representation scheme. An initial description for each frame is the collection of composed ribbons, which is obtained after filtering out most inadequate ribbons and grouping the remaining ribbons by using geometric constraints. Ribbon matching is then conducted between different frames and ribbons which match are retained. From the retained ribbons, the geometric constraints make integrated descriptions and the tracking of parts is established from the ribbon matching results. Since the ribbon matching allows one ribbon to be matched with two ribbons, an articulation which is not detected in one frame but is detected in another frame can be recovered. Experimental results are also shown.> Shoji Kurakake, Ramakant Nevatia |
ICPR (1) | 2 |
| 1992 | Issues in parallel tree search for object recognitionabstractThe authors describe the parallel implementation of a 3D object recognition algorithm. The algorithm is representative of methods utilized by various computer vision researchers and presents some interesting problems that are generally overlooked by parallel processing researchers that have studied tree search problems. They describe their objectives in developing the parallel implementation and discuss its performance. They also (briefly) discuss the affects that the parallel implementation has on the runtime characteristics of the algorithm.> Craig C. Reinhart, Ramakant Nevatia |
ICPR (4) | 2 |
| 1992 | Recovering building structures from stereoabstractAddresses the problem of extracting polyhedral building structures from a stereo pair of aerial intensity images. The authors describe a system that computes a hierarchy of descriptions such as segments, junctions, and links between junctions from each view, and matches these features at the different levels. Such high level features not only help reduce correspondence ambiguity during stereo matching, but also allow us to infer surface boundaries even though the boundaries may be broken because of noise and weak contrast. The authors hypothesize surface boundaries by examining global information such as continuity and coplanarity of linked edges in 3-D, rather than merely by looking at local depth information. When the walls of the buildings are visible, they also exploit the relationship among adjacent surfaces in a polyhedral object to help confirm the different levels of descriptions. The authors give some experimental results for aerial images taken from overhead views and oblique views.> Ronald Chung, Ramakant Nevatia |
WACV | 2 |
| 1992 | Perceptual Organization for Scene Segmentation and DescriptionabstractA data-driven system for segmenting scenes into objects and their components is presented. This segmentation system generates hierarchies of features that correspond to structural elements such as boundaries and surfaces of objects. The technique is based on perceptual organization, implemented as a mechanism for exploiting geometrical regularities in the shapes of objects as projected on images. Edges are recursively grouped on geometrical relationships into a description hierarchy ranging from edges to the visible surfaces of objects. These edge groupings, which are termed collated features, are abstract descriptors encoding structural information. The geometrical relationships employed are quasi-invariant over 2-D projections and are common to structures of most objects. Thus, collations have a high likelihood of corresponding to parts of objects. Collations serve as intermediate and high-level features for various visual processes. Applications of collations to stereo correspondence, object-level segmentation, and shape description are illustrated.> Rakesh Mohan, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1991 | Use of monocular groupings and occlusion analysis in a hierarchical stereo systemabstractA hierarchical stereo system is described that uses structural descriptions up to the surface level. Surface descriptions are computed from monocular images, by using a perceptual grouping technique. Occlusion can be a major problem in stereo analysis and is often not treated explicitly. An analysis is presented of occlusion effects in stereo, and it is shown how structural descriptions can be used to deal with them. Experimental results are given for scenes with curved objects and significant occlusions.> Ronald Chung, Ramakant Nevatia |
CVPR | 2 |
| 1991 | Recovering shape from contour for constant cross section generalized cylindersabstractThe authors analyze the properties of constant cross section generalized cylinders (CGCs), and derive the types of symmetries that the limb boundaries and cross sections of these objects produce on the image plane. The constraints on the 3-D shape of the objects are formulated based on the symmetries and from the geometry of the projection models. The methods that recover the 3-D shape from the image of their contours are discussed and recovered surfaces are shown for sample objects.> Fatih Ulupinar, Ramakant Nevatia |
CVPR | 2 |
| 1991 | Constraints for interpretation of line drawings under perspective projection
Fatih Ulupinar, Ramakant Nevatia |
CVGIP Image Underst. | 2 |
| 1990 | Shape from contour: straight homogeneous generalized conesabstractThe authors analyze the properties of straight homogeneous generalized cones (SHGCs) and derive the types of symmetries, that the limb boundaries and cross sections of these objects produce on the image plane. The constraints on the 3-D shape of the objects are formulated based on the symmetries and from the geometry of the projection models. Finally the methods that recover the 3-D shape from the image of their contours are discussed and recovered surfaces are shown for sample objects.> Fatih Ulupinar, Ramakant Nevatia |
ICCV | 2 |
| 1990 | Shape description from imperfect and incomplete dataabstractUsually, shape description systems assume that a scene has been segmented into objects and that object boundaries are given. This, however, is not realistic when working with intensity images; the resulting boundaries are fragmented and contain surface markings, and shadow and noise boundaries. A system is described which works with such input and computes shape descriptions of complex objects. Scene segmentation takes place through shape description. Generalized cones or, more precisely, their 2D analogs of ribbons are used as the basic shape representation scheme. Results for synthetic and real examples are shown. The output of the system is useful for object recognition, learning, further inference of 3D shape, grasping, and navigation.> Kashipati Rao, Ramakant Nevatia |
ICPR (1) | 2 |
| 1990 | Inferring shape from contour for curved surfacesabstractA technique based on analysis of symmetries in an image is proposed for inferring the 3-D shapes of surfaces of objects in it. This technique is analyzed and applied to zero-Gaussian-curvature surfaces. The method consists of deriving a number of constraints based on a few simple assumptions. The combination of constraints to give unique (or few) solutions is discussed. Experimental results on selected scenes are given and are shown to conform well with human perception.> Fatih Ulupinar, Ramakant Nevatia |
ICPR (1) | 2 |
| 1990 | Detecting runways in complex airport scenes
Andres Huertas, William Cole, Ramakant Nevatia |
Comput. Vis. Graph. Image Process. | 3 |
| 1989 | Segmentation and description based on perceptual organizationabstractThe authors present a description framework, motivated by perceptual organization, which consists of representations of the geometrical organizations of intensity discontinuities. The descriptors in this framework are called collated features, and are groupings identified by perceptual organization. The processes that operate on the image to obtain these descriptors and the visual processes that utilize them are discussed. The detection of collated features is robust to local problems. The structural information encoded in them aids various visual tasks such as object segmentation, correspondence processes (stereo, motion, and model matching), and shape inferences. Two primary grouping processes, cocurvilinearity and symmetry are applied to intensity edge contours to generate the collated features, including curves, symmetries, and ribbons. These collations can be used to segment into visible surfaces of objects and to describe the 2D shapes of those surfaces.> Rakesh Mohan, Ramakant Nevatia |
CVPR | 2 |
| 1989 | Using Generic Knowledge in Analysis of Aerial Scenes: A Case Study
Andres Huertas, William Cole, Ramakant Nevatia |
IJCAI | 3 |
| 1989 | Dissertation abstracts
Steven L. Tanimoto, Sargur N. Srihari, Martin D. Levine, Warren P. Seering, Ramakant Nevatia |
Mach. Vis. Appl. | 5 |
| 1989 | Recognizing 3-D Objects Using Surface DescriptionsabstractThe authors provide a complete method for describing and recognizing 3-D objects, using surface information. Their system takes as input dense range date and automatically produces a symbolic description of the objects in the scene in terms of their visible surface patches. This segmented representation may be viewed as a graph whose nodes capture information about the individual surface patches and whose links represent the relationships between them, such as occlusion and connectivity. On the basis of these relations, a graph for a given scene is decomposed into subgraphs corresponding to different objects. A model is represented by a set of such descriptions from multiple viewing angles, typically four to six. Models can therefore be acquired and represented automatically. Matching between the objects in a scene and the models is performed by three modules: the screener, in which the most likely candidate views for each object are found; the graph matcher, which compares the potential matching graphs and computes the 3-D transformation between them; and the analyzer, which takes a critical look at the results and proposes to split and merge object graphs.> Ting-Jun Fan, Gérard G. Medioni, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1989 | Stereo Error Detection, Correction, and EvaluationabstractAn algorithm is presented for error detection and correction of disparity, as a process separate from stereo matching, with the contention that matching is not necessarily the best way to utilize all the physical constraints characteristic to stereopsis. As a result of the bias in stereo research towards matching, vision tasks like surface interpolation and object modeling have to accept erroneous data from the stereo matchers without the benefits of any intervening stage of error correction. An algorithm which identifies all errors in disparity data that can be detected on the basis of figural continuity and corrects them is presented. The algorithm can be used as a postprocessor to any edged-based stereo matching algorithm, and can additionally be used to automatically provide quantitative evaluations on the performance of matching algorithms of this class.> Rakesh Mohan, Gérard G. Medioni, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1989 | Using Perceptual Organization to Extract 3-D StructuresabstractThe authors describe an approach to perceptual grouping for detecting and describing 3-D objects in complex images and apply it to the task of detecting and describing complex buildings in aerial images. They argue that representations of structural relationships in the arrangements of primitive image features, as detected by the perceptual organization process, are essential for analyzing complex imagery. They term these representations collated features. The choice of collated features is determined by the generic shape of the desired objects in the scene. The detection process for collated features is more robust than the local operations for region segmentation and contour tracing. The important structural information encoded in collated features aids various visual tasks such as object segmentation, correspondence processes, and shape description. The proposed method initially detects all reasonable feature groupings. A constraint satisfaction network is then used to model the complex interactions between the collations and select the promising ones. Stereo matching is performed on the collations to obtain height information. This aids in further reasoning on the collated features and results in the 3-D description of the desired objects.> Rakesh Mohan, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1988 | Using Symmetries For Analysis Of Shape From ContourabstractInference of 3-D shape from 2-D contours in a single image is an important problem in machine vision. We survey classes of techniques proposed in the past and provide a critical analysis. We propose two kinds of symmetries in figures, which we call parallel and mirror symmetries, give significant information about surface shape for a variety of objects. We show the constraints imposed by these symmetries and how to use them to infer 3-D shape. Our method is applicable to any zero-gaussian curvature surface, and also to a variety of doubly curved surfaces. One of our mathematical results is that for a cone, the surface shape can be constructed uniquely under very simple assumptions. We also show some preliminary results on extraction of symmetries from real images. Fatih Ulupinar, Ramakant Nevatia |
ICCV | 2 |
| 1988 | Matching 3-D objects using surface descriptionsabstractA method is developed to extract important curves, corresponding to physical boundaries of objects, from a range image. It is shown how to infer, from these curves, a segmentation of the scene into surface patches, and how to use these descriptions to establish correspondences between two scenes. In a first step, labeled curves corresponding to jump boundaries, creases, and limbs of objects are grouped into boundaries of regions. Each region is therefore described by its boundaries and by a polynomial approximation, which allows each path individually and also the complete scene to be reconstructed. In a second step, objects (or partial objects) are inferred from surface patches, and then two range images are at this partial object level. Graphs of objects are matched using a best-first search under three types of constraints: unary constraints between corresponding nodes, binary constraints between corresponding linked pairs of nodes, and constraints imposed by the computed geometric transformation. Substantial partial occlusion is allowed. The generality and robustness of this approach is illustrated by several examples.> Ting-Jun Fan, Gérard G. Medioni, Ramakant Nevatia |
ICRA | 3 |
| 1988 | Detecting buildings in aerial images
Andres Huertas, Ramakant Nevatia |
Comput. Vis. Graph. Image Process. | 2 |
| 1988 | Computing volume descriptions from sparse 3-D data
Kashipati Rao, Ramakant Nevatia |
Int. J. Comput. Vis. | 2 |
| 1987 | Detecting Runways in Aerial Images
Andres Huertas, William Cole, Ramakant Nevatia |
AAAI | 3 |
| 1987 | Segmented descriptions of 3-D surfacesabstractA method to segment and describe visible surfaces of three-dimensional (3-D) objects is presented by first segmenting the surfaces into simple surface patches and then using these patches and their boundaries to describe the 3-D surfaces. First, distinguished points are extracted which will comprise the edges of segmented surface patches, using the zero-crossings and extrema of curvature along a given direction. Two different methods are used: if the sensor provides relatively noise-free range images, the principal curvatures are computed at only one resolution, otherwise, a multiple scale approach is used and curvature is computed in four directions 45° apart to facilitate interscale tracking. These points are then grouped into curves and these curves are classified into different classes which correspond to significant physical properties such as jump boundaries, folds, and ridge lines (or smooth extrema). Then jump boundaries and folds are used to segment the surfaces into surface patches, and a simple surface is fitted to each patch to reconstruct the original objects. These descriptions not only make explicit most of the salient properties present in the original input, but are more suited to further processing, such as matching with a given model. The generality and robustness of this approach is illustrated on scene images with different available range sensors. Ting-Jun Fan, Gérard G. Medioni, Ramakant Nevatia |
IEEE J. Robotics Autom. | 3 |
| 1986 | Structural Analysis of Natural TexturesabstractMany textures can be described structurally, in terms of the individual textural elements and their spatial relationships. This paper describes a system to generate useful descriptions of natural textures in these terms. The basic approach is to determine an initial, partial description of the elements using edge features. This description controls the extraction of the texture elements. The elements are grouped by type, and spatial relationships between elements are computed. The descriptions are shown to be useful for recognition of the textures, and for reconstruction of periodic textures. Felicia M. Vilnrotter, Ramakant Nevatia, Keith E. Price |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1985 | Segment-based stereo matching
Gérard G. Medioni, Ramakant Nevatia |
Comput. Vis. Graph. Image Process. | 2 |
| 1984 | Matching Images Using Linear FeaturesabstractWe describe techniques for matching two images or an image and a map. This operation is basic for machine vision and is needed for the tasks of object recognition, change detection, map up-dating, passive navigation, and other tasks. Our system uses line-based descriptions, and matching is accomplished by a relaxation operation which computes most similar geometrical structures. A more efficient variation, called the ``kernel'' method, is also described. We give results on complex aerial images which contain many image differences, caused by varying sun position, different seasons, and imaging environments, and also structural changes caused by man-made alterations such as new construction. Gérard G. Medioni, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1984 | Visual inspection using linear features
Gérard G. Medioni, Ramakant Nevatia |
Pattern Recognit. | 3 |
| 1983 | Detection of Buildings in Aerial Images Using Shape and Shadows
Andres Huertas, Ramakant Nevatia |
IJCAI | 2 |
| 1982 | Locating Structures in Aerial ImagesabstractA technique for locating desired structures utilizing user specified information about properties of these structures and their relationships with other more easily extracted objects is described. An edge-based and region-based technique is used for scene segmentation. Experimental results of the processing of aerial pictures are presented. Ramakant Nevatia, Keith E. Price |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1979 | Describing Natural Textures
Ramakant Nevatia, Keith E. Price, Felicia M. Vilnrotter |
IJCAI | 1 |
| 1977 | Description and Recognition of Curved Objects
Ramakant Nevatia, Thomas O. Binford |
Artif. Intell. | 1 |
| 1976 | Locating Object Boundaries in Textured EnvironmentsabstractDetection of object boundaries is an important step in the analysis of an image. In the presence of a textured background, local edge operators generate many edges that do not correspond to the object boundaries. However, edges along the object boundaries link in elongated segments. An efficient algorithm to perform such linking is described and experimental results of a working program are presented. Ramakant Nevatia |
IEEE Trans. Computers | 1 |
| 1973 | Structured Descriptions of Complex Objects
Ramakant Nevatia, Thomas O. Binford |
IJCAI | 1 |