EDBT 2026 Demo / reviewers in the wild / expert
Harpreet Sawhney
dblp:86/689 · also Harpreet S. Sawhney
· DBLP profile ↗
92ranked-venue papers
18as first author
5since 2021 · last 2026
0009-0008-0487-4318ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 70 · 13 first-author · 4 since 2021Artificial intelligence and machine learning · 69 · 14 first-author · 4 since 2021Systems, architecture and hardware · 3Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
49 papers |
Video understanding and tracking · 33% 3D vision · 15% Vision and language · 8% | |
| Computer graphics and multimedia
27 papers |
Multimedia analysis and retrieval · 43% Image and video processing · 25% Virtual and augmented reality · 10% |
Topics — the 30 heaviest of 131, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Multi-agent systems
LLM-based multi-agent systems |
1.0 | 1 | 2026 | Schema-Guided Scene-Graph Reasoning Based on Multi-Agent Large Language Model System · AAAI 2026 |
Computer vision › Vision and language › visual reasoning
scene graph reasoning |
1.0 | 1 | 2026 | Schema-Guided Scene-Graph Reasoning Based on Multi-Agent Large Language Model System · AAAI 2026 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph reasoning
schema-guided reasoning |
1.0 | 1 | 2026 | Schema-Guided Scene-Graph Reasoning Based on Multi-Agent Large Language Model System · AAAI 2026 |
Computer vision › Video understanding and tracking › human motion prediction
hand motion prediction |
0.9 | 1 | 2025 | How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions · CVPR 2025 |
Computer vision › Face, body and person analysis › human pose estimation › articulated pose estimation
hand pose estimation |
0.9 | 1 | 2025 | How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions · CVPR 2025 |
Computer vision › Video understanding and tracking
action recognition |
0.7 | 1 | 2023 | Domain-Specific Priors and Meta Learning for Few-Shot First-Person Action Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.7 | 1 | 2023 | STEPs: Self-Supervised Key Step Extraction and Localization from Unlabeled Procedural Videos · ICCV 2023 |
Computer vision › Video understanding and tracking › action recognition › human action recognition
egocentric action recognition |
0.7 | 1 | 2023 | Domain-Specific Priors and Meta Learning for Few-Shot First-Person Action Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Computer vision › Video understanding and tracking › action recognition
few-shot action recognition |
0.7 | 1 | 2023 | Domain-Specific Priors and Meta Learning for Few-Shot First-Person Action Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot transfer |
0.7 | 1 | 2023 | Domain-Specific Priors and Meta Learning for Few-Shot First-Person Action Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Computer vision › Video understanding and tracking › activity recognition › procedural activity understanding
procedural video understanding |
0.7 | 1 | 2023 | STEPs: Self-Supervised Key Step Extraction and Localization from Unlabeled Procedural Videos · ICCV 2023 |
Computer vision › Video understanding and tracking
object tracking |
0.4 | 7 | 2010 | Vehicle detection and tracking in wide field-of-view aerial video · CVPR 2010 Geo-spatial aerial video processing for scene understanding and object tracking · CVPR 2008 Robust Object Matching for Persistent Tracking with Heterogeneous Features · IEEE Trans. Pattern Anal. Mach. Intell. 2007 |
Multimedia analysis and retrieval › event detection
complex event detection |
0.3 | 2 | 2013 | Semantic pooling for complex event detection · ACM Multimedia 2013 Evaluation of low-level features and their combinations for complex event detection in open source videos · CVPR 2012 |
Multimedia analysis and retrieval › event detection
video event detection |
0.3 | 2 | 2013 | Semantic pooling for complex event detection · ACM Multimedia 2013 Multimedia event recounting with concept based representation · ACM Multimedia 2012 |
Computer vision › Image recognition and object detection
pedestrian detection |
0.3 | 2 | 2014 | Pedestrian Detection in Low-Resolution Imagery by Learning Multi-scale Intrinsic Motion Structures (MIMS) · CVPR 2014 A real-time pedestrian detection system based on structure and appearance classification · ICRA 2010 |
Computer vision › 3D vision › human body modeling
contact map prediction |
0.3 | 1 | 2025 | How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions · CVPR 2025 |
Natural language and speech › Information extraction and text analysis › event extraction
event detection |
0.2 | 1 | 2016 | Zero-Shot Event Detection by Multimodal Distributional Semantic Embedding of Videos · AAAI 2016 |
Natural language and speech › Information extraction and text analysis › event extraction › event detection
zero-shot event detection |
0.2 | 1 | 2016 | Zero-Shot Event Detection by Multimodal Distributional Semantic Embedding of Videos · AAAI 2016 |
Virtual and augmented reality
augmented reality |
0.2 | 2 | 2023 | STEPs: Self-Supervised Key Step Extraction and Localization from Unlabeled Procedural Videos · ICCV 2023 Multi-View 3D Analysis with Applications for Augmented Reality and Enhanced Video Visualization · CVPR 2000 |
Computer vision › 3D vision
object matching |
0.2 | 3 | 2007 | Robust Object Matching for Persistent Tracking with Heterogeneous Features · IEEE Trans. Pattern Anal. Mach. Intell. 2007 PEET: Prototype Embedding and Embedding Transition for Matching Vehicles over Disparate Viewpoints · CVPR 2007 Unsupervised Learning of Discriminative Edge Measures for Vehicle Matching between Non-Overlapping Cameras · CVPR (1) 2005 |
Computer vision › Face, body and person analysis › person re-identification
vehicle re-identification |
0.2 | 3 | 2007 | Robust Object Matching for Persistent Tracking with Heterogeneous Features · IEEE Trans. Pattern Anal. Mach. Intell. 2007 PEET: Prototype Embedding and Embedding Transition for Matching Vehicles over Disparate Viewpoints · CVPR 2007 Vehicle Identification between Non-Overlapping Cameras without Direct Feature Matching · ICCV 2005 |
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning |
0.2 | 1 | 2023 | Domain-Specific Priors and Meta Learning for Few-Shot First-Person Action Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Computer vision › 3D vision
3d reconstruction |
0.2 | 6 | 2008 | Building segmentation for densely built urban regions using aerial LIDAR data · CVPR 2008 Dynamic Depth Recovery from Multiple Synchronized Video Streams · CVPR (1) 2001 Learning-Based Building Outline Detection from Multiple Aerial Images · CVPR (2) 2001 |
Computer vision › Video understanding and tracking › video analytics
aerial video analysis |
0.2 | 2 | 2010 | Vehicle detection and tracking in wide field-of-view aerial video · CVPR 2010 Geo-spatial aerial video processing for scene understanding and object tracking · CVPR 2008 |
Computer vision › Video understanding and tracking
motion detection |
0.2 | 1 | 2014 | Pedestrian Detection in Low-Resolution Imagery by Learning Multi-scale Intrinsic Motion Structures (MIMS) · CVPR 2014 |
Computer vision › Video understanding and tracking
multi-camera tracking |
0.2 | 2 | 2011 | Vehicle tracking across nonoverlapping cameras using joint kinematic and appearance features · CVPR 2011 Real-Time Wide Area Multi-Camera Stereo Tracking · CVPR (1) 2005 |
Computer vision › Image recognition and object detection
object detection |
0.2 | 3 | 2008 | Discovering class specific composite features through discriminative sampling with Swendsen-Wang Cut · CVPR 2008 Learning Exemplar-Based Categorization for the Detection of Multi-View Multi-Pose Objects · CVPR (2) 2006 Learning-Based Building Outline Detection from Multiple Aerial Images · CVPR (2) 2001 |
Computer vision › Video understanding and tracking
activity recognition |
0.2 | 1 | 2013 | Recognizing Activities via Bag of Words for Attribute Dynamics · CVPR 2013 |
Computer vision › Video understanding and tracking › temporal modeling
temporal structure modeling |
0.2 | 1 | 2013 | Recognizing Activities via Bag of Words for Attribute Dynamics · CVPR 2013 |
Robotics › Robot navigation and mapping
localization |
0.2 | 2 | 2008 | Real-time global localization with a pre-built visual landmark database · CVPR 2008 Ten-fold Improvement in Visual Odometry Using Landmark Matching · ICCV 2007 |
Methods — techniques the papers use, named apart from their topics
optical flow · 1.5depth · 1.3contrastive learning · 1.3clustering · 1.3large language model · 1.0iterative reasoning · 1.0code-writing · 1.0transformer decoder · 0.9codebook learning · 0.9VQ-VAE · 0.9gaze · 0.7low-level visual words · 0.2co-occurrence association · 0.2late fusion · 0.1early fusion · 0.1bag-of-words · 0.1quasi-rigid alignment · 0.1heterogeneous feature matching · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Schema-Guided Scene-Graph Reasoning Based on Multi-Agent Large Language Model SystemabstractScene graphs have emerged as a structured and serializable environment representation for grounded spatial reasoning with Large Language Models (LLMs). In this work, we propose SG2, an iterative Schema-Guided Scene-Graph reasoning framework based on multi-agent LLMs. The agents are grouped into two modules: a (1) Reasoner module for abstract task planning and graph information queries generation, and a (2) Retriever module for extracting corresponding graph information based on code-writing following the queries. Two modules collaborate iteratively, enabling sequential reasoning and adaptive attention to graph information. The scene graph schema, prompted to both modules, serves to not only streamline both reasoning and retrieval process, but also guide the cooperation between two modules. This eliminates the need to prompt LLMs with full graph data, reducing the chance of hallucination due to irrelevant information. Through experiments in multiple simulation environments, we show that our framework surpasses existing LLM-based approaches and baseline single-agent, tool-based Reason-while-Retrieve strategy in numerical Q&A and planning tasks. Yiye Chen, Harpreet Sawhney, Nicholas Gyde, Yanan Jian, Jack Saunders, Patricio A. Vela, Ben Lundell |
AAAI | 2 |
| 2025 | How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday InteractionsabstractWe tackle the novel problem of predicting 3D hand motion and contact maps (or Interaction Trajectories) given a single RGB view, action text, and a 3D contact point on the object as input. Our approach consists of (1) Interaction Codebook: a VQVAE model to learn a latent codebook of hand poses and contact points, effectively tokenizing interaction trajectories, (2) Interaction Predictor: a transformer-decoder module to predict the interaction trajectory from test time inputs by using an indexer module to retrieve a latent affordance from the learned codebook. To train our model, we develop a data engine that extracts 3D hand poses and contact trajectories from the diverse HoloAssist dataset. We evaluate our model on a benchmark that is 2.5-10× larger than existing works, in terms of diversity of objects and interactions observed, and test for generalization of the model across object categories, action categories, tasks, and scenes. Experimental results show the effectiveness of our approach over transformer & diffusion baselines across all settings. Ben Lundell, Dmitry Andreychuk, David Forsyth, Saurabh Gupta 0001, Harpreet Sawhney |
CVPR | 6 |
| 2023 | STEPs: Self-Supervised Key Step Extraction and Localization from Unlabeled Procedural VideosabstractWe address the problem of extracting key steps from un-labeled procedural videos, motivated by the potential of Augmented Reality (AR) headsets to revolutionize job training and performance. We decompose the problem into two steps: representation learning and key steps extraction. We propose a training objective, Bootstrapped Multi-Cue Contrastive (BMC2) loss to learn discriminative representations for various steps without any labels. Different from prior works, we develop techniques to train a light-weight temporal module which uses off-the-shelf features for self supervision. Our approach can seamlessly leverage information from multiple cues like optical flow, depth or gaze to learn discriminative features for key-steps, making it amenable for AR applications. We finally extract key steps via a tunable algorithm that clusters the representations and samples. We show significant improvements over prior works for the task of key step localization and phase classification. Qualitative results demonstrate that the extracted key steps are meaningful and succinctly represent various steps of the procedural tasks. Our code can be found at https://github.com/anshulbshah/STEPs. Anshul Shah 0001, Ben Lundell, Harpreet Sawhney, Rama Chellappa |
ICCV | 3 |
| 2023 | Self-supervised Learning with Local Contrastive Loss for Detection and Semantic SegmentationabstractWe present a self-supervised learning (SSL) method suitable for semi-global tasks such as object detection and semantic segmentation. We enforce local consistency between self-learned features that represent corresponding image locations of transformed versions of the same image, by minimizing a pixel-level local contrastive (LC) loss during training. LC-loss can be added to existing self-supervised learning methods with minimal overhead. We evaluate our SSL approach on two downstream tasks – object detection and semantic segmentation, using COCO, PASCAL VOC, and CityScapes datasets. Our method outperforms the existing state-of-the-art SSL approaches by 1.9% on COCO object detection, 1.4% on PASCAL VOC detection, and 0.6% on CityScapes segmentation. Ashraful Islam, Ben Lundell, Harpreet Sawhney, Sudipta Sinha, Peter Morales, Richard J. Radke |
WACV | 3 |
| 2023 | Domain-Specific Priors and Meta Learning for Few-Shot First-Person Action RecognitionabstractThe lack of large-scale real datasets with annotations makes transfer learning a necessity for video activity understanding. We aim to develop an effective method for few-shot transfer learning for first-person action classification. We leverage independently trained local visual cues to learn representations that can be transferred from a source domain, which provides primitive action labels, to a different target domain using only a handful of examples. Visual cues we employ include object-object interactions, hand grasps and motion within regions that are a function of hand locations. We employ a framework based on meta-learning to extract the distinctive and domain invariant components of the deployed visual cues. This enables transfer of action classification models across public datasets captured with diverse scene and action configurations. We present comparative results of our transfer learning methodology and report superior results over state-of-the-art action classification approaches for both inter-class and inter-dataset transfer. Huseyin Coskun, M. Zeeshan Zia, Bugra Tekin, Federica Bogo, Nassir Navab, Federico Tombari, Harpreet Sawhney |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2016 | Zero-Shot Event Detection by Multimodal Distributional Semantic Embedding of VideosabstractWe propose a new zero-shot Event-Detection method by Multi-modal Distributional Semantic embedding of videos. Our model embeds object and action concepts as well as other available modalities from videos into a distributional semantic space. To our knowledge, this is the first Zero-Shot event detection model that is built on top of distributional semantics and extends it in the following directions: (a) semantic embedding of multimodal information in videos (with focus on the visual modalities), (b) semantic embedding of concepts definitions, and (c) retrieve videos by free text event query (e.g., "changing a vehicle tire") based on their content. We first embed the video into the multi-modal semantic space and then measure the similarity between videos with the event query in free text form. We validated our method on the large TRECVID MED (Multimedia Event Detection) challenge. Using only the event title as a query, our method outperformed the state-the-art that uses big descriptions from 12.6\% to 13.5\% with MAP metric and from 0.73 to 0.83 with ROC-AUC metric. It is also an order of magnitude faster. Jingen Liu, Harpreet Sawhney, Ahmed M. Elgammal |
AAAI | 4 |
| 2015 | De-correlating CNN Features for Generative ClassificationabstractThe problem of training a classifier from a handful of positive examples, without having to supply class specific negatives is of great practical importance. The proposed approach to solving this problem builds on the idea of training LDA classifiers using only class specific foreground images and a large collection of unlabelled images, as described in [11]. While we adopt the LDA training methodology of [11], we depart from HOG features and work with those extracted from a Convolutional Neural Network (CNN) pre-trained on Image Net (Over feat). We combine Over feat features with the LDA training methodology to derive generative classifiers. When evaluated on a K-way classification problem, these classifiers are almost as good as those trained discriminatively using the same features. Unlike the HOG based approach of [11], our classifiers do not need any post-processing step of calibration, a step that requires positives and negatives. Finally, we show that in an instance retrieval setup, we can employ these generative classifiers to derive a novel query-expansion framework that achieves a significant performance boost by utilizing only the top ranked positive examples from an initial nearest-neighbor list. Chaitanya Desai, Jayan Eledath, Harpreet Sawhney, Mayank Bansal |
WACV | 3 |
| 2014 | Depth Extraction from Videos Using Geometric Context and Occlusion Boundaries
Syed Raza, Omar Javed, Aveek Das, Harpreet Sawhney, Irfan A. Essa |
BMVC | 4 |
| 2014 | Pedestrian Detection in Low-Resolution Imagery by Learning Multi-scale Intrinsic Motion Structures (MIMS)abstractDetecting pedestrians at a distance from large-format wide-area imagery is a challenging problem because of low ground sampling distance (GSD) and low frame rate of the imagery. In such a scenario, the approaches based on appearance cues alone mostly fail because pedestrians are only a few pixels in size. Frame-differencing and optical flow based approaches also give poor detection results due to noise, camera jitter and parallax in aerial videos. To overcome these challenges, we propose a novel approach to extract Multi-scale Intrinsic Motion Structure features from pedestrian's motion patterns for pedestrian detection. The MIMS feature encodes the intrinsic motion properties of an object, which are location, velocity and trajectory-shape invariant. The extracted MIMS representation is robust to noisy flow estimates. In this paper, we give a comparative evaluation of the proposed method and demonstrate that MIMS outperforms the state of the art approaches in identifying pedestrians from low resolution airborne videos. Jiejie Zhu, Omar Javed, Jingen Liu, Harpreet Sawhney |
CVPR | 6 |
| 2014 | Multimodal fusion using dynamic hybrid modelsabstractWe propose a novel hybrid model that exploits the strength of discriminative classifiers along with the representational power of generative models. Our focus is on detecting multimodal events in time varying sequences. Discriminative classifiers have been shown to achieve higher performances than the corresponding generative likelihood-based classifiers. On the other hand, generative models learn a rich informative space which allows for data generation and joint feature representation that discriminative models lack. We employ a deep temporal generative model for unsupervised learning of a shared representation across multiple modalities with time varying data. The temporal generative model takes into account short term temporal phenomena and allows for filling in missing data by generating data within or across modalities. The hybrid model involves augmenting the temporal generative model with a temporal discriminative model for event detection, and classification, which enables modeling long range temporal dynamics. We evaluate our approach on audio-visual datasets (AVEC, AVLetters, and CUAVE) and demonstrate its superiority compared to the state-of-the-art. Mohamed R. Amer, Behjat Siddiquie, Saad M. Khan, Ajay Divakaran, Harpreet Sawhney |
WACV | 5 |
| 2014 | Automatic 3D change detection for glaucoma diagnosisabstractImportant diagnostic criteria for glaucoma are changes in the 3D structure of the optic disc due to optic nerve damage. We propose an automatic approach for detecting these changes in 3D models reconstructed from fundus images of the same patient taken at different times. For each time session, only two uncalibated fundus images are required. The approach applies a 6-point algorithm to estimate relative camera pose assuming a constant camera focal length. To deal with the instability of 3D reconstruction associated with fundus images, our approach keeps multiple candidate reconstruction solutions for each image pair. The best 3D reconstruction is found by optimizing the 3D registration of all images after an iterative bundle adjustment that tolerates possible structure changes. The 3D structure changes are detected by evaluating the reprojection errors of feature points in image space. We validate the approach by comparing the diagnosis results with manual grading by human experts on a fundus image dataset. Vinutha Kallem, Mayank Bansal, Jayan Eledath, Harpreet Sawhney, Denise J. Pearson, Richard A. Stone |
WACV | 5 |
| 2013 | Recognizing Activities via Bag of Words for Attribute DynamicsabstractIn this work, we propose a novel video representation for activity recognition that models video dynamics with attributes of activities. A video sequence is decomposed into short-term segments, which are characterized by the dynamics of their attributes. These segments are modeled by a dictionary of attribute dynamics templates, which are implemented by a recently introduced generative model, the binary dynamic system~(BDS). We propose methods for learning a dictionary of BDSs from a training corpus, and for quantizing attribute sequences extracted from videos into these BDS code words. This procedure produces a representation of the video as a histogram of BDS code words, which is denoted the bag-of-words for attribute dynamics (BoWAD). An extensive experimental evaluation reveals that this representation outperforms other state-of-the-art approaches in temporal structure modeling for complex activity recognition. Weixin Li 0002, Harpreet Sawhney, Nuno Vasconcelos |
CVPR | 3 |
| 2013 | Affect analysis in natural human interaction using Joint Hidden Conditional Random FieldsabstractWe present a novel approach for multi-modal affect analysis in human interactions that is capable of integrating data from multiple modalities while also taking into account temporal dynamics. Our fusion approach, Joint Hidden Conditional Random Fields (JHRCFs), combines the advantages of purely feature level (early fusion) fusion approaches with late fusion (CRFs on individual modalities) to simultaneously learn the correlations between features from multiple modalities as well as their temporal dynamics. Our approach addresses major shortcomings of other fusion approaches such as the domination of other modalities by a single modality with early fusion and the loss of cross-modal information with late fusion. Extensive results on the AVEC 2011 dataset show that we outperform the state-of-the-art on the Audio Sub-Challenge, while achieving competitive performance on the Video Sub-Challenge and the Audiovisual Sub-Challenge. Behjat Siddiquie, Saad M. Khan, Ajay Divakaran, Harpreet Sawhney |
ICME | 4 |
| 2013 | Interactive Retinal Vessel Extraction by Integrating Vessel Tracing and Graph Search
Vinutha Kallem, Mayank Bansal, Jayan Eledath, Harpreet Sawhney, Karen Karp, Denise J. Pearson, Monte D. Mills, Graham E. Quinn, Richard A. Stone |
MICCAI (2) | 5 |
| 2013 | Semantic pooling for complex event detectionabstractComplex event detection is very challenging in open source such as You-Tube videos, which usually comprise very diverse visual contents involving various object, scene and action concepts. Not all of them, however, are relevant to the event. In other words, a video may contain a lot of "junk" information which is harmful for recognition. Hence, we propose a semantic pooling approach to tackle this issue. Unlike the conventional pooling over the entire video or specific spatial regions of a video, we employ a discriminative approach to acquire abstract semantic "regions" for pooling. For this purpose, we first associate low-level visual words with semantic concepts via their co-occurrence relationship. We then pool the low-level features separately according to their semantic information. The proposed semantic pooling strategy also provides a new mechanism for incorporating semantic concepts for low-level feature based event recognition. We evaluate our approach on TRECVID MED [1] dataset and the results show that semantic pooling consistently improves the performance compared with conventional pooling strategies. Jingen Liu, Ajay Divakaran, Harpreet Sawhney |
ACM Multimedia | 5 |
| 2013 | Video event recognition using concept attributesabstractWe propose to use action, scene and object concepts as semantic attributes for classification of video events in InTheWild content, such as YouTube videos. We model events using a variety of complementary semantic attribute features developed in a semantic concept space. Our contribution is to systematically demonstrate the advantages of this concept-based event representation (CBER) in applications of video event classification and understanding. Specifically, CBER has better generalization capability, which enables to recognize events with a few training examples. In addition, CBER makes it possible to recognize a novel event without training examples (i.e., zero-shot learning). We further show our proposed enhanced event model can further improve the zero-shot learning. Furthermore, CBER provides a straightforward way for event recounting/understanding. We use the TRECVID Multimedia Event Detection (MED11) open source event definitions and datasets as our test bed and show results on over 1400 hours of videos. Jingen Liu, Omar Javed, Saad Ali, Amir Tamrakar, Ajay Divakaran, Harpreet Sawhney |
WACV | 8 |
| 2013 | Image to LIDAR matching for geotagging in urban environmentsabstractWe present a novel method for matching ground-based query images to a georeferenced LIDAR 3D dataset acquired from an airborne platform in urban environments. We are addressing two main technical challenges: (i) different modalities between the query and the reference data (electro-optical vs. LIDAR) that impose unique challenges to the matching problem; (ii) very different viewing directions from which the query, respectively the LIDAR data were acquired. We make two main technical contributions in this paper. First, we present a method for automatically extracting features from LIDAR data that largely remain invariant to the projection in a 2D image and thus allow robust matching across modalities and change in viewpoint. Second, we describe a matching technique that finds the best 3D pose that relates the query input image to a rendered image of the 3D models. We present results of matching images to high-resolution LIDAR data covering five square kilometers over a city that demonstrate the power of the matching method proposed. Bogdan Matei, Nick Vander Valk, Harpreet Sawhney |
WACV | 5 |
| 2012 | Evaluation of low-level features and their combinations for complex event detection in open source videosabstractLow-level appearance as well as spatio-temporal features, appropriately quantized and aggregated into Bag-of-Words (BoW) descriptors, have been shown to be effective in many detection and recognition tasks. However, their effcacy for complex event recognition in unconstrained videos have not been systematically evaluated. In this paper, we use the NIST TRECVID Multimedia Event Detection (MED11 [1]) open source dataset, containing annotated data for 15 high-level events, as the standardized test bed for evaluating the low-level features. This dataset contains a large number of user-generated video clips. We consider 7 different low-level features, both static and dynamic, using BoW descriptors within an SVM approach for event detection. We present performance results on the 15 MED11 events for each of the features as well as their combinations using a number of early and late fusion strategies and discuss their strengths and limitations. Amir Tamrakar, Saad Ali, Jingen Liu, Omar Javed, Ajay Divakaran, Harpreet Sawhney |
CVPR | 8 |
| 2012 | Multimedia event recounting with concept based representationabstractMultimedia event detection has drawn a lot of attention in recent years. Given a recognized event, in this paper, we conduct a pilot study of the multimedia event recounting problem, which answers the question why this video is recognized as this event, i.e. what evidences this decision is made on. In order to provide a semantic recounting of the multimedia event, we adopt a concept-based event representation for learning a discriminative event model. Then, we present a recounting approach that exactly recovers the contribution of semantic evidence to the event classification decision. This approach can be applied on any additive discriminative classifiers. The promising result is shown on the MED11 dataset that contains 15 events in thousands of YouTube like videos. Jingen Liu, Ajay Divakaran, Harpreet Sawhney |
ACM Multimedia | 5 |
| 2011 | Vehicle tracking across nonoverlapping cameras using joint kinematic and appearance featuresabstractWe describe a vehicle tracking algorithm using input from a network of nonoverlapping cameras. Our algorithm is based on a novel statistical formulation that uses joint kinematic and image appearance information to link local tracks of the same vehicles into global tracks with longer persistence. The algorithm can handle significant spatial separation between the cameras and is robust to challenging tracking conditions such as high traffic density, or complex road infrastructure. In these cases, traditional tracking formulations based on MHT, or JPDA algorithms, may fail to produce track associations across cameras due to the weak predictive models employed. We make several new contributions in this paper. Firstly, we model kinematic constraints between any two local tracks using road networks and transit time distributions. The transit time distributions are calculated dynamically as convolutions of normalized transit time distributions that are learned and adapted separately for individual roads. Secondly, we present a complete statistical tracker formulation, which combines kinematic and appearance likelihoods within a multi-hypothesis framework. We have extensively evaluated the algorithm proposed using a network of ground-based cameras with narrow field of view. The tracking results obtained on a large ground-truthed dataset demonstrate the effectiveness of the algorithm proposed. Bogdan Matei, Harpreet Sawhney, Supun Samarasekera |
CVPR | 2 |
| 2011 | A LIDAR streaming architecture for mobile robotics with application to 3D structure characterizationabstractWe present a novel LIDAR streaming architecture for real-time, on-board processing using unmanned robots. We propose a two-level 3D data structure that allows pipelined and streaming processing of the 3D data as it arrives from a moving robot: (i) at the coarse level, the incoming 3D scans are stored in memory in a dense 3D voxel grid with a relatively large voxel size - this ensures buffering of the most recent data and the availability of sufficient 3D measurements within a specific processing volume at the next level; (ii) at the fine level, we employ a data chunking mechanism guided by the movement of the robot and a rolling dense 3D voxel grid for processing the data in the immediate vicinity of the robot, which enables reuse of previously computed features. The architecture proposed requires a very small memory footprint, minimal data copying, and allows a fast spatial access for 3D data, even at the finest resolutions. We illustrate the proposed streaming architecture on a real-time 3D structure characterization task for detecting doors and stairs in indoor environments and show qualitative results demonstrating the effectiveness of our approach. Mayank Bansal, Bogdan Matei, Ben Southall, Jayan Eledath, Harpreet Sawhney |
ICRA | 5 |
| 2011 | Geo-localization of street views with aerial image databasesabstractWe study the feasibility of solving the challenging problem of geolocalizing ground level images in urban areas with respect to a database of images captured from the air such as satellite and oblique aerial images. We observe that comprehensive aerial image databases are widely available while complete coverage of urban areas from the ground is at best spotty. As a result, localization of ground level imagery with respect to aerial collections is a technically important and practically significant problem. We exploit two key insights: (1) satellite image to oblique aerial image correspondences are used to extract building facades, and (2) building facades are matched between oblique aerial and ground images for geo-localization. Key contributions include: (1) A novel method for extracting building facades using building outlines; (2) Correspondence of building facades between oblique aerial and ground images without direct matching; and (3) Position and orientation estimation of ground images. We show results of ground image localization in a dense urban area. Mayank Bansal, Harpreet Sawhney, Kostas Daniilidis |
ACM Multimedia | 2 |
| 2010 | 3D model based vehicle classification in aerial imageryabstractWe present an approach that uses detailed 3D models to detect and classify objects into fine levels of vehicle categories. Unlike other approaches that use silhouette information to fit a 3D model, our approach uses complete appearance from the image. Each 3D model has a set of salient location markers that are determined a-priori. These salient locations represent a sub-sampling of 3D locations that make up the model. Scene conditions are simulated in the rendering of 3D models and the salient locations are used to bootstrap a HoG based feature classifier. HoG features are computed in both rendered and real scenes and a novel object match score the `Salient Feature Match Distribution Matrix' is computed. For each 3D model we also learn the patterns of misalignment with other vehicle types and use it as an additional cue for classification. Results are presented on a challenging aerial video dataset consisting of vehicle imagery from various viewpoints and environmental conditions. Saad M. Khan, Dennis Matthies, Harpreet Sawhney |
CVPR | 4 |
| 2010 | Vehicle detection and tracking in wide field-of-view aerial videoabstractThis paper presents a joint probabilistic relation graph approach to simultaneously detect and track a large number of vehicles in low frame rate aerial videos. Due to low frame rate, low spatial resolution and sheer number of moving objects, detection and tracking in wide area video poses unique challenges. In this paper, we explore vehicle behavior model from road structure and generate a set of constraints to regulate both object based vertex matching and pairwise edge matching schemes. The proposed relation graph approach then unifies these two matching schemes into a single cost minimization framework to produce a quadratic optimized association result. The experiments on hours of real videos demonstrate the graph matching framework with vehicle behavior model effectively improves tracking performance in large scale dense traffic scenarios. Jiangjian Xiao, Harpreet Sawhney, Feng Han 0002 |
CVPR | 3 |
| 2010 | A real-time pedestrian detection system based on structure and appearance classificationabstractWe present a real-time pedestrian detection system based on structure and appearance classification. We discuss several novel ideas that contribute to having low-false alarms and high detection rates, while at the same time achieving computational efficiency: (i) At the front end of our system we employ stereo to detect pedestrians in 3D range maps using template matching with a representative 3D shape model, and to classify other background objects in the scene such as buildings, trees and poles. The structure classification efficiently labels substantial amount of non-relevant image regions and guides the further computationally expensive process to focus on relatively small image parts; (ii)We improve the appearance-based classifiers based on HoG descriptors by performing template matching with 2D human shape contour fragments that results in improved localization and accuracy; (iii) We build a suite of classifiers tuned to specific distance ranges for optimized system performance. Our method is evaluated on publicly available datasets and is shown to match or exceed the performance of leading pedestrian detectors in terms of accuracy as well as achieving real-time computation (10 Hz), which makes it adequate for in-vehicle navigation platform. Mayank Bansal, Sang-Hack Jung, Bogdan Matei, Jayan Eledath, Harpreet Sawhney |
ICRA | 5 |
| 2009 | Recognition and volume estimation of food intake using a mobile deviceabstractWe present a system that improves accuracy of food intake assessment using computer vision techniques. Traditional dietetic method suffers from the drawback of either inaccurate assessment or complex lab measurement. Our solution is to use a mobile phone to capture images of foods, recognize food types, estimate their respective volumes and finally return quantitative nutrition information. Automated and accurate food recognition presents the following challenges. First, there exist a large variety of food types that people consume in everyday life. Second, a single category of food may contain large variations due to different ways of preparation. Also, diverse lighting conditions may lead to varying visual appearance of foods. All of these pose a challenge to the state of the art recognition approaches. Moreover, the low quality images captured using cellphones make the task of 3D reconstruction difficult. In this paper, we combine several vision techniques (visual recognition and 3D reconstruction) to achieve quantitative food intake estimation. Evaluation of both recognition and reconstruction is provided in the experimental results. Manika Puri, Ajay Divakaran, Harpreet Sawhney |
WACV | 5 |
| 2009 | Geo-Based Aerial Surveillance Video Processing for Scene Understanding and Object TrackingabstractThis paper presents an approach to extract semantic layers from aerial surveillance videos for scene understanding and object tracking. The input videos are captured by low flying aerial platforms and typically consist of strong parallax from non-ground-plane structures as well as moving objects. Our approach leverages the geo-registration between video frames and reference images (such as those available from Terraserver and Google satellite imagery) to establish a unique geo-spatial coordinate system for pixels in the video. The geo-registration process enables Euclidean 3D reconstruction with absolute scale unlike traditional monocular structure from motion where continuous scale estimation over long periods of time is an issue. Geo-registration also enables correlation of video data to other stored information sources such as GIS (Geo-spatial Information System) databases. In addition to the geo-registration and 3D reconstruction aspects, the other key contributions of this paper also include: (1) providing a reliable geo-based solution to estimate camera pose for 3D reconstruction, (2) exploiting appearance and 3D shape constraints derived from geo-registered videos for labeling of structures such as buildings, foliage, and roads for scene understanding, and (3) elimination of moving object detection and tracking errors using 3D parallax constraints and semantic labels derived from geo-registered videos. Experimental results on extended time aerial video data demonstrates the qualitative and quantitative aspects of our work. Jiangjian Xiao, Feng Han 0002, Harpreet Sawhney |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2008 | Matching vehicles under large pose transformations using approximate 3D models and piecewise MRF modelabstractWe propose a robust object recognition method based on approximate 3D models that can effectively match objects under large viewpoint changes and partial occlusion. The specific problem we solve is: given two views of an object, determine if the views are for the same or different object. Our domain of interest is vehicles, but the approach can be generalized to other man-made rigid objects. A key contribution of our approach is the use of approximate models with locally and globally constrained rendering to determine matching objects. We utilize a compact set of 3D models to provide geometry constraints and transfer appearance features for object matching across disparate viewpoints. The closest model from the set, together with its poses with respect to the data, is used to render an object both at pixel (local) level and region/part (global) level. Especially, symmetry and semantic part ownership are used to extrapolate appearance information. A piecewise Markov Random Field (MRF) model is employed to combine observations obtained from local pixel and global region level. Belief Propagation (BP) with reduced memory requirement is employed to solve the MRF model effectively. No training is required, and a realistic object image in a disparate viewpoint can be obtained from as few as just one image. Experimental results on vehicle data from multiple sensor platforms demonstrate the efficacy of our method. Yanlin Guo, Cen Rao, Supun Samarasekera, Janet Kim, Rakesh Kumar 0001, Harpreet Sawhney |
CVPR | 6 |
| 2008 | Discovering class specific composite features through discriminative sampling with Swendsen-Wang CutabstractThis paper proposes a novel approach to discover a set of class specific ldquocomposite featuresrdquo as the feature pool for the detection and classification of complex objects using AdaBoost. Each composite feature is constructed from the combination of multiple individual features. Unlike previous works that design features manually or with certain restrictions, the class specific features are selected from the space of all combinations of a set of individual features. To achieve this, we first establish an analogue between the problem of discriminative feature selection and generative image segmentation, and then draw discriminative samples from the combinatory space with a novel algorithm called discriminative generalized Swendsen-Wang cut. These samples form the initial pool of features, where AdaBoost is applied to learn a strong classifier combining the most discriminative composite features. We demonstrate the efficacy of our approach by comparing with existing detection algorithms for finding people in general pose. Feng Han 0002, Ying Shan, Harpreet Sawhney, Rakesh Kumar 0001 |
CVPR | 3 |
| 2008 | Building segmentation for densely built urban regions using aerial LIDAR dataabstractWe present a novel building segmentation system for densely built areas, containing thousands of buildings per square kilometer. We employ solely sparse LIDAR (Light/Laser Detection Ranging) 3D data, captured from an aerial platform, with resolution less than one point per square meter. The goal of our work is to create segmented and delineated buildings as well as structures on top of buildings without requiring scanning for the sides of buildings. Building segmentation is a critical component in many applications such as 3D visualization, robot navigation and cartography. LIDAR has emerged in recent years as a more robust alternative to 2D imagery because it acquires 3D structure directly, without the shortcomings of stereo in un- textured regions and at depth discontinuities. Our main technical contributions in this paper are: (i) a ground segmentation algorithm which can handle both rural regions, and heavily urbanized areas, where the ground is 20% or less of the data, (ii) a building segmentation technique, which is robust to buildings in close proximity to each other, sparse measurements and nearby structured vegetation clutter, and (Hi) an algorithm for estimating the orientation of a boundary contour of a building, based on minimizing the number of vertices in a rectilinear approximation to the building outline, which can cope with significant quantization noise in the outline measurements. We have applied the proposed building segmentation system to several urban regions with areas of hundreds of square kilometers each, obtaining average segmentation speeds of less than three minutes per km2on a standard Pentium processor. Extensive qualitative results obtained by overlaying the 3D segmented regions onto 2D imagery indicate accurate performance of our system. Bogdan Matei, Harpreet Sawhney, Supun Samarasekera, Janet Kim, Rakesh Kumar 0001 |
CVPR | 2 |
| 2008 | Geo-spatial aerial video processing for scene understanding and object trackingabstractThis paper presents an approach to extracting and using semantic layers from low altitude aerial videos for scene understanding and object tracking. The input video is captured by low flying aerial platforms and typically consists of strong parallax from non-ground-plane structures. A key aspect of our approach is the use of geo-registration of video frames to reference image databases (such as those available from Terraserver and Google satellite imagery) to establish a geo-spatial coordinate system for pixels in the video. Geo-registration enables Euclidean 3D reconstruction with absolute scale unlike traditional monocular structure from motion where continuous scale estimation over long periods of time is an issue. Geo-registration also enables correlation of video data to other stored information sources such as GIS (geo-spatial information system) databases. In addition to the geo-registration and 3D reconstruction aspects, the key contributions of this paper include: (1) exploiting appearance and 3D shape constraints derived from geo-registered videos for labeling of structures such as buildings, foliage, and roads for scene understanding, and (2) elimination of moving object detection and tracking errors using 3D parallax constraints and semantic labels derived from geo-registered videos. Experimental results on extended time aerial video data demonstrates the qualitative and quantitative aspects of our work. Jiangjian Xiao, Feng Han 0002, Harpreet Sawhney |
CVPR | 4 |
| 2008 | Real-time global localization with a pre-built visual landmark databaseabstractIn this paper, we study how to build a vision-based system for global localization with accuracies within 10cm. for robots and humans operating both indoors and outdoors over wide areas covering many square kilometers. In particular, we study the parameters of building a landmark database rapidly and utilizing that database online for real-time accurate global localization. Although the accuracy of traditional short-term motion based visual odometry systems has improved significantly in recent years, these systems alone cannot solve the drift problem over large areas. Landmark based localization combined with visual odometry is a viable solution to the large scale localization problem. However, a systematic study of the specification and use of such a landmark database has not been undertaken. We propose techniques to build and optimize a landmark database systematically and efficiently using visual odometry. First, topology inference is utilized to find overlapping images in the database. Second, bundle adjustment is used to refine the accuracy of each 3D landmark. Finally, the database is optimized to balance the size of the database with achievable accuracy. Once the landmark database is obtained, a new real-time global localization methodology that works both indoors and outdoors is proposed. We present results of our study on both synthetic and real datasets that help us determine critical design parameters for the landmark database and the achievable accuracies of our proposed system. Taragay Oskiper, Supun Samarasekera, Rakesh Kumar 0001, Harpreet Sawhney |
CVPR | 5 |
| 2008 | Special issue on video surveillance research in industry and academia
Harpreet Sawhney |
Mach. Vis. Appl. | 2 |
| 2008 | Toward a sentient environment: real-time wide area multiple human tracking with identities
Manoj Aggarwal, Thomas Germano, Ian Roth, Alexandar Knowles, Rakesh Kumar 0001, Harpreet Sawhney, Supun Samarasekera |
Mach. Vis. Appl. | 7 |
| 2008 | Unsupervised Learning of Discriminative Edge Measures for Vehicle Matching between Nonoverlapping CamerasabstractThis paper proposes a novel unsupervised algorithm learning discriminative features in the context of matching road vehicles between two non-overlapping cameras. The matching problem is formulated as a same-different classification problem, which aims to compute the probability of vehicle images from two distinct cameras being from the same vehicle or different vehicle(s). We employ a novel measurement vector that consists of three independent edge-based measures and their associated robust measures computed from a pair of aligned vehicle edge maps. The weight of each measure is determined by an unsupervised learning algorithm that optimally separates the same-different classes in the combined measurement space. This is achieved with a weak classification algorithm that automatically collects representative samples from same-different classes, followed by a more discriminative classifier based on Fisher' s Linear Discriminants and Gibbs Sampling. The robustness of the match measures and the use of unsupervised discriminant analysis in the classification ensures that the proposed method performs consistently in the presence of missing/false features, temporally and spatially changing illumination conditions, and systematic misalignment caused by different camera configurations. Extensive experiments based on real data of over 200 vehicles at different times of day demonstrate promising results. Ying Shan, Harpreet Sawhney, Rakesh Kumar 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Content-Based Matching of Videos Using Local Spatio-temporal Fingerprints
Gajinder Singh, Manika Puri, Jeffrey Lubin, Harpreet Sawhney |
ACCV (2) | 4 |
| 2007 | PEET: Prototype Embedding and Embedding Transition for Matching Vehicles over Disparate ViewpointsabstractThis paper presents a novel framework, prototype embedding and embedding transition (PEET), for matching objects, especially vehicles, that undergo drastic pose, appearance, and even modality changes. The problem of matching objects seen under drastic variations is reduced to matching embeddings of object appearances instead of matching the object images directly. An object appearance is first embedded in the space of a representative set of model prototypes (prototype embedding (PE)). Objects captured at disparate temporal and spatial sites are embedded in the space of prototypes that are rendered with the pose of the cameras at the respective sites. Low dimensional embedding vectors are subsequently matched. A significant feature of our approach is that no mapping function is needed to compute the distance between embedding vectors extracted from objects viewed from disparate pose and appearance changes, instead, an embedding transition (ET) scheme is utilized to implicitly realize the complex and non-linear mapping with high accuracy. The heterogeneous nature of matching between high-resolution and low-resolution image objects in PEET is discussed, and an unsupervised learning scheme based on the exploitation of the heterogeneous nature is developed to improve the overall matching performance of mixed resolution objects. The proposed approach has been applied to vehicular object classification and query application, and the extensive experimental results demonstrate the efficacy and versatility of the PEET framework. Yanlin Guo, Ying Shan, Harpreet Sawhney, Rakesh Kumar 0001 |
CVPR | 3 |
| 2007 | Ten-fold Improvement in Visual Odometry Using Landmark MatchingabstractOur goal is to create a visual odometry system for robots and wearable systems such that localization accuracies of centimeters can be obtained for hundreds of meters of distance traveled. Existing systems have achieved approximately a 1% to 5% localization error rate whereas our proposed system achieves close to 0.1% error rate, a ten-fold reduction. Traditional visual odometry systems drift over time as the frame-to-frame errors accumulate. In this paper, we propose to improve visual odometry using visual landmarks in the scene. First, a dynamic local landmark tracking technique is proposed to track a set of local landmarks across image frames and select an optimal set of tracked local landmarks for pose computation. As a result, the error associated with each pose computation is minimized to reduce the drift significantly. Second, a global landmark based drift correction technique is proposed to recognize previously visited locations and use them to correct drift accumulated during motion. At each visited location along the route, a set of distinctive visual landmarks is automatically extracted and inserted into a landmark database dynamically. We integrate the landmark based approach into a navigation system with 2 stereo pairs and a low-cost inertial measurement unit (IMU) for increased robustness. We demonstrate that a real-time visual odometry system using local and global landmarks can precisely locate a user within 1 meter over 1000 meters in unknown indoor/outdoor environments with challenging situations such as climbing stairs, opening doors, moving foreground objects etc.. Taragay Oskiper, Supun Samarasekera, Rakesh Kumar 0001, Harpreet Sawhney |
ICCV | 5 |
| 2007 | Robust Object Matching for Persistent Tracking with Heterogeneous FeaturesabstractThis paper addresses the problem of matching vehicles across multiple sightings under variations in illumination and camera poses. Since multiple observations of a vehicle are separated in large temporal and/or spatial gaps, thus prohibiting the use of standard frame-to-frame data association, we employ features extracted over a sequence during one time interval as a vehicle fingerprint that is used to compute the likelihood that two or more sequence observations are from the same or different vehicles. Furthermore, since our domain is aerial video tracking, in order to deal with poor image quality and large resolution and quality variations, our approach employs robust alignment and match measures for different stages of vehicle matching. Most notably, we employ a heterogeneous collection of features such as lines, points, and regions in an integrated matching framework. Heterogeneous features are shown to be important. Line and point features provide accurate localization and are employed for robust alignment across disparate views. The challenges of change in pose, aspect, and appearances across two disparate observations are handled by combining a novel feature-based quasi-rigid alignment with flexible matching between two or more sequences. However, since lines and points are relatively sparse, they are not adequate to delineate the object and provide a comprehensive matching set that covers the complete object. Region features provide a high degree of coverage and are employed for continuous frames to provide a delineation of the vehicle region for subsequent generation of a match measure. Our approach reliably delineates objects by representing regions as robust blob features and matching multiple regions to multiple regions using Earth Mover's Distance (EMD). Extensive experimentation under a variety of real-world scenarios and over hundreds of thousands of Confirmatory Identification (CID) trails has demonstrated about 95 percent accuracy in vehicle reacquisition with both visible and Infrared (IR) imaging cameras. Yanlin Guo, Steven C. Hsu, Harpreet Sawhney, Rakesh Kumar 0001, Ying Shan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2006 | Learning Exemplar-Based Categorization for the Detection of Multi-View Multi-Pose ObjectsabstractThis paper proposes a novel approach for multi-view multi-pose object detection using discriminative shapebased exemplars. The key idea underlying this method is motivated by numerous previous observations that manually clustering multi-view multi-pose training data into different categories and then combining the separately trained two-class classifiers greatly improved the detection performance. A novel computational framework is proposed to unify different processes of categorization, training individual classifier for each intra-class category, and training a strong classifier combining the individual classifiers. The individual processes employ a single objective function that is optimized using two nested AdaBoost loops. The outer AdaBoost loop is used to select discriminative exemplars and the inner AdaBoost is used to select discriminative features on the selected exemplars. The proposed approach replaces the manual time-consuming process of exemplar selection as well as addresses the problem of labeling ambiguity inherent in this process. Also, our approach fully complies with the standard AdaBoost-based object detection framework in terms of real-time implementation. Experiments on multi-view multi-pose people and vehicle data demonstrate the efficacy of the proposed approach. Ying Shan, Feng Han 0002, Harpreet Sawhney, Rakesh Kumar 0001 |
CVPR (2) | 3 |
| 2006 | Identification of Highly Similar 3D Objects Using Model Saliency
Bogdan Matei, Harpreet Sawhney, Clay Spence |
ECCV (4) | 2 |
| 2006 | Exploiting Model Similarity for Indexing and Matching to a Large Model Database
Bogdan Matei, Harpreet Sawhney |
ECCV (2) | 3 |
| 2006 | Bilateral Filtering-Based Optical Flow Estimation with Occlusion Detection
Jiangjian Xiao, Harpreet Sawhney, Cen Rao, Michael A. Isnardi |
ECCV (1) | 3 |
| 2006 | Rapid Object Indexing Using Locality Sensitive Hashing and Joint 3D-Signature Space EstimationabstractWe propose a new method for rapid 3D object indexing that combines feature-based methods with coarse alignment-based matching techniques. Our approach achieves a sublinear complexity on the number of models, maintaining at the same time a high degree of performance for real 3D sensed data that is acquired in largely uncontrolled settings. The key component of our method is to first index surface descriptors computed at salient locations from the scene into the whole model database using the Locality Sensitive Hashing (LSH), a probabilistic approximate nearest neighbor method. Progressively complex geometric constraints are subsequently enforced to further prune the initial candidates and eliminate false correspondences due to inaccuracies in the surface descriptors and the errors of the LSH algorithm. The indexed models are selected based on the MAP rule using posterior probability of the models estimated in the joint 3D-signature space. Experiments with real 3D data employing a large database of vehicles, most of them very similar in shape, containing 1,000,000 features from more than 365 models demonstrate a high degree of performance in the presence of occlusion and obscuration, unmodeled vehicle interiors and part articulations, with an average processing time between 50 and 100 seconds per query. Bogdan Matei, Ying Shan, Harpreet Sawhney, Rakesh Kumar 0001, Daniel F. Huber, Martial Hebert |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2006 | Shapeme Histogram Projection and Matching for Partial Object RecognitionabstractHistograms of shape signature or prototypical shapes, called shapemes, have been used effectively in previous work for 2D/3D shape matching and recognition. We extend the idea of shapeme histogram to recognize partially observed query objects from a database of complete model objects. We propose representing each model object as a collection of shapeme histograms and match the query histogram to this representation in two steps: 1) compute a constrained projection of the query histogram onto the subspace spanned by all the shapeme histograms of the model and 2) compute a match measure between the query histogram and the projection. The first step is formulated as a constrained optimization problem that is solved by a sampling algorithm. The second step is formulated under a Bayesian framework, where an implicit feature selection process is conducted to improve the discrimination capability of shapeme histograms. Results of matching partially viewed range objects with a 243 model database demonstrate better performance than the original shapeme histogram matching algorithm and other approaches. Ying Shan, Harpreet Sawhney, Bogdan Matei, Rakesh Kumar 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Vehicle Fingerprinting for Reacquisition and Tracking in VideosabstractVisual recognition of objects through multiple observations is an important component of object tracking. We address the problem of vehicle matching when multiple observations of a vehicle are separated in time such that frames of observations are not contiguous, thus prohibiting the use of standard frame-to-frame data association. We employ features extracted over a sequence during one time interval as a vehicle fingerprint that is used to compute the likelihood that two or more sequence observations are from the same or different vehicles. The challenges of change in pose, aspect and appearances across two disparate observations are handled by combining feature-based quasi-rigid alignment with flexible matching between two or more sequences. The current work uses the domain of vehicle tracking from aerial platforms where typically both the imaging platform and the vehicles are moving and the number of pixels on the object are limited to fairly low resolutions. Extensive evaluation with respect to ground truth is reported in the paper. Yanlin Guo, Steven C. Hsu, Ying Shan, Harpreet Sawhney, Rakesh Kumar 0001 |
CVPR (2) | 4 |
| 2005 | Unsupervised Learning of Discriminative Edge Measures for Vehicle Matching between Non-Overlapping CamerasabstractThis paper proposes a method for matching road vehicles between two non-overlapping cameras. The matching problem is formulated as a same-different classification problem: probability of two observations from two distinct cameras being from the same vehicle or from different vehicles. We employ a measurement vector consists of three independent edge-based measures and their associated robust measures computed from a pair of aligned vehicle edge maps. The weight of each match measure in the final decision is determined by a unsupervised learning process so that the same-different classification can be optimally separated in the combined measurement space. The robustness of the match measures and the use of discriminant analysis in the classification ensure that the proposed method performs better than existing edge-based approaches, especially in the presence of missing/false edges caused by shadows and different illumination conditions, and systematic misalignment caused by different camera configurations. Extensive experiments based on real data of over 200 vehicles at different times of day demonstrate promising results. Ying Shan, Harpreet Sawhney, Rakesh Kumar 0001 |
CVPR (1) | 2 |
| 2005 | Real-Time Wide Area Multi-Camera Stereo TrackingabstractWe present a fully integrated real-time system to track humans with a network of stereo sensors over a wide area. The processing includes single camera tracking and multi-camera fusion. Each single camera detects and tracks humans in its own view and a multi-camera fusion module combines all the local tracks of the same human into a global track. We propose stereo segmentation and tracking techniques to handle multiple humans moving in groups in cluttered environments. We have developed a ground-based fusion method for camera handoff using space-time constraint. We show results and performance evaluation on very challenging data from a 12-camera system. Manoj Aggarwal, Rakesh Kumar 0001, Harpreet Sawhney |
CVPR (1) | 4 |
| 2005 | Vehicle Identification between Non-Overlapping Cameras without Direct Feature MatchingabstractWe propose a novel method for identifying road vehicles between two nonoverlapping cameras. The problem is formulated as a same-different classification problem: probability of two vehicle images from two distinct cameras being from the same vehicle or from different vehicles. The key idea is to compute the probability without matching the two vehicle images directly, which is a process vulnerable to drastic appearance and aspect changes. We represent each vehicle image as an embedding amongst representative exemplars of vehicles within the same camera. The embedding is computed as a vector each of whose components is a nonmetric distance for a vehicle to an exemplar. The nonmetric distances are computed using robust matching of oriented edge images. A set of truthed training examples of same-different vehicle pairings across the two cameras is used to learn a classifier that encodes the probability distributions. A pair of the embeddings representing two vehicles across two cameras is then used to compute the same-different probability. In order for the vehicle exemplars to be representative for both cameras, we also propose a method for jointly selection of corresponding exemplars using the training data. Experiments on observations of over 400 vehicles under drastically illumination and camera conditions demonstrate promising results. Ying Shan, Harpreet Sawhney, Rakesh Kumar 0001 |
ICCV | 2 |
| 2005 | Clustering multiple image sequences with a sequence-to-sequence similarity measureabstractWe propose a novel similarity measure of two image sequences based on shapeme histograms. The idea of shapeme histogram has been used for single image/texture recognition, but is used here to solve the sequence-to-sequence matching problem. We develop techniques to represent each sequence as a set of shapeme histograms, which captures different variations of the object appearances within the sequence. These shapeme histograms are computed from the set of 2D invariant features that are stable across multiple images in the sequence, and therefore minimizes the effect of both background clutter, and 2D pose variations. We define sequence similarity measure as the similarity of the most similar pair of images from both sequences. This definition maximizes the chance of matching between two sequences of the same object, because it requires only part of the sequences being similar. We also introduce a weighting scheme to conduct an implicit feature selection process during the matching of two shapeme histograms. Experiments on clustering image sequences of tracked objects demonstrate the efficacy of the proposed method. Ying Shan, Harpreet Sawhney, Art Pope |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2004 | Linear Model Hashing and Batch RANSAC for Rapid and Accurate Object Recognition
Ying Shan, Bogdan Matei, Harpreet Sawhney, Rakesh Kumar 0001, Daniel F. Huber, Martial Hebert |
CVPR (2) | 3 |
| 2004 | Partial Object Matching with Shapeme Histograms
Ying Shan, Harpreet Sawhney, Bogdan Matei, Rakesh Kumar 0001 |
ECCV (3) | 2 |
| 2004 | Depth map compression for real-time view-based rendering
Bing-Bing Chai, Sriram Sethuraman, Harpreet Sawhney, Paul Hatrack |
Pattern Recognit. Lett. | 3 |
| 2002 | Is Super-Resolution with Optical Flow Feasible?
Wenyi Zhao, Harpreet Sawhney |
ECCV (1) | 2 |
| 2002 | Object Tracking with Bayesian Estimation of Dynamic Layer RepresentationsabstractDecomposing video frames into coherent 2D motion layers is a powerful method for representing videos. Such a representation provides an intermediate description that enables applications such as object tracking, video summarization and visualization, video insertion, and sprite-based video compression. Previous work on motion layer analysis has largely concentrated on two-frame or multi-frame batch formulations. The temporal coherency of motion layers and the domain constraints on shapes have not been exploited. This paper introduces a complete dynamic motion layer representation in which spatial and temporal constraints on shape, motion and layer appearance are modeled and estimated in a maximum a-posteriori (MAP) framework using the generalized expectation-maximization (EM) algorithm. In order to limit the computational complexity of tracking arbitrarily shaped layer ownership, we propose a shape prior that parameterizes the representation of shape and prevents motion layers from evolving into arbitrary shapes. In this work, a Gaussian shape prior is chosen to specifically develop a near-real-time tracker for vehicle tracking in aerial videos. However, the general idea of using a parametric shape representation as part of the state of a tracker is a powerful one that can be extended to other domains as well. Based on the dynamic layer representation, an iterative algorithm is developed for continuous object tracking over time. The proposed method has been successfully applied in an airborne vehicle tracking system. Its performance is compared with that of a correlation-based tracker and a motion change-based tracker to demonstrate the advantages of the new method. Examples of tracking when the backgrounds are cluttered and the vehicles undergo various rigid motions and complex interactions such as passing, turning, and stop-and-go demonstrate the strength of the complete dynamic layer representation. Harpreet Sawhney, Rakesh Kumar 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | Learning-Based Building Outline Detection from Multiple Aerial ImagesabstractThis paper presents a method for detecting building outlines using multiple aerial images. Since data-driven techniques may not be able to account for variability of building geometry and appearances, a key insight explored in this paper is a combination of model-based data driven front end with data driven learning in the back end for increased detection accuracy. The three main components of the detection algorithm are: (i) initialization. Image intensity and depth information are integrally used to efficiently detect buildings, and a robust rectilinear path finding algorithm is adopted to obtain good initial outlines. The initialization process involves the following steps: detecting location of buildings, determining the dominant orientations and knot points in the building outline and using these to fit the initial outline; (ii) learning. A compact set of building features are defined and learned from the well-delineated buildings, and a tree-based classifier is applied to the whole region to detect any missing buildings and obtain their rough outlines; and (iii) verification and refinement. Learned features are used to remove falsely detected buildings, and all outlines are refined by the deformation of rectilinear templates. The experiments, with improved detection rate and precise outlines, demonstrate the applicability of our algorithm. Yanlin Guo, Harpreet Sawhney, Rakesh Kumar 0001, Steven C. Hsu |
CVPR (2) | 2 |
| 2001 | Dynamic Depth Recovery from Multiple Synchronized Video StreamsabstractThis paper addresses the problem of extracting depth information of non-rigid dynamic 3D scenes from multiple synchronized video streams. Three main issues are discussed in this context. (i) temporally consistent depth estimation, (ii) sharp depth discontinuity estimation around object boundaries, and (iii) enforcement of the global visibility constraint. We present a framework in which the scene is modeled as a collection of 3D piecewise planar surface patches induced by color based image segmentation. This representation is continuously estimated using an incremental formulation in which the 3D geometric, motion, and global visibility constraints are enforced over space and time. The proposed algorithm optimizes a cost function that incorporates the spatial color consistency constraint and a smooth scene motion model. Harpreet Sawhney, Rakesh Kumar 0001 |
CVPR (1) | 2 |
| 2001 | A Global Matching Framework for Stereo Computation
Harpreet Sawhney, Rakesh Kumar 0001 |
ICCV | 2 |
| 2001 | Hybrid stereo camera: an IBR approach for synthesis of very high resolution stereoscopic image sequencesabstractThis paper introduces a novel application of IBR technology for efficient rendering of high quality CG and live action stereoscopic sequences. Traditionally, IBR has been applied to render novel views using image and depth based representations of the plenoptic functions. In this work, we present a restricted form of IBR in which lower resolution images for the views to be generated at a very high resolution are assumed to be available. Specifically, the paper addresses the problem of synthesizing stereo IMAX(R)1 3D motion picture images at a standard resolution of 4-6K. At such high resolutions, producing CG content is extremely time consuming and capturing live action requires bulky cameras. We propose a Hybrid Stereo Camera concept in which one view is rendered at the target high resolution but the other is rendered at a much lower resolution. Methods for synthesizing the second view sequence at the target resolution using image analysis and IBR techniques are the focus of this work. The high quality results from the techniques presented in this paper have been visually evaluated in the IMAX 3D large screen projection environment. The paper also highlights generalizations and extensions of the hybrid stereo camera concept. Harpreet Sawhney, Yanlin Guo, Keith J. Hanna, Rakesh Kumar 0001, Sean Adkins, Samuel Zhou |
SIGGRAPH | 1 |
| 2001 | Aerial video surveillance and exploitationabstractThere is growing interest in performing aerial surveillance using video cameras. Compared to traditional framing cameras, video cameras provide the capability to observe ongoing activity within a scene and to automatically control the camera to track the activity. However, the high data rates and relatively small field of view of video cameras present new technical challenges that must be overcome before such cameras can be widely used. In this paper, we present a framework and details of the key components for real-time, automatic exploitation of aerial video for surveillance applications. The framework involves separating an aerial video into the natural components corresponding to the scene. Three major components of the scene are the static background geometry, moving objects, and appearance of the static and dynamic components of the scene. In order to delineate videos into these scene components, we have developed real time, image-processing techniques for 2-D/3-D frame-to-frame alignment, change detection, camera control, and tracking of independently moving objects in cluttered scenes. The geo-location of video and tracked objects is estimated by registration of the video to controlled reference imagery, elevation maps, and site models. Finally static, dynamic and reprojected mosaics may be constructed for compression, enhanced visualization, and mapping applications. Rakesh Kumar 0001, Harpreet Sawhney, Supun Samarasekera, Steven C. Hsu, Yanlin Guo, Keith J. Hanna, Art Pope, Richard P. Wildes, David J. Hirvonen, Michael W. Hansen, Peter Burt |
Proc. IEEE | 2 |
| 2000 | Multi-View 3D Analysis with Applications for Augmented Reality and Enhanced Video VisualizationabstractThe article presents methods for 3D scene geometry recovery/refinement and pose estimation from motion imagery in two representative scenarios. First, we present a method for pose estimation and scene geometry recovery from extended sequences without prior knowledge of the scheme. Second, we discuss how to recover camera poses when a rough scene model is provided. We show how to extend and refine the scene model using the recovered poses. Finally, we present applications of the above techniques for 3D imagery manipulation such as enhanced visualization for video and 3D insertion of synthetic objects in the imagery. Yanlin Guo, Steven C. Hsu, Supun Samarasekera, Harpreet Sawhney, Rakesh Kumar 0001 |
CVPR | 4 |
| 2000 | Pose Estimation, Model Refinement, and Enhanced Visualization Using VideoabstractIn this paper we present methods for exploitation and enhanced visualization of video given a prior coarse untextured polyhedral model of a scene. Since it is necessary to estimate the 3D poses of the moving camera, we develop an algorithm where tracked features are used to predict the pose between frames and the predicted poses are refined by a coarse to fine process of aligning projected 3D model line segments to oriented image gradient energy pyramids. The estimated poses can be used to update the model with information derived from video, and to re-project and visualize the video from different points of view with a larger scene context. Via image registration, we update the placement of objects in the model and the 3D shape of new or erroneously modeled objects, then map video texture to the model. Experimental results are presented for long aerial and ground level videos of a large-scale urban scene. Stephen C. Hsu, Supun Samarasekera, Rakesh Kumar 0001, Harpreet Sawhney |
CVPR | 4 |
| 2000 | Dynamic Layer Representation with Applications to TrackingabstractA dynamic layer representation is proposed for tracking moving objects. Previous work on layered representations has largely concentrated on two-/multi-frame batch formulations, and tracking research has not addressed the issue of joint estimation of object motion ownership and appearance. The paper extends the estimation of layers in a dynamic scene to incremental estimation formulation and demonstrates how this naturally solves the tracking problem. The three components of the dynamic layer representation, namely, layer motion, ownership, and appearance, are estimated simultaneously over time in a MAP framework. In order to enforce a global shape constraint and to maintain the layer segmentation over time, a parametric segmentation prior is proposed. The generalized EM algorithm is employed to compute the optimal solution. We show the results on real-time tracking of multiple moving or static objects in a cluttered scene imaged from a moving aerial video camera. The moving objects may do complex motions, and have complex interactions such as passing. By using both the appearance and the segmentation information, many difficult tracking tasks are reliably handled. Harpreet Sawhney, Rakesh Kumar 0001 |
CVPR | 2 |
| 2000 | 3D Manipulation of Motion ImageryabstractWe present a set of automatic methods for the recovery and refinement of 3D scene geometry and camera poses from motion imagery. First, we present a two-frame "direct" method, which simultaneously estimates both relative pose between the cameras and 3D scene geometry using information from the images alone. Second, we discuss how to automatically estimate pose and scene geometry using extended sequences. Third, we present methods for recovery of pose when a scene model is known a priori. We show how the recovered pose can be use to extend and refine the scene model. Finally, we present applications of the above techniques for 3D imagery manipulation such as enhanced visualization of video, synthetic view generation and insertion of synthetic objects in the imagery. Rakesh Kumar 0001, Harpreet Sawhney, Yanlin Guo, Steven C. Hsu, Supun Samarasekera |
ICIP | 2 |
| 2000 | Terrain Reconstruction for Ground and Underwater RobotsabstractWe describe a new image-processing algorithm for estimating both the egomotion of an outdoor robotic platform and the structure of the surrounding terrain. The algorithm is based on correlation, and is embedded in an iterative, multi-resolution framework. As such, it is suited to outdoor ground-based and underwater scenes. Both single-camera rigs and multiple-camera rigs can be accommodated. The use of multiple synchronized cameras results in more rapid convergence of the iterative approach. We describe how the algorithm operates, and give examples of its application to three robotic domains: 1) autonomous mobility of ground-based outdoor robots, 2) reconnaissance tasks on ground-based vehicles, and 3) underwater robotics. Robert Mandelbaum, Garbis Salgian, Harpreet Sawhney, Michael W. Hansen |
ICRA | 3 |
| 2000 | Global matching criterion and color segmentation based stereoabstractIn this paper, we present a new analysis by synthesis computational framework for stereo vision. It is designed to achieve the following goals: (1) enforcing global visibility constraints, (2) obtaining reliable depth for depth boundaries and thin structures, (3) obtaining correct depth for textureless regions, and (4) hypothesizing correct depth for unmatched regions. The framework employs depth and visibility based rendering within a global matching criterion to compute depth in contrast with approaches that rely on local matching measures and relaxation. A color segmentation based depth representation guarantees smoothness in textureless regions. Hypothesizing depth from neighboring segments enables propagation of correct depth and produces reasonable depth values for unmatched region. A practical algorithm that integrates all these aspects is presented in this paper. Comparative experimental results are shown for real images. Results on new view rendering based on a single stereo pair are also demonstrated. Harpreet Sawhney |
WACV | 2 |
| 2000 | Independent Motion Detection in 3D ScenesabstractThis paper presents an algorithmic approach to the problem of detecting independently moving objects in 3D scenes that are viewed under camera motion. There are two fundamental constraints that can be exploited for the problem: 1) two/multiview camera motion constraint (for instance, the epipolar/trilinear constraint) and 2) shape constancy constraint. Previous approaches to the problem either use only partial constraints, or rely on dense correspondences or flow. We employ both the fundamental constraints in an algorithm that does not demand a priori availability of correspondences or flow. Our approach uses the plane-plus-parallax decomposition to enforce the two constraints. It is also demonstrated that for a class of scenes, called sparse 3D scenes in which genuine parallax and independent motions may be confounded, how the plane-plus-parallax decomposition allows progressive introduction, and verification of the fundamental constraints. Results of the algorithm on some difficult sparse 3D scenes are promising. Harpreet Sawhney, Yanlin Guo, Rakesh Kumar 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1999 | Correlation-based Estimation of Ego-Motion and Structure from Motion and StereoabstractThis paper describes a correlation-based, iterative, multi-resolution algorithm which estimates both scene structure and the motion of the camera rig through an environment from the stream(s) of incoming images. Both single-camera rigs and multiple-camera rigs can be accommodated. The use of multiple synchronized cameras results in more rapid convergence of the iterative approach. The algorithm uses a global ego-motion constraint to refine estimates of inter-frame camera rotation and translation. It uses local window-based correlation to refine the current estimate of scene structure. All analysis is performed at multiple resolutions. In order to combine, in a straightforward way, the correlation surfaces from multiple viewpoints and from multiple pixels in a support region, each pixel's correlation surface is modeled as a quadratic. This parameterization allows direct, explicit computation of incremental refinements for ego-motion and structure using linear algebra. Batches can be of arbitrary size, allowing a trade-off between accuracy and latency. Batches can also be daisy-chained for extended sequences. Results of the algorithm are shown on synthetic and real outdoor image sequences. Robert Mandelbaum, Garbis Salgian, Harpreet Sawhney |
ICCV | 3 |
| 1999 | Independent Motion Detection in 3D ScenesabstractPresents an algorithmic approach to the problem of detecting independently moving objects in 3D scenes that are viewed under camera motion. There are two fundamental constraints that can be exploited for the problem: (i) a two- (or multi-)view camera motion constraint (for instance, the epipolar/trilinear constraint), and (ii) a shape constancy constraint. Previous approaches to the problem either only used partial constraints or relied on dense correspondences or flow. We employ both of these fundamental constraints in an algorithm that does not demand a-priori availability of correspondences or flow. Our approach uses the plane-plus-parallax decomposition to enforce the two constraints. It is also demonstrated, for a class of scenes called sparse 3D scenes, in which genuine parallax and independent motions may be confounded, how the plane-plus-parallax decomposition allows progressive introduction and verification of the fundamental constraints. The results of applying the algorithm to some difficult sparse 3D scenes look promising. Harpreet Sawhney, Yanlin Guo, Jane C. Asmuth, Rakesh Kumar 0001 |
ICCV | 1 |
| 1999 | True Multi-Image Alignment and Its Application to Mosaicing and Lens Distortion CorrectionabstractMultiple images of a scene are related through 2D/3D view transformations and linear and nonlinear camera transformations. We present an algorithm for true multi-image alignment that does not rely on the measurements of a reference image being distortion free. The algorithm is developed to specifically align and mosaic images using parametric transformations in the presence of lens distortion. When lens distortion is present, none of the images can be assumed to be ideal. In our formulation, all the images are modeled as intensity measurements represented in their respective coordinate systems, each of which is related to an ideal coordinate system through an interior camera transformation and an exterior view transformation. The goal of the accompanying algorithm is to compute an image in the ideal coordinate system while solving for the transformations that relate the ideal system with each of the data images. Key advantages of the technique presented in this paper are: (i) no reliance on one distortion free image, (ii) ability to register images and compute coordinate transformations even when the multiple images are of an extended scene with no overlap between the first and last frame of the sequence, and (iii) ability to handle linear and nonlinear transformations within the same framework. Results of applying the algorithm are presented for the correction of lens distortion, and creation of video mosaics. Harpreet Sawhney, Rakesh Kumar 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1998 | Robust Video Mosaicing through Topology Inference and Local to Global Alignment
Harpreet Sawhney, Steven C. Hsu, Rakesh Kumar 0001 |
ECCV (2) | 1 |
| 1998 | Registration of video to geo-referenced imageryabstractThe ability to locate scenes and objects visible in aerial video imagery with their corresponding locations and coordinates in a reference coordinate system is important in visually-guided navigation, surveillance and monitoring systems. However, a key technical problem of locating objects and scenes in a video with their geo-coordinates needs to be solved in order to ascertain the geo-location of objects seen from the camera platform's current location. We present the key algorithms for the problem of accurate mapping between camera coordinates and geo-coordinates, called geo-spatial registration. Current systems for geo-location use the position and attitude information for the moving platform in some fixed world coordinates to locate the video frames in the reference database. However, the accuracy achieved is only of the order of 100s of pixels. Our approach utilizes the imagery and terrain information contained in the geo-spatial database to precisely align dynamic videos with the reference imagery and thus achieves a much higher accuracy. Applications of geo-spatial registration include aerial mapping, target location and tracking and enhanced visualization such as the overlay of textual/graphical annotations of objects of interest in the current video using the stored annotations in the reference database. Rakesh Kumar 0001, Harpreet Sawhney, Jane C. Asmuth, Art Pope, S. Hsu |
ICPR | 2 |
| 1998 | Multimedia applications of computer visionabstractThe technical foundation for many applications of computer vision to multimedia applications is efficient and robust image motion estimation. These algorithms enable the creation of algorithms for mosaic construction, registration of video to a database multi-sensor registration, 3D estimation and representation, and video content indexing and retrieval. Demonstrations on the above topics will be shown at WACV'98 by the Media Vision Group of Sarnoff Corporation. Jane C. Asmuth, Douglas Dixon, Keith J. Hanna, Steven C. Hsu, Rakesh Kumar 0001, Vince Paragano, Art Pope, Supun Samarasekera, Harpreet Sawhney |
WACV | 9 |
| 1998 | Influence of global constraints and lens distortion on pose and appearance recovery from a purely rotating cameraabstractGiven a video sequence acquired by an uncalibrated camera rotating about a fixed center of projection, it is desired to estimate the appearance of the scene in all possible directions, recover the camera orientations, and determine the constant focal length and radial lens distortion of the camera, without recourse to physical scene measurements. While parts of this problem have been studied before, there has not yet been a comprehensive algorithmic solution that accommodates a variety of trajectories nor a characterization of the prerequisites for successful reconstruction. This paper studies a computationally efficient algorithm that takes maximum advantage of all information available locally among pairs or small groups of frames as well as global closed cycle constraints. It works with any rotational trajectory, not just panning about a fixed axis. Analyzing the algorithm's error surface predicts the sensitivity of estimation to lens distortion and global constraints. Both the analysis and experiments with natural images demonstrate that it is crucial to use either lens distortion compensation or closed cycle constraints to get accurate parameter estimates and well-aligned mosaic images. Steven C. Hsu, Harpreet Sawhney |
WACV | 2 |
| 1998 | VideoBrushTM: experiences with consumer video mosaicingabstractWe present technical advances that were made in creating consumer level video mosaicing applications with inputs from hand held inexpensive cameras, and with near-real time response. In developing the VideoBrush/sup TM/ technologies, the goal was to give easily usable but robust video mosaicing capabilities in the hands of lay users. We highlight a number of problems that needed to be solved to enable this: 2D & 1D modes of scene scanning and alignment, progressive complexity of alignment models to provide real-time response to the user and multi-resolution blending for high quality mosaic output. The solution to these key technical challenges has led us to video mosaic technologies that can "scan" scenes as diverse as natural panoramic scenes to white boards and documents. Harpreet Sawhney, Rakesh Kumar 0001, Gary Gendel, James R. Bergen, Douglas Dixon, Vince Paragano |
WACV | 1 |
| 1997 | True Multi-Image Alignment and its Application to Mosaicing and Lens Distortion CorrectionabstractMultiple images of a scene are related through 2D/3D view transformations and linear and non-linear camera transformations. In all the traditional techniques to compute these transformations, especially the ones relying on direct intensity gradients, one image and its coordinate system have been assumed to be ideal and distortion free. In this paper, we present a formulation and an algorithm for true multi-image alignment that does not rely on the measurements of a reference image being distortion free. For instance, in the presence of lens distortion, none of the images can be assumed to be ideal. In our formulation, all the images are modeled as intensity measurements represented in their respective coordinate systems, each of which is related to an ideal coordinate system through an interior camera transformation and an exterior view transformation. The goal of the accompanying algorithm is to compute an image in the ideal coordinate system while solving for the transformations that relate the ideal system with each of the data images. Key advantages of the technique presented in this paper are: (i) no reliance on one distortion free image, (ii) ability to register images and compute coordinate transformations even when the multiple images are of an extended scene with no overlap between the first and last frame of the sequence, and (iii) ability to handle linear and non-linear transformations within the same framework. The new algorithm is evaluated in the context of two applications: (i) correction of lens distortion, and (ii) creation of video mosaics. Harpreet Sawhney, Rakesh Kumar 0001 |
CVPR | 1 |
| 1997 | Interactive content-based video indexing and browsingabstractIn this paper we present a framework for efficient representation, access, and manipulation of video data. Our approach is based on decomposing video information into its spatial (appearance), temporal (dynamics), geometric components. This derived information is organized into data representations that support non-linear browsing and efficient indexing to provide rapid access directly to the information of interest. Michal Irani, Harpreet Sawhney, Rakesh Kumar 0001, P. Anandan 0001 |
MMSP | 2 |
| 1996 | Compact Representations of Videos Through Dominant and Multiple Motion EstimationabstractAn explosion of on-line image and video data in digital form is already well underway. With the exponential rise in interactive information exploration and dissemination through the World-Wide Web (WWW), the major inhibitors of rapid access to on-line video data are costs and management of capture and storage, lack of real-time delivery, and nonavailability of content-based intelligent search and indexing techniques. The solutions for capture, storage, and delivery may be on the horizon or a little beyond. However, even with rapid delivery, the lack of efficient authoring and querying tools for visual content-based indexing may still inhibit as widespread a use of video information as that of text and traditional tabular data is currently. In order to be able to nonlinearly browse and index into videos through visual content, it is necessary to develop authoring tools that can automatically separate moving objects and significant components of the scene, and represent these in a compact form. Given that video data comes in torrents-almost a megabyte every 30th of a second-it will be highly inefficient to search for objects and scenes in every frame of a video. In this paper, we present techniques to automatically derive compact representations of scenes and objects from the motion information. Image motion is a significant cue in videos for the separation of scenes into their significant components and for the separation of moving objects. Motion analysis is useful in capturing the visual content of videos for indexing and browsing in two different ways. First, separation of the static scene from moving objects can be accomplished by employing dominant 2D/3D motion estimation methods. Alternatively, if the goal is to be able to represent the fixed scene too as a composition of significant structures and objects, then simultaneous multiple motion methods might be more appropriate. In either case, view-based summarized representations of the scene can be created by video compositing/mosaicing based on the estimated motions. We present robust algorithms for both kinds of representations: 1) dominant motion estimation based techniques which exploit a fairly common occurrence in videos that a mostly fixed background (scene) is imaged with or without independently moving objects, and 2) simultaneous multiple motion estimation and representation of motion video using layered representations. Ample examples of the representations achieved by each method are included in the paper. Harpreet Sawhney, Serge Ayer |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1995 | Layered Representation of Motion Video Using Robust Maximum-Likelihood Estimation of Mixture Models and MDL EncodingabstractRepresenting and modeling the motion and spatial support of multiple objects and surfaces from motion video sequences is an important intermediate step towards dynamic image understanding. One such representation, called layered representation, has recently been proposed. Although a number of algorithms have been developed for computing these representations, there has not been a consolidated effort into developing a precise mathematical formulation of the problem. This paper presents one such formulation based on maximum likelihood estimation (MLE) of mixture models and the minimum description length (MDL) encoding principle. The three major issues in layered motion representation are: (i) how many motion models adequately describe image motion, (ii) what are the motion model parameters, and (iii) what is the spatial support layer for each motion model.> Serge Ayer, Harpreet Sawhney |
ICCV | 2 |
| 1995 | Model-Based 2D&3D Dominant Motion Estimation for Mosaicing and Video RepresentationabstractIt is fairly common in video sequences that a mostly fixed background (scene) is imaged with or without objects. The dominant background changes in the image plane mostly due to camera operations and motion (zoom, pan, tilt, track etc.). We address the problem of computation of the dominant image transformation over time and demonstrate how this can be effectively used for efficient video representation through video mosaicing and image registration. We formulate the problem of dominant component estimation as that of model based robust estimation using M estimators with direct, multi resolution methods. In addition to 2D affine and plane projective models, that have been used in the past for describing image motion using direct methods, we also employ a true 3D model of motion and scene structure imaged with uncalibrated cameras. This model parameterizes the image motion as that due to a planar component and a parallax component. For rigid 3D scenes imaged under camera motion only, least squares (LS) methods with the plane and parallax parameterization are also presented. Furthermore, in the context of robust estimation, in contrast with previous approaches for similar problems, our algorithm employs an automatic computation of a scale parameter that is crucial in rejecting the non dominant components as outliers.> Harpreet Sawhney, Serge Ayer, Monika Gorkani |
ICCV | 1 |
| 1995 | Dominant and multiple motion estimation for video representationabstractThe major inhibitors of rapid access to online video data are costs and management of capture and storage, lack of high-speed real-time delivery and non-availability of content and context based intelligent search and indexing techniques. The solutions for capture, storage and delivery maybe on the horizon, however the lack of visual content based indexing of video and image information may still inhibit as widespread a use of this information modality as that of text or tabular data is currently. We present techniques for compact visual representation of video data that will be useful for visual content based presentation and indexing. Video data comes in torrents-almost a megabyte every 30th of a second-but also affords the exploitation of relatively smoothly changing information over time. The techniques presented exploit the motion information across video frames to represent the underlying scene in a compact visual form as it is seen across many slowly varying frames in a video. Two classes of techniques are presented: (i) dominant motion estimation based techniques which exploit a fairly common occurrence in videos that a mostly fixed background (scene) is imaged with or without independently moving objects, and (ii) simultaneous multiple motion estimation and representation of motion video using layered representations. Harpreet Sawhney, Serge Ayer, Monika Gorkani |
ICIP | 1 |
| 1995 | Fast Similarity Search in the Presence of Noise, Scaling, and Translation in Time-Series Databases
Rakesh Agrawal 0001, King-Ip Lin, Harpreet Sawhney, Kyuseok Shim |
VLDB | 3 |
| 1995 | Efficient Color Histogram Indexing for Quadratic Form Distance FunctionsabstractIn image retrieval based on color, the weighted distance between color histograms of two images, represented as a quadratic form, may be defined as a match measure. However, this distance measure is computationally expensive and it operates on high dimensional features (O(N)). We propose the use of low-dimensional, simple to compute distance measures between the color distributions, and show that these are lower bounds on the histogram distance measure. Results on color histogram matching in large image databases show that prefiltering with the simpler distance measures leads to significantly less time complexity because the quadratic histogram distance is now computed on a smaller set of images. The low-dimensional distance measure can also be used for indexing into the database.> James Lee Hafner, Harpreet Sawhney, William Equitz, Myron Flickner, Wayne Niblack |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1994 | 3D geometry from planar parallaxabstractDeriving 3D structure in a fixed object-centered coordinate system is an increasingly popular trend in shape from multiple views. For linear approximations to perspective projection (weak/para perspective), and for the case of image velocities, elegant linear methods have been devised for robust estimation. For reconstruction under arbitrary view transformations, linear projective methods using point correspondences have been suggested. In this paper, we formulate the problem of intrinsic 3D structure estimation through perspective projection using motion parallax, defined with respect to an arbitrary plane in the environment. It is shown that if an image coordinate system is warped using plane projective transformation with respect to a reference view, the residual image motion is dependent only on the epipoles and has a simple relation to the 3D structure. Our computational scheme avoids point/line correspondence and is based on hierarchical estimation and image warping working directly with spatio-temporal image intensities.> Harpreet Sawhney |
CVPR | 1 |
| 1994 | Efficient Color Histogram IndexingabstractIn image retrieval based on color, the weighted distance between color histograms of two images, represented as a quadratic form, may be defined as a match measure. However, this distance measure is computationally expensive (naively O(N/sup 2/) and at best O(N) in the number N of histogram bins) and it operates on high dimensional features (O(N)). We propose the use of low-dimensional, simple to compute distance measures between the color distributions, and show that these are lower bounds on the histogram distance measure. Results on color histogram matching in large image databases show that pre-filtering with the simpler distance measures leads to significantly less time complexity because the quadratic histogram distance is now computed on a smaller set of images. The low-dimensional distance measure can also be used for indexing into the database.> Harpreet Sawhney, James Lee Hafner |
ICIP (2) | 1 |
| 1994 | Simplifying motion and structure analysis using planar parallax and image warpingabstractRobust 3D motion and structure computation and segmentation has been the subject of an enormous body of work in reconstructive vision. For linear approximations to perspective projection (weak/para perspective), and for the case of image velocities, elegant linear methods have been devised for robust estimation. For reconstruction under arbitrary view transformations, linear projective methods using point correspondences have been suggested. In this paper, the authors present a formulation for 3D motion and structure analysis using motion parallax defined with respect to an arbitrary plane in the environment. It is shown that if an image coordinate system is warped using plane projective transformation with respect to a reference view, the residual image motion is dependent only on the epipoles and has a simple relation to the 3D structure. The author's computational scheme avoids point/line correspondence and is based on hierarchical estimation and image warping working directly with spatio-temporal image intensities. Results on real images demonstrate how this analysis simplifies ego and multiple motion analysis, and stable scene-centered 3D reconstruction. Harpreet Sawhney |
ICPR (1) | 1 |
| 1993 | Trackability as a cue for potential obstacle identification and 3-D description
Harpreet Sawhney, Allen R. Hanson |
Int. J. Comput. Vis. | 1 |
| 1993 | Image Description and 3-D Reconstruction From Image Trajectories of Rotational MotionabstractA new technique for reconstructing the 3-D structure and motion of a scene undergoing relative rotational motion with respect to the camera is discussed. Given image correspondences of point features tracked over many frames, a two-stage technique for reconstruction is presented. A grouping algorithm that exploits spatio-temporal constraints of the common motion to achieve a reliable description of discrete point correspondences as curved trajectories in the image plane is developed. In contrast, trajectories fitted to points independent of each other lead to arbitrary image descriptions and very inaccurate 3-D parameters. A new closed-form solution, under perspective projection, for the 3-D motion and location of points from the computed image trajectories is also presented. Both stages are applied to real image sequences with good results. This approach represents a first step in a longer-term research effort examining the role of explicit spatio-temporal organization in the interpretation of scenes from dynamic images.> Harpreet Sawhney, John Oliensis, Allen R. Hanson |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1992 | 3D model acquisition from monocular image sequencesabstractAutomatic building 3D models of objects and scenes from a sequence of 2D monocular images is approached by first building a partial model (possibly noisy) and then extending and refining it. The initial model is built by tracking and reconstructing shallow structures over a sequence of images using the constraint of affine trackability. This model is subsequently used to compute the pose that relates the model coordinate system and the camera coordinate system of the image frames in the sequence. The unmodeled 3D features (those not recovered by the shallow structure reconstruction) are tracked over the image sequence and their 3D locations recovered by a pseudotriangulation process. The triangulation process is also used to make new 3D measurements of the initial model points. These measurements are then fused with the previous estimates to refine the set of initial model points.> Rakesh Kumar 0001, Harpreet Sawhney, Allen R. Hanson |
CVPR | 2 |
| 1992 | Affine trackability aids obstacle detectionabstractPotential obstacles in the path of a mobile robot that can often be characterized as shallow (i.e., their extent in depth is small compared to their distance from the camera) are considered. The constraint of affine trackability is applied to automatic identification and 3-D reconstruction of shallow structures in realistic scenes. It is shown how this approach can handle independent object motion, occlusion, and motion discontinuity. Although the reconstructed structure is only a frontal plane approximation to the corresponding real structure, the robustness of depth of the approximation might be useful for obstacle avoidance, where the exact shape of an object may not be of consequence so long as collisions with it can be avoided.> Harpreet Sawhney, Allen R. Hanson |
CVPR | 1 |
| 1991 | Identification and 3D description of 'shallow' environmental structure in a sequence of imagesabstractA framework is presented for segmenting shallow structures from their background over a sequence of images. Shallowness is first quantified as affine describability. This is embedded in a tracking system within which hypothesized model structures undergo a cycle of prediction and model-matching. Structures emerge either as shallow or non-shallow based on their affine trackability. This paper rejects continuity heuristics for purely image motion in favor of temporal continuity defined as the consistency of generic 3-D models, namely shallow structures.> Harpreet Sawhney, Allen R. Hanson |
CVPR | 1 |
| 1990 | Description and reconstruction from image trajectories of rotational motionabstractA novel technique is presented for reconstructing the 3-D structure, and motion, of a scene undergoing relative rotational motion with respect to the camera. Given image correspondences of point features tracked over many frames, first, a grouping algorithm developed which exploits spatio-temporal constraints of the common motion to achieve a reliable description of discrete point correspondences as curved trajectories (general conics in the case of rotational motion) in the image plane. In contrast, trajectories fitted to points independent of each other lead to arbitrary image descriptions and very inaccurate 3-D parameters. Second, a novel closed-form solution, under perspective projection, for the 3-D motion and location of points from the computed image trajectories is presented. Both stages are applied to real image sequences with good results.> Harpreet Sawhney, John Oliensis, Allen R. Hanson |
ICCV | 1 |