Larry Davis 0001

dblp:d/LarrySDavis · also Larry S. Davis · DBLP profile ↗
← Back
455ranked-venue papers
29as first author
33since 2021 · last 2025
0000-0002-1656-1123ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 328 · 14 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 303 · 12 first-author · 26 since 2021Systems, architecture and hardware · 21 · 3 first-authorHuman-computer interaction and ubiquitous computing · 13 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 11 · 1 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2025 Text2Outfit: Controllable Outfit Generation With Multimodal Language Models
Yuanhao Zhai 0001, Yen-Liang Lin, Minxu Peng, Larry Davis 0001, Ashwin Chandramouli, Junsong Yuan 0001, David S. Doermann
ICCV4
2025 Fighting Malicious Media Data: A Survey on Tampering Detection and Deepfake Detection
abstract
Online media data, in the form of images and videos, are becoming mainstream communication channels. However, recent advances in deep learning (DL), particularly deep generative models, open the doors for producing perceptually convincing images and videos at a low cost, which not only poses a serious threat to the trustworthiness of digital information but also has severe societal implications. This motivates a growing interest in research in media tampering detection (TD), i.e., using DL techniques to examine whether media data have been maliciously manipulated. Depending on the content of the targeted images, media forgery could be divided into image tampering and Deepfake techniques. The former typically moves or erases the visual elements in ordinary images, while the latter manipulates the expressions and even the identity of human faces. Accordingly, the means of defense include image TD and Deepfake detection (DFD), which share a wide variety of properties. In this article, we provide a comprehensive review of the current media TD approaches and discuss the challenges and trends in this field for future research.
Zhenxin Li, Chao Zhang 0001, Jingjing Chen 0001, Zuxuan Wu, Larry Davis 0001, Yu-Gang Jiang 0001
Proc. IEEE6
2024 Leveraging Bitstream Metadata for Fast, Accurate, Generalized Compressed Video Quality Enhancement
abstract
Video compression is a central feature of the modern internet powering technologies from social media to video conferencing. While video compression continues to mature, for many compression settings, quality loss is still noticeable. These settings nevertheless have important applications to the efficient transmission of videos over bandwidth constrained or otherwise unstable connections. In this work, we develop a deep learning architecture capable of restoring detail to compressed videos which leverages the underlying structure and motion information embedded in the video bitstream. We show that this improves restoration accuracy compared to prior compression correction methods and is competitive when compared with recent deep-learning-based video compression methods on rate-distortion while achieving higher throughput. Furthermore, we condition our model on quantization data which is readily available in the bit-stream. This allows our single model to handle a variety of different compression quality settings which required an ensemble of models in prior work.
Max Ehrlich, Jon Barker, Namitha Padmanabhan, Larry Davis 0001, Andrew Tao, Bryan Catanzaro, Abhinav Shrivastava
WACV4
2024 GRIT: GAN Residuals for Paired Image-to-Image Translation
abstract
Current Image-to-Image translation (I2I) frameworks rely heavily on reconstruction losses, where the output needs to match a given ground truth image. An adversarial loss is commonly utilized as a secondary loss term, mainly to add more realism to the output. Compared to unconditional GANs, I2I translation frameworks have more supervisory signals, but still their output shows more artifacts and does not reach the same level of realism achieved by unconditional GANs. We study the performance gap, in terms of photo-realism, between I2I translation and unconditional GAN frameworks. Based on our observations, we propose a modified architecture and training objective to address this realism gap. Our proposal relaxes the role of reconstruction losses, to act as regularizers instead of doing all the heavy lifting which is common in current I2I frameworks. Furthermore, our proposed formulation decouples the optimization of reconstruction and adversarial objectives and removes pixel-wise constraints on the final output. This allows for a set of stochastic but realistic variations of any target output image. Our project page can be accessed at cs.umd.edu/~sakshams/grit.
Saksham Suri, Moustafa Meshry, Larry Davis 0001, Abhinav Shrivastava
WACV3
2024 Building an Open-Vocabulary Video CLIP Model With Better Architectures, Optimization and Data
abstract
Despite significant results achieved by Contrastive Language-Image Pretraining (CLIP) in zero-shot image recognition, limited effort has been made exploring its potential for zero-shot video recognition. This paper presents Open-VCLIP++, a simple yet effective framework that adapts CLIP to a strong zero-shot video classifier, capable of identifying novel actions and events during testing. Open-VCLIP++ minimally modifies CLIP to capture spatial-temporal relationships in videos, thereby creating a specialized video classifier while striving for generalization. We formally demonstrate that training Open-VCLIP++ is tantamount to continual learning with zero historical data. To address this problem, we introduce Interpolated Weight Optimization, a technique that leverages the advantages of weight interpolation during both training and testing. Furthermore, we build upon large language models to produce fine-grained video descriptions. These detailed descriptions are further aligned with video features, facilitating a better transfer of CLIP to the video domain. Our approach is evaluated on three widely used action recognition datasets, following a variety of zero-shot evaluation protocols. The results demonstrate that our method surpasses existing state-of-the-art techniques by significant margins. Specifically, we achieve zero-shot accuracy scores of 88.1%, 58.7%, and 81.2% on UCF, HMDB, and Kinetics-600 datasets respectively, outpacing the best-performing alternative methods by 8.5%, 8.2%, and 12.3%. We also evaluate our approach on the MSR-VTT video-text retrieval dataset, where it delivers competitive video-to-text and text-to-video retrieval performance, while utilizing substantially less fine-tuning data compared to other methods.
Zuxuan Wu, Zejia Weng, Wujian Peng, Xitong Yang, Ang Li 0001, Larry Davis 0001, Yu-Gang Jiang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 FlexNeRF: Photorealistic Free-viewpoint Rendering of Moving Humans from Sparse Views
abstract
We present FlexNeRF, a method for photorealistic free-viewpoint rendering of humans in motion from monocular videos. Our approach works well with sparse views, which is a challenging scenario when the subject is exhibiting fast/complex motions. We propose a novel approach which jointly optimizes a canonical time and pose configuration, with a pose-dependent motion field and pose-independent temporal deformations complementing each other. Thanks to our novel temporal and cyclic consistency constraints along with additional losses on intermediate representation such as segmentation, our approach provides high quality outputs as the observed views become sparser. We empirically demonstrate that our method significantly outperforms the state-of-the-art on public benchmark datasets as well as a self-captured fashion dataset. The project page is available at: https://flex-nerf.github.io/
Vinoj Jayasundara 0001, Nicolas Heron, Abhinav Shrivastava, Larry Davis 0001
CVPR5
2023 More Than Just Attention: Improving Cross-Modal Attentions with Contrastive Constraints for Image-Text Matching
abstract
Cross-modal attention mechanisms have been widely applied to the image-text matching task. They have achieved remarkable improvements thanks to their capability of learning fine-grained relevance across different modalities. However, the cross-modal attention models of existing methods could be sub-optimal and inaccurate because there is no direct supervision provided during the training process. In this work, we propose two novel training strategies, namely Contrastive Content Resourcing (CCR) and Contrastive Content Swapping (CCS) constraints, to address such limitations. These constraints supervise the training of cross-modal attention models in a contrastive learning manner without requiring explicit attention annotations. They are plug-in training strategies and can be generally integrated into existing cross-modal attention models. Additionally, we introduce three metrics, including Attention Precision, Recall, and F1-Score, to quantitatively measure the quality of learned attention models. We evaluate the proposed constraints by incorporating them into four state- of-the-art cross-modal attention-based image-text matching models. Experimental results on both Flickr30k and MS-COCO datasets demonstrate that integrating these constraints generally improves the model performance in terms of both retrieval performance and attention metrics.
Yuxiao Chen 0002, Long Zhao 0003, Larry Davis 0001, Dimitris N. Metaxas
WACV6
2023 Towards Transferable Adversarial Attacks on Image and Video Transformers
abstract
The transferability of adversarial examples across different convolutional neural networks (CNNs) makes it feasible to perform black-box attacks, resulting in security threats for CNNs. However, fewer endeavors have been made to investigate transferable attacks for vision transformers (ViTs), which achieve superior performance on various computer vision tasks. Unlike CNNs, ViTs establish relationships between patches extracted from inputs by the self-attention module. Thus, adversarial examples crafted on CNNs might hardly attack ViTs. To assess the security of ViTs comprehensively, we investigate the transferability across different ViTs in both untargetd and targeted scenarios. More specifically, we propose a Pay No Attention (PNA) attack, which ignores attention gradients during backpropagation to improve the linearity of backpropagation. Additionally, we introduce a PatchOut/CubeOut attack for image/video ViTs. They optimize perturbations within a randomly selected subset of patches/cubes during each iteration, preventing over-fitting to the white-box surrogate ViT model. Furthermore, we maximize the L2 norm of perturbations, ensuring that the generated adversarial examples deviate significantly from the benign ones. These strategies are designed to be harmoniously compatible. Combining them can enhance transferability by jointly considering patch-based inputs and the self-attention of ViTs. Moreover, the proposed combined attack seamlessly integrates with existing transferable attacks, providing an additional boost to transferability. We conduct experiments on ImageNet and Kinetics-400 for image and video ViTs, respectively. Experimental results demonstrate the effectiveness of the proposed method.
Zhipeng Wei 0001, Jingjing Chen 0001, Micah Goldblum, Zuxuan Wu, Tom Goldstein, Yu-Gang Jiang 0001, Larry Davis 0001
IEEE Trans. Image Process.7
2022 Rethinking Pseudo Labels for Semi-supervised Object Detection
abstract
Recent advances in semi-supervised object detection (SSOD) are largely driven by consistency-based pseudo-labeling methods for image classification tasks, producing pseudo labels as supervisory signals. However, when using pseudo labels, there is a lack of consideration in localization precision and amplified class imbalance, both of which are critical for detection tasks. In this paper, we introduce certainty-aware pseudo labels tailored for object detection, which can effectively estimate the classification and localization quality of derived pseudo labels. This is achieved by converting conventional localization as a classification task followed by refinement. Conditioned on classification and localization quality scores, we dynamically adjust the thresholds used to generate pseudo labels and reweight loss functions for each category to alleviate the class imbalance problem. Extensive experiments demonstrate that our method improves state-of-the-art SSOD performance by 1-2% AP on COCO and PASCAL VOC while being orthogonal and complementary to most existing methods. In the limited-annotation regime, our approach improves supervised baselines by up to 10% AP using only 1-10% labeled data from COCO.
Hengduo Li, Zuxuan Wu, Abhinav Shrivastava, Larry Davis 0001
AAAI4
2022 TAG: Boosting Text-VQA via Text-aware Visual Question-answer Generation
Jun Wang 0090, Mingfei Gao, Yuqian Hu, Ramprasaath R. Selvaraju, Chetan Ramaiah, Ran Xu 0001, Joseph F. JáJá, Larry Davis 0001
BMVC8
2022 Neural Space-Filling Curves
Hanyu Wang 0002, Kamal Gupta 0002, Larry Davis 0001, Abhinav Shrivastava
ECCV (7)3
2022 Responsible Disclosure of Generative Models Using Scalable Fingerprinting
Ning Yu 0006, Vladislav Skripniuk, Dingfan Chen, Larry Davis 0001, Mario Fritz
ICLR4
2022 M3DETR: Multi-representation, Multi-scale, Mutual-relation 3D Object Detection with Transformers
abstract
We present a novel architecture for 3D object detection, M3DETR, which combines different point cloud representations (raw, voxels, bird-eye view) with different feature scales based on multi-scale feature pyramids. M3DETR is the first approach that unifies multiple point cloud representations, feature scales, as well as models mutual relationships between point clouds simultaneously using transformers. We perform extensive ablation experiments that highlight the benefits of fusing representation and scale, and modeling the relationships. Our method achieves state-of-the-art performance on the KITTI 3D object detection dataset and Waymo Open Dataset. Results show that M3DETR improves the baseline significantly by 1.48% mAP for all classes on Waymo Open Dataset. In particular, our approach ranks 1ston the well-known KITTI 3D Detection Benchmark for both car and cyclist classes, and ranks 1ston Waymo Open Dataset with single frame point cloud input. Our code is available at: https://github.com/rayguan97/M3DETR.
Tianrui Guan, Jun Wang 0090, Shiyi Lan, Rohan Chandra, Zuxuan Wu, Larry Davis 0001, Dinesh Manocha
WACV6
2022 Scale Normalized Image Pyramids With AutoFocus for Object Detection
abstract
We present an efficient foveal framework to perform object detection. A scale normalized image pyramid (SNIP) is generated that, like human vision, only attends to objects within a fixed size range at different scales. Such a restriction of objects' size during training affords better learning of object-sensitive filters, and therefore, results in better accuracy. However, the use of an image pyramid increases the computational cost. Hence, we propose an efficient spatial sub-sampling scheme which only operates on fixed-size sub-regions likely to contain objects (as object locations are known during training). The resulting approach, referred to as Scale Normalized Image Pyramid with Efficient Resampling or SNIPER, yields up to 3× speed-up during training. Unfortunately, as object locations are unknown during inference, the entire image pyramid still needs processing. To this end, we adopt a coarse-to-fine approach, and predict the locations and extent of object-like regions which will be processed in successive scales of the image pyramid. Intuitively, it's akin to our active human-vision that first skims over the field-of-view to spot interesting regions for further processing and only recognizes objects at the right resolution. The resulting algorithm is referred to as AutoFocus and results in a 2.5-5× speed-up during inference when used with SNIP. Code: https://github.com/mahyarnajibi/SNIPER.
Mahyar Najibi, Abhishek Sharma 0001, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 A Dynamic Frame Selection Framework for Fast Video Recognition
abstract
We introduce AdaFrame, a conditional computation framework that adaptively selects relevant frames on a per-input basis for fast video recognition. AdaFrame, which contains a Long Short-Term Memory augmented with a global memory to provide context information, operates as an agent to interact with video sequences aiming to search over time which frames to use. Trained with policy search methods, at each time step, AdaFrame computes a prediction, decides where to observe next, and estimates a utility, i.e., expected future rewards, of viewing more frames in the future. Exploring predicted utilities at testing time, AdaFrame is able to achieve adaptive lookahead inference so as to minimize the overall computational cost without incurring a degradation in accuracy. We conduct extensive experiments on two large-scale video benchmarks, FCVID and ActivityNet. With a vanilla ResNet-101 model, AdaFrame achieves similar performance of using all frames while only requiring, on average, 8.21 and 8.65 frames on FCVID and ActivityNet, respectively. We also demonstrate AdaFrame is compatible with modern 2D and 3D networks for video recognition. Furthermore, we show, among other things, learned frame usage can reflect the difficulty of making prediction decisions both at instance-level within the same class and at class-level among different categories.
Zuxuan Wu, Hengduo Li, Caiming Xiong, Yu-Gang Jiang 0001, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 Deep Video Inpainting Detection
Peng Zhou 0009, Ning Yu 0006, Zuxuan Wu, Larry Davis 0001, Abhinav Shrivastava, Ser-Nam Lim
BMVC4
2021 Hierarchical Contrastive Motion Learning for Video Action Recognition
Xitong Yang, Xiaodong Yang 0001, Sifei Liu, Deqing Sun, Larry Davis 0001, Jan Kautz
BMVC5
2021 Knowledge Evolution in Neural Networks
abstract
Deep learning relies on the availability of a large corpus of data (labeled or unlabeled). Thus, one challenging unsettled question is: how to train a deep network on a relatively small dataset? To tackle this question, we propose an evolution-inspired training approach to boost performance on relatively small datasets. The knowledge evolution (KE) approach splits a deep network into two hypotheses: the fit-hypothesis and the reset-hypothesis. We iteratively evolve the knowledge inside the fit-hypothesis by perturbing the reset-hypothesis for multiple generations. This approach not only boosts performance, but also learns a slim network with a smaller inference cost. KE integrates seamlessly with both vanilla and residual convolutional networks. KE reduces both overfitting and the burden for data collection.We evaluate KE on various network architectures and loss functions. We evaluate KE using relatively small datasets (e.g., CUB-200) and randomly initialized deep net-works. KE achieves an absolute 21% improvement margin on a state-of-the-art baseline. This performance improvement is accompanied by a relative 73% reduction in inference cost. KE achieves state-of-the-art results on classification and metric learning benchmarks. Code available at http://bit.ly/3ulgwyb
Ahmed Taha 0001, Abhinav Shrivastava, Larry Davis 0001
CVPR3
2021 Efficient Object Embedding for Spliced Image Retrieval
abstract
Detecting spliced images is one of the emerging challenges in computer vision. Unlike prior methods that focus on detecting low-level artifacts generated during the manipulation process, we use an image retrieval approach to tackle this problem. When given a spliced query image, our goal is to retrieve the original image from a database of authentic images. To achieve this goal, we propose representing an image by its constituent objects based on the intuition that the finest granularity of manipulations is oftentimes at the object-level. We introduce a framework, object embeddings for spliced image retrieval (OE-SIR), that utilizes modern object detectors to localize object regions. Each region is then embedded and collectively used to represent the image. Further, we propose a student-teacher training paradigm for learning discriminative embeddings within object regions to avoid expensive multiple forward passes. Detailed analysis of the efficacy of different feature embedding models is also provided in this study. Extensive experimental results show that the OE-SIR achieves state-of-the-art performance in spliced image retrieval.
Bor-Chun Chen, Zuxuan Wu, Larry Davis 0001, Ser-Nam Lim
CVPR3
2021 SLADE: A Self-Training Framework for Distance Metric Learning
abstract
Most existing distance metric learning approaches use fully labeled data to learn the sample similarities in an embedding space. We present a self-training framework, SLADE, to improve retrieval performance by leveraging additional unlabeled data. We first train a teacher model on the labeled data and use it to generate pseudo labels for the unlabeled data. We then train a student model on both labels and pseudo labels to generate final feature embeddings. We use self-supervised representation learning to initialize the teacher model. To better deal with noisy pseudo labels generated by the teacher network, we design a new feature basis learning component for the student network, which learns basis functions of feature representations for unlabeled data. The learned basis vectors better measure the pairwise similarity and are used to select high-confident samples for training the student network. We evaluate our method on standard retrieval benchmarks: CUB-200, Cars-196 and In-shop. Experimental results demonstrate that with additional unlabeled data, our approach significantly improves the performance over the state-of-the-art methods.
Jiali Duan, Yen-Liang Lin, Son Dinh Tran, Larry Davis 0001, C.-C. Jay Kuo
CVPR4
2021 Learning Graphs for Knowledge Transfer With Limited Labels
abstract
Fixed input graphs are a mainstay in approaches that utilize Graph Convolution Networks (GCNs) for knowledge transfer. The standard paradigm is to utilize relationships in the input graph to transfer information using GCNs from training to testing nodes in the graph; for example, the semi-supervised, zero-shot, and few-shot learning setups. We propose a generalized framework for learning and improving the input graph as part of the standard GCN-based learning setup. Moreover, we use additional constraints between similar and dissimilar neighbors for each node in the graph by applying triplet loss on the intermediate layer output. We present results of semi-supervised learning on Citeseer, Cora, and Pubmed benchmarking datasets, and zero/few-shot action recognition on UCF101 and HMDB51 datasets, significantly outperforming current approaches. We also present qualitative results visualizing the graph connections that our approach learns to update.
Pallabi Ghosh, Nirat Saini, Larry Davis 0001, Abhinav Shrivastava
CVPR3
2021 The Lottery Ticket Hypothesis for Object Recognition
abstract
Recognition tasks, such as object recognition and key-point estimation, have seen widespread adoption in recent years. Most state-of-the-art methods for these tasks use deep networks that are computationally expensive and have huge memory footprints. This makes it exceedingly difficult to deploy these systems on low power embedded devices. Hence, the importance of decreasing the storage requirements and the amount of computation in such models is paramount. The recently proposed Lottery Ticket Hypothesis (LTH) states that deep neural networks trained on large datasets contain smaller subnetworks that achieve on par performance as the dense networks. In this work, we perform the first empirical study investigating LTH for model pruning in the context of object detection, instance segmentation, and keypoint estimation. Our studies reveal that lottery tickets obtained from Imagenet pretraining do not transfer well to the downstream tasks. We provide guidance on how to find lottery tickets with up to 80% overall sparsity on different sub-tasks without incurring any drop in the performance. Finally, we analyse the behavior of trained tickets with respect to various task attributes such as object size, frequency, and difficulty of detection.
Sharath Girish, Shishira R. Maiya, Kamal Gupta 0002, Hao Chen 0066, Larry Davis 0001, Abhinav Shrivastava
CVPR5
2021 2D or not 2D? Adaptive 3D Convolution Selection for Efficient Video Recognition
abstract
3D convolutional networks are prevalent for video recognition. While achieving excellent recognition performance on standard benchmarks, they operate on a sequence of frames with 3D convolutions and thus are computationally demanding. Exploiting large variations among different videos, we introduce Ada3D, a conditional computation framework that learns instance-specific 3D usage policies to determine frames and convolution layers to be used in a 3D network. These policies are derived with a two-head lightweight selection network conditioned on each input video clip. Then, only frames and convolutions that are selected by the selection network are used in the 3D model to generate predictions. The selection network is optimized with policy gradient methods to maximize a reward that encourages making correct predictions with limited computation. We conduct experiments on three video recognition benchmarks and demonstrate that our method achieves similar accuracies to state-of-the-art 3D models while requiring 20% − 50% less computation across different datasets. We also show that learned policies are transferable and Ada3D is compatible to different backbones and modern clip selection approaches. Our qualitative analysis indicates that our method allocates fewer 3D convolutions and frames for "static" inputs, yet uses more for motion-intensive clips.
Hengduo Li, Zuxuan Wu, Abhinav Shrivastava, Larry Davis 0001
CVPR4
2021 StEP: Style-Based Encoder Pre-Training for Multi-Modal Image Synthesis
abstract
We propose a novel approach for multi-modal Image-to-image (I2I) translation. To tackle the one-to-many relationship between input and output domains, previous works use complex training objectives to learn a latent embedding, jointly with the generator, that models the variability of the output domain. In contrast, we directly model the style variability of images, independent of the image synthesis task. Specifically, we pre-train a generic style encoder using a novel proxy task to learn an embedding of images, from arbitrary domains, into a low-dimensional style latent space. The learned latent space introduces several advantages over previous traditional approaches to multi-modal I2I translation. First, it is not dependent on the target dataset, and generalizes well across multiple domains. Second, it learns a more powerful and expressive latent space, which improves the fidelity of style capture and transfer. The proposed style pre-training also simplifies the training objective and speeds up the training significantly. Furthermore, we provide a detailed study of the contribution of different loss terms to the task of multi-modal I2I translation, and propose a simple alternative to VAEs to enable sampling from unconstrained latent spaces. Finally, we achieve state-of-the-art results on six challenging benchmarks with a simple training objective that includes only a GAN loss and a reconstruction loss.
Moustafa Meshry, Yixuan Ren, Larry Davis 0001, Abhinav Shrivastava
CVPR3
2021 Beyond Short Clips: End-to-End Video-Level Learning With Collaborative Memories
abstract
The standard way of training video models entails sampling at each iteration a single clip from a video and optimizing the clip prediction with respect to the video-level label. We argue that a single clip may not have enough temporal coverage to exhibit the label to recognize, since video datasets are often weakly labeled with categorical information but without dense temporal annotations. Furthermore, optimizing the model over brief clips impedes its ability to learn long-term temporal dependencies. To overcome these limitations, we introduce a collaborative memory mechanism that encodes information across multiple sampled clips of a video at each training iteration. This enables the learning of long-range dependencies beyond a single clip. We explore different design choices for the collaborative memory to ease the optimization difficulties. Our proposed framework is end-to-end trainable and significantly improves the accuracy of video classification at a negligible computational overhead. Through extensive experiments, we demonstrate that our framework generalizes to different video architectures and tasks, outperforming the state of the art on both action recognition (e.g., Kinetics-400 & 700, Charades, Something-Something-V1) and action detection (e.g., AVA v2.1 & v2.2).
Xitong Yang, Haoqi Fan 0001, Lorenzo Torresani, Larry Davis 0001
CVPR4
2021 LayoutTransformer: Layout Generation and Completion with Self-attention
abstract
We address the problem of scene layout generation for diverse domains such as images, mobile applications, documents, and 3D objects. Most complex scenes, natural or human-designed, can be expressed as a meaningful arrangement of simpler compositional graphical primitives. Generating a new layout or extending an existing layout re- quires understanding the relationships between these primitives. To do this, we propose LayoutTransformer, a novel framework that leverages self-attention to learn contextual relationships between layout elements and generate novel layouts in a given domain. Our framework allows us to generate a new layout either from an empty set or from an initial seed set of primitives, and can easily scale to support an arbitrary of primitives per layout. Furthermore, our analyses show that the model is able to automatically capture the semantic properties of the primitives. We propose simple improvements in both representation of layout primitives, as well as training methods to demonstrate competitive performance in very diverse data domains such as object bounding boxes in natural images (COCO bounding box), documents (PubLayNet), mobile applications (RICO dataset) as well as 3D shapes (Part-Net). Code and other materials will be made available at https://kampta.github.io/layout.
Kamal Gupta 0002, Justin Lazarow, Alessandro Achille, Larry Davis 0001, Vijay Mahadevan, Abhinav Shrivastava
ICCV4
2021 DiscoBox: Weakly Supervised Instance Segmentation and Semantic Correspondence from Box Supervision
abstract
We introduce DiscoBox, a novel framework that jointly learns instance segmentation and semantic correspondence using bounding box supervision. Specifically, we propose a self-ensembling framework where instance segmentation and semantic correspondence are jointly guided by a structured teacher in addition to the bounding box supervision. The teacher is a structured energy model incorporating a pairwise potential and a cross-image potential to model the pairwise pixel relationships both within and across the boxes. Minimizing the teacher energy simultaneously yields refined object masks and dense correspondences between intra-class objects, which are taken as pseudo-labels to supervise the task network and provide positive/negative correspondence pairs for dense contrastive learning. We show a symbiotic relationship where the two tasks mutually benefit from each other. Our best model achieves 37.9% AP on COCO instance segmentation, surpassing prior weakly supervised methods and is competitive to supervised methods. We also obtain state of the art weakly supervised results on PASCAL VOC12 and PF-PASCAL with real-time inference.
Shiyi Lan, Zhiding Yu, Christopher B. Choy, Subhashree Radhakrishnan, Guilin Liu, Yuke Zhu, Larry Davis 0001, Anima Anandkumar
ICCV7
2021 Learned Spatial Representations for Few-shot Talking-Head Synthesis
abstract
We propose a novel approach for few-shot talking-head synthesis. While recent works in neural talking heads have produced promising results, they can still produce images that do not preserve the identity of the subject in source images. We posit this is a result of the entangled representation of each subject in a single latent code that models 3D shape information, identity cues, colors, lighting and even background details. In contrast, we propose to factorize the representation of a subject into its spatial and style components. Our method generates a target frame in two steps. First, it predicts a discrete and dense spatial layout for the target image. Second, an image generator utilizes the predicted layout for spatial denormalization and synthesizes the target frame. We experimentally show that this disentangled representation leads to a significant improvement over previous methods, both quantitatively and qualitatively.
Moustafa Meshry, Saksham Suri, Larry Davis 0001, Abhinav Shrivastava
ICCV3
2021 Learning Realistic Human Reposing using Cyclic Self-Supervision with 3D Shape, Pose, and Appearance Consistency
abstract
Synthesizing images of a person in novel poses from a single image is a highly ambiguous task. Most existing approaches require paired training images; i.e. images of the same person with the same clothing in different poses. However, obtaining sufficiently large datasets with paired data is challenging and costly. Previous methods that forego paired supervision lack realism. We propose a self-supervised framework named SPICE (Self-supervised Person Image CrEation) that closes the image quality gap with supervised methods. The key insight enabling self-supervision is to exploit 3D information about the human body in several ways. First, the 3D body shape must remain unchanged when reposing. Second, representing body pose in 3D enables reasoning about self occlusions. Third, 3D body parts that are visible before and after reposing, should have similar appearance features. Once trained, SPICE takes an image of a person and generates a new image of that person in a new target pose. SPICE achieves state-of-the-art performance on the DeepFashion dataset, improving the FID score from 29.9 to 7.8 compared with previous unsupervised methods, and with performance similar to the state-of-the-art supervised method (6.4). SPICE also generates temporally coherent videos given an input image and a sequence of poses, despite being trained on static images only.
Soubhik Sanyal, Betty J. Mohler, Alex Vorobiov, Larry Davis 0001, Timo Bolkart, Javier Romero 0002, Matthew Loper, Michael J. Black
ICCV4
2021 Dual Contrastive Loss and Attention for GANs
abstract
Generative Adversarial Networks (GANs) produce impressive results on unconditional image generation when powered with large-scale image datasets. Yet generated images are still easy to spot especially on datasets with high variance (e.g. bedroom, church). In this paper, we propose various improvements to further push the boundaries in image generation. Specifically, we propose a novel dual contrastive loss and show that, with this loss, discriminator learns more generalized and distinguishable representations to incentivize generation. In addition, we revisit attention and extensively experiment with different attention blocks in the generator. We find attention to be still an important module for successful image generation even though it was not used in the recent state-of-the-art models. Lastly, we study different attention architectures in the discriminator, and propose a reference attention mechanism. By combining the strengths of these remedies, we improve the compelling state-of-the-art Fréchet Inception Distance (FID) by at least 17.5% on several benchmark datasets. We obtain even more significant improvements on compositional synthetic scenes (up to 47.5% in FID).
Ning Yu 0006, Guilin Liu, Aysegul Dundar, Andrew Tao, Bryan Catanzaro, Larry Davis 0001, Mario Fritz
ICCV6
2021 VideoLT: Large-scale Long-tailed Video Recognition
abstract
Label distributions in real-world are oftentimes long-tailed and imbalanced, resulting in biased models towards dominant labels. While long-tailed recognition has been extensively studied for image classification tasks, limited effort has been made for the video domain. In this paper, we introduce VideoLT, a large-scale long-tailed video recognition dataset, as a step toward real-world video recognition. VideoLT contains 256,218 untrimmed videos, annotated into 1,004 classes with a long-tailed distribution. Through extensive studies, we demonstrate that state-of-the-art methods used for long-tailed image recognition do not perform well in the video domain due to the additional temporal dimension in videos. This motivates us to propose FrameStack, a simple yet effective method for long-tailed video recognition. In particular, FrameStack performs sampling at the frame-level in order to balance class distributions, and the sampling ratio is dynamically determined using knowledge derived from the network during training. Experimental results demonstrate that FrameStack can improve classification performance without sacrificing the overall accuracy. Code and dataset are available at: https://github.com/17Skye17/VideoLT.
Xing Zhang 0013, Zuxuan Wu, Zejia Weng, Huazhu Fu, Jingjing Chen 0001, Yu-Gang Jiang 0001, Larry Davis 0001
ICCV7
2021 PSANet - subspace attention for personalized compatibility
abstract
Recommending sets of items that include both personalized and compatible items is crucial to personalized styling programs such as Amazon’s Personal Shopper. There is both an extensive literature on learning generic fashion compatibility and also on personalization in fashion. However, recommending pairs of items that the customer would like to wear together is still less studied as it involves learning a compatibility metric personalized to each customer. We propose a new framework (PSA-Net) to learn compatibility that is personalized to the customer - a customer dependent subspace learning framework where attention weights of subspaces are learnt using customer representations. We evaluate our approach on compatibility data provided directly by customers. Our approach outperforms the non-personalized approach in predicting compatibility preferences of customers. In other words, an approach that learns a common compatibility metric for all customers. In addition, we compare the significance of feedback collected directly from customers to that of data collected from human stylists in predicting compatibility for Amazon customers.
Meet Taraviya, Anurag Beniwal, Yen-Liang Lin, Larry Davis 0001
ICDM4
2021 A Coarse-to-Fine Framework for Resource Efficient Video Recognition
Zuxuan Wu, Hengduo Li, Yingbin Zheng, Caiming Xiong, Yu-Gang Jiang 0001, Larry Davis 0001
Int. J. Comput. Vis.6
2020 Depth Completion Using a View-constrained Deep Prior
abstract
Recent work has shown that the structure of convolutional neural networks (CNNs) induces a strong prior that favors natural images. This prior, known as a deep image prior (DIP), is an effective regularizer in inverse problems such as image denoising and inpainting. We extend the concept of the DIP to depth images. Given color images and noisy and incomplete target depth maps, we optimize a randomly-initialized CNN model to reconstruct a depth map restored by virtue of using the CNN network structure as a prior combined with a view-constrained photo-consistency loss. This loss is computed using images from a geometrically calibrated camera from nearby viewpoints. We apply this deep depth prior for inpainting and refining incomplete and noisy depth maps within both binocular and multi-view stereo pipelines. Our quantitative and qualitative evaluation shows that our refined depth maps are more accurate and complete, and after fusion, produces dense 3D models of higher quality.
Pallabi Ghosh, Vibhav Vineet, Larry Davis 0001, Abhinav Shrivastava, Sudipta N. Sinha, Neel Joshi
3DV3
2020 Recognizing Instagram Filtered Images with Feature De-Stylization
abstract
Deep neural networks have been shown to suffer from poor generalization when small perturbations are added (like Gaussian noise), yet little work has been done to evaluate their robustness to more natural image transformations like photo filters. This paper presents a study on how popular pretrained models are affected by commonly used Instagram filters. To this end, we introduce ImageNet-Instagram, a filtered version of ImageNet, where 20 popular Instagram filters are applied to each image in ImageNet. Our analysis suggests that simple structure preserving filters which only alter the global appearance of an image can lead to large differences in the convolutional feature space. To improve generalization, we introduce a lightweight de-stylization module that predicts parameters used for scaling and shifting feature maps to “undo” the changes incurred by filters, inverting the process of style transfer tasks. We further demonstrate the module can be readily plugged into modern CNN architectures together with skip connections. We conduct extensive studies on ImageNet-Instagram, and show quantitatively and qualitatively, that the proposed module, among other things, can effectively improve generalization by simply learning normalization parameters without retraining the entire network, thus recovering the alterations in the feature space caused by the filters.
Zhe Wu 0001, Zuxuan Wu, Larry Davis 0001
AAAI4
2020 Universal Adversarial Training
abstract
Standard adversarial attacks change the predicted class label of a selected image by adding specially tailored small perturbations to its pixels. In contrast, a universal perturbation is an update that can be added to any image in a broad class of images, while still changing the predicted class label. We study the efficient generation of universal adversarial perturbations, and also efficient methods for hardening networks to these attacks. We propose a simple optimization-based universal attack that reduces the top-1 accuracy of various network architectures on ImageNet to less than 20%, while learning the universal perturbation 13× faster than the standard method.To defend against these perturbations, we propose universal adversarial training, which models the problem of robust classifier generation as a two-player min-max game, and produces robust models with only 2× the cost of natural training. We also propose a simultaneous stochastic gradient method that is almost free of extra computation, which allows us to do universal adversarial training on ImageNet.
Ali Shafahi, Mahyar Najibi, Zheng Xu 0002, John Dickerson 0001, Larry Davis 0001, Tom Goldstein
AAAI5
2020 Generate, Segment, and Refine: Towards Generic Manipulation Segmentation
abstract
Detecting manipulated images has become a significant emerging challenge. The advent of image sharing platforms and the easy availability of advanced photo editing software have resulted in a large quantities of manipulated images being shared on the internet. While the intent behind such manipulations varies widely, concerns on the spread of false news and misinformation is growing. Current state of the art methods for detecting these manipulated images suffers from the lack of training data due to the laborious labeling process. We address this problem in this paper, for which we introduce a manipulated image generation process that creates true positives using currently available datasets. Drawing from traditional work on image blending, we propose a novel generator for creating such examples. In addition, we also propose to further create examples that force the algorithm to focus on boundary artifacts during training. Strong experimental results validate our proposal.
Peng Zhou 0009, Bor-Chun Chen, Xintong Han, Mahyar Najibi, Abhinav Shrivastava, Ser-Nam Lim, Larry Davis 0001
AAAI7
2020 M2KD: Incremental Learning via Multi-model and Multi-level Knowledge Distillation
Peng Zhou 0009, Long Mai, Jianming Zhang 0001, Ning Xu 0007, Zuxuan Wu, Larry Davis 0001
BMVC6
2020 SaccadeNet: A Fast and Accurate Object Detector
abstract
Object detection is an essential step towards holistic scene understanding. Most existing object detection algorithms attend to certain object areas once and then predict the object locations. However, scientists have revealed that human do not look at the scene in fixed steadiness. Instead, human eyes move around, locating informative parts to understand the object location. This active perceiving movement process is called saccade. In this paper, inspired by such mechanism, we propose a fast and accurate object detector called SaccadeNet. It contains four main modules, the Center Attentive Module, the Corner Attentive Module, the Attention Transitive Module, and the Aggregation Attentive Module, which allows it to attend to different informative object keypoints actively, and predict object locations from coarse to fine. The Corner Attentive Module is used only during training to extract more informative corner features which brings free-lunch performance boost. On the MS COCO dataset, we achieve the performance of 40.4% mAP at 28 FPS and 30.5% mAP at 118 FPS. Among all the real-time object detectors, our SaccadeNet achieves the best detection performance, which demonstrates the effectiveness of the proposed detection mechanism.
Shiyi Lan, Zhou Ren, Larry Davis 0001, Gang Hua 0001
CVPR4
2020 Learning From Noisy Anchors for One-Stage Object Detection
abstract
State-of-the-art object detectors rely on regressing and classifying an extensive list of possible anchors, which are divided into positive and negative samples based on their intersection-over-union (IoU) with corresponding ground-truth objects. Such a harsh split conditioned on IoU results in binary labels that are potentially noisy and challenging for training. In this paper, we propose to mitigate noise incurred by imperfect label assignment such that the contributions of anchors are dynamically determined by a carefully constructed cleanliness score associated with each anchor. Exploring outputs from both regression and classification branches, the cleanliness scores, estimated without incurring any additional computational overhead, are used not only as soft labels to supervise the training of the classification branch but also sample re-weighting factors for improved localization and classification accuracy. We conduct extensive experiments on COCO, and demonstrate, among other things, the proposed approach steadily improves RetinaNet by ~2% with various backbones.
Hengduo Li, Zuxuan Wu, Chen Zhu 0001, Caiming Xiong, Richard Socher, Larry Davis 0001
CVPR6
2020 Fashion Outfit Complementary Item Retrieval
abstract
Complementary fashion item recommendation is critical for fashion outfit completion. Existing methods mainly focus on outfit compatibility prediction but not in a retrieval setting. We propose a new framework for outfit complementary item retrieval. Specifically, a category-based subspace attention network is presented, which is a scalable approach for learning the subspace attentions. In addition, we introduce an outfit ranking loss that better models the item relationships of an entire outfit. We evaluate our method on the outfit compatibility, FITB and new retrieval tasks. Experimental results demonstrate that our approach outperforms state-of-the-art methods in both compatibility prediction and complementary item retrieval.
Yen-Liang Lin, Son Dinh Tran, Larry Davis 0001
CVPR3
2020 DOPS: Learning to Detect 3D Objects and Predict Their 3D Shapes
abstract
We propose DOPS, a fast single-stage 3D object detection method for LIDAR data. Previous methods often make domain-specific design decisions, for example projecting points into a bird-eye view image in autonomous driving scenarios. In contrast, we propose a general-purpose method that works on both indoor and outdoor scenes. The core novelty of our method is a fast, single-pass architecture that both detects objects in 3D and estimates their shapes. 3D bounding box parameters are estimated in one pass for every point, aggregated through graph convolutions, and fed into a branch of the network that predicts latent codes representing the shape of each detected object. The latent shape space and shape decoder are learned on a synthetic dataset and then used as supervision for the end-to-end training of the 3D object detection pipeline. Thus our model is able to extract shapes without access to ground-truth shape information in the target dataset. During experiments, we find that our proposed method achieves state-of-the-art results by~5% on object detection in ScanNet scenes, and it gets top results by 3.4% in the Waymo Open Dataset, while reproducing the shapes of detected cars.
Mahyar Najibi, Guangda Lai, Abhijit Kundu, Zhichao Lu, Vivek Rathod, Thomas A. Funkhouser, Caroline Pantofaru, David A. Ross, Larry Davis 0001, Alireza Fathi
CVPR9
2020 Deepstrip: High-Resolution Boundary Refinement
abstract
In this paper, we target refining the boundaries in high resolution images given low resolution masks. For memory and computation efficiency, we propose to convert the regions of interest into strip images and compute a boundary prediction in the strip domain. To detect the target boundary, we present a framework with two prediction layers. First, all potential boundaries are predicted as an initial prediction and then a selection layer is used to pick the target boundary and smooth the result. To encourage accurate prediction, a loss which measures the boundary distance in strip domain is introduced. In addition, we enforce a matching consistency and C0 continuity regularization to the network to reduce false alarms. Extensive experiments on both public and a newly created high resolution dataset strongly validate our approach.
Peng Zhou 0009, Brian L. Price, Scott Cohen, Gregg Wilensky, Larry Davis 0001
CVPR5
2020 Quantization Guided JPEG Artifact Correction
Max Ehrlich, Larry Davis 0001, Ser-Nam Lim, Abhinav Shrivastava
ECCV (8)2
2020 Consistency-Based Semi-supervised Active Learning: Towards Minimizing Labeling Cost
Mingfei Gao, Sercan Ö. Arik, Larry Davis 0001, Tomas Pfister
ECCV (10)5
2020 A Generic Visualization Approach for Convolutional Neural Networks
Ahmed Taha 0001, Xitong Yang, Abhinav Shrivastava, Larry Davis 0001
ECCV (17)4
2020 InfoFocus: 3D Object Detection for Autonomous Driving with Dynamic Information Modeling
Jun Wang 0090, Shiyi Lan, Mingfei Gao, Larry Davis 0001
ECCV (10)4
2020 Making an Invisibility Cloak: Real World Adversarial Attacks on Object Detectors
Zuxuan Wu, Ser-Nam Lim, Larry Davis 0001, Tom Goldstein
ECCV (4)3
2020 Inclusive GAN: Improving Data and Minority Coverage in Generative Models
Ning Yu 0006, Ke Li 0011, Peng Zhou 0009, Jitendra Malik, Larry Davis 0001, Mario Fritz
ECCV (22)5
2020 Stacked Spatio-Temporal Graph Convolutional Networks for Action Segmentation
abstract
We propose novel Stacked Spatio-Temporal Graph Convolutional Networks (Stacked-STGCN) for action segmentation, i.e., predicting and localizing a sequence of actions over long videos. We extend the Spatio-Temporal Graph Convolutional Network (STGCN) originally proposed for skeleton-based action recognition to enable nodes with different characteristics (e.g., scene, actor, object, action), feature descriptors with varied lengths, and arbitrary temporal edge connections to account for large graph deformation commonly associated with complex activities. We further introduce the stacked hourglass architecture to STGCN to leverage the advantages of an encoder-decoder design for improved generalization performance and localization accuracy. We explore various descriptors such as frame- level VGG, segment-level I3D, RCNN-based object, etc. as node descriptors to enable action segmentation based on joint inference over comprehensive contextual information. We show results on CAD120 (which provides pre-computed node features and edge weights for fair performance comparison across algorithms) as well as a more complex real- world activity dataset, Charades. Our Stacked-STGCN in general achieves improved performance over the state-of- the-art for both CAD120 and Charades. Moreover, due to its generic design, Stacked-STGCN can be applied to a wider range of applications that require structured inference over long sequences with heterogeneous data types and varied temporal extent.
Pallabi Ghosh, Larry Davis 0001, Ajay Divakaran
WACV3
2020 Boosting Standard Classification Architectures Through a Ranking Regularizer
abstract
We employ triplet loss as a feature embedding regularizer to boost classification performance. Standard architectures, like ResNet and Inception, are extended to support both losses with minimal hyper-parameter tuning. This promotes generality while fine-tuning pretrained networks. Triplet loss is a powerful surrogate for recently proposed embedding regularizers. Yet, it is avoided due to large batch-size requirement and high computational cost. Through our experiments, we re-assess these assumptions.During inference, our network supports both classification and embedding tasks without any computational overhead. Quantitative evaluation highlights a steady improvement on five fine-grained recognition datasets. Further evaluation on an imbalanced video dataset achieves significant improvement. Triplet loss brings feature embedding capabilities like nearest neighbor to classification models. Code available at http://bit.ly/2LNYEqL.
Ahmed Taha 0001, Yi-Ting Chen 0001, Teruhisa Misu, Abhinav Shrivastava, Larry Davis 0001
WACV5
2020 A Generic Improvement to Deep Residual Networks Based on Gradient Flow
abstract
Preactivation ResNets consistently outperforms the original postactivation ResNets on the CIFAR10/100 classification benchmark. However, these results surprisingly do not carry over to the standard ImageNet benchmark. First, we theoretically analyze this incongruity in terms of how the two variants differ in handling the propagation of gradients. Although identity shortcuts are critical in both variants for improving optimization and performance, we show that postactivation variants enable early layers to receive a diverse dynamic composition of gradients from effectively deeper paths in comparison to preactivation variants, enabling the network to make maximal use of its representational capacity. Second, we show that downsampling projections (while only a few in number) have a significantly detrimental effect on performance. We show that by simply replacing downsampling projections with identitylike dense-reshape shortcuts, the classification results of standard residual architectures such as ResNets, ResNeXts, and SE-Nets improve by up to 1.2% on ImageNet, without any increase in computational complexity (FLOPs).
Venkataraman Santhanam, Larry Davis 0001
IEEE Trans. Neural Networks Learn. Syst.2
2019 Soft Sampling for Robust Object Detection
Zhe Wu 0001, Navaneeth Bodla, Mahyar Najibi, Rama Chellappa, Larry Davis 0001
BMVC6
2019 Modeling Local Geometric Structure of 3D Point Clouds Using Geo-CNN
abstract
Recent advances in deep convolutional neural networks (CNNs) have motivated researchers to adapt CNNs to directly model points in 3D point clouds. Modeling local structure has been proven to be important for the success of convolutional architectures, and researchers exploited the modeling of local point sets in the feature extraction hierarchy. However, limited attention has been paid to explicitly model the geometric structure amongst points in a local region. To address this problem, we propose Geo-CNN, which applies a generic convolution-like operation dubbed as GeoConv to each point and its local neighborhood. Local geometric relationships among points are captured when extracting edge features between the center and its neighboring points. We first decompose the edge feature extraction process onto three orthogonal bases, and then aggregate the extracted features based on the angles between the edge vector and the bases. This encourages the network to preserve the geometric structure in Euclidean space throughout the feature extraction hierarchy. GeoConv is a generic and efficient operation that can be easily integrated into 3D point cloud analysis pipelines for multiple applications. We evaluate Geo-CNN on ModelNet40 and KITTI and achieve state-of-the-art performance.
Shiyi Lan, Ruichi Yu, Gang Yu 0002, Larry Davis 0001
CVPR4
2019 Explicit Bias Discovery in Visual Question Answering Models
abstract
Researchers have observed that Visual Question Answering (VQA ) models tend to answer questions by learning statistical biases in the data. For example, their answer to the question “What is the color of the grass?” is usually “Green”, whereas a question like “What is the title of the book?” cannot be answered by inferring statistical biases. It is of interest to the community to explicitly discover such biases, both for understanding the behavior of such models, and towards debugging them. Our work address this problem. In a database, we store the words of the question, answer and visual words corresponding to regions of interest in attention maps. By running simple rule mining algorithms on this database, we discover human-interpretable rules which give us unique insight into the behavior of such models. Our results also show examples of unusual behaviors learned by models in attempting VQA tasks.
Varun Manjunatha, Nirat Saini, Larry Davis 0001
CVPR3
2019 FA-RPN: Floating Region Proposals for Face Detection
abstract
We propose a novel approach for generating region proposals for performing face detection. Instead of classifying anchor boxes using features from a pixel in the convolutional feature map, we adopt a pooling-based approach for generating region proposals. However, pooling hundreds of thousands of anchors which are evaluated for generating proposals becomes a computational bottleneck during inference. To this end, an efficient anchor placement strategy for reducing the number of anchor-boxes is proposed. We then show that proposals generated by our network (Floating Anchor Region Proposal Network, FA-RPN) are better than RPN for generating region proposals for face detection. We discuss several beneficial features of FA-RPN proposals (which can be enabled without re-training) like iterative refinement, placement of fractional anchors and changing size/shape of anchors. Our face detector based on FA-RPN obtains 89.4% mAP with a ResNet-50 backbone on the WIDER dataset.
Mahyar Najibi, Larry Davis 0001
CVPR3
2019 AdaFrame: Adaptive Frame Selection for Fast Video Recognition
abstract
We present AdaFrame, a framework that adaptively selects relevant frames on a per-input basis for fast video recognition. AdaFrame contains a Long Short-Term Memory network augmented with a global memory that provides context information for searching which frames to use over time. Trained with policy gradient methods, AdaFrame generates a prediction, determines which frame to observe next, and computes the utility, i.e., expected future rewards, of seeing more frames at each time step. At testing time, AdaFrame exploits predicted utilities to achieve adaptive lookahead inference such that the overall computational costs are reduced without incurring a decrease in accuracy. Extensive experiments are conducted on two large-scale video benchmarks, FCVID and ActivityNet. AdaFrame matches the performance of using all frames with only 8.21 and 8.65 frames on FCVID and ActivityNet, respectively. We further qualitatively demonstrate learned frame usage can indicate the difficulty of making classification decisions; easier samples need fewer frames while harder ones require more, both at instance-level within the same class and at class-level among different categories.
Zuxuan Wu, Caiming Xiong, Chih-Yao Ma, Richard Socher, Larry Davis 0001
CVPR5
2019 STEP: Spatio-Temporal Progressive Learning for Video Action Detection
abstract
In this paper, we propose Spatio-TEmporal Progressive (STEP) action detector-a progressive learning framework for spatio-temporal action detection in videos. Starting from a handful of coarse-scale proposal cuboids, our approach progressively refines the proposals towards actions over a few steps. In this way, high-quality proposals (i.e., adhere to action movements) can be gradually obtained at later steps by leveraging the regression outputs from previous steps. At each step, we adaptively extend the proposals in time to incorporate more related temporal context. Compared to the prior work that performs action detection in one run, our progressive learning framework is able to naturally handle the spatial displacement within action tubes and therefore provides a more effective way for spatio-temporal modeling. We extensively evaluate our approach on UCF101 and AVA, and demonstrate superior detection results. Remarkably, we achieve mAP of 75.0% and 18.6% on the two datasets with 3 progressive steps and using respectively only 11 and 34 initial proposals.
Xitong Yang, Xiaodong Yang 0001, Ming-Yu Liu 0001, Fanyi Xiao, Larry Davis 0001, Jan Kautz
CVPR5
2019 MAN: Moment Alignment Network for Natural Language Moment Retrieval via Iterative Graph Adjustment
abstract
This research strives for natural language moment retrieval in long, untrimmed video streams. The problem is not trivial especially when a video contains multiple moments of interests and the language describes complex temporal dependencies, which often happens in real scenarios. We identify two crucial challenges: semantic misalignment and structural misalignment. However, existing approaches treat different moments separately and do not explicitly model complex moment-wise temporal relations. In this paper, we present Moment Alignment Network (MAN), a novel framework that unifies the candidate moment encoding and temporal structural reasoning in a single-shot feed-forward network. MAN naturally assigns candidate moment representations aligned with language semantics over different temporal locations and scales. Most importantly, we propose to explicitly model moment-wise temporal relations as a structured graph and devise an iterative graph adjustment network to jointly learn the best structure in an end-to-end manner. We evaluate the proposed approach on two challenging public benchmarks DiDeMo and Charades-STA, where our MAN significantly outperforms the state-of-the-art by a large margin.
Da Zhang 0001, Xiyang Dai, Xin Wang 0061, Yuan-Fang Wang, Larry Davis 0001
CVPR5
2019 WSLLN: Weakly Supervised Natural Language Localization Networks
abstract
Mingfei Gao, Larry Davis, Richard Socher, Caiming Xiong. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Mingfei Gao, Larry Davis 0001, Richard Socher, Caiming Xiong
EMNLP/IJCNLP (1)2
2019 Boundary-sensitive Network for Portrait Segmentation
abstract
Portrait segmentation has gained more and more attractions in recent years due to the popularity of selfie images. Compared to general semantic segmentation problems, portrait segmentation focuses on facial areas with higher requirements especially over the boundaries. To improve the performance of portrait segmentation, we propose a boundary-sensitive deep neural network (BSN) for better accuracy among the portrait boundaries. BSN introduces three novel techniques. First, an individual boundary-sensitive mask is proposed by dilating the contour line and assigning the boundary pixels with multi-class labels. Second, a global boundary-sensitive mask is employed as a position sensitive prior to further constrain the overall shape of the segmentation map. Third, we train a boundary-sensitive attribute classifier jointly with the segmentation network to reinforce the network with semantic boundary shape information. We have evaluated BSN on the state-of-the-art public portrait segmentation datasets, i.e., the PFCN dataset, as well as the portrait images collected from other three popular image segmentation datasets: COCO, COCO-Stuff, and PASCAL VOC. Our method achieves the superior quantitative and qualitative performance over state-of-the-arts on the evaluated datasets, especially obtains better visualization effect on the portrait boundary region.
Xianzhi Du, Xiaolong Wang 0006, Dawei Li 0006, Serafettin Tasci, Cameron Upright, Stephen Walsh, Larry Davis 0001
FG8
2019 Deep Residual Learning in the JPEG Transform Domain
abstract
We introduce a general method of performing Residual Network inference and learning in the JPEG transform domain that allows the network to consume compressed images as input. Our formulation leverages the linearity of the JPEG transform to redefine convolution and batch normalization with a tune-able numerical approximation for ReLu. The result is mathematically equivalent to the spatial domain network up to the ReLu approximation accuracy. A formulation for image classification and a model conversion algorithm for spatial domain networks are given as examples of the method. We show that the sparsity of the JPEG format allows for faster processing of images with little to no penalty in the network accuracy.
Max Ehrlich, Larry Davis 0001
ICCV2
2019 StartNet: Online Detection of Action Start in Untrimmed Videos
abstract
We propose StartNet to address Online Detection of Action Start (ODAS) where action starts and their associated categories are detected in untrimmed, streaming videos. Previous methods aim to localize action starts by learning feature representations that can directly separate the start point from its preceding background. It is challenging due to the subtle appearance difference near the action starts and the lack of training data. Instead, StartNet decomposes ODAS into two stages: action classification (using ClsNet) and start point localization (using LocNet). ClsNet focuses on per-frame labeling and predicts action score distributions online. Based on the predicted action scores of the past and current frames, LocNet conducts class-agnostic start detection by optimizing long-term localization rewards using policy gradient methods. The proposed framework is validated on two large-scale datasets, THUMOS'14 and ActivityNet. The experimental results show that StartNet significantly outperforms the state-of-the-art by 15%-30% p-mAP under the offset tolerance of 1-10 seconds on THUMOS'14, and achieves comparable performance on ActivityNet with 10 times smaller time offset.
Mingfei Gao, Larry Davis 0001, Richard Socher, Caiming Xiong
ICCV3
2019 FiNet: Compatible and Diverse Fashion Image Inpainting
abstract
Visual compatibility is critical for fashion analysis, yet is missing in existing fashion image synthesis systems. In this paper, we propose to explicitly model visual compatibility through fashion image inpainting. We present Fashion Inpainting Networks (FiNet), a two-stage image-to-image generation framework that is able to perform compatible and diverse inpainting. Disentangling the generation of shape and appearance to ensure photorealistic results, our framework consists of a shape generation network and an appearance generation network. More importantly, for each generation network, we introduce two encoders interacting with one another to learn latent codes in a shared compatibility space. The latent representations are jointly optimized with the corresponding generation network to condition the synthesis process, encouraging a diverse set of generated results that are visually compatible with existing fashion garments. In addition, our framework is readily extended to clothing reconstruction and fashion transfer. Extensive experiments on fashion synthesis quantitatively and qualitatively demonstrate the effectiveness of our method.
Xintong Han, Zuxuan Wu, Matthew R. Scott, Larry Davis 0001
ICCV5
2019 Cross-X Learning for Fine-Grained Visual Categorization
abstract
Recognizing objects from subcategories with very subtle differences remains a challenging task due to the large intra-class and small inter-class variation. Recent work tackles this problem in a weakly-supervised manner: object parts are first detected and the corresponding part-specific features are extracted for fine-grained classification. However, these methods typically treat the part-specific features of each image in isolation while neglecting their relationships between different images. In this paper, we propose Cross-X learning, a simple yet effective approach that exploits the relationships between different images and between different network layers for robust multi-scale feature learning. Our approach involves two novel components: (i) a cross-category cross-semantic regularizer that guides the extracted features to represent semantic parts and, (ii) a cross-layer regularizer that improves the robustness of multi-scale features by matching the prediction distribution across multiple layers. Our approach can be easily trained end-to-end and is scalable to large datasets like NABirds. We empirically analyze the contributions of different components of our approach and demonstrate its robustness, effectiveness and state-of-the-art performance on five benchmark datasets. Code is available at \url{https://github.com/cswluo/CrossX}.
Wei Luo 0006, Xitong Yang, Xianjie Mo, Yuheng Lu, Larry Davis 0001, Jun Li 0027, Jian Yang 0003, Ser-Nam Lim
ICCV5
2019 AutoFocus: Efficient Multi-Scale Inference
abstract
This paper describes AutoFocus, an efficient multi-scale inference algorithm for deep-learning based object detectors. Instead of processing an entire image pyramid, AutoFocus adopts a coarse to fine approach and only processes regions which are likely to contain small objects at finer scales. This is achieved by predicting category agnostic segmentation maps for small objects at coarser scales, called FocusPixels. FocusPixels can be predicted with high recall, and in many cases, they only cover a small fraction of the entire image. To make efficient use of FocusPixels, an algorithm is proposed which generates compact rectangular FocusChips which enclose FocusPixels. The detector is only applied inside FocusChips, which reduces computation while processing finer scales. Different types of error can arise when detections from FocusChips of multiple scales are combined, hence techniques to correct them are proposed. AutoFocus obtains an mAP of 47.9% (68.3% at 50% overlap) on the COCO test-dev set while processing 6.4 images per second on a Titan X (Pascal) GPU. This is 2.5× faster than our multi-scale baseline detector and matches its mAP. The number ofpixels processed in the pyramid can be reduced by 5× with a 1% drop in mAP. AutoFocus obtains more than 10% mAP gain compared to RetinaNet but runs at the same speed with the same ResNet-101 backbone.
Mahyar Najibi, Larry Davis 0001
ICCV3
2019 ACE: Adapting to Changing Environments for Semantic Segmentation
abstract
Deep neural networks exhibit exceptional accuracy when they are trained and tested on the same data distributions. However, neural classifiers are often extremely brittle when confronted with domain shift---changes in the input distribution that occur over time. We present ACE, a framework for semantic segmentation that dynamically adapts to changing environments over time. By aligning the distribution of labeled training data from the original source domain with the distribution of incoming data in a shifted domain, ACE synthesizes labeled training data for environments as it sees them. This stylized data is then used to update a segmentation model so that it performs well in new environments. To avoid forgetting knowledge from past environments, we introduce a memory that stores feature statistics from previously seen domains. These statistics can be used to replay images in any of the previously observed domains, thus preventing catastrophic forgetting. In addition to standard batch training using stochastic gradient decent (SGD), we also experiment with fast adaptation methods based on adaptive meta-learning. Extensive experiments are conducted on two datasets from SYNTHIA, the results demonstrate the effectiveness of the proposed approach when adapting to a number of tasks.
Zuxuan Wu, Xin Wang 0066, Joseph Gonzalez 0001, Tom Goldstein, Larry Davis 0001
ICCV5
2019 Temporal Recurrent Networks for Online Action Detection
abstract
Most work on temporal action detection is formulated as an offline problem, in which the start and end times of actions are determined after the entire video is fully observed. However, important real-time applications including surveillance and driver assistance systems require identifying actions as soon as each video frame arrives, based only on current and historical observations. In this paper, we propose a novel framework, the Temporal Recurrent Network (TRN), to model greater temporal context of each frame by simultaneously performing online action detection and anticipation of the immediate future. At each moment in time, our approach makes use of both accumulated historical evidence and predicted future information to better recognize the action that is currently occurring, and integrates both of these into a unified end-to-end architecture. We evaluate our approach on two popular online action detection datasets, HDD and TVSeries, as well as another widely used dataset, THUMOS'14. The results show that TRN significantly outperforms the state-of-the-art.
Mingfei Gao, Yi-Ting Chen 0001, Larry Davis 0001, David Crandall
ICCV4
2019 Attributing Fake Images to GANs: Learning and Analyzing GAN Fingerprints
abstract
Recent advances in Generative Adversarial Networks (GANs) have shown increasing success in generating photorealistic images. But they also raise challenges to visual forensics and model attribution. We present the first study of learning GAN fingerprints towards image attribution and using them to classify an image as real or GAN-generated. For GAN-generated images, we further identify their sources. Our experiments show that (1) GANs carry distinct model fingerprints and leave stable fingerprints in their generated images, which support image attribution; (2) even minor differences in GAN training can result in different fingerprints, which enables fine-grained model authentication; (3) fingerprints persist across different image frequencies and patches and are not biased by GAN artifacts; (4) fingerprint finetuning is effective in immunizing against five types of adversarial image perturbations; and (5) comparisons also show our learned fingerprints consistently outperform several baselines in a variety of setups.
Ning Yu 0006, Larry Davis 0001, Mario Fritz
ICCV2
2019 Layout-Induced Video Representation for Recognizing Agent-in-Place Actions
abstract
We address scene layout modeling for recognizing agent-in-place actions, which are actions associated with agents who perform them and the places where they occur, in the context of outdoor home surveillance. We introduce a novel representation to model the geometry and topology of scene layouts so that a network can generalize from the layouts observed in the training scenes to unseen scenes in the test set. This Layout-Induced Video Representation (LIVR) abstracts away low-level appearance variance and encodes geometric and topological relationships of places to explicitly model scene layout. LIVR partitions the semantic features of a scene into different places to force the network to learn generic place-based feature descriptions which are independent of specific scene layouts; then, LIVR dynamically aggregates features based on connectivities of places in each specific scene to model its layout. We introduce a new Agent-in-Place Action (APA) dataset to show that our method allows neural network models to generalize significantly better to unseen scenes.
Ruichi Yu, Ang Li 0001, Jingxiao Zheng, Vlad I. Morariu, Larry Davis 0001
ICCV6
2019 Unsupervised Super-Resolution of Satellite Imagery for High Fidelity Material Label Transfer
abstract
Urban material recognition in remote sensing imagery is a challenging problem due to the difficulty of obtaining human annotations, especially on low resolution satellite images. To this end, we propose an unsupervised domain adaptation-based approach using adversarial learning. We aim to harvest information from smaller quantities of high resolution data (source domain) and utilize the same to super-resolve low resolution imagery (target domain). This can potentially aid in semantic as well as material label transfer from a richly annotated source to a target domain.
Arthita Ghosh, Max Ehrlich, Larry Davis 0001, Rama Chellappa
IGARSS3
2019 Adversarial training for free!
abstract
Adversarial training, in which a network is trained on adversarial examples, is one of the few defenses against adversarial attacks that withstands strong attacks. Unfortunately, the high cost of generating strong adversarial examples makes standard adversarial training impractical on large-scale problems like ImageNet. We present an algorithm that eliminates the overhead cost of generating adversarial examples by recycling the gradient information computed when updating model parameters. Our "free" adversarial training algorithm achieves comparable robustness to PGD adversarial training on the CIFAR-10 and CIFAR-100 datasets at negligible additional cost compared to natural training, and can be 7 to 30 times faster than other strong adversarial training methods. Using a single workstation with 4 P100 GPUs and 2 days of runtime, we can train a robust model for the large-scale ImageNet classification task that maintains 40% accuracy against PGD attacks.
Ali Shafahi, Mahyar Najibi, Amin Ghiasi, Zheng Xu 0002, John Dickerson 0001, Christoph Studer, Larry Davis 0001, Gavin Taylor, Tom Goldstein
NeurIPS7
2019 LiteEval: A Coarse-to-Fine Framework for Resource Efficient Video Recognition
abstract
This paper presents LiteEval, a simple yet effective coarse-to-fine framework for resource efficient video recognition, suitable for both online and offline scenarios. Exploiting decent yet computationally efficient features derived at a coarse scale with a lightweight CNN model, LiteEval dynamically decides on-the-fly whether to compute more powerful features for incoming video frames at a finer scale to obtain more details. This is achieved by a coarse LSTM and a fine LSTM operating cooperatively, as well as a conditional gating module to learn when to allocate more computation. Extensive experiments are conducted on two large-scale video benchmarks, FCVID and ActivityNet, and the results demonstrate LiteEval requires substantially less computation while offering excellent classification accuracy for both online and offline predictions.
Zuxuan Wu, Caiming Xiong, Yu-Gang Jiang 0001, Larry Davis 0001
NeurIPS4
2019 TAN: Temporal Aggregation Network for Dense Multi-Label Action Recognition
abstract
We present Temporal Aggregation Network (TAN) which decomposes 3D convolutions into spatial and temporal aggregation blocks. By stacking spatial and temporal convolutions repeatedly, TAN forms a deep hierarchical representation for capturing spatio-temporal information in videos. Since we do not apply 3D convolutions in each layer but only apply temporal aggregation blocks once after each spatial downsampling layer in the network, we significantly reduce the model complexity. The use of dilated convolutions at different resolutions of the network helps in aggregating multi-scale spatio-temporal information efficiently. Experiments show that our model is well suited for dense multi-label action recognition, which is a challenging sub-topic of action recognition that requires predicting multiple action labels in each frame. We outperform state-of-the-art methods by 5% and 3% on the Charades and Multi-THUMOS dataset respectively.
Xiyang Dai, Joe Yue-Hei Ng, Larry Davis 0001
WACV4
2019 Weakly-Supervised Spatial Context Networks
abstract
We explore the power of spatial context as a self-supervisory signal for learning visual representations. In particular, we propose spatial context networks that learn to predict a representation of one image patch from another image patch, within the same image, conditioned on their real-valued relative spatial offset. Unlike auto-encoders, that aim to encode and reconstruct original image patches, our network aims to encode and reconstruct intermediate representations of the spatially offset patches. As such, the network learns a spatially conditioned contextual representation. By testing performance with various patch selection mechanisms we show that focusing on object-centric patches is important, and that using object proposal as a patch selection mechanism leads to the highest improvement in performance. Further, unlike auto-encoders, context encoders [21], or other forms of unsupervised feature learning, we illustrate that contextual supervision (with pre-trained model initialization) can improve on existing pre-trained model performance. We build our spatial context networks on top of standard VGG_19 and CNN_M architectures and, among other things, show that we can achieve improvements (with no additional explicit supervision) over the original ImageNet pre-trained VGG_19 and CNN_M models in object categorization and detection on VOC2007.
Zuxuan Wu, Larry Davis 0001, Leonid Sigal
WACV2
2019 Truncated Cauchy Non-Negative Matrix Factorization
abstract
Non-negative matrix factorization (NMF) minimizes the euclidean distance between the data matrix and its low rank approximation, and it fails when applied to corrupted data because the loss function is sensitive to outliers. In this paper, we propose a Truncated CauchyNMF loss that handle outliers by truncating large errors, and develop a Truncated CauchyNMF to robustly learn the subspace on noisy datasets contaminated by outliers. We theoretically analyze the robustness of Truncated CauchyNMF comparing with the competing models and theoretically prove that Truncated CauchyNMF has a generalization bound which converges at a rate of order , where is the sample size. We evaluate Truncated CauchyNMF by image clustering on both simulated and real datasets. The experimental results on the datasets containing gross corruptions validate the effectiveness and robustness of Truncated CauchyNMF for learning robust subspaces.
Naiyang Guan, Tongliang Liu, Yangmuzi Zhang, Dacheng Tao, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2019 Award winning papers from the 23rd International Conference on Pattern Recognition (ICPR)
Larry Davis 0001, Alberto Del Bimbo, Brian C. Lovell
Pattern Recognit. Lett.1
2019 Introduction to the Special Section on Deep Learning for Visual Surveillance
abstract
We are now living in an era of visual information where data is unceasingly generated and pushed into consumption at astounding rates. A remarkable portion of this sensory input comes in the form of videos streaming from large-scale surveillance infrastructures as well as consumer-grade monitoring systems. The sheer amount of ground-based, aerial and mobile video surveillance data demands fittingly competent, accurate, effective techniques to extract useful cues and provide assistance for detection, prevention, and intervention tasks in traffic, safety, security, defense, forensic, health, biology, ethology, and retail space management applications.
Fatih Porikli, Larry Davis 0001, Qi Wang 0009, Yi Li 0025, Carlo S. Regazzoni
IEEE Trans. Circuits Syst. Video Technol.2
2018 Deception Detection in Videos
abstract
We present a system for covert automated deception detection using information available in a video. We study the importance of different modalities like vision, audio and text for this task. On the vision side, our system uses classifiers trained on low level video features which predict human micro-expressions. We show that predictions of high-level micro-expressions can be used as features for deception prediction. Surprisingly, IDT (Improved Dense Trajectory) features which have been widely used for action recognition, are also very good at predicting deception in videos. We fuse the score of classifiers trained on IDT features and high-level micro-expressions to improve performance. MFCC (Mel-frequency Cepstral Coefficients) features from the audio domain also provide a significant boost in performance, while information from transcripts is not very beneficial for our system. Using various classifiers, our automated system obtains an AUC of 0.877 (10-fold cross-validation) when evaluated on subjects which were not part of the training set. Even though state-of-the-art methods use human annotations of micro-expressions for deception detection, our fully automated approach outperforms them by 5%. When combined with human annotations of micro-expressions, our AUC improves to 0.922. We also present results of a user-study to analyze how well do average humans perform on this task, what modalities they use for deception detection and how they perform if only one modality is accessible.
Zhe Wu 0001, Larry Davis 0001, V. S. Subrahmanian
AAAI3
2018 Dynamic Zoom-In Network for Fast Object Detection in Large Images
abstract
We introduce a generic framework that reduces the computational cost of object detection while retaining accuracy for scenarios where objects with varied sizes appear in high resolution images. Detection progresses in a coarse-to-fine manner, first on a down-sampled version of the image and then on a sequence of higher resolution regions identified as likely to improve the detection accuracy. Built upon reinforcement learning, our approach consists of a model (R-net) that uses coarse detection results to predict the potential accuracy gain for analyzing a region at a higher resolution and another model (Q-net) that sequentially selects regions to zoom in. Experiments on the Caltech Pedestrians dataset show that our approach reduces the number of processed pixels by over 50% without a drop in detection accuracy. The merits of our approach become more significant on a high resolution test set collected from YFCC100M dataset, where our approach maintains high detection performance while reducing the number of processed pixels by about 70% and the detection time by over 50%.
Mingfei Gao, Ruichi Yu, Ang Li 0001, Vlad I. Morariu, Larry Davis 0001
CVPR5
2018 VITON: An Image-Based Virtual Try-On Network
abstract
We present an image-based VIirtual Try-On Network (VITON) without using 3D information in any form, which seamlessly transfers a desired clothing item onto the corresponding region of a person using a coarse-to-fine strategy. Conditioned upon a new clothing-agnostic yet descriptive person representation, our framework first generates a coarse synthesized image with the target clothing item overlaid on that same person in the same pose. We further enhance the initial blurry clothing area with a refinement network. The network is trained to learn how much detail to utilize from the target clothing item, and where to apply to the person in order to synthesize a photo-realistic image in which the target item deforms naturally with clear visual patterns. Experiments on our newly collected Zalando dataset demonstrate its promise in the image-based virtual try-on task over state-of-the-art generative models.
Xintong Han, Zuxuan Wu, Zhe Wu 0001, Ruichi Yu, Larry Davis 0001
CVPR5
2018 An Analysis of Scale Invariance in Object Detection ­ SNIP
abstract
An analysis of different techniques for recognizing and detecting objects under extreme scale variation is presented. Scale specific and scale invariant design of detectors are compared by training them with different configurations of input data. By evaluating the performance of different network architectures for classifying small objects on ImageNet, we show that CNNs are not robust to changes in scale. Based on this analysis, we propose to train and test detectors on the same scales of an image-pyramid. Since small and large objects are difficult to recognize at smaller and larger scales respectively, we present a novel training scheme called Scale Normalization for Image Pyramids (SNIP) which selectively back-propagates the gradients of object instances of different sizes as a function of the image scale. On the COCO dataset, our single model performance is 45.7% and an ensemble of 3 networks obtains an mAP of 48.3%. We use off-the-shelf ImageNet-1000 pre-trained models and only train with bounding box supervision. Our submission won the Best Student Entry in the COCO 2017 challenge. Code will be made available at http://bit.ly/2yXVg4c.
Larry Davis 0001
CVPR2
2018 R-FCN-3000 at 30fps: Decoupling Detection and Classification
abstract
We propose a modular approach towards large-scale real-time object detection by decoupling objectness detection and classification. We exploit the fact that many object classes are visually similar and share parts. Thus, a universal objectness detector can be learned for class-agnostic object detection followed by fine-grained classification using a (non)linear classifier. Our approach is a modification of the R-FCN architecture to learn shared filters for performing localization across different object classes. We trained a detector for 3000 object classes, called R-FCN-3000, that obtains an mAP of 34.9% on the ImageNet detection dataset. It outperforms YOLO-9000 by 18% while processing 30 images per second. We also show that the objectness learned by R-FCN-3000 generalizes to novel classes and the performance increases with the number of training object classes - supporting the hypothesis that it is possible to learn a universal objectness detector. Code will be made available.
Hengduo Li, Abhishek Sharma 0001, Larry Davis 0001
CVPR4
2018 Learning a Discriminative Filter Bank Within a CNN for Fine-Grained Recognition
abstract
Compared to earlier multistage frameworks using CNN features, recent end-to-end deep approaches for fine-grained recognition essentially enhance the mid-level learning capability of CNNs. Previous approaches achieve this by introducing an auxiliary network to infuse localization information into the main classification network, or a sophisticated feature encoding method to capture higher order feature statistics. We show that mid-level representation learning can be enhanced within the CNN framework, by learning a bank of convolutional filters that capture class-specific discriminative patches without extra part or bounding box annotations. Such a filter bank is well structured, properly initialized and discriminatively learned through a novel asymmetric multi-stream architecture with convolutional filter supervision and a non-random layer initialization. Experimental results show that our approach achieves state-of-the-art on three publicly available fine-grained recognition datasets (CUB-200-2011, Stanford Cars and FGVC-Aircraft). Ablation studies and visualizations are provided to understand our approach.
Yaming Wang, Vlad I. Morariu, Larry Davis 0001
CVPR3
2018 BlockDrop: Dynamic Inference Paths in Residual Networks
abstract
Very deep convolutional neural networks offer excellent recognition results, yet their computational expense limits their impact for many real-world applications. We introduce BlockDrop, an approach that learns to dynamically choose which layers of a deep network to execute during inference so as to best reduce total computation without degrading prediction accuracy. Exploiting the robustness of Residual Networks (ResNets) to layer dropping, our framework selects on-the-fly which residual blocks to evaluate for a given novel image. In particular, given a pretrained ResNet, we train a policy network in an associative reinforcement learning setting for the dual reward of utilizing a minimal number of blocks while preserving recognition accuracy. We conduct extensive experiments on CIFAR and ImageNet. The results provide strong quantitative and qualitative evidence that these learned policies not only accelerate inference but also encode meaningful visual information. Built upon a ResNet-101 model, our method achieves a speedup of 20% on average, going as high as 36% for some images, while maintaining the same 76.4% top-1 accuracy on ImageNet.
Zuxuan Wu, Tushar Nagarajan, Abhishek Kumar 0001, Steven Rennie, Larry Davis 0001, Kristen Grauman, Rogério Feris
CVPR5
2018 NISP: Pruning Networks Using Neuron Importance Score Propagation
abstract
To reduce the significant redundancy in deep Convolutional Neural Networks (CNNs), most existing methods prune neurons by only considering the statistics of an individual layer or two consecutive layers (e.g., prune one layer to minimize the reconstruction error of the next layer), ignoring the effect of error propagation in deep networks. In contrast, we argue that for a pruned network to retain its predictive power, it is essential to prune neurons in the entire neuron network jointly based on a unified goal: minimizing the reconstruction error of important responses in the "final response layer" (FRL), which is the second-to-last layer before classification. Specifically, we apply feature ranking techniques to measure the importance of each neuron in the FRL, formulate network pruning as a binary integer optimization problem, and derive a closed-form solution to it for pruning neurons in earlier layers. Based on our theoretical analysis, we propose the Neuron Importance Score Propagation (NISP) algorithm to propagate the importance scores of final responses to every neuron in the network. The CNN is pruned by removing neurons with least importance, and it is then fine-tuned to recover its predictive power. NISP is evaluated on several datasets with multiple CNN models and demonstrated to achieve significant acceleration and compression with negligible accuracy loss.
Ruichi Yu, Ang Li 0001, Chun-Fu Chen 0001, Jui-Hsin Lai, Vlad I. Morariu, Xintong Han, Mingfei Gao, Ching-Yung Lin, Larry Davis 0001
CVPR9
2018 Learning Rich Features for Image Manipulation Detection
abstract
Image manipulation detection is different from traditional semantic object detection because it pays more attention to tampering artifacts than to image content, which suggests that richer features need to be learned. We propose a two-stream Faster R-CNN network and train it end-to-end to detect the tampered regions given a manipulated image. One of the two streams is an RGB stream whose purpose is to extract features from the RGB image input to find tampering artifacts like strong contrast difference, unnatural tampered boundaries, and so on. The other is a noise stream that leverages the noise features extracted from a steganalysis rich model filter layer to discover the noise inconsistency between authentic and tampered regions. We then fuse features from the two streams through a bilinear pooling layer to further incorporate spatial co-occurrence of these two modalities. Experiments on four standard image manipulation datasets demonstrate that our two-stream framework outperforms each individual stream, and also achieves state-of-the-art performance compared to alternative methods with robustness to resizing and compression.
Peng Zhou 0009, Xintong Han, Vlad I. Morariu, Larry Davis 0001
CVPR4
2018 Weakly-Supervised Video Summarization Using Variational Encoder-Decoder and Web Prior
Sijia Cai, Wangmeng Zuo, Larry Davis 0001, Lei Zhang 0006
ECCV (14)3
2018 C-WSL: Count-Guided Weakly Supervised Localization
Mingfei Gao, Ang Li 0001, Ruichi Yu, Vlad I. Morariu, Larry Davis 0001
ECCV (1)5
2018 DCAN: Dual Channel-Wise Alignment Networks for Unsupervised Scene Adaptation
Zuxuan Wu, Xintong Han, Yen-Liang Lin, Mustafa Gökhan Uzunbas, Tom Goldstein, Ser-Nam Lim, Larry Davis 0001
ECCV (5)7
2018 SNIPER: Efficient Multi-Scale Training
abstract
We present SNIPER, an algorithm for performing efficient multi-scale training in instance level visual recognition tasks. Instead of processing every pixel in an image pyramid, SNIPER processes context regions around ground-truth instances (referred to as chips) at the appropriate scale. For background sampling, these context-regions are generated using proposals extracted from a region proposal network trained with a short learning schedule. Hence, the number of chips generated per image during training adaptively changes based on the scene complexity. SNIPER only processes 30% more pixels compared to the commonly used single scale training at 800x1333 pixels on the COCO dataset. But, it also observes samples from extreme resolutions of the image pyramid, like 1400x2000 pixels. As SNIPER operates on resampled low resolution chips (512x512 pixels), it can have a batch size as large as 20 on a single GPU even with a ResNet-101 backbone. Therefore it can benefit from batch-normalization during training without the need for synchronizing batch-normalization statistics across GPUs. SNIPER brings training of instance level recognition tasks like object detection closer to the protocol for image classification and suggests that the commonly accepted guideline that it is important to train on high resolution images for instance level visual recognition tasks might not be correct. Our implementation based on Faster-RCNN with a ResNet-101 backbone obtains an mAP of 47.6% on the COCO dataset for bounding box detection and can process 5 images per second during inference with a single GPU. Code is available at https://github.com/MahyarNajibi/SNIPER/ .
Mahyar Najibi, Larry Davis 0001
NeurIPS3
2018 ActionFlowNet: Learning Motion Representation for Action Recognition
abstract
We present a data-efficient representation learning approach to learn video representation with small amount of labeled data. We propose a multitask learning model ActionFlowNet to train a single stream network directly from raw pixels to jointly estimate optical flow while recognizing actions with convolutional neural networks, capturing both appearance and motion in a single model. Our model effectively learns video representation from motion information on unlabeled videos. Our model significantly improves action recognition accuracy by a large margin (23.6%) compared to state-of-the-art CNN-based unsupervised representation learning methods trained without external large scale data and additional optical flow input. Without pretraining on large external labeled datasets, our model, by well exploiting the motion information, achieves competitive recognition accuracy to the models trained with large labeled datasets such as ImageNet and Sport-1M.
Joe Yue-Hei Ng, Jan Neumann, Larry Davis 0001
WACV4
2018 Temporal Difference Networks for Video Action Recognition
abstract
Deep convolutional neural networks have been great success for image based recognition tasks. However, it is still unclear how to model the temporal evolution of videos effectively by deep networks. While recent deep models for videos show improvement by incorporating optical flow or aggregating high level appearance across frames, they focus on modeling either the long term temporal relations or short term motion. We propose Temporal Difference Networks (TDN) that model both long term relations and short term motion from videos. We leverage a simple but effective motion representation: difference of CNN features in our network and jointly modeling the motion at multiple scales in a single CNN. It achieves state-of-the-art performance on three different video classification benchmarks, showing the effectiveness of our approach to learn temporal relations in videos.
Joe Yue-Hei Ng, Larry Davis 0001
WACV2
2018 Face-MagNet: Magnifying Feature Maps to Detect Small Faces
abstract
In this paper, we introduce the Face Magnifier Network (Face-MageNet), a face detector based on the Faster-RCNN framework which enables the flow of discriminative information of small scale faces to the classifier without any skip or residual connections. To achieve this, Face-MagNet deploys a set of ConvTranspose, also known as deconvolution, layers in the Region Proposal Network (RPN) and another set before the Region of Interest (RoI) pooling layer to facilitate detection of finer faces. In addition, we also design, train, and evaluate three other well-tuned architectures that represent the conventional solutions to the scale problem: context pooling, skip connections, and scale partitioning. Each of these three networks achieves comparable results to the state-of-the-art face detectors. With extensive experiments, we show that Face-MagNet based on a VGG16 architecture achieves better results than the recently proposed ResNet101-based HR [7] method on the task of face detection on WIDER [25] dataset and also achieves similar results on the hard set as our other method SSH [17].
Pouya Samangouei, Rama Chellappa, Mahyar Najibi, Larry Davis 0001
WACV4
2018 ReMotENet: Efficient Relevant Motion Event Detection for Large-Scale Home Surveillance Videos
abstract
This paper addresses the problem of detecting relevant motion caused by objects of interest (e.g., person and vehicles) in large scale home surveillance videos. The traditional method usually consists of two separate steps, i.e., detecting moving objects with background subtraction running on the camera, and filtering out nuisance motion events with deep learning based object detection and tracking running on cloud. The method is extremely slow, and does not fully leverage the spatial-temporal redundancies with a pre-trained off-the-shelf object detector. To dramatically speedup relevant motion event detection and improve its performance, we propose a novel network for relevant motion event detection, ReMotENet, which is a unified, end-to-end data-driven method using spatial-temporal attention-based 3D ConvNets to jointly model the appearance and motion of objects-of-interest in a video. Re-MotENet parses an entire video clip in one forward pass of a neural network to achieve significant speedup, which exploits the properties of home surveillance videos, and enhances 3D ConvNets with a spatial-temporal attention model and frame differencing to encourage the network to focus on the relevant moving objects. Experiments demonstrate that our method can achieve comparable or event better performance than the object detection based method but with three to four orders of magnitude speedup (up to 20k) on GPU devices. Our network is efficient, compact and light-weight. It can detect relevant motion on a 15s surveillance video clip within 4-8 milliseconds on a GPU and a fraction of second (0.17-0.39s) on a CPU with a model size of less than 1MB.
Ruichi Yu, Larry Davis 0001
WACV3
2018 Multi-Task Learning with Low Rank Attribute Embedding for Multi-Camera Person Re-Identification
abstract
We propose Multi-Task Learning with Low Rank Attribute Embedding (MTL-LORAE) to address the problem of person re-identification on multi-cameras. Re-identifications on different cameras are considered as related tasks, which allows the shared information among different tasks to be explored to improve the re-identification accuracy. The MTL-LORAE framework integrates low-level features with mid-level attributes as the descriptions for persons. To improve the accuracy of such description, we introduce the low-rank attribute embedding, which maps original binary attributes into a continuous space utilizing the correlative relationship between each pair of attributes. In this way, inaccurate attributes are rectified and missing attributes are recovered. The resulting objective function is constructed with an attribute embedding error and a quadratic loss concerning class labels. It is solved by an alternating optimization strategy. The proposed MTL-LORAE is tested on four datasets and is validated to outperform the existing methods with significant margins.
Chi Su, Fan Yang 0016, Shiliang Zhang, Qi Tian 0001, Larry Davis 0001, Wen Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2018 Learning structured ordinal measures for video based face recognition
Ran He 0001, Tieniu Tan, Larry Davis 0001, Zhenan Sun
Pattern Recognit.3
2018 Preface of Special Issue on Data Representation and Representation Learning for Video Analysis
William Robson Schwartz, Larry Davis 0001
Pattern Recognit. Lett.2
2017 FASON: First and Second Order Information Fusion Network for Texture Recognition
abstract
Deep networks have shown impressive performance on many computer vision tasks. Recently, deep convolutional neural networks (CNNs) have been used to learn discriminative texture representations. One of the most successful approaches is Bilinear CNN model that explicitly captures the second order statistics within deep features. However, these networks cut off the first order information flow in the deep network and make gradient back-propagation difficult. We propose an effective fusion architecture - FASON that combines second order information flow and first order information flow. Our method allows gradients to back-propagate through both flows freely and can be trained effectively. We then build a multi-level deep architecture to exploit the first and second order information within different convolutional layers. Experiments show that our method achieves improvements over state-of-the-art methods on several benchmark datasets.
Xiyang Dai, Joe Yue-Hei Ng, Larry Davis 0001
CVPR3
2017 Fast-At: Fast Automatic Thumbnail Generation Using Deep Neural Networks
Seyed A. Esmaeili, Larry Davis 0001
CVPR3
2017 The Amazing Mysteries of the Gutter: Drawing Inferences Between Panels in Comic Book Narratives
abstract
Visual narrative is often a combination of explicit information and judicious omissions, relying on the viewer to supply missing details. In comics, most movements in time and space are hidden in the gutters between panels. To follow the story, readers logically connect panels together by inferring unseen actions through a process called closure. While computers can now describe the content of natural images, in this paper we examine whether they can understand the closure-driven narratives conveyed by stylized artwork and dialogue in comic book panels. We collect a dataset, COMICS, that consists of over 1.2 million panels (120 GB) paired with automatic textbox transcriptions. An in-depth analysis of COMICS demonstrates that neither text nor image alone can tell a comic book story, so a computer must understand both modalities to keep up with the plot. We introduce three cloze-style tasks that ask models to predict narrative and character-centric aspects of a panel given n preceding panels as context. Various deep neural architectures underperform human baselines on these tasks, suggesting that COMICS contains fundamental challenges for both vision and language.
Mohit Iyyer, Varun Manjunatha, Anupam Guha, Yogarshi Vyas, Jordan L. Boyd-Graber, Hal Daumé III, Larry Davis 0001
CVPR7
2017 Generating Holistic 3D Scene Abstractions for Text-Based Image Retrieval
abstract
Spatial relationships between objects provide important information for text-based image retrieval. As users are more likely to describe a scene from a real world perspective, using 3D spatial relationships rather than 2D relationships that assume a particular viewing direction, one of the main challenges is to infer the 3D structure that bridges images with users text descriptions. However, direct inference of 3D structure from images requires learning from large scale annotated data. Since interactions between objects can be reduced to a limited set of atomic spatial relations in 3D, we study the possibility of inferring 3D structure from a text description rather than an image, applying physical relation models to synthesize holistic 3D abstract object layouts satisfying the spatial constraints present in a textual description. We present a generic framework for retrieving images from a textual description of a scene by matching images with these generated abstract object layouts. Images are ranked by matching object detection outputs (bounding boxes) to 2D layout candidates (also represented by bounding boxes) which are obtained by projecting the 3D scenes with sampled camera directions. We validate our approach using public indoor scene datasets and show that our method outperforms baselines built upon object occurrence histograms and learned 2D pairwise relations.
Ang Li 0001, Jin Sun 0011, Joe Yue-Hei Ng, Ruichi Yu, Vlad I. Morariu, Larry Davis 0001
CVPR6
2017 Generalized Deep Image to Image Regression
abstract
We present a Deep Convolutional Neural Network architecture which serves as a generic image-to-image regressor that can be trained end-to-end without any further machinery. Our proposed architecture, the Recursively Branched Deconvolutional Network (RBDN), develops a cheap multi-context image representation very early on using an efficient recursive branching scheme with extensive parameter sharing and learnable upsampling. This multi-context representation is subjected to a highly non-linear locality preserving transformation by the remainder of our network comprising of a series of convolutions/deconvolutions without any spatial downsampling. The RBDN architecture is fully convolutional and can handle variable sized images during inference. We provide qualitative/quantitative results on 3 diverse tasks: relighting, denoising and colorization and show that our proposed RBDN architecture obtains comparable results to the state-of-the-art on each of these tasks when used off-the-shelf without any post processing or task-specific architectural modifications.
Venkataraman Santhanam, Vlad I. Morariu, Larry Davis 0001
CVPR3
2017 Soft-NMS - Improving Object Detection with One Line of Code
abstract
Non-maximum suppression is an integral part of the object detection pipeline. First, it sorts all detection boxes on the basis of their scores. The detection box M with the maximum score is selected and all other detection boxes with a significant overlap (using a pre-defined threshold) with M are suppressed. This process is recursively applied on the remaining boxes. As per the design of the algorithm, if an object lies within the predefined overlap threshold, it leads to a miss. To this end, we propose Soft-NMS, an algorithm which decays the detection scores of all other objects as a continuous function of their overlap with M. Hence, no object is eliminated in this process. Soft-NMS obtains consistent improvements for the coco-style mAP metric on standard datasets like PASCAL VOC2007 (1.7% for both R-FCN and Faster-RCNN) and MS-COCO (1.3% for R-FCN and 1.1% for Faster-RCNN) by just changing the NMS algorithm without any additional hyper-parameters. Using Deformable-RFCN, Soft-NMS improves state-of-the-art in object detection from 39.8% to 40.9% with a single model. Further, the computational complexity of Soft-NMS is the same as traditional NMS and hence it can be efficiently implemented. Since Soft-NMS does not require any extra training and is simple to implement, it can be easily integrated into any object detection pipeline. Code for Soft-NMS is publicly available on GitHub http://bit.ly/2nJLNMu.
Navaneeth Bodla, Rama Chellappa, Larry Davis 0001
ICCV4
2017 Temporal Context Network for Activity Localization in Videos
abstract
We present a Temporal Context Network (TCN) for precise temporal localization of human activities. Similar to the Faster-RCNN architecture, proposals are placed at equal intervals in a video which span multiple temporal scales. We propose a novel representation for ranking these proposals. Since pooling features only inside a segment is not sufficient to predict activity boundaries, we construct a representation which explicitly captures context around a proposal for ranking it. For each temporal segment inside a proposal, features are uniformly sampled at a pair of scales and are input to a temporal convolutional neural network for classification. After ranking proposals, non-maximum suppression is applied and classification is performed to obtain final detections. TCN outperforms state-of-the-art methods on the ActivityNet dataset and the THU-MOS14 dataset.
Xiyang Dai, Guyue Zhang, Larry Davis 0001, Yan Qiu Chen
ICCV4
2017 Automatic Spatially-Aware Fashion Concept Discovery
abstract
This paper proposes an automatic spatially-aware concept discovery approach using weakly labeled image-text data from shopping websites. We first fine-tune GoogleNet by jointly modeling clothing images and their corresponding descriptions in a visual-semantic embedding space. Then, for each attribute (word), we generate its spatiallyaware representation by combining its semantic word vector representation with its spatial representation derived from the convolutional maps of the fine-tuned network. The resulting spatially-aware representations are further used to cluster attributes into multiple groups to form spatiallyaware concepts (e.g., the neckline concept might consist of attributes like v-neck, round-neck, etc). Finally, we decompose the visual-semantic embedding space into multiple concept-specific subspaces, which facilitates structured browsing and attribute-feedback product retrieval by exploiting multimodal linguistic regularities. We conducted extensive experiments on our newly collected Fashion200K dataset, and results on clustering quality evaluation and attribute-feedback product retrieval task demonstrate the effectiveness of our automatically discovered spatially-aware concepts.
Xintong Han, Zuxuan Wu, Phoenix X. Huang, Menglong Zhu, Larry Davis 0001
ICCV8
2017 SSH: Single Stage Headless Face Detector
abstract
We introduce the Single Stage Headless (SSH) face detector. Unlike two stage proposal-classification detectors, SSH detects faces in a single stage directly from the early convolutional layers in a classification network. SSH is headless. That is, it is able to achieve state-of-the-art results while removing the “head” of its underlying classification network - i.e. all fully connected layers in the VGG-16 which contains a large number of parameters. Additionally, instead of relying on an image pyramid to detect faces with various scales, SSH is scale-invariant by design. We simultaneously detect faces with different scales in a single forward pass of the network, but from different layers. These properties make SSH fast and light-weight. Surprisingly, with a headless VGG-16, SSH beats the ResNet-101-based state-of-the-art on the WIDER dataset. Even though, unlike the current state-of-the-art, SSH does not use an image pyramid and is 5X faster. Moreover, if an image pyramid is deployed, our light-weight network achieves state-of-the-art on all subsets of the WIDER dataset, improving the AP by 2.5%. SSH also reaches state-of-the-art results on the FDDB and Pascal-Faces datasets while using a small input size, leading to a speed of 50 frames/second on a GPU.
Mahyar Najibi, Pouya Samangouei, Rama Chellappa, Larry Davis 0001
ICCV4
2017 Visual Relationship Detection with Internal and External Linguistic Knowledge Distillation
abstract
Understanding the visual relationship between two objects involves identifying the subject, the object, and a predicate relating them. We leverage the strong correlations between the predicate and the hsubj; obji pair (both semantically and spatially) to predict predicates conditioned on the subjects and the objects. Modeling the three entities jointly more accurately reflects their relationships compared to modeling them independently, but it complicates learning since the semantic space of visual relationships is huge and training data is limited, especially for longtail relationships that have few instances. To overcome this, we use knowledge of linguistic statistics to regularize visual model learning. We obtain linguistic knowledge by mining from both training annotations (internal knowledge) and publicly available text, e.g., Wikipedia (external knowledge), computing the conditional probability distribution of a predicate given a (subj, obj) pair. As we train the visual model, we distill this knowledge into the deep model to achieve better generalization. Our experimental results on the Visual Relationship Detection (VRD) and Visual Genome datasets suggest that with this linguistic knowledge distillation, our model outperforms the stateof- the-art methods significantly, especially when predicting unseen relationships (e.g., recall improved from 8.45% to 19.17% on VRD zero-shot testing set).
Ruichi Yu, Ang Li 0001, Vlad I. Morariu, Larry Davis 0001
ICCV4
2017 Towards Unified Data and Lifecycle Management for Deep Learning
abstract
Deep learning has improved state-of-the-art results in many important fields, and has been the subject of much research in recent years, leading to the development of several systems for facilitating deep learning. Current systems, however, mainly focus on model building and training phases, while the issues of data management, model sharing, and lifecycle management are largely ignored. Deep learning modeling lifecycle generates a rich set of data artifacts, e.g., learned parameters and training logs, and it comprises of several frequently conducted tasks, e.g., to understand the model behaviors and to try out new models. Dealing with such artifacts and tasks is cumbersome and largely left to the users. This paper describes our vision and implementation of a data and lifecycle management system for deep learning. First, we generalize model exploration and model enumeration queries from commonly conducted tasks by deep learning modelers, and propose a high-level domain specific language (DSL), inspired by SQL, to raise the abstraction level and thereby accelerate the modeling process. To manage the variety of data artifacts, especially the large amount of checkpointed float parameters, we design a novel model versioning system (dlv), and a read-optimized parameter archival storage system (PAS) that minimizes storage footprint and accelerates query workloads with minimal loss of accuracy. PAS archives versioned models using deltas in a multi-resolution fashion by separately storing the less significant bits, and features a novel progressive query (inference) evaluation algorithm. Third, we develop e cient algorithms for archiving versioned models using deltas under co-retrieval constraints. We conduct extensive experiments over several real datasets from computDeep learning has improved state-of-the-art results in many important fields, and has been the subject of much research in recent years, leading to the development of several systems for facilitating deep learning. Current systems, however, mainly focus on model building and training phases, while the issues of data management, model sharing, and lifecycle management are largely ignored. Deep learning modeling lifecycle generates a rich set of data artifacts, e.g., learned parameters and training logs, and it comprises of several frequently conducted tasks, e.g., to understand the model behaviors and to try out new models. Dealing with such artifacts and tasks is cumbersome and largely left to the users. This paper describes our vision and implementation of a data and lifecycle management system for deep learning. First, we generalize model exploration and model enumeration queries from commonly conducted tasks by deep learning modelers, and propose a high-level domain specific language (DSL), inspired by SQL, to raise the abstraction level and thereby accelerate the modeling process. To manage the variety of data artifacts, especially the large amount of checkpointed float parameters, we design a novel model versioning system (dlv), and a read-optimized parameter archival storage system (PAS) that minimizes storage footprint and accelerates query workloads with minimal loss of accuracy. PAS archives versioned models using deltas in a multi-resolution fashion by separately storing the less significant bits, and features a novel progressive query (inference) evaluation algorithm. Third, we develop efficient algorithms for archiving versioned models using deltas under co-retrieval constraints. We conduct extensive experiments over several real datasets from computer vision domain to show the efficiency of the proposed techniques.
Hui Miao 0001, Ang Li 0001, Larry Davis 0001, Amol Deshpande
ICDE3
2017 ModelHub: Deep Learning Lifecycle Management
abstract
Deep learning has improved the state-of-the-art results in many domains, leading to the development of several systems for facilitating deep learning. Current systems, however, mainly focus on model building and training phases, while the issues of lifecycle management are largely ignored. Deep learning modeling lifecycle contains a rich set of artifacts and frequently conducted tasks, dealing with them is cumbersome and left to the users. To address these issues in a comprehensive manner, we demonstrate ModelHub, which includes a novel model versioning system (dlv), a domain-specific language for searching through model space (DQL), and a hosted service (ModelHub).
Hui Miao 0001, Ang Li 0001, Larry Davis 0001, Amol Deshpande
ICDE3
2017 Learning Fashion Compatibility with Bidirectional LSTMs
abstract
The ubiquity of online fashion shopping demands effective recommendation services for customers. In this paper, we study two types of fashion recommendation: (i) suggesting an item that matches existing components in a set to form a stylish outfit (a collection of fashion items), and (ii) generating an outfit with multimodal (images/text) specifications from a user. To this end, we propose to jointly learn a visual-semantic embedding and the compatibility relationships among fashion items in an end-to-end fashion. More specifically, we consider a fashion outfit to be a sequence (usually from top to bottom and then accessories) and each item in the outfit as a time step. Given the fashion items in an outfit, we train a bidirectional LSTM (Bi-LSTM) model to sequentially predict the next item conditioned on previous ones to learn their compatibility relationships. Further, we learn a visual-semantic space by regressing image features to their semantic representations aiming to inject attribute and category information as a regularization for training the LSTM. The trained network can not only perform the aforementioned recommendations effectively but also predict the compatibility of a given outfit. We conduct extensive experiments on our newly collected Polyvore dataset, and the results provide strong qualitative and quantitative evidence that our framework outperforms alternative methods.
Xintong Han, Zuxuan Wu, Yu-Gang Jiang 0001, Larry Davis 0001
ACM Multimedia4
2017 LSVC2017: Large-Scale Video Classification Challenge
abstract
Recognizing visual contents in unconstrained videos has become a very important problem for many applications, such as Web video search and recommendation, smart advertising, robotics, etc. This workshop and challenge aims at exploring new challenges and approaches for large-scale video classification with large number of classes from open source videos in a realistic setting, based upon an extension of Fudan-Columbia Video Dataset (FCVID). This newly collected dataset contains over 8000 hours of video data from YouTube and Flicker, annotated into 500 categories. We hope this dataset can stimulate innovative research on this challenging and important problem.
Zuxuan Wu, Yu-Gang Jiang 0001, Larry Davis 0001, Shih-Fu Chang
ACM Multimedia3
2017 Fused DNN: A Deep Neural Network Fusion Approach to Fast and Robust Pedestrian Detection
abstract
We propose a deep neural network fusion architecture for fast and robust pedestrian detection. The proposed network fusion architecture allows for parallel processing of multiple networks for speed. A single shot deep convolutional network is trained as a object detector to generate all possible pedestrian candidates of different sizes and occlusions. This network outputs a large variety of pedestrian candidates to cover the majority of ground-truth pedestrians while also introducing a large number of false positives. Next, multiple deep neural networks are used in parallel for further refinement of these pedestrian candidates. We introduce a soft-rejection based network fusion method to fuse the soft metrics from all networks together to generate the final confidence scores. Our method performs better than existing state-of-the-arts, especially when detecting small-size and occluded pedestrians. Furthermore, we propose a method for integrating pixel-wise semantic segmentation network into the network fusion architecture as a reinforcement to the pedestrian detector. The approach outperforms state-of-the-art methods on most protocols on Caltech Pedestrian dataset, with significant boosts on several protocols. It is also faster than all other methods.
Xianzhi Du, Mostafa El-Khamy, Larry Davis 0001
WACV4
2017 Learning Discriminative Features via Label Consistent Neural Network
abstract
Deep Convolutional Neural Networks (CNN) enforce supervised information only at the output layer, and hidden layers are trained by back propagating the prediction error from the output layer without explicit supervision. We propose a supervised feature learning approach, Label Consistent Neural Network, which enforces direct supervision in late hidden layers in a novel way. We associate each neuron in a hidden layer with a particular class label and encourage it to be activated for input signals from the same class. More specifically, we introduce a label consistency regularization called "discriminative representation error" loss for late hidden layers and combine it with classification error loss to build our overall objective function. This label consistency constraint alleviates the common problem of gradient vanishing and tends to faster convergence, it also makes the features derived from late hidden layers discriminative enough for classification even using a simple k-NN classifier. Experimental results demonstrate that our approach achieves state-of-the-art performances on several public datasets for action and object category recognition.
Zhuolin Jiang, Yaming Wang, Larry Davis 0001, Walter Andrews, Viktor Rozgic
WACV3
2017 Attributes driven tracklet-to-tracklet person re-identification using latent prototypes space mapping
Chi Su, Shiliang Zhang, Fan Yang 0016, Guangxiao Zhang, Qi Tian 0001, Wen Gao 0001, Larry Davis 0001
Pattern Recognit.7
2017 Joint Human Detection and Head Pose Estimation via Multistream Networks for RGB-D Videos
abstract
We propose a multistream multitask deep network for joint human detection and head pose estimation in RGB-D videos. To achieve high accuracy, we jointly utilize appearance, shape, and motion information as inputs. Based on the depth information, we generate scale invariant proposals, which are then fed into a novel contextual region of interest pooling (CRP) layer in our deep network. This CRP has two branches to deal with contextual information for each subject. The proposed method outperforms state-of-the-art approaches on three public datasets.
Guyue Zhang, Jun Liu 0036, Hengduo Li, Yan Qiu Chen, Larry Davis 0001
IEEE Signal Process. Lett.5
2017 VRFP: On-the-Fly Video Retrieval Using Web Images and Fast Fisher Vector Products
abstract
On-the-fly video retrieval using Web images and fast Fisher Vector products (VRFP) is a real-time video retrieval framework based on short text input queries, which obtains weakly labeled training images from the Web after the query is known. The retrieved Web images representing the query and each database video are treated as unordered collections of images, and each collection is represented using a single Fisher Vector built on CNN features. Our experiments show that a Fisher Vector is robust to noise present in Web images and compares favorably in terms of accuracy to other standard representations. While a Fisher Vector can be constructed efficiently for a new query, matching against the test set is slow due to its high dimensionality. To perform matching in real time, we present a lossless algorithm that accelerates the inner product computation between high-dimensional Fisher Vectors. We prove that the expected number of multiplications required decreases quadratically with the sparsity of Fisher Vectors. We are not only able to construct and apply query models in real time, but with the help of a simple reranking scheme, we also outperform state-of-the-art automatic retrieval methods by a significant margin on TRECVID MED13 (3.5%), MED14 (1.3%), and CCV datasets (5.2%). We also provide a direct comparison on standard datasets between two different paradigms for automatic video retrieval: zero-shot learning and on-the-fly retrieval.
Xintong Han, Vlad I. Morariu, Larry Davis 0001
IEEE Trans. Multim.4
2016 Knowledge Transfer with Interactive Learning of Semantic Relationships
abstract
We propose a novel learning framework for object categorization with interactive semantic feedback. In this framework, a discriminative categorization model improves through human-guided iterative semantic feedbacks. Specifically, the model identifies the most helpful relational semantic queries to discriminatively refine the model. The user feedback on whether the relationship is semantically valid or not is incorporated back into the model, in the form of regularization, and the process iterates. We validate the proposed model in a few-shot multi-class classification scenario, where we measure classification performance on a set of ‘target’ classes, with few training instances, by leveraging and transferring knowledge from ‘anchor’ classes, that contain larger set of labeled instances.
Sung Ju Hwang, Leonid Sigal, Larry Davis 0001
AAAI4
2016 Supervised Incremental Hashing
Bahadir Ozdemir, Mahyar Najibi, Larry Davis 0001
BMVC3
2016 The Role of Context Selection in Object Detection
Ruichi Yu, Xi Chen 0016, Vlad I. Morariu, Larry Davis 0001
BMVC4
2016 Learning Temporal Regularity in Video Sequences
abstract
Perceiving meaningful activities in a long video sequence is a challenging problem due to ambiguous definition of 'meaningfulness' as well as clutters in the scene. We approach this problem by learning a generative model for regular motion patterns (termed as regularity) using multiple sources with very limited supervision. Specifically, we propose two methods that are built upon the autoencoders for their ability to work with little to no supervision. We first leverage the conventional handcrafted spatio-temporal local features and learn a fully connected autoencoder on them. Second, we build a fully convolutional feed-forward autoencoder to learn both the local features and the classifiers as an end-to-end learning framework. Our model can capture the regularities from multiple datasets. We evaluate our methods in both qualitative and quantitative ways - showing the learned regularity of videos in various aspects and demonstrating competitive performance on anomaly detection datasets as an application.
Mahmudul Hasan 0003, Jan Neumann, Amit K. Roy-Chowdhury, Larry Davis 0001
CVPR5
2016 G-CNN: An Iterative Grid Based Object Detector
abstract
We introduce G-CNN, an object detection technique based on CNNs which works without proposal algorithms. G-CNN starts with a multi-scale grid of fixed bounding boxes. We train a regressor to move and scale elements of the grid towards objects iteratively. G-CNN models the problem of object detection as finding a path from a fixed grid to boxes tightly surrounding the objects. G-CNN with around 180 boxes in a multi-scale grid performs comparably to Fast R-CNN which uses around 2K bounding boxes generated with a proposal technique. This strategy makes detection faster by removing the object proposal stage as well as reducing the number of boxes to be processed.
Mahyar Najibi, Mohammad Rastegari, Larry Davis 0001
CVPR3
2016 Mining Discriminative Triplets of Patches for Fine-Grained Classification
abstract
Fine-grained classification involves distinguishing between similar sub-categories based on subtle differences in highly localized regions, therefore, accurate localization of discriminative regions remains a major challenge. We describe a patch-based framework to address this problem. We introduce triplets of patches with geometric constraints to improve the accuracy of patch localization, and automatically mine discriminative geometrically-constrained triplets for classification. The resulting approach only requires object bounding boxes. Its effectiveness is demonstrated using four publicly available fine-grained datasets, on which it outperforms or achieves comparable performance to the state-of-the-art in classification.
Yaming Wang, Vlad I. Morariu, Larry Davis 0001
CVPR4
2016 Modeling Context Between Objects for Referring Expression Understanding
Varun K. Nagaraja, Vlad I. Morariu, Larry Davis 0001
ECCV (4)3
2016 Weakly Supervised Learning of Heterogeneous Concepts in Videos
Sohil Shah, Kuldeep Kulkarni, Arijit Biswas, Ankit Gandhi, Om Deshmukh, Larry Davis 0001
ECCV (6)6
2016 Semantic Binary Codes
abstract
Fast Image Retrieval is required for many applications like Image Search and Shopping, especially for large datasets. Hashing addresses this problem by learning compact binary codes for images and using them as direct addresses into hash tables. In practice, using binary codes as addresses does not guarantee fast retrieval, as similar images are not mapped to the same binary code(address). We address this problem by presenting an efficient supervised hashing method that aims to explicitly map all images from the same class to a unique binary code to obtain fast retrieval. We refer to the binary codes of the images as 'Semantic Binary Codes' and the unique code for all same class images as 'Class Binary Code'. We formulate this intuitive objective 'directly' by minimizing the squared error criterion between the semantic binary codes and the corresponding class binary codes. We further propose a Deep Semantic Binary Code model that utilizes the class binary codes and show that we significantly outperform the state-of-the-art. We also propose a new class-based Hamming metric that dramatically reduces the retrieval times for larger databases and also improves the performance of the method by large margins.
Sravanthi Bondugula, Larry Davis 0001
ICMR2
2016 Object detection in 20 questions
abstract
We propose a novel general strategy for object detection. Instead of passively evaluating all object detectors at all possible locations in an image, we develop a divide-and-conquer approach by actively and sequentially evaluating contextual cues related to the query based on the scene and previous evaluations — like playing a "20 Questions" game — to decide where to search for the object. We formulate the problem as a Markov Decision Process and learn a search policy by reinforcement learning. To demonstrate the efficacy of our generic algorithm, we apply the 20 questions approach in the recent framework of simultaneous object detection and segmentation. Experimental results on the Pascal VOC dataset show that our algorithm reduces about 45.3% of the object proposals and 36% of average evaluation time while achieving better average precision compared to exhaustive search.
Xi Stephen Chen, He He 0001, Larry Davis 0001
WACV3
2016 Special Issue on Individual and Group Activities in Video Event Analysis
Liang Wang 0001, Ioannis Patras, Jian Zhang 0002, Greg Mori, Larry Davis 0001
Comput. Vis. Image Underst.5
2016 Guest Editorial: Large Scale Visual Media Geo-Localization
Riad I. Hammoud, Josef Sivic, Larry Davis 0001, Marc Pollefeys
Int. J. Comput. Vis.3
2016 Multi-Directional Multi-Level Dual-Cross Patterns for Robust Face Recognition
abstract
To perform unconstrained face recognition robust to variations in illumination, pose and expression, this paper presents a new scheme to extract "Multi-Directional Multi-Level Dual-Cross Patterns" (MDML-DCPs) from face images. Specifically, the MDML-DCPs scheme exploits the first derivative of Gaussian operator to reduce the impact of differences in illumination and then computes the DCP feature at both the holistic and component levels. DCP is a novel face image descriptor inspired by the unique textural structure of human faces. It is computationally efficient and only doubles the cost of computing local binary patterns, yet is extremely robust to pose and expression variations. MDML-DCPs comprehensively yet efficiently encodes the invariant characteristics of a face image from multiple levels into patterns that are highly discriminative of inter-personal differences but robust to intra-personal variations. Experimental results on the FERET, CAS-PERL-R1, FRGC 2.0, and LFW databases indicate that DCP outperforms the state-of-the-art local descriptors (e.g., LBP, LTP, LPQ, POEM, tLBP, and LGXP) for both face identification and face verification tasks. More impressively, the best performance is achieved on the challenging LFW and FRGC 2.0 databases by deploying MDML-DCPs in a simple recognition scheme.
Changxing Ding, Dacheng Tao, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Joint Image Clustering and Labeling by Matrix Factorization
abstract
We propose a novel algorithm to cluster and annotate a set of input images jointly, where the images are clustered into several discriminative groups and each group is identified with representative labels automatically. For these purposes, each input image is first represented by a distribution of candidate labels based on its similarity to images in a labeled reference image database. A set of these label-based representations are then refined collectively through a non-negative matrix factorization with sparsity and orthogonality constraints; the refined representations are employed to cluster and annotate the input images jointly. The proposed approach demonstrates performance improvements in image clustering over existing techniques, and illustrates competitive image labeling accuracy in both quantitative and qualitative evaluation. In addition, we extend our joint clustering and labeling framework to solving the weakly-supervised image classification problem and obtain promising results.
Seunghoon Hong, Jan Feyereisl, Bohyung Han, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2016 A non-parametric approach to extending generic binary classifiers for multi-classification
Venkataraman Santhanam, Vlad I. Morariu, David Harwood, Larry Davis 0001
Pattern Recognit.4
2015 An end-to-end system for content-based video retrieval using behavior, actions, and appearance with interactive query refinement
abstract
We describe a system for content-based retrieval from large surveillance video archives, using behavior, action and appearance of objects. Objects are detected, tracked, and classified into broad categories. Their behavior and appearance are characterized by action detectors and descriptors, which are indexed in an archive. Queries can be posed as video exemplars, and the results can be refined through relevance feedback. The contributions of our system include the fusion of behavior and action detectors with appearance for matching; the improvement of query results through interactive query refinement (IQR), which learns a discriminative classifier online based on user feedback; and reasonable performance on low resolution, poor quality video. The system operates on video from ground cameras and aerial platforms, both RGB and IR. Performance is evaluated on publicly-available surveillance datasets, showing that subtle actions can be detected under difficult conditions, with reasonable improvement from IQR.
Anthony Hoogs, A. G. Amitha Perera, Roderic Collins, Arslan Basharat, Keith Fieldhouse, Chuck Atkins, Linus Sherrill, Benjamin Boeckel, Russell Blue, Matthew Woehlke, C. Greco, Zhaohui Sun, Eran Swears, Naresh P. Cuntoor, J. Luck, B. Drew, D. Hanson, D. Rowley, J. Kopaz, T. Rude, D. Keefe, Amit Srivastava, Saurabh Khanwalkar, Chia-Chih Chen, Jake K. Aggarwal, Larry Davis 0001, Yaser Yacoob, Dong Liu 0001, Shih-Fu Chang, Bi Song, Amit K. Roy-Chowdhury, Kenneth Sullivan, Jelena Tesic, Shivkumar Chandrasekaran, B. S. Manjunath, K. Reddy, Mubarak Shah, K. Chang, Tsuhan Chen, Mita Desai
AVSS27
2015 Searching for Objects using Structure in Indoor Scenes
abstract
To identify the location of objects of a particular class, a passive computer vision system generally processes all the regions in an image to finally output few regions. However, we can use structure in the scene to search for objects without processing the entire image. We propose a search technique that sequentially processes image regions such that the regions that are more likely to correspond to the query class object are explored earlier. We frame the problem as a Markov decision process and use an imitation learning algorithm to learn a search strategy. Since structure in the scene is essential for search, we work with indoor scene images as they contain both unary scene context information and object-object context in the scene. We perform experiments on the NYU-depth v2 dataset and show that the unary scene context features alone can achieve a significantly high average precision while processing only 20-25\% of the regions for classes like bed and sofa. By considering object-object context along with the scene context features, the performance is further improved for classes like counter, lamp, pillow and sofa.
Varun K. Nagaraja, Vlad I. Morariu, Larry Davis 0001
BMVC3
2015 Class consistent multi-modal fusion with binary features
abstract
Many existing recognition algorithms combine different modalities based on training accuracy but do not consider the possibility of noise at test time. We describe an algorithm that perturbs test features so that all modalities predict the same class. We enforce this perturbation to be as small as possible via a quadratic program (QP) for continuous features, and a mixed integer program (MIP) for binary features. To efficiently solve the MIP, we provide a greedy algorithm and empirically show that its solution is very close to that of a state-of-the-art MIP solver. We evaluate our algorithm on several datasets and show that the method outperforms existing approaches.
Ashish Shrivastava 0001, Mohammad Rastegari, Rama Chellappa, Larry Davis 0001
CVPR5
2015 Selective Encoding for Recognizing Unreliably Localized Faces
abstract
Most existing face verification systems rely on precise face detection and registration. However, these two components are fallible under unconstrained scenarios (e.g., mobile face authentication) due to partial occlusions, pose variations, lighting conditions and limited view-angle coverage of mobile cameras. We address the unconstrained face verification problem by encoding face images directly without any explicit models of detection or registration. We propose a selective encoding framework which injects relevance information (e.g., foreground/background probabilities) into each cluster of a descriptor codebook. An additional selector component also discards distractive image patches and improves spatial robustness. We evaluate our framework using Gaussian mixture models and Fisher vectors on challenging face verification datasets. We apply selective encoding to Fisher vector features, which in our experiments degrade quickly with inaccurate face localization, our framework improves robustness with no extra test time computation. We also apply our approach to mobile based active face authentication task, demonstrating its utility in real scenarios.
Ang Li 0001, Vlad I. Morariu, Larry Davis 0001
ICCV3
2015 Selecting Relevant Web Trained Concepts for Automated Event Retrieval
abstract
Complex event retrieval is a challenging research problem, especially when no training videos are available. An alternative to collecting training videos is to train a large semantic concept bank a priori. Given a text description of an event, event retrieval is performed by selecting concepts linguistically related to the event description and fusing the concept responses on unseen videos. However, defining an exhaustive concept lexicon and pre-training it requires vast computational resources. Therefore, recent approaches automate concept discovery and training by leveraging large amounts of weakly annotated web data. Compact visually salient concepts are automatically obtained by the use of concept pairs or, more generally, n-grams. However, not all visually salient n-grams are necessarily useful for an event query -- some combinations of concepts may be visually compact but irrelevant -- and this drastically affects performance. We propose an event retrieval algorithm that constructs pairs of automatically discovered concepts and then prunes those concepts that are unlikely to be helpful for retrieval. Pruning depends both on the query and on the specific video instance being evaluated. Our approach also addresses calibration and domain adaptation issues that arise when applying concept detectors to unseen videos. We demonstrate large improvements over other vision based systems on the TRECVID MED 13 dataset.
Xintong Han, Zhe Wu 0001, Vlad I. Morariu, Larry Davis 0001
ICCV5
2015 Multi-Task Learning with Low Rank Attribute Embedding for Person Re-Identification
abstract
We propose a novel Multi-Task Learning with Low Rank Attribute Embedding (MTL-LORAE) framework for person re-identification. Re-identifications from multiple cameras are regarded as related tasks to exploit shared information to improve re-identification accuracy. Both low level features and semantic/data-driven attributes are utilized. Since attributes are generally correlated, we introduce a low rank attribute embedding into the MTL formulation to embed original binary attributes to a continuous attribute space, where incorrect and incomplete attributes are rectified and recovered to better describe people. The learning objective function consists of a quadratic loss regarding class labels and an attribute embedding error, which is solved by an alternating optimization procedure. Experiments on three person re-identification datasets have demonstrated that MTL-LORAE outperforms existing approaches by a large margin and produces state-of-the-art results.
Chi Su, Fan Yang 0016, Shiliang Zhang, Qi Tian 0001, Larry Davis 0001, Wen Gao 0001
ICCV5
2015 SHOE: Sibling Hashing with Output Embeddings
abstract
We present a supervised binary encoding scheme for image retrieval that learns projections by taking into account similarity between classes obtained from output embeddings. Our motivation is that binary hash codes learned in this way improve the visual quality of retrieval results by ranking related (or ``sibling'') class images before unrelated class images. We employ a sequential greedy optimization that learns relationship aware projections by minimizing the difference between inner products of binary codes and output embedding vectors. We develop a joint optimization framework to learn projections which improve the accuracy of supervised hashing over the current state of the art with respect to standard and sibling evaluation metrics. We further obtain discriminative features learned from correlations of kernelized input CNN features and output embeddings, which significantly boosts performance. Experiments are performed on three datasets: CUB-2011, SUN-Attribute and ImageNet ILSVRC 2010, where we show significant improvement in sibling performance metrics over state-of-the-art supervised hashing techniques, while maintaining performance with respect to standard metrics.
Sravanthi Bondugula, Varun Manjunatha, Larry Davis 0001, David S. Doermann
ACM Multimedia3
2015 Collaborative Fashion Recommendation: A Functional Tensor Factorization Approach
abstract
With the rapid expansion of online shopping for fashion products, effective fashion recommendation has become an increasingly important problem. In this work, we study the problem of personalized outfit recommendation, i.e. automatically suggesting outfits to users that fit their personal fashion preferences. Unlike existing recommendation systems that usually recommend individual items, we suggest sets of items, which interact with each other, to users. We propose a functional tensor factorization method to model the interactions between user and fashion items. To effectively utilize the multi-modal features of the fashion items, we use a gradient boosting based method to learn nonlinear functions to map the feature vectors from the feature space into some low dimensional latent space. The effectiveness of the proposed algorithm is validated through extensive experiments on real world user data from a popular fashion-focused social network.
Yang Hu 0006, Xi Yi, Larry Davis 0001
ACM Multimedia3
2015 Hierarchical Spherical Hashing for Compressing High Dimensional Vectors
abstract
We present a hierarchical approach to compress large dimensional vectors using hyper spherical hashing functions. We provide a practical solution for learning hyper spherical hashing functions by partitioning the vectors and learning hyper spheres in subspaces. Our method is an efficient way to preserve the hashing properties of sub-space hashing functions to generate the full-hashing functions in a divide and conquer fashion. We demonstrate the performance of our approach on the ILSVRC2010 Validation dataset and two large scale datasets: ILSVRC2010 Train and Holidays+Flickr 1M with high dimensional representations of size 128000, 25600 and 12800 respectively. Our results highlight the compact nature of hyper spherical hashing functions which significantly outperform the state-of-the art methods at compression ratios of 512, 256 and 128. Furthermore, we boost the retrieval performance by introducing an as symetric distance for spherical hashing functions.
Sravanthi Bondugula, Larry Davis 0001
WACV2
2015 Clauselets: Leveraging Temporally Related Actions for Video Event Analysis
abstract
We propose clause lets, sets of concurrent actions and their temporal relationships, and explore their application to video event analysis. We train clause lets in two stages. We initially train first level clause let detectors that find a limited set of actions in particular qualitative temporal configurations based on Allen's interval relations. In the second stage, we apply the first level detectors to training videos, and discriminatively learn temporal patterns between activations that involve more actions over longer durations and lead to improved second level clause let models. We demonstrate the utility of clause lets by applying them to the task of "in-the-wild" video event recognition on the TRECVID MED 11 dataset. Not only do clause lets achieve state-of-the-art results on this task, but qualitative results suggest that they may also lead to semantically meaningful descriptions of videos in terms of detected actions and their temporal relationships.
Hyungtae Lee, Vlad I. Morariu, Larry Davis 0001
WACV3
2015 Unsupervised Feature Extraction Inspired by Latent Low-Rank Representation
abstract
Latent Low-Rank Representation (Lat LRR) has the empirical capability of identifying "salient" features. However, the reason behind this feature extraction effect is still not understood. Its optimization leads to non-unique solutions and has high computational complexity, limiting its potential in practice. We show that Lat LRR learns a transformation matrix which suppresses the most significant principal components corresponding to the largest singular values while preserving the details captured by the components with relatively smaller singular values. Based on this, we propose a novel feature extraction method which directly designs the transformation matrix and has similar behavior to Lat LRR. Our method has a simple analytical solution and can achieve better performance with little computational cost. The effectiveness and efficiency of our method are validated on two face recognition datasets.
Yaming Wang, Vlad I. Morariu, Larry Davis 0001
WACV3
2015 Re-ranking by Multi-feature Fusion with Diffusion for Image Retrieval
abstract
We present a re-ranking algorithm for image retrieval by fusing multi-feature information. We utilize pair wise similarity scores between images to exploit the underlying relationships among images. The initial ranked list for a query from each feature is represented as an undirected graph, where edge strength comes from feature-specific image similarity. Graphs from multiple features are combined by a mixture Markov model. In addition, we utilize a probabilistic model based on the statistics of similarity scores of similar and dissimilar image pairs to determine the weight for each graph. The weight for a feature is query specific, where the ranked lists of different queries receive different weights. Our approach for calculating weights is data-driven and does not require any learning. A diffusion process is then applied to the fused graph to reduce noise and achieve better retrieval performance. Experiments demonstrate that our approach significantly improves performance over baseline methods and outperforms many state-of-the-art retrieval methods.
Fan Yang 0016, Bogdan Matei, Larry Davis 0001
WACV3
2015 Learning predictable binary codes for face indexing
Ran He 0001, Yinghao Cai, Tieniu Tan, Larry Davis 0001
Pattern Recognit.4
2014 PSPGC: Part-Based Seeds for Parametric Graph-Cuts
Xintong Han, Zhe Wu 0001, Larry Davis 0001
ACCV (3)4
2014 Submodular Reranking with Multiple Feature Modalities for Image Retrieval
Fan Yang 0016, Zhuolin Jiang, Larry Davis 0001
ACCV (1)3
2014 Jointly Learning Dictionaries and Subspace Structure for Video-Based Face Recognition
Guangxiao Zhang, Ran He 0001, Larry Davis 0001
ACCV (3)3
2014 Planar Structure Matching under Projective Uncertainty for Geolocation
Ang Li 0001, Vlad I. Morariu, Larry Davis 0001
ECCV (7)3
2014 Jointly Optimizing 3D Model Fitting and Fine-Grained Classification
Yen-Liang Lin, Vlad I. Morariu, Winston H. Hsu, Larry Davis 0001
ECCV (4)4
2014 Toward Sparse Coding on Cosine Distance
abstract
Sparse coding is a regularized least squares solution using the L1or L0constraint, based on the Euclidean distance between original and reconstructed signals with respect to a predefined dictionary. The Euclidean distance, however, is not a good metric for many feature descriptors, especially histogram features, e.g. many visual features including SIFT, HOG, LBP and Bag-of-visual-words. In contrast, cosine distance is a more appropriate metric for such features. To leverage the benefit of the cosine distance in sparse coding, we formulate a new sparse coding objective function based on approximate cosine distance by constraining a norm of the reconstructed signal to be close to the norm of the original signal. We evaluate our new formulation on three computer vision datasets (UCF101 Action dataset, AR dataset and Extended YaleB dataset) and show improvements over the Euclidean distance based objective.
Hyunjong Cho, Jungsuk Kwac, Larry Davis 0001
ICPR4
2014 Toward a non-intrusive, physio- behavioral biometric for smartphones
abstract
Biometric authentication relies on an individual's inner characteristics and traits. We propose an active authentication system on a mobile device that relies on two biometric modalities: 3D gestures and face recognition. The novelty of our approach is to combine 3D gesture and face recognition in a nonintrusive and unconstrained environment; the active authentication system is running in the background while the user is performing his/her main task.
Esther Vasiete, Yan Chen 0033, Ian Char, Tom Yeh, Vishal M. Patel, Larry Davis 0001, Rama Chellappa
Mobile HCI6
2014 Multi-Modal Image Retrieval for Complex Queries using Small Codes
abstract
We propose a unified framework for image retrieval capable of handling complex and descriptive queries of multiple modalities in a scalable manner. A novel aspect of our approach is that it supports query specification in terms of objects, attributes and spatial relationships, thereby allowing for substantially more complex and descriptive queries. We allow these complex queries to be specified in three different modalities - images, sketches and structured textual descriptions. Furthermore, we propose a unique multi-modal hashing algorithm capable of mapping queries of different modalities to the same binary representation, enabling efficient and scalable image retrieval based on multi-modal queries. Extensive experimental evaluation shows that our approach outperforms the state-of-the-art image retrieval and hashing techniques on the MSRC and SUN09 datasets by about 100%, while the performance on a dataset of 1M images, from Flickr, demonstrates its scalability.
Behjat Siddiquie, Brandyn White, Abhishek Sharma 0001, Larry Davis 0001
ICMR4
2014 A Probabilistic Framework for Multimodal Retrieval using Integrative Indian Buffet Process
Bahadir Ozdemir, Larry Davis 0001
NIPS2
2014 Object co-labeling in multiple images
abstract
We introduce a new problem called object co-labeling where the goal is to jointly annotate multiple images of the same scene which do not have temporal consistency. We present an adaptive framework for joint segmentation and recognition to solve this problem. We propose an objective function that considers not only appearance but also appearance and context consistency across images of the scene. A relaxed form of the cost function is minimized using an efficient quadratic programming solver. Our approach improves labeling performance compared to labeling each image individually. We also show the application of our co-labeling framework to other recognition problems such as label propagation in videos and object recognition in similar scenes. Experimental results demonstrates the efficacy of our approach.
Xi Chen 0016, Larry Davis 0001
WACV3
2014 Interactive video segmentation using occlusion boundaries and temporally coherent superpixels
abstract
We propose an interactive video segmentation system built on the basis of occlusion and long term spatio-temporal structure cues. User supervision is incorporated in a superpixel graph clustering framework that differs crucially from prior art in that it modifies the graph according to the output of an occlusion boundary detector. Working with long temporal intervals (up to 100 frames) enables our system to significantly reduce annotation effort with respect to state of the art systems. Even though the segmentation results are less than perfect, they are obtained efficiently and can be used in weakly supervised learning from video or for video content description. We do not rely on a discriminative object appearance model and allow extracting multiple foreground objects together, saving user time if more than one object is present. Additional experiments with unsupervised clustering based on occlusion boundaries demonstrate the importance of this cue for video segmentation and thus validate our system design.
Radu Dondera, Vlad I. Morariu, Yulu Wang, Larry Davis 0001
WACV4
2014 Composite Discriminant Factor analysis
abstract
We propose a linear dimensionality reduction method, Composite Discriminant Factor (CDF) analysis, which searches for a discriminative but compact feature subspace that can be used as input to classifiers that suffer from problems such as multi-collinearity or the curse of dimensionality. The subspace selected by CDF maximizes the performance of the entire classification pipeline, and is chosen from a set of candidate subspaces that are each discriminative. Our method is based on Partial Least Squares (PLS) analysis, and can be viewed as a generalization of the PLS1 algorithm, designed to increase discrimination in classification tasks. We demonstrate our approach on the UCF50 action recognition dataset, two object detection datasets (INRIA pedestrians and vehicles from aerial imagery), and machine learning datasets from the UCI Machine Learning repository. Experimental results show that the proposed approach improves significantly in terms of accuracy over linear SVM, and also over PLS in terms of compactness and efficiency, while maintaining or improving accuracy.
Vlad I. Morariu, Ejaz Ahmed 0002, Venkataraman Santhanam, David Harwood, Larry Davis 0001
WACV5
2014 Online discriminative dictionary learning for visual tracking
abstract
Dictionary learning has been applied to various computer vision problems, such as image restoration, object classification and face recognition. In this work, we propose a tracking framework based on sparse representation and online discriminative dictionary learning. By associating dictionary items with label information, the learned dictionary is both reconstructive and discriminative, which better distinguishes target objects from the background. During tracking, the best target candidate is selected by a joint decision measure. Reliable tracking results and augmented training samples are accumulated into two sets to update the dictionary. Both online dictionary learning and the proposed joint decision measure are important for the final tracking performance. Experiments show that our approach outperforms several recently proposed trackers.
Fan Yang 0016, Zhuolin Jiang, Larry Davis 0001
WACV3
2014 Special issue on background modeling for foreground detection in real-world dynamic scenes
Thierry Bouwmans, Jordi Gonzàlez 0001, Caifeng Shan, Massimo Piccardi, Larry Davis 0001
Mach. Vis. Appl.5
2014 Screen-based active user authentication
Mohammed E. Fathy 0001, Vishal M. Patel, Tom Yeh, Yangmuzi Zhang, Rama Chellappa, Larry Davis 0001
Pattern Recognit. Lett.6
2013 System and algorithms on detection of objects embedded in perspective geometry using monocular cameras
abstract
In this work, we present a framework to detect objects embedded in complex perspective geometry. Our goal is to accurately identify objects such as people standing in balconies or windows on building facades of surrounding buildings. Compared to traditional computer vision work focused on activity analysis from a horizontal view, our framework provides a solution for the application domain of mobile surveillance in urban areas. A novel solution for a monocular camera is formulated by tightly coupling various computational modules including geometric analysis, segmentation, scale estimation, and object detection. In particular, our proposed approach alleviates the effect of the perspective geometry and corresponding distortion in object appearance effectively, and provides accurate scale priors to eliminate unlikely object detection hypotheses. The experimental results on collected video dataset show that the proposed approach is more accurate than traditional detection approaches based on brute-force scanning windows.
Yiliang Xu, Sangmin Oh, Fan Yang 0016, Zhuolin Jiang, Naresh P. Cuntoor, Anthony Hoogs, Larry Davis 0001
AVSS7
2013 Discriminative Tensor Sparse Coding for Image Classification
abstract
A novel approach to learn a discriminative dictionary over a tensor sparse model is presented. A structural incoherence constraint between dictionary atoms from different classes is introduced to promote discriminating information into the dictionary. The incoherence term encourages dictionary atoms to be as independent as possible. In addition, we incorporate classification error into the objective function of dictionary learning. The dictionary is learned in a supervised setting to make it useful for classification. A linear multi-class classifier and the dictionary are learned simultaneously during the training phase. Our approach is evaluated on three types of public databases, including texture, digit, and face databases. Experimental results demonstrate the effectiveness of our approach. 1
Yangmuzi Zhang, Zhuolin Jiang, Larry Davis 0001
BMVC3
2013 Adding Unlabeled Samples to Categories by Learned Attributes
abstract
We propose a method to expand the visual coverage of training sets that consist of a small number of labeled examples using learned attributes. Our optimization formulation discovers category specific attributes as well as the images that have high confidence in terms of the attributes. In addition, we propose a method to stably capture example-specific attributes for a small sized training set. Our method adds images to a category from a large unlabeled image pool, and leads to significant improvement in category recognition accuracy evaluated on a large-scale dataset, Image Net.
Mohammad Rastegari, Ali Farhadi, Larry Davis 0001
CVPR4
2013 Representing Videos Using Mid-level Discriminative Patches
abstract
How should a video be represented? We propose a new representation for videos based on mid-level discriminative spatio-temporal patches. These spatio-temporal patches might correspond to a primitive human action, a semantic object, or perhaps a random but informative spatio-temporal patch in the video. What defines these spatio-temporal patches is their discriminative and representative properties. We automatically mine these patches from hundreds of training videos and experimentally demonstrate that these patches establish correspondence across videos and align the videos for label transfer techniques. Furthermore, these patches can be used as a discriminative vocabulary for action classification where they demonstrate state-of-the-art performance on UCF50 and Olympics datasets.
Abhinav Gupta 0001, Mikel Rodriguez, Larry Davis 0001
CVPR4
2013 Submodular Salient Region Detection
abstract
The problem of salient region detection is formulated as the well-studied facility location problem from operations research. High-level priors are combined with low-level features to detect salient regions. Salient region detection is achieved by maximizing a sub modular objective function, which maximizes the total similarities (i.e., total profits) between the hypothesized salient region centers (i.e., facility locations) and their region elements (i.e., clients), and penalizes the number of potential salient regions (i.e., the number of open facilities). The similarities are efficiently computed by finding a closed-form harmonic solution on the constructed graph for an input image. The saliency of a selected region is modeled in terms of appearance and spatial location. By exploiting the sub modularity properties of the objective function, a highly efficient greedy-based optimization algorithm can be employed. This algorithm is guaranteed to be at least a (e - 1)/e 0.632-approximation to the optimum. Experimental results demonstrate that our approach outperforms several recently proposed saliency detection approaches.
Zhuolin Jiang, Larry Davis 0001
CVPR2
2013 Learning Structured Low-Rank Representations for Image Classification
abstract
An approach to learn a structured low-rank representation for image classification is presented. We use a supervised learning method to construct a discriminative and reconstructive dictionary. By introducing an ideal regularization term, we perform low-rank matrix recovery for contaminated training data from all categories simultaneously without losing structural information. A discriminative low-rank representation for images with respect to the constructed dictionary is obtained. With semantic structure information and strong identification capability, this representation is good for classification tasks even using a simple linear multi-classifier. Experimental results demonstrate the effectiveness of our approach.
Yangmuzi Zhang, Zhuolin Jiang, Larry Davis 0001
CVPR3
2013 Sampling for unsupervised domain adaptive object detection
abstract
We explore the problem of extreme class imbalance present when performing fully unsupervised domain adaptation for object detection. The main challenge arises from the fact that images in unconstrained settings are mostly occupied by the background (negative class). Therefore, random sampling will not typically result in a sufficient number of positive samples from the target domain, which is required by domain adaptation methods. Motivated by traditional semi-supervised learning algorithms that aim for better classification using both labeled and unlabeled data, we propose a variation of co-learning technique that automatically constructs a more balanced set of samples from the target domain. We evaluate the effectiveness of our approach using a vehicle detection task in an urban surveillance dataset. Furthermore, we compare the performance of our technique with two other approaches-one based on unbiased learning on multiple training data sets and the other on self-learning.
Fatemeh Mirrashed, Vlad I. Morariu, Larry Davis 0001
ICIP3
2013 Predictable Dual-View Hashing
abstract
We propose a Predictable Dual-View Hashing (PDH) algorithm which embeds proximity of data samples in the original spaces. We create a cross-view hamming space with the ability to compare information from previously incomparable domains with a notion of ‘predictability’. By performing comparative experimental analysis on two large datasets, PASCAL-Sentence and SUN-Attribute, we demonstrate the superiority of our method to the state-of-the-art dual-view binary code learning algorithms.
Mohammad Rastegari, Shobeir Fakhraei, Hal Daumé III, Larry Davis 0001
ICML (3)5
2013 Domain adaptive object detection
abstract
We study the use of domain adaptation and transfer learning techniques as part of a framework for adaptive object detection. Unlike recent applications of domain adaptation work in computer vision, which generally focus on image classification, we explore the problem of extreme class imbalance present when performing domain adaptation for object detection. The main difficulty caused by this imbalance is that test images contain millions or billions of negative image subwindows but just a few image subwindows containing positive instances, which makes it difficult to adapt to changes in the positive classes present new domains by simple techniques such as random sampling. We propose an initial approach to addressing this problem and apply our technique to vehicle detection in a challenging urban surveillance dataset, demonstrating the performance of our approach with various amounts of supervision, including the fully unsupervised case.
Fatemeh Mirrashed, Vlad I. Morariu, Behjat Siddiquie, Rogério Feris, Larry Davis 0001
WACV5
2013 A unified tree-based framework for joint action localization, recognition and segmentation
Zhuolin Jiang, Zhe Lin 0001, Larry Davis 0001
Comput. Vis. Image Underst.3
2013 A data-driven detection optimization framework
William Robson Schwartz, Victor C. de Melo, Hélio Pedrini, Larry Davis 0001
Neurocomputing4
2013 Label Consistent K-SVD: Learning a Discriminative Dictionary for Recognition
abstract
A label consistent K-SVD (LC-KSVD) algorithm to learn a discriminative dictionary for sparse coding is presented. In addition to using class labels of training data, we also associate label information with each dictionary item (columns of the dictionary matrix) to enforce discriminability in sparse codes during the dictionary learning process. More specifically, we introduce a new label consistency constraint called "discriminative sparse-code error" and combine it with the reconstruction error and the classification error to form a unified objective function. The optimal solution is efficiently obtained using the K-SVD algorithm. Our algorithm learns a single overcomplete dictionary and an optimal linear classifier jointly. The incremental dictionary learning algorithm is presented for the situation of limited memory resources. It yields dictionaries so that feature points with the same class labels have similar sparse codes. Experimental results demonstrate that our algorithm outperforms many recently proposed sparse-coding techniques for face, action, scene, and object category recognition under the same learning conditions.
Zhuolin Jiang, Zhe Lin 0001, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Tracking People's Hands and Feet Using Mixed Network AND/OR Search
abstract
We describe a framework that leverages mixed probabilistic and deterministic networks and their AND/OR search space to efficiently find and track the hands and feet of multiple interacting humans in 2D from a single camera view. Our framework detects and tracks multiple people's heads, hands, and feet through partial or full occlusion; requires few constraints (does not require multiple views, high image resolution, knowledge of performed activities, or large training sets); and makes use of constraints and AND/OR Branch-and-Bound with lazy evaluation and carefully computed bounds to efficiently solve the complex network that results from the consideration of interperson occlusion. Our main contributions are: 1) a multiperson part-based formulation that emphasizes extremities and allows for the globally optimal solution to be obtained in each frame, and 2) an efficient and exact optimization scheme that relies on AND/OR Branch-and-Bound, lazy factor evaluation, and factor cost sensitive bound computation. We demonstrate our approach on three datasets: the public single person HumanEva dataset, outdoor sequences where multiple people interact in a group meeting scenario, and outdoor one-on-one basketball videos. The first dataset demonstrates that our framework achieves state-of-the-art performance in the single person setting, while the last two demonstrate robustness in the presence of partial and full occlusion and fast nontrivial motion.
Vlad I. Morariu, David Harwood, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 Discriminative Dictionary Learning with Pairwise Constraints
Huimin Guo, Zhuolin Jiang, Larry Davis 0001
ACCV (1)3
2012 Qualitative Pose Estimation by Discriminative Deformable Part Models
Hyungtae Lee, Vlad I. Morariu, Larry Davis 0001
ACCV (2)3
2012 Online Semi-Supervised Discriminative Dictionary Learning for Sparse Representation
Guangxiao Zhang, Zhuolin Jiang, Larry Davis 0001
ACCV (1)3
2012 On partial least squares in head pose estimation: How to simultaneously deal with misalignment
abstract
Head pose estimation is a critical problem in many computer vision applications. These include human computer interaction, video surveillance, face and expression recognition. In most prior work on heads pose estimation, the positions of the faces on which the pose is to be estimated are specified manually. Therefore, the results are reported without studying the effect of misalignment. We propose a method based on partial least squares (PLS) regression to estimate pose and solve the alignment problem simultaneously. The contributions of this paper are two-fold: 1) we show that the kernel version of PLS (kPLS) achieves better than state-of-the-art results on the estimation problem and 2) we develop a technique to reduce misalignment based on the learned PLS factors.
Murad Al Haj, Jordi Gonzàlez 0001, Larry Davis 0001
CVPR3
2012 Submodular dictionary learning for sparse coding
abstract
A greedy-based approach to learn a compact and discriminative dictionary for sparse representation is presented. We propose an objective function consisting of two components: entropy rate of a random walk on a graph and a discriminative term. Dictionary learning is achieved by finding a graph topology which maximizes the objective function. By exploiting the monotonicity and submodularity properties of the objective function and the matroid constraint, we present a highly efficient greedy-based optimization algorithm. It is more than an order of magnitude faster than several recently proposed dictionary learning approaches. Moreover, the greedy algorithm gives a near-optimal solution with a (1/2)-approximation bound. Our approach yields dictionaries having the property that feature points from the same class have very similar sparse codes. Experimental results demonstrate that our approach outperforms several recently proposed dictionary learning techniques for face, action and object category recognition.
Zhuolin Jiang, Guangxiao Zhang, Larry Davis 0001
CVPR3
2012 A flow model for joint action recognition and identity maintenance
abstract
We propose a framework that performs action recognition and identity maintenance of multiple targets simultaneously. Instead of first establishing tracks using an appearance model and then performing action recognition, we construct a network flow-based model that links detected bounding boxes across video frames while inferring activities, thus integrating identity maintenance and action recognition. Inference in our model reduces to a constrained minimum cost flow problem, which we solve exactly and efficiently. By leveraging both appearance similarity and action transition likelihoods, our model improves on state-of-the-art results on action recognition for two datasets.
Sameh Khamis, Vlad I. Morariu, Larry Davis 0001
CVPR3
2012 Covariance discriminative learning: A natural and efficient approach to image set classification
abstract
We propose a novel discriminative learning approach to image set classification by modeling the image set with its natural second-order statistic, i.e. covariance matrix. Since nonsingular covariance matrices, a.k.a. symmetric positive definite (SPD) matrices, lie on a Riemannian manifold, classical learning algorithms cannot be directly utilized to classify points on the manifold. By exploring an efficient metric for the SPD matrices, i.e., Log-Euclidean Distance (LED), we derive a kernel function that explicitly maps the covariance matrix from the Riemannian manifold to a Euclidean space. With this explicit mapping, any learning method devoted to vector space can be exploited in either its linear or kernel formulation. Linear Discriminant Analysis (LDA) and Partial Least Squares (PLS) are considered in this paper for their feasibility for our specific problem. We further investigate the conventional linear subspace based set modeling technique and cast it in a unified framework with our covariance matrix based modeling. The proposed method is evaluated on two tasks: face recognition and object categorization. Extensive experimental results show not only the superiority of our method over state-of-the-art ones in both accuracy and efficiency, but also its stability to two real challenges: noisy set data and varying set size.
Ruiping Wang 0001, Huimin Guo, Larry Davis 0001, Qionghai Dai
CVPR3
2012 Combining Per-frame and Per-track Cues for Multi-person Action Recognition
Sameh Khamis, Vlad I. Morariu, Larry Davis 0001
ECCV (1)3
2012 Multi-scale shared features for cascade object detection
abstract
We introduce an efficient computational framework to extract multi-scale feature descriptors. The framework is based on sharing of descriptor elements across the image and scale space to minimize redundant computation. Any type of local patch or grid-based features can be computed through this framework for capturing coarse-to-fine object appearances. We apply it to human detection by boosting a strong soft cascade classifier. Our experiments demonstrate that the proposed descriptors achieve superior performance both in computational efficiency and detection accuracy.
Zhe Lin 0001, Gang Hua 0001, Larry Davis 0001
ICIP3
2012 Unsupervised model selection for view-invariant object detection in surveillance environments
Behjat Siddiquie, Rogério Feris, Ankur Datta, Larry Davis 0001
ICPR4
2012 A complementary local feature descriptor for face identification
abstract
In many descriptors, spatial intensity transforms are often packed into a histogram or encoded into binary strings to be insensitive to local misalignment and compact. Discriminative information, however, might be lost during the process as a trade-off. To capture the lost pixel-wise local information, we propose a new feature descriptor, Circular Center Symmetric-Pairs of Pixels (CCS-POP). It concatenates the symmetric pixel differences centered at a pixel position along various orientations with various radii; it is a generalized form of Local Binary Patterns, its variants and Pairs-of-Pixels (POP). Combining CCS-POP with existing descriptors achieves better face identification performance on FRGC Ver. 1.0 and FERET datasets compared to state-of-the-art approaches.
William Robson Schwartz, Huimin Guo, Larry Davis 0001
WACV4
2012 Class consistent k-means: Application to face and action recognition
Zhuolin Jiang, Zhe Lin 0001, Larry Davis 0001
Comput. Vis. Image Underst.3
2012 Robust pose invariant face recognition using coupled latent space discriminant analysis
Abhishek Sharma 0001, Murad Al Haj, Larry Davis 0001, David Jacobs 0001
Comput. Vis. Image Underst.4
2012 Density-Based Multifeature Background Subtraction with Support Vector Machine
abstract
Background modeling and subtraction is a natural technique for object detection in videos captured by a static camera, and also a critical preprocessing step in various high-level computer vision applications. However, there have not been many studies concerning useful features and binary segmentation algorithms for this problem. We propose a pixelwise background modeling and subtraction technique using multiple features, where generative and discriminative techniques are combined for classification. In our algorithm, color, gradient, and Haar-like features are integrated to handle spatio-temporal variations for each pixel. A pixelwise generative background model is obtained for each feature efficiently and effectively by Kernel Density Approximation (KDA). Background subtraction is performed in a discriminative manner using a Support Vector Machine (SVM) over background likelihood vectors for a set of features. The proposed algorithm is robust to shadow, illumination changes, spatial variations of background. We compare the performance of the algorithm with other density-based methods using several different feature combinations and modeling techniques, both quantitatively and qualitatively.
Bohyung Han, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 Recognizing Human Actions by Learning and Matching Shape-Motion Prototype Trees
abstract
A shape-motion prototype-based approach is introduced for action recognition. The approach represents an action as a sequence of prototypes for efficient and flexible action matching in long video sequences. During training, an action prototype tree is learned in a joint shape and motion space via hierarchical K-means clustering and each training sequence is represented as a labeled prototype sequence; then a look-up table of prototype-to-prototype distances is generated. During testing, based on a joint probability model of the actor location and action prototype, the actor is tracked while a frame-to-prototype correspondence is established by maximizing the joint probability, which is efficiently performed by searching the learned prototype tree; then actions are recognized using dynamic prototype sequence matching. Distance measures used for sequence matching are rapidly obtained by look-up table indexing, which is an order of magnitude faster than brute-force computation of frame-to-frame distances. Our approach enables robust action matching in challenging situations (such as moving cameras, dynamic backgrounds) and allows automatic alignment of action sequences. Experimental results demonstrate that our approach achieves recognition rates of 92.86 percent on a large gesture data set (with dynamic backgrounds), 100 percent on the Weizmann action data set, 95.77 percent on the KTH action data set, 88 percent on the UCF sports data set, and 87.27 percent on the CMU action data set.
Zhuolin Jiang, Zhe Lin 0001, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 Face Identification Using Large Feature Sets
abstract
With the goal of matching unknown faces against a gallery of known people, the face identification task has been studied for several decades. There are very accurate techniques to perform face identification in controlled environments, particularly when large numbers of samples are available for each face. However, face identification under uncontrolled environments or with a lack of training data is still an unsolved problem. We employ a large and rich set of feature descriptors (with more than 70,000 descriptors) for face identification using partial least squares to perform multichannel feature weighting. Then, we extend the method to a tree-based discriminative structure to reduce the time required to evaluate probe samples. The method is evaluated on Facial Recognition Technology (FERET) and Face Recognition Grand Challenge (FRGC) data sets. Experiments show that our identification method outperforms current state-of-the-art results, particularly for identifying faces acquired across varying conditions.
William Robson Schwartz, Huimin Guo, Larry Davis 0001
IEEE Trans. Image Process.4
2011 AVSS 2011 demo session: A large-scale benchmark dataset for event recognition in surveillance video
abstract
Summary form only given. We present a concept for automatic construction site monitoring by taking into account 4D information (3D over time), that is acquired from highly-overlapping digital aerial images. On the one hand today's maturity of flying micro aerial vehicles (MAVs) enables a low-cost and an efficient image acquisition of high-quality data that maps construction sites entirely from many varying viewpoints. On the other hand, due to low-noise sensors and high redundancy in the image data, recent developments in 3D reconstruction workflows have benefited the automatic computation of accurate and dense 3D scene information. Having both an inexpensive high-quality image acquisition and an efficient 3D analysis workflow enables monitoring, documentation and visualization of observed sites over time with short intervals. Relating acquired 4D site observations, composed of color, texture, geometry over time, largely supports automated methods toward full scene understanding, the acquisition of both the change and the construction site's progress.
Sangmin Oh, Anthony Hoogs, A. G. Amitha Perera, Naresh P. Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, Jake K. Aggarwal, Hyungtae Lee, Larry Davis 0001, Eran Swears, Xiaoyang Wang 0001, Kishore K. Reddy, Mubarak Shah, Carl Vondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba 0001, Bi Song, Anesco Fong, Amit K. Roy-Chowdhury, Mita Desai
AVSS10
2011 Local Response Context Applied to Pedestrian Detection
William Robson Schwartz, Larry Davis 0001, Hélio Pedrini
CIARP2
2011 Piecing together the segmentation jigsaw using context
abstract
We present an approach to jointly solve the segmentation and recognition problem using a multiple segmentation framework. We formulate the problem as segment selection from a pool of segments, assigning each selected segment a class label. Previous multiple segmentation approaches used local appearance matching to select segments in a greedy manner. In contrast, our approach formulates a cost function based on contextual information in conjunction with appearance matching. This relaxed cost function formulation is minimized using an efficient quadratic programming solver and an approximate solution is obtained by discretizing the relaxed solution. Our approach improves labeling performance compared to other segmentation based recognition approaches.
Xi Chen 0016, Abhinav Gupta 0001, Larry Davis 0001
CVPR4
2011 Learning a discriminative dictionary for sparse coding via label consistent K-SVD
abstract
A label consistent K-SVD (LC-KSVD) algorithm to learn a discriminative dictionary for sparse coding is presented. In addition to using class labels of training data, we also associate label information with each dictionary item (columns of the dictionary matrix) to enforce discriminability in sparse codes during the dictionary learning process. More specifically, we introduce a new label consistent constraint called `discriminative sparse-code error' and combine it with the reconstruction error and the classification error to form a unified objective function. The optimal solution is efficiently obtained using the K-SVD algorithm. Our algorithm learns a single over-complete dictionary and an optimal linear classifier jointly. It yields dictionaries so that feature points with the same class labels have similar sparse codes. Experimental results demonstrate that our algorithm outperforms many recently proposed sparse coding techniques for face and object category recognition under the same learning conditions.
Zhuolin Jiang, Zhe Lin 0001, Larry Davis 0001
CVPR3
2011 Multi-agent event recognition in structured scenarios
abstract
We present a framework for the automatic recognition of complex multi-agent events in settings where structure is imposed by rules that agents must follow while performing activities. Given semantic spatio-temporal descriptions of what generally happens (i.e., rules, event descriptions, physical constraints), and based on video analysis, we determine the events that occurred. Knowledge about spatio-temporal structure is encoded using first-order logic using an approach based on Allen's Interval Logic, and robustness to low-level observation uncertainty is provided by Markov Logic Networks (MLN). Our main contribution is that we integrate interval-based temporal reasoning with probabilistic logical inference, relying on an efficient bottom-up grounding scheme to avoid combinatorial explosion. Applied to one-on-one basketball, our framework detects and tracks players, their hands and feet, and the ball, generates event observations from the resulting trajectories, and performs probabilistic logical inference to determine the most consistent sequence of events. We demonstrate our approach on 1hr (100,000 frames) of outdoor videos.
Vlad I. Morariu, Larry Davis 0001
CVPR2
2011 A large-scale benchmark dataset for event recognition in surveillance video
abstract
We introduce a new large-scale video dataset designed to assess the performance of diverse visual event recognition algorithms with a focus on continuous visual event recognition (CVER) in outdoor areas with wide coverage. Previous datasets for action recognition are unrealistic for real-world surveillance because they consist of short clips showing one action by one individual [15, 8]. Datasets have been developed for movies [11] and sports [12], but, these actions and scene conditions do not apply effectively to surveillance videos. Our dataset consists of many outdoor scenes with actions occurring naturally by non-actors in continuously captured videos of the real world. The dataset includes large numbers of instances for 23 event types distributed throughout 29 hours of video. This data is accompanied by detailed annotations which include both moving object tracks and event examples, which will provide solid basis for large-scale evaluation. Additionally, we propose different types of evaluation modes for visual recognition tasks and evaluation metrics along with our preliminary experimental results. We believe that this dataset will stimulate diverse aspects of computer vision research and help us to advance the CVER tasks in the years ahead.
Sangmin Oh, Anthony Hoogs, A. G. Amitha Perera, Naresh P. Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, Jake K. Aggarwal, Hyungtae Lee, Larry Davis 0001, Eran Swears, Xiaoyang Wang 0001, Kishore K. Reddy, Mubarak Shah, Carl Vondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba 0001, Bi Song, Anesco Fong, Amit K. Roy-Chowdhury, Mita Desai
CVPR10
2011 Image ranking and retrieval based on multi-attribute queries
abstract
We propose a novel approach for ranking and retrieval of images based on multi-attribute queries. Existing image retrieval methods train separate classifiers for each word and heuristically combine their outputs for retrieving multiword queries. Moreover, these approaches also ignore the interdependencies among the query terms. In contrast, we propose a principled approach for multi-attribute retrieval which explicitly models the correlations that are present between the attributes. Given a multi-attribute query, we also utilize other attributes in the vocabulary which are not present in the query, for ranking/retrieval. Furthermore, we integrate ranking and retrieval within the same formulation, by posing them as structured prediction problems. Extensive experimental evaluation on the Labeled Faces in the Wild(LFW), FaceTracer and PASCAL VOC datasets show that our approach significantly outperforms several state-of-the-art ranking and retrieval methods.
Behjat Siddiquie, Rogério Feris, Larry Davis 0001
CVPR3
2011 Face verification using large feature sets and one shot similarity
abstract
We present a method for face verification that combines Partial Least Squares (PLS) and the One-Shot similarity model[28]. First, a large feature set combining shape, texture and color information is used to describe a face. Then PLS is applied to reduce the dimensionality of the feature set with multi-channel feature weighting. This provides a discriminative facial descriptor. PLS regression is used to compute the similarity score of an image pair by One-Shot learning. Given two feature vector representing face images, the One-Shot algorithm learns discriminative models exclusively for the vectors being compared. A small set of unlabeled images, not containing images belonging to the people being compared, is used as a reference (negative) set. The approach is evaluated on the Labeled Face in the Wild (LFW) benchmark and shows very comparable results to the state-of-the-art methods (achieving 86.12% classification accuracy) while maintaining simplicity and good generalization ability.
Huimin Guo, William Robson Schwartz, Larry Davis 0001
IJCB3
2011 Birdlets: Subordinate categorization using volumetric primitives and pose-normalized appearance
abstract
Subordinate-level categorization typically rests on establishing salient distinctions between part-level characteristics of objects, in contrast to basic-level categorization, where the presence or absence of parts is determinative. We develop an approach for subordinate categorization in vision, focusing on an avian domain due to the fine-grained structure of the category taxonomy for this domain. We explore a pose-normalized appearance model based on a volumetric poselet scheme. The variation in shape and appearance properties of these parts across a taxonomy provides the cues needed for subordinate categorization. Training pose detectors requires a relatively large amount of training data per category when done from scratch; using a subordinate-level approach, we exploit a pose classifier trained at the basic-level, and extract part appearance and shape information to build subordinate-level models. Our model associates the underlying image pattern parameters used for detection with corresponding volumetric part location, scale and orientation parameters. These parameters implicitly define a mapping from the image pixels into a pose-normalized appearance space, removing view and pose dependencies, facilitating fine-grained categorization from relatively few training examples.
Ryan Farrell, Om Oza, Ning Zhang 0014, Vlad I. Morariu, Trevor Darrell, Larry Davis 0001
ICCV6
2011 Action recognition using Partial Least Squares and Support Vector Machines
abstract
We introduce an action recognition approach based on Partial Least Squares (PLS) and Support Vector Machines (SVM). We extract very high dimensional feature vectors representing spatio-temporal properties of actions and use multiple PLS regressors to find relevant features that distinguish amongst action classes. Finally, we use a multi-class SVM to learn and classify those relevant features. We applied our approach to INRIA's IXMAS dataset. Experimental results show that our method is superior to other methods applied to the IXMAS dataset.
Samah Ramadan, Larry Davis 0001
ICIP2
2011 A novel feature descriptor based on the shearlet transform
abstract
Problems such as image classification, object detection and recognition rely on low-level feature descriptors to represent visual information. Several feature extraction methods have been proposed, including the Histograms of Oriented Gradients (HOG), which captures edge information by analyzing the distribution of intensity gradients and their directions. In addition to directions, the analysis of edge at different scales provides valuable information. Shearlet transforms provide a general framework for analyzing and representing data with anisotropic information at multiple scales. As a consequence, signal singularities, such as edges, can be precisely detected and located in images. Based on the idea of employing histograms to estimate the distribution of edge orientations and on the accurate multi-scale analysis provided by shearlet transforms, we propose a feature descriptor called Histograms of Shearlet Coefficients (HSC). Experimental results comparing HOG with HSC show that HSC provides significantly better results for the problems of texture classification and face identification.
William Robson Schwartz, Ricardo Dutra da Silva, Larry Davis 0001, Hélio Pedrini
ICIP3
2011 Creating contextual help for GUIs using screenshots
abstract
Contextual help is effective for learning how to use GUIs by showing instructions and highlights on the actual interface rather than in a separate viewer. However, end-users and third-party tech support typically cannot create contextual help to assist other users because it requires programming skill and source code access. We present a creation tool for contextual help that allows users to apply common computer skills-taking screenshots and writing simple scripts. We perform pixel analysis on screenshots to make this tool applicable to a wide range of applications and platforms without source code access. We evaluated the tool's usability with three groups of participants: developers, in-structors, and tech support. We further validated the applicability of our tool with 60 real tasks supported by the tech support of a university campus.
Tom Yeh, Tsung-Hsiang Chang, Bo Xie 0001, Greg Walsh, Ivan Watkins, Krist Wongsuphasawat, Man Huang, Larry Davis 0001, Benjamin B. Bederson
UIST8
2011 A case for query by image and text content: searching computer help using screenshots and keywords
abstract
The multimedia information retrieval community has dedicated extensive research effort to the problem of content-based image retrieval (CBIR). However, these systems find their main limitation in the difficulty of creating pictorial queries. As a result, few systems offer the option of querying by visual examples, and rely on automatic concept detection and tagging techniques to provide support for searching visual content using textual queries.
Tom Yeh, Brandyn White, José San Pedro, Boris Katz, Larry Davis 0001
WWW5
2011 Multi-Camera Tracking with Adaptive Resource Allocation
Bohyung Han, Seong-Wook Joo, Larry Davis 0001
Int. J. Comput. Vis.3
2011 Predicate Logic Based Image Grammars for Complex Pattern Recognition
Vinay D. Shet, Maneesh Kumar Singh 0001, Claus Bahlmann, Visvanathan Ramesh, Jan Neumann, Larry Davis 0001
Int. J. Comput. Vis.6
2011 Vehicle Detection Using Partial Least Squares
abstract
Detecting vehicles in aerial images has a wide range of applications, from urban planning to visual surveillance. We describe a vehicle detector that improves upon previous approaches by incorporating a very large and rich set of image descriptors. A new feature set called Color Probability Maps is used to capture the color statistics of vehicles and their surroundings, along with the Histograms of Oriented Gradients feature and a simple yet powerful image descriptor that captures the structural characteristics of objects named Pairs of Pixels. The combination of these features leads to an extremely high-dimensional feature set (approximately 70,000 elements). Partial Least Squares is first used to project the data onto a much lower dimensional sub-space. Then, a powerful feature selection analysis is employed to improve the performance while vastly reducing the number of features that must be calculated. We compare our system to previous approaches on two challenging data sets and show superior performance.
Aniruddha Kembhavi, David Harwood, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2010 Learning What and How of Contextual Models for Scene Labeling
Abhinav Gupta 0001, Larry Davis 0001
ECCV (4)3
2010 Why Did the Person Cross the Road (There)? Scene Understanding Using Probabilistic Logic Models and Common Sense Reasoning
Aniruddha Kembhavi, Tom Yeh, Larry Davis 0001
ECCV (2)3
2010 A Robust and Scalable Approach to Face Identification
William Robson Schwartz, Huimin Guo, Larry Davis 0001
ECCV (6)3
2010 Shape-Based Human Detection and Segmentation via Hierarchical Part-Template Matching
abstract
We propose a shape-based, hierarchical part-template matching approach to simultaneous human detection and segmentation combining local part-based and global shape-template-based schemes. The approach relies on the key idea of matching a part-template tree to images hierarchically to detect humans and estimate their poses. For learning a generic human detector, a pose-adaptive feature computation scheme is developed based on a tree matching approach. Instead of traditional concatenation-style image location-based feature encoding, we extract features adaptively in the context of human poses and train a kernel-SVM classifier to separate human/nonhuman patterns. Specifically, the features are collected in the local context of poses by tracing around the estimated shape boundaries. We also introduce an approach to multiple occluded human detection and segmentation based on an iterative occlusion compensation scheme. The output of our learned generic human detector can be used as an initial set of human hypotheses for the iterative optimization. We evaluate our approaches on three public pedestrian data sets (INRIA, MIT-CBCL, and USC-B) and two crowded sequences from Caviar Benchmark and Munich Airport data sets.
Zhe Lin 0001, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 Understanding videos, constructing plots learning a visually grounded storyline model from annotated videos
abstract
Analyzing videos of human activities involves not only recognizing actions (typically based on their appearances), but also determining the story/plot of the video. The storyline of a video describes causal relationships between actions. Beyond recognition of individual actions, discovering causal relationships helps to better understand the semantic meaning of the activities. We present an approach to learn a visually grounded storyline model of videos directly from weakly labeled data. The storyline model is represented as an AND-OR graph, a structure that can compactly encode storyline variation across videos. The edges in the AND-OR graph correspond to causal relationships which are represented in terms of spatio-temporal constraints. We formulate an Integer Programming framework for action recognition and storyline extraction using the storyline model and visual groundings learned from training data.
Abhinav Gupta 0001, Praveen Srinivasan, Jianbo Shi, Larry Davis 0001
CVPR4
2009 Multiple instance fFeature for robust part-based object detection
abstract
Feature misalignment in object detection refers to the phenomenon that features which fire up in some positive detection windows do not fire up in other positive detection windows. Most often it is caused by pose variation and local part deformation. Previous work either totally ignores this issue, or naively performs a local exhaustive search to better position each feature. We propose a learning framework to mitigate this problem, where a boosting algorithm is performed to seed the position of the object part, and a multiple instance boosting algorithm further pursues an aggregated feature for this part, namely multiple instance feature. Unlike most previous boosting based object detectors, where each feature value produces a single classification result, the value of the proposed multiple instance feature is the Noisy-OR integration of a bag of classification results. Our approach is applied to the task of human detection and is tested on two popular benchmarks. The proposed approach brings significant improvement in performance, i.e., smaller number of features used in the cascade and better detection accuracy.
Zhe Lin 0001, Gang Hua 0001, Larry Davis 0001
CVPR3
2009 Incremental Multiple Kernel Learning for object recognition
abstract
A good training dataset, representative of the test images expected in a given application, is critical for ensuring good performance of a visual categorization system. Obtaining task specific datasets of visual categories is, however, far more tedious than obtaining a generic dataset of the same classes. We propose an Incremental Multiple Kernel Learning (IMKL) approach to object recognition that initializes on a generic training database and then tunes itself to the classification task at hand. Our system simultaneously updates the training dataset as well as the weights used to combine multiple information sources. We demonstrate our system on a vehicle classification problem in a video stream overlooking a traffic intersection. Our system updates itself with images of vehicles in poses more commonly observed in the scene, as well as with image patches of the background, leading to an increase in performance. A considerable change in the kernel combination weights is observed as the system gathers scene specific training data over time. The system is also seen to adapt itself to the illumination change in the scene as day transitions to night.
Aniruddha Kembhavi, Behjat Siddiquie, Roland Miezianko, Scott McCloskey, Larry Davis 0001
ICCV5
2009 Recognizing actions by shape-motion prototype trees
abstract
A prototype-based approach is introduced for action recognition. The approach represents an action as a sequence of prototypes for efficient and flexible action matching in long video sequences. During training, first, an action prototype tree is learned in a joint shape and motion space via hierarchical k-means clustering; then a lookup table of prototype-to-prototype distances is generated. During testing, based on a joint likelihood model of the actor location and action prototype, the actor is tracked while a frame-to-prototype correspondence is established by maximizing the joint likelihood, which is efficiently performed by searching the learned prototype tree; then actions are recognized using dynamic prototype sequence matching. Distance matrices used for sequence matching are rapidly obtained by look-up table indexing, which is an order of magnitude faster than brute-force computation of frame-to-frame distances. Our approach enables robust action matching in very challenging situations (such as moving cameras, dynamic backgrounds) and allows automatic alignment of action sequences. Experimental results demonstrate that our approach achieves recognition rates of 91.07% on a large gesture dataset (with dynamic backgrounds), 100% on the Weizmann action dataset and 95.77% on the KTH action dataset.
Zhe Lin 0001, Zhuolin Jiang, Larry Davis 0001
ICCV3
2009 Human detection using partial least squares analysis
abstract
Significant research has been devoted to detecting people in images and videos. In this paper we describe a human detection method that augments widely used edge-based features with texture and color information, providing us with a much richer descriptor set. This augmentation results in an extremely high-dimensional feature space (more than 170,000 dimensions). In such high-dimensional spaces, classical machine learning algorithms such as SVMs are nearly intractable with respect to training. Furthermore, the number of training samples is much smaller than the dimensionality of the feature space, by at least an order of magnitude. Finally, the extraction of features from a densely sampled grid structure leads to a high degree of multicollinearity. To circumvent these data characteristics, we employ Partial Least Squares (PLS) analysis, an efficient dimensionality reduction technique, one which preserves significant discriminative information, to project the data onto a much lower dimensional subspace (20 dimensions, reduced from the original 170,000). Our human detection system, employing PLS analysis over the enriched descriptor set, is shown to outperform state-of-the-art techniques on three varied datasets including the popular INRIA pedestrian dataset, the low-resolution gray-scale DaimlerChrysler pedestrian dataset, and the ETHZ pedestrian dataset consisting of full-length videos of crowded scenes.
William Robson Schwartz, Aniruddha Kembhavi, David Harwood, Larry Davis 0001
ICCV4
2009 Object detection via boosted deformable features
abstract
It is a common practice to model an object for detection tasks as a boosted ensemble of many models built on features of the object. In this context, features are defined as subregions with fixed relative locations and extents with respect to the object's image window. We introduce using deformable features with boosted ensembles. A deformable features adapts its location depending on the visual evidence in order to match the corresponding physical feature. Therefore, deformable features can better handle deformable objects. We empirically show that boosted ensembles of deformable features perform significantly better than boosted ensembles of fixed features for human detection.
Mohamed A. Hussein 0001, Fatih Porikli, Larry Davis 0001
ICIP3
2009 Concurrent transition and shot detection in football videos using Fuzzy Logic
abstract
Shot detection is a fundamental step in video processing and analysis that should be achieved with high degree of accuracy. In this paper, we introduce a unified algorithm for shot detection in sports video using fuzzy logic as a powerful inference mechanism. Fuzzy logic overcomes the problems of hard cut thresholds and the need to large training data used in previous work. The proposed algorithm integrates many features like color histogram, edgeness, intensity variance, etc. Membership functions to represent different features and transitions between shots have been developed to detect different shot boundary and transition types. We address the detection of cut, fade, dissolve, and wipe shot transitions. The results show that our algorithm achieves high degree of accuracy.
Mohammed A. Refaey, Khaled M. F. Elsayed, Sanaa M. Hanafy, Larry Davis 0001
ICIP4
2009 Assigning cameras to subjects in video surveillance systems
abstract
We consider the problem of tracking multiple agents moving amongst obstacles, using multiple cameras. Given an environment with obstacles, and many people moving through it, we construct a separate narrow field of view video for as many people as possible, by stitching together video segments from multiple cameras over time. We employ a novel approach to assign cameras to people as a function of time, with camera switches when needed. The problem is modeled as a bipartite graph and the solution corresponds to a maximum matching. As people move, the solution is efficiently updated by computing an augmenting path rather than by solving for a new matching. This reduces computation time by an order of magnitude. In addition, solving for the shortest augmenting path minimizes the number of camera switches at each update. When not all people can be covered by the available cameras, we cluster as many people as possible into small groups, then assign cameras to groups using a minimum cost matching algorithm. We test our method using numerous runs from different simulators.
Hazem El-Alfy, David Jacobs 0001, Larry Davis 0001
ICRA3
2009 Combining multiple kernels for efficient image classification
abstract
We investigate the problem of combining multiple feature channels for the purpose of efficient image classification. Discriminative kernel based methods, such as SVMs, have been shown to be quite effective for image classification. To use these methods with several feature channels, one needs to combine base kernels computed from them. Multiple kernel learning is an effective method for combining the base kernels. However, the cost of computing the kernel similarities of a test image with each of the support vectors for all feature channels is extremely high. We propose an alternate method, where training data instances are selected, using AdaBoost, for each of the base kernels. A composite decision function, which can be evaluated by computing kernel similarities with respect to only these chosen instances, is learnt. This method significantly reduces the number of kernel computations required during testing. Experimental results on the benchmark UCI datasets, as well as on a challenging painting dataset, are included to demonstrate the effectiveness of our method.
Behjat Siddiquie, Shiv Vitaladevuni, Larry Davis 0001
WACV3
2009 Probabilistic fusion-based parameter estimation for visual tracking
Bohyung Han, Larry Davis 0001
Comput. Vis. Image Underst.2
2009 Segmentation using Appearance of Mesostructure Roughness
Yaser Yacoob, Larry Davis 0001
Int. J. Comput. Vis.2
2009 Observing Human-Object Interactions: Using Spatial and Functional Compatibility for Recognition
abstract
Interpretation of images and videos containing humans interacting with different objects is a daunting task. It involves understanding scene/event, analyzing human movements, recognizing manipulable objects, and observing the effect of the human movement on those objects. While each of these perceptual tasks can be conducted independently, recognition rate improves when interactions between them are considered. Motivated by psychological studies of human perception, we present a Bayesian approach which integrates various perceptual tasks involved in understanding human-object interactions. Previous approaches to object and action recognition rely on static shape/appearance feature matching and motion analysis, respectively. Our approach goes beyond these traditional approaches and applies spatial and functional constraints on each of the perceptual elements for coherent semantic interpretation. Such constraints allow us to recognize objects and actions when the appearances are not discriminative enough. We also demonstrate the use of such constraints in recognition of actions from static images without using any motion information.
Abhinav Gupta 0001, Aniruddha Kembhavi, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2009 Visual Tracking by Continuous Density Propagation in Sequential Bayesian Filtering Framework
abstract
Particle filtering is frequently used for visual tracking problems since it provides a general framework for estimating and propagating probability density functions for nonlinear and non-Gaussian dynamic systems. However, this algorithm is based on a Monte Carlo approach and the cost of sampling and measurement is a problematic issue, especially for high-dimensional problems. We describe an alternative to the classical particle filter in which the underlying density function has an analytic representation for better approximation and effective propagation. The techniques of density interpolation and density approximation are introduced to represent the likelihood and the posterior densities with Gaussian mixtures, where all relevant parameters are automatically determined. The proposed analytic approach is shown to perform more efficiently in sampling in high-dimensional space. We apply the algorithm to real-time tracking problems and demonstrate its performance on real video sequences as well as synthetic examples.
Bohyung Han, Ying Zhu 0006, Dorin Comaniciu, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2009 A Comprehensive Evaluation Framework and a Comparative Study for Human Detectors
abstract
We introduce a framework for evaluating human detectors that considers the practical application of a detector on a full image using multisize sliding-window scanning. We produce detection error tradeoff (DET) curves relating the miss detection rate and the false-alarm rate computed by deploying the detector on cropped windows and whole images, using, in the latter, either image resize or feature resize. Plots for cascade classifiers are generated based on confidence scores instead of on variation of the number of layers. To assess a method's overall performance on a given test, we use the average log miss rate (ALMR) as an aggregate performance score. To analyze the significance of the obtained results, we conduct 10-fold cross-validation experiments. We applied our evaluation framework to two state-of-the-art cascade-based detectors on the standard INRIA person dataset and a local dataset of near-infrared images. We used our evaluation framework to study the differences between the two detectors on the two datasets with different evaluation methods. Our results show the utility of our framework. They also suggest that the descriptors used to represent features and the training window size are more important in predicting the detection performance than the nature of the imaging process, and that the choice between resizing images or features can have serious consequences.
Mohamed E. Hussein 0001, Fatih Porikli, Larry Davis 0001
IEEE Trans. Intell. Transp. Syst.3
2008 Context and observation driven latent variable model for human pose estimation
abstract
Current approaches to pose estimation and tracking can be classified into two categories: generative and discriminative. While generative approaches can accurately determine human pose from image observations, they are computationally expensive due to search in the high dimensional human pose space. On the other hand, discriminative approaches do not generalize well, but are computationally efficient. We present a hybrid model that combines the strengths of the two in an integrated learning and inference framework. We extend the Gaussian process latent variable model (GPLVM) to include an embedding from observation space (the space of image features) to the latent space. GPLVM is a generative model, but the inclusion of this mapping provides a discriminative component, making the model observation driven. Observation Driven GPLVM (OD-GPLVM) not only provides a faster inference approach, but also more accurate estimates (compared to GPLVM) in cases where dynamics are not sufficient for the initialization of search in the latent space. We also extend OD-GPLVM to learn and estimate poses from parameterized actions/gestures. Parameterized gestures are actions which exhibit large systematic variation in joint angle space for different instances due to difference in contextual variables. For example, the joint angles in a forehand tennis shot are function of the height of the ball (Figure 2). We learn these systematic variations as a function of the contextual variables. We then present an approach to use information from scene/objects to provide context for human pose estimation for such parameterized actions.
Abhinav Gupta 0001, Trista Pei-Chun Chen, Francine Chen 0001, Don Kimber, Larry Davis 0001
CVPR5
2008 Kernel integral images: A framework for fast non-uniform filtering
abstract
Integral images are commonly used in computer vision and computer graphics applications. Evaluation of box filters via integral images can be performed in constant time, regardless of the filter size. Although Heckbert (1986) extended the integral image approach for more complex filters, its usage has been very limited, in practice. In this paper, we present an extension to integral images that allows for application of a wide class of non-uniform filters. Our approach is superior to Heckbertpsilas in terms of precision requirements and suitability for parallelization. We explain the theoretical basis of the approach and instantiate two concrete examples: filtering with bilinear interpolation, and filtering with approximated Gaussian weighting. Our experiments show the significant speedups we achieve, and the higher accuracy of our approach compared to Heckbertpsilas.
Mohamed E. Hussein 0001, Fatih Porikli, Larry Davis 0001
CVPR3
2008 Action recognition using ballistic dynamics
abstract
We present a Bayesian framework for action recognition through ballistic dynamics. Psycho-kinesiological studies indicate that ballistic movements form the natural units for human movement planning. The framework leads to an efficient and robust algorithm for temporally segmenting videos into atomic movements. Individual movements are annotated with person-centric morphological labels called ballistic verbs. This is tested on a dataset of interactive movements, achieving high recognition rates. The approach is also applied on a gesture recognition task, improving a previously reported recognition rate from 84% to 92%. Consideration of ballistic dynamics enhances the performance of the popular Motion History Image feature. We also illustrate the approachpsilas general utility on real-world videos. Experiments indicate that the method is robust to view, style and appearance variations.
Shiv Vitaladevuni, Vili Kellokumpu, Larry Davis 0001
CVPR3
2008 Beyond Nouns: Exploiting Prepositions and Comparative Adjectives for Learning Visual Classifiers
Abhinav Gupta 0001, Larry Davis 0001
ECCV (1)2
2008 A Pose-Invariant Descriptor for Human Detection and Segmentation
Zhe Lin 0001, Larry Davis 0001
ECCV (4)2
2008 Event Modeling and Recognition Using Markov Logic Networks
Son Dinh Tran, Larry Davis 0001
ECCV (2)2
2008 Human detection using iterative feature selection and logistic principal component analysis
abstract
We present a fast feature selection algorithm suitable for object detection applications where the image being tested must be scanned repeatedly to detected the object of interest at different locations and scales. The algorithm iteratively estimates the belongness probability of image pixels to foreground of the image. To prove the validity of the algorithm, we apply it to a human detection problem. The edge map is filtered using a feature selection algorithm. The filtered edge map is then projected onto an eigen space of human shapes to determine if the image contains a human. Since the edge maps are binary in nature, Logistic Principal Component Analysis is used to obtain the eigen human shape space. Experimental results illustrate the accuracy of the human detector.
Wael Abd-Almageed, Larry Davis 0001
ICRA2
2008 A "Shape Aware" Model for semi-supervised Learning of Objects and its Context
abstract
Integrating semantic and syntactic analysis is essential for document analysis. Using an analogous reasoning, we present an approach that combines bag-of-words and spatial models to perform semantic and syntactic analysis for recognition of an object based on its internal appearance and its context. We argue that while object recognition requires modeling relative spatial locations of image features within the object, a bag-of-word is sufficient for representing context. Learning such a model from weakly labeled data involves labeling of features into two classes: foreground(object) or ''informative'' background(context). labeling. We present a ''shape-aware'' model which utilizes contour information for efficient and accurate labeling of features in the image. Our approach iterates between an MCMC-based labeling and contour based labeling of features to integrate co-occurrence of features and shape similarity.
Abhinav Gupta 0001, Jianbo Shi, Larry Davis 0001
NIPS3
2008 Automatic online tuning for fast Gaussian summation
abstract
Many machine learning algorithms require the summation of Gaussian kernel functions, an expensive operation if implemented straightforwardly. Several methods have been proposed to reduce the computational complexity of evaluating such sums, including tree and analysis based methods. These achieve varying speedups depending on the bandwidth, dimension, and prescribed error, making the choice between methods difficult for machine learning tasks. We provide an algorithm that combines tree methods with the Improved Fast Gauss Transform (IFGT). As originally proposed the IFGT suffers from two problems: (1) the Taylor series expansion does not perform well for very low bandwidths, and (2) parameter selection is not trivial and can drastically affect performance and ease of use. We address the first problem by employing a tree data structure, resulting in four evaluation methods whose performance varies based on the distribution of sources and targets and input parameters such as desired accuracy and bandwidth. To solve the second problem, we present an online tuning approach that results in a black box method that automatically chooses the evaluation method and its parameters to yield the best performance for the input data, desired accuracy, and bandwidth. In addition, the new IFGT parameter selection approach allows for tighter error bounds. Our approach chooses the fastest method at negligible additional cost, and has superior performance in comparisons with previous approaches.
Vlad I. Morariu, Balaji Vasan Srinivasan, Vikas C. Raykar, Ramani Duraiswami, Larry Davis 0001
NIPS5
2008 Tracking Down Under: Following the Satin Bowerbird
abstract
Socio biologists collect huge volumes of video to study animal behavior (our collaborators work with 30,000 hours of video). The scale of these datasets demands the development of automated video analysis tools. Detecting and tracking animals is a critical first step in this process. However, off-the-shelf methods prove incapable of handling videos characterized by poor quality, drastic illumination changes, non-stationary scenery and foreground objects that become motionless for long stretches of time. We improve on existing approaches by taking advantage of specific aspects of this problem: by using information from the entire video we are able to find animals that become motionless for long intervals of time; we make robust decisions based on regional features; for different parts of the image, we tailor the selection of model features, choosing the features most helpful in differentiating the target animal from the background in that part of the image. We evaluate our method, achieving almost 83% tracking accuracy on a more than 200,000 frame dataset of Satin Bowerbird courtship videos.
Aniruddha Kembhavi, Ryan Farrell, Yuancheng Luo, David Jacobs 0001, Ramani Duraiswami, Larry Davis 0001
WACV6
2008 A General Method for Sensor Planning in Multi-Sensor Systems: Extension to Random Occlusion
Anurag Mittal, Larry Davis 0001
Int. J. Comput. Vis.2
2008 Constraint Integration for Efficient Multiview Pose Estimation with Self-Occlusions
abstract
Automatic initialization and tracking of human pose is an important task in visual surveillance. We present a part-based approach that incorporates a variety of constraints in a unified framework. These constraints include the kinematic constraints between parts that are physically connected to each other, the occlusion of one part by another and the high correlation between the appearance of certain parts, such as the arms. The location probability distribution of each part is determined by evaluating appropriate likelihood measures. The graphical (non-tree) structure representing the interdependencies between parts is utilized to "connect" such part distributions via nonparametric belief propagation. Methods are also developed to perform this optimization efficiently in the large space of pose configurations.
Abhinav Gupta 0001, Anurag Mittal, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2008 Sequential Kernel Density Approximation and Its Application to Real-Time Visual Tracking
abstract
Visual features are commonly modeled with probability density functions in computer vision problems, but current methods such as a mixture of Gaussians and kernel density estimation suffer from either the lack of flexibility, by fixing or limiting the number of Gaussian components in the mixture, or large memory requirement, by maintaining a non-parametric representation of the density. These problems are aggravated in real-time computer vision applications since density functions are required to be updated as new data becomes available. We present a novel kernel density approximation technique based on the mean-shift mode finding algorithm, and describe an efficient method to sequentially propagate the density modes over time. While the proposed density representation is memory efficient, which is typical for mixture densities, it inherits the flexibility of non-parametric methods by allowing the number of components to be variable. The accuracy and compactness of the sequential kernel density approximation technique is illustrated by both simulations and experiments. Sequential kernel density approximation is applied to on-line target appearance modeling for visual tracking, and its performance is demonstrated on a variety of videos.
Bohyung Han, Dorin Comaniciu, Ying Zhu 0006, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2007 Task Scheduling in Large Camera Networks
Ser-Nam Lim, Larry Davis 0001, Anurag Mittal
ACCV (1)2
2007 Simultaneous Appearance Modeling and Segmentation for Matching People Under Occlusion
Zhe Lin 0001, Larry Davis 0001, David S. Doermann, Daniel DeMenthon
ACCV (2)2
2007 Objects in Action: An Approach for Combining Action Understanding and Object Perception
abstract
Analysis of videos of human-object interactions involves understanding human movements, locating and recognizing objects and observing the effects of human movements on those objects. While each of these can be conducted independently, recognition improves when interactions between these elements are considered. Motivated by psychological studies of human perception, we present a Bayesian approach which unifies the inference processes involved in object classification and localization, action understanding and perception of object reaction. Traditional approaches for object classification and action understanding have relied on shape features and movement analysis respectively. By placing object classification and localization in a video interpretation framework, we can localize and classify objects which are either hard to localize due to clutter or hard to recognize due to lack of discriminative features. Similarly, by applying context on human movements from the objects on which these movements impinge and the effects of these movements, we can segment and recognize actions which are either too subtle to perceive or too hard to recognize using motion features alone.
Abhinav Gupta 0001, Larry Davis 0001
CVPR2
2007 Bilattice-based Logical Reasoning for Human Detection
abstract
The capacity to robustly detect humans in video is a critical component of automated visual surveillance systems. This paper describes a bilattice based logical reasoning approach that exploits contextual information and knowledge about interactions between humans, and augments it with the output of different low level detectors for human detection. Detections from low level parts-based detectors are treated as logical facts and used to reason explicitly about the presence or absence of humans in the scene. Positive and negative information from different sources, as well as uncertainties from detections and logical rules, are integrated within the bilattice framework. This approach also generates proofs or justifications for each hypothesis it proposes. These justifications (or lack thereof) are further employed by the system to explain and validate, or reject potential hypotheses. This allows the system to explicitly reason about complex interactions between humans and handle occlusions. These proofs are also available to the end user as an explanation of why the system thinks a particular hypothesis is actually a human. We employ a boosted cascade of gradient histograms based detector to detect individual body parts. We have applied this framework to analyze the presence of humans in static images from different datasets.
Vinay D. Shet, Jan Neumann, Visvanathan Ramesh, Larry Davis 0001
CVPR4
2007 Multimodal Tracking for Smart Videoconferencing and Video Surveillance
abstract
Many applications require the ability to track the 3-D motion of the subjects. We build a particle filter based framework for multimodal tracking using multiple cameras and multiple microphone arrays. In order to calibrate the resulting system, we propose a method to determine the locations of all microphones using at least five loudspeakers and under assumption that for each loudspeaker there exists a microphone very close to it. We derive the maximum likelihood (ML) estimator, which reduces to the solution of the non-linear least squares problem. We verify the correctness and robustness of the multimodal tracker and of the self-calibration algorithm both with Monte-Carlo simulations and on real data from three experimental setups.
Dmitry N. Zotkin, Vikas C. Raykar, Ramani Duraiswami, Larry Davis 0001
CVPR4
2007 Learning Higher-order Transition Models in Medium-scale Camera Networks
abstract
We present a Bayesian framework for learning higher- order transition models in video surveillance networks. Such higher-order models describe object movement between cameras in the network and have a greater predictive power for multi-camera tracking than camera adjacency alone. These models also provide inherent resilience to camera failure, filling in gaps left by single or even multiple non-adjacent camera failures. Our approach to estimating higher-order transition models relies on the accurate assignment of camera observations to the underlying trajectories of objects moving through the network. We addresses this data association problem by gathering the observations and evaluating alternative partitions of the observation set into individual object trajectories. Searching the complete partition space is intractable, so an incremental approach is taken, iteratively adding observations and pruning unlikely partitions. Partition likelihood is determined by the evaluation of a probabilistic graphical model. When the algorithm has considered all observations, the most likely (MAP) partition is taken as the true object trajectories. From these recovered trajectories, the higher-order statistics we seek can be derived and employed for tracking. The partitioning algorithm we present is parallel in nature and can be readily extended to distributed computation in medium-scale smart camera networks.
Ryan Farrell, David S. Doermann, Larry Davis 0001
ICCV3
2007 COST: An Approach for Camera Selection and Multi-Object Inference Ordering in Dynamic Scenes
abstract
Development of multiple camera based vision systems for analysis of dynamic objects such as humans is challenging due to occlusions and similarity in the appearance of a person with the background and other people- visual "confusion". Since occlusion and confusion depends on the presence of other people in the scene, it leads to a dependency structure where there are often loops in the resulting Bayesian network. While approaches such as loopy belief propagation can be used for inference, they are computationally expensive and convergence is not guaranteed in many situations. We present a unified approach, COST, that reasons about such dependencies and yields an order for the inference of each person in a group of people and a set of cameras to be used for inferences for a person. Using the probabilistic distribution of the positions and appearances of people, COST performs visibility and confusion analysis for each part of each person and computes the amount of information that can be computed with and without more accurate estimation of the positions of other people. We present an optimization problem to select set of cameras and inference dependencies for each person which attempts to minimize the computational cost under given performance constraints. Results show the efficiency of COST in improving the performance of such systems and reducing the computational resources required.
Abhinav Gupta 0001, Anurag Mittal, Larry Davis 0001
ICCV3
2007 Probabilistic Fusion Tracking Using Mixture Kernel-Based Bayesian Filtering
abstract
Even though sensor fusion techniques based on particle filters have been applied to object tracking, their implementations have been limited to combining measurements from multiple sensors by the simple product of individual likelihoods. Therefore, the number of observations is increased as many times as the number of sensors, and the combined observation may become unreliable through blind integration of sensor observations—especially if some sensors are too noisy and non-discriminative. We describe a methodology to model interactions between multiple sensors and to estimate the current state by using a mixture of Bayesian filters—one filter for each sensor, where each filter makes a different level of contribution to estimate the combined posterior in a reliable manner. In this framework, an adaptive particle arrangement system is constructed in which each particle is allocated to only one of the sensors for observation and a different number of samples is assigned to each sensor using prior distribution and partial observations. We apply this technique to visual tracking in logical and physical sensor fusion frameworks, and demonstrate its effectiveness through tracking results.
Bohyung Han, Seong-Wook Joo, Larry Davis 0001
ICCV3
2007 Hierarchical Part-Template Matching for Human Detection and Segmentation
abstract
Local part-based human detectors are capable of handling partial occlusions efficiently and modeling shape articulations flexibly, while global shape template-based human detectors are capable of detecting and segmenting human shapes simultaneously. We describe a Bayesian approach to human detection and segmentation combining local part-based and global template-based schemes. The approach relies on the key ideas of matching a part-template tree to images hierarchically to generate a reliable set of detection hypotheses and optimizing it under a Bayesian MAP framework through global likelihood re-evaluation and fine occlusion analysis. In addition to detection, our approach is able to obtain human shapes and poses simultaneously. We applied the approach to human detection and segmentation in crowded scenes with and without background subtraction. Experimental results show that our approach achieves good performance on images and video sequences with severe occlusion.
Zhe Lin 0001, Larry Davis 0001, David S. Doermann, Daniel DeMenthon
ICCV2
2007 An Interactive Approach to Pose-Assisted and Appearance-based Segmentation of Humans
abstract
An interactive human segmentation approach is described. Given regions of interest provided by users, the approach iteratively estimates segmentation via a generalized EM algorithm. Specifically, it encodes both spatial and color information in a nonparametric kernel density estimator, and incorporates local MRF constraints and global pose inferences to propagate beliefs over image space iteratively to determine a coherent segmentation. This ensures the segmented humans resemble the shapes of human poses. Additionally, a layered occlusion model and a probabilistic occlusion reasoning method are proposed to handle segmentation of multiple humans in occlusion. The approach is tested on a wide variety of images containing single or multiple occluded humans, and the segmentation performance is evaluated quantitatively.
Zhe Lin 0001, Larry Davis 0001, David S. Doermann, Daniel DeMenthon
ICCV2
2007 Robust Object Trackinng wvith Regional Affine Invariant Features
abstract
We present a tracking algorithm based on motion analysis of regional affine invariant image features. The tracked object is represented with a probabilistic occupancy map. Using this map as support, regional features are detected and probabilistically matched across frames. The motion of pixels is then established based on the feature motion. The object occupancy map is in turn updated according to the pixel motion consistency. We describe experiments to measure the sensitivities of our approach to inaccuracy in initialization, and compare it with other approaches.
Son Dinh Tran, Larry Davis 0001
ICCV2
2007 Segmentation using Meta-texture Saliency
abstract
We address segmentation of an image into patches that have an underlying salient surface-roughness. Three intrinsic images are derived: reflectance, shading and meta- texture images. A constructive approach is proposed for computing a meta-texture image by preserving, equalizing and enhancing the underlying surface-roughness across color, brightness and illumination variations. We evaluate the performance on sample images and illustrate quantitatively that different patches of the same material, in an image, are normalized in their statistics despite variations in color, brightness and illumination. Finally, segmentation by line-based boundary-detection is proposed and results are provided and compared to known algorithms.
Yaser Yacoob, Larry Davis 0001
ICCV2
2007 Multi-scale video cropping
abstract
We consider the problem of cropping surveillance videos. This process chooses a trajectory that a small sub-window can take through the video, selecting the most important parts of the video for display on a smaller monitor. We model the information content of the video simply, by whether the image changes at each pixel. Then we show that we can find the globally optimal trajectory for a cropping window by using a shortest path algorithm. In practice, we can speed up this process without affecting the results, by stitching together trajectories computed over short intervals. This also reduces system latency. We then show that we can use a second shortest path formulation to find good cuts from one trajectory to another, improving coverage of interesting events in the video. We describe additional techniques to improve the quality and efficiency of the algorithm, and show results on surveillance videos.
Hazem El-Alfy, David Jacobs 0001, Larry Davis 0001
ACM Multimedia3
2007 Pedestrian Detection via Periodic Motion Analysis
Yang Ran, Isaac Weiss, Qinfen Zheng, Larry Davis 0001
Int. J. Comput. Vis.4
2007 Human appearance modeling for matching across video sequences
David Harwood, Kyongil Yoon, Larry Davis 0001
Mach. Vis. Appl.4
2006 Density Estimation Using Mixtures of Mixtures of Gaussians
Wael Abd-Almageed, Larry Davis 0001
ECCV (4)2
2006 Multi-camera Tracking and Segmentation of Occluded People on Ground Plane Using Search-Guided Particle Filtering
Kyungnam Kim, Larry Davis 0001
ECCV (3)2
2006 Multivalued Default Logic for Identity Maintenance in Visual Surveillance
Vinay D. Shet, David Harwood, Larry Davis 0001
ECCV (4)3
2006 3D Surface Reconstruction Using Graph Cuts with Surface Constraints
Son Dinh Tran, Larry Davis 0001
ECCV (2)2
2006 Real-Time Human Detection, Tracking, and Verification in Uncontrolled Camera Motion Environments
abstract
In environments where a camera is installed on a freely moving platform, e.g. a vehicle or a robot, object detection and tracking becomes much more difficult. In this paper, we presents a real time system for human detection, tracking, and verification in such challenging environments. To deliver a robust performance, the system integrates several computer vision algorithms to perform its function: a human detection algorithm, an object tracking algorithm, and a motion analysis algorithm. To utilize the available computing resources to the maximum possible extent, each of the system components is designed to work in a separate thread that communicates with the other threads through shared data structures. The focus of this paper is more on the implementation issues than on the algorithmic issues of the system. Object oriented design was adopted to abstract algorithmic details away from the system structure.
Mohamed E. Hussein 0001, Wael Abd-Almageed, Yang Ran, Larry Davis 0001
ICVS4
2006 Parametric Hand Tracking for Recognition of Virtual Drawings
abstract
A hand tracking system for recognition of virtual spatial drawings is presented. Using a stereo camera, the 3D position of the hand in space is estimated. Then, by tracking the central region of the hand in 3D and estimating a virtual plane in space, the intended drawing of the user is recognized. Experimental results demonstrate the accuracy and effectiveness of this technique. The system can be used to communicate drawings and alphabets to a computer where a classifier can transform the drawn alphabets into interpretable characters.
Afshin Sepehri, Yaser Yacoob, Larry Davis 0001
ICVS3
2006 Tracking Articulating Objects from Ground Vehicles using Mixtures of Mixtures
abstract
An algorithm for tracking articulating objects from moving camera platforms is presented. Mixtures of mixtures are used to model the appearance of the object and the background. The state of the object is tracked using a particle filter. Egomotion information are estimated and used to set the state variance of the particle filter. Results of tracking human objects from an unmanned ground vehicle are used to evaluate the tracking algorithm
Wael Abd-Almageed, Mohamed E. Hussein 0001, Larry Davis 0001
IROS3
2006 Edge affinity for pose-contour matching
V. Shiv Naga Prasad, Larry Davis 0001, Son Dinh Tran, Ahmed M. Elgammal
Comput. Vis. Image Underst.2
2006 Editorial
Aaron F. Bobick, Rama Chellappa, Larry Davis 0001
Int. J. Comput. Vis.3
2006 3D Structure Recovery and Unwarping of Surfaces Applicable to Planes
Nail A. Gumerov, Ali Zandifar, Ramani Duraiswami, Larry Davis 0001
Int. J. Comput. Vis.4
2006 Segmentation of Planar Objects and Their Shadows in Motion Sequences
Yaser Yacoob, Larry Davis 0001
Int. J. Comput. Vis.2
2006 Appearance-based person recognition using color/path-length profile
Kyongil Yoon, David Harwood, Larry Davis 0001
J. Vis. Commun. Image Represent.3
2006 Constructing task visibility intervals for video surveillance
Ser-Nam Lim, Larry Davis 0001, Anurag Mittal
Multim. Syst.2
2006 Detection and Analysis of Hair
abstract
We develop computational models for measuring hair appearance for comparing different people. The models and methods developed have applications to person recognition and image indexing. An automatic hair detection algorithm is described and results reported. A multidimensional representation of hair appearance is presented and computational algorithms are described. Results on a data set of 524 subjects are reported. Identification of people using hair attributes is compared to eigenface-based recognition along with a joint, eigenface-hair-based identification.
Yaser Yacoob, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 VidMAP: video monitoring of activity with Prolog
abstract
This paper describes the architecture of a visual surveillance system that combines real time computer vision algorithms with logic programming to represent and recognize activities involving interactions amongst people, packages and the environments through which they move. The low level computer vision algorithms log primitive events of interest as observed facts, while the higher level Prolog based reasoning engine uses these facts in conjunction with predefined rules to recognize various activities in the input video streams. The system is illustrated in action on a multi-camera surveillance scenario that includes both security and safety violations.
Vinay D. Shet, David Harwood, Larry Davis 0001
AVSS3
2005 Kernel-Based Bayesian Filtering for Object Tracking
abstract
Particle filtering provides a general framework for propagating probability density functions in nonlinear and non-Gaussian systems. However, the algorithm is based on a Monte Carlo approach and sampling is a problematic issue, especially for high dimensional problems. This paper presents a new kernel-based Bayesian filtering framework, which adopts an analytic approach to better approximate and propagate density functions. In this framework, the techniques of density interpolation and density approximation are introduced to represent the likelihood and the posterior densities by Gaussian mixtures, where all parameters such as the number of mixands, their weight, mean, and covariance are automatically determined. The proposed analytic approach is shown to perform sampling more efficiently in high dimensional space. We apply our algorithm to real-time tracking problems, and demonstrate its performance on real video sequences as well as synthetic examples.
Bohyung Han, Ying Zhu 0006, Dorin Comaniciu, Larry Davis 0001
CVPR (1)4
2005 Fast Illumination-Invariant Background Subtraction Using Two Views: Error Analysis, Sensor Placement and Applications
abstract
Background modeling and subtraction to detect new or moving objects in a scene is an important component of many intelligent video applications. Compared to a single camera, the use of multiple cameras leads to better handling of shadows, specularities and illumination changes due to the utilization of geometric information. Although the result of stereo matching can be used as the feature for detection, it has been shown that the detection process can be made much faster by a simple subtraction of the intensities observed at stereo-generated conjugate pairs in the two views. The methodology however, suffers from false and missed detections due to some geometric considerations. In this paper, we perform a detailed analysis of such errors. Then, we propose a sensor configuration that eliminates false detections. Algorithms are also proposed that effectively eliminate most detection errors due to missed detections, specular reflections and objects being geometrically close to the background. Experiments on several scenes illustrate the utility and enhanced performance of the proposed approach compared to existing techniques.
Ser-Nam Lim, Anurag Mittal, Larry Davis 0001, Nikos Paragios
CVPR (1)3
2005 Efficient Mean-Shift Tracking via a New Similarity Measure
abstract
The mean shift algorithm has achieved considerable success in object tracking due to its simplicity and robustness. It finds local minima of a similarity measure between the color histograms or kernel density estimates of the model and target image. The most typically used similarity measures are the Bhattacharyya coefficient or the Kullback-Leibler divergence. In practice, these approaches face three difficulties. First, the spatial information of the target is lost when the color histogram is employed, which precludes the application of more elaborate motion models. Second, the classical similarity measures are not very discriminative. Third, the sample-based classical similarity measures require a calculation that is quadratic in the number of samples, making real-time performance difficult. To deal with these difficulties we propose a new, simple-to-compute and more discriminative similarity measure in spatial-feature spaces. The new similarity measure allows the mean shift algorithm to track more general motion models in an integrated way. To reduce the complexity of the computation to linear order we employ the recently proposed improved fast Gauss transform. This leads to a very efficient and robust nonparametric spatial-feature tracking algorithm. The algorithm is tested on several image sequences and shown to achieve robust and reliable frame-rate tracking.
Changjiang Yang, Ramani Duraiswami, Larry Davis 0001
CVPR (1)3
2005 Reliable Segmentation of Pedestrians in Moving Scenes
abstract
This paper describes a periodic motion based pedestrian segmentation algorithm for videos acquired from moving platforms. Given a sequence of bounding boxes containing the detected and tracked walking human, the goal is to analyze the low D structure by considering every object sample as a point in the high D manifold space and use the learned structure for segmentation. In this work, we introduce a novel bottom-up learning approach. We represent the human stride as a cascade of models with increasing parameter numbers. These parameters describe the dynamics of pedestrians from coarse to fine. By applying the learned manifold structure, we can predict the location of body parts, especially legs, with high accuracy at every frame. The segmentation in consecutive images is done by EM clustering. With the accuracy for prediction using the twin-pendulum model, EM is more likely to converge to global maximums. Experimental results for real videos are presented The algorithm has demonstrated a reliable performance for videos acquired from moving platforms.
Yang Ran, Qinfen Zheng, Isaac Weiss, Larry Davis 0001
ICASSP (2)4
2005 On-Line Density-Based Appearance Modeling for Object Tracking
abstract
Object tracking is a challenging problem in real-time computer vision due to variations of lighting condition, pose, scale, and view-point over time. However, it is exceptionally difficult to model appearance with respect to all of those variations in advance; instead, on-line update algorithms are employed to adapt to these changes. We present a new on-line appearance modeling technique which is based on sequential density approximation. This technique provides accurate and compact representations using Gaussian mixtures, in which the number of Gaussians is automatically determined. This procedure is performed in linear time at each time step, which we prove by amortized analysis. Features for each pixel and rectangular region are modeled together by the proposed sequential density approximation algorithm, and the target model is updated in scale robustly. We show the performance of our method by simulations and tracking in natural videos
Bohyung Han, Larry Davis 0001
ICCV2
2005 Detecting Rotational Symmetries
abstract
We present an algorithm for detecting multiple rotational symmetries in natural images. Given an image, its gradient magnitude field is computed, and information from the gradients is spread using a diffusion process in the form of a gradient vector flow (GVF) field. We construct a graph whose nodes correspond to pixels in tire image, connecting points that are likely to be rotated versions of one another The n-cycles present in tire graph are made to vote for C/sub n/ symmetries, their votes being weighted by the errors in transformation between GVF in the neighborhood of the voting points, and the irregularity of the n-sided polygons formed by the voters. The votes are accumulated at tire centroids of possible rotational symmetries, generating a confidence map for each order of symmetry. We tested the method with several natural images.
V. Shiv Naga Prasad, Larry Davis 0001
ICCV2
2005 Detection, Analysis and Matching of Hair
abstract
We develop computational models for measuring hair appearance for comparing different people. The models and methods developed have applications to person recognition and face image indexing. An automatic hair detection algorithm is described and results reported. A multidimensional representation of hair appearance is presented and computational algorithms are described. Results on a dataset of 524 subjects are reported. Identification of people using hair attributes is compared to eigenface-based recognition along with a joint, eigenface-hair based identification.
Yaser Yacoob, Larry Davis 0001
ICCV2
2005 Fast Multiple Object Tracking via a Hierarchical Particle Filter
abstract
A very efficient and robust visual object tracking algorithm based on the particle filter is presented. The method characterizes the tracked objects using color and edge orientation histogram features. While the use of more features and samples can improve the robustness, the computational load required by the particle filter increases. To accelerate the algorithm while retaining robustness we adopt several enhancements in the algorithm. The first is the use of integral images for efficiently computing the color features and edge orientation histograms, which allows a large amount of particles and a better description of the targets. Next, the observation likelihood based on multiple features is computed in a coarse-to-fine manner, which allows the computation to quickly focus on the more promising regions. Quasi-random sampling of the particles allows the filter to achieve a higher convergence rate. The resulting tracking algorithm maintains multiple hypotheses and offers robustness against clutter or short period occlusions. Experimental results demonstrate the efficiency and effectiveness of the algorithm for single and multiple object tracking.
Changjiang Yang, Ramani Duraiswami, Larry Davis 0001
ICCV3
2005 Closely Coupled Object Detection and Segmentation
abstract
We propose a closely coupled object detection and segmentation algorithm for enhancing both processes in a cooperative and iterative manner. Figure-ground segmentation reduces the effect of background clutter on template matching; the matched template provides shape constraints on segmentation. More precisely, we estimate the probability of each pixel belonging to the foreground by a weighted sum of the estimates based on shape and color alone. The weight on the shape-based estimate is related to the probability that a familiar object is present and is updated dynamically so that we enforce shape constraints only where the object is present. Experiments on detecting people in images of cluttered scenes demonstrate that the proposed algorithm improves both segmentation and detection. More accurate object boundaries are extracted; higher object detection rates and lower false alarm rates are achieved than performing the two processes separately or sequentially.
Larry Davis 0001
ICCV2
2005 Extracting regions of symmetry
abstract
This paper presents an approach for extending the normalized-cut (n-cut) segmentation algorithm to find symmetric regions present in natural images. We use an existing algorithm to quickly detect possible symmetries present in an image. The detected symmetries are then individually verified using the modified n-cut algorithm to eliminate spurious detections. The weights of the n-cut algorithm are modified so as to include both symmetric and spatial affinities. A global parameter is defined to model the tradeoff between spatial coherence and symmetry. Experimental results indicate that symmetric quality measure for a region segmented by our algorithm is a good indicator for the significance of the principal axis of symmetry.
Abhinav Gupta 0001, V. Shiv Naga Prasad, Larry Davis 0001
ICIP (3)3
2005 Robust observations for object tracking
abstract
It is a difficult task to find an observation model that will perform well for long-term visual tracking. In this paper, we propose an adaptive observation enhancement technique based on likelihood images, which are derived from multiple visual features. The most discriminative likelihood image is extracted by principal component analysis (PCA) and incrementally updated frame by frame to reduce temporal tracking error. In the particle filter framework, the feasibility of each sample is computed using this most discriminative likelihood image before the observation process. Integral image is employed for efficient computation of the feasibility of each sample. We illustrate how our enhancement technique contributes to more robust observations through demonstrations.
Bohyung Han, Larry Davis 0001
ICIP (2)2
2005 Pedestrian classification from moving platforms using cyclic motion pattern
abstract
This paper describes an efficient pedestrian detection system for videos acquired from moving platforms. Given a detected and tracked object as a sequence of images within a bounding box, we describe the periodic signature of its motion pattern using a twin-pendulum model. Then a principle gait angle is extracted in every frame providing gait phase information. By estimating the periodicity from the phase data using a digital phase locked loop (dPLL), we quantify the cyclic pattern of the object, which helps us to continuously classify it as a pedestrian. Past approaches have used shape detectors applied to a single image or classifiers based on human body pixel oscillations, but ours is the first to integrate a global cyclic motion model and periodicity analysis. Novel contributions of this paper include: i) development of a compact shape representation of cyclic motion as a signature for a pedestrian, ii) estimation of gait period via a feedback loop module, and iii) implementation of a fast online pedestrian classification system which operates on videos acquired from moving platforms.
Yang Ran, Qinfen Zheng, Isaac Weiss, Larry Davis 0001, Wael Abd-Almageed
ICIP (2)4
2005 Segmentation and appearance model building from an image sequence
abstract
In this paper we explore the problem of accurately segmenting a person from a video given only approximate location of that person. Unlike previous work which assumes that the appearance model is known in advance, we developed an iterative expectation-sampling (ES) algorithm for solving segmentation and appearance modeling simultaneously The appearance model is encoded with a kernel-based PDF defined in a joint color/path-length space. This appearance model remains unchanged during a short time period, although the object can articulate. Thus, we can perform the ES iteration not only for a single frame but also for an image sequence. The algorithm is iterative, but simple, efficient and gives visually good results.
Larry Davis 0001
ICIP (1)2
2005 A video-based framework for the analysis of presentations/posters
Ali Zandifar, Ramani Duraiswami, Larry Davis 0001
Int. J. Document Anal. Recognit.3
2004 Window-Based, Discontinuity Preserving Stereo
Motilal Agrawal, Larry Davis 0001
CVPR (1)2
2004 Incremental Density Approximation and Kernel-Based Bayesian Filtering for Object Tracking
Bohyung Han, Dorin Comaniciu, Ying Zhu 0006, Larry Davis 0001
CVPR (1)4
2004 Structure of Applicable Surfaces from Single Views
Nail A. Gumerov, Ali Zandifar, Ramani Duraiswami, Larry Davis 0001
ECCV (3)4
2004 Visibility Analysis and Sensor Planning in Dynamic Environments
Anurag Mittal, Larry Davis 0001
ECCV (1)2
2004 Flexible layout and optimal cancellation of the orthonormality error for spherical microphone arrays
abstract
This paper describes an approach to achieving a flexible layout of microphones on the surface of a spherical microphone array for beamforming. Our approach achieves orthonormality of spherical harmonics to higher order for relatively distributed layouts. This gives great flexibility in microphone layout on the spherical surface. One direct advantage is that it makes it much easier to build a real world system, such as those with cable outlets and a mounting base, with minimal effects on the performance. Simulation results are presented.
Zhiyun Li, Ramani Duraiswami, Elena Grassi, Larry Davis 0001
ICASSP (4)4
2004 Object tracking by adaptive feature extraction
Bohyung Han, Larry Davis 0001
ICIP2
2004 Background modeling and subtraction by codebook construction
abstract
We present a new fast algorithm for background modeling and subtraction. Sample background values at each pixel are quantized into codebooks which represent a compressed form of background model for a long image sequence. This allows us to capture structural background variation due to periodic-like motion over a long period of time under limited memory. Our method can handle scenes containing moving backgrounds or illumination variations (shadows and highlights), and it achieves robust detection for compressed videos. We compared our method with other multimode modeling techniques.
Kyungnam Kim, Thanarat H. Chalidabhongse, David Harwood, Larry Davis 0001
ICIP4
2004 A fine-structure image/video quality measure using local statistics
abstract
An objective no-reference measure is presented to assess line-structure image/video quality. It was designed to measure image/video quality for video surveillance applications, especially for background modeling and foreground object detection. The proposed measure using local statistics reflects image degradation well in terms of noise and blur. The experimental results on a background subtraction algorithm validate the usefulness of the proposed measure, by showing its correlation with the algorithm's performance.
Kyungnam Kim, Larry Davis 0001
ICIP2
2004 Uncalibrated stereo rectification for automatic 3d surveillance
abstract
We describe a stereo rectification method suitable for automatic 3D surveillance. We take advantage of the fact that in a typical urban scene, there is ordinarily a small number of dominant planes. Given two views of the scene, we align a dominant plane in one view with the other. Conjugate epipolar lines between the reference view and plane-aligned image become geometrically identical and can be added to the rectified image pair line by line. Selecting conjugate epipolar lines to cover the whole image is simplified since they are geometrically identical. In addition, the polarities of conjugate epipolar lines are automatically preserved by plane alignment, which simplifies stereo matching.
Ser-Nam Lim, Anurag Mittal, Larry Davis 0001, Nikos Paragios
ICIP3
2004 Multi-level fast multipole method for thin plate spline evaluation
Ali Zandifar, Ser-Nam Lim, Ramani Duraiswami, Nail A. Gumerov, Larry Davis 0001
ICIP5
2004 A Method for Designing Marker-Based Tracking Probes
abstract
Many tracking systems utilize collections of fiducial markers arranged in rigid configurations, called tracking probes, to determine the pose of objects within an environment. In this paper, we present a technique for designing tracking probes called the viewpoints algorithm. The algorithm is generally applicable to tracking systems that use at least three fiduciary marks to determine the pose of an object. The algorithm is used to create an integrated, head-mounted display tracking probe. The predicted accuracy of this probe was 0.032 /spl plusmn/ 0.02 degrees in orientation and 0.09 /spl plusmn/ 0.07 mm in position. The measured accuracy of the probe was 0.028 /spl plusmn/ 0.01 degrees in orientation and 0.11 /spl plusmn/ 0.01 mm in position. These results translate to a predicted, static positional overlay error of a virtual object presented at 1m of less than 0.5 mm. The algorithm is part of a larger framework for designing tracking probes based upon performance goals and environmental constraints.
Larry Davis 0001, Felix G. Hamza-Lup, Jannick P. Rolland
ISMAR1
2004 Recording and Reproducing High Order Surround Auditory Scenes for Mixed and Augmented Reality
abstract
Virtual reality systems are largely based on computer graphics and vision technologies. However, sound also plays an important role in human's interaction with the surrounding environment, especially for the visually impaired people. In this paper, we develop the theory of recording and reproducing real-world surround auditory scenes in high orders using specially designed microphone and loudspeaker arrays. It is complementary to vision-based technologies in creating mixed and augmented realities. Design examples and simulations are presented.
Zhiyun Li, Ramani Duraiswami, Larry Davis 0001
ISMAR3
2004 Efficient Kernel Machines Using the Improved Fast Gauss Transform
abstract
The computation and memory required for kernel machines with N train- ing samples is at least O(N 2). Such a complexity is significant even for moderate size problems and is prohibitive for large datasets. We present an approximation technique based on the improved fast Gauss transform to reduce the computation to O(N ). We also give an error bound for the approximation, and provide experimental results on the UCI datasets.
Changjiang Yang, Ramani Duraiswami, Larry Davis 0001
NIPS3
2004 In Memory of Azriel Rosenfeld
Larry Davis 0001
Int. J. Comput. Vis.1
2004 Rendering localized spatial audio in a virtual auditory space
abstract
High-quality virtual audio scene rendering is required for emerging virtual and augmented reality applications, perceptual user interfaces, and sonification of data. We describe algorithms for creation of virtual auditory spaces by rendering cues that arise from anatomical scattering, environmental scattering, and dynamical effects. We use a novel way of personalizing the head related transfer functions (HRTFs) from a database, based on anatomical measurements. Details of algorithms for HRTF interpolation, room impulse response creation, HRTF selection from a database, and audio scene presentation are presented. Our system runs in real time on an office PC without specialized DSP hardware.
Dmitry N. Zotkin, Ramani Duraiswami, Larry Davis 0001
IEEE Trans. Multim.3
2003 A Scalable Image-Based Multi-Camera Visual Surveillance System
abstract
We describe the design of a scalable and wide coverage visual surveillance system. Scalability (the ability to add and remove cameras easily during system operation with minimal overhead and system degradation) is achieved by utilizing only image-based information for camera control. We show that when a pan-tilt-zoom camera pans and tilts, a given image point moves in a circular and a linear trajectory, respectively. We create a scene model using a plan view of the scene. The scene model makes it easy for us to handle occlusion prediction and schedule video acquisition tasks subject to visibility constraints. We describe a maximum weight matching algorithm to assign cameras to tasks that meet the visibility constraints. The system is illustrated both through simulations and real video from a 6-camera configuration.
Ser-Nam Lim, Larry Davis 0001, Ahmed M. Elgammal
AVSS2
2003 Human Body Pose Estimation Using Silhouette Shape Analysis
abstract
We describe a system for human body pose estimation from multiple views that is fast and completely automatic. The algorithm works in the presence of multiple people by decoupling the problems of pose estimation of different people. The pose is estimated based on a likelihood function that integrates information from multiple views and thus obtains a globally optimal solution. Other characteristics that make our method more general than previous work include: (1) no manual initialization; (2) no specification of the dimensions of the 3D structure; (3) no reliance on some learned poses or patterns of activity; (4) insensitivity to edges and clutter in the background and within the foreground. The algorithm has applications in surveillance and promising results have been obtained.
Anurag Mittal, Larry Davis 0001
AVSS3
2003 Probabilistic Tracking in Joint Feature-Spatial Spaces
abstract
In this paper, we present a probabilistic framework for tracking regions based on their appearance. We exploit the feature-spatial distribution of a region representing an object as a probabilistic constraint to track that region over time. The tracking is achieved by maximizing a similarity-based objective function over transformation space given a nonparametric representation of the joint feature-spatial distribution. Such a representation imposes a probabilistic constraint on the region feature distribution coupled with the region structure, which yields an appearance tracker that is robust to small local deformations and partial occlusion. We present the approach for the general form of joint feature-spatial distributions and apply it to tracking with different types of image features including row intensity, color and image gradient.
Ahmed M. Elgammal, Ramani Duraiswami, Larry Davis 0001
CVPR (1)3
2003 Learning Dynamics for Exemplar-based Gesture Recognition
abstract
This paper addresses the problem of capturing the dynamics for exemplar-based recognition systems. Traditional HMM provides a probabilistic tool to capture system dynamics and in exemplar paradigm, HMM states are typically coupled with the exemplars. Alternatively, we propose a non-parametric HMM approach that uses a discrete HMM with arbitrary states (decoupled from exemplars) to capture the dynamics over a large exemplar space where a nonparametric estimation approach is used to model the exemplar distribution. This reduces the need for lengthy and non-optimal training of the HMM observation model. We used the proposed approach for view-based recognition of gestures. The approach is based on representing each gesture as a sequence of learned body poses (exemplars). The gestures are recognized through a probabilistic framework for matching these body poses and for imposing temporal constraints between different poses using the proposed non-parametric HMM.
Ahmed M. Elgammal, Vinay D. Shet, Yaser Yacoob, Larry Davis 0001
CVPR (1)4
2003 Pitch and timbre manipulations using cortical representation of sound
abstract
The sound received at the ears is processed by humans using signal processing that separates the signal along intensity, pitch and timbre dimensions. Conventional Fourier-based signal processing, while endowed with fast algorithms, is unable to represent a signal easily along the lines of these attributes. We use a recently proposed cortical representation (Elhilali, M. et al., Speech Communications, 2002) to represent and manipulate sound. We briefly overview algorithms for obtaining, manipulating and inverting cortical representation of a sound and describe algorithms for manipulating signal pitch and timbre separately. The algorithms are first used to create the sound of an instrument between a "guitar" and a "trumpet". Applications to creating maximally separable sounds in auditory user interfaces are discussed.
Dmitry N. Zotkin, Shihab A. Shamma, Powen Ru, Ramani Duraiswami, Larry Davis 0001
ICASSP (5)5
2003 Camera calibration using spheres: A semi-definite programming approach
abstract
Vision algorithms utilizing camera networks with a common field of view are becoming increasingly feasible and important. Calibration of such camera networks is a challenging and cumbersome task. The current approaches for calibration using planes or a known 3D target may not be feasible as these objects may not be simultaneously visible in all the cameras. In this paper, we present a new algorithm to calibrate cameras using occluding contours of spheres. In general, an occluding contour of a sphere projects to an ellipse in the image. Our algorithm uses the projection of the occluding contours of three spheres and solves for the intrinsic parameters and the locations of the spheres. The problem is formulated in the dual space and the parameters are solved for optimally and efficiently using semidefinite programming. The technique is flexible, accurate and easy to use. In addition, since the contour of a sphere is simultaneously visible in all the cameras, our approach can greatly simplify calibration of multiple cameras with a common field of view. Experimental results from computer simulated data and real world data, both for a single camera and multiple cameras, are presented.
Motilal Agrawal, Larry Davis 0001
ICCV2
2003 Improved Fast Gauss Transform and Efficient Kernel Density Estimation
abstract
Evaluating sums of multivariate Gaussians is a common computational task in computer vision and pattern recognition, including in the general and powerful kernel density estimation technique. The quadratic computational complexity of the summation is a significant barrier to the scalability of this algorithm to practical applications. The fast Gauss transform (FGT) has successfully accelerated the kernel density estimation to linear running time for low-dimensional problems. Unfortunately, the cost of a direct extension of the FGT to higher-dimensional problems grows exponentially with dimension, making it impractical for dimensions above 3. We develop an improved fast Gauss transform to efficiently estimate sums of Gaussians in higher dimensions, where a new multivariate expansion scheme and an adaptive space subdivision technique dramatically improve the performance. The improved FGT has been applied to the mean shift algorithm achieving linear computational complexity. Experimental results demonstrate the efficiency and effectiveness of our algorithm.
Changjiang Yang, Ramani Duraiswami, Nail A. Gumerov, Larry Davis 0001
ICCV4
2003 Mean-shift analysis using quasiNewton methods
abstract
Mean-shift analysis is a general nonparametric clustering technique based on density estimation for the analysis of complex feature spaces. The algorithm consists of a simple iterative procedure that shifts each of the feature points to the nearest stationary point along the gradient directions of the estimated density function. It has been successfully applied to many applications such as segmentation and tracking. However, despite its promising performance, there are applications for which the algorithm converges too slowly to be practical. We propose and implement an improved version of the mean-shift algorithm using quasiNewton methods to achieve higher convergence rates. Another benefit of our algorithm is its ability to achieve clustering even for very complex and irregular feature-space topography. Experimental results demonstrate the efficiency and effectiveness of our algorithm.
Changjiang Yang, Ramani Duraiswami, Daniel DeMenthon, Larry Davis 0001
ICIP (2)4
2003 Image-based pan-tilt camera control in a multi-camera surveillance environment
abstract
In automated surveillance systems with multiple cameras, the system must be able to position the cameras accurately. Each camera must be able to pan-tilt such that an object detected in the scene is in a vantage position in the camera's image plane and subsequently capture images of that object. Typically, camera calibration is required. We propose an approach that uses only image-based information. Each camera is assigned a pan-tilt zero-position. Position of an object detected in one camera is related to the other cameras by homographies between the zero-positions while different pan-tilt positions of the same camera are related in the form of projective rotations. We then derive that the trajectories in the image plane corresponding to these projective rotations are approximately circular for pan and linear for tilt. The camera control technique is subsequently tested in a working prototype.
Ser-Nam Lim, Ahmed M. Elgammal, Larry Davis 0001
ICME3
2003 Using computer vision to generate customized spatial audio
abstract
Creating high quality virtual spatial audio over headphones requires real-time head tracking, personalized head-related transfer functions (HRTFs) and customized room response models. While there are expensive solutions to address these issues based on costly head trackers, measured personalized HRTFs and room responses, these are not suitable for widespread or easy deployment and use. We report on the development of a system that uses computer vision to produce customizable models for both the HRTF and the room response, and to achieve head-tracking. The system uses relatively inexpensive cameras and widely available personal computers. Computer-vision based anthropometric measurements of the head, torso, and the external ears are used for HRTF customization. For low-frequency HRTF customization we employ a simple head-and-torso model developed recently [V. R. Algazi et al., 2002]. For high frequency customization we employ measured pinna characteristics as an index into a database of HRTFs [D. N. Zotkin et al., 2002]. For head tracking we employ an online implementation of the POSIT algorithm [D. DeMenthon and L. Davis, 1995] along with active markers to compute head pose in real-time. The system provides an enhanced virtual listening experience at low cost.
Ankur Mohan, Ramani Duraiswami, Dmitry N. Zotkin, Daniel DeMenthon, Larry Davis 0001
ICME5
2003 Pitch and timbre manipulations using cortical representation of sound
abstract
The sound receiver at the ears is processed by humans using signal processing that separate the signal along intensity, pitch and timbre dimensions. Conventional Fourier-based signal processing, while endowed with fast algorithms, is unable to easily represent signal along these attributes. In this paper we use a cortical representation to represent the manipulate sound. We briefly overview algorithms for obtaining, manipulating and inverting cortical representation of sound and describe algorithms for manipulating signal pitch and timbre separately. The algorithms are first used to create sound of an instrument between a guitar and a trumpet. Applications to creating maximally separable sounds in auditory user interfaces are discussed.
Dmitry N. Zotkin, Shihab A. Shamma, Powen Ru, Ramani Duraiswami, Larry Davis 0001
ICME5
2003 Predicting Accuracy in Pose Estimation for Marker-based Tracking
abstract
Tracking is a necessity for interactive virtual environments. Marker-based tracking solutions involve the placement of fiducials in a rigid configuration on the object(s) to be tracked, called a tracking probe. The realization that tracking performance is linked to probe performance necessitates investigation into the design of tracking probes for proponents of marker-based tracking. A challenge involved with probe design is predicting the accuracy of a tracking probe. We present a method for predicting the accuracy of a tracking probe based upon a first-order propagation of the errors associated with the markers on the probe. Results for two sample tracking probes show excellent agreement between measured and predicted errors.
Larry Davis 0001, Eric Clarkson, Jannick P. Rolland
ISMAR1
2003 M2Tracker: A Multi-View Approach to Segmenting and Tracking People in a Cluttered Scene
Anurag Mittal, Larry Davis 0001
Int. J. Comput. Vis.2
2003 Efficient Kernel Density Estimation Using the Fast Gauss Transform with Applications to Color Modeling and Tracking
abstract
Many vision algorithms depend on the estimation of a probability density function from observations. Kernel density estimation techniques are quite general and powerful methods for this problem, but have a significant disadvantage in that they are computationally intensive. In this paper, we explore the use of kernel density estimation with the fast Gauss transform (FGT) for problems in vision. The FGT allows the summation of a mixture of ill Gaussians at N evaluation points in O(M+N) time, as opposed to O(MN) time for a naive evaluation and can be used to considerably speed up kernel density estimation. We present applications of the technique to problems from image segmentation and tracking and show that the algorithm allows application of advanced statistical techniques to solve practical vision problems in real-time with today's computers.
Ahmed M. Elgammal, Ramani Duraiswami, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2002 M2Tracker: A Multi-view Approach to Segmenting and Tracking People in a Cluttered Scene Using Region-Based Stereo
Anurag Mittal, Larry Davis 0001
ECCV (1)2
2002 Gesture recognition using a probabilistic framework for pose matching
abstract
This paper presents an approach for view-based recognition of gestures. The approach is based on representing each gesture as a sequence of learned body poses. The gestures are recognized through a probabilistic framework for matching these body poses and for imposing temporal constrains between different poses. Matching individual poses to image data is performed using a probabilistic formulation for edge matching to obtain a likelihood measurement for each individual pose. The paper introduces a weighted matching scheme for edge templates that emphasize discriminating features in the matching. The weighting does not require establishing correspondences between the different pose models. The probabilistic framework also imposes temporal constrains between different pose through a learned Hidden Markov Model (HMM)for each gesture.
Ahmed M. Elgammal, Vhay Shet, Yaser Yacoob, Larry Davis 0001
ICARCV4
2002 Creation of virtual auditory spaces
abstract
High-quality virtual audio scene rendering is a must for emerging virtual/augmented reality applications and for perceptual user interfaces. We describe algorithms for creation of virtual auditory spaces using measured and non-individualized HRTFs and head tracking. Details of algorithms for HRTF interpolation, room impulse response creation, and audio scene presentation are presented. Tests show that individuals externalize well, and find our interface natural. The system runs in real time with latency of less than 30 ms on an office PC without specialized DSP.
Dmitry N. Zotkin, Ramani Duraiswami, Larry Davis 0001
ICASSP3
2002 A Video Based Interface to Textual Information for the Visually Impaired
abstract
We describe the development of an interface to textual information for the visually impaired that uses video, image processing, optical-character-recognition (OCR) and text-to-speech (TTS). The video provides a sequence of low resolution images in which text must be detected, rectified and converted into high resolution rectangular blocks that are capable of being analyzed via off-the-shelf OCR. To achieve this, various problems related to feature detection, mosaicing, auto-focus, zoom, and systems integration were solved in the development of the system.
Ali Zandifar, Ramani Duraiswami, Antoine Chahine, Larry Davis 0001
ICMI4
2002 Dynamic superimposition of synthetic objects on rigid and simple-deformable real objects
Yann Argotti, Larry Davis 0001, Valerie Outters, Jannick P. Rolland
Comput. Graph.2
2002 Background and foreground modeling using nonparametric kernel density estimation for visual surveillance
abstract
Automatic understanding of events happening at a site is the ultimate goal for many visual surveillance systems. Higher level understanding of events requires that certain lower level computer vision tasks be performed. These may include detection of unusual motion, tracking targets, labeling body parts, and understanding the interactions between people. To achieve many of these tasks, it is necessary to build representations of the appearance of objects in the scene. This paper focuses on two issues related to this problem. First, we construct a statistical representation of the scene background that supports sensitive detection of moving objects in the scene, but is robust to clutter arising out of natural scene variations. Second, we build statistical representations of the foreground regions (moving objects) that support their tracking and support occlusion reasoning. The probability density functions (pdfs) associated with the background and foreground are likely to vary from image to image and will not in general have a known parametric form. We accordingly utilize general nonparametric kernel density estimation techniques for building these statistical representations of the background and the foreground. These techniques estimate the pdf directly from the data without any assumptions about the underlying distributions. Example results from applications are presented.
Ahmed M. Elgammal, Ramani Duraiswami, David Harwood, Larry Davis 0001
Proc. IEEE4
2001 A Probabilistic Framework for Surface Reconstruction from Multiple Images
abstract
The paper presents a novel probabilistic framework for 3D surface reconstruction from multiple stereo images. The method works on a discrete voxelized representation of the scene. An iterative scheme is used to estimate the probability that a scene point lies on the true 3D surface. The novelty of our approach lies in the ability to model and recover surfaces which may be occluded in some views. This is done by explicitly estimating the probabilities that a 3D scene point is visible in a particular view from the set of given images. This relies on the fact that for a point on a lambertian surface, if the pixel intensities of its projection along two views differ, then the point is necessarily occluded in one of the views. We present results of surface reconstruction from both real and synthetic image sets.
Motilal Agrawal, Larry Davis 0001
CVPR (2)2
2001 Efficient Non-Parametric Adaptive Color Modeling Using Fast Gauss Transform
abstract
Modeling the color distribution of a homogeneous region is used extensively for object tracking and recognition applications. The color distribution of an object represents a feature that is robust to partial occlusion, scaling and object deformation. A variety of parametric and non-parametric statistical techniques have been used to model color distributions. In this paper we present a non-parametric color modeling approach based on kernel density estimation as well as a computational framework for efficient density estimation. Theoretically, our approach is general since kernel density estimators can converge to any density shape with sufficient samples. Therefore, this approach is suitable to model the color distribution of regions with patterns and mixture of colors. Since kernel density estimation techniques are computationally expensive, the paper introduces the use of the fast Gauss transform for efficient computation of the color densities. We show that this approach can be used successfully for color-based segmentation of body parts as well as segmentation of many people under occlusion.
Ahmed M. Elgammal, Ramani Duraiswami, Larry Davis 0001
CVPR (2)3
2001 Active speech source localization by a dual coarse-to-fine search
abstract
Accurate and fast localization of multiple speech sound sources is a significant problem in videoconferencing systems. Based on the observation that the wavelengths of the sound from a speech source are comparable to the dimensions of the space being searched, and that the source is broadband, we develop an efficient search strategy that finds the source(s) in a given space. The search is made efficient by using coarse-to-fine strategies in both space and frequency. The algorithm is shown to be robust compared to typical delay-based estimators and fast enough for real-time implementation. Its performance can be further improved by using constraints from computer vision.
Ramani Duraiswami, Dmitry N. Zotkin, Larry Davis 0001
ICASSP3
2001 Probabilistic Framework for Segmenting People Under Occlusion
abstract
In this paper we address the problem of segmenting foreground regions corresponding to a group of people given models of their appearance that were initialized before occlusion. We present a general framework that uses maximum likelihood estimation to estimate the best arrangement for people in terms of 2D translation that yields a segmentation for the foreground region. Given the segmentation result we conduct occlusion reasoning to recover relative depth information and we show how to utilize this depth information in the same segmentation framework. We also present a more practical solution for the segmentation problem that is online to avoid searching an exponential space of hypothesis. The person model is based on segmenting the body into regions in order to spatially localize the color-features corresponding to the way people are dressed. Modeling these regions involves modeling their appearance (color distributions) as well us their spatial distribution with respect to the body. We use a non-parametric approach bused on kernel density estimation to represent the color distribution of each region and therefore we do not restrict the clothing to be of uniform color instead it can be any mixture of colors and/or patterns. We also present a method to automatically initialize these models and learn them before the occlusion.
Ahmed M. Elgammal, Larry Davis 0001
ICCV2
2001 Multimodal Tracking For Smart Videoconferencing
abstract
Many applications require the ability to track the 3-D motion of the subjects. We build a particle filter based framework for multimodal tracking using multiple cameras and multiple microphone arrays. In order to calibrate the resulting system, we propose a method to determine the locations of all microphones using at least five loudspeakers and under assumption that for each loudspeaker there exists a microphone very close to it. We derive the maximum likelihood (ML) estimator, which reduces to the solution of the non-linear least squares problem. We verify the correctness and robustness of the multimodal tracker and of the self-calibration algorithm both with Monte-Carlo simulations and on real data from three experimental setups. 1.
Dmitry N. Zotkin, Ramani Duraiswami, Harsh Nanda, Larry Davis 0001
ICME4
2001 Technologies for Augmented Reality: Calibration for Real-Time Superimposition on Rigid and Simple-Deformable Real Objects
Yann Argotti, Valerie Outters, Larry Davis 0001, Ami Sun, Jannick P. Rolland
MICCAI3
2001 Backpack: Detection of People Carrying Objects Using Silhouettes
Ismail Haritaoglu, Ross Cutler, David Harwood, Larry Davis 0001
Comput. Vis. Image Underst.4
2000 Robust Periodic Motion and Motion Symmetry Detection
abstract
We describe a robust technique for detecting nonstationary periodic motion from a moving and static camera. We also describe a robust technique for discriminating motion symmetries (periodic motion classification), which we apply to classifying running humans (bipeds) and canines (quadrupeds). The system has been implemented to run in real-time (30 Hz) on standard PC workstations.
Ross Cutler, Larry Davis 0001
CVPR2
2000 Non-parametric Model for Background Subtraction
Ahmed M. Elgammal, David Harwood, Larry Davis 0001
ECCV (2)3
2000 Quasi-Random Sampling for Condensation
Vasanth Philomin, Ramani Duraiswami, Larry Davis 0001
ECCV (2)3
2000 A Probabilistic Framework for Rigid and Non-Rigid Appearance Based Tracking and Recognition
abstract
This paper describes an unified probabilistic framework for appearance-based tracking of rigid and non-rigid objects. A spatio-temporal dependent shape-texture eigenspace and mixture of diagonal Gaussians are learned in a hidden Markov model (HMM)-like structure to better constrain the model and for recognition purposes. Particle filtering is used to track the object while switching between different shape/texture models. This framework allows recognition and temporal segmentation of activities. Additionally an automatic stochastic initialization is proposed, the number of states in the HMM are selected based on the Akaike information criterion and comparison with deterministic tracking for 2D models is discussed. Preliminary results of eye tracking, lip tracking and temporal segmentation of mouth events are presented.
Fernando De la Torre, Yaser Yacoob, Larry Davis 0001
FG3
2000 Tracking Humans from a Moving Platform
abstract
Research at the Computer Vision Laboratory at the University of Maryland has focussed on developing algorithms and systems that can look at humans and recognize their activities in near real-time. Our earlier implementation while quite successful, was restricted to applications with a fixed camera. In this paper we present some recent work that removes this restriction. Such systems are required for machine vision from moving platforms such as robots, intelligent vehicles, and unattended large field of regard cameras with a small field of view. Our approach is based on the use of a deformable shape model for humans coupled with a novel variant of the condensation algorithm that uses quasi-random sampling for efficiency. This allows the use of simple motion models which results in algorithm robustness, enabling us to handle unknown camera/human motion with unrestricted camera viewing angles. We present the details of our human tracking algorithms and some examples from pedestrian tracking and automated surveillance.
Larry Davis 0001, Vasanth Philomin, Ramani Duraiswami
ICPR1
2000 A Fast Background Scene Modeling and Maintenance for Outdoor Surveillance
abstract
We describe fast background scene modeling and maintenance techniques for real time visual surveillance system for tracking people in an outdoor environment. It operates on monocular gray scale video imagery or on video imagery from an infrared camera. The system learns and models background scene statistically to detect foreground objects, even when the background is not completely stationary (e.g. motion of tree branches) using shape and motion cues. Also, a background maintenance model is proposed for preventing false positives, such as, illumination changes (the sun being blocked by clouds causing changes in brightness), or false negative, such as, physical changes (person detection while he is getting out of the parked car). Experimental results demonstrate robustness and real-time performance of the algorithm.
Ismail Haritaoglu, David Harwood, Larry Davis 0001
ICPR3
2000 An Appearance-Based Body Model for Multiple People Tracking
abstract
We describe an appearance-based human body model for tracking multiple people when they are interaction with each others causing significant occlusion amongst them, or when they re-enter the scene. The proposed model allows real time surveillance systems understand "who is who" after multiple people interactions, partial or total occlusions, or when a person "reappeal" in the scene. It combines the grayscale textural appearance and expected shape information together in a 2D dynamic template. Experimental results demonstrates the robustness and real time performance of the proposed model.
Ismail Haritaoglu, David Harwood, Larry Davis 0001
ICPR3
2000 An audio-video front-end for multimedia applications
abstract
Applications such as video gaming, virtual reality, multimodal user interfaces and videoconferencing, require systems that can locate and track persons in a room through a combination of visual and audio cues, enhance the sound that they produce, and perform identification. We describe the development of a particular multimodal sensor fusion system that is portable, runs in real time and achieves these objectives. The system employs novel algorithms for acoustical source location, video-based person tracking and overall system control, which are also described.
Dmitry N. Zotkin, Ramani Duraiswami, Larry Davis 0001, Ismail Haritaoglu
SMC3
2000 Learned Models for Estimation of Rigid and Articulated Human Motion from Stationary or Moving Camera
Yaser Yacoob, Larry Davis 0001
Int. J. Comput. Vis.2
2000 Real-time multiple vehicle detection and tracking from a moving vehicle
Margrit Betke, Esin Haritaoglu, Larry Davis 0001
Mach. Vis. Appl.3
2000 Robust Real-Time Periodic Motion Detection, Analysis, and Applications
abstract
We describe new techniques to detect and analyze periodic motion as seen from both a static and a moving camera. By tracking objects of interest, we compute an object's self-similarity as it evolves in time. For periodic motion, the self-similarity measure is also periodic and we apply time-frequency analysis to detect and characterize the periodic motion. The periodicity is also analyzed robustly using the 2D lattice structures inherent in similarity matrices. A real-time system has been implemented to track and classify objects using periodicity. Examples of object classification (people, running dogs, vehicles), person counting, and nonstationary periodicity are provided.
Ross Cutler, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 W4: Real-Time Surveillance of People and Their Activities
abstract
W/sup 4/ is a real time visual surveillance system for detecting and tracking multiple people and monitoring their activities in an outdoor environment. It operates on monocular gray-scale video imagery, or on video imagery from an infrared camera. W/sup 4/ employs a combination of shape analysis and tracking to locate people and their parts (head, hands, feet, torso) and to create models of people's appearance so that they can be tracked through interactions such as occlusions. It can determine whether a foreground region contains multiple people and can segment the region into its constituent people and track them. W/sup 4/ can also determine whether people are carrying objects, and can segment objects from their silhouettes, and construct appearance models for them so they can be identified in subsequent frames. W/sup 4/ can recognize events between people and objects, such as depositing an object, exchanging bags, or removing an object. It runs at 25 Hz for 320/spl times/240 resolution images on a 400 MHz dual-Pentium II PC.
Ismail Haritaoglu, David Harwood, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
1999 Real-Time Periodic Motion Detection, Analysis, and Applications
abstract
We describe a new technique to detect and analyze periodic motion as seen from both a static and moving camera. By tracking objects of interest, we compute an object's self-similarity as it evolves in time. For periodic motion, the self-similarity measure is also periodic, and we apply time-frequency analysis to detect and characterize the periodic motion. A real-time system has been implemented to track and classify objects using periodicity. Examples of object classification, person counting, and non-stationary periodicity are provided.
Ross Cutler, Larry Davis 0001
CVPR2
1999 Backpack: Detection of People Carrying Objects using Silhouettes
abstract
We described a video-rate surveillance algorithm to detect and track people from a stationary camera, and to determine if they are carrying objects or moving unencumbered. The contribution of the paper is the shape analysis algorithm that both determines if a person is carrying an object and segments the object from the person so that it can be tracked, e.g., during an exchange of objects between two people. As the object is segmented an appearance model of the object is constructed. The method combines periodic motion estimation with static symmetry analysis of the silhouettes of a person in each frame of the sequence. Experimental results demonstrate robustness and real-time performance of the proposed algorithm.
Ismail Haritaoglu, Ross Cutler, David Harwood, Larry Davis 0001
ICCV4
1999 Estimation of Composite Object and Camera Image Motion
abstract
An approach for estimating composite independent object and camera image motions is proposed. The approach employs spatio-temporal flow models learned through observing typical movements of the object, to decompose image motion into independent object and camera motions. The spatio-temporal flow models of the object motion are represented as a set of orthogonal flow bases that are learned using principal component analysis of instantaneous flow measurements from a stationary camera. These models are then employed in scenes with a moving camera to extract motion trajectories relative to those learned. The performance of the algorithm is demonstrated on several image sequences of rigid and articulated bodies in motion.
Yaser Yacoob, Larry Davis 0001
ICCV2
1999 Tracking Rigid Motion using a Compact-Structure Constraint
abstract
An approach for tracking the motion of a rigid object using parameterized flow models and a compact-structure constraint is proposed. While polynomial parameterized flow models have been shown to be effective in tracking the rigid motion of planar objects, these models are inappropriate for tracking moving objects that change appearance revealing their 3D structure. We extend these models by adding a structure-compactness constraint that accounts for image motion that deviates from a planar structure. The constraint is based on the assumption that object structure variations are limited with respect to planar object projection onto the image plane and therefore can be expressed as a direct constraint on the image motion. The performance of the algorithm is demonstrated on several long image sequences of rigidly moving objects.
Yaser Yacoob, Larry Davis 0001
ICCV2
1999 Detection of Independent Motion Using Directional Motion Estimation
Sándor Fejes, Larry Davis 0001
Comput. Vis. Image Underst.2
1999 Temporal Multi-Scale Models for Flow and Acceleration
Yaser Yacoob, Larry Davis 0001
Int. J. Comput. Vis.2
1998 Visual Surveillance of Human Activity
Larry Davis 0001, Sándor Fejes, David Harwood, Yaser Yacoob, Ismail Haritaoglu, Michael J. Black
ACCV (2)1
1998 Interpretation of Complex Scenes Using Bayesian Networks
Mark F. Westling, Larry Davis 0001
ACCV (2)2
1998 W4: A Real Time System for Detecting and Tracking People
Ismail Haritaoglu, David Harwood, Larry Davis 0001
CVPR3
1998 W4S: A real-time system detecting and tracking people in 2 1/2D
Ismail Haritaoglu, David Harwood, Larry Davis 0001
ECCV (1)3
1998 W4: Who? When? Where? What? A Real Time System for Detecting and Tracking People
Ismail Haritaoglu, David Harwood, Larry Davis 0001
FG3
1998 What Can Projections of Flow Fields Tell Us About Visual Motion
abstract
The dimensionality of visual motion analysis can be reduced by analyzing projections of flow vector fields. In contrast to motion vector fields, these projections exhibit simple geometric properties which are invariant to the scene structure and depend only on the camera motion. Using these properties, structure and motion can be either completely or partially decoupled. We estimate motion parameters from projections of flow fields by using robust techniques, implemented an a reclusive observer model. The model is applicable to general camera motion and to large field of view and requires no point correspondence. We demonstrate our projection method on the problem of detecting independently moving objects from a moving camera. Using the projection approach, the problem can be reduced to a one-dimensional optimization process which involves robust line-fitting and outlier detection. Instantaneous detection measurements are integrated temporally using tracking and spatially applying grouping of coherently moving points.
Sándor Fejes, Larry Davis 0001
ICCV2
1998 Learned Temporal Models of Image Motion
abstract
An approach for learning and estimating temporal-flow models from image sequences is proposed. The temporal-flow models are represented as a set of orthogonal temporal-flow bases that are learned using principal component analysis of instantaneous flow measurements. Spatial constraints on the temporal-flow are also developed for modeling the motion of regions in rigid and coordinated motion. The performance of these models is demonstrated on several long image sequences of rigid and articulated bodies in motion.
Yaser Yacoob, Larry Davis 0001
ICCV2
1998 View-based detection and analysis of periodic motion
abstract
We describe a technique that detects periodic motion. Assuming a static camera, we first segment moving objects from the background. By tracking objects of interest, we compute the object's self-similarity as it evolves in time. For periodic motion, the self-similarity metric is periodic, and is Fourier analyzed to detect and characterize periodicity. Examples on real image sequences are given.
Ross Cutler, Larry Davis 0001
ICPR2
1998 Ghost: a human body part labeling system using silhouettes
abstract
Ghost is a real time system for estimating human body posture and detecting body parts in monochromatic imagery. It constructs a silhouette based body model to determine the location of the body parts while people are in generic postures. It combines hierarchical body pose estimation, a convex hull analysis of the silhouette, and a partial mapping from the body parts to the silhouette segments using a distance transform method that does not violate the topology of the human body. Experimental results demonstrate robustness and real-time performance of the proposed algorithm.
Ismail Haritaoglu, David Harwood, Larry Davis 0001
ICPR3
1998 Appearance-based automatic target recognition in overhead LADAR range imagery
abstract
This paper presents an algorithm for automatic target recognition in, overhead LADAR range imagery using "appearance-based" models. The detection is accomplished using a two stage Hough transform where the target is modelled as a small set of boundary pixels with associated local properties. The first stage Hough transform keeps tracks of all the edge points matched by the model points. The second stage Hough transform then efficiently identifies the candidate space by incrementing the accumulator array along precomputed (for each model point) swaths. A robust estimator, the MAD, is then used to do the template matching directly on the second stage solutions which are already very accurate. Finally, we illustrate our algorithm on some real LADAR images and discuss the error rates.
Vasanth Philomin, David Harwood, Larry Davis 0001
ICPR3
1998 The Design and Evaluation of a High-Performance Earth Science Database
Carter Shock, Chialin Chang, Bongki Moon, Anurag Acharya 0001, Larry Davis 0001, Joel H. Saltz, Alan Sussman
Parallel Comput.5
1997 Temporal Multi-scale Models for Flow and Acceleration
abstract
A model for computing image flow in image sequences containing a very wide range of instantaneous flows is proposed. This model integrates the spatio-temporal image derivatives from multiple temporal scales to provide both reliable and accurate instantaneous flow estimates. The integration employs robust regression and automatic scale weighting in a generalized brightness constancy framework. In addition to instantaneous flow estimation the model supports recovery of dense estimates of image acceleration and can be readily combined with parameterized flow and acceleration models. A demonstration of performance on image sequences of typical human actions taken with a high frame-rate camera, is given.
Yaser Yacoob, Larry Davis 0001
CVPR2
1996 3-D model-based tracking of humans in action: a multi-view approach
abstract
We present a vision system for the 3-D model-based tracking of unconstrained human movement. Using image sequences acquired simultaneously from multiple views, we recover the 3-D body pose at each time instant without the use of markers. The pose-recovery problem is formulated as a search problem and entails finding the pose parameters of a graphical human model whose synthesized appearance is most similar to the actual appearance of the real human in the multi-view images. The models used for this purpose are acquired from the images. We use a decomposition approach and a best-first technique to search through the high dimensional pose parameter space. A robust variant of chamfer matching is used as a fast similarity measure between synthesized and real edge images. We present initial tracking results from a large new Humans-in-Action (HIA) database containing more than 2500 frames in each of four orthogonal views. They contain subjects involved in a variety of activities, of various degrees of complexity, ranging from the more simple one-person hand waving to the challenging two-person close interaction in the Argentine Tango.
Dariu Gavrila, Larry Davis 0001
CVPR2
1996 Computing 3-D head orientation from a monocular image sequence
abstract
An approach for estimating 3D head orientation in a monocular image sequence is proposed. The approach employs recently developed image-based parameterized tracking for face and face features to locate the area in which a sub-pixel parameterized shape estimation of the eye's boundary is performed. This involves tracking of five points (four at the eye corners and the fifth is the lip of the nose). The authors describe an approach that relies on the coarse structure of the face to compute orientation relative to the camera plane. Our approach employs projective invariance of the cross-ratios of the eye corners and anthropometric statistics to estimate the head yaw, roll and pitch. Analytical and experimental results are reported.
Thanarat H. Chalidabhongse, Yaser Yacoob, Larry Davis 0001
FG3
1996 Recognition of head gestures using hidden Markov models
abstract
This paper explores the use of hidden Markov models (HMMs) for the recognition of head gestures. A gesture corresponds to a particular pattern of head movement. The facial plane is tracked using a parameterized model and the temporal sequence of three image rotation parameters are used to describe four gestures. A dynamic vector quantization scheme was implemented to transform the parameters into suitable input data for the HMMs. Each model was trained by the iterative Baum-Welch procedure using 28 sequences taken from 5 persons. Experimental results from a different data set (33 new sequences from 6 other persons) demonstrate the effectiveness of this approach.
Carlos Hitoshi Morimoto, Yaser Yacoob, Larry Davis 0001
ICPR3
1996 Object recognition by fast hypothesis generation and reasoning about object interactions
abstract
We present a two-step approach for recognizing multiple 3-D objects in single 2-D images. In the first step, hypotheses of object instances are generated using a memory-based technique. This technique relies on an array, which is computed off-line, that associates a large number of object poses with corresponding image features. During actual recognition, the array serves as a discrete approximation of the inverse projection function, and each image feature returns a set of poses that are accumulated by a generalized Hough transform. In the second step, the configuration of hypotheses that best interprets the image is calculated using a Bayesian network. The network represents both visual effects, such as the creation and occlusion of image features, and physical constraints, such as object interference.
Mark F. Westling, Larry Davis 0001
ICPR2
1996 Iterative Pose Estimation Using Coplanar Feature Points
Denis Oberkampf, Daniel DeMenthon, Larry Davis 0001
Comput. Vis. Image Underst.3
1996 Recognizing Human Facial Expressions From Long Image Sequences Using Optical Flow
abstract
An approach to the analysis and representation of facial dynamics for recognition of facial expressions from image sequences is presented. The algorithms utilize optical flow computation to identify the direction of rigid and nonrigid motions that are caused by human facial expressions. A mid-level symbolic representation motivated by psychological considerations is developed. Recognition of six facial expressions, as well as eye blinking, is demonstrated on a large set of image sequences.
Yaser Yacoob, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
1996 Parallel algorithms for image enhancement and segmentation by region growing, with an experimental study
David A. Bader, Joseph F. JáJá, David Harwood, Larry Davis 0001
J. Supercomput.4
1996 An improved radial basis function network for visual autonomous road following
abstract
We have developed a radial basis function network (RBFN) for visual autonomous road following. Preliminary testing of the RBFN was done using a driving simulator, and the RBFN was then installed on an actual vehicle at Carnegie Mellon University for testing in an outdoor road-following application. In our first attempts, the RBFN had some success, but it experienced some significant problems such as jittery control and driving failure. Several improvements have been made to the original RBFN architecture to overcome these problems in simulation and more importantly in actual road following, and the improvements are described in this paper.
Mark Rosenblum, Larry Davis 0001
IEEE Trans. Neural Networks2
1996 Human expression recognition from motion using a radial basis function network architecture
abstract
In this paper a radial basis function network architecture is developed that learns the correlation of facial feature motion patterns and human expressions. We describe a hierarchical approach which at the highest level identifies expressions, at the mid level determines motion of facial features, and at the low level recovers motion directions. Individual expression networks were trained to recognize the "smile" and "surprise" expressions. Each expression network was trained by viewing a set of sequences of one expression for many subjects. The trained neural network was then tested for retention, extrapolation, and rejection ability. Success rates were 88% for retention, 88% for extrapolation, and 83% for rejection.
Mark Rosenblum, Yaser Yacoob, Larry Davis 0001
IEEE Trans. Neural Networks3
1995 Model-based object pose in 25 lines of code
Daniel DeMenthon, Larry Davis 0001
Int. J. Comput. Vis.2
1995 Egomotion analysis based on the Frenet-Serret motion model
Zoran Duric, Azriel Rosenfeld, Larry Davis 0001
Int. J. Comput. Vis.3
1995 Texture classification by center-symmetric auto-correlation, using Kullback discrimination of distributions
David Harwood, Timo Ojala, Matti Pietikäinen, Shalom Kelman, Larry Davis 0001
Pattern Recognit. Lett.5
1994 Site model supported monitoring of aerial images
abstract
Image monitoring, the process of locating and identifying significant changes or new activities, is one of the most important imagery exploitation tasks. A site model supported image monitoring system which utilizes image understanding techniques driven by an underlying site model is presented. In our approach, we first register the image to be monitored to an existing site model, which is constructed using the RADIUS Common Development Environment; the regions of interest are then delineated based on site information, camera acquisition parameters, and goals of the image analyst; object extraction is then done using constraints on size, shape, orientation, and shadow of the target object derived from known information about image resolution, 3-D shape of the object, camera viewing and illuminant directions. The results of object detection are used for monitoring changes.>
C. L. Lin, Qinfen Zheng, Rama Chellappa, Larry Davis 0001, Xiaopeng Zhang 0006
CVPR4
1994 Computing spatio-temporal representations of human faces
abstract
An approach for analysis and representation of facial dynamics for recognition of facial expressions from image sequences is proposed. The algorithms we develop utilize optical flow computation to identify the direction of rigid and non-rigid motions that are caused by human, facial expressions. A mid-level symbolic representation that is motivated by linguistic and psychological considerations is developed. Recognition of six facial expressions, as well as eye blinking, on a large set of image sequences is reported.>
Yaser Yacoob, Larry Davis 0001
CVPR2
1994 High performance computing for land cover dynamics
abstract
Presents the overall goals of the authors' research program on the application of high performance computing to remote sensing applications, specifically applications in land cover dynamics. This involves developing scalable and portable programs for a variety of image and map data processing applications, eventually integrated with new models for parallel I/O of large scale images and maps. After an overview of the multiblock PARTI run time support system, the authors explain extensions made to that system to support image processing applications, and then present an example involving multiresolution image processing. Results of running the parallel code on both a TMC CM5 and an Intel Paragon are discussed.
Rahul Parulekar, Larry Davis 0001, Rama Chellappa, Joel H. Saltz, Alan Sussman, John Townshend
ICPR (3)2
1994 Recognizing facial expressions by spatio-temporal analysis
abstract
An approach for analysis and representation of facial dynamics for recognition of facial expressions from image sequences is proposed. The algorithms the authors develop utilize optical flow computation to identify the direction of rigid and non-rigid motions that are caused by human facial expressions. A mid-level symbolic representation that is motivated by linguistic and psychological considerations is developed. Recognizing six facial expressions, as well as eye blinking, are demonstrated on a collection of image sequences.
Yaser Yaccob, Larry Davis 0001
ICPR (1)2
1994 Parallel search for the interpretation of aerial images
abstract
Abstract In this paper, we present a parallel search scheme for model‐based interpretation of aerial images, following a focus‐of‐attention paradigm. Interpretation is performed using the gray level image of an aerial scene and its segmentation into connected components of almost constant gray level. Candidate objects are generated from the window as connected combinations of its components. Each candidate is matched against the model by checking if the model constraints are satisfied by the parameters computed from the region. The problem of candidate generation and matching is posed as searching in the space of combinations of connected components in the image, with finding an (optimally) successful region as the goal. Our implementation exploits parallelism at multiple levels by parallelizing the management of the open list and other control tasks as well as the task of model matching. We discuss and present the implementation of the interpretation system on a Connection Machine CM‐2. The implementation reported a successful match in a few hundred milliseconds whenever they existed.
P. J. Narayanan, Larry Davis 0001
Concurr. Pract. Exp.2
1994 Parallel Terrain Triangulation
abstract
Digital Elevation Models are considered in relation to their use in a parallel computing environment. In particular, the problem of approximating terrain surface through a Triangulated Irregular Network (TIN) is analysed. A parallel algorithm is presented that builds a TIN based on Delaunay triangulation, by selecting a sparse subset of points from a dense regular grid of sampled data. An implementation of the algorithm on a CM-2 is described and experimental results are shown.
Enrico Puppo, Larry Davis 0001, Daniel DeMenthon, Y. Ansel Teng
Int. J. Geogr. Inf. Sci.2
1993 Iterative pose estimation using coplanar points
abstract
A new method is presented for the computation of the position and orientation of a camera with respect to a known object, using four or more coplanar feature points. Starting under the scaled orthographic projection approximation, this method iteratively refines up to two different pose estimates, and provides associated quality measures. When the distance of the object to the camera is large, or when the accuracy of the feature point extraction is low, both pose estimates are plausible.>
Denis Oberkampf, Daniel DeMenthon, Larry Davis 0001
CVPR3
1993 Early vision processing using a multi-stage diffusion process
abstract
The use of a multistage diffusion process in the early processing of range data is examined. The input range data are interpreted as occupying a volume in 3-D space. Each diffusion stage simulates the process of diffusing part of the boundary of the volume into the volume. The outcome of the process can be used for both discontinuity detection and segmentation into shape homogeneous regions. The process is applied to synthetic noise-free and noisy step, roof, and valley edges as well as to real range images.>
Yaser Yacoob, Larry Davis 0001
CVPR2
1993 Labeling of human face components from range data
abstract
An approach to labeling the components of human faces from range images is proposed. The components of interest are those humans usually find significant for recognition. To cope with the nonrigidity of faces, a qualitative approach is used. The preprocessing stage employs a multi-stage diffusion process to identify convexity and concavity points. These points are grouped into components and qualitative reasoning about possible interpretations of the components is performed. Consistency of hypothesized interpretations is carried out using context-based reasoning. Experimental results on real images of several faces are provided.>
Yaser Yacoob, Larry Davis 0001
CVPR2
1993 Egomotion analysis based on the Frenet-Serret motion model
abstract
A new model, Frenet-Serret motion, is proposed for the motion of an observer in a stationary environment. This model relates the motion parameters of the observer to the curvature and torsion of the path along which the observer moves. Screw-motion equations for Frenet-Serret motion are derived and employed for geometrical analysis of the motion. Normal flow is used to derive constraints on the rotational and translational velocity of the observer and to compute egomotion by intersecting these constraints in the manner proposed by Z. Duric/spl acute/ and Y. Aloimonos (1989). The accuracy of egomotion estimation is analyzed for different combinations of observer motion and feature distance. The authors explain the advantages of controlling feature distance to analyze egomotion and derive the constraints on depth which make either rotation or translation dominant in the perceived normal flow field. The results of experiments on real image sequences are presented.>
Zoran Duric, Azriel Rosenfeld, Larry Davis 0001
ICCV3
1993 Region-to-region visibility analysis using data parallel machines
abstract
Abstract We propose an algorithm for solving region‐to‐region visibility problems on digital terrain models using data parallel machines. Since global communication is the bottleneck in this kind of algorithm, the algorithm we propose focuses on the reduction of global communication. The algorithm analyses a strip of the source region at a time and sweeps through the source strip by strip. At most four sweeps are needed for the analysis. By exploring the coherence properties in the processor structure, global communication is minimized and complexity is substantially improved. Furthermore, all global write operations are exclusive and concurrency in global read operations is minimized. Since the problem size is usually large, we also designed rules of decomposition to efficiently handle the cases where the required number of processors is greater than available. The algorithm has been implemented on a Connection Machine CM‐2, and results of computational experiments are presented.
Y. Ansel Teng, Daniel DeMenthon, Larry Davis 0001
Concurr. Pract. Exp.3
1993 A Parallel Algorithm for the Visibility of a Simple Polygon Using Scan Operations
Ling Tony Chen, Larry Davis 0001
CVGIP Graph. Model. Image Process.2
1993 Efficient Parallel Processing of Image Contours
abstract
Describes two parallel algorithms for ranking the pixels on a curve in O (log N) time using either an EREW or CREW PRAM model. The algorithms accomplish this with N processors for a square root N* square root N image. After applying such an algorithm to an image, it is possible to move the pixels from a curve into processors having consecutive addresses. This is important because one can subsequently apply many algorithms to the curve (such as piecewise linear approximation algorithms or point in polygon tests) using segmented scan operations (i.e. parallel prefix operations). Scan operations can be executed in logarithmic time on many interconnection networks, such as hypercube, tree, butterfly, and shuffle exchange machines as well as on the EREW PRAM. The algorithms were implemented on the hypercube structured Connection Machine, and various performance tests were conducted.>
Ling Tony Chen, Larry Davis 0001, Clyde P. Kruskal
IEEE Trans. Pattern Anal. Mach. Intell.2
1993 Parallel curve matching on the Connection Machine
Ling Tony Chen, Larry Davis 0001
Pattern Recognit. Lett.2
1993 Stealth terrain navigation
abstract
A method for solving visibility-based terrain path planning problems using massively parallel hypercube machines is proposed. A typical example is to find a path that is hidden from moving adversaries. This kind of problem can be generalized as a time-varying constrained path planning problem and is proven to be computationally hard. An approximation based on both temporal and, spatial sampling is proposed. Since a 2-D grid cell representation of terrain can be embedded into a hypercube with extra links for fast communication, the method can be very efficient when implemented on hypercube machines. The time complexity is in general O(T*E*log N) using O(N) processors, where T is the number of temporal samples, E is the number of adversary agents, and N is the number of grid cells on the terrain. It is also shown that the method can be applied to several realistic problems with a variety of path optimizations. All algorithms have been implemented on the Connection Machine CM-2 and results of experiments are presented.>
Y. Ansel Teng, Daniel DeMenthon, Larry Davis 0001
IEEE Trans. Syst. Man Cybern.3
1992 Computational ground and airborne localization over rough terrain
abstract
An approach for autonomous localization of ground vehicles on natural terrain is proposed. The localization problem is solved using measurements including attitude, heading, and distances to specific environmental points. The algorithm utilizes random acquisition of distance measurements to prune the possible location(s) of the viewer. The approach is also applicable to airborne localization.>
Yaser Yacoob, Larry Davis 0001
CVPR2
1992 Model-Based Object Pose in 25 Lines of Code
Daniel DeMenthon, Larry Davis 0001
ECCV2
1992 Probabilistic navigation methods for uncertain and dynamic environments
abstract
The trajectory planning problem for mobile robots in unknown dynamic workspaces is posed as an optimization problem with optimality criteria the probability of not colliding with the obstacles and the probability of accessing an operational position with respect to a moving target object. The authors study a formal computational framework in which such probabilities can be derived for elementary robot displacements.>
Philippe Burlina, Daniel DeMenthon, Larry Davis 0001
ICPR (1)3
1992 Rank order filtering on SIMD machines
abstract
Rank order filters form an important class of low level image operations that have widespread applications in image smoothing, texture analysis, etc. In the paper, the authors study several ways of computing rank order filters on processor array architectures. They also present a replicated data algorithm for efficient processing of small images on relatively large processor arrays. Results of implementing the algorithms on a Connection Machine CM-2 and a Mas-Par MP-1 are presented.>
P. J. Narayanan, Larry Davis 0001
ICPR (4)2
1992 Ground and airborne localization over rough terrain using random environmental range-measurements
abstract
The authors propose an approach for autonomous localization of ground systems on natural terrain. The localization problem is solved using measurements including altitude, heading and distances to specific environmental points. The algorithm utilizes random acquisition of distance measurements to prune the possible location(s) of the viewer. The proposed approach is also applicable to airborne localization. Experiments on a 512*512 terrain are provided.>
Yaser Yacoob, Larry Davis 0001
ICPR (1)2
1992 Navigation with uncertainty: reaching a goal in a high collision risk region
abstract
The authors describe a computational framework in which a probabilistic method for noisy sensor-based robotic navigation in dynamic environments can be devised. The aim of the method is to generate an optimal trajectory by considering as optimality criteria the probability of not colliding with the obstacles and the probability of accessing an operational position with respect to a moving target object. A formal framework in which the probability of collision associated with an elementary robot displacement can be calculated is discussed. Estimates on the obstacle kinematic parameters and measures of confidence on these estimates are used to produce the probability of collision associated with any robot displacement. The probability of collision is derived in two steps: a stochastic model is defined in the kinematic state space of the obstacles and collision events are given a simple geometric characterization in this state space.>
Philippe Burlina, Daniel DeMenthon, Larry Davis 0001
ICRA3
1992 Replicated data algorithms in image processing
P. J. Narayanan, Larry Davis 0001
CVGIP Image Underst.2
1992 Replicated Image Algorithms and Their Analyses on SIMD Machines
abstract
Data parallel processing on processor array architectures has gained popularity in data intensive applications, such as image processing and scientific computing, as massively parallel processor array machines became feasible commercially. The data parallel paradigm of assigning one processing element to each data element results in an inefficient utilization of a large processor array when a relatively small data structure is processed on it. The large degree of parallelism of a massively parallel processor array machine does not result in a faster solution to a problem involving relatively small data structures than the modest degree of parallelism of a machine that is just as large as the data structure. We presented data replication technique to speed up the processing of small data structures on large processor arrays. In this paper, we present replicated data algorithms for digital image convolutions and median filtering, and compare their performance with conventional data parallel algorithms for the same on three popular array interconnection networks, namely, the 2-D mesh, the 3-D mesh, and the hypercube.
P. J. Narayanan, Larry Davis 0001
Int. J. Pattern Recognit. Artif. Intell.2
1992 Stealth Terrain Navigation for Multi-Vehicle Path Planning
abstract
In this paper, we propose a method for solving visibility-based terrain path planning problems for groups of vehicles using data parallel machines. The discussion focuses on path planning for two groups of vehicles so that they move in a bounding overwatch manner. Furthermore, the planned paths for the vehicles themselves are subject to intervisibility constraints, configuration constraints, and different terrain traversabilities due to variations in terrain type and slope. A spatial-temporal sampling approach is adopted to discretize the solution space and facilitate fast computation on a data parallel machine. One of the key computations in the planning is the region-to-region visibility analysis, which is computationally expensive but essential to the choice of subgoals to carry out reconnaissance activities. A parallel algorithm for this analysis is developed. By reducing the communication complexity, our algorithm achieves much faster running time than traditional methods. The algorithms are implemented on a Connection Machine CM-2, and the experimental results show that the planning system effectively generates good paths.
Y. Ansel Teng, Daniel DeMenthon, Larry Davis 0001
Int. J. Pattern Recognit. Artif. Intell.3
1992 Exact and Approximate Solutions of the Perspective-Three-Point Problem
abstract
Model-based pose estimation techniques that match image and model triangles require large numbers of matching operations in real-world applications. The authors show that by using approximations to perspective, 2D lookup tables can be built for each of the triangles of the models. An approximation called 'weak perspective' has been applied previously to this problem; the authors consider two other perspective approximations: paraperspective and orthoperspective. These approximations produce lower errors for off-center image features than weak perspective.>
Daniel DeMenthon, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
1992 Surface reconstruction (of rough terrain) in range image shadows
Behzad Kamgar-Parsi, P. J. Narayanan, Larry Davis 0001
Pattern Recognit. Lett.3
1991 Textural analysis of range images
Sunil Arya, Daniel DeMenthon, Peter Meer, Larry Davis 0001
Pattern Recognit. Lett.4
1991 Parallel algorithms for testing if a point is inside a closed curve
Ling Tony Chen, Larry Davis 0001
Pattern Recognit. Lett.2
1991 Parallel calculation of 3-D pose of a known object in a single view
Kari Pehkonen, David Harwood, Larry Davis 0001
Pattern Recognit. Lett.3
1990 Inverse Perspective Of A Triangle: New Exact And Approximate Solutions
Daniel DeMenthon, Larry Davis 0001
ECCV2
1990 Speedup Analysis of Centralized Parallel Heuristic Search Algorithms
Shie-rei Huang, Larry Davis 0001
ICPP (3)2
1990 Connection machine vision-Replicated data structures
abstract
The problem of efficiently processing small data structures on massively parallel single-instruction multiple-data machines using replication methods is discussed. The problem stems from considerations of both multiresolution vision systems and focus of attention vision systems. A general framework for developing replicated algorithms, based on the four steps of embedding, distribution, decomposition, and collection, is described. A simple example is provided based on computing the histogram of a gray-level image. Replicated chain processing is discussed, and an efficient algorithm for ranking the elements in a chain in log (n) time on a concurrent write parallel random access machine is presented.>
Larry Davis 0001, Ling Tony Chen, P. J. Narayanan
ICPR (2)1
1990 New exact and approximate solutions of the three-point perspective problem
abstract
An exact method for computing the position of a triangle in space from its image is presented. Also presented is an approximate method based on orthoperspective, an approximation of perspective which produces lower errors for off-center triangle images than scaled orthographic projection. A comparison is made of exact and approximate solutions for the triangle pose. This comparison gives the relative combinations of image and triangle characteristics which are likely to generate the largest errors. Model-based pose estimation techniques which match image and model triangles require large numbers of matching operations in real-world applications. It is shown that the approximate model can be used to build lookup tables for each of the triangles of a model and that they speed up the estimation of an object pose.>
Daniel DeMenthon, Larry Davis 0001
ICRA2
1990 Reconstruction of a road by local image matches and global 3D optimization
abstract
A method is presented for reconstructing a 3-D road from a single image. It finds the images of opposite points of the road. Opposite points are points which face each other on the opposite sides of the road; the images of these points are called matching points. For points chosen from one side of the road image, the algorithm finds all the matching point candidates on the other side, based on local properties of a road. However, these solutions do not necessarily satisfy the global properties of a typical road. A dynamic programming algorithm is applied to reject the candidates which do not fit the global road. A benchmark using synthetic roads is described. It shows that the roads reconstructed by the proposed method match the actual roads better than those reconstructed by two other road reconstruction algorithms. Experiments with 50 road images taken by the autonomous land vehicle (ALV) showed that the method is robust with real-world data and that the reconstructions are fairly consistent with road profiles obtained by fusion between range images and video images.>
Daniel DeMenthon, Larry Davis 0001
ICRA2
1990 Fast range scanner using an optic RAM
abstract
A range scanner which calculates ranges by triangulation between the incident angles of laser stripes and the positions of their images on a camera sensor was developed for robotic applications. It uses a solid-state image sensor called Optic RAM instead of a CCD sensor. This sensor chip has three desirable characteristics for the position detection of laser stripes in images: it thresholds the image, detecting only the brighter stripes in binary form; it is an image memory; and pixel values can be addressed randomly in the image. Thus, the design does not require an A/D converter or a frame buffer and is consequently inexpensive. For improved performance, only the image region next to the previous stripe location is searched, and a 64-KB lookup table stored in RAM is indexed by incident laser angles and stripe addresses to output range data. A 128*256 range image is produced in about 20 s. This is reasonably fast considering that this process requires analyzing 256 images. The speed bottleneck is the low sensitivity of the optic RAM chip, which requires a long exposure time per frame (60 ms), corresponding to half the standard video frame rate. Simple calibration methods using planar patterns of parallel lines are presented.>
Tsutomu Ito, Daniel DeMenthon, Larry Davis 0001
ICRA3
1990 Efficient Algorithms for Obstacle Detection Using Range Data
Phillip A. Veatch, Larry Davis 0001
Comput. Vis. Graph. Image Process.2