Stan Sclaroff

dblp:s/StanSclaroff · DBLP profile ↗
← Back
168ranked-venue papers
12as first author
19since 2021 · last 2026
0000-0002-0711-4313ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 143 · 11 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 106 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-authorComputer networks · 1
YearPublicationVenuePosition
2026 Generalized prompt-driven zero-shot domain adaptive segmentation with feature rectification and semantic modulation
Jinyi Li, Longyu Yang, Donghyun Kim 0006, Kuniaki Saito, Kate Saenko, Stan Sclaroff, Xiaofeng Zhu 0001, Ping Hu 0001
Comput. Vis. Image Underst.6
2024 Video Frame Interpolation With Many-to-Many Splatting and Spatial Selective Refinement
abstract
In this work, we first propose a fully differentiable Many-to-Many (M2M) splatting framework to interpolate frames efficiently. Given a frame pair, we estimate multiple bidirectional flows to directly forward warp the pixels to the desired time step before fusing any overlapping pixels. In doing so, each source pixel renders multiple target pixels and each target pixel can be synthesized from a larger area of visual context, establishing a many-to-many splatting scheme with robustness to undesirable artifacts. For each input frame pair, M2M has a minuscule computational overhead when interpolating an arbitrary number of in-between frames, hence achieving fast multi-frame interpolation. However, directly warping and fusing pixels in the intensity domain is sensitive to the quality of motion estimation and may suffer from less effective representation capacity. To improve interpolation accuracy, we further extend an M2M++ framework by introducing a flexible Spatial Selective Refinement (SSR) component, which allows for trading computational efficiency for interpolation quality and vice versa. Instead of refining the entire interpolated frame, SSR only processes difficult regions selected under the guidance of an estimated error map, thereby avoiding redundant computation. Evaluation on multiple benchmark datasets shows that our method is able to improve the efficiency while maintaining competitive video interpolation quality, and it can be adjusted to use more or less compute as needed.
Ping Hu 0001, Simon Niklaus, Lu Zhang 0053, Stan Sclaroff, Kate Saenko
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 DualCoOp++: Fast and Effective Adaptation to Multi-Label Recognition With Limited Annotations
abstract
Multi-label image recognition in the low-label regime is a task of great challenge and practical significance. Previous works have focused on learning the alignment between textual and visual spaces to compensate for limited image labels, yet may suffer from reduced accuracy due to the scarcity of high-quality multi-label annotations. In this research, we leverage the powerful alignment between textual and visual features pretrained with millions of auxiliary image-text pairs. We introduce an efficient and effective framework calledEvidence-guided Dual Context Optimization(DualCoOp++), which serves as a unified approach for addressing partial-label and zero-shot multi-label recognition. InDualCoOp++we separately encode evidential, positive, and negative contexts for target classes as parametric components of the linguistic input (i.e., prompts). The evidential context aims to discover all the related visual content for the target class, and serves as guidance to aggregate positive and negative contexts from the spatial domain of the image, enabling better distinguishment between similar categories. Additionally, we introduce a Winner-Take-All module that promotes inter-class interaction during training, while avoiding the need for extra parameters and costs. AsDualCoOp++imposes minimal additional learnable overhead on the pretrained vision-language framework, it enables rapid adaptation to multi-label recognition tasks with limited annotations and even unseen classes. Experiments on standard multi-label recognition benchmarks across two challenging low-label settings demonstrate the superior performance of our approach compared to state-of-the-art methods.
Ping Hu 0001, Ximeng Sun, Stan Sclaroff, Kate Saenko
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Practical Disruption of Image Translation Deepfake Networks
abstract
By harnessing the latest advances in deep learning, image-to-image translation architectures have recently achieved impressive capabilities. Unfortunately, the growing representational power of these architectures has prominent unethical uses. Among these, the threats of (1) face manipulation ("DeepFakes") used for misinformation or pornographic use (2) "DeepNude" manipulations of body images to remove clothes from individuals, etc. Several works tackle the task of disrupting such image translation networks by inserting imperceptible adversarial attacks into the input image. Nevertheless, these works have limitations that may result in disruptions that are not practical in the real world. Specifically, most works generate disruptions in a white-box scenario, assuming perfect knowledge about the image translation network. The few remaining works that assume a black-box scenario require a large number of queries to successfully disrupt the adversary's image translation network. In this work we propose Leaking Transferable Perturbations (LTP), an algorithm that significantly reduces the number of queries needed to disrupt an image translation network by dynamically re-purposing previous disruptions into new query efficient disruptions.
Nataniel Ruiz, Sarah Adel Bargal, Cihang Xie, Stan Sclaroff
AAAI4
2022 Many-to-many Splatting for Efficient Video Frame Interpolation
abstract
Motion-based video frame interpolation commonly relies on optical flow to warp pixels from the inputs to the desired interpolation instant. Yet due to the inherent challenges of motion estimation (e.g. occlusions and discontinuities), most state-of-the-art interpolation approaches require subsequent refinement of the warped result to generate satisfying outputs, which drastically decreases the efficiency for multi-frame interpolation. In this work, we propose a fully differentiable Many-to-Many (M2M) splatting framework to interpolate frames efficiently. Specifically, given a frame pair, we estimate multiple bidirectional flows to directly forward warp the pixels to the desired time step, and then fuse any overlapping pixels. In doing so, each source pixel renders multiple target pixels and each target pixel can be synthesized from a larger area of visual context. This establishes a many-to-many splatting scheme with robustness to artifacts like holes. Moreover, for each input frame pair, M2M only performs motion estimation once and has a minuscule computational overhead when interpolating an arbitrary number of in-between frames, hence achieving fast multi-frame interpolation. We conducted extensive experiments to analyze M2M, and found that it significantly improves the efficiency while maintaining high effectiveness.
Ping Hu 0001, Simon Niklaus, Stan Sclaroff, Kate Saenko
CVPR3
2022 Simulated Adversarial Testing of Face Recognition Models
abstract
Most machine learning models are validated and tested on fixed datasets. This can give an incomplete picture of the capabilities and weaknesses of the model. Such weaknesses can be revealed at test time in the real world. The risks involved in such failures can be loss of profits, loss of time or even loss of life in certain critical applications. In order to alleviate this issue, simulators can be controlled in a finegrained manner using interpretable parameters to explore the semantic image manifold. In this work, we propose a framework for learning how to test machine learning algorithms using simulators in an adversarial manner in order to find weaknesses in the model before deploying it in critical scenarios. We apply this method in a face recognition setup. We show that certain weaknesses of models trained on real data can be discovered using simulated samples. Using our proposed method, we can find adversarial synthetic faces that fool contemporary face recognition models. This demonstrates the fact that these models have weaknesses that are not measured by commonly used validation datasets. We hypothesize that this type of adversarial examples are not isolated, but usually lie in connected spaces in the latent space of the simulator. We present a method to find these adversarial regions as opposed to the typical adversarial points found in the adversarial example literature.
Nataniel Ruiz, Adam Kortylewski, Weichao Qiu, Cihang Xie, Sarah Adel Bargal, Alan L. Yuille, Stan Sclaroff
CVPR7
2022 A Unified Framework for Domain Adaptive Pose Estimation
Donghyun Kim 0006, Kaihong Wang, Kate Saenko, Margrit Betke, Stan Sclaroff
ECCV (33)5
2022 A Broad Study of Pre-training for Domain Generalization and Adaptation
Donghyun Kim 0006, Kaihong Wang, Stan Sclaroff, Kate Saenko
ECCV (33)3
2022 Finding Differences Between Transformers and ConvNets Using Counterfactual Simulation Testing
abstract
Modern deep neural networks tend to be evaluated on static test sets. One shortcoming of this is the fact that these deep neural networks cannot be easily evaluated for robustness issues with respect to specific scene variations. For example, it is hard to study the robustness of these networks to variations of object scale, object pose, scene lighting and 3D occlusions. The main reason is that collecting real datasets with fine-grained naturalistic variations of sufficient scale can be extremely time-consuming and expensive. In this work, we present Counterfactual Simulation Testing, a counterfactual framework that allows us to study the robustness of neural networks with respect to some of these naturalistic variations by building realistic synthetic scenes that allow us to ask counterfactual questions to the models, ultimately providing answers to questions such as "Would your classification still be correct if the object were viewed from the top?" or "Would your classification still be correct if the object were partially occluded by another object?". Our method allows for a fair comparison of the robustness of recently released, state-of-the-art Convolutional Neural Networks and Vision Transformers, with respect to these naturalistic variations. We find evidence that ConvNext is more robust to pose and scale variations than Swin, that ConvNext generalizes better to our simulated domain and that Swin handles partial occlusion better than ConvNext. We also find that robustness for all networks improves with network scale and with data scale and variety. We release the Naturalistic Variation Object Dataset (NVD), a large simulated dataset of 272k images of everyday objects with naturalistic variations such as object pose, scale, viewpoint, lighting and occlusions. Project page: https://counterfactualsimulation.github.io
Nataniel Ruiz, Sarah Adel Bargal, Cihang Xie, Kate Saenko, Stan Sclaroff
NeurIPS5
2022 Revisiting Image-Language Networks for Open-Ended Phrase Detection
abstract
Most existing work that grounds natural language phrases in images starts with the assumption that the phrase in question is relevant to the image. In this paper we address a more realistic version of the natural language grounding task where we must both identify whether the phrase is relevant to an image and localize the phrase. This can also be viewed as a generalization of object detection to an open-ended vocabulary, introducing elements of few- and zero-shot detection. We propose an approach for this task that extends Faster R-CNN to relate image regions and phrases. By carefully initializing the classification layers of our network using canonical correlation analysis (CCA), we encourage a solution that is more discerning when reasoning between similar phrases, resulting in over double the performance compared to a naive adaptation on three popular phrase grounding datasets, Flickr30K Entities, ReferIt Game, and Visual Genome, with test-time phrase vocabulary sizes of 5K, 32K, and 159K, respectively.
Bryan A. Plummer, Kevin J. Shih, Svetlana Lazebnik, Stan Sclaroff, Kate Saenko
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Leveraging Geometric Structure for Label-Efficient Semi-Supervised Scene Segmentation
abstract
Label-efficient scene segmentation aims to achieve effective per-pixel classification with reduced labeling effort. Recent approaches for this task focus on leveraging unlabelled images by formulating consistency regularization or pseudo labels for individual pixels. Yet most of these methods ignore the 3D geometric structures naturally conveyed by image scenes, which is free for enhancing training segmentation models with better discrimination of image details. In this work, we present a novel Geometric Structure Refinement (GSR) framework to explicitly exploit the geometric structures of image scenes to enhance the semi-supervised training of segmentation models. In the training phase, we generate initial dense pseudo labels based on fast and coarse annotations, and then utilize the free unsupervised 3D reconstruction of the image scene to calibrate the dense pseudo labels with more reliable details. With the calibrated pseudo groundtruth, we are able to conveniently train any existing image segmentation models without increasing the costs of annotations or modifying the models' architectures. Moreover, we explore different strategies for allocating labeling effort in semi-supervised scene segmentation, and find that a combination of finely-labeled samples and coarsely-labeled samples performs better than the traditional dense-fine only annotations. Extensive experiments on datasets including Cityscapes and KITTI are conducted to evaluate our proposed methods. The results demonstrate that GSR can be easily applied to boost the performance of existing models like PSPNet, DeepLabv3+, etc with reduced annotations. With half of the annotation effort, GSR achieves 99% of the accuracy of its fully supervised state-of-the-art counterparts.
Ping Hu 0001, Stan Sclaroff, Kate Saenko
IEEE Trans. Image Process.2
2021 Siamese Natural Language Tracker: Tracking by Natural Language Descriptions With Siamese Trackers
abstract
We propose a novel Siamese Natural Language Tracker (SNLT), which brings the advancements in visual tracking to the tracking by natural language (NL) descriptions task. The proposed SNLT is applicable to a wide range of Siamese trackers, providing a new class of baselines for the tracking by NL task and promising future improvements from the advancements of Siamese trackers. The carefully designed architecture of the Siamese Natural Language Region Proposal Network (SNL-RPN), together with the Dynamic Aggregation of vision and language modalities, is introduced to perform the tracking by NL task. Empirical results over tracking benchmarks with NL annotations show that the proposed SNLT improves Siamese trackers by 3 to 7 percentage points with a slight tradeoff of speed. The proposed SNLT outperforms all NL trackers to-date and is competitive among state-of-the-art real-time trackers on LaSOT benchmarks while running at 50 frames per second on a single GPU. Code for this work is available at https://github.com/fredfung007/snlt.
Qi Feng 0004, Vitaly Ablavsky, Qinxun Bai, Stan Sclaroff
CVPR4
2021 Leveraging Affect Transfer Learning for Behavior Prediction in an Intelligent Tutoring System
abstract
In this work, we propose a video-based transfer learning approach for predicting problem outcomes of students working with an intelligent tutoring system (ITS). By analyzing a student's face and gestures, our method predicts the outcome of a student answering a problem in an ITS from a video feed. Our work is motivated by the reasoning that the ability to predict such outcomes enables tutoring systems to adjust interventions, such as hints and encouragement, and to ultimately yield improved student learning. We collected a large labeled dataset of student interactions with an intelligent online math tutor consisting of 68 sessions, where 54 individual students solved 2,749 problems. We will release this dataset publicly upon publication of this paper. It will be available at https://www.cs.bu.edu/faculty/betke/research/learning/. Working with this dataset, our transfer-learning challenge was to design a representation in the source domain of pictures obtained “in the wild” for the task of facial expression analysis, and transferring this learned representation to the task of human behavior prediction in the domain of webcam videos of students in a classroom environment. We developed a novel facial affect representation and a user-personalized training scheme that unlocks the potential of this representation. We designed several variants of a recurrent neural network that models the temporal structure of video sequences of students solving math problems. Our final model, named ATL-BP for Affect Transfer Learning for Behavior Prediction, achieves a relative increase in mean F -score of 50 % over the state-of-the-art method on this new dataset.
Nataniel Ruiz, Hao Yu 0014, Danielle Allessio, Mona Jalal, Ajjen Joshi, Tom Murray 0001, John J. Magee, Jacob Whitehill, Vitaly Ablavsky, Ivon Arroyo, Beverly P. Woolf, Stan Sclaroff, Margrit Betke
FG12
2021 CDS: Cross-Domain Self-supervised Pre-training
abstract
We present a two-stage pre-training approach that improves the generalization ability of standard single-domain pre-training. While standard pre-training on a single large dataset (such as ImageNet) can provide a good initial representation for transfer learning tasks, this approach may result in biased representations that impact the success of learning with new multi-domain data (e.g., different artistic styles) via methods like domain adaptation. We propose a novel pre-training approach called Cross-Domain Self-supervision (CDS), which directly employs unlabeled multi-domain data for downstream domain transfer tasks. Our approach uses self-supervision not only within a single domain but also across domains. In-domain instance discrimination is used to learn discriminative features on new data in a domain-adaptive manner, while cross-domain matching is used to learn domain-invariant features. We apply our method as a second pre-training step (after ImageNet pre-training), resulting in a significant target accuracy boost to diverse domain transfer tasks compared to standard one-stage pre-training.
Donghyun Kim 0006, Kuniaki Saito, Tae-Hyun Oh, Bryan A. Plummer, Stan Sclaroff, Kate Saenko
ICCV5
2021 Learning Cross-Modal Contrastive Features for Video Domain Adaptation
abstract
Learning transferable and domain adaptive feature representations from videos is important for video-relevant tasks such as action recognition. Existing video domain adaptation methods mainly rely on adversarial feature alignment, which has been derived from the RGB image space. However, video data is usually associated with multi-modal information, e.g., RGB and optical flow, and thus it remains a challenge to design a better method that considers the cross-modal inputs under the cross-domain adaptation setting. To this end, we propose a unified framework for video domain adaptation, which simultaneously regularizes cross-modal and cross-domain feature representations. Specifically, we treat each modality in a domain as a view and leverage the contrastive learning technique with properly designed sampling strategies. As a result, our objectives regularize feature spaces, which originally lack the connection across modalities or have less alignment across domains. We conduct experiments on domain adaptive action recognition benchmark datasets, i.e., UCF, HMDB, and EPIC-Kitchens, and demonstrate the effectiveness of our components against state-of-the-art algorithms.
Donghyun Kim 0006, Yi-Hsuan Tsai, Bingbing Zhuang, Xiang Yu 0002, Stan Sclaroff, Kate Saenko, Manmohan Krishna Chandraker
ICCV5
2021 Tune it the Right Way: Unsupervised Validation of Domain Adaptation via Soft Neighborhood Density
abstract
Unsupervised domain adaptation (UDA) methods can dramatically improve generalization on unlabeled target domains. However, optimal hyper-parameter selection is critical to achieving high accuracy and avoiding negative transfer. Supervised hyper-parameter validation is not possible without labeled target data, which raises the question: How can we validate unsupervised adaptation techniques in a realistic way? We first empirically analyze existing criteria and demonstrate that they are not very effective for tuning hyper-parameters. Intuitively, a well-trained source classifier should embed target samples of the same class nearby, forming dense neighborhoods in feature space. Based on this assumption, we propose a novel unsupervised validation criterion that measures the density of soft neighborhoods by computing the entropy of the similarity distribution between points. Our criterion is simpler than competing validation methods, yet more effective; it can tune hyper-parameters and the number of training iterations in both image classification and semantic segmentation models. The code used for the paper will be available at https://github.com/VisionLearningGroup/SND.
Kuniaki Saito, Donghyun Kim 0006, Piotr Teterwak, Stan Sclaroff, Trevor Darrell, Kate Saenko
ICCV4
2021 Distillation Multiple Choice Learning for Multimodal Action Recognition
abstract
In this work, we address the problem of learning an ensemble of specialist networks using multimodal data, while considering the realistic and challenging scenario of possible missing modalities at test time. Our goal is to leverage the complementary information of multiple modalities to the benefit of the ensemble and each individual network. We introduce a novel Distillation Multiple Choice Learning framework for multimodal data, where different modality networks learn in a cooperative setting from scratch, strengthening one another. The modality networks learned using our method achieve significantly higher accuracy than if trained separately, due to the guidance of other modalities. We evaluate this approach on three video action recognition benchmark datasets. We obtain state-of-the-art results in comparison to other approaches that work with missing modalities at test time.
Nuno C. Garcia, Sarah Adel Bargal, Vitaly Ablavsky, Pietro Morerio, Vittorio Murino, Stan Sclaroff
WACV6
2021 Excitation Dropout: Encouraging Plasticity in Deep Neural Networks
Andrea Zunino, Sarah Adel Bargal, Pietro Morerio, Jianming Zhang 0001, Stan Sclaroff, Vittorio Murino
Int. J. Comput. Vis.5
2021 Guided Zoom: Zooming into Network Evidence to Refine Fine-Grained Model Decisions
abstract
In state-of-the-art deep single-label classification models, the top- k (k=2,3,4, ...) accuracy is usually significantly higher than the top-1 accuracy. This is more evident in fine-grained datasets, where differences between classes are quite subtle. Exploiting the information provided in the top k predicted classes boosts the final prediction of a model. We propose Guided Zoom, a novel way in which explainability could be used to improve model performance. We do so by making sure the model has "the right reasons" for a prediction. The reason/evidence upon which a deep neural network makes a prediction is defined to be the grounding, in the pixel space, for a specific class conditional probability in the model output. Guided Zoom examines how reasonable the evidence used to make each of the top- k predictions is. Test time evidence is deemed reasonable if it is coherent with evidence used to make similar correct decisions at training time. This leads to better informed predictions. We explore a variety of grounding techniques and study their complementarity for computing evidence. We show that Guided Zoom results in an improvement of a model's classification accuracy and achieves state-of-the-art classification performance on four fine-grained classification datasets. Our code is available at https://github.com/andreazuna89/Guided-Zoom.
Sarah Adel Bargal, Andrea Zunino, Vitali Petsiuk, Jianming Zhang 0001, Kate Saenko, Vittorio Murino, Stan Sclaroff
IEEE Trans. Pattern Anal. Mach. Intell.7
2020 MULE: Multimodal Universal Language Embedding
abstract
Existing vision-language methods typically support two languages at a time at most. In this paper, we present a modular approach which can easily be incorporated into existing vision-language methods in order to support many languages. We accomplish this by learning a single shared Multimodal Universal Language Embedding (MULE) which has been visually-semantically aligned across all languages. Then we learn to relate MULE to visual data as if it were a single language. Our method is not architecture specific, unlike prior work which typically learned separate branches for each language, enabling our approach to easily be adapted to many vision-language methods and tasks. Since MULE learns a single language branch in the multimodal model, we can also scale to support many languages, and languages with fewer annotations can take advantage of the good representation learned from other (more abundant) language data. We demonstrate the effectiveness of our embeddings on the bidirectional image-sentence retrieval task, supporting up to four languages in a single model. In addition, we show that Machine Translation can be used for data augmentation in multilingual learning, which, combined with MULE, improves mean recall by up to 20.2% on a single language compared to prior work, with the most significant gains seen on languages with relatively few annotations. Our code is publicly available1.
Donghyun Kim 0006, Kuniaki Saito, Kate Saenko, Stan Sclaroff, Bryan A. Plummer
AAAI4
2020 Temporally Distributed Networks for Fast Video Semantic Segmentation
abstract
We present TDNet, a temporally distributed network designed for fast and accurate video semantic segmentation. We observe that features extracted from a certain high-level layer of a deep CNN can be approximated by composing features extracted from several shallower sub-networks. Leveraging the inherent temporal continuity in videos, we distribute these sub-networks over sequential frames. Therefore, at each time step, we only need to perform a lightweight computation to extract a sub-features group from a single sub-network. The full features used for segmentation are then recomposed by application of a novel attention propagation module that compensates for geometry deformation between frames. A grouped knowledge distillation loss is also introduced to further improve the representation power at both full and sub-feature levels. Experiments on Cityscapes, CamVid, and NYUD-v2 demonstrate that our method achieves state-of-the-art accuracy with significantly faster speed and lower latency.
Ping Hu 0001, Fabian Caba Heilbron, Oliver Wang, Zhe Lin 0001, Stan Sclaroff, Federico Perazzi
CVPR5
2020 Uncertainty-Aware Learning for Zero-Shot Semantic Segmentation
abstract
Zero-shot semantic segmentation (ZSS) aims to classify pixels of novel classes without training examples available. Recently, most ZSS methods focus on learning the visual-semantic correspondence to transfer knowledge from seen classes to unseen classes at the pixel level. Yet, few works study the adverse effects caused by the noisy and outlying training samples in the seen classes. In this paper, we identify this challenge and address it with a novel framework that learns to discriminate noisy samples based on Bayesian uncertainty estimation. Specifically, we model the network outputs with Gaussian and Laplacian distributions, with the variances accounting for the observation noise as well as the uncertainty of input samples. Learning objectives are then derived with the estimated variances playing as adaptive attenuation for individual samples in training. Consequently, our model learns more attentively from representative samples of seen classes while suffering less from noisy and outlying ones, thus providing better reliability and generalization toward unseen categories. We demonstrate the effectiveness of our framework through comprehensive experiments on multiple challenging benchmarks, and show that our method achieves significant accuracy improvement over previous approaches for large open-set segmentation.
Ping Hu 0001, Stan Sclaroff, Kate Saenko
NeurIPS2
2020 Universal Domain Adaptation through Self Supervision
abstract
Unsupervised domain adaptation methods traditionally assume that all source categories are present in the target domain. In practice, little may be known about the category overlap between the two domains. While some methods address target settings with either partial or open-set categories, they assume that the particular setting is known a priori. We propose a more universally applicable domain adaptation approach that can handle arbitrary category shift, called Domain Adaptative Neighborhood Clustering via Entropy optimization (DANCE). Our approach combines two novel ideas: First, as we cannot fully rely on source categories to learn features discriminative for the target, we propose a novel neighborhood clustering technique to learn the structure of the target domain in a self-supervised way. Second, we use entropy-based feature alignment and rejection to align target features with the source, or reject them as unknown categories based on their entropy. We show through extensive experiments that DANCE outperforms baselines across open-set, open-partial and partial domain adaptation settings.
Kuniaki Saito, Donghyun Kim 0006, Stan Sclaroff, Kate Saenko
NeurIPS3
2020 Real-time Visual Object Tracking with Natural Language Description
abstract
In this work, we argue that conditioning on the natural language (NL) description of a target provides information for longer-term invariance, and thus helps cope with typical tracking challenges. However, deriving a formulation to combine the strengths of appearance-based tracking with the language modality is not straightforward. Therefore, we propose a novel deep tracking-by-detection formulation that can take advantage of NL descriptions. Regions that are related to the given NL description are generated by a proposal network during the detection stage of the tracker. Our LSTM based tracker then predicts the update of the target from regions proposed by the NL based detection stage. Our method runs at over 30 fps on a single GPU. In benchmarks, our method is competitive with state of the art trackers that employ bounding boxes for initialization, while it outperforms all other trackers on targets given unambiguous and precise language annotations. When conditioned on NL descriptions only, our model doubles the performance of the previous best attempt [25].
Qi Feng 0004, Vitaly Ablavsky, Qinxun Bai, Guorong Li, Stan Sclaroff
WACV5
2020 DIPNet: Dynamic Identity Propagation Network for Video Object Segmentation
abstract
Many recent methods for semi-supervised Video Object Segmentation (VOS) have achieved good performance by exploiting the annotated first frame via one-shot fine-tuning or mask propagation. However, heavily relying on the first frame may weaken the robustness for VOS, since video objects can show large variations through time. In this work, we propose a Dynamic Identity Propagation Network (DIPNet) that adaptively propagates and accurately segments the video objects over time. To achieve this, DIPNet factors the VOS task at each time step into a dynamic propagation phase and a spatial segmentation phase. The former utilizes a novel identity representation to adaptively propagate objects’ reference information over time, which enhances the robustness to videos’ temporal variations. The segmentation phase uses the propagated information to tackle the object segmentation as an easier static image problem that can be optimized via light-weight fine-tuning on the first frame, thus reducing the computational cost. As a result, by optimizing these two components to complement each other, we can achieve a robust system for VOS. Evaluations on four benchmark datasets show that DIPNet provides state-of-the-art performance with time efficiency.
Ping Hu 0001, Jun Liu 0036, Gang Wang 0012, Vitaly Ablavsky, Kate Saenko, Stan Sclaroff
WACV6
2020 Multi-way Encoding for Robustness
abstract
Deep models are state-of-the-art for many computer vision tasks including image classification and object detection. However, it has been shown that deep models are vulnerable to adversarial examples. We highlight how one-hot encoding directly contributes to this vulnerability and propose breaking away from this widely-used, but highly-vulnerable mapping. We demonstrate that by leveraging a different output encoding, multi-way encoding, we decorre-late source and target models, making target models more secure. Our approach makes it more difficult for adversaries to find useful gradients for generating adversarial attacks. We present robustness for black-box and white-box attacks on four benchmark datasets: MNIST, CIFAR-10, CIFAR-100, and SVHN. The strength of our approach is also presented in the form of an attack for model watermarking, raising challenges in detecting stolen models.
Donghyun Kim 0006, Sarah Adel Bargal, Jianming Zhang 0001, Stan Sclaroff
WACV4
2019 Multilevel Language and Vision Integration for Text-to-Clip Retrieval
abstract
We address the problem of text-based activity retrieval in video. Given a sentence describing an activity, our task is to retrieve matching clips from an untrimmed video. To capture the inherent structures present in both text and video, we introduce a multilevel model that integrates vision and language features earlier and more tightly than prior work. First, we inject text features early on when generating clip proposals, to help eliminate unlikely clips and thus speed up processing and boost performance. Second, to learn a fine-grained similarity metric for retrieval, we use visual features to modulate the processing of query sentences at the word level in a recurrent neural network. A multi-task loss is also employed by adding query re-generation as an auxiliary task. Our approach significantly outperforms prior work on two challenging benchmarks: Charades-STA and ActivityNet Captions.
Huijuan Xu 0001, Kun He 0003, Bryan A. Plummer, Leonid Sigal, Stan Sclaroff, Kate Saenko
AAAI5
2019 Guided Zoom: Questioning Network Evidence for Fine-grained Classification
Sarah Adel Bargal, Andrea Zunino, Vitali Petsiuk, Jianming Zhang 0001, Kate Saenko, Vittorio Murino, Stan Sclaroff
BMVC7
2019 Deep Metric Learning to Rank
abstract
We propose a novel deep metric learning method by revisiting the learning to rank approach. Our method, named FastAP, optimizes the rank-based Average Precision measure, using an approximation derived from distance quantization. FastAP has a low complexity compared to existing methods, and is tailored for stochastic gradient descent. To fully exploit the benefits of the ranking formulation, we also propose a new minibatch sampling scheme, as well as a simple heuristic to enable large-batch training. On three few-shot image retrieval datasets, FastAP consistently outperforms competing methods, which often involve complex optimization heuristics or costly model ensembles.
Fatih Çakir, Kun He 0003, Xide Xia, Brian Kulis, Stan Sclaroff
CVPR5
2019 Affect-driven Learning Outcomes Prediction in Intelligent Tutoring Systems
abstract
Equipping an Intelligent Tutoring System (ITS) with the ability to interpret affective signals from students could potentially improve the learning experience of students by enabling the tutor to monitor the students' progress and provide timely interventions as well as present appropriate affective reactions via a virtual tutor. Most ITSs equipped with affect modeling capabilities attempt to predict the emotional state of users. However, the focus in this work is instead on trying to directly predict the learning outcomes of students from a stream of video capturing the students faces as they work on a set of math problems. Using facial features extracted from a video stream, we train classifiers to directly predict the success or failure of a student's attempt to answer a question while the student has just begun to work on the problem. In this work, we first introduce a novel dataset of student interactions with MathSpring, a popular ITS. We provide an exploratory analysis of the different problem outcome classes using typical facial action unit activations. We develop baseline models to predict the problem outcome labels of students solving math problems and discuss how early problem outcome labels can be forecasted and utilized to provide possible interventions.
Ajjen Joshi, Danielle Allessio, John J. Magee, Jacob Whitehill, Ivon Arroyo, Beverly P. Woolf, Stan Sclaroff, Margrit Betke
FG7
2019 Language Features Matter: Effective Language Representations for Vision-Language Tasks
abstract
Shouldn't language and vision features be treated equally in vision-language (VL) tasks? Many VL approaches treat the language component as an afterthought, using simple language models that are either built upon fixed word embeddings trained on text-only data or are learned from scratch. We conclude that language features deserve more attention, which has been informed by experiments which compare different word embeddings, language models, and embedding augmentation steps on five common VL tasks: image-sentence retrieval, image captioning, visual question answering, phrase grounding, and text-to-clip retrieval. Our experiments provide some striking results; an average embedding language model outperforms a LSTM on retrieval-style tasks; state-of-the-art representations such as BERT perform relatively poorly on vision-language tasks. From this comprehensive set of experiments we can propose a set of best practices for incorporating the language component of vision-language tasks. To further elevate language features, we also show that knowledge in vision-language problems can be transferred across tasks to gain performance with multi-task training. This multi-task training is applied to a new Graph Oriented Vision-Language Embedding (GrOVLE), which we adapt from Word2Vec using WordNet and an original visual-language graph built from Visual Genome, providing a ready-to-use vision-language embedding: http://ai.bu.edu/grovle.
Andrea Burns, Reuben Tan, Kate Saenko, Stan Sclaroff, Bryan A. Plummer
ICCV4
2019 Semi-Supervised Domain Adaptation via Minimax Entropy
abstract
Contemporary domain adaptation methods are very effective at aligning feature distributions of source and target domains without any target supervision. However, we show that these techniques perform poorly when even a few labeled examples are available in the target domain. To address this semi-supervised domain adaptation (SSDA) setting, we propose a novel Minimax Entropy (MME) approach that adversarially optimizes an adaptive few-shot model. Our base model consists of a feature encoding network, followed by a classification layer that computes the features' similarity to estimated prototypes (representatives of each class). Adaptation is achieved by alternately maximizing the conditional entropy of unlabeled target data with respect to the classifier and minimizing it with respect to the feature encoder. We empirically demonstrate the superiority of our method over many baselines, including conventional feature alignment and few-shot methods, setting a new state of the art for SSDA. Our code is available at http://cs-people.bu.edu/keisaito/research/MME.html.
Kuniaki Saito, Donghyun Kim 0006, Stan Sclaroff, Trevor Darrell, Kate Saenko
ICCV3
2019 Cost-Aware Fine-Grained Recognition for IoTs Based on Sequential Fixations
abstract
We consider the problem of fine-grained classification on an edge camera device that has limited power. The edge device must sparingly interact with the cloud to minimize communication bits to conserve power, and the cloud upon receiving the edge inputs returns a classification label. To deal with fine-grained classification, we adopt the perspective of sequential fixation with a foveated field-of-view to model cloud-edge interactions. We propose a novel deep reinforcement learning-based foveation model, DRIFT, that sequentially generates and recognizes mixed-acuity images. Training of DRIFT requires only image-level category labels and encourages fixations to contain task-relevant information, while maintaining data efficiency. Specifically, we train a foveation actor network with a novel Deep Deterministic Policy Gradient by Conditioned Critic and Coaching(DDPGC3) algorithm. In addition, we propose to shape the reward to provide informative feedback after each fixation to better guide RL training. We demonstrate the effectiveness of DRIFT on this task by evaluating on five fine-grained classification benchmark datasets, and show that the proposed approach achieves state-of-the-art performance with over 3X reduction in transmitted pixels.
Hanxiao Wang 0001, Venkatesh Saligrama, Stan Sclaroff, Vitaly Ablavsky
ICCV3
2019 Generalized Majorization-Minimization
abstract
Non-convex optimization is ubiquitous in machine learning. Majorization-Minimization (MM) is a powerful iterative procedure for optimizing non-convex functions that works by optimizing a sequence of bounds on the function. In MM, the bound at each iteration is required to touch the objective function at the optimizer of the previous bound. We show that this touching constraint is unnecessary and overly restrictive. We generalize MM by relaxing this constraint, and propose a new optimization framework, named Generalized Majorization-Minimization (G-MM), that is more flexible. For instance, G-MM can incorporate application-specific biases into the optimization procedure without changing the objective function. We derive G-MM algorithms for several latent variable models and show empirically that they consistently outperform their MM counterparts in optimizing non-convex objectives. In particular, G-MM algorithms appear to be less sensitive to initialization.
Sobhan Naderi Parizi, Kun He 0003, Reza Aghajani, Stan Sclaroff, Pedro F. Felzenszwalb
ICML4
2019 Hashing with Mutual Information
abstract
Binary vector embeddings enable fast nearest neighbor retrieval in large databases of high-dimensional objects, and play an important role in many practical applications, such as image and video retrieval. We study the problem of learning binary vector embeddings under a supervised setting, also known as hashing. We propose a novel supervised hashing method based on optimizing an information-theoretic quantity, mutual information. We show that optimizing mutual information can reduce ambiguity in the induced neighborhood structure in the learned Hamming space, which is essential in obtaining high retrieval performance. To this end, we optimize mutual information in deep neural networks with minibatch stochastic gradient descent, with a formulation that maximally and efficiently utilizes available supervision. Experiments on four image retrieval benchmarks, including ImageNet, confirm the effectiveness of our method in learning high-quality binary embeddings for nearest neighbor retrieval.
Fatih Çakir, Kun He 0003, Sarah Adel Bargal, Stan Sclaroff
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Hashing as Tie-Aware Learning to Rank
abstract
Hashing, or learning binary embeddings of data, is frequently used in nearest neighbor retrieval. In this paper, we develop learning to rank formulations for hashing, aimed at directly optimizing ranking-based evaluation metrics such as Average Precision (AP) and Normalized Discounted Cumulative Gain (NDCG). We first observe that the integer-valued Hamming distance often leads to tied rankings, and propose to use tie-aware versions of AP and NDCG to evaluate hashing for retrieval. Then, to optimize tie-aware ranking metrics, we derive their continuous relaxations, and perform gradient-based optimization with deep neural networks. Our results establish the new state-of-the-art for image retrieval by Hamming ranking in common benchmarks.
Kun He 0003, Fatih Çakir, Sarah Adel Bargal, Stan Sclaroff
CVPR4
2018 Local Descriptors Optimized for Average Precision
abstract
Extraction of local feature descriptors is a vital stage in the solution pipelines for numerous computer vision tasks. Learning-based approaches improve performance in certain tasks, but still cannot replace handcrafted features in general. In this paper, we improve the learning of local feature descriptors by optimizing the performance of descriptor matching, which is a common stage that follows descriptor extraction in local feature based pipelines, and can be formulated as nearest neighbor retrieval. Specifically, we directly optimize a ranking-based retrieval performance metric, Average Precision, using deep neural networks. This general-purpose solution can also be viewed as a listwise learning to rank approach, which is advantageous compared to recent local ranking approaches. On standard benchmarks, descriptors learned with our formulation achieve state-of-the-art results in patch verification, patch retrieval, and image matching.
Kun He 0003, Stan Sclaroff
CVPR3
2018 Excitation Backprop for RNNs
Sarah Adel Bargal, Andrea Zunino, Donghyun Kim 0006, Jianming Zhang 0001, Vittorio Murino, Stan Sclaroff
CVPR6
2018 Hashing with Binary Matrix Pursuit
Fatih Çakir, Kun He 0003, Stan Sclaroff
ECCV (5)3
2018 Context-Sensitive Prediction of Facial Expressivity Using Multimodal Hierarchical Bayesian Neural Networks
abstract
Objective automated affect analysis systems can be applied to quantify the progression of symptoms in neurodegenerative diseases such as Parkinson's Disease (PD). PD hampers the ability of patients to emote by decreasing the mobility of their facial musculature, a phenomenon known as ``facial masking.'' In this work, we focus on building a system that can predict an accurate score of active facial expressivity in people suffering from Parkinson's disease using features extracted from both video and audio. An ideal automated system should be able to mimic the ability of human experts to take into account contextual information while making these predictions. For example, patients exhibit different emotions with varying intensities when speaking about positive and negative experiences. We utilize a hierarchical Bayesian neural network framework to enable the learning of model parameters that subtly adapt to pre-defined notions of context, such as the gender of the patient or the valence of the expressed sentiment. We evaluate our formulation on a dataset of 772 20-second video clips of Parkinson's disease patients and demonstrate that training a context-specific hierarchical Bayesian framework yields an improvement in model performance in both multiclass classification and regression settings compared to baseline models trained on all data pooled together.
Ajjen Joshi, Soumya Ghosh, Sarah Gunnery, Linda Tickle-Degnen, Stan Sclaroff, Margrit Betke
FG5
2018 Predicting Foreground Object Ambiguity and Efficiently Crowdsourcing the Segmentation(s)
Danna Gurari, Kun He 0003, Jianming Zhang 0001, Mehrnoosh Sameki, Suyog Dutt Jain, Stan Sclaroff, Margrit Betke, Kristen Grauman
Int. J. Comput. Vis.7
2018 Space-Time Tree Ensemble for Action Recognition and Localization
Shugao Ma, Jianming Zhang 0001, Stan Sclaroff, Nazli Ikizler-Cinbis, Leonid Sigal
Int. J. Comput. Vis.3
2018 Top-Down Neural Attention by Excitation Backprop
Jianming Zhang 0001, Sarah Adel Bargal, Zhe Lin 0001, Jonathan Brandt, Xiaohui Shen, Stan Sclaroff
Int. J. Comput. Vis.6
2017 Personalizing Gesture Recognition Using Hierarchical Bayesian Neural Networks
abstract
Building robust classifiers trained on data susceptible to group or subject-specific variations is a challenging pattern recognition problem. We develop hierarchical Bayesian neural networks to capture subject-specific variations and share statistical strength across subjects. Leveraging recent work on learning Bayesian neural networks, we build fast, scalable algorithms for inferring the posterior distribution over all network weights in the hierarchy. We also develop methods for adapting our model to new subjects when a small number of subject-specific personalization data is available. Finally, we investigate active learning algorithms for interactively labeling personalization data in resource-constrained scenarios. Focusing on the problem of gesture recognition where inter-subject variations are commonplace, we demonstrate the effectiveness of our proposed techniques. We test our framework on three widely used gesture recognition datasets, achieving personalization performance competitive with the state-of-the-art.
Ajjen Joshi, Soumya Ghosh, Margrit Betke, Stan Sclaroff, Hanspeter Pfister
CVPR4
2017 MIHash: Online Hashing with Mutual Information
abstract
Learning-based hashing methods are widely used for nearest neighbor retrieval, and recently, online hashing methods have demonstrated good performance-complexity trade-offs by learning hash functions from streaming data. In this paper, we first address a key challenge for online hashing: the binary codes for indexed data must be recomputed to keep pace with updates to the hash functions. We propose an efficient quality measure for hash functions, based on an information-theoretic quantity, mutual information, and use it successfully as a criterion to eliminate unnecessary hash table updates. Next, we also show how to optimize the mutual information objective using stochastic gradient descent. We thus develop a novel hashing method, MIHash, that can be used in both online and batch settings. Experiments on image retrieval benchmarks (including a 2.5M image dataset) confirm the effectiveness of our formulation, both in reducing hash table recomputations and in learning high-quality hash functions.
Fatih Çakir, Kun He 0003, Sarah Adel Bargal, Stan Sclaroff
ICCV4
2017 Online supervised hashing
Fatih Çakir, Sarah Adel Bargal, Stan Sclaroff
Comput. Vis. Image Underst.3
2017 Salient Object Subitizing
Jianming Zhang 0001, Shugao Ma, Mehrnoosh Sameki, Stan Sclaroff, Margrit Betke, Zhe Lin 0001, Xiaohui Shen, Brian L. Price, Radomír Mech
Int. J. Comput. Vis.4
2017 Comparing random forest approaches to segmenting and classifying gestures
Ajjen Joshi, Camille Monnier, Margrit Betke, Stan Sclaroff
Image Vis. Comput.4
2017 Do less and achieve more: Training CNNs for action recognition utilizing action images from the Web
Shugao Ma, Sarah Adel Bargal, Jianming Zhang 0001, Leonid Sigal, Stan Sclaroff
Pattern Recognit.5
2016 Learning Activity Progression in LSTMs for Activity Detection and Early Detection
abstract
In this work we improve training of temporal deep models to better learn activity progression for activity detection and early detection tasks. Conventionally, when training a Recurrent Neural Network, specifically a Long Short Term Memory (LSTM) model, the training loss only considers classification error. However, we argue that the detection score of the correct activity category, or the detection score margin between the correct and incorrect categories, should be monotonically non-decreasing as the model observes more of the activity. We design novel ranking losses that directly penalize the model on violation of such monotonicities, which are used together with classification loss in training of LSTM models. Evaluation on ActivityNet shows significant benefits of the proposed ranking losses in both activity detection and early detection tasks.
Shugao Ma, Leonid Sigal, Stan Sclaroff
CVPR3
2016 Unconstrained Salient Object Detection via Proposal Subset Optimization
abstract
We aim at detecting salient objects in unconstrained images. In unconstrained images, the number of salient objects (if any) varies from image to image, and is not given. We present a salient object detection system that directly outputs a compact set of detection windows, if any, for an input image. Our system leverages a Convolutional-Neural-Network model to generate location proposals of salient objects. Location proposals tend to be highly overlapping and noisy. Based on the Maximum a Posteriori principle, we propose a novel subset optimization framework to generate a compact set of detection windows out of noisy proposals. In experiments, we show that our subset optimization formulation greatly enhances the performance of our system, and our system attains 16-34% relative improvement in Average Precision compared with the state-of-the-art on three challenging salient object datasets.
Jianming Zhang 0001, Stan Sclaroff, Zhe Lin 0001, Xiaohui Shen, Brian L. Price, Radomír Mech
CVPR2
2016 Top-Down Neural Attention by Excitation Backprop
Jianming Zhang 0001, Zhe Lin 0001, Jonathan Brandt, Xiaohui Shen, Stan Sclaroff
ECCV (4)5
2016 Differential Geometric Regularization for Supervised Learning of Classifiers
abstract
We study the problem of supervised learning for both binary and multiclass classification from a unified geometric perspective. In particular, we propose a geometric regularization technique to find the submanifold corresponding to an estimator of the class probability P(y|\vec x). The regularization term measures the volume of this submanifold, based on the intuition that overfitting produces rapid local oscillations and hence large volume of the estimator. This technique can be applied to regularize any classification function that satisfies two requirements: firstly, an estimator of the class probability can be obtained; secondly, first and second derivatives of the class probability estimator can be calculated. In experiments, we apply our regularization technique to standard loss functions for classification, our RBF-based implementation compares favorably to widely used regularization methods for both binary and multiclass classification.
Qinxun Bai, Steven Rosenberg, Zheng Wu 0003, Stan Sclaroff
ICML4
2016 Discovering useful parts for pose estimation in sparsely annotated datasets
abstract
Our work introduces a novel way to increase pose estimation accuracy by discovering parts from unannotated regions of training images. Discovered parts are used to generate more accurate appearance likelihoods for traditional part-based models like Pictorial Structures and its derivatives. Our experiments on images of a hawkmoth in flight show that our proposed approach significantly improves over existing work for this application, while also being more generally applicable. Our proposed approach localizes landmarks at least twice as accurately as a baseline based on a Mixture of Pictorial Structures (MPS) model. Our unique High-Resolution Moth Flight (HRMF) dataset is made publicly available with annotations.
Mikhail Breslav, Tyson L. Hedrick, Stan Sclaroff, Margrit Betke
WACV3
2016 Poselet-Based Contextual Rescoring for Human Pose Estimation via Pictorial Structures
Antonio Hernández-Vela, Stan Sclaroff, Sergio Escalera
Int. J. Comput. Vis.2
2016 Exploiting Surroundedness for Saliency Detection: A Boolean Map Approach
abstract
We demonstrate the usefulness of surroundedness for eye fixation prediction by proposing a Boolean Map based Saliency model (BMS). In our formulation, an image is characterized by a set of binary images, which are generated by randomly thresholding the image's feature maps in a whitened feature space. Based on a Gestalt principle of figure-ground segregation, BMS computes a saliency map by discovering surrounded regions via topological analysis of Boolean maps. Furthermore, we draw a connection between BMS and the Minimum Barrier Distance to provide insight into why and how BMS can properly captures the surroundedness cue via Boolean maps. The strength of BMS is verified by its simplicity, efficiency and superior performance compared with 10 state-of-the-art methods on seven eye tracking benchmark datasets.
Jianming Zhang 0001, Stan Sclaroff
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Space-time tree ensemble for action recognition
abstract
Human actions are, inherently, structured patterns of body movements. We explore ensembles of hierarchical spatio-temporal trees, discovered directly from training data, to model these structures for action recognition. The hierarchical spatio-temporal trees provide a robust mid-level representation for actions. However, discovery of frequent and discriminative tree structures is challenging due to the exponential search space, particularly if one allows partial matching. We address this by first building a concise action vocabulary via discriminative clustering. Using the action vocabulary we then utilize tree mining with subsequent tree clustering and ranking to select a compact set of highly discriminative tree patterns. We show that these tree patterns, alone, or in combination with shorter patterns (action words and pairwise patterns) achieve state-of-the-art performance on two challenging datasets: UCF Sports and HighFive. Moreover, trees learned on HighFive are used in recognizing two action classes in a different dataset, Hollywood3D, demonstrating the potential for cross-dataset generality of the trees our approach discovers.
Shugao Ma, Leonid Sigal, Stan Sclaroff
CVPR3
2015 Salient Object Subitizing
abstract
People can immediately and precisely identify that an image contains 1, 2, 3 or 4 items by a simple glance. The phenomenon, known as Subitizing, inspires us to pursue the task of Salient Object Subitizing (SOS), i.e. predicting the existence and the number of salient objects in a scene using holistic cues. To study this problem, we propose a new image dataset annotated using an online crowdsourcing marketplace. We show that a proposed subitizing technique using an end-to-end Convolutional Neural Network (CNN) model achieves significantly better than chance performance in matching human labels on our dataset. It attains 94% accuracy in detecting the existence of salient objects, and 42–82% accuracy (chance is 20%) in predicting the number of salient objects (1, 2, 3, and 4+), without resorting to any object localization process. Finally, we demonstrate the usefulness of the proposed subitizing technique in two computer vision applications: salient object detection and object proposal.
Jianming Zhang 0001, Shugao Ma, Mehrnoosh Sameki, Stan Sclaroff, Margrit Betke, Zhe Lin 0001, Xiaohui Shen, Brian L. Price, Radomír Mech
CVPR4
2015 Adaptive Hashing for Fast Similarity Search
abstract
With the staggering growth in image and video datasets, algorithms that provide fast similarity search and compact storage are crucial. Hashing methods that map the data into Hamming space have shown promise, however, many of these methods employ a batch-learning strategy in which the computational cost and memory requirements may become intractable and infeasible with larger and larger datasets. To overcome these challenges, we propose an online learning algorithm based on stochastic gradient descent in which the hash functions are updated iteratively with streaming data. In experiments with three image retrieval benchmarks, our online algorithm attains retrieval accuracy that is comparable to competing state-of-the-art batch-learning solutions, while our formulation is orders of magnitude faster and being online it is adaptable to the variations of the data. Moreover, our formulation yields improved retrieval performance over a recently reported online hashing technique, Online Kernel Hashing.
Fatih Çakir, Stan Sclaroff
ICCV2
2015 Minimum Barrier Salient Object Detection at 80 FPS
abstract
We propose a highly efficient, yet powerful, salient object detection method based on the Minimum Barrier Distance (MBD) Transform. The MBD transform is robust to pixel-value fluctuation, and thus can be effectively applied on raw pixels without region abstraction. We present an approximate MBD transform algorithm with 100X speedup over the exact algorithm. An error bound analysis is also provided. Powered by this fast MBD transform algorithm, the proposed salient object detection method runs at 80 FPS, and significantly outperforms previous methods with similar speed on four large benchmark datasets, and achieves comparable or better performance than state-of-the-art methods. Furthermore, a technique based on color whitening is proposed to extend our method to leverage the appearance-based backgroundness cue. This extended version further improves the performance, while still being one order of magnitude faster than all the other leading methods.
Jianming Zhang 0001, Stan Sclaroff, Zhe Lin 0001, Xiaohui Shen, Brian L. Price, Radomír Mech
ICCV2
2015 Online supervised hashing
abstract
Fast similarity search is becoming more and more critical given the ever growing sizes of datasets. Hashing approaches provide both fast search mechanisms and compact indexing structures to address this critical need. In image retrieval problems where labeled training data is available, supervised hashing methods prevail over un-supervised methods. However, most supervised hashing methods are batch-learners; this hinders their ability to adapt to changes as a dataset grows and diversifies. In this work, we propose an online supervised hashing technique that is based on Error Correcting Output Codes. Given an incoming stream of training data with corresponding labels, our method learns and adapts its hashing functions in a discriminative manner. Our method makes no assumption about the number of possible class labels, and accommodates new classes as they are presented in the incoming data stream. In experiments with three image retrieval benchmarks, the proposed method yields state-of-the-art retrieval performance as measured in Mean Average Precision, while also being orders-of-magnitude faster than competing batch methods for supervised hashing.
Fatih Çakir, Stan Sclaroff
ICIP2
2015 Hankelet-based dynamical systems modeling for 3D action recognition
abstract
This paper proposes to model an action as the output of a sequence of atomic Linear Time Invariant (LTI) systems. The sequence of LTI systems generating the action is modeled as a Markov chain , where a Hidden Markov Model (HMM) is used to model the transition from one atomic LTI system to another. In turn, the LTI systems are represented in terms of their Hankel matrices. For classification purposes, the parameters of a set of HMMs (one for each action class) are learned via a discriminative approach. This work proposes a novel method to learn the atomic LTI systems from training data , and analyzes in detail the action representation in terms of a sequence of Hankel matrices. Extensive evaluation of the proposed approach on two publicly available datasets demonstrates that the proposed method attains state-of-the-art accuracy in action classification from the 3D locations of body joints (skeleton).
Liliana Lo Presti, Marco La Cascia, Stan Sclaroff, Octavia I. Camps
Image Vis. Comput.3
2015 Scale and Rotation Invariant Matching Using Linearly Augmented Trees
abstract
We propose a novel linearly augmented tree method for efficient scale and rotation invariant object matching. The proposed method enforces pairwise matching consistency defined on trees, and high-order constraints on all the sites of a template. The pairwise constraints admit arbitrary metrics while the high-order constraints use L1 norms and therefore can be linearized. Such a linearly augmented tree formulation introduces hyperedges and loops into the basic tree structure. But, different from a general loopy graph, its special structure allows us to relax and decompose the optimization into a sequence of tree matching problems that are efficiently solvable by dynamic programming. The proposed method also works on continuous scale and rotation parameters; we can match with a scale up to any large value with the same efficiency. Our experiments on ground truth data and a variety of real images and videos show that the proposed method is efficient, accurate and reliable.
Hao Jiang 0007, Tai-Peng Tian, Stan Sclaroff
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Gesture Modeling by Hanklet-Based Hidden Markov Model
Liliana Lo Presti, Marco La Cascia, Stan Sclaroff, Octavia I. Camps
ACCV (3)3
2014 Contextual Rescoring for Human Pose Estimation
Antonio Hernández-Vela, Sergio Escalera, Stan Sclaroff
BMVC3
2014 Adaptive Structured Pooling for Action Recognition
Svebor Karaman, Lorenzo Seidenari, Shugao Ma, Alberto Del Bimbo, Stan Sclaroff
BMVC5
2014 Parameterizing Object Detectors in the Continuous Pose Space
Kun He 0003, Leonid Sigal, Stan Sclaroff
ECCV (4)3
2014 MEEM: Robust Tracking via Multiple Experts Using Entropy Minimization
Jianming Zhang 0001, Shugao Ma, Stan Sclaroff
ECCV (6)3
2014 A Bayesian Framework for Online Classifier Ensemble
abstract
We propose a Bayesian framework for recursively estimating the classifier weights in online learning of a classifier ensemble. In contrast with past methods, such as stochastic gradient descent or online boosting, our framework estimates the weights in terms of evolving posterior distributions. For a specified class of loss functions, we show that it is possible to formulate a suitably defined likelihood function and hence use the posterior distribution as an approximation to the global empirical loss minimizer. If the stream of training data is sampled from a stationary process, we can also show that our framework admits a superior rate of convergence to the expected loss minimizer than is possible with standard stochastic gradient descent. In experiments with real-world datasets, our formulation often performs better than online boosting algorithms.
Qinxun Bai, Henry Lam, Stan Sclaroff
ICML3
2014 Supervised hashing with error correcting codes
abstract
One widely-used solution to expedite similarity search of multimedia data is to construct hash functions to map the data into a Hamming space where linear search is known to be fast and often sublinear solutions perform well. In this paper, we propose a Boosting based formulation for supervised learning of the hash functions that is based on Error Correcting Codes. This approach allows us to apply established theoretical results for Boosting in our analysis of our hashing solution. Specifically, we show that the training accuracy in Boosting can be considered as a lower bound on the (empirical) Mean Average Precision (mAP) score. In experiments with three image retrieval benchmarks, the proposed formulation yields significant improvement in mAP over state-of-the-art supervised hashing methods, while using fewer bits in the hash codes.
Fatih Çakir, Stan Sclaroff
ACM Multimedia2
2014 3D pose estimation of bats in the wild
abstract
Vision-based methods have gained popularity as a tool for helping to analyze the behavior of bats. Though, for bats in the wild, there are still no tools capable of estimating and subsequently analyzing articulated 3D bat pose. We propose a model-based multi-view articulated 3D bat pose estimation framework for this novel problem. Key challenges include the large search space associated with articulated 3D pose, the ambiguities that arise from 2D projections of 3D bodies, and the low resolution image data we have available. Our method uses multi-view camera geometry and temporal constraints to reduce the state space of possible articulated 3D bat poses and finds an optimal set using a Markov Random Field based model. Our experiments use real video data of flying bats and gold-standard annotations by a bat biologist. Our results show, for the first time in the literature, articulated 3D pose estimates being generated automatically for video sequences of bats flying in the wild. The average differences in body orientation and wing joint angles, between estimates produced by our method and those based on gold-standard annotations, ranged from 16° - 21° (i.e., ≈ 17% - 23%) for orientation and 14° - 26° (i.e., ≈ 7%- 14%) for wing joint angles.
Mikhail Breslav, Nathan W. Fuller, Stan Sclaroff, Margrit Betke
WACV3
2014 Best of Automatic Face and Gesture Recognition 2013
Stan Sclaroff, Lijun Yin 0001
Image Vis. Comput.2
2013 Decoding Children's Social Behavior
abstract
We introduce a new problem domain for activity recognition: the analysis of children's social and communicative behaviors based on video and audio data. We specifically target interactions between children aged 1-2 years and an adult. Such interactions arise naturally in the diagnosis and treatment of developmental disorders such as autism. We introduce a new publicly-available dataset containing over 160 sessions of a 3-5 minute child-adult interaction. In each session, the adult examiner followed a semi-structured play interaction protocol which was designed to elicit a broad range of social behaviors. We identify the key technical challenges in analyzing these behaviors, and describe methods for decoding the interactions. We present experimental results that demonstrate the potential of the dataset to drive interesting research questions, and show preliminary results for multi-modal activity recognition.
James M. Rehg, Gregory D. Abowd, Agata Rozga, Mario Romero, Mark A. Clements, Stan Sclaroff, Irfan A. Essa, Opal Y. Ousley, Yin Li 0003, Chanho Kim, Hrishikesh Rao 0001, Jonathan C. Kim, Liliana Lo Presti, Jianming Zhang 0001, Denis Lantsman, Jonathan Bidwell, Zhefan Ye
CVPR6
2013 Randomized Ensemble Tracking
abstract
We propose a randomized ensemble algorithm to model the time-varying appearance of an object for visual tracking. In contrast with previous online methods for updating classifier ensembles in tracking-by-detection, the weight vector that combines weak classifiers is treated as a random variable and the posterior distribution for the weight vector is estimated in a Bayesian manner. In essence, the weight vector is treated as a distribution that reflects the confidence among the weak classifiers used to construct and adapt the classifier ensemble. The resulting formulation models the time-varying discriminative ability among weak classifiers so that the ensembled strong classifier can adapt to the varying appearance, backgrounds, and occlusions. The formulation is tested in a tracking-by-detection implementation. Experiments on 28 challenging benchmark videos demonstrate that the proposed method can achieve results comparable to and often better than those of state-of-the-art approaches.
Qinxun Bai, Zheng Wu 0003, Stan Sclaroff, Margrit Betke, Camille Monnier
ICCV3
2013 Action Recognition and Localization by Hierarchical Space-Time Segments
abstract
We propose Hierarchical Space-Time Segments as a new representation for action recognition and localization. This representation has a two-level hierarchy. The first level comprises the root space-time segments that may contain a human body. The second level comprises multi-grained space-time segments that contain parts of the root. We present an unsupervised method to generate this representation from video, which extracts both static and non-static relevant space-time segments, and also preserves their hierarchical and temporal relationships. Using simple linear SVM on the resultant bag of hierarchical space-time segments representation, we attain better than, or comparable to, state-of-the-art action recognition performance on two challenging benchmark datasets and at the same time produce good action localization results.
Shugao Ma, Jianming Zhang 0001, Nazli Ikizler-Cinbis, Stan Sclaroff
ICCV4
2013 Saliency Detection: A Boolean Map Approach
abstract
A novel Boolean Map based Saliency (BMS) model is proposed. An image is characterized by a set of binary images, which are generated by randomly thresholding the image's color channels. Based on a Gestalt principle of figure-ground segregation, BMS computes saliency maps by analyzing the topological structure of Boolean maps. BMS is simple to implement and efficient to run. Despite its simplicity, BMS consistently achieves state-of-the-art performance compared with ten leading methods on five eye tracking datasets. Furthermore, BMS is also shown to be advantageous in salient object detection.
Jianming Zhang 0001, Stan Sclaroff
ICCV2
2013 ChaLearn multi-modal gesture recognition 2013: grand challenge and workshop summary
abstract
We organized a Grand Challenge and Workshop on Multi-Modal Gesture Recognition.
Sergio Escalera, Jordi Gonzàlez 0001, Xavier Baró, Miguel Reyes, Isabelle Guyon, Vassilis Athitsos, Hugo Jair Escalante, Leonid Sigal, Antonis A. Argyros, Cristian Sminchisescu, Richard Bowden, Stan Sclaroff
ICMI12
2012 Online Multi-person Tracking by Tracker Hierarchy
abstract
Tracking-by-detection is a widely used paradigm for multi-person tracking but is affected by variations in crowd density, obstacles in the scene, varying illumination, human pose variation, scale changes, etc. We propose an improved tracking-by-detection framework for multi-person tracking where the appearance model is formulated as a template ensemble updated online given detections provided by a pedestrian detector. We employ a hierarchy of trackers to select the most effective tracking strategy and an algorithm to adapt the conditions for trackers' initialization and termination. Our formulation is online and does not require calibration information. In experiments with four pedestrian tracking benchmark datasets, our formulation attains accuracy that is comparable to, or better than, the state-of-the-art pedestrian trackers that must exploit calibration information and operate offline.
Jianming Zhang 0001, Liliana Lo Presti, Stan Sclaroff
AVSS3
2012 People in Motion: Pose, Action and Communication
Stan Sclaroff
BMVC1
2012 Scale resilient, rotation invariant articulated object matching
abstract
A novel method is proposed for matching articulated objects in cluttered videos. The method needs only a single exemplar image of the target object. Instead of using a small set of large parts to represent an articulated object, the proposed model uses hundreds of small units to represent walks along paths of pixels between key points on an articulated object. Matching directly on dense pixels is key to achieving reliable matching when motion blur occurs. The proposed method fits the model to local image properties, conforms to structure constraints, and remembers the steps taken along a pixel path. The model formulation handles variations in object scaling, rotation and articulation. Recovery of the optimal pixel walks is posed as a special shortest path problem, which can be solved efficiently via dynamic programming. Further speedup is achieved via factorization of the path costs. An efficient method is proposed to find multiple walks and simultaneously match multiple key points. Experiments show that the proposed method is efficient and reliable and can be used to match articulated objects in fast motion videos with strong clutter and blurry imagery.
Hao Jiang 0007, Tai-Peng Tian, Kun He 0003, Stan Sclaroff
CVPR4
2012 Coupling detection and data association for multiple object tracking
abstract
We present a novel framework for multiple object tracking in which the problems of object detection and data association are expressed by a single objective function. The framework follows the Lagrange dual decomposition strategy, taking advantage of the often complementary nature of the two subproblems. Our coupling formulation avoids the problem of error propagation from which traditional “detection-tracking approaches” to multiple object tracking suffer. We also eschew common heuristics such as “nonmaximum suppression” of hypotheses by modeling the joint image likelihood as opposed to applying independent likelihood assumptions. Our coupling algorithm is guaranteed to converge and can handle partial or even complete occlusions. Furthermore, our method does not have any severe scalability issues but can process hundreds of frames at the same time. Our experiments involve challenging, notably distinct datasets and demonstrate that our method can achieve results comparable to those of state-of-art approaches, even without a heavily trained object detector.
Zheng Wu 0003, Ashwin Thangali, Stan Sclaroff, Margrit Betke
CVPR3
2012 Contextual Object Detection Using Set-Based Classification
Ramazan Gokberk Cinbis, Stan Sclaroff
ECCV (6)2
2012 Detecting Reduplication in Videos of American Sign Language
Zoya Gavrilov, Stan Sclaroff, Carol Neidle, Sven J. Dickinson
LREC2
2012 Divide, Conquer and Coordinate: Globally Coordinated Switching Linear Dynamical System
abstract
The goal of this work is to learn a parsimonious and informative representation for high-dimensional time series. Conceptually, this comprises two distinct yet tightly coupled tasks: learning a low-dimensional manifold and modeling the dynamical process. These two tasks have a complementary relationship as the temporal constraints provide valuable neighborhood information for dimensionality reduction and, conversely, the low-dimensional space allows dynamics to be learned efficiently. Solving these two tasks simultaneously allows important information to be exchanged mutually. If nonlinear models are required to capture the rich complexity of time series, then the learning problem becomes harder as the nonlinearities in both tasks are coupled. A divide, conquer, and coordinate method is proposed. The solution approximates the nonlinear manifold and dynamics using simple piecewise linear models. The interactions and coordinations among the linear models are captured in a graphical model. The model structure setup and parameter learning are done using a variational Bayesian approach, which enables automatic Bayesian model structure selection, hence solving the problem of overfitting. By exploiting the model structure, efficient inference and learning algorithms are obtained without oversimplifying the model of the underlying dynamical process. Evaluation of the proposed framework with competing approaches is conducted in three sets of experiments: dimensionality reduction and reconstruction using synthetic time series, video synthesis using a dynamic texture database, and human motion synthesis, classification, and tracking on a benchmark data set. In all experiments, the proposed approach provides superior performance.
Rui Li 0053, Tai-Peng Tian, Stan Sclaroff
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 Web-Based Classifiers for Human Action Recognition
abstract
Action recognition in uncontrolled videos is a challenging task, where it is relatively hard to find the large amount of required training videos to model all the variations of the domain. This paper addresses this challenge and proposes a generic method for action recognition. The idea is to use images collected from the Web to learn representations of actions and leverage this knowledge to automatically annotate actions in videos. For this purpose, we first use an incremental image retrieval procedure to collect and clean up the necessary training set for building the human pose classifiers. Our approach is unsupervised in the sense that it requires no human intervention other than the text querying to an internet search engine. Its benefits are two-fold: 1) we can improve retrieval of action images, and 2) we can collect a large generic database of action poses, which can then be used in tagging videos. We present experimental evidence that using action images collected from the Web, annotating actions in the videos is possible. Additionally, we explore how the Web-based pose classifiers can be utilized in conjunction with limited labelled videos. We propose to use “ordered pose pairs” (OPP) for encoding the temporal ordering of poses in our action model, and show that considering the temporal ordering of pose pairs can increase the action recognition accuracy. We also show that by selecting the keyposes with the help of Web-based classifiers, the classification time can be reduced. Our experiments demonstrate that, with or without available video data, the pose models learned from the Web can improve the performance of the action recognition systems.
Nazli Ikizler-Cinbis, Stan Sclaroff
IEEE Trans. Multim.2
2012 Path Modeling and Retrieval in Distributed Video Surveillance Databases
abstract
We propose a framework for querying a distributed database of video surveillance data in order to retrieve a set of likely paths of a person moving in the area under surveillance. In our framework, each camera of the surveillance system locally processes the data and stores video sequences in a storage unit and the metadata for each detected person in the distributed database. A pedestrian's path is formulated as a dynamic Bayesian network (DBN) to model the dependencies between subsequent observations of the person as he makes his way through the camera network. We propose a tool by which the analyst can pose queries about where a certain person appeared while moving in the site during a specified temporal window. The DBN is used in an algorithm that finds potentially relevant metadata records from the distributed databases and then assembles these into probable paths that the person took in the camera network. Finally, the system presents the analyst with the retrieved set of likely paths in ranked order. The computational complexity for our method is quadratic in the number of camera nodes and linear in the number of moving persons. Experiments were carried out on simulated data to test the system with large distributed databases and in a real setting in which six databases store the data from six video cameras. The simulations confirm that our method provides good results with varying numbers of cameras and persons moving in the network. In a real setting, the method reconstructs paths across the camera network with approximatively 75% accuracy at rank 1.
Liliana Lo Presti, Stan Sclaroff, Marco La Cascia
IEEE Trans. Multim.2
2012 Event prediction in a hybrid camera network
abstract
Given a hybrid camera layout—one containing, for example, static and active cameras—and people moving around following established traffic patterns, our goal is to predict a subset of cameras, respective camera parameter settings, and future time windows that will most likely lead to success the vision tasks, such as, face recognition when a camera observes an event of interest. We propose an adaptive probabilistic model that accrues temporal camera correlations over time as the cameras report observed events. No extrinsic, intrinsic, or color calibration of cameras is required. We efficiently obtain the camera parameter predictions using a modified Sequential Monte Carlo method. We demonstrate the performance of the model in an example face detection scenario in both simulated and real environment experiments, using several active cameras.
Ugur Murat Erdem, Stan Sclaroff
ACM Trans. Sens. Networks2
2011 Scale and rotation invariant matching using linearly augmented trees
abstract
We propose a novel linearly augmented tree method for efficient scale and rotation invariant object matching. The proposed method enforces pairwise matching consistency defined on trees, and high-order constraints on all the sites of a template. The pairwise constraints admit arbitrary metrics while the high-order constraints use L1 norms and therefore can be linearized. Such a linearly augmented tree formulation introduces hyperedges and loops into the basic tree structure, but different from a general loopy graph, its special structure allows us to relax and decompose the optimization into a sequence of tree matching problems efficiently solvable by dynamic programming. The proposed method also works on continuous scale and rotation parameters; we can match with a scale up to any large number with the same efficiency. Our experiments on ground truth data and a variety of real images and videos show that the proposed method is efficient, accurate and reliable.
Hao Jiang 0007, Tai-Peng Tian, Stan Sclaroff
CVPR3
2011 Exploiting phonological constraints for handshape inference in ASL video
abstract
Handshape is a key linguistic component of signs, and thus, handshape recognition is essential to algorithms for sign language recognition and retrieval. In this work, linguistic constraints on the relationship between start and end handshapes are leveraged to improve handshape recognition accuracy. A Bayesian network formulation is proposed for learning and exploiting these constraints, while taking into consideration inter-signer variations in the production of particular handshapes. A Variational Bayes formulation is employed for supervised learning of the model parameters. A non-rigid image alignment algorithm, which yields improved robustness to variability in handshape appearance, is proposed for computing image observation likelihoods in the model. The resulting handshape inference algorithm is evaluated using a dataset of 1500 lexical signs in American Sign Language (ASL), where each lexical sign is produced by three native ASL signers.
Ashwin Thangali, Joan P. Nash, Stan Sclaroff, Carol Neidle
CVPR3
2011 Learning parameterized histogram kernels on the simplex manifold for image and action classification
abstract
State-of-the-art image and action classification systems often employ vocabulary-based representations. The classification accuracy achieved with such vocabulary-based representations depends significantly on the chosen histogram-distance. In particular, when the decision function is a support-vector-machine (SVM), the classification accuracy depends on the chosen histogram kernel. In this paper we focus on smoothly-parameterized kernels in the space of histograms, such as, but not limited to, kernels that are derived from smoothly-parameterized histogram-distance functions. We learn parameters of histogram kernels so that the SVM accuracy is improved. This is accomplished by simultaneously maximizing the SVM's geometric margin and minimizing an estimate of its generalization error. We validate our approach on a previously-published two-class synthetic dataset and three real-world multi-class datasets: Oxford5K, KTH, and UCF. On these datasets our approach yields results that compare favorably to or exceed the state of the art.
Vitaly Ablavsky, Stan Sclaroff
ICCV2
2011 Layered Graphical Models for Tracking Partially Occluded Objects
abstract
We propose a representation for scenes containing relocatable objects that can cause partial occlusions of people in a camera's field of view. In many practical applications, relocatable objects tend to appear often; therefore, models for them can be learned offline and stored in a database. We formulate an occluder-centric representation, called a graphical model layer, where a person's motion in the ground plane is defined as a first-order Markov process on activity zones, while image evidence is aggregated in 2D observation regions that are depth-ordered with respect to the occlusion mask of the relocatable object. We represent real-world scenes as a composition of depth-ordered, interacting graphical model layers, and account for image evidence in a way that handles mutual overlap of the observation regions and their occlusions by the relocatable objects. These layers interact: Proximate ground-plane zones of different model instances are linked to allow a person to move between the layers, and image evidence is shared between the observation regions of these models. We demonstrate our formulation in tracking pedestrians in the vicinity of parked vehicles. Our results compare favorably with a sprite-learning algorithm, with a pedestrian tracker based on deformable contours, and with pedestrian detectors.
Vitaly Ablavsky, Stan Sclaroff
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Learning a Family of Detectors via Multiplicative Kernels
abstract
Object detection is challenging when the object class exhibits large within-class variations. In this work, we show that foreground-background classification (detection) and within-class classification of the foreground class (pose estimation) can be jointly learned in a multiplicative form of two kernel functions. Model training is accomplished via standard SVM learning. When the foreground object masks are provided in training, the detectors can also produce object segmentations. A tracking-by-detection framework to recover foreground state in video sequences is also proposed with our model. The advantages of our method are demonstrated on tasks of object detection, view angle estimation, and tracking. Our approach compares favorably to existing methods on hand and vehicle detection tasks. Quantitative tracking results are given on sequences of moving vehicles and human faces.
Ashwin Thangali, Vitaly Ablavsky, Stan Sclaroff
IEEE Trans. Pattern Anal. Mach. Intell.4
2010 Fast globally optimal 2D human detection with loopy graph models
abstract
This paper presents an algorithm for recovering the globally optimal 2D human figure detection using a loopy graph model. This is computationally challenging because the time complexity scales exponentially in the size of the largest clique in the graph. The proposed algorithm uses Branch and Bound (BB) to search for the globally optimal solution. The algorithm converges rapidly in practice and this is due to a novel method for quickly computing tree based lower bounds. The key idea is to recycle the dynamic programming (DP) tables associated with the tree model to look up the tree based lower bound rather than recomputing the lower bound from scratch. This technique is further sped up using Range Minimum Query data structures to provide O(1) cost for computing the lower bound for most iterations of the BB algorithm. The algorithm is evaluated on the Iterative Parsing dataset and it is shown to run fast empirically.
Tai-Peng Tian, Stan Sclaroff
CVPR2
2010 Object, Scene and Actions: Combining Multiple Features for Human Action Recognition
Nazli Ikizler-Cinbis, Stan Sclaroff
ECCV (1)2
2010 Fast Multi-aspect 2D Human Detection
Tai-Peng Tian, Stan Sclaroff
ECCV (3)2
2010 Object Recognition and Localization Via Spatial Instance Embedding
abstract
We propose an approach for improving object recognition and localization using spatial kernels together with instance embedding. Our approach treats each image as a bag of instances (image features) within a multiple instance learning framework, where the relative locations of the instances are considered as well as the appearance similarity of the localized image features. The introduced spatial kernel augments the recognition power of the instance embedding in an intuitive and effective way, providing increased localization performance. We test our approach over two object datasets and present promising results.
Nazli Ikizler-Cinbis, Stan Sclaroff
ICPR2
2010 3D Human Motion Tracking with a Coordinated Mixture of Factor Analyzers
abstract
A major challenge in applying Bayesian tracking methods for tracking 3D human body pose is the high dimensionality of the pose state space. It has been observed that the 3D human body pose parameters typically can be assumed to lie on a low-dimensional manifold embedded in the high-dimensional space. The goal of this work is to approximate the low-dimensional manifold so that a low-dimensional state vector can be obtained for efficient and effective Bayesian tracking. To achieve this goal, a globally coordinated mixture of factor analyzers is learned from motion capture data. Each factor analyzer in the mixture is a “locally linear dimensionality reducer” that approximates a part of the manifold. The global parametrization of the manifold is obtained by aligning these locally linear pieces in a global coordinate system. To enable automatic and optimal selection of the number of factor analyzers and the dimensionality of the manifold, a variational Bayesian formulation of the globally coordinated mixture of factor analyzers is proposed. The advantages of the proposed model are demonstrated in a multiple hypothesis tracker for tracking 3D human body pose. Quantitative comparisons on benchmark datasets show that the proposed method produces more accurate 3D pose estimates over time than those obtained from two previously proposed Bayesian tracking methods.
Rui Li 0053, Tai-Peng Tian, Stan Sclaroff, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.3
2010 Matching Trajectories between Video Sequences by Exploiting a Sparse Projective Invariant Representation
abstract
Identifying correspondences between trajectory segments observed from nonsynchronized cameras is important for reconstruction of the complete trajectory of moving targets in a large scene. Such a reconstruction can be obtained from motion data by comparing the trajectory segments and estimating both the spatial and temporal alignments. Exhaustive testing of all possible correspondences of trajectories over a temporal window is only viable in the cases with a limited number of moving targets and large view overlaps. Therefore, alternative solutions are required for situations with several trajectories that are only partially visible in each view. In this paper, we propose a new method that is based on view-invariant representation of trajectories, which is used to produce a sparse set of salient points for trajectory segments observed in each view. Only the neighborhoods at these salient points in the view--invariant representation are then used to estimate the spatial and temporal alignment of trajectory pairs in different views. It is demonstrated that, for planar scenes, the method is able to recover with good precision and efficiency both spatial and temporal alignments, even given relatively small overlap between views and arbitrary (unknown) temporal shifts of the cameras. The method also provides the same capabilities in the case of trajectories that are only locally planar, but exhibit some nonplanarity at a global level.
Walter Nunziati, Stan Sclaroff, Alberto Del Bimbo
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 Reducing JointBoost-based multiclass classification to proximity search
abstract
Boosted one-versus-all (OVA) classifiers are commonly used in multiclass problems, such as generic object recognition, biometrics-based identification, or gesture recognition. JointBoost is a recently proposed method where OVA classifiers are trained jointly and are forced to share features. JointBoost has been demonstrated to lead both to higher accuracy and smaller classification time, compared to using OVA classifiers that were trained independently and without sharing features. However, even with the improved efficiency of JointBoost, the time complexity of OVA-based multiclass recognition is still linear to the number of classes, and can lead to prohibitively large running times in domains with a very large number of classes. In this paper, it is shown that JointBoost-based recognition can be reduced, at classification time, to nearest neighbor search in a vector space. Using this reduction, we propose a simple and easy-to-implement vector indexing scheme based on principal component analysis (PCA). In our experiments, the proposed method achieves a speedup of two orders of magnitude over standard JointBoost classification, in a hand pose recognition system where the number of classes is close to 50,000, with negligible loss in classification accuracy. Our method also yields promising results in experiments on the widely used FRGC-2 face recognition dataset, where the number of classes is 535.
Alexandra Stefan, Vassilis Athitsos, Stan Sclaroff
CVPR4
2009 Learning actions from the Web
abstract
This paper proposes a generic method for action recognition in uncontrolled videos. The idea is to use images collected from the Web to learn representations of actions and use this knowledge to automatically annotate actions in videos. Our approach is unsupervised in the sense that it requires no human intervention other than the text querying. Its benefits are two-fold: 1) we can improve retrieval of action images, and 2) we can collect a large generic database of action poses, which can then be used in tagging videos. We present experimental evidence that using action images collected from the Web, annotating actions is possible.
Nazli Ikizler-Cinbis, Ramazan Gokberk Cinbis, Stan Sclaroff
ICCV3
2009 Is a detector only good for detection?
abstract
A common design of an object recognition system has two steps, a detection step followed by a foreground within-class classification step. For example, consider face detection by a boosted cascade of detectors followed by face ID recognition via one-vs-all (OVA) classifiers. Another example is human detection followed by pose recognition. Although the detection step can be quite fast, the foreground within-class classification process can be slow and becomes a bottleneck. In this work, we formulate a filter-and-refine scheme, where the binary outputs of the weak classifiers in a boosted detector are used to identify a small number of candidate foreground state hypotheses quickly via Hamming distance or weighted Hamming distance. The approach is evaluated in three applications: face recognition on the FRGC V2 data set, hand shape detection and parameter estimation on a hand data set and vehicle detection and view angle estimation on a multi-view vehicle data set. On all data sets, our approach has comparable accuracy and is at least five times faster than the brute force approach.
Stan Sclaroff
ICCV2
2009 Mining frequent arrangements of temporal intervals
Panagiotis Papapetrou, George Kollios, Stan Sclaroff, Dimitrios Gunopulos
Knowl. Inf. Syst.3
2009 A Unified Framework for Gesture Recognition and Spatiotemporal Gesture Segmentation
abstract
Within the context of hand gesture recognition, spatiotemporal gesture segmentation is the task of determining, in a video sequence, where the gesturing hand is located and when the gesture starts and ends. Existing gesture recognition methods typically assume either known spatial segmentation or known temporal segmentation, or both. This paper introduces a unified framework for simultaneously performing spatial segmentation, temporal segmentation, and recognition. In the proposed framework, information flows both bottom-up and top-down. A gesture can be recognized even when the hand location is highly ambiguous and when information about when the gesture begins and ends is unavailable. Thus, the method can be applied to continuous image streams where gestures are performed in front of moving, cluttered backgrounds. The proposed method consists of three novel contributions: a spatiotemporal matching algorithm that can accommodate multiple candidate hand detections in every frame, a classifier-based pruning framework that enables accurate and early rejection of poor matches to gesture models, and a subgesture reasoning algorithm that learns which gesture models can falsely match parts of other longer gestures. The performance of the approach is evaluated on two challenging applications: recognition of hand-signed digits gestured by users wearing short-sleeved shirts, in front of a cluttered background, and retrieval of occurrences of signs of interest in a video database containing continuous, unsegmented signing in American Sign Language (ASL).
Jonathan Alon, Vassilis Athitsos, Stan Sclaroff
IEEE Trans. Pattern Anal. Mach. Intell.4
2009 Sign Language Spotting with a Threshold Model Based on Conditional Random Fields
abstract
Sign language spotting is the task of detecting and recognizing signs in a signed utterance, in a set vocabulary. The difficulty of sign language spotting is that instances of signs vary in both motion and appearance. Moreover, signs appear within a continuous gesture stream, interspersed with transitional movements between signs in a vocabulary and nonsign patterns (which include out-of-vocabulary signs, epentheses, and other movements that do not correspond to signs). In this paper, a novel method for designing threshold models in a conditional random field (CRF) model is proposed which performs an adaptive threshold for distinguishing between signs in a vocabulary and nonsign patterns. A short-sign detector, a hand appearance-based sign verification method, and a subsign reasoning method are included to further improve sign language spotting accuracy. Experiments demonstrate that our system can spot signs from continuous data with an 87.0 percent spotting rate and can recognize signs from isolated data with a 93.5 percent recognition rate versus 73.5 percent and 85.4 percent, respectively, for CRFs without a threshold model, short-sign detection, subsign reasoning, and hand appearance-based sign verification. Our system can also achieve a 15.0 percent sign error rate (SER) from continuous data and a 6.4 percent SER from isolated data versus 76.2 percent and 14.5 percent, respectively, for conventional CRFs.
Hee-Deok Yang, Stan Sclaroff, Seong-Whan Lee
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Layered graphical models for tracking partially-occluded objects
abstract
Partial occlusions are commonplace in a variety of real world computer vision applications: surveillance, intelligent environments, assistive robotics, autonomous navigation, etc. While occlusion handling methods have been proposed, most methods tend to break down when confronted with numerous occluders in a scene. In this paper, a layered image-plane representation for tracking people through substantial occlusions is proposed. An image-plane representation of motion around an object is associated with a pre-computed graphical model, which can be instantiated efficiently during online tracking. A global state and observation space is obtained by linking transitions between layers. A reversible jump Markov chain Monte Carlo approach is used to infer the number of people and track them online. The method outperforms two state-of-the-art methods for tracking over extended occlusions, given videos of a parking lot with numerous vehicles and a laboratory with many desks and workstations.
Vitaly Ablavsky, Ashwin Thangali, Stan Sclaroff
CVPR3
2008 Multiplicative kernels: Object detection, segmentation and pose estimation
abstract
Object detection is challenging when the object class exhibits large within-class variations. In this work, we show that foreground-background classification (detection) and within-class classification of the foreground class (pose estimation) can be jointly learned in a multiplicative form of two kernel functions. One kernel measures similarity for foreground-background classification. The other kernel accounts for latent factors that control within-class variation and implicitly enables feature sharing among foreground training samples. Detector training can be accomplished via standard SVM learning. The resulting detectors are tuned to specific variations in the foreground class. They also serve to evaluate hypotheses of the foreground state. When the foreground parameters are provided in training, the detectors can also produce parameter estimate. When the foreground object masks are provided in training, the detectors can also produce object segmentation. The advantages of our method over past methods are demonstrated on data sets of human hands and vehicles.
Ashwin Thangali, Vitaly Ablavsky, Stan Sclaroff
CVPR4
2008 Tracking with Dynamic Hidden-State Shape Models
Zheng Wu 0003, Margrit Betke, Jingbin Wang, Vassilis Athitsos, Stan Sclaroff
ECCV (1)5
2008 Benchmark Databases for Video-Based Automatic Sign Language Recognition
Philippe Dreuw, Carol Neidle, Vassilis Athitsos, Stan Sclaroff, Hermann Ney
LREC4
2008 Multi-scale 3D scene flow from binocular stereo sequences
Rui Li 0053, Stan Sclaroff
Comput. Vis. Image Underst.2
2008 BoostMap: An Embedding Method for Efficient Nearest Neighbor Retrieval
abstract
This paper describes BoostMap, a method for efficient nearest neighbor retrieval under computationally expensive distance measures. Database and query objects are embedded into a vector space, in which distances can be measured efficiently. Each embedding is treated as a classifier that predicts for any three objects X, A, B whether X is closer to A or to B. It is shown that a linear combination of such embeddingbased classifiers naturally corresponds to an embedding and a distance measure. Based on this property, the BoostMap method reduces the problem of embedding construction to the classical boosting problem of combining many weak classifiers into an optimized strong classifier. The classification accuracy of the resulting strong classifier is a direct measure of the amount of nearest neighbor structure preserved by the embedding. An important property of BoostMap is that the embedding optimization criterion is equally valid in both metric and non-metric spaces. Performance is evaluated in databases of hand images, handwritten digits, and time series. In all cases, BoostMap significantly improves retrieval efficiency with small losses in accuracy compared to brute-force search. Moreover, BoostMap significantly outperforms existing nearest neighbor retrieval methods, such as Lipschitz embeddings, FastMap, and VP-trees.
Vassilis Athitsos, Jonathan Alon, Stan Sclaroff, George Kollios
IEEE Trans. Pattern Anal. Mach. Intell.3
2008 Detecting Objects of Variable Shape Structure With Hidden State Shape Models
abstract
This paper proposes a method for detecting object classes that exhibit variable shape structure in heavily cluttered images. The term "variable shape structure" is used to characterize object classes in which some shape parts can be repeated an arbitrary number of times, some parts can be optional, and some parts can have several alternative appearances. Hidden State Shape Models (HSSMs), a generalization of Hidden Markov Models (HMMs), are introduced to model object classes of variable shape structure using a probabilistic framework. A polynomial inference algorithm automatically determines object location, orientation, scale and structure by finding the globally optimal registration of model states with the image features, even in the presence of clutter. Experiments with real images demonstrate that the proposed method can localize objects of variable shape structure with high accuracy. For the task of hand shape localization and structure identification, the proposed method is significantly more accurate than previously proposed methods based on chamfer-distance matching. Furthermore, by integrating simple temporal constraints, the proposed method gains speed-ups of more than an order of magnitude, and produces highly accurate results in experiments on non-rigid hand motion tracking.
Jingbin Wang, Vassilis Athitsos, Stan Sclaroff, Margrit Betke
IEEE Trans. Pattern Anal. Mach. Intell.3
2007 Parameter Sensitive Detectors
abstract
Object detection can be challenging when the object class exhibits large variations. One commonly-used strategy is to first partition the space of possible object variations and then train separate classifiers for each portion. However, with continuous spaces the partitions tend to be arbitrary since there are no natural boundaries (for example, consider the continuous range of human body poses). In this paper, a new formulation is proposed, where the detectors themselves are associated with continuous parameters, and reside in a parameterized function space. There are two advantages of this strategy. First, a-priori partitioning of the parameter space is not needed; the detectors themselves are in a parameterized space. Second, the underlying parameters for object variations can be learned from training data in an unsupervised manner. In profile face detection experiments, at a fixed false alarm number of 90, our method attains a detection rate of 75% vs. 70% for the method of Viola-Jones. In hand shape detection, at a false positive rate of 0.1%, our method achieves a detection rate of 99.5% vs. 98% for partition based methods. In pedestrian detection, our method reduces the miss detection rate by a factor of three at a false positive rate of 1%, compared with the method of Dalal-Triggs.
Ashwin Thangali, Vitaly Ablavsky, Stan Sclaroff
CVPR4
2007 ClassMap: Efficient Multiclass Recognition via Embeddings
abstract
In many computer vision applications, such as face recognition and hand pose estimation, we need systems that can recognize a very large number of classes. Large margin classification methods, such as AdaBoost and SVMs, often provide competitive accuracy rates, but at the cost of evaluating a large number of binary classifiers. We propose an embedding-based method for efficient multiclass recognition. In our method, patterns and classes are mapped to vectors in such a way that patterns and their associated classes tend to get mapped close to each other. This way, given a test pattern, a small set of candidate classes can be identified efficiently using simple vector comparisons. In experiments with 3D hand pose recognition (2430 classes) and face recognition (535 classes), our method is between 3 and 28 times faster compared to evaluating all binary classifiers, with negligible or no loss in classification accuracy.
Vassilis Athitsos, Alexandra Stefan, Stan Sclaroff
ICCV4
2007 Simultaneous Learning of Nonlinear Manifold and Dynamical Models for High-dimensional Time Series
abstract
The goal of this work is to learn a parsimonious and informative representation for high-dimensional time series. Conceptually, this comprises two distinct yet tightly coupled tasks: learning a low-dimensional manifold and modeling the dynamical process. These two tasks have a complementary relationship as the temporal constraints provide valuable neighborhood information for dimensionality reduction and conversely, the low-dimensional space allows dynamics to be learnt efficiently. Solving these two tasks simultaneously allows important information to be exchanged mutually. If nonlinear models are required to capture the rich complexity of time series, then the learning problem becomes harder as the nonlinearities in both tasks are coupled. The proposed solution approximates the nonlinear manifold and dynamics using piecewise linear models. The interactions among the linear models are captured in a graphical model. By exploiting the model structure, efficient inference and learning algorithms are obtained without oversimplifying the model of the underlying dynamical process. Evaluation of the proposed framework with competing approaches is conducted in three sets of experiments: dimensionality reduction and reconstruction using synthetic time series, video synthesis using a dynamic texture database, and human motion synthesis, classification and tracking on a benchmark data set. In all experiments, the proposed approach provides superior performance.
Rui Li 0053, Tai-Peng Tian, Stan Sclaroff
ICCV3
2007 Query-sensitive embeddings
abstract
A common problem in many types of databases is retrieving the most similar matches to a query object. Finding these matches in a large database can be too slow to be practical, especially in domains where objects are compared using computationally expensive similarity (or distance) measures. Embedding methods can significantly speed-up retrieval by mapping objects into a vector space, where distances can be measured rapidly using a Minkowski metric. In this article we present a novel way to improve embedding quality. In particular, we propose to construct embeddings that use a query-sensitive distance measure for the target space of the embedding. This distance measure is used to compare those vectors that the query and database objects are mapped to. The term “query-sensitive” means that the distance measure changes, depending on the current query object. We demonstrate theoretically that using a query-sensitive distance measure increases the modeling power of embeddings and allows them to capture more of the structure of the original space. We also demonstrate experimentally that query-sensitive embeddings can significantly improve retrieval performance. In experiments with an image database of handwritten digits and a time-series database, the proposed method outperforms existing state-of-the-art non-Euclidean indexing methods, meaning that it provides significantly better tradeoffs between efficiency and retrieval accuracy.
Vassilis Athitsos, Marios Hadjieleftheriou, George Kollios, Stan Sclaroff
ACM Trans. Database Syst.4
2006 Detecting Instances of Shape Classes That Exhibit Variable Structure
Vassilis Athitsos, Jingbin Wang, Stan Sclaroff, Margrit Betke
ECCV (1)3
2006 Monocular Tracking of 3D Human Motion with a Coordinated Mixture of Factor Analyzers
Rui Li 0053, Ming-Hsuan Yang 0001, Stan Sclaroff, Tai-Peng Tian
ECCV (2)3
2006 Automated camera layout to satisfy task-specific and floor plan-specific coverage requirements
Ugur Murat Erdem, Stan Sclaroff
Comput. Vis. Image Underst.2
2006 Combining Generative and Discriminative Models in a Framework for Articulated Pose Estimation
Rómer Rosales, Stan Sclaroff
Int. J. Comput. Vis.2
2005 Look there! Predicting where to look for motion in an active camera network
abstract
A framework is proposed that answers the following question: if a moving object is observed by one camera in a pan-tilt-zoom (PTZ) camera network, what other camera(s) might be foveated on that object within a predefined time window, and what would be the corresponding PTZ parameter settings? No calibration is assumed, and there are no restrictions on camera placement or initial parameter settings. The framework accrues a predictive model over time. To start out, the cameras follow randomized "tours" in discretized PTZ space. If a moving object is detected in the field of view of more than one camera at a particular instant or within a predefined time window, then the model is updated to record the cameras' associations and the corresponding parameter settings. As more and more moving objects are observed, the model adapts and the most frequent associations are discovered. The formulation also allows for verification of its predictions, and reinforces its correct predictions. The system is demonstrated in observing people in an office environment with a three PTZ camera network.
Ugur Murat Erdem, Stan Sclaroff
AVSS2
2005 View registration using interesting segments of planar trajectories
abstract
We introduce a method for recovering the spatial and temporal alignment between two or more views of objects moving over a ground plane. Existing approaches either assume that the streams are globally synchronized, so that only solving the spatial alignment is needed, or that the temporal misalignment is small enough so that exhaustive search can be performed. In contrast, our approach can recover both the spatial and temporal alignment, regardless of their magnitude. We compute for each trajectory a number of interesting segments, and we use their description to form putative matches between trajectories. Each pair of corresponding interesting segments induces a temporal alignment, and defines an interval of common support across two views of an object that is used to recover the spatial alignment. Interesting segments and their descriptors are defined using algebraic projective invariants measured along the trajectories. Similarity between interesting segments is computed taking into account the statistics of such invariants. Candidate alignment parameters are verified by checking the consistency, in terms of the symmetric transfer error, of all the putative pairs of corresponding interesting segments. Experiments are conducted with two different sets of data, one with two views of an outdoor scene featuring moving people and cars, and one with four views of a laboratory sequence featuring moving radio-controlled cars.
Walter Nunziati, Jonathan Alon, Stan Sclaroff, Alberto Del Bimbo
AVSS3
2005 Efficient Nearest Neighbor Classification Using a Cascade of Approximate Similarity Measures
abstract
This paper proposes a method for efficient nearest neighbor classification in non-Euclidean spaces with computationally expensive similarity/distance measures. Efficient approximations of such measures are obtained using the BoostMap algorithm, which produces embeddings into a real vector space. A modification to the BoostMap algorithm is proposed, which uses an optimization cost that is more appropriate when our goal is classification accuracy as opposed to nearest neighbor retrieval accuracy. Using the modified algorithm, multiple approximate nearest neighbor classifiers are obtained, that provide a wide range of trade-offs between accuracy and efficiency. The approximations are automatically combined to form a cascade classifier, which applies the slower and more accurate approximations only to the hardest cases. The proposed method is experimentally evaluated in the domain of handwritten digit recognition using shape context matching. Results on the MNIST database indicate that a speed-up of two to three orders of magnitude is gained over brute force search, with minimal losses in classification accuracy.
Vassilis Athitsos, Jonathan Alon, Stan Sclaroff
CVPR (1)3
2005 Online and Offline Character Recognition Using Alignment to Prototypes
abstract
Nearest neighbor classifiers are simple to implement, yet they can model complex non-parametric distributions, and provide state-of-the-art recognition accuracy in OCR databases. At the same time, they may be too slow for practical character recognition, especially when they rely on similarity measures that require computationally expensive pair-wise alignments between characters. This paper proposes an efficient method for computing an approximate similarity score between two characters based on their exact alignment to a small number of prototypes. The proposed method is applied to both online and offline character recognition, where similarity is based on widely used and computationally expensive alignment methods, i.e., dynamic time warping and the Hungarian method respectively. In both cases significant recognition speedup is obtained at the expense of only a minor increase in recognition error.
Jonathan Alon, Vassilis Athitsos, Stan Sclaroff
ICDAR3
2005 Tracking, Analysis, and Recognition of Human Gestures in Video
abstract
An overview of research in automated gesture spotting, tracking and recognition by the Image and Video Computing Group at Boston University is given. Approaches for localization and tracking human hands in video, estimation of hand shape and upper body pose; tracking head and facial motion, as well as efficient spotting and recognition of specific gestures in video streams are summarized. Methods for efficient dimensionality reduction of gesture time series, boosting of classifiers for nearest neighbor search in pose space, and model-based pruning of gesture alignment hypotheses are described. Algorithms are demonstrated in three domains: American sign language, hand signals like those employed by flight-directors on airport runways, and gesture-based interfaces for severely disabled users. The methods described are general and can be applied in other domains that require efficient detection and analysis of patterns in time-series, images or video.
Stan Sclaroff, Margrit Betke, George Kollios, Jonathan Alon, Vassilis Athitsos, Rui Li 0053, John J. Magee, Tai-Peng Tian
ICDAR1
2005 Discovering Frequent Arrangements of Temporal Intervals
abstract
In this paper we study a new problem in temporal pattern mining: discovering frequent arrangements of temporal intervals. We assume that the database consists of sequences of events, where an event occurs during a time-interval. The goal is to mine arrangements of event intervals that appear frequently in the database. There are many applications where these type of patterns can be useful, including data network, scientific, and financial applications. Efficient methods to find frequent arrangements of temporal intervals using both breadth first and depth first search techniques are described. The performance of the proposed algorithms is evaluated and compared with other approaches on real datasets (American sign language streams and network data) and large synthetic datasets.
Panagiotis Papapetrou, George Kollios, Stan Sclaroff, Dimitrios Gunopulos
ICDM3
2005 Query-Sensitive Embeddings
abstract
A common problem in many types of databases is retrieving the most similar matches to a query object. Finding those matches in a large database can be too slow to be practical, especially in domains where objects are compared using computationally expensive similarity (or distance) measures. This paper proposes a novel method for approximate nearest neighbor retrieval in such spaces. Our method is embedding-based, meaning that it constructs a function that maps objects into a real vector space. The mapping preserves a large amount of the proximity structure of the original space, and it can be used to rapidly obtain a short list of likely matches to the query. The main novelty of our method is that it constructs, together with the embedding, a query-sensitive distance measure that should be used when measuring distances in the vector space. The term "query-sensitive" means that the distance measure changes depending on the current query object. We report experiments with an image database of handwritten digits, and a time-series database. In both cases, the proposed method outperforms existing state-of-the-art embedding methods, meaning that it provides significantly better trade-offs between efficiency and retrieval accuracy.
Vassilis Athitsos, Marios Hadjieleftheriou, George Kollios, Stan Sclaroff
SIGMOD Conference4
2004 BoostMap: A Method for Efficient Approximate Similarity Rankings
Vassilis Athitsos, Jonathan Alon, Stan Sclaroff, George Kollios
CVPR (2)3
2004 Fully automatic, real-time detection of facial gestures from generic video
abstract
A technique for the detection of facial gestures from low resolution video sequences is presented. The technique builds upon the automatic 3D head tracker formulation of [M. La Cascia et al., 2000]. The tracker is based on the registration of a texture-mapped cylindrical model. Facial gesture analysis is performed in the texture map by assuming that the residual registration error can be modeled as a linear combination of facial motion templates. Two formulations are proposed and tested. In one formulation, the head and facial motion are estimated in a single, combined linear system. In the other formulation, head motion and then facial motion are estimated in a two-step process. The two-step approach significantly yields better accuracy in facial gesture analysis. The system is demonstrated in detecting two types of facial gestures: "mouth opening" and "eyebrows raising." On a dataset with lots of head motion, the two-step algorithm achieved a recognition accuracy of 70% for the "mouth opening" and an accuracy of 66% for "eyebrows raising" gestures. The algorithm can reliably track and classify facial gestures without any user intervention and runs in real-time.
Marco La Cascia, Lorenzo Valenti, Stan Sclaroff
MMSP3
2004 Deformable model-guided region split and merge of image regions
Lifeng Liu, Stan Sclaroff
Image Vis. Comput.2
2004 Skin Color-Based Video Segmentation under Time-Varying Illumination
abstract
A novel approach for real-time skin segmentation in video sequences is described. The approach enables reliable skin segmentation despite wide variation in illumination during tracking. An explicit second order Markov model is used to predict evolution of the skin-color (HSV) histogram over time. Histograms are dynamically updated based on feedback from the current segmentation and predictions of the Markov model. The evolution of the skin-color distribution at each frame is parameterized by translation, scaling, and rotation in color space. Consequent changes in geometric parameterization of the distribution are propagated by warping and resampling the histogram. The parameters of the discrete-time dynamic Markov model are estimated using Maximum Likelihood Estimation and also evolve over time. The accuracy of the new dynamic skin color segmentation algorithm is compared to that obtained via a static color model. Segmentation accuracy is evaluated using labeled ground-truth video sequences taken from staged experiments and popular movies. An overall increase in segmentation accuracy of up to 24 percent is observed in 17 out of 21 test sequences. In all but one case, the skin-color classification rates for our system were higher, with background classification rates comparable to those of the static segmentation.
Leonid Sigal, Stan Sclaroff, Vassilis Athitsos
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Discovering Clusters in Motion Time-Series Data
abstract
An approach is proposed for clustering time-series data. The approach can be used to discover groupings of similar object motions that were observed in a video collection. A finite mixture of hidden Markov models (HMMs) is fitted to the motion data using the expectation maximization (EM) framework. Previous approaches for HMM-based clustering employ a k-means formulation, where each sequence is assigned to only a single HMM. In contrast, the formulation presented in this paper allows each sequence to belong to more than a single HMM with some probability, and the hard decision about the sequence class membership can be deferred until a later time when such a decision is required. Experiments with simulated data demonstrate the benefit of using this EM-based approach when there is more "overlap" in the processes generating the data. Experiments with real data show the promising potential of HMM-based motion clustering in a number of applications.
Jonathan Alon, Stan Sclaroff, George Kollios, Vladimir Pavlovic 0001
CVPR (1)2
2003 Estimating 3D Hand Pose from a Cluttered Image
abstract
A method is proposed that can generate a ranked list of plausible three-dimensional hand configurations that best match an input image. Hand pose estimation is formulated as an image database indexing problem, where the closest matches for an input hand image are retrieved from a large database of synthetic hand images. In contrast to previous approaches, the system can function in the presence of clutter, thanks to two novel clutter-tolerant indexing methods. First, a computationally efficient approximation of the image-to-model chamfer distance is obtained by embedding binary edge images into a high-dimensional Euclidean space. Second, a general-purpose, probabilistic line matching method identifies those line segment correspondences between model and input images that are the least likely to have occurred by chance. The performance of this clutter tolerant approach is demonstrated in quantitative experiments with hundreds of real hand images.
Vassilis Athitsos, Stan Sclaroff
CVPR (2)2
2003 Stochastic Refinement of the Visual Hull to Satisfy Photometric and Silhouette Consistency Constraints
abstract
An iterative method for reconstructing a 3D polygonal mesh and color texture map from multiple views of an object is presented. In each iteration, the method first estimates a texture map given the current shape estimate. The texture map and its associated residual error image are obtained via maximum a posteriori estimation and reprojection of the multiple views into texture space. Next, the surface shape is adjusted to minimize residual error in texture space. The surface is deformed towards a photometrically-consistent solution via a series of 1D epipolar searches at randomly selected surface points. The texture space formulation has improved computational complexity over standard image-based error approaches, and allows computation of the reprojection error and uncertainty for any point on the surface. Moreover, shape adjustments can be constrained such that the recovered model's silhouette matches those of the input images. Experiments with real world imagery demonstrate the validity of the approach.
John Isidoro, Stan Sclaroff
ICCV2
2003 Segmenting Foreground Objects from a Dynamic Textured Background via a Robust Kalman Filter
abstract
The algorithm presented aims to segment the foreground objects in video (e.g., people) given time-varying, textured backgrounds. Examples of time-varying backgrounds include waves on water, clouds moving, trees waving in the wind, automobile traffic, moving crowds, escalators, etc. We have developed a novel foreground-background segmentation algorithm that explicitly accounts for the nonstationary nature and clutter-like appearance of many dynamic textures. The dynamic texture is modeled by an autoregressive moving average model (ARMA). A robust Kalman filter algorithm iteratively estimates the intrinsic appearance of the dynamic texture, as well as the regions of the foreground objects. Preliminary experiments with this method have demonstrated promising results.
Stan Sclaroff
ICCV2
2003 A framework for heading-guided recognition of human activity
Rómer Rosales, Stan Sclaroff
Comput. Vis. Image Underst.2
2003 Active blobs: region-based, deformable appearance models
Stan Sclaroff, John Isidoro
Comput. Vis. Image Underst.1
2002 Index trees for accelerating deformable template matching
Lifeng Liu, Stan Sclaroff
Pattern Recognit. Lett.2
2001 Estimating 3D Body Pose using Uncalibrated Cameras
abstract
An approach for estimating 3D body pose from multiple, uncalibrated views is proposed. First, a mapping from image features to 2D body joint locations is computed using a statistical framework that yields a set of several body pose hypotheses. The concept of a "virtual camera" is introduced that makes this mapping invariant to translation, image-plane rotation, and scaling of the input. As a consequence, the calibration matrices (intrinsics) of the virtual cameras can be considered completely known, and their poses are known up to a single angular displacement parameter Given pose hypotheses obtained in the multiple virtual camera views, the recovery of 3D body pose and camera relative orientations is formulated as a stochastic optimization problem. An Expectation-Maximization algorithm is derived that can obtain the locally most likely (self-consistent) combination of body pose hypotheses. Performance of the approach is evaluated with synthetic sequences as well as real video sequences of human motion.
Rómer Rosales, Matheen Siddiqui, Jonathan Alon, Stan Sclaroff
CVPR (1)4
2001 Region Segmentation via Deformable Model-Guided Split and Merge
Lifeng Liu, Stan Sclaroff
ICCV2
2001 3D Hand Pose Reconstruction Using Specialized Mappings
Rómer Rosales, Vassilis Athitsos, Leonid Sigal, Stan Sclaroff
ICCV4
2001 Medical image segmentation and retrieval via deformable models
abstract
A new method based on deformable shape models for medical image segmentation is described. Experiments for blood cell micrographs have been conducted to verify the accuracy of the shape model-based segmentation and object shape description method. The cell segmentation method does not require user input for initialization. Coherence information between cells is utilized via a globally consistent cost function. The proposed segmentation method can be used in automated analysis for images of stained blood smear and segmentation of other medical structures. A method for shape population-based retrieval is also described. Results of population-based image queries for a database of blood cell micrographs are shown.
Lifeng Liu, Stan Sclaroff
ICIP (3)2
2001 Learning Body Pose via Specialized Maps
abstract
A nonlinear supervised learning model, the Specialized Mappings Architecture (SMA), is described and applied to the estimation of human body pose from monocular images. The SMA consists of several specialized forward mapping functions and an inverse map(cid:173) ping function. Each specialized function maps certain domains of the input space (image features) onto the output space (body pose parameters). The key algorithmic problems faced are those of learning the specialized domains and mapping functions in an op(cid:173) timal way, as well as performing inference given inputs and knowl(cid:173) edge of the inverse function. Solutions to these problems employ the EM algorithm and alternating choices of conditional indepen(cid:173) dence assumptions. Performance of the approach is evaluated with synthetic and real video sequences of human motion.
Rómer Rosales, Stan Sclaroff
NIPS2
2001 Deformable Shape Detection and Description via Model-Based Region Grouping
abstract
A method for deformable shape detection and recognition is described. Deformable shape templates are used to partition the image into a globally consistent interpretation, determined in part by the minimum description length principle. Statistical shape models enforce the prior probabilities on global, parametric deformations for each object class. Once trained, the system autonomously segments deformed shapes from the background, while not merging them with adjacent objects or shadows. The formulation can be used to group image regions obtained via any region segmentation algorithm, e.g., texture, color, or motion. The recovered shape models can be used directly in object recognition. Experiments with color imagery are reported.
Stan Sclaroff, Lifeng Liu
IEEE Trans. Pattern Anal. Mach. Intell.1
2001 Corrections to 'Deformable Shape Detection and Description via Model-Based Region Grouping'
Stan Sclaroff, Lifeng Liu
IEEE Trans. Pattern Anal. Mach. Intell.1
2000 Recursive Estimation of Motion and Planar Structure
abstract
A specialized formulation of Azarbayejani and Pentland's (1995) framework for recursive recovery of motion, structure and focal length from feature correspondences tracked through an image sequence is presented. The specialized formulation addresses the case where all tracked points lie on a plane. This planarity constraint reduces the dimension of the original state vector, and consequently the number of feature points needed to estimate the state. Experiments with synthetic data and real imagery illustrate the system performance. The experiments confirm that the specialized formulation provides improved accuracy, stability to observation noise, and rate of convergence in estimation for the case where the tracked points lie on a plane.
Jonathan Alon, Stan Sclaroff
CVPR2
2000 Inferring Body Pose without Tracking Body Parts
abstract
A novel approach for estimating articulated body posture and motion from monocular video sequences is proposed. Human pose is defined as the instantaneous two dimensional configuration (i.e. the projection onto the image plane) of a single articulated body in terms of the position of a predetermined sets of joints. First, statistical segmentation of the human bodies from the background is performed and low-level visual features are found given the segmented body shape. The goal is to be able to map these generally low level visual features to body configurations. The system estimates different mappings, each one with a specific cluster in the visual feature space. Given a set of body motion sequences for training, unsupervised clustering is obtained via the Expectation Maximization algorithm. For each of the clusters, a function is estimated to build the mapping between low-level features to 2D pose. Given new visual features, a mapping from each cluster is performed to yield a set of possible poses. From this set, the system selects the most likely pose given the learned probability distribution and the visual feature of the proposed approach is characterized using real and artificially generated body postures, showing promising results.
Rómer Rosales, Stan Sclaroff
CVPR2
2000 Estimation and Prediction of Evolving Color Distributions for Skin Segmentation under Varying Illumination
abstract
A model approach for real time skin segmentation in video sequences is described. The approach enables reliable skin segmentation despite wide variation in illumination during tracking. An explicit second order Markov model is used to predict evolution of the skin color (HSV) histogram over time. Histograms are dynamically updated based on feedback from the current segmentation and based on predictions of the Markov model. The evolution of the skin color distribution at each frame is parameterized by translation, scaling and rotation in color space. Consequent changes in geometric parameterization of the distribution are propagated by warping and re-sampling the histogram. The parameters of the discrete time dynamic Markov model are estimated using maximum likelihood estimation, and also evolve over time. Quantitative evaluation of the method was conducted on labeled ground-truth video sequences taken from popular movies.
Leonid Sigal, Stan Sclaroff, Vassilis Athitsos
CVPR2
2000 Learning and Synthesizing Human Body Motion and Posture
abstract
A novel approach is presented for estimating human body posture and motion from a video sequence. Human pose is defined as the instantaneous image plane configuration of a single articulated body in terms of the position of a predetermined set of joints. First, statistical segmentation of the human bodies from the background is performed and low-level visual features are found given the segmented body shape. The goal is to be able to map these visual features to body configurations. Given a set of body motion sequences for training, a set of clusters is built in which each has statistically similar configurations. This unsupervised task is done using the expectation maximization algorithm. Then, for each of the clusters, a neural network is trained to build this mapping. Clustering body configurations improves the mapping accuracy. Given new visual features, a mapping from each cluster is performed providing a set of possible poses. From this set, the most likely pose is extracted given the learned probability distribution and the visual feature similarity between hypothesis and input. Performance of the system is characterized using a new set of known body postures, showing promising results.
Rómer Rosales, Stan Sclaroff
FG2
2000 Fast, Reliable Head Tracking under Varying Illumination: An Approach Based on Registration of Texture-Mapped 3D Models
abstract
A technique for 3D head tracking under varying illumination is proposed. The head is modeled as a texture mapped cylinder. Tracking is formulated as an image registration problem in the cylinder's texture map image. The resulting dynamic texture map provides a stabilized view of the face that can be used as input to many existing 2D techniques for face recognition, facial expressions analysis, lip reading, and eye tracking. To solve the registration problem with lighting variation and head motion, the residual registration error is modeled as a linear combination of texture warping templates and orthogonal illumination templates. Fast stable online tracking is achieved via regularized weighted least-squares error minimization. The regularization tends to limit potential ambiguities that arise in the warping and illumination templates. It enables stable tracking over extended sequences. Tracking does not require a precise initial model fit; the system is initialized automatically using a simple 2D face detector. It is assumed that the target is facing the camera in the first frame. The formulation uses texture mapping hardware. The nonoptimized implementation runs at about 15 frames per second on a SGI O2 graphic workstation. Extensive experiments evaluating the effectiveness of the formulation are reported. The sensitivity of the technique to illumination, regularization parameters, errors in the initial positioning, and internal camera parameters are analyzed. Examples and applications of tracking are reported.
Marco La Cascia, Stan Sclaroff, Vassilis Athitsos
IEEE Trans. Pattern Anal. Mach. Intell.2
1999 Fast, Reliable Head Tracking under Varying Illumination
abstract
An improved technique for 3D head tracking under varying illumination conditions is proposed. The head is modeled as a texture mapped cylinder. Tracking is formulated as an image registration problem in the cylinder's texture map image. To solve the registration problem in the presence of lighting variation and head motion, the residual error of registration is modeled as a linear combination of texture warping templates and orthogonal illumination templates. Fast and stable on-line tracking is then achieved via regularized weighted least squares minimization of the registration error. The regularization term tends to limit potential ambiguities that arise in the warping and illumination templates. Tracking does not require a precise initial of the model; the system is initialized automatically using a simple 2D face detector. The only assumption is that the target is facing the camera in the first frame of the sequence. Experiments in tracking are reported.
Marco La Cascia, Stan Sclaroff
CVPR2
1999 Deformable Shape Detection and Description via Model-Based Region
abstract
A method for deformable shape detection and recognition is described. Deformable shape templates are used to partition the image into a globally consistent interpretation, determined in part by the minimum description length principle. Statistical shape models enforce the prior probabilities on global, parametric deformations for each object class. Once trained, the system autonomously segments deformed shapes from the background, while not merging them with adjacent objects or shadows. The formulation can be used to group image regions based on any image homogeneity predicate; e.g., texture, color or motion. The recovered shape models can be used directly in object recognition. Experiments with color imagery are reported.
Lifeng Liu, Stan Sclaroff
CVPR2
1999 3D Trajectory Recovery for Tracking Multiple Objects and Trajectory Guided Recognition of Actions
abstract
A mechanism is proposed that integrates low-level (image processing), mid-level (recursive 3D trajectory estimation), and high level (action recognition) processes. It is assumed that the system observes multiple moving objects via a single, uncalibrated video camera. A novel extended Kalman filter formulation is used in estimating the relative 3D motion trajectories up to a scale factor. The recursive estimation process provides a prediction and error measure that is exploited in higher-level stages of action recognition. Conversely, higher-level mechanisms provide feedback that allows the system to reliable segment and maintain the tracking of moving objects before, during, and after occlusion. The 3D trajectory, occlusion, and segmentation information are utilized in extracting stabilized views of the moving object. Trajectory-guided recognition (TGR) is proposed as a new and efficient method for adaptive classification of action. The TGR approach is demonstrated using "motion history images" that are then recognized via a mixture of Gaussian classifier. The system was tested in recognizing various dynamic human outdoor activities; e.g., running, walking, roller blading, and cycling. Experiments with synthetic data sets are used to evaluate stability of the trajectory estimator with respect to noise.
Rómer Rosales, Stan Sclaroff
CVPR2
1999 Unifying Textual and Visual Cues for Content-Based Image Retrieval on the World Wide Web
abstract
A system is proposed that combines textual and visual statistics in a single index vector for content-based search of a WWW image database. Textual statistics are captured in vector form using latent semantic indexing based on text in the containing HTML document. Visual statistics are captured in vector form using color and orientation histograms. By using an integrated approach, it becomes possible to take advantage of possible statistical couplings between the content of the document (latent semantic content) and the contents of images (visual statistics). The combined approach allows improved performance in conducting content-based search. Search performance experiments are reported for a database containing 350,000 images collected from the WWW.
Stan Sclaroff, Marco La Cascia, Saratendu Sethi, Leonid Taycher
Comput. Vis. Image Underst.1
1998 Active Voodoo Dolls: A Vision Based Input Device for Nonrigid Control
abstract
A vision based technique for nonrigid control is presented that can be used for animation and video game applications. The user grasps a soft, squishable object in front of a camera that can be moved and deformed in order to specify motion. Active Blobs, a nonrigid tracking technique is used to recover the position, rotation and nonrigid deformations of the object. The resulting transformations can be applied to a texture mapped mesh, thus allowing the user to control it interactively. Our use of texture mapping hardware in tracking makes the system responsive enough for interactive animation and video game character control.
John Isidoro, Stan Sclaroff
CA2
1998 Head Tracking via Robust Registration in Texture Map Images
abstract
A novel method for 3D head tracking in the presence of large head rotations and facial expression changes is described. Tracking is formulated in terms of color image registration in the texture map of a 3D surface model. Model appearance is recursively updated via image mosaicking in the texture map as the head orientation varies. The resulting dynamic texture map provides a stabilized view of the face that can be used as input to many existing 2D techniques for face recognition, facial expressions analysis, lip reading, and eye tracking. Parameters are estimated via a robust minimization procedure; this provides robustness to occlusions, wrinkles, shadows and specular highlights. The system was tested on a variety of sequences taken with low quality, uncalibrated video cameras. Experimental results are reported.
Marco La Cascia, John Isidoro, Stan Sclaroff
CVPR3
1998 Active Blobs
abstract
A new region-based approach to nonrigid motion tracking is described. Shape is defined in terms of a deformable triangular mesh that captures object shape plus a color texture map that captures object appearance. Photometric variations are also modeled. Nonrigid shape registration and motion tracking are achieved by posing the problem as an energy-based, robust minimization procedure. The approach provides robustness to occlusions, wrinkles, shadows, and specular highlights. The formulation is tailored to rake advantage of texture mapping hardware available in many workstations, PCs, and game consoles. This enables nonrigid tracking at speeds approaching video rate.
Stan Sclaroff, John Isidoro
ICCV1
1998 Characterization of Neuropathological Shape Deformations
abstract
We present a framework for analyzing the shape deformation of structures within the human brain. A mathematical model is developed describing the deformation of any brain structure whose shape is affected by both gross and detailed physical processes. Using our technique, the total shape deformation is decomposed into analytic modes of variation obtained from finite element modeling, and statistical modes of variation obtained from sample data. Our method is general, and can be applied to many problems where the goal is to separate out important from unimportant shape variation across a class of objects. In this paper, we focus on the analysis of diseases that affect the shape of brain structures. Because the shape of these structures is affected not only by pathology but also by overall brain shape, disease discrimination is difficult. By modeling the brain's elastic properties, we are able to compensate for some of the nonpathological modes of shape variation. This allows us to experimentally characterize modes of variation that are indicative of disease processes. We apply our technique to magnetic resonance images of the brains of individuals with schizophrenia, Alzheimer's disease, and normal-pressure hydrocephalus, as well as to healthy volunteers. Classification results are presented.
Alex Pentland, Stan Sclaroff, Ron Kikinis
IEEE Trans. Pattern Anal. Mach. Intell.3
1997 Deformable prototypes for encoding shape categories in image databases
abstract
An image database search method is described that uses strain energy from prototypes to represent shape categories. Rather than directly comparing a candidate shape with all entries in a database, shapes are ordered in terms of non-rigid deformations that relate them to a small subset of representative prototypes. Shape correspondences are obtained via modal matching, a decomposition for matching, describing, and comparing shapes despite sensor variations and non-rigid deformations. Deformation is decomposed into an ordered basis of orthogonal principal components. This allows selective invariance to in-plane rotation, translation, and scaling, and quasi-invariance to affine deformations. Retrieval accuracy and stability are evaluated in experiments with 2-D image databases.
Stan Sclaroff
Pattern Recognit.1
1996 Photobook: Content-based manipulation of image databases
Alex Pentland, Rosalind W. Picard, Stan Sclaroff
Int. J. Comput. Vis.3
1995 Modal Matching for Correspondence and Recognition
abstract
Modal matching is a new method for establishing correspondences and computing canonical descriptions. The method is based on the idea of describing objects in terms of generalized symmetries, as defined by each object's eigenmodes. The resulting modal description is used for object recognition and categorization, where shape similarities are expressed as the amounts of modal deformation energy needed to align the two objects. In general, modes provide a global-to-local ordering of shape deformation and thus allow for selecting which types of deformations are used in object alignment and comparison. In contrast to previous techniques, which required correspondence to be computed with an initial or prototype shape, modal matching utilizes a new type of finite element formulation that allows for an object's eigenmodes to be computed directly from available image information. This improved formulation provides greater generality and accuracy, and is applicable to data of any dimensionality. Correspondence results with 2D contour and point feature data are shown, and recognition experiments with 2D images of hand tools and airplanes are described.>
Stan Sclaroff, Alex Pentland
IEEE Trans. Pattern Anal. Mach. Intell.1
1994 Visually guided animation
abstract
We are interested in being able to take classic film characters, or video of current-day personalities, and produce computer models and animations of them by automatic analysis of the video or film footage. In this paper we survey our progress toward producing such automatic modeling and animation systems.>
Alex Pentland, Trevor Darrell, Irfan A. Essa, Ali Azarbayejani, Stan Sclaroff
CA5
1994 Object representation for object recognition
abstract
This paper discusses some representation issues and challenges involved in object recognition. It is intended as a step toward assessing current object representation schemes and proposing design and evaluation criteria for future ones.>
Jean Ponce, Ruzena Bajcsy, Dimitris N. Metaxas, Thomas O. Binford, David A. Forsyth, Martial Hebert, Katsushi Ikeuchi, Avinash C. Kak, Linda G. Shapiro, Stan Sclaroff, Alex Pentland, George C. Stockman
CVPR10
1993 A modal framework for correspondence and description
abstract
The authors describe a framework for establishing correspondence, computing canonical descriptions, and recognizing objects that is based on the idea of describing objects by their generalized symmetries, as defined by the object's free vibration modes. A technique given by A. Pentland and S. Scarloff (1991) described objects in terms of the modes of some prototype shape. In contrast, this new method computes the object's modes directly from available image information. This results in greater generality and accuracy, and is applicable to data of any dimensionality. For the purposes of illustration, a detailed mathematical formulation of the method is given for 2-D problems, and it is demonstrated on gray-scale image and contour data.>
Stan Sclaroff, Alex Pentland
ICCV1
1992 A Unified Approach for Physical and Geometric Modeling for Graphics and Animation
abstract
Abstract We present a unified approach for geometric and physical modeling using implicit functions, for application to graphics and animation. This method extends previously proposed techniques, and allows the standard finite element method to be directly combined with geometric modeling, resulting in quick calculation of an object's mass and stiffness matrices, and its vibration modes and frequencies. Because the approach is based on an implicitfunction representation, it allows very fast collision detection and characterization. Examples of complex physical and geometric modeling are presented.
Irfan A. Essa, Stan Sclaroff, Alex Pentland
Comput. Graph. Forum2
1991 Closed-form solutions for physically-based shape modeling and recognition
abstract
An efficient, physically based solution for recovering a 3-D solid model from collections of 3-D surface measurements is presented. Given a sufficient number of independent measurements, the solution is overconstrained and unique except for rotational symmetries. A physically based object recognition method that allows simple, closed-form comparisons of recovered 3-D solid models is given. The performance of these methods is evaluated using both synthetic and real laser rangefinder data.>
Stan Sclaroff, Alex Pentland
CVPR1
1991 Generalized implicit functions for computer graphics
abstract
We describe a method of generalizing implicit functions by use of modal deformations and displacement maps. Modal deformations, also known as free vibration modes, are used to describe the overall shape of a solid, while displacement maps provide local and fine surface detail by offsetting the surface of the solid along its surface normals. The advantage of this approach to geometric description is that collision detection and dynamic simulation become simple and inexpensive even for complex shapes. In addition, we outline an efficient method for fitting such models to three dimensional point data.
Stan Sclaroff, Alex Pentland
SIGGRAPH1
1991 Closed-Form Solutions for Physically Based Shape Modeling and Recognition
abstract
The authors present a closed-form, physically based solution for recovering a three-dimensional (3-D) solid model from collections of 3-D surface measurements. Given a sufficient number of independent measurements, the solution is overconstrained and unique except for rotational symmetries. The proposed approach is based on the finite element method (FEM) and parametric solid modeling using implicit functions. This approach provides both the convenience of parametric modeling and the expressiveness of the physically based mesh formulation and, in addition, can provide great accuracy at physical simulation. A physically based object-recognition method that allows simple, closed-form comparisons of recovered 3-D solid models is presented. The performance of these methods is evaluated using both synthetic range data with various signal-to-noise ratios and using laser rangefinder data.>
Alex Pentland, Stan Sclaroff
IEEE Trans. Pattern Anal. Mach. Intell.2
1990 Segmentation by minimal description
abstract
The authors formulate the segmentation task as a search for a set of descriptions which minimally encodes a scene. A novel framework for cooperative robust estimation is used to estimate descriptions that locally provide the most savings in encoding an image. A modified Hopfield-Tank networks finds the subset of these descriptions which best describes an entire scene, accounting for occlusion and transparent overlap among individual descriptions. Using a part-based 3-D shape model the authors have implemented a system that is able to successfully segment images into their constituent structure.>
Trevor Darrell, Stan Sclaroff, Alex Pentland
ICCV2