EDBT 2026 Demo / reviewers in the wild / expert
Greg Mori
dblp:m/GregMori
· DBLP profile ↗
119ranked-venue papers
7as first author
13since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 110 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 73 · 5 first-author · 6 since 2021Systems, architecture and hardware · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Embodied Human Activity RecognitionabstractWe study how to utilize the mobility of an embodied agent to improve its ability to recognize human activities. We introduce the embodied human activity recognition problem, where an agent moves in a 3D environment to recognize the category of ongoing human activities. The agent must make movement decisions based on its egocentric observations acquired up to the current time, with the goal of choosing movements to obtain new views that lead to accurate human activity recognition. Towards this goal, we propose a reinforcement learning approach that learns a policy controlling the agent’s movements over time. We evaluate our approach with two realistic human activity datasets. Results show that our approach can learn to move effectively to achieve high performance in recognizing human activities. Sha Hu 0003, Greg Mori |
WACV | 3 |
| 2022 | TD-GEN: Graph Generation Using Tree DecompositionabstractWe propose TD-GEN, a graph generation framework based on tree decomposition, and introduce a reduced upper bound on the maximum number of decisions needed for graph generation. The framework includes a permutation invariant tree generation model which forms the backbone of graph generation. Tree nodes are supernodes, each representing a cluster of nodes in the graph. Graph nodes and edges are incrementally generated inside the clusters by traversing the tree supernodes, respecting the structure of the tree decomposition, and following node sharing decisions between the clusters. Further, we discuss the shortcomings of the standard evaluation criteria based on statistical properties of the generated graphs. We propose to compare the generalizability of models based on expected likelihood. Empirical results on a variety of standard graph generation datasets demonstrate the superior performance of our method. Hamed Shirzad, Hossein Hajimirsadeghi, Amir H. Abdi, Greg Mori |
AISTATS | 4 |
| 2022 | Rethinking Learning Approaches for Long-Term Action Anticipation
Megha Nawhal, Akash Abdu Jyothi, Greg Mori |
ECCV (34) | 3 |
| 2022 | Filtered-CoPhy: Unsupervised Learning of Counterfactual Physics in Pixel Space
Steeven Janny, Fabien Baradel, Natalia Neverova, Madiha Nadri Wolf, Greg Mori, Christian Wolf 0001 |
ICLR | 5 |
| 2022 | RankSim: Ranking Similarity Regularization for Deep Imbalanced RegressionabstractData imbalance, in which a plurality of the data samples come from a small proportion of labels, poses a challenge in training deep neural networks. Unlike classification, in regression the labels are continuous, potentially boundless, and form a natural ordering. These distinct features of regression call for new techniques that leverage the additional information encoded in label-space relationships. This paper presents the RankSim (ranking similarity) regularizer for deep imbalanced regression, which encodes an inductive bias that samples that are closer in label space should also be closer in feature space. In contrast to recent distribution smoothing based approaches, RankSim captures both nearby and distant relationships: for a given data sample, RankSim encourages the sorted list of its neighbors in label space to match the sorted list of its neighbors in feature space. RankSim is complementary to conventional imbalanced learning techniques, including re-weighting, two-stage training, and distribution smoothing, and lifts the state-of-the-art performance on three imbalanced regression benchmarks: IMDB-WIKI-DIR, AgeDB-DIR, and STS-B-DIR. Greg Mori, Frederick Tung |
ICML | 2 |
| 2022 | Monotonicity regularization: Improved penalties and novel applications to disentangled representation learning and robust classificationabstractWe study settings where gradient penalties are used alongside risk minimization with the goal of obtaining predictors satisfying different notions of monotonicity. Specifically, we present two sets of contributions. In the first part of the paper, we show that different choices of penalties define the regions of the input space where the property is observed. As such, previous methods result in models that are monotonic only in a small volume of the input space. We thus propose an approach that uses mixtures of training instances and random points to populate the space and enforce the penalty in a much larger region. As a second set of contributions, we introduce regularization strategies that enforce other notions of monotonicity in different settings. In this case, we consider applications, such as image classification and generative modeling, where monotonicity is not a hard constraint but can help improve some aspects of the model. Namely, we show that inducing monotonicity can be beneficial in applications such as: (1) allowing for controllable data generation, (2) defining strategies to detect anomalous data, and (3) generating explanations for predictions. Our proposed approaches do not introduce relevant computational overhead while leading to efficient procedures that provide extra benefits over baseline models. Joao Monteiro, Mohamed Osama Ahmed, Hossein Hajimirsadeghi, Greg Mori |
UAI | 4 |
| 2021 | Variational Selective Autoencoder: Learning from Partially-Observed Heterogeneous DataabstractLearning from heterogeneous data poses challenges such as combining data from various sources and of different types. Meanwhile, heterogeneous data are often associated with missingness in real-world applications due to heterogeneity and noise of input sources. In this work, we propose the variational selective autoencoder (VSAE), a general framework to learn representations from partially-observed heterogeneous data. VSAE learns the latent dependencies in heterogeneous data by modeling the joint distribution of observed data, unobserved data, and the imputation mask which represents how the data are missing. It results in a unified model for various downstream tasks including data generation and imputation. Evaluation on both low-dimensional and high-dimensional heterogeneous datasets for these two tasks shows improvement over state-of-the-art models. Hossein Hajimirsadeghi, Jiawei He 0001, Thibaut Durand, Greg Mori |
AISTATS | 5 |
| 2021 | Learning to Predict Convolutional Filters with Guidance for Conditional Image Generation
Lei Chen 0023, Mengyao Zhai, Greg Mori |
BMVC | 3 |
| 2021 | MUSE: Feature Self-Distillation with Mutual Information and Self-Information
Ye Yu 0003, Gaurav Mittal, Greg Mori |
BMVC | 4 |
| 2021 | Learning Discriminative Prototypes With Dynamic Time WarpingabstractDynamic Time Warping (DTW) is widely used for temporal data processing. However, existing methods can neither learn the discriminative prototypes of different classes nor exploit such prototypes for further analysis. We propose Discriminative Prototype DTW (DP-DTW), a novel method to learn class-specific discriminative prototypes for temporal recognition tasks. DP-DTW shows superior performance compared to conventional DTWs on time series classification benchmarks1. Combined with end-to-end deep learning, DP-DTW can handle challenging weakly supervised action segmentation problems and achieves state of the art results on standard benchmarks. Moreover, detailed reasoning on the input video is enabled by the learned action prototypes. Specifically, an action-based video summarization can be obtained by aligning the input sequence with action prototypes. Xiaobin Chang, Frederick Tung, Greg Mori |
CVPR | 3 |
| 2021 | Hyper-LifelongGAN: Scalable Lifelong Learning for Image Conditioned GenerationabstractDeep neural networks are susceptible to catastrophic forgetting: when encountering a new task, they can only remember the new task and fail to preserve its ability to accomplish previously learned tasks. In this paper, we study the problem of lifelong learning for generative models and propose a novel and generic continual learning framework Hyper-LifelongGAN which is more scalable compared with state-of-the-art approaches. Given a sequence of tasks, the conventional convolutional filters are factorized into the dynamic base filters which are generated using task specific filter generators, and deterministic weight matrix which linearly combines the base filters and is shared across different tasks. Moreover, the shared weight matrix is multiplied by task specific coefficients to introduce more flexibility in combining task specific base filters differently for different tasks. Attributed to the novel architecture, the proposed method can preserve or even improve the generation quality at a low cost of parameters. We validate Hyper-LifelongGAN on diverse image-conditioned generation tasks, extensive ablation studies and comparisons with state-of-the-art models are carried out to show that the proposed approach can address catastrophic forgetting effectively. Mengyao Zhai, Lei Chen 0023, Greg Mori |
CVPR | 3 |
| 2021 | Neural fidelity warping for efficient robot morphology designabstractWe consider the problem of optimizing a robot morphology to achieve the best performance for a target task, under computational resource limitations. The evaluation process for each morphological design involves learning a controller for the design, which can consume substantial time and computational resources. To address the challenge of expensive robot morphology evaluation, we present a continuous multi-fidelity Bayesian Optimization framework that efficiently utilizes computational resources via low-fidelity evaluations. We identify the problem of non-stationarity over fidelity space. Our proposed fidelity warping mechanism can learn representations of learning epochs and tasks to model non-stationary covariances between continuous fidelity evaluations which prove challenging for off-the-shelf stationary kernels. Various experiments demonstrate that our method can utilize the low-fidelity evaluations to efficiently search for the optimal robot morphology, outperforming state-of-the-art methods. Sha Hu 0003, Zeshi Yang, Greg Mori |
ICRA | 3 |
| 2021 | Continuous Latent Process FlowsabstractPartial observations of continuous time-series dynamics at arbitrary time stamps exist in many disciplines. Fitting this type of data using statistical models with continuous dynamics is not only promising at an intuitive level but also has practical benefits, including the ability to generate continuous trajectories and to perform inference on previously unseen time stamps. Despite exciting progress in this area, the existing models still face challenges in terms of their representational power and the quality of their variational approximations. We tackle these challenges with continuous latent process flows (CLPF), a principled architecture decoding continuous latent processes into continuous observable processes using a time-dependent normalizing flow driven by a stochastic differential equation. To optimize our model using maximum likelihood, we propose a novel piecewise construction of a variational posterior process and derive the corresponding variational lower bound using trajectory re-weighting. Our ablation studies demonstrate the effectiveness of our contributions in various inference tasks on irregular time grids. Comparisons to state-of-the-art baselines show our model's favourable performance on both synthetic and real-world time-series data. Ruizhi Deng, Marcus A. Brubaker, Greg Mori, Andreas M. Lehrmann |
NeurIPS | 3 |
| 2020 | House-GAN: Relational Generative Adversarial Networks for Graph-Constrained House Layout Generation
Nelson Nauata, Kai-Hung Chang, Chin-Yi Cheng, Greg Mori, Yasutaka Furukawa |
ECCV (1) | 4 |
| 2020 | Generating Videos of Zero-Shot Compositions of Actions and Objects
Megha Nawhal, Mengyao Zhai, Andreas M. Lehrmann, Leonid Sigal, Greg Mori |
ECCV (12) | 5 |
| 2020 | Piggyback GAN: Efficient Lifelong Learning for Image Conditioned Generation
Mengyao Zhai, Lei Chen 0023, Jiawei He 0001, Megha Nawhal, Frederick Tung, Greg Mori |
ECCV (21) | 6 |
| 2020 | CoPhy: Counterfactual Learning of Physical Dynamics
Fabien Baradel, Natalia Neverova, Julien Mille, Greg Mori, Christian Wolf 0001 |
ICLR | 4 |
| 2020 | Relational Graph Learning for Crowd NavigationabstractWe present a relational graph learning approach for robotic crowd navigation using model-based deep reinforcement learning that plans actions by looking into the future. Our approach reasons about the relations between all agents based on their latent features and uses a Graph Convolutional Network to encode higher-order interactions in each agent's state representation, which is subsequently leveraged for state prediction and value estimation. The ability to predict human motion allows us to perform multi-step lookahead planning, taking into account the temporal evolution of human crowds. We evaluate our approach against a state-of-the-art baseline for crowd navigation and ablations of our model to demonstrate that navigation with our approach is more efficient, results in fewer collisions, and avoids failure cases involving oscillatory and freezing behaviors. Changan Chen, Sha Hu 0003, Payam Nikdel, Greg Mori, Manolis Savva |
IROS | 4 |
| 2020 | Modeling Continuous Stochastic Processes with Dynamic Normalizing FlowsabstractNormalizing flows transform a simple base distribution into a complex target distribution and have proved to be powerful models for data generation and density estimation. In this work, we propose a novel type of normalizing flow driven by a differential deformation of the continuous-time Wiener process. As a result, we obtain a rich time series model whose observable process inherits many of the appealing properties of its base process, such as efficient computation of likelihoods and marginals. Furthermore, our continuous treatment provides a natural framework for irregular time series with an independent arrival process, including straightforward interpolation. We illustrate the desirable properties of the proposed model on popular stochastic processes and demonstrate its superior flexibility to variational RNN and latent ODE baselines in a series of experiments on synthetic and real-world data. Ruizhi Deng, Bo Chang 0002, Marcus A. Brubaker, Greg Mori, Andreas M. Lehrmann |
NeurIPS | 4 |
| 2020 | Adapting Grad-CAM for Embedding NetworksabstractThe gradient-weighted class activation mapping (Grad-CAM) method can faithfully highlight important regions in images for deep model prediction in image classification, image captioning and many other tasks. It uses the gradients in back-propagation as weights (grad-weights) to explain network decisions. However, applying Grad-CAM to embedding networks raises significant challenges because embedding networks are trained by millions of dynamically paired examples (e.g. triplets). To overcome these challenges, we propose an adaptation of the Grad-CAM method for embedding networks. First, we aggregate grad-weights from multiple training examples to improve the stability of Grad-CAM. Then, we develop an efficient weight-transfer method to explain decisions for any image without back-propagation. We extensively validate the method on the standard CUB200 dataset in which our method produces more accurate visual attention than the original Grad-CAM method. We also apply the method to a house price estimation application using images. The method produces convincing qualitative results, showcasing the practicality of our approach. Lei Chen 0023, Hossein Hajimirsadeghi, Greg Mori |
WACV | 4 |
| 2020 | Guest Editorial: Special Issue on ACCV 2018
C. V. Jawahar, Hongdong Li, Greg Mori, Konrad Schindler |
Int. J. Comput. Vis. | 3 |
| 2020 | Structured Label Inference for Visual UnderstandingabstractVisual data such as images and videos contain a rich source of structured semantic labels as well as a wide range of interacting components. Visual content could be assigned with fine-grained labels describing major components, coarse-grained labels depicting high level abstractions, or a set of labels revealing attributes. Such categorization over different, interacting layers of labels evinces the potential for a graph-based encoding of label information. In this paper, we exploit this rich structure for performing graph-based inference in label space for a number of tasks: multi-label image and video classification and action detection in untrimmed videos. We consider the use of the Bidirectional Inference Neural Network (BINN) and Structured Inference Neural Network (SINN) for performing graph-based inference in label space and propose a Long Short-Term Memory (LSTM) based extension for exploiting activity progression on untrimmed videos. The methods were evaluated on (i) the Animal with Attributes (AwA), Scene Understanding (SUN) and NUS-WIDE datasets for multi-label image classification, (ii) the first two releases of the YouTube-8M large scale dataset for multi-label video classification, and (iii) the THUMOS'14 and MultiTHUMOS video datasets for action detection. Our results demonstrate the effectiveness of structured label inference in these challenging tasks, achieving significant improvements against baselines. Nelson Nauata, Hexiang Hu, Guang-Tong Zhou, Zhiwei Deng, Zicheng Liao, Greg Mori |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | Deep Neural Network Compression by In-Parallel Pruning-QuantizationabstractDeep neural networks enable state-of-the-art accuracy on visual recognition tasks such as image classification and object detection. However, modern networks contain millions of learned connections, and the current trend is towards deeper and more densely connected architectures. This poses a challenge to the deployment of state-of-the-art networks on resource-constrained systems, such as smartphones or mobile robots. In general, a more efficient utilization of computation resources would assist in deployment scenarios from embedded platforms to computing clusters running ensembles of networks. In this paper, we propose a deep network compression algorithm that performs weight pruning and quantization jointly, and in parallel with fine-tuning. Our approach takes advantage of the complementary nature of pruning and quantization and recovers from premature pruning errors, which is not possible with two-stage approaches. In experiments on ImageNet, CLIP-Q (Compression Learning by In-Parallel Pruning-Quantization) improves the state-of-the-art in network compression on AlexNet, VGGNet, GoogLeNet, and ResNet. We additionally demonstrate that CLIP-Q is complementary to efficient network architecture design by compressing MobileNet and ShuffleNet, and that CLIP-Q generalizes beyond convolutional networks by compressing a memory network for visual question answering. Frederick Tung, Greg Mori |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Learning a Deep ConvNet for Multi-Label Classification With Partial LabelsabstractDeep ConvNets have shown great performance for single-label image classification (e.g. ImageNet), but it is necessary to move beyond the single-label classification task because pictures of everyday life are inherently multi-label. Multi-label classification is a more difficult task than single-label classification because both the input images and output label spaces are more complex. Furthermore, collecting clean multi-label annotations is more difficult to scale-up than single-label annotations. To reduce the annotation cost, we propose to train a model with partial labels i.e. only some labels are known per image. We first empirically compare different labeling strategies to show the potential for using partial labels on multi-label datasets. Then to learn with partial labels, we introduce a new classification loss that exploits the proportion of known labels per example. Our approach allows the use of the same training settings as when learning with all the annotations. We further explore several curriculum learning based strategies to predict missing labels. Experiments are performed on three large-scale multi-label datasets: MS COCO, NUS-WIDE and Open Images. Thibaut Durand, Nazanin Mehrasa, Greg Mori |
CVPR | 3 |
| 2019 | A Variational Auto-Encoder Model for Stochastic Point ProcessesabstractWe propose a novel probabilistic generative model for action sequences. The model is termed the Action Point Process VAE (APP-VAE), a variational auto-encoder that can capture the distribution over the times and categories of action sequences. Modeling the variety of possible action sequences is a challenge, which we show can be addressed via the APP-VAE's use of latent representations and non-linear functions to parameterize distributions over which event is likely to occur next in a sequence and at what time. We empirically validate the efficacy of APP-VAE for modeling action sequences on the MultiTHUMOS and Breakfast datasets. Nazanin Mehrasa, Akash Abdu Jyothi, Thibaut Durand, Jiawei He 0001, Leonid Sigal, Greg Mori |
CVPR | 6 |
| 2019 | VideoWhiz: Non-Linear Interactive Overviews for Recipe Videos
Megha Nawhal, Jacqueline B. Lang, Greg Mori, Parmit K. Chilana |
Graphics Interface | 3 |
| 2019 | LayoutVAE: Stochastic Scene Layout Generation From a Label SetabstractRecently there is an increasing interest in scene generation within the research community. However, models used for generating scene layouts from textual description largely ignore plausible visual variations within the structure dictated by the text. We propose LayoutVAE, a variational autoencoder based framework for generating stochastic scene layouts. LayoutVAE is a versatile modeling framework that allows for generating full image layouts given a label set, or per label layouts for an existing image given a new label. In addition, it is also capable of detecting unusual layouts, potentially providing a way to evaluate layout generation problem. Extensive experiments on MNIST-Layouts and challenging COCO 2017 Panoptic dataset verifies the effectiveness of our proposed framework. Akash Abdu Jyothi, Thibaut Durand, Jiawei He 0001, Leonid Sigal, Greg Mori |
ICCV | 5 |
| 2019 | Similarity-Preserving Knowledge DistillationabstractKnowledge distillation is a widely applicable technique for training a student neural network under the guidance of a trained teacher network. For example, in neural network compression, a high-capacity teacher is distilled to train a compact student; in privileged learning, a teacher trained with privileged data is distilled to train a student without access to that data. The distillation loss determines how a teacher's knowledge is captured and transferred to the student. In this paper, we propose a new form of knowledge distillation loss that is inspired by the observation that semantically similar inputs tend to elicit similar activation patterns in a trained network. Similarity-preserving knowledge distillation guides the training of a student network such that input pairs that produce similar (dissimilar) activations in the teacher network produce similar (dissimilar) activations in the student network. In contrast to previous distillation methods, the student is not required to mimic the representation space of the teacher, but rather to preserve the pairwise similarities in its own representation space. Experiments on three public datasets demonstrate the potential of our approach. Frederick Tung, Greg Mori |
ICCV | 2 |
| 2019 | Lifelong GAN: Continual Learning for Conditional Image GenerationabstractLifelong learning is challenging for deep neural networks due to their susceptibility to catastrophic forgetting. Catastrophic forgetting occurs when a trained network is not able to maintain its ability to accomplish previously learned tasks when it is trained to perform new tasks. We study the problem of lifelong learning for generative models, extending a trained network to new conditional generation tasks without forgetting previous tasks, while assuming access to the training data for the current task only. In contrast to state-of-the-art memory replay based approaches which are limited to label-conditioned image generation tasks, a more generic framework for continual learning of generative models under different conditional image generation settings is proposed in this paper. Lifelong GAN employs knowledge distillation to transfer learned knowledge from previous networks to the new network. This makes it possible to perform image-conditioned generation tasks in a lifelong learning setting. We validate Lifelong GAN for both image-conditioned and label-conditioned generation tasks, and provide qualitative and quantitative results to show the generality and effectiveness of our method. Mengyao Zhai, Lei Chen 0023, Frederick Tung, Jiawei He 0001, Megha Nawhal, Greg Mori |
ICCV | 6 |
| 2019 | Variational Autoencoders with Jointly Optimized Latent Dependency Structure
Jiawei He 0001, Joseph Marino, Greg Mori, Andreas M. Lehrmann |
ICLR (Poster) | 4 |
| 2018 | Adaptive Appearance Rendering
Mengyao Zhai, Ruizhi Deng, Lei Chen 0023, Zhiwei Deng, Greg Mori |
BMVC | 6 |
| 2018 | CLIP-Q: Deep Network Compression Learning by In-Parallel Pruning-QuantizationabstractDeep neural networks enable state-of-the-art accuracy on visual recognition tasks such as image classification and object detection. However, modern deep networks contain millions of learned weights; a more efficient utilization of computation resources would assist in a variety of deployment scenarios, from embedded platforms with resource constraints to computing clusters running ensembles of networks. In this paper, we combine network pruning and weight quantization in a single learning framework that performs pruning and quantization jointly, and in parallel with fine-tuning. This allows us to take advantage of the complementary nature of pruning and quantization and to recover from premature pruning errors, which is not possible with current two-stage approaches. Our proposed CLIP-Q method (Compression Learning by In-Parallel Pruning-Quantization) compresses AlexNet by 51-fold, GoogLeNet by 10-fold, and ResNet-50 by 15-fold, while preserving the uncompressed network accuracies on ImageNet. Frederick Tung, Greg Mori |
CVPR | 2 |
| 2018 | Object Level Visual Reasoning in Videos
Fabien Baradel, Natalia Neverova, Christian Wolf 0001, Julien Mille, Greg Mori |
ECCV (13) | 5 |
| 2018 | Constraint-Aware Deep Neural Network Compression
Changan Chen, Frederick Tung, Naveen Vedula, Greg Mori |
ECCV (8) | 4 |
| 2018 | Probabilistic Video Generation Using Holistic Attribute Control
Jiawei He 0001, Andreas M. Lehrmann, Joseph Marino, Greg Mori, Leonid Sigal |
ECCV (5) | 4 |
| 2018 | Hierarchical Relational Networks for Group Activity Recognition and Retrieval
Mostafa S. Ibrahim, Greg Mori |
ECCV (3) | 2 |
| 2018 | Sparsely Aggregated Convolutional Networks
Ligeng Zhu, Ruizhi Deng, Michael Maire, Zhiwei Deng, Greg Mori |
ECCV (12) | 5 |
| 2018 | Probabilistic Neural Programmed Networks for Scene GenerationabstractIn this paper we address the text to scene image generation problem. Generative models that capture the variability in complicated scenes containing rich semantics is a grand goal of image generation. Complicated scene images contain rich visual elements, compositional visual concepts, and complicated relations between objects. Generative models, as an analysis-by-synthesis process, should encompass the following three core components: 1) the generation process that composes the scene; 2) what are the primitive visual elements and how are they composed; 3) the rendering of abstract concepts into their pixel-level realizations. We propose PNP-Net, a variational auto-encoder framework that addresses these three challenges: it flexibly composes images with a dynamic network structure, learns a set of distribution transformers that can compose distributions based on semantics, and decodes samples from these distributions into realistic images. Zhiwei Deng, Yifang Fu, Greg Mori |
NeurIPS | 4 |
| 2018 | Generic Tubelet Proposals for Action LocalizationabstractWe develop a novel framework for action localization in videos. We propose the Tube Proposal Network (TPN), which can generate generic, class-independent, video-level tubelet proposals in videos. The generated tubelet proposals can be utilized in various video analysis tasks, including recognizing and localizing actions in videos. In particular, we integrate these generic tubelet proposals into a unified temporal deep network for action classification. Compared with other methods, our generic tubelet proposal method is accurate, general, and is fully differentiable under a smoothL1 loss function. We demonstrate the performance of our algorithm on the standard UCF-Sports, J-HMDB21, and UCF-101 datasets. Our class-independent TPN outperforms other tubelet generation methods, and our unified temporal deep network achieves state-of-the-art localization results on all three datasets. Jiawei He 0001, Zhiwei Deng, Mostafa S. Ibrahim, Greg Mori |
WACV | 4 |
| 2018 | Scaling Human-Object Interaction Recognition Through Zero-Shot LearningabstractRecognizing human object interactions (HOI) is an important part of distinguishing the rich variety of human action in the visual world. While recent progress has been made in improving HOI recognition in the fully supervised setting, the space of possible human-object interactions is large and it is impractical to obtain labeled training data for all interactions of interest. In this work, we tackle the challenge of scaling HOI recognition to the long tail of categories through a zero-shot learning approach. We introduce a factorized model for HOI detection that disentangles reasoning on verbs and objects, and at test-time can therefore produce detections for novel verb-object pairs. We present experiments on the recently introduced large-scale HICODET dataset, and show that our model is able to both perform comparably to state-of-the-art in fully-supervised HOI detection, while simultaneously achieving effective zeroshot detection of new HOI categories. Liyue Shen, Serena Yeung-Levy, Judy Hoffman, Greg Mori, Li Fei-Fei 0001 |
WACV | 4 |
| 2018 | Guest Editorial
Lamberto Ballan, Shih-Fu Chang, Gang Hua 0001, Thomas Mensink, Greg Mori, Rahul Sukthankar |
Comput. Vis. Image Underst. | 5 |
| 2018 | Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos
Serena Yeung-Levy, Olga Russakovsky, Ning Jin 0002, Mykhaylo Andriluka, Greg Mori, Li Fei-Fei 0001 |
Int. J. Comput. Vis. | 5 |
| 2017 | Fine-Pruning: Joint Fine-Tuning and Compression of a Convolutional Network with Bayesian Optimization
Frederick Tung, Srikanth Muralidharan, Greg Mori |
BMVC | 3 |
| 2017 | Factorized Variational Autoencoders for Modeling Audience Reactions to MoviesabstractMatrix and tensor factorization methods are often used for finding underlying low-dimensional patterns from noisy data. In this paper, we study non-linear tensor factorization methods based on deep variational autoencoders. Our approach is well-suited for settings where the relationship between the latent representation to be learned and the raw data representation is highly complex. We apply our approach to a large dataset of facial expressions of movie-watching audiences (over 16 million faces). Our experiments show that compared to conventional linear factorization methods, our method achieves better reconstruction of the data, and further discovers interpretable latent factors. Zhiwei Deng, Rajitha Navarathna, Peter Carr 0001, Stephan Mandt, Yisong Yue, Iain A. Matthews, Greg Mori |
CVPR | 7 |
| 2017 | Learning to Learn from Noisy Web VideosabstractUnderstanding the simultaneously very diverse and intricately fine-grained set of possible human actions is a critical open problem in computer vision. Manually labeling training videos is feasible for some action classes but doesnt scale to the full long-tailed distribution of actions. A promising way to address this is to leverage noisy data from web queries to learn new actions, using semi-supervised or webly-supervised approaches. However, these methods typically do not learn domain-specific knowledge, or rely on iterative hand-tuned data labeling policies. In this work, we instead propose a reinforcement learning-based formulation for selecting the right examples for training a classifier from noisy web search results. Our method uses Q-learning to learn a data labeling policy on a small labeled training dataset, and then uses this to automatically label noisy web data for new visual concepts. Experiments on the challenging Sports-1M action recognition benchmark as well as on additional fine-grained and newly emerging action classes demonstrate that our method is able to learn good labeling policies for noisy data and use this to learn accurate visual concept classifiers. Serena Yeung-Levy, Vignesh Ramanathan, Olga Russakovsky, Liyue Shen, Greg Mori, Li Fei-Fei 0001 |
CVPR | 5 |
| 2017 | Multi-Instance Classification by Max-Margin Training of Cardinality-Based Markov NetworksabstractWe propose a probabilistic graphical framework for multi-instance learning (MIL) based on Markov networks. This framework can deal with different levels of labeling ambiguity (i.e., the portion of positive instances in a bag) in weakly supervised data by parameterizing cardinality potential functions. Consequently, it can be used to encode different cardinality-based multi-instance assumptions, ranging from the standard MIL assumption to more general assumptions. In addition, this framework can be efficiently used for both binary and multiclass classification. To this end, an efficient inference algorithm and a discriminative latent max-margin learning algorithm are introduced to train and test the proposed multi-instance Markov network models. We evaluate the performance of the proposed framework on binary and multi-class MIL benchmark datasets as well as two challenging computer vision tasks: cyclist helmet recognition and human group activity recognition. Experimental results verify that encoding the degree of ambiguity in data can improve classification performance. Hossein Hajimirsadeghi, Greg Mori |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Structure Inference Machines: Recurrent Neural Networks for Analyzing Relations in Group Activity RecognitionabstractRich semantic relations are important in a variety of visual recognition problems. As a concrete example, group activity recognition involves the interactions and relative spatial relations of a set of people in a scene. State of the art recognition methods center on deep learning approaches for training highly effective, complex classifiers for interpreting images. However, bridging the relatively low-level concepts output by these methods to interpret higher-level compositional scenes remains a challenge. Graphical models are a standard tool for this task. In this paper, we propose a method to integrate graphical models and deep neural networks into a joint framework. Instead of using a traditional inference method, we use a sequential inference modeled by a recurrent neural network. Beyond this, the appropriate structure for inference can be learned by imposing gates on edges between nodes. Empirical results on group activity recognition demonstrate the potential of this model to handle highly structured learning tasks. Zhiwei Deng, Arash Vahdat, Hexiang Hu, Greg Mori |
CVPR | 4 |
| 2016 | Learning Structured Inference Neural Networks with Label RelationsabstractImages of scenes have various objects as well as abundant attributes, and diverse levels of visual categorization are possible. A natural image could be assigned with fine-grained labels that describe major components, coarse-grained labels that depict high level abstraction, or a set of labels that reveal attributes. Such categorization at different concept layers can be modeled with label graphs encoding label information. In this paper, we exploit this rich information with a state-of-art deep learning framework, and propose a generic structured model that leverages diverse label relations to improve image classification performance. Our approach employs a novel stacked label prediction neural network, capturing both inter-level and intra-level label semantics. We evaluate our method on benchmark image datasets, and empirical results illustrate the efficacy of our model. Hexiang Hu, Guang-Tong Zhou, Zhiwei Deng, Zicheng Liao, Greg Mori |
CVPR | 5 |
| 2016 | A Hierarchical Deep Temporal Model for Group Activity RecognitionabstractIn group activity recognition, the temporal dynamics of the whole activity can be inferred based on the dynamics of the individual people representing the activity. We build a deep model to capture these dynamics based on LSTM (long short-term memory) models. To make use of these observations, we present a 2-stage deep temporal model for the group activity recognition problem. In our model, a LSTM model is designed to represent action dynamics of individual people in a sequence and another LSTM model is designed to aggregate person-level information for whole activity understanding. We evaluate our model over two datasets: the Collective Activity Dataset and a new volleyball dataset. Experimental results demonstrate that our proposed model improves group activity recognition performance compared to baseline methods. Mostafa S. Ibrahim, Srikanth Muralidharan, Zhiwei Deng, Arash Vahdat, Greg Mori |
CVPR | 5 |
| 2016 | End-to-End Learning of Action Detection from Frame Glimpses in VideosabstractIn this work we introduce a fully end-to-end approach for action detection in videos that learns to directly predict the temporal bounds of actions. Our intuition is that the process of detecting actions is naturally one of observation and refinement: observing moments in video, and refining hypotheses about when an action is occurring. Based on this insight, we formulate our model as a recurrent neural network-based agent that interacts with a video over time. The agent observes video frames and decides both where to look next and when to emit a prediction. Since backpropagation is not adequate in this non-differentiable setting, we use REINFORCE to learn the agent's decision policy. Our model achieves state-of-the-art results on the THUMOS'14 and ActivityNet datasets while observing only a fraction (2% or less) of the video frames. Serena Yeung-Levy, Olga Russakovsky, Greg Mori, Li Fei-Fei 0001 |
CVPR | 3 |
| 2016 | Unsupervised learning of supervoxel embeddings for video SegmentationabstractWe present an algorithm for learning a feature representation for video segmentation. Standard video segmentation algorithms utilize similarity measurements in order to group related pixels. The contribution of our paper is an unsupervised method for learning the feature representation used for this similarity. The feature representation is defined over video supervoxels. An embedding framework learns a feature mapping for supervoxels in an unsupervised fashion such that supervoxels with similar context have similar embeddings. Based on the learned representation, we can merge similar supervoxels into spatio-temporal segments. Experimental results demonstrate the effectiveness of this learned supervoxel embedding on standard benchmark data. Mehran Khodabandeh, Srikanth Muralidharan, Arash Vahdat, Nazanin Mehrasa, Eduardo M. Pereira, Shin'ichi Satoh 0001, Greg Mori |
ICPR | 7 |
| 2016 | Special Issue on Individual and Group Activities in Video Event Analysis
Liang Wang 0001, Ioannis Patras, Jian Zhang 0002, Greg Mori, Larry Davis 0001 |
Comput. Vis. Image Underst. | 4 |
| 2015 | Deep Structured Models For Group Activity RecognitionabstractThis thesis presents a deep neural-network-based hierarchical graphical model for individual and group activity recognition in surveillance scenes. As the first step, deep networks are used to recognize activities of individual people in a scene. Then, a neural network-based hierarchical graphical model refines the predicted labels for each activity by considering dependencies between different classes. Similar to the inference mechanism in a probabilistic graphical model, the refinement step mimics a message-passing encoded into a deep neural network architecture. We show that this approach can be effective in group activity recognition. The deep graphical model improves recognition rates over baseline methods. Zhiwei Deng, Mengyao Zhai, Lei Chen 0023, Srikanth Muralidharan, Mehrsan Javan Roshtkhari, Greg Mori |
BMVC | 7 |
| 2015 | Multi-Task Transfer Methods to Improve One-Shot Learning for Multimedia Event DetectionabstractLearning a model for complex video event detection from only one positive sam-ple is a challenging and important problem in practice, yet seldom has been addressed. This paper proposes a new one-shot learning method based on multi-task learning to ad-dress this problem. Information from external relevant events is utilized to overcome the paucity of positive samples for the given event. Relevant events are identified implicitly and are emphasized more in the training. Moreover, a new dataset focusing on personal video search is collected. Experiments on both TRECVid Multimedia Event Detection video set and the new dataset verify the efficacy of the proposed methods. 1 Wang Yan, Jordan Yap, Greg Mori |
BMVC | 3 |
| 2015 | Visual recognition by counting instances: A multi-instance cardinality potential kernelabstractMany visual recognition problems can be approached by counting instances. To determine whether an event is present in a long internet video, one could count how many frames seem to contain the activity. Classifying the activity of a group of people can be done by counting the actions of individual people. Encoding these cardinality relationships can reduce sensitivity to clutter, in the form of irrelevant frames or individuals not involved in a group activity. Learned parameters can encode how many instances tend to occur in a class of interest. To this end, this paper develops a powerful and flexible framework to infer any cardinality relation between latent labels in a multi-instance model. Hard or soft cardinality relations can be encoded to tackle diverse levels of ambiguity. Experiments on tasks such as human activity recognition, video event detection, and video summarization demonstrate the effectiveness of using cardinality relations for improving recognition results. Hossein Hajimirsadeghi, Wang Yan, Arash Vahdat, Greg Mori |
CVPR | 4 |
| 2015 | Learning Ensembles of Potential Functions for Structured Prediction with Latent VariablesabstractMany visual recognition tasks involve modeling variables which are structurally related. Hidden conditional random fields (HCRFs) are a powerful class of models for encoding structure in weakly supervised training examples. This paper presents HCRF-Boost, a novel and general framework for learning HCRFs in functional space. An algorithm is proposed to learn the potential functions of an HCRF as a combination of abstract nonlinear feature functions, expressed by regression models. Consequently, the resulting latent structured model is not restricted to traditional log-linear potential functions or any explicit parameterization. Further, functional optimization helps to avoid direct interactions with the possibly large parameter space of nonlinear models and improves efficiency. As a result, a complex and flexible ensemble method is achieved for structured prediction which can be successfully used in a variety of applications. We validate the effectiveness of this method on tasks such as group activity recognition, human action recognition, and multi-instance learning of video events. Hossein Hajimirsadeghi, Greg Mori |
ICCV | 2 |
| 2015 | Learning Temporal Embeddings for Complex Video AnalysisabstractIn this paper, we propose to learn temporal embeddings of video frames for complex video analysis. Large quantities of unlabeled video data can be easily obtained from the Internet. These videos possess the implicit weak label that they are sequences of temporally and semantically coherent images. We leverage this information to learn temporal embeddings for video frames by associating frames with the temporal context that they appear in. To do this, we propose a scheme for incorporating temporal context based on past and future frames in videos, and compare this to other contextual representations. In addition, we show how data augmentation using multi-resolution samples and hard negatives helps to significantly improve the quality of the learned embeddings. We evaluate various design decisions for learning temporal embeddings, and show that our embeddings can improve performance for multiple video tasks such as retrieval, classification, and temporal order recovery in unconstrained Internet video. Vignesh Ramanathan, Kevin D. Tang, Greg Mori, Li Fei-Fei 0001 |
ICCV | 3 |
| 2015 | Clustered Exemplar-SVM: Discovering sub-categories for visual recognitionabstractWe present a novel algorithm for image classification that is targeted to capture class variability. A single model is often not sufficient to represent a category since categories can vary from large semantic classes to fine-grained sub-categories. Instead, we develop a representation based on discovering visually similar sub-categories within a given class. We introduce a novel Clustered Exemplar SVM classifier which incorporates data-driven and exemplar focused discovery. Semi-supervised learning is employed for training each C-eSVM classifier. We evaluate our approach on two datasets and demonstrate the efficacy of our method over standard Exemplar SVM. Nataliya Shapovalova, Greg Mori |
ICIP | 2 |
| 2015 | Discriminative key-component models for interaction detection and recognition
Yasaman S. Sefidgar, Arash Vahdat, Stephen Se, Greg Mori |
Comput. Vis. Image Underst. | 4 |
| 2014 | Discovering Video Clusters from Visual Features and Noisy Tags
Arash Vahdat, Guang-Tong Zhou, Greg Mori |
ECCV (6) | 3 |
| 2014 | Integrating multi-modal interfaces to command UAVsabstractWe present an integrated human-robot interaction system that enables a user to select and command a team of two Unmanned Aerial Vehicles (UAV) using voice, touch, face engagement and hand gestures. This system integrates multiple human [multi]-robot interaction interfaces as well as a navigation and mapping algorithm in a coherent semi-realistic scenario. The task of the UAVs is to explore and map a simulated Mars environment. Valiallah Monajjemi, Shokoofeh Pourmehr, Seyed Abbas Sadat, Fei Zhan, Jens Wawerla, Greg Mori, Richard Vaughan 0001 |
HRI | 6 |
| 2014 | "You are green": a touch-to-name interaction in an integrated multi-modal multi-robot HRI systemabstractWe present a multi-modal multi-robot interaction whereby a user can identify an individual or a group of robots using haptic stimuli, and name them using a voice command (e.g."You two are green"). Subsequent commands can be addressed to the same robot(s) by name (e.g. "Green! Take off!"). We demonstrate this as part of a real-world integrated system in which a user commands teams of autonomous robots in a coordinated exploration task. Shokoofeh Pourmehr, Valiallah Monajjemi, Seyed Abbas Sadat, Fei Zhan, Jens Wawerla, Greg Mori, Richard Vaughan 0001 |
HRI | 6 |
| 2014 | Multimedia event detection with multimodal feature fusion and temporal concept localization
Sangmin Oh, Scott McCloskey, Ilseo Kim, Arash Vahdat, Kevin J. Cannons, Hossein Hajimirsadeghi, Greg Mori, A. G. Amitha Perera, Megha Pandey, Jason J. Corso |
Mach. Vis. Appl. | 7 |
| 2013 | A Max-Margin Riffled Independence Model for Image Tag RankingabstractWe propose Max-Margin Riffled Independence Model (MMRIM), a new method for image tag ranking modeling the structured preferences among tags. The goal is to predict a ranked tag list for a given image, where tags are ordered by their importance or relevance to the image content. Our model integrates the max-margin formalism with riffled independence factorizations proposed in [10], which naturally allows for structured learning and efficient ranking. Experimental results on the SUN Attribute and Label Me datasets demonstrate the superior performance of the proposed model compared with baseline tag ranking methods. We also apply the predicted rank list of tags to several higher-level computer vision applications in image understanding and retrieval, and demonstrate that MMRIM significantly improves the accuracy of these applications. Tian Lan 0006, Greg Mori |
CVPR | 2 |
| 2013 | Learning Class-to-Image Distance with Object MatchingsabstractWe conduct image classification by learning a class-to-image distance function that matches objects. The set of objects in training images for an image class are treated as a collage. When presented with a test image, the best matching between this collage of training image objects and those in the test image is found. We validate the efficacy of the proposed model on the PASCAL 07 and SUN 09 datasets, showing that our model is effective for object classification and scene classification tasks. State-of-the-art image classification results are obtained, and qualitative results demonstrate that objects can be accurately matched. Guang-Tong Zhou, Tian Lan 0006, Weilong Yang, Greg Mori |
CVPR | 4 |
| 2013 | From Subcategories to Visual Composites: A Multi-level Framework for Object DetectionabstractThe appearance of an object changes profoundly with pose, camera view and interactions of the object with other objects in the scene. This makes it challenging to learn detectors based on an object-level label (e.g., "car"). We postulate that having a richer set of labelings (at different levels of granularity) for an object, including finer-grained subcategories, consistent in appearance and view, and higher order composites - contextual groupings of objects consistent in their spatial layout and appearance, can significantly alleviate these problems. However, obtaining such a rich set of annotations, including annotation of an exponentially growing set of object groupings, is simply not feasible. We propose a weakly-supervised framework for object detection where we discover subcategories and the composites automatically with only traditional object-level category labels as input. To this end, we first propose an exemplar-SVM-based clustering approach, with latent SVM refinement, that discovers a variable length set of discriminative subcategories for each object class. We then develop a structured model for object detection that captures interactions among object subcategories and automatically discovers semantically meaningful and discriminatively relevant visual composites. We show that this model produces state-of-the-art performance on UIUC phrase object detection benchmark. Tian Lan 0006, Michalis Raptis, Leonid Sigal, Greg Mori |
ICCV | 4 |
| 2013 | Compositional Models for Video Event Detection: A Multiple Kernel Learning Latent Variable ApproachabstractWe present a compositional model for video event detection. A video is modeled using a collection of both global and segment-level features and kernel functions are employed for similarity comparisons. The locations of salient, discriminative video segments are treated as a latent variable, allowing the model to explicitly ignore portions of the video that are unimportant for classification. A novel, multiple kernel learning (MKL) latent support vector machine (SVM) is defined, that is used to combine and re-weight multiple feature types in a principled fashion while simultaneously operating within the latent variable framework. The compositional nature of the proposed model allows it to respond directly to the challenges of temporal clutter and intra-class variation, which are prevalent in unconstrained internet videos. Experimental results on the TRECVID Multimedia Event Detection 2011 (MED11) dataset demonstrate the efficacy of the method. Arash Vahdat, Kevin J. Cannons, Greg Mori, Sangmin Oh, Ilseo Kim |
ICCV | 3 |
| 2013 | Handling Uncertain Tags in Visual RecognitionabstractGathering accurate training data for recognizing a set of attributes or tags on images or videos is a challenge. Obtaining labels via manual effort or from weakly-supervised data typically results in noisy training labels. We develop the FlipSVM, a novel algorithm for handling these noisy, structured labels. The FlipSVM models label noise by "flipping" labels on training examples. We show empirically that the FlipSVM is effective on images-and-attributes and video tagging datasets. Arash Vahdat, Greg Mori |
ICCV | 2 |
| 2013 | A robust integrated system for selecting and commanding multiple mobile robotsabstractWe describe a system whereby multiple humans and mobile robots interact robustly using a combination of sensing and signalling modalities. Extending our previous work on selecting an individual robot from a population by face-engagement, we show that reaching toward a robot - a specialization of pointing - can be used to designate a particular robot for subsequent one-on-one interaction. To achieve robust operation despite frequent sensing problems, the robots use three phases of human detection and tracking, and emit audio cues to solicit interaction and guide the behaviour of the human. A series of real-world trials demonstrates the practicality of our approach. Shokoofeh Pourmehr, Valiallah Monajjemi, Jens Wawerla, Richard Vaughan 0001, Greg Mori |
ICRA | 5 |
| 2013 | HRI in the sky: Creating and commanding teams of UAVs with a vision-mediated gestural interfaceabstractExtending our previous work in real-time vision-based Human Robot Interaction (HRI) with multi-robot systems, we present the first example of creating, modifying and commanding teams of UAVs by an uninstrumented human. To create a team the user focuses attention on an individual robot by simply looking at it, then adds or removes it from the current team with a motion-based hand gesture. Another gesture commands the entire team to begin task execution. Robots communicate among themselves by wireless network to ensure that no more than one robot is focused, and so that the whole team agrees that it has been commanded. Since robots can be added and removed from the team, the system is robust to incorrect additions. A series of trials with two and three very low-cost UAVs and off-board processing demonstrates the practicality of our approach. Valiallah Monajjemi, Jens Wawerla, Richard Vaughan 0001, Greg Mori |
IROS | 4 |
| 2013 | "You two! Take off!": Creating, modifying and commanding groups of robots using face engagement and indirect speech in voice commandsabstractWe present a multimodal system for creating, modifying and commanding groups of robots from a population. Extending our previous work on selecting an individual robot from a population by face engagement, we show that we can dynamically create groups of a desired number of robots by speaking the number we desire, e.g. “You three”, and looking at the robots we intend to form the group. We evaluate two different methods of detecting which robots are intended by the user, and show that an iterated election performs well in our setting. We also show that teams can be modified by adding and removing individual robots: “And you. Not you”. The success of the system is examined for different spatial configurations of robots with respect to each other and the user to find the proper workspace of selection methods. Shokoofeh Pourmehr, Valiallah Monajjemi, Richard Vaughan 0001, Greg Mori |
IROS | 4 |
| 2013 | Segmental multi-way local pooling for video recognitionabstractIn this work, we address the problem of complex event detection on unconstrained videos. We introduce a novel multi-way feature pooling approach which leverages segment-level information. The approach is simple and widely applicable to diverse audio-visual features. Our approach uses a set of clusters discovered via unsupervised clustering of segment-level features. Depending on feature characteristics, not only scene-based clusters but also motion/audio-based clusters can be incorporated. Then, every video is represented with multiple descriptors, where each descriptor is designed to relate to one of the pre-built clusters. For classification, intersection kernel SVMs are used where the kernel is obtained by combining multiple kernels computed from corresponding per-cluster descriptor pairs. Evaluation on TRECVID'11 MED dataset shows a significant improvement by the proposed approach beyond the state-of-the-art. Ilseo Kim, Sangmin Oh, Arash Vahdat, Kevin J. Cannons, A. G. Amitha Perera, Greg Mori |
ACM Multimedia | 6 |
| 2013 | Action is in the Eye of the Beholder: Eye-gaze Driven Model for Spatio-Temporal Action LocalizationabstractWe propose a new weakly-supervised structured learning approach for recognition and spatio-temporal localization of actions in video. As part of the proposed approach we develop a generalization of the Max-Path search algorithm, which allows us to efficiently search over a structured space of multiple spatio-temporal paths, while also allowing to incorporate context information into the model. Instead of using spatial annotations, in the form of bounding boxes, to guide the latent model during training, we utilize human gaze data in the form of a weak supervisory signal. This is achieved by incorporating gaze, along with the classification, into the structured loss within the latent SVM learning framework. Experiments on a challenging benchmark dataset, UCF-Sports, show that our model is more accurate, in terms of classification, and achieves state-of-the-art results in localization. In addition, we show how our model can produce top-down saliency maps conditioned on the classification label and localized latent paths. Nataliya Shapovalova, Michalis Raptis, Leonid Sigal, Greg Mori |
NIPS | 4 |
| 2013 | Latent Maximum Margin ClusteringabstractWe present a maximum margin framework that clusters data using latent variables. Using latent representations enables our framework to model unobserved information embedded in the data. We implement our idea by large margin learning, and develop an alternating descent algorithm to effectively solve the resultant non-convex optimization problem. We instantiate our latent maximum margin clustering framework with tag-based video clustering tasks, where each video is represented by a latent tag model describing the presence or absence of video tags. Experimental results obtained on three standard datasets show that the proposed method outperforms non-latent maximum margin clustering as well as conventional clustering approaches. Guang-Tong Zhou, Tian Lan 0006, Arash Vahdat, Greg Mori |
NIPS | 4 |
| 2013 | Multiple Instance Learning by Discriminative Training of Markov Networks
Hossein Hajimirsadeghi, Jinling Li, Greg Mori, Mohamed H. Zaki, Tarek Sayed |
UAI | 3 |
| 2013 | Optimizing Nondecomposable Loss Functions in Structured PredictionabstractWe develop an algorithm for structured prediction with nondecomposable performance measures. The algorithm learns parameters of Markov Random Fields (MRFs) and can be applied to multivariate performance measures. Examples include performance measures such as Fβ score (natural language processing), intersection over union (object category segmentation), Precision/Recall at k (search engines), and ROC area (binary classifiers). We attack this optimization problem by approximating the loss function with a piecewise linear function. The loss augmented inference forms a Quadratic Program (QP), which we solve using LP relaxation. We apply this approach to two tasks: object class-specific segmentation and human action retrieval from videos. We show significant improvement over baseline approaches that either use simple loss functions or simple scoring functions on the PASCAL VOC and H3D Segmentation datasets, and a nursing home action recognition dataset. Mani Ranjbar, Tian Lan 0006, Yang Wang 0003, Stephen N. Robinovitch, Ze-Nian Li, Greg Mori |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2012 | Social roles in hierarchical models for human activity recognitionabstractWe present a hierarchical model for human activity recognition in entire multi-person scenes. Our model describes human behaviour at multiple levels of detail, ranging from low-level actions through to high-level events. We also include a model of social roles, the expected behaviours of certain people, or groups of people, in a scene. The hierarchical model includes these varied representations, and various forms of interactions between people present in a scene. The model is trained in a discriminative max-margin framework. Experimental results demonstrate that this model can improve performance at all considered levels of detail, on two challenging datasets. Tian Lan 0006, Leonid Sigal, Greg Mori |
CVPR | 3 |
| 2012 | Complex loss optimization via dual decompositionabstractWe describe a novel max-margin parameter learning approach for structured prediction problems under certain non-decomposable performance measures. Structured prediction is a common approach in many vision problems. Non-decomposable performance measures are also commonplace. However, efficient general methods for learning parameters against non-decomposable performance measures do not exist. In this paper we develop such a method, based on dual decomposition, that is applicable to a large class of non-decomposable performance measures. We exploit dual decomposition to factorize the original hard problem into two smaller problems and show how to optimize each factor efficiently. We show experimentally that the proposed approach significantly outperforms alternatives, which either sacrifice the model structure or approximate the performance measure, and is an order of magnitude faster than a previous approach with comparable results. Mani Ranjbar, Arash Vahdat, Greg Mori |
CVPR | 3 |
| 2012 | Image Retrieval with Structured Object Queries Using Latent Ranking SVM
Tian Lan 0006, Weilong Yang, Yang Wang 0003, Greg Mori |
ECCV (6) | 4 |
| 2012 | Similarity Constrained Latent Support Vector Machine: An Application to Weakly Supervised Action Classification
Nataliya Shapovalova, Arash Vahdat, Kevin J. Cannons, Tian Lan 0006, Greg Mori |
ECCV (7) | 5 |
| 2012 | Multiple instance real boosting with aggregation functions
Hossein Hajimirsadeghi, Greg Mori |
ICPR | 2 |
| 2012 | Kernel Latent SVM for Visual RecognitionabstractLatent SVMs (LSVMs) are a class of powerful tools that have been successfully applied to many applications in computer vision. However, a limitation of LSVMs is that they rely on linear models. For many computer vision tasks, linear models are suboptimal and nonlinear models learned with kernels typically perform much better. Therefore it is desirable to develop the kernel version of LSVM. In this paper, we propose kernel latent SVM (KLSVM) -- a new learning framework that combines latent SVMs and kernel methods. We develop an iterative training algorithm to learn the model parameters. We demonstrate the effectiveness of KLSVM using three different applications in visual recognition. Our KLSVM formulation is very general and can be applied to solve a wide range of applications in computer vision and machine learning. Weilong Yang, Yang Wang 0003, Arash Vahdat, Greg Mori |
NIPS | 4 |
| 2012 | A large margin framework for single camera offline tracking with hybrid cues
Bahman Yari Saeed Khanloo, Ferdinand Stefanus, Mani Ranjbar, Ze-Nian Li, Nicolas Saunier, Tarek Sayed, Greg Mori |
Comput. Vis. Image Underst. | 7 |
| 2012 | Discriminative Latent Models for Recognizing Contextual Group ActivitiesabstractIn this paper, we go beyond recognizing the actions of individuals and focus on group activities. This is motivated from the observation that human actions are rarely performed in isolation; the contextual information of what other people in the scene are doing provides a useful cue for understanding high-level activities. We propose a novel framework for recognizing group activities which jointly captures the group activity, the individual person actions, and the interactions among them. Two types of contextual information, group-person interaction and person-person interaction, are explored in a latent variable framework. In particular, we propose three different approaches to model the person-person interaction. One approach is to explore the structures of person-person interaction. Differently from most of the previous latent structured models, which assume a predefined structure for the hidden layer, e.g., a tree structure, we treat the structure of the hidden layer as a latent variable and implicitly infer it during learning and inference. The second approach explores person-person interaction in the feature level. We introduce a new feature representation called the action context (AC) descriptor. The AC descriptor encodes information about not only the action of an individual person in the video, but also the behavior of other people nearby. The third approach combines the above two. Our experimental results demonstrate the benefit of using contextual information for disambiguating group activities. Tian Lan 0006, Yang Wang 0003, Weilong Yang, Stephen N. Robinovitch, Greg Mori |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2011 | Max-margin Latent Dirichlet Allocation for Image Classification and AnnotationabstractWe present the max-margin latent Dirichlet allocation, a max-margin variant of supervised topic models, for image classification and annotation. Our model for image classification (called MMLDA c) integrates discriminative classification with generative topic models. Our model for image annotation (called MMLDA a) extends MMLDA c to the case of multi-label problems, where each image can be associated with more than one annotation terms. We derive efficient learning algorithms for both models and demonstrate experimentally the advantages of our proposed models over other baseline methods. 1 Yang Wang 0003, Greg Mori |
BMVC | 2 |
| 2011 | Latent Boosting for Action RecognitionabstractIn this thesis, we present work towards addressing a grand challenge of computer vision, human action recognition and detection. In particular, we focus on the problem of recognizing and detecting the actions of a person from a video sequence. To recognize human actions in a video, a typical approach involves first detecting and tracking people, followed by classification. However, accurate tracking is challenging, and the state-of-art tracking methods are not reliable. Since accurate tracking is not a direct end-goal of action recognition, we consider tracking as a latent variable and train a model focused on action recognition. We propose a novel learning algorithm for training models with latent variables in a boosting framework. Moreover, we show that the algorithm can be used to train an action recognition model in which the tracking trajectory of a person is a latent variable. This new model outperforms baselines on a variety of datasets. Zhi Feng Huang, Weilong Yang, Yang Wang 0003, Greg Mori |
BMVC | 4 |
| 2011 | Selecting and commanding groups in a multi-robot vision based systemabstractWe present a novel method for a human user to select groups of robots without using any external instruments. We use computer vision techniques to read hand gestures from a user and use the hand gesture information to select single or multiple robots from a population and assign them to a task. To select robots the user simply draws a circle in the air around the robots that the user wants to command. Once the user selects the group of robots, he or she can send them to a location by pointing to a target location. Brian Milligan, Greg Mori, Richard Vaughan 0001 |
HRI | 2 |
| 2011 | Discriminative figure-centric models for joint action localization and recognitionabstractIn this paper we develop an algorithm for action recognition and localization in videos. The algorithm uses a figure-centric visual word representation. Different from previous approaches it does not require reliable human detection and tracking as input. Instead, the person location is treated as a latent variable that is inferred simultaneously with action recognition. A spatial model for an action is learned in a discriminative fashion under a figure-centric representation. Temporal smoothness over video sequences is also enforced. We present results on the UCF-Sports dataset, verifying the effectiveness of our model in situations where detection and tracking of individuals is challenging. Tian Lan 0006, Yang Wang 0003, Greg Mori |
ICCV | 3 |
| 2011 | Hidden Part Models for Human Action Recognition: Probabilistic versus Max MarginabstractWe present a discriminative part-based approach for human action recognition from video sequences using motion features. Our model is based on the recently proposed hidden conditional random field (HCRF) for object recognition. Similarly to HCRF for object recognition, we model a human action by a flexible constellation of parts conditioned on image observations. Differently from object recognition, our model combines both large-scale global features and local patch features to distinguish various actions. Our experimental results show that our model is comparable to other state-of-the-art approaches in action recognition. In particular, our experimental results demonstrate that combining large-scale global features and local patch features performs significantly better than directly applying HCRF on local patches alone. We also propose an alternative for learning the parameters of an HCRF model in a max-margin framework. We call this method the max-margin hidden conditional random field (MMHCRF). We demonstrate that MMHCRF outperforms HCRF in human action recognition. In addition, MMHCRF can handle a much broader range of complex hidden structures arising in various problems in computer vision. Yang Wang 0003, Greg Mori |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Recognizing human actions from still images with latent posesabstractWe consider the problem of recognizing human actions from still images. We propose a novel approach that treats the pose of the person in the image as latent variables that will help with recognition. Different from other work that learns separate systems for pose estimation and action recognition, then combines them in an ad-hoc fashion, our system is trained in an integrated fashion that jointly considers poses and actions. Our learning objective is designed to directly exploit the pose information for action recognition. Our experimental results demonstrate that by inferring the latent poses, we can improve the final action recognition results. Weilong Yang, Yang Wang 0003, Greg Mori |
CVPR | 3 |
| 2010 | Optimizing Complex Loss Functions in Structured Prediction
Mani Ranjbar, Greg Mori, Yang Wang 0003 |
ECCV (2) | 2 |
| 2010 | A Discriminative Latent Model of Object Classes and Attributes
Yang Wang 0003, Greg Mori |
ECCV (5) | 2 |
| 2010 | Selecting and commanding individual robots in a vision-based multi-robot systemabstractThis video presents a computer vision based system for interaction between a single human and multiple robots. Face contact and motion-based gestures are used as two different non-verbal communication channels; a user first selects a particular robot by simply looking at it, then assigns it a task by waving his or her hand. Alex Couture-Beil, Richard Vaughan 0001, Greg Mori |
HRI | 3 |
| 2010 | Beyond Actions: Discriminative Models for Contextual Group ActivitiesabstractWe propose a discriminative model for recognizing group activities. Our model jointly captures the group activity, the individual person actions, and the interactions among them. Two new types of contextual information, group-person interaction and person-person interaction, are explored in a latent variable framework. Different from most of the previous latent structured models which assume a predefined structure for the hidden layer, e.g. a tree structure, we treat the structure of the hidden layer as a latent variable and implicitly infer it during learning and inference. Our experimental results demonstrate that by inferring this contextual information together with adaptive structures, the proposed model can significantly improve activity recognition performance. Tian Lan 0006, Yang Wang 0003, Weilong Yang, Greg Mori |
NIPS | 4 |
| 2010 | A Discriminative Latent Model of Image Region and Object Tag CorrespondenceabstractWe propose a discriminative latent model for annotating images with unaligned object-level textual annotations. Instead of using the bag-of-words image representation currently popular in the computer vision community, our model explicitly captures more intricate relationships underlying visual and textual information. In particular, we model the mapping that translates image regions to annotations. This mapping allows us to relate image regions to their corresponding annotation terms. We also model the overall scene label as latent information. This allows us to cluster test images. Our training data consist of images and their associated annotations. But we do not have access to the ground-truth region-to-annotation mapping or the overall scene label. We develop a novel variant of the latent SVM framework to model them as latent variables. Our experimental results demonstrate the effectiveness of the proposed model compared with other baseline methods. Yang Wang 0003, Greg Mori |
NIPS | 2 |
| 2009 | Efficient Human Action Detection Using a Transferable Distance Function
Weilong Yang, Yang Wang 0003, Greg Mori |
ACCV (2) | 3 |
| 2009 | Stacks of convolutional Restricted Boltzmann Machines for shift-invariant feature learningabstractIn this paper we present a method for learning class-specific features for recognition. Recently a greedy layer-wise procedure was proposed to initialize weights of deep belief networks, by viewing each layer as a separate restricted Boltzmann machine (RBM). We develop the convolutional RBM (C-RBM), a variant of the RBM model in which weights are shared to respect the spatial structure of images. This framework learns a set of features that can generate the images of a specific object class. Our feature extraction model is a four layer hierarchy of alternating filtering and maximum subsampling. We learn feature parameters of the first and third layers viewing them as separate C-RBMs. The outputs of our feature extraction hierarchy are then fed as input to a discriminative classifier. It is experimentally demonstrated that the extracted features are effective for object detection, using them to obtain performance comparable to the state of the art on handwritten digit recognition and pedestrian detection. Mohammad Norouzi 0002, Mani Ranjbar, Greg Mori |
CVPR | 3 |
| 2009 | Max-margin hidden conditional random fields for human action recognitionabstractWe present a new method for classification with structured latent variables. Our model is formulated using the max-margin formalism in the discriminative learning literature. We propose an efficient learning algorithm based on the cutting plane method and decomposed dual optimization. We apply our model to the problem of recognizing human actions from video sequences, where we model a human action as a global root template and a constellation of several “parts”. We show that our model outperforms another similar method that uses hidden conditional random fields, and is comparable to other state-of-the-art approaches. More importantly, our proposed work is quite general and can potentially be applied in a wide variety of vision problems that involve various complex, interdependent latent structures. Yang Wang 0003, Greg Mori |
CVPR | 2 |
| 2009 | A Rate Distortion Approach for Semi-Supervised Conditional Random FieldsabstractWe propose a novel information theoretic approach for semi-supervised learning of conditional random fields. Our approach defines a training objective that combines the conditional likelihood on labeled data and the mutual information on unlabeled data. Different from previous minimum conditional entropy semi-supervised discriminative learning methods, our approach can be naturally cast into the rate distortion theory framework in information theory. We analyze the tractability of the framework for structured prediction and present a convergent variational training algorithm to defy the combinatorial explosion of terms in the sum over label configurations. Our experimental results show that the rate distortion approach outperforms standard $l_2$ regularization and minimum conditional entropy regularization on both multi-class classification and sequence labeling problems. Yang Wang 0003, Gholamreza Haffari, Greg Mori |
NIPS | 4 |
| 2009 | Second and Third Canadian Conferences on Computer and Robot Vision
Robert Sim, Greg Mori, Ioannis M. Rekleitis |
Image Vis. Comput. | 2 |
| 2009 | BearCam: automated wildlife monitoring at the arctic circle
Jens Wawerla, Shelley Marshall, Greg Mori, Kristina Rothley, Payam Sabzmeydani |
Mach. Vis. Appl. | 3 |
| 2009 | Human Action Recognition by Semilatent Topic ModelsabstractWe propose two new models for human action recognition from video sequences using topic models. Video sequences are represented by a novel "bag-of-words" representation, where each frame corresponds to a "word." Our models differ from previous latent topic models for visual recognition in two major aspects: first of all, the latent topics in our models directly correspond to class labels; second, some of the latent variables in previous topic models become observed in our case. Our models have several advantages over other latent topic models used in visual recognition. First of all, the training is much easier due to the decoupling of the model parameters. Second, it alleviates the issue of how to choose the appropriate number of latent topics. Third, it achieves much better performance by utilizing the information provided by the class labels in the training set. We present action classification results on five different data sets. Our results are either comparable to, or significantly better than previously published results on these data sets. Yang Wang 0003, Greg Mori |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Action recognition by learning mid-level motion featuresabstractThis paper presents a method for human action recognition based on patterns of motion. Previous approaches to action recognition use either local features describing small patches or large-scale features describing the entire human figure. We develop a method constructing mid-level motion features which are built from low-level optical flow information. These features are focused on local regions of the image sequence and are created using a variant of AdaBoost. These features are tuned to discriminate between different classes of action, and are efficient to compute at run-time. A battery of classifiers based on these mid-level features is created and used to classify input sequences. State-of-the-art results are presented on a variety of standard datasets. Alireza Fathi, Greg Mori |
CVPR | 2 |
| 2008 | Multiple Tree Models for Occlusion and Spatial Constraints in Human Pose Estimation
Yang Wang 0003, Greg Mori |
ECCV (3) | 2 |
| 2008 | Boosting with incomplete informationabstractIn real-world machine learning problems, it is very common that part of the input feature vector is incomplete: either not available, missing, or corrupted. In this paper, we present a boosting approach that integrates features with incomplete information and those with complete information to form a strong classifier. By introducing hidden variables to model missing information, we form loss functions that combine fully labeled data with partially labeled data to effectively learn normalized and unnormalized models. The primal problems of the proposed optimization problems with these loss functions are provided to show their close relationship and the motivations behind them. We use auxiliary functions to bound the change of the loss functions and derive explicit parameter update rules for the learning algorithms. We demonstrate encouraging results on two real-world problems --- visual object recognition in computer vision and named entity recognition in natural language processing --- to show the effectiveness of the proposed boosting approach. Gholamreza Haffari, Yang Wang 0003, Greg Mori, Feng Jiao |
ICML | 4 |
| 2008 | Learning a discriminative hidden part model for human action recognitionabstractWe present a discriminative part-based approach for human action recognition from video sequences using motion features. Our model is based on the recently proposed hidden conditional random field~(hCRF) for object recognition. Similar to hCRF for object recognition, we model a human action by a flexible constellation of parts conditioned on image observations. Different from object recognition, our model combines both large-scale global features and local patch features to distinguish various actions. Our experimental results show that our model is comparable to other state-of-the-art approaches in action recognition. In particular, our experimental results demonstrate that combining large-scale global features and local patch features performs significantly better than directly applying hCRF on local patches alone. Yang Wang 0003, Greg Mori |
NIPS | 2 |
| 2008 | Human Pose Estimation with Rotated Geometric BlurabstractWe consider the problem of estimating the pose of a human figure in a single image. Our method uses an exemplar-matching framework, where a test image is matched to a database of exemplars upon which body joint positions have been marked. We find the best matching exemplar for a test image by employing a variant of an existing deformable template matching framework. A hierarchical correspondence process is developed to improve the efficiency of the existing framework. Quantitative results on the CMUMoBo dataset verify the effectiveness of our approach. Bo Chen 0019, Greg Mori |
WACV | 3 |
| 2008 | Action-Based Multifield Video VisualizationabstractOne challenge in video processing is to detect actions and events, known or unknown, in video streams dynamically. This paper proposes a visualization solution, where a video stream is depicted as a series of snapshots at a relatively sparse interval, and detected actions are highlighted with continuous abstract illustrations. The combined imagery and illustrative visualization conveys multi-field information in a manner similar to electrocardiograms (ECG) and seismographs. We thus name this type of video visualization as VideoPerpetuoGram (VPG). In this paper, we describe a system that handles the aw and processed information of the video stream in a multi-field visualization pipeline. As examples, we consider the needs for highlighting several types of processed information, including detected actions in video streams, and estimated relationship between recognized objects. We examine the effective means for depicting multi-field information in VPG, and support our choice of visual mappings through a survey. Our GPU implementation facilitates the VPG-specific viewing specification through a sheared object space, as well as volume bricking and combinational rendering of volume data and glyphs. Ralf Peter Botchen, Sven Bachthaler, Fabian Schick, Min Chen 0001, Greg Mori, Daniel Weiskopf, Thomas Ertl |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2007 | Detecting Pedestrians by Learning Shapelet FeaturesabstractIn this paper, we address the problem of detecting pedestrians in still images. We introduce an algorithm for learning shapelet features, a set of mid-level features. These features are focused on local regions of the image and are built from low-level gradient information that discriminates between pedestrian and non-pedestrian classes. Using Ad-aBoost, these shapelet features are created as a combination of oriented gradient responses. To train the final classifier, we use AdaBoost for a second time to select a subset of our learned shapelets. By first focusing locally on smaller feature sets, our algorithm attempts to harvest more useful information than by examining all the low-level features together. We present quantitative results demonstrating the effectiveness of our algorithm. In particular, we obtain an error rate 14 percentage points lower (at 10-6FPPW) than the previous state of the art detector of Dalal and Triggs on the INRIA dataset. Payam Sabzmeydani, Greg Mori |
CVPR | 2 |
| 2007 | Human Pose Estimation using Motion ExemplarsabstractWe present a motion exemplar approach for finding body configuration in monocular videos. A motion correlation technique is employed to measure the motion similarity at various space-time locations between the input video and stored video templates. These observations are used to predict the conditional state, distributions of exemplars and joint positions. Exemplar sequence selection and joint position estimation are then solved with approximate inference using Gibbs sampling and gradient ascent. The presented approach is able to find joint positions accurately for people with textured clothing. Results are presented on a dataset containing slow, fast and incline walk videos of various people from different view angles. The results demonstrate an overall improvement compared to previous methods. Alireza Fathi, Greg Mori |
ICCV | 2 |
| 2006 | Unsupervised Discovery of Action ClassesabstractIn this paper we consider the problem of describing the action being performed by human figures in still images. We will attack this problem using an unsupervised learning approach, attempting to discover the set of action classes present in a large collection of training images. These action classes will then be used to label test images. Our approach uses the coarse shape of the human figures to match pairs of images. The distance between a pair of images is computed using a linear programming relaxation technique. This is a computationally expensive process, and we employ a fast pruning method to enable its use on a large collection of images. Spectral clustering is then performed using the resulting distances. We present clustering and image labeling results on a variety of datasets. Yang Wang 0003, Hao Jiang 0007, Mark S. Drew, Ze-Nian Li, Greg Mori |
CVPR (2) | 5 |
| 2006 | Recovering 3D Human Body Configurations Using Shape ContextsabstractThe problem we consider in this paper is to take a single two-dimensional image containing a human figure, locate the joint positions, and use these to estimate the body configuration and pose in three-dimensional space. The basic approach is to store a number of exemplar 2D views of the human body in a variety of different configurations and viewpoints with respect to the camera. On each of these stored views, the locations of the body joints (left elbow, right knee, etc.) are manually marked and labeled for future use. The input image is then matched to each stored view, using the technique of shape context matching in conjunction with a kinematic chain-based deformation model. Assuming that there is a stored view sufficiently similar in configuration and pose, the correspondence process will succeed. The locations of the body joints are then transferred from the exemplar view to the test shape. Given the 2D joint locations, the 3D body configuration and pose are then estimated using an existing algorithm. We can apply this technique to video by treating each frame independently--tracking just becomes repeated recognition. We present results on a variety of data sets. Greg Mori, Jitendra Malik |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Guiding Model Search Using SegmentationabstractIn this paper we show how segmentation as preprocessing paradigm can be used to improve the efficiency and accuracy of model search in an image. We operationalize this idea using an over-segmentation of an image into superpixels. The problem domain we explore is human body pose estimation from still images. The superpixels prove useful in two ways. First, we restrict the joint positions in our human body model to lie at centers of superpixels, which reduces the size of the model search space. In addition, accurate support masks for computing features on half-limbs of the body model are obtained by using agglomerations of superpixels as half limb segments. We present results on a challenging dataset of people in sports news images Greg Mori |
ICCV | 1 |
| 2005 | Efficient Shape Matching Using Shape ContextsabstractWe demonstrate that shape contexts can be used to quickly prune a search for similar shapes. We present two algorithms for rapid shape retrieval: representative shape contexts, performing comparisons based on a small number of shape contexts, and shapemes, using vector quantization in the space of shape contexts to obtain prototypical shape pieces. Greg Mori, Serge J. Belongie, Jitendra Malik |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Recovering Human Body Configurations: Combining Segmentation and Recognition
Greg Mori, Xiaofeng Ren, Alexei A. Efros, Jitendra Malik |
CVPR (2) | 1 |
| 2003 | Recognizing Objects in Adversarial Clutter: Breaking a Visual CAPTCHAabstractIn this paper we explore object recognition in clutter. We test our object recognition techniques on Gimpy and EZ-Gimpy, examples of visual CAPTCHAs. A CAPTCHA ("Completely Automated Public Turing test to Tell Computers and Humans Apart") is a program that can generate and grade tests that most humans can pass, yet current computer programs can't pass. EZ-Gimpy, currently used by Yahoo, and Gimpy are CAPTCHAs based on word recognition in the presence of clutter. These CAPTCHAs provide excellent test sets since the clutter they contain is adversarial; it is designed to confuse computer programs. We have developed efficient methods based on shape context matching that can identify the word in an EZ-Gimpy image with a success rate of 92%, and the requisite 3 words in a Gimpy image 33% of the time. The problem of identifying words in such severe clutter provides valuable insight into the more general problem of object recognition in scenes. The methods that we present are instances of a framework designed to tackle this general problem. Greg Mori, Jitendra Malik |
CVPR (1) | 1 |
| 2003 | Recognizing Action at a DistanceabstractOur goal is to recognize human action at a distance, at resolutions where a whole person may be, say, 30 pixels tall. We introduce a novel motion descriptor based on optical flow measurements in a spatiotemporal volume for each stabilized human figure, and an associated similarity measure to be used in a nearest-neighbor framework. Making use of noisy optical flow measurements is the key challenge, which is addressed by treating optical flow not as precise pixel displacements, but rather as a spatial pattern of noisy measurements which are carefully smoothed and aggregated to form our spatiotemporal motion descriptor. To classify the action being performed by a human figure in a query sequence, we retrieve nearest neighbor(s) from a database of stored, annotated video sequences. We can also use these retrieved exemplars to transfer 2D/3D skeletons onto the figures in the query sequence, as well as two forms of data-based action synthesis "do as I do" and "do as I say". Results are demonstrated on ballet, tennis as well as football datasets. Alexei A. Efros, Alexander C. Berg, Greg Mori, Jitendra Malik |
ICCV | 3 |
| 2002 | Estimating Human Body Configurations Using Shape Context Matching
Greg Mori, Jitendra Malik |
ECCV (3) | 1 |
| 2001 | Shape contexts enable efficient retrieval of similar shapesabstractIn this paper we demonstrate that a recently introduced shape descriptor, the "shape context", can be used to quickly prune a search for similar shapes. Our representation for a shape is a discrete set of n points sampled from its internal and external contours. For each of these points, the shape context is a histogram of the relative positions of the n - 1 remaining points. We present two methods for rapid shape retrieval: one that does comparisons based on a small number of shape contexts and another that uses vector quantization in the space of shape contexts. We verify the discriminative power of these methods with tests on the Columbia (COIL-100) 3D object database and the Snodgrass and Vanderwart line drawings. The shape context-based methods are shown to quickly produce an accurate shortlist of candidates suitable for a more exact matching engine in spite of pose variation and occlusion. Greg Mori, Serge J. Belongie, Jitendra Malik |
CVPR (1) | 1 |