Vladimir Pavlovic 0001

dblp:98/2506 · also Vladimir I. Pavlovic · DBLP profile ↗
← Back
141ranked-venue papers
13as first author
30since 2021 · last 2025
0000-0003-3979-1236ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 102 · 9 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 86 · 8 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-authorDatabases, data management, data science and information retrieval · 8 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Hallucinatory Image Tokens: A Training-Free EAZY Approach to Detecting and Mitigating Object Hallucinations in LVLMs
abstract
Despite their remarkable potential, Large Vision-Language Models (LVLMs) still face challenges with object hallucination, a problem where their generated outputs mistakenly incorporate objects that do not actually exist. Although most works focus on addressing this issue within the language-model backbone, our work shifts the focus to the image input source, investigating how specific image tokens contribute to hallucinations. Our analysis reveals a striking finding: a small subset of image tokens with high attention scores are the primary drivers of object hallucination. By removing these hallucinatory image tokens (only 1.5% of all image tokens), the issue can be effectively mitigated. This finding holds consistently across different models and datasets. Building on this insight, we introduce EAZY, a novel, training-free method that automatically identifies and Eliminates hAllucinations by Zeroing out hallucinatorY image tokens. We utilize EAZY for unsupervised object hallucination detection, achieving 15% improvement compared to previous methods. Additionally, EAZY demonstrates remarkable effectiveness in mitigating hallucinations while preserving model utility and seamlessly adapting to various LVLM architectures.
Liwei Che, Tony Qingze Liu, Weiyi Qin, Ruixiang Tang, Vladimir Pavlovic 0001
ICCV6
2025 GenVP: Generating Visual Puzzles with Contrastive Hierarchical VAEs
abstract
Raven’s Progressive Matrices (RPMs) is an established benchmark to examine the ability to perform high-level abstract visual reasoning (AVR). Despite the current success of algorithms that solve this task, humans can generalize beyond a given puzzle and create new puzzles given a set of rules, whereas machines remain locked in solving a fixed puzzle from a curated choice list. We propose Generative Visual Puzzles (GenVP), a framework to model the entire RPM generation process, a substantially more challenging task. Our model’s capability spans from generating multiple solutions for one specific problem prompt to creating complete new puzzles out of the desired set of rules. Experiments on five different datasets indicate that GenVP achieves state-of-the-art (SOTA) performance both in puzzle-solving accuracy and out-of-distribution (OOD) generalization in 22 out of 24 OOD scenarios. Further, compared to SOTA generative approaches, which struggle to solve RPMs when the feasible solution space increases, GenVP efficiently generalizes to these challenging scenarios. Moreover, our model demonstrates the ability to produce a wide range of complete RPMs given a set of abstract rules by effectively capturing the relationships between abstract rules and visual object properties.
Kalliopi Basioti, Pritish Sahu, Tony Qingze Liu, Zihao Xu 0001, Hao Wang 0014, Vladimir Pavlovic 0001
ICLR6
2025 SODA: Spectral Orthogonal Decomposition Adaptation for Diffusion Models
abstract
daptation (SODA), which balances computational efficiency and representation capacity. Extensive evaluations on text-to-image diffusion models demonstrate SODA's effectiveness, offering a spectrum-aware alternative to existing fine-tuning methods.
Xinxi Zhang, Song Wen 0001, Ligong Han, Felix Juefei-Xu, Akash Srivastava, Junzhou Huang, Vladimir Pavlovic 0001, Hao Wang 0014, Molei Tao, Dimitris N. Metaxas
WACV7
2025 Guest Editorial: When Multimedia Meets Food: Multimedia Computing for Food Data Analysis and Applications
Weiqing Min, Shuqiang Jiang, Petia Radeva, Vladimir Pavlovic 0001, Chong-Wah Ngo, Kiyoharu Aizawa, Wanqing Li 0001
IEEE Trans. Multim.4
2024 Learning from Synthetic Human Group Activities
abstract
The study of complex human interactions and group activities has become a focal point in human-centric computer vision. However, progress in related tasks is often hindered by the challenges of obtaining large-scale labeled datasets from real-world scenarios. To address the limitation, we introduce M3 Act, a synthetic data generator for multi-view multi-group multi-person human atomic actions and group activities. Powered by Unity Engine, M3 Act features mul-tiple semantic groups, highly diverse and photorealistic images, and a comprehensive set of annotations, which facilitates the learning of human-centered tasks across single-person, multi-person, and multi-group conditions. We demonstrate the advantages of M3 Act across three core experiments. The results suggest our synthetic dataset can significantly improve the performance of several downstream methods and replace real-world datasets to reduce cost. Notably, M3 Act improves the state-of-the-art MOTRv2 on DanceTrack dataset, leading to a hop on the leaderboard from 10thto 2ndplace. Moreover, M3 Act opens new research for controllable 3D group activity generation. We define multiple metrics and propose a competitive baseline for the novel task. Our code and data are available at our project page: http://cjerry1243.github.io/M3Act.
Che-Jui Chang, Danrui Li, Deep Patel, Parth Goel, Honglu Zhou, Seonghyeon Moon, Samuel S. Sohn, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia
CVPR9
2024 CIC-BART-SSA: Controllable Image Captioning with Structured Semantic Augmentation
Kalliopi Basioti, Mohamed Ashraf Abdelsalam, Federico Fancellu, Vladimir Pavlovic 0001, Afsaneh Fazly
ECCV (66)4
2024 Box2Flow: Instance-Based Action Flow Graphs from Videos
Jiatong Li 0001, Kalliopi Basioti, Vladimir Pavlovic 0001
ICPR (27)3
2024 TrajDiffuse: A Conditional Diffusion Model for Environment-Aware Trajectory Prediction
Tony Qingze Liu, Danrui Li, Samuel S. Sohn, Sejong Yoon, Mubbasir Kapadia, Vladimir Pavlovic 0001
ICPR (29)6
2023 Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural Videos
abstract
Following along how-to videos requires alternating focus between understanding procedural video instructions and performing them. Examining how to support these continuous context switches for the user has been largely unexplored. In this paper, we describe a user study with thirty participants who performed an hour-long cooking task while interacting with a wizard-of-oz hands-free interactive system that is aware of both their cooking progress and environment contexts. Through analysis of the session scripts, we identify a dichotomy between participant query differences and workflow alignment similarities, under-studied interactions that require AI functionality beyond video navigation alone, and queries that call for multimodal sensing of a user’s environment. By understanding the assistant experience through the participants’ interactions, we identify design implications for a smart assistant that can discern a user’s task completion flow and personal characteristics, accommodate requests within and external to the task domain, and support nonvoice-based queries.
Georgianna Lin, Jin Yi Li, Afsaneh Fazly, Vladimir Pavlovic 0001, Khai N. Truong
CHI4
2023 MSI: Maximize Support-Set Information for Few-Shot Segmentation
abstract
FSS (Few-shot segmentation) aims to segment a target class using a small number of labeled images (support set). To extract information relevant to the target class, a dominant approach in best performing FSS methods removes background features using a support mask. We observe that this feature excision through a limiting support mask introduces an information bottleneck in several challenging FSS cases, e.g., for small targets and/or inaccurate target boundaries. To this end, we present a novel method (MSI), which maximizes the support-set information by exploiting two complementary sources of features to generate super correlation maps. We validate the effectiveness of our approach by instantiating it into three recent and strong FSS methods. Experimental results on several publicly available FSS benchmarks show that our proposed method consistently improves performance by visible margins and leads to faster convergence. Our code and trained models are available at: https://github.com/moonsh/MSI-Maximize-Support-Set-Information
Seonghyeon Moon, Samuel S. Sohn, Honglu Zhou, Sejong Yoon, Vladimir Pavlovic 0001, Muhammad Haris Khan, Mubbasir Kapadia
ICCV5
2023 ALWOD: Active Learning for Weakly-Supervised Object Detection
abstract
Object detection (OD), a crucial vision task, remains challenged by the lack of large training datasets with precise object localization labels. In this work, we propose ALWOD, a new framework that addresses this problem by fusing active learning (AL) with weakly and semi-supervised object detection paradigms. Because the performance of AL critically depends on the model initialization, we propose a new auxiliary image generator strategy that utilizes an extremely small labeled set, coupled with a large weakly tagged set of images, as a warm-start for AL. We then propose a new AL acquisition function, another critical factor in AL success, that leverages the student-teacher OD pair disagreement and uncertainty to effectively propose the most informative images to annotate. Finally, to complete the AL loop, we introduce a new labeling task delegated to human annotators, based on selection and correction of model-proposed detections, which is both rapid and effective in labeling the informative images. We demonstrate, across several challenging benchmarks, that ALWOD significantly narrows the gap between the ODs trained on few partially labeled but strategically selected image instances and those that rely on the fullylabeled data. Our code is publicly available on https://github.com/seqam-lab/ALWOD.
Yuting Wang 0004, Velibor Ilic, Jiatong Li 0001, Branislav Kisacanin, Vladimir Pavlovic 0001
ICCV5
2023 NP-SemiSeg: When Neural Processes meet Semi-Supervised Semantic Segmentation
abstract
Semi-supervised semantic segmentation involves assigning pixel-wise labels to unlabeled images at training time. This is useful in a wide range of real-world applications where collecting pixel-wise labels is not feasible in time or cost. Current approaches to semi-supervised semantic segmentation work by predicting pseudo-labels for each pixel from a class-wise probability distribution output by a model. If this predicted probability distribution is incorrect, however, it leads to poor segmentation results which can have knock-on consequences in safety critical systems, like medical images or self-driving cars. It is, therefore, important to understand what a model does not know, which is mainly achieved by uncertainty quantification. Recently, neural processes (NPs) have been explored in semi-supervised image classification, and they have been a computationally efficient and effective method for uncertainty quantification. In this work, we move one step forward by adapting NPs to semi-supervised semantic segmentation, resulting in a new model called NP-SemiSeg. We experimentally evaluated NP-SemiSeg on the public benchmarks PASCAL VOC 2012 and Cityscapes, with different training settings, and the results verify its effectiveness.
Daniela Massiceti, Xiaolin Hu 0001, Vladimir Pavlovic 0001, Thomas Lukasiewicz
ICML4
2023 D2F2WOD: Learning Object Proposals for Weakly-Supervised Object Detection via Progressive Domain Adaptation
abstract
Weakly-supervised object detection (WSOD) models attempt to leverage image-level annotations in lieu of accurate but costly-to-obtain object localization labels. This oftentimes leads to substandard object detection and lo-calization at inference time. To tackle this issue, we propose D2F2WOD, a Dual-Domain Fully-to-Weakly Supervised Object Detection framework that leverages synthetic data, annotated with precise object localization, to supplement a natural image target domain, where only image-level labels are available. In its warm-up domain adaptation stage, the model learns a fully-supervised object detector (FSOD) to improve the precision of the object proposals in the target domain, and at the same time learns target-domain-specific and detection-aware proposal features. In its main WSOD stage, a WSOD model is specifically tuned to the target domain. The feature extractor and the object proposal generator of the WSOD model are built upon the fine-tuned FSOD model. We test D2F2WOD on five dual-domain image benchmarks. The results show that our method results in consistently improved object detection and localization compared with state-of-the-art methods.
Yuting Wang 0004, Ricardo Guerrero, Vladimir Pavlovic 0001
WACV3
2023 Heterogeneous Crowd Simulation Using Parametric Reinforcement Learning
abstract
Agent-based synthetic crowd simulation affords the cost-effective large-scale simulation and animation of interacting digital humans. Model-based approaches have successfully generated a plethora of simulators with a variety of foundations. However, prior approaches have been based on statically defined models predicated on simplifying assumptions, limited video-based datasets, or homogeneous policies. Recent works have applied reinforcement learning to learn policies for navigation. However, these approaches may learn static homogeneous rules, are typically limited in their generalization to trained scenarios, and limited in their usability in synthetic crowd domains. In this article, we present a multi-agent reinforcement learning-based approach that learns a parametric predictive collision avoidance and steering policy. We show that training over a parameter space produces a flexible model across crowd configurations. That is, our goal-conditioned approach learns a parametric policy that affords heterogeneous synthetic crowds. We propose a model-free approach without centralization of internal agent information, control signals, or agent communication. The model is extensively evaluated. The results show policy generalization across unseen scenarios, agent parameters, and out-of-distribution parameterizations. The learned model has comparable computational performance to traditional methods. Qualitatively the model produces both expected (laminar flow, shuffling, bottleneck) and unexpected (side-stepping) emergent qualitative behaviours, and quantitatively the approach is performant across measures of movement quality.
Kaidong Hu, M. Brandon Haworth, Glen Berseth, Vladimir Pavlovic 0001, Petros Faloutsos, Mubbasir Kapadia
IEEE Trans. Vis. Comput. Graph.4
2022 Cross-Modal Coherence for Text-to-Image Retrieval
abstract
Common image-text joint understanding techniques presume that images and the associated text can universally be characterized by a single implicit model. However, co-occurring images and text can be related in qualitatively different ways, and explicitly modeling it could improve the performance of current joint understanding models. In this paper, we train a Cross-Modal Coherence Model for text-to-image retrieval task. Our analysis shows that models trained with image–text coherence relations can retrieve images originally paired with target text more often than coherence-agnostic models. We also show via human evaluation that images retrieved by the proposed coherence-aware model are preferred over a coherence-agnostic baseline by a huge margin. Our findings provide insights into the ways that different modalities communicate and the role of coherence relations in capturing commonsense inferences in text and imagery.
Malihe Alikhani, Fangda Han, Hareesh Ravi, Mubbasir Kapadia, Vladimir Pavlovic 0001, Matthew Stone
AAAI5
2022 Variational Continual Proxy-Anchor for Deep Metric Learning
abstract
The recent proxy-anchor method achieved outstanding performance in deep metric learning, which can be acknowledged to its data efficient loss based on hard example mining, as well as far lower sampling complexity than pair-based approaches. In this paper we extend the proxy-anchor method by posing it within the continual learning framework, motivated from its batch-expected loss form (instead of instance-expected, typical in deep learning), which can potentially incur the catastrophic forgetting of historic batches. By regarding each batch as a task in continual learning, we adopt the Bayesian variational continual learning approach to derive a novel loss function. Interestingly the resulting loss has two key modifications to the original proxy-anchor loss: i) we inject noise to the proxies when optimizing the proxy-anchor loss, and ii) we encourage momentum update to avoid abrupt model changes. As a result, the learned model achieves higher test accuracy than proxy-anchor due to the robustness to noise in data (through model perturbation during training), and the reduced batch forgetting effect. We demonstrate the improved results on several benchmark datasets.
Minyoung Kim 0001, Ricardo Guerrero, Hai Xuan Pham, Vladimir Pavlovic 0001
AISTATS4
2022 Visual Semantic Parsing: From Images to Abstract Meaning Representation
abstract
Mohamed Ashraf Abdelsalam, Zhan Shi, Federico Fancellu, Kalliopi Basioti, Dhaivat Bhatt, Vladimir Pavlovic, Afsaneh Fazly. Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL). 2022.
Mohamed Ashraf Abdelsalam, Federico Fancellu, Kalliopi Basioti, Dhaivat Bhatt, Vladimir Pavlovic 0001, Afsaneh Fazly
CoNLL6
2022 MUSE-VAE: Multi-Scale VAE for Environment-Aware Long Term Trajectory Prediction
abstract
Accurate long-term trajectory prediction in complex scenes, where multiple agents (e.g., pedestrians or vehicles) interact with each other and the environment while attempting to accomplish diverse and often unknown goals, is a challenging stochastic forecasting problem. In this work, we propose MUSEVAE, a new probabilistic modeling framework based on a cascade of Conditional VAEs, which tackles the long-term, uncertain trajectory prediction task using a coarse-to-fine multi-factor forecasting architecture. In its Macro stage, the model learns a joint pixel-space representation of two key factors, the underlying environment and the agent movements, to predict the long and short term motion goals. Conditioned on them, the Micro stage learns a fine-grained spatio-temporal representation for the prediction of individual agent trajectories. The VAE backbones across the two stages make it possible to naturally account for the joint uncertainty at both levels of granularity. As a result, MUSEVAE offers diverse and simultaneously more accurate predictions compared to the current state-of-the-art. We demonstrate these assertions through a comprehensive set of experiments on nuScenes and SDD benchmarks as well as PFSD, a new synthetic dataset, which challenges the forecasting ability of models on complex agent-environment interaction scenarios.
Mihee Lee, Samuel S. Sohn, Seonghyeon Moon, Sejong Yoon, Mubbasir Kapadia, Vladimir Pavlovic 0001
CVPR6
2022 HM: Hybrid Masking for Few-Shot Segmentation
Seonghyeon Moon, Samuel S. Sohn, Honglu Zhou, Sejong Yoon, Vladimir Pavlovic 0001, Muhammad Haris Khan, Mubbasir Kapadia
ECCV (20)5
2022 NP-Match: When Neural Processes meet Semi-Supervised Learning
abstract
Semi-supervised learning (SSL) has been widely explored in recent years, and it is an effective way of leveraging unlabeled data to reduce the reliance on labeled data. In this work, we adjust neural processes (NPs) to the semi-supervised image classification task, resulting in a new method named NP-Match. NP-Match is suited to this task for two reasons. Firstly, NP-Match implicitly compares data points when making predictions, and as a result, the prediction of each unlabeled data point is affected by the labeled data points that are similar to it, which improves the quality of pseudolabels. Secondly, NP-Match is able to estimate uncertainty that can be used as a tool for selecting unlabeled samples with reliable pseudo-labels. Compared with uncertainty-based SSL methods implemented with Monte Carlo (MC) dropout, NP-Match estimates uncertainty with much less computational overhead, which can save time at both the training and the testing phases. We conducted extensive experiments on four public datasets, and NP-Match outperforms state-of-theart (SOTA) results or achieves competitive results on them, which shows the effectiveness of NPMatch and its potential for SSL.
Thomas Lukasiewicz, Daniela Massiceti, Xiaolin Hu 0001, Vladimir Pavlovic 0001, Alexandros Neophytou
ICML5
2022 DAReN: A Collaborative Approach Towards Visual Reasoning And Disentangling
abstract
Computational learning approaches to solving visual reasoning tests, such as Raven’s Progressive Matrices (RPM), critically depend on the ability to identify the visual concepts used in the test (i.e., the representation) as well as the latent rules based on those concepts (i.e., the reasoning). However, learning of representation and reasoning is a challenging and ill-posed task, often approached in a stage-wise manner (first representation, then reasoning). In this work, we propose an end-to-end joint representation-reasoning learning framework, which leverages a weak form of inductive bias to improve both tasks together. Specifically, we introduce a general generative graphical model for RPMs, GM-RPM, and apply it to solve the reasoning test. We accomplish this using a novel learning framework Disentangling based Abstract Reasoning Network (DAReN) based on the principles of GM-RPM. We perform an empirical evaluation of DAReN over several benchmark datasets. DAReN shows consistent improvement over state-of-the-art (SOTA) models on both the reasoning and the disentanglement tasks. This demonstrates the strong correlation between disentangled latent representation and the ability to solve abstract visual reasoning tasks.
Pritish Sahu, Kalliopi Basioti, Vladimir Pavlovic 0001
ICPR3
2022 Harnessing Fourier Isovists and Geodesic Interaction for Long-Term Crowd Flow Prediction
abstract
With the rise in popularity of short-term Human Trajectory Prediction (HTP), Long-Term Crowd Flow Prediction (LTCFP) has been proposed to forecast crowd movement in large and complex environments. However, the input representations, models, and datasets for LTCFP are currently limited. To this end, we propose Fourier Isovists, a novel input representation based on egocentric visibility, which consistently improves all existing models. We also propose GeoInteractNet (GINet), which couples the layers between a multi-scale attention network (M-SCAN) and a convolutional encoder-decoder network (CED). M-SCAN approximates a super-resolution map of where humans are likely to interact on the way to their goals and produces multi-scale attention maps. The CED then uses these maps in either its encoder's inputs or its decoder's attention gates, which allows GINet to produce super-resolution predictions with substantially higher accuracy than existing models even with Fourier Isovists. In order to evaluate the scalability of models to large and complex environments, which the only existing LTCFP dataset is unsuitable for, a new synthetic crowd dataset with both real and synthetic environments has been generated. In its nascent state, LTCFP has much to gain from our key contributions. The Supplementary Materials, dataset, and code are available at sssohn.github.io/GeoInteractNet.
Samuel S. Sohn, Seonghyeon Moon, Honglu Zhou, Mihee Lee, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia
IJCAI6
2022 SAViR-T: Spatially Attentive Visual Reasoning with Transformers
Pritish Sahu, Kalliopi Basioti, Vladimir Pavlovic 0001
ECML/PKDD (3)3
2022 A2X: An end-to-end framework for assessing agent and environment interactions in multimodal human trajectory prediction
Samuel S. Sohn, Mihee Lee, Seonghyeon Moon, Gang Qiao, Muhammad Usman 0010, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia
Comput. Graph.7
2022 Learning Continuous Facial Actions From Speech for Real-Time Animation
abstract
Speech conveys not only the verbal communication, but also emotions, manifested as facial expressions of the speaker. In this article, we present deep learning frameworks that directly infer facial expressions from just speech signals. Specifically, the time-varying contextual non-linear mapping between audio stream and micro facial movements is realized by our proposed recurrent neural networks to drive a 3D blendshape face model in real-time. Our models not only activate appropriate facial action units (AUs), defined as 3D expression blendshapes in the FaceWarehouse database, to depict different utterance generating actions in the form of lip movements, but also, without any assumption, automatically estimate emotional intensity of the speaker and reproduces her ever-changing affective states by adjusting strength of related facial unit activations. In the baseline models, conventional handcrafted acoustic features are utilized to predict facial actions. Furthermore, we show that it is more advantageous to learn meaningful acoustic feature representation from speech spectrograms with convolutional nets, which subsequently improves the accuracy of facial action synthesis. Experiments on diverse audiovisual corpora of different actors across a wide range of facial actions and emotional states show promising results of our approaches. Being speaker-independent, our generalized models are readily applicable to various tasks in human-machine interaction and animation.
Hai Xuan Pham, Yuting Wang 0004, Vladimir Pavlovic 0001
IEEE Trans. Affect. Comput.3
2021 CHEF: Cross-modal Hierarchical Embeddings for Food Domain Retrieval
abstract
Despite the abundance of multi-modal data, such as image-text pairs, there has been little effort in understanding the individual entities and their different roles in the construction of these data instances. In this work, we endeavour to discover the entities and their corresponding importance in cooking recipes automatically as a visual-linguistic association problem. More specifically, we introduce a novel cross-modal learning framework to jointly model the latent representations of images and text in the food image-recipe association and retrieval tasks. This model allows one to discover complex functional and hierarchical relationships between images and text, and among textual parts of a recipe including title, ingredients and cooking instructions. Our experiments show that by making use of efficient tree-structured Long Short-Term Memory as the text encoder in our computational cross-modal retrieval framework, we are not only able to identify the main ingredients and cooking actions in the recipe descriptions without explicit supervision, but we can also learn more meaningful feature representations of food recipes, appropriate for challenging cross-modal retrieval and recipe adaption tasks.
Hai Xuan Pham, Ricardo Guerrero, Vladimir Pavlovic 0001, Jiatong Li 0001
AAAI3
2021 Multi-attribute Pizza Generator: Cross-domain Attribute Control with Conditional StyleGAN
Fangda Han, Guoyao Hao, Ricardo Guerrero, Vladimir Pavlovic 0001
BMVC4
2021 A2X: An Agent and Environment Interaction Benchmark for Multimodal Human Trajectory Prediction
abstract
In recent years, human trajectory prediction (HTP) has garnered attention in computer vision literature. Although this task has much in common with the longstanding task of crowd simulation, there is little from crowd simulation that has been borrowed, especially in terms of evaluation protocols. The key difference between the two tasks is that HTP is concerned with forecasting multiple steps at a time and capturing the multimodality of real human trajectories. A majority of HTP models are trained on the same few datasets, which feature small, transient interactions between real people and little to no interaction between people and the environment. Unsurprisingly, when tested on crowd egress scenarios, these models produce erroneous trajectories that accelerate too quickly and collide too frequently, but the metrics used in HTP literature cannot convey these particular issues. To address these challenges, we propose (1) the A2X dataset, which has simulated crowd egress and complex navigation scenarios that compensate for the lack of agent-to-environment interaction in existing real datasets, and (2) evaluation metrics that convey model performance with more reliability and nuance. A subset of these metrics are novel multiverse metrics, which are better-suited for multimodal models than existing metrics. The dataset is available at: https://mubbasir.github.io/HTP-benchmark/.
Samuel S. Sohn, Mihee Lee, Seonghyeon Moon, Gang Qiao, Muhammad Usman 0010, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia
MIG7
2021 Cross-modal Retrieval and Synthesis (X-MRS): Closing the Modality Gap in Shared Subspace Learning
abstract
Computational food analysis (CFA) naturally requires multi-modal evidence of a particular food, e.g., images, recipe text, etc. A key to making CFA possible is multi-modal shared representation learning, which aims to create a joint representation of the multiple views (text and image) of the data. In this work we propose a method for food domain cross-modal shared representation learning that preserves the vast semantic richness present in the food data. Our proposed method employs an effective transformer-based multilingual recipe encoder coupled with a traditional image embedding architecture. Here, we propose the use of imperfect multilingual translations to effectively regularize the model while at the same time adding functionality across multiple languages and alphabets. Experimental analysis on the public Recipe1M dataset shows that the representation learned via the proposed method significantly outperforms the current state-of-the-arts (SOTA) on retrieval tasks. Furthermore, the representational power of the learned representation is demonstrated through a generative food image synthesis model conditioned on recipe embeddings. Synthesized images can effectively reproduce the visual appearance of paired samples, indicating that the learned representation captures the joint semantics of both the textual recipe and its visual content, thus narrowing the modality gap.
Ricardo Guerrero, Hai Xuan Pham, Vladimir Pavlovic 0001
ACM Multimedia3
2021 Learning Disentangled Factors from Paired Data in Cross-Modal Retrieval: An Implicit Identifiable VAE Approach
abstract
We tackle the problem of learning the underlying disentangled latent factors that are shared between the paired bi-modal data in cross-modal retrieval. Typically the data in both modalities are complex, structured, and high dimensional (e.g., image and text), for which the conventional deep auto-encoding latent variable models such as the Variational Autoencoder (VAE) often suffer from difficulty of accurate decoder training or realistic synthesis. In this paper we propose a novel idea of the implicit decoder, which completely removes the ambient data decoding module from a latent variable model, via implicit encoder inversion that is achieved by Jacobian regularization of the low-dimensional embedding function. Motivated from the recent Identifiable-VAE (IVAE) model, we modify it to incorporate the query modality data as conditioning auxiliary input, which allows us to prove that the true parameters of the model can be identifiable under some regularity conditions. Tested on various datasets where the true factors are fully/partially available, our model is shown to identify the factors accurately, significantly outperforming conventional latent variable models.
Minyoung Kim 0001, Ricardo Guerrero, Vladimir Pavlovic 0001
ACM Multimedia3
2020 Laying the Foundations of Deep Long-Term Crowd Flow Prediction
Samuel S. Sohn, Honglu Zhou, Seonghyeon Moon, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia
ECCV (29)5
2020 Ordinal-Content VAE: Isolating Ordinal-Valued Content Factors in Deep Latent Variable Models
abstract
In deep representational learning, it is often desired to isolate a particular factor (termed content) from other factors (referred to as style). What constitutes the content is typically specified by users through explicit labels in the data, while all unlabeled/unknown factors are regarded as style. Recently, it has been shown that such content-labeled data can be effectively exploited by modifying the deep latent factor models (e.g., VAE) such that the style and content are well separated in the latent representations. However, the approach assumes that the content factor is categorical-valued (e.g., subject ID in face image data, or digit class in the MNIST dataset). In certain situations, the content is ordinal-valued, that is, the values the content factor takes are ordered rather than categorical, making content-labeled VAEs, including the latent space they infer, suboptimal. In this paper, we propose a novel extension of VAE that imposes a partially ordered set (poset) structure in the content latent space, while simultaneously making it aligned with the ordinal content values. To this end, instead of the iid Gaussian latent prior adopted in prior approaches, we introduce a conditional Gaussian spacing prior model. This model admits a tractable joint Gaussian prior, but also effectively places negligible density values on the content latent configurations that violate the poset constraint. To evaluate this model, we consider two specific ordinal structured problems: estimating a subject's age in a face image and elucidating the calorie amount in a food meal image. We demonstrate significant improvements in content-style separation over previous non-ordinal approaches.
Minyoung Kim 0001, Vladimir Pavlovic 0001
ICDM2
2020 Picture-to-Amount (PITA): Predicting Relative Ingredient Amounts from Food Images
abstract
Increased awareness of the impact of food consumption on health and lifestyle today has given rise to novel data-driven food analysis systems. Although these systems may recognize the ingredients, a detailed analysis of their amounts in the meal, which is paramount for estimating the correct nutrition, is usually ignored. In this paper, we study the novel and challenging problem of predicting the relative amount of each ingredient from a food image. We propose PITA, the Picture-to-Amount deep learning architecture to solve the problem. More specifically, we predict the ingredient amounts using a domain-driven Wasserstein loss from image-to-recipe cross-modal embeddings learned to align the two views of food data. Experiments on a dataset of recipes collected from the Internet show the model generates promising results and improves the baselines on this challenging task. A demo of our system and our data is available at: foodai.cs.rutgers.edu.
Jiatong Li 0001, Fangda Han, Ricardo Guerrero, Vladimir Pavlovic 0001
ICPR4
2020 Recursive Inference for Variational Autoencoders
abstract
Inference networks of traditional Variational Autoencoders (VAEs) are typically amortized, resulting in relatively inaccurate posterior approximation compared to instance-wise variational optimization. Recent semi-amortized approaches were proposed to address this drawback; however, their iterative gradient update procedures can be computationally demanding. In this paper, we consider a different approach of building a mixture inference model. We propose a novel recursive mixture estimation algorithm for VAEs that iteratively augments the current mixture with new components so as to maximally reduce the divergence between the variational and the true posteriors. Using the functional gradient approach, we devise an intuitive learning criteria for selecting a new mixture component: the new component has to improve the data likelihood (lower bound) and, at the same time, be as divergent from the current mixture distribution as possible, thus increasing representational diversity. Although there have been similar approaches recently, termed boosted variational inference (BVI), our methods differ from BVI in several aspects, most notably that ours deal with recursive inference in VAEs in the form of amortized inference, while BVI is developed within the standard VI framework, leading to a non-amortized single optimization instance, inappropriate for VAEs. A crucial benefit of our approach is that the inference at test time needs a single feed-forward pass through the mixture inference network, making it significantly faster than the semi-amortized approaches. We show that our approach yields higher test data likelihood than the state-of-the-arts on several benchmark datasets.
Minyoung Kim 0001, Vladimir Pavlovic 0001
NeurIPS2
2020 CookGAN: Meal Image Synthesis from Ingredients
abstract
In this work we propose a new computational framework, based on generative deep models, for synthesis of photo-realistic food meal images from textual list of its ingredients. Previous works on synthesis of images from text typically rely on pre-trained text models to extract text features, followed by generative neural networks (GAN) aimed to generate realistic images conditioned on the text features. These works mainly focus on generating spatially compact and well-defined categories of objects, such as birds or flowers, but meal images are significantly more complex, consisting of multiple ingredients whose appearance and spatial qualities are further modified by cooking methods. To generate real-like meal images from ingredients, we propose Cook Generative Adversarial Networks (CookGAN), CookGAN first builds an attention-based ingredients-image association model, which is then used to condition a generative neural network tasked with synthesizing meal images. Furthermore, a cycle-consistent constraint is added to further improve image quality and control appearance. Experiments show our model is able to generate meal images corresponding to the ingredients.
Fangda Han, Ricardo Guerrero, Vladimir Pavlovic 0001
WACV3
2020 Predicting Crowd Egress and Environment Relationships to Support Building Design Optimization
Kaidong Hu, Sejong Yoon, Vladimir Pavlovic 0001, Petros Faloutsos, Mubbasir Kapadia
Comput. Graph.3
2020 Unsupervised Multi-Target Domain Adaptation: An Information Theoretic Approach
abstract
Unsupervised domain adaptation (uDA) models focus on pairwise adaptation settings where there is a single, labeled, source and a single target domain. However, in many real-world settings one seeks to adapt to multiple, but somewhat similar, target domains. Applying pairwise adaptation approaches to this setting may be suboptimal, as they fail to leverage shared information among multiple domains. In this work, we propose an information theoretic approach for domain adaptation in the novel context of multiple target domains with unlabeled instances and one source domain with labeled instances. Our model aims to find a shared latent space common to all domains, while simultaneously accounting for the remaining private, domain-specific factors. Disentanglement of shared and private information is accomplished using a unified information-theoretic approach, which also serves to establish a stronger link between the latent representations and the observed data. The resulting model, accompanied by an efficient optimization algorithm, allows simultaneous adaptation from a single source to multiple target domains. We test our approach on three challenging publicly-available datasets, showing that it outperforms several popular domain adaptation methods.
Behnam Gholami, Pritish Sahu, Ognjen Rudovic, Konstantinos Bousmalis, Vladimir Pavlovic 0001
IEEE Trans. Image Process.5
2020 Guest Editorial Special Issue on Structured Multi-Output Learning: Modeling, Algorithm, Theory, and Applications
abstract
Structured multioutput learning is a topic in artificial intelligence that considers multiple structured outputs prediction for a given input. The output may involve structured objects in the form of sequence, string, tree, lattice, or graph and has values that are characterized by diverse data types, such as binary, nominal, ordinal, and real-valued variables. Such learning problems arise in a variety of real-world applications, ranging from document classification, computer emulation, sensor network analysis, concept-based information retrieval, and human action/causal induction to video analysis, image annotation/retrieval, gene function prediction, and brain science. As many complex real-world scenarios can be posed as a structured multioutput learning problem, their importance and popularity have been increasing steadily.
Weiwei Liu 0003, Xiaobo Shen 0001, Yew-Soon Ong, Ivor W. Tsang, Chen Gong 0002, Vladimir Pavlovic 0001
IEEE Trans. Neural Networks Learn. Syst.6
2019 Unsupervised Visual Domain Adaptation: A Deep Max-Margin Gaussian Process Approach
abstract
For unsupervised domain adaptation, the target domain error can be provably reduced by having a shared input representation that makes the source and target domains indistinguishable from each other. Very recently it has been shown that it is not only critical to match the marginal input distributions, but also align the output class distributions. The latter can be achieved by minimizing the maximum discrepancy of predictors. In this paper, we take this principle further by proposing a more systematic and effective way to achieve hypothesis consistency using Gaussian processes (GP). The GP allows us to induce a hypothesis space of classifiers from the posterior distribution of the latent random functions, turning the learning into a large-margin posterior separation problem, significantly easier to solve than previous approaches based on adversarial minimax optimization. We formulate a learning objective that effectively influences the posterior to minimize the maximum discrepancy. This is shown to be equivalent to maximizing margins and minimizing uncertainty of the class predictions in the target domain. Empirical results demonstrate that our approach leads to state-to-the-art performance superior to existing methods on several challenging benchmarks for domain adaptation.
Minyoung Kim 0001, Pritish Sahu, Behnam Gholami, Vladimir Pavlovic 0001
CVPR4
2019 Bayes-Factor-VAE: Hierarchical Bayesian Deep Auto-Encoder Models for Factor Disentanglement
abstract
We propose a family of novel hierarchical Bayesian deep auto-encoder models capable of identifying disentangled factors of variability in data. While many recent attempts at factor disentanglement have focused on sophisticated learning objectives within the VAE framework, their choice of a standard normal as the latent factor prior is both suboptimal and detrimental to performance. Our key observation is that the disentangled latent variables responsible for major sources of variability, the relevant factors, can be more appropriately modeled using long-tail distributions. The typical Gaussian priors are, on the other hand, better suited for modeling of nuisance factors. Motivated by this, we extend the VAE to a hierarchical Bayesian model by introducing hyper-priors on the variances of Gaussian latent priors, mimicking an infinite mixture, while maintaining tractable learning and inference of the traditional VAEs. This analysis signifies the importance of partitioning and treating in a different manner the latent dimensions corresponding to relevant factors and nuisances. Our proposed models, dubbed Bayes-Factor-VAEs, are shown to outperform existing methods both quantitatively and qualitatively in terms of latent disentanglement across several challenging benchmark tasks.
Minyoung Kim 0001, Yuting Wang 0004, Pritish Sahu, Vladimir Pavlovic 0001
ICCV4
2019 Efficient Deep Gaussian Process Models for Variable-Sized Inputs
abstract
Deep Gaussian processes (DGP) have appealing Bayesian properties, can handle variable-sized data, and learn deep features. Their limitation is that they do not scale well with the size of the data. Existing approaches address this using a deep random feature (DRF) expansion model, which makes inference tractable by approximating DGPs. However, DRF is not suitable for variable-sized input data such as trees, graphs, and sequences. We introduce the GP-DRF, a novel Bayesian model with an input layer of GPs, followed by DRF layers. The key advantage is that the combination of GP and DRF leads to a tractable model that can both handle a variable-sized input as well as learn deep long-range dependency structures of the data. We provide a novel efficient method to simultaneously infer the posterior of GP's latent vectors and infer the posterior of DRF's internal weights and random frequencies. Our experiments show that GP-DRF outperforms the standard GP model and DRF model across many datasets. Furthermore, they demonstrate that GP-DRF enables improved uncertainty quantification compared to GP and DRF alone, with respect to a Bhattacharyya distance assessment.
Issam H. Laradji, Mark Schmidt 0001, Vladimir Pavlovic 0001, Minyoung Kim 0001
IJCNN3
2019 Scenario Generalization of Data-driven Imitation Models in Crowd Simulation
abstract
Crowd simulation, the study of the movement of multiple agents in complex environments, presents a unique application domain for machine learning. One challenge in crowd simulation is to imitate the movement of expert agents in highly dense crowds. An imitation model could substitute an expert agent if the model behaves as good as the expert. This will bring many exciting applications. However, we believe no prior studies have considered the critical question of how training data and training methods affect imitators when these models are applied to novel scenarios. In this work, a general imitation model is represented by applying either the Behavior Cloning (BC) training method or a more sophisticated Generative Adversarial Imitation Learning (GAIL) method, on three typical types of data domains: standard benchmarks for evaluating crowd models, random sampling of state-action pairs, and egocentric scenarios that capture local interactions. Simulated results suggest that (i) simpler training methods are overall better than more complex training methods, (ii) training samples with diverse agent-agent and agent-obstacle interactions are beneficial for reducing collisions when the trained models are applied to new scenarios. We additionally evaluated our models in their ability to imitate real world crowd trajectories observed from surveillance videos. Our findings indicate that models trained on representative scenarios generalize to new, unseen situations observed in real human crowds.
Gang Qiao, Honglu Zhou, Mubbasir Kapadia, Sejong Yoon, Vladimir Pavlovic 0001
MIG5
2019 Cartoonish sketch-based face editing in videos using identity deformation transfer
Long Zhao 0003, Fangda Han, Xi Peng 0005, Mubbasir Kapadia, Vladimir Pavlovic 0001, Dimitris N. Metaxas
Comput. Graph.6
2019 Visibility Constrained Generative Model for Depth-Based 3D Facial Pose Tracking
abstract
In this paper, we propose a generative framework that unifies depth-based 3D facial pose tracking and face model adaptation on-the-fly, in the unconstrained scenarios with heavy occlusions and arbitrary facial expression variations. Specifically, we introduce a statistical 3D morphable model that flexibly describes the distribution of points on the surface of the face model, with an efficient switchable online adaptation that gradually captures the identity of the tracked subject and rapidly constructs a suitable face model when the subject changes. Moreover, unlike prior art that employed ICP-based facial pose estimation, to improve robustness to occlusions, we propose a ray visibility constraint that regularizes the pose based on the face model's visibility with respect to the input point cloud. Ablation studies and experimental results on Biwi and ICT-3DHP datasets demonstrate that the proposed framework is effective and outperforms completing state-of-the-art depth-based methods.
Lu Sheng, Jianfei Cai 0001, Tat-Jen Cham, Vladimir Pavlovic 0001, King Ngi Ngan
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 Copula Ordinal Regression Framework for Joint Estimation of Facial Action Unit Intensity
abstract
Joint modeling of the intensity of multiple facial action units (AUs) from face images is challenging due to the large number of AUs (30+) and their intensity levels (6). This is in part due to the lack of suitable models that can efficiently handle such a large number of outputs/classes simultaneously, but also due to the lack of suitable data the models on. For this reason, majority of the methods resort to independent classifiers for the AU intensity. This is suboptimal for at least two reasons: the facial appearance of some AUs changes depending on the intensity of other AUs, and some AUs co-occur more often than others. To this end, we propose the Copula regression approach for modeling multivariate ordinal variables. Our model accounts for ordinal structure in output variables and their non-linear dependencies via copula functions modeled as cliques of a conditional random fields. The copula ordinal regression model achieves the joint learning and inference of intensities of multiple AUs, while being computationally tractable. We demonstrate the effectiveness of our approach on three challenging datasets of naturalistic facial expressions and we show that the estimation of target AU intensities improves especially in the case of (a) noisy image features, (b) head-pose variations and (c) imbalanced training data. Lastly, we show that the proposed approach consistently outperforms (i) independent modeling of AU intensities and (ii) the state-of-the-art approach for the target task and (iii) deep convolutional neural networks.
Robert Walecki, Ognjen Rudovic, Vladimir Pavlovic 0001, Maja Pantic
IEEE Trans. Affect. Comput.3
2018 The Role of Data-Driven Priors in Multi-Agent Crowd Trajectory Estimation
abstract
Resource constraints frequently complicate multi-agent planning problems. Existing algorithms for resource-constrained, multi-agent planning problems rely on the assumption that the constraints are deterministic. However, frequently resource constraints are themselves subject to uncertainty from external influences. Uncertainty about constraints is especially challenging when agents must execute in an environment where communication is unreliable, making on-line coordination difficult. In those cases, it is a significant challenge to find coordinated allocations at plan time depending on availability at run time. To address these limitations, we propose to extend algorithms for constrained multi-agent planning problems to handle stochastic resource constraints. We show how to factorize resource limit uncertainty and use this to develop novel algorithms to plan policies for stochastic constraints. We evaluate the algorithms on a search-and-rescue problem and on a power-constrained planning domain where the resource constraints are decided by nature. We show that plans taking into account all potential realizations of the constraint obtain significantly better utility than planning for the expectation, while causing fewer constraint violations.
Gang Qiao, Sejong Yoon, Mubbasir Kapadia, Vladimir Pavlovic 0001
AAAI4
2018 End-to-end Learning for 3D Facial Animation from Speech
abstract
We present a deep learning framework for real-time speech-driven 3D facial animation from speech audio. Our deep neural network directly maps an input sequence of speech spectrograms to a series of micro facial action unit intensities to drive a 3D blendshape face model. In particular, our deep model is able to learn the latent representations of time-varying contextual information and affective states within the speech. Hence, our model not only activates appropriate facial action units at inference to depict different utterance generating actions, in the form of lip movements, but also, without any assumption, automatically estimates emotional intensity of the speaker and reproduces her ever-changing affective states by adjusting strength of related facial unit activations. For example, in a happy speech, the mouth opens wider than normal, while other facial units are relaxed; or both eyebrows raise higher in a surprised state. Experiments on diverse audiovisual corpora of different actors across a wide range of facial actions and emotional states show promising results of our approach. Being speaker-independent, our generalized model is readily applicable to various tasks in human-machine interaction and animation.
Hai Xuan Pham, Yuting Wang 0004, Vladimir Pavlovic 0001
ICMI3
2018 Variational Inference for Gaussian Process Models for Survival Analysis
Minyoung Kim 0001, Vladimir Pavlovic 0001
UAI2
2017 Probabilistic Temporal Subspace Clustering
abstract
Subspace clustering is a common modeling paradigm used to identify constituent modes of variation in data with locally linear structure. These structures are common to many problems in computer vision, including modeling time series of complex human motion. However classical subspace clustering algorithms learn the relationships within a set of data without considering the temporal dependency and then use a separate clustering step (e.g., spectral clustering) for final segmentation. Moreover, these, frequently optimization-based, algorithms assume that all observations have complete features. In contrast in real-world applications, some features are often missing, which results in incomplete data and substantial performance degeneration of these approaches. In this paper, we propose a unified non-parametric generative framework for temporal subspace clustering to segment data drawn from a sequentially ordered union of subspaces that deals with the missing features in a principled way. The non-parametric nature of our generative model makes it possible to infer the number of subspaces and their dimension automatically from data. Experimental results on human action datasets demonstrate that the proposed model consistently outperforms other state-of-the-art subspace clustering approaches.
Behnam Gholami, Vladimir Pavlovic 0001
CVPR2
2017 A Generative Model for Depth-Based Robust 3D Facial Pose Tracking
abstract
We consider the problem of depth-based robust 3D facial pose tracking under unconstrained scenarios with heavy occlusions and arbitrary facial expression variations. Unlike the previous depth-based discriminative or data-driven methods that require sophisticated training or manual intervention, we propose a generative framework that unifies pose tracking and face model adaptation on-the-fly. Particularly, we propose a statistical 3D face model that owns the flexibility to generate and predict the distribution and uncertainty underlying the face model. Moreover, unlike prior arts employing the ICP-based facial pose estimation, we propose a ray visibility constraint that regularizes the pose based on the face models visibility against the input point cloud, which augments the robustness against the occlusions. The experimental results on Biwi and ICT-3DHP datasets reveal that the proposed framework is effective and outperforms the state-of-the-art depth-based methods.
Lu Sheng, Jianfei Cai 0001, Tat-Jen Cham, Vladimir Pavlovic 0001, King Ngi Ngan
CVPR4
2017 Deep Structured Learning for Facial Action Unit Intensity Estimation
abstract
We consider the task of automated estimation of facial expression intensity. This involves estimation of multiple output variables (facial action units - AUs) that are structurally dependent. Their structure arises from statistically induced co-occurrence patterns of AU intensity levels. Modeling this structure is critical for improving the estimation performance, however, this performance is bounded by the quality of the input features extracted from face images. The goal of this paper is to model these structures and estimate complex feature representations simultaneously by combining conditional random field (CRF) encoded AU dependencies with deep learning. To this end, we propose a novel Copula CNN deep learning approach for modeling multivariate ordinal variables. Our model accounts for ordinal structure in output variables and their non-linear dependencies via copula functions modeled as cliques of a CRF. These are jointly optimized with deep CNN feature encoding layers using a newly introduced balanced batch iterative training algorithm. We demonstrate the effectiveness of our approach on the task of AU intensity estimation on two benchmark datasets. We show that joint learning of the deep features and the target output structure results in significant performance gains compared to existing structured deep models and deep models for analysis of facial expressions.
Robert Walecki, Ognjen Rudovic, Vladimir Pavlovic 0001, Björn W. Schuller, Maja Pantic
CVPR3
2017 PUnDA: Probabilistic Unsupervised Domain Adaptation for Knowledge Transfer Across Visual Categories
abstract
This paper introduces a probabilistic latent variable model to address unsupervised domain adaptation problems. Specifically, we tackle the task of categorization of visual input from different domains by learning projections from each domain to a latent (shared) space jointly with the classifier in the latent space, which simultaneously minimizes the domain disparity while maximizing the classifier's discriminative power. Furthermore, the non-parametric nature of our adaptation model makes it possible to infer the latent space dimension automatically from data. We also develop a novel regularized Variational Bayes (VB) algorithm for efficient estimation of the model parameters. We compare the proposed model with the state-of-the-art methods for the tasks of visual domain adaptation using both handcrafted and deep-net features. Our experiments show that even with a simple softmax classifier, our model outperforms several state-of-the-art methods that take advantage of more sophisticated classification schemes.
Behnam Gholami, Ognjen Rudovic, Vladimir Pavlovic 0001
ICCV3
2017 Characterizing the relationship between environment layout and crowd movement using machine learning
abstract
Crowd simulations facilitate the study of how an environment layout impacts the movement and behavior of its inhabitants. However, simulations are computationally expensive, which make them infeasible when used as part of interactive systems (e.g., Computer-Assisted Design software). Machine learning models, such as neural networks (NN), can learn observed behaviors from examples, and can potentially offer a rational prediction of a crowd's behavior efficiently. To this end, we propose a method to predict the aggregate characteristics of crowd dynamics using regression neural networks (NN). We parametrize the environment, the crowd distribution and the steering method to serve as inputs to the NN models, while a number of common performance measures serve as the output. Our preliminary experiments show that our approach can help users evaluate a large number of environments efficiently.
Weining Liu, Vladimir Pavlovic 0001, Kaidong Hu, Petros Faloutsos, Sejong Yoon, Mubbasir Kapadia
MIG2
2017 Variable-state Latent Conditional Random Field models for facial expression analysis
Robert Walecki, Ognjen Rudovic, Vladimir Pavlovic 0001, Maja Pantic
Image Vis. Comput.3
2017 Using 3D face priors for depth recovery
Chongyu Chen, Hai Xuan Pham, Vladimir Pavlovic 0001, Jianfei Cai 0001, Guangming Shi, Yuefang Gao
J. Vis. Commun. Image Represent.3
2016 Robust Real-Time 3D Face Tracking from RGBD Videos under Extreme Pose, Depth, and Expression Variation
abstract
We introduce a novel end-to-end real-time pose-robust 3D face tracking framework from RGBD videos, which is capable of tracking head pose and facial actions simultaneously in unconstrained environment without intervention or pre-calibration from a user. In particular, we emphasize tracking the head pose from profile to profile and improving tracking performance in challenging instances, where the tracked subject is at a considerably large distance from the camera and the quality of data deteriorates severely. To achieve these goals, the tracker is guided by an efficient multi-view 3D shape regressor, trained upon generic RGB datasets, which is able to predict model parameters despite large head rotations or tracking range. Specifically, the shape regressor is made aware of the head pose by inferring the possibility of particular facial landmarks being visible through a joint regression-classification local random forest framework, and piecewise linear regression models effectively map visibility features into shape parameters. In addition, the regressor is combined with a joint 2D+3D optimization that sparsely exploits depth information to further refine shape parameters to maintain tracking accuracy over time. The result is a robust on-line RGBD 3D face tracker that can model extreme head poses and facial expressions accurately in challenging scenes, which are demonstrated in our extensive experiments.
Hai Xuan Pham, Vladimir Pavlovic 0001
3DV2
2016 Decentralized Approximate Bayesian Inference for Distributed Sensor Network
abstract
Bayesian models provide a framework for probabilistic modelling of complex datasets. Many such models are computationally demanding, especially in the presence of large datasets. In sensor network applications, statistical (Bayesian) parameter estimation usually relies on decentralized algorithms, in which both data and computation are distributed across the nodes of the network. In this paper we propose a framework for decentralized Bayesian learning using Bregman Alternating Direction Method of Multipliers (B-ADMM). We demonstrate the utility of our framework, with Mean Field Variational Bayes (MFVB) as the primitive for distributed affine structure from motion (SfM).
Behnam Gholami, Sejong Yoon, Vladimir Pavlovic 0001
AAAI3
2016 Fast ADMM Algorithm for Distributed Optimization with Adaptive Penalty
abstract
We propose new methods to speed up convergence of the Alternating Direction Method of Multipliers (ADMM), a common optimization tool in the context of large scale and distributed learning. The proposed method accelerates the speed of convergence by automatically deciding the constraint penalty needed for parameter consensus in each iteration. In addition, we also propose an extension of the method that adaptively determines the maximum number of iterations to update the penalty. We show that this approach effectively leads to an adaptive, dynamic network topology underlying the distributed optimization. The utility of the new penalty update schemes is demonstrated on both synthetic and real data, including an instance of the probabilistic matrix factorization task known as the structure from motion problem.
Changkyu Song, Sejong Yoon, Vladimir Pavlovic 0001
AAAI3
2016 Copula Ordinal Regression for Joint Estimation of Facial Action Unit Intensity
abstract
Joint modeling of the intensity of facial action units (AUs) from face images is challenging due to the large number of AUs (30+) and their intensity levels (6). This is in part due to the lack of suitable models that can efficiently handle such a large number of outputs/classes simultaneously, but also due to the lack of labelled target data. For this reason, majority of the methods proposed so far resort to independent classifiers for the AU intensity. This is suboptimal for at least two reasons: the facial appearance of some AUs changes depending on the intensity of other AUs, and some AUs co-occur more often than others. Encoding this is expected to improve the estimation of target AU intensities, especially in the case of noisy image features, head-pose variations and imbalanced training data. To this end, we introduce a novel modeling framework, Copula Ordinal Regression (COR), that leverages the power of copula functions and CRFs, to detangle the probabilistic modeling of AU dependencies from the marginal modeling of the AU intensity. Consequently, the COR model achieves the joint learning and inference of intensities of multiple AUs, while being computationally tractable. We show on two challenging datasets of naturalistic facial expressions that the proposed approach consistently outperforms (i) independent modeling of AU intensities, and (ii) the state-of the-art approach for the target task.
Robert Walecki, Ognjen Rudovic, Vladimir Pavlovic 0001, Maja Pantic
CVPR3
2016 A Shape-Based Approach for Salient Object Detection Using Deep Learning
Jongpil Kim, Vladimir Pavlovic 0001
ECCV (4)2
2016 A shape preserving approach for salient object detection using convolutional neural networks
abstract
Determining visual saliency is one of the fundamental problems in computer vision as the saliency not only identifies the most informative parts of a visual scene but may also reduce computational complexity by filtering out irrelevant segments of the scene. In this paper, we propose a novel saliency object detection method that combines a shape-preserving saliency prediction driven by a convolutional neural network with the mid and low-level region preserving image information. Our model learns a saliency shape dictionary, which is subsequently used to train a CNN to predict the salient class of a target region and estimate the full but coarse saliency map of the target image. The map is then refined using image specific low-to-mid level information. Performance evaluation on popular benchmark datasets shows that the proposed method outperforms existing state-of-the-art methods in saliency detection.
Jongpil Kim, Vladimir Pavlovic 0001
ICPR2
2016 Discovering characteristic landmarks on ancient coins using convolutional networks
abstract
We propose a novel method to find characteristic landmarks and recognize ancient Roman imperial coins using deep convolutional neural networks (CNNs) combined with expert-designed domain hierarchies. We first propose a new framework to recognize the Roman coin which exploits the hierarchical knowledge structure embedded in the coin domain, which we combine with the CNN-based category classifiers. We next formulate an optimization problem to discover class-specific salient coin regions. Analysis of discovered salient regions confirms that they are largely consistent with human expert annotations. Experimental results show that the proposed framework is able to effectively recognize the ancient Roman coins as well as successfully identify landmarks in a general fine-grained classification problem. For this research, we have collected a new Roman coin dataset where all coins are annotated and consist of obverse (head) and reverse (tail) images.
Jongpil Kim, Vladimir Pavlovic 0001
ICPR2
2016 Robust real-time performance-driven 3D face tracking
abstract
We introduce a novel robust hybrid 3D face tracking framework from RGBD video streams, which is capable of tracking head pose and facial actions without pre-calibration or intervention from a user. In particular, we emphasize on improving the tracking performance in instances where the tracked subject is at a large distance from the cameras, and the quality of point cloud deteriorates severely. This is accomplished by the combination of a flexible 3D shape regressor and the joint 2D+3D optimization on shape parameters. Our approach fits facial blendshapes to the point cloud of the human head, while being driven by an efficient and rapid 3D shape regressor trained on generic RGB datasets. As an on-line tracking system, the identity of the unknown user is adapted on-the-fly resulting in improved 3D model reconstruction and consequently better tracking performance. The result is a robust RGBD face tracker capable of handling a wide range of target scene depths, whose performances are demonstrated in our extensive experiments better than those of the state-of-the-arts.
Hai Xuan Pham, Vladimir Pavlovic 0001, Jianfei Cai 0001, Tat-Jen Cham
ICPR2
2016 Robust time-series retrieval using probabilistic adaptive segmental alignment
Shahriar Shariat, Vladimir Pavlovic 0001
Knowl. Inf. Syst.2
2015 Multi-cue Structure Preserving MRF for Unconstrained Video Segmentation
abstract
Video segmentation is a stepping stone to understanding video context. Video segmentation enables one to represent a video by decomposing it into coherent regions which comprise whole or parts of objects. However, the challenge originates from the fact that most of the video segmentation algorithms are based on unsupervised learning due to expensive cost of pixelwise video annotation and intra-class variability within similar unconstrained video classes. We propose a Markov Random Field model for unconstrained video segmentation that relies on tight integration of multiple cues: vertices are defined from contour based superpixels, unary potentials from temporally smooth label likelihood and pairwise potentials from global structure of a video. Multi-cue structure is a breakthrough to extracting coherent object regions for unconstrained videos in absence of supervision. Our experiments on VSB100 dataset show that the proposed model significantly outperforms competing state-of-the-art algorithms. Qualitative analysis illustrates that video segmentation result of the proposed model is consistent with human perception of objects.
Saehoon Yi, Vladimir Pavlovic 0001
ICCV2
2015 Context-Sensitive Dynamic Ordinal Regression for Intensity Estimation of Facial Action Units
abstract
Modeling intensity of facial action units from spontaneously displayed facial expressions is challenging mainly because of high variability in subject-specific facial expressiveness, head-movements, illumination changes, etc. These factors make the target problem highly context-sensitive. However, existing methods usually ignore this context-sensitivity of the target problem. We propose a novel Conditional Ordinal Random Field (CORF) model for context-sensitive modeling of the facial action unit intensity, where the W5+ (who, when, what, where, why and how) definition of the context is used. While the proposed model is general enough to handle all six context questions, in this paper we focus on the context questions: who (the observed subject), how (the changes in facial expressions), and when (the timing of facial expressions and their intensity). The context questions who and howare modeled by means of the newly introduced context-dependent covariate effects, and the context question when is modeled in terms of temporal correlation between the ordinal outputs, i.e., intensity levels of action units. We also introduce a weighted softmax-margin learning of CRFs from data with skewed distribution of the intensity levels, which is commonly encountered in spontaneous facial data. The proposed model is evaluated on intensity estimation of pain and facial action units using two recently published datasets (UNBC Shoulder Pain and DISFA) of spontaneously displayed facial expressions. Our experiments show that the proposed model performs significantly better on the target tasks compared to the state-of-the-art approaches. Furthermore, compared to traditional learning of CRFs, we show that the proposed weighted learning results in more robust parameter estimation from the imbalanced intensity data.
Ognjen Rudovic, Vladimir Pavlovic 0001, Maja Pantic
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Depth Recovery with Face Priors
Chongyu Chen, Hai Xuan Pham, Vladimir Pavlovic 0001, Jianfei Cai 0001, Guangming Shi
ACCV (4)3
2014 Ancient Coin Recognition Based on Spatial Coding
abstract
Roman coins play an important role to understand the Roman empire because they convey rich information about key historical events of the time. Moreover, as large amounts of coins are daily traded over the Internet, it becomes necessary to develop automatic coin recognition systems to prevent illegal trades. In this paper, we propose an automatic recognition method for ancient Roman coins. The proposed method exploits the structure of the coin by using a spatially local coding method. Results show that the proposed method outperforms traditional rigid spatial structure models such as the spatial pyramid.
Jongpil Kim, Vladimir Pavlovic 0001
ICPR2
2014 Hybrid On-Line 3D Face and Facial Actions Tracking in RGBD Video Sequences
abstract
In this paper, we propose a hybrid model-based tracker for simultaneous tracking of 3D head pose and facial actions in sequences of texture and depth frames. Our tracker utilizes a generic wireframe model, the Candide-3, to represent facial deformations. This wireframe model is initially fit into the first frame by an Iterative Closest Point algorithm. Given the result after the first frame, our tracking algorithm combines both Iterative Closest Point technique and Appearance Model for head pose and facial actions tracking. The tracker is capable of adapting on-line to the changes in appearance of the target and thus the prior training process is avoided. Furthermore, the tracking system works automatically without any intervention from human operators.
Hai Xuan Pham, Vladimir Pavlovic 0001
ICPR2
2014 Pose Invariant Activity Classification for Multi-floor Indoor Localization
abstract
Smartphone based indoor localization caught massive interest of the localization community in recent years. Combining pedestrian dead reckoning obtained using the phone's inertial sensors with the Graph SLAM (Simultaneous Localization and Mapping) algorithm is one of the most effective approaches to reconstruct the entire pedestrian trajectory given a set of visited landmarks during movement. A key to Graph SLAM-based localization is the detection of reliable landmarks, which are typically identified using visual cues or via NFC tags or QR codes. Alternatively, human activity can be classified to detect organic landmarks such as visits to stairs and elevators while in movement. We provide a novel human activity classification framework that is invariant to the pose of the smartphone. Pose invariant features allow robust observation no matter how a user puts the phone in the pocket. In addition, activity classification obtained by an SVM (Support Vector Machine) is used in a Bayesian framework with an HMM (Hidden Markov Model) that improves the activity inference based on temporal smoothness. Furthermore, the HMM jointly infers activity and floor information, thus providing multi-floor indoor localization. Our experiments show that the proposed framework detects landmarks accurately and enables multi-floor indoor localization from the pocket using Graph SLAM.
Saehoon Yi, Piotr Mirowski, Tin Kam Ho, Vladimir Pavlovic 0001
ICPR4
2014 Dynamic Probabilistic CCA for Analysis of Affective Behavior and Fusion of Continuous Annotations
abstract
Fusing multiple continuous expert annotations is a crucial problem in machine learning and computer vision, particularly when dealing with uncertain and subjective tasks related to affective behavior. Inspired by the concept of inferring shared and individual latent spaces in Probabilistic Canonical Correlation Analysis (PCCA), we propose a novel, generative model that discovers temporal dependencies on the shared/individual spaces (Dynamic Probabilistic CCA, DPCCA). In order to accommodate for temporal lags, which are prominent amongst continuous annotations, we further introduce a latent warping process, leading to the DPCCA with Time Warpings (DPCTW) model. Finally, we propose two supervised variants of DPCCA/DPCTW which incorporate inputs (i.e., visual or audio features), both in a generative (SG-DPCCA) and discriminative manner (SD-DPCCA). We show that the resulting family of models (i) can be used as a unifying framework for solving the problems of temporal alignment and fusion of multiple annotations in time, (ii) can automatically rank and filter annotations based on latent posteriors or other model statistics, and (iii) that by incorporating dynamics, modeling annotation-specific biases, noise estimation, time warping and supervision, DPCTW outperforms state-of-the-art methods for both the aggregation of multiple, yet imperfect expert annotations as well as the alignment of affective behavior.
Mihalis A. Nicolaou, Vladimir Pavlovic 0001, Maja Pantic
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 A New Adaptive Segmental Matching Measure for Human Activity Recognition
abstract
The problem of human activity recognition is a central problem in many real-world applications. In this paper we propose a fast and effective segmental alignment-based method that is able to classify activities and interactions in complex environments. We empirically show that such model is able to recover the alignment that leads to improved similarity measures within sequence classes and hence, raises the classification performance. We also apply a bounding technique on the histogram distances to reduce the computation of the otherwise exhaustive search.
Shahriar Shariat, Vladimir Pavlovic 0001
ICCV2
2013 Relative spatial features for image memorability
abstract
Recent studies in image memorability showed that the memorability of an image is a measurable quantity and is closely correlated with semantic attributes. However, the intrinsic characteristics of memorability are not yet fully understood. It has been reported that in contrast to a popular belief unusualness or aesthetic beauty of the image may not be positively correlated with the image memorability. This counter-intuitive characteristic of memorability hinders a better understanding of image memorability and its applicability. In this paper, we investigate two new spatial features that are closely correlated with the image memorability yet intuitively explainable. We propose the Weighted Object Area (WOA) that jointly considers the location and size of objects and the Relative Area Rank (RAR) that captures the relative unusualness of the size of objects. We empirically demonstrate their useful correlation with the image memorability. Results show that both WOA and RAR can improve the memorability prediction. In addition, we provide evidence that the RAR can effectively capture object-centric unusualness of size.
Jongpil Kim, Sejong Yoon, Vladimir Pavlovic 0001
ACM Multimedia3
2012 Structured Learning for Multiple Object Tracking
abstract
Adaptive tracking-by-detection methods use previous tracking results to generate a new training set for object appearance, and update the current model to predict the object location in subsequent frames. Such approaches are typically bootstarpped by manual or semi-automatic initialization in the first several frames. However, most adaptive tracking-bydetection methods focus on tracking of a single object or multiple unrelated objects. Although one can trivially engage several single object trackers to track multiple objects, such solution is frequently suboptimal because it does not utilize the inter-object constraints or the obejct layout information [2]. We propose in this paper an adaptive tracking-by-detection method for multiple objects, inspired by recent work in [1] and [2]. The constraints for structured Support Vector Machine (SVM) in [1] are modified to localize multiple objects simultaneously with both appearance and layout information. Moreover, additional binary constraints are introduced to detect the existences of respective objects and to prevent possible model drift. Thus the method can handle frequent occlusion in multiple object tracking, as well as objects entering or leaving the scene. Those binary constraints make the optimization problem significantly different from the original Structured SVM [3]. The inter-object constraints, embedded in a linear programming technique similar to [2] for optimal position assignment, are applied to diminish false detections. In single object tracking case, given a set of frames {x1,x2, . . . ,xn} indexed by time, and the corresponding set of labeling, i.e. bounding box, {y1,y2, . . . ,yn}, structured SVM tries to find a model f (x,y), such that the task of predicting object location in a testing frame x could be conquered by maximizing: f (x,y) = 〈w,Ψ(x|y)〉, (1)
Wang Yan, Xiaoye Han, Vladimir Pavlovic 0001
BMVC3
2012 Multi-output Laplacian dynamic ordinal regression for facial expression recognition and intensity estimation
abstract
Automated facial expression recognition has received increased attention over the past two decades. Existing works in the field usually do not encode either the temporal evolution or the intensity of the observed facial displays. They also fail to jointly model multidimensional (multi-class) continuous facial behaviour data; binary classifiers - one for each target basic-emotion class - are used instead. In this paper, intrinsic topology of multidimensional continuous facial affect data is first modeled by an ordinal manifold. This topology is then incorporated into the Hidden Conditional Ordinal Random Field (H-CORF) framework for dynamic ordinal regression by constraining H-CORF parameters to lie on the ordinal manifold. The resulting model attains simultaneous dynamic recognition and intensity estimation of facial expressions of multiple emotions. To the best of our knowledge, the proposed method is the first one to achieve this on both deliberate as well as spontaneous facial affect data.
Ognjen Rudovic, Vladimir Pavlovic 0001, Maja Pantic
CVPR2
2012 Dynamic Probabilistic CCA for Analysis of Affective Behaviour
Mihalis A. Nicolaou, Vladimir Pavlovic 0001, Maja Pantic
ECCV (7)2
2012 Analysis of Causality in Stock Market Data
abstract
Analyzing the changes in volatility is an important aspect in financial data analysis leading to effective estimation of risk and discovering underlying causes of such changes. While there is a rich literature in estimating implied and stochastic volatility in financial time series using traditional econometric methods, the application of machine learning methods such as sparse regression with temporal smoothness constraints is still in its infancy. In this paper, we propose a sparse, smooth regularized regression model to infer the volatility of the data while explicitly accounting for dependencies between different companies. Using real stock market data, we construct dynamic time varying graphs for different sectors of companies to further analyze how the volatility dependency between companies within sectors vary over time. We also show how our model captures the fluctuations in volatility over different economic conditions such as financial crisis periods. Further, based on these regression estimates we show how the proposed model assists in discovering useful correlations with external factors such as oil price, inflation, S&P500 index and also with various domestic trend indices.
Chathra Hendahewa, Vladimir Pavlovic 0001
ICMLA (1)2
2012 Attribute rating for classification of visual objects
Jongpil Kim, Vladimir Pavlovic 0001
ICPR2
2012 Sparse Granger causality graphs for human action classification
Saehoon Yi, Vladimir Pavlovic 0001
ICPR2
2012 Efficient evaluation of large sequence kernels
abstract
Classification of sequences drawn from a finite alphabet using a family of string kernels with inexact matching (e.g., spectrum or mismatch) has shown great success in machine learning. However, selection of optimal mismatch kernels for a particular task is severely limited by inability to compute such kernels for long substrings (k-mers) with potentially many mismatches (m). In this work we introduce a new method that allows us to exactly evaluate kernels for large k, m and arbitrary alphabet size. The task can be accomplished by first solving the more tractable problem for small alphabets, and then trivially generalizing to any alphabet using a small linear system of equations. This makes it possible to explore a larger set of kernels with a wide range of kernel parameters, opening a possibility to better model selection and improved performance of the string kernels. To investigate the utility of large (k,m) string kernels, we consider several sequence classification problems, including protein remote homology detection, fold prediction, and music classification. Our results show that increased k-mer lengths with larger substitutions can improve classification performance.
Pavel P. Kuksa, Vladimir Pavlovic 0001
KDD2
2012 Distributed Probabilistic Learning for Camera Networks with Missing Data
abstract
Probabilistic approaches to computer vision typically assume a centralized setting, with the algorithm granted access to all observed data points. However, many problems in wide-area surveillance can benefit from distributed modeling, either because of physical or computational constraints. Most distributed models to date use algebraic approaches (such as distributed SVD) and as a result cannot explicitly deal with missing data. In this work we present an approach to estimation and learning of generative probabilistic models in a distributed context where certain sensor data can be missing. In particular, we show how traditional centralized models, such as probabilistic PCA and missing-data PPCA, can be learned when the data is distributed across a network of sensors. We demonstrate the utility of this approach on the problem of distributed affine structure from motion. Our experiments suggest that the accuracy of the learned probabilistic structure and motion models rivals that of traditional centralized factorization methods while being able to handle challenging situations such as missing or noisy observations.
Sejong Yoon, Vladimir Pavlovic 0001
NIPS2
2012 Generalized Similarity Kernels for Efficient Sequence Classification
abstract
String kernel-based machine learning methods have yielded great success in practical tasks of structured/sequential data analysis. They often exhibit state-of-the-art performance on tasks such as document topic elucidation, music genre classification, protein superfamily and fold prediction. However, typical string kernel methods rely on symbolic Hamming-distance based matching which may not necessarily reflect the underlying (e.g., physical) similarity between sequence fragments. In this work we propose a novel computational framework that uses general similarity metrics S(·, ·) and distance-preserving embeddings with string kernels to improve sequence classification. In particular, we consider two approaches that allow one either to incorporate non-Hamming similarity S(·, ·) into similarity evaluation by matching only the features that are similar according to S(·, ·) or to retain actual (approximate) similarity/distance scores in similarity evaluation. An embedding step, a distance-preserving bit-string mapping, is used to effectively capture similarity between otherwise symbolically different sequence elements. We show that it is possible to retain computational efficiency of string kernels while using this more “precise” measure of similarity. We then demonstrate that on a number of sequence classification tasks such as music, and biological sequence classification, the new method can substantially improve upon state-of-the-art string kernel baselines.
Pavel P. Kuksa, Vladimir Pavlovic 0001
SDM3
2011 Isotonic CCA for sequence alignment and activity recognition
abstract
This paper presents an approach for sequence alignment based on canonical correlation analysis(CCA). We show that a novel set of constraints imposed on traditional CCA leads to canonical solutions with the time warping property, i.e., non-decreasing monotonicity in time. This formulation generalizes the more traditional dynamic time warping (DTW) solutions to cases where the alignment is accomplished on arbitrary subsequence segments, optimally determined from data, instead on individual sequence samples. We then introduce a robust and efficient algorithm to find such alignments using non-negative least squares reductions. Experimental results show that this new method, when applied to MOCAP activity recognition problems, can yield improved recognition accuracy.
Shahriar Shariat, Vladimir Pavlovic 0001
ICCV2
2011 A Belief Propagation algorithm for bias field estimation and image segmentation
abstract
Intensity-based image segmentation is often plagued by the spatial intensity inhomogeneities (or non-uniformities) that are caused by the imperfection of the imaging devices and the varying operating conditions, also known as the bias field. We present a graphical model representation of the joint segmentation and bias field estimation problem and propose an iterative solver based on the Belief Propagation (BP) algorithm. The intractable joint inference problem of the original graphical model is decoupled into two MRF-MAP estimation problems and solved by a discrete-valued BP and a Gaussian BP, respectively and iteratively. We validate our method using both simulated and real data and show its connection to some of the classical filtering-based approaches.
Rui Huang 0001, Nong Sang, Vladimir Pavlovic 0001, Dimitris N. Metaxas
ICIP3
2011 Sequence classification via large margin hidden Markov models
Minyoung Kim 0001, Vladimir Pavlovic 0001
Data Min. Knowl. Discov.2
2011 Central Subspace Dimensionality Reduction Using Covariance Operators
abstract
We consider the task of dimensionality reduction informed by real-valued multivariate labels. The problem is often treated as Dimensionality Reduction for Regression (DRR), whose goal is to find a low-dimensional representation, the central subspace, of the input data that preserves the statistical correlation with the targets. A class of DRR methods exploits the notion of inverse regression (IR) to discover central subspaces. Whereas most existing IR techniques rely on explicit output space slicing, we propose a novel method called the Covariance Operator Inverse Regression (COIR) that generalizes IR to nonlinear input/output spaces without explicit target slicing. COIR's unique properties make DRR applicable to problem domains with high-dimensional output data corrupted by potentially significant amounts of noise. Unlike recent kernel dimensionality reduction methods that employ iterative nonconvex optimization, COIR yields a closed-form solution. We also establish the link between COIR, other DRR techniques, and popular supervised dimensionality reduction methods, including canonical correlation analysis and linear discriminant analysis. We then extend COIR to semi-supervised settings where many of the input points lack their labels. We demonstrate the benefits of COIR on several important regression problems in both fully supervised and semi-supervised settings.
Minyoung Kim 0001, Vladimir Pavlovic 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Prediction of Protein Functions with Gene Ontology and Interspecies Protein Homology Data
abstract
Accurate computational prediction of protein functions increasingly relies on network-inspired models for the protein function transfer. This task can become challenging for proteins isolated in their own network or those with poor or uncharacterized neighborhoods. Here, we present a novel probabilistic chain-graph-based approach for predicting protein functions that builds on connecting networks of two (or more) different species by links of high interspecies sequence homology. In this way, proteins are able to "exchange" functional information with their neighbors-homologs from a different species. The knowledge of interspecies relationships, such as the sequence homology, can become crucial in cases of limited information from other sources of data, including the protein-protein interactions or cellular locations of proteins. We further enhance our model to account for the Gene Ontology dependencies by linking multiple but related functional ontology categories within and across multiple species. The resulting networks are of significantly higher complexity than most traditional protein network models. We comprehensively benchmark our method by applying it to two largest protein networks, the Yeast and the Fly. The joint Fly-Yeast network provides substantial improvements in precision, accuracy, and false positive rate over networks that consider either of the sources in isolation. At the same time, the new model retains the computational efficiency similar to that of the simpler networks.
Antonina Mitrofanova, Vladimir Pavlovic 0001, Bud Mishra
IEEE ACM Trans. Comput. Biol. Bioinform.2
2010 Structured Output Ordinal Regression for Dynamic Facial Emotion Intensity Prediction
Minyoung Kim 0001, Vladimir Pavlovic 0001
ECCV (3)2
2010 Reinforcement Learning for Robust and Efficient Real-World Tracking
abstract
In this paper we present a new approach for combining several independent trackers into one robust real-time tracker. Unlike previous work that employ multiple tracking objectives used in unison, our tracker manages to determine an optimal sequence of individual trackers given the characteristics present in the video and the desire to achieve maximally efficient tracking. This allows for the selection of fast less-robust trackers when little movement is sensed, while using more robust but computationally intensive trackers in more dynamic scenes. We test this approach on the problem of real-world face tracking. Results show that this approach is a viable method for combining several independent trackers into one robust real-time tracker capable of tracking faces in varied lighting conditions, video resolutions, and with occlusions.
Andre Cohen, Vladimir Pavlovic 0001
ICPR2
2010 Spatial Representation for Efficient Sequence Classification
abstract
We present a general, simple feature representation of sequences that allows efficient inexact matching, comparison and classification of sequential data. This approach, recently introduced for the problem of biological sequence classification, exploits a novel multi-scale representation of strings. The new representation leads to discovery of very efficient algorithms for string comparison, independent of the alphabet size. We show that these algorithms can be generalized to handle a wide gamut of sequence classification problems in diverse domains such as the music and text sequence classification. The presented algorithms offer low computational cost and highly scalable implementations across different application domains. The new method demonstrates order-of-magnitude running time improvements over existing state-of-the-art approaches while matching or exceeding their predictive accuracy.
Pavel P. Kuksa, Vladimir Pavlovic 0001
ICPR2
2010 Hidden Conditional Ordinal Random Fields for Sequence Classification
Minyoung Kim 0001, Vladimir Pavlovic 0001
ECML/PKDD (2)2
2010 Semi-supervised Abstraction-Augmented String Kernel for Multi-level Bio-Relation Extraction
Pavel P. Kuksa, Yanjun Qi, Ronan Collobert, Jason Weston, Vladimir Pavlovic 0001, Xia Ning
ECML/PKDD (2)6
2010 Efficient motif finding algorithms for large-alphabet inputs
abstract
BACKGROUND: We consider the problem of identifying motifs, recurring or conserved patterns, in the biological sequence data sets. To solve this task, we present a new deterministic algorithm for finding patterns that are embedded as exact or inexact instances in all or most of the input strings. RESULTS: The proposed algorithm (1) improves search efficiency compared to existing algorithms, and (2) scales well with the size of alphabet. On a synthetic planted DNA motif finding problem our algorithm is over 10× more efficient than MITRA, PMSPrune, and RISOTTO for long motifs. Improvements are orders of magnitude higher in the same setting with large alphabets. On benchmark TF-binding site problems (FNP, CRP, LexA) we observed reduction in running time of over 12×, with high detection accuracy. The algorithm was also successful in rapidly identifying protein motifs in Lipocalin, Zinc metallopeptidase, and supersecondary structure motifs for Cadherin and Immunoglobin families. CONCLUSIONS: Our algorithm reduces computational complexity of the current motif finding algorithms and demonstrate strong running time improvements over existing exact algorithms, especially in important and difficult cases of large-alphabet sequences.
Pavel P. Kuksa, Vladimir Pavlovic 0001
BMC Bioinform.2
2010 Baselines for Image Annotation
Ameesh Makadia, Vladimir Pavlovic 0001, Sanjiv Kumar
Int. J. Comput. Vis.2
2009 Fast Motif Selection for Biological Sequences
abstract
We consider the problem of identifying motifs, recurring or conserved patterns, in the sets of biological sequences. To solve this task, we present new deterministic and exact algorithms for finding patterns that are embedded as exact or inexact instances in all or most of the input strings. The proposed algorithms (1) improve search efficiency compared to existing exact algorithms by focusing search on a selected set of potential motif instances, and (2) scale well with the input length and the size of alphabet. Our algorithms are orders of magnitude faster than existing exact algorithms for common pattern identification. We evaluate our algorithms on benchmark motif finding problems and real applications in biological sequence analysis and show that they exhibit significant running time improvements compared to the state-of-the-art approaches.
Pavel P. Kuksa, Vladimir Pavlovic 0001
BIBM2
2009 Efficient use of unlabeled data for protein sequence classification: a comparative study
abstract
BACKGROUND: Recent studies in computational primary protein sequence analysis have leveraged the power of unlabeled data. For example, predictive models based on string kernels trained on sequences known to belong to particular folds or superfamilies, the so-called labeled data set, can attain significantly improved accuracy if this data is supplemented with protein sequences that lack any class tags-the unlabeled data. In this study, we present a principled and biologically motivated computational framework that more effectively exploits the unlabeled data by only using the sequence regions that are more likely to be biologically relevant for better prediction accuracy. As overly-represented sequences in large uncurated databases may bias the estimation of computational models that rely on unlabeled data, we also propose a method to remove this bias and improve performance of the resulting classifiers. RESULTS: Combined with state-of-the-art string kernels, our proposed computational framework achieves very accurate semi-supervised protein remote fold and homology detection on three large unlabeled databases. It outperforms current state-of-the-art methods and exhibits significant reduction in running time. CONCLUSION: The unlabeled sequences used under the semi-supervised setting resemble the unpolished gemstones; when used as-is, they may carry unnecessary features and hence compromise the classification accuracy but once cut and polished, they improve the accuracy of the classifiers considerably.
Pavel P. Kuksa, Pai-Hsi Huang, Vladimir Pavlovic 0001
BMC Bioinform.3
2009 Efficient alignment-free DNA barcode analytics
abstract
BACKGROUND: In this work we consider barcode DNA analysis problems and address them using alternative, alignment-free methods and representations which model sequences as collections of short sequence fragments (features). The methods use fixed-length representations (spectrum) for barcode sequences to measure similarities or dissimilarities between sequences coming from the same or different species. The spectrum-based representation not only allows for accurate and computationally efficient species classification, but also opens possibility for accurate clustering analysis of putative species barcodes and identification of critical within-barcode loci distinguishing barcodes of different sample groups. RESULTS: New alignment-free methods provide highly accurate and fast DNA barcode-based identification and classification of species with substantial improvements in accuracy and speed over state-of-the-art barcode analysis methods. We evaluate our methods on problems of species classification and identification using barcodes, important and relevant analytical tasks in many practical applications (adverse species movement monitoring, sampling surveys for unknown or pathogenic species identification, biodiversity assessment, etc.) On several benchmark barcode datasets, including ACG, Astraptes, Hesperiidae, Fish larvae, and Birds of North America, proposed alignment-free methods considerably improve prediction accuracy compared to prior results. We also observe significant running time improvements over the state-of-the-art methods. CONCLUSION: Our results show that newly developed alignment-free methods for DNA barcoding can efficiently and with high accuracy identify specimens by examining only few barcode features, resulting in increased scalability and interpretability of current computational approaches to barcoding.
Pavel P. Kuksa, Vladimir Pavlovic 0001
BMC Bioinform.2
2009 Discriminative Learning for Dynamic State Prediction
abstract
We consider the problem of predicting a sequence of real-valued multivariate states that are correlated by some unknown dynamics, from a given measurement sequence. Although dynamic systems such as the State-Space Models are popular probabilistic models for the problem, their joint modeling of states and observations, as well as the traditional generative learning by maximizing a joint likelihood may not be optimal for the ultimate prediction goal. In this paper, we suggest two novel discriminative approaches to the dynamic state prediction: 1) learning generative state-space models with discriminative objectives and 2) developing an undirected conditional model. These approaches are motivated by the success of recent discriminative approaches to the structured output classification in discrete-state domains, namely, discriminative training of Hidden Markov Models and Conditional Random Fields (CRFs). Extending CRFs to real multivariate state domains generally entails imposing density integrability constraints on the CRF parameter space, which can make the parameter learning difficult. We introduce an efficient convex learning algorithm to handle this task. Experiments on several problem domains, including human motion and robot-arm state estimation, indicate that the proposed approaches yield high prediction accuracy comparable to or better than state-of-the-art methods.
Minyoung Kim 0001, Vladimir Pavlovic 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 On the Role of Local Matching for Efficient Semi-supervised Protein Sequence Classification
abstract
Recent studies in protein sequence analysis have leveraged the power of unlabeled data. For example, the profile and mismatch neighborhood kernels have shown significant improvements over classifiers estimated under the fully supervised setting. In this study, we present a principled and biologically motivated framework that more effectively exploits the unlabeled data by only utilizing regions that are more likely to be biologically relevant for better prediction accuracy. As overly-represented sequences in large uncurated databases may bias kernel estimations that rely on unlabeled data, we also propose a method to remove this bias and improve performance of resulting classifiers.Combined with a computationally efficient sparse family of string kernels, our proposed framework achieves state-of-the-art accuracy in semi-supervised protein remote homology detection on three large unlabeled databases.
Pavel P. Kuksa, Pai-Hsi Huang, Vladimir Pavlovic 0001
BIBM3
2008 Integrative Protein Function Transfer Using Factor Graphs and Heterogeneous Data Sources
abstract
We propose a novel approach for predicting protein functions of an organism by coupling sequence homology and PPI data between two (or more) species with multi-functional Gene Ontology information into a single computational model. Instead of using a network of one organism in isolation, we join networks of different species by inter-species sequence homology links of sufficient similarity. As a consequence, the knowledge of a protein's function is acquired not only from one species' network alone, but also through homologous links to the networks of different species. We apply our method to two largest protein networks, Yeast (Saccharomyces cerevisiae) and Fly (Drosophila melanogaster). Our joint Fly-Yeast network displays statistically significant improvements in precision, accuracy, and false positive rate over networks that consider either of the sources in isolation, while retaining the computational efficiency of the simpler models.
Antonina Mitrofanova, Vladimir Pavlovic 0001, Bud Mishra
BIBM2
2008 Face tracking and recognition with visual constraints in real-world videos
abstract
We address the problem of tracking and recognizing faces in real-world, noisy videos. We track faces using a tracker that adaptively builds a target model reflecting changes in appearance, typical of a video setting. However, adaptive appearance trackers often suffer from drift, a gradual adaptation of the tracker to non-targets. To alleviate this problem, our tracker introduces visual constraints using a combination of generative and discriminative models in a particle filtering framework. The generative term conforms the particles to the space of generic face poses while the discriminative one ensures rejection of poorly aligned targets. This leads to a tracker that significantly improves robustness against abrupt appearance changes and occlusions, critical for the subsequent recognition phase. Identity of the tracked subject is established by fusing pose-discriminant and person-discriminant features over the duration of a video sequence. This leads to a robust video-based face recognizer with state-of-the-art recognition performance. We test the quality of tracking and face recognition on real-world noisy videos from YouTube as well as the standard Honda/UCSD database. Our approach produces successful face tracking results on over 80% of all videos without video or person-specific parameter tuning. The good tracking performance induces similarly high recognition rates: 100% on Honda/UCSD and over 70% on the YouTube set containing 35 celebrities in 1500 sequences.
Minyoung Kim 0001, Sanjiv Kumar, Vladimir Pavlovic 0001, Henry A. Rowley
CVPR3
2008 Dimensionality reduction using covariance operator inverse regression
abstract
We consider the task of dimensionality reduction for regression (DRR) whose goal is to find a low dimensional representation of input covariates, while preserving the statistical correlation with output targets. DRR is particularly suited for visualization of high dimensional data as well as the efficient regressor design with a reduced input dimension. In this paper we propose a novel nonlinear method for DRR that exploits the kernel Gram matrices of input and output. While most existing DRR techniques rely on the inverse regression, our approach removes the need for explicit slicing of the output space using covariance operators in RKHS. This unique property make DRR applicable to problem domains with high dimensional output data with potentially significant amounts of noise. Although recent kernel dimensionality reduction algorithms make use of RKHS covariance operators to quantify conditional dependency between the input and the targets via the dimension-reduced input, they are either limited to a transduction setting or linear input subspaces and restricted to non-closed-form solutions. In contrast, our approach provides a closed-form solution to the nonlinear basis functions on which any new input point can be easily projected. We demonstrate the benefits of the proposed method in a comprehensive set of evaluations on several important regression problems that arise in computer vision.
Minyoung Kim 0001, Vladimir Pavlovic 0001
CVPR2
2008 A New Baseline for Image Annotation
Ameesh Makadia, Vladimir Pavlovic 0001, Sanjiv Kumar
ECCV (3)2
2008 Fast protein homology and fold detection with sparse spatial sample kernels
abstract
In this work we present a new string similarity feature, the sparse spatial sample (SSS). An SSS is a set of short substrings at specific spatial displacements contained in the original string. Using this feature we induce the SSS kernel (SSSK) which measures the agreement in the SSS content between pairs of strings. The SSSK yields better prediction performance at substantially reduced computational cost than existing algorithms for sequence classification tasks. We show that on the task of predicting the functional and structural classes of proteins, the SSSK results in state-of-the-art performance across several benchmark sets in both supervised and semi-supervised learning settings. The results have immediate practical value for accurate protein superfamily and fold classification and may be similarly extended to other sequence modeling domains.
Pavel P. Kuksa, Pai-Hsi Huang, Vladimir Pavlovic 0001
ICPR3
2008 Scalable Algorithms for String Kernels with Inexact Matching
abstract
We present a new family of linear time algorithms based on sufficient statistics for string comparison with mismatches under the string kernels framework. Our algorithms improve theoretical complexity bounds of existing approaches while scaling well with respect to the sequence alphabet size, the number of allowed mismatches and the size of the dataset. In particular, on large alphabets with loose mismatch constraints our algorithms are several orders of magnitude faster than the existing algorithms for string comparison under the mismatch similarity measure. We evaluate our algorithms on synthetic data and real applications in music genre classification, protein remote homology detection and protein fold prediction. The scalability of the algorithms allows us to consider complex sequence transformations, modeled using longer string features and larger numbers of mismatches, leading to a state-of-the-art performance with significantly reduced running times.
Pavel P. Kuksa, Pai-Hsi Huang, Vladimir Pavlovic 0001
NIPS3
2008 Boosted Bayesian network classifiers
Yushi Jing, Vladimir Pavlovic 0001, James M. Rehg
Mach. Learn.2
2007 Discriminative Learning of Dynamical Systems for Motion Tracking
abstract
We introduce novel discriminative learning algorithms for dynamical systems. Models such as conditional random fields or maximum entropy Markov models outperform the generative hidden Markov models in sequence tagging problems in discrete domains. However, continuous state domains introduce a set of constraints that can prevent direct application of these traditional models. Instead, we suggest to learn generative dynamic models with discriminative cost functionals. For linear dynamical systems, the proposed methods provide significantly lower prediction error than the standard maximum likelihood estimator, often comparable to nonlinear models. As a result, the models with lower representational capacity but computationally more tractable than nonlinear models can be used for accurate and efficient state estimation. We evaluate the generalization performance of our methods on the 3D human pose tracking problem from monocular videos. The experiments indicate that the discriminative learning can lead to improved accuracy of pose estimation with no increase in computational cost of tracking.
Minyoung Kim 0001, Vladimir Pavlovic 0001
CVPR2
2007 Embedded Profile Hidden Markov Models for Shape Analysis
abstract
An ideal shape model should be both invariant to global transformations and robust to local distortions. In this paper we present a new shape modeling framework that achieves both efficiently. A shape instance is described by a curvature-based shape descriptor. A Profile Hidden Markov Model (PHMM) is then built on such descriptors to represent a class of similar shapes. PHMMs are a particular type of Hidden Markov Models (HMMs) with special states and architecture that can tolerate considerable shape contour perturbations, including rigid and non-rigid deformations, occlusions, and missing parts. The sparseness of the PHMM structure provides efficient inference and learning algorithms for shape modeling and analysis. To capture the global characteristics of a class of shapes, the PHMM parameters are further embedded into a subspace that models long term spatial dependencies. The new framework can be applied to a wide range of problems, such as shape matching/registration, classification/recognition, etc. Our experimental results demonstrate the effectiveness and robustness of this new model in these different settings.
Rui Huang 0001, Vladimir Pavlovic 0001, Dimitris N. Metaxas
ICCV2
2007 Conditional State Space Models for Discriminative Motion Estimation
abstract
We consider the problem of predicting a sequence of real-valued multivariate states from a given measurement sequence. Its typical application in computer vision is the task of motion estimation. State Space Models are widely used generative probabilistic models for the problem. Instead of jointly modeling states and measurements, we propose a novel discriminative undirected graphical model which conditions the states on the measurements while exploiting the sequential structure of the problem. The major benefits of this approach are: (1) It focuses on the ultimate prediction task while avoiding probably unnecessary effort in modeling the measurement density, (2) It relaxes generative models' assumption that the measurements are independent given the states, and (3) The proposed inference algorithm takes linear time in the measurement dimension as opposed to the cubic time for Kalman filtering, which allows us to incorporate large numbers of measurement features. We show that the parameter learning can be cast as an instance of convex optimization. We also provide efficient convex optimization methods based on theorems from linear algebra. The performance of the proposed model is evaluated on both synthetic data and the human body pose estimation from silhouette videos.
Minyoung Kim 0001, Vladimir Pavlovic 0001
ICCV2
2007 A recursive method for discriminative mixture learning
abstract
We consider the problem of learning density mixture models for classification. Traditional learning of mixtures for density estimation focuses on models that correctly represent the density at all points in the sample space. Discriminative learning, on the other hand, aims at representing the density at the decision boundary. We introduce a novel discriminative learning method for mixtures of generative models. Unlike traditional discriminative learning methods that often resort to computationally demanding gradient search optimization, the proposed method is highly efficient as it reduces to generative learning of individual mixture components on weighted data. Hence it is particularly suited to domains with complex component models, such as hidden Markov models or Bayesian networks in general, that are usually too complex for effective gradient search. We demonstrate the benefits of the proposed method in a comprehensive set of evaluations on time-series sequence classification problems.
Minyoung Kim 0001, Vladimir Pavlovic 0001
ICML2
2007 Fast Kernel Methods for SVM Sequence Classifiers
Pavel P. Kuksa, Vladimir Pavlovic 0001
WABI2
2007 Special issue on vision for human-computer interaction
Mathias Kölsch, Vladimir Pavlovic 0001, Branislav Kisacanin, Thomas S. Huang
Comput. Vis. Image Underst.2
2007 Outlier rejection in high-dimensional deformable models
Christian Vogler, Siome Goldenstein, Jorge Stolfi, Vladimir Pavlovic 0001, Dimitris N. Metaxas
Image Vis. Comput.4
2006 Discriminative Learning of Mixture of Bayesian Network Classifiers for Sequence Classification
abstract
A mixture of Bayesian Network Classifiers(BNC) has a potential to yield superior classification and generative performance to a single BNC model. We introduce novel discriminative learning methods for mixtures of BNCs. Unlike a single BNC model where the discriminative learning resorts to a gradient search, we can exploit the properties of a mixture to alleviate the complex learning task. The proposed method adds mixture components recursively via functional gradient boosting while maximizing the conditional likelihood. This method is highly efficient as it reduces to generative learning of a base BNC model on weighed data. The proposed approach is particularly suited to sequence classification problems where the kernels in the base model are usually too complex for effective gradient search. We demonstrate the improved classification performance of the proposed methods in an extensive set of evaluations on time-series sequence data, including human motion classification problems.
Minyoung Kim 0001, Vladimir Pavlovic 0001
CVPR (1)2
2006 Impact of Dynamics on Subspace Embedding and Tracking of Sequences
abstract
In this paper we study the role of dynamics in dimensionality reduction problems applied to sequences. We propose a new family of marginal auto-regressive (MAR) models that describe the space of all stable auto-regressive sequences, regardless of their specific dynamics. We apply the MAR class of models as sequence priors in probabilistic sequence subspace embedding problems. In particular, we consider a Gaussian process latent variable approach to dimensionality reduction and show that the use of MAR priors may lead to better estimates of sequence subspaces than the ones obtained by traditional non-sequential priors. We then propose a learning method for estimating nonlinear dynamic system (NDS) models that utilizes the new MAR priors. The utility of the proposed methods is demonstrated on several synthetic datasets as well as on the task of tracking 3D articulated figures in monocular image sequences.
Kooksang Moon, Vladimir Pavlovic 0001
CVPR (1)2
2006 A Profile Hidden Markov Model Framework for Modeling and Analysis of Shape
abstract
In this paper we propose a new framework for modeling 2D shapes. A shape is first described by a sequence of local features (e.g., curvature) of the shape boundary. The resulting description is then used to build a profile hidden Markov model (PHMM) representation of the shape. PHMMs are a particular type of hidden Markov models (HMMs) with special states and architecture that can tolerate considerable shape contour perturbations, including rigid and non-rigid deformations, occlusions and missing contour parts. Different from traditional HMM-based shape models, the sparseness of the PHMM structure allows efficient inference and learning algorithms for shape modeling and analysis. The new framework can be applied to a wide range of problems, from shape matching and classification to shape segmentation. Our experimental results show the effectiveness and robustness of this new approach in the three application domains.
Rui Huang 0001, Vladimir Pavlovic 0001, Dimitris N. Metaxas
ICIP2
2006 Protein classification using probabilistic chain graphs and the Gene Ontology structure
abstract
MOTIVATION: Probabilistic graphical models have been developed in the past for the task of protein classification. In many cases, classifications obtained from the Gene Ontology have been used to validate these models. In this work we directly incorporate the structure of the Gene Ontology into the graphical representation for protein classification. We present a method in which each protein is represented by a replicate of the Gene Ontology structure, effectively modeling each protein in its own 'annotation space'. Proteins are also connected to one another according to different measures of functional similarity, after which belief propagation is run to make predictions at all ontology terms. RESULTS: The proposed method was evaluated on a set of 4879 proteins from the Saccharomyces Genome Database whose interactions were also recorded in the GRID project. Results indicate that direct utilization of the Gene Ontology improves predictive ability, outperforming traditional models that do not take advantage of dependencies among functional terms. Average increase in accuracy (precision) of positive and negative term predictions of 27.8% (2.0%) over three different similarity measures and three subontologies was observed. AVAILABILITY: C/C++/Perl implementation is available from authors upon request.
Steven Carroll, Vladimir Pavlovic 0001
Bioinform.2
2005 Efficient discriminative learning of Bayesian network classifier via boosted augmented naive Bayes
abstract
The use of Bayesian networks for classification problems has received significant recent attention. Although computationally efficient, the standard maximum likelihood learning method tends to be suboptimal due to the mismatch between its optimization criteria (data likelihood) and the actual goal for classification (label prediction). Recent approaches to optimizing the classification performance during parameter or structure learning show promise, but lack the favorable computational properties of maximum likelihood learning. In this paper we present the Boosted Augmented Naive Bayes (BAN) classifier. We show that a combination of discriminative data-weighting with generative training of intermediate models can yield a computationally efficient method for discriminative parameter learning and structure selection. 1.
Yushi Jing, Vladimir Pavlovic 0001, James M. Rehg
ICML2
2004 A Graphical Model Framework for Coupling MRFs and Deformable Models
Rui Huang 0001, Vladimir Pavlovic 0001, Dimitris N. Metaxas
CVPR (2)2
2004 Model-Based Motion Clustering Using Boosted Mixture Modeling
Vladimir Pavlovic 0001
CVPR (1)1
2003 Discovering Clusters in Motion Time-Series Data
abstract
An approach is proposed for clustering time-series data. The approach can be used to discover groupings of similar object motions that were observed in a video collection. A finite mixture of hidden Markov models (HMMs) is fitted to the motion data using the expectation maximization (EM) framework. Previous approaches for HMM-based clustering employ a k-means formulation, where each sequence is assigned to only a single HMM. In contrast, the formulation presented in this paper allows each sequence to belong to more than a single HMM with some probability, and the hard decision about the sequence class membership can be deferred until a later time when such a decision is required. Experiments with simulated data demonstrate the benefit of using this EM-based approach when there is more "overlap" in the processes generating the data. Experiments with real data show the promising potential of HMM-based motion clustering in a number of applications.
Jonathan Alon, Stan Sclaroff, George Kollios, Vladimir Pavlovic 0001
CVPR (1)4
2003 RankGene: identification of diagnostic genes based on expression data
abstract
Abstract Summary: RankGene is a program for analyzing gene expression data and computing diagnostic genes based on their predictive power in distinguishing between different types of samples. The program integrates into one system a variety of popular ranking criteria, ranging from the traditional t-statistic to one-dimensional support vector machines. This flexibility makes RankGene a useful tool in gene expression analysis and feature selection. Availability: http://genomics10.bu.edu/yangsu/rankgene Contact: [email protected] * To whom correspondence should be addressed.
T. M. Murali 0001, Vladimir Pavlovic 0001, Michael Schaffer, Simon Kasif
Bioinform.3
2003 Guest Editors' Introduction to the Special Section on Graphical Models in Computer Vision
abstract
THE last 10 years have witnessed rapid growth in the popularity of graphical models, most notably Bayesian networks, as a tool for representing, learning, and computing complex probability distributions. Graphical models provide an explicit representation of the statistical dependencies between the components of a complex probability model, effectively marrying probability theory and graph theory. As Jordan puts it in [2], graphical models are “a natural tool for dealing with two problems that occur throughout applied mathematics and engineering—uncertainty and complexity—and, in particular, they are playing an increasingly important role in the design and analysis of machine learning algorithms.” Graphical models provide powerful computational support for the Bayesian approach to computer vision, which has become a standard framework for addressing vision problems. Many familiar tools from the vision literature, such as Markov random fields, hidden Markov models, and the Kalman filter, are instances of graphical models. More importantly, the graphical models formalism makes it possible to generalize these tools and develop novel statistical representations and associated algorithms for inference and learning. The history of graphical models in computer vision follows closely that of graphical models in general. Research by Pearl [3] and Lauritzen [4] in the late 1980s played a seminal role in introducing this formalism to areas of AI and statistical learning. Not long after, the formalism spread to fields such as statistics, systems engineering, information theory, pattern recognition, and, among others, computer vision. One of the earliest occurrences of graphical models in the vision literature was a paper by Binford et al. [1]. The paper described the use of Bayesian inference in a hierarchical probability model to match 3D object models to groupings of curves in a single image. The following year marked the publication of Pearl’s influential book [3] on graphical models. Since then, many technical papers have been published in IEEE journals and conference proceedings that address different aspects and applications of graphical models in computer vision. Our goal in organizing this special section was to demonstrate the breadth of applicability of the graphical models formalism to vision problems. Our call for papers in February 2002 produced 16 submissions. After a careful review process, we selected six papers for publication, including five regular papers, and one short paper. These papers reflect the state-of-the-art in the use of graphical models in vision problems that range from low-level image understanding to high-level scene interpretation. We believe these papers will appeal both to vision researchers who are actively engaged in the use of graphical models and machine learning researchers looking for a challenging application domain. The first paper in this section is “Stereo Matching Using Belief Propagation” by J. Sun, N.-N. Zheng, and H.-Y. Shum. The authors describe a new stereo algorithm based on loopy belief propagation, a powerful inference technique for complex graphical models in which exact inference is intractable. They formulate the dense stereo matching problem as MAP estimation on coupled Markov random fields and obtain promising results on standard test data sets. One of the benefits of this formulation, as the authors demonstrate, is the ease with which it can be extended to handle multiview stereo matching. In their paper “Statistical Cue Integration of DAG Deformable Models” S.K. Goldenstein, C. Vogler, and D. Metaxas describe a scheme for combining different sources of information into estimates of the parameters of a deformable model. They use a DAG representation of the interdependencies between the nodes in a deformable model. This framework supports the efficient integration of information from edges and other cues using the machinery of affine arithmetic and the propagation of uncertainties. They present experimental results for a face tracking application. Y. Song, L. Goncalves, and P. Perona describe, in their paper “Unsupervised Learning of Human Motion,” a method for learning probabilistic models of human motion from video sequences in cluttered scenes. Two key advantages of their method are its unsupervised nature, which can mitigate the need for tedious hand labeling of data, and the utilization of graphical model constraints to reduce the search space when fitting a human figure model. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 25, NO. 7, JULY 2003 785
James M. Rehg, Vladimir Pavlovic 0001, Thomas S. Huang, William T. Freeman
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Boosted learning in dynamic Bayesian networks for multimodal speaker detection
abstract
Bayesian network models provide an attractive framework for multimodal sensor fusion. They combine an intuitive graphical representation with efficient algorithms for inference and learning. However, the unsupervised nature of standard parameter learning algorithms for Bayesian networks can lead to poor performance in classification tasks. We have developed a supervised learning framework for Bayesian networks, which is based on the Adaboost algorithm of Schapire and Freund. Our framework covers static and dynamic Bayesian networks with both discrete and continuous states. We have tested our framework in the context of a novel multimodal HCI application: a speech-based command and control interface for a Smart Kiosk. We provide experimental evidence for the utility of our boosted learning approach.
Ashutosh Garg 0001, Vladimir Pavlovic 0001, James M. Rehg
Proc. IEEE2
2002 A Bayesian framework for combining gene predictions
abstract
Abstract Motivation: Gene identification and gene discovery in new genomic sequences is one of the most timely computational questions addressed by bioinformatics scientists. This computational research has resulted in several systems that have been used successfully in many whole-genome analysis projects. As the number of such systems grows the need for a rigorous way to combine the predictions becomes more essential. Results: In this paper we provide a Bayesian network framework for combining gene predictions from multiple systems. The framework allows us to treat the problem as combining the advice of multiple experts. Previous work in the area used relatively simple ideas such as majority voting. We introduce, for the first time, the use of hidden input/output Markov models for combining gene predictions. We apply the framework to the analysis of the Adh region in Drosophila that has been carefully studied in the context of gene finding and used as a basis for the GASP competition. The main challenge in combination of gene prediction programs is the fact that the systems are relying on similar features such as cod on usage and as a result the predictions are often correlated. We show that our approach is promising to improve the prediction accuracy and provides a systematic and flexible framework for incorporating multiple sources of evidence into gene prediction systems. Availability: Software can be made available on request from the authors. Contact: [email protected] * Part of this research was presented at Computational Genomics 2000, Baltimore, MD, November 2000. Portions of this research were conducted at Compaq Computer Corporation, Cambridge Research Laboratory, Cambridge, MA.
Vladimir Pavlovic 0001, Ashutosh Garg 0001, Simon Kasif
Bioinform.1
2000 Impact of Dynamic Model Learning on Classification of Human Motion
abstract
The human figure exhibits complex and rich dynamic behavior that is both nonlinear and time-varying. However, most work on tracking and analysis of figure motion has employed either generic or highly specific hand-tailored dynamic models superficially coupled with hidden Markov models (HMMs) of motion regimes. Recently, an alternative class of learned dynamic models known as switching linear dynamic systems (SLDSs) has been cast in the framework of dynamic Bayesian networks (DBNs) and applied to analysis and tracking of the human figure. In this paper we further study the impact of learned SLDS models on analysis and tracking of human motion and contrast them to the more common HMM models. We develop a novel approximate structured variational inference algorithm for SLDS, a globally convergent DBN inference scheme, and compare it with standard SLDS inference techniques. Experimental results on learning and analysis of figure dynamics from video data indicate the significant potential of the SLDS approach.
Vladimir Pavlovic 0001, James M. Rehg
CVPR1
2000 Multimodal Speaker Detection Using Error Feedback Dynamic Bayesian Networks
abstract
Design and development of novel human-computer interfaces poses a challenging problem: actions and intentions of users have to be inferred from sequences of noisy and ambiguous multi-sensory data such as video and sound. Temporal fusion of multiple sensors has been efficiently formulated using dynamic Bayesian networks (DBNs) which allows the power of statistical inference and learning to be combined with contextual knowledge of the problem. Unfortunately simple learning methods can cause such appealing models to fail when the data exhibits complex behavior. We formulate a learning framework for DBNs based on error-feedback and statistical boosting theory. We apply this framework to the problem of audio/visual speaker detection in an interactive kiosk environment using "off-the-shelf" visual and audio sensors (face, skin, texture, mouth motion, and silence detectors). Detection results obtained in this setup demonstrate superiority of our learning framework over that of the classical ML learning in DBNs.
Vladimir Pavlovic 0001, James M. Rehg, Ashutosh Garg 0001, Thomas S. Huang
CVPR1
2000 Audio-Visual Speaker Detection Using Dynamic Bayesian Networks
abstract
The development of human-computer interfaces poses a challenging problem: actions and intentions of different users have to be inferred from sequences of noisy and ambiguous sensory data. Temporal fusion of multiple sensors can be efficiently formulated using dynamic Bayesian networks (DBN). The DBN framework allows the power of statistical inference and learning to be combined with contextual knowledge of the problem. We demonstrate the use of DBN in tackling the problem of audio/visual speaker detection. "Off-the-shelf" visual and audio sensors (face, skin, texture, mouth motion, and silence detectors) are optimally fused along with contextual information in a DBN architecture that infers instances when an individual is speaking. Results obtained in the setup of an actual human-machine interaction system (Genie Casino Kiosk) demonstrate superiority of our approach over that of static, context-free fusion architecture.
Ashutosh Garg 0001, Vladimir Pavlovic 0001, James M. Rehg
FG2
2000 Multimodal Speaker Detection Using Input/Output Dynamic Bayesian Networks
Vladimir Pavlovic 0001, Ashutosh Garg 0001, James M. Rehg
ICMI1
2000 Learning Switching Linear Models of Human Motion
abstract
The human figure exhibits complex and rich dynamic behavior that is both nonlinear and time-varying. Effective models of human dynamics can be learned from motion capture data using switching linear dynamic system (SLDS) models. We present results for human motion synthe(cid:173) sis, classification, and visual tracking using learned SLDS models. Since exact inference in SLDS is intractable, we present three approximate in(cid:173) ference algorithms and compare their performance. In particular, a new variational inference algorithm is obtained by casting the SLDS model as a Dynamic Bayesian Network. Classification experiments show the superiority of SLDS over conventional HMM's for our problem domain.
Vladimir Pavlovic 0001, James M. Rehg, John MacCormick
NIPS1
1999 Time-Series Classification Using Mixed-State Dynamic Bayesian Networks
abstract
We present a novel mixed-state dynamic Bayesian network (DBN) framework for modeling and classifying time-series data such as object trajectories. A hidden Markov model (HMM) of discrete actions is coupled with a linear dynamical system (LDS) model of continuous trajectory motion. This combination allows us to model both the discrete and continuous causes of trajectories such as human gestures. The model is derived using a rich theoretical corpus from the Bayesian network literature. This allows us to use an approximate structured variational inference technique to solve the otherwise intractable inference of action and system states. Using the same DBN framework we show how to learn the mixed-state model parameters from data. Experiments show that with high statistical confidence the mixed-state DBNs perform favorably when compared to decoupled HMM/LDS models on the task of recognizing human gestures made with a computer mouse.
Vladimir Pavlovic 0001, Brendan J. Frey, Thomas S. Huang
CVPR1
1999 A Dynamic Bayesian Network Approach to Figure Tracking using Learned Dynamic Models
abstract
The human figure exhibits complex and rich dynamic behavior that is both nonlinear and time-varying. However most work on tracking and synthesizing figure motion has employed either simple, generic dynamic models or highly specific hand-tailored ones. Recently, a broad class of learning and inference algorithms for time-series models have been successfully cast in the framework of dynamic Bayesian networks (DBNs). This paper describes a novel DBN-based switching linear dynamic system (SLDS) model and presents its application to figure motion analysis. A key feature of our approach is an approximate Viterbi inference technique for overcoming the intractability of exact inference in mixed-state DBNs. We present experimental results for learning figure dynamics from video data and show promising initial results for tracking, interpolation, synthesis, and classification using learned models.
Vladimir Pavlovic 0001, James M. Rehg, Tat-Jen Cham, Kevin Murphy 0002
ICCV1
1999 Variational Learning in Mixed-State Dynamic Graphical Models
Vladimir Pavlovic 0001, Brendan J. Frey, Thomas S. Huang
UAI1
1998 Multimodal Tracking and Classification of Audio-Visual Features
abstract
The surge of interest in multimedia and multimodal interfaces has prompted the need for novel estimation and classification techniques for data from different but coupled modalities. Unimodal techniques ported to this domain have only exhibited limited success. We propose a new framework for feature prediction and classification based on multimodal knowledge-constrained hidden Markov models (HMMs). The classical role of HMMs as statistical classifiers is enhanced by their new role as multimodal feature predictors. Moreover, by fusing the multimodal formulation with higher level knowledge we allow the influence of such knowledge to be reflected in feature prediction as well as in feature classification.
Vladimir Pavlovic 0001
ICIP (1)1
1998 Toward multimodal human-computer interface
abstract
Recent advances in various signal processing technologies, coupled with an explosion in the available computing power, have given rise to a number of novel human-computer interaction (HCI) modalities: speech, vision-based gesture recognition, eye tracking, electroencephalograph, etc. Successful embodiment of these modalities into an interface has the potential of easing the HCI bottleneck that has become noticeable with the advances in computing and communication. It has also become increasingly evident that the difficulties encountered in the analysis and interpretation of individual sensing modalities may be overcome by integrating them into a multimodal human-computer interface. We examine several promising directions toward achieving multimodal HCI. We consider some of the emerging novel input modalities for HCI and the fundamental issues in integrating them at various levels, from early signal level to intermediate feature level to late decision level. We discuss the different computational approaches that may be applied at the different levels of modality integration. We also briefly review several demonstrated multimodal HCI systems and applications. Despite all the recent developments, it is clear that further research is needed for interpreting and fitting multiple sensing modalities in the context of HCI. This research can benefit from many disparate fields of study that increase our understanding of the different human communication modalities and their potential role in HCI.
Rajeev Sharma, Vladimir Pavlovic 0001, Thomas S. Huang
Proc. IEEE2
1997 A Visual Computing Environment for Very Large Scale Biomolecular Modeling
abstract
Knowledge of the complex molecular structures of living cells is being accumulated at a tremendous rate. Key technologies enabling this success have been, high performance computing and powerful molecular graphics applications, but the technology is beginning to seriously lag behind challenges posed by the size and number of new structures and by the emerging opportunities in drug design and genetic engineering. A visual computing environment is being developed which permits interactive modeling of biopolymers by linking a 3D molecular graphics program with an efficient molecular dynamics simulation program executed on remote high-performance parallel computers. The system will be ideally suited for distributed computing environments, by utilizing both local 3D graphics facilities and the peak capacity of high-performance computers for the purpose of interactive biomolecular modeling. To create an interactive 3D environment three input methods will be explored: (1) a six degree of freedom "mouse" for controlling the space shared by the model and the user; (2) voice commands monitored through a microphone and recognized by a speech recognition interface; (3) hand gestures, detected through cameras and interpreted using computer vision techniques. Controlling 3D graphics connected to real time simulations and the use of voice with suitable language semantics, as well as hand gestures, promise great benefits for many types of problem solving environments. Our focus on structural biology takes advantage of existing sophisticated software, provides concrete objectives, defines a well-posed domain of tasks and offers a well-developed vocabulary for spoken communication.
Michael Zeller, James C. Phillips, Andrew Dalke, William Humphrey, Klaus Schulten, Thomas S. Huang, Vladimir Pavlovic 0001, Yunxin Zhao, Zion Lo, Stephen M. Chu, Rajeev Sharma
ASAP7
1997 Integration of Audio/Visual Information for Use in Human-Computer Intelligent Interaction
abstract
Human-computer intelligent interaction (HCII) in virtual environments is a rapidly developing field. Natural human communication is multi-modal, however, most modern computer interfaces rely exclusively on one mode of interaction. We employ a novel approach to integrating multiple modes of human-computer communication. By using auditory and visual features at different levels of integration we explore optimal ways of combining these modalities.
Vladimir Pavlovic 0001, G. A. Berry, Thomas S. Huang
ICIP (1)1
1997 Visual Interpretation of Hand Gestures for Human-Computer Interaction: A Review
abstract
The use of hand gestures provides an attractive alternative to cumbersome interface devices for human-computer interaction (HCI). In particular, visual interpretation of hand gestures can help in achieving the ease and naturalness desired for HCI. This has motivated a very active research area concerned with computer vision-based analysis and interpretation of hand gestures. We survey the literature on visual interpretation of hand gestures in the context of its role in HCI. This discussion is organized on the basis of the method used for modeling, analyzing, and recognizing gestures. Important differences in the gesture interpretation approaches arise depending on whether a 3D model of the human hand or an image appearance model of the human hand is used. 3D hand models offer a way of more elaborate modeling of hand gestures but lead to computational hurdles that have not been overcome given the real-time requirements of HCI. Appearance-based models lead to computationally efficient "purposive" approaches that work well under constrained situations but seem to lack the generality desirable for HCI. We also discuss implemented gestural systems as well as other potential applications of vision-based gesture recognition. Although the current progress is encouraging, further theoretical as well as computational advances are needed before gestures can be widely used for HCI. We discuss directions of future research in gesture recognition, including its integration with other natural modes of human-computer interaction.
Vladimir Pavlovic 0001, Rajeev Sharma, Thomas S. Huang
IEEE Trans. Pattern Anal. Mach. Intell.1
1996 Invited Speech: "Gestural Interface to a visual computing Environment for Molecular biologists"
abstract
In recent years there has been tremendous progress in 3-D, immersive display and virtual reality (VR) technologies. Scientific visualization af data is one of many applications that has benefited from this progress. To fully exploit the potential of these applications in the new environment there is a need for "natural" interfaces that allow the manipulation of such displays without burdensome attachments. This paper describes the use of visual hand gesture analysis enhanced with speech recognition for developing a bimodal gesture/speech interface for controlling a 3-D display. The interface augments an existing application, VMD, which is a VR visual computing environment for molecular biologists. The free hand gestures are used for manipulating the 3-D graphical display together with a set of speech commands. We concentrate on the visual gesture analysis techniques used in developing this interface. The dual modality of gesture/speech is found to greatly aid the interaction capability.
Vladimir Pavlovic 0001, Rajeev Sharma, Thomas S. Huang
FG1
1996 Transform image coding based on joint adaptation of filter banks and tree structures
abstract
Recent work on filter banks and related expansions has revealed an interesting insight: different filter bank trees can be regarded as different ways of constructing orthonormal bases for linear signal expansion. In particular, fast algorithms for finding best bases in an operational rate-distortion sense have been successfully used in image coding. Independently of this work, recent research has also explored the design of filter banks that optimize energy compaction for a single signal or a class of signals. In this paper we integrate these two different but complementary approaches to best-basis design and propose an image coder in which subband filter banks, tree structure and quantizers are chosen so as to optimize rate-distortion performance. These optimal filter banks, tree structure and quantizers represent side information. They are selected from a codebook designed from training data, using a rate-distortion criterion.
Pierre Moulin, Kannan Ramchandran, Vladimir Pavlovic 0001
ICIP (2)3
1996 Speech/gesture interface to a visual computing environment for molecular biologists
abstract
Recent progress in 3-D, immersive display and virtual reality (VR) technologies has made possible many exciting applications, for example interactive visualization of complex scientific data. To fully exploit this potential there is a need for "natural" interfaces that allow the manipulation of such displays without cumbersome attachments. In this paper we describe the use of visual hand gesture analysis and speech recognition for developing a speech/gesture interface for controlling a 3-D display. The interface enhances an existing application, VMD, which is a VR visual computing environment far molecular biologists. The free hand gestures are used for manipulating the 3-D graphical display together with a set of speech commands. We describe the visual gesture analysis and the speech analysis techniques used in developing this interface. The dual modality of speech/gesture is found to greatly aid the interaction capability.
Rajeev Sharma, Thomas S. Huang, Vladimir Pavlovic 0001, Yunxin Zhao, Zion Lo, Stephen M. Chu, Klaus Schulten
ICPR3