Irfan A. Essa

dblp:e/IrfanAEssa · also Irfan Essa · DBLP profile ↗
← Back
146ranked-venue papers
8as first author
39since 2021 · last 2025
0000-0002-6236-2969ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 91 · 5 first-author · 19 since 2021Artificial intelligence and machine learning · 84 · 5 first-author · 28 since 2021Human-computer interaction and ubiquitous computing · 24 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Systems, architecture and hardware · 5 · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Security and privacy · 2Computer networks · 1
YearPublicationVenuePosition
2025 Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition - And Ways to Overcome Them
abstract
Cross-modal contrastive pre-training between natural language and other modalities, e.g., vision and audio, has demonstrated astonishing performance and effectiveness across a diverse variety of tasks and domains. In this paper, we investigate whether such natural language supervision can be used for wearable sensor based Human Activity Recognition (HAR), and discover that--surprisingly--it performs substantially worse than standard end-to-end training and self-supervision. We identify the primary causes for this as: sensor heterogeneity and the lack of rich, diverse text descriptions of activities. To mitigate their impact, we also develop strategies and assess their effectiveness through an extensive experimental evaluation. These strategies lead to significant increases in activity recognition, bringing performance closer to supervised and self-supervised training, while also enabling the recognition of unseen activities and cross modal retrieval of videos. Overall, our work paves the way for better sensor-language learning, ultimately leading to the development of foundational models for HAR using wearables.
Harish Haresamudram, Apoorva Beedu, Mashfiqui Rabbi, Sankalita Saha, Irfan A. Essa, Thomas Plötz
AAAI5
2025 AfriMed-QA: A Pan-African, Multi-Specialty, Medical Question-Answering Benchmark Dataset
abstract
Charles Nimo, Tobi Olatunji, Abraham Toluwase Owodunni, Tassallah Abdullahi, Emmanuel Ayodele, Mardhiyah Sanni, Ezinwanne C. Aka, Folafunmi Omofoye, Foutse Yuehgoh, Timothy Faniran, Bonaventure F. P. Dossou, Moshood O. Yekini, Jonas Kemp, Katherine A Heller, Jude Chidubem Omeke, Chidi Asuzu Md, Naome A Etori, Aïmérou Ndiaye, Ifeoma Okoh, Evans Doe Ocansey, Wendy Kinara, Michael L. Best, Irfan Essa, Stephen Edward Moore, Chris Fourie, Mercy Nyamewaa Asiedu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Charles Nimo, Tobi Olatunji, Abraham Toluwase Owodunni, Tassallah Abdullahi, Emmanuel Ayodele, Mardhiyah Sanni, Ezinwanne C. Aka, Folafunmi Omofoye, Foutse Yuehgoh, Timothy Faniran, Bonaventure F. P. Dossou, Moshood Yekini, Jonas Kemp, Katherine A. Heller, Jude Chidubem Omeke, Chidi Asuzu MD, Naome A. Etori, Aimérou Ndiaye, Ifeoma Okoh, Evans Doe Ocansey, Wendy Kinara, Michael L. Best, Irfan A. Essa, Stephen Edward Moore, Chris Fourie, Mercy Asiedu
ACL (1)23
2025 Cropper: Vision-Language Model for Image Cropping through In-Context Learning
abstract
The goal of image cropping is to identify visually appealing crops in an image. Conventional methods are trained on specific datasets and fail to adapt to new requirements. Recent breakthroughs in large vision-language models (VLMs) enable visual in-context learning without explicit training. However, downstream tasks with VLMs remain under explored. In this paper, we propose an effective approach to leverage VLMs for image cropping. First, we propose an efficient prompt retrieval mechanism for image cropping to automate the selection of in-context examples. Second, we introduce an iterative refinement strategy to iteratively enhance the predicted crops. The proposed framework, we refer to as Cropper, is applicable to a wide range of cropping tasks, including free-form cropping, subject-aware cropping, and aspect ratio-aware cropping. Extensive experiments demonstrate that Cropper significantly outperforms state-of-the-art methods across several benchmarks.
Jijun Jiang, Zhuofang Li, Junjie Ke, Yinxiao Li, Junfeng He, Steven Hickson, Katie Datsenko, Sangpil Kim, Ming-Hsuan Yang 0001, Irfan A. Essa, Feng Yang 0008
CVPR12
2025 Calibrated Multi-Preference Optimization for Aligning Diffusion Models
abstract
Aligning text-to-image (T2I) diffusion models with preference optimization is valuable for human-annotated datasets, but the heavy cost of manual data collection limits scalability. Using reward models offers an alternative, however, current preference optimization methods fall short in exploiting the rich information, as they only consider pairwise preference distribution. Furthermore, they lack generalization to multi-preference scenarios and struggle to handle inconsistencies between rewards. To address this, we present Calibrated Preference Optimization (CaPO), a novel method to align T2I diffusion models by incorporating the general preference from multiple reward models without human annotated data. The core of our approach involves a reward calibration method to approximate the general preference by computing the expected win-rate against the samples generated by the pretrained models. Additionally, we propose a frontier-based pair selection method that effectively manages the multi-preference distribution by selecting pairs from Pareto frontiers. Finally, we use regression loss to fine-tune diffusion models to match the difference between calibrated rewards of a selected pair. Experimental results show that CaPO consistently outperforms prior methods, such as Direct Preference Optimization (DPO), in both single and multi-reward settings validated by evaluation on T2I benchmarks, including GenEval and T2I-Compbench.
Kyungmin Lee, Xiahong Li, Qifei Wang, Junfeng He, Junjie Ke, Ming-Hsuan Yang 0001, Irfan A. Essa, Jinwoo Shin, Feng Yang 0008, Yinxiao Li
CVPR7
2025 Africa Health Check: Probing Cultural Bias in Medical LLMs
abstract
Large language models (LLMs) are increasingly deployed in global healthcare, yet their outputs often reflect Western-centric training data and omit indigenous medical systems and region-specific treatments.This study investigates cultural bias in instruction-tuned medical LLMs using a curated dataset of African traditional herbal medicine.We evaluate model behavior across two complementary tasks, namely, multiple-choice questions and fill-in-the-blank completions, designed to capture both treatment preferences and responsiveness to cultural context.To quantify outcome preferences and prompt influences, we apply two complementary metrics: Cultural Bias Score (CBS) and Cultural Bias Attribution (CBA).Our results show that while prompt adaptation can reduce inherent bias and enhance cultural alignment, models vary in how responsive they are to contextual guidance.Persistent default to allopathic 1 (Western) treatments in zero-shot scenarios suggest that many biases remain embedded in model training.These findings underscore the need for culturally informed evaluation strategies to guide the development of AI systems that equitably serve diverse global health contexts.By releasing our dataset and providing a dual-metric evaluation approach, we offer practical tools for developing more culturally aware and clinically grounded AI systems for healthcare settings in the Global South.
Charles Nimo, Shuheng Liu 0002, Irfan A. Essa, Michael L. Best
EMNLP3
2025 Mamba Fusion: Learning Actions Through Questioning
abstract
Video Language Models (VLMs) are crucial for generalizing across diverse tasks and using language cues to enhance learning. While transformer-based architectures have been the de facto in vision-language training, they face challenges like quadratic computational complexity, high GPU memory usage, and difficulty with long-term dependencies. To address these limitations, we introduce MambaVL, a novel model that leverages recent advancements in selective state space modality fusion to efficiently capture long-range dependencies and learn joint representations for vision and language data. MambaVL utilizes a shared state transition matrix across both modalities, allowing the model to capture a more comprehensive understanding of the actions in the scene. Furthermore, we propose a question-answering task that helps guide the model toward relevant cues. These questions provide critical information about actions, objects, and environmental context, leading to enhanced performance. As a result, MambaVL achieves state-of-the-art performance in action recognition on the Epic-Kitchens-100 dataset and outperforms baseline methods in action anticipation. The code is available at https://github.com/Dongzhikang/MambaVL.
Apoorva Beedu, Zhikang Dong, Jason Sheinkopf, Irfan A. Essa
ICASSP4
2025 Text Descriptions of Actions and Objects Improve Action Anticipation
abstract
Anticipating future actions is a highly challenging task due to the diversity and scale of potential future actions; yet, additional and complementary information from different modalities help narrow down plausible action choices. Going beyond typical sources such as video and audio, we primarily explore how text descriptions of actions and objects leads to more accurate action anticipation, as they provide additional contextual cues, e.g., about the environment and its contents. We propose Multi-modal Contrastive Anticipative Transformer (M-CAT), which is trained in two stages, where the model first learns to align video and other modalities with descriptions of future actions, and is subsequently fine-tuned to predict future actions. Through extensive experimental evaluation, we demonstrate that M-CAT outperforms baselines on the EpicKitchens datasets, and show that explicit incorporation of object and action information via their text descriptions leads to more effective action anticipation. Code available at https://github.com/ApoorvaBeedu/M-CAT.
Apoorva Beedu, Harish Haresamudram, Irfan A. Essa
ICASSP3
2025 Enabling Controllable, Identity Preserving, Non-Rigid Edits in Human-Centric Images
abstract
We approach the problem of inserting a person into a novel scene and controlling their pose via text guidance. Given an image of a person, a masked image of a scene, and a text description of the target pose, our model generates realistic, highly controllable images. We validate the robustness of our model’s true-to-text accuracy and identity preservation via a user study on in-the-wild images. In addition, we present a novel dataset containing pairs of frames from human-centric and action-rich videos, with text captions of the difference in human pose between frames. We also explore the challenges of controllable identity preservation for in-the-wild scenes and the failure modes of similar models. Our methods achieve a 10% increase in pose adherence ([email protected]) over comparable methods without compromising visual fidelity, and show a clear qualitative improvement.
Nikolai Warner, Jack Kolb, Meera Hahn, Jonathan Huang, Vighnesh Birodkar, Irfan A. Essa
ICIP6
2024 Prompt-Free Diffusion: Taking "Text" Out of Text-to-Image Diffusion Models
abstract
Text-to-image (T2I) research has grown explosively in the past year, owing to the large-scale pre-trained diffusion models and many emerging personalization and editing approaches. Yet, one pain point persists: the text prompt engineering, and searching high-quality text prompts for customized results is more art than science. Moreover, as commonly argued: “an image is worth a thousand words” - the attempt to describe a desired image with texts often ends up being ambiguous and cannot comprehensively cover delicate visual details, hence necessitating more additional controls from the visual domain. In this paper, we take a bold step forward: taking “Text” out of a pretrained T2I diffusion model, to reduce the burdensome prompt engineering efforts for users. Our proposed frame-work, Prompt-Free Diffusion, relies on only visual inputs to generate new images: it takes a reference image as “context”, an optional image structural conditioning, and an initial noise, with absolutely no text prompt. The core architecture behind the scene is Semantic Context Encoder (SeeCoder), substituting the commonly used CLIP-based or LLM-based text encoder. The reusability of SeeCoder also makes it a convenient drop-in component: one can also pre-train a SeeCoder in one T2I model and reuse it for another. Through extensive experiments, Prompt-Free Diffusion is experimentally found to (i) outperform prior exemplar-based image synthesis approaches; (ii) perform on par with state-of-the-art T2I models using prompts following the best practice; and (iii) be naturally extensible to other downstream applications such as anime figure generation and virtual try-on, with promising quality. Our code and models will be open-sourced.
Xingqian Xu, Zhangyang Wang, Gao Huang 0001, Irfan A. Essa, Humphrey Shi
CVPR5
2024 Photorealistic Video Generation with Diffusion Models
Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei 0001, Irfan A. Essa, Lu Jiang 0004, José Lezama
ECCV (79)7
2024 Parrot: Pareto-Optimal Multi-reward Reinforcement Learning Framework for Text-to-Image Generation
Yinxiao Li, Junjie Ke, Innfarn Yoo, Han Zhang 0010, Qifei Wang, Fei Deng 0001, Glenn Entis, Junfeng He, Gang Li 0021, Sangpil Kim, Irfan A. Essa, Feng Yang 0008
ECCV (38)13
2024 Language Model Beats Diffusion - Tokenizer is key to visual generation
abstract
While Large Language Models (LLMs) are the dominant models for generative tasks in language, they do not perform as well as diffusion models on image and video generation. To effectively use LLMs for visual generation, one crucial component is the visual tokenizer that maps pixel-space inputs to discrete tokens appropriate for LLM learning. In this paper, we introduce \modelname{}, a video tokenizer designed to generate concise and expressive tokens for both videos and images using a common token vocabulary. Equipped with this new tokenizer, we show that LLMs outperform diffusion models on standard image and video generation benchmarks including ImageNet and Kinetics. In addition, we demonstrate that our tokenizer surpasses the previously top-performing video tokenizer on two more tasks: (1) video compression comparable to the next-generation video codec (VCC) according to human evaluations, and (2) learning effective representations for action recognition tasks.
Lijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng 0003, Agrim Gupta, Xiuye Gu, Alex Hauptmann 0001, Boqing Gong, Ming-Hsuan Yang 0001, Irfan A. Essa, David A. Ross, Lu Jiang 0004
ICLR13
2024 VideoPoet: A Large Language Model for Zero-Shot Video Generation
abstract
We present VideoPoet, a language model capable of synthesizing high-quality video from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs – including images, videos, text, and audio. The training protocol follows that of Large Language Models (LLMs), consisting of two stages: pretraining and task-specific adaptation. During pretraining, VideoPoet incorporates a mixture of multimodal generative objectives within an autoregressive Transformer framework. The pretrained LLM serves as a foundation that can be adapted for a range of video generation tasks. We present empirical results demonstrating the model’s state-of-the-art capabilities in zero-shot video generation, specifically highlighting the ability to generate high-fidelity motions. Project page: http://sites.research.google/videopoet/
Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng 0003, Joshua V. Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso Martinez, David Minnen, Mikhail Sirotenko, Kihyuk Sohn, Hartwig Adam, Ming-Hsuan Yang 0001, Irfan A. Essa, Huisheng Wang, David A. Ross, Bryan Seybold, Lu Jiang 0004
ICML27
2024 BayRnTune: Adaptive Bayesian Domain Randomization via Strategic Fine-tuning
abstract
Domain randomization (DR), which entails training a policy with randomized dynamics, has proven to be a simple yet effective algorithm for reducing the gap between simulation and the real world. However, DR often requires careful tuning of randomization parameters. Methods like Bayesian Domain Randomization (Bayesian DR) and Active Domain Randomization (Adaptive DR) address this issue by automating parameter range selection using real-world experience. While effective, these algorithms often require long computation time, as a new policy is trained from scratch every iteration. In this work, we propose Adaptive Bayesian Domain Randomization via Strategic Fine-tuning (BayRnTune), which inherits the spirit of BayRn but aims to significantly accelerate the learning processes by fine-tuning from previously learned policy. This idea leads to a critical question: which previous policy should we use as a prior during fine-tuning? We investigated four different fine-tuning strategies and compared them against baseline algorithms in five simulated environments, ranging from simple benchmark tasks to more complex legged robot environments. Our analysis demonstrates that our method yields better rewards in the same amount of timesteps compared to vanilla domain randomization or Bayesian DR.
Tianle Huang, Nitish Sontakke, K. Niranjan Kumar, Irfan A. Essa, Stefanos Nikolaidis, Dennis W. Hong, Sehoon Ha
IROS4
2024 FineStyle: Fine-grained Controllable Style Personalization for Text-to-image Models
abstract
Few-shot fine-tuning of text-to-image (T2I) generation models enables people to create unique images in their own style using natural languages without requiring extensive prompt engineering. However, fine-tuning with only a handful, as little as one, of image-text paired data prevents fine-grained control of style attributes at generation. In this paper, we present FineStyle, a few-shot fine-tuning method that allows enhanced controllability for style personalized text-to-image generation. To overcome the lack of training data for fine-tuning, we propose a novel concept-oriented data scaling that amplifies the number of image-text pair, each of which focuses on different concepts (e.g., objects) in the style reference image. We also identify the benefit of parameter-efficient adapter tuning of key and value kernels of cross-attention layers. Extensive experiments show the effectiveness of FineStyle at following fine-grained text prompts and delivering visual quality faithful to the specified style, measured by CLIP scores and human raters.
Gong Zhang 0011, Kihyuk Sohn, Meera Hahn, Humphrey Shi, Irfan A. Essa
NeurIPS5
2023 Integrating Noisy Knowledge into Language Representations for E-Commerce Applications
abstract
Integrating structured knowledge into language model representations increases recall of domain-specific information useful for downstream tasks. Matching between knowledge graph entities and text entity mentions can be easily performed when entity names are unique or there exists entity linking data. When extending this setting to new domains, newly mined knowledge contains ambiguous and incorrect information, with no explicit linking information. In such settings, we design a framework to robustly link relevant knowledge to input texts as an intermediate modeling step while performing end-to-end domain fine-tuning tasks. This is done by first computing the similarity of the existing task labels with candidate knowledge triplets to generate relevance labels. We use these labels to train a relevance model, which predicts the relevance of the inserted triplets to the original text. This relevance model is integrated within a language model, leading to our Knowledge Relevance BERT (KR-BERT) framework. We test KR-BERT for entity linking tasks on a real-world e-commerce dataset as well as a public linking task, where we show performance improvements over strong baselines.
Karan Samel, Irfan A. Essa
IEEE Big Data5
2023 Text and Click inputs for unambiguous open vocabulary instance segmentation
Vighnesh Birodkar, Jonathan Huang, Meera Hahn, Irfan A. Essa, Nikolai Warner
BMVC4
2023 Slide Gestalt: Automatic Structure Extraction in Slide Decks for Non-Visual Access
abstract
Presentation slides commonly use visual patterns for structural navigation, such as titles, dividers, and build slides. However, screen readers do not capture such intention, making it time-consuming and less accessible for blind and visually impaired (BVI) users to linearly consume slides with repeated content. We present Slide Gestalt, an automatic approach that identifies the hierarchical structure in a slide deck. Slide Gestalt computes the visual and textual correspondences between slides to generate hierarchical groupings. Readers can navigate the slide deck from the higher-level section overview to the lower-level description of a slide group or individual elements interactively with our UI. We derived side consumption and authoring practices from interviews with BVI readers and sighted creators and an analysis of 100 decks. We performed our pipeline with 50 real-world slide decks and a large dataset. Feedback from eight BVI participants showed that Slide Gestalt helped navigate a slide deck by anchoring content more efficiently, compared to using accessible slides.
Yi-Hao Peng, Peggy Chi, Anjuli Kannan, Meredith Ringel Morris, Irfan A. Essa
CHI5
2023 MaskSketch: Unpaired Structure-guided Masked Image Generation
abstract
Recent conditional image generation methods produce images of remarkable diversity, fidelity and realism. However, the majority of these methods allow conditioning only on labels or text prompts, which limits their level of control over the generation result. In this paper, we introduce MaskSketch, an image generation method that allows spatial conditioning of the generation result using a guiding sketch as an extra conditioning signal during sampling. MaskSketch utilizes a pretrained masked generative transformer, requiring no model training or paired supervision, and works with input sketches of different levels of abstraction. We show that intermediate self-attention maps of a masked generative transformer encode important structural information of the input image, such as scene layout and object shape, and we propose a novel sampling method based on this observation to enable structure-guided generation. Our results show that MaskSketch achieves high image realism and fidelity to the guiding structure. Evaluated on standard benchmark datasets, MaskSketch outperforms state-of-the-art methods for sketch-to-image translation, as well as unpaired image-to-image translation approaches. The code can be found on our project website: https://masksketch.github.io/
Dina Bashkirova, José Lezama, Kihyuk Sohn, Kate Saenko, Irfan A. Essa
CVPR5
2023 Visual Prompt Tuning for Generative Transfer Learning
abstract
Learning generative image models from various domains efficiently needs transferring knowledge from an image synthesis model trained on a large dataset. We present a recipe for learning vision transformers by generative knowledge transfer. We base our framework on generative vision transformers representing an image as a sequence of visual tokens with the autoregressive or non-autoregressive transformers. To adapt to a new domain, we employ prompt tuning, which prepends learnable tokens called prompts to the image token sequence and introduces a new prompt design for our task. We study on a variety of visual domains with varying amounts of training images. We show the effectiveness of knowledge transfer and a significantly better image generation quality.11https://github.com/google-research/generative_transfer
Kihyuk Sohn, Huiwen Chang, José Lezama, Luisa Polania, Han Zhang 0010, Yuan Hao, Irfan A. Essa, Lu Jiang 0004
CVPR7
2023 MAGVIT: Masked Generative Video Transformer
abstract
We introduce the MAsked Generative VIdeo Transformer, MAGVIT, to tackle various video synthesis tasks with a single model. We introduce a 3D tokenizer to quantize a video into spatial-temporal visual tokens and propose an embedding method for masked video token modeling to facilitate multi-task learning. We conduct extensive experiments to demonstrate the quality, efficiency, and flexibility of MAGVIT. Our experiments show that (i) MAGVIT performs favorably against state-of-the-art approaches and establishes the best-published FVD on three video generation benchmarks, including the challenging Kinetics-600. (ii) MAGVIT outperforms existing methods in inference time by two orders of magnitude against diffusion models and by 60x against autoregressive models. (iii) A single MAGVIT model supports ten diverse generation tasks and generalizes across videos from different visual domains. The source code and trained models will be released to the public at https://magvit.cs.cmu.edu.
Lijun Yu, Yong Cheng 0003, Kihyuk Sohn, José Lezama, Han Zhang 0010, Huiwen Chang, Alex Hauptmann 0001, Ming-Hsuan Yang 0001, Yuan Hao, Irfan A. Essa, Lu Jiang 0004
CVPR10
2023 Discrete Predictor-Corrector Diffusion Models for Image Synthesis
José Lezama, Tim Salimans, Lu Jiang 0004, Huiwen Chang, Jonathan Ho, Irfan A. Essa
ICLR6
2023 Emergence of Maps in the Memories of Blind Navigation Agents
Erik Wijmans, Manolis Savva, Irfan A. Essa, Stefan Lee, Ari S. Morcos, Dhruv Batra
ICLR3
2023 StyleDrop: Text-to-Image Synthesis of Any Style
abstract
Pre-trained large text-to-image models synthesize impressive images with an appropriate use of text prompts. However, ambiguities inherent in natural language, and out-of-distribution effects make it hard to synthesize arbitrary image styles, leveraging a specific design pattern, texture or material. In this paper, we introduce *StyleDrop*, a method that enables the synthesis of images that faithfully follow a specific style using a text-to-image model. StyleDrop is extremely versatile and captures nuances and details of a user-provided style, such as color schemes, shading, design patterns, and local and global effects. StyleDrop works by efficiently learning a new style by fine-tuning very few trainable parameters (less than 1\% of total model parameters), and improving the quality via iterative training with either human or automated feedback. Better yet, StyleDrop is able to deliver impressive results even when the user supplies only a *single* image specifying the desired style. An extensive study shows that, for the task of style tuning text-to-image models, StyleDrop on Muse convincingly outperforms other methods, including DreamBooth and textual inversion on Imagen or Stable Diffusion. More results are available at our project website: [https://styledrop.github.io](https://styledrop.github.io).
Kihyuk Sohn, Lu Jiang 0004, Jarred Barber, Kimin Lee, Nataniel Ruiz, Dilip Krishnan, Huiwen Chang, Yuanzhen Li, Irfan A. Essa, Michael Rubinstein, Yuan Hao, Glenn Entis, Irina Blok, Daniel Castro Chin
NeurIPS9
2023 SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs
abstract
In this work, we introduce Semantic Pyramid AutoEncoder (SPAE) for enabling frozen LLMs to perform both understanding and generation tasks involving non-linguistic modalities such as images or videos. SPAE converts between raw pixels and interpretable lexical tokens (or words) extracted from the LLM's vocabulary. The resulting tokens capture both the rich semantic meaning and the fine-grained details needed for visual reconstruction, effectively translating the visual content into a language comprehensible to the LLM, and empowering it to perform a wide array of multimodal tasks. Our approach is validated through in-context learning experiments with frozen PaLM 2 and GPT 3.5 on a diverse set of image understanding and generation tasks. Our method marks the first successful attempt to enable a frozen LLM to generate image content while surpassing state-of-the-art performance in image understanding tasks, under the same setting, by over 25%.
Lijun Yu, Yong Cheng 0003, Zhiruo Wang 0001, Wolfgang Macherey, Yanping Huang, David A. Ross, Irfan A. Essa, Yonatan Bisk, Ming-Hsuan Yang 0001, Kevin Murphy 0002, Alex Hauptmann 0001, Lu Jiang 0004
NeurIPS8
2023 Investigating Enhancements to Contrastive Predictive Coding for Human Activity Recognition
abstract
The dichotomy between the challenging nature of obtaining annotations for activities, and the more straightforward nature of data collection from wearables, has resulted in significant interest in the development of techniques that utilize large quantities of unlabeled data for learning representations. Contrastive Predictive Coding (CPC) is one such method, learning effective representations by leveraging properties of time-series data to setup a contrastive future timestep prediction task. In this work, we propose enhancements to CPC, by systematically investigating the encoder architecture, the aggregator network, and the future timestep prediction, resulting in a fully con-volutional architecture. Across sensor positions and activities, our method shows substantial improvements on four of six target datasets, demonstrating its ability to empower a wide range of application scenarios. Further, in the presence of very limited labeled data, our technique significantly outperforms both supervised and self-supervised baselines, positively impacting situations where collecting only a few seconds of labeled data may be possible. This is promising, as CPC does not require specialized data transformations or reconstructions for learning effective representations.
Harish Haresamudram, Irfan A. Essa, Thomas Plötz
PERCOM2
2022 BLT: Bidirectional Layout Transformer for Controllable Layout Generation
Xiang Kong, Lu Jiang 0004, Huiwen Chang, Han Zhang 0010, Yuan Hao, Haifeng Gong, Irfan A. Essa
ECCV (17)7
2022 Improved Masked Image Generation with Token-Critic
José Lezama, Huiwen Chang, Lu Jiang 0004, Irfan A. Essa
ECCV (23)4
2022 Discrete Representations Strengthen Vision Transformer Robustness
Chengzhi Mao, Lu Jiang 0004, Mostafa Dehghani 0001, Carl Vondrick, Rahul Sukthankar, Irfan A. Essa
ICLR6
2022 Graph-based Cluttered Scene Generation and Interactive Exploration using Deep Reinforcement Learning
abstract
We introduce a novel method to teach a robotic agent to interactively explore cluttered yet structured scenes, such as kitchen pantries and grocery shelves, by leveraging the physical plausibility of the scene. We propose a novel learning framework to train an effective scene exploration policy to discover hidden objects with minimal interactions. First, we define a novel scene grammar to represent structured clutter. Then we train a Graph Neural Network (GNN) based Scene Generation agent using deep reinforcement learning (deep RL), to manipulate this Scene Grammar to create a diverse set of stable scenes, each containing multiple hidden objects. Given such cluttered scenes, we then train a Scene Exploration agent, using deep RL, to uncover hidden objects by interactively rearranging the scene. We show that our learned agents hide and discover significantly more objects than the baselines. We present quantitative results that prove the generalization capabilities of our agents. We also demonstrate sim-to-real transfer by successfully deploying the learned policy on a real UR10 robot to explore real-world cluttered scenes. The supplemental video can be found at: https://www.youtube.com/watch?v=T2Jo7wwaXss.
K. Niranjan Kumar, Irfan A. Essa, Sehoon Ha
ICRA2
2022 Tackling Hate Speech in Low-resource Languages with Context Experts
abstract
Given Myanmar’s historical and socio-political context, hate speech spread on social media have escalated into offline unrest and violence. This paper presents findings from our remote study on the automatic detection of hate speech online in Myanmar. We argue that effectively addressing this problem will require community-based approaches that combine the knowledge of context experts with machine learning tools that can analyze the vast amount of data produced. To this end, we develop a systematic process to facilitate this collaboration covering key aspects of data collection, annotation, and model validation strategies. We highlight challenges in this area stemming from small and imbalanced datasets, the need to balance non-glamorous data work and stakeholder priorities, and closed data-sharing practices. Stemming from these findings, we discuss avenues for further work in developing and deploying hate speech detection systems for low-resource languages.
Daniel Nkemelu, Harshil Shah, Michael L. Best, Irfan A. Essa
ICTD4
2022 VER: Scaling On-Policy RL Leads to the Emergence of Navigation in Embodied Rearrangement
abstract
We present Variable Experience Rollout (VER), a technique for efficiently scaling batched on-policy reinforcement learning in heterogenous environments (where different environments take vastly different times to generate rollouts) to many GPUs residing on, potentially, many machines. VER combines the strengths of and blurs the line between synchronous and asynchronous on-policy RL methods (SyncOnRL and AsyncOnRL, respectively). Specifically, it learns from on-policy experience (like SyncOnRL) and has no synchronization points (like AsyncOnRL) enabling high throughput.We find that VER leads to significant and consistent speed-ups across a broad range of embodied navigation and mobile manipulation tasks in photorealistic 3D simulation environments. Specifically, for PointGoal navigation and ObjectGoal navigation in Habitat 1.0, VER is 60-100% faster (1.6-2x speedup) than DD-PPO, the current state of art for distributed SyncOnRL, with similar sample efficiency. For mobile manipulation tasks (open fridge/cabinet, pick/place objects) in Habitat 2.0 VER is 150% faster (2.5x speedup) on 1 GPU and 170% faster (2.7x speedup) on 8 GPUs than DD-PPO. Compared to SampleFactory (the current state-of-the-art AsyncOnRL), VER matches its speed on 1 GPU, and is 70% faster (1.7x speedup) on 8 GPUs with better sample efficiency.We leverage these speed-ups to train chained skills for GeometricGoal rearrangement tasks in the Home Assistant Benchmark (HAB). We find a surprising emergence of navigation in skills that do not ostensible require any navigation. Specifically, the Pick skill involves a robot picking an object from a table. During training the robot was always spawned close to the table and never needed to navigate. However, we find that if base movement is part of the action space, the robot learns to navigate then pick an object in new environments with 50% success, demonstrating surprisingly high out-of-distribution generalization.
Erik Wijmans, Irfan A. Essa, Dhruv Batra
NeurIPS2
2022 Synthesis-Assisted Video Prototyping From a Document
abstract
Video productions commonly start with a script, especially for talking head videos that feature a speaker narrating to the camera. When the source materials come from a written document – such as a web tutorial, it takes iterations to refine content from a text article to a spoken dialogue, while considering visual compositions in each scene. We propose Doc2Video, a video prototyping approach that converts a document to interactive scripting with a preview of synthetic talking head videos. Our pipeline decomposes a source document into a series of scenes, each automatically creating a synthesized video of a virtual instructor. Designed for a specific domain – programming cookbooks, we apply visual elements from the source document, such as a keyword, a code snippet or a screenshot, in suitable layouts. Users edit narration sentences, break or combine sections, and modify visuals to prototype a video in our Editing UI. We evaluated our pipeline with public programming cookbooks. Feedback from professional creators shows that our method provided a reasonable starting point to engage them in interactive scripting for a narrated instructional video.
Peggy Chi, Christian Früh, Brian Colonna, Vivek Kwatra, Irfan A. Essa
UIST6
2022 Sharing Decoders: Network Fission for Multi-task Pixel Prediction
abstract
We examine the benefits of splitting encoder-decoders for multitask learning and showcase results on three tasks (semantics, surface normals, and depth) while adding very few FLOPS per task. Current hard parameter sharing methods for multi-task pixel-wise labeling use one shared encoder with separate decoders for each task. We generalize this notion and term the splitting of encoder-decoder architectures at different points as fission. Our ablation studies on fission show that sharing most of the decoder layers in multi-task encoder-decoder networks results in improvement while adding far fewer parameters per task. Our proposed method trains faster, uses less memory, results in better accuracy, and uses significantly fewer floating point operations (FLOPS) than conventional multi-task methods, with additional tasks only requiring 0.017% more FLOPS than the single-task network. We show results with a real-time model on a Pixel phone with released source code.
Steven Hickson, Karthik Raveendran, Irfan A. Essa
WACV3
2021 Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views
abstract
We study the task of semantic mapping – specifically, an embodied agent (a robot or an egocentric AI assistant) is given a tour of a new environment and asked to build an allocentric top-down semantic map (‘what is where?’) from egocentric observations of an RGB-D camera with known pose (via localization sensors). Importantly, our goal is to build neural episodic memories and spatio-semantic representations of 3D spaces that enable the agent to easily learn subsequent tasks in the same space – navigating to objects seen during the tour (‘Find chair’) or answering questions about the space (‘How many chairs did you see in the house?’). Towards this goal, we present Semantic MapNet (SMNet), which consists of: (1) an Egocentric Visual Encoder that encodes each egocentric RGB-D frame, (2) a Feature Projector that projects egocentric features to appropriate locations on a floor-plan, (3) a Spatial Memory Tensor of size floor-plan length×width×feature-dims that learns to accumulate projected egocentric features, and (4) a Map Decoder that uses the memory tensor to produce semantic top-down maps. SMNet combines the strengths of (known) projective camera geometry and neural representation learning. On the task of semantic mapping in the Matterport3D dataset, SMNet significantly outperforms competitive baselines by 4.01−16.81% (absolute) on mean-IoU and 3.81−19.69% (absolute) on Boundary-F1 metrics. Moreover, we show how to use the spatio-semantic allocentric representations build by SMNet for the task of ObjectNav and Embodied Question Answering. Project page: https://vincentcartillier.github.io/smnet.html.
Vincent Cartillier, Zhile Ren, Stefan Lee, Irfan A. Essa, Dhruv Batra
AAAI5
2021 Unsupervised Discovery of Actions in Instructional Videos
A. J. Piergiovanni, Anelia Angelova, Michael S. Ryoo, Irfan A. Essa
BMVC4
2021 Automatic Generation of Two-Level Hierarchical Tutorials from Instructional Makeup Videos
abstract
We present a multi-modal approach for automatically generating hierarchical tutorials from instructional makeup videos. Our approach is inspired by prior research in cognitive psychology, which suggests that people mentally segment procedural tasks into event hierarchies, where coarse-grained events focus on objects while fine-grained events focus on actions. In the instructional makeup domain, we find that objects correspond to facial parts while fine-grained steps correspond to actions on those facial parts. Given an input instructional makeup video, we apply a set of heuristics that combine computer vision techniques with transcript text analysis to automatically identify the fine-level action steps and group these steps by facial part to form the coarse-level events. We provide a voice-enabled, mixed-media UI to visualize the resulting hierarchy and allow users to efficiently navigate the tutorial (e.g., skip ahead, return to previous steps) at their own pace. Users can navigate the hierarchy at both the facial-part and action-step levels using click-based interactions and voice commands. We demonstrate the effectiveness of segmentation algorithms and the resulting mixed-media UI on a variety of input makeup videos. A user study shows that users prefer following instructional makeup videos in our mixed-media format to the standard video UI and that they find our format much easier to navigate.
Anh Truong, Peggy Chi, David Salesin, Irfan A. Essa, Maneesh Agrawala
CHI4
2021 Text as Neural Operator: Image Manipulation by Text Instruction
abstract
n recent years, text-guided image manipulation has gained increasing attention in the multimedia and computer vision community. The input to conditional image generation has evolved from image-only to multimodality. In this paper, we study a setting that allows users to edit an image with multiple objects using complex text instructions to add, remove, or change the objects. The inputs of the task are multimodal including (1) a reference image and (2) an instruction in natural language that describes desired modifications to the image. We propose a GAN-based method to tackle this problem. The key idea is to treat text as neural operators to locally modify the image feature. We show that the proposed model performs favorably against recent strong baselines on three public datasets. Specifically, it generates images of greater fidelity and semantic relevance, and when used as a image query, leads to better retrieval performance.
Hung-Yu Tseng, Lu Jiang 0004, Weilong Yang, Honglak Lee, Irfan A. Essa
ACM Multimedia6
2021 Automatic Instructional Video Creation from a Markdown-Formatted Tutorial
abstract
We introduce HowToCut, an automatic approach that converts a Markdown-formatted tutorial into an interactive video that presents the visual instructions with a synthesized voiceover for narration. HowToCut extracts instructional content from a multimedia document that describes a step-by-step procedure. Our method selects and converts text instructions to a voiceover. It makes automatic editing decisions to align the narration with edited visual assets, including step images, videos, and text overlays. We derive our video editing strategies from an analysis of 125 web tutorials and apply Computer Vision techniques to the assets. To enable viewers to interactively navigate the tutorial, HowToCut’s conversational UI presents instructions in multiple formats upon user commands. We evaluated our automatically-generated video tutorials through user studies (N=20) and validated the video quality via an online survey (N=93). The evaluation shows that our method was able to effectively create informative and useful instructional videos from a web tutorial document for both reviewing and following.
Peggy Chi, Nathan Frey, Katrina Panovich, Irfan A. Essa
UIST4
2020 Neural Design Network: Graphic Layout Generation with Constraints
Hsin-Ying Lee 0001, Lu Jiang 0004, Irfan A. Essa, Phuong B. Le, Haifeng Gong, Ming-Hsuan Yang 0001, Weilong Yang
ECCV (3)3
2020 DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames
Erik Wijmans, Abhishek Kadian, Ari S. Morcos, Stefan Lee, Irfan A. Essa, Devi Parikh, Manolis Savva, Dhruv Batra
ICLR5
2020 Automatic Video Creation From a Web Page
abstract
Creating marketing videos from scratch can be challenging, especially when designing for multiple platforms with different viewing criteria. We present URL2Video, an automatic approach that converts a web page into a short video given temporal and visual constraints. URL2Video captures quality materials and design styles extracted from a web page, including fonts, colors, and layouts. Using constraint programming, URL2Video's design engine organizes the visual assets into a sequence of shots and renders to a video with user-specified aspect ratio and duration. Creators can review the video composition, modify constraints, and generate video variation through a user interface. We learned the design process from designers and compared our automatically generated results with their creation through interviews and an online survey. The evaluation shows that URL2Video effectively extracted design elements from a web page and supported designers by bootstrapping the video creation process.
Peggy Chi, Katrina Panovich, Irfan A. Essa
UIST4
2019 Audio Visual Scene-Aware Dialog
abstract
We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history of previous turns in the dialog. To answer successfully, agents must ground concepts from the question in the video while leveraging contextual cues from the dialog history. To benchmark this task, we introduce the Audio Visual Scene-Aware Dialog (AVSD) Dataset. For each of more than 11,000 videos of human actions from the Charades dataset, our dataset contains a dialog about the video, plus a final summary of the video by one of the dialog participants. We train several baseline systems for this task and evaluate the performance of the trained models using both qualitative and quantitative metrics. Our results indicate that models must utilize all the available inputs (video, audio, question, and dialog history) to perform best on this dataset.
Huda AlAmri, Vincent Cartillier, Abhishek Das 0002, Jue Wang 0010, Anoop Cherian, Irfan A. Essa, Dhruv Batra, Tim K. Marks, Chiori Hori, Stefan Lee, Devi Parikh
CVPR6
2019 Embodied Question Answering in Photorealistic Environments With Point Cloud Perception
abstract
To help bridge the gap between internet vision-style problems and the goal of vision for embodied perception we instantiate a large-scale navigation task -- Embodied Question Answering [1] in photo-realistic environments (Matterport 3D). We thoroughly study navigation policies that utilize 3D point clouds, RGB images, or their combination. Our analysis of these models reveals several key findings. We find that two seemingly naive navigation baselines, forward-only and random, are strong navigators and challenging to outperform, due to the specific choice of the evaluation setting presented by [1]. We find a novel loss-weighting scheme we call Inflection Weighting to be important when training recurrent models for navigation with behavior cloning and are able to out perform the baselines with this technique. We find that point clouds provide a richer signal than RGB images for learning obstacle avoidance, motivating the use (and continued study) of 3D deep learning models for embodied navigation.
Erik Wijmans, Samyak Datta, Oleksandr Maksymets, Abhishek Das 0002, Georgia Gkioxari, Stefan Lee, Irfan A. Essa, Devi Parikh, Dhruv Batra
CVPR7
2019 End-to-end Audio Visual Scene-aware Dialog Using Multimodal Attention-based Video Features
abstract
In order for machines interacting with the real world to have conversations with users about the objects and events around them, they need to understand dynamic audiovisual scenes. The recent revolution of neural network models allows us to combine various modules into a single end-to-end differentiable network. As a result, Audio Visual Scene-Aware Dialog (AVSD) systems for real-world applications can be developed by integrating state-of-the-art technologies from multiple research areas, including end-to-end dialog technologies, visual question answering (VQA) technologies, and video description technologies. In this paper, we introduce a new data set of dialogs about videos of human behaviors, as well as an end-to-end Audio Visual Scene-Aware Dialog (AVSD) model, trained using this new data set, that generates responses in a dialog about a video. By using features that were developed for multimodal attention-based video description, our system improves the quality of generated dialog about dynamic video scenes.
Chiori Hori, Huda AlAmri, Jue Wang 0010, Gordon Wichern, Takaaki Hori, Anoop Cherian, Tim K. Marks, Vincent Cartillier, Raphael Gontijo Lopes, Abhishek Das 0002, Irfan A. Essa, Dhruv Batra, Devi Parikh
ICASSP11
2019 A Data-Driven Predictive Model of Individual-Specific Effects of FES on Human Gait Dynamics
abstract
Modeling individual-specific gait dynamics based on kinematic data could aid development of gait rehabilitation robotics by enabling robots to predict the user's gait kinematics with and without external inputs, such as mechanical or electrical perturbations. Here we address a current limitation of data-driven gait models, which do not yet predict human gait dynamics nor responses to perturbations. We used Switched Linear Dynamical Systems (SLDS) to model joint angle kinematic data from healthy individuals walking on a treadmill during normal gait and during gait perturbed by functional electrical stimulation (FES) to the ankle muscles. Our SLDS models were able to generate joint angle trajectories in each of four gait phases, as well as across an entire gait cycle, given initial conditions and gait phase information. Because the SLDS dynamics matrices encoded significant coupling across joints that differed across indivdiuals, we compared the SLDS predictions to that of a kinematic model, where the joint angles were independent. Joint angle trajectories generated by SLDS and kinematic models were similar over time horizons of a few milliseconds, but SLDS models provided better predictions of gait kinematics over time horizons of up to a second. We also demonstrated that SLDS models can infer and predict individual-specific responses to FES during swing phase. As such, SLDS models may be a promising approach for online estimation and control of and human gait dynamics, allowing robotic control strategies to be tailored to an individual's specific gait coordination patterns.
Luke Drnach, Jessica L. Allen, Irfan A. Essa, Lena H. Ting
ICRA3
2019 Video Jigsaw: Unsupervised Learning of Spatiotemporal Context for Video Action Recognition
abstract
We propose a self-supervised learning method to jointly reason about spatial and temporal context for video recognition. Recent self-supervised approaches have used spatial context [9, 34] as well as temporal coherency [32] but a combination of the two requires extensive preprocessing such as tracking objects through millions of video frames [59] or computing optical flow to determine frame regions with high motion [30]. We propose to combine spatial and temporal context in one self-supervised framework without any heavy preprocessing. We divide multiple video frames into grids of patches and train a network to solve jigsaw puzzles on these patches from multiple frames. So the network is trained to correctly identify the position of a patch within a video frame as well as the position of a patch over time. We also propose a novel permutation strategy that outperforms random permutations while significantly reducing computational and memory constraints. We use our trained network for transfer learning tasks such as video activity recognition and demonstrate the strength of our approach on two benchmark video action recognition datasets without using a single frame from these datasets for unsupervised pretraining of our proposed video jigsaw network.
Unaiza Ahsan, Rishi Madhok, Irfan A. Essa
WACV3
2019 Eyemotion: Classifying Facial Expressions in VR Using Eye-Tracking Cameras
abstract
One of the main challenges of social interaction in virtual reality settings is that head-mounted displays occlude a large portion of the face, blocking facial expressions and thereby restricting social engagement cues among users. We present an algorithm to automatically infer expressions by analyzing only a partially occluded face while the user is engaged in a virtual reality experience. Specifically, we show that images of the user's eyes captured from an IR gaze-tracking camera within a VR headset are sufficient to infer a subset of facial expressions without the use of any fixed external camera. Using these inferences, we can generate dynamic avatars in real-time which function as an expressive surrogate for the user. We propose a novel data collection pipeline as well as a novel approach for increasing CNN accuracy via personalization. Our results show a mean accuracy of 74% (F1 of 0.73) among 5 'emotive' expressions and a mean accuracy of 70% (F1 of 0.68) among 10 distinct facial action units, outperforming human raters.
Steven Hickson, Nick Dufour, Avneesh Sud, Vivek Kwatra, Irfan A. Essa
WACV5
2018 Surgical Activity Recognition in Robot-Assisted Radical Prostatectomy Using Deep Learning
Aneeq Zia, Andrew Hung, Irfan A. Essa, Anthony M. Jarc
MICCAI (4)3
2018 rtCaptcha: A Real-Time CAPTCHA Based Liveness Detection System
Erkam Uzun, Simon P. Chung, Irfan A. Essa, Wenke Lee
NDSS3
2017 One-Shot Learning for Semantic Segmentation
Amirreza Shaban, Shray Bansal, Zhen Liu 0019, Irfan A. Essa, Byron Boots
BMVC4
2017 Selfie-Presentation in Everyday Life: A Large-Scale Characterization of Selfie Contexts on Instagram
Julia Deeb-Swihart, Christopher Polack, Eric Gilbert, Irfan A. Essa
ICWSM4
2017 Towards using visual attributes to infer image sentiment of social events
abstract
Widespread and pervasive adoption of smartphones has led to instant sharing of photographs that capture events ranging from mundane to life-altering happenings. We propose to capture sentiment information of such social event images leveraging their visual content. Our method extracts an intermediate visual representation of social event images based on the visual attributes that occur in the images going beyond sentiment-specific attributes. We map the top predicted attributes to sentiments and extract the dominant emotion associated with a picture of a social event. Unlike recent approaches, our method generalizes to a variety of social events and even to unseen events, which are not available at training time. We demonstrate the effectiveness of our approach on a challenging social event image dataset and our method outperforms state-of-the-art approaches for classifying complex event images into sentiments.
Unaiza Ahsan, Munmun De Choudhury, Irfan A. Essa
IJCNN3
2017 Complex Event Recognition from Images with Few Training Examples
abstract
We propose to leverage concept-level representations for complex event recognition in photographs given limited training examples. We introduce a novel framework to discover event concept attributes from the web and use that to extract semantic features from images and classify them into social event categories with few training examples. Discovered concepts include a variety of objects, scenes, actions and event sub-types, leading to a discriminative and compact representation for event images. Web images are obtained for each discovered event concept and we use (pretrained) CNN features to train concept classifiers. Extensive experiments on challenging event datasets demonstrate that our proposed method outperforms several baselines using deep CNN features directly in classifying images into events with limited training examples. We also demonstrate that our method achieves the best overall accuracy on a dataset with unseen event categories using a single training example.
Unaiza Ahsan, Chen Sun 0002, James Hays, Irfan A. Essa
WACV4
2017 Computer Vision in Sports
Thomas B. Moeslund, Graham A. Thomas, Adrian Hilton 0001, Peter Carr 0001, Irfan A. Essa
Comput. Vis. Image Underst.5
2016 Leveraging Contextual Cues for Generating Basketball Highlights
abstract
The massive growth of sports videos has resulted in a need for automatic generation of sports highlights that are comparable in quality to the hand-edited highlights produced by broadcasters such as ESPN. Unlike previous works that mostly use audio-visual cues derived from the video, we propose an approach that additionally leverages contextual cues derived from the environment that the game is being played in. The contextual cues provide information about the excitement levels in the game, which can be ranked and selected to automatically produce high-quality basketball highlights. We introduce a new dataset of 25 NCAA games along with their play-by-play stats and the ground-truth excitement data for each basket. We explore the informativeness of five different cues derived from the video and from the environment through user studies. Our experiments show that for our study participants, the highlights produced by our system are comparable to the ones produced by ESPN for the same games.
Vinay Bettadapura, Caroline Pantofaru, Irfan A. Essa
ACM Multimedia3
2016 Discovering picturesque highlights from egocentric vacation videos
abstract
We present an approach for identifying picturesque highlights from large amounts of egocentric video data. Given a set of egocentric videos captured over the course of a vacation, our method analyzes the videos and looks for images that have good picturesque and artistic properties. We introduce novel techniques to automatically determine aesthetic features such as composition, symmetry and color vibrancy in egocentric videos and rank the video frames based on their photographic qualities to generate highlights. Our approach also uses contextual information such as GPS, when available, to assess the relative importance of each geographic location where the vacation videos were shot. Furthermore, we specifically leverage the properties of egocentric videos to improve our highlight detection. We demonstrate results on a new egocentric vacation dataset which includes 26.5 hours of videos taken over a 14 day vacation that spans many famous tourist destinations and also provide results from a user-study to access our results.
Vinay Bettadapura, Irfan A. Essa
WACV3
2015 A practical approach for recognizing eating moments with wrist-mounted inertial sensing
abstract
Recognizing when eating activities take place is one of the key challenges in automated food intake monitoring. Despite progress over the years, most proposed approaches have been largely impractical for everyday usage, requiring multiple on-body sensors or specialized devices such as neck collars for swallow detection. In this paper, we describe the implementation and evaluation of an approach for inferring eating moments based on 3-axis accelerometry collected with a popular off-the-shelf smartwatch. Trained with data collected in a semi-controlled laboratory setting with 20 subjects, our system recognized eating moments in two free-living condition studies (7 participants, 1 day; 1 participant, 31 days), with F-scores of 76.1% (66.7% Precision, 88.8% Recall), and 71.3% (65.2% Precision, 78.6% Recall). This work represents a contribution towards the implementation of a practical, automated system for everyday food intake monitoring, with applicability in areas ranging from health research and food journaling.
Edison Thomaz, Irfan A. Essa, Gregory D. Abowd
UbiComp2
2015 Inferring Meal Eating Activities in Real World Settings from Ambient Sounds: A Feasibility Study
abstract
Dietary self-monitoring has been shown to be an effective method for weight-loss, but it remains an onerous task despite recent advances in food journaling systems. Semi-automated food journaling can reduce the effort of logging, but often requires that eating activities be detected automatically. In this work we describe results from a feasibility study conducted in-the-wild where eating activities were inferred from ambient sounds captured with a wrist-mounted device; twenty participants wore the device during one day for an average of 5 hours while performing normal everyday activities. Our system was able to identify meal eating with an F-score of 79.8% in a person-dependent evaluation, and with 86.6% accuracy in a person-independent evaluation. Our approach is intended to be practical, leveraging off-the-shelf devices with audio sensing capabilities in contrast to systems for automated dietary assessment based on specialized sensors.
Edison Thomaz, Cheng Zhang 0011, Irfan A. Essa, Gregory D. Abowd
IUI3
2015 Automated Assessment of Surgical Skills Using Frequency Analysis
Aneeq Zia, Yachna Sharma, Vinay Bettadapura, Eric L. Sarin, Mark A. Clements, Irfan A. Essa
MICCAI (1)6
2015 Egocentric Field-of-View Localization Using First-Person Point-of-View Devices
abstract
We present a technique that uses images, videos and sensor data taken from first-person point-of-view devices to perform egocentric field-of-view (FOV) localization. We define egocentric FOV localization as capturing the visual information from a person's field-of-view in a given environment and transferring this information onto a reference corpus of images and videos of the same space, hence determining what a person is attending to. Our method matches images and video taken from the first-person perspective with the reference corpus and refines the results using the first-person's head orientation information obtained using the device sensors. We demonstrate single and multi-user egocentric FOV localization in different indoor and outdoor environments with applications in augmented reality, event understanding and studying social interactions.
Vinay Bettadapura, Irfan A. Essa, Caroline Pantofaru
WACV2
2015 Leveraging Context to Support Automated Food Recognition in Restaurants
abstract
The pervasiveness of mobile cameras has resulted in a dramatic increase in food photos, which are pictures reflecting what people eat. In this paper, we study how taking pictures of what we eat in restaurants can be used for the purpose of automating food journaling. We propose to leverage the context of where the picture was taken, with additional information about the restaurant, available online, coupled with state-of-the-art computer vision techniques to recognize the food being consumed. To this end, we demonstrate image-based recognition of foods eaten in restaurants by training a classifier with images from restaurant's online menu databases. We evaluate the performance of our system in unconstrained, real-world settings with food images taken in 10 restaurants across 5 different types of food (American, Indian, Italian, Mexican and Thai).
Vinay Bettadapura, Edison Thomaz, Aman Parnami, Gregory D. Abowd, Irfan A. Essa
WACV5
2015 Semantic Instance Labeling Leveraging Hierarchical Segmentation
abstract
Most of the approaches for indoor RGBD semantic labeling focus on using pixels or super pixels to train a classifier. In this paper, we implement a higher level segmentation using a hierarchy of super pixels to obtain a better segmentation for training our classifier. By focusing on meaningful segments that conform more directly to objects, regardless of size, we train a random forest of decision trees as a classifier using simple features such as the 3D size, LAB color histogram, width, height, and shape as specified by a histogram of surface normal's. We test our method on the NYU V2 depth dataset, a challenging dataset of cluttered indoor environments. Our experiments using the NYU V2 depth dataset show that our method achieves state of the art results on both a general semantic labeling introduced by the dataset (floor, structure, furniture, and objects) and a more object specific semantic labeling. We show that training a classifier on a segmentation from a hierarchy of super pixels yields better results than training directly on super pixels, patches, or pixels as in previous work.
Steven Hickson, Irfan A. Essa, Henrik I. Christensen
WACV2
2015 Finding Temporally Consistent Occlusion Boundaries in Videos Using Geometric Context
abstract
We present an algorithm for finding temporally consistent occlusion boundaries in videos to support segmentation of dynamic scenes. We learn occlusion boundaries in a pair wise Markov random field (MRF) framework. We first estimate the probability of an spatio-temporal edge being an occlusion boundary by using appearance, flow, and geometric features. Next, we enforce occlusion boundary continuity in a MRF model by learning pair wise occlusion probabilities using a random forest. Then, we temporally smooth boundaries to remove temporal inconsistencies in occlusion boundary estimation. Our proposed framework provides an efficient approach for finding temporally consistent occlusion boundaries in video by utilizing causality, redundancy in videos, and semantic layout of the scene. We have developed a dataset with fully annotated ground-truth occlusion boundaries of over 30 videos (~5000 frames). This dataset is used to evaluate temporal occlusion boundaries and provides a much needed baseline for future studies. We perform experiments to demonstrate the role of scene layout, and temporal information for occlusion reasoning in dynamic scenes.
S. Hussain Raza 0001, Ahmad Humayun, Irfan A. Essa, Matthias Grundmann 0002
WACV3
2014 Depth Extraction from Videos Using Geometric Context and Occlusion Boundaries
Syed Raza, Omar Javed, Aveek Das, Harpreet Sawhney, Irfan A. Essa
BMVC6
2014 Efficient Hierarchical Graph-Based Segmentation of RGBD Videos
abstract
We present an efficient and scalable algorithm for segmenting 3D RGBD point clouds by combining depth, color, and temporal information using a multistage, hierarchical graph-based approach. Our algorithm processes a moving window over several point clouds to group similar regions over a graph, resulting in an initial over-segmentation. These regions are then merged to yield a dendrogram using agglomerative clustering via a minimum spanning tree algorithm. Bipartite graph matching at a given level of the hierarchical tree yields the final segmentation of the point clouds by maintaining region identities over arbitrarily long periods of time. We show that a multistage segmentation with depth then color yields better results than a linear combination of depth and color. Due to its incremental processing, our algorithm can process videos of any length and in a streaming pipeline. The algorithm's ability to produce robust, efficient segmentation is demonstrated with numerous experimental results on challenging sequences from our own as well as public RGBD data sets.
Steven Hickson, Stanley T. Birchfield, Irfan A. Essa, Henrik I. Christensen
CVPR3
2014 Measuring Child Visual Attention using Markerless Head Tracking from Color and Depth Sensing Cameras
abstract
A child's failure to respond to his or her name being called is an early warning sign for autism and response to name is currently assessed as a part of standard autism screening and diagnostic tools. In this paper, we explore markerless child head tracking as an unobtrusive approach for automatically predicting child response to name. Head turns are used as a proxy for visual attention. We analyzed 50 recorded response to name sessions with the goal of predicting if children, ages 15 to 30 months, responded to name calls by turning to look at an examiner within a defined time interval. The child's head turn angles and hand annotated child name call intervals were extracted from each session. Human assisted tracking was employed using an overhead Kinect camera, and automated tracking was later employed using an additional forward facing camera as a proof-of-concept. We explore two distinct analytical approaches for predicting child responses, one relying on rule-based approached and another on random forest classification. In addition, we derive child response latency as a new measurement that could provide researchers and clinicians with finer grain quantitative information currently unavailable in the field due to human limitations. Finally we reflect on steps for adapting our system to work in less constrained natural settings.
Jonathan Bidwell, Irfan A. Essa, Agata Rozga, Gregory D. Abowd
ICMI2
2014 A visualization framework for team sports captured using multiple static cameras
Raffay Hamid, Ramkrishan K. Kumar, Jessica K. Hodgins, Irfan A. Essa
Comput. Vis. Image Underst.4
2013 Beyond Sentiment: The Manifold of Human Emotions
abstract
Sentiment analysis predicts the presence of positive or negative emotions in a text document. In this paper we consider higher dimensional extensions of the sentiment concept, which represent a richer set of human emotions. Our approach goes beyond previous work in that our model contains a continuous manifold rather than a finite set of human emotions. We investigate the resulting model, compare it to psychological observations, and explore its predictive capabilities. Besides obtaining significant improvements over a baseline without manifold, we are also able to visualize different notions of positive sentiment in different domains.
Seungyeon Kim 0001, Fuxin Li, Guy Lebanon, Irfan A. Essa
AISTATS4
2013 Augmenting Bag-of-Words: Data-Driven Discovery of Temporal and Structural Information for Activity Recognition
abstract
We present data-driven techniques to augment Bag of Words (BoW) models, which allow for more robust modeling and recognition of complex long-term activities, especially when the structure and topology of the activities are not known a priori. Our approach specifically addresses the limitations of standard BoW approaches, which fail to represent the underlying temporal and causal information that is inherent in activity streams. In addition, we also propose the use of randomly sampled regular expressions to discover and encode patterns in activities. We demonstrate the effectiveness of our approach in experimental evaluations where we successfully recognize activities and detect anomalies in four complex datasets.
Vinay Bettadapura, Grant Schindler, Thomas Plötz, Irfan A. Essa
CVPR4
2013 Geometric Context from Videos
abstract
We present a novel algorithm for estimating the broad 3D geometric structure of outdoor video scenes. Leveraging spatio-temporal video segmentation, we decompose a dynamic scene captured by a video into geometric classes, based on predictions made by region-classifiers that are trained on appearance and motion features. By examining the homogeneity of the prediction, we combine predictions across multiple segmentation hierarchy levels alleviating the need to determine the granularity a priori. We built a novel, extensive dataset on geometric context of video to evaluate our method, consisting of over 100 ground-truth annotated outdoor videos with over 20,000 frames. To further scale beyond this dataset, we propose a semi-supervised learning framework to expand the pool of labeled data with high confidence predictions obtained from unlabeled data. Our system produces an accurate prediction of geometric context of video achieving 96% accuracy across main geometric classes.
S. Hussain Raza 0001, Matthias Grundmann 0002, Irfan A. Essa
CVPR3
2013 Decoding Children's Social Behavior
abstract
We introduce a new problem domain for activity recognition: the analysis of children's social and communicative behaviors based on video and audio data. We specifically target interactions between children aged 1-2 years and an adult. Such interactions arise naturally in the diagnosis and treatment of developmental disorders such as autism. We introduce a new publicly-available dataset containing over 160 sessions of a 3-5 minute child-adult interaction. In each session, the adult examiner followed a semi-structured play interaction protocol which was designed to elicit a broad range of social behaviors. We identify the key technical challenges in analyzing these behaviors, and describe methods for decoding the interactions. We present experimental results that demonstrate the potential of the dataset to drive interesting research questions, and show preliminary results for multi-modal activity recognition.
James M. Rehg, Gregory D. Abowd, Agata Rozga, Mario Romero, Mark A. Clements, Stan Sclaroff, Irfan A. Essa, Opal Y. Ousley, Yin Li 0003, Chanho Kim, Hrishikesh Rao 0001, Jonathan C. Kim, Liliana Lo Presti, Jianming Zhang 0001, Denis Lantsman, Jonathan Bidwell, Zhefan Ye
CVPR7
2013 Technological approaches for addressing privacy concerns when recognizing eating behaviors with wearable cameras
abstract
First-person point-of-view (FPPOV) images taken by wearable cameras can be used to better understand people's eating habits. Human computation is a way to provide effective analysis of FPPOV images in cases where algorithmic approaches currently fail. However, privacy is a serious concern. We provide a framework, the privacy-saliency matrix, for understanding the balance between the eating information in an image and its potential privacy concerns. Using data gathered by 5 participants wearing a lanyard-mounted smartphone, we show how the framework can be used to quantitatively assess the effectiveness of four automated techniques (face detection, image cropping, location filtering and motion filtering) at reducing the privacy-infringing content of images while still maintaining evidence of eating behaviors throughout the day.
Edison Thomaz, Aman Parnami, Jonathan Bidwell, Irfan A. Essa, Gregory D. Abowd
UbiComp4
2013 Post-processing approach for radiometric self-calibration of video
abstract
We present a novel data-driven technique for radiometric self-calibration of video from an unknown camera. Our approach self-calibrates radiometric variations in video, and is applied as a post-process; there is no need to access the camera, and in particular it is applicable to internet videos. This technique builds on empirical evidence that in video the camera response function (CRF) should be regarded time variant, as it changes with scene content and exposure, instead of relying on a single camera response function. We show that a time-varying mixture of responses produces better accuracy and consistently reduces the error in mapping intensity to irradiance when compared to a single response model. Furthermore, our mixture model counteracts the effects of possible nonlinear exposure-dependent intensity perturbations and white-balance changes caused by proprietary camera firmware. We further show how radiometrically calibrated video improves the performance of other video analysis algorithms, enabling a video segmentation algorithm to be invariant to exposure and gain variations over the sequence. We validate our data-driven technique on videos from a variety of cameras and demonstrate the generality of our approach by applying it to internet video.
Matthias Grundmann 0002, Chris McClanahan, Sing Bing Kang, Irfan A. Essa
ICCP4
2013 Detecting insider threats in a real corporate database of computer usage activity
abstract
This paper reports on methods and results of an applied research project by a team consisting of SAIC and four universities to develop, integrate, and evaluate new approaches to detect the weak signals characteristic of insider threats on organizations' information systems. Our system combines structural and semantic information from a real corporate database of monitored activity on their users' computers to detect independently developed red team inserts of malicious insider activities. We have developed and applied multiple algorithms for anomaly detection based on suspected scenarios of malicious insider behavior, indicators of unusual activities, high-dimensional statistical patterns, temporal sequences, and normal graph evolution. Algorithms and representations for dynamic graph processing provide the ability to scale as needed for enterprise-level deployments on real-time data streams. We have also developed a visual language for specifying combinations of features, baselines, peer groups, time periods, and algorithms to detect anomalies suggestive of instances of insider threat behavior. We defined over 100 data features in seven categories based on approximately 5.5 million actions per day from approximately 5,500 users. We have achieved area under the ROC curve values of up to 0.979 and lift values of 65 on the top 50 user-days identified on two months of real data.
Ted E. Senator, Henry G. Goldberg, Alex Memory, William T. Young, Bradley Rees, Robert Pierce, Daniel Huang 0003, Matthew Reardon, David A. Bader, Edmond Chow, Irfan A. Essa, Joshua Jones, Vinay Bettadapura, Polo Chau, Oded Green, Oguz Kaya, Anita Zakrzewska, Erica Briscoe, Rudolph Louis Mappus IV, Robert McColl, Lora Weiss, Thomas G. Dietterich, Alan Fern, Weng-Keen Wong, Shubhomoy Das, Andrew Emmott, Jed Irvine, Jay-Yoon Lee, Danai Koutra, Christos Faloutsos, Daniel D. Corkill, Lisa Friedland, Amanda Gentzel, David D. Jensen
KDD11
2012 Detecting regions of interest in dynamic scenes with camera motions
abstract
We present a method to detect the regions of interests in moving camera views of dynamic scenes with multiple moving objects. We start by extracting a global motion tendency that reflects the scene context by tracking movements of objects in the scene. We then use Gaussian process regression to represent the extracted motion tendency as a stochastic vector field. The generated stochastic field is robust to noise and can handle a video from an uncalibrated moving camera. We use the stochastic field for predicting important future regions of interest as the scene evolves dynamically. We evaluate our approach on a variety of videos of team sports and compare the detected regions of interest to the camera motion generated by actual camera operators. Our experimental results demonstrate that our approach is computationally efficient and provides better predictions than previously proposed RBF-based approaches.
Dongryeol Lee, Irfan A. Essa
CVPR3
2012 Recognizing water-based activities in the home through infrastructure-mediated sensing
abstract
Activity recognition in the home has been long recognized as the foundation for many desirable applications in fields such as home automation, sustainability, and healthcare. However, building a practical home activity monitoring system remains a challenge. Striking a balance between cost, privacy, ease of installation and scalability continues to be an elusive goal. In this paper, we explore infrastructure-mediated sensing combined with a vector space model learning approach as the basis of an activity recognition system for the home. We examine the performance of our single-sensor water-based system in recognizing eleven high-level activities in the kitchen and bathroom, such as cooking and shaving. Results from two studies show that our system can estimate activities with overall accuracy of 82.69% for one individual and 70.11% for a group of 23 participants. As far as we know, our work is the first to employ infrastructure-mediated sensing for inferring high-level human activities in a home setting.
Edison Thomaz, Vinay Bettadapura, Gabriel Reyes, Megha Sandesh, Grant Schindler, Thomas Plötz, Gregory D. Abowd, Irfan A. Essa
UbiComp8
2012 Orientation-aware scene understanding for mobile cameras
abstract
We present a novel approach that allows anyone to quickly teach their smartphone how to understand the visual world around them. We achieve this visual scene understanding by leveraging a camera-phone's inertial sensors to lead to both a faster and more accurate automatic labeling of the regions of an image into semantic classes (e.g. sky, tree, building). We focus on letting a user train our system from scratch while out in the real world by annotating image regions in situ as training images are captured on a mobile device, making it possible to recognize new environments and new semantic classes on the fly. We show that our approach outperforms existing methods, while at the same time performing data collection, annotation, feature extraction, and image segment classification all on the same mobile device.
Grant Schindler, Irfan A. Essa
UbiComp3
2012 Calibration-free rolling shutter removal
abstract
We present a novel algorithm for efficient removal of rolling shutter distortions in uncalibrated streaming videos. Our proposed method is calibration free as it does not need any knowledge of the camera used, nor does it require calibration using specially recorded calibration sequences. Our algorithm can perform rolling shutter removal under varying focal lengths, as in videos from CMOS cameras equipped with an optical zoom. We evaluate our approach across a broad range of cameras and video sequences demonstrating robustness, scaleability, and repeatability. We also conducted a user study, which demonstrates preference for the output of our algorithm over other state-of-the art methods. Our algorithm is computationally efficient, easy to parallelize, and robust to challenging artifacts introduced by various cameras with differing technologies.
Matthias Grundmann 0002, Vivek Kwatra, Irfan A. Essa
ICCP4
2012 Linguistic transfer of human assembly tasks to robots
abstract
We demonstrate the automatic transfer of an assembly task from human to robot. This work extends efforts showing the utility of linguistic models in verifiable robot control policies by now performing real visual analysis of human demonstrations to automatically extract a policy for the task. This method tokenizes each human demonstration into a sequence of object connection symbols, then transforms the set of sequences from all demonstrations into an automaton, which represents the task-language for assembling a desired object. Finally, we combine this assembly automaton with a kinematic model of a robot arm to reproduce the demonstrated task.
Neil Dantam, Irfan A. Essa, Mike Stilman
IROS2
2011 Auto-directed video stabilization with robust L1 optimal camera paths
abstract
We present a novel algorithm for automatically applying constrainable, L1-optimal camera paths to generate stabilized videos by removing undesired motions. Our goal is to compute camera paths that are composed of constant, linear and parabolic segments mimicking the camera motions employed by professional cinematographers. To this end, our algorithm is based on a linear programming framework to minimize the first, second, and third derivatives of the resulting camera path. Our method allows for video stabilization beyond the conventional filtering of camera paths that only suppresses high frequency jitter. We incorporate additional constraints on the path of the camera directly in our algorithm, allowing for stabilized and retargeted videos. Our approach accomplishes this without the need of user interaction or costly 3D reconstruction of the scene, and works as a post-process for videos from any camera or from an online source.
Matthias Grundmann 0002, Vivek Kwatra, Irfan A. Essa
CVPR3
2011 Gaussian process regression flow for analysis of motion trajectories
abstract
Recognition of motions and activities of objects in videos requires effective representations for analysis and matching of motion trajectories. In this paper, we introduce a new representation specifically aimed at matching motion trajectories. We model a trajectory as a continuous dense flow field from a sparse set of vector sequences using Gaussian Process Regression. Furthermore, we introduce a random sampling strategy for learning stable classes of motions from limited data. Our representation allows for incrementally predicting possible paths and detecting anomalous events from online trajectories. This representation also supports matching of complex motions with acceleration changes and pauses or stops within a trajectory. We use the proposed approach for classifying and predicting motion trajectories in traffic monitoring domains and test on several data sets. We show that our approach works well on various types of complete and incomplete trajectories from a variety of video data sets with different frame rates.
Dongryeol Lee, Irfan A. Essa
ICCV3
2011 Guest Editors' Introduction to the Special Section on Award-Winning Papers from the IEEE Conference on Computer Vision and Pattern Recognition 2009 (CVPR 2009)
abstract
. Kaiming He, Jian Sun, and Xiaoou Tang, “Single Image Haze Removal Using Dark Channel Prior,” winner of the Best Paper Award (sponsored by Microsoft). . Anat Levin, Yair Weiss, Fredo Durand, and Bill Freeman, “Understanding and evaluating blind deconvolution algorithms,” winner of the Best Paper-Honorable Mention Award (sponsored by Honeywell). . Ce Liu, Jenny Yuen, and Antonio Torralba, “Nonparametric Scene Parsing: Label Transfer via Dense Scene Alignment,” winner of the Best Student Paper Award (sponsored by MERL). . Olivier Duchenne, Francis Bach, In So Kweon, and Jean Ponce, “A Tensor-Based Algorithm for HighOrder Graph Matching,” winner of the Best Student Paper-Honorable Mention Award (sponsored by Hewlett-Packard). CVPR 2009 received 1,464 complete submissions by the 20 November 2008 deadline. We worked with 46 Area Chairs, who are well-respected members of the computer vision community, and 749 reviewers to select 61 papers as Orals and 322 papers as Posters. Twenty of these accepted papers were recommended to the Awards Committee for consideration. The Awards Committee consisted of five senior members of the vision community, three of whom were Area Chairs. The committee members had no conflicts with the candidate papers. The four award papers were selected after three phases of reviewing. These award papers were presented in the only singletrack session of the main conference. The authors of the award-winning papers and honorable mentions were invited to submit an extended version of their paper as a journal submission to TPAMI. The papers were reviewed by expert reviewers in the field, following the usual TPAMI procedure. All of the papers were accepted, in most cases after relatively minor revisions. CVPR 2009 also presented the Longuet-Higgins Prize for fundamental contributions in computer vision that have withstood the test of time. The awards committee that selected the CVPR 2009 awards was also asked to assist in the selection of this award. The prize (sponsored by IBM) was awarded to the following two papers that were published in CVPR 1999:
Irfan A. Essa, Sing Bing Kang, Marc Pollefeys
IEEE Trans. Pattern Anal. Mach. Intell.1
2011 Bilayer Segmentation of Webcam Videos Using Tree-Based Classifiers
abstract
This paper presents an automatic segmentation algorithm for video frames captured by a (monocular) webcam that closely approximates depth segmentation from a stereo camera. The frames are segmented into foreground and background layers that comprise a subject (participant) and other objects and individuals. The algorithm produces correct segmentations even in the presence of large background motion with a nearly stationary foreground. This research makes three key contributions: First, we introduce a novel motion representation, referred to as "motons," inspired by research in object recognition. Second, we propose estimating the segmentation likelihood from the spatial context of motion. The estimation is efficiently learned by random forests. Third, we introduce a general taxonomy of tree-based classifiers that facilitates both theoretical and experimental comparisons of several known classification algorithms and generates new ones. In our bilayer segmentation algorithm, diverse visual cues such as motion, motion context, color, contrast, and spatial priors are fused by means of a conditional random field (CRF) model. Segmentation is then achieved by binary min-cut. Experiments on many sequences of our videochat application demonstrate that our algorithm, which requires no initialization, is effective in a variety of scenes, and the segmentation results are comparable to those obtained by stereo systems.
Pei Yin, Antonio Criminisi, John M. Winn, Irfan A. Essa
IEEE Trans. Pattern Anal. Mach. Intell.4
2010 Discontinuous seam-carving for video retargeting
abstract
We introduce a new algorithm for video retargeting that uses discontinuous seam-carving in both space and time for resizing videos. Our algorithm relies on a novel appearance-based temporal coherence formulation that allows for frame-by-frame processing and results in temporally discontinuous seams, as opposed to geometrically smooth and continuous seams. This formulation optimizes the difference in appearance of the resultant retargeted frame to the optimal temporally coherent one, and allows for carving around fast moving salient regions. Additionally, we generalize the idea of appearance-based coherence to the spatial domain by introducing piece-wise spatial seams. Our spatial coherence measure minimizes the change in gradients during retargeting, which preserves spatial detail better than minimization of color difference alone. We also show that per-frame saliency (gradient-based or feature-based) does not always produce desirable retargeting results and propose a novel automatically computed measure of spatio-temporal saliency. As needed, a user may also augment the saliency by interactive region-brushing. Our retargeting algorithm processes the video sequentially, making it conducive for streaming applications.
Matthias Grundmann 0002, Vivek Kwatra, Irfan A. Essa
CVPR4
2010 Efficient hierarchical graph-based video segmentation
abstract
We present an efficient and scalable technique for spatiotemporal segmentation of long video sequences using a hierarchical graph-based algorithm. We begin by over-segmenting a volumetric video graph into space-time regions grouped by appearance. We then construct a “region graph” over the obtained segmentation and iteratively repeat this process over multiple levels to create a tree of spatio-temporal segmentations. This hierarchical approach generates high quality segmentations, which are temporally coherent with stable region boundaries, and allows subsequent applications to choose from varying levels of granularity. We further improve segmentation quality by using dense optical flow to guide temporal connections in the initial graph. We also propose two novel approaches to improve the scalability of our technique: (a) a parallel out-of-core algorithm that can process volumes much larger than an in-core algorithm, and (b) a clip-based processing algorithm that divides the video into overlapping clips in time, and segments them successively while enforcing consistency. We demonstrate hierarchical segmentations on video shots as long as 40 seconds, and even support a streaming mode for arbitrarily long videos, albeit without the ability to process them hierarchically.
Matthias Grundmann 0002, Vivek Kwatra, Irfan A. Essa
CVPR4
2010 Player localization using multiple static cameras for sports visualization
abstract
We present a novel approach for robust localization of multiple people observed using multiple cameras. We use this location information to generate sports visualizations, which include displaying a virtual offside line in soccer games, and showing players' positions and motion patterns. Our main contribution is the modeling and analysis for the problem of fusing corresponding players' positional information as finding minimum weight K-length cycles in complete K-partite graphs. To this end, we use a dynamic programming based approach that varies over a continuum of being maximally to minimally greedy in terms of the number of paths explored at each iteration. We present an end-to-end sports visualization framework that employs our proposed algorithm-class. We demonstrate the robustness of our framework by testing it on 60,000 frames of soccer footage captured over 5 different illumination conditions, play types, and team attire.
Raffay Hamid, Ramkrishan K. Kumar, Matthias Grundmann 0002, Irfan A. Essa, Jessica K. Hodgins
CVPR5
2010 Motion fields to predict play evolution in dynamic sport scenes
abstract
Videos of multi-player team sports provide a challenging domain for dynamic scene analysis. Player actions and interactions are complex as they are driven by many factors, such as the short-term goals of the individual player, the overall team strategy, the rules of the sport, and the current context of the game. We show that constrained multi-agent events can be analyzed and even predicted from video. Such analysis requires estimating the global movements of all players in the scene at any time, and is needed for modeling and predicting how the multi-agent play evolves over time on the field. To this end, we propose a novel approach to detect the locations of where the play evolution will proceed, e.g. where interesting events will occur, by tracking player positions and movements over time. We start by extracting the ground level sparse movement of players in each time-step, and then generate a dense motion field. Using this field we detect locations where the motion converges, implying positions towards which the play is evolving. We evaluate our approach by analyzing videos of a variety of complex soccer plays.
Matthias Grundmann 0002, Ariel Shamir, Iain A. Matthews, Jessica K. Hodgins, Irfan A. Essa
CVPR6
2010 Fluid Simulation with Articulated Bodies
abstract
We present an algorithm for creating realistic animations of characters that are swimming through fluids. Our approach combines dynamic simulation with data-driven kinematic motions (motion capture data) to produce realistic animation in a fluid. The interaction of the articulated body with the fluid is performed by incorporating joint constraints with rigid animation and by extending a solid/fluid coupling method to handle articulated chains. Our solver takes as input the current state of the simulation and calculates the angular and linear accelerations of the connected bodies needed to match a particular motion sequence for the articulated body. These accelerations are used to estimate the forces and torques that are then applied to each joint. Based on this approach, we demonstrate simulated swimming results for a variety of different strokes, including crawl, backstroke, breaststroke, and butterfly. The ability to have articulated bodies interact with fluids also allows us to generate simulations of simple water creatures that are driven by simple controllers.
Nipun Kwatra, Christopher Wojtan, Mark T. Carlson, Irfan A. Essa, Peter J. Mucha, Greg Turk
IEEE Trans. Vis. Comput. Graph.4
2009 Videolyzer: quality analysis of online informational video for bloggers and journalists
abstract
Tools to aid people in making sense of the information quality of online informational video are essential for media consumers seeking to be well informed. Our application, Videolyzer, addresses the information quality problem in video by allowing politically motivated bloggers or journalists to analyze, collect, and share criticisms of the information quality of online political videos. Our interface innovates by providing a fine-grained and tightly coupled interaction paradigm between the timeline, the time-synced transcript, and annotations. We also incorporate automatic textual and video content analysis to suggest areas of interest for further assessment by a person. We present an evaluation of Videolyzer looking at the user experience, usefulness, and behavior around the novel features of the UI as well as report on the collaborative dynamic of the discourse generated with the tool.
Nicholas Diakopoulos, Sergio Goldenberg, Irfan A. Essa
CHI3
2009 Learning the basic units in American Sign Language using discriminative segmental feature selection
abstract
The natural language for most deaf signers in the United States is American Sign Language (ASL). ASL has internal structure like spoken languages, and ASL linguists have introduced several phonemic models. The study of ASL phonemes is not only interesting to linguists, but also useful for scalability in recognition by machines. Since machine perception is different than human perception, this paper learns the basic units for ASL directly from data. Comparing with previous studies, our approach computes a set of data-driven units (fenemes) discriminatively from the results of segmental feature selection. The learning iterates the following two steps: first apply discriminative feature selection segmentally to the signs, and then tie the most similar temporal segments to re-train. Intuitively, the sign parts indistinguishable to machines are merged to form basic units, which we call ASL fenemes. Experiments on publicly available ASL recognition data show that the extracted data-driven fenemes are meaningful, and recognition using those fenemes achieves improved accuracy at reduced model complexity.
Pei Yin, Thad Starner, Harley Hamilton, Irfan A. Essa, James M. Rehg
ICASSP4
2009 Augmenting Aerial Earth Maps with dynamic information
abstract
We introduce methods for augmenting aerial visualizations of Earth (from services like Google Earth or Microsoft Virtual Earth) with dynamic information obtained from videos. Our goal is to make Augmented Aerial Earth Maps that visualize an alive and dynamic scene within a city. We propose different approaches for analyzing videos of cities with pedestrians and cars, under differing conditions and then created augmented Aerial Earth Maps (AEMs) with live and dynamic information. We further extend our visualizations to include analysis of natural phenomenon (specifically clouds) and add this information to the AEMs adding to the visual reality.
Sangmin Oh, Jeonggyu Lee, Irfan A. Essa
ISMAR4
2009 Human video textures
abstract
This paper describes a data-driven approach for generating photorealistic animations of human motion. Each animation sequence follows a user-choreographed path and plays continuously by seamlessly transitioning between different segments of the captured data. To produce these animations, we capitalize on the complementary characteristics of motion capture data and video. We customize our capture system to record motion capture data that are synchronized with our video source. Candidate transition points in video clips are identified using a new similarity metric based on 3-D marker trajectories and their 2-D projections into video. Once the transitions have been identified, a video-based motion graph is constructed. We further exploit hybrid motion and video data to ensure that the transitions are seamless when generating animations. Motion capture marker projections serve as control points for segmentation of layers and nonrigid transformation of regions. This allows warping and blending to generate seamless in-between frames for animation. We show a series of choreographed animations of walks and martial arts scenes as validation of our approach.
Matthew Flagg, Atsushi Nakazawa, Qiushuang Zhang, Sing Bing Kang, Young Kee Ryu, Irfan A. Essa, James M. Rehg
SI3D6
2009 A novel sequence representation for unsupervised analysis of human activities
Raffay Hamid, Siddhartha Maddi, Amos Y. Johnson, Aaron F. Bobick, Irfan A. Essa, Charles L. Isbell Jr.
Artif. Intell.5
2008 Computational photography and video: interacting and creating with videos and images
abstract
Digital image capture, processing, and sharing has become pervasive in our society. This has had significant impact on how we create novel scenes, how we share our experiences, and how we interact with images and videos. In this talk, I will present an overview of series of ongoing efforts in the analysis of images and videos for rendering novel scenes. First I will discuss (in brief) our work on Video Textures, where repeating information is extracted to generate extended sequences of videos. I will then describe some our extensions to this approach that allows for controlled generation of animations of video sprites. We have developed various learning and optimization techniques that allow for video-based animations of photorealistic characters. Using these sets of approaches as a foundation, then I will show how new images and videos can be generated. I will show examples of Photorealistic and Non-photorealistic Renderings of Scenes (Videos and Images) and how these methods support the media reuse culture, so common these days with user generated content. Time permitting, I will also share some of our efforts on video annotation and how we have taken some of these new concepts of video analysis to undergraduate classrooms.
Irfan A. Essa
AVI1
2008 Discriminative feature selection for hidden Markov models using Segmental Boosting
abstract
We address the feature selection problem for hidden Markov models (HMMs) in sequence classification. Temporal correlation in sequences often causes difficulty in applying feature selection tech niques. Inspired by segmental k-means segmentation (SKS) [B. Juang and L. Rabiner, 1990], we propose Segmentally Boosted HMMs (SBHMMs), where the state-optimized features are constructed in a segmental and discriminative manner. The contributions are twofold. First, we introduce a novel feature selection algorithm, where the temporal dynamics are decoupled from the static learning procedure by assuming that the sequential data are piecewise independent and identically distributed. Second, we show that the SBHMM consistently improves traditional HMM recognition in various domains. The reduction of error compared to traditional HMMs ranges from 17% to 70% in American Sign Language recognition, human gait identification, lip reading, and speech recognition.
Pei Yin, Irfan A. Essa, Thad Starner, James M. Rehg
ICASSP2
2008 3D Shape Context and Distance Transform for action recognition
abstract
We propose the use of 3D (2D+time) Shape Context to recognize the spatial and temporal details inherent in human actions. We represent an action in a video sequence by a 3D point cloud extracted by sampling 2D silhouettes over time. A non-uniform sampling method is introduced that gives preference to fast moving body parts using a Euclidean 3D Distance Transform. Actions are then classified by matching the extracted point clouds. Our proposed approach is based on a global matching and does not require specific training to learn the model. We test the approach thoroughly on two publicly available datasets and compare to several state-of-the-art methods. The achieved classification accuracy is on par with or superior to the best results reported to date.
Matthias Grundmann 0002, Franziska Meier, Irfan A. Essa
ICPR3
2008 Audio Puzzler: piecing together time-stamped speech transcripts with a puzzle game
abstract
We have developed an audio-based casual puzzle game which produces a time-stamped transcription of spoken audio as a by-product of play. Our evaluation of the game indicates that it is both fun and challenging. The transcripts generated using the game are more accurate than those produced using a standard automatic transcription system and the time-stamps of words are within several hundred milliseconds of ground truth.
Nicholas Diakopoulos, Kurt Luther, Irfan A. Essa
ACM Multimedia3
2007 Discovering Multivariate Motifs using Subsequence Density Estimation and Greedy Mixture Learning
David Minnen, Charles L. Isbell Jr., Irfan A. Essa, Thad Starner
AAAI3
2007 Tree-based Classifiers for Bilayer Video Segmentation
abstract
This paper presents an algorithm for the automatic segmentation of monocular videos into foreground and background layers. Correct segmentations are produced even in the presence of large background motion with nearly stationary foreground. There are three key contributions. The first is the introduction of a novel motion representation, "motons", inspired by research in object recognition. Second, we propose learning the segmentation likelihood from the spatial context of motion. The learning is efficiently performed by Random Forests. The third contribution is a general taxonomy of tree-based classifiers, which facilitates theoretical and experimental comparisons of several known classification algorithms, as well as spawning new ones. Diverse visual cues such as motion, motion context, colour, contrast and spatial priors are fused together by means of a conditional random field (CRF) model. Segmentation is then achieved by binary min-cut. Our algorithm requires no initialization. Experiments on many video-chat type sequences demonstrate the effectiveness of our algorithm in a variety of scenes. The segmentation results are comparable to those obtained by stereo systems.
Pei Yin, Antonio Criminisi, John M. Winn, Irfan A. Essa
CVPR4
2007 Incorporating Phase Information for Source Separation via Spectrogram Factorization
abstract
Spectrogram factorization methods have been proposed for single channel source separation and audio analysis. Typically, the mixture signal is first converted into a time-frequency representation such as the short-time Fourier transform (STFT). The phase information is thrown away and this spectrogram matrix is then factored into the sum of rank-one source spectrograms. This approach incorrectly assumes the mixture spectrogram is the sum of the source spectrograms. In fact, the mixture spectrogram depends on the phase of the source STFTs. We investigate the consequences of this common assumption and introduce an approach that leverages a probabilistic representation of phase to improve the separation results.
R. Mitchell Parry, Irfan A. Essa
ICASSP (2)2
2007 Structure from Statistics - Unsupervised Activity Analysis using Suffix Trees
abstract
Models of activity structure for unconstrained environments are generally not available a priori. Recent representational approaches to this end are limited by their computational complexity, and ability to capture activity structure only up to some fixed temporal scale. In this work, we propose Suffix Trees as an activity representation to efficiently extract structure of activities by analyzing their constituent event-subsequences over multiple temporal scales. We empirically compare Suffix Trees with some of the previous approaches in terms of feature cardinality, discriminative prowess, noise sensitivity and activity-class discovery. Finally, exploiting properties of Suffix Trees, we present a novel perspective on anomalous subsequences of activities, and propose an algorithm to detect them in linear-time. We present comparative results over experimental data, collected from a kitchen environment to demonstrate the competence of our proposed framework.
Raffay Hamid, Siddhartha Maddi, Aaron F. Bobick, Irfan A. Essa
ICCV4
2007 Detecting Subdimensional Motifs: An Efficient Algorithm for Generalized Multivariate Pattern Discovery
abstract
Discovering recurring patterns in time series data is a fundamental problem for temporal data mining. This paper addresses the problem of locating subdimensional motifs in real-valued, multivariate time series, which requires the simultaneous discovery of sets of recurring patterns along with the corresponding relevant dimensions. While many approaches to motif discovery have been developed, most are restricted to categorical data, univariate time series, or multivariate data in which the temporal patterns span all of the dimensions. In this paper, we present an expected linear-time algorithm that addresses a generalization of multivariate pattern discovery in which each motif may span only a subset of the dimensions. To validate our algorithm, we discuss its theoretical properties and empirically evaluate it using several data sets including synthetic data and motion capture data collected by an on-body iner- tial sensor.
David Minnen, Charles L. Isbell Jr., Irfan A. Essa, Thad Starner
ICDM3
2007 Improving Activity Discovery with Automatic Neighborhood Estimation
David Minnen, Thad Starner, Irfan A. Essa, Charles L. Isbell Jr.
IJCAI3
2007 A Boosted Segmentation Method for Surgical Workflow Analysis
Nicolas Padoy, Tobias Blum, Irfan A. Essa, Hubertus Feußner, Marie-Odile Berger, Nassir Navab
MICCAI (1)3
2006 Element-Free Elastic Models for Volume Fitting and Capture
abstract
We present a new method of fitting an element-free volumetric model to a sequence of deforming surfaces of a moving object. Given a sequence of visual hulls, we iteratively fit an element-free elastic model to the visual hull in order to extract the optimal pose of the captured volume. The fitting of the volumetric model is acheived by minimizing a combination of elastic potential energy, a surface distance measure, and a self-intersection penalty for each frame. A unique aspect of our work is that the model is mesh free - since the model is represented as a point cloud, it is easy to construct, manipulate and update the model as needed. Additionally, linear elasicity with rotation compensation makes it possible to handle local deformations and large rotations of body parts much more efficiently than other volume fitting approaches. Our experimental results for volume fitting and capture in a multi-view camera setting demonstrate the robustness of element-free elastic models against noise and self-occlusions.
Jaeil Choi, Andrzej Szymczak, Greg Turk, Irfan A. Essa
CVPR (2)4
2006 Learning Temporal Sequence Model from Partially Labeled Data
abstract
Graphical models are often used to represent and recognize activities. Purely unsupervised methods (such as HMMs) can be trained automatically but yield models whose internal structure - the nodes - are difficult to interpret semantically. Manually constructed networks typically have nodes corresponding to sub-events, but the programming and training of these networks is tedious and requires extensive domain expertise. In this paper, we propose a semi-supervised approach in which a manually structured, Propagation Network (a form of a DBN) is initialized from a small amount of fully annotated data, and then refined by an EM-based learning method in an unsupervised fashion. During node refinement (the M step) a boosting-based algorithm is employed to train the evidence detectors of individual nodes. Experiments on a variety of data types - vision and inertial measurements - in several tasks demonstrate the ability to learn from as little as one fully annotated example accompanied by a small number of positive but non-annotated training examples. The system is applied to both recognition and anomaly detection tasks.
Aaron F. Bobick, Irfan A. Essa
CVPR (2)3
2006 Source Detection Using Repetitive Structure
abstract
Blind source separation algorithms typically require that the number of sources are known in advance. However, it is often the case that the number of sources change over time and that the total number is not known. Existing source separation techniques require source number estimation methods to determine how many sources are active within the mixture signals. These methods typically operate on the covariance matrix of mixture recordings and require fewer active sources than mixtures. When sources do not overlap in the time-frequency domain, more sources than mixtures may be detected and then separated. However, separating more sources than mixtures when sources overlap in time and frequency poses a particularly difficult problem. This paper addresses the issue of source detection when more sources than sensors overlap in time and frequency. We show that repetitive structure in the form of time-time correlation matrices can reveal when each source is active
R. Mitchell Parry, Irfan A. Essa
ICASSP (4)2
2006 Interactive mosaic generation for video navigation
abstract
Navigation through large multimedia collections that include videos and images still remains cumbersome. In this paper, we introduce a novel method to visualize and navigate through the collection by creating a mosaic image that visually represents the compilation. This image is generated by a labeling-based layout algorithm using various sizes of sample tile images from the collection. Each tile represents both the photographs and video files representing scenes selected by matching algorithms. This generated mosaic image provides a new way for thematic video and visually summarizes the videos. Users can generate these mosaics with some predefined themes and layouts, or base it on the results of their queries. Our approach supports automatic generation of these layouts by using meta-information such as color, time-line and existence of faces or manually generated annotated information from existing systems (e.g., the Family Video Archive).
Irfan A. Essa, Gregory D. Abowd
ACM Multimedia2
2006 Videotater: an approach for pen-based digital video segmentation and tagging
abstract
The continuous growth of media databases necessitates development of novel visualization and interaction techniques to support management of these collections. We present Videotater, an experimental tool for a Tablet PC that supports the efficient and intuitive navigation, selection, segmentation, and tagging of video. Our veridical representation immediately signals to the user where appropriate segment boundaries should be placed and allows for rapid review and refinement of manually or automatically generated segments. Finally, we explore a distribution of modalities in the interface by using multiple timeline representations, pressure sensing, and a tag painting/erasing metaphor with the pen.
Nicholas Diakopoulos, Irfan A. Essa
UIST2
2005 Video-based nonphotorealistic and expressive illustration of motion
abstract
We present a semi-automatic approach for adding expressive renderings to images and videos that highlight motions and movement. Our technique relies on motion analysis of video where the motion information from the image sequence is used to add expressive information. The first step in our approach is to extract a moving region of the video by segmenting and then grouping regions of compatible motions. In the second step, a user can interactively choose or refine a grouping region that represents the moving object of interest. In the third and final stage, the user can apply various visual effects such as a temporal-flare, time-lapse, and particle-effects. We have implemented a prototype system that can be used to illustrate and expressively render motions in videos and images, with simple user interaction. Our system can deal with most translational and rotational motions without a need for a fixed background.
Irfan A. Essa
Computer Graphics International2
2005 Tracking Multiple Objects through Occlusions
abstract
We present an approach for tracking varying number of objects through both temporally and spatially significant occlusions. Our method builds on the idea of object permanence to reason about occlusions. To this end, tracking is performed at both the region level and the object level. At the region level, a customized genetic algorithm is used to search for optimal region tracks. This limits the scope of object trajectories. At the object level, each object is located based on adaptive appearance models, spatial distributions and inter-occlusion relationships. The proposed architecture is capable of tracking objects even in the presence of long periods of full occlusions. We demonstrate the viability of this approach by experimenting on several videos of a user interacting with a variety of objects on a desktop.
Yan Huang 0017, Irfan A. Essa
CVPR (2)2
2005 Tracking Multiple Objects through Occlusions
abstract
Visual tracking of multiple objects with continuous interactions is important for many applications such as surveillance and effective user interfaces. Complex interactions between objects results in both temporally and spatially significant occlusions, making multi-object tracking a challenging problem. We present an approach for tracking varying number of objects through significant occlusions. Our approach builds on the idea of object permanence. That is, partially or fully occluded objects, even though not observable, still exist in the close proximity of their occluders. To demonstrate the viability of this approach, we have concentrated on tracking hands and objects as a person interacts with multiple objects on a desktop. We experimented with 5 indoor video sequences and present the results of two representative ones: lego sequence and shell game.
Yan Huang 0017, Irfan A. Essa
CVPR (2)2
2005 Unsupervised Activity Discovery and Characterization From Event-Streams
Rafay Hammid, Siddhartha Maddi, Amos Y. Johnson, Aaron F. Bobick, Irfan A. Essa, Charles L. Isbell Jr.
UAI5
2005 Mediating photo collage authoring
abstract
The medium of collage supports the visualization of meaningful event summaries using photographs. It can however be rather tedious to author a collage from a large collection of photographs. In this work we present an approach that supports efficient construction of a collage by assisting the user with an automatic layout procedure that can be controlled at a high level. Our layout method utilizes a pre-designed template which consists of cells for photos and annotations applied to these cells. The layout is then filled by matching the metadata of photos to the annotations in the cells using an optimization algorithm. The user exercises flexibility in the authoring process by (a) maintaining high-level control through the types of constraints applied and (b) leveraging visual emphases supported by the layout algorithm. The user can of course provide fine-grained control of the final collage through direct manipulation. Off-loading the tedium of collage construction to a user controlled yet automated process clears the way for rapidly generating different views of the same album and could also support the increased sharing of digital photos in the form of compact collages.
Nicholas Diakopoulos, Irfan A. Essa
UIST2
2005 Experiences with optimizing two stream-based applications for cluster execution
Yavor Angelov, Umakishore Ramachandran, Kenneth M. Mackenzie, James M. Rehg, Irfan A. Essa
J. Parallel Distributed Comput.5
2005 Texture optimization for example-based synthesis
abstract
We present a novel technique for texture synthesis using optimization. We define a Markov Random Field (MRF)-based similarity metric for measuring the quality of synthesized texture with respect to a given input sample. This allows us to formulate the synthesis problem as minimization of an energy function, which is optimized using an Expectation Maximization (EM)-like algorithm. In contrast to most example-based techniques that do region-growing, ours is a joint optimization approach that progressively refines the entire texture. Additionally, our approach is ideally suited to allow for controllable synthesis of textures. Specifically, we demonstrate controllability by animating image textures using flow fields. We allow for general two-dimensional flow fields that may dynamically change over time. Applications of this technique include dynamic texturing of fluid animations and texture-based flow visualization.
Vivek Kwatra, Irfan A. Essa, Aaron F. Bobick, Nipun Kwatra
ACM Trans. Graph.2
2004 Propagation Networks for Recognition of Partially Ordered Sequential Action
Yan Huang 0017, David Minnen, Aaron F. Bobick, Irfan A. Essa
CVPR (2)5
2004 Asymmetrically Boosted HMM for Speech Reading
Pei Yin, Irfan A. Essa, James M. Rehg
CVPR (2)2
2004 Novel Skeletal Representation for Articulated Creatures
Gabriel J. Brostow, Irfan A. Essa, Drew Steedly, Vivek Kwatra
ECCV (3)2
2004 Parameterized Authentication
Michael J. Covington, Mustaque Ahamad, Irfan A. Essa, H. Venkateswaran
ESORICS3
2003 Expectation Grammars: Leveraging High-Level Expectations for Activity Recognition
abstract
Video-based recognition and prediction of a temporally extended activity can benefit from a detailed description of high-level expectations about the activity. Stochastic grammars allow for an efficient representation of such expectations and are well-suited for the specification of temporally well-ordered activities. In this paper, we extend stochastic grammars by adding event parameters, state checks, and sensitivity to an internal scene model. We present an implemented system that uses human-specified grammars to recognize a person performing the Towers of Hanoi task from a video sequence by analyzing object interaction events. Experimental results from several videos show robust recognition of the full task and its constituent sub-tasks even though no appearance models of the objects in the video are provided. These experiments include videos of the task performed with different shaped objects and with distracting and extraneous interactions.
David Minnen, Irfan A. Essa, Thad Starner
CVPR (2)2
2003 Mandatory human participation: a new authentication scheme for building secure systems
abstract
Mandatory human participation (MHP) is a novel authentication scheme that asks the question "are you human?" (Instead of "who are you?"), and upon the correct answer to this question, can prove a principal to be a human being instead of a computer program. MHP helps solve old and new problems in computer security that existing security measures cannot address properly, including password (or PIN number) guessing attacks and application-level denial of service. A key component of this "are you human?" authentication process is a character morphing algorithm that transforms a character string into its graphical form in such a way that a human being won't have any problem recognizing the original string, while a computer program (e.g., an optical character recognition program), will not be able to decipher it or make a correct guess with nonnegligible probability. The basic idea of the MHP scheme is to ask an agent to recognize the string before its login attempts or transaction requests can be honored. Here a protocol is needed to send a puzzle to an agent, check if the answer supplied by the agent is correct, and most importantly make sure that the agent cannot cheat in the process. A number of system and security issues that relate to the protocol need to be addressed for the protocol to be secure, efficient, robust, and user-friendly. The MHP scheme contributes to the foundation of the computer security by faithfully implementing novel security semantics, "human," which existing cryptographic measures cannot express accurately. As many real-world security applications involve the interaction between a human and a computer, which naturally contains "human" as a part of its protocol semantics, we believe that the MHP scheme will find many new applications in the future.
Jun (Jim) Xu, Richard J. Lipton, Irfan A. Essa, Minho Sung
ICCCN3
2003 Spectral Partitioning for Structure from Motion
abstract
We propose a spectral partitioning approach for large-scale optimization problems, specifically structure from motion. In structure from motion, partitioning methods reduce the problem into smaller and better conditioned subproblems which can be efficiently optimized. Our partitioning method uses only the Hessian of the reprojection error and its eigenvector. We show that partitioned systems that preserve the eigenvectors corresponding to small eigenvalues result in lower residual error when optimized. We create partitions by clustering the entries of the eigenvectors of the Hessian corresponding to small eigenvalues. This is a more general technique than relying on domain knowledge and heuristics such as bottom-up structure from motion approaches. Simultaneously, it takes advantage of more information than generic matrix partitioning algorithms.
Drew Steedly, Irfan A. Essa, Frank Dellaert
ICCV2
2003 Perceptual user interfaces using vision-based eye tracking
abstract
We present a multi-camera vision-based eye tracking method to robustly locate and track user's eyes as they interact with an application. We propose enhancements to various vision-based eye-tracking approaches, which include (a) the use of multiple cameras to estimate head pose and increase coverage of the sensors and (b) the use of probabilistic measures incorporating Fisher's linear discriminant to robustly track the eyes under varying lighting conditions in real-time. We present experiments and quantitative results to demonstrate the robustness of our eye tracking in two application prototypes.
Ravikrishna Ruddarraju, Antonio Haro, Kristine S. Nagel, Quan T. Tran, Irfan A. Essa, Gregory D. Abowd, Elizabeth D. Mynatt
ICMI5
2003 Presenting Movement in a Computer-Based Dance Tutor
abstract
This article addresses how to present movement information to learners as part of a larger project on developing a nonconventional computational system that teaches ballet. The requirements of such a system are first described, and then discoveries regarding the first requirement, presenting movement to a user, are discussed. Background research regarding how people learn movement, hypotheses concerning presenting movement with computer animation versus videotape, and an experiment testing those hypotheses are presented. The experiment required individuals to perform movements after viewing them in one of the formats. Each participant viewed a movement sequence multiple times and then was evaluated on his or her performance of that movement by two expert judges. Animations resulted in higher performance ratings for individuals with some previous dance experience. Format did not affect performance for other learners. This result implies that domain knowledge interacts with presentation format in learning ballet. These results will influence the design and implementation of a computer-based dance tutor under development, and they point to several interesting research directions, including exploring the effects of multimodal sensory presentations and prior knowledge in learning movement.
Katherine E. Sukel, Richard Catrambone, Irfan A. Essa, Gabriel J. Brostow
Int. J. Hum. Comput. Interact.3
2003 Graphcut textures: image and video synthesis using graph cuts
abstract
In this paper we introduce a new algorithm for image and video texture synthesis. In our approach, patch regions from a sample image or video are transformed and copied to the output and then stitched together along optimal seams to generate a new (and typically larger) output. In contrast to other techniques, the size of the patch is not chosen a-priori , but instead a graph cut technique is used to determine the optimal patch region for any given offset between the input and output texture. Unlike dynamic programming, our graph cut technique for seam optimization is applicable in any dimension. We specifically explore it in 2D and 3D to perform video texture synthesis in addition to regular image synthesis. We present approximative offset search techniques that work well in conjunction with the presented patch size optimization. We show results for synthesizing regular, random, and natural images and videos. We also demonstrate how this method can be used to interactively merge different images to generate new scenes.
Vivek Kwatra, Arno Schödl, Irfan A. Essa, Greg Turk, Aaron F. Bobick
ACM Trans. Graph.3
2001 Depth Layers from Occlusions
abstract
We present a method to extract relative depth information from an uncalibrated monocular video sequence. Our method detects occlusions caused by an object moving in a static scene to infer relative depth relationships between scene parts. Our approach does not rely on any strong assumptions about the object or the scene to aid in this segmentation into layers. In general, the problem of building relative depth relationships from occlusion events is underconstrained, even in the absence of observation noise. A minimum description length algorithm is used to reliably calculate layer opacities and their depth relationships in the absence of hard constraints. Our approach extends previously published approaches that are restricted to work with a certain type of moving object or require strong image edges to allow for an a-priori segmentation of the scene. We also discuss ideas on how to extend our algorithm to make use of a richer set of observations.
Arno Schödl, Irfan A. Essa
CVPR (1)2
2001 Propagation of Innovative Information in Non-Linear Least-Squares Structure from Motion
abstract
We present a new technique that improves upon existing structure from motion (SFM) methods. We propose a SFM algorithm that is both recursive and optimal. Our method incorporates innovative information from new frames into an existing solution without optimizing every camera pose and scene structure parameter. To do this, we incrementally optimize larger subsets of parameters until the error is minimized. These additional parameters are included in the optimization by tracing connections between points and frames. In many cases, the complexity of adding a frame is much smaller than full bundle adjustment of all the parameters. Our algorithm is best described us incremental bundle adjustment as it allows new information to be added to art existing non-linear least-squares solution.
Drew Steedly, Irfan A. Essa
ICCV2
2001 Image-based motion blur for stop motion animation
abstract
Stop motion animation is a well-established technique where still pictures of static scenes are taken and then played at film speeds to show motion. A major limitation of this method appears when fast motions are desired; most motion appears to have sharp edges and there is no visible motion blur. Appearance of motion blur is a strong perceptual cue, which is automatically present in live-action films, and synthetically generated in animated sequences. In this paper, we present an approach for automatically simulating motion blur. Ours is wholly a post-process, and uses image sequences, both stop motion or raw video, as input. First we track the frame-to-frame motion of the objects within the image plane. We then integrate the scene's appearance as it changed over a period of time. This period of time corresponds to shutter speed in live-action filming, and gives us interactive control over the extent of the induced blur. We demonstrate a simple implementation of our approach as it applies to footage of different motions and to scenes of varying complexity. Our photorealistic renderings of these input sequences approximate the effect of capturing moving objects on film that is exposed for finite periods of time.
Gabriel J. Brostow, Irfan A. Essa
SIGGRAPH2
2000 Detecting and Tracking Eyes by Using Their Physiological Properties, Dynamics, and Appearance
abstract
Reliable detection and tracking of eyes is an important requirement for attentive user interfaces. In this paper, we present a methodology for detecting eyes robustly in indoor environments in real-time. We exploit the physiological properties and appearance of eyes as well as head/eye motion dynamics. Infrared lighting is used to capture the physiological properties of eyes, Kalman trackers are used to model eye/head dynamics, and a probabilistic based appearance model is used to represent eye appearance. By combining three separate modalities, with specific enhancements within each modality, our approach allows eyes to be treated as robust features that can be used for other higher-level processing.
Antonio Haro, Myron Flickner, Irfan A. Essa
CVPR3
2000 Machine Learning for Video-Based Rendering
abstract
We present techniques for rendering and animation of realistic scenes by analyzing and training on short video sequences. This work extends the new paradigm for computer animation, video tex(cid:173) tures, which uses recorded video to generate novel animations by replaying the video samples in a new order. Here we concentrate on video sprites, which are a special type of video texture. In video sprites, instead of storing whole images, the object of inter(cid:173) est is separated from the background and the video samples are stored as a sequence of alpha-matted sprites with associated veloc(cid:173) ity information. They can be rendered anywhere on the screen to create a novel animation of the object. We present methods to cre(cid:173) ate such animations by finding a sequence of sprite samples that is both visually smooth and follows a desired path. To estimate visual smoothness, we train a linear classifier to estimate visual similarity between video samples. If the motion path is known in advance, we use beam search to find a good sample sequence. We can specify the motion interactively by precomputing the sequence cost function using Q-Iearning.
Arno Schödl, Irfan A. Essa
NIPS2
2000 Video textures
abstract
This paper introduces a new type of medium, called a video texture, which has qualities somewhere between those of a photograph and a video. A video texture provides a continuous infinitely varying stream of images. While the individual frames of a video texture may be repeated from time to time, the video sequence as a whole is never repeated exactly. Video textures can be used in place of digital photos to infuse a static image with dynamic qualities and explicit actions. We present techniques for analyzing a video clip to extract its structure, and for synthesizing a new, similar looking video of arbitrary length. We combine video textures with view morphing techniques to obtain 3D video textures. We also introduce video-based animation, in which the synthesis of video textures can be guided by a user through high-level interactive controls. Applications of video textures and their extensions include the display of dynamic scenes on web pages, the creation of dynamic backdrops for special effects and games, and the interactive control of video-based animation.
Arno Schödl, Richard Szeliski, David Salesin, Irfan A. Essa
SIGGRAPH4
1999 Motion based Decompositing of Video
abstract
We present a method to decompose video sequences into layers that represent the relative depths of complex scenes. Our method combines spatial information with temporal occlusions to determine relative depths of these layers. Spatial information is obtained through edge detection and a customized contour completion algorithm. Activity in a scene is used to extract temporal occlusion events, which are in turn, used to classify objects as occluders or occludes. The path traversed by the moving objects determines the segmentation of the scene. Several examples of decompositing and compositing of video are shown. This approach can be applied in the pre-processing of sequences for compositing or tracking purposes and to determine the approximate 3D structure of a scene.
Gabriel J. Brostow, Irfan A. Essa
ICCV2
1999 Exploiting Human Actions and Object Context for Recognition Tasks
abstract
Our goal is to exploit human motion and object context to perform action recognition and object classification. Towards this end, we introduce a framework for recognizing actions and objects by measuring image-, object- and action-based information from video. Hidden Markov models are combined with object context to classify hand actions, which are aggregated by a Bayesian classifier to summarize activities. We also use Bayesian methods to differentiate the class of unknown objects by evaluating detected actions along with low-level, extracted object features. Our approach is appropriate for locating and classifying objects under a variety of conditions including full occlusion. We show experiments where both familiar and previously unseen objects are recognized using action and context information.
Darnell J. Moore, Irfan A. Essa, Monson H. Hayes III
ICCV2
1997 Coding, Analysis, Interpretation, and Recognition of Facial Expressions
abstract
We describe a computer vision system for observing facial motion by using an optimal estimation optical flow method coupled with geometric, physical and motion-based dynamic models describing the facial structure. Our method produces a reliable parametric representation of the face's independent muscle action groups, as well as an accurate estimate of facial motion. Previous efforts at analysis of facial expression have been based on the facial action coding system (FACS), a representation developed in order to allow human psychologists to code expression from static pictures. To avoid use of this heuristic coding scheme, we have used our computer vision system to probabilistically characterize facial motion and muscle activation in an experimental population, thus deriving a new, more accurate, representation of human facial expressions that we call FACS+. Finally, we show how this method can be used for coding, analysis, interpretation, and recognition of facial expressions.
Irfan A. Essa, Alex Pentland
IEEE Trans. Pattern Anal. Mach. Intell.1
1996 Modeling, Tracking and Interactive Animation of Faces and Heads Using Input from Video
abstract
We describe tools that use measurements from video for the extraction of facial modeling and animation parameters, head tracking, and real time interactive facial animation. These tools share common goals but rely on varying details of physical and geometric modeling and in their input measurement system. Accurate facial modeling involves fine details of geometry and muscle coarticulation. By coupling pixel by pixel measurements of surface motion to a physically based face model and a muscle control model, we have been able to obtain detailed spatio temporal records of both the displacement of each point on the facial surface and the muscle control required to produce the observed facial motion. We discuss the importance of this visually extracted representation in terms of realistic facial motion synthesis. A similar method that uses an ellipsoidal model of the head coupled with detailed estimates of visual motion allows accurate tracking of head motion in 3D. Additionally, by coupling sparse, fast visual measurements with our physically based model via an interpolation process, we have produced a real time interactive facial animation/mimicking system.
Irfan A. Essa, Sumit Basu, Trevor Darrell, Alex Pentland
CA1
1996 Vision-Based HCI - What's Next and What are the Difficult Problems?
Irfan A. Essa
FG1
1996 Motion regularization for model-based head tracking
abstract
This paper describes a method for the robust tracking of rigid head motion from video. This method uses a 3D ellipsoidal model of the head and interprets the optical flow in terms of the possible rigid motions of the model. This method is robust to large angular and translational motions of the head and is not subject to the singularities of a 2D model. The method has been successfully applied to heads with a variety of shapes, hair styles, etc. This method also has the advantage of accurately capturing the 3D motion parameters of the head. This accuracy is shown through comparison with a ground truth synthetic sequence (a rendered 3D animation of a model head). In addition, the ellipsoidal model is robust to small variations in the initial fit, enabling the automation of the model initialization. Lastly, due to its consideration of the entire 3D aspect of the head, the tracking is very stable over a large number of frames. This robustness extends even to sequences with very low frame rates and noisy camera images.
Sumit Basu, Irfan A. Essa, Alex Pentland
ICPR2
1996 Task-Specific Gesture Analysis in Real-Time Using Interpolated Views
abstract
Hand and face gestures are modeled using an appearance-based approach in which patterns are represented as a vector of similarity scores to a set of view models defined in space and time. These view models are learned from examples using unsupervised clustering techniques. A supervised teaming paradigm is then used to interpolate view scores into a task-dependent coordinate system appropriate for recognition and control tasks. We apply this analysis to the problem of context-specific gesture interpolation and recognition, and demonstrate real-time systems which perform these tasks.
Trevor Darrell, Irfan A. Essa, Alex Pentland
IEEE Trans. Pattern Anal. Mach. Intell.2
1995 Facial Expression Recognition Using a Dynamic Model and Motion Energy
abstract
Previous efforts at facial expression recognition have been based on the Facial Action Coding System (FACS), a representation developed in order to allow human psychologists to code expression from static facial "mugshots." We develop new more accurate representations for facial expression by building a video database of facial expressions and then probabilistically characterizing the facial muscle activation associated with each expression using a detailed physical model of the skin and muscles. This produces a muscle based representation of facial motion, which is then used to recognize facial expressions in two different ways. The first method uses the physics based model directly, by recognizing expressions through comparison of estimated muscle activations. The second method uses the physics based model to generate spatio temporal motion energy templates of the whole face for each different expression. These simple, biologically plausible motion energy "templates" are then used for recognition. Both methods show substantially greater accuracy at expression recognition than has been previously achieved.>
Irfan A. Essa, Alex Pentland
ICCV1
1994 Visually guided animation
abstract
We are interested in being able to take classic film characters, or video of current-day personalities, and produce computer models and animations of them by automatic analysis of the video or film footage. In this paper we survey our progress toward producing such automatic modeling and animation systems.>
Alex Pentland, Trevor Darrell, Irfan A. Essa, Ali Azarbayejani, Stan Sclaroff
CA3
1994 A vision system for observing and extracting facial action parameters
abstract
We describe a computer vision system for observing the "action units" of a face using video sequences as input. The visual observation (sensing) is achieved by using an optimal estimation optical flow method coupled with a geometric and a physical (muscle) model describing the facial structure. This modeling results in a time-varying spatial patterning of facial shape and a parametric representation of the independent muscle action groups, responsible for the observed facial motions. These muscle action patterns may then be used for analysis, interpretation, and synthesis. Thus, by interpreting facial motions within a physics-based optimal estimation framework, a new control model of facial movement is developed. The newly extracted action units (which we name "FACS+") are both physics and geometry-based, and extend the well-known FACS parameters for facial expressions by adding temporal information and non-local spatial patterning of facial motion.>
Irfan A. Essa, Alex Pentland
CVPR1
1994 Correlation and Interpolation Networks for Real-time Expression Analysis/Synthesis
abstract
We describe a framework for real-time tracking of facial expressions that uses neurally-inspired correlation and interpolation methods. A distributed view-based representation is used to characterize facial state, and is computed using a replicated correlation network. The ensemble response of the set of view correlation scores is input to a network based interpolation method, which maps perceptual state to motor control states for a simulated 3-D face model. Activation levels of the motor state correspond to muscle activations in an anatomically derived model. By integrating fast and robust 2-D processing with 3-D models, we obtain a system that is able to quickly track and interpret complex facial motions in real-time.
Trevor Darrell, Irfan A. Essa, Alex Pentland
NIPS2
1992 A Unified Approach for Physical and Geometric Modeling for Graphics and Animation
abstract
Abstract We present a unified approach for geometric and physical modeling using implicit functions, for application to graphics and animation. This method extends previously proposed techniques, and allows the standard finite element method to be directly combined with geometric modeling, resulting in quick calculation of an object's mass and stiffness matrices, and its vibration modes and frequencies. Because the approach is based on an implicitfunction representation, it allows very fast collision detection and characterization. Examples of complex physical and geometric modeling are presented.
Irfan A. Essa, Stan Sclaroff, Alex Pentland
Comput. Graph. Forum1
1990 The ThingWorld modeling system: virtual sculpting by modal forces
abstract
We describe a real-time solid modeling system that is based on the physical analogy of forming clay by applying forces. The system is implemented by simulating real materials as they react to user-supplied forces. Unlike other physically-based modeling approaches, the Thingworld system allows the user to restrict forming action to simple global deformations during the initial roughing in phase of modeling, and then later concern themselves with detailing. The Thingworld system also allows users to automatically model existing objects by using measurements taken from the object's surface. These measurements are used to generate artificial forces that mold the computer model much as a human would mold a clay model. Timed examples for constructing solid models are shown.
Stanley E. Scharoff, Alex Pentland, Irfan A. Essa, Martin Friedmann, Bradley Horowitz
I3D3