Ajay Divakaran

dblp:73/467 · DBLP profile ↗
← Back
71ranked-venue papers
7as first author
12since 2021 · last 2025
0000-0003-0371-5346ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 58 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 18 · 10 since 2021Human-computer interaction and ubiquitous computing · 4Databases, data management, data science and information retrieval · 2 · 1 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Punching Bag vs. Punching Person: Motion Transferability in Videos
abstract
Action recognition models demonstrate strong generalization, but can they effectively transfer high-level motion concepts across diverse contexts, even within similar distributions? For example, can a model recognize the broad action "punching" when presented with an unseen variation such as "punching person"? To explore this, we introduce a motion transferability framework with three datasets: (1) Syn-TA, a synthetic dataset with 3D object motions; (2) Kinetics400-TA; and (3) Something-Something-v2-TA, both adapted from natural video datasets. We evaluate 13 state-of-the-art models on these benchmarks and observe a significant drop in performance when recognizing high-level actions in novel contexts. Our analysis reveals: 1) Multimodal models struggle more with fine-grained unknown actions than with coarse ones; 2) The bias-free Syn-TA proves as challenging as real-world datasets, with models showing greater performance drops in controlled settings; 3) Larger models improve transferability when spatial cues dominate but struggle with intensive temporal reasoning, while reliance on object and background cues hinders generalization. We further explore how disentangling coarse and fine motions can improve recognition in temporally challenging datasets. We believe this study establishes a crucial benchmark for assessing motion transferability in action recognition. Datasets and relevant code: https://github.com/raiyaan-abdullah/Motion-Transfer.
Raiyaan Abdullah, Jared Claypoole, Michael Cogswell, Ajay Divakaran, Yogesh S. Rawat
ICCV4
2025 A Video is Worth 10, 000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval
abstract
Existing long video retrieval systems are trained and tested in the paragraph-to- video retrieval regime, where ev-ery long video is described by a single long paragraph. This neglects the richness and variety of possible valid de-scriptions of a video, which could range anywhere from moment-by-moment detail to a single phrase summary. To provide a more thorough evaluation of the capabilities of long video retrieval systems, we propose a pipeline that leverages state-of-the-art large language models to care-fully generate a diverse set of synthetic captions for long videos. We validate this pipeline's fidelity via rigorous hu-man inspection. We use synthetic captionsfrom this pipeline to perform a benchmark of a representative set of video language models using long video datasets, and show that the models struggle on shorter captions. We show that finetuning on this data can both mitigate these issues (+2.8% R@ 1 over SOTA on ActivityNet with diverse captions), and even improve performance on standard paragraph-to-video re-trieval (+ 1.0% R@1 on ActivityNet). We also use synthetic data from our pipeline as query expansion in the zero-shot setting (+3.4% R@ 1 on ActivityNet). We derive insights by analyzing failure cases for retrieval with short captions.
Matthew Gwilliam, Michael Cogswell, Meng Ye 0002, Karan Sikka, Abhinav Shrivastava, Ajay Divakaran
WACV6
2024 DRESS : Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback
abstract
We present DRESS , a large vision language model (LVLM) that innovatively exploits Natural Language feedback (NLF) from Large Language Models to enhance its alignment and interactions by addressing two key limitations in the state-of-the-art LVLMs. First, prior LVLMs generally rely only on the instruction finetuning stage to enhance alignment with human preferences. Without incorporating extra feedback, they are still prone to generate unhelpful, hallucinated, or harmful responses. Second, while the visual instruction tuning data is generally structured in a multi-turn dialogue format, the connections and dependencies among consecutive conversational turns are weak. This reduces the capacity for effective multi-turn interactions. To tackle these, we propose a novel categorization of the NLF into two key types: critique and refinement. The critique NLF identifies the strengths and weaknesses of the responses and is used to align the LVLMs with human preferences. The refinement NLF offers concrete suggestions for improvement and is adopted to improve the interaction ability of the LVLMs- which focuses on LVLMs' ability to refine responses by incorporating feedback in multi-turn interactions. To address the non-differentiable nature of NLF, we generalize conditional reinforcement learning for training. Our experimental results demonstrate that DRESS can generate more helpful (9.76%), honest (11.52%), and harmless (21.03%) responses, and more effectively learn from feedback during multi-turn interactions compared to SOTA LVLMs.
Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji 0001, Ajay Divakaran
CVPR5
2024 Pelican: Correcting Hallucination in Vision-LLMs via Claim Decomposition and Program of Thought Verification
abstract
Large Visual Language Models (LVLMs) struggle with hallucinations in visual instruction following task(s). These issues hinder their trustworthiness and real-world applicability. We propose Pelican – a novel framework designed to detect and mitigate hallucinations through claim verification. Pelican first decomposes the visual claim into a chain of sub-claims based on first-order predicates. These sub-claims consists of (predicate, question) pairs and can be conceptualized as nodes of a computational graph. We then use use Program-of-Thought prompting to generate Python code for answering these questions through flexible composition of external tools. Pelican improves over prior work by introducing (1) intermediate variables for precise grounding of object instances, and (2) shared computation for answering the sub-question to enable adaptive corrections and inconsistency identification. We finally use reasoning abilities of LLM to verify the correctness of the the claim by considering the consistency and confidence of the (question, answer) pairs from each sub-claim. Our experiments demonstrate consistent performance improvements over various baseline LVLMs and existing hallucination mitigation approaches across several benchmarks.
Pritish Sahu, Karan Sikka, Ajay Divakaran
EMNLP3
2024 Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models
abstract
Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji, Ajay Divakaran. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji 0001, Ajay Divakaran
NAACL-HLT5
2023 Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
abstract
The recent growth in the consumption of online media by children during early childhood necessitates data-driven tools enabling educators to filter out appropriate educational content for young learners. This paper presents an approach for detecting educational content in online videos. We focus on two widely used educational content classes: literacy and math. For each class, we choose prominent codes (sub-classes) based on the Common Core Standards. For example, literacy codes include ‘letter names’, ‘letter sounds’, and math codes include ‘counting’, ‘sorting’. We pose this as a finegrained multilabel classification problem as videos can contain multiple types of educational content and the content classes can get visually similar (e.g., ‘letter names’vs ‘letter sounds’). We propose a novel class prototypes based supervised contrastive learning approach that can handle fine-grained samples associated with multiple labels. We learn a class prototype for each class and a loss function is employed to minimize the distances between a class prototype and the samples from the class. Similarly, distances between a class prototype and the samples from other classes are maximized. As the alignment between visual and audio cues are crucial for effective comprehension, we consider a multimodal transformer network to capture the interaction between visual and audio cues in videos while learning the embedding for videos. For evaluation, we present a dataset, APPROVE, employing educational videos from YouTube labeled with fine-grained education classes by education researchers. APPROVE consists of 193 hours of expert-annotated videos with 19 classes. The proposed approach outperforms strong baselines on APPROVE and other benchmarks such as Youtube-8M, and COIN. The dataset is available at https://nusci.csl.sri.com/project/APPROVE.
Rohit Gupta 0012, Claire Christensen, Sujeong Kim, Sarah Gerard, Madeline Cincebeaux, Ajay Divakaran, Todd Grindal, Mubarak Shah
CVPR7
2023 Multilingual Content Moderation: A Case Study on Reddit
abstract
Meng Ye, Karan Sikka, Katherine Atwell, Sabit Hassan, Ajay Divakaran, Malihe Alikhani. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Meng Ye 0002, Karan Sikka, Katherine Atwell, Sabit Hassan, Ajay Divakaran, Malihe Alikhani
EACL5
2023 TIJO: Trigger Inversion with Joint Optimization for Defending Multimodal Backdoored Models
abstract
We present a Multimodal Backdoor Defense technique TIJO (Trigger Inversion using Joint Optimization). Recent work [48] has demonstrated successful backdoor attacks on multimodal models for the Visual Question Answering task. Their dual-key backdoor trigger is split across two modalities (image and text), such that the backdoor is activated if and only if the trigger is present in both modalities. We propose TIJO that defends against dual-key attacks through a joint optimization that reverse-engineers the trigger in both the image and text modalities. This joint optimization is challenging in multimodal models due to the disconnected nature of the visual pipeline which consists of an offline feature extractor, whose output is then fused with the text using a fusion module. The key insight enabling the joint optimization in TIJO is that the trigger inversion needs to be carried out in the object detection box feature space as opposed to the pixel space. We demonstrate the effectiveness of our method on the TrojVQA benchmark, where TIJO improves upon the state-of-the-art unimodal methods from an AUC of 0.6 to 0.92 on multimodal dual-key back-doors. Furthermore, our method also improves upon the unimodal baselines on unimodal backdoors. We present ablation studies and qualitative results to provide insights into our algorithm such as the critical importance of overlaying the inverted feature triggers on all visual features during trigger inversion. The prototype implementation of TIJO is available at https://github.com/SRI-CSL/TIJO.
Indranil Sur, Karan Sikka, Matthew Walmer, Kaushik Koneripalli, Ajay Divakaran, Susmit Jha
ICCV7
2023 Predicting Information Pathways Across Online Communities
abstract
The problem of community-level information pathway prediction (CLIPP) aims at predicting the transmission trajectory of content across online communities. A successful solution to CLIPP holds significance as it facilitates the distribution of valuable information to a larger audience and prevents the proliferation of misinfor- mation. Notably, solving CLIPP is non-trivial as inter-community relationships and influence are unknown, information spread is multi-modal, and new content and new communities appear over time. In this work, we address CLIPP by collecting large-scale, multi-modal datasets to examine the diffusion of online YouTube videos on Reddit. We analyze these datasets to construct community influence graphs (CIGs) and develop a novel dynamic graph frame- work, INPAC (Information Pathway Across Online Communities), which incorporates CIGs to capture the temporal variability and multi-modal nature of video propagation across communities. Ex- perimental results in both warm-start and cold-start scenarios show that INPAC outperforms seven baselines in CLIPP. Our code and datasets are available at https://github.com/claws-lab/INPAC
Yiqiao Jin, Yeon-Chang Lee, Kartik Sharma, Meng Ye 0002, Karan Sikka, Ajay Divakaran, Srijan Kumar
KDD6
2022 Detecting Out-Of-Context Objects Using Graph Contextual Reasoning Network
abstract
This paper presents an approach for detecting out-of-context (OOC) objects in images. Given an image with a set of objects, our goal is to determine if an object is inconsistent with the contextual relations and detect the OOC object with a bounding box. In this work, we consider common contextual relations such as co-occurrence relations, the relative size of an object with respect to other objects, and the position of the object in the scene. We posit that contextual cues are useful to determine object labels for in-context objects and inconsistent context cues are detrimental to determining object labels for out-of-context objects. To realize this hypothesis, we propose a graph contextual reasoning network (GCRN) to detect OOC objects. GCRN consists of two separate graphs to predict object labels based on the contextual cues in the image: 1) a representation graph to learn object features based on the neighboring objects and 2) a context graph to explicitly capture contextual cues from the neighboring objects. GCRN explicitly captures the contextual cues to improve the detection of in-context objects and identify objects that violate contextual relations. In order to evaluate our approach, we create a large-scale dataset by adding OOC object instances to the COCO images. We also evaluate on recent OCD benchmark. Our results show that GCRN outperforms competitive baselines in detecting OOC objects and correctly detecting in-context objects. Code and data: https://nusci.csl.sri.com/project/trinity-ooc
Manoj Acharya, Kaushik Koneripalli, Susmit Jha, Christopher Kanan, Ajay Divakaran
IJCAI6
2022 Challenges in Procedural Multimodal Machine Comprehension: A Novel Way To Benchmark
abstract
We focus on Multimodal Machine Reading Comprehension (M3C) where a model is expected to answer questions based on given passage (or context), and the context and the questions can be in different modalities. Previous works such as RecipeQA have proposed datasets and cloze-style tasks for evaluation. However, we identify three critical biases stemming from the question-answer generation process and memorization capabilities of large deep models. These biases makes it easier for a model to overfit by relying on spurious correlations or naive data patterns. We propose a systematic framework to address these biases through three Control-Knobs that enable us to generate a test bed of datasets of progressive difficulty levels. We believe that our benchmark (referred to as Meta- RecipeQA) will provide, for the first time, a fine grained estimate of a model’s generalization capabilities. We also propose a general M3C model that is used to realize several prior SOTA models and motivate a novel hierarchical transformer based reasoning network (HTRN). We perform a detailed evaluation of these models with different language and visual features on our benchmark. We observe a consistent improvement with HTRN over SOTA (~ 18% in Visual Cloze task and ~ 13% in average over all the tasks). We also observe a drop in performance across all the models when testing on RecipeQA and proposed Meta–RecipeQA (e.g. 83.6% versus 67.1% for HTRN), which shows that the proposed dataset is relatively less biased. We conclude by highlighting the impact of the control knobs with some quantitative results.
Pritish Sahu, Karan Sikka, Ajay Divakaran
WACV3
2021 Confidence Calibration for Domain Generalization under Covariate Shift
abstract
Existing calibration algorithms address the problem of covariate shift via unsupervised domain adaptation. However, these methods suffer from the following limitations: 1) they require unlabeled data from the target domain, which may not be available at the stage of calibration in real-world applications and 2) their performance depends heavily on the disparity between the distributions of the source and target domains. To address these two limitations, we present novel calibration solutions via domain generalization. Our core idea is to leverage multiple calibration domains to reduce the effective distribution disparity between the target and calibration domains for improved calibration transfer without needing any data from the target domain. We provide theoretical justification and empirical experimental results to demonstrate the effectiveness of our proposed algorithms. Compared against state-of-the-art calibration methods designed for domain adaptation, we observe a decrease of 8.86 percentage points in expected calibration error or, equivalently, an increase of 35 percentage points in improvement ratio for multi-class classification on the Office-Home dataset.
Yunye Gong, Thomas G. Dietterich, Ajay Divakaran, Melinda T. Gervasio
ICCV5
2020 Stacked Spatio-Temporal Graph Convolutional Networks for Action Segmentation
abstract
We propose novel Stacked Spatio-Temporal Graph Convolutional Networks (Stacked-STGCN) for action segmentation, i.e., predicting and localizing a sequence of actions over long videos. We extend the Spatio-Temporal Graph Convolutional Network (STGCN) originally proposed for skeleton-based action recognition to enable nodes with different characteristics (e.g., scene, actor, object, action), feature descriptors with varied lengths, and arbitrary temporal edge connections to account for large graph deformation commonly associated with complex activities. We further introduce the stacked hourglass architecture to STGCN to leverage the advantages of an encoder-decoder design for improved generalization performance and localization accuracy. We explore various descriptors such as frame- level VGG, segment-level I3D, RCNN-based object, etc. as node descriptors to enable action segmentation based on joint inference over comprehensive contextual information. We show results on CAD120 (which provides pre-computed node features and edge weights for fair performance comparison across algorithms) as well as a more complex real- world activity dataset, Charades. Our Stacked-STGCN in general achieves improved performance over the state-of- the-art for both CAD120 and Charades. Moreover, due to its generic design, Stacked-STGCN can be applied to a wider range of applications that require structured inference over long sequences with heterogeneous data types and varied temporal extent.
Pallabi Ghosh, Larry Davis 0001, Ajay Divakaran
WACV4
2019 Integrating Text and Image: Determining Multimodal Document Intent in Instagram Posts
abstract
Julia Kruk, Jonah Lubin, Karan Sikka, Xiao Lin, Dan Jurafsky, Ajay Divakaran. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Julia Kruk, Jonah Lubin, Karan Sikka, Daniel Jurafsky, Ajay Divakaran
EMNLP/IJCNLP (1)6
2019 Sunny and Dark Outside?! Improving Answer Consistency in VQA through Entailed Question Generation
abstract
Arijit Ray, Karan Sikka, Ajay Divakaran, Stefan Lee, Giedrius Burachas. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Arijit Ray, Karan Sikka, Ajay Divakaran, Stefan Lee, Giedrius Burachas
EMNLP/IJCNLP (1)3
2019 Can You Explain That? Lucid Explanations Help Human-AI Collaborative Image Retrieval
abstract
While there have been many proposals on making AI algorithms explainable, few have attempted to evaluate the impact of AI-generated explanations on human performance in conducting human-AI collaborative tasks. To bridge the gap, we propose a Twenty-Questions style collaborative image retrieval game, Explanation-assisted Guess Which (ExAG), as a method of evaluating the efficacy of explanations (visual evidence or textual justification) in the context of Visual Question Answering (VQA). In our proposed ExAG, a human user needs to guess a secret image picked by the VQA agent by asking natural language questions to it. We show that overall, when AI explains its answers, users succeed more often in guessing the secret image correctly. Notably, a few correct explanations can readily improve human performance when VQA answers are mostly incorrect as compared to no-explanation games. Furthermore, we also show that while explanations rated as “helpful” significantly improve human performance, “incorrect” and “unhelpful” explanations can degrade performance as compared to no-explanation games. Our experiments, therefore, demonstrate that ExAG is an effective means to evaluate the efficacy of AI-generated explanation on a human-AI collaborative task.
Arijit Ray, Rakesh Kumar 0001, Ajay Divakaran, Giedrius Burachas
HCOMP4
2019 Align2Ground: Weakly Supervised Phrase Grounding Guided by Image-Caption Alignment
abstract
We address the problem of grounding free-form textual phrases by using weak supervision from image-caption pairs. We propose a novel end-to-end model that uses caption-to-image retrieval as a downstream task to guide the process of phrase localization. Our method, as a first step, infers the latent correspondences between regions-of-interest (RoIs) and phrases in the caption and creates a discriminative image representation using these matched RoIs. In the subsequent step, this learned representation is aligned with the caption. Our key contribution lies in building this "caption-conditioned" image encoding, which tightly couples both the tasks and allows the weak supervision to effectively guide visual grounding. We provide extensive empirical and qualitative analysis to investigate the different components of our proposed model and compare it with competitive baselines. For phrase localization, we report an improvement of 4.9% and 1.3% (absolute) over the prior state-of-the-art on the VisualGenome and Flickr30k Entities datasets. We also report results that are at par with the state-of-the-art on the downstream caption-to-image retrieval task on COCO and Flickr30k datasets.
Samyak Datta, Karan Sikka, Karuna Ahuja, Devi Parikh, Ajay Divakaran
ICCV6
2018 Zero-Shot Object Detection
Ankan Bansal, Karan Sikka, Gaurav Sharma 0004, Rama Chellappa, Ajay Divakaran
ECCV (1)5
2018 Deep Multimodal Fusion: A Hybrid Approach
Mohamed R. Amer, Timothy J. Shields, Behjat Siddiquie, Amir Tamrakar, Ajay Divakaran, Sek M. Chai
Int. J. Comput. Vis.5
2016 Multimodal analytics to study collaborative problem solving in pair programming
abstract
Collaborative problem solving (CPS) is seen as a key skill in K-12 education---in computer science as well as other subjects. Efforts to introduce children to computing rely on pair programming as a way of having young learners engage in CPS. Characteristics of quality collaboration are joint exploring or understanding, joint representation, and joint execution. We present a data driven approach to assessing and elucidating collaboration through modeling of multimodal student behavior and performance data.
Shuchi Grover, Marie A. Bienkowski, Amir Tamrakar, Behjat Siddiquie, David A. Salter, Ajay Divakaran
LAK6
2015 The Tower Game Dataset: A multimodal dataset for analyzing social interaction predicates
abstract
We introduce the Tower Game Dataset for computational modeling of social interaction predicates. Existing research in affective computing has focused primarily on recognizing the emotional and mental state of a human based on external behaviors. Recent research in the social science community argues that engaged and sustained social interactions require the participants to jointly coordinate their verbal and non-verbal behaviors. With this as our guiding principle, we collected the Tower Game Dataset consisting of multimodal recordings of two players participating in a tower building game, in the process communicating and collaborating with each other. The format of the game was specifically chosen as it elicits spontaneous communication from the participants through social interaction predicates such as joint attention and entrainment. The dataset will be made public and we believe that it will foster new research in the area of computational social interaction modeling.
David A. Salter, Amir Tamrakar, Behjat Siddiquie, Mohamed R. Amer, Ajay Divakaran, Brian Lande, Darius Mehri
ACII5
2015 Audio-based affect detection in web videos
abstract
We present a new technique for detecting audio concepts in web content as well outline the technique's applications to video sequence parsing. Our focus is primarily on affective concepts and in order to study them we have collected a new dataset, consisting of videos where a speaker is persuading a crowd, called “Rallying a Crowd”. We develop new classifiers for graded levels of arousal in speech as well as crowd noise and music and demonstrate their effectiveness on web content. These techniques achieve high detection accuracy (58.2%) for affective concepts on this new dataset and outperform (36.8%) state-of-the-art techniques (33.1%) for semantic concepts on a previously collected dataset. We also develop a new audio sequence segmentation technique which enables us to rapidly classify subsections of test sequence audio into the aforementioned audio classes. We are thus able to robustly address the detection of affective concepts in highly variable web content as well as the computational challenge of quick classification so as to enable web scale processing.
Dave Chisholm, Behjat Siddiquie, Ajay Divakaran, Elizabeth Shriberg
ICME3
2015 Exploiting Multimodal Affect and Semantics to Identify Politically Persuasive Web Videos
abstract
We introduce the task of automatically classifying politically persuasive web videos and propose a highly effective multi-modal approach for this task. We extract audio, visual, and textual features that attempt to capture affect and semantics in the audio-visual content and sentiment in the viewers' comments. We demonstrate that each of the feature modalities can be used to classify politically persuasive content, and that fusing them leads to the best performance. We also perform experiments to examine human accuracy and inter-coder reliability for this task and show that our best automatic classifier slightly outperforms average human performance. Finally we show that politically persuasive videos generate more strongly negative viewer comments than non-persuasive videos and analyze how affective content can be used to predict viewer reactions.
Behjat Siddiquie, Dave Chisholm, Ajay Divakaran
ICMI3
2015 2nd Workshop on Computational Models of Social Interactions: Human-Computer-Media Communication (HCMC2015)
abstract
Communicating ideas and information from and to humans is a very important subject. In our daily life, human interact with variety of entities, such as, other humans, machines, media. Constructive interactions are needed for good communication, which would result in successful outcomes, such as answering a query, learning a new skill, getting a service done, and communicating emotions. Each of these entities invokes a set of signals. Current research has focused on analyzing one entity's signals with no respect to the other entities in a unidirectional manner. The computer vision community focused on detection, classification and recognition of humans and their poses and gestures progressing onto actions, activities, and events but it does not go beyond that. The signal processing community focused on emotion recognition from facial expressions or audio or both combined. The HCI community focused on making easier interfaces for machines to ease their usage. The goal of this workshop is to bring multiple disciplines together, to process human directed signals holistically, in a bidirectional manner, rather than isolation. This workshop is positioned to display this rich domain of applications, which will provide the necessary next boost for these technologies. At the same time, it seeks to ground computational models on theory that would help achieve the technology goals. This would allow us to leverage decades of research in different fields and to spur interdisciplinary research thereby opening up new problem domains for the multimedia community.
Mohamed R. Amer, Ajay Divakaran, Shih-Fu Chang, Nicu Sebe
ACM Multimedia2
2014 Emotion detection in speech using deep networks
abstract
We propose a novel staged hybrid model for emotion detection in speech. Hybrid models exploit the strength of discriminative classifiers along with the representational power of generative models. Discriminative classifiers have been shown to achieve higher performances than the corresponding generative likelihood-based classifiers. On the other hand, generative models learn a rich informative representations. Our proposed hybrid model consists of a generative model, which is used for unsupervised representation learning of short term temporal phenomena and a discriminative model, which is used for event detection and classification of long range temporal dynamics. We evaluate our approach on multiple audio-visual datasets (AVEC, VAM, and SPD) and demonstrate its superiority compared to the state-of-the-art.
Mohamed R. Amer, Behjat Siddiquie, Colleen Richey, Ajay Divakaran
ICASSP4
2014 Multimodal fusion using dynamic hybrid models
abstract
We propose a novel hybrid model that exploits the strength of discriminative classifiers along with the representational power of generative models. Our focus is on detecting multimodal events in time varying sequences. Discriminative classifiers have been shown to achieve higher performances than the corresponding generative likelihood-based classifiers. On the other hand, generative models learn a rich informative space which allows for data generation and joint feature representation that discriminative models lack. We employ a deep temporal generative model for unsupervised learning of a shared representation across multiple modalities with time varying data. The temporal generative model takes into account short term temporal phenomena and allows for filling in missing data by generating data within or across modalities. The hybrid model involves augmenting the temporal generative model with a temporal discriminative model for event detection, and classification, which enables modeling long range temporal dynamics. We evaluate our approach on audio-visual datasets (AVEC, AVLetters, and CUAVE) and demonstrate its superiority compared to the state-of-the-art.
Mohamed R. Amer, Behjat Siddiquie, Saad M. Khan, Ajay Divakaran, Harpreet Sawhney
WACV4
2013 Dynamic Pooling for Complex Event Recognition
abstract
The problem of adaptively selecting pooling regions for the classification of complex video events is considered. Complex events are defined as events composed of several characteristic behaviors, whose temporal configuration can change from sequence to sequence. A dynamic pooling operator is defined so as to enable a unified solution to the problems of event specific video segmentation, temporal structure modeling, and event detection. Video is decomposed into segments, and the segments most informative for detecting a given event are identified, so as to dynamically determine the pooling operator most suited for each sequence. This dynamic pooling is implemented by treating the locations of characteristic segments as hidden information, which is inferred, on a sequence-by-sequence basis, via a large-margin classification rule with latent variables. Although the feasible set of segment selections is combinatorial, it is shown that a globally optimal solution to the inference problem can be obtained efficiently, through the solution of a series of linear programs. Besides the coarse-level location of segments, a finer model of video structure is implemented by jointly pooling features of segment-tuples. Experimental evaluation demonstrates that the resulting event detector has state-of-the-art performance on challenging video datasets.
Weixin Li 0002, Ajay Divakaran, Nuno Vasconcelos
ICCV3
2013 Affect analysis in natural human interaction using Joint Hidden Conditional Random Fields
abstract
We present a novel approach for multi-modal affect analysis in human interactions that is capable of integrating data from multiple modalities while also taking into account temporal dynamics. Our fusion approach, Joint Hidden Conditional Random Fields (JHRCFs), combines the advantages of purely feature level (early fusion) fusion approaches with late fusion (CRFs on individual modalities) to simultaneously learn the correlations between features from multiple modalities as well as their temporal dynamics. Our approach addresses major shortcomings of other fusion approaches such as the domination of other modalities by a single modality with early fusion and the loss of cross-modal information with late fusion. Extensive results on the AVEC 2011 dataset show that we outperform the state-of-the-art on the Audio Sub-Challenge, while achieving competitive performance on the Video Sub-Challenge and the Audiovisual Sub-Challenge.
Behjat Siddiquie, Saad M. Khan, Ajay Divakaran, Harpreet Sawhney
ICME3
2013 Semantic pooling for complex event detection
abstract
Complex event detection is very challenging in open source such as You-Tube videos, which usually comprise very diverse visual contents involving various object, scene and action concepts. Not all of them, however, are relevant to the event. In other words, a video may contain a lot of "junk" information which is harmful for recognition. Hence, we propose a semantic pooling approach to tackle this issue. Unlike the conventional pooling over the entire video or specific spatial regions of a video, we employ a discriminative approach to acquire abstract semantic "regions" for pooling. For this purpose, we first associate low-level visual words with semantic concepts via their co-occurrence relationship. We then pool the low-level features separately according to their semantic information. The proposed semantic pooling strategy also provides a new mechanism for incorporating semantic concepts for low-level feature based event recognition. We evaluate our approach on TRECVID MED [1] dataset and the results show that semantic pooling consistently improves the performance compared with conventional pooling strategies.
Jingen Liu, Ajay Divakaran, Harpreet Sawhney
ACM Multimedia4
2013 Video event recognition using concept attributes
abstract
We propose to use action, scene and object concepts as semantic attributes for classification of video events in InTheWild content, such as YouTube videos. We model events using a variety of complementary semantic attribute features developed in a semantic concept space. Our contribution is to systematically demonstrate the advantages of this concept-based event representation (CBER) in applications of video event classification and understanding. Specifically, CBER has better generalization capability, which enables to recognize events with a few training examples. In addition, CBER makes it possible to recognize a novel event without training examples (i.e., zero-shot learning). We further show our proposed enhanced event model can further improve the zero-shot learning. Furthermore, CBER provides a straightforward way for event recounting/understanding. We use the TRECVID Multimedia Event Detection (MED11) open source event definitions and datasets as our test bed and show results on over 1400 hours of videos.
Jingen Liu, Omar Javed, Saad Ali, Amir Tamrakar, Ajay Divakaran, Harpreet Sawhney
WACV6
2012 Evaluation of low-level features and their combinations for complex event detection in open source videos
abstract
Low-level appearance as well as spatio-temporal features, appropriately quantized and aggregated into Bag-of-Words (BoW) descriptors, have been shown to be effective in many detection and recognition tasks. However, their effcacy for complex event recognition in unconstrained videos have not been systematically evaluated. In this paper, we use the NIST TRECVID Multimedia Event Detection (MED11 [1]) open source dataset, containing annotated data for 15 high-level events, as the standardized test bed for evaluating the low-level features. This dataset contains a large number of user-generated video clips. We consider 7 different low-level features, both static and dynamic, using BoW descriptors within an SVM approach for event detection. We present performance results on the 15 MED11 events for each of the features as well as their combinations using a number of early and late fusion strategies and discuss their strengths and limitations.
Amir Tamrakar, Saad Ali, Jingen Liu, Omar Javed, Ajay Divakaran, Harpreet Sawhney
CVPR6
2012 How to put it into words - using random forests to extract symbol level descriptions from audio content for concept detection
abstract
This paper presents a system that uses symbolic representations of audio concepts as words for the descriptions of audio tracks, that enable it to go beyond the state of the art, which is audio event classification of a small number of audio classes in constrained settings, to large-scale classification in the wild. These audio words might be less meaningful for an annotator but they are descriptive for computer algorithms. We devise a random-forest vocabulary learning method with an audio word weighting scheme based on TF-IDF and TD-IDD, so as to combine the computational simplicity and accurate multi-class classification of the random forest with the data-driven discriminative power of the TF-IDF/TD-IDD methods. The proposed random forest clustering with text-retrieval methods significantly outperforms two state-of-the-art methods on the dry-run set and the full set of the TRECVID MED 2010 dataset.
Po-Sen Huang, Robert Mertens 0003, Ajay Divakaran, Gerald Friedland, Mark Hasegawa-Johnson
ICASSP3
2012 Multimedia event recounting with concept based representation
abstract
Multimedia event detection has drawn a lot of attention in recent years. Given a recognized event, in this paper, we conduct a pilot study of the multimedia event recounting problem, which answers the question why this video is recognized as this event, i.e. what evidences this decision is made on. In order to provide a semantic recounting of the multimedia event, we adopt a concept-based event representation for learning a discriminative event model. Then, we present a recounting approach that exactly recovers the contribution of semantic evidence to the event classification decision. This approach can be applied on any additive discriminative classifiers. The promising result is shown on the MED11 dataset that contains 15 events in thousands of YouTube like videos.
Jingen Liu, Ajay Divakaran, Harpreet Sawhney
ACM Multimedia4
2011 On the Applicability of Speaker Diarization to Audio Concept Detection for Multimedia Retrieval
abstract
Recently, audio concepts emerged as a useful building block in multimodal video retrieval systems. Information like "this file contains laughter", "this file contains engine sounds" or "this file contains slow music" can significantly improve purely visual based retrieval. The weak point of current approaches to audio concept detection is that they heavily rely on human annotators. In most approaches, audio material is manually inspected to identify relevant concepts. Then instances that contain examples of relevant concepts are selected -- again manually -- and used to train concept detectors. This approach comes with two major disadvantages: (1) it leads to rather abstract audio concepts that hardly cover the audio domain at hand and (2) the way human annotators identify audio concepts likely differs from the way a computer algorithm clusters audio data -- introducing additional noise in training data. This paper explores whether unsupervized audio segementation systems can be used to identify useful audio concepts by analyzing training data automatically and whether these audio concepts can be used for multimedia document classification and retrieval. A modified version of the ICSI (International Computer Science Institute) speaker diarization system finds segments in an audio track that have similar perceptual properties and groups these segments. This article provides an in-depth analysis on the statistic properties of similar acoustic segments identified by the diarization system in a predefined document set and the theoretical fitness of this approach to discern one document class from another.
Robert Mertens 0003, Po-Sen Huang, Luke R. Gottlieb, Gerald Friedland, Ajay Divakaran
ISM5
2009 Recognition and volume estimation of food intake using a mobile device
abstract
We present a system that improves accuracy of food intake assessment using computer vision techniques. Traditional dietetic method suffers from the drawback of either inaccurate assessment or complex lab measurement. Our solution is to use a mobile phone to capture images of foods, recognize food types, estimate their respective volumes and finally return quantitative nutrition information. Automated and accurate food recognition presents the following challenges. First, there exist a large variety of food types that people consume in everyday life. Second, a single category of food may contain large variations due to different ways of preparation. Also, diverse lighting conditions may lead to varying visual appearance of foods. All of these pose a challenge to the state of the art recognition approaches. Moreover, the low quality images captured using cellphones make the task of 3D reconstruction difficult. In this paper, we combine several vision techniques (visual recognition and 3D reconstruction) to achieve quantitative food intake estimation. Evaluation of both recognition and reconstruction is provided in the experimental results.
Manika Puri, Ajay Divakaran, Harpreet Sawhney
WACV4
2008 Speech denoising using nonnegative matrix factorization with priors
abstract
We present a technique for denoising speech using nonnegative matrix factorization (NMF) in combination with statistical speech and noise models. We compare our new technique to standard NMF and to a state-of-the-art Wiener filter implementation and show improvements in speech quality across a range of interfering noise types.
Kevin W. Wilson, Bhiksha Raj, Paris Smaragdis, Ajay Divakaran
ICASSP4
2007 An SVM Framework for Genre-Independent Scene Change Detection
abstract
We present a novel genre-independent SVM framework for detecting scene changes in broadcast video. Our framework works on content from a diverse range of genres by allowing sets of features, extracted from both audio and video streams, to be combined and compared automatically without the use of explicit thresholds. For ground truth, we use hand-labeled video scene boundaries from a wide variety of broadcast genres to generate positive and negative samples for the SVM. Our experiments include high-and low-level audio features such as semantic histograms and distances between Gaussian models, as well as video features such as shot cut positions. We evaluate the importance of these measures in a structured frame-work, with performance comparisons obtained via ROC curves. We achieve over 70% detection rate for 10% false positive rate on our corpus of over 7.5 hours of data collected from news, talk shows, sitcoms, dramas, music videos, and how-to shows.
Naveen Goela, Kevin W. Wilson, Feng Niu, Ajay Divakaran, Isao Otsuka
ICME4
2006 Generative Process Tracking for Audio Analysis
abstract
The problem of generative process tracking involves detecting and adapting to changes in the underlying generative process that creates a time series of observations. It has been widely used for visual background modelling to adaptively track the generative process that generates the pixel intensities. In this paper, we extend this idea to audio background modelling and show its applications in surveillance domain. We adaptively learn the parameters of the generative audio background process and detect foreground events. We have tested the effectiveness of the proposed algorithms using synthetic time series data and show its performance on elevator audio surveillance
Regunathan Radhakrishnan, Ajay Divakaran
ICASSP (5)2
2006 Broadcast Video Program Summarization using Face Tracks
abstract
We present a novel video summarization and skimming technique using face detection on broadcast video programs. We take the faces in video as our primary target as they constitute the focus of most consumer video programs. We detect face tracks in video and define face-scene fragments based on start and end of face tracks. We define a fast-forward skimming method using frames selected from fragments, thus covering all the faces and their interactions in the video program. We also define novel constraints for a smooth and visually representative summary, and construct longer but smoother summaries
Kadir A. Peker, Isao Otsuka, Ajay Divakaran
ICME3
2006 Sports Program Boundary Detection
abstract
In the recent years, consumer devices that can record broadcast video have become prevalent. Such devices rely on the program guide information about a program start time and end time for recording. However, sports broadcasts can run over the specified time sometimes. In this paper, we propose a framework based on an audio classification framework to correctly detect the end of a sport broadcast so as to enable complete recording of the game. Our experimental results show that the proposed algorithm can detect sports program boundaries with a high accuracy
Regunathan Radhakrishnan, Ajay Divakaran, Isao Otsuka
ICME2
2005 Layered dynamic mixture model for pattern discovery in asynchronous multi-modal streams [video applications]
abstract
We propose a layered dynamic mixture model for asynchronous multi-modal fusion for unsupervised pattern discovery in video. The lower layer of the model uses generative temporal structures such as a hierarchical hidden Markov model to convert the audiovisual streams into mid-level labels, it also models the correlations in text with probabilistic latent semantic analysis. The upper layer fuses the statistical evidence across diverse modalities with a flexible meta-mixture model that assumes loose temporal correspondence. Evaluation on a large news database shows that multi-modal clusters have better correspondence to news topics than audio-visual clusters alone; novel analysis techniques suggest that meaningful clusters occur when the prediction of salient features by the model concurs with those shown in the story clusters.
Lexing Xie, Lyndon S. Kennedy, Shih-Fu Chang, Ajay Divakaran, Huifang Sun, Ching-Yung Lin
ICASSP (2)4
2005 Highlights extraction from sports video based on an audio-visual marker detection framework
abstract
We propose to use a visual object (e.g., the baseball catcher) detection algorithm to find local, semantic objects in video frames in addition to an audio classification algorithm to find semantic audio objects in the audio track for sports highlights extraction. The highlight candidates are then further grouped into finer-resolution highlight segments, using color or motion information. During the grouping phase, many of the false alarms can be correctly identified and eliminated. Our experimental results with baseball, soccer and golf video are promising.
Ziyou Xiong, Regunathan Radhakrishnan, Ajay Divakaran, Thomas S. Huang
ICME3
2004 Video mining: pattern discovery versus pattern recognition
abstract
We examine the significance of video mining as pattern discovery in multimedia content. We examine the underlying issue of pattern discovery versus pattern recognition, since most past work has not drawn such a sharp distinction. We argue that while the term "pattern discovery" implies a purely unsupervised approach, in practice a mixture of unsupervised and supervised techniques will have to be used. We compare conventional data mining with video mining and observe that a key difference is in the multilayered semantics of multimedia content. We then identify significant challenges posed by video mining.
Ajay Divakaran, Kadir A. Peker, Shih-Fu Chang, Regunathan Radhakrishnan, Lexing Xie
ICIP1
2004 Discovering meaningful multimedia patterns with audio-visual concepts and associated text
abstract
The work presents the first effort to automatically annotate the semantic meanings of temporal video patterns obtained through unsupervised discovery processes. This problem is interesting in domains where neither perceptual patterns nor semantic concepts have simple structures. The patterns in video are modeled with hierarchical hidden Markov models (HHMM), with efficient algorithms to learn the parameters, the model complexity and the relevant features; the meanings are contained in words of the speech transcript of the video. The pattern-word association is obtained via cooccurrence analysis and statistical machine translation models. Promising results are obtained through extensive experiments on 20+ hours of TRECVID news videos: video patterns that associate with distinct topics such as el-nino and politics are identified; the HHMM temporal structure model compares favorably to a nontemporal clustering algorithm.
Lexing Xie, Lyndon S. Kennedy, Shih-Fu Chang, Ajay Divakaran, Huifang Sun, Ching-Yung Lin
ICIP4
2004 Towards maximizing the end-user experience
abstract
Increasing bandwidth and diverse devices have led to a proliferation of options for the consumer of digital video. The user can easily be overwhelmed by the diversity of devices, bandwidth and content. In our view, maximizing the end-user experience involves first helping the user to easily search for and retrieve desired content, and then delivering the content to him by seamlessly adapting to his device and available bandwidth. The first part is therefore based on semantic criteria while the second part is based on purely signal-based criteria. We propose that scalable content summarization combined with dynamic content adaptation, and unobtrusive processing of user preferences comprise the key technologies for maximizing the end-user experience.
Ajay Divakaran, Anthony Vetro, T. Kan
ICME1
2004 Adaptive fast playback-based video skimming using a compressed-domain visual complexity measure
abstract
We present a novel compressed domain measure of spatio-temporal activity or visual complexity of a video segment. The visual complexity measure indicates how fast a video segment can be played within human perceptual limits. We present an adaptive "smart fast-forward" based video skimming method where the playback speed is varied based on the visual complexity. Alternatively, spatio-temporal smoothing is used to reduce visual complexity for an acceptable playback at a given playback speed. The complexity measure and the skimming method are based on early vision principles, thus they are applicable across a wide range of content type and applications. It is best suited for low temporal compression instant skims. It preserves the temporal continuity and eliminates the risk of missing an important event. It can be extended to include semantic inputs such as face or event detection, or can be a presentation end to semantic summarization.
Kadir A. Peker, Ajay Divakaran
ICME2
2004 Time series analysis and segmentation using eigenvectors for mining semantic audio label sequences
abstract
Pattern discovery from video has promising applications in summarizing different genre types, including surveillance and sports. After pattern discovery, a summary of the video can be constructed from a combination of usual and unusual patterns, depending on the application domain. Previously, we used an unsupervised label mining approach to extract highlight moments from soccer videos (Radhakrishan, R. et al., IEEE Pacific-Rim Conf. on Multimedia, 2003). We now formulate the problem of pattern discovery from semantic audio labels as a time series clustering problem and propose a new unsupervised mining framework based on segmentation theory using eigenvectors of the affinity matrix. We test the validity of the technique using synthetically generated label sequences as well as label sequences from broadcast sports video. Our sports highlights extraction accuracy is comparable to that achieved in our previous work.
Regunathan Radhakrishnan, Ziyou Xiong, Ajay Divakaran, T. Kan
ICME3
2004 Effective and efficient sports highlights extraction using the minimum description length criterion in selecting GMM structures
abstract
In fitting the training data with Gaussian mixture models (GMMs) of appropriate structures using the MDL (minimum description length) criterion, we are able to improve audio classification accuracy with a large margin. With the MDL-GMMs, we are also able to greatly improve the accuracy in extracting sports highlights. Since we have focused on audio domain processing, it enables us to extract highlights very quickly. We have demonstrated the importance of a better understanding of model structures in such a pattern recognition task.
Ziyou Xiong, Regunathan Radhakrishnan, Ajay Divakaran, Thomas S. Huang
ICME3
2004 MPEG-7 meta-data enhanced encoder system for embedded systems
abstract
We describe a MPEG-7 Meta-Data enhanced Audio-Visual Encoder system that targets DVD recorders. We extract features in the compressed domain with both video and audio, which allows us to add the meta-data extraction without altering the hardware architecture of the encoder core. Our feature extraction algorithms are simple, and thus implementable through a simple combination of software and hardware on the integrated DVD chip. The primary application of the meta-data is video summarization, which enables rapid browsing of stored video by the end user. The simplicity of our summarization and feature extraction algorithms enables incorporation of the powerful functionality of smart content navigation through content summarization, into the DVD recorder at a low cost.
Kohtaro Asai, Hirofumi Nishikawa, Daiki Kudo, Ajay Divakaran
VCIP4
2004 Framework for measurement of the intensity of motion activity of video segments
Kadir A. Peker, Ajay Divakaran
J. Vis. Commun. Image Represent.2
2004 Structure analysis of soccer video with domain knowledge and hidden Markov models
Lexing Xie, Shih-Fu Chang, Ajay Divakaran, Huifang Sun
Pattern Recognit. Lett.4
2003 Comparing MFCC and MPEG-7 audio features for feature extraction, maximum likelihood HMM and entropic prior HMM for sports audio classification
abstract
We present a comparison of 6 methods for classification of sports audio. For feature extraction, we have two choices: MPEG-7 audio features and Mel-scale frequency cepstrum coefficients (MFCC). For classification, we also have two choices: maximum likelihood hidden Markov models (ML-HMM) and entropic prior HMMs (EP-HMM). EP-HMMs, in turn, have two variations: with and without trimming of the model parameters. We thus have 6 possible methods, each of which corresponds to a combination. Our results show that all the combinations achieve classification accuracy of around 90% with the best and the second best being, respectively, MPEG-7 features with EP-HMM and MFCC with ML-HMM.
Ziyou Xiong, Regunathan Radhakrishnan, Ajay Divakaran, Thomas S. Huang
ICASSP (5)3
2003 Audio events detection based highlights extraction from baseball, golf and soccer games in a unified framework
abstract
We have developed a unified framework to extract highlights from three sports - baseball, golf and soccer - by detecting some of the common audio events that are directly indicative of highlights. We used MPEG-7 audio features and entropic prior hidden Markov models (HMM) for feature extraction and classification, respectively, to recognize these common audio events. Together with pre- and post-processing techniques using general sports knowledge, we have been able to generate promising results dealing with an audio track that is dominated by audio mixtures and noisy background.
Ziyou Xiong, Regunathan Radhakrishnan, Ajay Divakaran, Thomas S. Huang
ICASSP (5)3
2003 Feature selection for unsupervised discovery of statistical temporal structures in video
abstract
In this paper, we present algorithms for automatic feature selection for of structure discovery from video sequences. Feature selection in this scenario is hard because of the absence of class labels to evaluate against, and the temporal correlation among samples that prevents the direct estimation of posterior probabilities of the cluster given the sequence. The overall problem of structure discovery is formulated as simultaneously finding the statistical descriptions of structure and locating segments that matches the descriptions. Under Markov assumptions among events, structures in the video are modelled with hierarchical hidden Markov models, with efficient algorithms to jointly learn the model parameters and the optimal model complexity. Feature selection iterates between a wrapper step that partitions the large feature pool into consistent subsets, and a filter step that eliminate redundancy within these subsets, respectively. The feature subsets are then ranked according to the normalized Bayesian Information criteria, and the learning results from these ranked subsets can be evaluated and interpreted by a human observer. Results on soccer and baseball videos show that the automatically selected feature set coincides with those selected with domain knowledge and intuition, while achieving a correspondence comparable to that of supervised learning against manually labelled ground truth.
Lexing Xie, Shih-Fu Chang, Ajay Divakaran, Huifang Sun
ICIP (1)3
2003 Generation of sports highlights using motion activity in combination with a common audio feature extraction framework
abstract
In our past work we have used temporal patterns of motion activity to extract sports highlights. We have also used audio classification based approaches to develop a common audio-based platform for feature extraction that works across three different sports. In this paper, we combine the two aforementioned complementary approaches so as to get higher accuracy. We propose a framework for mining the semantic audio-visual labels in order to detect "interesting" events. Our results show that the proposed techniques work well across our three sports of interest, soccer, golf and baseball.
Zixiang Xiong, Regunathan Radhakrishnan, Ajay Divakaran
ICIP (1)3
2003 Multi-camera calibration, object tracking and query generation
abstract
An automatic object tracking and video summarization method for multi-camera systems with a large number of non-overlapping field-of-view cameras is explained. In this framework, video sequences are stored for each object as opposed to storing a sequence for each camera. Object-based representation enables annotation of video segments, and extraction of content semantics for further analysis. We also present a novel solution to the inter-camera color calibration problem. The transitive model function enables effective compensation for lighting changes and radiometric distortions for large-scale systems. After initial calibration, objects are tracked at each camera by background subtraction and mean-shift analysis. The correspondence of objects between different cameras is established by using a Bayesian belief network. This framework empowers the user to get a concise response to queries such as "which locations did an object visit on Monday and what did it do there?".
Fatih Porikli, Ajay Divakaran
ICME2
2003 Unsupervised discovery of multilevel statistical video structures using hierarchical hidden Markov models
abstract
Structure elements in a time sequence (e.g. video) are repetitive segments with consistent deterministic or stochastic characteristics. While most existing work in detecting structures follows a supervised paradigm, we propose a fully unsupervised statistical solution in this paper. We present a unified approach to structure discovery from long video sequences as simultaneously finding the statistical descriptions of structure and locating segments that matches the descriptions. We model the multilevel statistical structure as hierarchical hidden Markov models, and present efficient algorithms for learning both the parameters and the model structure. When tested on a specific domain, soccer video, the unsupervised learning scheme achieves very promising results: it automatically discovers the statistical descriptions of high-level structures, and at the same time achieves even slightly better accuracy in detecting discovered structures in unlabelled videos than a supervised approach designed with domain knowledge and trained with comparable hidden Markov models.
Lexing Xie, Shih-Fu Chang, Ajay Divakaran, Huifang Sun
ICME3
2003 Comparing MFCC and MPEG-7 audio features for feature extraction, maximum likelihood HMM and entropic prior HMM for sports audio classification
abstract
We present a comparison of 6 methods for classification of sports audio. For the feature extraction we have two choices: MPEG-7 audio features and Mel-scale frequency cepstrum coefficients (MFCC). For the classification we also have two choices: maximum likelihood hidden Markov models (ML-HMM) and entropic prior HMM (EP-HMM). EP-HMM, in turn, has two variations: with and without trimming of the model parameters. We thus have 6 possible methods, each of which corresponds to a combination. Our results show that all the combinations achieve classification accuracy of around 90% with the best and the second best being MPEG-7 features with EP-HMM and MFCC with ML-HMM.
Ziyou Xiong, Regunathan Radhakrishnan, Ajay Divakaran, Thomas S. Huang
ICME3
2003 Audio events detection based highlights extraction from baseball, golf and soccer games in a unified framework
abstract
We developed a unified framework to extract highlights from three sports: baseball, golf and soccer by detecting some of the common audio events that are directly indicative of highlights. We used MPEG-7 audio features and entropic prior hidden Markov models (HMM) as the audio features and classifier respectively to recognize these common audio events. Together with pre- and post-processing techniques using general sports knowledge, we have been able to generate promising results dealing with the audio track that is dominated by audio mixtures and noisy background.
Ziyou Xiong, Regunathan Radhakrishnan, Ajay Divakaran, Thomas S. Huang
ICME3
2003 Survey of compressed-domain features used in audio-visual indexing and analysis
Hualu Wang, Ajay Divakaran, Anthony Vetro, Shih-Fu Chang, Huifang Sun
J. Vis. Commun. Image Represent.2
2002 Structure analysis of soccer video with hidden Markov models
abstract
In this paper, we present algorithms for parsing the structure of produced soccer programs. The problem is important in the context of a personalized video streaming and browsing system. While prior work focuses on the detection of special events such as goals or corner kicks, this paper is concerned with generic structural elements of the game. We begin by defining two mutually exclusive states of the game, play and break based on the rules of soccer. We select a domain-tuned feature set, dominant color ratio and motion intensity, based on the special syntax and content characteristics of soccer videos. Each state of the game has a stochastic structure that is modeled with a set of hidden Markov models. Finally, standard dynamic programming techniques are used to obtain the maximum likelihood segmentation of the game into the two states. The system works well, with 83.5% classification accuracy and good boundary timing from extensive tests over diverse data sets.
Lexing Xie, Shih-Fu Chang, Ajay Divakaran, Huifang Sun
ICASSP3
2002 Motion activity-based extraction of key-frames from video shots
abstract
We describe a key-frame extraction technique based on the intuition that the higher the motion the more the number of key-frames required for summarization. We verify experimentally that the intensity of motion activity directly indicates the summarizability of the video segment, by using the MPEG-7 motion activity descriptor (see Jeannin, S. and Divakaran, A., IEEE Trans. Circuits and Systems for Video Tech., vol.11, no.6, p.720-4, 2001) and the fidelity measure described by H.S. Chang et al. (see IEEE Trans. Circuits and Systems for Video Tech., vol.9, no.8, p.1269-79, 1999). We obtain the key-frames by dividing the shot in parts of equal cumulative motion activity, and then selecting the frame located at the half-way point of each sub-segment. Furthermore, we establish an empirical relationship between the motion activity of a segment and the required number of key-frames. We thus provide a unique and rapid way to find the required number of key-frames and compute them. Our scheme is much faster than conventional color-based key-frame extraction schemes since it relies on simple computation and compressed domain extraction. It is close to the theoretical optimum in accuracy.
Ajay Divakaran, Kadir A. Peker, Regunathan Radhakrishnan
ICIP (1)1
2002 Representation of motion activity in hierarchical levels for video indexing and filtering
abstract
A method for video indexing and filtering based on motion activity characteristics in hierarchical levels is proposed. To extract motion activity information, an MPEG (MPEG-1/2) video is first adaptively segmented into hierarchical levels with fixed percentage of original video length based on P-frame macroblock motion information. Three motion activity characteristics - motion intensity which represents the degree of change in motion, motion intensity histogram which represents the temporal statistics of motion intensity, and spatial descriptor which represents the spatial attribute of motion, are then computed to represent different levels of video. The descriptors from different levels are used selectively in different steps of video indexing and filtering. Experimental results show the proposed method is fast and effective, and provides a powerful video indexing and filtering tool.
Xinding Sun, Ajay Divakaran, B. S. Manjunath
ICIP (1)2
2001 An Overview of MPEG-7 Motion Descriptors and Their Applications
Ajay Divakaran
CAIP1
2001 Constant pace skimming and temporal sub-sampling of video using motion activity
abstract
We describe a "constant pace" framework for video summarization via fast playback or temporal sub-sampling. The pace of the summary serves as a parameter that enables production of a video summary of any desired length. Earlier, we showed that the intensity of motion activity (or pace) of a video sequence is a good indication of its "summarizability". Here we build on this notion by adapting the playback frame-rate or the temporal subsampling rate to the pace. Either the less active parts of the sequence are played back at a faster frame rate or the less active parts of the sequence are sub-sampled more heavily than are the more active parts, so as to produce a summary with constant pace. The basic idea is to skip over the less interesting parts of the video. Our technique is computationally simple and gives satisfactory results with surveillance and entertainment video.
Ajay Divakaran, Kadir A. Peker, Huifang Sun
ICIP (3)1
2001 A Novel Pair-Wise Comparison Based Analytical Framework For Automatic Measurement Of Intensity Of Motion Activity Of Video Segments
abstract
We present a novel pair-wise comparison based psychovisual and analytical framework for automatic measurement of motion activity in video sequences. In [1] we constructed a test-set of video segments and a ground truth, based on subjective tests with naive subjects. We presented automatically extractable descriptors of motion activity computed from MPEG block motion vectors, based on different hypotheses about subjective perception of motion activity. We tested the average error performance of the descriptors. In this paper we test the performance of the descriptors against the ground truth using pair-wise comparison of video segments. We show that all the descriptors perform well and that the MPEG-7 motion activity descriptor, based on variance of motion vector magnitudes, is one of the best in performance. We determine the common limitations of the proposed lowlevel motion activity descriptors. We find that some of the proposed descriptors are significantly less susceptible to the aforementioned limitations.
Kadir A. Peker, Ajay Divakaran
ICME2
2001 Algorithms And System For Segmentation And Structure Analysis In Soccer Video
abstract
In this paper, we present a novel system and effective algorithms for soccer video segmentation. The output, about whether the ball is in play, reveals high-level structure of the content. The first step is to classify each sample frame into 3 kinds of view using a unique domain-specific feature, grass-area-ratio. Here the grass value and classification rules are learned and automatically adjusted to each new clip. Then heuristic rules are used in processing the view label sequence, and obtain play/break status of the game. The results provide good basis for detailed content analysis in next step. We also show that lowlevel features and mid-level view classes can be combined to extract more information about the game, via the example of detecting grass orientation in the field. The results are evaluated under different metrics intended for different applications; the best result in segmentation is 86.5%. 1.
Lexing Xie, Shih-Fu Chang, Ajay Divakaran, Anthony Vetro, Huifang Sun
ICME4
2001 MPEG-7 visual motion descriptors
abstract
This paper describes tools and techniques for representing motion information in the context of MPEG-7 standardization for multimedia description interfaces. It first gives an overview of the current organization of the set of MPEG-7 motion descriptions, then illustrates this by presenting two of them, motion activity and motion trajectory, in more detail. It explains how to extract them from the content, how to express them in a compact way, and illustrates their use in concrete applications scenarios.
Sylvie Jeannin, Ajay Divakaran
IEEE Trans. Circuits Syst. Video Technol.2
2000 A Region Based Descriptor for Spatial Distribution of Motion Activity for Compressed Video
abstract
In this paper we present a new descriptor for the spatial distribution of motion activity in video sequences. We construct a histogram of areas of distinct regions (or "blobs") of "motion active" regions over the entire video shot. We carry out another thresholding process on the histogram to get our descriptor, which is a histogram normalized with respect to the average size of the blobs, and thus normalized with respect to frame size. We get similar precision-recall performance to the spatial activity descriptor in the current MPEG-7 experimental model. We are also able to successfully capture the effects of camera motion as well as the effects of non-camera motion in distinct uncorrelated parts of our descriptor. Since the feature extraction is in the compressed domain and simple, it is extremely fast. We find that our descriptor enables fast and accurate indexing of video.
Ajay Divakaran, Kadir A. Peker, Huifang Sun
ICIP1
1997 Video compression by mean-corrected motion compensation of partial quadtrees
abstract
This paper describes Iterated Systems' submission to the MPEG-4 committee in January 1996. A system for compressing video is presented which expresses predictive frames of a sequence by means of motion vectors applied to variable-size blocks with a constant intensity adjustment. The motion vectors are organized into a partial quadtree, which allows incomplete splitting of blocks into zero to four quadrants. The motion vectors, mean intensity offsets, and partial quadtree structure are arrived at by a joint optimization process which may be carried out in a bottom-up fashion. A form of generalized overlapped block motion compensation applicable at all block sizes is presented, and a method for encoding segmented video is shown. The performance of the algorithm is compared with that of Telenor's H.263 for three of the MPEG-4 test sequences which show that the former is comparable or superior to the latter for low bit rates.
Steve Calzone, Keshi Chen, Chih-Chwen Chuang, Ajay Divakaran, Simant Dube, Lyman Hurd, Jarkko Kari 0001, Gang Liang, Fu-Huei Lin, John Muller, Hawley K. Rising III
IEEE Trans. Circuits Syst. Video Technol.4
1995 Information-theoretic performance of quadrature mirror filters
abstract
Most existing quadrature mirror filters (QMFs) closely match the derived closed-form expression for an efficient class of QMFs. We use the closed-form expressions to derive the relationship between information-theoretic loss and the frequency selectivity of the QMF, by calculating first-order entropy as well as rate-distortion theoretic performance of a two band QMF system. We find that practical QMFs do not suffer a significant information-theoretic loss with first-order autoregressive Gaussian sources. With second-order autoregressive sources we find that practical QMFs suffer a notable information-theoretic loss when the bandwidth of the source is extremely narrow, but incur a small loss when the bandwidth is wider. We suggest that our results broadly apply to higher order autoregressive sources as well.
Ajay Divakaran, William A. Pearlman
IEEE Trans. Inf. Theory1