Karan Sikka

dblp:119/1499 · DBLP profile ↗
← Back
29ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0002-0187-5322ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorSecurity and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A Video is Worth 10, 000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval
abstract
Existing long video retrieval systems are trained and tested in the paragraph-to- video retrieval regime, where ev-ery long video is described by a single long paragraph. This neglects the richness and variety of possible valid de-scriptions of a video, which could range anywhere from moment-by-moment detail to a single phrase summary. To provide a more thorough evaluation of the capabilities of long video retrieval systems, we propose a pipeline that leverages state-of-the-art large language models to care-fully generate a diverse set of synthetic captions for long videos. We validate this pipeline's fidelity via rigorous hu-man inspection. We use synthetic captionsfrom this pipeline to perform a benchmark of a representative set of video language models using long video datasets, and show that the models struggle on shorter captions. We show that finetuning on this data can both mitigate these issues (+2.8% R@ 1 over SOTA on ActivityNet with diverse captions), and even improve performance on standard paragraph-to-video re-trieval (+ 1.0% R@1 on ActivityNet). We also use synthetic data from our pipeline as query expansion in the zero-shot setting (+3.4% R@ 1 on ActivityNet). We derive insights by analyzing failure cases for retrieval with short captions.
Matthew Gwilliam, Michael Cogswell, Meng Ye 0002, Karan Sikka, Abhinav Shrivastava, Ajay Divakaran
WACV4
2024 DRESS : Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback
abstract
We present DRESS , a large vision language model (LVLM) that innovatively exploits Natural Language feedback (NLF) from Large Language Models to enhance its alignment and interactions by addressing two key limitations in the state-of-the-art LVLMs. First, prior LVLMs generally rely only on the instruction finetuning stage to enhance alignment with human preferences. Without incorporating extra feedback, they are still prone to generate unhelpful, hallucinated, or harmful responses. Second, while the visual instruction tuning data is generally structured in a multi-turn dialogue format, the connections and dependencies among consecutive conversational turns are weak. This reduces the capacity for effective multi-turn interactions. To tackle these, we propose a novel categorization of the NLF into two key types: critique and refinement. The critique NLF identifies the strengths and weaknesses of the responses and is used to align the LVLMs with human preferences. The refinement NLF offers concrete suggestions for improvement and is adopted to improve the interaction ability of the LVLMs- which focuses on LVLMs' ability to refine responses by incorporating feedback in multi-turn interactions. To address the non-differentiable nature of NLF, we generalize conditional reinforcement learning for training. Our experimental results demonstrate that DRESS can generate more helpful (9.76%), honest (11.52%), and harmless (21.03%) responses, and more effectively learn from feedback during multi-turn interactions compared to SOTA LVLMs.
Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji 0001, Ajay Divakaran
CVPR2
2024 Pelican: Correcting Hallucination in Vision-LLMs via Claim Decomposition and Program of Thought Verification
abstract
Large Visual Language Models (LVLMs) struggle with hallucinations in visual instruction following task(s). These issues hinder their trustworthiness and real-world applicability. We propose Pelican – a novel framework designed to detect and mitigate hallucinations through claim verification. Pelican first decomposes the visual claim into a chain of sub-claims based on first-order predicates. These sub-claims consists of (predicate, question) pairs and can be conceptualized as nodes of a computational graph. We then use use Program-of-Thought prompting to generate Python code for answering these questions through flexible composition of external tools. Pelican improves over prior work by introducing (1) intermediate variables for precise grounding of object instances, and (2) shared computation for answering the sub-question to enable adaptive corrections and inconsistency identification. We finally use reasoning abilities of LLM to verify the correctness of the the claim by considering the consistency and confidence of the (question, answer) pairs from each sub-claim. Our experiments demonstrate consistent performance improvements over various baseline LVLMs and existing hallucination mitigation approaches across several benchmarks.
Pritish Sahu, Karan Sikka, Ajay Divakaran
EMNLP2
2024 SayNav: Grounding Large Language Models for Dynamic Planning to Navigation in New Environments
abstract
Semantic reasoning and dynamic planning capabilities are crucial for an autonomous agent to perform complex navigation tasks in unknown environments. It requires a large amount of common-sense knowledge, that humans possess, to succeed in these tasks. We present SayNav, a new approach that leverages human knowledge from Large Language Models (LLMs) for efficient generalization to complex navigation tasks in unknown large-scale environments. SayNav uses a novel grounding mechanism, that incrementally builds a 3D scene graph of the explored environment as inputs to LLMs, for generating feasible and contextually appropriate high-level plans for navigation. The LLM-generated plan is then executed by a pre-trained low-level planner, that treats each planned step as a short-distance point-goal navigation sub-task. SayNav dynamically generates step-by-step instructions during navigation and continuously refines future steps based on newly perceived information. We evaluate SayNav on multi-object navigation (MultiON) task, that requires the agent to utilize a massive amount of human knowledge to efficiently search multiple different objects in an unknown environment. We also introduce a benchmark dataset for MultiON task employing ProcTHOR framework that provides large photo-realistic indoor environments with variety of objects. SayNav achieves state-of-the-art results and even outperforms an oracle based baseline with strong ground-truth assumptions by more than 8% in terms of success rate, highlighting its ability to generate dynamic plans for successfully locating objects in large-scale new environments. The code, benchmark dataset and demonstration videos are accessible at https://www.sri.com/ics/computer-vision/saynav.
Abhinav Rajvanshi, Karan Sikka, Bhoram Lee, Han-Pang Chiu, Alvaro Velasquez
ICAPS2
2024 Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models
abstract
Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji, Ajay Divakaran. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji 0001, Ajay Divakaran
NAACL-HLT2
2023 Multilingual Content Moderation: A Case Study on Reddit
abstract
Meng Ye, Karan Sikka, Katherine Atwell, Sabit Hassan, Ajay Divakaran, Malihe Alikhani. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Meng Ye 0002, Karan Sikka, Katherine Atwell, Sabit Hassan, Ajay Divakaran, Malihe Alikhani
EACL2
2023 TIJO: Trigger Inversion with Joint Optimization for Defending Multimodal Backdoored Models
abstract
We present a Multimodal Backdoor Defense technique TIJO (Trigger Inversion using Joint Optimization). Recent work [48] has demonstrated successful backdoor attacks on multimodal models for the Visual Question Answering task. Their dual-key backdoor trigger is split across two modalities (image and text), such that the backdoor is activated if and only if the trigger is present in both modalities. We propose TIJO that defends against dual-key attacks through a joint optimization that reverse-engineers the trigger in both the image and text modalities. This joint optimization is challenging in multimodal models due to the disconnected nature of the visual pipeline which consists of an offline feature extractor, whose output is then fused with the text using a fusion module. The key insight enabling the joint optimization in TIJO is that the trigger inversion needs to be carried out in the object detection box feature space as opposed to the pixel space. We demonstrate the effectiveness of our method on the TrojVQA benchmark, where TIJO improves upon the state-of-the-art unimodal methods from an AUC of 0.6 to 0.92 on multimodal dual-key back-doors. Furthermore, our method also improves upon the unimodal baselines on unimodal backdoors. We present ablation studies and qualitative results to provide insights into our algorithm such as the critical importance of overlaying the inverted feature triggers on all visual features during trigger inversion. The prototype implementation of TIJO is available at https://github.com/SRI-CSL/TIJO.
Indranil Sur, Karan Sikka, Matthew Walmer, Kaushik Koneripalli, Ajay Divakaran, Susmit Jha
ICCV2
2023 Predicting Information Pathways Across Online Communities
abstract
The problem of community-level information pathway prediction (CLIPP) aims at predicting the transmission trajectory of content across online communities. A successful solution to CLIPP holds significance as it facilitates the distribution of valuable information to a larger audience and prevents the proliferation of misinfor- mation. Notably, solving CLIPP is non-trivial as inter-community relationships and influence are unknown, information spread is multi-modal, and new content and new communities appear over time. In this work, we address CLIPP by collecting large-scale, multi-modal datasets to examine the diffusion of online YouTube videos on Reddit. We analyze these datasets to construct community influence graphs (CIGs) and develop a novel dynamic graph frame- work, INPAC (Information Pathway Across Online Communities), which incorporates CIGs to capture the temporal variability and multi-modal nature of video propagation across communities. Ex- perimental results in both warm-start and cold-start scenarios show that INPAC outperforms seven baselines in CLIPP. Our code and datasets are available at https://github.com/claws-lab/INPAC
Yiqiao Jin, Yeon-Chang Lee, Kartik Sharma, Meng Ye 0002, Karan Sikka, Ajay Divakaran, Srijan Kumar
KDD5
2022 Dual-Key Multimodal Backdoors for Visual Question Answering
abstract
The success of deep learning has enabled advances in multimodal tasks that require non-trivial fusion of multiple input domains. Although multimodal models have shown potential in many problems, their increased complexity makes them more vulnerable to attacks. A Backdoor (or Trojan) attack is a class of security vulnerability wherein an attacker embeds a malicious secret behavior into a network (e.g. targeted misclassification) that is activated when an attacker-specified trigger is added to an input. In this work, we show that multimodal networks are vulnerable to a novel type of attack that we refer to as Dual-Key Multimodal Backdoors. This attack exploits the complex fusion mechanisms used by state-of-the-art networks to embed backdoors that are both effective and stealthy. Instead of using a single trigger, the proposed attack embeds a trigger in each of the input modalities and activates the malicious behavior only when both the triggers are present. We present an extensive study of multimodal backdoors on the Visual Question Answering (VQA) task with multiple architectures and visual feature backbones. A major challenge in embedding backdoors in VQA models is that most models use visual features extracted from a fixed pretrained object detector. This is challenging for the attacker as the detector can distort or ignore the visual trigger entirely, which leads to models where backdoors are over-reliant on the language trigger. We tackle this problem by proposing a visual trigger optimization strategy designed for pretrained object detectors. Through this method, we create Dual-Key Backdoors with over a 98% attack success rate while only poisoning 1% of the training data. Finally, we release TrojVQA, a large collection of clean and trojan VQA models to enable research in defending against multimodal backdoors.
Matthew Walmer, Karan Sikka, Indranil Sur, Abhinav Shrivastava, Susmit Jha
CVPR2
2022 Challenges in Procedural Multimodal Machine Comprehension: A Novel Way To Benchmark
abstract
We focus on Multimodal Machine Reading Comprehension (M3C) where a model is expected to answer questions based on given passage (or context), and the context and the questions can be in different modalities. Previous works such as RecipeQA have proposed datasets and cloze-style tasks for evaluation. However, we identify three critical biases stemming from the question-answer generation process and memorization capabilities of large deep models. These biases makes it easier for a model to overfit by relying on spurious correlations or naive data patterns. We propose a systematic framework to address these biases through three Control-Knobs that enable us to generate a test bed of datasets of progressive difficulty levels. We believe that our benchmark (referred to as Meta- RecipeQA) will provide, for the first time, a fine grained estimate of a model’s generalization capabilities. We also propose a general M3C model that is used to realize several prior SOTA models and motivate a novel hierarchical transformer based reasoning network (HTRN). We perform a detailed evaluation of these models with different language and visual features on our benchmark. We observe a consistent improvement with HTRN over SOTA (~ 18% in Visual Cloze task and ~ 13% in average over all the tasks). We also observe a drop in performance across all the models when testing on RecipeQA and proposed Meta–RecipeQA (e.g. 83.6% versus 67.1% for HTRN), which shows that the proposed dataset is relatively less biased. We conclude by highlighting the impact of the control knobs with some quantitative results.
Pritish Sahu, Karan Sikka, Ajay Divakaran
WACV2
2021 MISA: Online Defense of Trojaned Models using Misattributions
abstract
Recent studies have shown that neural networks are vulnerable to Trojan attacks, where a network is trained to respond to specially crafted trigger patterns in the inputs in specific and potentially malicious ways. This paper proposes MISA, a new online approach to detect Trojan triggers for neural networks at inference time. Our approach is based on a novel notion called misattributions, which captures the anomalous manifestation of a Trojan activation in the feature space. Given an input image and the corresponding output prediction, our algorithm first computes the model’s attribution on different features. It then statistically analyzes these attributions to ascertain the presence of a Trojan trigger. Across a set of benchmarks, we show that our method can effectively detect Trojan triggers for a wide variety of trigger patterns, including several recent ones for which there are no known defenses. Our method achieves 96% AUC for detecting images that include a Trojan trigger without any assumptions on the trigger pattern.
Panagiota Kiourti, Wenchao Li 0001, Karan Sikka, Susmit Jha
ACSAC4
2020 RGB2LIDAR: Towards Solving Large-Scale Cross-Modal Visual Localization
abstract
We study an important, yet largely unexplored problem of large-scale cross-modal visual localization by matching ground RGB images to a geo-referenced aerial LIDAR 3D point cloud (rendered as depth images). Prior works were demonstrated on small datasets and did not lend themselves to scaling up for large-scale applications. To enable large-scale evaluation, we introduce a new dataset containing over 550K pairs (covering 143 km2 area) of RGB and aerial LIDAR depth images. We propose a novel joint embedding based method that effectively combines the appearance and semantic cues from both modalities to handle drastic cross-modal variations. Experiments on the proposed dataset show that our model achieves a strong result of a median rank of 5 in matching across a large test set of 50K location pairs collected from a 14km^2 area. This represents a significant advancement over prior works in performance and scale. We conclude with qualitative results to highlight the challenging nature of this task and the benefits of the proposed model. Our work provides a foundation for further research in cross-modal visual localization.
Niluthpol Chowdhury Mithun, Karan Sikka, Han-Pang Chiu, Supun Samarasekera, Rakesh Kumar 0001
ACM Multimedia2
2019 Semantically-Aware Attentive Neural Embeddings for 2D Long-Term Visual Localization
Zachary Seymour, Karan Sikka, Han-Pang Chiu, Supun Samarasekera, Rakesh Kumar 0001
BMVC2
2019 Integrating Text and Image: Determining Multimodal Document Intent in Instagram Posts
abstract
Julia Kruk, Jonah Lubin, Karan Sikka, Xiao Lin, Dan Jurafsky, Ajay Divakaran. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Julia Kruk, Jonah Lubin, Karan Sikka, Daniel Jurafsky, Ajay Divakaran
EMNLP/IJCNLP (1)3
2019 Sunny and Dark Outside?! Improving Answer Consistency in VQA through Entailed Question Generation
abstract
Arijit Ray, Karan Sikka, Ajay Divakaran, Stefan Lee, Giedrius Burachas. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Arijit Ray, Karan Sikka, Ajay Divakaran, Stefan Lee, Giedrius Burachas
EMNLP/IJCNLP (1)2
2019 Align2Ground: Weakly Supervised Phrase Grounding Guided by Image-Caption Alignment
abstract
We address the problem of grounding free-form textual phrases by using weak supervision from image-caption pairs. We propose a novel end-to-end model that uses caption-to-image retrieval as a downstream task to guide the process of phrase localization. Our method, as a first step, infers the latent correspondences between regions-of-interest (RoIs) and phrases in the caption and creates a discriminative image representation using these matched RoIs. In the subsequent step, this learned representation is aligned with the caption. Our key contribution lies in building this "caption-conditioned" image encoding, which tightly couples both the tasks and allows the weak supervision to effectively guide visual grounding. We provide extensive empirical and qualitative analysis to investigate the different components of our proposed model and compare it with competitive baselines. For phrase localization, we report an improvement of 4.9% and 1.3% (absolute) over the prior state-of-the-art on the VisualGenome and Flickr30k Entities datasets. We also report results that are at par with the state-of-the-art on the downstream caption-to-image retrieval task on COCO and Flickr30k datasets.
Samyak Datta, Karan Sikka, Karuna Ahuja, Devi Parikh, Ajay Divakaran
ICCV2
2018 Zero-Shot Object Detection
Ankan Bansal, Karan Sikka, Gaurav Sharma 0004, Rama Chellappa, Ajay Divakaran
ECCV (1)2
2018 Discriminatively Trained Latent Ordinal Model for Video Classification
abstract
We address the problem of video classification for facial analysis and human action recognition. We propose a novel weakly supervised learning method that models the video as a sequence of automatically mined, discriminative sub-events (e.g., onset and offset phase for "smile", running and jumping for "highjump"). The proposed model is inspired by the recent works on Multiple Instance Learning and latent SVM/HCRF - it extends such frameworks to model the ordinal aspect in the videos, approximately. We obtain consistent improvements over relevant competitive baselines on four challenging and publicly available video based facial analysis datasets for prediction of expression, clinical pain and intent in dyadic conversations, and on three challenging human action datasets. We also validate the method with qualitative results and show that they largely support the intuitions behind the method.
Karan Sikka, Gaurav Sharma 0004
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 AdaScan: Adaptive Scan Pooling in Deep Convolutional Neural Networks for Human Action Recognition in Videos
abstract
We propose a novel method for temporally pooling frames in a video for the task of human action recognition. The method is motivated by the observation that there are only a small number of frames which, together, contain sufficient information to discriminate an action class present in a video, from the rest. The proposed method learns to pool such discriminative and informative frames, while discarding a majority of the non-informative frames in a single temporal scan of the video. Our algorithm does so by continuously predicting the discriminative importance of each video frame and subsequently pooling them in a deep learning framework. We show the effectiveness of our proposed pooling method on standard benchmarks where it consistently improves on baseline pooling methods, with both RGB and optical flow based Convolutional networks. Further, in combination with complementary video representations, we show results that are competitive with respect to the state-of-the-art results on two challenging and publicly available benchmark datasets.
Amlan Kar, Nishant Rai, Karan Sikka, Gaurav Sharma 0004
CVPR3
2017 Deep active object recognition by joint label and action prediction
Mohsen Malmir, Karan Sikka, Deborah Forster, Ian R. Fasel, Javier R. Movellan, Garrison W. Cottrell
Comput. Vis. Image Underst.2
2016 LOMo: Latent Ordinal Model for Facial Analysis in Videos
abstract
We study the problem of facial analysis in videos. We propose a novel weakly supervised learning method that models the video event (expression, pain etc.) as a sequence of automatically mined, discriminative sub-events (e.g. onset and offset phase for smile, brow lower and cheek raise for pain). The proposed model is inspired by the recent works on Multiple Instance Learning and latent SVM/HCRF - it extends such frameworks to model the ordinal or temporal aspect in the videos, approximately. We obtain consistent improvements over relevant competitive baselines on four challenging and publicly available video based facial analysis datasets for prediction of expression, clinical pain and intent in dyadic conversations. In combination with complimentary features, we report state-of-the-art results on these datasets.
Karan Sikka, Gaurav Sharma 0004, Marian Stewart Bartlett
CVPR1
2015 Deep Q-learning for Active Recognition of GERMS: Baseline performance on a standardized dataset for active learning
abstract
Mohsen Malmir1 http://mplab.ucsd.edu/~mmalmir/ Karan Sikka1 http://mplab.ucsd.edu/~ksikka/ Deborah Forster1 [email protected] Javier Movellan2 http://www.emotient.com/ Garrison W. Cottrell3 http://cseweb.ucsd.edu/~gary/ 1 Machine Perception Lab. University of California San Diego, San Diego, CA, USA 2 Emotient, Inc. 4435 Eastgate Mall, Suite 320, San Diego, CA, USA 3 Computer Science and Engineering Dept. University of California San Diego, San Diego, CA, USA
Mohsen Malmir, Karan Sikka, Deborah Forster, Javier R. Movellan, Garison Cottrell
BMVC2
2015 Joint Clustering and Classification for Multiple Instance Learning
Karan Sikka, Ritwik Giri, Marian Stewart Bartlett
BMVC1
2014 Emotion Recognition In The Wild Challenge 2014: Baseline, Data and Protocol
abstract
The Second Emotion Recognition In The Wild Challenge (EmotiW) 2014 consists of an audio-video based emotion classification challenge, which mimics the real-world conditions. Traditionally, emotion recognition has been performed on data captured in constrained lab-controlled like environment. While this data was a good starting point, such lab controlled data poorly represents the environment and conditions faced in real-world situations. With the exponential increase in the number of video clips being uploaded online, it is worthwhile to explore the performance of emotion recognition methods that work `in the wild'. The goal of this Grand Challenge is to carry forward the common platform defined during EmotiW 2013, for evaluation of emotion recognition methods in real-world conditions. The database in the 2014 challenge is the Acted Facial Expression In Wild (AFEW) 4.0, which has been collected from movies showing close-to-real-world conditions. The paper describes the data partitions, the baseline method and the experimental protocol.
Abhinav Dhall, Roland Göcke, Jyoti Joshi, Karan Sikka, Tom Gedeon
ICMI4
2014 Facial Expression Analysis for Estimating Pain in Clinical Settings
abstract
Pain assessment is vital for effective pain management in clinical settings. It is generally obtained via patient's self-report or observer's assessment. Both of these approaches suffer from several drawbacks such as unavailability of self-report, idiosyncratic use and observer bias. This work aims at developing automated machine learning based approaches for estimating pain in clinical settings. We propose to use facial expression information to accomplish current goals since previous studies have demonstrated consistency between facial behavior and experienced pain. Moreover, with recent advances in computer vision it is possible to design algorithms for identifying spontaneous expressions such as pain in more naturalistic conditions.
Karan Sikka
ICMI1
2014 A discriminative parts based model approach for fiducial points free and shape constrained head pose normalisation in the wild
abstract
Continuous Confidence Map Based Normalisation: While continuous head pose normalisation is not the goal of this paper, we demonstrate as a proof of concept that it is possible to extend the current method for continuous head pose normalisation. For dealing with faces in videos [1], continuous head pose normalisation is required. [2] argue that the appearance of a part does not changes with a subtle pose change, therefore a detector for part i in pose angle p can be shared for the same part i for a pose angle p + δ. Further experiments in [2] showed that sharing based models and independent model have comparable performance. However, sharing based models are faster upto ten times as compared to the independent models [2]. The confidence maps based methods (CM-HPNPSand CM-HPNPI) can be extended from discrete to continuous by sharing part-specific regression models R, which are shared among neighboring pose angles.
Abhinav Dhall, Karan Sikka, Gwen Littlewort, Roland Göcke, Marian Stewart Bartlett
WACV2
2014 A discriminative parts based model approach for fiducial points free and shape constrained head pose normalisation in the wild
abstract
This paper proposes a method for parts-based view-invariant head pose normalisation, which works well even in difficult real-world conditions. Handling pose is a classical problem in facial analysis. Recently, parts-based models have shown promising performance for facial landmark points detection `in the wild'. Leveraging on the success of these models, the proposed data-driven regression framework computes a constrained normalised virtual frontal head pose. The response maps of a discriminatively trained part detector are used as texture information. These sparse texture maps are projected from non-frontal to frontal pose using block-wise structured regression. Finally, a facial kinematic shape constraint is achieved by applying a shape model. The advantages of the proposed approach are: a) no explicit dependence on the outputs of a facial parts detector and, thus, avoiding any error propagation owing to their failure; (b) the application of a shape prior on the reconstructed frontal maps provides an anatomically constrained facial shape; and c) modelling head pose as a mixture-of-parts model allows the framework to work without any prior pose information. Experiments are performed on the Multi-PIE and the `in the wild' SFEW databases. The results demonstrate the effectiveness of the proposed method.
Abhinav Dhall, Karan Sikka, Gwen Littlewort, Roland Göcke, Marian Stewart Bartlett
WACV2
2014 Classification and weakly supervised pain localization using multiple segment representation
Karan Sikka, Abhinav Dhall, Marian Stewart Bartlett
Image Vis. Comput.1
2013 Multiple kernel learning for emotion recognition in the wild
abstract
We propose a method to automatically detect emotions in unconstrained settings as part of the 2013 Emotion Recognition in the Wild Challenge [16], organized in conjunction with the ACM International Conference on Multimodal Interaction (ICMI 2013). Our method combines multiple visual descriptors with paralinguistic audio features for multimodal classification of video clips. Extracted features are combined using Multiple Kernel Learning and the clips are classified using an SVM into one of the seven emotion categories: Anger, Disgust, Fear, Happiness, Neutral, Sadness and Surprise. The proposed method achieves competitive results, with an accuracy gain of approximately 10% above the challenge baseline.
Karan Sikka, Karmen Dykstra, Suchitra Sathyanarayana, Gwen Littlewort, Marian Stewart Bartlett
ICMI1